Inside Cognition Labs: Devin v2 and the bet on long-horizon agents
A profile. Cognition's $25B valuation rests on a thesis the rest of the field is still arguing about — that the next moat in coding agents is the ability to hold a task open for hours, not minutes.
FILED FROM SAN FRANCISCO — Cognition Labs has spent the last twelve months doing something almost no other coding-agent company has been willing to do publicly: arguing, in print and in product, that the right unit of work for an autonomous engineer is not a pull request but a session.
The argument is structural. A pull-request-sized task is a unit of work an autonomous system can finish in a couple of minutes — fix the failing test, refactor the function, add the missing migration. The model layer is, by mid-2026, good enough that the PR-sized task is increasingly commoditized; Claude Code, Cursor's background agents, OpenAI Codex, GitHub Copilot Workspace, and the larger field of vibe-coding tools all compete in roughly the same minutes-long window. Cognition's bet is that the next moat is the hours-long window — the kind of task an experienced engineer would walk away from for a coffee and come back to find substantially progressed.
This is the bet behind Devin v2. The product released in late August keeps a project state, retains its own scratchpad memory across sub-tasks, and is built around what the company internally calls "the day-long task primitive." The CEO, Scott Wu, told the Bulletin in late August that the framing was deliberate: "the day-long unit is the one our customers measure their senior engineers against. If we can hold a task open for that horizon, the unit economics of autonomy land in the right place."
The unit economics are, in fact, the part of the Cognition story most worth watching. Devin is reportedly at $73 million ARR by June 2025 (SiliconANGLE), up from $1M nine months earlier. The company closed at $10.2B in September 2025 (TechCrunch). By spring of 2026 the company was in talks at $25B (Bloomberg). The trajectory is steeper than almost any other autonomy-layer company the Bulletin tracks, including Lovable (which reached $400M ARR by February 2026, on a vibe-coding rather than long-horizon thesis).
What the Bulletin's read of the company keeps returning to is the cultural artifact underneath the product. Cognition was unusual from the start in its hiring posture — competitive-programming background, IOI and IMO medalists, an unusual concentration of people who think about correctness rather than throughput. The company's engineering culture, as several former employees have described it to us, prioritizes long-horizon reasoning at the expense of short-horizon iteration. That cultural posture is the part of Cognition that does not show up on the cap table but that the Bulletin thinks does the most work in explaining the product's distinctive shape.
The Windsurf acquisition
The mid-2025 acquisition of Windsurf (TechFundingNews) is the move that makes the long-horizon thesis credible at scale. Windsurf brought a competent IDE-shaped product into the Cognition stack. The combined posture — long-horizon agent at the back, real engineer-facing IDE at the front — gives Cognition something the competing coding-agent products have struggled to build: a single product that can hold the engineer's full attention during the short-horizon work and recover the session into the long-horizon work without losing context.
The structural reason this matters is that the long-horizon bet is not just a model bet. It is a memory bet, a session-state bet, a tool-use bet, and a UI bet. Cognition has been spending the last twelve months engineering all four. The Bulletin's view is that this is the part of the field that will produce the most durable moats once the foundation-model layer commoditizes further.
What the field is doing differently
Cursor is the dominant competitor in the editor-shaped product window. Lovable owns the vibe-coding category. Claude Code is the model-vendor reference implementation. None of them is positioned as a long-horizon competitor; all of them treat the long-horizon problem as a marketing claim rather than a product primitive.
This is not a criticism of those products. The PR-sized task is a useful unit of work, and a product that does the PR-sized task well will keep finding buyers. The Bulletin's editorial position is that the long-horizon task is a structurally different category, that Cognition is currently the only company building for it seriously, and that the competitive dynamics in eighteen months will look quite different from the way they look now.
What we are watching
Three things, specifically.
The first is whether Devin's $73M ARR continues its current trajectory. If it does, the long-horizon thesis acquires a commercial floor that the field will have to respond to. If it stalls, the thesis becomes harder to argue.
The second is whether Cognition's hiring posture survives the scale-up. Competitive-programming background and long-horizon reasoning do not necessarily survive a thousand-person company. The Bulletin will be watching for the first public signals of cultural drift.
The third is whether the $25B talked-about number closes higher or lower than the floor. The base-rate behavior of mega-rounds in this category is that the talked-about number is the floor. If Cognition lands closer to $30B or above, the field will treat the long-horizon thesis as validated. If it lands materially below, the thesis is harder to argue.
We will report on each of these as they resolve.
Inside Cognition Labs: Devin v2 and the bet on long-horizon agents · Idris Aksoy · The Web4 Bulletin · 2025-09-12
Retrieved 2026-06-12 · Permalink: https://web4bulletin.com/articles/inside-cognition-labs-devin-v2-and-long-horizon-agents/