Cognition has brought Fusion to Devin Desktop and CLI. A frontier lead plans and reviews while a lower-cost sidekick works in a separate context; published benchmark savings are useful pilot inputs, not a guaranteed bill reduction.
RELATED BUYER PROFILEReview Devin Fusion pricing, access and test-pending evidencePricing · access · strengths · limitations →CONTINUE THE DECISIONRun a same-repository lead-and-sidekick pilotEvidence · workflow · next action →CONTINUE THE DECISIONCalculate complete cost per accepted code changeEvidence · workflow · next action →CONTINUE THE DECISIONBuild a controlled AI stack for a software teamEvidence · workflow · next action →What you need to know
- Fusion uses two active agents rather than making one initial model-routing decision: a lead owns planning and review while a sidekick executes delegated work
- Cognition recommends Fable 5.1 with its new SWE-2 model, while the CLI also shows an Astra and SWE-2 pairing
- Published savings vary by benchmark and configuration; they do not guarantee lower cost or equal quality on a buyer's own codebase
The published evaluation boundary
What Cognition released
Cognition announced on 11 September that Devin Fusion is available in Devin Desktop and Devin CLI. The harness asks the user to choose two models. A frontier lead owns the plan, interprets ambiguity, reviews the work and can take control back; a lower-cost sidekick explores the repository, edits files and runs tests from its own persistent context. The recommended pairing is Claude Fable 5.1 as lead and Cognition SWE-2 as sidekick, while Cognition also publishes results for GPT-6 Astra with SWE-2.
This is orchestration, not a universal model ranking
Fusion is designed to avoid a single up-front routing decision. The lead remains involved throughout the task, sends bounded briefs to the sidekick and reviews returned work. That can preserve expensive reasoning for planning and judgement while moving more mechanical work to a cheaper model. It also creates a more complex system: two contexts, delegation quality, review loops and recovery behaviour all affect the result. A strong lead-sidekick pair on one benchmark is not evidence that either model is universally best.
What the published numbers actually show
Cognition's five-row table reports cost reductions from 11% to 46% for selected Fusion configurations against the same lead model without Fusion. Scores are sometimes close, sometimes higher and sometimes lower. On Terminal-Bench 4, for example, the reported Astra Fusion score falls from 55.6 to 50.0 while cost falls from $10.08 to $6.06. On SWE-Atlas QnA, the Fable Fusion score is reported slightly higher while cost is lower. Cognition says it worked with Artificial Analysis and Vals AI on the evaluations. HubAI treats these as disclosed launch evidence, not an independent HubAI test or a savings guarantee.
Price per token is not price per accepted change
The release makes a useful commercial point: a stronger sidekick can use fewer turns and create less correction work, so the lower token rate is not always the cheaper system. Buyers should calculate the complete cost of an accepted repository change: lead and sidekick inference, retries, failed runs, developer review, CI time and any repair after the task. Cognition's current pricing page lists Devin Pro at $20 per month, Max at $200, and Teams at $80 per month plus $40 per full developer seat. Paid plans include quotas; additional usage can be purchased at API pricing. Exact task cost still varies by model, effort and complexity.
SWE-2 access is promotional, not permanently free
Cognition's pricing page says SWE-2 use is free in Devin Desktop and CLI for Pro, Max and Teams subscribers through 10 October 2026. The public Free plan has a light quota and limited model availability, so it should not be described as a full Fusion trial. The company has not promised that the promotional SWE-2 terms will continue after that date. UK teams should confirm current billing, data-processing and organisation controls in the live plan before procurement.
HubAI buyer verdict
Fusion deserves a shortlist place for teams already comparing premium coding agents, because it turns model selection into a visible lead-and-sidekick operating design. Do not migrate from Claude Code, Codex, Copilot or a fixed Devin model on benchmark savings alone. Freeze five representative repository tasks, define an accepted change before each run, and compare total inference, elapsed time, tests, reviewer corrections, reverted changes and security interventions. Keep Fusion only if the completed-work economics improve without weakening quality or control.
HUBAI VIEWEngineering teams should compare accepted-change cost, review burden and failure recovery on their own repository—not choose a harness from token price or a vendor-arranged benchmark alone.
Buyer decision signal: New multi-model harness · Desktop and CLI
What to verify next
1Confirm the Devin plan, included quota, overage rules and SWE-2 promotion end date
2Choose lead and sidekick models explicitly and record their effort settings
3Run the same five repository tasks through Fusion and the current single-agent baseline
4Measure accepted changes, retries, reviewer corrections, elapsed time and complete task cost
5Test what happens when the sidekick misunderstands a brief or the lead fails to catch an error
6Review local credentials, shell permissions, network access, hooks and team-enforced rules
7Verify UK billing and contracted data-processing terms rather than inferring them from a public page
8Recheck price and access after the 10 October 2026 promotion
Read the evidence
Capabilities, availability and prices can change. HubAI keeps analysis separate from the underlying official material.
01Cognition: Introducing Fusion in Devin Desktop and CLIOpen source ↗02Devin CLI: current access, model pairings and controlsOpen source ↗03Devin: current plans, quotas and SWE-2 promotionOpen source ↗04Cognition: Introducing SWE-2Open source ↗
Business
Business
Models
Safety