Introducing SWE-2, our closest model yet to the frontier.
On leading evals, it scores on par with recent frontier models – at up to 70% lower cost.
We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
Cognition
@cognition
SWE-2 is post-trained on Kimi-K3 and proves that our RL recipe continues to scale on stronger base models.
On FrontierCode, SWE-2 achieves a score of 50.0%. It beats SWE-1.7, Grok 4.6, and GPT 5.6 Sol while matching Fable 5.1 at 64% lower cost.
On FrontierCode, SWE-2 achieves a score of 50.0%. It beats SWE-1.7, Grok 4.6, and GPT 5.6 Sol while matching Fable 5.1 at 64% lower cost.

3:21 PM UTC · Sep 10, 2026 · 85.8K Views
92574362
Cognition
@cognition
SWE-2 is our first model to support effort levels. For that, we made significant improvements to our length penalty recipe - read more in our blog post!
In a single RL run, we are able to push the Pareto curve while preserving its shape: medium effort becomes both cheaper & smarter; max effort learns to use more tokens & turns to achieve the highest scores.
In a single RL run, we are able to push the Pareto curve while preserving its shape: medium effort becomes both cheaper & smarter; max effort learns to use more tokens & turns to achieve the highest scores.
3:21 PM UTC · Sep 10, 2026 · 53.9K Views
11039733
Cognition
@cognition
The result is a model that is way more efficient than SWE-1.7.
On FrontierCode 1.1 Main, SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average.
We observe that SWE-2 learns to explore in a more focused way before making changes.
On FrontierCode 1.1 Main, SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average.
We observe that SWE-2 learns to explore in a more focused way before making changes.

3:21 PM UTC · Sep 10, 2026 · 37.7K Views
2528213
Cognition
@cognition
As a result, SWE-2 is our most capable model yet for software engineering tasks – while achieving market-leading price performance.
Our team at Cognition uses it for feature development, debugging, and even complex visualizations of novel math:
x.com/penlume/status…
Our team at Cognition uses it for feature development, debugging, and even complex visualizations of novel math:
x.com/penlume/status…
penlu @penlume
a visualization of the flow field of the navier-stokes finite-time blowup solution. it's a good model sir
(singularity at t = 1 with viscosity matching water, unit distance ~ 1 mm. leading and first order background terms)
x.com/OpenAI/status/…
(singularity at t = 1 with viscosity matching water, unit distance ~ 1 mm. leading and first order background terms)
x.com/OpenAI/status/…
3:21 PM UTC · Sep 10, 2026 · 53.4K Views
1223416
Cognition
@cognition
SWE-2 is available today in Devin across Desktop and CLI. We’re making it free for all Pro, Max & Teams subscribers for the next month.
Read more about how we trained SWE-2:
cognition.com/blog/swe-2
Read more about how we trained SWE-2:
cognition.com/blog/swe-2

cognition.com
Introducing SWE-2: Pushing the Pareto Frontier
Today we’re introducing SWE-2, our most advanced coding model yet. SWE-2 delivers highly competitive agentic coding performance across multiple effort…
3:21 PM UTC · Sep 10, 2026 · 111.8K Views
3526466148
