top of page

Cognition Releases SWE-1.7 Close to GPT-5.5 and Opus Levels

Jul 10
2 min read

Updated: Jul 20

Cognition released SWE-1.7, its latest coding model trained on Kimi K2.7 and tuned through an expanded reinforcement-learning pipeline.

The new model posted 42.3 percent on FrontierCode 1.1 Main. That score sits between Kimi K2.7 Code at 30.1 percent and GPT-5.5 at 43.0 percent. Opus 4.8 still leads at 46.5 percent. On Terminal-Bench 2.1 the score reached 81.5 percent and on SWE-Bench Multilingual the score reached 77.8 percent.

New scores shift the coding-model leaderboard

FrontierCode 1.1 Main measures performance on multi-file, multi-hour engineering tasks. The 12-point jump from the base Kimi model shows the added reinforcement-learning steps delivered measurable progress. The same improvements lifted Terminal-Bench 2.1 and the multilingual SWE-Bench track.

Cognition now ships SWE-1.7 inside Devin on web, desktop, and CLI clients. Inference runs on Cerebras hardware at 1000 tokens per second.

Engineering teams feel pressure from tighter cost curves

Teams that previously chose GPT-5.5 or Opus for long-running agents now have a lower-cost alternative that lands within a few points on the same benchmarks. Procurement teams must decide whether the remaining gap justifies paying premium rates or whether SWE-1.7 covers their typical workload.

The gap to Opus stays at roughly four points on FrontierCode. That margin keeps Opus relevant for the hardest instances, yet the pricing difference changes the default choice for daily work.

Cost and latency set the real contest

SWE-1.7 reached its scores after targeted upgrades in infrastructure, training stability, data quality, and long-context handling. The company reports the same stack that supports inference at 1000 tokens per second also lowers per-token cost compared with earlier frontier models.

Competitors still hold leads on raw capability, yet the combination of near-frontier scores and lower inference cost creates the immediate pressure. Companies that sell high-margin coding agents must now defend against a cheaper option that already covers most engineering workflows.

Remaining gaps keep the choice case-by-case

The four-point spread on FrontierCode 1.1 Main still leaves some long-horizon tasks where Opus performs better. In addition, the new model has not yet accumulated the same volume of independent audits that GPT-5.5 and Opus carry.

Users who require maximum reliability on the hardest instances will continue to run the current leaders while routing routine work to SWE-1.7. Mixed deployments are the pattern most teams are testing first.

Three signals to track over the next quarter

Watch whether SWE-Bench Multilingual scores move above 80 percent in the next release. A larger jump would indicate the reinforcement-learning pipeline scales beyond the current benchmarks.

Watch adoption data inside Devin. If weekly active coding sessions rise sharply, the cost advantage is translating into real usage.

Watch whether GPT-5.5 or Opus respond with price cuts or new agent features aimed at the same async-task segment. Any move on those fronts will show how seriously the established models view the new competitor.

Developers and engineering leads who manage agent spend should test SWE-1.7 on their actual task mix this month. The narrow gap to the current leaders and the reported cost reduction make it the clearest pricing benchmark the category has seen so far.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page