top of page

Artificial Analysis Coding Agent Index Puts Claude First, but Cost Reorders the Leaders

6 days ago
13 min read

Artificial Analysis put Claude Sonnet 5.5 first with 68 points, but its new coding-agent results reveal a much less comfortable cost story.

The Artificial Analysis Coding Agent Index places Claude Code with Sonnet 5.5 at maximum effort ahead of Gemini 4 Argon and GPT-6.1 Sol. Yet the winner consumes far more time, tokens, and API spending per task than its closest new rivals.

That difference changes the practical decision for development teams. Claude leads the composite benchmark, while Codex with GPT-6.1 Sol offers nearly comparable results using a fraction of the measured resources. Gemini 4 Argon sits between them, combining a strong aggregate score with a notable advantage on long-horizon software tasks.

The original benchmark post presents the three releases as new leaders. The underlying results support their importance, but not a simple three-place podium. Other configurations, including Claude Opus 5.5, also appear near the top.

More importantly, each result belongs to a model, reasoning setting, and agent harness. The leaderboard does not test abstract model intelligence alone. It tests complete coding systems such as Claude Code, Codex, and Antigravity CLI.

That distinction is the center of the story. Teams are no longer choosing only the model with the highest score. They are choosing how much additional time and compute one more benchmark point is worth.

What Changed in the Artificial Analysis Coding Agent Index

The latest results separate benchmark leadership from operational efficiency more clearly than previous model comparisons.

Artificial Analysis evaluates coding agents on end-to-end work rather than isolated code-completion questions. Its index combines repository modification, terminal operation, and codebase comprehension into one score.

Claude Code running Sonnet 5.5 at maximum effort leads with 68 points. Its component results are 72 percent on DeepSWE v1.1, 66 percent on Terminal-Bench 4.0, and 67 percent on SWE-Atlas-QnA.

Antigravity CLI running Gemini 4 Argon scores 64. It reaches 79 percent on DeepSWE, 56 percent on Terminal-Bench, and 56 percent on repository questions.

Codex with GPT-6.1 Sol at xhigh effort scores 63. That configuration posts 73 percent on DeepSWE, 55 percent on Terminal-Bench, and 61 percent on SWE-Atlas-QnA.

Those totals put only five points between Claude and Sol. However, Artificial Analysis measured the Claude configuration using about 8.7 times as many tokens per task. It also ran for nearly six times longer.

Measured API cost shows an even wider divide. The maximum-effort Claude configuration costs about 13.6 times as much per task as xhigh Sol in Codex.

Gemini 4 Argon lands between those extremes. Its measured cost is approximately 5.6 times Sol’s, while its index advantage is one point. It consumes about 4.3 times as many tokens and takes more than twice as long.

The benchmark comparison therefore tells two stories. Claude has the highest composite score, while Sol delivers the strongest measured score-to-cost relationship among these three configurations.

Artificial Analysis also reports several effort settings for Sonnet 5.5. That detail matters because maximum effort is not Claude Code’s default setting.

Sonnet 5.5 at xhigh effort scores 63, matching Sol’s headline score. It uses more than twice Sol’s tokens and has a measured cost more than three times higher.

At high effort, Sonnet scores 55. At medium effort, the default in Claude Code, it scores 46. These results demonstrate how much the reasoning budget changes the product being evaluated.

The same pattern appears within the Sol results. GPT-6.1 Sol at medium effort scores 61, only two points below its xhigh configuration. Its measured cost and runtime also fall substantially.

Maximum reasoning does not automatically produce the best outcome. Sol’s xhigh result exceeds its maximum-effort result by three points in the published evaluation.

That counterintuitive result is a reminder that agent benchmarks contain variance. More inference time can help, but it can also create longer trajectories, unnecessary tool calls, or unproductive reconsideration.

The headline winner remains Claude Code with maximum-effort Sonnet 5.5. The more consequential change is that buyers can now see how steeply its last five points are priced.

The Benchmark Measures Systems, Not Models Alone

A coding-agent score reflects the interaction between a model, its harness, its tools, and its reasoning budget.

The Coding Agent Index v1.5 uses three equally weighted components. According to the published index methodology, each component tests a different part of software work.

DeepSWE v1.1 contains 113 long-horizon tasks. Agents must modify existing repositories, while separate verifier environments judge whether their committed patches pass.

Terminal-Bench 4.0 contains 66 tasks covering areas such as software engineering, machine learning, security, and system administration. Agents work through command-line environments, then test suites grade their results.

SWE-Atlas-QnA contains 124 repository questions. These tasks measure whether an agent can trace unfamiliar code and accurately explain its behavior.

Each task receives three attempts. Artificial Analysis calculates pass-at-one results for each component, then gives the three components equal weight in the composite index.

This structure is broader than a conventional code-generation test. It rewards agents that can inspect repositories, choose tools, operate terminals, maintain context, and recover from mistakes.

It also makes the harness important. A harness is the software layer that connects a model with files, terminals, instructions, context management, and tool execution.

Claude Code, Codex, and Antigravity CLI do not expose identical workflows. They may package context differently, encourage different tool-use patterns, or impose different limits.

Consequently, the benchmark cannot establish that Sonnet 5.5 is always a better coding model than GPT-6.1 Sol. It establishes that one tested Claude Code configuration scored higher than one tested Codex configuration.

The distinction becomes visible in the component scores. Gemini 4 Argon leads the trio on DeepSWE with 79 percent, despite ranking below Claude overall.

Sol narrowly exceeds maximum-effort Sonnet on DeepSWE. Claude builds its overall lead through Terminal-Bench and SWE-Atlas-QnA, where it gains larger margins.

The results therefore describe different capability profiles. Gemini appears strongest on the benchmark’s long-horizon repository changes. Claude looks more balanced across terminal work and repository comprehension.

Sol remains competitive in all three while using fewer measured resources. It does not win an included component against both rivals, but it avoids a severe weakness.

That balance matters for production use. A team maintaining a large repository may value patch completion more than repository question answering. Another team may need reliable terminal work across varied environments.

The index’s single number helps readers scan the field. It should not replace the component results when selecting a tool for a defined workload.

Artificial Analysis also pools token, cost, and time data across the same benchmark suite. Missing telemetry is excluded from the relevant average instead of being treated as zero.

Its cost calculation accounts for ordinary input, cached input, cache writes, reasoning, and output tokens when providers price those categories separately. It represents pay-per-token API cost, not subscription pricing.

The reported cost also excludes several operational expenses. It does not include engineering integration, human review, environment setup, security controls, or the consequences of a faulty patch.

Those exclusions do not weaken the comparison. They define what it can answer: how much model usage the evaluated agents consumed under this test.

They also explain why the cheapest benchmark run will not always produce the cheapest accepted pull request. A weaker result can create additional review, correction, and rerun costs.

Teams should therefore evaluate both direct inference cost and cost per successful outcome. The published index supplies useful ingredients, but it does not calculate that full business measure.

Claude Sonnet 5.5 Wins Performance at a Steep Efficiency Premium

Claude’s lead is real within this benchmark, but maximum effort turns a modest score advantage into a large resource commitment.

Anthropic released Sonnet 5.5 on September 28, 2026. The company positions it as a faster, lower-cost complement to Opus 5.5 for scoped daily tasks, debugging, and document creation.

Anthropic’s model release details emphasize adjustable effort. Lower settings prioritize speed and economy, while higher settings give the model more time to reason and check its work.

The Artificial Analysis results show both sides of that design. Moving Sonnet from medium to maximum effort raises its index from 46 to 68.

That 22-point improvement is substantial. It is accompanied by roughly 21 times more tokens, more than ten times the runtime, and nearly 23 times the measured API cost.

Maximum effort also produces very long agent trajectories. Artificial Analysis records about 266 turns per task and 27.7 million total tokens for the leading configuration.

Those figures do not mean every real task will consume the same resources. They show the average behavior across a demanding benchmark suite with hundreds of task attempts.

They also illuminate how the model wins. The best-performing configuration is not simply producing a smarter answer from the same budget. It is spending much more time interacting with its environment.

That strategy pays off in the composite score. Claude leads Gemini by four points and Sol by five. It also posts the trio’s best results on two of the three component benchmarks.

The premium becomes harder to justify when Claude’s lower effort settings enter the comparison. Sonnet at xhigh matches Sol’s 63-point score but consumes more tokens, time, and measured spending.

At high effort, Claude falls eight points behind Sol xhigh. Its resource use is closer to Sol’s, but the performance gap becomes meaningful.

At medium effort, Claude becomes much cheaper and faster than its maximum configuration. However, its score sits 17 points below Sol xhigh and 18 points below Gemini.

There is no contradiction between these results. Anthropic lets users buy more test-time reasoning, and the benchmark shows that additional reasoning can raise task completion.

The tradeoff concerns scale. An individual developer may accept a long, expensive run for a difficult migration. A company processing thousands of routine changes faces a different calculation.

The best setting may also vary during one workflow. A team can use medium effort for exploration, high effort for implementation, and maximum effort only for stubborn failures.

That routing strategy would preserve access to Claude’s peak capability without applying its highest resource budget to every ticket. It requires measurement and clear escalation rules.

Claude’s benchmark victory is therefore most relevant to tasks where completion quality dominates all other constraints. Examples include difficult cross-repository fixes, fragile migrations, or incidents with high failure costs.

It is less decisive for high-volume maintenance. Dependency updates, small refactors, test generation, and routine bug fixes often reward acceptable quality at predictable cost.

This is why the Artificial Analysis Coding Agent Index should not become a purchasing shortcut. The 68-point result represents a ceiling configuration, not an automatic default.

GPT-6.1 Sol and Gemini 4 Argon Pressure Claude From Different Directions

Sol challenges Claude on efficiency, while Argon challenges it on long-horizon repository work.

OpenAI introduced GPT-6.1 Sol on September 29, one day after Anthropic released Sonnet 5.5. Google followed with Gemini 4 Argon on September 30.

The timing gave Artificial Analysis three new frontier configurations to compare within days. Their benchmark positions reveal more differentiation than their launch descriptions suggest.

OpenAI describes Sol as a near-flagship model for coding, computer use, and professional work at lower cost. Its Sol model card supports five reasoning settings from low through maximum.

In the Coding Agent Index, xhigh is Sol’s best tested setting. It scores 63, while maximum effort scores 60.

That result undermines the assumption that the largest reasoning budget is always safest. It suggests teams should benchmark effort settings instead of selecting the highest label by default.

Sol’s main advantage is consistency per unit of resource. The xhigh configuration finishes an average task in about 15.5 minutes and consumes 3.2 million tokens.

Maximum-effort Claude needs about 90 minutes and 27.7 million tokens. Gemini requires about 34.5 minutes and 13.7 million tokens.

Sol also performs competitively on each component. Its 73 percent DeepSWE result exceeds Claude’s 72 percent, though it trails Gemini’s 79 percent.

Its Terminal-Bench and repository-question scores remain below Claude. Those deficits create the five-point composite gap.

For many organizations, that gap will be acceptable. Sol’s lower resource use permits more attempts, wider deployment, or additional verification within the same budget.

The comparison does not establish that Sol is universally more economical. Provider pricing can change, caching patterns differ, and internal workloads may produce different token distributions.

It does establish a strong hypothesis worth testing. If a team’s tasks resemble the benchmark, Codex with Sol may deliver a better cost-performance balance than maximum-effort Claude.

Gemini 4 Argon creates a different kind of pressure. Google introduced Argon as a model for sustained reasoning across complex professional workflows.

Google’s Argon announcement describes internal uses involving code migration, memory optimization, research, and cybersecurity. Those examples remain company claims unless independently reproduced.

The Coding Agent Index adds third-party evidence for one part of that story. Argon’s 79 percent DeepSWE result is the strongest among the three highlighted systems.

That outcome aligns with Google’s focus on long-horizon work. It suggests Argon deserves attention for extended repository changes, even though Claude leads the overall index.

Argon’s weaker repository-question score pulls down its composite result. Its 56 percent result trails Sol by five points and Claude by eleven.

The model also lacks Sol’s measured efficiency. Argon gains one aggregate point over Sol while requiring more than four times as many tokens per task.

That does not make the configuration irrational. A higher DeepSWE completion rate may outweigh resource use for organizations facing difficult implementation work.

The important question is workload matching. Sol appears attractive as an efficient generalist, while Argon offers a stronger signal on long-duration repository modification.

Claude remains the balanced performance leader at its most aggressive setting. The market pressure comes from rivals that make different portions of that lead less valuable.

This is a healthier competitive picture than one universal ranking. It gives engineering teams distinct options instead of three nearly interchangeable model brands.

It also raises the importance of maintaining portable workflows. Teams should avoid tying prompts, review practices, and context preparation to one model unless the benefit is measurable.

A searchable record of requirements, decisions, and prior changes can make those comparisons more consistent. Teams can use an engineering knowledge base to preserve that context across agent trials.

The goal is not to change models every week. It is to make switching and evaluation possible when the performance frontier moves.

What the Numbers Do Not Prove

A five-point benchmark lead does not guarantee better code, safer deployment, or lower total engineering cost inside a real organization.

Artificial Analysis publishes more methodological detail than many leaderboard operators. Its component tasks, attempt counts, scoring methods, and efficiency definitions are documented.

Even so, a benchmark remains a sample. It cannot represent every language, repository shape, dependency environment, security policy, or review standard.

The index weights its three components equally. A real company rarely values repository questions, terminal operations, and patch completion in exactly equal proportions.

One organization may spend most of its time on TypeScript services with extensive tests. Another may maintain embedded C code, data pipelines, or regulated financial systems.

Their internal ranking can differ from the public leaderboard. A model that excels on DeepSWE may still struggle with proprietary frameworks or poorly documented legacy code.

Pass-at-one scoring also compresses important quality differences. Two patches can both pass an automated verifier while differing in maintainability, security, readability, or architectural fit.

The reverse can happen as well. A useful partial solution may fail one verifier condition and receive the same binary outcome as an unusable attempt.

SWE-Atlas-QnA introduces another dependency. Artificial Analysis uses an automated judge to decide whether repository answers satisfy all required criteria.

Automated judging supports evaluation at scale. It can still inherit ambiguity, model bias, or grading errors, especially on explanations with several valid formulations.

The benchmark’s pooled averages also hide dispersion. Average cost does not reveal whether most tasks are predictable while a small group creates extremely long trajectories.

That variance matters for budgeting. A service can tolerate a moderate average while still suffering from individual runs that consume excessive tokens or occupy environments for hours.

Agent behavior can also change after product updates. Tool selection, context compression, retry logic, and hidden system instructions may shift without a new public model name.

For that reason, the benchmark should be treated as a dated measurement. It is not a permanent property of Claude Code, Codex, Antigravity CLI, or their underlying models.

Maximum-effort Claude illustrates the risk of reading a ceiling result as a default experience. The benchmarked configuration is far more resource-intensive than Claude Code’s medium default.

The index also compares pay-per-token API spending. Subscription limits, negotiated enterprise rates, regional processing, and internal infrastructure can alter a team’s actual economics.

Human costs are missing too. A slower agent may be acceptable if it works asynchronously. A faster agent may be more valuable when a developer is waiting for feedback.

Review burden is another unresolved variable. An inexpensive patch that needs extensive inspection can cost more overall than an expensive patch accepted after a short review.

Security deserves similar caution. None of the headline scores alone establishes that an agent follows least-privilege access, resists malicious repository instructions, or avoids leaking sensitive context.

Google has limited Argon’s initial availability while conducting staged safety work. That rollout means its public usage evidence may remain thinner than benchmark attention implies.

Vendor claims also require careful attribution. Anthropic, OpenAI, and Google each highlight favorable evaluation results from different suites and settings.

Those results can be accurate without being directly comparable. Different harnesses, task sets, budgets, and scoring rules often produce different leaders.

The Artificial Analysis benchmark improves comparability by running configurations in one framework. It cannot remove every difference introduced by proprietary agents and model interfaces.

Engineering leaders should reproduce a small internal trial before standardizing. A useful test set includes completed tickets, known failure cases, and representative repository constraints.

Reviewers should grade correctness, unnecessary changes, security, test coverage, explanation quality, and time to acceptance. Token spending should be recorded beside those outcomes.

The resulting metric should be accepted work per dollar or accepted work per engineer-hour. A public composite score can guide candidate selection, but it cannot replace that measurement.

Three Signals Will Decide Whether Claude’s Lead Matters

The next phase will be decided by default-setting performance, accepted-change economics, and benchmark stability across updates.

The first signal is performance at practical effort settings. Maximum configurations attract headlines, but defaults shape most daily use.

Sonnet 5.5 at medium effort scores far below its maximum result. Sol loses only two points when moving from xhigh to medium in the published data.

If Anthropic narrows that default-setting gap, Claude’s 68-point ceiling will become more relevant to ordinary teams. If the gap persists, Sol’s efficiency case will strengthen.

The second signal is cost per accepted change. Public benchmarks currently measure per-task API spending, not the complete path from request to merged code.

Teams should watch whether vendors or independent evaluators publish review-adjusted results. These should include reruns, human correction time, and regressions discovered after verification.

Claude’s premium becomes easier to defend if its patches require less review. Sol’s advantage becomes stronger if its lower inference use does not create additional correction work.

Argon could lead this measure on complex repository changes if its DeepSWE strength transfers into production. Its aggregate score alone cannot answer that question.

The third signal is ranking stability. Coding agents change through model updates, harness revisions, tool policies, and context-management improvements.

A stable leader should retain its position across repeated runs and benchmark versions. Large changes after small system updates would reduce confidence in narrow score differences.

Artificial Analysis already publishes component results, efficiency metrics, and methodology revisions. Future reruns will show whether the five-point spread represents durable separation or temporary configuration effects.

Development teams do not need to wait for a perfect benchmark. They can make a bounded decision now.

Start with a representative internal task set. Compare Claude at more than one effort level against Sol and Argon where access permits.

Keep the agent permissions, repository snapshot, and success criteria consistent. Record wall time, tokens, failures, review time, and whether the final change was accepted.

Use a high-effort configuration only when the task justifies escalation. Routine work should begin with the least expensive setting that meets the team’s acceptance threshold.

Recheck the comparison after major model or harness updates. The Artificial Analysis Coding Agent Index is useful precisely because the frontier is moving.

For now, its message is clear. Claude Sonnet 5.5 owns the highest published score among the three new configurations, but it does not own every practical definition of first place.

GPT-6.1 Sol offers a compelling efficiency profile, while Gemini 4 Argon leads the trio on long-horizon repository work. The right choice depends on which result a team values.

Would your organization pay a large resource premium for five additional index points, or fund more attempts and verification with Sol? Test that question against your own merged work before choosing a default.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page