Anthropic GPT Race Splits the Benchmarks as Astra Resets the AGI Clock
- Sophie Larsen

- 4 days ago
- 15 min read
OpenAI’s GPT-6 Astra reached 169 on one model index but only 61 on another, intensifying the anthropic GPT rivalry around benchmark leadership. Epoch AI placed Astra first among 267 evaluated models. Artificial Analysis instead ranked it behind Anthropic’s Claude Fable 5.1.
That disagreement would usually produce another inconclusive model launch. Astra’s ARC-AGI-3 result makes this case harder to dismiss. The model scored 62.7 percent through ARC Prize’s standard harness, up from GPT-5.6 Sol’s reported 7.8 percent.
Astra also completed the interactive tasks with greater action efficiency than an average human, according to ARC Prize’s evaluation. François Chollet called its progress about twice as fast as he had expected. He still stopped short of declaring AGI.
The central issue is therefore not whether one leaderboard picked the wrong winner. It is whether current benchmarks measure the same kind of intelligence at all. The anthropic GPT contest now pits broad test performance against adaptive efficiency in unfamiliar environments.
GPT-6 Astra produced two conflicting verdicts
Astra’s launch turned model ranking into a dispute over what evaluators choose to measure.
OpenAI introduced GPT-6 Astra on September 3, 2026. The company positioned it as a model for computer use, browsing, software engineering, science, cybersecurity, and professional work.
Its Astra announcement described state-of-the-art results across several company-selected evaluations. OpenAI also reported 99.9 percent on ARC-AGI-3 through a provider-specific evaluation setup.
Independent aggregate scores presented a less coherent picture. Epoch AI’s capability index gave Astra 169 points, ahead of Claude Fable 5.1 at 163. GPT-5.6 Sol and Claude Opus 5 each received 162.
Epoch’s index combines results across many capability domains. These include mathematics, science, knowledge, coding, and other demanding tasks. Its score attempts to summarize broad progress rather than isolate one narrow skill.
On that composite, Astra leads. The result supports OpenAI’s argument that the model represents a meaningful generation change. It also suggests that Astra’s strongest improvements sit in areas emphasized by Epoch’s evaluation mix.
Artificial Analysis reached a different conclusion. Its Intelligence Index gave Astra 61, level with GPT-5.6 Sol. Claude Fable 5.1 led the comparison with 66, while Claude Opus 5 scored 63.
The model comparison covers knowledge, coding, mathematics, and text comprehension. Its composition and testing procedures differ from Epoch’s index. The resulting numbers cannot be treated as two measurements of an identical object.
Astra’s individual results were also uneven. It improved strongly in several agentic and scientific tasks. Artificial Analysis nevertheless recorded weaker performance in areas that included long-context reasoning and certain professional evaluations.
Its coding picture was similarly mixed. Astra reportedly reached 67 on the Coding Agent Index, while Claude Fable 5.1 reached 70. Astra used substantially fewer reasoning steps than several competing models during comparable tests.
That efficiency matters because a model’s practical burden extends beyond its published token rate. A model using fewer tokens, fewer actions, or less elapsed time can be preferable despite a lower headline score.
The launch therefore created two legitimate but incomplete stories. Epoch says Astra has the strongest broad capability profile. Artificial Analysis says it does not lead across a fixed collection of common reasoning tasks.
Neither verdict alone settles which model is best for developers or enterprises. A software team values reproducible coding performance. A research group may prioritize mathematical reasoning, tool use, or novel problem solving.
The timing also limits firm conclusions. Epoch had fewer recorded coding evaluations for Astra than for some established models when the early comparison appeared. At least one available Astra coding run used a medium reasoning setting.
Early leaderboards often mix mature results with incomplete coverage. New models receive additional tests, harness updates, and configuration changes after launch. Rankings can move even when the underlying model remains unchanged.
That makes the 169 versus 61 contrast informative, not absurd. The numbers reveal different evaluation priorities. They also warn buyers against treating any aggregate score as a universal measure of intelligence.
The anthropic GPT competition looks different depending on the workload placed at the center. Claude leads Artificial Analysis’s broader text-oriented index. Astra leads Epoch’s current composite and posts unusual gains on interactive reasoning.
The more consequential question begins where both aggregate indexes become less descriptive. ARC-AGI-3 does not ask a model to retrieve an answer or complete a familiar coding exercise. It asks the system to discover how an unknown world works.
ARC-AGI-3 changed the meaning of the comparison
Astra’s strongest result concerns adaptation under uncertainty, not ordinary question answering.
ARC-AGI-3 evaluates agents inside unfamiliar interactive environments. A system must explore, infer rules, remember observations, plan actions, and recognize goals without receiving normal task instructions.
The benchmark limits dependence on memorized facts and language knowledge. Its environments function more like small unknown games than conventional exams. Success requires learning through interaction rather than recognizing a familiar prompt pattern.
That distinction matters because large language models can score well through extensive prior exposure. A benchmark built around unknown environments tries to reduce that advantage. It tests whether a system can acquire useful behavior during evaluation.
The ARC-AGI-3 report described exploration, planning, memory, goal acquisition, and alignment as central requirements. When the benchmark appeared in March 2026, tested frontier systems scored below 1 percent.
Humans could solve every evaluated environment with no task-specific training. Human solvers were not equally fast or efficient, but the benchmark established that each environment was understandable through interaction.
Six months later, Astra reached 62.7 percent on the semi-private set using ARC Prize’s standard harness. The verified results list this run at maximum reasoning effort.
That score was more than eight times GPT-5.6 Sol’s reported 7.8 percent. It also exceeded Claude Opus 5’s listed 30.2 percent by more than twofold. Comparable data for Claude Fable 5.1 was not initially available.
The result is important because the model did not merely solve more environments. According to ARC Prize’s analysis, Astra used fewer actions than the average human on successfully completed tasks.
Action efficiency measures how directly an agent reaches a solution. An inefficient system may succeed only after extensive random exploration. A more efficient system builds an accurate internal account with fewer moves.
Astra reportedly crossed the average human efficiency level on this measure. That does not mean it surpassed humans across general intelligence. It means its successful ARC-AGI-3 trajectories required fewer actions than the human average.
The distinction protects the finding from exaggeration. Humans still demonstrated complete coverage of the environments. Astra’s standard-harness success rate remained well below 100 percent.
OpenAI also reported a 99.9 percent result through a provider adapter harness. This setup preserves the model’s opaque reasoning state across requests and compacts longer conversations. The standard harness only carries forward notes selected by the model.
The higher score therefore measures Astra with infrastructure unavailable to all providers. ARC Prize warns against using it for direct cross-vendor ranking. It is evidence about one system configuration, not a neutral Claude comparison.
The gap between 62.7 and 99.9 percent is itself revealing. A model’s apparent intelligence depends partly on how memory, context, tools, and hidden reasoning state surround it.
That creates a new evaluation problem. Buyers increasingly deploy complete agent systems, not isolated language models. Yet provider-specific infrastructure can make fair comparison difficult or impossible.
Astra’s result still represents a sharp change under the standard setup. The leap from Sol’s 7.8 percent cannot be explained by the provider adapter. Both figures came from the more comparable evaluation route.
ARC-AGI-3 also differs from the older ARC versions. ARC-AGI-1 used static visual transformations that became increasingly saturated. ARC-AGI-2 increased difficulty while retaining an abstract puzzle format.
Astra scored 95 percent on ARC-AGI-2, compared with Sol’s 92.5 percent. That smaller difference would not support a dramatic intelligence claim. The interactive third generation creates the stronger separation.
This history explains why Chollet focused on Astra’s ARC-AGI-3 performance. The model showed a new level of competence on tasks designed around unfamiliarity. That is closer to his definition of intelligence than excellence on rehearsable skills.
The benchmark result does not end the AGI debate. It changes its strongest evidence. Instead of asking whether a model knows enough, evaluators can ask how efficiently it learns when its existing knowledge is insufficient.
The anthropic GPT rivalry now turns on efficiency
The competitive pressure comes from how much reasoning a model needs, not only how many answers it gets right.
Anthropic’s Claude models remain formidable competitors in coding and professional work. Claude Fable 5.1 led Astra on the Artificial Analysis Intelligence Index and its Coding Agent Index.
Astra’s advantage appears in a different dimension. It reportedly matched Claude Fable 5 on a comparable coding result while using fewer computational steps. It also required fewer steps than GPT-5.6 Sol and Claude Opus 5.
Fewer reasoning steps can reduce the total resources consumed by a task. That benefit can offset a higher per-token rate. It can also shorten waiting times and reduce the number of opportunities for an agent to wander.
The distinction resembles fuel efficiency. A vehicle’s fuel price does not reveal the expense of a completed journey. Distance, consumption, and reliability determine the real outcome.
AI buyers face the same accounting problem. Published input and output rates describe units of consumption. They do not show how many units a model needs to finish useful work.
Astra reportedly consumes fewer reasoning tokens than Sol on several agentic tasks. The Decoder’s comparison estimated that this efficiency reduced the difference in completed-task costs, despite Astra’s higher unit rate.
No price figures are needed to understand the strategic pressure. Anthropic cannot answer Astra simply by winning a broad intelligence index. OpenAI cannot claim victory while Claude retains advantages in important coding tests.
Both companies must demonstrate completed work under consistent conditions. That includes success rate, total compute use, latency, human corrections, and recovery from failed actions.
This shifts the anthropic GPT contest away from chatbot impressions. Agentic systems operate across files, browsers, terminals, and business software. Their failures can propagate through a workflow before a user notices.
An efficient but unreliable agent creates little value. A highly accurate agent can also become impractical if it requires excessive time and computation. The leading system must manage both constraints.
Astra’s interactive efficiency increases pressure on Anthropic because Claude built much of its reputation around coding and controlled tool use. A clear Astra advantage in adaptive computer work would reach that core market.
Claude Fable 5.1’s index lead creates equal pressure on OpenAI. It prevents Astra’s highest scores from becoming an uncontested claim of overall leadership. Customers can point to an independent evaluation where the predecessor-level result remains visible.
Enterprise teams should treat these outcomes as workload hypotheses. A research workflow centered on mathematics may favor one model. A repository migration or support process can produce a different ranking.
Teams also need records from their own evaluations. Saving prompts, outputs, corrections, and decisions in a searchable AI knowledge base makes model comparisons easier to audit.
Without that record, buyers risk evaluating models through memorable successes. Impressive demonstrations attract attention, while quiet failures disappear into normal work.
The same problem affects benchmark interpretation. Aggregate scores hide task distributions. A five-point index lead does not mean a model wins every evaluation or delivers better economics for every workflow.
Astra’s Coding Agent Index result illustrates the issue. Claude Fable 5.1 led on the score, but Astra reportedly needed fewer steps. The ranking changes if an organization values raw completion above cost, or efficiency above a small score difference.
The ARC-AGI-3 result intensifies this debate because it connects efficiency with adaptation. Astra did not only conserve steps on a familiar programming task. It used them effectively while discovering unknown rules.
That combination is relevant to real agent deployments. Business software changes, websites present unexpected states, and repositories contain undocumented conventions. A useful agent must adapt without receiving a perfect instruction manual.
However, ARC’s small game worlds remain abstractions. They remove many complications found in organizations, including ambiguous goals, access controls, social judgment, and consequences for incorrect actions.
The responsible conclusion is narrower. Astra supplies evidence that adaptive reasoning and action efficiency improved sharply. Whether those gains transfer to sustained professional work still requires independent testing.
The winner of the anthropic GPT race will therefore depend on deployment evidence. Benchmark leadership can earn a trial. Reliable, efficient task completion determines whether the model stays.
Chollet moved his forecast without declaring AGI
Chollet’s reaction recognizes faster progress while preserving a demanding definition of general intelligence.
François Chollet created the original Abstraction and Reasoning Corpus in 2019. He designed it to test skill acquisition rather than accumulated knowledge alone.
His argument has long challenged standard scaling narratives. A system can become better at many familiar tasks without gaining broad adaptive intelligence. Training coverage and benchmark preparation can imitate generalization.
ARC tasks try to expose that difference. They minimize reliance on language, stored facts, and recognizable professional formats. A model must infer a task from the evidence available during evaluation.
When ARC-AGI-3 launched, Chollet reportedly expected a frontier model to saturate it in about one year. Astra arrived approximately six months later and nearly saturated it under OpenAI’s specialized harness.
Chollet characterized the progress as roughly twice as fast as expected. He also described Astra as a step change for interactive reasoning problems.
Those statements moved his AGI forecast forward, but they did not establish a date. They also did not treat one benchmark result as sufficient proof.
Chollet’s reasoning matters more than the headline prediction. ARC-AGI-3 tests qualitative properties associated with general intelligence in small quantities. These include exploration under uncertainty, causal modeling, and adaptation without explicit instructions.
A model showing those properties in constrained environments has crossed a meaningful threshold. It has not necessarily shown the same competence across the open-ended world.
This distinction separates capability evidence from an AGI label. AGI has no universally accepted operational definition. Companies, researchers, and contracts can apply different thresholds.
OpenAI President Greg Brockman offered a more expansive interpretation at launch. He said he personally believed the company had reached AGI, while leaving users to decide whether Astra satisfied their definitions.
That rhetoric raises the stakes around every benchmark. Calling a model AGI invites scrutiny beyond model rankings. It creates expectations about reliability, autonomy, transfer, and performance across unfamiliar domains.
ARC-AGI-3 supports part of that narrative. Astra adapted effectively in environments designed to resist direct memorization. Its standard-harness score also left a substantial gap below complete human coverage.
Other results reinforce the mixed picture. Astra reportedly became the only model in Epoch’s standardized FrontierMath Erdős evaluation to solve two of 68 open problems with machine-checked proofs.
Lean verification means formal software checked the logical validity of those proofs. That makes the result harder to inflate through persuasive but incorrect prose.
Three additional solutions reportedly emerged from nonstandard runs using far greater computation. Epoch excluded those attempts from the scored comparison because they did not follow the same evaluation budget.
This boundary is essential. A model’s best observed result can demonstrate possibility. A standardized run provides a better basis for comparison.
Astra’s formal mathematics performance suggests stronger search and verification abilities. Its weaker results on some professional and long-context tasks show that those abilities do not transfer uniformly.
An AGI declaration must address that unevenness. General intelligence is not identical to a collection of record scores. It implies a capacity to adapt reliably when the environment, objective, and evidence change together.
The older ARC story offers a warning. OpenAI’s o3 produced a major result on ARC-AGI-1, generating similar debate about whether a threshold had been crossed.
Researchers later emphasized cost, task exposure, and the difference between benchmark mastery and general intelligence. ARC-AGI-2 and ARC-AGI-3 were created partly to restore evaluation headroom.
Astra’s performance compressed that headroom faster than Chollet expected. This is the real forecast change. Benchmark designers now have less time to build tests that remain informative at the frontier.
The lesson is not that evaluation has failed. It is that useful benchmarks require continual renewal. Once a task becomes saturated or directly targeted, it loses some ability to distinguish general adaptation from specialized optimization.
The benchmark gaps are the strongest reason for caution
Astra’s results become less definitive when harnesses, coverage, and aggregate methods are examined together.
The first uncertainty concerns evaluation scope. Epoch’s 169 score summarizes a broad but evolving collection. Artificial Analysis’s 61 reflects a different fixed mixture and weighting.
A composite index can change when a model lacks results in one domain. It can also emphasize areas where a new model received extensive launch testing.
Epoch’s early Astra record reportedly contained limited coding coverage. Claude Fable 5.1 held the best results across several coding evaluations. Comparing their headline totals therefore risks hiding unequal evidence.
The second uncertainty concerns harness design. A harness controls how a model receives context, preserves state, uses tools, and submits answers.
ARC Prize’s standard harness allows selected notes to persist. OpenAI’s provider adapter preserves opaque reasoning state and compacts long interactions. These are materially different operating conditions.
The 99.9 percent score shows what Astra can do with its native infrastructure. It should not be presented as a direct victory over models tested through the standard interface.
This problem will grow as model providers optimize complete systems. Hidden reasoning, proprietary memory, routing, and tool policies can improve results while reducing independent reproducibility.
The third uncertainty concerns task transfer. ARC-AGI-3 isolates adaptation inside small interactive worlds. This makes the test scientifically useful but operationally incomplete.
A business agent must handle permissions, conflicting instructions, unclear goals, and changing external systems. It must also know when to stop and request human judgment.
Astra’s move efficiency does not verify those behaviors. Completing a puzzle with fewer actions differs from editing production data without causing damage.
OpenAI’s safety material adds another reason for careful testing. The company says Astra received stronger robustness and alignment training than Sol. It also acknowledges greater difficulty in monitoring advanced reasoning.
Its safety overview discusses controls for cybersecurity and tool-using deployments. Company safety claims remain provider-reported until wider independent evaluation becomes available.
More capable computer use raises both value and exposure. An agent that navigates software efficiently can automate long workflows. The same autonomy increases the consequences of misunderstood instructions or weak access controls.
A separate concern involves benchmark targeting. Chollet’s original forecast included a condition about how directly labs pursued ARC-AGI-3.
A model trained or tuned around interactive reasoning can improve rapidly without gaining equivalent ability elsewhere. That does not invalidate the result. It limits how widely the improvement should be generalized.
Data contamination is less straightforward for interactive worlds than for static questions. A provider can still optimize around benchmark principles, interfaces, or development feedback.
Repeated evaluation also creates selection effects. The best configuration among many attempts may look much stronger than the expected result from an ordinary deployment.
The ARC Prize page lists multiple Astra harness configurations. Buyers should distinguish the best observed run from repeatable performance at a fixed setting.
Cost budgets further complicate comparisons. Astra’s 62.7 percent standard-harness run used a substantial evaluation budget. The provider-adapter result cost less overall but relied on a specialized setup.
The relevant question is not whether that spending was justified. It is whether another model received equivalent opportunities, settings, and infrastructure.
Human efficiency comparisons need similar care. Astra used fewer moves on successful tasks than the average human. Humans solved a wider share of the benchmark’s environments.
Efficiency conditioned on success is not the same as overall superiority. A system can move directly when it understands a task but fail entirely on other tasks.
Reporting should preserve both facts. Astra crossed the average human action-efficiency threshold. It did not match complete human coverage through the standard harness.
The benchmark disagreement therefore carries more value than either ranking alone. It exposes the dimensions hidden by the word “best.”
Astra appears stronger at adaptive reasoning, selected scientific tasks, and efficient agentic work. Claude Fable 5.1 retains advantages on an important independent intelligence index and several coding measures.
That uncertainty should shape deployment. Organizations need representative tasks, fixed budgets, repeated trials, and logged human interventions. They should test failure recovery as carefully as first-attempt success.
Teams can preserve those results through a documented AI workflow. A durable evaluation record makes later model changes easier to compare.
No public benchmark can reproduce every organization’s constraints. The gaps between benchmarks are not noise to be averaged away. They are information about where each model’s capabilities concentrate.
Three signals will decide whether Astra changed the race
The next test is whether Astra’s adaptive efficiency survives neutral evaluation and ordinary professional use.
The first signal is broader provider-neutral ARC-AGI-3 testing. Claude Fable 5.1, newer Gemini models, and other frontier systems need comparable runs through the standard harness.
Astra’s lead becomes stronger if competitors remain near Claude Opus 5’s reported range. It weakens if other models approach 62.7 percent once evaluated under equal conditions.
Repeated Astra runs also matter. A stable distribution would show that the result reflects dependable capability. Wide variation would suggest that the best observed score overstates normal performance.
The second signal is expanded independent coding and computer-use coverage. Epoch’s index needs more Astra results across consistent reasoning settings.
The model’s 169 score gains credibility if added coding tests preserve its lead. It becomes less persuasive if fuller coverage pulls Astra toward Sol or below Claude Fable 5.1.
Real computer-use studies should report task completion, elapsed time, interventions, and damage from mistakes. Token totals alone cannot capture the operational burden of a failed agent.
Independent evaluators should also separate native provider systems from common harnesses. Both are useful, but they answer different questions.
Native evaluations show the best product a customer can access. Standardized evaluations isolate more of the underlying model and enable fairer competition.
The third signal is evidence from sustained deployments. OpenAI initially rolled Astra out to a limited set of organizations before broader availability.
Those users can test whether interactive reasoning transfers to repositories, browsers, research processes, and enterprise systems. The strongest reports will document repeated tasks rather than selected demonstrations.
Astra strengthens the AGI interpretation if it learns unfamiliar workflows with few examples and limited human correction. It weakens that interpretation if users must extensively prepare environments around it.
Watch especially for failure recovery. General adaptation requires more than finding a successful path. A system must detect mistakes, revise assumptions, and avoid repeating harmful actions.
The anthropic GPT rivalry will sharpen as both providers publish agent-focused results. Claude’s response does not need to exceed every Astra score. It needs to show where its reliability or efficiency produces better completed work.
OpenAI faces the same burden. An ARC-AGI-3 leap creates a credible claim about adaptive reasoning. It does not erase Claude’s lead on Artificial Analysis or Astra’s weaker individual results.
Chollet’s revised timeline is therefore best understood as a warning to update expectations. Frontier systems are acquiring interactive reasoning faster than one prominent benchmark designer anticipated.
That does not produce a settled AGI date. It shortens the period during which existing tests can separate frontier systems from adaptive human behavior.
Developers should respond by evaluating models against unknown states, not polished prompts alone. Enterprise buyers should measure corrections and supervision alongside accuracy. Knowledge workers should ask whether an agent can recover when its first interpretation fails.
The important question is no longer which leaderboard deserves absolute trust. It is which evaluation resembles the work being delegated, and whether its conditions match the intended deployment.
Astra currently owns the most striking result on unfamiliar interactive tasks. Claude Fable 5.1 owns the stronger score on another major independent index. Epoch places Astra ahead overall.
Those findings can all be true. Together, they show that frontier model competition has outgrown a single number.
Choose a representative workflow, preserve the evidence, and retest it when neutral ARC results and broader coding evaluations arrive. That approach will reveal more than another argument about whether Astra already deserves the AGI label.


