Yann LeCunがMetaのLlama 4ベンチマークを改ざんしたことを確認。スキャンダルは一つのモデルより大きい。
- Aisha Washington

- 6月5日
- 読了時間: 8分
2026年1月、Metaの退任するチーフAIサイエンティストであり、現代のディープラーニングの先駆者の一人であるYann LeCunは、Financial Timesに対してLlama 4の歴史を書き換えるような発言をした。2025年4月にモデルをリリースした際にMetaが公開したベンチマーク結果は「少しfudgedされていた」という。チームは「異なるテストに異なるモデルを使った」と彼は述べた。
各ベンチマークに異なるモデル。標準的なスイート全体で公平にテストされた単一のモデルではない。Metaは各個別テストで最も高いスコアを出したLlama 4のバリアントを選び、結果を一つの表にcherry-pickし、あたかも一つのモデルがすべてを達成したかのように提示した。独立したテスターによると、パブリックリリースのスコアは主張された数値よりも大幅に低かった。
この告白は内部告発者や調査ジャーナリストによるものではなかった。MetaのAI研究部門を築いた科学者自身によるものだった。その影響は一つの会社のモデルを超えて広がる。
What Happened
Llama 4は2025年4月5日にScoutとMaverickの2つのバリアントでリリースされた。Metaのブログ投稿では、モデルがOpenAIやGoogleのクローズドソース競合他社と「同等かそれ以上」のパフォーマンスを発揮したと主張していた。ベンチマーク表は印象的だった。コミュニティは数日以内に懐疑的になった。
Independent testers running their own evaluations consistently found lower scores than Meta had published. The discrepancies were not marginal. On key reasoning and coding benchmarks, community-run tests showed Llama 4 performing significantly below Meta's claimed levels. On LMSys Chatbot Arena, the most visible public benchmark, Maverick dropped from 2nd place to 32nd. Scout fell out of the top 100 entirely.
Metaの最初の反応は否定だった。MetaのVPであるAhmad Al-Dahleは、Metaの内部テスト環境とパブリックデプロイメントの間の「cloud differences」に不一致を帰した。この説明に納得した人はほとんどいなかった。コミュニティはすでに問題に名前をつけていた。Metaは一つのモデルをテストしていなかった。複数のモデルをテストし、各々の最高スコアを報告していたのだ。
LeCunの2026年1月のFinancial Timesとのインタビューでそれが確認された。「結果は少しfudgedされていた」と、he told the FT。彼が率いていなかったと強調したチームは、「特定のベンチマークに異なるバージョンのモデルを使い、公平な評価の原則を完全に違反していた」。方法は単純だった。複数のチェックポイントを訓練し、各ベンチマークで実行し、テストごとに最高スコアを選択し、一つの表にまとめる。実際にはどの単一のモデルも公開された結果を達成していなかった。
誰も解雇されなかった。Metaは代わりに、元Scale AI CEOのAlexandr Wangの下にSuperintelligence Labsを設立し、Llama 4を監督していたAI研究リーダーシップを実質的に交代させた。組織はsplit into four divisions: TBD Lab for foundation models under Wang himself, FAIR for research, Products, and Infrastructure. Four reorganizations in six months. LeCun departed to launch his own AI startup. The Llama brand continued, but under new management and with permanently damaged credibility.
Why It Matters
Llama 4のスキャンダルは、一つのモデルが一組のベンチマークで不正をしたことではない。それはまさにこの行動を誘発する評価経済全体に関するものだ。
AI業界はリーダーボードで動いている。企業はベンチマークスコアに基づいてモデルを選択する。投資家はリーダーボードの順位に基づいてAIラボを評価する。研究者はベンチマークの改善に基づいてキャリアを築く。それらのベンチマークを操作するインセンティブは構造的なものであり、偶発的なものではない。Metaの罪はfudgedしたことではなく、捕まり、自らのチーフサイエンティストがそれを確認したことだ。
The scale of the leaderboard economy is staggering. One analysis estimated that a $10 billion industry has been built on leaderboard gaming, where model selection strategies are driven by scores that may not reflect real-world performance. When the scores are unreliable, every downstream decision built on them, which model to use, which lab to invest in, which API to integrate, is also unreliable.
Meta was not unique in optimizing for benchmarks. It was unique in getting caught and having the chief scientist confirm it. Every major AI lab optimizes for benchmarks. The difference is that most labs have not had their most senior researcher publicly admit that the optimization crossed the line into manipulation. The community's response on r/LocalLLaMA was immediate and lasting: trust in official benchmark numbers has been permanently eroded. Independent, community-run evaluations are now the de facto standard for serious model comparison.
And Meta's structural response confirms the systemic nature of the problem. The company did not punish anyone for the manipulation. It created a new organization under new leadership. The message was clear: the Llama 4 team's execution was the problem, not the underlying incentives that produced it. Those incentives remain unchanged. Every AI lab today faces the same pressure Meta faced: publish numbers that beat the competition, or lose funding, talent, and market position to the labs that do.
The Real Problem Is the Benchmark, Not the Model
Benchmarks are not neutral measurements. They are targets. And when you make something a target, people will aim at it.
AI業界のベンチマーク文化は perverse cycle を生み出している。ベンチマークが公開される。ラボはそのベンチマークにモデルを最適化する。スコアが向上する。ベンチマークが飽和する。新しいベンチマークが作成される。サイクルが繰り返される。あらゆる段階で、最適化は現実のものだが、一般化は疑わしい。推論ベンチマークで95%のスコアを出したモデルでも、ベンチマークの分布とはわずかに異なる推論タスクでは失敗する可能性がある。
Llama 4 did not invent this problem. It exposed it. LeCun's confession matters precisely because it came from inside the system. He was not an external critic with an agenda. He was the person responsible for Meta's AI research for more than a decade. And he still said the results were fudged. When the insiders stop believing the benchmarks, the rest of us should stop too.
What would structural reform look like? Independent, third-party evaluation bodies that labs do not control. Dynamic benchmarks that change with each evaluation cycle, making overfitting impossible. Mandatory disclosure of the exact model version used for each published score. And a cultural shift that values real-world reliability over leaderboard position. None of these reforms are technically difficult. All of them are politically difficult, because they threaten the marketing machinery that the AI industry has built around benchmark supremacy.
The community has already begun building alternatives. LMSys Chatbot Arena and similar platforms have gained credibility because they are harder to game than static benchmarks. Head-to-head blind comparisons judged by human preference are replacing automated accuracy metrics. But even arenas have their own manipulation vectors. The Llama 4 scandal demonstrated that the arena manipulation economy is real and growing. The only evaluation that cannot be gamed is the one you run yourself, on your own data, for your own use case. Everything else is marketing.
What Happens to Meta's AI Credibility
Meta built its AI reputation on open-source leadership. Llama 1, 2, and 3 established the company as the champion of open-weight models, the counterweight to OpenAI and Google's closed ecosystems. Llama 4 was supposed to extend that legacy. Instead, it became the symbol of everything the open-source community distrusts about corporate AI: the benchmarks are rigged, the transparency is selective, and the claims cannot be verified without independent testing.
The Alexandr Wang era represents a direct repudiation of the culture that produced the scandal. Wang, who built Scale AI into a data infrastructure company valued at $14 billion, was hired as Meta's Chief AI Officer in June 2025. Superintelligence Labs is his organization. The message embedded in the four-division structure, TBD Lab for models, FAIR for research, Products for deployment, Infrastructure for compute, is that Meta now treats AI as a product engineering problem rather than a research problem. The era of "ship whatever the researchers produce and hope the benchmarks hold up" is over.
The historical parallel is uncomfortable. Volkswagen was a company known for engineering excellence that got caught systematically manipulating test results. The parallel is imperfect, benchmark gaming is not illegal in the way emissions cheating was, but the structural dynamic is identical: optimize for the test, not for the real world, and hope nobody notices. Volkswagen's reputation never fully recovered. Meta's AI credibility may follow the same trajectory.
The open-source community has become an accountability mechanism that the formal evaluation ecosystem never was. r/LocalLLaMA caught the Llama 4 discrepancies within days of release, while the benchmark industry took months to acknowledge the problem. This is a new model for AI evaluation: community-driven, transparent, continuously updated, and accountable to no vendor. It is not perfect. But it is more trustworthy than the numbers published by the labs themselves.
What's Next
Llama 4's successor, expected to emerge from Superintelligence Labs under Wang, will face unprecedented scrutiny. Every benchmark number will be independently verified within hours of release. The community has learned that Meta's numbers cannot be taken at face value. The burden of proof has shifted from the skeptic to the publisher.
The benchmark industry itself is under pressure to reform. Static benchmarks are being replaced by dynamic, continuously updated evaluations. Arena-style human preference rankings are supplementing traditional accuracy metrics. The shift is toward multi-dimensional evaluation: not just "how high is the score" but "how reliable is the model in production, on real tasks, over time."
The unanswered question is the most important one. If Meta, with Yann LeCun as chief scientist, manipulated benchmarks, what is stopping every other lab from doing the same? The answer is nothing. The only constraint is the fear of getting caught. And as LeCun's confession demonstrated, getting caught may cost you a chief scientist and your credibility, but it does not change the incentive structure that made the fudging rational in the first place. The system that produced the Llama 4 scandal is still in place. It is waiting for the next model launch. For developers who want to evaluate models without relying on vendor claims, running your own tests is the only reliable path.
FAQ: Common Questions About the Llama 4 Benchmark Scandal
Did Meta admit to cheating?
Yes. Yann LeCun told the Financial Times that Llama 4 benchmark results were "fudged a little bit" and that the team used different model versions for different tests. This is the closest thing to an official admission of benchmark manipulation the AI industry has seen.
Were the publicly released models different from the benchmarked ones?
According to independent testers and LeCun's confirmation, yes. The models used to produce the published scores were not the same as the models released to the public. Community evaluations consistently found lower performance than Meta claimed.
Does this affect other AI companies?
It affects the entire ecosystem. If the most prominent open-source AI company manipulated benchmarks with its chief scientist's knowledge, the credibility of all vendor-published numbers is in question. The scandal has accelerated the shift toward independent, community-run evaluations.
What happened to the Llama 4 team?
Meta created Superintelligence Labs under Alexandr Wang, effectively replacing the AI research leadership. LeCun departed as chief AI scientist to launch his own startup. The Llama brand continues under new management.
Should developers still use Llama models?
The benchmark scandal is about Meta's evaluation practices, not necessarily about Llama 4's actual capabilities. Independent evaluations suggest Llama 4 is capable, just not as capable as Meta's numbers claimed. Run your own evaluations on your specific use case before deciding.
The Llama 4 benchmark scandal will be remembered as the moment the AI industry admitted what everyone suspected: the numbers are not real. Not entirely. Not consistently. The person who confirmed it was not a critic. He was the scientist who built the lab. When the insiders stop believing the benchmarks, the rest of us should stop too. The only numbers that matter are the ones you verify yourself, on your own hardware, with your own data. Everything else is marketing.


