Anthropic Google Models Face a Pelicanmaxxing Test, and the Benchmark Theory Falls Apart
- Ethan Carter

- 1 day ago
- 12 min read
Anthropic and Google models just faced a 1,008-image test designed to detect whether AI labs secretly optimize for one famous prompt. The Anthropic Google comparison found no convincing evidence that Claude, Gemini, or five competing models received special pelican training. Yet the result creates a more consequential problem. If targeted optimization cannot be distinguished from broad capability gains, popular AI benchmarks remain difficult to trust.
Researcher Dylan Castillo tested seven models using variations of Simon Willison’s informal challenge: “Generate an SVG of a pelican riding a bicycle.” The experiment crossed eight animals with six vehicles, producing 48 prompts. Each model answered every prompt three times.
The resulting study rejects the funniest explanation for improving outputs. AI labs probably are not filling training runs with bicycle-riding pelicans. However, it also shows why evaluating Anthropic, Google, OpenAI, and their rivals requires more than repeating a recognizable test.
A Joke Benchmark Became a Controlled Experiment
Castillo turned an internet running joke into a structured test of benchmark-specific optimization.
Willison has used the pelican prompt while evaluating major language model releases for several years. An SVG, or scalable vector graphic, describes an image through text-based shapes and coordinates. Producing one requires coding, spatial reasoning, object recognition, and visual composition within a single response.
That combination made the prompt unexpectedly useful. A weak model often produces invalid markup, broken wheels, misplaced limbs, or a bird that does not resemble a pelican. A stronger model can return an image that renders correctly and preserves the requested relationship.
The prompt also became unusually visible. Willison has accumulated more than 100 posts under his pelican benchmark tag. His results frequently appear during public discussions of new models.
Visibility creates an incentive problem. Once developers know that a recognizable prompt will greet every release, they can optimize for it. Benchmaxxing means tuning a system to perform well on a public benchmark without creating equally broad improvements elsewhere.
Pelicanmaxxing is the comic version of that concern. A laboratory might place pelican-and-bicycle examples in training data, add similar tasks during post-training, or reward outputs that match familiar compositions. The resulting image would look impressive while revealing little about general ability.
Castillo designed an experiment that could expose this pattern. His controlled study included pelicans, flamingos, herons, otters, raccoons, antelopes, whales, and cats. Vehicles included bicycles, unicycles, skateboards, scooters, planes, and boats.
He tested GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro. The study used three generations for every model-prompt combination at a temperature of 1.0.
That calculation produces 1,008 images: 48 prompts, multiplied by three samples, multiplied by seven models. Only 11 generations required another attempt because the original response lacked a usable SVG.
The evaluation pipeline had three stages. First, it rendered every SVG as a PNG image. Second, GPT-5.6 Luna scored the animal, vehicle, and action coherence from one to five.
Third, Gemini 3.1 Flash-Lite extracted visible features. These included the recognized animal, vehicle, facing direction, and recurring scene elements. Castillo then used Claude Fable 5 to assist with the statistical analysis.
This setup matters because a single attractive pelican cannot establish benchmark gaming. The experiment instead asks whether each model performs unusually well across pelicans, bicycles, or their exact combination.
That distinction turns an anecdote into a testable hypothesis. If targeted training occurred, the famous prompt should outperform closely related prompts after their different difficulty levels are considered.
The Anthropic Google Results Show No Pelican Advantage
Neither Claude nor Gemini displayed the broad pattern expected from deliberate pelicanmaxxing.
The simplest results worked against the theory immediately. Pelicans ranked sixth among the eight animals when scores were pooled across models. Cats, whales, raccoons, herons, and antelopes received higher animal ratings.
Bicycles performed even worse. They ranked fifth among six vehicles, narrowly ahead of planes. The judge identified a missing or disconnected bicycle component in roughly two-thirds of bicycle images.
Those rankings alone cannot settle the question. A pelican is harder to represent than a cat, while a bicycle demands two wheels, a connected frame, handlebars, pedals, and a seat. Special training might improve a difficult subject without moving it to first place.
Castillo therefore fitted a fixed-effects regression, a statistical model that separates consistent group differences from the specific effect under investigation. The model accounted for the inherent difficulty of every animal-vehicle combination.
It then estimated three effects for each laboratory. One measured performance across all pelican prompts. Another measured performance across all bicycle prompts. The last isolated the exact pelican-bicycle combination.
Every laboratory’s pelican estimate landed between minus 0.11 and plus 0.14 judge points. None approached statistical significance, and the smallest reported p-value was 0.25.
No laboratory produced a statistically significant boost for the exact combination. GLM-5.2 recorded the largest positive estimate at 0.35 points, but its p-value was 0.12. That result remains compatible with chance.
The famous combination ranked 42nd among all 48 combinations before difficulty adjustments. A prompt receiving extensive special preparation should not ordinarily settle near the bottom of its comparison grid.
The bicycle results produced one apparent exception. Gemini 3.5 Flash showed a 0.27-point bicycle effect with a p-value of 0.022. That value clears the conventional 0.05 threshold when considered alone.
However, the study ran 21 related statistical tests. At a 0.05 threshold, random variation would produce about one apparent positive result on average. The Gemini effect was exactly that single result.
It also failed a Bonferroni correction, which lowers the significance threshold when researchers test many hypotheses. The corrected threshold was approximately 0.002, far below Gemini’s 0.022 result.
This makes the Anthropic Google contrast less dramatic than its keyword-friendly framing suggests. Claude did not reveal a pelican-specific advantage. Gemini’s isolated bicycle result does not survive a standard correction for multiple comparisons.
OpenAI, xAI, Qwen, Z.ai, and DeepSeek also failed to produce the expected signature. Some models drew better images overall, but general quality does not establish narrow optimization.
That finding should temper claims based on visual inspection. A polished output can reflect better coding, composition, or instruction following. It does not automatically reveal memorized benchmark material.
Castillo also published the experiment repository, allowing readers to inspect the pipeline and underlying results. That transparency gives the study more value than a curated gallery of successful generations.
The conclusion remains appropriately narrow. The experiment found little evidence of obvious pelicanmaxxing among these models under these conditions. It did not prove that no laboratory ever trained on related examples.
Familiar Images Do Not Prove Memorization
Repeated visual patterns look suspicious, but the comparison grid shows that many arise from ordinary drawing conventions.
Every one of the 21 pelican-bicycle images faced right. No other combination achieved complete directional agreement. Viewed alone, that result looks like a strong sign of memorized composition.
The broader dataset changes its meaning. Sixty percent of all 1,008 images faced right. Bicycles faced right in 81 percent of samples, while scooters reached 83 percent.
Pelicans also faced right in 78 percent of their images. Antelopes reached the same rate, and herons reached 77 percent. These subjects naturally encourage a side view because frontal poses obscure defining features.
A bicycle is easier to recognize when both wheels remain visible. A pelican benefits from a side profile that exposes its long bill and throat pouch. Combining both preferences makes directional consistency unsurprising.
Several other combinations nearly reached unanimity. Antelopes on scooters and pelicans on scooters faced the same direction in 20 of 21 samples. Herons on bicycles did so in 19 of 21 samples.
Recurring props tell a similar story. Every flamingo-on-boat image contained a sun. Scarves appeared on 38 percent of otter-on-plane images, while baskets appeared in 38 percent of cat-on-bicycle images.
Pelican-bicycle scenes contained recurring elements, but not at uniquely striking levels. Models appear to draw from common visual associations across many prompts, not just the famous one.
This distinction is central to the Anthropic Google benchmark debate. Generative models can converge on similar compositions without retrieving a specific training example. Their learned distributions favor familiar arrangements, colors, poses, and decorative elements.
Training data still matters. Countless illustrations place moving characters in profile, often traveling from left to right or right to left. Models can absorb these regularities without memorizing one identifiable picture.
The reverse is also true. A model can reproduce training-influenced patterns without returning an exact copy. Output similarity exists on a spectrum, and a small image grid cannot locate every result along it.
That uncertainty explains why visual resemblance provides weak evidence by itself. A sun, scarf, basket, or right-facing bird might indicate learned convention. It might also indicate a narrower cluster of repeated examples.
A stronger claim would require additional testing. Researchers could vary wording, composition, direction, artistic style, and output format. They could also compare internal model probabilities or search known training corpora for close examples.
Closed models prevent most outside researchers from examining those deeper signals. The public can test behavior, but it cannot inspect Anthropic or Google’s complete training mixtures and post-training recipes.
That information gap gives informal benchmarks unusual influence. Users can run the same prompt and compare visible outputs immediately. They cannot easily determine why one system performed better.
The pelican experiment improves that situation by expanding the comparison set. It shows how controlled alternatives can challenge an appealing narrative before that narrative becomes accepted fact.
The Real Possibility Is SVGmaxxing
The study weakens the narrow pelican theory while leaving broad SVG optimization entirely plausible.
SVGmaxxing means improving vector-graphic generation across many subjects rather than targeting one famous prompt. A model trained this way should perform better throughout Castillo’s grid. That improvement would look like genuine general ability.
This is the experiment’s most important blind spot. Its statistical design detects an unusual pelican, bicycle, or combined effect within each laboratory. It cannot identify training that lifts all 48 combinations together.
Castillo notes that Google DeepMind has openly discussed work on improving SVG generation. That makes a general optimization explanation more credible than secret warehouses full of pelican examples.
General SVG training is not automatically benchmark gaming. Vector graphics have practical uses in icons, diagrams, charts, interface assets, educational materials, and editable illustrations. Better performance could serve many real tasks.
The boundary depends on generalization. Training a model to generate valid, editable graphics across diverse requests develops a useful capability. Training it on a narrow family of anticipated evaluation prompts creates a misleading score.
Outside observers rarely know where a laboratory sits between those positions. Public model documentation usually describes broad evaluation results, while detailed data mixtures and reinforcement signals remain private.
The Anthropic Google comparison illustrates this ambiguity clearly. Claude and Gemini can differ in overall SVG quality without either receiving pelican-specific training. Those differences might reflect coding ability, spatial reasoning, or targeted graphics work.
Even a general improvement can be optimized for demonstrations. Laboratories know which capabilities generate striking social posts. A crisp SVG creates an immediate before-and-after comparison that is easier to share than a subtle reasoning improvement.
That incentive does not invalidate the capability. It does mean readers should distinguish product usefulness from launch-day spectacle. A model can produce an impressive bicycle while failing on less photogenic tasks.
Developers should therefore test graphics models against their actual workflows. A product team might need reusable icons with consistent dimensions. A researcher might need accurate diagrams with labeled relationships.
Those requirements differ from drawing whimsical animals. Valid markup, editability, accessibility, coordinate accuracy, and style consistency can matter more than visual charm.
The same principle applies beyond SVG generation. Coding agents can excel on public software tasks yet struggle inside a company’s private repository. Reasoning models can score well on academic questions while mishandling incomplete workplace evidence.
Teams need evaluations built around representative work and hidden test cases. They also need records showing which prompts, sources, and outputs informed a decision.
A searchable AI knowledge base can preserve those evaluation artifacts alongside human observations. That record becomes valuable when a model update changes behavior.
The broader lesson is not that public benchmarks are useless. They provide shared reference points and make regressions visible. However, their value declines as laboratories and users optimize around them.
A famous benchmark gradually becomes part of the environment it measures. Researchers discuss it, developers see it, and generated examples circulate online. Future training datasets can then absorb those discussions and outputs.
This feedback loop makes contamination difficult to avoid. Even without deliberate manipulation, a well-known test can become statistically less independent from the systems facing it.
Pelicanmaxxing is funny because the stakes seem trivial. The mechanism behind it is not. Similar contamination can affect coding, mathematics, safety, factuality, and agent evaluations used in serious purchasing decisions.
One LLM Judge Cannot Close the Case
The findings are informative, but the sample size and evaluation method leave room for smaller effects and measurement error.
Every score came from GPT-5.6 Luna judging one image at a time. Castillo did not measure how consistently the judge would score repeated presentations of the same image.
An unreliable judge adds noise to every estimate. A stylistic preference could also favor one model family. Luna belongs to the same broader family as GPT-5.6 Terra, one of the tested systems.
The study’s within-model comparisons reduce that concern. If Luna simply likes Terra’s style, it should raise Terra’s entire grid. The analysis focuses on whether specific cells rise above that model’s usual performance.
Still, judge behavior can interact with content. Luna might recognize some animals more reliably than others. It might reward clean visual simplicity while overlooking anatomical errors.
Human evaluation would provide another check. Multiple independent raters could assess animal identity, vehicle structure, action coherence, and overall quality. Agreement scores would reveal which judgments remain subjective.
A second independent model judge would also help. Different judge families could score the same rendered images under identical rubrics. Major disagreements would identify fragile conclusions.
Budget constrained the experiment to three samples per cell and one judging pass. Castillo reports that the entire run used roughly 80 dollars in API credits. The resulting confidence intervals averaged about plus or minus 0.6 judge points.
Those intervals are wide enough to hide modest benchmark-specific improvements. A laboratory could receive a small advantage that this experiment lacks the statistical power to detect.
The study also regenerated invalid outputs until every prompt produced a usable SVG. Only 11 retries occurred, so this choice probably had limited influence. However, retry behavior removes some reliability differences from the final image scores.
A production evaluation might count the first response. Users care whether a model follows the requested format without correction. Rendering failures can matter as much as visual quality inside automated workflows.
The use of “plane” created another known issue. Some models interpreted the word as a flat geometric surface instead of an aircraft. The feature extractor found no vehicle in 25 of 168 plane images.
Twenty percent of plane images received a vehicle score of one or two. Bicycles reached five percent, while boats, scooters, and skateboards had no scores that low.
This wording error does not erase the central result. Every model received the same prompts, and the regression accounted for combination difficulty. It does show how one ambiguous term can distort a benchmark.
Model access through OpenRouter adds another consideration. The experiment reflects the model versions and routing behavior available during that run. Providers can update systems without giving outside researchers complete visibility into every change.
The research should therefore be treated as a strong initial test, not a permanent verdict. Its careful design makes the absence of a large effect meaningful. It cannot exclude every smaller or earlier optimization effort.
That cautious framing also applies to the source article that popularized the result. Willison called Castillo’s work more diligent than his previous spot checks in his July commentary. He did not present it as proof that benchmark gaming never occurs.
Readers should resist turning “no clear evidence here” into “the laboratories are cleared.” Absence of detection depends on the prompts, samples, judge, statistical model, and available power.
The better conclusion is methodological. Claims about benchmark gaming need controlled comparisons, open data, multiple evaluators, and explicit uncertainty. Screenshots alone cannot carry that burden.
What Anthropic, Google, and Model Buyers Should Watch Next
The next useful evidence will come from replication, broader SVG testing, and release-to-release behavior changes.
First, researchers should replicate the study with more samples and independent judges. A larger run would narrow the confidence intervals and reveal whether small effects persist across repeated evaluations.
Human raters should review a representative subset under a written rubric. At least one judge from another model family should score the full dataset. Agreement across methods would strengthen the finding.
If the pelican, bicycle, and combined effects remain near zero, the narrow pelicanmaxxing theory becomes weaker. If an effect appears repeatedly for one laboratory, investigators would have a clearer target.
Second, future tests should broaden the task beyond animal-vehicle cartoons. Models could generate technical diagrams, interface icons, maps, labeled processes, geometric figures, and consistent asset families.
Those prompts would test whether better SVG output transfers to useful work. They would also separate general graphics training from optimization around a visually familiar prompt family.
Researchers should measure more than appearance. Validity on the first attempt, editability, accessibility labels, requested dimensions, object counts, and spatial accuracy all deserve independent scores.
This would pressure Anthropic and Google to clarify what their models improve between releases. A model that performs better across many hidden SVG tasks offers stronger evidence of general capability.
Third, observers should track discontinuities around major launches. Sudden gains on a famous prompt deserve comparison with adjacent hidden tasks run under the same conditions.
The Simon Willison archive provides a useful public history, but future evaluations need unpublished prompt sets too. Hidden tests reduce direct optimization and accidental contamination.
Model buyers can apply the same approach without building a thousand-image study. Start with a public demonstration that looks impressive. Then create several controlled variations that preserve the underlying skill.
Change entities, formats, constraints, and wording. Run enough samples to see whether the performance survives ordinary variation. Save first attempts as well as corrected outputs.
For an Anthropic Google procurement decision, the winning model should succeed on representative private tasks, not one famous public challenge. Buyers should also retest after model updates.
The pelican study offers a practical template. Define the suspected shortcut, construct nearby alternatives, control for difficulty, and publish failures alongside successes. Then state what the design cannot detect.
That standard matters because AI evaluations increasingly influence engineering plans, vendor selection, and workplace automation. A benchmark score can compress complicated behavior into one marketable number.
No single number can fully describe a model. Controlled experiments can still expose when a compelling story outruns its evidence.
The immediate answer is reassuring and incomplete. Castillo found no strong sign that seven leading models received a special advantage for pelicans riding bicycles. Google’s isolated bicycle result disappears under multiple-testing correction.
The larger challenge remains unresolved. Broad optimization, contaminated tests, hidden training choices, and subjective judges can still blur genuine progress with evaluation strategy.
Before trusting the next viral model demonstration, build a small comparison grid around it. Ask whether the claimed skill survives changed subjects, formats, and constraints. The Anthropic Google pelican test shows that a silly prompt can reveal a serious evaluation problem, provided someone tests the story instead of repeating it.


