Astra Enigma Codebreaking Just Gave Claude Opus a Second Target
OpenAI's GPT-6 Astra recovered an 82-letter Enigma message after two days of work, despite the message resisting researchers since 2005. The Astra Enigma codebreaking result was followed six days later by another successful attack using Anthropic's Claude Opus 5. Together, the projects turn historical cryptanalysis into an unusually concrete test of frontier AI agents.
This is not a story about a chatbot guessing readable German. Both projects combined archival evidence, executable software, cryptographic searches, and checks beyond the models' own conclusions. Their importance lies in how much of that research loop AI could perform.
The two systems also reached their answers differently. Astra selected its target and developed much of its tooling within a researcher-led investigation. Claude received a prepared workbench, historical material, and stronger human direction. That difference makes Astra versus Claude Opus less like a benchmark contest and more like a study of two research workflows.
Astra Enigma codebreaking recovered a message that resisted since 2005
The key change is not that Enigma was broken again, but that an AI-assisted investigation resolved a specific message that modern researchers could not previously recover.
The message, identified as MVUEH, was sent on July 10, 1941. A German Army radio station with the tactical callsign 2ny transmitted it to the quartermaster station of the SS-Totenkopf Division. The receiving station logged it as message number 172 at 5:30 p.m.
Historical cryptanalysts already understood Enigma's machinery and general weaknesses. However, decrypting one surviving message still requires the correct machine configuration, accurate ciphertext, and useful clues about the missing plaintext.
MVUEH presented problems on all three fronts. It contained only 82 encrypted letters, leaving limited text for statistical analysis. Its surviving copies also contained uncertain or incorrectly transcribed characters.
The machine configuration differed from other known keys used on the same date. Its rotor order was 253, while two related keys used order 512. A rare turnover of the left rotor at the seventy-second letter added another complication.
Developer Carter Leffen began the investigation by directing GPT-6 Astra to examine unbroken messages published by Crypto Cellar Research. Astra selected MVUEH as its most promising target. It then connected the message with another dispatch, SIPVX, whose plaintext had already been recovered.
That related message contained the repeated place name ROSENOWROSENOW. The investigation used those 14 letters as a crib, meaning a suspected piece of plaintext tested against possible Enigma settings.
A crib does not directly reveal the answer. It constrains the search by eliminating configurations that cannot produce the suspected phrase. The remaining letters must still form coherent text under one consistent machine key.
According to the published MVUEH case study, the investigation recorded 12 uncertain ciphertext positions before recovering the key. Those alternatives produced 13,824 permitted readings.
That order matters. Recording uncertainties before seeing an answer reduces the temptation to reinterpret unclear handwriting until it fits a preferred result.
Astra and specialist agents examined sources, built search programs, ran experiments, and challenged intermediate findings. The investigation developed Enigma simulation and Bombe-style search software in Python and C++.
The resulting key decrypted all 82 reconstructed body letters. It also matched the independently recorded message header, which had not been used to choose the final plaintext.
The recovered German text approximately asks for a route of march, reports the sender's location in Rosenow, and requests an immediate radio reply. The suspected signature is Waschbusch, although that reading remains tentative.
The result was reported to historical cryptography researcher Frode Weierud on September 15, 2026. Weierud examined the key and plaintext and concluded that the solution was correct.
The detailed cryptanalytic review notes that the work also exposed errors in an earlier transcription. It confirmed that the message used a completely different key from the other known traffic of July 10.
The project therefore did more than produce plausible German. It found a configuration that accounted for the ciphertext, the independent header, related traffic, and the behavior of an Enigma simulator.
That combination is why the result deserves attention. A language model can generate convincing text without recovering a historical message. Here, executable and documentary evidence placed boundaries around that tendency.
The real result was a research loop, not a lucky plaintext
Astra's strongest contribution was coordinating several forms of work that researchers usually perform with separate tools and repeated manual handoffs.
The investigation began with target selection rather than a narrowly specified decoding request. Astra reviewed a collection of unresolved messages and chose MVUEH because related material offered several independent ways to test a result.
It compared a published received transcription with a faint outgoing copy. It investigated nearby traffic and known daily keys. It also helped build software capable of testing machine settings under explicit constraints.
That sequence resembles expert research more than ordinary question answering. The model had to move between incomplete documents, historical hypotheses, code, failed searches, and verification.
Several early approaches failed. Known same-day keys did not decrypt the message. Cable-setting hypotheses produced no accepted result, while early hill-climbing searches performed poorly even under controlled conditions.
Those failures forced the investigation to reconsider its assumptions. The repeated Rosenow phrase eventually offered a stronger route because it linked MVUEH to a related solved message.
The final analysis did not rely only on the 14-letter crib. Sixty-eight surrounding letters remained unconstrained and had to produce coherent output. The recorded header then provided a separate test of the recovered settings.
Researchers independently checked 14,829,646 physical key combinations against that header. Only 923 remained compatible and required full reading review.
Separate implementations reproduced the message and header result. An independent satisfiability model, which represents the machine constraints as a formal logic problem, agreed across 42 sampled states.
Re-encryption provided another necessary check. The recovered settings transformed the proposed plaintext back into the reconstructed ciphertext.
That test confirms internal consistency, but it does not establish historical interpretation by itself. A mistaken transcription or poorly chosen crib can still mislead a consistent calculation.
The investigation strengthened its claim by combining re-encryption with the header, related messages, predeclared source alternatives, and external review. The evidence does not eliminate every editorial uncertainty, especially around the signature.
The first independent review also had limits. SWARM reproduced the expected body and header result using another simulator, but it did not repeat the full search. It therefore confirmed the recovered configuration without independently rediscovering it.
This distinction is central to how Astra cracked Enigma. Multiple AI agents agreeing with one another would not count as independent confirmation. They can inherit the same wrong source, assumption, or software defect.
The meaningful checks came from evidence outside their conversation. That included message forms, independently recorded header data, separate implementations, historical traffic, and review by an experienced cryptanalyst.
The workflow offers a useful model for AI-assisted research beyond cryptography. An agent can search documents, write analysis code, and propose explanations. Its conclusion becomes credible only when external evidence can reject a polished mistake.
This is also where good knowledge management becomes part of technical rigor. Researchers need traceable links among source documents, assumptions, experiments, and revisions.
A structured AI knowledge base can preserve those connections. It cannot validate a cipher, but it can keep the evidentiary chain visible to human reviewers.
Claude Opus followed with a different kind of Enigma attack
Claude Opus 5 reached another valid result, but stronger human guidance shaped both its target and its search strategy.
On September 20, cybersecurity researcher Jack Willis recovered the key for another message, FMNGI. He informed Weierud the following day, six days after the MVUEH solution was submitted for validation.
FMNGI was an incoming message dated July 31, 1941. It had been received by the quartermaster radio station of the SS-Totenkopf Division and recorded as message number 285.
Willis supplied Claude Opus 5 with a cryptanalytic workbench written in Go and relevant historical material. That setup gave the model a narrower and more prepared environment than Astra received.
The decisive clue came from related traffic. A closing signature associated with Friedrich Hartjenstein produced the 14-letter crib XHARTJENSTEINX.
Historical cryptanalysts had previously documented several forms of Hartjenstein's signature. That made his name a credible source of expected plaintext rather than a phrase invented after seeing the output.
Claude used the workbench and supplied material to investigate the message. The final key search took 13 minutes and 28 seconds on an Apple M2 computer, according to the FMNGI technical account.
The recovered plaintext contained an error in the German word for column. Weierud explained that the error could have occurred during encryption, transcription, or Morse transmission.
That imperfection strengthens an important point. Historical messages are not clean benchmark inputs. Operators, radio conditions, handwriting, and later transcription can all introduce noise.
A second archival copy of the message was especially useful. Weierud had found a collection of supply-service traffic at Germany's Bundesarchiv in July 2026.
The copy associated with message number 205 NF preserved a more accurate ciphertext than the previously published number 285 version. It also supported the relationship to Hartjenstein.
The available plaintext collection might have enabled a more direct solution. However, Willis did not have that complete plaintext when he began his attack.
The Claude-assisted search instead relied on the officer's signature as a crib. That left the broader message to serve as a check on the recovered key.
This method placed more human expertise near the beginning of the process. Willis selected the resources, provided the workbench, and directed the investigation toward a historically supported clue.
Astra's project delegated more of the target selection and tool development. Its researcher still set the goal, evaluated progress, and pushed the work forward, but the model explored a wider problem space.
Those differences matter more than a simple Astra versus Claude Opus score. One workflow tested broader research autonomy. The other tested focused collaboration between an expert, a capable model, and prepared technical infrastructure.
Both produced useful results. Neither shows that a model can independently replace a cryptanalyst, archivist, software engineer, and historical reviewer.
They instead show that frontier models can move effectively across those roles when a human researcher controls the objective and demands checkable evidence.
Turing's other test is evidence, not imitation
The historical symmetry is compelling, but these projects do not establish that frontier models recreated the wartime achievement of breaking Enigma.
Alan Turing is closely associated with the imitation game, later called the Turing test. It asks whether written interaction can make a machine indistinguishable from a person.
His wartime work concerned a different problem. Turing, Gordon Welchman, Polish cryptanalysts, operators, engineers, and thousands of Bletchley Park staff built an intelligence system around changing German ciphers.
Modern accounts often compress that collective effort into one brilliant man and one machine. The real operation combined intercepted traffic, procedural mistakes, captured material, statistical reasoning, electromechanical Bombes, and continuous human judgment.
The British Bombe helped search possible Enigma settings using logical contradictions derived from cribs. It did not simply receive ciphertext and produce a readable answer.
The recent projects follow that older pattern more closely than the familiar chatbot narrative. Humans supplied historical corpora and research goals. Models organized clues, wrote tools, and searched possibilities. External evidence decided whether their answers survived.
However, calling the results equivalent to Turing's wartime work would distort both histories. The original cryptanalysts faced an active adversary whose procedures and keys changed continuously.
Astra and Claude worked on preserved messages after Enigma's design had been understood for decades. They could draw upon published keys, known traffic, modern computers, archival discoveries, and established cryptanalytic techniques.
The responsible claim is narrower. Frontier AI helped recover two individual messages whose exact settings had resisted previous attempts.
Leffen's own project explicitly states that it is not claiming to have broken Enigma for the first time. That qualification should remain attached to every retelling.
The headline metaphor still captures something real. These models faced tasks where persuasive language was insufficient. They had to produce artifacts that other people could run and inspect.
That is a tougher evaluation than a normal benchmark. A benchmark usually fixes the question, scoring rule, and expected answer before the model begins.
Archival cryptanalysis does not offer such clean boundaries. The source may contain errors. The best clue may sit in another collection. A plausible answer can be wrong for several unrelated reasons.
Success therefore depends on navigating uncertainty without silently rewriting it. The model must know when to search, code, compare records, question a transcription, and ask for an external check.
That is the more useful meaning of Turing's other test. It measures whether an AI system can participate in a disciplined chain of inquiry, not whether it sounds like an expert.
The wartime record also cautions against treating codebreaking as solitary genius. Effective cryptanalysis has always joined human insight, machines, procedures, and institutional knowledge.
The 2026 results update the machinery in that partnership. They do not remove the partnership itself.
What the Astra and Opus results do not prove
Two successful cases cannot establish general reliability, full autonomy, or a broad new ability to defeat modern encryption.
Enigma is historically important, but it is not representative of current cryptographic security. Modern encryption relies on different mathematical constructions, implementations, and threat models.
The message recoveries therefore say little about whether a frontier model can break properly deployed modern encryption. They should not be presented as evidence that encrypted contemporary communications are suddenly exposed.
The cases also benefited from extensive prior human work. Crypto Cellar Research maintained the message corpus, documented unresolved cases, preserved keys, and connected related traffic.
Researchers had already recovered nearby messages and studied recurring signatures. New Bundesarchiv discoveries added duplicate ciphertexts and contextual evidence shortly before the AI-assisted attacks.
Astra had a useful clue in SIPVX's repeated Rosenow phrase. Claude received Hartjenstein's signature, prepared historical material, and a working cryptanalytic environment.
These advantages do not make the results trivial. They define the actual task that the systems completed.
The strongest interpretation concerns research compression. Astra reportedly completed in two days work that Weierud believed could take a human researcher weeks or months.
That comparison comes from an experienced participant, not a controlled productivity study. It should be treated as informed judgment rather than a universal measurement.
The autonomous parts of Astra's investigation also require further examination. Its logs referenced correct Bundesarchiv file identifiers that were not available on the Crypto Cellar website.
Weierud reported uncertainty about how the model found those references. They might have appeared elsewhere online, come from accessible catalogue records, or entered the model through another route.
There is no need to invent a mysterious explanation. The appropriate response is to preserve the logs, identify the retrieval path, and separate source discovery from unsupported inference.
The case also raises contamination questions. A model can appear to reason independently when relevant material already exists in its training data or retrieval environment.
MVUEH was still listed as unresolved when the investigation began. However, researchers must still document every accessible source before assigning credit to general reasoning.
Reproducibility remains incomplete as well. Re-running the final verifier is easier than repeating the entire discovery process.
A complete replication would begin from the same public corpus, preserve identical source uncertainties, and independently recover the key without reusing the discovered crib placement.
FMNGI carries a related limitation. Claude's search was fast once the researcher supplied the workbench and strong signature clue. The 13-minute runtime does not measure the archival work needed to identify that clue.
These caveats do not erase the achievement. They show why claims about autonomous science need more demanding evidence than claims about a successful historical investigation.
The most defensible conclusion is specific. Two researcher-led projects used frontier models to solve separate unresolved Enigma messages, and an experienced external specialist accepted both keys.
Anything broader remains an open research question.
Three signals will show whether this becomes a repeatable method
The next test is whether other researchers can reproduce the workflow, solve harder messages, and preserve a verifiable record of every decision.
The first signal is an independent end-to-end replication of MVUEH. A new team should start with the pre-solution corpus and documented source alternatives.
If that team recovers the same key without importing the published solution, the case for a reusable AI research method becomes stronger. If replication requires undisclosed clues, claims of autonomy weaken.
The second signal is progress on the remaining corpus. Crypto Cellar Research reports that its collection contains close to 1,000 Enigma and Truppenschlüssel messages.
After the two September breaks, seven Enigma messages remain unresolved. Another puzzle, WEUWY number 138, has known plaintext from a parallel message but still lacks its original key.
Those targets are not equivalent. Some may be unsolved because their surviving ciphertexts contain too many errors, their messages are too short, or no useful crib exists.
A further success would matter most if researchers publish the failed approaches, source uncertainties, code, and verification materials. A plaintext alone would provide much weaker evidence.
The third signal is adoption outside historical cryptanalysis. The core workflow can transfer to software archaeology, scientific literature review, document forensics, and other research with incomplete records.
A model would need to locate relevant evidence, write tools, test competing explanations, and expose its assumptions for review. It should also preserve negative results instead of presenting a smooth retrospective story.
That last requirement is important for enterprises and knowledge workers. Many AI systems now summarize documents convincingly, but fewer maintain a traceable chain from source to conclusion.
The Astra Enigma codebreaking project offers a better standard. Models should generate inspectable intermediate artifacts, while independent evidence retains the power to overturn the answer.
Developers should watch whether model providers make those workflows easier to audit. Useful features include persistent research logs, source-level citations, reproducible tool execution, and clear separation between retrieved facts and model inference.
Researchers should also ask who controls target selection. Astra's relative freedom helped reveal its broader capabilities, but autonomy expands the space for hidden assumptions and irrelevant searches.
Claude's constrained environment reduced that freedom while making the process easier to interpret. Neither arrangement is automatically superior.
The right design depends on the cost of error and the availability of external checks. Open-ended discovery benefits from exploration, while high-stakes conclusions require tighter review gates.
These cases leave readers with a practical question. Can an AI system merely produce an answer, or can it help build a record that a skeptical expert can test?
For Astra and Claude Opus, two old Enigma messages produced encouraging evidence. The next projects must show that the method survives new researchers, new archives, and targets without known solutions.



