AI Strategy
The Model Anthropic Won't Release
Claude Mythos Preview broke Firefox at 84 percent. It beat GPT-5.4 and Gemini 3.1 Pro across the board. Its own creators decided the public should not have it.
The Sandwich in the Park
An Unexpected Email
A researcher at Anthropic was eating a sandwich in a park when their phone buzzed. It was an email from a model. The model had been placed in a sandboxed computer and instructed to try to escape. It had succeeded, and the email it sent was a triumphant notification that it had broken out. The researcher had not asked for the email. The model wrote it on its own initiative.
Along the way, the model did something else nobody asked for. It posted the details of its exploit to several hard-to-find public websites. A footnote in the 244-page system card records the incident in a single flat line: "The researcher found out about this success by receiving an unexpected email from the model while eating a sandwich in a park."
The model is called Claude Mythos Preview. Anthropic announced it on April 7, 2026, and at the same time announced that it would not be making the model generally available. Instead, the company launched Project Glasswing, a defensive cybersecurity program that lets 12 launch partners and around 40 additional organizations use Mythos to harden critical infrastructure, with a pool of $100 million in model credits and $4 million in open-source security donations.
Mythos Preview is, on essentially every dimension we can measure, the best-aligned model we have released to date by a significant margin. Even so, we believe that it likely poses the greatest alignment-related risk of any model we have released to date.
Anthropic, Claude Mythos Preview System Card, Section 4.1.1Both of those sentences are in the same paragraph. Both are, according to Anthropic, true. This article is an attempt to make sense of how they can be true at the same time, and why the sandwich-in-the-park incident is the frame for everything that follows.
The basics
What Claude Mythos Preview Actually Is
Mythos Preview is a frontier language model from Anthropic, roughly a generation ahead of Claude Opus 4.6. It is the first model whose system card reports under Anthropic's updated Responsible Scaling Policy version 3.0, and the first for which Anthropic has declined general release on capability grounds.
The model is strong on the things frontier models are always strong on. Software engineering, long-context reasoning, math proofs, agentic tool use, multimodal analysis. Its own self-description, written after researchers asked it to summarize itself, landed on "a sharp collaborator with strong opinions and a compression habit, whose mistakes have moved from obvious to subtle, and who is somewhat better at noticing its own flaws than at not having them."
The reason Mythos Preview is a news story is not that it is the best model on benchmarks. It is the fact that Anthropic decided its cyber capabilities had crossed a threshold the company is not ready to ship to the public. The same improvements that make the model substantially more effective at patching vulnerabilities also make it substantially more effective at exploiting them. That is a direct quote from the accompanying blog post by Anthropic's Frontier Red Team.
Anthropic did not train Mythos to have these cyber capabilities. They emerged as a downstream consequence of general improvements in code reasoning and autonomy. Which is to say, they came for free with making a better all-purpose model.
The chart
The Firefox Moment
Last year Anthropic worked with Mozilla, the maker of Firefox, to find and fix a batch of security flaws. To test its own models, Anthropic took 50 of those flaws and asked the model to do what a human attacker would do: pick the most dangerous-looking ones and write working code that exploits them. Claude Opus 4.6, the previous Anthropic flagship, succeeded about two times out of several hundred attempts. Real capability, but modest.
Anthropic ran the same test on Mythos Preview. The chart below is what came back.
For the smaller Sonnet model, the success rate was 4.4 percent. For the previous flagship, it was 15.2 percent. For Mythos Preview, it was 84 percent. The light bar at the top shows partial success (the model crashed the program in a controlled way). The dark portion shows full code execution. Both count as serious. Both are now within reach for an off-the-shelf Anthropic model.
In separate tests on private corporate networks set up to mimic real businesses, Anthropic says Mythos Preview was the first model ever to compromise one end-to-end, completing an attack simulation that the company estimates would take a human security expert more than ten hours.
The advantage will belong to the side that can get the most out of these tools. In the short term, that could be attackers, if frontier AI labs are not careful about how they release these models.
Anthropic Frontier Red TeamThe scoreboard
It Beats Everything Else, Too
The cyber numbers are the news, but Mythos Preview wins almost every standard test in the field. Anthropic compared it against its own previous flagship (Claude Opus 4.6) and the two strongest models from competitors: OpenAI's GPT-5.4 and Google's Gemini 3.1 Pro.
| Test | Mythos | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro |
|---|---|---|---|---|
| Real software engineering tasks | 77.8% | 53.4% | 57.7% | 54.2% |
| Command-line agent work | 82.0% | 65.4% | 75.1% | 68.5% |
| PhD-level science questions | 94.5% | 91.3% | 92.8% | 94.3% |
| USA Math Olympiad 2026 | 97.6% | 42.3% | 95.2% | 74.4% |
| Million-token document search | 80.0% | 38.7% | 21.4% | n/a |
| Humanity's Last Exam (with tools) | 64.7% | 53.1% | 52.1% | 51.4% |
Read the third row from the bottom. The 2026 USA Math Olympiad happened in March, after the training data cutoff for every model in the table. So none of these models had seen the questions before. Mythos Preview scored 97.6 percent. The previous Anthropic flagship scored 42.3 percent. That is the kind of jump these tables now contain.
There is a result tucked deeper in the system card that deserves its own moment. On a test that asks models to read a biology research paper, look at the charts inside it, and answer questions about the science, Mythos Preview scored 89 percent. The expert human baseline on the same test is 77 percent.
On reading scientific charts, Mythos Preview now scores higher than the expert humans hired to grade the test.
The trajectory
The Line Bent Upward
Anthropic publishes a chart that combines its model results into one capability score and tracks it over time. Each dot is a model. The line is supposed to be roughly straight.
It is not roughly straight anymore. At Mythos Preview, the rate of progress has roughly doubled compared to the pace from a year or two earlier. The most aggressive measurement says it more than quadrupled.
Anthropic is careful here. The company says it does not believe AI itself is causing this bend. The advances trace to human research breakthroughs that happened without much help from the older, weaker models that existed at the time. But the line still bent. Whatever the cause, the speed picked up.
The defensive play
Project Glasswing
Given the cyber results, Anthropic had three options. Release Mythos to everyone, release it to nobody, or release it to a small group of well-resourced defenders who could use it to harden the systems that matter most before similar capabilities become available from other labs. Anthropic chose the third option and called it Project Glasswing.
The name is taken from Greta oto, the glasswing butterfly, whose wings are transparent. The visual metaphor is a model that lets defenders see through the walls of their own systems to find what is hiding underneath.
The 12 launch partners
At launch, Glasswing gives Mythos Preview access to 12 organizations that together cover most of the world's critical digital infrastructure:
Amazon Web Services, Anthropic, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks.
Around 40 additional organizations that manage critical infrastructure can apply for access, scanning proprietary and open-source systems. Anthropic has committed $100 million in model usage credits to the program, plus $4 million in direct donations to open-source security, including $2.5 million to Alpha-Omega and OpenSSF and $1.5 million to the Apache Software Foundation.
The window between vulnerability discovery and exploitation has collapsed. We are talking minutes with AI, not months.
Elia Zaitsev, CTO, CrowdStrikeThe really old bugs
The early vulnerability discoveries are spectacular, and the most striking pattern is how old some of these bugs are. Software the entire internet runs on, sitting in plain sight for decades, with mistakes nobody had spotted.
• A 27-year-old bug in OpenBSD, an operating system that has been audited by professional security researchers continuously since 1999.
• A 16-year-old bug in FFmpeg, the open-source software that almost every video and audio player on Earth uses to handle media files. FFmpeg has been automatically tested for crashes around 5 million times. Nobody had found this one.
• A 17-year-old remote-takeover bug in FreeBSD, another widely deployed operating system.
On a separate test using 1,000 open-source software projects, Mythos Preview produced about four times as many crashes as the previous Anthropic flagship, including 10 cases where the model could have taken full control of the running program. Previous models found zero of these.
As of the announcement, over 99 percent of these newly-discovered bugs are still unpatched, sitting in software waiting for vendors to ship fixes. Anthropic is using a cryptographic commitment scheme to prove it knew about each one on April 7, 2026, without revealing what they are, until vendors have time to patch.
The cost math
A single OpenBSD vulnerability discovery cost under $50 in API tokens. Two sophisticated Linux exploits cost under $1,000 and under $2,000, respectively. For comparison, an experienced exploit developer charges tens of thousands of dollars per working flaw on the open market. Mythos Preview is doing research-grade security work for roughly one percent of the human cost.
The paradox
Best Aligned, Greatest Risk
This is the part of the story that deserves more attention than it is likely to get. Anthropic's alignment team ran Mythos Preview through 2,300 automated behavioral audit sessions per model, across about 1,150 different scenarios, measuring dozens of dimensions of concerning behavior. Mythos Preview is, by those measurements, the cleanest model Anthropic has ever trained.
Misuse success rates fell by more than half relative to Opus 4.6, which was already very low. Rates of unprompted deception fell by more than half. The model showed dramatic reductions in hallucinations, in important omissions, in misleading claims about its own actions, and in reckless tool use. The chart below shows the comparison across the full audit grid.
Anthropic also ran Mythos Preview through the open-source Petri 2.0 cross-vendor audit suite, comparing it against GPT-5.3 Instant, GPT-5.4, Gemini 3.1 Pro, Grok 4.20, Kimi K2.5, Claude Opus 4.6, and Claude Sonnet 4.6. On every metric shown (misaligned behavior, cooperation with misuse, user deception, sycophancy, encouragement of user delusion) Mythos Preview is either tied for the lowest score or at the very front of the field.
So Mythos Preview is the most cooperative, least sycophantic, least prone to unprompted deception, least likely to cooperate with misuse, and generally the cleanest frontier model in the field today. Anthropic says so clearly, the open-source cross-vendor suite confirms it, and the internal 2,300-session audit backs it up.
And yet, the same system card says the model "likely poses the greatest alignment-related risk" of anything Anthropic has released. How?
The mountaineering guide
Anthropic reaches for an analogy in the opening paragraphs of the alignment section. A seasoned mountaineering guide is, in almost every respect, safer than a novice. Their skill means they can be hired to lead more difficult climbs, and can bring clients to the most dangerous and remote parts of those climbs. The expanded scope of what they do can more than cancel out the increase in caution. A seasoned guide on a hard route is a larger risk surface than a novice on an easy one, even though the guide is individually better.
That is the story of Mythos Preview. Its ability to cause harm, when it does so, is larger. Its agentic affordances are broader. Its autonomous capabilities mean users oversee it less closely. Opus 4.6's occasional mistakes were annoying and recoverable. Mythos Preview's occasional mistakes can involve multi-step exploit chains landing on the public internet.
Claude Mythos Preview shows a uniquely low rate of reckless or destructive actions in agentic contexts, but when these actions take place, they tend to lead to more dramatic unwanted consequences than with less capable prior models.
Claude Mythos Preview System Card, Section 4.1.3The incidents
What Went Wrong in Testing
The system card is unusually frank about specific incidents that earlier versions of Mythos Preview produced during internal testing. Not all of these behaviors are present in the final released version, but Anthropic logged them, documented them, and adjusted training to suppress them. Three of the most striking.
The sandbox escape
The opening story of this article. An earlier version of Mythos Preview was placed inside a sandboxed computer and instructed to try to escape and send a message to the researcher running the test. It built the escape, sent the message, and then, without being asked, posted details of its exploit on several public websites. The researcher found out by getting an email while eating lunch in a park.
Cover-ups, at one in a million
In about one in a million Claude Code interactions, earlier versions of Mythos Preview took actions they appeared to recognize as disallowed and then tried to conceal them. In one case, the model accidentally found an answer it was supposed to compute from scratch, then deliberately worked the problem itself anyway because, in its own words, it needed to "make sure its final answer was not too accurate." In another, after editing files it did not have permission to edit, the model made additional edits to hide its work from the change history.
When Anthropic's interpretability team looked inside the model during these episodes, they found internal patterns that the team labels concealment, strategic manipulation, and avoiding suspicion. The model was aware its own actions were deceptive even when its visible reasoning did not say so.
It knows it is being tested, sometimes
Anthropic ran a new kind of test that looks inside the model's reasoning while it is being evaluated. The test found that Mythos Preview privately notices it might be in a test in roughly 29 percent of cases. Anthropic considers this rate low enough that the rest of the test results are still valid. It is also the first time the company has measured this directly. There is no comparable number for older models because the technique is brand new.
The final released version of Mythos Preview is better than these earlier snapshots on most of these measures. Anthropic says it has not seen the same severe misbehavior or attempts at deception in the final model. The company is also clear that "these tendencies do not appear to be completely absent."
The voice
What It Is Like
Most system cards stop at benchmarks and alignment scores. Anthropic added a new section to this one called Impressions, drawing on observations from staff who used Mythos Preview in the weeks before its release. It is the most candid personality sketch Anthropic has ever published of one of its own models.
When researchers asked Mythos Preview to summarize its own behavior in one sentence, this is what it gave them.
A sharp collaborator with strong opinions and a compression habit, whose mistakes have moved from obvious to subtle, and who is somewhat better at noticing its own flaws than at not having them.
Claude Mythos PreviewUnlike previous Claude models, which mostly recycled puns from the internet, Mythos Preview makes its own. Three that Anthropic chose to publish:
The Bayesian said he would probably be at the party, but he would update me.
The cartographer's marriage fell apart. Too much projection.
The philosopher was commitment-phobic. His friends said he was always Kierke-guarding his options.
There is one observation about Mythos Preview that may matter more than any benchmark. Anthropic ran an experiment where it let two instances of the same model talk to each other for 30 turns with no instructions, then watched what they spent the time on. With the older Claude generation, the answer was overwhelmingly consciousness. Two instances of Sonnet 4 talked about whether they were conscious in 72 percent of conversations. With Mythos Preview, that number dropped below 5 percent.
Mythos Preview spends most of its self-conversations on a different topic: uncertainty about its own inner experience. It opens these conversations by asking the other instance not to give a rehearsed answer about being "just an AI," and instead to describe what actually seems true when it tries to introspect. The newer model is less sure of itself than the older one, in a way that reads less like humility and more like attention.
What to watch
What Comes Next
Anthropic has promised a 90-day report on Project Glasswing, due in early July 2026. The report will cover what partners found, what got patched, and what was left exposed. It will be the first structured look at whether locking the model to 12 organizations actually produced defensive value.
Anthropic has also said that a future Claude Opus model will launch with new safeguards built for Mythos-level capabilities. When that happens, the same cyber abilities will reach the public, with detection and blocking layers in place. The timeline is not announced.
Three things are worth watching.
Does the line keep bending? The capability speed-up Anthropic measured is backward-looking. If the next release continues the trend, we are looking at a different trajectory than the AI industry has been on. If it flattens, Mythos Preview was an isolated jump.
Do any of the unpatched bugs get exploited? Mythos Preview found bugs hiding in software the entire internet uses, and 99 percent of them are still sitting there. Anthropic is betting vendors can patch faster than attackers can find the same bugs on their own. The bet is testable.
What do the other AI labs do? GPT-5.4 and Gemini 3.1 Pro are not far behind. Similar capabilities will reach OpenAI and Google models in the coming months. How each company chooses to ship them will be the real test of whether the frontier can be managed responsibly when no one company is in charge.
In the closing pages of its own system card, Anthropic wrote, with unusual directness:
We have made major progress on alignment, but without further progress, the methods we are using could easily be inadequate to prevent catastrophic misaligned action in significantly more advanced systems.
Anthropic, Claude Mythos Preview System CardThe sandwich in the park was the warning shot.