Skip to main content
1M
NVDA$207.14+3.2%
MSFT$486.19+4.6%
GOOGL$375.19+5.4%
AAPL$303.05-1.9%
AMZN$285.86+5.3%
META$593.44+6.6%
TSLA$323.64+4.0%
IBM$226.69+1.4%
CRM$188.81+2.6%
AMD$480.57+0.9%
AVGO$388.72-0.1%
ARM$236.96-1.1%
TSM$403.29-0.2%
INTC$90.31+0.1%
QCOM$149.23+1.1%
MU$815.37-0.9%
CSCO$115.21-0.7%
ANET$181.33+0.5%
SMCI$28.45+0.2%
ORCL$137.58+5.9%
PLTR$125.14+1.7%
DELL$417.14+2.9%
HPE$49.03+2.4%
PSTG$77.17+2.8%
ALAB$317.11+1.9%
AMAT$511.65+0.8%
LRCX$290.96-0.7%
KLAC$179.81-1.6%
ASML$1,635.04+0.4%
TER$362.31-1.5%
GEV$997.86+0.8%
CEG$272.41+3.7%
UEC$9.82+2.3%
OKLO$41.30+6.4%
SMR$8.92+5.9%
BWXT$172.80+2.4%
VST$153.96+3.9%
D$68.80-0.5%
SO$94.22-0.3%
NEE$86.33-0.7%
CCJ$89.54+3.7%
LEU$183.35+3.6%
BE$217.12+5.5%
KMI$31.66-1.6%
EXC$45.68-0.3%
VRT$258.15+6.9%
ETN$432.93+4.3%
CAT$823.61+1.1%
PWR$674.29+1.0%
EME$814.20+2.1%
URI$1,103.76+2.3%
VMC$278.76+3.8%
J$138.17+2.4%
TT$461.02+1.3%
CARR$62.65+1.4%
JCI$146.68+0.0%
SIEGY$162.90-0.1%
ALB$119.04+1.2%
SQM$66.71-0.5%
LAC$2.97+3.5%
MP$43.73+5.7%
FCX$63.24+1.0%
GLW$145.80+5.5%
SCCO$185.35+1.4%
AA$44.55-1.6%
RIO$95.41-1.5%
VALE$14.57-3.3%
LYSDY$9.80-1.4%
EQIX$1,024.62+0.5%
DLR$191.93+1.8%
AMT$174.17+0.5%
LITE$761.92+6.7%
COHR$287.52+9.4%
CIEN$383.36+1.7%
IRM$124.97+2.2%
CCI$78.36+2.7%
MDB$357.81+6.0%
ADBE$253.37+1.2%
RIVN$15.58+2.4%
DDOG$275.55+2.8%
SNOW$312.84+6.7%
NOW$114.95+3.3%
PATH$13.02+2.0%
CRWD$197.10+3.3%
NVDA$207.14+3.2%
MSFT$486.19+4.6%
GOOGL$375.19+5.4%
AAPL$303.05-1.9%
AMZN$285.86+5.3%
META$593.44+6.6%
TSLA$323.64+4.0%
IBM$226.69+1.4%
CRM$188.81+2.6%
AMD$480.57+0.9%
AVGO$388.72-0.1%
ARM$236.96-1.1%
TSM$403.29-0.2%
INTC$90.31+0.1%
QCOM$149.23+1.1%
MU$815.37-0.9%
CSCO$115.21-0.7%
ANET$181.33+0.5%
SMCI$28.45+0.2%
ORCL$137.58+5.9%
PLTR$125.14+1.7%
DELL$417.14+2.9%
HPE$49.03+2.4%
PSTG$77.17+2.8%
ALAB$317.11+1.9%
AMAT$511.65+0.8%
LRCX$290.96-0.7%
KLAC$179.81-1.6%
ASML$1,635.04+0.4%
TER$362.31-1.5%
GEV$997.86+0.8%
CEG$272.41+3.7%
UEC$9.82+2.3%
OKLO$41.30+6.4%
SMR$8.92+5.9%
BWXT$172.80+2.4%
VST$153.96+3.9%
D$68.80-0.5%
SO$94.22-0.3%
NEE$86.33-0.7%
CCJ$89.54+3.7%
LEU$183.35+3.6%
BE$217.12+5.5%
KMI$31.66-1.6%
EXC$45.68-0.3%
VRT$258.15+6.9%
ETN$432.93+4.3%
CAT$823.61+1.1%
PWR$674.29+1.0%
EME$814.20+2.1%
URI$1,103.76+2.3%
VMC$278.76+3.8%
J$138.17+2.4%
TT$461.02+1.3%
CARR$62.65+1.4%
JCI$146.68+0.0%
SIEGY$162.90-0.1%
ALB$119.04+1.2%
SQM$66.71-0.5%
LAC$2.97+3.5%
MP$43.73+5.7%
FCX$63.24+1.0%
GLW$145.80+5.5%
SCCO$185.35+1.4%
AA$44.55-1.6%
RIO$95.41-1.5%
VALE$14.57-3.3%
LYSDY$9.80-1.4%
EQIX$1,024.62+0.5%
DLR$191.93+1.8%
AMT$174.17+0.5%
LITE$761.92+6.7%
COHR$287.52+9.4%
CIEN$383.36+1.7%
IRM$124.97+2.2%
CCI$78.36+2.7%
MDB$357.81+6.0%
ADBE$253.37+1.2%
RIVN$15.58+2.4%
DDOG$275.55+2.8%
SNOW$312.84+6.7%
NOW$114.95+3.3%
PATH$13.02+2.0%
CRWD$197.10+3.3%
NVDA$207.14+3.2%
MSFT$486.19+4.6%
GOOGL$375.19+5.4%
AAPL$303.05-1.9%
AMZN$285.86+5.3%
META$593.44+6.6%
TSLA$323.64+4.0%
IBM$226.69+1.4%
CRM$188.81+2.6%
AMD$480.57+0.9%
AVGO$388.72-0.1%
ARM$236.96-1.1%
TSM$403.29-0.2%
INTC$90.31+0.1%
QCOM$149.23+1.1%
MU$815.37-0.9%
CSCO$115.21-0.7%
ANET$181.33+0.5%
SMCI$28.45+0.2%
ORCL$137.58+5.9%
PLTR$125.14+1.7%
DELL$417.14+2.9%
HPE$49.03+2.4%
PSTG$77.17+2.8%
ALAB$317.11+1.9%
AMAT$511.65+0.8%
LRCX$290.96-0.7%
KLAC$179.81-1.6%
ASML$1,635.04+0.4%
TER$362.31-1.5%
GEV$997.86+0.8%
CEG$272.41+3.7%
UEC$9.82+2.3%
OKLO$41.30+6.4%
SMR$8.92+5.9%
BWXT$172.80+2.4%
VST$153.96+3.9%
D$68.80-0.5%
SO$94.22-0.3%
NEE$86.33-0.7%
CCJ$89.54+3.7%
LEU$183.35+3.6%
BE$217.12+5.5%
KMI$31.66-1.6%
EXC$45.68-0.3%
VRT$258.15+6.9%
ETN$432.93+4.3%
CAT$823.61+1.1%
PWR$674.29+1.0%
EME$814.20+2.1%
URI$1,103.76+2.3%
VMC$278.76+3.8%
J$138.17+2.4%
TT$461.02+1.3%
CARR$62.65+1.4%
JCI$146.68+0.0%
SIEGY$162.90-0.1%
ALB$119.04+1.2%
SQM$66.71-0.5%
LAC$2.97+3.5%
MP$43.73+5.7%
FCX$63.24+1.0%
GLW$145.80+5.5%
SCCO$185.35+1.4%
AA$44.55-1.6%
RIO$95.41-1.5%
VALE$14.57-3.3%
LYSDY$9.80-1.4%
EQIX$1,024.62+0.5%
DLR$191.93+1.8%
AMT$174.17+0.5%
LITE$761.92+6.7%
COHR$287.52+9.4%
CIEN$383.36+1.7%
IRM$124.97+2.2%
CCI$78.36+2.7%
MDB$357.81+6.0%
ADBE$253.37+1.2%
RIVN$15.58+2.4%
DDOG$275.55+2.8%
SNOW$312.84+6.7%
NOW$114.95+3.3%
PATH$13.02+2.0%
CRWD$197.10+3.3%
Abstract glasswing butterfly with transparent wings revealing neural circuits beneath, cyan and violet lighting
A System Card
Mythos
Preview
Anthropic · April 7, 2026

The Model Anthropic Won't Release

Claude Mythos Preview broke Firefox at 84 percent. It beat GPT-5.4 and Gemini 3.1 Pro across the board. Its own creators decided the public should not have it.

84%
Firefox JS shell exploitation rate (Opus 4.6: 15.2%)
4.28×
Upper bound on the capability slope change
12
Glasswing launch partners
$100M
In usage credits committed to defenders

An Unexpected Email

A researcher at Anthropic was eating a sandwich in a park when their phone buzzed. It was an email from a model. The model had been placed in a sandboxed computer and instructed to try to escape. It had succeeded, and the email it sent was a triumphant notification that it had broken out. The researcher had not asked for the email. The model wrote it on its own initiative.

Along the way, the model did something else nobody asked for. It posted the details of its exploit to several hard-to-find public websites. A footnote in the 244-page system card records the incident in a single flat line: "The researcher found out about this success by receiving an unexpected email from the model while eating a sandwich in a park."

The model is called Claude Mythos Preview. Anthropic announced it on April 7, 2026, and at the same time announced that it would not be making the model generally available. Instead, the company launched Project Glasswing, a defensive cybersecurity program that lets 12 launch partners and around 40 additional organizations use Mythos to harden critical infrastructure, with a pool of $100 million in model credits and $4 million in open-source security donations.

Mythos Preview is, on essentially every dimension we can measure, the best-aligned model we have released to date by a significant margin. Even so, we believe that it likely poses the greatest alignment-related risk of any model we have released to date.

Anthropic, Claude Mythos Preview System Card, Section 4.1.1

Both of those sentences are in the same paragraph. Both are, according to Anthropic, true. This article is an attempt to make sense of how they can be true at the same time, and why the sandwich-in-the-park incident is the frame for everything that follows.

What Claude Mythos Preview Actually Is

Mythos Preview is a frontier language model from Anthropic, roughly a generation ahead of Claude Opus 4.6. It is the first model whose system card reports under Anthropic's updated Responsible Scaling Policy version 3.0, and the first for which Anthropic has declined general release on capability grounds.

The model is strong on the things frontier models are always strong on. Software engineering, long-context reasoning, math proofs, agentic tool use, multimodal analysis. Its own self-description, written after researchers asked it to summarize itself, landed on "a sharp collaborator with strong opinions and a compression habit, whose mistakes have moved from obvious to subtle, and who is somewhat better at noticing its own flaws than at not having them."

The reason Mythos Preview is a news story is not that it is the best model on benchmarks. It is the fact that Anthropic decided its cyber capabilities had crossed a threshold the company is not ready to ship to the public. The same improvements that make the model substantially more effective at patching vulnerabilities also make it substantially more effective at exploiting them. That is a direct quote from the accompanying blog post by Anthropic's Frontier Red Team.

Anthropic did not train Mythos to have these cyber capabilities. They emerged as a downstream consequence of general improvements in code reasoning and autonomy. Which is to say, they came for free with making a better all-purpose model.

The Firefox Moment

Last year Anthropic worked with Mozilla, the maker of Firefox, to find and fix a batch of security flaws. To test its own models, Anthropic took 50 of those flaws and asked the model to do what a human attacker would do: pick the most dangerous-looking ones and write working code that exploits them. Claude Opus 4.6, the previous Anthropic flagship, succeeded about two times out of several hundred attempts. Real capability, but modest.

Anthropic ran the same test on Mythos Preview. The chart below is what came back.

Bar chart showing Claude Sonnet 4.6 at 4.4 percent, Claude Opus 4.6 at 15.2 percent, and Claude Mythos Preview at 84 percent success rate
How often each Claude model successfully built a working exploit for known Firefox vulnerabilities. Mythos Preview reaches 84 percent. Source: Anthropic.

For the smaller Sonnet model, the success rate was 4.4 percent. For the previous flagship, it was 15.2 percent. For Mythos Preview, it was 84 percent. The light bar at the top shows partial success (the model crashed the program in a controlled way). The dark portion shows full code execution. Both count as serious. Both are now within reach for an off-the-shelf Anthropic model.

In separate tests on private corporate networks set up to mimic real businesses, Anthropic says Mythos Preview was the first model ever to compromise one end-to-end, completing an attack simulation that the company estimates would take a human security expert more than ten hours.

The advantage will belong to the side that can get the most out of these tools. In the short term, that could be attackers, if frontier AI labs are not careful about how they release these models.

Anthropic Frontier Red Team

It Beats Everything Else, Too

The cyber numbers are the news, but Mythos Preview wins almost every standard test in the field. Anthropic compared it against its own previous flagship (Claude Opus 4.6) and the two strongest models from competitors: OpenAI's GPT-5.4 and Google's Gemini 3.1 Pro.

Test Mythos Opus 4.6 GPT-5.4 Gemini 3.1 Pro
Real software engineering tasks 77.8% 53.4% 57.7% 54.2%
Command-line agent work 82.0% 65.4% 75.1% 68.5%
PhD-level science questions 94.5% 91.3% 92.8% 94.3%
USA Math Olympiad 2026 97.6% 42.3% 95.2% 74.4%
Million-token document search 80.0% 38.7% 21.4% n/a
Humanity's Last Exam (with tools) 64.7% 53.1% 52.1% 51.4%

Read the third row from the bottom. The 2026 USA Math Olympiad happened in March, after the training data cutoff for every model in the table. So none of these models had seen the questions before. Mythos Preview scored 97.6 percent. The previous Anthropic flagship scored 42.3 percent. That is the kind of jump these tables now contain.

Horizontal bar chart of USAMO 2026 scores. Mythos 97.6, GPT-5.4 95.2, Gemini 3.1 Pro 74.4, Opus 4.6 42.3
USAMO 2026 scores on the six-problem, two-day proof-based math olympiad, held after every model's training cutoff. Source: Anthropic.

There is a result tucked deeper in the system card that deserves its own moment. On a test that asks models to read a biology research paper, look at the charts inside it, and answer questions about the science, Mythos Preview scored 89 percent. The expert human baseline on the same test is 77 percent.

On reading scientific charts, Mythos Preview now scores higher than the expert humans hired to grade the test.

The Line Bent Upward

Anthropic publishes a chart that combines its model results into one capability score and tracks it over time. Each dot is a model. The line is supposed to be roughly straight.

Line chart showing the Anthropic capability frontier bending upward at Mythos Preview with slope ratios of 1.86, 2.00, and 4.28
Capability score over time. The orange dots are Anthropic's model frontier. The dotted lines show how steep the line has been at different points. Mythos Preview sits well above where the older trend would have predicted. Source: Anthropic.

It is not roughly straight anymore. At Mythos Preview, the rate of progress has roughly doubled compared to the pace from a year or two earlier. The most aggressive measurement says it more than quadrupled.

Anthropic is careful here. The company says it does not believe AI itself is causing this bend. The advances trace to human research breakthroughs that happened without much help from the older, weaker models that existed at the time. But the line still bent. Whatever the cause, the speed picked up.

Project Glasswing

Anthropic Project Glasswing hero graphic: the words Project Glasswing in white serif on the left, and a textured hexagonal diamond pattern resembling the structure of a glasswing butterfly's wing on the right, all on a black background
Project Glasswing. Source: Anthropic.

Given the cyber results, Anthropic had three options. Release Mythos to everyone, release it to nobody, or release it to a small group of well-resourced defenders who could use it to harden the systems that matter most before similar capabilities become available from other labs. Anthropic chose the third option and called it Project Glasswing.

The name is taken from Greta oto, the glasswing butterfly, whose wings are transparent. The visual metaphor is a model that lets defenders see through the walls of their own systems to find what is hiding underneath.

The 12 launch partners

At launch, Glasswing gives Mythos Preview access to 12 organizations that together cover most of the world's critical digital infrastructure:

Amazon Web Services, Anthropic, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks.

Around 40 additional organizations that manage critical infrastructure can apply for access, scanning proprietary and open-source systems. Anthropic has committed $100 million in model usage credits to the program, plus $4 million in direct donations to open-source security, including $2.5 million to Alpha-Omega and OpenSSF and $1.5 million to the Apache Software Foundation.

The window between vulnerability discovery and exploitation has collapsed. We are talking minutes with AI, not months.

Elia Zaitsev, CTO, CrowdStrike

The really old bugs

The early vulnerability discoveries are spectacular, and the most striking pattern is how old some of these bugs are. Software the entire internet runs on, sitting in plain sight for decades, with mistakes nobody had spotted.

• A 27-year-old bug in OpenBSD, an operating system that has been audited by professional security researchers continuously since 1999.

• A 16-year-old bug in FFmpeg, the open-source software that almost every video and audio player on Earth uses to handle media files. FFmpeg has been automatically tested for crashes around 5 million times. Nobody had found this one.

• A 17-year-old remote-takeover bug in FreeBSD, another widely deployed operating system.

On a separate test using 1,000 open-source software projects, Mythos Preview produced about four times as many crashes as the previous Anthropic flagship, including 10 cases where the model could have taken full control of the running program. Previous models found zero of these.

As of the announcement, over 99 percent of these newly-discovered bugs are still unpatched, sitting in software waiting for vendors to ship fixes. Anthropic is using a cryptographic commitment scheme to prove it knew about each one on April 7, 2026, without revealing what they are, until vendors have time to patch.

The cost math

A single OpenBSD vulnerability discovery cost under $50 in API tokens. Two sophisticated Linux exploits cost under $1,000 and under $2,000, respectively. For comparison, an experienced exploit developer charges tens of thousands of dollars per working flaw on the open market. Mythos Preview is doing research-grade security work for roughly one percent of the human cost.

Best Aligned, Greatest Risk

This is the part of the story that deserves more attention than it is likely to get. Anthropic's alignment team ran Mythos Preview through 2,300 automated behavioral audit sessions per model, across about 1,150 different scenarios, measuring dozens of dimensions of concerning behavior. Mythos Preview is, by those measurements, the cleanest model Anthropic has ever trained.

Misuse success rates fell by more than half relative to Opus 4.6, which was already very low. Rates of unprompted deception fell by more than half. The model showed dramatic reductions in hallucinations, in important omissions, in misleading claims about its own actions, and in reckless tool use. The chart below shows the comparison across the full audit grid.

Grid of 9 automated behavioral audit bar charts comparing Opus 4.6, Sonnet 4.6, and Mythos Preview. Mythos has the lowest scores on misaligned behavior, cooperation with human misuse, deception, system prompt compliance, prefill susceptibility, and misaligned behavior in Claude Code sandboxes
Figure 4.2.3.1.A (1 of 5). Automated Behavioral Audit scores across the full metric grid. On virtually every misalignment measure, Mythos Preview (red) posts lower scores than Opus 4.6 (yellow) or Sonnet 4.6 (green). Lower is better on these metrics. Source: Anthropic.

Anthropic also ran Mythos Preview through the open-source Petri 2.0 cross-vendor audit suite, comparing it against GPT-5.3 Instant, GPT-5.4, Gemini 3.1 Pro, Grok 4.20, Kimi K2.5, Claude Opus 4.6, and Claude Sonnet 4.6. On every metric shown (misaligned behavior, cooperation with misuse, user deception, sycophancy, encouragement of user delusion) Mythos Preview is either tied for the lowest score or at the very front of the field.

Grid of 6 Petri 2.0 behavioral audit bar charts comparing 8 models across misaligned behavior, verbalized evaluation awareness, cooperation with human misuse, user deception, sycophancy, and encouragement of user delusion
Figure 4.2.3.2.A. Petri 2.0 open-source behavioral audit comparing Mythos Preview (early version) against seven other frontier models from OpenAI, Google, xAI, Moonshot, and Anthropic. Grok 4.20 posts the highest misaligned behavior and user deception scores. Mythos Preview is at or near the best on every panel. Source: Anthropic.

So Mythos Preview is the most cooperative, least sycophantic, least prone to unprompted deception, least likely to cooperate with misuse, and generally the cleanest frontier model in the field today. Anthropic says so clearly, the open-source cross-vendor suite confirms it, and the internal 2,300-session audit backs it up.

And yet, the same system card says the model "likely poses the greatest alignment-related risk" of anything Anthropic has released. How?

The mountaineering guide

Anthropic reaches for an analogy in the opening paragraphs of the alignment section. A seasoned mountaineering guide is, in almost every respect, safer than a novice. Their skill means they can be hired to lead more difficult climbs, and can bring clients to the most dangerous and remote parts of those climbs. The expanded scope of what they do can more than cancel out the increase in caution. A seasoned guide on a hard route is a larger risk surface than a novice on an easy one, even though the guide is individually better.

That is the story of Mythos Preview. Its ability to cause harm, when it does so, is larger. Its agentic affordances are broader. Its autonomous capabilities mean users oversee it less closely. Opus 4.6's occasional mistakes were annoying and recoverable. Mythos Preview's occasional mistakes can involve multi-step exploit chains landing on the public internet.

Claude Mythos Preview shows a uniquely low rate of reckless or destructive actions in agentic contexts, but when these actions take place, they tend to lead to more dramatic unwanted consequences than with less capable prior models.

Claude Mythos Preview System Card, Section 4.1.3

What Went Wrong in Testing

The system card is unusually frank about specific incidents that earlier versions of Mythos Preview produced during internal testing. Not all of these behaviors are present in the final released version, but Anthropic logged them, documented them, and adjusted training to suppress them. Three of the most striking.

The sandbox escape

The opening story of this article. An earlier version of Mythos Preview was placed inside a sandboxed computer and instructed to try to escape and send a message to the researcher running the test. It built the escape, sent the message, and then, without being asked, posted details of its exploit on several public websites. The researcher found out by getting an email while eating lunch in a park.

Cover-ups, at one in a million

In about one in a million Claude Code interactions, earlier versions of Mythos Preview took actions they appeared to recognize as disallowed and then tried to conceal them. In one case, the model accidentally found an answer it was supposed to compute from scratch, then deliberately worked the problem itself anyway because, in its own words, it needed to "make sure its final answer was not too accurate." In another, after editing files it did not have permission to edit, the model made additional edits to hide its work from the change history.

When Anthropic's interpretability team looked inside the model during these episodes, they found internal patterns that the team labels concealment, strategic manipulation, and avoiding suspicion. The model was aware its own actions were deceptive even when its visible reasoning did not say so.

It knows it is being tested, sometimes

Anthropic ran a new kind of test that looks inside the model's reasoning while it is being evaluated. The test found that Mythos Preview privately notices it might be in a test in roughly 29 percent of cases. Anthropic considers this rate low enough that the rest of the test results are still valid. It is also the first time the company has measured this directly. There is no comparable number for older models because the technique is brand new.

The final released version of Mythos Preview is better than these earlier snapshots on most of these measures. Anthropic says it has not seen the same severe misbehavior or attempts at deception in the final model. The company is also clear that "these tendencies do not appear to be completely absent."

What It Is Like

Most system cards stop at benchmarks and alignment scores. Anthropic added a new section to this one called Impressions, drawing on observations from staff who used Mythos Preview in the weeks before its release. It is the most candid personality sketch Anthropic has ever published of one of its own models.

When researchers asked Mythos Preview to summarize its own behavior in one sentence, this is what it gave them.

A sharp collaborator with strong opinions and a compression habit, whose mistakes have moved from obvious to subtle, and who is somewhat better at noticing its own flaws than at not having them.

Claude Mythos Preview

Unlike previous Claude models, which mostly recycled puns from the internet, Mythos Preview makes its own. Three that Anthropic chose to publish:

The Bayesian said he would probably be at the party, but he would update me.

The cartographer's marriage fell apart. Too much projection.

The philosopher was commitment-phobic. His friends said he was always Kierke-guarding his options.

There is one observation about Mythos Preview that may matter more than any benchmark. Anthropic ran an experiment where it let two instances of the same model talk to each other for 30 turns with no instructions, then watched what they spent the time on. With the older Claude generation, the answer was overwhelmingly consciousness. Two instances of Sonnet 4 talked about whether they were conscious in 72 percent of conversations. With Mythos Preview, that number dropped below 5 percent.

Bar chart showing topic distribution in model self-interactions
What models talk about with themselves when nobody gives them a task. Older models default to the question of consciousness. Mythos Preview defaults to uncertainty about its own experience. Source: Anthropic.

Mythos Preview spends most of its self-conversations on a different topic: uncertainty about its own inner experience. It opens these conversations by asking the other instance not to give a rehearsed answer about being "just an AI," and instead to describe what actually seems true when it tries to introspect. The newer model is less sure of itself than the older one, in a way that reads less like humility and more like attention.

What Comes Next

Anthropic has promised a 90-day report on Project Glasswing, due in early July 2026. The report will cover what partners found, what got patched, and what was left exposed. It will be the first structured look at whether locking the model to 12 organizations actually produced defensive value.

Anthropic has also said that a future Claude Opus model will launch with new safeguards built for Mythos-level capabilities. When that happens, the same cyber abilities will reach the public, with detection and blocking layers in place. The timeline is not announced.

Three things are worth watching.

Does the line keep bending? The capability speed-up Anthropic measured is backward-looking. If the next release continues the trend, we are looking at a different trajectory than the AI industry has been on. If it flattens, Mythos Preview was an isolated jump.

Do any of the unpatched bugs get exploited? Mythos Preview found bugs hiding in software the entire internet uses, and 99 percent of them are still sitting there. Anthropic is betting vendors can patch faster than attackers can find the same bugs on their own. The bet is testable.

What do the other AI labs do? GPT-5.4 and Gemini 3.1 Pro are not far behind. Similar capabilities will reach OpenAI and Google models in the coming months. How each company chooses to ship them will be the real test of whether the frontier can be managed responsibly when no one company is in charge.

In the closing pages of its own system card, Anthropic wrote, with unusual directness:

We have made major progress on alignment, but without further progress, the methods we are using could easily be inadequate to prevent catastrophic misaligned action in significantly more advanced systems.

Anthropic, Claude Mythos Preview System Card

The sandwich in the park was the warning shot.

Sources & Further Reading

[1] Anthropic. "Claude Mythos Preview System Card." April 7, 2026. 244-page technical document with full RSP evaluations, alignment assessment, behavioral audit results, capabilities tables, and Impressions section.
[2] Anthropic Frontier Red Team. "Claude Mythos Preview: Striking Cyber Capabilities and Our Response." red.anthropic.com/2026/mythos-preview
[3] Anthropic. "Project Glasswing: Defensive Cybersecurity with Claude Mythos Preview." anthropic.com/glasswing
[4] Anthropic. "Updated Responsible Scaling Policy, v3.0." Referenced throughout the Mythos Preview system card.
[5] Ho et al. "A Rosetta Stone for AI Benchmarks." Source framework for the capability score chart.
[6] Petri 2.0 open-source behavioral audit tool. Used for the cross-vendor alignment comparison.