Read that carefully. The engineer closest to the product is not proud of the model. He is proud of the decision to preview it. “Responsibly preview” is the corporate phrasing for a constraint the safety framing is wrapped around. And the constraint, once you run the numbers, is not primarily about safety.
It is about hardware.
Anthropic unveiled Project Glasswing today alongside AWS, Apple, Google, Microsoft, NVIDIA, Cisco, CrowdStrike, Palo Alto Networks, JPMorgan Chase, and the Linux Foundation. The vehicle is a new frontier model, Claude Mythos Preview, which in weeks of red-teaming autonomously surfaced thousands of zero-day vulnerabilities across every major operating system and web browser. It found a 27-year-old remote-crash bug in OpenBSD and a 16-year-old FFmpeg flaw that automated fuzzers had hit five million times without catching.
But the most important document Anthropic published today isn’t the announcement. It’s the pricing page.
The Price Sheet Is a BOM
Mythos Preview is listed at $25 input / $125 output per million tokens. Opus 4.6, the model Anthropic actually ships today, sits at $5 / $25. The ratio is a clean 5×. Not 4.7×. Not 5.3×. Exactly five times per token, at both ends of the sheet.
Frontier labs round preview pricing to the nearest clean multiple of their flagship because it is the cleanest way to signal compute cost without admitting it. When a lab prints 5× on a preview model, they are telling you the unit economics of serving one Mythos query are approximately five times the unit economics of serving one Opus 4.6 query on the hardware they actually have access to today.
That 5× is the first leaked artifact from a bill of materials most of this industry pretends doesn’t exist. The HBM4 stack cost. The CoWoS substrate. The optical interconnect between racks. The power budget per cabinet. Every physical constraint across the AI buildout stack: silicon, memory, interconnect, packaging, power. Collapsed into a single number on a public pricing page. The API price sheet is the BOM wearing software clothing.
Now run the demand side.
If Anthropic opened Mythos to every paying Claude user tomorrow, workload routing would migrate toward capability on contact. Every Claude Code session, every Cowork task, every agentic workflow currently running on Sonnet or Opus would reach for Mythos the moment it delivers better results. A conservative estimate is that 30 to 50 percent of current inference demand would shift to a model roughly three times more expensive per task. The per-token price is 5×, but Mythos uses fewer tokens per task because it reasons more efficiently. Per-token price is the BOM signal. Per-task cost is what actually drives the fleet math.
That alone implies a 1.6 to 2.5× step-up in total inference compute from the existing user base. But capability does not replace demand one-for-one. It expands the universe of things users attempt. Agentic sessions run longer because the model can actually finish the task. Workflows that were too unreliable to deploy in production become viable. Context windows fill up because the model can use them. Stack the demand expansion on top of the per-query cost delta and a realistic serving estimate lands closer to seven times Anthropic’s current inference fleet.
And that is before counting the continuous agentic security workloads Glasswing is about to unleash across every kernel, every browser, and every piece of critical infrastructure on Earth. Those workloads do not run interactively. They run 24/7. That is not a workload category. It is a new compute category entirely.
This is why Mythos is gated to roughly 50 partners. Anthropic cannot physically build seven times their current inference fleet in any reasonable timeframe. Nobody can. Not yet.
The Benchmarks Are a Memory Test
Look at where Mythos actually improves versus Opus 4.6. SWE-bench Verified jumps from 80.8% to 93.9%. SWE-bench Pro goes from 53.4% to 77.8%. Terminal-Bench 2.0 moves from 65.4% to 82.0%. CyberGym from 66.6% to 83.1%.
These are not reasoning benchmarks in the traditional sense. GPQA Diamond, the closest proxy to pure knowledge recall, barely moves, from 91.3% to 94.6%. The real step-function is in agentic coding and long-horizon tool use. Those workloads share a single underlying requirement: holding enormous context in working memory. Entire codebases. Multi-hour tool-use traces. KV caches that grow continuously across hundreds of turns without collapse.
That is not a compute problem. That is a memory bandwidth problem. Long-context agentic inference is fundamentally bound by how fast you can stream the KV cache in and out of HBM at every generation step. Double the context window and you do not double the FLOPs, you double the bytes-per-token that have to cross the memory bus. The compute units sit idle waiting for data. This is the mechanical constraint I laid out in The Memory Wars, and Mythos is the first public model clearly scaled into the regime where it binds hardest.
Which is why the 5× pricing makes sense mechanically. A model with a materially larger effective context window, longer reasoning traces, and deeper tool-use loops does not cost 1.5× more to serve. It costs several multiples more, because every incremental token burns HBM bandwidth on memory hardware already sold out for the next three years.
The Co-Design Confession
Mythos is the first public evidence that frontier AI capability is no longer separable from the silicon, memory, and interconnect it runs on. The pricing, the gating, the benchmark pattern, and the engineer’s careful language are all consistent with a single interpretation: Mythos is not a software release gated by compute. It is a hardware-constrained model whose price tag is the silicon talking through the API.
This inverts the argument I made earlier this year. In The Fourth Piece, I argued that inference economics collapse when you co-design the model against purpose-built silicon that bypasses the memory wall. Mythos at 5× Opus pricing is the inverse signal. It is a model so bandwidth-hungry and so context-heavy that Anthropic cannot arbitrage its way out via dataflow tricks or architectural shortcuts yet. Either Mythos is running on GPUs hitting the memory wall at full tilt, or it is running on TPUs through the reported $30 billion Google deal and the pricing reflects TPU scarcity. Either way, the 5× is hardware talking.
Mythos is the first internal model a frontier lab has publicly admitted it cannot ship. OpenAI and Google DeepMind almost certainly have their own. The compute-constrained regime is not Anthropic-specific. It is industry-wide.
And Anthropic just put numbers on it. The day before Glasswing, they announced an expanded partnership with Google and Broadcom for multiple gigawatts of next-generation TPU capacity coming online starting 2027. The CFO, Krishna Rao, disclosed the revenue math that makes the buildout rational: Anthropic’s run-rate revenue has crossed $30 billion, up from roughly $9 billion at the end of 2025. That is a 3.3× increase in four months, arguably one of the fastest enterprise revenue ramps in software history. Business customers spending over $1 million annualized have doubled in under eight weeks, from 500 to more than 1,000.
Read the sequence. Multi-gigawatt TPU deal on Monday to telegraph the hardware response. Glasswing on Tuesday to telegraph the capability that justifies it. $30 billion run-rate to telegraph the demand underneath both. This is the mechanical reality behind every hyperscaler capex guide that the sell side keeps calling a bubble. It is not a bubble. It is a disclosed constraint with a disclosed solution and a disclosed customer base willing to pay for it..
The Broadcom inclusion is the Fourth Piece thesis playing out in contract form. Broadcom is the ASIC partner behind Google’s TPU co-design, the custom silicon arm that has quietly become the second most important chip company in the AI stack. Anthropic signing them directly, rather than only through Google Cloud, is a co-design signal: the model, the TPU, and the custom silicon around the TPU are all being designed as one system.
The Bear Case
Two things could weaken this read.
The first is that Mythos may be a deliberately oversized research artifact rather than a model intended for production. Anthropic’s system card language suggests the gating is as much about offensive-cyber uplift risk as it is about compute economics. If that framing is dominant internally, the 5× pricing may reflect a risk-adjusted premium rather than cost-to-serve, and the implied hardware delta is smaller than the math above suggests. A smaller delta means a smaller fleet expansion and a smaller downstream demand signal for the supply chain.
The second is that efficiency gains could partially absorb the capability jump. Speculative decoding, better KV cache management, mixture-of-experts routing, and continued per-token cost declines at the silicon level all compound in the direction of “serving Mythos costs less next year than it does today.” If those gains arrive faster than demand expands, the 7× fleet estimate compresses toward 2 to 3×. Still a meaningful step-up, but not the regime shift the bull case implies.
Neither risk changes the direction of the trade. Both change the slope.
So What?
Stop reading the benchmarks. Start reading the pricing. Every frontier model release from here forward will ship with a price sheet that functions as a leaked BOM, and the BOM is the investment thesis. The companies capturing the margin are the ones sitting at the physical chokepoints where demand meets capacity: the foundries, the memory oligopoly, the optical component suppliers with real InP and EML capacity, and the power-generation names whose backlog is measured in years rather than quarters.
Boris Cherny is right that Mythos should feel terrifying. But the terrifying part isn’t what the model can do. The terrifying part is that a frontier lab just quietly confirmed, on the record, that they have built the most capable model in the world and cannot afford to run it. The constraint is not intelligence, or talent, or capital. It is atoms.
We are not late to this trade. We are not even at the starting line.
Coming Up
My OFC 2026 recap ships this week. The Lumentum laser moat, the Tower manufacturing backbone, the CPO reliability data, the NVIDIA $6 billion optical supply chain commitment, and a Credo reassessment are all in there. Paid subscribers get it first.
Next week I’ll be at the Jefferies Private Internet Conference, meeting directly with Ayar Labs and several other private companies building the AI infrastructure stack. Two weeks after that, I’ll be at Google Cloud Next on April 23rd, with meetings lined up across the hyperscaler and custom silicon side. Field reporting and on-the-ground thesis updates will flow to paid subscribers in real time.
And if you missed it, The Token Dollar lays out the framework that made today’s Glasswing announcement legible in real time. I’ve heard great feedback from fellow subscribers on that piece.
If you’re attending either conference and want to connect, reply to this email or reach out on X @benitoz.
Related BEP Research
The Memory Wars: Why NVIDIA’s 2028 Architecture Ends the AI Chip Competition
The Packaging Paradox: Why CoWoS, Not 2nm, Is the Real AI Bottleneck
The Token Explosion: Why GTC 2026 Was About Infrastructure, Not Chips
Resources
Disclosure: The author holds positions in NVIDIA and related semiconductor investments, including LITE, CRDO, ALAB, LSCC, TSEM, ORCL, and BE. This is investment research, not advice. Do your own work.








Thanks for this different take. Very useful. Would it be possible for you to take a deep dive on the Perkinamine polymer developed by Lightwave Logic?
Reading this, and your previous posts, I wonder if there's any particular reason that MU isn't among the tickers listed in your disclaimer.