A billion agents, and most of them are standing still. Illustration.
What Stalled Our Agent Was a Human Check
Our research combines my firsthand agent tests with a fleet model built from the task up.
I’ve been using Grok’s agent for over a month, and Muse since it came out. I tried doing a claim form for a class action settlement on both. I gave up on Grok’s agent because I couldn’t even click the Cloudflare human test. Muse took about five minutes, because as it went through each screen of filling out the information, there was a human check from Cloudflare. I had to physically verify that I was a human, but it eventually finished the task.
A settlement claim is exactly the find-lost-money errand consumer agents are sold on, and exactly the kind of site that checks for humans, because stopping fake claims is its job. The obstacle we could see was a human check on every screen. The test shows the friction an authorized user encounters; it does not measure the compute behind the task.
On Monday the market went the other way. Meta’s Muse set off a rally in the CPU names. Per Yahoo Finance, Arm rose 13% to $312.46 (Meta is the lead partner on Arm’s AGI CPU), Intel 12% to $121.38, Qualcomm 9.3%, and AMD about 9%, which took it past $1 trillion in market value for the first time. Meta itself rose 11% to $741.25, and Wells Fargo raised its target to $796. AMD became a trillion-dollar company on a CPU story about agents for everyone.
We named that slow part in January. As we wrote in The Verification Gap: Who Audits the Agent Swarm?: “We’re building infrastructure to deploy millions of agents without infrastructure to verify they’re working correctly. The compute is ready. The control plane isn’t.” That piece imagined 500 agents inside one bank. A consumer agent is one per person, loose on the open web, and the open web was built to stop bots.
Item two on that piece’s checklist, hierarchical verification, ended on this: “Critical actions require human-in-the-loop approval. The verification intensity must scale with consequence.” The claim form got the check we reserved for critical actions, on every screen, because the site could not tell an authorized agent from a scraper. An enterprise agent borrows an employee’s login. A consumer agent has nothing to borrow, so the check lands on the user’s own hands.
The second test was on my phone. I upgraded to the new Siri AI, yet it still doesn’t have agentic app control. It won’t press buttons or do computer control, like how these models are running on virtual CPUs. I think that is needed to really take it to the next level. I asked it to disarm my home alarm, and it answered: “I can’t interact with elements on your screen directly. You’ll need to tap the ‘Press to Disarm’ button in the ADT Pulse app yourself.”
Both readings hold. A phone assistant needs a controlled way to act inside other apps. That could run locally or in the cloud; our fleet model sizes cloud browser sessions. And a disarm is a critical action, so a Siri that hands it back is item two working: the claim form over-checked a routine step, and the phone checked the one step that deserved it. The next level needs app control and a way to vouch for the critical steps.
The test raised two questions: what does an agent cost to run, and what lets it finish the job? Our fleet model tackles the first. The human checks, Amazon’s block and the new merchant partnerships expose the second. For investors, the opportunity depends on which problem a company gets paid to solve.
We call the second problem The Gap Went Retail. January’s enterprise question was whether the agent’s work was correct. The consumer version also asks whether it is authorized to act. A checkpoint could become a business, but only if someone pays for it; the partner path may keep that value inside the platform and merchant.
We have been running our own Amazon test for weeks: Grok’s agent, told to buy an Amazon-exclusive NeeDoh, a $19.99 squishy toy that comes back in stock one or two at a time. Its latest reports, in its own words. At 11:32 AM it “Hit Buy Now same run, but it sold out before Place Order.” At 1:02 PM it “Tried Buy Now twice from my computer (your Mac was offline); both times checkout said unavailable.” At 3:05 PM, “Buy Now hit checkout, qty dropped to 0, then the page went unavailable.” No charge any time, and every report ends “Still watching.” The agent calls it a stock race. Amazon doesn’t like bots. Its reports describe a stock race at the last click; they do not measure compute use. Two details matter for the CPU math. The agent runs on the owner’s Mac when it can and falls back to its own computer when it can’t, and it says orders land more often from the Mac. And an agent that has been watching for weeks is not the six-minute errand in our base case.

This week showed all three ways it can go. Amazon blocked Muse from its retail site starting September 20. Per GeekWire and Bloomberg, Meta did not tell Amazon, the agent does not identify itself when it browses, and it appears to capture and store customer credentials; shoppers see pop-ups saying Muse violates Amazon’s terms. Amazon has moved against Google’s and OpenAI’s shopping agents before and sued Perplexity last November. That is the open-web path closing at the biggest merchant. The claim form is the path left open but checked. And Expedia announced the third, posting on X: “Here at Expedia, we love planning trips. Soon, your personal AI agent will too. We’re joining @Muse: tell it where you’re headed and it can work with Expedia to sort your hotels, and everything in between.” On that path Meta and the merchant settle who the agent is in advance, and a direct integration can replace browser interaction.
That path filled up inside two days. On Monday, Shopify said Muse will browse Shopify-powered stores and check out through Shop Pay, and per Yahoo Finance Mark Zuckerberg wrote “More partnerships like this coming soon.” PayPal announced its own. On Tuesday, after Expedia, Instacart posted on X: “Instacart is coming to @Muse! Soon, you can connect Instacart, say “Taco Tuesday,” and turn the idea into a cart from your favorite store. Then check out and get groceries delivered to your door.” In the week the biggest merchant blocked Muse, four companies signed on to it.
Expedia chose the partner path the day after Muse’s chip rally. A direct integration can put authorization inside the platform and merchant; the commercial terms remain undisclosed.
Blocked, checked or partnered: which path the large merchants choose decides who gets paid for verification. “Doesn’t identify itself” is the exact problem Cloudflare’s signed agents are built to solve, a cryptographic signature that lets a site tell an authorized, user-directed agent from a bot. That is the mechanism that could turn a human check on every screen into a pass for a trusted agent.
Our January rule scaled verification to stakes. Ben’s open-web test repeatedly asked for a human; direct partnerships offer another route. The diagram is a hypothesis about who could get paid, not disclosed agent revenue.
Zuckerberg’s superintelligence push has a business model. Meta already uses information about people’s activity to personalize advertising, and its published policy extends that to AI interactions. The price of a free agent may be paid in data and attention: the more useful the assistant becomes, the more it can learn about what a user wants. For Meta, a request for help can also be a commercial signal.
A consumer agent sits idle most of the day, and when it does transact, someone has to vouch for it.
I personally didn’t give Muse access to my email. In an Oppenheimer survey of 1,500 US consumers around the launch, reported by Yahoo Tech, 8% said they would trust Meta with their passwords, against 30% for Google, 23% for Apple and 16% for ChatGPT, and 58% said they would give them to no AI agent at all. Meta’s answer is architecture. Per TechCrunch, Muse runs on its own “dedicated, secure computer with its own browser,” and a separate agent Meta calls Sentinel runs on the same virtual machine, kept apart from Muse. That is a checkpoint Meta built inside its own platform, and it is one virtual machine per user. Our fleet model gives an idle agent no machine. If Meta kept every user’s machine warm between tasks at our 2GB each, a billion users would hold about 2 exabytes of memory; suspended to disk, they hold almost none. Meta has not said which.
A Billion Muse Agents Are a Small CPU Line
This is where the software and hardware yin and yang becomes co-design. The agent’s software determines how the hardware has to be set up. CPUs run tools and browser sessions; GPUs or other accelerators serve the models. Memory, storage and interconnects tie that work together. An agent that sleeps between tasks creates a different workload from one that stays active, while keeping state in memory differs from suspending it to storage. The hardware architecture, in turn, shapes what the software can do efficiently. We have to evaluate that integrated system. A billion registered agents is not, by itself, a hardware specification.
We built our own fleet model, and a billion Muse users in our base case buy about $1.3 billion of processors a year. It is a model, not Muse telemetry; Meta has published no usage data. Users times tasks per day times active minutes per task gives the share of the day an agent runs a browser. We provision for a peak three times the average, give each active browser session one CPU thread and 2GB of memory, run servers at 70% CPU, and let idle agents hold no machine. The server is a single-socket, 256-thread box priced at the AMD EPYC 9755‘s $10,931 one-thousand-unit list price, and the fleet is built over three years. It runs out of CPU before memory under these assumptions. A billion daily users doing five tasks each is an adoption-upside scenario, not a forecast. The dollar figure covers processor purchases only; it excludes memory, the rest of the server, power and operations.
Four cases at a billion users. Light, two three-minute tasks a day, is 0.42% concurrency and about $0.25 billion a year of CPUs. Base, five six-minute tasks, is 2.1% and about $1.3 billion a year, $3.8 billion installed. Heavy, ten twelve-minute tasks, is 8.3% and $5.1 billion a year, which crosses our own $5 billion line. Extreme, twenty twenty-minute tasks, is 27.8% and about $17 billion a year; that is a coding agent’s duty cycle, not a shopper’s. The claim form is the reality check: about five minutes with a check on every screen, right at the base case’s six, and friction that holds sessions open pushes toward Heavy. Base works out to about 45 million cores and Heavy to about 179 million, against the roughly 120 million Rene Haas told Arm’s March event one agentic gigawatt needs.
At the July Muse Spark 1.1 API price used in our model, the Base scenario implies about $533 billion a year of model calls. That is a pricing yardstick, not a forecast of Meta’s spending. A 90% discount brings it to about $53 billion, roughly 42 times the $1.27 billion processor-purchase line. These are different measures: model-call prices include more than hardware, while the CPU figure is chip-only. We also hold tokens per task constant as task duration changes, so the ratio across scenarios depends on that assumption. The subscription-price sensitivity below brings the gap much closer.
Our Base scenario implies $1.27 billion a year of processor purchases and $53 billion of model calls at 90% below the July API list price. The 42x ratio compares chip purchases with a service-price yardstick; it is not Meta’s internal cost ratio.
Even the CPU Dollars Go to the Best Token Factory
We priced the fleet at the EPYC 9755 because it is a public yardstick, not because we think AMD wins the socket. NVIDIA used the same chip as its yardstick in July. In NVIDIA Vera: When CPU Latency Becomes GPU Economics we reported NVIDIA’s estimate of “Vera 925, EPYC Turin 9755 898.” “That is 176 cores outscoring 256.” Vera also showed “12.7 GB/s of memory bandwidth per core against Turin’s 3.1,” all “vendor-run, estimated on a non-SPEC-compliant reference system.” The line that matters here: “Tokens per second prices the green bars only. An agent task is billed end to end.”
On August 3 we asked AMD for its side in An Open Letter to Lisa Su: The Score With No Spec Sheet: four questions on the Venice-versus-Vera claims, noting that NVIDIA had published the whole SPEC suite while AMD had published one estimated number. AMD never responded. Our August 24 audit put the other side plainly: “Vera is the first CPU built for that job,” with NVIDIA’s own benches, not independent ones, putting Vera Rubin NVL72 at up to 30 times Blackwell’s throughput per megawatt on agentic coding.
Vera Rubin is probably the better agentic CPU and total system that works for agents. At the end of the day, you have to have a token factory, and it’s tokens per megawatt.
We said it about chips in April, in The Chip Is Dead, Long Live The Factory: “A custom ASIC that beats NVIDIA on cost-per-GPU-hour but delivers a fraction of the tokens per megawatt is not a bargain.” A CPU socket works the same way. The API-price scenario puts the larger spending opportunity in the model-serving system. The size of that opportunity depends on realized token costs and work completed. That is the tension in Monday’s tape: AMD crossed $1 trillion on a CPU story, and in our model the CPU is the small line. AMD’s real answer is Helios, where our July call put meaningful volume in the second half of 2027.
Our March CPU Call Now Carries a Duty Cycle
On March 25, the day Arm jumped about 20% on its first production chip, we wrote in The Agentic CPU that “Agentic AI doesn’t reduce CPU demand. It explodes it.” In May we corrected a unit error in the earlier CPU-to-GPU comparison: about 9:1 at rack level and 2.6:1 at chip level. Our May always-on inference thesis remains the route toward the heavier-use cases. Muse narrows the call again: agentic AI drives CPU demand where agents run most of the hour, in coding, research and reinforcement learning, which is our Extreme case. Our Base consumer scenario instead assumes agents work a small slice of the day; actual usage could be heavier. A user count only matters once you multiply it by that slice.
One call narrowed, two caveats now bind. The agentic CPU call carries a duty cycle, and our July warning that Muse can run off non-NVIDIA silicon limits what Muse means for NVIDIA.
Agents Burn Tokens, and the Price Sheets Split
An agent will essentially just be blowing a lot of tokens, so expect token pricing to come down. This month’s two frontier releases split on it. Anthropic’s Claude Opus 5.5, released September 22, lists at $4 in and $20 out per million tokens, 20% below Opus 5, and Anthropic estimates it costs about 40% less to run. OpenAI’s GPT-6 Astra lists at $10 and $50 on OpenAI’s price page, up from GPT-5.5’s $5 and $30, sold on a lower cost per task because it uses fewer tokens. Both now sell on cost per task, and their per-token prices went opposite ways. The evidence is mixed: one frontier vendor cut its token price and another raised it. Our hypothesis is that routine tasks face price pressure while differentiated capability can command a premium. Cost per completed task is the comparison that will test it. Meta priced the commodity end in July, when Bloomberg reported Muse Spark 1.1 at $1.25 in and $4.25 out. On one illustrative web task, the model calls cost 20 to 70 times a browser session at Cloudflare’s $0.09 a browser-hour.
Meta’s subscription plans show why the API-price multiple should not be read as a forecast. Muse is free up to a usage limit, and Meta’s help page lists Power at $20 a month for 500 million Muse tokens a week, and Maximum at $100 for 3 billion. At full use, Power works out to about $0.0092 per million Muse tokens. Our Base scenario assumes 200,000 input and 10,000 output model tokens per task, about 7.35 million a week per user. If a Muse token equaled a model token, valuing that workload at the full-use rate would imply about $3.5 billion a year, roughly 2.8 times the CPU line. Meta has not established that equivalence. This is a conditional sensitivity, not a bill a user pays or an estimate of Meta’s internal cost. The token assumptions and realized cost can change the size of the opportunity substantially.
On one illustrative web task the browser box costs about a cent and a half, and the model calls cost 20 to 70 times that at list prices. Halve the tokens and it is still 10 to 35 times. An agent is mostly a token bill.
Below the paywall: where we stand on NVIDIA, Arm, Cloudflare, Expedia and Meta; the strongest case against our model; and the specific evidence that would change our view by March 31, 2027.









