“With an open model like NVIDIA Nemotron, a LangChain harness, the NVIDIA OpenShell runtime, and a company’s own data, every enterprise can build custom agents that understand its business, use its tools, and turn knowledge into action.”
That’s Jensen Huang on Wednesday, launching the NemoClaw Deep Agents Blueprint with LangChain. In the fireside released alongside it, with LangChain’s Harrison Chase, he went further than the product pitch: “most companies are built on business processes. … In the future most companies will be built on harnesses.” That’s his answer for how agents actually reach the enterprise. A fast open model, Nemotron, does the specialist work. A harness runs the loop around it. And the company deploying it owns the stack.
It’s been one of the busiest weeks the AI world has put together this year, and almost none of the coverage connected the events that mattered. So here’s the week read as one story. Pour yourself a coffee; this one runs long.
Huang sells the silicon every piece of that stack runs on, so his pitch arrives with a vendor discount attached. It also happens to be the key to everything else that shipped. Between July 7 and 9, NVIDIA published its case that the CPU, of all things, is now an AI part, and the blueprint moved LangChain’s own agent eval roughly 10x on cost through configuration alone. Cursor and xAI shipped a model trained on a harness’s telemetry, OpenAI shipped a model family that leads its pitch with token efficiency per task, and Meta charged for a model for the first time in its history.
Cover them separately and you get five product stories. I think they’re one. Somewhere between Tuesday and Thursday, AI stopped selling talk and started selling finished work. The unit that matters now is not the token. It is the completed task, and every stage of the stack, from config files to a CPU, just started competing on it.
The Model Talks. The Harness Does the Work.
Strip the jargon and a large language model is a brilliant mind that can only talk. It can’t run code, clone a repository, execute a test suite, or click a button. Everything that turns talk into work comes from the machinery wrapped around the model: the system prompt, the tool definitions, the middleware passing messages, the logic that spawns sub-agents, the sandbox where the work executes. That machinery is the harness.
You’ve probably used one this week without the word for it. Claude Code, Cursor, and Codex CLI are harnesses. Each wraps a model in its own prompts, tools, retry logic, and context management, and sells the package as a coding agent. The model is the engine; the harness decides which tools it reaches for, when it retries a failed test, what context it keeps, and when it stops. Anyone who’s swapped the same model between two of those tools has felt the difference. One setup one-shots the refactor; the other burns twenty minutes calling the wrong tools on identical weights. Adel El Hallak, NVIDIA’s VP of Product for AI, posted the cleanest definition of the week on LinkedIn Thursday morning, “An agent is more than a model. It is a model plus a harness,” and named the practice of tuning the scaffolding instead of the weights. “We call that harness engineering.”
Part of what the market has been pricing as model capability is harness capability. Tokens are an input price. A completed task consumes token bursts, but also tool calls, repository clones, test runs, sandbox time, failed branches, and retries, and the harness decides how much of that is waste. Dollars per completed task prices the whole loop, retries included.

This is what I meant in The Compute Confession when I argued that a frontier price sheet is a leaked bill of materials. A token price leaks the hardware bill underneath the model; a task price leaks the model, the harness, and every wasted call between them.
Hardly anything referees this in public. There’s no fixed task set and no agreed success threshold, so every vendor scores itself.
Every Launch This Week Was Selling Tasks
Start with the incumbent. OpenAI opened GPT-5.6 Sol, Terra, and Luna on Thursday at unchanged or lower token prices; Sol launched at GPT-5.5’s exact sheet, $5 in and $30 out per million tokens. The pitch went to the new unit instead, with Sol billed as 54% more token efficient on agentic coding tasks. Altman told CNBC: “Every enterprise now is thinking about spend and the value they’re getting in exchange for AI, and this is what we really want to do.”
Meta ran the same play from the price end. Muse Spark 1.1 is the first model Meta has ever charged for, and per Bloomberg’s coverage it lands at $1.25 in and $4.25 out per million tokens with a 1M-token context window, 4x below GPT-5.5 on input. Coverage of Meta’s eval tables, vendor numbers all, shows the aim. It sits near the top on tool use (JobBench 54.7 against 48.4 for Opus 4.8; MCP Atlas 88.1) and a step back on one-shot coding (SWE-Bench Pro 61.5 against Opus’s 69.2). That price sheet is built for task volume, not benchmark trophies.
And Grok 4.5 only makes sense in the task lens. Cursor and xAI announced it as the first jointly trained model between a frontier lab and a harness company. (The lab’s own launch page styles the entity “SpaceXAI”; I’m writing xAI until the consolidation is confirmed.) The training corpus is the tell: trillions of tokens of Cursor telemetry, priced at $2 in and $6 out per launch coverage. A tuned profile bends the harness to the model. Grok 4.5 bends the model to the harness. Either direction, the thing being optimized is finished work per dollar.
The other two launches were infrastructure, and they bookend the story. LangChain and NVIDIA’s blueprint shipped a per-model harness profile as pure config, no fine-tuning, carrying a 10x cost claim on their own eval that I take apart below. In the fireside, Harrison Chase names that comparator: Claude Opus, at 87 to Nemotron's 86, ten times the cost.And NVIDIA’s Vera CPU posts landed July 7, two days before any of the model news, pitched in NVIDIA’s own press language as a shift “from cores per dollar to tokens per dollar.” A CPU launch marketed in AI-economics terms is itself a tell, and the engineering behind it is the strongest confirmation of the whole thesis.
That’s the free half: the vocabulary and the news. Behind the paywall is the part that makes it investable. The mechanical proof of the Denominator Stack, including the exact figures on the config file that moved an agent eval an order of magnitude with the weights untouched, and the enterprise that measured the same lever on its own multi-million-line codebase. The only third-party board pricing coding agents in dollars per task, what its same-day rerun of GPT-5.6 did to the open-versus-closed narrative, and why the loser of that reading isn’t who you’d guess. The purpose-built-silicon read on NVIDIA’s Vera CPU. The bounded NVDA claim, and the one company that may end up selling the loop to everyone who can’t build it. Four ranked risks, including the one aimed at my own framing, and the watch items for the next two weeks.
The Referee Ratified the New Metric Same-Day
Every model launch markets efficiency, and two things separate this one. The token price stayed put while the whole pitch moved to relative task cost, and the pitch got checked within hours by a board OpenAI doesn’t run. DeepSWE, a third-party benchmark from Datacurve, is built on 113 long-horizon coding tasks written from scratch across 91 real repositories, with hand-written verifiers, every model run on the same frozen harness so the models are comparable. Late Thursday it re-ran its leaderboard with the new family.
Sol tops the board at 73% pass@1 for $8.39 a task, against Claude Fable 5’s 70% at $21.63. Terra matches Fable 5’s pass rate at $4.95, and Luna matches GPT-5.5’s 67% at $3.03 versus $7.23. OpenAI’s launch copy claimed a sixteenth of Fable 5’s cost for its cheap tiers; the board measured roughly a quarter at matching accuracy, deflated but directionally intact. The rows also cut against the week’s easiest hot take, because the closed incumbent now tops the buyer’s metric among the cheapest per success. Whatever repriced this week, it wasn’t open models eating the closed frontier.

The buyers are measuring too. On July 8, Databricks published an internal benchmark built from real merged PRs against its multi-million-line codebase and concluded that “token costs are often a poor indicator of overall task costs.” The cleanest cell is Sonnet 5, about 1.7x cheaper per token than Opus 4.8 and still more expensive per task, $2.09 against $1.94 at a lower completion rate, because it consumed 1.9x more tokens getting there. Cheap tokens bought expensive work.
The Harness Went on Sale, and Open Source Got Hired
The quietest launch of the week carries the biggest repricing. On July 8, LangChain and NVIDIA shipped the NemoClaw Deep Agents Blueprint, and inside it a harness profile for Nemotron 3 Ultra, a config layer that adapts the scaffolding to one model’s quirks. No fine-tuning, no serving changes. On LangChain’s own agent eval suite, the tuned profile posts a 0.86 aggregate score at $4.48 per run; the next closest performing model cost $43.48, “roughly 10x lower inference cost” in LangChain’s phrase. Vendor eval, vendor suite, three parties with product to sell. Discount all of it and the mechanism stands. Config alone moved the task number an order of magnitude, weights untouched, which means part of the closed-model premium was scaffolding all along, and scaffolding now ships as a file.
Databricks corroborated the lever from neutral ground the same day, running the same model at the same thinking effort through two harnesses, Claude Code and Codex versus the lighter Pi. Cost per task differed by “more than 2x in some cases” while “quality remained the same,” mostly because Pi sent about 3x less context per turn. Harness choice rather than a tuned profile, but the same lever, measured by a buyer. DeepSWE can’t see any of this, since it freezes the harness by design; the harness has no public referee yet, though buyers have started grading it privately; Databricks published its methodology the same week.
Who runs which model inside the loop is settling too. In production, the tiers compose, a frontier model orchestrating with open models executing underneath. NVIDIA says its own research agent won Deep Research Bench with Opus planning over Nemotron sub-agents, Coinbase reported cutting AI spend nearly in half by routing execution to open-weight defaults, and Databricks’ board landed GLM 5.2 statistically tied with Opus 4.8 on quality at $1.28 a task against $1.94. That composition is the argument of The Open Source War Is Over, published Tuesday on X. Neither side won; both got hired into the same loop, and the seam where the hiring decision gets made is the harness router. El Hallak said it on the BEP Research podcast back in March: “All agents will be a system of models.”
The nation-scale version arrived the same week. Palantir is taking NVIDIA’s open models to frontier grade for the US government, agencies keeping the weights, and Karp put the buyer’s frame plainly on CNBC: “I wanna own the GPUs. I wanna own my data. I wanna own the model.” The All-In panel’s read of the same deal was that most foreign nations will run open source on NVIDIA hardware. Sovereign work is picking the open executor tier, and it mints on NVIDIA silicon either way.
Vera Is the Part NVIDIA Built for the Loop
Now look at where the loop physically runs. The doing in an agent task, the tool calls, repo clones, test suites, and sandbox spin-ups, is CPU work; the GPU fires in bursts between those steps. When the industry priced itself in dollars per million tokens, the GPU carried the whole story. Price by completed task and a large share of the work never touches the GPU at all.
This is the mechanical constraint I laid out in The Agentic CPU in March: “Agentic AI doesn’t reduce CPU demand. It explodes it.” Vera is NVIDIA acting on that arithmetic in silicon, three and a half months later.
The engineering reads like a spec sheet for loop stitching. NVIDIA’s claims run to 88 in-house Olympus cores on a monolithic compute die, 1.8x sustained per-core performance under full socket load, and more than 3x per-core memory bandwidth at less than half the power of a traditional x86 datacenter CPU; single-threaded speed and sandbox concurrency are exactly what the stitching starves for. Perplexity, the one named third party in NVIDIA’s own post, cloned a repository and ran its test suite in sandboxes about 1.5x faster than x86, and started concurrent sandboxes up to 1.9x faster. And NVIDIA’s own press release re-denominates the business in as many words, “AI factory economics are shifting from cores per dollar to tokens per dollar.”
The dates say NVIDIA saw this coming. Nemotron, the free open executor model, shipped in December. In March a product VP sat on my microphone describing agents as systems of models, and the CPU the loop runs on arrived in July with agents on the label. A company reacting to a metric change doesn’t have the part ready the same week; NVIDIA had already built for it.
Call it the Denominator Stack: five sellers at five stages, config to silicon, all competing on the buyer’s task cost inside three days, each with its own eval and its own units. The hard cells across the board are the price sheets and the DeepSWE rows; the rest is vendors grading their own homework.

The equity read is narrower than the excitement. NVIDIA doesn’t bill Grok 4.5, GPT-5.6, or Muse Spark, and all three can be served off non-NVIDIA silicon. What the week supports is the plumbing claim: NVIDIA sells a lever at more stages of the agent xloop than anyone else, the GPU for the token bursts, Vera for the stitching, the harness profile for config, the OpenShell runtime for the sandbox. Who bills the models running through the loop is a different question, and mostly the answer is someone else. I’ve held NVDA since 2016 and I’m not adjusting anything on this week’s news; what changed is the shape of the claim I’m willing to make for it.

The forward question is who sells the loop to everyone who can’t build it. DoorDash, Coinbase, and Databricks are the sophisticated tail; they staffed their own routers and built their own benchmarks. Most enterprises never will. If the harness decides task cost and the composition decides the bill, the other 95 percent buy the loop as a product, governance included, and ServiceNow is positioned as exactly that layer, “the layer that every model depends on but no model can replace” (The Orchestration Layer). Maybe the cleanest way to play the loop is the company selling it, not the companies inside it. That’s a question I’m taking into the paid follow-up, not a call I’m making today, and the same bound applies: ServiceNow bills none of the models either.
Where This Breaks
The stages cannibalize each other. A 54% token cut, a 4x model price cut, and a 10x config gain don’t multiply into a compound saving. They compete for the same savings line in the same enterprise budget, and a buyer who captures one may stop shopping for the rest.
The denominator has no general referee, and the two most dramatic multiples of the week, the config gain and the silicon gain, are publicly unauditable. A metric this loose invites gaming, and vendors game loose metrics every time.
Value capture is the soft spot. The assumption that the unlocked loop volume lands on NVIDIA iron is the load-bearing one, and it’s contestable. Muse Spark runs on Meta’s own buildout, and TPU, Trainium, and Groq serve agent loops today. The sovereign builds cut both ways, open volume weakening closed-lab capture while minting on NVIDIA silicon, which refines the risk without answering it. If the volume pools on someone else’s silicon, the loop thesis survives but the ticker translation doesn’t.
There’s a fourth risk, and it’s mine: curation. Token-efficiency claims, price cuts, and CPU launches happen constantly, and the line through these five is my drawing. Two of the five price sheets are still written in tokens; the translation into the buyer’s unit is this piece’s work, not the sellers’ announcement.
Underneath all four sits the capability floor, because the denominator is completed tasks. Divide each DeepSWE cost by its pass rate: Sol’s $8.39 becomes $11.49 per successful task, up 37%; Terra’s $4.95 becomes $7.07, up 43%; Opus 4.8’s $13.22 becomes $22.41, up 70%; Sonnet 5’s $26.40 becomes $48.89, up 85%. Failure inflates every bill, and it inflates the low pass rates most. A cheap task that fails is the most expensive task there is, because a human gets paid to clean it up.

The first two risks change the slope of the story. Value capture can change its direction. What flips me: loop volume visibly pooling on non-NVIDIA silicon, or the harness lever failing to reproduce. One profile is a demo; three at the same quality is a production line, and LangChain has shipped one.
So What?
Five takeaways, and one chain to keep. The harness drives the agent, the CPU carries the harness’s work, and the completed task is what the buyer pays for. Every link in that chain got a launch this week, and NVIDIA sells a lever at more of those links than anyone else.
1. AI started selling finished work. Five launches in three days all competed on the buyer’s cost per completed task, a third-party board repriced the frontier in that unit within hours, and an enterprise published its own task-cost benchmark the same week. Sellers converged on the denominator, and the buyers were already measuring in it. The general referee seat is still open, and whoever fills it inherits pricing power.
2. The harness is the lever, and it just became a product. LangChain turned harness capability into a config file worth 10x on its own eval; Databricks measured 2x from harness choice on its own codebase. The moat question in agents moves from who has the best weights to who owns the loop around them.
3. The open source war ended in a hiring decision. Production composes the tiers, frontier orchestrating, open executing, and the sovereign version of the same pattern is forming at nation scale. The fight that matters now is over the router that makes the hiring decision.
4. Vera is what it looks like when the silicon vendor believes the thesis. The doing in an agent task is CPU work, NVIDIA built a CPU for it, and its own press language now reads “cores per dollar to tokens per dollar.” NVDA lands inside the plumbing claim only; I hold the position and I am not extending the claim past the plumbing. Of the four risks above, value capture is the one I’d underwrite first.
5. The forward question is who sells the loop. The sophisticated tail builds its own routers and benchmarks; the other 95 percent will buy the loop as a product, and ServiceNow is positioned as that layer. Watch three things: the next two LangChain profiles (does the config lever reproduce), a Grok or Nemotron row on DeepSWE (the neutral-ground test), and the neocloud disclosures (where the loop volume actually pools).
I don’t know which stage wins the savings claim in an actual enterprise negotiation, and I’m skeptical of anyone who claims to this week. The sellers have already picked their unit.
Price the loop, not the model.
Coming Up
Ill be at AMD Advancing AI 2026 later this month in San Francisco
Related BEP Research
The Reasoning Tax: Why GPT-5.4 Just Validated the Memory Wars
AI for the Rest of the World: Why Muse Spark Changes the Meta Thesis
The Orchestration Layer Is the New Platform War: NVIDIA’s AI Agent Strategy (GTC 2026 podcast)
Resources
NVIDIA Developer: Vera CPU boosts AI factory throughput for agentic workloads (Jul 7, 2026)
NVIDIA: Vera, the max single-threaded CPU at scale (Ian Buck, Jul 7, 2026; Perplexity figures)
NVIDIA Newsroom: Vera, the CPU for agents (”cores per dollar to tokens per dollar”)
DeepSWE leaderboard v1.1 by Datacurve (113 tasks, 91 repos; updated Jul 9, 2026, read same day)
CNBC: Altman on GPT-5.6 Sol and enterprise spend (Jul 9, 2026)
Bloomberg: Meta starts charging for AI with Muse Spark 1.1 (Jul 9, 2026; pricing source)
xAI: Grok 4.5 launch post (page styles the entity SpaceXAI; pricing per launch coverage)
Ben Pouladian: The Open Source War Is Over (X article, Jul 7, 2026)
Disclosure: Long NVDA, NOW, LITE, CRDO, TSEM, LSCC, ALAB, WOLF, SMCI, BE, and ORCL (2027 LEAPS). This is not investment advice.




Phenomenal -- really liked the call-outs about the actual true cost of some of the coding work acknowledging the system won't get it right the first time and you need a human to intervene