That is Jensen Huang in March, working the crowd at the Nebius booth at GTC, telling the company’s CRO that “Nebius will take care of you.” Nebius posted the clip with its own two-word caption underneath, “We will.”
On Tuesday, at Nebius’s first Inflection forum in San Francisco, I watched what that promise looks like as a business model. The product story was agents. The investment story was duration.
The market already understands the first neocloud species: CoreWeave, the contracted utility. Borrow heavily, lock in long contracts, satisfy lenders, scale with NVIDIA’s roadmap.
Nebius is trying to be the second species: the merchant generator. Fund with more equity, customer prepayments, backstop contracts, and converts; avoid five-year price locks; sell scarce GPU capacity short into a rising market.
Same GPU shortage. Opposite duration bet.
CoreWeave is a contracted utility. Nebius is a merchant generator.
The market color I keep replaying came out of the hallway. I posted it Tuesday evening
I heard this from a credible source I cannot name. I am not treating hallway color as a signed contract. I am treating it as a live price signal from people close enough to the shortage to know where the bids are. The important part is the structure: premium pricing, short duration, no spare capacity, and buyers still taking the capacity. That is the merchant book in one anecdote.
Alex Heath’s Sources heard the same market at the same event, with a number attached: Google reportedly paying SpaceX $920 million a month for Blackwell capacity, a high premium on an effectively 90-day rolling commitment.
The hallway quote is not the thesis by itself. It is the price signal.
Below the paywall, I’ll walk through the actual underwriting question: whether Nebius has built a merchant GPU book that can reprice into scarcity, or whether it is just another capital-hungry GPU landlord arriving before the next supply wave.
The Workload Without an Incumbent
Start with the keynote. One line has stuck with me since.
“Agentic AI will be a new workload for everyone. No decade of accumulated experience. Everyone starts from scratch.”
That is a claim about moats, dressed as a product pitch. The standing assumption in cloud investing is that AWS, Azure, and GCP carry an unassailable advantage, twenty years of operational scar tissue. The keynote’s claim is that the next workload class resets that clock to zero. If agents really do behave differently from humans, with parallel API calls and continuous multi-step execution, then nobody has a decade of experience serving them, Amazon included.
Nebius wants to be the AWS of true AI, and it said so on stage. It is not pitching models or apps; it wants to own the AI infrastructure stack from rack to token to agent, and the more of the stack it controls, the more it can optimize cost per token. The product slide names the layers SCALE, BUILD, RUN, with Agent Echo, the new agentic layer, sitting on top.

A hallway conversation put ground truth under the vertical-integration claim. Devang Sachdev, ex-NVIDIA, Twilio, and Snorkel AI, who led Nebius’s last two acquisitions, told me the data center heritage runs through Andrey Korolenko, the Head of Infrastructure who spent years building Yandex data centers, and that the racks are proprietary, designed in house. In the company’s telling, owning the rack design squeezes out more tokens per watt than anyone else. That is a company claim relayed by an employee and I weight it as one, but it points at the variable the whole stack is built to win. Tokens per watt is the floor under cost per token. Customers arrive on bare metal, GPUs by the hour, in verticals from robotics to financials.
The two deals Devang led each buy margin above hosting. Eigen AI is a model-optimization team that squeezes more tokens out of every NVIDIA GPU, the same cost-per-token axis as the proprietary racks; Tavily builds real-time search infrastructure for agents, the capability the agent layer was missing. Eigen was announced May 1 at roughly $643 million, Tavily on February 10 at $275 million and up to $400 million on milestones. Over $900 million of announced M&A puts real money behind the stack climb.

I introduced the NeoCloud framework in The NeoCloud Hypothesis: a cloud provider whose “primary business model is deploying NVIDIA GPUs at scale”, with no competing silicon program and no legacy estate in the way. Nebius fits, and I owe readers a correction: the original piece mapped CoreWeave and Oracle, and Nebius was not on my board. It should have been. Consider this piece the fix. What it adds is a layer neither CoreWeave nor Oracle built, a developer platform that climbs from bare metal to tokens to agents. CoreWeave’s edge is deployment speed on NVIDIA’s roadmap; Oracle’s is the database-to-GPU pipeline; Nebius is betting the agentic workload class is new enough that a platform built for it from scratch beats the retrofit.
Agent Echo is the layer I keep coming back to. The keynote described it as a stateful runtime rather than stateless model serving, with orchestration, durable execution across tasks that run minutes to hours, full trace observability, cost accounting, and permissioning with spend caps. Whoever operates the runtime owns the traces, and they become the raw material for optimizing customer agents the way inference endpoints get tuned today. If everyone starts from scratch, the operator accumulating the most agent traces stops starting from scratch first.
The hardware corroborates the workload shift. One customer reportedly runs on the order of 100,000 CPUs against roughly 3,000 GPUs, with early Vera CPU adoption and CPU attach rising; agent workloads are CPU-heavy in a way GPU-by-the-hour hosting does not capture. Whether any of it monetizes still hangs on the trust problem in bear three.
The Ramp
None of the positioning matters unless the buildout under it is real. Nebius went from 10 MW of running power to more than 200 MW in under two years, a 20x ramp, and the stated target is 0.81 GW running by the end of 2026, roughly 4x from here in six months. Contracted capacity sits at 3 GW today and is guided to 4 GW-plus by year end, with two thirds of that 4 GW on owned land and owned power shell rather than leases.

All stated company targets, not my estimates. The customer proof points were the checkable part: P99 inference latency of 400 to 500 milliseconds that one customer called phenomenal, training time cut from months to hours, and an unnamed customer’s 70 percent time reduction on tasks that took two to three weeks. Those numbers are the utilization side of the merchant model; capacity that performs stays sold out between repricings. The scale claim, hundreds of thousands of GPUs on a path the company describes in millions, would put Nebius alongside only the three hyperscalers among publicly available clouds.
The Merchant Model: Equity Buys the Right to Sell Short
The financing structure is built backwards from CoreWeave on purpose — the most consequential structural choice I heard all day. The capital staircase: $2 million of reservation capital 23 months ago, roughly 1 percent of total capex; billions raised last year as the data center build began; tens of billions reserved this year for the gigawatt-scale GPU buildout; hundreds of billions now targeted, structured through customer prepayments, backstop contracts that cheapen construction financing, and convertible notes. Those are order-of-magnitude figures as stated on stage, and announcements were teased as coming very soon; I am treating the staircase as a stated plan until they print.

Nebius is deliberately financed with more equity than debt so that it does not need long-term contracts to satisfy lenders, which frees it to sell short-term reservations into a market where GPU prices are rising monthly, capturing the upside that a long contract locks in for the customer.
The session described demand as unlimited, every new megawatt sold the moment it lands, supply the only binding constraint. Company statements at the company’s own event, and the unlimited assumption gets its own entry in the bear case. The checkable version is narrower — customers now request reservations out of fear of under-supply, not for pricing certainty.
CoreWeave’s lenders require duration. The weighted average contract length stretched to five years, the backlog sits in the tens of billions, and the debt-to-equity ratio printed near 894 percent at the Q4 report. That structure is rational for a debt-funded builder betting that locking today’s GPU price for five years is good business. Duration is the underwriting variable that splits the two books.

In The Token Dollar, I argued that the neocloud financing loop runs borrow dollars, buy GPUs, mint tokens, sell tokens for dollars, service the carry, roll the principal, with the GPU fleet as collateral. Nebius is running the same loop with one leg swapped, equity instead of carry. No carry means no lender forcing duration onto the book; the token side of the loop floats with the market price of compute.
The market validation the session reached for was a recent on-demand deal, capacity sold at what was described on stage as fairly insane prices with no long-term obligation. That is the stage version of the hallway price signal up top.
The margin philosophy was profitability from day one, not revenue multiples. When I pushed on the bare-metal-to-full-stack revenue multiplier in my notes, the question was raised on stage and not answered with a number. That non-answer is the honest gap in the story, and it is the number I want before underwriting the stack-depth claim.
Token Deflation Is the Business Model
The investor session says GPU prices are rising monthly. The customer panel says half of enterprise tasks are migrating to open models at a tenth the cost. Both are true, and for the same reason. Buyers have learned AI is expensive and useful at once, so demand keeps exploding while everyone turns price sensitive.
Roman Chernin, Nebius’s co-founder, posted the reconciliation himself three weeks before the event. From his tweet of May 22: “Start with frontier models → grow usage → collect data → tune cheaper models -> get margins.” He pointed at Anthropic’s revenue and token growth as the preview for open-source tokens. In this plan, per-token deflation is the margin engine. An investor reacting to the post called it Layer 4, where the greatest margin sits — tuned open models serving a myriad of users efficiently.
Amir Efrati of The Information, who moderated the customer panel, framed the moment as a rerun of the public cloud’s adolescence, when AWS customers woke up $20 million over budget and an entire cost-optimization industry was born inside a decade. The stance has shifted from “let’s try AI” to “let’s maximize AI efficiency, but with clear, documented ROI,” and every tool serving that stance routes inference spend toward open models on cheap infrastructure. The panel was not unanimous on the premise; Cognition CEO Scott Wu argued the GPUs are “expensive, but they’re not that expensive” next to the humans they stand in for, and that the real shift is buyers measuring outcomes rather than consumption.
Roughly 80 to 90 percent of enterprise AI tasks do not need the smartest model, open-source models handle 50 to 60 percent of tasks at roughly 10x cheaper and faster, and the hardest 10 to 20 percent still warrant frontier models. The tooling response is converging on AI gateways, one control point with full visibility and free model switching. Databricks runs all internal work through one tool, per VP Nikita Shamgunov, and the visibility surfaced bottlenecks that were never model calls at all, like stacked PRs waiting on CI/CD.

Nebius put a price on the routing table from its own stage. Anubhav Maheshwari, the company’s VP of ecosystem strategy, demoed a healthcare compliance agent that cost $637 per run on a closed frontier model; rerouted onto open-weight models, DeepSeek and then NVIDIA’s Nemotron, the same run cost $24 and finished in 15 minutes instead of an hour. Heath’s Sources has the writeup.
Seen from the supply side, the migration is the same story. The typical customer journey starts on frontier APIs, hits an economics or data-sensitivity wall, fine-tunes an open-source model, and shifts the workload. The open-source frontier lag is roughly six months. Shopify was the concrete example. Latency-sensitive inference stays on GCP near the control plane; everything else gets distributed across providers on price, with fine-tuned models doing merchant search. The cost reckoning is not shrinking AI budgets. It is redistributing them down the price curve, toward open-model inference at scale, the middle layer of Nebius’s stack. I think Nebius wins in exactly that world, where large enterprises understand the use cases but are picky about which models they use for what, and want full control over their data and, above all, their costs.
The orchestrator-worker pattern is the concrete shape of the migration. NVIDIA’s own internal practice, as I understand it, runs Opus as the orchestrator and Nemotron as the worker. The expensive frontier model plans and delegates; the cheap tuned open model executes the volume. My base case is that we end up with more models, and more specialized ones, rather than one god-model carrying every workload. Anthropic’s terms of service now restrict using Claude for frontier-LLM development work, per SemiAnalysis reporting, and per its model card the new Fable model is hardened against distillation and LLM-development use, pushing exactly that work toward open models and the clouds that tune and serve them.
Every voice in that story has an interest in it. The first disinterested signal is showing up in spend data. Citadel Securities’ “Tokenomics” note, built on Silicon Data’s LLM Expenditure Index, reads the index’s recent decline as substitution away from expensive frontier usage toward cheaper models, and cites reports of unexpectedly large token bills. That is exactly the world Nebius wants — and the scarce asset in it is still GPU capacity. When a sellside macro desk charts that shift, the migration has graduated from conference talk to a data print.

Now run the declining index against the merchant book. Tokens deflate per task, total volume grows faster than the deflation under the Jevons dynamic the panel described, and every tuned open model still runs on GPUs somebody has to rack and power. What Nebius sells short-term is capacity, not tokens, so token prices and reservation prices can fall and rise at the same time: the cheaper the work gets, the more of it gets done, and the megawatts stay sold out.
Nebius's CRO opened the day pitching the room on graduating from tokenmaxxing to valuemaxxing, making tokens count rather than counting them. The infrastructure underneath is indifferent to which slogan wins. Nebius Token Factory bills for the GPUs that mint tokens, so a falling price per token pulls more volume onto the same megawatts. That is the tokenmaxxing factory, and the floor under its book is the total number of tokens produced, whatever any single one ends up selling for.
This is also why I am skeptical of the bitcoin-miner AI pivots. Power and shells get you into hosting, not into the token factory, and hosting margins compress as the gigawatts land. If margins migrate up the stack, the miners become the control group for what happens when you own megawatts but not the platform. My read: many of those pivots do not survive the compression, and the failures will get blamed on AI demand when the real problem is stack position.
What This Means for NBIS, CRWV, and NVDA
The split is no longer hypothetical. The newest, fastest-ramping entrant chose the merchant structure, funded it with equity on purpose, and says it is selling every megawatt the moment it lands. That is a market signal. The people closest to GPU pricing are betting it stays supply-constrained through the window that matters. The open-source migration is part of the same bet; per-token deflation grows the workload base that keeps capacity scarce.
NBIS is the purest public expression of the merchant structure and the only neocloud selling a developer platform up to the agent layer. A merchant book marks to the GPU price faster than any contracted backlog, in both directions. The demand filling it skews toward production inference tied to live products with users, a higher-quality book than speculative training reservations. I am watching the financing announcements and the full-stack revenue multiplier before underwriting it properly.
CRWV should now be read through duration. The five-year backlog is the defensive asset in a falling price curve and the capped asset in a rising one.
NVDA wins either way, but the short-duration book is the better tell on pricing, because it marks the GPU price to market every month while a contracted book hides it for years. The scramble underneath is already visible; everyone at the compute layer is trying to learn how much Vera Rubin supply it can actually get, and how much its competitors are getting, a sorting process the hallways described as cutthroat. And the hyperscalers’ real exposure is whether the agentic workload class grows up on someone else’s platform.
The Bear Case and the Tripwire
Four risks, ranked. The first can change the direction of the thesis. The rest change the slope.
Bear one: the merchant book has no floor. Short-duration reservations capture a rising GPU price curve, and they reprice within a quarter if the curve rolls over. A Rubin-era supply wave or a hyperscaler custom-silicon ramp that eases scarcity would hit Nebius’s revenue line faster than any neocloud carrying a five-year backlog; so would plain macro softness. The hallway consensus at the event put dates on that risk: constrained until sometime next year, possible softening mid-2027, and some attendees pushed relief out to late 2027 or 2028. The merchant book has to earn its premium inside that window.
Bear two: the merchant structure is not proprietary. Anyone with access to equity capital can copy the balance sheet, and if the curve keeps rising, every well-funded GPU cloud can shorten its book and chase the same repricing. The harder half to copy is the stack above the financing, the half the Eigen and Tavily checks just bought, and the bitcoin-miner pivots are the live control group: power and shells with no platform, competing for the same hosting dollars. If Nebius’s margins over the next year look like hosting rather than Layer 4, the merchant premium was never defensible and the company is one well-financed landlord among many.
Bear three: Agent Echo is a product for an unsolved problem. The differentiated top layer monetizes only if enterprises deploy agents at scale, and that needs a security and verification layer nobody has shipped. Nebius’s own security session was candid: agentic security is unsolved industry-wide, the preferred approach is identity-based access control at the agent runtime, and the panel conceded the industry has worked the problem for three years without cracking it. If enterprise agent deployment stalls on trust, the agentic layer monetizes late and the stack-depth premium compresses toward bare metal.
Bear four: the staircase needs open capital markets. Hundreds of billions is a target, not a balance sheet. The whole staircase works while credit is loose and GPU collateral is bid. A spread blowout forces an ugly choice: sign the long contracts the model was built to avoid, or slow the buildout while competitors with contracted backlogs keep building. Mostly slope, but at the tail it bends direction, because the merchant edge dies the day a lender demands duration.
And the tripwire, so this is checkable later: the first quarter Nebius’s realized reservation price declines sequentially, or the first three-year-plus contract it signs. Either one says the merchant thesis is breaking.
I do not hold Nebius today. I expect that to change soon, and I will disclose when it does.
Underwrite the duration of the book, not the size of the buildout.
Related BEP Research
Elsewhere: Alex Heath covered the same event from the product side in Sources, including the panel debate this piece draws on and the $920M-a-month Google-SpaceX detail. If you follow the AI race, his newsletter belongs in your inbox.
Disclosure: Long NVDA, LITE, CRDO, TSEM, LSCC, ALAB. Long ORCL 2027 LEAPS. Long BE. Long WOLF speculative. No position in NBIS at publication; I expect to initiate one soon and will disclose when I do. Positions current as of publish. This is opinion and analysis, not investment advice.




