GTC 2026 token economics: the inference inflection and $1T claim
Watch the original video · 139 min
This segment runs from 0:43:00 to 1:07:56. It is the demand side of the keynote. Jensen explains why AI compute demand jumped in the last two years, raises his order outlook to $1 trillion, and introduces the metric he wants every CEO to manage by: tokens per watt. Watch it if you analyze NVIDIA as an investment, or if you run an inference business and want to understand the numbers NVIDIA will use to sell you hardware. The hardware that follows is covered in our Vera Rubin notes.
Key takeaways
- Three shifts in two years drove demand: generative AI with ChatGPT, reasoning with o1 and o3, then agents with Claude Code (0:46:18). Each one multiplied tokens per task.
- Stage claim: compute per task up about 10,000x and usage about 100x, which Jensen rounds to a million-fold rise in demand (0:51:12). These are his estimates, not measured figures.
- The demand outlook doubled to at least $1 trillion through 2027, up from the $500 billion figure he gave in 2025 (0:53:36).
- Tokens per watt is the new headline metric. NVIDIA cites SemiAnalysis InferenceX: GB300 NVL72 delivers up to 50x the performance per watt of H200 at 35x lower cost per token (1:00:53).
- Software alone moved a customer 7x. Jensen says Fireworks AI went from about 700 to nearly 5,000 tokens per second on the same hardware after NVIDIA software updates (1:06:21). This figure is not in NVIDIA’s official releases.
Chapter notes
0:43:00 – 0:46:18 AI-native companies and $150 billion of venture money
The section opens with a crowded slide of AI-native customers, from OpenAI and Anthropic to companies most viewers have never heard of. Jensen’s point is the shape of their spending. He cites about $150 billion of venture investment into these startups and says it is the first time investment rounds routinely reach billions of dollars, because each company needs huge amounts of compute from day one.
He compares this to past platform shifts. The PC, the internet and mobile cloud each produced a new generation of large companies, such as Google, Amazon and Meta from the last round. NVIDIA’s claim is that it has reinvented computing again, so a new set of giants will form on top of it. That is a narrative device, but it also explains why NVIDIA invests in and partners with so many startups: they are tomorrow’s largest buyers.
0:46:18 – 0:52:23 Three shifts and the inference inflection
Jensen lists three changes. Generative AI (late 2022, exploding in 2023) turned computing from retrieval into generation. Reasoning models (o1, then o3) made AI trustworthy enough for real use, but generated far more output tokens while thinking. Agents (he names Claude Code as the first) stopped answering questions and started doing work: reading files, writing code, compiling, testing and iterating. He adds that every NVIDIA software engineer now uses one or more coding agents, such as Claude Code, Codex or Cursor.
The inflection, in his framing, is that all three need inference. Thinking, reading and acting each generate tokens. He sums it up by declaring that the inference inflection has arrived. His math is 10,000x more compute per task and 100x more usage over two years. He rounds that up to a million-fold demand increase.
Treat the multipliers as directional. The important claim for investors is qualitative. If agents keep running for longer and longer tasks, compute demand grows with the amount of work AI does, not with how often new models are trained. That makes demand less lumpy than the training-driven cycle of 2023–2024.
0:52:23 – 0:59:00 The $1 trillion outlook and who is buying
Jensen recalls a $500 billion figure for high-confidence Blackwell and Rubin demand through 2026. He says this was “last year”, and that figure was given at GTC Washington, D.C. in October 2025. Now, a few months later, he says he sees at least $1 trillion through 2027. He calls even that conservative.
The wording matters. On stage it is demand and orders through 2027. NVIDIA’s live blog says at least $1 trillion in revenue from 2025 through 2027. Neither is a backlog figure from a financial filing, so treat it as management guidance.
He then breaks down the customer base. About 60% of NVIDIA’s business is the top five hyperscalers. A large share of that is their own internal AI, such as recommender systems and search moving from older methods to deep learning and LLMs. The other 40% is spread across regional, sovereign, enterprise and industrial clouds, robotics, edge and supercomputing. His argument is that this spread makes demand resilient. A skeptic would note that 60% concentration in five buyers is also a risk.
He also says Anthropic and Meta’s superintelligence lab have chosen NVIDIA. He claims NVIDIA now runs every kind of AI model across every industry and every location, from cloud to on-premises to any country.
0:59:00 – 1:05:29 Tokens per watt and the SemiAnalysis results
This is the heart of the segment. Jensen recalls the risk NVIDIA took in moving from the 8-GPU Hopper box to the 72-GPU Grace Blackwell NVL72 rack while Hopper was still selling well. He lists what made that pay off: NVFP4 (a 4-bit format he says loses no accuracy for inference and can also be used for training), Dynamo, TensorRT-LLM and new kernels, plus a multi-billion-dollar internal supercomputer, DGX Cloud, used just to optimize that software.
How to read the chart: the horizontal axis is interactivity, meaning tokens per second for each user. Jensen links it to intelligence, because faster generation lets you run bigger models and think longer in the same response time. The vertical axis on the left is tokens per watt, which he calls the factory’s output. Because a 1 GW site never becomes 2 GW, output per watt is revenue.
Jensen says that last year he claimed 35x better performance per watt for Grace Blackwell over Hopper, when Moore’s law would suggest about 1.5x. He jokes that SemiAnalysis’s Dylan Patel accused him of sandbagging, since the measured gap was 50x. NVIDIA’s own blog on the InferenceX data says up to 50x higher throughput per megawatt and 35x lower cost per token against Hopper, so the slide’s figures match an official NVIDIA source.
Two caveats. First, the benchmark is one model (DeepSeek R1) at one input/output shape, and the 50x gap is read at a specific interactivity level where H200 is near the end of its curve. Second, the “Competition” line on the slide is not named. Check the public InferenceX dashboard if you need the exact comparison.
Jensen’s cost argument follows. A 1 GW facility costs about $40 billion amortized over 15 years before any computers are installed. So a cheaper but slower chip can still give a higher cost per token. In his words, with the wrong architecture even free hardware is not cheap enough. NVIDIA shows a “Token King” belt here, a reference to the “Inference King” label on the slide.
1:05:29 – 1:07:56 Fireworks and the token factory
Jensen describes NVIDIA as vertically integrated but horizontally open: it builds the full stack, then hands it to inference providers. His example is Fireworks AI. He says its CEO Lin Qiao was in the audience, and that Fireworks grew token output 100x last year. After NVIDIA updated its algorithms and software on the same systems, Fireworks went from about 700 to nearly 5,000 tokens per second, roughly 7x.
The lesson he draws is the one to remember from this whole segment. Every cloud, AI company and enterprise will soon judge itself by the performance of its token factory, because tokens are the new commodity and power is the limit. The rest of the keynote is built on that frame. See tokens per watt explained for a plain-language version.
What changed since GTC 2025
At GTC 2025, Jensen introduced the AI factory and token framing. He called the moment a “$1 trillion computing inflection point” for data-center build-out. NVIDIA’s March 2025 releases claimed Blackwell delivered 40x Hopper’s AI factory performance with Dynamo. Blackwell Ultra (GB300 NVL72) was announced with 1.5x more AI performance than GB200 NVL72, a 50x revenue opportunity versus Hopper, and availability from H2 2025.
| Topic | GTC 2025 | GTC 2026 |
|---|---|---|
| Big number | “$1 trillion computing inflection” in data-center spending | At least $1 trillion of Blackwell + Rubin demand through 2027 |
| Blackwell vs. Hopper | 40x AI factory performance (official claim) | Jensen recalls “35x per watt”; SemiAnalysis data shows up to 50x per watt |
| Blackwell Ultra | Announced for H2 2025 | Shipping; it is the GB300 NVL72 in the SemiAnalysis chart |
| Framing | AI factories, tokens, reasoning | Inference inflection, agents, tokens per watt as the CEO metric |
Two shifts are worth noting. The 2025 “$1 trillion” referred to industry-wide data-center spending. The 2026 “$1 trillion” is NVIDIA’s own Blackwell and Rubin demand, a much stronger claim. And the 2025 performance claims were NVIDIA’s own projections, while in 2026 NVIDIA leans on a third-party benchmark. Jensen’s on-stage recollection of “35x” last year does not match the 40x figure in the 2025 live blog. The live blog also tied that 40x to Dynamo, not to performance per watt alone.
Sources: GTC 2025 live updates, Blackwell Ultra release, NVIDIA blog on InferenceX data.
Skip list
- 0:43:00 – 0:43:50 Scrolling customer-logo slide. Jensen admits it is too small to read.
- 0:48:41 – 0:49:12 Anecdote about NVIDIA engineers’ coding tools; the point is summarized above.
- 0:53:00 – 0:53:26 Pause for audience reaction to the $500 billion recap.
- 1:04:58 – 1:05:29 The “Token King” belt gag.
Glossary
- Inference — running a trained model to produce output tokens, as opposed to training it.
- Tokens per watt — tokens produced per unit of power; the output measure of a power-capped AI factory.
- Interactivity (TPS/user) — generation speed seen by one user, in tokens per second.
- NVFP4 — NVIDIA’s 4-bit floating-point format, used to raise throughput and efficiency on Blackwell and later.
- SemiAnalysis InferenceX — a continuously running third-party benchmark of inference performance across accelerators and frameworks.
- Hyperscaler — one of the largest cloud providers, such as AWS, Microsoft Azure or Google Cloud.
FAQ
What is the "inference inflection" Jensen Huang talks about?
It is his claim that AI has moved from mostly training models to mostly running them. Reasoning models and coding agents generate far more tokens per task, so compute demand now grows with usage, not just with model training. He argues that every step of an agent's work, from reading to thinking to acting, is an inference call.
Did NVIDIA say it has $1 trillion in orders?
On stage Jensen said he sees at least $1 trillion of Blackwell and Rubin demand through 2027 and called that conservative. NVIDIA's own live blog phrases it as at least $1 trillion in revenue from 2025 through 2027. It is a forecast of visible demand, not booked revenue.
Why does tokens per watt matter more than raw performance?
AI data centers are capped by the power they can get, so a 1 GW site stays 1 GW. The number of tokens it can produce per watt sets how much it can sell. Jensen also notes a 1 GW site costs about $40 billion before any computers go in, so the hardware choice decides the return on that fixed cost.
What did the SemiAnalysis benchmark show?
NVIDIA's slide, citing SemiAnalysis InferenceX on DeepSeek R1, shows GB300 NVL72 with up to 50x higher performance per watt and 35x lower cost per token than H200. NVIDIA's blog on the same data says up to 50x higher throughput per megawatt and 35x lower cost per token than Hopper.