Tokens per Watt Explained: GTC 2026's AI Factory Revenue Metric

Updated 2026-10-01

Tokens per watt is the amount of AI output a system produces per unit of power. In practice it is quoted as tokens per second per megawatt. At GTC 2026 NVIDIA made it the main measure of its hardware, replacing FLOPS. NVIDIA’s technical blog calls it “the rate at which power is converted into revenue-generating intelligence.”

Why it became the headline number

The argument starts with a constraint, not a chip. At 1:01:13, Jensen Huang points out that a data center is limited by the power it can draw. A 1 GW site stays a 1 GW site. If power is the fixed input and tokens are the product being sold, then tokens per watt is the factory’s output rate, and output rate times price is revenue. He repeats the point at 1:20:18 with the throughput-versus-speed chart. The DSX film at 1:42:28 states it outright: an unused watt is lost revenue.

This framing also explains a line Jensen uses at 1:03:29: with the wrong architecture, hardware is too expensive even when it is free. A gigawatt facility has large fixed costs before any servers go in; Jensen put it at about $40 billion over 15 years, a figure not found in NVIDIA’s written materials. When the building and the power contract dominate cost, the system that produces the most tokens from that power wins, even at a higher sticker price.

The numbers NVIDIA attached to it

Claim Where Official source says
Grace Blackwell NVL72 vs Hopper: ~50x per watt, ~35x lower token cost 1:03:29 50x per MW and 35x lower token cost on DeepSeek-R1
Vera Rubin + Groq LPX: up to 35x per MW 1:10:05 35x per MW vs Blackwell for trillion-parameter models
1 GW factory output: 2M to 700M tokens/s in two years (350x) 1:35:26 Rubin at about 700K tokens/s per MW, which equals 700M at 1 GW

The blog adds a longer trend: throughput per megawatt rose about 1,000,000x over six architectures, from under 1 token per second per MW on Kepler in 2012 to about 700K on Rubin.

How it differs from FLOPS and TOPS

FLOPS and TOPS measure the peak math a chip can do. They say nothing about whether memory bandwidth, interconnect or software keep that math busy. Tokens per watt is measured on a whole system running a real model. That makes it closer to what an operator buys. The drawback is that it carries hidden settings: which model, which precision, how long the inputs and outputs are, and how fast each user is served.

The last setting matters most. On NVIDIA’s charts, the x-axis is interactivity, meaning tokens per second for one user. Serving each user faster lowers total throughput. A system can lead by 50x at one point on the curve and by much less at another.

How to read a vendor’s claim

The token economics page shows how NVIDIA turns this metric into revenue tiers. The disaggregated inference explainer shows one way to push the curve further right. The full keynote context is in our chapter notes.

FAQ

What does tokens per watt mean?

It is the number of AI output tokens a system produces for each unit of power it draws, usually quoted as tokens per second per megawatt. Because a data center's power supply is fixed, this number caps how much a site can produce and sell.

Is tokens per watt better than FLOPS for comparing AI chips?

For inference it is closer to what operators pay for, because it includes memory, networking and software, not just peak math. But it is only meaningful with a model, a precision and a per-user speed attached, and those choices can move the number a lot.