GTC 2026 Keynote Notes: All 15 Chapters With Timestamps

Updated 2026-10-01 · Video: NVIDIA, published 2026-03-16

These notes cover the whole GTC 2026 keynote, from 0:09 to 2:18:56, in 15 chapters. Each chapter says what was announced, how to read the numbers, and which deeper page to open next. They are for engineers, investors and operators who want the argument without sitting through two hours and nineteen minutes. For the short version, start at the event overview.

Key takeaways

  1. The unit of the talk is tokens per watt, not FLOPS. From 1:01:13, Jensen describes every data center as a power-limited token factory whose revenue is set by throughput per megawatt. Explainer.
  2. Demand guidance doubled. At 53:32 he moved from roughly $500 billion of Blackwell and Rubin demand through 2026 to at least $1 trillion through 2027. NVIDIA’s own blog words this as $1 trillion in revenue from 2025 through 2027.
  3. Inference is now split across two chip types. At 1:31:55, Dynamo runs decode attention on Rubin GPUs and the feed-forward layers on Groq 3 LPX. Groq LPX.
  4. Vera Rubin is in full production. The roadmap at 1:36:29 adds Rubin Ultra, LP35, Feynman, LP40 and a CPU called Rosa. Roadmap.
  5. Software is pitched as a demand engine. OpenClaw at 1:46:32, the Nemotron coalition at 2:00:41 and robotaxis at 2:06:47 all end in more tokens consumed.

Chapter notes

0:09 – 6:02 Opening: three platforms, five layers

After the opening film, Jensen sets out the structure. NVIDIA, he says, has three platforms: CUDA-X libraries, its systems, and a new one called the AI factory. He then describes the industry as a five-layer stack of land and power, chips, platforms, models and applications. The framing matters because it lets NVIDIA claim a role in every layer except land. It also explains why a talk nominally about GPUs spends time on power grids and agent software.

6:02 – 16:15 CUDA at 20, GeForce and DLSS 5

The CUDA anniversary segment is a business argument told as history. A large install base attracts developers, developers produce new algorithms, and new algorithms open new markets. Jensen’s evidence that the flywheel works is that six-year-old Ampere GPUs still see cloud rental prices rise. DLSS 5, shown at 13:33, blends rendered 3D geometry with generative AI. The general idea is structured data steering a probabilistic model, and it returns later in the talk. DLSS 5 and CUDA at 20.

16:15 – 30:00 Data processing and cloud partners

This is the least flashy segment and possibly the most practical one. NVIDIA wants SQL engines and vector search to run on GPUs, using cuDF for tables and cuVS for embeddings. Its pitch is that AI agents will query enterprise data far more often than people do. The IBM watsonx.data case (Nestlé, 5x faster, 83% lower cost) and a Google Cloud case (Snap, about 80% lower cost) are presented as savings stories. The cloud tour that follows makes one point: NVIDIA brings customers to the clouds, so the clouds keep buying.

30:00 – 43:00 Vertical integration and CUDA-X

Jensen argues that “accelerated computing” really means application acceleration. A general CPU speedup has run out, so gains come one domain library at a time. That is why NVIDIA builds the chip, the system and the libraries, then licenses the software stack to anyone. The industry tour runs from finance to telecom. He says the conference released about 70 libraries and 40 models. The CUDA-X film at 38:31 is well made, but it adds no new information.

43:00 – 54:40 AI-native companies and the inference inflection

Here is the demand thesis. Generative AI, then reasoning models, then agentic coding tools each multiplied the tokens generated per task. Jensen names Claude Code, Codex and Cursor as tools in daily use at NVIDIA. He says compute per task rose about 10,000x and usage about 100x over two years. These are illustrative figures rather than measured ones. The concrete claim is at 53:32: at least $1 trillion through 2027. Token economics.

54:40 – 1:07:56 Tokens per watt and the token factory

This is the conceptual core of the keynote. A 1 GW site stays a 1 GW site, so revenue depends on tokens produced per watt at a given response speed.

Slide titled NVIDIA Extreme Co-Design Revolutionized Token Cost, with a tokens-per-watt curve showing GB300 NVL72 about 50x above H200 and a token-cost curve showing 35x lower cost
1:01:13 — The left panel plots tokens per watt against interactivity on DeepSeek R1, with GB300 NVL72 about 50x above H200. The right panel turns the same data into about 35x lower cost per token. NVIDIA's own blog repeats both figures.

Read the left curve as a menu, not a single score. Moving right buys faster answers per user but lowers total output, so a vendor’s headline multiple depends on where along the x-axis it is measured. Jensen’s $40 billion figure for a 1 GW facility over 15 years supports his line that wrong hardware is costly even when it is free. That figure was not found in NVIDIA’s written materials. Explainer.

1:07:56 – 1:20:00 Vera Rubin platform

A short history from DGX-1 to Blackwell NVL72 leads into Vera Rubin. NVIDIA calls it seven chips and five rack types: the NVL72 compute rack, a Vera CPU rack, Groq 3 LPX, Spectrum-6 Ethernet with co-packaged optics, and BlueField-4 STX storage. Two details deserve attention. The CPU is built for single-thread speed because agents wait on tool calls. Storage was redesigned around KV cache traffic. Rubin Ultra, shown at 1:17:34, moves to vertical Kyber racks with 144 GPUs in one NVLink domain. Vera Rubin platform.

1:20:00 – 1:29:18 Token tiers and the case for Groq

Jensen reuses last year’s throughput-versus-speed chart and adds prices. He sketches a free tier, $3 and $6 per million tokens, a $45 premium tier and a $150 tier for very fast, long-context work. With power split evenly across tiers, he says Blackwell and Vera Rubin each produce about 5x the revenue of the previous generation. The catch arrives at 1:27:56: beyond roughly 400 tokens per second per user, NVL72 runs out of memory bandwidth. That limit is the opening for Groq. Token economics.

1:29:18 – 1:34:40 Groq LPX and disaggregated decode

Diagram of NVIDIA Dynamo routing prefill and decode attention to a Vera Rubin NVL72 rack holding the KV cache, exchanging activations with a Groq 3 LPX rack that runs decode FFN and emits tokens
1:31:55 — The Vera Rubin rack keeps the KV cache and runs prefill plus decode attention. Activations then hop to the Groq rack for the feed-forward step. This is the single slide that explains how two very different chips share one model.

A Groq LP30 chip holds about 500 MB of SRAM, while a Rubin GPU has 288 GB of HBM. Neither chip suits the whole job. The fix is to split each decode step: attention over the large KV cache stays on the GPUs, and the feed-forward layers run on LPUs at SRAM speed. Jensen’s practical advice was to put about 25% of a coding-heavy data center on Groq and leave the rest on Vera Rubin. He gave Q3 for LPX shipments, and NVIDIA’s release says the second half of 2026. Groq LPX · Disaggregated inference, explained.

1:34:40 – 1:40:06 Production status and roadmap

Jensen says Azure has its first Vera Rubin rack running. He also says the supply chain can produce thousands of systems a week, which works out to gigawatts of AI factories per month. His headline figure is token output for a 1 GW factory rising from 2 million to 700 million tokens per second in two years, or 350x. That is consistent with NVIDIA’s blog figure of about 700K tokens per second per MW for Rubin.

Roadmap slide titled NVIDIA Extreme Co-Design Delivering X-Factors Every Year, showing Blackwell, Rubin and Feynman generations with Oberon and Kyber racks
1:36:29 — The roadmap runs from Blackwell through Rubin to Feynman, with Oberon and Kyber rack formats side by side. The point is that copper and co-packaged optics both stay on the plan rather than one replacing the other.

Roadmap.

1:40:06 – 1:46:32 Omniverse DSX and AI factories

DSX is a digital-twin blueprint for designing and operating gigawatt sites. It covers simulation (DSX Sim), operating data (Exchange), grid-aware power (Flex) and dynamic power capping (Max-Q). Jensen estimates there is about 2x of wasted capacity to recover in a typical site. That is a loose claim, but it is consistent with the talk’s logic: when power is fixed, idle watts are lost revenue. The Vera Rubin Space-1 orbital computer at 1:45:40 is a research direction, not a product.

1:46:32 – 1:56:37 OpenClaw, NemoClaw and OpenShell

Jensen calls OpenClaw an operating system for agents. It has resources, tools, a file system, scheduling and sub-agents. He compares it to Linux, HTML and Kubernetes, and he says every company now needs an OpenClaw strategy. The enterprise problem is that an agent can read sensitive data, run code and talk to the outside world. NemoClaw is NVIDIA’s reference stack for that problem. It uses OpenShell to enforce policy engines, network guardrails and privacy routing. NemoClaw and OpenClaw.

1:56:37 – 2:05:35 Open models and the Nemotron coalition

NVIDIA lists six open model families: Nemotron, Cosmos, GR00T, Alpamayo, BioNeMo and Earth-2. It promises steady new versions, including Nemotron 4. The coalition announced at 2:00:41 includes Mistral, Perplexity, Cursor, LangChain and Thinking Machines Lab. The commercial logic is plain: a company running custom agents needs a model it can tune, and NVIDIA would rather supply that model than leave the slot to a closed lab. Jensen’s idea of a token budget per engineer belongs to the same argument. Nemotron coalition.

2:05:35 – 2:15:20 Physical AI, robotaxis and Olaf

BYD, Hyundai, Nissan and Geely join the robotaxi platform, and NVIDIA deepens its work with Uber. Jensen puts the four new partners’ combined output at about 18 million cars a year. The robotics section rests on one premise: real-world data will never cover the edge cases, so simulation through Isaac Lab, Newton and Cosmos generates the rest. Disney’s Olaf robot, which learned to walk in simulation, makes that premise easy to see. Physical AI.

2:15:20 – 2:18:56 Closing

The closing is a recap song over generated footage. It contains nothing new.

What changed since GTC 2025

Skip list

Glossary

FAQ

How long is the GTC 2026 keynote?

The official recording runs 2 hours 18 minutes 56 seconds. Roughly 15 minutes of that is films, thank-yous and the closing song, so the substance fits in about two hours.

Where in the GTC 2026 keynote does Jensen announce Vera Rubin and Groq LPX?

The Vera Rubin segment starts at 1:07:56 and the Groq LPX pairing is explained from 1:29:18 to 1:34:40. The roadmap follows at 1:36:29.

Did NVIDIA say Vera Rubin is shipping?

Jensen said Vera Rubin is in full production and that Microsoft Azure already has its first rack running. NVIDIA's press release says partner systems become available from the second half of 2026, and Groq LPX is expected around the third quarter.