Hot Chips 2026: Rubin, BlueField-4, TPU v8 and OpenAI Silicon at Stanford
Quick summary
Hot Chips 2026 lands August 23-25 at Stanford Memorial Auditorium. NVIDIA Rubin, BlueField-4, Google eighth-gen TPU, OpenAI chip talk, and Meta MTIA headline the agentic AI hardware week before Nvidia earnings.
If your traffic dropped
Check which pages lost clicks in Google Search Console, then run Core Web Vitals on those URLs.
Read next
- OpenAI Jalapeño Chip Cuts Inference Costs 50% vs NvidiaOpenAI unveiled Jalapeño on June 24, 2026 — its first custom chip, built with Broadcom and TSMC — promising 50% cheaper LLM inference than Nvidia GPUs.
- Nvidia Earnings Aug 26 2026: Q2 FY27 Developer Preview GuideNvidia reports Q2 fiscal 2027 on August 26, 2026 after the close. Q1 hit $81.6B revenue, $91B guided. What developers should watch on Blackwell, Rubin, and China.
Hot Chips 2026 runs Sunday, August 23 through Tuesday, August 25, 2026 at Memorial Auditorium, Stanford University, Palo Alto, California. The official advance program is live — and this year's agenda is dominated by agentic AI infrastructure: GPUs, DPUs, custom ASICs, and networking built for multi-step inference, not single-shot chat completions.
If you build ML platforms, FinOps models, or on-prem GPU clusters, Hot Chips is the week to align hardware roadmaps with software architecture — four days before Nvidia Q2 FY27 earnings on August 26.
Hot Chips 2026 Schedule at a Glance
| Day | Date | Headline sessions |
|---|---|---|
| Tutorials | Sun Aug 23 | HBM evolution (Micron, Samsung, SK Hynix), RISC-V + NVIDIA GPU interoperability |
| Day 1 | Mon Aug 24 | NVIDIA Rubin GPU, AMD MI400, Intel Crescent Island, NVIDIA Vera CPU |
| Day 2 | Tue Aug 25 AM | BlueField-4, Spectrum-X, Thor Ultra NIC (Broadcom) |
| Day 2 | Tue Aug 25 PM | Meta MTIA, Microsoft MAIA 200, Cerebras, Google TPU v8, OpenAI chips |
Monday Aug 24: GPU Session — Rubin and the Agentic AI Stack
The GPU session (4:45–6:45 p.m. PDT) is the core developer block:
| Talk | Presenter | Why it matters |
|---|---|---|
| NVIDIA Rubin GPU: Driving the Era of Agentic AI | Manas Mandal, Rajballav Dash, Ravi Manyam (NVIDIA) | Next-gen accelerator after Blackwell; ties to Vera CPU co-design |
| AMD Instinct MI400 Series GPU Architecture | Alan Smith, Maiyuran Subramaniam (AMD) | Primary Nvidia data-center competitor roadmap |
| System Architecture of the AMD MI400 Series GPU | Steve Scott et al. (AMD) | Rack-scale design for inference/training |
| Crescent Island: GPU Designed for Agentic AI Inference | Sumit Mohan, Hong Jiang (Intel) | Intel's inference-focused GPU play |
Nvidia's Q1 FY27 release already named the Vera Rubin platform — Vera CPU, Rubin GPU, and BlueField-4 STX for context memory. Hot Chips is where architecture details go public. For background on HBM and memory bottlenecks, see our Rubin HBM4 analysis.
Earlier Monday, the CPU session includes NVIDIA Vera CPU (Jonathan Evans, Polychronis Xekalakis) — relevant because agentic workloads split compute between GPU inference and CPU tool orchestration.
Tuesday Aug 25: Networking — BlueField-4 and Spectrum-X
The Networking & Interconnect session (11:00 a.m.–1:00 p.m.) addresses a problem most developers underestimate: agentic AI multiplies network traffic.
| Talk | Presenter |
|---|---|
| Thor Ultra: Ethernet NIC for AI & HPC | Hemal Shah (Broadcom) |
| NVIDIA BlueField-4 Processor Powers the AI Factory Operating System | Idan Burstein (NVIDIA) |
| NVIDIA Spectrum-X Multiplane Network Architecture | Gilad Shainer (NVIDIA) |
| Massively Parallel Optical I/O | Mike Wiemer (Mojo Vision) |
Nvidia's technical blog on BlueField-4 (verified specs for this guide):
- BlueField-4 DPU: up to 800 Gb/s Ethernet or InfiniBand, 64-core Grace CPU, PCIe Gen6, LPDDR5X memory, DOCA programmability
- vs BlueField-3: 2x networking bandwidth, up to 6x compute, 4x memory capacity, 3x+ memory bandwidth
- BlueField-4 STX: Vera CPU + ConnectX-9 SuperNIC, up to 1.6 Tb/s Spectrum-X Ethernet, NVMe + DOCA Memos for KV cache management
- CMX context memory platform: shareable flash tier for KV cache between GPU memory and shared storage
Developer implication: When one user prompt triggers dozens of model and tool calls, networking and KV cache placement become inference bottlenecks — not FLOPs. Platform teams should plan DPU offload alongside GPU procurement.
Tuesday Aug 25: AI Sessions — TPU v8, OpenAI, Meta, Microsoft
AI Session 1 (2:15–4:15 p.m.):
| Talk | Company |
|---|---|
| MTIA: Meta's custom AI Silicon Program | Meta (Srinagesh Loke, Cindy Chen, Jatinder Singh) |
| LPX: Heterogeneous GPU-LPU | NVIDIA |
| Cerebras rack-scale WSE | Cerebras |
| MAIA 200 data-center AI system | Microsoft |
AI Session 2 (4:45–6:15 p.m.) — the highest-signal block for multi-cloud developers:
| Talk | Presenters |
|---|---|
| SN50 RDU dataflow | SambaNova |
| The Eighth Generation TPU Family: Two Chips Optimized for Training and Serving in the Agentic Era | Norman Jouppi, Sridhar Lakshmanamurthy (Google) |
| You Can Just Build Things … Chips | Richard Ho, Ravi Narayanaswami, Chris Leary (OpenAI) |
Google's talk title confirms two TPU chips — one biased toward training, one toward serving — explicitly framed for the agentic era. Deep dive: our TPU v8 vs Rubin comparison.
OpenAI's session follows Jalapeño (June 2026 inference ASIC). Hot Chips likely covers broader silicon strategy — see our OpenAI Hot Chips analysis.
Meta's MTIA program spans MTIA 400 through 500 with GenAI inference optimizations — complements Meta's ranking/recommendation silicon story.
Sunday Tutorials: Memory and RISC-V
Tutorial 1 (Memory technology) covers HBM evolution from Samsung, SK Hynix, Micron, and Meta/D-Matrix on generative inference — directly relevant to Rubin and TPU v8 supply chains.
Tutorial 2 (RISC-V) includes "RISC-V profile and platform for interoperability with NVIDIA GPUs" (Frans Sijstermans, NVIDIA) — signal that AI factories may standardize host CPU ISAs around accelerator compatibility.
Our Analysis: What to Prioritize If You Cannot Attend
Watch the Rubin + BlueField-4 block first. Agentic inference is a system problem — GPU, DPU, KV cache storage, and Spectrum-X fabric co-designed. Buying GPUs without networking/DPU planning repeats the 2023 "we have H100s but no NVLink" mistake.
Treat OpenAI and Google talks as FinOps inputs. Custom silicon and TPU v8 serving chips change per-token economics on GCP and OpenAI infrastructure — even if you never touch hardware.
Cross-reference Nvidia earnings (Aug 26). Hot Chips architecture + Q2 financials together tell you whether 2027 cluster refreshes should target Rubin, stay on Blackwell, or diversify to TPU/MTIA paths.
Developer Prep Checklist
- ] Bookmark [hotchips.org/advance-program for slide drops post-conference
- [ ] Map your cloud provider's 2026 instance roadmap against Rubin, TPU v8, and MAIA 200
- [ ] Evaluate whether agentic workloads need DPU/KV-cache tier (BlueField-4 STX / CMX pattern)
- [ ] Schedule architecture review the week of August 25 with platform + FinOps teams
- ] Link Hot Chips findings to [LLM API pricing models
Key Takeaways
- Hot Chips 2026: August 23–25, 2026 at Stanford — tutorials Sunday, conference Mon–Tue.
- Headline talks: NVIDIA Rubin GPU, BlueField-4, Spectrum-X, Google eighth-gen TPU family, OpenAI custom silicon, Meta MTIA, AMD MI400, Microsoft MAIA 200.
- BlueField-4: up to 800 Gb/s, 64-core Grace CPU, DOCA Memos for KV cache — infrastructure is part of inference.
- Agentic AI is the unifying theme across GPU, DPU, networking, and custom ASIC sessions.
- For developers: Attend virtually via post-conference slides or plan Q4 architecture around Rubin + networking co-design.
- What to watch: OpenAI chip talk (Tue 4:45 p.m.), Google TPU v8 (same session), Nvidia earnings Aug 26 for financial confirmation.
Related Reading
- Nvidia Q2 FY27 earnings preview
- Google TPU v8 vs Nvidia Rubin
- OpenAI Hot Chips silicon
- MCP 2026 stateless enterprise spec
Sources
FAQ
Frequently Asked Questions
When is Hot Chips 2026?
Hot Chips 2026 runs Sunday, August 23 through Tuesday, August 25, 2026 at Memorial Auditorium, Stanford University, Palo Alto, California. Tutorials are on Sunday; the main conference sessions are Monday and Tuesday. The advance program is published at hotchips.org.
What are the main AI chip talks at Hot Chips 2026?
Key sessions include NVIDIA Rubin GPU (Monday GPU session), NVIDIA BlueField-4 and Spectrum-X (Tuesday networking), Google eighth-generation TPU family for training and serving (Tuesday AI session 2), OpenAI custom silicon talk titled You Can Just Build Things … Chips (same session), Meta MTIA program, AMD MI400, Microsoft MAIA 200, and Intel Crescent Island for agentic inference.
What is NVIDIA BlueField-4?
BlueField-4 is NVIDIA data processing unit for AI factories with up to 800 Gb/s Ethernet or InfiniBand, a 64-core Grace CPU, PCIe Gen6, and DOCA software. BlueField-4 STX is a storage processor for the CMX context memory platform, managing KV cache with up to 1.6 Tb/s Spectrum-X connectivity. NVIDIA positions it as the infrastructure operating system for agentic AI workloads.
Why does Hot Chips 2026 matter for AI developers?
Hot Chips reveals architecture direction for GPUs, DPUs, custom ASICs, and networking that determines cloud instance families, inference latency, KV cache economics, and multi-cloud strategy. The 2026 program focuses on agentic AI — workloads that combine GPU inference, CPU orchestration, tool calls, and high-bandwidth networking.
How does Hot Chips 2026 relate to Nvidia earnings?
Hot Chips 2026 ends August 25; Nvidia reports Q2 FY27 earnings August 26, 2026. Architecture details from Rubin and BlueField-4 talks provide technical context for financial metrics on data-center revenue, networking growth, and Rubin platform timing discussed on the earnings call.
Free Weekly Briefing
The AI & Dev Briefing
One honest email a week — what actually matters in AI and software engineering. No noise, no sponsored content. Read by developers across 30+ countries.
No spam. Unsubscribe anytime.
More on AI Chips
All posts →OpenAI Jalapeño Chip Cuts Inference Costs 50% vs Nvidia
OpenAI unveiled Jalapeño on June 24, 2026 — its first custom chip, built with Broadcom and TSMC — promising 50% cheaper LLM inference than Nvidia GPUs.
Nvidia Earnings Aug 26 2026: Q2 FY27 Developer Preview Guide
Nvidia reports Q2 fiscal 2027 on August 26, 2026 after the close. Q1 hit $81.6B revenue, $91B guided. What developers should watch on Blackwell, Rubin, and China.
Google TPU v8 vs Nvidia Rubin: Agentic AI Infrastructure Guide
Google TPU v8 eighth generation debuts at Hot Chips Aug 25, 2026 — two chips for training and serving. How it compares to NVIDIA Rubin for agentic AI developers on GCP vs CUDA.
OpenAI Hot Chips 2026: Custom Silicon After Jalapeño Explained
OpenAI presents You Can Just Build Things … Chips at Hot Chips Aug 25, 2026. After Jalapeño inference ASIC, what custom silicon means for API costs and Nvidia dependency.
Free Tool
Will AI replace your job?
4 questions. Get a personalised developer risk score based on your stack, role, and what you actually build day to day.
Check Your AI Risk Score →Written by
Software Engineer based in Delhi, India. Writes about AI models, semiconductor supply chains, and tech geopolitics — covering the intersection of infrastructure and global events. 1016+ posts cited by ChatGPT, Perplexity, and Gemini. Read in 167 countries.
