Hot Chips 2026: Rubin, BlueField-4, TPU v8 and OpenAI Silicon at Stanford

Abhishek GautamAbhishek Gautam12 min read
Hot Chips 2026: Rubin, BlueField-4, TPU v8 and OpenAI Silicon at Stanford

Quick summary

Hot Chips 2026 lands August 23-25 at Stanford Memorial Auditorium. NVIDIA Rubin, BlueField-4, Google eighth-gen TPU, OpenAI chip talk, and Meta MTIA headline the agentic AI hardware week before Nvidia earnings.

If your traffic dropped

Check which pages lost clicks in Google Search Console, then run Core Web Vitals on those URLs.

Hot Chips 2026 runs Sunday, August 23 through Tuesday, August 25, 2026 at Memorial Auditorium, Stanford University, Palo Alto, California. The official advance program is live — and this year's agenda is dominated by agentic AI infrastructure: GPUs, DPUs, custom ASICs, and networking built for multi-step inference, not single-shot chat completions.

If you build ML platforms, FinOps models, or on-prem GPU clusters, Hot Chips is the week to align hardware roadmaps with software architecture — four days before Nvidia Q2 FY27 earnings on August 26.

Hot Chips 2026 Schedule at a Glance

DayDateHeadline sessions
TutorialsSun Aug 23HBM evolution (Micron, Samsung, SK Hynix), RISC-V + NVIDIA GPU interoperability
Day 1Mon Aug 24NVIDIA Rubin GPU, AMD MI400, Intel Crescent Island, NVIDIA Vera CPU
Day 2Tue Aug 25 AMBlueField-4, Spectrum-X, Thor Ultra NIC (Broadcom)
Day 2Tue Aug 25 PMMeta MTIA, Microsoft MAIA 200, Cerebras, Google TPU v8, OpenAI chips

Monday Aug 24: GPU Session — Rubin and the Agentic AI Stack

The GPU session (4:45–6:45 p.m. PDT) is the core developer block:

TalkPresenterWhy it matters
NVIDIA Rubin GPU: Driving the Era of Agentic AIManas Mandal, Rajballav Dash, Ravi Manyam (NVIDIA)Next-gen accelerator after Blackwell; ties to Vera CPU co-design
AMD Instinct MI400 Series GPU ArchitectureAlan Smith, Maiyuran Subramaniam (AMD)Primary Nvidia data-center competitor roadmap
System Architecture of the AMD MI400 Series GPUSteve Scott et al. (AMD)Rack-scale design for inference/training
Crescent Island: GPU Designed for Agentic AI InferenceSumit Mohan, Hong Jiang (Intel)Intel's inference-focused GPU play

Nvidia's Q1 FY27 release already named the Vera Rubin platform — Vera CPU, Rubin GPU, and BlueField-4 STX for context memory. Hot Chips is where architecture details go public. For background on HBM and memory bottlenecks, see our Rubin HBM4 analysis.

Earlier Monday, the CPU session includes NVIDIA Vera CPU (Jonathan Evans, Polychronis Xekalakis) — relevant because agentic workloads split compute between GPU inference and CPU tool orchestration.

Tuesday Aug 25: Networking — BlueField-4 and Spectrum-X

The Networking & Interconnect session (11:00 a.m.–1:00 p.m.) addresses a problem most developers underestimate: agentic AI multiplies network traffic.

TalkPresenter
Thor Ultra: Ethernet NIC for AI & HPCHemal Shah (Broadcom)
NVIDIA BlueField-4 Processor Powers the AI Factory Operating SystemIdan Burstein (NVIDIA)
NVIDIA Spectrum-X Multiplane Network ArchitectureGilad Shainer (NVIDIA)
Massively Parallel Optical I/OMike Wiemer (Mojo Vision)

Nvidia's technical blog on BlueField-4 (verified specs for this guide):

  • BlueField-4 DPU: up to 800 Gb/s Ethernet or InfiniBand, 64-core Grace CPU, PCIe Gen6, LPDDR5X memory, DOCA programmability
  • vs BlueField-3: 2x networking bandwidth, up to 6x compute, 4x memory capacity, 3x+ memory bandwidth
  • BlueField-4 STX: Vera CPU + ConnectX-9 SuperNIC, up to 1.6 Tb/s Spectrum-X Ethernet, NVMe + DOCA Memos for KV cache management
  • CMX context memory platform: shareable flash tier for KV cache between GPU memory and shared storage

Developer implication: When one user prompt triggers dozens of model and tool calls, networking and KV cache placement become inference bottlenecks — not FLOPs. Platform teams should plan DPU offload alongside GPU procurement.

Tuesday Aug 25: AI Sessions — TPU v8, OpenAI, Meta, Microsoft

AI Session 1 (2:15–4:15 p.m.):

TalkCompany
MTIA: Meta's custom AI Silicon ProgramMeta (Srinagesh Loke, Cindy Chen, Jatinder Singh)
LPX: Heterogeneous GPU-LPUNVIDIA
Cerebras rack-scale WSECerebras
MAIA 200 data-center AI systemMicrosoft

AI Session 2 (4:45–6:15 p.m.) — the highest-signal block for multi-cloud developers:

TalkPresenters
SN50 RDU dataflowSambaNova
The Eighth Generation TPU Family: Two Chips Optimized for Training and Serving in the Agentic EraNorman Jouppi, Sridhar Lakshmanamurthy (Google)
You Can Just Build Things … ChipsRichard Ho, Ravi Narayanaswami, Chris Leary (OpenAI)

Google's talk title confirms two TPU chips — one biased toward training, one toward serving — explicitly framed for the agentic era. Deep dive: our TPU v8 vs Rubin comparison.

OpenAI's session follows Jalapeño (June 2026 inference ASIC). Hot Chips likely covers broader silicon strategy — see our OpenAI Hot Chips analysis.

Meta's MTIA program spans MTIA 400 through 500 with GenAI inference optimizations — complements Meta's ranking/recommendation silicon story.

Sunday Tutorials: Memory and RISC-V

Tutorial 1 (Memory technology) covers HBM evolution from Samsung, SK Hynix, Micron, and Meta/D-Matrix on generative inference — directly relevant to Rubin and TPU v8 supply chains.

Tutorial 2 (RISC-V) includes "RISC-V profile and platform for interoperability with NVIDIA GPUs" (Frans Sijstermans, NVIDIA) — signal that AI factories may standardize host CPU ISAs around accelerator compatibility.

Our Analysis: What to Prioritize If You Cannot Attend

Watch the Rubin + BlueField-4 block first. Agentic inference is a system problem — GPU, DPU, KV cache storage, and Spectrum-X fabric co-designed. Buying GPUs without networking/DPU planning repeats the 2023 "we have H100s but no NVLink" mistake.

Treat OpenAI and Google talks as FinOps inputs. Custom silicon and TPU v8 serving chips change per-token economics on GCP and OpenAI infrastructure — even if you never touch hardware.

Cross-reference Nvidia earnings (Aug 26). Hot Chips architecture + Q2 financials together tell you whether 2027 cluster refreshes should target Rubin, stay on Blackwell, or diversify to TPU/MTIA paths.

Developer Prep Checklist

Key Takeaways

  • Hot Chips 2026: August 23–25, 2026 at Stanford — tutorials Sunday, conference Mon–Tue.
  • Headline talks: NVIDIA Rubin GPU, BlueField-4, Spectrum-X, Google eighth-gen TPU family, OpenAI custom silicon, Meta MTIA, AMD MI400, Microsoft MAIA 200.
  • BlueField-4: up to 800 Gb/s, 64-core Grace CPU, DOCA Memos for KV cache — infrastructure is part of inference.
  • Agentic AI is the unifying theme across GPU, DPU, networking, and custom ASIC sessions.
  • For developers: Attend virtually via post-conference slides or plan Q4 architecture around Rubin + networking co-design.
  • What to watch: OpenAI chip talk (Tue 4:45 p.m.), Google TPU v8 (same session), Nvidia earnings Aug 26 for financial confirmation.

Related Reading

Sources

FAQ

Frequently Asked Questions

When is Hot Chips 2026?

Hot Chips 2026 runs Sunday, August 23 through Tuesday, August 25, 2026 at Memorial Auditorium, Stanford University, Palo Alto, California. Tutorials are on Sunday; the main conference sessions are Monday and Tuesday. The advance program is published at hotchips.org.

What are the main AI chip talks at Hot Chips 2026?

Key sessions include NVIDIA Rubin GPU (Monday GPU session), NVIDIA BlueField-4 and Spectrum-X (Tuesday networking), Google eighth-generation TPU family for training and serving (Tuesday AI session 2), OpenAI custom silicon talk titled You Can Just Build Things … Chips (same session), Meta MTIA program, AMD MI400, Microsoft MAIA 200, and Intel Crescent Island for agentic inference.

What is NVIDIA BlueField-4?

BlueField-4 is NVIDIA data processing unit for AI factories with up to 800 Gb/s Ethernet or InfiniBand, a 64-core Grace CPU, PCIe Gen6, and DOCA software. BlueField-4 STX is a storage processor for the CMX context memory platform, managing KV cache with up to 1.6 Tb/s Spectrum-X connectivity. NVIDIA positions it as the infrastructure operating system for agentic AI workloads.

Why does Hot Chips 2026 matter for AI developers?

Hot Chips reveals architecture direction for GPUs, DPUs, custom ASICs, and networking that determines cloud instance families, inference latency, KV cache economics, and multi-cloud strategy. The 2026 program focuses on agentic AI — workloads that combine GPU inference, CPU orchestration, tool calls, and high-bandwidth networking.

How does Hot Chips 2026 relate to Nvidia earnings?

Hot Chips 2026 ends August 25; Nvidia reports Q2 FY27 earnings August 26, 2026. Architecture details from Rubin and BlueField-4 talks provide technical context for financial metrics on data-center revenue, networking growth, and Rubin platform timing discussed on the earnings call.

Free Weekly Briefing

The AI & Dev Briefing

One honest email a week — what actually matters in AI and software engineering. No noise, no sponsored content. Read by developers across 30+ countries.

No spam. Unsubscribe anytime.

Free Tool

Will AI replace your job?

4 questions. Get a personalised developer risk score based on your stack, role, and what you actually build day to day.

Check Your AI Risk Score →

Written by

Software Engineer based in Delhi, India. Writes about AI models, semiconductor supply chains, and tech geopolitics — covering the intersection of infrastructure and global events. 1016+ posts cited by ChatGPT, Perplexity, and Gemini. Read in 167 countries.