Hot Chips 2026: five bets on serving agents cheaply
Intel, IBM, Meta and SiFive each showed a different answer to serving model calls cheaply. Two are already in production, one is a 2027 Xeon.
Intel, IBM, Meta and SiFive all presented server silicon at Hot Chips 2026 this week. Each pitched the same design brief: move a lot of model calls per watt and per rack. Nvidia went further and renamed the metric it optimizes for.
There’s a cost structure behind that. An agent doesn’t send one prompt and stop, it fires dozens of tool calls and re-reads a long context on every hop, so the bill gets dominated by memory capacity and interconnect rather than peak matrix throughput. Nvidia’s Vera Rubin NVL72 deck listed tokens per watt, time to first token and mean time between interruptions as the numbers that matter now, under a slide reading “Agentic AI changes the optimization target: Token Revenue.” Four other vendors turned up with four different ways to move those numbers.
Five chips, one design brief
The conference program spread them across three sessions on Monday and Tuesday at Stanford. Intel took two of the five slots, one for a CPU and one for a GPU, which is informative on its own: the company hasn’t decided whether serving agents is a core-count problem or a memory-capacity problem, so it brought an answer to each. IBM’s answer is a mainframe core that also speaks Arm. Meta’s is an accelerator with the network cards moved onto the package. SiFive’s is a 2U box you can rack this quarter.
| Chip | What it is | Why it’s notable | Link |
|---|---|---|---|
| Intel Xeon 7 “Diamond Rapids” | Up to 256 performance cores on Intel 18A-P | 1.28 GB of last-level cache and about 1.6 TB/s of memory bandwidth, on a 2027 calendar | ServeTheHome |
| Intel Crescent Island | Inference GPU, 32 Xe3P cores, 256 XMX engines | Up to 480 GB of LPDDR5X inside a 350-watt air-cooled PCIe slot | Intel Newsroom |
| IBM Z and LinuxONE next core | Dual-ISA core, 11 cores, 5.7 GHz base, 2nm | Executes z/Architecture and AArch64 natively in the same core | ServeTheHome |
| Meta MTIA 300 | Recommendation training accelerator | 12 on-package 800 Gbps RDMA NICs, 1.2 TB/s of I/O that never crosses PCIe | Engineering at Meta |
| SiFive BigSky SF-2U870 | 2U RISC-V server, 32 P870-D cores | First rackable RISC-V box that boots Ubuntu 26.04 unmodified | Chips and Cheese |
Intel Xeon 7, codename Diamond Rapids
Intel’s flagship next Xeon tops out at 256 performance cores and 1.28 GB of last-level cache, fed by 16 memory channels that reach 12,800 MT/s with MRDIMMs for roughly 1.6 TB/s of bandwidth, with 128 lanes of PCIe Gen6 and CXL 3.0 on top. It’s built on Intel 18A-P, adds AMX and AVX 10.2 for on-CPU matrix work, and swaps Intel’s proprietary EMIB bridges for UCIe-S over substrate copper, per ServeTheHome’s session write-up. Akhilesh Kumar and Krishnakanth Sistla presented it in the second CPU session. The catch is the calendar. ServeTheHome reads the disclosures as a second-half-2027 part, so Diamond Rapids is a planning input rather than a purchase.
Intel Crescent Island
Intel’s other talk was titled “Crescent Island: GPU Designed for Agentic AI Inference,” which spares everyone the trouble of guessing the strategy. Sumit Mohan and Hong Jiang showed a 350-watt air-cooled PCIe card carrying 32 Xe3P cores, 256 XMX engines and up to 480 GB of LPDDR5X, per Intel’s newsroom post. Skipping HBM is the whole bet: LPDDR5X gives up bandwidth to buy capacity per dollar and per watt. Intel’s own reference card carries 160 GB and leaves the 480 GB builds to ODM partners. Customer sampling started in the second half of this year, with volume in 2027.
IBM’s dual-ISA mainframe core
Christian Zoellin, an IBM distinguished engineer, showed the first core that executes z/Architecture and AArch64 natively in the same silicon. The part has 11 cores at a 5.7 GHz base on a 2nm process, with SMT available to both instruction sets and 2,792 AArch64 instructions implemented, including Arm v9.3 with SVE and SVE2. A little-endian Arm implementation now sits beside a big-endian mainframe one. Cache figures look absurd next to x86: 36 MB of private L2 per core, 432 MB of virtual L3, 3.5 GB of virtual L4. A separate inference chipset adds 16 active AI cores and 96 GB of HBM3e at roughly 4 TB/s. The point is running Arm-native AI code beside transaction code without buying a second box, and ServeTheHome has the slide detail.
Meta MTIA 300
Srinagesh Loke, Cindy Chen and Jatinder Singh gave Tuesday’s talk, “Meta’s Custom AI Silicon: From Recommendation to Dual-Mandate.” Meta published the MTIA 300 details the same week: two network chiplets, each holding six custom 800 Gbps RDMA NICs, for 12 NICs and 1.2 TB/s of I/O that never crosses a PCIe bus. Its 16 message engines run collectives autonomously through HCCL, a communication library co-designed with the hardware, hitting 940 GB/s inside a rack. Meta reports large matrix multiplies running concurrently with collectives lose under 0.5% of compute throughput, against more than 20% on general-purpose GPUs. A 150-billion-parameter recommendation model on 40 accelerators communicated 3.9x faster than the equivalent GPU cluster. It trains ranking models, not frontier language models, and nobody outside Meta can buy one.
SiFive BigSky SF-2U870
SiFive put RISC-V in a rack. The BigSky SF-2U870 is a standard 19-inch 2U server with 32 P870-D cores, 256 GB of DDR5-5600, a 3.84 TB U.2 NVMe drive, 64 lanes of PCIe Gen5 and a 10/25 GbE OCP 3.0 NIC, and it takes a double-wide 450-watt GPU. Because the cores implement the RVA23 profile, Ubuntu 26.04 LTS and Red Hat Enterprise Linux 10 boot unmodified, which George Cozma at Chips and Cheese calls the fix for one of the RISC-V ecosystem’s largest weaknesses. He also caught SiFive’s own materials disagreeing with each other on clock speed, listing both 2.0 and 2.2 GHz, and on the drive configuration. Chairman and CEO Patrick Little’s line was blunter: “RISC-V in the datacenter isn’t a distant aspiration any more, it is happening right now.” Treat it as a porting and validation box for hyperscalers and chipmakers, available in limited quantities.
Our pick: Crescent Island
Crescent Island is the one to watch, because it’s the only part here that changes what an ordinary enterprise rack can hold. Fitting 480 GB behind a single 350-watt air-cooled PCIe slot means a 70-billion-parameter model in BF16 with roughly 340 GB left over for KV cache, in a chassis that needs neither a liquid loop nor an NVLink domain. Nvidia’s Rubin racks want a 45 C liquid inlet and 800 VDC power delivery. MTIA 300 is the better piece of engineering, and moving the NICs onto the package is the smartest idea anyone showed all week, but you can’t buy one at any price. Diamond Rapids is a host CPU on a 2027 schedule.
The risk is software. LPDDR5X trades bandwidth for capacity, so Crescent Island will lose badly to HBM4 parts on prefill-heavy work, and Intel still has to convince serving stacks to target Xe3P at all. Capacity per watt looks like the right read on agent traffic. Whether Intel’s software can collect on it is a separate question.
What this means for you
Nothing here is a purchase decision this quarter except SiFive’s box, and that one is a validation platform. What changed is the question you put to a vendor. Peak FLOPS used to be it. For agent traffic the numbers that decide your bill are memory capacity per watt, time to first token under concurrency, and how much of your spend is KV cache rather than compute.
Measure that split now, before the 2027 parts land. Decode-heavy and context-heavy workloads are exactly what a 480 GB LPDDR5X card is built for, and an HBM part would be overkill. Heavy prefill flips it. Nobody agrees on which case dominates yet, which is why Nvidia still takes the overwhelming majority of data-center AI spend while AMD buys startups that hardwire one model into silicon and Anthropic stands up its own chip team. Intel says Crescent Island samples went out in the second half of this year, with volume in 2027. That’s the first date any of this stops being a slide.
Share this article
Quick reference
- HBM
- High-bandwidth memory, stacks of DRAM bonded together for huge throughput. It's the memory AI accelerators like Nvidia's GPUs depend on, and the supply now competes with phones and PCs.
- RISC-V
- An open, royalty-free instruction set architecture, the rulebook a chip's cores run. Unlike Arm or x86 it has no license fee, so anyone can design a CPU around it.
Sources
- Intel Outlines Architectures for Agentic AI at Hot Chips 2026 — Intel Newsroom
- Intel Diamond Rapids the 2027 Intel Xeon at Hot Chips 2026 — ServeTheHome
- IBM Z and LinuxONE Dual ISA Processor and AI Acceleration at Hot Chips 2026 — ServeTheHome
- MTIA 300: Meta's First Training Chip with Built-in NICs and Communication-Offloading Engines — Engineering at Meta
- SiFive's First Server Platform — Chips and Cheese
- NVIDIA Vera Rubin NVL72 Rack at Hot Chips 2026 — ServeTheHome
- Hot Chips 2026 program — Hot Chips
- SiFive Drives RISC-V Into The Data Center With BigSky Server Platform — HotHardware
- IBM's first dual-ISA core natively executes ARM and z/Architecture in the same core — Tom's Hardware