AMD and Cerebras Split AI Inference Between Helios and Wafer-Scale Chips | eWeek

AMD and Cerebras Split AI Inference Between Helios and Wafer-Scale Chips

AMD Cerebras disaggregated inference running across enterprise data center infrastructure

AMD and Cerebras plan to divide AI inference workloads across specialized hardware to improve latency and infrastructure efficiency. Image: AMD

Verfasst von
eWEEK Staff
eWEEK Staff
Jul 26, 2026
3 minute read
eWeek Inhalte und Produktempfehlungen sind redaktionell unabhängig. Wir können Geld verdienen, wenn Sie auf Links zu unseren Partnern klicken. Mehr erfahren

AMD and Cerebras are splitting a single AI inference request across two processor architectures. Announced July 23, the planned service will use AMD Helios systems to process prompts and Cerebras’ Wafer-Scale Engine to generate output tokens, with availability expected through Cerebras Cloud in the second half of 2026.

The design targets coding assistants, real-time copilots, live agents and other applications where response delays can compound across repeated model calls. AMD and Cerebras have not disclosed pricing, supported models or end-to-end production results, so the architecture remains unproven outside company modeling.

How the two-stage workflow operates

Large language model inference has two main stages. During prefill, the system processes the prompt and creates the key-value, or KV, cache used to produce a response. During decode, the model reads that stored state and generates output one token at a time.

AMD Helios will handle prefill and large context windows, while Cerebras’ Wafer-Scale Engine will perform the memory-bandwidth-intensive decode stage. The companies describe the design as a single integrated workflow in their partnership announcement, although the service is not yet generally available.

Prefill benefits from parallel compute capacity. Decode repeatedly accesses model data while generating tokens, making memory bandwidth and latency especially important. Cerebras integrates SRAM on its wafer-scale processor to keep more data close to the compute cores.

Cerebras is pursuing a similar disaggregated inference service with AWS, using Trainium for prefill, CS-3 systems for decode and Amazon’s Elastic Fabric Adapter to move the KV cache between them. AMD and Cerebras have not identified the interconnect for their service, and independent reporting notes that public end-to-end benchmarks are still unavailable.

Advertisement

Performance claims still need independent testing

AMD and Cerebras project that the combined configuration could deliver up to five times more tokens per second per watt than a Cerebras-only setup. The estimate comes from internal modeling with the Kimi K2.6 1-trillion-parameter model, not an independent comparison with Nvidia hardware or a standalone AMD deployment.

The companies also have not released pricing, total-cost-of-ownership data or service-level commitments. Enterprises should not treat the five-times efficiency projection as a purchasing benchmark. They will need to measure latency, cost per completed request and model-state transfer overhead under their own prompt lengths and traffic patterns.

Helios is a liquid-cooled reference design based on the Open Compute Project’s Open Rack Wide standard. It combines 72 Instinct MI455X GPUs with EPYC processors and Pensando networking, supporting AMD’s rack-scale AI infrastructure push. OEM and ODM partners, rather than AMD itself, will sell systems based on the blueprint.

The Cerebras deal follows Anthropic’s commitment to deploy up to 2 gigawatts of AMD Instinct capacity through a multiyear infrastructure agreement. The first gigawatt is expected online in the first half of 2027.

The deal enters a growing market for specialized AI inference clouds. A firm launch date, supported-model list, pricing, and independent measurements of latency, throughput, and energy use will provide the next proof points. Those results will show whether splitting inference across AMD and Cerebras hardware reduces production costs and latency — or merely shifts complexity into the network and orchestration layer.

Read more: A planned $5 billion TPU cloud venture would give enterprises another route to specialized AI compute outside standard public-cloud offerings.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Eigentum von TechnologyAdvice. © 2026 TechnologyAdvice. Alle Rechte vorbehalten

Werbetreibenden-Offenlegung: Einige der auf dieser Website erscheinenden Produkte stammen von Unternehmen, von denen TechnologyAdvice eine Vergütung erhält. Diese Vergütung kann beeinflussen, wie und wo Produkte auf dieser Website erscheinen, einschließlich beispielsweise der Reihenfolge, in der sie erscheinen. TechnologyAdvice schließt nicht alle Unternehmen oder alle auf dem Marktplatz verfügbaren Produkttypen ein.