AMD and Cerebras are splitting a single AI inference request across two processor architectures. Announced July 23, the planned service will use AMD Helios systems to process prompts and Cerebras’ Wafer-Scale Engine to generate output tokens, with availability expected through Cerebras Cloud in the second half of 2026.
The design targets coding assistants, real-time copilots, live agents and other applications where response delays can compound across repeated model calls. AMD and Cerebras have not disclosed pricing, supported models or end-to-end production results, so the architecture remains unproven outside company modeling.
How the two-stage workflow operates
Large language model inference has two main stages. During prefill, the system processes the prompt and creates the key-value, or KV, cache used to produce a response. During decode, the model reads that stored state and generates output one token at a time.
AMD Helios will handle prefill and large context windows, while Cerebras’ Wafer-Scale Engine will perform the memory-bandwidth-intensive decode stage. The companies describe the design as a single integrated workflow in their partnership announcement, although the service is not yet generally available.
Prefill benefits from parallel compute capacity. Decode repeatedly accesses model data while generating tokens, making memory bandwidth and latency especially important. Cerebras integrates SRAM on its wafer-scale processor to keep more data close to the compute cores.
Cerebras is pursuing a similar disaggregated inference service with AWS, using Trainium for prefill, CS-3 systems for decode and Amazon’s Elastic Fabric Adapter to move the KV cache between them. AMD and Cerebras have not identified the interconnect for their service, and independent reporting notes that public end-to-end benchmarks are still unavailable.
Performance claims still need independent testing
AMD and Cerebras project that the combined configuration could deliver up to five times more tokens per second per watt than a Cerebras-only setup. The estimate comes from internal modeling with the Kimi K2.6 1-trillion-parameter model, not an independent comparison with Nvidia hardware or a standalone AMD deployment.
The companies also have not released pricing, total-cost-of-ownership data or service-level commitments. Enterprises should not treat the five-times efficiency projection as a purchasing benchmark. They will need to measure latency, cost per completed request and model-state transfer overhead under their own prompt lengths and traffic patterns.
Helios is a liquid-cooled reference design based on the Open Compute Project’s Open Rack Wide standard. It combines 72 Instinct MI455X GPUs with EPYC processors and Pensando networking, supporting AMD’s rack-scale AI infrastructure push. OEM and ODM partners, rather than AMD itself, will sell systems based on the blueprint.
The Cerebras deal follows Anthropic’s commitment to deploy up to 2 gigawatts of AMD Instinct capacity through a multiyear infrastructure agreement. The first gigawatt is expected online in the first half of 2027.
The deal enters a growing market for specialized AI inference clouds. A firm launch date, supported-model list, pricing, and independent measurements of latency, throughput, and energy use will provide the next proof points. Those results will show whether splitting inference across AMD and Cerebras hardware reduces production costs and latency — or merely shifts complexity into the network and orchestration layer.
Read more: A planned $5 billion TPU cloud venture would give enterprises another route to specialized AI compute outside standard public-cloud offerings.


