AMD is expanding its AI infrastructure strategy through a new partnership with Cerebras Systems that supports a shift toward specialized chips designed for AI inference workloads.
Announced at AMD’s Advancing AI 2026 event, the agreement combines AMD’s Helios rack-scale systems with Cerebras’ Wafer-Scale Engine to create a disaggregated inference platform that assigns different stages of AI inference to hardware optimized for each task. Rather than relying on a single type of processor to execute the full workload, the architecture separates prompt processing from token generation, allowing each system to perform the work for which it is best suited.
Under the joint platform, AMD Helios systems handle prompt processing and large context windows, workloads that require high throughput across large volumes of data. Cerebras’ Wafer-Scale Engine is responsible for token generation, the stage that demands extremely high memory bandwidth and rapid response times. The companies connect the two systems into a unified inference workflow rather than treating them as separate deployments.
Cerebras plans to deploy AMD Helios systems throughout its own data centers, with the combined platform becoming available through Cerebras Cloud during the second half of 2026. Customers will also be able to configure AMD Helios systems with Cerebras wafer-scale processors.
“OpenAI hasn’t been able to serve GPT-5.6 Sol on Cerebras at scale. The AMD and Cerebras partnership can take OpenAI’s 10 million agent users off the waitlist for fast, premium inference,” Brendan Burke, Research Director at The Futurum Group, told Techstrong.ai
“Wafer-scale SRAM is the fastest decode engine in production and also the scarcest, so burning wafer cycles on prefill wastes the most valuable silicon in the rack. Disaggregation solves it. Helios racks can build long-context prefill, hand Cerebras wafers a populated KV cache, and let each engine run the workload it was built for.”
Furthermore, Burke added, “The vendors claim up to 5x tokens per second per watt versus wafer-only serving. If that holds in production, OpenAI serves multiples more users on the same wafer footprint, and high-speed Codex becomes a genuine experiential differentiator.”
Workload Specialization
AMD CEO Lisa Su said AI infrastructure is moving toward greater workload specialization, with different processors handling different stages of AI execution instead of relying on a single architecture.
To support this approach, the two companies claim the combined platform can deliver up to five times more tokens per second per watt than competing approaches. Power consumption has become a primary concern for hyperscalers as data centers encounter limits in available electricity and cooling capacity. Improving performance per watt has become nearly as important as increasing raw compute performance.
The announcement also reinforces AMD’s broader AI infrastructure push. The company agreed to provide up to two gigawatts of AI computing capacity for Anthropic, with the first gigawatt expected to become operational in 2027. AMD also committed to invest up to $5 billion in the AI model developer, strengthening its position among companies supplying infrastructure for foundation AI model builders.
The partnership helps Cerebras expand momentum following its agreement with OpenAI announced earlier this year, a contract valued at more than $10 billion to deliver 750 megawatts of AI compute capacity through 2028.

