Cerebras reports 5x inference throughput gain from disaggregation – Unite.AI
Cerebras Systems said on October 1, 2026 that it increased inference speed by 5x in early results using a technique called disaggregation, with the same number of Cerebras systems and with no loss in token generation speed. The revelation came in Disaggregated Inference From the Ground Up, a company blog post by Isaac Tai and Zhenwei Gao that opens a planned series on the topic.
The post frames the series for readers who have heard the term disaggregation, or the statement that precompilation is computation-bound and decoding is memory-bound, and have wondered what they actually mean. He builds the explanation from the ground up, starting with how accelerators balance arithmetic with data movement.
Precompilation and decoding place different demands on the hardware
The post defines arithmetic intensity as the number of floating-point operations divided by the number of bytes transferred between an accelerator’s memory and compute units. In one of his examples, adding two arrays performs 1 FLOP for every 6 bytes moved, an arithmetic intensity of 0.167 FLOPs per byte, and that ratio remains constant as the arrays grow. Matrix multiplication behaves differently: each output value is constructed from an entire row of one input and an entire column of the other, so the loaded values contribute to multiple outputs, and the arithmetic intensity grows with the size of the input.
Inference, the post explains, is a chain of such matrix multiplications between a model’s fixed weights and its input tokens, and works in two stages with different intensity profiles. During precompilation, the entire prompt is processed in parallel as one large array operation, and real-world prompts can contain thousands or even hundreds of thousands of tokens. During decoding, tokens are generated one at a time, and the model relies on the KV cache, which stores keys and computed values for previous tokens so that they are reused rather than recomputed.
Both phases still need to move all model weights, potentially hundreds of gigabytes or terabytes, into compute units for each generated token. The post shows that memory movement remains nearly constant while arithmetic intensity decreases after prefilling, and cites this gap as a reason why adding raw computing power doesn’t necessarily cause tokens to arrive faster during decoding.
Disaggregation breaks the inference into separate pools
In production, an inference server typically handles many requests at once, and when precompilation and decoding are performed on the same hardware, compute-intensive precompilation can block active decoding requests. Schedulers must then choose between receiving new requests quickly on their first token, keeping active responses streaming smoothly, and maximizing total throughput. Batching allows concurrent requests to share the work of reading model weights, but larger batches may require more time for each decoding pass, so total throughput may increase as each user receives tokens more slowly.
The post describes disaggregation as a systems design pattern that runs the two phases in separate hardware pools. Once phases are separated, operators can allocate hardware, set batching policies, and prioritize latency or throughput for each phase independently: a system with strict time-to-first-token objectives can reserve more capacity for precompilation, while one built around smooth streaming can provide decoding with a larger or more tightly scheduled pool. The pools can also be individually sized.
The separation introduces a new requirement. After precompilation creates the KV cache, the request-specific state must be transferred to the decoding pool, where it is loaded into memory before generation can continue, while the model weights are already loaded into both pools. The post notes that the move adds networking and coordination overhead, that both pools can sit idle if capacities don’t match demand, and that the added latency depends on whether the cache moves across co-located machines or between regions. He argues that disaggregation is more compelling at scale, where the gains from independently sizing and scheduling pools can outweigh switching and operational costs, and that it changes the control interface of the utility system rather than simply leveling the stream.
Heterogeneous hardware, first results and partnerships
Cerebras said it is leading the development of heterogeneous disaggregation, combining multiple chip types into a single inference system and assigning different hardware to memory- or compute-related segments. The post contrasts the company’s wafer-scale design, which deploys SRAM along with compute across the entire wafer, with GPUs, which stage model data from high-bandwidth memory through on-chip memories and smaller caches.
A chart published in the Peak Memory Bandwidth post, dated September 10, 2026, lists Cerebras WSE-3 on-chip SRAM at 21,000 TB/s per wafer, along with an unnamed on-chip SRAM accelerator at 150 TB/s and HBM4 GPU at 23.3 and 22 TB/s. The graph cautions that SRAM data sums local memory bandwidth on a processor while HBM data measures traffic from off-chip memory, so the data describes different levels of memory rather than measured token speeds.
The post also reproduces artificial analysis data from September 10, 2026 for GPT-oss-120B performing high reasoning with 10,000 input tokens. It lists Cerebras at 1,669 output tokens per second, SambaNova at 708, Groq at 475, Microsoft Azure at 319, Nebius at 294, and Baseten at 293.
Cerebras said that in a traditional aggregated system, increasing capacity meant deploying more hardware, and that by leveraging partner accelerators to handle timely processing, it increased capacity 5x in early testing with the same WSE impact. The company said it has announced partnerships with several hardware partners to bring more ultrafast tokens to market, and a diagram in the post shows AWS Trainium and AMD Helios Instinct GPU systems among the precompilation hardware options powering a Cerebras decoding pool.
The post identifies agent-based applications as a compelling solution for heterogeneous disaggregation, as they often involve long, multi-round workflows where context grows through model calls and delays at each compound step. Cerebras said future installments will cover the hardware and software stacks involved and the economic tradeoffs of implementing disaggregated inference at scale.



Post Comment