On August 19, 2026, Cerebras Systems launched the CS-4, its fourth-generation wafer-scale AI accelerator system, with an overclocked WSE-3 Turbo engine that delivers twice the compute, twice the memory bandwidth, and twice the I/O bandwidth of the prior generation, while keeping the same 46,225 mm² silicon die, the same 900,000 cores, and the same 44 GB of on-wafer SRAM that the company has shipped since the WSE-3 in 2024. The headline number is 30x faster inference than GPU systems on the same models, which is the right framing for the inference-compute war of 2026. The right read of the launch is that Cerebras has stopped competing with Nvidia on training and committed fully to inference, that the wafer-scale bet has matured from research project into a hyperscaler product, and that the inference ASIC market is now a real category with multiple credible entrants.

This guide works through the CS-4 hardware, the inference economics, the OpenAI and AWS partnerships, the competitive landscape (Etched, Groq, hyperscaler custom silicon), and what an operator should conclude about wafer-scale AI compute in 2026-2028. The numbers come from Cerebras's August 2026 launch materials, third-party benchmarks from Artificial Analysis, and the OpenAI Cerebras deal coverage from January 2026.

Memory bandwidth, 2026 AI accelerators Higher is better. Log scale shown. Cerebras SRAM is the standout. 3.35 TB/s Nvidia H100 ~22 TB/s Nvidia Rubin 43.2 PB/s Cerebras WSE-3T (single wafer) 130 PB/s Cerebras CS-4 rack (3 wafers) Source: Cerebras CS-4 launch materials, The Register, The Next Platform, August 2026
Memory bandwidth comparison, 2026 AI accelerators. The Cerebras WSE-3T delivers 43.2 PB/s per wafer, roughly 2,000x the bandwidth of an Nvidia Rubin GPU. The full CS-4 rack combines three wafers for 130 PB/s of total bandwidth, the highest of any production AI accelerator in 2026.
Cerebras's wafer-scale engine, the largest chip in commercial production, used for AI training and inference at the major US hyperscalers. (Cerebras)
Cerebras's wafer-scale engine, the largest chip in commercial production, used for AI training and inference at the major US hyperscalers. (Cerebras)

The CS-4 at a glance

FieldWSE-3 (March 2024)WSE-3 Turbo (Aug 2026)Change
Wafer area46,225 mm²46,225 mm²Same
Process nodeTSMC 5nmTSMC 5nmSame
Transistor count4 trillion4 trillionSame
Core count900,000900,000Same
SRAM capacity44 GB44 GBSame
Sparse FP16125 PFLOPS250 PFLOPS2x
Dense FP1612.5 PFLOPS25 PFLOPS2x
Memory bandwidth21.6 PB/s43.2 PB/s2x
I/O bandwidth1.2 Tbps2.4 Tbps2x
Estimated clock1.4 GHz2.8 GHz2x
TDP (wafer)15 kW33 kW est.2.2x
TDP (system)23 kW46 kW est.2x

The comparison shows the two companies are solving different problems. Cerebras is the high-end, hyperscaler-deployed decode accelerator with proven scale (the OpenAI deal). Etched is the focused, transformer-only inference ASIC with a smaller and earlier customer book (Jane Street as the only named customer). The two companies are not really competing for the same buyers in 2026, though they may converge over time as Etched scales.

The limitations of the wafer-scale bet

Three real constraints on the CS-4 that the launch materials do not emphasize. First, the SRAM capacity has not increased meaningfully since the WSE-2 launched five years ago. 132 GB of SRAM per rack is genuinely a lot, but it is roughly the same memory capacity as a single high-end Nvidia GPU from 2024 (the H100 NVL had 188 GB of HBM, the B200 has 192 GB of HBM3E). For models above 100B parameters, the CS-4 has to shard across multiple racks, and the latency advantage narrows. Cerebras's roadmap, per the August 2026 launch, is to add SRAM capacity in the WSE-4 generation, likely via 3D SRAM stacking. The CS-4 is the last generation at 44 GB per wafer.

Second, the power draw. Each WSE-3T wafer is rated at 33 kW, and a full CS-4 system at 120-140 kW. That is roughly the same as Nvidia's NVL72 rack (240-250 kW) but the CS-4 is a single-purpose system versus Nvidia's general-purpose rack. The 46 kW per backpack requires liquid cooling and dense power delivery, which limits deployment to hyperscaler-grade data centers with the right infrastructure. Most enterprise customers cannot deploy the CS-4 in their existing data centers.

Third, the ecosystem lock-in. Cerebras has its own software stack, its own model router, its own integration with prefill partners. The customer buys into a complete inference platform, not a standalone chip. This is similar to the Etched pitch (transformer-only inference as a complete platform) and similar to the Nvidia CUDA lock-in. The downside for the customer is the cost of switching if Cerebras's economics shift or if a competitor closes the gap. The upside for Cerebras is durable customer relationships and recurring revenue.

What an operator should conclude

The CS-4 launch is the canonical example of a credible AI infrastructure category maturing in 2026. Inference-only ASICs (Cerebras, Etched, Groq, the hyperscalers' custom silicon) are now a real product category, not a research project. The cost per token at scale is the metric that matters, and the right architecture for high-throughput inference is heterogeneous, with a prefill accelerator (Nvidia GPU, AMD GPU, AWS Trainium, Google TPU) and a decode accelerator (Cerebras, Etched, Groq, custom silicon). The general-purpose GPU-only architecture is being squeezed from both ends: Nvidia's Rubin GPUs are getting better at inference (the bandwidth story), and the dedicated inference ASICs are getting better at the decode phase (the throughput story).

Three concrete takeaways. First, if you are buying inference capacity, the cost per token is dropping fast, and the right architecture is a heterogeneous stack with a specialized decode accelerator. Second, if you are investing in AI infrastructure, the inference-versus-training split is real and durable, and the right thesis is that inference will be the larger market by 2030 (it already is, by some measures, at 60-65 percent of AI compute spend). Third, if you are building AI infrastructure, the planning horizon for inference is shorter than for training. Inference economics are moving 6-12 months, training economics are moving 18-24 months. Plan accordingly.

Frequently asked questions

What is the Cerebras CS-4

The CS-4 is Cerebras's fourth-generation wafer-scale AI accelerator, launched August 19, 2026. It contains three WSE-3 Turbo processors per rack with 250 PFLOPS of sparse FP16 compute per wafer, 132 GB of SRAM per rack, and 43.2 PB/s of memory bandwidth per wafer. The CS-4 is positioned as a decode accelerator for high-throughput LLM inference.

What is the WSE-3 Turbo

The WSE-3 Turbo is the second-generation implementation of the WSE-3 wafer-scale engine, with twice the compute, twice the memory bandwidth, and twice the I/O bandwidth of the 2024 WSE-3, achieved by overclocking the same silicon from 1.4 GHz to 2.8 GHz. The wafer area, core count, and SRAM capacity are unchanged.

Who are Cerebras's customers

OpenAI (750 MW deployment deal through 2028, valued at more than $10 billion), AWS (disaggregated inference with Trainium), AMD (Helios integration), and enterprise customers through the Cerebras Cloud. The company also has a long-standing relationship with G42 in the UAE and with several pharmaceutical and genomics customers.

How much faster is CS-4 than GPU inference

Up to 30x faster on the model set Cerebras published. Specifically, 4,400 tokens per second per user on gpt-oss-120b on a single CS-4, versus ~350 tokens per second on the fastest GPU-based inference service. The advantage narrows for models above 100B parameters where sharding is required.

Can the CS-4 do training

Not as a primary use case. The WSE-3 was originally a training accelerator, but the company has shifted strategic focus to inference. Training runs on the CS-4 are technically possible but economically inferior to dedicated training infrastructure (Nvidia, AMD, hyperscaler custom silicon). The CS-4 is positioned as a decode accelerator, with prefill and training running on partner hardware.

What is disaggregated inference

A two-phase inference architecture where the compute-heavy prompt processing (prefill) runs on one accelerator type (typically GPU or XPU) and the bandwidth-constrained token generation (decode) runs on a different accelerator type (Cerebras, Etched, Groq). The two phases are linked by a high-bandwidth interconnect. The pattern reduces total cost per token by matching the right accelerator to each phase.

Sources