The Fastest AI Just Got Faster. Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to deploy hyperscale capacity. It is the architecture for frontier AI. Three WSE-3 Turbos per System Each wafer delivers up to 2x the speed of the previous generation Faster Wafer I/O Scales massive models and enables heterogeneous, disaggregated inference Nexus Rack-Scale Platform Enables rapid deployment in hyperscale datacenters Up to 30x faster than GPUs Powered by WSE-3 Turbo, CS-4 delivers up to 30x faster inference compared to GPU systems, setting a new record for the fastest inference available in production. Higher ultrafast throughput The CS-4 solution shifts the inference Pareto frontier, delivering up to 10x more throughput per watt than CS-3 while generating tokens up to 30x faster than production GPU systems. The result is a system designed to deliver both throughput and interactivity. Frontier-ready architecture By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters, preserving interactive decode performance at unprecedented scale. THE NEXUS RACK-SCALE PLATFORM A MODULAR RACK DESIGN TO ENABLE FASTER DEPLOYMENTS CS-4 is the first iteration of the new Cerebras Nexus Platform Architecture. It is built around a modular concept with three foundational elements: Compute, Power, and I/O – each with significant innovation to simplify manufacturing, deployment, maintenance, and upgrades. Modular compute backpack design Cerebras has fundamentally re-imagined the server. Each Wafer-Scale Backpack is a self-contained assembly that folds the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a compact 3D package with 50% fewer components. This design simplifies manufacturing and reduces deployment time from days to hours. High-density power delivery With power delivery just 0.5 millimeters away from the processor - roughly 100x closer than the roughly 50mm of conventional GPU boards - CS-4 nearly eliminates board-level power loss. This enables the delivery of twice as much power to the WSE-3T, enabling higher operating frequencies and faster token generation. Next-gen wafer I/O interface CS-4 introduces a new programmable I/O subsystem that doubles I/O bandwidth and reduces latency, benefitting both aggregated and disaggregated solutions. The Wafer I/O Module also enables wafers to be linked within and across racks without a switch, for wafer-to-wafer latency as low as two microseconds that is key to interactivity for models with tens of trillions of parameters. THE FASTEST WAFER YET CS-4 runs on WSE-3 Turbo—the world’s largest and fastest AI processor. Its four trillion transistors and 900,000 AI cores deliver 250 PFLOPS of compute and 43.2 petabytes per second of memory bandwidth. Twice the compute. Twice the bandwidth. Less than half the latency. A massive leap in AI speed and throughput. CS-4 by the numbers First CS-4 shipments begin this quarter. Bring the fastest AI to your data center. Get started Datasheet FAQ What is Cerebras CS-4? How is CS4 different from CS-3? How fast is CS-4? How much throughput does CS-4 provide? Why is CS-4 well suited for agentic AI? What is a Wafer-Scale Backpack? How does the Nexus Platform Architecture simplify hyperscale deployment? What models and inference architectures does CS-4 support?
Cerebras CS-4 芯片
为什么值得读Cerebras CS-4:晶圆级 AI 芯片。
English在原文网站阅读 ↗
如果可以重来,你还愿意花这段时间读它吗?
· 匿名阅读记录只用于改进推荐