What 30x Faster AI Inference Changes for Builders
A new AI chip from Cerebras claims to run models 30 times faster than standard GPU servers — and that speed jump could make AI agents and real-time tools much cheaper to build.
- 30×
- faster than GPU inference
- 4,400+
- tokens per second on a 120B model
- 10×
- more output per watt than previous chip
- Q3 2026
- first shipments begin
The speed of AI just got a lot faster. On Tuesday, a company called Cerebras unveiled a new machine — the CS-4 — built specifically to run AI models. Their claim: it runs AI 30 times faster than the best GPU systems on the market today. That kind of speed gap is big enough to change what AI products can actually do.
The CS-4 stacks three enormous custom chips into one machine. Each chip has 900,000 computing cores and fast memory baked directly into the chip itself. The result: on a 120-billion-parameter model — a large but common size for serious AI tools — the CS-4 outputs more than 4,400 pieces of text per second. The fastest GPU-based service today manages around 350.
For anyone building with AI, speed and cost are the real limits. A slow model frustrates users. An expensive one kills margins. Many teams today run smaller, cheaper models not because they work better, but because running a big model fast costs too much. Hardware like the CS-4 — if the numbers hold up outside a demo — starts to change those tradeoffs.
Cerebras says first machines ship this quarter. If cloud providers pick up the CS-4, they could offer developers faster and cheaper inference. If you are building AI agents, real-time tools, or anything that processes lots of text, watch cloud inference prices over the next two quarters. That is where this hardware story turns into a business story.