Infrastructure
750 Tokens per Second at the Frontier: How Cerebras Wafer-Scale Powers GPT-5.6 Sol Ultrafast
OpenAI and Cerebras have unveiled Ultrafast, a tier running GPT-5.6 Sol at up to 750 tokens per second. Here is how wafer-scale design sidesteps the GPU memory bottleneck — and what the numbers still need to prove.
Nemotron 3.5 Lightning and the Routing Layer
A purpose-built MoE executor that activates only 3 billion parameters per token, paired with NeMo Switchyard routing, is redrawing the cost structure of agentic workloads — though the speed claims remain vendor-reported for now.