One AI request may soon begin on one computer and finish on another. WHAT?!
AMD and Cerebras are separating the two phases of LLM inference: Helios processes prompts and long context, while Cerebras generates tokens. They claim up to 5x more tokens per second per watt, although the figure is based on internal modeling.
Meanwhile, OpenAI reportedly cut inference costs for one segment of ChatGPT by more than half through an undisclosed optimization. The inference race is shifting from installing more chips to extracting more useful work from them.
Attention Span explains how the race to make AI inference faster and cheaper is reshaping hardware and the software that orchestrates it.
👉 Subscribe for high-signal AI mechanics
👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv
👉 More analysis: TuringPost.com
👉 Interviews: @realturingpost
Sources and further reading
The Information on OpenAI's reported inference optimization: https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half
OpenAI on GPT-5.6 inference and agent-harness efficiency: https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/
AMD and Cerebras announcement: https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution
AMD Helios architecture: https://www.amd.com/en/blogs/2026/amd-launches-helios-the-highest-performing-rackscale-ai-infrastructure-solution.html
AMD Helios networking: https://www.amd.com/en/blogs/2026/amd-helios-resilient-scale-up-networking-for-ai.html
Cerebras WSE-3: https://www.cerebras.ai/chip
Cerebras CS-3 datasheet: https://cdn.sanity.io/files/e4qjo92p/production/0d73d528371618c0372fcb9de9b3c0da703adf9e.pdf
Cerebras inference architecture: https://www.cerebras.ai/blog/introducing-cerebras-inference-ai-at-instant-speed
Kimi K2.6 model card: https://huggingface.co/moonshotai/Kimi-K2.6
AWS and Cerebras disaggregated inference: https://www.aboutamazon.com/news/aws/aws-cerebras-ai-inference
AMD and vLLM MORI-IO disaggregation test: https://vllm.ai/blog/2026-04-07-moriio-kv-connector
PagedAttention: https://doi.org/10.1145/3600006.3613165
Splitwise: https://arxiv.org/abs/2311.18677
DistServe: https://arxiv.org/abs/2401.09670
TetriInfer: https://arxiv.org/abs/2401.11181
Mooncake: https://www.usenix.org/conference/fast25/presentation/qin
NVIDIA Dynamo: https://www.nvidia.com/dynamo/
NVIDIA Groq 3 LPX: https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/
CDC 6600: https://www.cisl.ucar.edu/ncar-supercomputing-history/cdc6600
#ArtificialIntelligence #AI #MachineLearning #LLM #Inference #OpenAI #AMD #Cerebras #AIInfrastructure #FutureOfAI