Turing Post
Avsnitt

Inference Race: OpenAI Cut Inference Costs in Half. AMD and Cerebras Split AI

Dela

One AI request may soon begin on one computer and finish on another. WHAT?!


AMD and Cerebras are separating the two phases of LLM inference: Helios processes prompts and long context, while Cerebras generates tokens. They claim up to 5x more tokens per second per watt, although the figure is based on internal modeling.


Meanwhile, OpenAI reportedly cut inference costs for one segment of ChatGPT by more than half through an undisclosed optimization. The inference race is shifting from installing more chips to extracting more useful work from them.


Attention Span explains how the race to make AI inference faster and cheaper is reshaping hardware and the software that orchestrates it.


👉 Subscribe for high-signal AI mechanics 

👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv 

👉 More analysis: TuringPost.com 

👉 Interviews: @realturingpost


Sources and further reading

The Information on OpenAI's reported inference optimization: https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half

OpenAI on GPT-5.6 inference and agent-harness efficiency: https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/

AMD and Cerebras announcement: https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution

AMD Helios architecture: https://www.amd.com/en/blogs/2026/amd-launches-helios-the-highest-performing-rackscale-ai-infrastructure-solution.html

AMD Helios networking: https://www.amd.com/en/blogs/2026/amd-helios-resilient-scale-up-networking-for-ai.html

Cerebras WSE-3: https://www.cerebras.ai/chip

Cerebras CS-3 datasheet: https://cdn.sanity.io/files/e4qjo92p/production/0d73d528371618c0372fcb9de9b3c0da703adf9e.pdf

Cerebras inference architecture: https://www.cerebras.ai/blog/introducing-cerebras-inference-ai-at-instant-speed

Kimi K2.6 model card: https://huggingface.co/moonshotai/Kimi-K2.6

AWS and Cerebras disaggregated inference: https://www.aboutamazon.com/news/aws/aws-cerebras-ai-inference

AMD and vLLM MORI-IO disaggregation test: https://vllm.ai/blog/2026-04-07-moriio-kv-connector

PagedAttention: https://doi.org/10.1145/3600006.3613165

Splitwise: https://arxiv.org/abs/2311.18677

DistServe: https://arxiv.org/abs/2401.09670

TetriInfer: https://arxiv.org/abs/2401.11181

Mooncake: https://www.usenix.org/conference/fast25/presentation/qin 

NVIDIA Dynamo: https://www.nvidia.com/dynamo/

NVIDIA Groq 3 LPX: https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/

CDC 6600: https://www.cisl.ucar.edu/ncar-supercomputing-history/cdc6600


#ArtificialIntelligence #AI #MachineLearning #LLM #Inference #OpenAI #AMD #Cerebras #AIInfrastructure #FutureOfAI

Podden och tillhörande omslagsbild på den här sidan tillhör Turing Post. Innehållet i podden är skapat av Turing Post och inte av, eller tillsammans med, Poddtoppen.