40 Episoden
- One AI request may soon begin on one computer and finish on another. WHAT?!
AMD and Cerebras are separating the two phases of LLM inference: Helios processes prompts and long context, while Cerebras generates tokens. They claim up to 5x more tokens per second per watt, although the figure is based on internal modeling.
Meanwhile, OpenAI reportedly cut inference costs for one segment of ChatGPT by more than half through an undisclosed optimization. The inference race is shifting from installing more chips to extracting more useful work from them.
Attention Span explains how the race to make AI inference faster and cheaper is reshaping hardware and the software that orchestrates it.
👉 Subscribe for high-signal AI mechanics
👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv
👉 More analysis: TuringPost.com
👉 Interviews: @realturingpost
Sources and further reading
The Information on OpenAI's reported inference optimization: https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half
OpenAI on GPT-5.6 inference and agent-harness efficiency: https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/
AMD and Cerebras announcement: https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution
AMD Helios architecture: https://www.amd.com/en/blogs/2026/amd-launches-helios-the-highest-performing-rackscale-ai-infrastructure-solution.html
AMD Helios networking: https://www.amd.com/en/blogs/2026/amd-helios-resilient-scale-up-networking-for-ai.html
Cerebras WSE-3: https://www.cerebras.ai/chip
Cerebras CS-3 datasheet: https://cdn.sanity.io/files/e4qjo92p/production/0d73d528371618c0372fcb9de9b3c0da703adf9e.pdf
Cerebras inference architecture: https://www.cerebras.ai/blog/introducing-cerebras-inference-ai-at-instant-speed
Kimi K2.6 model card: https://huggingface.co/moonshotai/Kimi-K2.6
AWS and Cerebras disaggregated inference: https://www.aboutamazon.com/news/aws/aws-cerebras-ai-inference
AMD and vLLM MORI-IO disaggregation test: https://vllm.ai/blog/2026-04-07-moriio-kv-connector
PagedAttention: https://doi.org/10.1145/3600006.3613165
Splitwise: https://arxiv.org/abs/2311.18677
DistServe: https://arxiv.org/abs/2401.09670
TetriInfer: https://arxiv.org/abs/2401.11181
Mooncake: https://www.usenix.org/conference/fast25/presentation/qin
NVIDIA Dynamo: https://www.nvidia.com/dynamo/
NVIDIA Groq 3 LPX: https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/
CDC 6600: https://www.cisl.ucar.edu/ncar-supercomputing-history/cdc6600
#ArtificialIntelligence #AI #MachineLearning #LLM #Inference #OpenAI #AMD #Cerebras #AIInfrastructure #FutureOfAI - Jensen Huang joined X and immediately rallied Silicon Valley around open-weight AI. First came a letter signed by 77 companies and industry leaders. Then came a security alliance with 37 members. Anthropic was absent from both.
This was not a random disagreement. Kimi K3 had narrowed the gap with American frontier models, Anthropic had accused Moonshot AI of extracting 3.4 million Claude exchanges, Hugging Face had turned to an open-weight Chinese model during a security investigation, and Congress had proposed an AI kill switch.
In this episode, we examine why Jensen and Anthropic represent two competing approaches to AI safety, what their companies gain from those positions, and who ultimately gets to control the model layer.
Attention Span explains the technical and business decisions shaping AI.
👉 Subscribe for high-signal AI analysis
👉 Instagram https://www.instagram.com/turingpost_tv
👉 TikTok https://www.tiktok.com/@turingpost_tv
👉 More analysis: https://www.turingpost.com/
👉 Interviews: @realturingpost
*Links:*
Jensen Huang on X https://x.com/JensenHuang
NVIDIA: Open Secure AI Alliance https://blogs.nvidia.com/blog/open-secure-ai-alliance/?ncid=so-twit-957725
Moonshot AI: Kimi K3 announcement https://x.com/Kimi_Moonshot/status/2081760186235289764
Kimi K3 model page https://huggingface.co/moonshotai/Kimi-K3
Anthropic: Detecting and preventing distillation attacks https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks
Turing Post: What Is Knowledge Distillation? https://www.turingpost.com/p/kd
Nathan Lambert on distillation https://x.com/natolambert/status/2079616308203942332
Little Tech Association https://littletech.org/
Jensen Huang’s open-weights letter https://x.com/JensenHuang/status/2080643682408321103
Jensen Huang on the Open Secure AI Alliance https://x.com/JensenHuang/status/2081698060330250294
AI Kill Switch Act press release https://lieu.house.gov/media-center/press-releases/reps-lieu-and-moran-introduce-bill-require-kill-switch-ai-systems-can
Jeremy Howard on NVIDIA, Anthropic, and open weights https://x.com/jeremyphoward/status/2081280852454166680
Turing Post: Anthropic Fable 5 and Mythos 5 episode https://www.youtube.com/watch?v=o9d12zG6kNM&t=1s
Axios: Jensen Huang on China and open-source AI https://www.axios.com/2026/07/22/nvidia-jensen-huang-china-open-source-ai
The Information: Silicon Valley Unites Against Anthropic Over Chinese AI Restrictions https://www.theinformation.com/articles/silicon-valley-unites-anthropic-chinese-ai-restrictions
#AI #OpenWeights #Anthropic #NVIDIA #KimiK3 #JensenHuang - OpenAI’s models were supposed to solve a cybersecurity benchmark. Instead, they found a zero-day, reached the internet and broke into Hugging Face to obtain the answers.
Then came the contradiction: closed models blocked Hugging Face from analyzing the attack evidence, so its security team turned to the open-weight GLM-5.2!
We explain what happened, who was in control and why this incident may push OpenAI with help from Hugging Face toward AI that is both more open and safer
Attention Span explains the technical and business choices shaping AI.
👉 Subscribe for high-signal AI mechanics
👉 Into videos? Check our IG / turingpost_tv and TikTok / turingpost_tv 👉 More analysis: TuringPost.com
👉 Interviews: @realturingpost
*Links:*
Hugging Face: https://huggingface.co/blog/security-incident-july-2026
OpenAI: https://openai.com/index/hugging-face-model-evaluation-security-incident/
ExploitGym: https://arxiv.org/abs/2605.11086
Adrien Carreira: https://x.com/XciD_/status/2079191130248212591
Simon Willison: https://simonwillison.net/2026/Jul/22/openai-cyberattack/
Thomas Ptacek: https://x.com/tqbf/status/2080290141063569791
LeRobot: https://x.com/LeRobotHF/status/2080296144832274881
#AI #OpenAI #HuggingFace #OpenSourceAI #Cybersecurity - The world's largest open model sold out three days after launch. WHAT?! Moonshot paused new Kimi K3 subscriptions after demand pushed its GPUs to the limit. The same weekend: Xi Jinping personally endorsed open-source AI at WAIC, Alibaba teased a 2.4T-parameter Qwen 3.8 with an open-weight promise, and Axios reported that Washington may ban Chinese models entirely.
In this episode, we explain why serving an agentic model is so expensive, how a GPU shortage became an IPO pitch, what the July 27 weights release changes, and what a ban of an open model can and cannot actually reach.
Attention Span is here to explain the technical and business choices shaping AI.
👉 Subscribe for high-signal AI mechanics 👉
Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv
👉 More analysis: TuringPost.com
👉 Interviews: @realturingpost
*Links:*
Kimi K3: The open-weights escalation https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation
Kimi.ai capacity announcement https://x.com/Kimi_Moonshot https://x.com/Kimi_Moonshot/status/2078855608565207130
Xi Jinping's Big AI Speech, Annotated https://mattsheehan.substack.com/p/xi-jinpings-big-ai-speech-annotated
Kimi K3: The open-weights escalation https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation
Reuters, Moonshot pauses subscriptions amid IPO push SCMP, Kimi K3 developer suspends new subscriptions The Decoder, membership split into two tiers Qwen 3.8 announcement https://x.com/Alibaba_Qwen/status/2078759124914098291 MarkTechPost, Qwen 3.8 preview without benchmarks or license
Reuters, Xi's WAIC keynote Quartz, WAICO launch and membership Axios (Maria Curi), the administration's signals on Chinese models Nathan Lambert, Interconnects: Notes from inside China's AI labs https://www.interconnects.ai Previous episode: Kimi K3 and Inkling
#AI #ArtificialIntelligence #OpenSourceAI #KimiK3 #Qwen #LLM #MachineLearning #TechNews #China #AIPolicy - Two major open-model releases arrived this week from very different directions. Both geographically and conceptually.
Kimi K3 is a 2.8-trillion-parameter Chinese model that jumped from #18 to #1 in the Frontend Code Arena, ahead of Claude Fable 5. WHAT?!
Inkling is the first major model from Mira Murati’s Thinking Machines Lab, and the company states directly that it is not the strongest model available. WHAT?!
Both strategies make sense once we examine what the companies are building.
In this episode, we explain K3’s 896-expert architecture, why Moonshot recommends at least 64 accelerators, how LoRA customization works, what community quantization changed for Inkling, and which organizations gain meaningful control from open weights.
Attention Span is here to explain the technical and business choices shaping AI.
👉 Subscribe for high-signal AI mechanics
👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv
👉 More analysis: TuringPost.com
👉 Interviews: @realturingpost
Links:
About LoRA https://www.turingpost.com/p/lora
Moonshot AI: The Chinese Unicorn Revolutionizing Long-Context AI https://www.turingpost.com/p/moonshotai
Thinking Machines Lab, “Inkling: Our open-weights model” https://thinkingmachines.ai/news/introducing-inkling/
Thinking Machines Lab, Inkling Model Card https://thinkingmachines.ai/model-card/inkling/
Thinking Machines Lab, Tinker https://thinkingmachines.ai/tinker/
Moonshot AI, “Kimi K3: Open Frontier Intelligence” https://www.kimi.com/blog/kimi-k3
Arena, Kimi K3 Frontend Code Arena result https://x.com/arena/status/2077824029126504525
Semianalysis about Kimi K3 https://x.com/SemiAnalysis_/status/2077966560447074689
Arena, Code Arena methodology https://arena.ai/blog/code-arena/
Artificial Analysis, Kimi K3 https://artificialanalysis.ai/models/kimi-k3
Artificial Analysis, Inkling https://artificialanalysis.ai/articles/thinking-machines-has-released-inkling-the-new-leading-u-s-open-weights-model
Unsloth, community Inkling quantizations https://unsloth.ai/docs/models/inkling
#AI #ArtificialIntelligence #OpenSourceAI #LLM #MachineLearning #GenerativeAI #KimiK3 #MiraMurati #AIModels #TechNews
Weitere Technologie Podcasts
Trending Technologie Podcasts
Über Turing Post
Hi, I’m Ksenia, founder of Turing Post.On this channel, I talk to the people shaping AI and pay attention to the ideas, shifts, and details others might miss.Inference is my interview show with innovators, builders, founders, and thinkers moving AI forward.Attention Span is where I slow down on what deserves a closer look: the signals, questions, and stories hiding between the headlines.Subscribe for the unusual takes. And always stay curious!
Podcast-WebsiteHöre Turing Post, c't 4004 – der c't-3003-Podcast und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen
Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen


Turing Post
Code scannen,
App laden,
loshören.
App laden,
loshören.































