51 Episoden
- What if a model finds an optimization that surprises the engineers who have spent years working on that system? What, exactly, has it understood?
I brought that question to OpenAI’s Phil Tillet and Matt Ferrari, whose work involves making AI cheaper and more accessible. They’re increasingly doing that work with the models themselves.
Matt talks about research ideas his team used to dismiss because the engineering would be too complicated. Now they can give a model years of earlier research and ask it to explore what might work. They’re using models to help improve the smaller models behind speculative decoding, including the training process itself. So I ask whether the model is also helping choose the ideas, and how much that expands what they’re willing to try.
I also ask something familiar to anyone who uses these systems: why does the same model sometimes feel different? Phil explains that even changing the order of floating-point calculations can introduce differences in its behavior. That puts a very concrete problem behind our conversation about understanding: an optimization can make the system faster and still change something you wanted to preserve. We get into how they catch those changes, and what happens when the failure is something nobody thought to test for.
*We talk about:*
When models became useful for engineering decisions.
Why OpenAI’s models needed more control than Triton gave them.
What happens inside the system after you send a request.
How model-assisted kernel improvements helped cut Sol’s serving costs by 20%.
Letting models investigate bugs independently, and deciding when to step in.
Why successfully optimizing something can still be a waste of effort.
Whether models need an internal representation of how a computer system behaves.
How AI assistance opens up experiments that engineers previously couldn’t justify attempting.
*Chapters:*
*Follow on*: https://www.turingpost.com/
*Did you like the episode? You know the drill:*
📌 Subscribe here and here (https://www.turingpost.com/subscribe) for more conversations with the builders shaping real-world AI.
💬 Leave a comment
👍 Like it
🫶 Thank you for watching and sharing!
*Guests:*
Philippe (Phil) Tillet created Triton, a programming language that makes efficient GPU programming more accessible. He joined OpenAI as an intern in 2019, before it had an API or a product, and spent years improving training efficiency. His interests extend from compilers and kernels to Bertrand Russell and philosophy of mind.
Matthew (Matt) Ferrari works on inference efficiency at OpenAI, across request routing, load balancing, debugging and speculative decoding. His fascination with optimization began in school, when GPU programming changed his understanding of how fast an algorithm could run. Today, he brings that curiosity to the entire system serving a model.
#openai #inference #optimization - NVIDIA was once so hostile to open source that Linus Torvalds gave the company the finger. Today, it maintains open Linux modules, releases hundreds of models and datasets, and has reportedly agreed to buy Hugging Face for $12.9 billion. WHAT?!
The change makes sense once we examine what NVIDIA learned from nearly dying with NV1, spending years searching for CUDA’s market, and watching researchers discover deep learning on gaming GPUs.
In this episode, we follow that strategy from NV1 and CUDA to NVIDIA’s reported $12.9 billion acquisition of Hugging Face, and ask whether the company has found the most profitable model for open-source AI.
*Watch it.*
👉 Subscribe for high-signal AI analysis
👉 Instagram https://www.instagram.com/turingpost_tv
👉 TikTok https://www.tiktok.com/@turingpost_tv
👉 More analysis: https://www.turingpost.com/
👉 Interviews: @realturingpost
Attention Span is here to explain the technical and business choices shaping AI.
Links:
NVIDIA FY2026 Form 10-K https://www.sec.gov/Archives/edgar/data/1045810/000104581026000021/nvda-20260125.htm
Interview with Spencer Huang (Nvidia) https://www.youtube.com/watch?v=NEv9EnD7JVU&t=1231
Interview with Clem Delangue (Hugging Face) https://www.youtube.com/watch?v=DfJV722V1WY
NVIDIA Open Source https://opensource.nvidia.com/en-us
NVIDIA on Hugging Face https://huggingface.co/nvidia
Reuters on the reported $12.9B agreement https://www.reuters.com/technology/nvidia-talks-acquire-hugging-face-13-billion-deal-business-insider-reports-2026-08-27/
NVIDIA Open GPU Kernel Modules https://github.com/NVIDIA/open-gpu-kernel-modules
Turing Post’s history of computer vision and AlexNet https://www.turingpost.com/p/cvhistory6
#NVIDIA #OpenSourceAI #HuggingFace #AI #GitHub - Demis Hassabis, Yann LeCun, Fei-Fei Li – they all talk about “building a world model.” Some of them are dedicating their professional lives to it!
But do they mean the same?
So before joining the World Models workshop at Chicago Booth, I wanted to answer a basic question: what do researchers mean by a world model, and how many different ideas are sitting under this name?
World models are absolutely fascinating area of research with its GPT moment still in the nearest future.
This episode is based on the current research and provides a comprehensive overview of three broad approaches: generating future observations, predicting inside learned representations such as JEPA, and learning only what a planner needs to make decisions.
*Watch it.*
👉 Subscribe for high-signal AI analysis
👉 Instagram https://www.instagram.com/turingpost_tv
👉 TikTok https://www.tiktok.com/@turingpost_tv
👉 More analysis: https://www.turingpost.com/
👉 Interviews: @realturingpost
Attention Span is here to show you AI isn’t magic. Sometimes the best way to understand a model is to change the background to purple and see what breaks.
*Links:*
Demis Hassabis on world models
https://www.youtube.com/watch?v=sZaM6MadDZU
Yann LeCun on world models
https://www.youtube.com/watch?v=8sS9UJzb_t4
Fei-Fei Li on large world models
https://www.youtube.com/watch?v=pNYVckbCFuk
Beyond LLMs: JEPA and the Road to AGI – the main milestones so far https://www.youtube.com/watch?v=z0fh0SY3VWc
stable-worldmodel https://github.com/galilai-group/stable-worldmodel/issues/153
VideoPhy-2, a benchmark https://arxiv.org/pdf/2503.06800
Physion-Eval https://arxiv.org/html/2603.19607v1
What Is JEPA? LeCun Architecture & World Models https://www.turingpost.com/p/jepa
#WorldModels #AI #MachineLearning #YannLeCun #FeiFeiLi #DemisHassabis #JEPA #PhysicalAI #TuringPost #AttentionSpan - OpenCode began as an open-source coding agent. Now it is selling model access, negotiating directly with suppliers and preparing to reserve its own GPU capacity. That puts it on a collision course with OpenRouter, the model marketplace Stripe has agreed to acquire for a reported $8 billion.
This episode follows this new shift in the industry and what Ox Alpha showed about the value of distribution: the company controlling the workflow may influence which models win long before a developer opens the model menu.
*Watch it.*
👉 Subscribe for high-signal AI analysis
👉 Instagram https://www.instagram.com/turingpost_tv
👉 TikTok https://www.tiktok.com/@turingpost_tv
👉 Interviews: @realturingpost
Attention Span is the video side of Turing Post. The newsletter goes to 115,000+ people who work on this stuff: https://www.turingpost.com
#OpenCode #OpenRouter #AIAgents #CodingAgents #AIInfrastructure
Sources and further reading
OpenRouter is joining Stripe https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/
OpenCode https://opencode.ai/
Ox Alpha, Explained Without the Hype https://www.youtube.com/watch?v=tN8xiPoareo&t=16s
OpenCode Zen https://opencode.ai/docs/zen/
GLM-5.3-Flash, formerly Ox Alpha, usage data https://opencode.ai/data/zhipuai/glm-5.3-flash
Dax Raad on OpenCode’s direction https://x.com/thdxr/status/2093161006226612377
Dax Raad on inference economics https://x.com/thdxr/status/2093161006226612377
Dax Raad on OpenCode’s buying power https://x.com/thdxr/status/2092844520119345160
Jay V on OpenCode’s token volume https://x.com/snowmaker/status/2080667637861011924 - An anonymous model called Ox Alpha appeared on OpenRouter and OpenCode on August 20 with a million-token context window, video input, and a price of zero. Within four days it had processed tens of trillions of tokens, and the internet had spent those same four days trying to work out who built it.
In this episode:
how you fingerprint a model you know nothing about,
why the evidence points at Z.ai's unreleased multimodal GLM,
what the 113-task benchmark runs really show versus the viral 80 percent,
the three contradictory data policies governing your prompts,
and the thought I keep coming back to – that the platform a model launches on is becoming as decisive as the lab that trained it.
*Watch it.*
👉 Subscribe for high-signal AI analysis
👉 Instagram https://www.instagram.com/turingpost_tv
👉 TikTok https://www.tiktok.com/@turingpost_tv
👉 Interviews: @realturingpost
Attention Span is the video side of Turing Post. The newsletter goes to 115,000+ people who work on this stuff: https://www.turingpost.com
Sources and further reading
Ox Alpha vs GLM-5.3 on OpenRouter: https://openrouter.ai/compare/stealth/ox-alpha/z-ai/glm-5.3
Ox Alpha on OpenCode https://opencode.ai/data/unknown/ox-alpha
OpenCode Zen documentation https://dev.opencode.ai/docs/zen
OpenRouter Stealth Model Terms https://openrouter.ai/terms/stealth
The Tokenizer Is a Fingerprint by Joseph Elstner https://isimplifyme.com/whitepapers/the-tokenizer-is-a-fingerprint
DeepSWE result https://x.com/winkey_h/status/2090814178810306874/photo/1
58.4% run, MatchaOnMuffins/oxalpha https://github.com/MatchaOnMuffins/oxalpha/blob/main/README.md
64.6% run, jyeric/ox-alpha-deepswe https://github.com/jyeric/ox-alpha-deepswe/blob/main/README.md
Community fingerprinting summary: https://cellcog.ai/blog/what-is-ox-alpha/
Prediction market on the reveal: https://manifold.markets/Sketchy/who-is-behind-ox-alpha-the-mysterio
#OxAlpha #OpenRouter #OpenCode #GLM #AIcoding #stealthmodel
Weitere Technologie Podcasts
Trending Technologie Podcasts
Über Turing Post
Hi, I’m Ksenia, founder of Turing Post.On this channel, I talk to the people shaping AI and pay attention to the ideas, shifts, and details others might miss.Inference is my interview show with innovators, builders, founders, and thinkers moving AI forward.Attention Span is where I slow down on what deserves a closer look: the signals, questions, and stories hiding between the headlines.Subscribe for the unusual takes. And always stay curious!
Podcast-WebsiteHöre Turing Post, White Hat Café und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen
Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen


Turing Post
Code scannen,
App laden,
loshören.
App laden,
loshören.



























