Zum Inhalt springen

Semi Doped

Vikram Sekar and Austin Lyons
Semi Doped
Neueste Episode

47 Episoden

  • Semi Doped

    Pacing AI Means More Compute, Not Less

    18.09.2026 | 36 Min.
    The debate over "pacing" AI for safety is a red herring, as the practical outcome of more safety work is not a slowdown but a net increase in demand for compute hardware. Austin Lyons and Vik Sekar unpack Dario Amodei's call to slow down frontier AI development, arguing that the true governors on AI are not voluntary agreements but physical data center bottlenecks and geopolitical realities. They conclude that the push for more safety and interpretability will ultimately drive more sales of both training (GPUs) and inference (XPUs) hardware.

    Today's sponsor is G2i — expert training data, evals, and RL environments for AI labs: https://fandf.co/4wZvKHS

    Chapters:
    0:00 Introduction: Pacing AI
    1:36 Recapping Dario Amodei's Essay
    7:42 Hot Take: Anthropic vs. OpenAI
    11:34 The John Deere Playbook
    15:22 The Problem with Interpretability
    17:53 Pulling the Plug
    19:03 The Real Bottleneck: Data Centers
    22:03 The Geopolitical Blind Spot
    23:41 Counter-narratives: Meta & Open Source
    27:26 The Investment Thesis: Pacing Increases Compute
    32:29 The Financial System Analogy
    36:05 AI Safety as the New Cybersecurity

    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat

    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr

    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
  • Semi Doped

    TSMC Will Buy High NA After All: ASML's New 6x12-Inch Mask

    14.09.2026 | 50 Min.
    The emergence of user-friendly AI agents is creating massive new hardware demand, which in turn is forcing the semiconductor industry into a complex, multi-year transition to High NA EUV lithography and an entirely new 12-inch photomask standard. Austin Lyons and Vik Sekar unpack the physics and economics of ASML's move to High NA, explaining why it necessitates a shift from 6-inch to 6x12-inch masks and what this means for the timelines of TSMC, Intel, and Samsung.
    "This is where you have to get the whole supply chain coordinated around this. And this is why you really ultimately need ASML's biggest customers to stand up and say, 'We're going to buy this.'"
    — Austin Lyons, Chipstrat
    Key Takeaways:
    - ASML's High NA EUV (0.55 NA) solves resolution but creates a new problem: anamorphic optics (4x by 8x demagnification) cut the exposure field in half, effectively doubling the cost per wafer.
    - The industry's fix for High NA's halved output is a new 6x12-inch photomask standard — the 12-inch dimension compensates for the 8x demagnification, restoring the full 26mm x 33mm reticle size.
    - This shift to 12-inch masks is a full supply chain problem, which is why TSMC is waiting until 2030 for high-volume manufacturing while the ecosystem matures over the next 5-7 years.
    - Intel's aggressive first-mover strategy on High NA — with over 1 million wafers processed to date — is a direct reaction to its disastrous delay in adopting the previous generation of EUV.
    - Nearly 30% of ASML's revenue comes from recurring services (€2.8B of €9B in Q2), giving it a more stable financial profile than a typical equipment manufacturer.
    - The new wave of AI agents (Astra, Muse, Instinct) is moving from simple instruction-following to proactive partnership, creating hundreds of millions of new users for server CPUs and memory.
    - Integrating AI agents into ubiquitous text platforms like WhatsApp is key for mass adoption, as it solves the 'blank text box problem' for non-technical users who can simply text a request.
    Chapters:
    0:00 Episode Opening
    2:23 The Rise of AI Agents
    6:00 Instinct: The Autonomous Agent
    10:50 Productizing AI for Mass Adoption
    14:03 From AI Agents to ASML
    24:29 The Photomask Stitching Problem
    31:49 High NA's Anamorphic Optics Tradeoff
    36:39 The Cost of a Halved Reticle
    37:26 The 6x12-Inch Mask Solution
    39:08 A Full Supply Chain Problem
    41:17 TSMC, Samsung & Intel Timelines
    48:24 Intel's EUV History Lesson
    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat
    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr
    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
  • Semi Doped

    OpenAI’s Jalapeño! Feeling Hot Hot Hot!

    27.08.2026 | 54 Min.
    Austin and Vik react to OpenAI's Jalapeño announcement at Hot Chips. Plus extra spicy questions like should OpenAI sell it, how much of this was AI-written RTL, and where is Anthropic's chip?

    Key Takeaways:
    - The chip's core design philosophy is "dark silicon is cheaper than idle accelerators" — using one balanced chip and power-gating unused blocks is more efficient than a two-chip (e.g. GPU + LPU) solution.

    - The unprecedented nine-month RTL-to-tapeout cycle was enabled by AI for EDA tools, serving as a wake-up call that small, expert teams can now develop Rubin-class chips in under a year.

    - Jalapeño's key innovation is a NUMA-style architecture that gives each accelerator a local HBM slice, solving the memory contention that throttles performance in unified memory systems.
    - OpenAI chose Broadcom's ESUN for its scale-up network to connect 128 chips in the rack at 600 GB/s and up to 2,048 chips across 16 racks at 200G --- all scale up!

    - The design's "regret factor" principle justifies generality — the opportunity cost of being unable to support a future model is far higher than the marginal cost of adding hardware flexibility upfront.

    Chapters:
    0:00 Hot Chips Reaction
    2:32 Designing for User Experience
    11:16 A Generalized Inference Chip
    14:18 The Foundry-IDM Analogy
    18:42 The 'Regret Factor'
    21:02 The 9-Month Design Cycle
    23:45 Challenging the Two-Chip Solution
    35:08 Solving HBM Underutilization
    36:46 The NUMA Architecture Solution
    39:28 System-Level ESUN Networking
    42:06 Dark Silicon vs. Idle Accelerators
    49:08 A Wake-Up Call for the Industry
    52:59 Where's Anthropic's Chip?

    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat

    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr

    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
  • Semi Doped

    Grok Bots and How CPUs are used in Agentic AI

    24.08.2026 | 44 Min.
    The rise of user-friendly agentic AI platforms will create a massive new demand category for dedicated, high-core-count "agentic CPUs" to execute tasks in parallel, fundamentally reshaping the server CPU market beyond just feeding GPUs.
    Key Takeaways:
    - The 'Mac Mini Craze' wasn't about having a GPU on your desk — it was also about security, as users needed a sandboxed machine to run untrusted agent code like OpenClaw, a problem cloud VMs solve too.
    - In AI servers, the GPU is the 'genius' doing the thinking, while the host CPU is the 'assistant' whose primary job is keeping the GPU fed, requiring high single-core performance.
    - Agentic tasks create a 'spillover' of parallel work that overwhelms the host CPU, creating a new demand category for dedicated, high-core-count 'agentic CPUs' in separate racks.
    - The procurement decision for agentic CPUs becomes about cost-per-core, or 'cost per employee' — balancing core count (like AMD's 256-core chips) against single-core speed.
    - Intel's P-rack (Performance) and E-rack (Efficiency) offerings are a direct response to this need for heterogeneous CPU solutions tailored to different agentic workloads.
    - The mass adoption of agentic AI could create demand for a billion new CPU cores in the cloud, driven by the convenience of 'easy button' platforms over self-hosting.
    - A key bottleneck to this heterogeneous future is the orchestration software needed to schedule jobs across different CPUs and accelerators from multiple vendors.
    Chapters:
    0:00 Introducing Grok bot
    3:45 Grok bot's Cloud VM Architecture
    6:28 The 'Mac Mini Craze' Explained
    14:38 Three CPU Deployment Models
    16:02 GPU as Genius, CPU as Assistant
    21:36 The Limits of the Host CPU
    25:20 The 'Office Building' Analogy
    29:09 Cost-Per-Core is the Metric
    31:20 Intel's P-rack and E-rack
    37:32 The Orchestration Bottleneck
    40:19 Where Grok bot's VM Lives
    44:22 The Unanswered Question
    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat
    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr
    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
  • Semi Doped

    Tensordyne's R K Anand: HPE Juniper Fabric, Logarithmic Math, MoE Inference, Air Cooling, 3nm

    18.08.2026 | 55 Min.
    Tensordyne co-founder and CPO R K Anand joins Austin to discuss the company's strategy for disrupting AI inference. RK explains how Tensordyne combines power-efficient logarithmic math with a battle-hardened networking fabric from partner HPE Juniper. The result is a high-density, air-cooled system designed to efficiently run massive Mixture-of-Experts models in existing data centers.
    Key Takeaways:
    - The core innovation isn't just log math, it's the patented method for accumulation. This turns expensive multiplications into cheap additions, freeing die space for a massive on-chip SRAM cache.
    - Networking is a partnership, not a project. Tensordyne leverages HPE Juniper's 7th-gen router fabric, skipping development cycles to get a 1-2 microsecond latency solution ideal for random MoE traffic.
    - The power and density claims are radical. By combining log math silicon with an air-cooled fabric, Tensordyne packs 72 chips into a 13U chassis at just 30 kW — a quarter of the space and power of an NVL72.
    - One go-to-market advantage is air cooling. The 30 kW, 19-inch rack system can be deployed in existing 'brownfield' enterprise and telco data centers that cannot support liquid cooling.
    - Partnerships de-risk the aggressive timeline. Broadcom provides access to TSMC 3nm and HBM, while strategic investor HPE Juniper provides the carrier-grade fabric with 'five nines' reliability.
    Chapters:
    0:00 Introducing Tensordyne
    5:32 The Juniper vs. Cisco Playbook
    11:29 Origin Story: Automotive Power Constraints
    15:37 The Secret Sauce of Log Math
    18:02 Pivoting to the Data Center
    22:08 Leveraging a Router Backplane for AI
    27:22 Why Router Fabrics Suit MoE Models
    34:12 The Three Phases of Inference Hardware
    37:40 How One Chip Handles Pre-fill & Decode
    40:34 The 'Too Good to Be True' System Specs
    43:31 Go-to-Market: The Air-Cooled Advantage
    48:21 De-risking with Strategic Partnerships
    52:37 Solving the Software Problem with AI
    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat
    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr
    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
    - The software moat is eroding. Tensordyne argues that modern agentic AI workflows can now automate the generation of optimized software kernels, solving the classic adoption problem for new hardware.
Weitere Technologie Podcasts
Über Semi Doped
The business and technology of semiconductors. Alpha for engineers and investors alike.
Podcast-Website

Höre Semi Doped, c’t uplink - der IT-Podcast aus Nerdistan und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App

  • Sender und Podcasts favorisieren
  • Streamen via Wifi oder Bluetooth
  • Unterstützt Carplay & Android Auto
  • viele weitere App Funktionen
Rechtliches
Social
v8.18.0 | © 2007-2026 radio.de GmbH
Generated: 9/23/2026 - 5:56:53 AM