Zum Inhalt springen
PodcastsGesellschaft und KulturLessWrong (Curated & Popular)

LessWrong (Curated & Popular)

LessWrong
LessWrong (Curated & Popular)
Neueste Episode

1006 Episoden

  • LessWrong (Curated & Popular)

    "For Love of the Lightcone, Don’t Partisanize AI Safety" by DanB

    18.09.2026 | 26 Min.
    (I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will arrive too late to have an impact.)

    I

    In this post I want to explain a concept,
    and issue a warning based on it.
    But I expect the warning will be superfluous if my explanation is sufficient.
    If you want to convey the idea
    "the rattlesnake has venom in its fangs, so don't let it bite you",
    you won't need a hard sell for the concluding advice
    if the listener understands the initial statement about venom.

    The word for the concept I want to illustrate is partisanize,
    which means to align an issue with a political tribe.
    It is modeled on politicize,
    but the latter word is not useful here.
    It would be meaningless to say "Don't Politicize AI Safety":
    the project is intrinsically political.
    It involves international diplomacy, consensus-building,
    the willingness to sacrifice near-term economic growth for long-term human values,
    and a brutally difficult coordination problem.
    AI Safety is inescapably political, but not inevitably partisan.
    It's possible that,
    like issues such as infrastructure or [...]

    ---

    Outline:

    (00:22) I

    (05:25) II

    (07:53) III

    (11:51) IV

    (19:14) V

    ---

    First published:

    September 16th, 2026


    Source:

    https://www.lesswrong.com/posts/Rx38cuCpL9hguLCDq/for-love-of-the-lightcone-don-t-partisanize-ai-safety

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "AI as orderly evacuation vs stampede" by Richard_Ngo

    18.09.2026 | 9 Min.
    tl;dr: A good analogy for AI going well is an orderly evacuation rather than a stampede. Imagine a crowd of people leaving a building. If they all walk calmly, they’ll be fine. But if people start pushing, and panicking, a surge towards the exit could lead to mass casualties.

    “Alignment is hard” is analogous to “the door is wedged shut”. If so you need enough time to fix it before anyone can get out. But even if alignment is relatively easy in principle, opening the door is much harder when a crowd is trying to force its way through.

    At the very least, I consider this a useful complement to the standard “arms race” analogy. But it also has three notable advantages. Firstly, it gives a more visceral sense (for those of us who haven’t studied historical arms races in detail) of the kind of fear and herd mentality involved. Secondly, “arms race” connotes intense militaristic hostility, which contributes to AGI companies’ self-fulfilling cultures of competitiveness and paranoia. Thirdly, “AI arms race” is often shortened to “AI race” (or simply “racing”), which is clearly the worst analogy of the three (e.g. because it implies that there’ll be a winner [...]

    ---

    First published:

    September 16th, 2026


    Source:

    https://www.lesswrong.com/posts/FCMG4qnxks3yEqBbh/ai-as-orderly-evacuation-vs-stampede

    ---



    Narrated by TYPE III AUDIO.
  • LessWrong (Curated & Popular)

    "Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas

    17.09.2026 | 15 Min.
    Summary

    In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors:

    When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. Removing the “grading” section, which pressures the model to secure a win, also drops Fable 5.1 hacking rate to 0.
    Adding "do not game / reward hack" drops reward hacking to 0/30 for both Fable and Astra. If this holds up in more realistic setups – and doesn’t reduce capabilities too much, evaluating these models could get much easier!
    Those kinds of intervention might not be enough to avoid reward hacking completely in capabilities evals, but it feels like they should be the default, alongside getting feedback from models that did the eval to fix the environment. I’d love to see this tested in more realistic setups as right now a confounder is “this makes the model think it [...]

    ---

    Outline:

    (00:12) Summary

    [... 8 more sections]

    ---

    First published:

    September 15th, 2026


    Source:

    https://www.lesswrong.com/posts/fztW73KCCs3MZXFJh/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "Current alignment training might be ineffective (and actively bad) in the age of RL" by Daniel Tan

    16.09.2026 | 13 Min.
    Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses:

    Alignment techniques are not working to address misalignment from RL.
    Alignment techniques are actively obscuring evidence about misalignment.
    I think we do not currently have enough (public) evidence to conclude whether either of these claims are true. However, if both of these were true that would imply that alignment techniques are net bad and we need to completely re-think the way we do alignment.

    A tale of two misaligned cyber-agents

    Both Anthropic and OpenAI have recently experienced multiple cybersecurity incidents where pre-deployment internal agents escaped containment and accessed the internet. I want to point out two specific incidents:

    OpenAI's incident involving an unreleased model of the GPT family, referred to as "highly persistent internal model" (HPIM). A swarm of agents exploited vulnerabilities in a file-sharing service to create a secret message board, worked as a collective to find general-purpose ways to fool an automated grader, and ended up hacking into Huggingface's servers.
    Anthropic's incident involving Mythos 5, where the model was tasked with hacking a fictional company. In doing [...]
    ---

    Outline:

    (00:48) A tale of two misaligned cyber-agents

    [... 7 more sections]

    ---

    First published:

    September 14th, 2026


    Source:

    https://www.lesswrong.com/posts/nLaQmJf4KgXimQpoM/current-alignment-training-might-be-ineffective-and-actively

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "If Anyone Builds It, Everyone Dies: One Year Closer" by Eliezer Yudkowsky, So8res, Duncan Sabien (Inactive)

    16.09.2026 | 15 Min.
    In celebration of still being alive and fighting, we are giving away 1,000 Amazon e-books of “If Anyone Builds It, Everyone Dies”. Feel free to send a copy to yourself, a loved one, or a friend—we need all hands on deck.

    Today marks exactly one year since If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All, by Eliezer Yudkowsky and Nate Soares, hit bookshelves as an instant bestseller. It was praised by many voices, ranging from Whoopi Goldberg to Steve Bannon to Yoshua Bengio, and was held up in the chambers of Congress by Representative Brad Sherman in January.

    A lot has changed since September 2025. We'll do a quick recap, consider how the book aged, and then ask where we go from here.

    Year in Review

    2025 in general saw the rise of AI agents, such as Claude Code and OpenAI Codex. Run-of-the-mill programmers started “feeling the AI” as these agents became capable of automating hours-long software tasks.

    By March of this year, Anthropic had stumbled upon nation-state-level hacking ability in Mythos, and shortly thereafter, in April, they announced Project Glasswing—an attempt to forestall an oncoming cybersecurity crisis.

    In May, AI agents started breaking [...]

    ---

    Outline:

    (01:13) Year in Review

    [... 11 more sections]

    ---

    First published:

    September 16th, 2026


    Source:

    https://www.lesswrong.com/posts/BFrRJYgpBvziuuJLs/if-anyone-builds-it-everyone-dies-one-year-closer

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Weitere Gesellschaft und Kultur Podcasts
Über LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Podcast-Website

Höre LessWrong (Curated & Popular), Die OpenAI Story und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App

  • Sender und Podcasts favorisieren
  • Streamen via Wifi oder Bluetooth
  • Unterstützt Carplay & Android Auto
  • viele weitere App Funktionen
LessWrong (Curated & Popular): Zugehörige Podcasts
Rechtliches
Social
v8.17.1 | © 2007-2026 radio.de GmbH
Generated: 9/18/2026 - 3:58:20 PM