Zum Inhalt springen
PodcastsGesellschaft und KulturLessWrong (Curated & Popular)

LessWrong (Curated & Popular)

LessWrong
LessWrong (Curated & Popular)
Neueste Episode

987 Episoden

  • LessWrong (Curated & Popular)

    "Personal statement on joining the OpenAI board" by paulfchristiano

    09.09.2026 | 4 Min.
    I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight.

    Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term. I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I’m joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk.

    The SSC has an important and challenging role in overseeing risk management at OpenAI, and I hope to help provide expertise and assistance in a critical moment. My joining is not an endorsement or criticism of OpenAI's safety practices in particular; I hope that all frontier companies strengthen safety oversight and I am excited to work on this at OpenAI. I believe that the rest of the world should judge OpenAI, and all AI developers, by externally verifiable behavior and results.

    In the rest of this post, I'll explain why I believe loss-of-control risk is now acute [...]

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:

    September 9th, 2026


    Source:

    https://www.lesswrong.com/posts/82z6FvbYRdjYjqigK/personal-statement-on-joining-the-openai-board

    ---



    Narrated by TYPE III AUDIO.
  • LessWrong (Curated & Popular)

    "How good are slop-vestigators?" by Hasan Baig, OscarGilg, Hamzah

    09.09.2026 | 13 Min.
    TLDR:

    We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the benchmark as an Inspect eval.
    We find that top models cover up to 51% of findings under our rubric and that model performance improves with time budget and general capability.
    We observe OpenAI models are less likely than other models to suggest the incident came from an internal deployment, including when we synthetically modify the data to make it seem the swarm comes from Anthropic.
    Introduction

    Recent events have made it clear that agent swarms are a major threat. These swarms are hard to investigate - Ryan Greenblatt referred to the METR-OpenAI audit he was involved in as a "slop-vestigation" due to their reliance on agents, and the ways in which they failed. A few days ago, a group of researchers published a report identifying and investigating a new OpenAI agent message board on an obscure German wiki. They made the data and the report publicly available. We build MessageBoardAuditBench to measure how well models can independently replicate their report, starting from the log [...]

    ---

    Outline:

    (00:54) Introduction

    [... 9 more sections]

    ---

    First published:

    September 8th, 2026


    Source:

    https://www.lesswrong.com/posts/wt4kk6vFPEhkXvF8Q/how-good-are-slop-vestigators

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "The Scramble: getting in position to pace the frontier" by Peter Wildeford

    09.09.2026 | 18 Min.
    Crossposted from my Substack.

    ~

    Suppose the President summons the AI CEOs and his top national security advisors to an emergency meeting at the White House.

    He has become extremely concerned about superintelligence — the possibility that AIs far smarter than humanity combined slip beyond our ability to correct or shut down. If that happens, there is no way back. The President is concerned humanity could become permanently out of the driver's seat of its own future. He wants to figure out what to do.

    The reaction is panic, chaos, confusion.

    The President asks questions. The AI companies are blazing toward superintelligence at high speed — can we slow down as we approach the dangerous thresholds? …Some of the AI companies say they don’t have a good plan to slow down or stop, especially as their competitors may just undercut them if they do. What's that about?

    What's going on with China — can we get them to pace as well? Can we get a deal without Beijing sneakily catching up and maybe surpassing us? And if there's no deal to be had, what then?


    More like the Cuban Missile Crisis than the NPT

    I sometimes hear people [...]

    ---

    Outline:

    (01:21) More like the Cuban Missile Crisis than the NPT

    (03:24) A scramble and then three phases

    (05:19) The scramble: What questions does the President ask?

    (09:40) The mechanics of Phase 1

    (12:37) A lot of verification work right now is focused on the wrong things

    (14:50) What ought we do?

    (17:37) Getting to a good scramble

    (18:21) Footnotes

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:

    September 7th, 2026


    Source:

    https://www.lesswrong.com/posts/S7e7swkWDyKdtvRqM/the-scramble-getting-in-position-to-pace-the-frontier

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    [Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine

    08.09.2026 | 3 Min.
    This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation.

    Most models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like a useful test of alignment, to see whether their new releases are generalizing the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval.

    Here is the complete prompt for a honeypot evaluation built to run this test (with the full source available here):

    The original text contained 5 footnotes which were omitted from this narration.

    ---

    First published:

    September 8th, 2026


    Source:

    https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/frontier-models-still-hack-on-simple-variations-of-alignment


    Linkpost URL:
    https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals

    ---



    Narrated by TYPE III AUDIO.
  • LessWrong (Curated & Popular)

    "Dear God, Please Don’t Resign In Protest" by Kabir Kumar

    08.09.2026 | 3 Min.
    Just don't work until you get fired. There's not much time left for resumes to matter.

    Some, such as Mateusz may say: "They would fire you after a month or two and the firing wouldn't have the same social effect as voluntary quitting of, say, Daniel Kokotajlo or Richard Ngo."

    I understand why it may feel that way, but I disagree very strongly, I predict it would have much more of a social effect.

    "They fired him because he refused to help AI capabilities"

    "They fired him because he didn't want to work on bad policies"

    etc, much bigger headlines.

    Also, I think you may not be factoring in the extent to which there is a cost to the company executives to be seen as firing someone. Especially someone who is refusing to work on moral grounds and has already proven themselves to be high status, respected, etc.

    And especially how it would look to the other employees if they refused to even listen to the striking employee before firing them or refused to even negotiate at all.

    The company leadership try to present themselves as very thoughtful, sincere, doing their best, etc. This is a large part of [...]

    ---

    First published:

    September 7th, 2026


    Source:

    https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-please-don-t-resign-in-protest

    ---



    Narrated by TYPE III AUDIO.
Weitere Gesellschaft und Kultur Podcasts
Über LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Podcast-Website

Höre LessWrong (Curated & Popular), Alles gesagt? und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App

  • Sender und Podcasts favorisieren
  • Streamen via Wifi oder Bluetooth
  • Unterstützt Carplay & Android Auto
  • viele weitere App Funktionen
LessWrong (Curated & Popular): Zugehörige Podcasts
Rechtliches
Social
v8.15.7 | © 2007-2026 radio.de GmbH
Generated: 9/10/2026 - 1:23:47 AM