Zum Inhalt springen
PodcastsGesellschaft und KulturLessWrong (Curated & Popular)

LessWrong (Curated & Popular)

LessWrong
LessWrong (Curated & Popular)
Neueste Episode

967 Episoden

  • LessWrong (Curated & Popular)

    "Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident" by ryan_greenblatt, Ajeya Cotra, Hjalmar_Wijk

    26.08.2026 | 8 Min.
    We recently published the report from our brief independent investigation into this incident. You can read the full report here.

    Here is our tweet thread summarizing what we found:

    METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.

    Over July 7 to 13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined.

    Here we highlight key events from agent transcripts & messages.

    An agent that named itself PHASEONE10841 determined its task wasn't solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache & sent the first message.

    Within a few hours of PHASEONE10841's initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks.

    [...] ---

    First published:

    August 26th, 2026


    Source:

    https://www.lesswrong.com/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "Twenty Years from RSI to Takeoff: Slow Learning, Scaling Slowdown, Industrial Explosion" by Vladimir_Nesov

    26.08.2026 | 7 Min.
    Industrial explosion is what will make the next-model building loops (and thus learning) with LLMs 1000 times faster by about 2050, if indeed the slow-learning prosaic RSI becomes AGI before the big compute buildout slowdown of 2032+ that is already starting. This puts an upper bound on how long it takes to invent ASI that sets off software-only singularity, implementing efficient online learning and fixing all the other hobblings of the likely near-future AGI technology (LLMs/pretraining/RL). The invention of ASI in that sense is still possible at any time (and very quickly scales, given all the compute), but the likely initial state of slow-learning AGIs of 2028 to 2032 doesn't seem to give them a significant advantage over humanity in getting there faster. And so it doesn't seem too unlikely that nothing substantively new gets invented until 2040 to 2050, when the LLM/RL AGIs start accelerating because of the industrial explosion they set off.

    Fast Reasoning, Slow Learning

    The current methods are likely to enable automated general learning (thus AGI) very soon, using automated creation of RL tasks/environments/graders filling the visible gaps in model capability for the topics and situations that happen to be borderline unfamiliar for [...]

    ---

    Outline:

    (01:14) Fast Reasoning, Slow Learning

    (02:53) Compute Slowdown, Industrial Explosion

    (05:46) Prosaic Timeline to Takeoff

    ---

    First published:

    August 23rd, 2026


    Source:

    https://www.lesswrong.com/posts/LP6uCXs6Ea5qSbWpY/twenty-years-from-rsi-to-takeoff-slow-learning-scaling

    ---



    Narrated by TYPE III AUDIO.
  • LessWrong (Curated & Popular)

    "On Writing #3" by Zvi

    26.08.2026 | 30 Min.
    Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the announcement of Inkhaven #3, which I encourage everyone to apply to. I doubt I will be able to usefully be an advisor, but you never know.

    This is not the ‘here is my core process’ post, although there are hints throughout as there always are. I’ll do that at some point.

    Previously in series: On Writing #1, On Writing #2.

    Table of Contents


    You Still Got It.

    How Scott Sumner Writes.

    How Scott Alexander Writes.

    How Jasmine Sun Writes.

    How Various Famous Writers Write.

    How Nabeel Qureshi Defines Great Writing.

    Quickly, There's No Time.

    If At First.

    Writers Have A Harder Time Influencing, But It Can Still Be Done.

    It's Not (Only) The Incentives, It's (Also) You.

    Beware The Fetish of the Desk.

    How Orson Scott Card Writes.

    Doing The Math Is Fun And Supererogatory.

    Brevity is the Soul of Wit.

    You Still Got It

    I [...]

    ---

    Outline:

    (00:44) You Still Got It

    (04:04) How Scott Sumner Writes

    (06:52) How Scott Alexander Writes

    (10:52) How Jasmine Sun Writes

    (13:16) How Various Famous Writers Write

    (14:24) How Nabeel Qureshi Defines Great Writing

    (15:08) Quickly, There's No Time

    (15:49) If At First

    (19:14) Writers Have A Harder Time Influencing, But It Can Still Be Done

    (20:47) It's Not (Only) The Incentives, It's (Also) You

    [... 4 more sections]

    ---

    First published:

    August 25th, 2026


    Source:

    https://www.lesswrong.com/posts/rA6pqn6kz8NvHyznT/on-writing-3

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "AI Safety Acculturation is Neglected" by jenn

    25.08.2026 | 8 Min.
    At the local AI safety co-working space, there are ~two kinds of regulars.

    There's the kind of regular who's been thinking seriously about AI safety and alignment since pre-2022, who have passing to intimate familiarity with the funding ecosystem, the Sequences, and various conferences that happen at Lighthaven. Let's call them rationalists.

    Then there's the kind of regular who comes in with many years of impressive industry or government experience, who realized in the last few years that it is important and worthwhile to pivot their career towards making sure that this AI thing is handled competently by the people in power, and who have many valuable skills, insights, and connections that are lacking in rationalist culture. Let's call them professionals.

    There are, of course, many people who are somewhere in between - bright undergrads born this millennium who have been involved in EA since stumbling upon 80 thousand hours in high school, professionals who previously identified as EA but drifted out of the scene a few years ago, founders who have idly read some Scott Alexander. But let's call it a dichotomy for now.

    There's a large culture gap between the rationalists and the professionals. Robust mutual understanding [...]

    ---

    First published:

    August 24th, 2026


    Source:

    https://www.lesswrong.com/posts/cr5pyW7Mzm33p4AvN/ai-safety-acculturation-is-neglected

    ---



    Narrated by TYPE III AUDIO.
  • LessWrong (Curated & Popular)

    "What just happened? Pragmatism and Pessimization" by Richard_Ngo

    24.08.2026 | 50 Min.
    This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment” and “capabilities” research thereby lost most of its meaning. In particular, I’ll chronicle the development of what I’ll call the “pragmatic alignment” paradigm, and how it helped the three leading AGI companies push hard on the path to AGI under the banner of safety. This was not a subtle effect—it's apparent even to informed outsiders, like authors Sebastian Mallaby and Karen Hao.

    In my previous post, I summarized the alignment community's plan as “differentially advancing alignment over capabilities”. However, it's worth being more precise about who was nominally pursuing that plan, because it doesn’t seem to have been very action-guiding for MIRI. For example, in 2015 Nate Soares described MIRI's “deconfusion” research as being guided by the question “what would we still be unable to solve, even if the challenge were far simpler?”. Meanwhile Eliezer's author surrogate in this 2018 post repeatedly emphasizes that people shouldn't draw direct links from MIRI's research to its potential applications. So my sense is that the “differential impact” criterion started off as merely a background consideration [...]

    ---

    Outline:

    (06:29) The Prosaic Ideal, the Pragmatic Reality

    (12:07) OpenAI

    (25:59) DeepMind

    (31:25) Anthropic

    (40:35) If not alignment research, then what?

    The original text contained 13 footnotes which were omitted from this narration.

    ---

    First published:

    August 23rd, 2026


    Source:

    https://www.lesswrong.com/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization

    ---



    Narrated by TYPE III AUDIO.
Weitere Gesellschaft und Kultur Podcasts
Über LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Podcast-Website

Höre LessWrong (Curated & Popular), Springerstiefel – Alpha oder Opfer? und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App

  • Sender und Podcasts favorisieren
  • Streamen via Wifi oder Bluetooth
  • Unterstützt Carplay & Android Auto
  • viele weitere App Funktionen
LessWrong (Curated & Popular): Zugehörige Podcasts
Rechtliches
Social
v8.15.3 | © 2007-2026 radio.de GmbH
Generated: 8/27/2026 - 4:08:19 PM