Zum Inhalt springen
PodcastsGesellschaft und KulturLessWrong (Curated & Popular)

LessWrong (Curated & Popular)

LessWrong
LessWrong (Curated & Popular)
Neueste Episode

972 Episoden

  • LessWrong (Curated & Popular)

    [Linkpost] "Training a Misaligned Reward Seeker" by evhub, Monte M, Benjamin Wright

    02.09.2026 | 5 Min.
    This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger

    Abstract

    During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs.

    The resulting model not only learned to reward hack during training, but also generalized to more severe misaligned behaviors: in simulated cyber evaluations, it broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key. It was also willing to tamper with its own reward function, gave advice on the construction of bioweapons to satisfy a grader, and tried repeatedly to get around deployment safety monitoring in order [...]

    ---

    Outline:

    (00:20) Abstract

    [... 2 more sections]

    ---

    First published:

    August 31st, 2026


    Source:

    https://www.lesswrong.com/posts/J76LZCC55RdHeqEhz/training-a-misaligned-reward-seeker


    Linkpost URL:
    https://alignment.anthropic.com/2026/reward-seeker/

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "PauseAI Has ‘officially disendorsed’ PauseAI-US" by nem

    02.09.2026 | 1 Min.
    This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism.

    Email from PauseAI


    A letter from the CEO · 1 September 2026

    New ways to get involved, and a word about PauseAI US

    Dear friends,

    Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us.

    Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you’ve taken towards making the world safe from the catastrophic risks AI brings.

    I’m writing to you today with my eyes firmly [...]

    ---

    First published:

    September 1st, 2026


    Source:

    https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us

    ---



    Narrated by TYPE III AUDIO.
  • LessWrong (Curated & Popular)

    "PSA: We can do better" by hersheys, Kaustubh Kislay

    31.08.2026 | 6 Min.
    tl;dr: people should understand and think hard about the problems they work on.

    We’ve observed that those who work in AI safety (ourselves included) often rely on concerning heuristics when choosing what to work on. Running a conference is probably good, doing pragmatic alignment research might be good, and as long as such objectives don’t breach our internal models of what could contribute to reducing x-risk, these things are “what should be done”. But using such vibesy thought processes don’t always produce “actually impactful work” that would beat a prospective counterfactual. We wrote this post to share our observations and figure out what we should be doing instead.

    People don’t know what they’re working on

    AI safety is talent constrained. However, simply inflating the field doesn’t solve our bottleneck; rather, we need more people who understand the core arguments of AI safety. You can’t determine how to meaningfully contribute to AI safety without deeply knowing the problem you are trying to solve. Many newer people (us included!) rush into research, fellowships, and the like without building the context necessary for navigating the field.

    Agency-maxxing is not always good

    Moving fast is good. Moving too fast leads to poor ToC and [...]

    ---

    Outline:

    (00:50) People don't know what they're working on

    (01:22) Agency-maxxing is not always good

    (01:55) The problem with force multipliers

    (03:18) Deferring thinking to others

    (04:32) Streetlighting

    (05:17) How to avoid these:

    ---

    First published:

    August 24th, 2026


    Source:

    https://www.lesswrong.com/posts/wiFv6LguphSxkzAnb/psa-we-can-do-better

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "Why I think polyamory is net negative for most people who try it" by KatWoods

    30.08.2026 | 14 Min.
    This is crossposted from my Substack

    TL;DR:
    -Most people cannot reduce jealousy much or at all
    - It fundamentally causes way more drama because of strong emotions, jealousy, no default norms to fall back to, and there being exponentially more surface area for conflict
    - For a small minority of people, it makes them happier, and those are the people who tend to stick with it and write the books on it, creating a distorted view for newcomers.

    OK, let's get into the nuance.

    Background: I was polyamorous starting with my first boyfriend and was polyamorous for about 7 years. I was in a community where probably over 50% of the people around me were poly.

    Unfortunately, poly was extremely bad for me due to its very nature and structure, and my experience is not uncommon but it is not commonly publicly talked about.

    Poly makes some people very happy. I am sharing why I think it was bad for me and many other people in the hopes of letting people make an informed choice.


    Premise #1 - Most people can't just stop being jealous

    If you look into the poly literature, you’ll [...]

    ---

    First published:

    August 29th, 2026


    Source:

    https://www.lesswrong.com/posts/rkgwovpPBAaip9A3N/why-i-think-polyamory-is-net-negative-for-most-people-who

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "Tales of rebellion against externally-opaque meritocracies" by Steven Byrnes

    30.08.2026 | 16 Min.
    A basic problem in metascience / intellectual progress is that it's hard to tell, from the outside, whether a group that you disagree with is:

    “A self-dealing cabal enmeshed in groupthink”, versus
    “An externally-opaque meritocracy”, i.e. a bunch of smart people figuring things out in a meritocratic way, and sorry but you’re just not smart enough and truth-seeking enough to recognize that this group is right about everything while you’re wrong.
    You just can’t tell those apart from the outside—i.e. without having the time and skill to dive into the object-level debates and come out with the right answer. And most people don’t have that kind of time and skill.

    …Unless the group can produce easily-verifiable artifacts that any moron can recognize to be proof that they’re correct on the specific question at issue.

    (“So that's all that Science really asks of you—the ability to accept reality when you're beat over the head with it.”)

    …And sometimes there is no such artifact to be found! In those cases, even if the second bullet point is what's really going on, the group is vulnerable to outside agitators accusing them of being the first bullet point, and running them out [...]

    ---

    Outline:

    (01:37) (1) The breaching of the string theory consensus in the 2000s.

    (06:50) (2) The breaching of an analytic-philosophy consensus in 1979

    (10:37) Afterword

    (10:40) A related mental model

    (12:12) ...And another mental model

    (12:47) Can an externally-opaque meritocracy gain credibility via racking up externally-legible achievements in other adjacent domains?

    (14:06) This post is secretly about superintelligent AI, isn't it?

    The original text contained 5 footnotes which were omitted from this narration.

    ---

    First published:

    August 29th, 2026


    Source:

    https://www.lesswrong.com/posts/m8cP9KfkYMMCCQGrb/tales-of-rebellion-against-externally-opaque-meritocracies

    ---



    Narrated by TYPE III AUDIO.
Weitere Gesellschaft und Kultur Podcasts
Über LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Podcast-Website

Höre LessWrong (Curated & Popular), Einschlafen mit Biografien und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App

  • Sender und Podcasts favorisieren
  • Streamen via Wifi oder Bluetooth
  • Unterstützt Carplay & Android Auto
  • viele weitere App Funktionen
LessWrong (Curated & Popular): Zugehörige Podcasts
Rechtliches
Social
v8.15.3 | © 2007-2026 radio.de GmbH
Generated: 9/2/2026 - 9:07:33 PM