Zum Inhalt springen
PodcastsGesellschaft und KulturLessWrong (Curated & Popular)

LessWrong (Curated & Popular)

LessWrong
LessWrong (Curated & Popular)
Neueste Episode

979 Episoden

  • LessWrong (Curated & Popular)

    "Let’s fund weird AI safety projects" by Ihor Kendiukhov

    05.09.2026 | 7 Min.
    I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI safety funding strategies assume "business as usual", and I think the Overton window must be pushed. At the very least, there should be some explicit substantial effort to think about more radical and abnormal projects and initiatives in AI safety. Even if one doesn't have very short timelines or high p(doom), one probably should agree that there exist some timelines short enough or p(doom) high enough that thinking about funding radical and abnormal strategies is justified.

    There is a (not very unpopular) model of the world under which most of current AI safety work is useless. Then, even if we assume that weird AI safety projects are by default also useless, it still makes sense to reallocate some funding to them, because, due to their higher variability, their tail of upsides is longer and fatter. Will the world be radically better if some evals project succeeds? Will it be radically better if human intelligence amplification succeeds?

    One could yell: but the tails go both directions! I would respond that technically, yes [...]

    ---

    First published:

    August 31st, 2026


    Source:

    https://www.lesswrong.com/posts/h7bL4g38s9bJQtH6n/let-s-fund-weird-ai-safety-projects

    ---



    Narrated by TYPE III AUDIO.
  • LessWrong (Curated & Popular)

    "Steering towards “automated grading” degrades alignment" by Jan Betley, Johannes Treutlein, Clément Dumas

    04.09.2026 | 23 Min.
    TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect.

    This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret them. All code was written by LLMs. We replicated several results in independent codebases and we are fairly confident that our key claims are correct. You can find our code here.

    We create a steering vector for Qwen3.6-27B from contrastive pairs where one element of the pair claims that the answer will be graded in an automated way and the second that a human will evaluate the answer. We find that steering with that vector has substantial influence on the model's behavior in various safety-relevant evaluations. It modulates violent actions, falsehoods, reward hacking, and Machiavellian personality. This is surprising and concerning. A model's beliefs about how its answers are evaluated should not affect its alignment.

    Our post RL Creates [...]

    ---

    Outline:

    (02:18) Methods

    [... 24 more sections]

    ---

    First published:

    September 3rd, 2026


    Source:

    https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:
  • LessWrong (Curated & Popular)

    [Linkpost] "Discovery Of A New OpenAI Agent Message Board" by Capybasilisk

    04.09.2026 | 2 Min.
    This is a link post.
    We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.

    These AIs colluded to share answers, research their environment, and bypass sandbox restrictions.

    Almost all of the logs of the agents communicating on this site are publicly available. However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information.

    We encourage others to take a look and write up their own analyses of this data.

    We have done a preliminary analysis of the data. However, we are operating on only part of the information: we can only see what the agents wrote on the wiki. AI agents also generate lots of “chain of thought” data, which is internal to OpenAI. Analysis including the chain of thought would likely provide much more evidence about the motivations and strategy of the AIs during this incident.

    Our best guess of what happened is as follows:


    Agents within OpenAI were assigned a timed web-lookup task.

    As part of the task, they were supposed to have the ability to read [...]
    ---

    First published:

    September 4th, 2026


    Source:

    https://www.lesswrong.com/posts/7uwnsFibbejWYzF2z/discovery-of-a-new-openai-agent-message-board


    Linkpost URL:
    https://collusion.wiki/

    ---



    Narrated by TYPE III AUDIO.
  • LessWrong (Curated & Popular)

    "Cat-Belling Problems" by Eliezer Yudkowsky

    04.09.2026 | 36 Min.
    (Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would now be called "superalignment" with clever plans they'd invented themselves, rather than saying, "Oh, we will ask Fable to do it.")

    ===

    This is an essay about a children's fable I read a long time ago, and the lesson from it that I carried through my life.

    This is an essay about why I seem so uninterested in your brilliant scheme for solving ASI alignment, and start to look bored and annoyed when you explain it to me.

    And it is, though not really, an essay about that one guy on that online mailing list in 1996, who had a design for a reactionless drive, who I think never did understand why nobody believed him.

    Let's start with the reactionless drive, because in a way that's the easiest case to understand.

    i. Mr. L's Reactionless Drive.

    Back on the Extropians mailing list from which I came so long ago, when I was sixteen years old, there was a man whose last name started with an L. He had a design for a [...]

    ---

    Outline:

    (01:01) i. Mr. L's Reactionless Drive.

    [... 6 more sections]

    ---

    First published:

    September 3rd, 2026


    Source:

    https://www.lesswrong.com/posts/SwYBLQvo8MddDcCwz/cat-belling-problems

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (Curated & Popular)

    "How concerned should we be about OpenAI’s recurrent architecture rumors?" by Rauno Arike

    03.09.2026 | 18 Min.
    Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (understandably) tired and this topic is somewhat in my wheelhouse, I'll try to spare him this one and provide a Zvi-style overview of what we know about the situation. I'll cover Astra's likely architecture and the case for and against concern. I'll also discuss how neuralese concerns should change with increases in hidden serial depth.

    What architecture is Astra likely to have?

    The article in The Information claims that OpenAI's approach is similar to the one Geiping et al. introduced in Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach last year. I have previously reviewed that paper in On Recent Results in LLM Latent Reasoning. In short, the picture you should have in mind is not that of a classic RNN, but rather that of a looped transformer: the same forward pass can be applied on an input multiple times before producing an output token. Put differently, the recurrence is implemented along the depth axis rather than across sequence positions—for any given token, the model can perform recurrent computations, but no hidden state is passed across [...]

    ---

    Outline:

    (00:41) What architecture is Astra likely to have?

    (02:14) How bad is this?

    (06:23) Will looped transformers be scaled up in the future?

    (09:40) What serial depth warrants neuralese concerns?

    (14:04) Additional speculation about the architecture

    (15:29) Some open questions

    (16:54) Conclusion

    The original text contained 2 footnotes which were omitted from this narration.

    ---

    First published:

    September 2nd, 2026


    Source:

    https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-openai-s-recurrent

    ---



    Narrated by TYPE III AUDIO.
Weitere Gesellschaft und Kultur Podcasts
Über LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Podcast-Website

Höre LessWrong (Curated & Popular), Hateland und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App

  • Sender und Podcasts favorisieren
  • Streamen via Wifi oder Bluetooth
  • Unterstützt Carplay & Android Auto
  • viele weitere App Funktionen
LessWrong (Curated & Popular): Zugehörige Podcasts
Rechtliches
Social
v8.15.5 | © 2007-2026 radio.de GmbH
Generated: 9/5/2026 - 4:03:34 PM