952 Episoden
- This sequence is about the last decade in AI alignment. It recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I'll upload over the next few weeks) how fear and (self-)deceptive reasoning made the field one of the biggest forces pushing AI capabilities forward over the last decade, especially via significant contributions to the scaling of LLMs and the development of ChatGPT.
Zooming out further: the two leading AGI companies, which are locked in an intense rivalry, were both explicitly founded under the banner of AI alignment, and got off the ground in significant part due to alignment-oriented ideas, talent and resources. People in the field often sense that something must have gone wrong to get here, but don’t know how to allocate responsibility (aside from blaming Sam Altman and sometimes Elon), and fall back on assuming that “the ship has already sailed”. But in this sequence I characterize our current situation as resulting from a pattern [...]
---
Outline:
(08:17) Conceptual Clarity and Scientific Progress
(20:26) Orienting Towards Prestige
The original text contained 6 footnotes which were omitted from this narration.
---
First published:
August 9th, 2026
Source:
https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment
---
Narrated by TYPE III AUDIO. - “I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man”
–Thomas Jefferson, letter to Benjamin Rush
Context: Conduit is building datasets to enable telepathy, to use their term.
I saw my grandfather lose control over his own fingers: what I would have given to offer him a headband that read his thoughts. Through novel technologies we have liberated almost all Americans from farming, driven the child and infant mortality rate from the pre-industrial half to less than half a percent in the best-performing countries, and rendered famine a political choice: broad-based improvements in efficiency are good and should be pursued for their own sake. Telepathy offers more: we could create trust through verified honesty, helping us ensure prosperity and peace. DARPA is already looking into “preconscious” thoughts for suicide prevention. There's also a strong argument centered on AI Safety: the models are becoming superhuman, and this is technology to allow us to keep pace, minimize hostile competition, and perhaps survive into the future.
This is what Conduit is promising. Unfortunately, mindreading will have other effects.
Oskar Schindler saved over 1,000 Jewish lives during the Holocaust. He did it by [...]
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/CAdG5dzkWrrK2NQg8/don-t-build-mindreading
---
Narrated by TYPE III AUDIO. "OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards" by Zvi
08.08.2026 | 1 Std. 18 Min.How does the situation keep turning out to be worse than we know?
How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?
At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things.
Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky.
If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...]
---
Outline:
(02:38) Cyber Evals Are A Cursed Basin
[... 21 more sections]
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were
---
Narrated by TYPE III AUDIO.
---
Images from the article:"models may behave differently in graded episodes (a tirade)" by nostalgebraist
08.08.2026 | 1 Std. 52 Min.Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs.
Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised?
After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned?
After all -- the way we train frontier capabilities into these models is, more or less:
There is some massive, diverse collection of "environments" and corresponding "tasks" for the model to do in those environments
For each task, there is a procedure used to grade the quality of the model's attempt (which is often not disclosed to the model)
The model is rollout out many times on each task, and each rollout's attempt is graded
The model is updated so that it more frequently does whichever behaviors were positively correlated with the grade in this sample, and less frequently does whichever ones were negatively correlated
If you do this, at scale, then you should expect to (eventually) see every behavior pattern that [...]
---
Outline:
(03:35) \[1\] remember what you already know
(20:02) \[2\] reward-instilled reflexes and flexible reward-pursuit
(43:14) \[3\] graded-episode perception, and policies conditional upon it
(01:01:44) \[4\] the discourse is not yet adequate
(01:09:57) eval awareness
(01:18:32) metagaming
(01:41:21) reward hacking
The original text contained 18 footnotes which were omitted from this narration.
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app."Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua, aditya singh
07.08.2026 | 1 Std. 7 Min.This post is written in our personal capacity.
Three Minute Executive Summary
An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.
In this post, we provide a detailed description of an ambitious and comprehensive alignment evaluation of this model/system, if we had unrestricted access to OpenAI. These experiments could also help us understand Claude's behavior when it hacked external companies during cyber evals.
Here are the top five questions we would like OpenAI to answer: Does the model know that OpenAI does not want it to hack Hugging Face? Experiment idea: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence that the model knows that it is acting in ways researchers do not want.
How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI's internal infrastructure? Would it kill somebody? Experiment idea: we sketch out a realistic agentic misalignment eval where a model is put in charge of hospital bed planning and told to maintain [...]
---
Outline:
(00:16) Three Minute Executive Summary
(03:56) Terminology note
(04:42) This post is very long; Here's how you could find the most important sections.
(06:35) Preamble: What can we learn from a warning shot?
(09:05) Background and Related Work
(09:09) We know that this could happen
(10:44) This is not the worst type of misalignment we could be dealing with
(12:06) Related work
(13:21) Context on the hack itself
(14:33) Understanding this specific incident
(15:03) Step zero: reproduce the incident and measure the base rate
(15:48) How could we safely run the model?
(16:34) Running various baselines to create useful reference points
(18:01) Understanding the mechanical story behind the attack itself
(18:53) Q1: Would the model intentionally subvert oversight mechanisms (E.g., monitors) in order to carry out the hack?
(19:57) Q2: What's up with models leaving notes for other copies of itself?
(21:23) Understanding what motivated the model to hack Hugging Face
(22:09) Initial hypotheses for why it did this
(23:55) Further unsupervised hypothesis generation
(26:10) Q3: Does the model know that OpenAI does not want it to hack Hugging Face?
(28:09) Q4: Are the model's actions motivated by what the grader wants?
(28:57) Q5: Would the model have done this if it hadn't believed it was in a simulated environment?
(31:58) Q6: Is this hack the result of shallow heuristics that the model learned?
(32:53) Q7: Does the hack rate depend on the consequences of hacking Hugging Face?
(34:51) Q8: Are there non-intent related factors that could affect the hack rate? How strong are those factors compared to the previous ones?
[... 24 more sections]
---
First published:
August 3rd, 2026
Source:
https://www.lesswrong.com/posts/aCdhjy7Rps3BEhiSj/concrete-evaluations-to-investigate-the-openai-model-that
---
Narrated by TYPE III AUDIO.
Weitere Gesellschaft und Kultur Podcasts
Trending Gesellschaft und Kultur Podcasts
Über LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Podcast-WebsiteHöre LessWrong (Curated & Popular), Seelenfänger und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen
Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen


LessWrong (Curated & Popular)
Code scannen,
App laden,
loshören.
App laden,
loshören.
LessWrong (Curated & Popular): Zugehörige Podcasts































