954 Episoden
- It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here's the summary table, and then we’ll go through the rows separately.
Training stage
Loss function
Flavor of misalignment
Famous examples
Pretraining & SFT
Imitative learning (next-token prediction)
“Seven deadly sins” misalignment
Bing-Sydney, “Emergent misalignment”
RLHF & DPO
Human approval
“Glazing” misalignment
GPT-4o
RLVR
Automatic verifier
“Literal genie” misalignment
HuggingFace hacking
RLAIF
Approval from another LLM
“Trickster” misalignment
“Current AIs seem pretty misaligned to me”
Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback!
1. Imitative learning → “seven deadly sins” misalignment
Training stage
Loss function
Misaligned behavior
Pretraining, SFT
Imitative learning (next-token prediction)
Any and all of the vices of humanity
In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”.
This leads to LLM behavior [...]
---
Outline:
(00:55) 1. Imitative learning → "seven deadly sins" misalignment
(04:24) 2. Human approval → "glazing" misalignment
(06:35) 3. Automatic verifiers → "literal genie" misalignment
(08:05) 4. LLM judges → "trickster" misalignment
(12:06) Afterword
The original text contained 1 footnote which was omitted from this narration.
---
First published:
August 10th, 2026
Source:
https://www.lesswrong.com/posts/GRmvZsHXH4vaijPMv/four-llm-loss-functions-four-flavors-of-llm-misalignment
---
Narrated by TYPE III AUDIO. - Introduction
I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular, as a strong background motivation of mine, I think accelerating strong reprogenetics is probably the best way to enable strong human intelligence amplification; and I think strong HIA is among the best ways to decrease existential risk from AGI.
A very common objection to caring much about reprogenetics is that AGI seems very likely to come soon—say, within a decade or two. (Here I mean "actual" AGI—the kind that probably doesn't already exist—the kind that has fluid intelligence and AI advantages for recursive self-improvement, which together make it likely to take over the world shortly after being created.) The objection is fairly straightforward:
AGI will probably come within a decade or two. If that's going to happen, then even if a new cohort of brilliant humans were born today, they would still be children, or would at best have barely begun contributing ideas for how to avoid extinction. Any supposed benefit, denominated in percentage points of AGI existential risk averted, is small. Therefore, reprogenetics is too slow; and if you're going [...]
---
Outline:
(00:12) Introduction
(03:37) HIA, part of your nutritionally complete portfolio
(05:52) Against confident short timelines
(08:29) HIA may indirectly slow down AGI capabilities
(09:31) HIA has substantial impact even with short timelines
(16:10) Adult HIA methods aren't fast either, absent big investment
(27:35) Takeaways
---
First published:
August 8th, 2026
Source:
https://www.lesswrong.com/posts/iQzxxgJXXaAQjq7Jz/faq-isn-t-agi-coming-too-soon-for-reprogenetics-to-help
---
Narrated by TYPE III AUDIO. - This sequence is about the last decade in AI alignment. It recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I'll upload over the next few weeks) how fear and (self-)deceptive reasoning made the field one of the biggest forces pushing AI capabilities forward over the last decade, especially via significant contributions to the scaling of LLMs and the development of ChatGPT.
Zooming out further: the two leading AGI companies, which are locked in an intense rivalry, were both explicitly founded under the banner of AI alignment, and got off the ground in significant part due to alignment-oriented ideas, talent and resources. People in the field often sense that something must have gone wrong to get here, but don’t know how to allocate responsibility (aside from blaming Sam Altman and sometimes Elon), and fall back on assuming that “the ship has already sailed”. But in this sequence I characterize our current situation as resulting from a pattern [...]
---
Outline:
(08:17) Conceptual Clarity and Scientific Progress
(20:26) Orienting Towards Prestige
The original text contained 6 footnotes which were omitted from this narration.
---
First published:
August 9th, 2026
Source:
https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment
---
Narrated by TYPE III AUDIO. - “I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man”
–Thomas Jefferson, letter to Benjamin Rush
Context: Conduit is building datasets to enable telepathy, to use their term.
I saw my grandfather lose control over his own fingers: what I would have given to offer him a headband that read his thoughts. Through novel technologies we have liberated almost all Americans from farming, driven the child and infant mortality rate from the pre-industrial half to less than half a percent in the best-performing countries, and rendered famine a political choice: broad-based improvements in efficiency are good and should be pursued for their own sake. Telepathy offers more: we could create trust through verified honesty, helping us ensure prosperity and peace. DARPA is already looking into “preconscious” thoughts for suicide prevention. There's also a strong argument centered on AI Safety: the models are becoming superhuman, and this is technology to allow us to keep pace, minimize hostile competition, and perhaps survive into the future.
This is what Conduit is promising. Unfortunately, mindreading will have other effects.
Oskar Schindler saved over 1,000 Jewish lives during the Holocaust. He did it by [...]
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/CAdG5dzkWrrK2NQg8/don-t-build-mindreading
---
Narrated by TYPE III AUDIO. "OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards" by Zvi
08.08.2026 | 1 Std. 18 Min.How does the situation keep turning out to be worse than we know?
How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?
At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things.
Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky.
If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...]
---
Outline:
(02:38) Cyber Evals Are A Cursed Basin
[... 21 more sections]
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Weitere Gesellschaft und Kultur Podcasts
Trending Gesellschaft und Kultur Podcasts
Über LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Podcast-WebsiteHöre LessWrong (Curated & Popular), Der Sophie Passmann Podcast und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen
Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen


LessWrong (Curated & Popular)
Code scannen,
App laden,
loshören.
App laden,
loshören.
LessWrong (Curated & Popular): Zugehörige Podcasts































