991 Episoden
- Longtime lurker, first-time poster.
I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the essay is what Dawn Song et al. call self-sovereign agents, or AI agents that are independent actors in the world. At the end of the essay, I say that I feel I haven’t spoken about this topic over my 2.5 years of writing with sufficient candor, and that I think this critique applies to others in the AI policy community–particularly the parts of it that tend to manifest themselves in Washington, Sacramento, and Albany–in other words, the parts of the AI safety world that are most involved in hands-on AI policy work. I attribute this primarily to a desire to remain “within the Overton Window,” or to not sound “crazy” within the halls of power, and I assert that others in my profession have made this same calculation.
I believe–and have believed for three years–that self-sovereign AI as I describe it in my essay would likely happen on our current trajectory. That being said, the essay takes pains to distinguish between “self-sovereign” AI and truly “rogue” [...]
---
First published:
September 10th, 2026
Source:
https://www.lesswrong.com/posts/y9TNHfgDwh6vw7Ert/the-locally-optimal-discursive-posture
---
Narrated by TYPE III AUDIO. "Proposal for tracking the effects of architecture on monitorability" by ryan_greenblatt, Alek Westover, Lukas Finnveden
11.09.2026 | 11 Min.Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought or communication (we’ll refer to this property as “monitorability” going forward). As companies begin to explore such architectures, we believe it is important to transparently share evidence about how monitorability varies with architecture and training method. To inform the scientific debate on how to make tradeoffs between performance and monitorability, we believe AI companies should:
Regularly report externally verified information about the degree to which their architectures may allow for latent reasoning or communication. Companies should publicly disclose enough information about architectures to allow external scientists to determine whether they could potentially enable models to perform much more complex reasoning without this reasoning appearing in the chain of thought (“latent reasoning”) or allow for latent communication between different instances of a model.
Following GDM, we propose measuring opaque serial depth as a minimally-invasive proxy for the degree to which an architecture may enable latent reasoning, though companies could provide sufficient architecture transparency in other ways. We propose that companies work with third-party evaluators to produce independently verified reports of the rough distribution of [...]
---
Outline:
(05:57) Appendix: A sketch of what stress tests of CoT monitorability could look like
(06:31) Testing monitorability in control settings
(07:57) Testing monitorability on deployment-time misbehaviors
(08:55) Testing qualitative monitorability on (hopefully realistic) model organisms
The original text contained 13 footnotes which were omitted from this narration.
---
First published:
September 10th, 2026
Source:
https://www.lesswrong.com/posts/hLPGv8QjPcNLtDp3A/proposal-for-tracking-the-effects-of-architecture-on
---
Narrated by TYPE III AUDIO.- I’ve tried various times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good core questions). This post gives the deepest version of that question I’ve found thus far: how should you relate to the parts of the world you can’t directly model or control?
Let me explain further in terms of a distinction between two perspectives. From the third person perspective you think of yourself as “outside” the world, looking in. You’re a good Bayesian, in that you have a set of mutually exclusive collectively exhaustive hypotheses. You choose actions by multiplying your credences by your utilities over those hypotheses, and you treat those actions as the only way you influence the world.
Some problems with the third person perspective (aka Cartesian or dualistic agency) were described in Scott and Abram's sequence on embedded agency. One crucial issue is that most realistic environments contain other agents which are modeling you back, which means that your thoughts might affect the world via channels that aren’t just your actions. Game theory somewhat mitigates this problem, but only in the very specific case where all [...]
---
Outline:
(05:08) Rationality of reward
(09:09) Letters from spirits
(12:21) Languages as Schelling points
(15:43) Actions and entanglements
The original text contained 1 footnote which was omitted from this narration.
---
First published:
September 1st, 2026
Source:
https://www.lesswrong.com/posts/pYFBD2SnqiWkuNns5/explaining-knightianism-on-one-foot
---
Narrated by TYPE III AUDIO. - TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
Epistemic status: Heavily LLM-dependent research, and the precise results are somewhat sensitive to researcher decisions, but I’ve done enough sanity checks that I’d be surprised if the core claims were misleading
One of the most striking things in the Astra report was the massive jump UK AISI found in no-CoT reasoning abilities. I was somewhat suspicious, given the size of the jump, and the many ways this kind of measurement can be misleading. Conveniently, I’ve independently been making my own no CoT reasoning benchmark and tried it on there!
Unfortunately, it replicates. Astra is a massive jump, and disproportionately for no CoT reasoning:
No CoT Reasoning Index (NCRI) vs Epoch Capability Index (ECI) - NCRI represents ability without verbal reasoning, ECI represents overall model capability[2]. 10 NCRI points is a doubling of the odds of solving a problem. Astra represents a significant increase in NCRI, beyond what its overall capability improvements predict, though recent models were also [...]
---
Outline:
(01:39) Executive Summary
[... 9 more sections]
---
First published:
September 9th, 2026
Source:
https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-do-a-concerning-amount-with-no-chain-of-thought
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight.
Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term. I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I’m joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk.
The SSC has an important and challenging role in overseeing risk management at OpenAI, and I hope to help provide expertise and assistance in a critical moment. My joining is not an endorsement or criticism of OpenAI's safety practices in particular; I hope that all frontier companies strengthen safety oversight and I am excited to work on this at OpenAI. I believe that the rest of the world should judge OpenAI, and all AI developers, by externally verifiable behavior and results.
In the rest of this post, I'll explain why I believe loss-of-control risk is now acute [...]
The original text contained 1 footnote which was omitted from this narration.
---
First published:
September 9th, 2026
Source:
https://www.lesswrong.com/posts/82z6FvbYRdjYjqigK/personal-statement-on-joining-the-openai-board
---
Narrated by TYPE III AUDIO.
Weitere Gesellschaft und Kultur Podcasts
Trending Gesellschaft und Kultur Podcasts
Über LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Podcast-WebsiteHöre LessWrong (Curated & Popular), Springerstiefel – Alpha oder Opfer? und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen
Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen


LessWrong (Curated & Popular)
Code scannen,
App laden,
loshören.
App laden,
loshören.
LessWrong (Curated & Popular): Zugehörige Podcasts































