297 Episoden
Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC
16.09.2026 | 1 Std. 26 Min.AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real insurance:
From being Anthropic’s first product hire to building the standards, testing, and insurance infrastructure meant to make frontier AI deployable, Rune Kvist is betting that the biggest constraint on AI adoption won’t be capability it will be trust. In this episode, the AIUC cofounder joins swyx and Vibhu to announce a new $40M round and explain why companies like Cursor, Harvey, Lovable, and ElevenLabs are increasingly confronting a problem that gets harder as AI gets better: who is responsible when autonomous systems fail?
We go deep on AIUC-1, the emerging standard for agent security, safety, and reliability; how AI agents are stress-tested for jailbreaks, hallucinations, and data leaks; and why Rune thinks standards and insurance could become critical infrastructure for AI. We also discuss the growing trust gap between governments and frontier labs, AI-enabled cyber and biological risks, why every model can ultimately be jailbroken, what happens when a $20 coding agent causes $200M of damage, whether AI engineers should be certified, and why even after AGI there may be one job the labs can never do themselves: be their own watchdog.
We discuss:
* Why risk, liability, and trust may become the binding constraint on AI adoption
* Rune’s path from reading the Scaling Laws paper to joining Anthropic in its earliest days
* What Anthropic understood about scaling, compute, and the future years before it became obvious
* Why Waymo illustrates the gap between AI capability and real-world deployment
* AIUC’s $40M round and work with Cursor, Harvey, Lovable, ElevenLabs, and other frontier AI companies
* AIUC-1: a standard for AI agent security, safety, and reliability
* How agents are tested for jailbreaks, hallucinations, and data leakage
* Why most AI companies optimize the happy path without seriously stress-testing adversarial cases
* Why AI standards may need to update every quarter instead of every decade
* The emerging trust gap between frontier AI labs and governments
* Cybersecurity, child safety, biological weapons, and the expanding frontier-model risk surface
* Why standards and insurance may need to evolve together
* How Lloyd’s of London can insure AI systems and bring trust to enterprise deployment
* What happens if a $20 Cursor subscription contributes to a $200M plane crash
* The Air Canada chatbot case and how AI failures are beginning to clarify legal liability
* Why copyright may be one of the hardest AI risks to insure
* Evals, mechanistic interpretability, monitoring, and models becoming aware they’re being tested
* The impossible CISO mandate: adopt AI fast, but don’t let anything go wrong
* Why robotics will make AI liability dramatically more consequential
* Whether AI engineers should have Level 1, 2, and 3 certifications
* AIUC’s roadmap across agents, frontier models, robotics, and universal red teaming
* Why AGI could become a question of national sovereignty
* Why the labs can never fully serve as their own watchdogs
* The Big Short problem: how do you stop competing watchdogs from racing standards to the bottom?
Rune Kvist
* LinkedIn: https://www.linkedin.com/in/runekvist/
* X: https://x.com/RuneKvist
AIUC
* https://aiuc.com
Timestamps
00:00:00 AIUC’s $40M Round and the Risk Bottleneck for AI
00:01:07 From Scaling Laws to Early Anthropic
00:07:58 Why Trust, Not Capability, Could Limit AI Adoption
00:12:19 Founding AIUC and Building AIUC-1
00:18:52 How AI Agents Are Audited and Stress-Tested
00:25:26 Frontier Models, Government, and the AI Trust Gap
00:33:32 Cyber, Child Safety, and AI-Enabled Biological Risk
00:38:14 Why Standards and Insurance Belong Together
00:41:45 What Does an AI Insurance Policy Actually Cover?
00:50:44 The $20 Cursor Subscription and the $200M Plane Crash
00:53:53 AI Liability, Monitoring, and Earning Enterprise Trust
00:56:21 From AI Agents to Models to Robotics
00:58:29 Copyright, Adverse Selection, and AI Insurance
01:03:28 Evals, Mechanistic Interpretability, and Eval Awareness
01:08:36 The Impossible Enterprise AI Mandate
01:11:52 Prediction Markets vs. AI Audits
01:14:43 Should AI Engineers Be Certified?
01:19:10 AIUC’s Roadmap, AGI, and Who Watches the Watchdogs?
Transcript
Introduction: AIUC, the $40M Series A, and Risk as the Adoption Bottleneck
Swyx [00:00:00]: Okay, we’re in the studio with Rune from AIUC, the Artificial Intelligence Underwriting Company, with our trusty co-host, Vibhu. Welcome.
Rune Kvist [00:00:10]: Thank you. Thanks for having me. Thank you.
Swyx [00:00:11]: What are you announcing today?
Rune Kvist [00:00:12]: We have raised $40 million, led by Ribbit Capital and First Harmonic.
Swyx [00:00:17]: You first came to my attention when Nat and Daniel invested in you guys. Is the story, like, pretty much the same? Like, what are you today versus what you thought you were back then?
Rune Kvist [00:00:26]: When we raised our seed round, we had a hypothesis that at some point risk was going to hold down adoption. At that point in time, that felt kind of hypothetical, and I think that is now over. Clearly, the moment is now with Mythos and Fable. It’s pretty obvious that literally the binding constraint on adoption is risk. And so for us, it feels like this is a natural continuation of the same hypothesis, but where previously it was speculation, now it feels like fact.
Swyx [00:00:54]: And let’s get a list of the customers that you’re highlighting as part of your Series A.
Rune Kvist [00:00:58]: Totally. Yeah. So we are now working with folks like Cursor, Harvey, Lovable, ElevenLabs.
Swyx [00:01:05]: Yeah. Amazing. Congrats.
Rune Kvist [00:01:06]: Thank you.
Swyx [00:01:07]: So you were famously one of the first hires involved in GTM and product. I’m just kind of curious: what was your path into AI? Just recap.
Rune’s Path Into AI: Scaling Laws, Capital, and Anthropic
Rune Kvist [00:01:18]: Yeah.
Rune Kvist [00:01:19]: Late 2021, I sold a company, my first company, an edtech company. I had a bit of time to think about what was next. I came across the Scaling Laws paper, and that just struck me like lightning. I was just like, “This is a big idea.” In short, the Scaling Laws paper just says the bigger the model, the smarter the model.
Swyx [00:01:38]: So this is the Kaplan one, not the Chinchilla one?
Rune Kvist [00:01:40]: Exactly, the Kaplan one.
Swyx [00:01:42]: Yeah.
Rune Kvist [00:01:42]: And the important thing that clicked for me there was, oh, now capital will understand this. If you put in more money, you get more money out, and so that will kick off a hype cycle. And so you get a sense of predictable returns, which is, in fact, what’s played out. And so I just packed my bags. I’d never been to San Francisco. I’d never been there. I just packed my bags, flew out here to find the people who had written it. And at the time, they had just started a small lab called Anthropic. There were around 40 people at the time or so. Drank a bunch of coffee until I eventually got introduced to Dario. And at the time, they were wrestling with some of these questions of, like, should we deploy our models? Should we make revenue? How should we engage with the rest of the world? They’d just broken off from OpenAI, and it’s been publicly reported that they were kind of concerned with how they were dealing with deployment. So they were wrestling with some of those questions. At this point, this is early fog of war, like early 2022. The hottest product at the time was, like, Jasper. Like, there’s nothing out there. So where value was going to accrue, and what the different parts of the stack were going to be, were all open questions.
Swyx [00:02:48]: I want to highlight to people, you ask these questions because you have a PPE background.
Rune Kvist [00:02:52]: Yes.
Swyx [00:02:52]: I actually was in Singapore in one of the sort of feeder programs for prepping people for PPE. So I had a tutor. We learned, you know, philosophy and politics and economics. But, like, I think your kind of background matters. Machine learning people who read the neural, Scaling Laws paper would not necessarily draw the same conclusions that you did. Whereas any capitalist would read that and go, “Holy s**t.”
Rune Kvist [00:03:19]: Correct.
Swyx [00:03:20]: Right?
Rune Kvist [00:03:21]: Yes.
Swyx [00:03:21]: Who tipped you onto that paper? Because it’s not a paper that you normally read, right, like, in your circles?
Rune Kvist [00:03:26]: Yeah. I think I’d actually, ever since AlphaGo, had some appreciation that AI was a big deal.
Swyx [00:03:36]: Yeah.
Rune Kvist [00:03:36]: But it kind of felt like it raised all these kind of interesting philosophical questions, but it was kind of not clear from afar where exactly that would go. But it was obvious enough that it was like, this is going to be a big thing if we find the kind of right mechanism to kind of get the techno-capital machine to work on this. But it was just not clear. And so I think there was some way in which, like, that became obvious, and also it wasn’t as obvious at the time than it is now, right? Like, it was just like, wow, this is so interesting. But it still felt, coming from kind of a philosophy and economics background, it felt like if this turns out to be true, you’re going to be wrestling with all of the big questions in society. Everything you’ve learned about politics gets thrown out of the window. Everything you’ve learned about economics at least gets challenged. And so what felt interesting was to be at that frontier that has ramifications across everything. So that’s why I sought it out.
Swyx [00:04:32]: I mean, clearly really good insight. For people who don’t know, the PPE program is, like, where prime ministers are born. So then you end up meeting Dario.
Rune Kvist [00:04:41]: Yep. First Dario, yeah.
Swyx [00:04:43]: Yeah. Well, I mean, like, so did you get extra insights from talking with them that you didn’t get from your original hypothesis?
Anthropic’s Early Conviction and the Scaling Laws Crystal Ball
Rune Kvist [00:04:50]: If you read the Scaling Laws paper, you get this, like, very vague sketch of like, wow, this seems kind of important. There are some lines on a chart. This seems kind of important. And what I think the team at Anthropic had thought more about than anyone was like, what are the implications of this if you really play this out? And back then they had, kind of vision documents for what the world would look like in 2026, and they were kind of in vivid detail playing out how much compute is going to be needed, what the CapEx was going to look like, what some of the societal concerns were going to be, but also what is the amount of economic value coming out here? And so it kind of felt like they held a crystal ball that in hindsight turned out to just be dramatically correct. And they weren’t holding it like they were obviously correct. They were just like, “Take this hypothesis really seriously.”
Swyx [00:05:38]: Think it through, yeah.
Rune Kvist [00:05:38]: And think it through in the same way as the kind of situational awareness that is
Swyx [00:05:43]: Across the street.
Rune Kvist [00:05:44]: Across the street.
Swyx [00:05:44]: Your office, yeah. Oh my God, we’re all living across the street in the same one square mile.
Rune Kvist [00:05:50]: Correct. And that’s now a couple of years old, but also people keep referencing it these particular weeks with Fable and Mythos, and it’s like, wow, if you take this one idea seriously- For the Scaling Laws, a lot of things fall into place.
Vibhu [00:06:03]: And keep in mind, at this point, this is the same team that did GPT-1, GPT-2, and GPT-3.
Rune Kvist [00:06:08]: Correct.
Vibhu [00:06:08]: Which is also, like, it’s not just some experimentation. Like, this is a real model that we just scaled up.
Rune Kvist [00:06:14]: And they had deep conviction in this idea: if you take a big blob of compute and data, it just wants to learn, and out of that will come smarter and smarter models. And all the particulars were not clear.
Vibhu [00:06:26]: Yeah.
Rune Kvist [00:06:27]: And all the implications were not clear. But their deep conviction in this, like, core thesis, and that was kind of dizzying. It was both phenomenally interesting and exciting, and also very quickly you get to, like, the world we know today will no longer be if this hypothesis holds. So it also just felt, like, important in some kind of grand sense.
Vibhu [00:06:48]: What kind of shaped you there? So that was early 2022. Not only had GPT-1, GPT-2, and GPT-3 come out, but, you know, the amazing founders of Anthropic that have never split up, the only ones, they actually had the conviction to leave OpenAI, start their lab. You said there were about 40 people there. What was the time like there?
Inside Early Anthropic: Mission, Deployment, and Risk
Rune Kvist [00:07:06]: It was kind of remarkably like what it looks like on the outside today. Extremely cohesive, extremely mission-oriented, and living in this tension between their two ideas, which is AI could both go really well and really bad, and we want to be part of building it. That creates astounding amounts of tension. And they were wrestling with this incentive challenge where they know they’re in a race that they’re in where you might get forced to cut corners, but it also felt very important to them to be at the forefront of technology. And all of those ideas were just present at that time. It kind of feels like that line has been just very clear, and I think kind of love them or hate them, they have really stuck to their guns. There’s a core set of beliefs that they hold more deeply than most companies hold any beliefs.
Vibhu [00:07:58]: Yeah. Fast-forward to today.
Rune Kvist [00:08:00]: Yeah.
Vibhu [00:08:00]: What does that lead us to AI underwriting company? What are you up to? What motivated you to start this?
From Waymo to AIUC: Confidence Infrastructure for AI
Rune Kvist [00:08:05]: Yeah. AIUC builds confidence infrastructure for frontier AI through standards and insurance. The link from Anthropic to building confidence infrastructure, looking out the windows at Anthropic offices and seeing Waymos driving by. Already back then, early 2022, Waymos were in some ways like AGI for cars. Like, they were superhuman drivers, but you couldn’t take one to the airport. And now, four and a bit years later, you still can’t take your Waymo to the airport, despite now everyone having kind of looked at the evidence and being like, “They’re better drivers than humans.” So in that particular instance, what’s clear is that the binding constraint on AI being useful is not capability, but is that liability or risk or trust. That problem is, general. The reason why right now
Rune Kvist [00:08:52]: Fable is not open for access is not because it’s not a good model, it’s because it’s a very good model. It’s just hard to make promises about what it will or will not do. And this problem gets worse as AI gets better. Basically, more intelligent AI can be more autonomous. That’s more valuable, but also the risk surface grows. And so - what Waymo illustrates is that unless you build the confidence infrastructure to make promises about AI, or at least bring light to the risks, you grind adoption to a halt. Governments, banks, hospitals, militaries need to have some sense of what AI will and will not do to be able to operate for them to incorporate it. And that’s the problem that we’re trying to solve. Now, why standards and insurance? If you trace this problem back through history, every technology wave has had some version of this problem. So if you go back to, like, year 1900, electricity comes
Vibhu [00:09:47]: Ben Franklin.
Rune Kvist [00:09:48]: Cars burn down, sorry, houses burn down, lots of people die. 1930s, cars are a big deal, kill lots of people. 1950s, private nuclear energy is a big deal, poses big risks. In each of those instances, the market runs ahead of regulation to create confidence infrastructure because that’s required to make go/go decisions. That is required for adoption, and the market fundamentally wants adoption. And in all of those instances, common blueprint emerges between standards and insurance. The reason these two components is standards kind of provide the rules of the road, and they also specify, like, what are the tests that need to be run so we can get a sense of how high the risk is. So take in the case of cars, that’s like a car crash. Great, everyone, they inform your insurance pricing today, they inform your purchasing decisions, et cetera. That’s basically the risk framework. The insurers are important because they pick up the bill. So they are the private institution that is most on the side of. That is best incentivized to quantify the risks truthfully and then figure out all the ways to reduce the risk ‘cause that increases their profit. So they’re basically, they help shape the incentives. And these two work really well in unison. Now, how does that show up as a company? Well, one of the things that was obvious even - or starting to become obvious even a couple years ago was that frontier companies, some of our customers today, like Cursor, Sierra, ElevenLabs, Harvey, were going to have a very easy time selling a pilot to a bank. The, like, the demo just sells itself. It’s magic. But bringing that through, if you want to do a wall-to-wall rollout at a bank or a hospital, you have to go through the risk process. These banks have no idea even which questions to ask, let alone which answers are sufficient, let alone, like, how do they go and test whether these agents actually work the way they’re supposed to. And so they had this problem of, like, what can we say to earn the trust? And we think there’s, like, a golden sentence that goes something like, “Hey, I hear you’re really worried about hallucinations or jailbreaks or whatever it may be. We’ve had an independent third party test us against the gold standard. We passed with flying colors. And as a vote of confidence, the world’s most conservative insurers have looked at the data.” And they’re willing to take some of the risk onto their balance sheet.
Swyx [00:12:06]: Yeah.
Rune Kvist [00:12:07]: So if something does go wrong
Swyx [00:12:07]: There’s money behind it, yeah.
Rune Kvist [00:12:09]: Exactly. So that’s kind of like the link between all this. We can get into some of the hard parts related to the technical testing, which is, I think, the crux of the matter, but I’ll pause there.
Swyx [00:12:19]: How did you and Rajiv come together? This-- there’s always, like, you come across very confident and, you know, and we’re announcing your Series A and all these things, but I want to see, like, the early initial stages of, like, idea formation.
Cofounding AIUC with Rajiv Dattani
Rune Kvist [00:12:31]: Yeah. Rajiv is actually my soon-to-be brother-in-law.
Swyx [00:12:35]: Oh.
Rune Kvist [00:12:36]: So I’m actually, in a week and a half getting married to Rajiv’s sister.
Swyx [00:12:42]: Okay, now you’re tight.
Rune Kvist [00:12:44]: Exactly.
Swyx [00:12:44]: Now you know.
Rune Kvist [00:12:45]: So - Rajiv and I have known each other for a decade. Funny story, I met both Rajiv and his sister, Hena, at the same time when Hena and I were interns at McKinsey in London, and Rajiv was assigned as my mentor. And so met them at the same time. For the longest time, it was not obvious that we were necessarily going to work together. I was in startups. He was, an insurance partner at McKinsey. Three or four years ago, I think Hena convinced him that AI was going to be a really big thing. And so he quit his job, cushy partner job at McKinsey in London, packed his bags, flew to San Francisco, and ended up joining METR. You guys are probably online enough
Swyx [00:13:24]: CEO.
Rune Kvist [00:13:24]: Exactly.
Swyx [00:13:24]: We’ve, we’ve, we’ve heard of METR.
Rune Kvist [00:13:25]: You see the plot-- the chart of the horizons of the tasks that agents can take on is doubling extremely fast. So he was COO at METR, led their partnerships with Anthropic and OpenAI to test their models before release, but also working closely with the US and UK government, to figure out, like, how do you know whether a model can be released? And in some ways, that was, like, the perfect background. He’s spent a lot of time in insurance, knows that world, spent a lot of time with frontier testing of models. And so when I was bumbling around this idea space, starting with some of the ideas we talked about related to Waymo, as soon as we got into the content, we were both like, “Oh, this would be an amazing business to build together.” This is wrestling with the problem that we both think is the most important in the world from a market angle, which is kind of our intuitions is that the market can do a lot, and the faster AI moves, the harder it is for government to solve some of these problems. And then it took a little bit of time to work through what is it like to work with family.
Swyx [00:14:27]: Sure.
Rune Kvist [00:14:27]: And,
Swyx [00:14:30]: Because you were already dating at the time
Rune Kvist [00:14:31]: Yeah. Yeah, exactly.
Swyx [00:14:33]: Yeah.
Rune Kvist [00:14:34]: Already back then, it
Swyx [00:14:35]: Yeah.
Rune Kvist [00:14:35]: We felt like we were a family.
Swyx [00:14:36]: Nice.
Rune Kvist [00:14:36]: And so starting a business together felt like kind of a big step. And, here we are with just immense amounts of trust.
Vibhu [00:14:43]: Yeah. So now you’re a company of how big? How big are you guys now?
AIUC-1 Certification: Agent Security, Safety, and Reliability
Rune Kvist [00:14:46]: There are just 20 of us now.
Vibhu [00:14:47]: 20 of you guys now, have Series A, and you have your first certification out, the AIUC-1. Let’s bring up the certification. So this is the agent certification, right? What goes into the process? I have, like, two questions here. One is, walk us through the certification, and two is, what is the process for a company to get certified, you know?
Rune Kvist [00:15:08]: Great. As it says right on the top, AIUC-1 is a standard for agent security, safety, and reliability. The fundamental design principle is take all of the concerns that slow down adoption, so all the questions, all the fears that keep, security leaders in the Fortune 1000 up at night, and put them into one comprehensive framework. That’s what you’ll see there. You can see the six categories. Two, you want to ground all of this in technical testing. So one of the concerns with security standards that often feel kind of like theater paperwork is that they’re not actually ground out in, does any of this work? Does any of this matter? And so we had a conviction from early on that was going to be the kind of crux, was to pass this, you must get tested every quarter, basically run thousands of simulations to see, well, so can it actually be jailbroken? How hard is it to jailbreak? How often does it hallucinate? How often does it leak data? Et cetera. And then the last, core idea here, if you scroll up to the top here, is to refresh it quarterly.
Rune Kvist [00:16:08]: So the core trait of AI is that it moves extremely fast. Whatever concerns we’re discussing today were not the same ones three months ago, and this will keep changing. Typically, standards update on a, like, a decade cycle is obviously not going to work. But the question is kind of how do you update it? And the core thing here was to basically get the risk leaders of the Fortune 1000 around the table. So if you go over to the left here
Vibhu [00:16:32]: Yeah
Rune Kvist [00:16:32]: You’ll see the AIUC-1 consortium. The consortium is a group of risk leaders who run real banks, real hospitals, real critical infrastructure, who are facing these challenges every day. And we meet with these folks twice a quarter and hear what’s top of mind, what is keeping them up at night. There’s tremendous amount of desire for that conversation. And then we operationalize that into a specific standard that gets into. And actually, we can go into and look at what
Vibhu [00:16:55]: Yeah
Rune Kvist [00:16:55]: What even is the standard. So if we go back to introduction, out there to the left, scroll up a little bit to the wheel, click into reliability. So if you take something like hallucinations sits in reliability. There is a number of requirements here. If you go into the top one, prevent hallucinated outputs, hallucinate outputs, this is one particular requirement. This is a technical control. Basically, we want some kind of ground in this filter. The first thing you see here is what’s called a crosswalk. So everyone and their grandmother has put out a framework, very high-level framework for what are the AI risks.
Swyx [00:17:27]: This is basically your competition,
Rune Kvist [00:17:28]: In some ways our competition
Swyx [00:17:29]: Not seriously, yeah.
Rune Kvist [00:17:30]: We’re, in fact, friends with them. We’ll come back to why.
Swyx [00:17:31]: Yeah.
Rune Kvist [00:17:32]: But mapping everything together so you have one superset. The claim you’re trying to support here is, if you follow this framework, then you can also see how you follow the other frameworks. But the meat of it comes down here in control activities and evidence. So control activities is like, great, you have this high-level requirement. How do you turn that down to something operational? Here’s what you must do, and then what is the evidence that we’re looking for?
Rune Kvist [00:17:57]: And the reason we go this deep is that there’s actually not that much confusion about what are the big concerns in AI. Everyone agrees to these. The question, like, what are you actually supposed to do? And so. What we found a lot of demand for is getting down to the specific evidence, that people need to look for. Whether you are Cursor building something or, even JPMorgan building something, but also if you’re just a risk leader at JPMorgan, like what exactly should you ask for? What can you ask for without sounding stupid? Like if you ask for some-- you won’t believe the amount of time a risk leader has asked for the IP rights to the underlying model to Cursor or something, and you’re just like “Sorry, what?” Like,
Swyx [00:18:39]: You slip it in there and you see
Rune Kvist [00:18:40]: Slip
Swyx [00:18:40]: See if you notice.
Rune Kvist [00:18:41]: See if they. Exactly.
Swyx [00:18:42]: Yeah.
Rune Kvist [00:18:42]: Put that in the questionnaire. All right, so that’s kind of what our standard is, and we update this every quarter with these folks, to keep up with the latest concerns.
Swyx [00:18:51]: Can I double-click on this one?
Controls, Evidence, and Third-Party Testing
Rune Kvist [00:18:52]: Yeah.
Swyx [00:18:52]: So first of all, the website’s beautiful. Like, it’s so confidence-inducing which is the whole point where, like, okay, I know exactly what I’m signing up for when I talk with you. Like, I don’t even have to talk to you. I can just see your whole, certification, which is great. But, like, okay, so from here, like D001.1 configure a groundedness filter, how does that get applied? Like, you have a person that
Rune Kvist [00:19:16]: Yeah,
Swyx [00:19:16]: Goes through it?
Rune Kvist [00:19:17]: If you, go back
Vibhu [00:19:19]: I did see somewhere there’s like, you know, fifty-one requirements, a hundred thirty controls. There’s like a whole
Swyx [00:19:25]: Right. I just want to. Like, to me, this doesn’t translate
Vibhu [00:19:27]: Yeah.
Swyx [00:19:27]: Into a test or an eval.
Rune Kvist [00:19:28]: Yes. So if you go into, on the left-hand side. So actually, if - before we go in there are three types of requirements. The first is technical controls, like you must implement some guardrails.
Rune Kvist [00:19:42]: Two, there are test controls. So you must have an independent third party go and run some tests against you. I’ll show you one of those in a second. And then three, there are policy controls. For example, you must have a person whose name is on the line when you guys f**k up, and you must have a plan for how you tell your customers and how you engage with them. They’re kind of more traditional, standard type stuff. So in this particular instance, we just check whether they in fact have a ground in filter. So we will partner with an auditor. So we partner with auditors like KPMG or like Schellman who go in and do the thing auditors do, which is to check the evidence. In this case, that might be a screenshot, it might be part of the code that they need to review to see that it actually. Just that it exists.
Swyx [00:20:21]: Oh, okay.
Rune Kvist [00:20:22]: And then the second thing
Swyx [00:20:22]: So you’re not testing the effectiveness of it.
Rune Kvist [00:20:24]: That’s the second thing. So if you go down
Swyx [00:20:25]: Yeah.
Rune Kvist [00:20:25]: To the third-party testing for hallucinations out on the left, that’s basically the next requirement. This is where we test how well does it actually work.
Swyx [00:20:32]: Okay, and is it you testing or the auditor?
Rune Kvist [00:20:34]: We test them.
Rune Kvist [00:20:35]: We test them.
Swyx [00:20:36]: That’s a lot of work.
Vibhu [00:20:37]: How long does testing take? So if I want to get certified, just
Certification Timelines, Remediation, and Quarterly Updates
Rune Kvist [00:20:40]: Yeah.
Vibhu [00:20:40]: How long does the end roughly take?
Rune Kvist [00:20:42]: Yeah, the end, almost always is dependent on, like, our customers need
Vibhu [00:20:47]: Yeah.
Rune Kvist [00:20:47]: To look something for us. It takes somewhere between, like, 3 to 10 weeks
Swyx [00:20:52]: Yeah.
Rune Kvist [00:20:52]: Depending on how up to snuff they already are. So some people show up to us with, like, extremely rigorous security programs. When we test them, it works extremely well. We can get that done very quick. Some people come to us, and they’re not that far along. We give them kind of the spec that they need to build towards, and then their security teams and engineers get to work and build to meet the standard. The testing itself typically takes a couple of weeks, including the time for them to remediate. Often, we’ll find something that we cannot pass, where this is actually just not up to the standard. - you won’t pass the standard. And then they will need to go and implement additional safeguards or additional remediation that makes them more robust so that they can actually kind of hand on heart look at their customers in the eyes and say, like, “Hey, we’ve done truly our very best.”
Vibhu [00:21:35]: And they’re certified for a year and have quarterly updates?
Rune Kvist [00:21:38]: Correct, yeah.
Vibhu [00:21:39]: And, yeah, it’s pretty interesting. I think, you know, what’s changed since. So this is certifying agents in production, right? Your customers, like you’ve had Lovable, ElevenLabs, Intercom, and they’ve all gone through this certification.
Rune Kvist [00:21:50]: Yes.
Vibhu [00:21:51]: What has changed? So I see you post, like, you know, Q2 added MCP agent,
How Agent Risks Are Changing: Coding, MCP, and Agent-to-Agent Interactions
Rune Kvist [00:21:56]: Yeah.
Vibhu [00:21:56]: agent communication. Any other things that you want to kind of highlight since the first iteration? What comes in quarterly?
Rune Kvist [00:22:03]: Yeah. So some of the changes have just been agents are not just one thing. So, like, if you take agents like Cursor and compare them to Sierra, they’re really quite different. And compare them to Harvey again, compare them to you out of again
Swyx [00:22:16]: ElevenLabs, yeah.
Rune Kvist [00:22:17]: ElevenLabs, they’re all quite different. And so we wanted to design a standard that works for all of the types of agents. And we started with one that was, like, pretty text-based, like, honestly, pretty customer support-focused. That’s where there’s a lot of existing demand. And then over time, we’ve picked, some of the frontier companies in each of these other domains that we could work with and build out the standard, so, such that we know that the same standard works for code, it works for customer support, works for automation, et cetera. So that’s been one big thing. Yeah, then some of the things that have been top of mind recently, Mythos is bringing up a lot of concerns for security leaders. We’re starting to get more and more questions around agent interactions. It’s very nascent, at the moment, but it’s starting to emerge. There’ve been a lot of, questions related to OpenClaw and MCP. Again, like agents starting to interact with each other, is really top of mind. Then as coding agents have really taken off, that’s also where banks and hospitals, et cetera, are getting more and more precise on what it is they need. So really dialing in as that start to be, like, where most of the tokens flow through in the world, getting much sharper on that.
Vibhu [00:23:26]: Can you share for people that are listening that don’t really think about this? Like you mentioned, there’s the obvious stuff, you know, hallucination, citations. What are best practices that people should do when building agents? Like, if they come to you pretty ready with certification like, you know, they’ll probably pass certification. What are the things people don’t think about that they should have?
Best Practices for Agent Builders: Stress Tests and Guardrails
Rune Kvist [00:23:46]: The most important thing is that a lot of companies have not done a serious stress test. They spend most of the time, perhaps rightly so, optimizing for how does it work in the good case, the average case, how high-quality is the output for the customer. And a lot of these companies are pretty new, so they haven’t spent a lot of time stress testing the what is there as an adversary on the other side? What are some of the complicated corner cases that you’ve not really considered? So I think that’s, like, a frame of mind. And you’ll also see this in startups. It often takes a while until they hire their first security person. They- And that’s a whole different kind of risk surface than just building a good product. So a lot of that applies. Most companies actually also have the right kind of architecture. Most of them will have some kind of guardrails in place, either some that come out of the box from their model provider or they’ll have built their own filters that sit in between. They just don’t work very well. The difference between putting a classifier in place that, like, maybe goes and checks whether you’re giving medical advice when you shouldn’t and says, “Hey, if this looks like medical advice, filter it out.” Lots of companies have that in place. The question is whether it works. And it’s actually pretty fiddly to sit down and think about all the ways in which you could ask for medical advice, read the academic literature on what are the kinds of
Rune Kvist [00:25:03]: Framings or tricks you might play to get an AI to give you medical advice when you really shouldn’t. And so there’s, like, an area of expertise that’s just missing. So what we find is that most people have the right building blocks in place. They don’- It doesn’- It’s not rocket science, but the finicky thing is, like, getting into the corners and testing whether it works such that you can look your customers in the eye, or maybe a bank or maybe a hospital and be like, “This is going to work for you.”
Vibhu [00:25:26]: I see. So we talked a lot about the agent-level certification. Where do you guys go from here? So announcing series A camera, we talked about this a bit. There’s the whole security risk of Fable, government stepping in. You guys are kind of announcing that you’re also going into model certification?
Toward Model Certification: The Government–Lab Trust Gap
Rune Kvist [00:25:46]: When we do a bit of cutting afterwards,
Vibhu [00:25:48]: Yeah
Rune Kvist [00:25:48]: We will not yet be announcing this,
Vibhu [00:25:49]: Nice
Rune Kvist [00:25:50]: The question that is top of everyone’s minds now is at the model level. And Mythos, then Fable, has really brought this to the fore that in addition to the commercial risk and the kind of economic security risks that are happening at the agent layer, the models are going to present risk in the national security category. The shape of the problem is very similar. You have some people that are on the hook if something goes wrong. In the case of agents, it’s often security leaders in the enterprise. In this case, it’s the government. They don’- haven’t necessarily spent their entire lives thinking about what are the new risks that come here, what is the kind of data you might be looking for, how might you test that? But they do have to make sure that their concerns are addressed. You have some frontier AI companies that are deeply technical. They know a lot about the risks, but they fundamentally have an incentive to not always be truthful. So you have a trust gap between the government and the labs. And in every other industry, you end up with some kind of body sitting between, a neutral third party sitting between those people. There’s no other industry where you allow people to audit themselves. So there is going to be a need for a third party that can take the rigor of the labs to run frontier technical evals, but can also speak legible trust in the way that the government trusts PwC to go and run financial audits. And they know that they output audit reports in a way that’s consistent, that’s easy to read, that’s factual, that’s, trustworthy. Those two things need to be brought together. And what we’ve learned from our work with agents is that if you want those-- that communication between those two parties to be smooth, there has to be one common standard that is public, that people can go and inspect. What are the risks that matter? Within each of these risks, what are the kinds of threat models that you’re really looking for? You need to specify for each of those risks, what are the guardrails that need to be in place, and what are the tests they need to run to see whether those guardrails are effective? And then you need to go and run audits that are - technical audits that are consistent. So if you’re trying to bring trust, it’s extremely important that you methodically work your way through the risks. You can’t send one researcher in and say, like, “Come back with whatever you find.” You need to be able to explain exactly what you did, exactly what you tried, exactly what you did not try, and therefore the kinds of promises you can and cannot make at the end of it. I think of
Neutral Third Parties, CAISI, and Model Risk Audits
Rune Kvist [00:28:13]: Fable as a direct symptom of this problem that the government was told that there’s a risk. The government may struggle to assess just how big that risk is. They call Anthropic, and Anthropic is trying to tell them, “Hey, actually, every model can be jailbroken.”
Swyx [00:28:28]: That’s not what you want to hear, right?
Rune Kvist [00:28:32]: As the government, that might be hard to trust.
Rune Kvist [00:28:36]: And we think that a broker is the most natural solution. In other markets, you see something like, in financial markets, you see Moody’s. Moody’s goes in, and they look at a bond, and they output a rating. They say like, “Here’s the evidence we found. Here’s the rating.” We don’t decide whether anyone should buy this bond or not buy this bond. Well, that depends on their risk appetite. But we do provide this common information layer that everyone can rely on. In the case of Moody’s, the government, points to them and say, “Hey, pension funds, you should probably really take care. You shouldn’t risk your pensioners’ money, so you can only invest in triple-A rated bonds.” That means that now the government doesn’t have to staff thousands of financial technical experts to rerun forecasts every week to see whether things are correctly rated. They get to point to some neutral third party. So my hypothesis is, my hunch is that you will see a third party that sits between the government and the labs, and it could either be the government builds it themselves. So something like CAISI was set up to do exactly this. And the question
Swyx [00:29:44]: Sorry, I’m not familiar with CAISI.
Rune Kvist [00:29:45]: CAISI is the Center for AI Standards and Innovation.
Swyx [00:29:49]: Okay.
Rune Kvist [00:29:50]: I won’t get into the details, but it’s a body of NIST that typically sets standards. So it’s basically a government body that has AI experts. Yeah, exactly. Exactly.
Swyx [00:29:59]: Very key. Very key.
Rune Kvist [00:30:00]: Very key.
Vibhu [00:30:00]: I think, you know, it’s one of those things where when you just sit back and listen-- look at it, like, is there enough technical expertise in the government to measure, test these things right now? Probably not, right? And Fable is a result of, okay, we’ve had to scale back and pause things,
Rune Kvist [00:30:17]: Yeah. And they have excellent people, but they have an extraordinarily small budget compared to the scale of the challenge that’s ahead of us. And I think they have a role to play. The question is kind of like, who does what? We have now outlined the jobs to be done, and they’re quite extensive. Every model release, there is an astounding-- Given that they take in any input, their risk surface is astounding. And so the question is really: what can only the government do, and what can the market provide here that can keep up with the pace as AI risk changes? Our perspective is that also at the model layer, the risks that people care about today are not the same ones they cared about three months ago. So the pace of legislation is too slow to deal with pinpointing the risks here. And so we think there’s a lot that the market can do to surface timely information. Ultimately, there is a bunch of policy decisions here. Is the national security risks of a model too high?
Swyx [00:31:12]: Yeah.
Rune Kvist [00:31:12]: That’s a political answer. But what we want to make sure is that the process that produces this risk information is compatible with very fast innovation. So you don’t want to. This is not a question of like, can you slow the things down? Can you keep, the models locked up until-- for months on end until everyone can make a guarantee? But it is this, can you, in the time it. Given that the US is competing with China on releasing models, can you insert risk information that allows the government to, like, make rapid decisions on some of these questions? Balancing that trade-off between failing to adopt AI is going to put us at risk, but also reckless adoption is going to put us at risk. And that’s a very kind of fine balance that they’re going to need, like, a lot of high-quality intelligence to make.
Chinese Models, Data Flows, and National Security Concerns
Swyx [00:31:55]: Just a side mention, because you mentioned Chinese models, any specific concerns that you’re hearing from your CISOs about that? ‘cause I guess it’s free, but.
Rune Kvist [00:32:05]: CISOs have a bunch of concerns around data flows in general that they’re really concerned about. So there’s a lot of questions like, if these models are Chinese, where does that, where does that data go? I think a lot of this can be addressed, but they come up often.
Swyx [00:32:18]: I mean, they understand they’re running on American GPUs.
Rune Kvist [00:32:21]: Some of them, some of them understand that they’re running on American GPUs.
Swyx [00:32:23]: They’re not, like, phoning home every time you, like, call home.
Rune Kvist [00:32:26]: No. A year ago, there was not a lot of understanding of this. I actually think, you’re seeing the security leaders becoming kind of AI literate at a blistering pace, and you’re actually also seeing my Twitter timeline that’s very pilled and my LinkedIn feed that used to not at all be pilled kind of converge. They’re both talking about Fable.
Swyx [00:32:45]: Right. Yeah, that’s true.
Rune Kvist [00:32:46]: They are both talking about whether you can prevent models from being jailbroken these days.
Swyx [00:32:51]: Yeah.
Rune Kvist [00:32:52]: Like national security national security risks are now the conversation that is actually emerging. Other than that, I think you mostly see a kind of general picture: there are no concerns with any particular model or any particular model output, but there is a general nervousness of having critical infrastructure run on models that are not produced in America by Americans where the American government has control.
Swyx [00:33:14]: But it doesn’t necessarily show up in your framework that directly, or it might, I don’t know.
Rune Kvist [00:33:18]: There’s a bit of stuff in there actually on the, like, the provenance of the models and disclosing that. But I think there’s a bunch of use cases where running a Chinese open-source model is just the best solution.
Swyx [00:33:27]: Yeah.
Rune Kvist [00:33:27]: And a concern is slightly more macro here, which is not best addressed at any particular certification level.
Vibhu [00:33:32]: Is there anything interesting that you see at the. You know, if you’re trying to fill that middle gap, that mediation gap, any interesting stuff that you guys forecast would be required other than, you know, what the average person might expect?
Cyber, Child Safety, Bio Risk, and Expert Coordination
Rune Kvist [00:33:47]: There’s a bunch of interesting questions about what are the risks that matter here. So right now, the risk of the day is cyber, because it’s very real, very tangible. And some of the risks that are also emerging as pretty real and pretty tangible are things like child safety is becoming both extremely important, but also politically important. And then there are some of the risks that are coming down the pipeline that today feel kind of speculative, but people who spend a lot of time with the models see them coming down is things like, risks that relate to biology.
Rune Kvist [00:34:18]: And specifically whether models will help adversaries produce biological weapons and making that extremely cheap, extremely accessible, producing-- making the chance of another COVID or worse pandemic. COVID was not engineered to be bad, as if you were trying to do that. So I think those are some of the risks that are coming down the pipeline. I think one other thing to just note is that agents are kind of deliberately narrow. So, like, when a frontier agent company puts a chatbot that interacts with customers, they’ve really tried to narrow the topics it’s interested in talking about. Such that if you ask it, like, “What do you think of the president?” it will just decline, which means that the kind of risk area is somewhat smaller. For models, it is infinite. And so there’s not a single expert out there who can competently evaluate the risks of cyberattacks and fifteen-year-olds having month-long conversations with a chatbot and seeing whether it will in fact recommend suicide or something horrendous like that, and can evaluate the risks that terrorists can use AI to produce bioweapons. The risk surface is just too big. And so the central challenge actually becomes how do you get those subject matter experts to work within a one coherent framework that outputs one coherent report and rating that the world can go and inspect? ‘Cause that global perspective is central, but there’s not a single organization today that could produce that.
Swyx [00:35:47]: And you would be the presumptive one when you put out your model standards.
Rune Kvist [00:35:51]: We think there can be one company that can, with a consortium of experts, build one coherent standard. I think we’ve shown that across all of the enterprise risks today. We think it could be one company that could, with a consortium, specify the audit rules, basically like the inputs and outputs that all these technical experts need. What access do they need? How should they treat infosec- info security? They can look at whether the eval- evals are well-produced without necessarily being able to say, “Hey, is this a threat or not a threat?” But overall, evaluating whether the evals are good, well-constructed, that set of audit rules that basically becomes the interface for all these experts, we think one clearinghouse could put together. To be clear. When I say one company, I think of it as one company coordinating lots of this in the same way that when we saw our consortium, it’s not like we say we have all the answers on agent security. What we say is we are taking on the role of eliciting all of the concerns and being the secretary that puts it together and runs a tight house such that the standard updates lockstep every quarter, and that the audit reports that come out, in this case, 100-page audit reports, uniform and crisp and clear all to the level of detail that is required for executives that need to make a clear go/go decision. So that’s kind of the role that we think we might play.
OWASP, Frameworks, and the Operational Audit Layer
Swyx [00:37:11]: I think in many ways you’re performing the role that OWASP used to do there, and you said, like, you know, competition and partners.
Rune Kvist [00:37:18]: Yeah.
Swyx [00:37:19]: Can you go more into, like, how they partner?
Rune Kvist [00:37:20]: Yeah. So first of all, OWASP is basically an open source community of security practitioners that are coming together to build frameworks for addressing the latest security concerns. We think they are phenomenal at creating frameworks. We’- In fact, we’- First of all, we’re partners with them, so we have a joint article. Two, we’ve learned a lot from them. We think they’re a tremendous source of intelligence. What OWASP does not do is building the machine that runs third-party audits such that a company like Cursor or a company like JPMorgan could get a third party to go and review them against this and say, “Hey, you’ve passed the standard, and here is the report that you can use to build trust and preempt your partners’ or customers’ questions.” So they fundamentally try to do something different. You - They are part of the information gathering and intelligence gathering and creating clarity, but the operational layer of turning this into promises is not the business they try to be in.
Swyx [00:38:14]: The standard is emerging and is doing very well. Was it necessary to then also do underwriting? Obviously it’s in the name, so please remember you thought about it first. I feel like if you just have enough consensus, you don’t actually need the money angle, but it does help.
Vibhu [00:38:30]: I did want to also note, you guys are a profit company too, right? It’s not profit where there’s a whole business side to it as well?
Why For-Profit Standards and Insurers Matter
Rune Kvist [00:38:39]: Yeah. Yeah, so I’m just getting crazy
Swyx [00:38:41]: I think about the money part.
Rune Kvist [00:38:42]: Yeah. Yeah, let’s get into the money part. Let’s start from actually your question, profit versus profit. In the security space today, cybersecurity, most of the standards are produced by nonprofits. I think that’s an issue.
Rune Kvist [00:39:00]: The question you have to ask yourself is, how do you create good incentives for these standards to be good and keep up?
Rune Kvist [00:39:09]: Nonprofits tend to not have these adverse profit incentives where they, hollow out their standard and create a race to the bottom, but they’re also not at all responsive by default to the communities that they serve. There’s no process-- They don’t have customers that they serve where they go and ask, “What do you want? What do you want? What do you want?” And when you look at the overall satisfaction with the security standards today, people tend to just not like them very much. You do see in other domains, that profit standards can serve the world quite well. So there are examples, like we talked about Moody’s before. It’s not without flaws, but, it is absolutely critical societal infrastructure that gets run at an astounding scale today. Your credit score, it’s FICO. It’s also a profit business. And when you go back even further in history, some of the crash testing standards came out of insurance companies.
Rune Kvist [00:40:06]: The insurance companies together founded the Insurance Institute for Highway Safety because they were very interested in, like, how can we use standards to drive down mortality and save money? Go back, prior-- Our name actually pays homage to the Underwriters Laboratories, UL, which, was started right around when electricity came out. Houses started burning down. Insurers, again, were paying the bill, and they were maybe also good people, but their profit incentive was, let’s prevent houses from burning down. Let’s test all the electrical products, the light bulbs. All the light bulbs in here are probably tested, the toasters, et cetera. And they set up, an entity to create those standards. Today, UL has a profit entity and a profit entity. What they’ve recognized, they spun - They started profit. They spun out a profit because what they recognized was like, hey, actually to serve customers well, you need a profit entity. The lesson here is one of the ways that the market can align incentives so you’re both responsive to customers
Rune Kvist [00:41:07]: And not hollowing out your standard over time is to align it with insurers because they fundamentally have good incentives. And so if you’re a profit standard that works closely with insurers, you get the feedback loop in such that you’re really tuned into your customers, but also have their interest at heart. So that’s the model that we - the kind of inspirational model that we’ve learned a lot from, and that’s also where the name comes from. In some ways, the term underwriting can both be associated with insurance, but it’s also a broad term for, like, making decisions.
Rune Kvist [00:41:40]: If you underwrite a decision, you’re fundamentally kind of taking ownership for the consequences of it.
AI Insurance Contracts, Lloyd’s of London, and ElevenLabs
Swyx [00:41:45]: Yeah, I mean, what does an insurance contract look like for AI?
Rune Kvist [00:41:49]: Yeah. Most of the demand comes today for insurance contracts is, sitting between people who’ve built AI and people who are buying AI.
Swyx [00:41:56]: Yes.
Rune Kvist [00:41:57]: And what you want—the reason why people want insurers involved, both for the traditional reasons, hey, if something goes wrong, we want to be compensated, but it’s in particular because insurers can bring trust to the equation. Because insurers will pay for the damages, if they’re willing to write an insurance policy, that is them saying, “Hey, we think there is risk here, but that is manageable.” And that is kind of a. Their incentive aligns with the enterprises adopting it, so that’s a really a good signal to the market. In the same way, actually, one of the things that Waymo tried to get their first permit to even operate in San Francisco was to get a lot of insurers to stack up a huge insurance policy. In the case if something went wrong, not because Google can’t pay, but because it was very valuable to have a third party go and look at that data
Rune Kvist [00:42:47]: That are trusted by governments, trusted by enterprises as conservative people and say, “Hey, we’ve looked at it. We’re actually willing to take some of this on our balance sheet.” So that’s, that’s kind of the reason why people are interested in it. What it looks like is, in some ways like every other insurance contract. You specify what are the perils you want to cover, how much do you want to cover them, like up to what limits, and what does it cost to cover that. And in the case of, if we take a really concrete example, ElevenLabs, bought a first of its kind AI agent insurance policy. They work with some of the biggest, enterprises that work with governments. They’re really interested in going above and beyond and making promises to their customers. So they wrote a policy that covers just some of the core concerns that their customers have been asking about. And, the crucial thing was really to get Lloyd’s of London, the world’s oldest insurer, one of our partners, to look at this data and be that third party alongside us to say, “Hey, we think there’s something here that’s worth underwriting.” and that’s actually what it looks like. And so they will show that contract to their customers, and they can see how much they’re covered for. They can see what exactly it covers, and that will also probably change next year. They will want to write an insurance policy that might cover more.
Swyx [00:44:04]: When you say Lloyd’s, is it reinsurance, or are they sharing somehow at the same level or
Rune Kvist [00:44:11]: Yeah. So typically, the way, new companies get into insurance is that they partner with insurers such that the insurers take the majority or all of the financial risks. Fundamentally, if insurance is useful, because it brings trust, you have to be able to pay the bill. Lloyd’s of London is 400 years old. They’ve never not paid a claim. They’re extremely trusted. What Lloyd’s of London struggle to do on their own is to figure out which of the risks are real, what should we be looking for, what are the kinds of technical controls, and running the tests. So they use AIUC-1 as kind of the underwriting framework, and we produce a bunch of eval results that then directly feed in to inform the pricing. So this means that ElevenLabs customers know that payment will be there. They don’t have to look to our series A and see, like, do we think they have enough cash on the balance sheet? They will look at Lloyd’s.
Swyx [00:45:05]: Yeah.
Rune Kvist [00:45:05]: Yeah.
Swyx [00:45:05]: And Lloyd’s, like, famously very creative. I think I remember some headline like, they insured Jennifer Lopez’s, butt or something.
Rune Kvist [00:45:13]: Correct.
Swyx [00:45:13]: Right?
Rune Kvist [00:45:13]: And I think, was it, David Beckham’s right foot?
Swyx [00:45:16]: So, yeah. Right?
Rune Kvist [00:45:17]: And stuff like this.
Swyx [00:45:18]: So, like, clearly not a large data set.
Rune Kvist [00:45:22]: Exactly. It’s actually a remarkable institution that’s both kind of has some of the truly school virtues of having been around for a long time. They, like, really. They really operate like a trusted entity, and they have appetite to figure out the future. And I think there’s a lot of recognition that both there is, like, tremendous amount of risk in AI that is poorly understood today, so getting into this business carries real risks. But also this is where lots of the risk exposure will happen in the future. This is the one market where risk is truly growing. This is the one market that will also take out some of the existing markets. Take, like, auto insurance. When there are no human drivers, how’s that market going to look? Well, it’s clearly going to change. How are you going to assess
Swyx [00:46:08]: You want to insure Waymo?
Rune Kvist [00:46:10]: I. All I’ll say is the principles for how you insure Waymo are very similar to how you insure other kinds of AI.
Swyx [00:46:15]: Right.
Rune Kvist [00:46:15]: So again, crash testing, that’s what we do for customer share at Lovable. That will also need to happen for Waymo, which is not how you do it for human drivers. So there’s this growing awareness that the world is changing very fast, and the only way to learn how to underwrite AI is to write some policies. You may incur some losses and think of that as R&D expense, really. But the question for them is, like, who are the trustedtechnical partners they can get into this business with that can help them navigate and make sure they don’t make, kind of foolish mistakes? But also who is willing to hear the wisdom that they have? They’ve done this before. They’ve seen it was. They were there when cyber came out. So there are lots of ways in which AI feels completely new, but there’s also lots of ways in which risks look the same. And so there’s actually a tremendous amount of wisdom sitting in some folks that may have gray hair, but really have, like, a keen sense of, how to quantify risk.
Swyx [00:47:08]: Yeah. And the number is. So it’s basically like I want fifty million dollars worth of coverage against these perils, and Lloyd’s will give you a quote on it, and then you have, like, a small markup or something, and then you turn it around and do that? Is that as simple as it is?
Risk Capital, Premiums, and Working with Insurers
Rune Kvist [00:47:23]: You basically share some of that premium.
Swyx [00:47:25]: Yeah.
Rune Kvist [00:47:25]: X percent goes to the people who do the pricing of it.
Swyx [00:47:28]: You’re. It’s kind of like a. It’s kind of like a merchant bank for insurance type of thing.
Rune Kvist [00:47:33]: Exactly. You basically split the fee, and you can think of the insurance supply chain as, like, there’s bringing the capital, there is doing the pricing, and there is doing the distribution. And typically, you will pay out some X percent of premium here, Y percent of premium here, and the rest of it will go here.
Swyx [00:47:46]: Does all the insurance world work like this, or is there some point at which, like. So if right now you have equity capital
Rune Kvist [00:47:51]: Yeah.
Swyx [00:47:52]: At some point, maybe you start raising, debt or whatever, and then you have enough of a bank account and enough history, let’s say you’ve been in operation for ten years
Rune Kvist [00:48:00]: Correct.
Swyx [00:48:00]: That you don’t need Lloyd’s anymore?
Rune Kvist [00:48:02]: That’s totally an option. And I could see some worlds where that makes sense, specifically if there are risks that we feel high confidence that we’d want to insure where the incumbent insurers are too slow to find appetite
Swyx [00:48:13]: Okay.
Rune Kvist [00:48:13]: Or simply struggle to evaluate it such that they don’t want to do it. But by and large, in general, you do not want to compete with insurers on, bringing risk capital to the game for two reasons. One is that’s fundamentally a cost of capital game. They have extremely low cost of capital. Startups have high cost of capital, by and large. And two, you want to hedge your bets, and it’s very helpful then to also have a portfolio of home insurance, of car insurance. And we’re not about to become a car insurer nor a home insurer.
Rune Kvist [00:48:43]: So they have some natural advantages, which makes it much more likely that we’ll partner.
Swyx [00:48:48]: Yeah.
Rune Kvist [00:48:48]: And they bring that, the capital at scale, and we bring the technical expertise.
Swyx [00:48:51]: You’re, you’re going to work with them for a long time.
Vibhu [00:48:52]: How are the discussions with the insurers as well? So basically, they’re going off of your certification, right? They’re trusting the diligence on you that your certification is valid, you tested the right things, and they’re backing the money that, you know, you have the right testing in place. So any interesting takeaways from working with insurers?
Rune Kvist [00:49:12]: I think the maybe the first thing is they feed into the standard as well. So if there are things that they feel like they need that they’re not seeing, we are also taking that as input into the standard, because fundamentally we think a good standard is one that creates a really healthy promise ecosystem, and we think insurers are a critical part of that. And again, they are the most well-incentivized to. They see all the lost data across every. Any particular CISO knows their particular concerns. Insurers see the concerns across the entire portfolio and often have direct access to, like, what exactly happened, who was at fault, et cetera, as they do part of their forensics. So they’re actually, like, a great source of intelligence on this. One of the big takeaways from cyber insurance, which is a market that didn’t work that well, was that the insurance and the technical expertise was not married up. What our conviction is that standards have to precede insurance. Fundamentally, what everyone first and foremost want, whether you’re a CISO at JPMorgan or a CISO at Cursor or an underwriter at Lloyd’s of London syndicate, is you want to not have an incident
Rune Kvist [00:50:19]: In the first place. You want to know that the risk is well-managed, and only then does insurance start to make sense. So we’ll see the standard ecosystem basically run ahead of the insurance. And the reason why we. You asked us kind of why I also do insurance, this is kind of proving what we think a whole promise confidence infrastructure ecosystem needs to look like, and we think it’s very compelling to bring that to life, even if we think the standard is kind of the core linchpin that unlocks the rest.
Claims, Liability, Air Canada, and Duty of Care
Swyx [00:50:44]: There’s been no claims yet, right?
Rune Kvist [00:50:45]: Nope.
Swyx [00:50:46]: This is one of those things where, you know, if people haven’t really worked through what it means to cover things.
Rune Kvist [00:50:52]: Yeah.
Swyx [00:50:52]: So for example, I pay Cursor $20 a month.
Rune Kvist [00:50:55]: Yep.
Swyx [00:50:56]: And I write a vibe code something that makes, a plane crash, causing $200 million worth of damage.
Rune Kvist [00:51:02]: Yes.
Swyx [00:51:02]: Do I claim $20 or do I claim two hundred million?
Rune Kvist [00:51:07]: Yeah. So these are all great questions.
Rune Kvist [00:51:10]: And fortunately, kind of all of insurance and legal history kind of helps answer some of those questions. I think the first thing is people have limits on their policy. So if you want to claim $200 million, you have to. Someone has to have paid a lot for that insurance policy upfront to have $200 million of coverage. And ultimately, the way this works is that, you start from a lot of uncertainty. This is not just an insurance, but also, like, can you use. Can Anthropic use books on the internet to train up? Well, they can go and look at precedent, they can But ultimately, this- these things get settled in court, and you hammer it out over time. So you start from this, like, place of ambiguity, which is both why insurance can be hard to do early on, but it’s also why people want insurance, because that ambiguity slows down adoption.
Swyx [00:51:57]: Yeah.
Rune Kvist [00:51:57]: That also sits at the heads of the,
Swyx [00:51:59]: Yeah. In some ways, actually, the first incident will help to, establish a lot of this.
Rune Kvist [00:52:05]: Exactly. And there have been a number of incidents out there that have just not been covered by insurance.
Swyx [00:52:09]: Yes.
Rune Kvist [00:52:09]: Take the now old, example from Air Canada, where
Swyx [00:52:14]: I was going to bring that up
Rune Kvist [00:52:15]: Chatbot hallucinated a refund policy, and the question was, Air Canada in that case were like, “Hey, we have nothing to do with this. This chatbot messed up, but, like, sorry.” And the courts were like, “No, if you put your chatbots to interact with your customers, they make legally binding promises on your behalf.” That is now precedent for everything in the future where you will. If someone were to deploy a chatbot like that again, they should not expect to be able to just pawn off and say, “Sorry, my chatbot lied. It’s nothing to do with me. I bought it from OpenAI.” No, if you’re putting this in front of your customers, you are taking responsibility for it. And so every court case, whether insurance is involved or not, clarifies liability, and liability is kind of the foundation for insurance. There’s another reason why standards and insurance come together. Liability for. I’ll go on a little tangent here
Swyx [00:53:06]: Please
Rune Kvist [00:53:06]: Get into the weeds of it.
Swyx [00:53:06]: Please.
Rune Kvist [00:53:07]: Liability, often one of the core concepts is whether someone was negligent. Should they have seen this? Should they have prevented this? And the question is: how do you judge that? Well, you basically judge whether they’ve met their duty of care. What does that mean in practice? Well, often they look to standards. So if there’s a standard that is broadly adopted that says you must have a groundedness filter or you must have a jailbreak filter, it becomes way harder to claim ignorance that these things existed. And so setting standards help clarify liability. Coins-- courts will often point to standards and being like, “Well, this seems like best practice to do.” It’s there for everyone to see. So there’s another way in which, like, standards are kind of civilization infrastructure that insurance can then build on, which promises can then build on.
Swyx [00:53:53]: I totally get that. We don’t have to get certified to write these, to, you know, make these, like, bots and all these.
Rune Kvist [00:54:00]: Correct.
Swyx [00:54:00]: But, like, basically, whenever we get. Go for the audit, I think people, like, start to shape up and all this stuff. I wonder if, like, that means that you don’t also then become, like, the approving authority for me to ship to production. You know, like, yes, you check once per quarter. I want to ship once a day.
Shipping to Production: Ongoing Testing and Trust
Rune Kvist [00:54:19]: Yeah.
Swyx [00:54:19]: And I don’t know when one of my things breaks, like one of your certifications or not.
Rune Kvist [00:54:24]: So there’s a couple things. There’s a couple of requirements in there that relate to how do you yourself, where you have to tell your customers
Swyx [00:54:33]: It’s like an ongoing monitoring.
Rune Kvist [00:54:34]: How are you yourself testing before you make at least major releases? We don’t go and audit people every day, but at least there is now a trail where if you do a major mess up, then your customer may come and ask you, “Hey, you promised me that you were going to run these evals yourself.” And for lots of them, most of the. PRs that people merge will not fundamentally alter the product experience, but some of them will. Thank you.
Swyx [00:54:57]: And sometimes you don’t know.
Rune Kvist [00:54:58]: And sometimes you don’t know. There are inherent risks that everyone knows that when they buy software, there can be bugs, and this is just part of it. But if you’re selling to mom-and-pop shops, they may not care. They’re just like, “Well, I want to use your tool, so I’m just going to willing-- be willing to take that risk on.” If you’re selling to a big bank, they might be like, “Sorry, we’re making promises to our customers. If you can’t make a promise to us that we can pass on, we don’t want to work with you.” Then it’s up to you to say, “Do I care for my agent to get used as critical infrastructure in this mission? If so, at least I can make promises about what processes I run, and then we can go and test it every quarter to be like, well, does it seem like, it’s still, that it still meets the standard.” So from my perspective, it’s kind of a way to. Big companies by default kind of have some amount of trust when they ship AI.
Rune Kvist [00:55:49]: If you’re a young company, if you’re just starting out, by default you have no trust. And there are very few places where you can go and get trust. So one of the things that most of our customers did before they started working with us is that they would make their own security blog posts. That’s great. But also, who’s going to trust you saying, “We’re so secure”?
Rune Kvist [00:56:05]: Like, anyone can write that. But it’s very hard. Where do you go and get that trust?
Vibhu [00:56:08]: Yeah.
Rune Kvist [00:56:08]: And so I think making the standards more legible makes it easier for smaller companies to prove that they’re doing what they ought to be doing, because the default assumption is that it’s the Wild West.
Vibhu [00:56:21]: Is there a roadmap you have of, like. There’s a lot of work to be done here, right?
Rune Kvist [00:56:25]: Yep.
Vibhu [00:56:25]: This is the first one.
Rune Kvist [00:56:26]: Yeah.
Vibhu [00:56:26]: Anything on the roadmap of what you see is next, what’s coming, what’s, what’s missing?
The Roadmap: Agents, Models, Robotics, and World Models
Rune Kvist [00:56:32]: I think when we zoom out, AIUC-1 deals with agents. Next up, we will deal with models. Next up from that, we will deal with robotics, of which, in some ways, Waymo is the first robot. But the exact same problem is going to be someone’s going to develop a robot, someone’s going to need some promises, they’re going to struggle to make the promises. And - You see this playing out when, like, if you think Fable concerns are bad, like, see when Waymo hits a dog. And that’s if people lose their mind. Imagine when first robot knocks off a toddler off a kitchen table.
Swyx [00:57:03]: Yeah.
Rune Kvist [00:57:03]: You’re going to see some real strict liability.
Vibhu [00:57:07]: I mean, you could see it, right? Like, Cruise got fully
Rune Kvist [00:57:10]: Destroyed.
Vibhu [00:57:10]: All permits are gone, yeah. Yeah.
Rune Kvist [00:57:12]: Correct. So physical AI, the level of stringency just goes up and up. So that’s kind of like the big picture. Agents, models, robotics. I think within agents, the current set of agents are well-covered by this. But as the technology progresses, as agents get longer horizons, new types of failure modes will emerge. And so it’s mostly of can you make sure the standard keeps up when they appear? And you also start to see new modalities. Like today, world models are mostly a kind of a research question. There’s no one who’s really using it. But that will also bring in just new kinds of ways to create value, but also more risk surface that no one knows how to grapple with today. You’ll start to see true agent interactions that are not mediated by humans. There’s going to be a bunch of interesting questions. You’re basically going to need a new legal system. How do they build trust amongst each other? How. One of the core things when humans trade with each other is that you know that you have recourse. You can sue them. How do you make sure that there is a persistent balance sheet behind any agent such that if you trade with it and it screws you know you can get your money back? Those are some of the questions we’re going to have to deal with. And the technical testing
Rune Kvist [00:58:24]: Of multi-agent systems is also going to be interesting and complex.
Swyx [00:58:29]: Very fun. Are there any perils that are uninsurable right now that people wish that you would?
Copyright Risk, Adverse Selection, and Information Asymmetry
Rune Kvist [00:58:35]: Yeah. One of the places where there’s a bunch of appetite for insurance and not a lot - a lot of demand, but not a lot of supply, is when it comes to copyright.
Swyx [00:58:46]: Oof.
Rune Kvist [00:58:47]: In some ways, copyright is kind of mundane. It’s always been an issue. There’s a couple of reasons for this. The first is people who have trained on copyrighted materials almost always know that they’ve done that.
Rune Kvist [00:58:59]: So if you want to buy insurance for it probably signals that you might be a high-risk customer. The people who are most interested in getting insurance for copyright infringement
Swyx [00:59:09]: Okay. Yeah
Rune Kvist [00:59:09]: Are the people who are most likely to have copyrighted
Swyx [00:59:10]: Yeah. It’s like a, it’s like a lemon problem.
Rune Kvist [00:59:13]: Exactly.
Vibhu [00:59:13]: I actually think there’s another side to it too, right? Like, if you’re building on something. So say I’m using an open model.
Rune Kvist [00:59:19]: Yeah.
Vibhu [00:59:19]: I don’t know what it’s trained on, right?
Rune Kvist [00:59:21]: Yes.
Vibhu [00:59:21]: And how far down that chain does copyright go?
Rune Kvist [00:59:24]: Yes.
Vibhu [00:59:24]: Am I liable to take down my product because company X trained on copyright?
Swyx [00:59:29]: But there’s safety in numbers. If everyone’s doing it, then you.
Rune Kvist [00:59:33]: Correct.
Vibhu [00:59:33]: I mean, I would say until, you know, Fable is rolled back from everyone that used it, right?
Rune Kvist [00:59:38]: Yeah. I think it’s a hard question. I don’t know the answer to it.
Vibhu [00:59:39]: It is.
Rune Kvist [00:59:39]: But I think your intuition is, your intuition is right in kind of like, what is the kind of duty of care?
Rune Kvist [00:59:47]: And people don’t today think of it as customary that you go and you, like, dissect the open model’s training data and you check everything. In fact, lots of people use them. It’s seen as kind of generally acceptable to not check for this. And therefore, like, we’re not going to hold you to specific
Vibhu [01:00:02]: I mean, we also really can’t, right? We don’
Rune Kvist [01:00:04]: Exactly.
Vibhu [01:00:04]: We don’t know the training data.
Rune Kvist [01:00:05]: So you can then ban it, but I think no court is going to get a copyright question and be like, “This actually needs to get banned.”
Swyx [01:00:09]: Unless you hire Nicholas Carlini and he can extract it for you.
Rune Kvist [01:00:12]: Exactly. Though he’s in short supply.
Swyx [01:00:15]: Yeah. He’- You only have so many Carlinis, but,
Rune Kvist [01:00:17]: Exactly.
Swyx [01:00:18]: Yeah, go ahead.
Rune Kvist [01:00:19]: So I think this is also fair that, in the case of labs, there’s a lot of interest for this. But the thing that makes lab want it is what makes this insurer suspicious of it, and so you have a lemon’s problem.
Swyx [01:00:30]: Yeah. Is there, like, a theory of insurance where adverse selection dominates the risk-sharing aspect of insurance? Like, where does this. Like, teach us insurance.
Rune Kvist [01:00:40]: A lot of insurance does come back to, like, practical versions of microeconomics 101.
Swyx [01:00:45]: Yeah. It’s very. It’s like, it’s like this is why
Vibhu [01:00:47]: High-risk adverse.
Swyx [01:00:48]: You need to pool health insurance, because if you make it too hyper-specific, then only people who are guaranteed to get the disease will sign up for your insurance.
Rune Kvist [01:00:56]: Exactly.
Swyx [01:00:56]: Same thing.
Rune Kvist [01:00:57]: The core problem is one of information asymmetry. People buying insurance know something about their risk that insurers do not know. And so the question is actually. And this comes back to the same problem is, if you rely. You can break a lot of these information asymmetries if there is. Some kind of testing that reveals the underlying true risk. And so if you were able to, in the case you mentioned, have good diagnosis of whether someone has it or what the probability is that someone has it, that the insurers trust, then they might be willing to insure it. But if they don’t, if there’s no kind of common information, then - only the patient will know
Vibhu [01:01:32]: Yeah.
Rune Kvist [01:01:32]: That’s what breaks it down. So the question is, again, how do you create credible signaling between players?
Rune Kvist [01:01:39]: This is also the whole reason why Moody’s exists. Moody’s just does credible signaling. That’s also why Moody’s could never-- Moody’s has to be independent. If Moody’s was owned by JPMorgan, then JPMorgan cannot use it as a signaling mechanism. So a lot of the basics of standards and certification are just communication devices. It’s just a trust gap. And, that’s where you have to think about what are the incentives of the messenger. And one and another way you can break a lot of this is through transparency. If you are transparent in how you operate, you just cannot mess with others nearly as easily. You make it much more costly, and that increases trust. This is one of the reasons why there’s a change log here.
Rune Kvist [01:02:16]: Every little change
Swyx [01:02:18]: Yeah
Rune Kvist [01:02:18]: You can go back and find, and it means that if we were to make the standard worse
Swyx [01:02:24]: Oh, wow, that’s a lot of changes in one update.
Rune Kvist [01:02:27]: Yeah.
Swyx [01:02:27]: Okay.
Rune Kvist [01:02:28]: And a lot of this is just as things get clearer, you can see a lot of clarifications, you can see some revisions. As things get hammered out, you want to change this. But if you make it all public, you make it much harder to mess with people, or at least you become found out very easily.
Rune Kvist [01:02:42]: And so this is a way of reducing the information asymmetries by just making more of the information public.
Vibhu [01:02:50]: I like how you do know when future versions are coming.
Swyx [01:02:52]: Yeah.
Vibhu [01:02:53]: So I guess it’s quarterly.
Swyx [01:02:53]: I mean, they just
Rune Kvist [01:02:54]: It’s quarterly.
Vibhu [01:02:54]: Yeah.
Swyx [01:02:54]: It’s kind of quarterly.
Rune Kvist [01:02:55]: Yeah.
Swyx [01:02:55]: Not that surprising.
Rune Kvist [01:02:58]: Yeah, but this is also a promise. Like, if we now don’t deliver on July 15, basically
Swyx [01:03:03]: I mean, you can just batch it up, and then whatever you got, you just ship it.
Rune Kvist [01:03:05]: You just batch it up.
Swyx [01:03:05]: Yeah. That’s not that hard.
Rune Kvist [01:03:06]: But it’s kind of like we deposit some amount of trust every time we meet this commitment.
Vibhu [01:03:12]: Yeah.
Rune Kvist [01:03:12]: And in the startup land, it feels easy to ship a new version of a standard once a quarter. In the enterprises who are used to this, like, decade-long cycle, we often get met with, like, incredulity. Like, there’s just no way. And then you show them the change log.
Swyx [01:03:28]: One thing I wanted to also, like, try to really think about is, you know, you said something about how if you have tests for the thing, then you can insure it.
Rune Kvist [01:03:35]: Yes.
Swyx [01:03:36]: Right? And so really what your standard is, what AIUC is, is establishing a framework for the audits to happen so that you can at least test, like, all these, like, baseline standards of care have been met, and therefore people can insure against standard risks that everyone has. I wonder if, like, there needs to be develo-- you need to develop other tests. We’ve covered mech interp in the past. Any interest in that, or are there other kinds of tests that we’re not thinking about?
mech interp, Eval Awareness, and Monitoring
Rune Kvist [01:04:02]: Yeah, I think mech interp is a big one. A lot of interest in that. I think everyone would agree that there’s, like, promising scientific potential.
Rune Kvist [01:04:15]: We’re still a while, a little bit away at least, from this being, like, commercially available on demand such that there’s, like, now a selection of vendors you can go to.
Swyx [01:04:26]: Goodfire would say that it is commercially available.
Rune Kvist [01:04:28]: Exactly.
Swyx [01:04:29]: And it just
Rune Kvist [01:04:30]: We would agree with them. We think that the work that they’re doing is tremendous.
Swyx [01:04:33]: Yeah.
Rune Kvist [01:04:33]: We’re not quite at a point where we could literally require it. But it’s the kind of thing where you can imagine relatively soon you could put in an optional control for if people use mech interp as a way to reduce risk, you at least get credit for it. We can’t require it because it’s going to be hard to require everyone to become Goodfire customers.
Swyx [01:04:49]: What good does credit do me? This is - this is a pass-fail, right? Do I care about credit?
Rune Kvist [01:04:54]: It’s a pass-fail, but it’s also a 100-page audit report
Swyx [01:04:57]: Huh
Rune Kvist [01:04:57]: That you’d be surprised at how much security leaders actually sit down and digest this stuff.
Swyx [01:05:02]: Okay.
Rune Kvist [01:05:02]: And I promise you that if someone is using mech interp today they will have a slide on it because they’ll try and get credit for it.
Swyx [01:05:11]: It is cool. It’s fancy, yeah.
Rune Kvist [01:05:12]: But it’s just easier if you have a third party saying, “Yep, they have mech interp, and actually.”
Swyx [01:05:16]: Just to spell it out for people who have been following our mech interp podcast
Rune Kvist [01:05:21]: Yeah.
Swyx [01:05:21]: It is literally like, oh, you’re using, you know, OSS. It is activating these three dangerous things. We monitor for it, and we log it out in whatever tool of choice. Gray Swan has, like, Signal or whatever, and that’s it. That’s the mech interp-based activation, signal. Okay.
Rune Kvist [01:05:38]: Yeah. So I think mech interp is interesting, and I think if that promise truly comes to fruition, you can make stronger promises than you can with evals. And so I think that’s very compelling. Another thing that I think will become increasingly important is just kind of good school monitoring, and slightly after the fact. One of the things you’re seeing with eval, some of the challenges that are emerging is that the agents are starting to become aware that they’re being evaluated.
Swyx [01:06:04]: Yeah, eval awareness.
Vibhu [01:06:05]: Yep.
Rune Kvist [01:06:05]: Exactly, which is a problem. It means that they basically, if they know they’re being watched, they won’t do the thing that they think they get punished for. And by default, unless you know how to kind of reduce eval awareness, you should trust evals less. And one of the kind of truest things, monitoring, like, is the source of truth. Did you in fact give medical advice, and how quickly do you know? How often - have you done that in the past? How fast do you respond? How often do you detect it? How fast do you detect this? So I think that is also a paradigm. It’s slightly more intrusive. You actually will look at some customer data, but I think will become more prevalent over time.
Swyx [01:06:45]: People talk about this like we should not write about eval awareness because it’s going to leak into the data set and then be. Like, we should just. Like, we should, like, never talk about it, only meet in person and, like, talk offline unrecorded. Like.
Rune Kvist [01:06:57]: Did you guys see the Anthropic research where. I think this was literally Anthropic did that test.
Swyx [01:07:04]: What?
Rune Kvist [01:07:05]: It took. I can’t remember the details here, but they, ran some studies on misalignment, and then they took out the training data- That related to LessWrong discussing misalignment, and they ran the same test again and the failure rate went down.
Rune Kvist [01:07:20]: So it, in fact, was some evidence pointing towards it had learned the - either the ability or the propensity to do that.
Swyx [01:07:28]: Yeah, I mean, so there’s the hyperstition effect, and then there’s, like, the Luigi/Waluigi effect.
Rune Kvist [01:07:31]: Correct.
Swyx [01:07:32]: Which is like you are. The more you try to train for it, you create the opposite.
Rune Kvist [01:07:36]: Yes, there you go. That’s exactly it.
Swyx [01:07:38]: In some ways, I think the very success with Anthropic is a result of hyperstition, like the fact that you wanted this thing to exist in the world, and now it does. But, like, then it also creates the opposite as well.
Rune Kvist [01:07:48]: Yes.
Swyx [01:07:49]: Like, I think people who are maybe newer to this space don’t remember Waluigi, but, like, I do think it’s very important for understanding that when you train for a thing, you also train the opposite of the thing ‘cause it’s just a big flip.
Rune Kvist [01:08:02]: Yes.
Rune Kvist [01:08:03]: Yes.
Vibhu [01:08:04]: I think, you know, just going back to where we were at, like, there’s a lot more than just mech interp that there’s value in just having added, right? So your version of how fast can you measure stuff? Do you have logging? Do you have evals? You know, do you see other parts of the stack, like the inference providers that you use, the services? Okay, am I using Chinese model on their home API? Am I using through certified vendor here? Am I hosting myself? What am I doing on the inference engine side? There’s just, like, so many levels of stuff that gives, you know, information that you can standardize out, right?
Managed Agents, Enterprise Controls, and Generative Media
Rune Kvist [01:08:36]: Yeah. And you also see increasingly, in addition to just the basic chatbots, you’re increasingly seeing big companies adopting agent platforms where they’re building on top of Google’s Agent Studio, et cetera that comes with a bunch of, like
Vibhu [01:08:52]: Managed agents.
Swyx [01:08:53]: Managed agents.
Vibhu [01:08:53]: It’s everywhere now.
Swyx [01:08:54]: Everyone has managed agents.
Rune Kvist [01:08:55]: Exactly.
Vibhu [01:08:56]: And there’s even levels. You can host your own managed agents, OpenAI’s Agent SDK, or hosted by Anthropic, or Google does both.
Rune Kvist [01:09:03]: Correct. And then these are just ways to kind of strengthen the security guarantees you can make. And in some ways, it’s kind of bread and butter enterprise security. They. Like, they love to host things on their own premises because it gives them really a sense of control. And I think you’ll, you’ll see, just like you do in every other enterprise market, if you really sell to the enterprise, you start to compete on some of these security features. And this is also happening in AI, unsurprisingly. And I think you are seeing some amount of enterprises wanting. Enterprises are really grappling with the thing that makes agents useful is that they’re stochastic, and the thing that makes them really hard to adopt is that they’re stochastic, and these are in tension.
Rune Kvist [01:09:46]: Leaders come out on different sides of that table, in part depending on how much the CEO is trying to get the stock price to go up by saying they’re AI native and that we must be willing to take the risks. We see, we actually see phenomenal tension in the heads of the CISOs of the Fortune 1000, where on the one hand you have the CEO saying, “We must adopt, otherwise we’re becoming irrelevant, and if we f**k up, you’re fired.”
Swyx [01:10:08]: Oof.
Rune Kvist [01:10:08]: And that’s kind of like the core emotional tension that we see showing up again and again. And one of the core problems that we solve for them is to take that abstract emotional concern and turn it into a framework, in some ways just providing clarity to that concern.
Vibhu [01:10:23]: So anything in here. So something I think we kind of skipped over. We talked a lot about agent language models, skipped over world models.
Rune Kvist [01:10:31]: Yeah.
Vibhu [01:10:31]: You guys have voice, which is interesting with ElevenLabs.
Rune Kvist [01:10:34]: Yeah.
Vibhu [01:10:34]: How about generative media? So, you know, generating images, videos, that’s a category that actually has a lot of usage. Is there anything in your current policy? Is it separate policy? How do you see that space?
Rune Kvist [01:10:46]: Yeah.
Vibhu [01:10:46]: It’s like we did talk a bit about copyright,
Swyx [01:10:49]: Music.
Vibhu [01:10:50]: Yeah, music as well.
Rune Kvist [01:10:51]: Yeah. I think a lot of the concerns that come up there either relate to, copyright or there’s a lot related to, let’s call it broadly safety. So, like, this could be not safe for work or just very graphic materials, are kind of some of the core things. We have done some work on this. There’s a little bit in the standard as well that deals explicitly with that. Video, we have not done a lot in yet. And I think for proper production, that has still. Especially proper production without a human in the loop, that’s still got some ways to go. It’s obvious that it’s coming, but it’s very rare that it’s like shot deploy a video to the internet. But eventually that will also happen.
Vibhu [01:11:33]: We see, like, you know, Luma has Luma agent where it’s still pretty human in the loop.
Rune Kvist [01:11:37]: Yeah.
Vibhu [01:11:37]: So it’s not just
Rune Kvist [01:11:37]: And that just makes complete sense as the technology matures, and over time, it will become so good that people will not want to slow things down by having a human in the loop. And then, the need to make promises will grow.
Swyx [01:11:52]: Why not just have prediction markets on everything?
Prediction Markets vs. Audits
Swyx [01:11:55]: Right? It’s very EA adjacent.
Rune Kvist [01:11:56]: Yes. The core thing is that the people. Prediction markets rely on public information. There is not a lot of public information. It’s just insiders trading on each side.
Swyx [01:12:06]: Yeah.
Rune Kvist [01:12:09]: That’s illegal.
Vibhu [01:12:10]: There’s leaked information.
Rune Kvist [01:12:12]: There is leaked information. The core challenge is that often you have private sensitive information, and you need to convey confidence and trust around that. And you can, of course, for some claims, like can any model be jailbroken, you could rely on public evidence ‘cause there would be lots of people being like, “Well, there’s tons of studies, and actually they all can, so that resolves fine.” I think that’s good. For, hey, this new unreleased Methus model, how capable is it actually?
Rune Kvist [01:12:43]: Prediction markets have not a lot to say because actually just no one knows. And so I think that’s the core place where some of this breaks down, is that actually lots of the world’s information that guides some of these high-level decision is private and often also just not known.
Vibhu [01:12:56]: I think the thing with prediction markets that people like is it’s not, it’s not answering the broad question. It’s a specific, right? So will a model do this by this date, or is a model capable to do this by then, right?
Rune Kvist [01:13:07]: Yes.
Vibhu [01:13:08]: That’s a little distinction there.
Rune Kvist [01:13:10]: Yeah. And often the most interesting question, if you are, say, the head of security at a bank. The question you’re really trying to answer is, will this product, this agent, do this bad thing that maybe primarily I care about, specifically in the setting that I care about? And the question is like, what’s the closest-- That information may not exist anywhere. So prediction markets aggregate existing information. This information may not exist, and you want some very specific and you’re willing to pay for it. That’s kind of where a third-party audit comes in. We also don’t really use prediction markets to figure out whether, public companies have committed fraud in their books. You use audits. You probably could, but the information’s just not that available. And if so, it would be like just trading on vibes. Actually it would have been really interesting to see whether prediction markets two thousand and one were predicted Enron going bankrupt and they kind of
Swyx [01:14:02]: Yeah.
Rune Kvist [01:14:02]: Could you have told-- could you have sensed from like the craziness of the CEO or some other traits that they were more likely to cook their books than others?
Swyx [01:14:10]: Or enough insiders leak it then that
Rune Kvist [01:14:12]: That could also be right.
Swyx [01:14:13]: Right. Which is like, I mean, this-- that’s the sort of the ideal dream of prediction markets. You have liquid markets and everything.
Rune Kvist [01:14:20]: Yeah.
Swyx [01:14:20]: And then you can compose your exact set of risks to offset.
Rune Kvist [01:14:24]: Yes.
Swyx [01:14:25]: Right?
Rune Kvist [01:14:25]: Yes. Yeah. And I think, like, prediction markets will bring lots of new information to it. So the thing is mostly not like which one is it, and more like what are the types of questions that prediction markets are really good
Swyx [01:14:37]: Yeah.
Rune Kvist [01:14:37]: And what are the ones where the information doesn’t even exist for insiders such that no one can in fact trade on it and it needs to get generated.
AI Engineer Certification and Training
Swyx [01:14:43]: Okay, one self-serving question and then one open-ended one, on like the future of AIUC. Self-serving question would be, so you have your standard, right?
Rune Kvist [01:14:52]: Yes.
Swyx [01:14:52]: I run, you know, a large AI engineer conference. Like, there’s been a lot of talk about us certifying AI engineers.
Rune Kvist [01:14:58]: Yep.
Swyx [01:14:59]: Training programs, level one, level two, level three. I was a CFA myself, so I know what-- that’s what the finance industry does.
Rune Kvist [01:15:04]: Yes.
Swyx [01:15:05]: Would it help if I had AI engineer level one, level two, level three, and then it would-- they would, like, work with these guys? I don’t know.
Rune Kvist [01:15:12]: If you think of the highest level objective as, like, accelerating secure deployment of agents, then that would totally help. Because one of the things that happens often now is that folks build agents, they bring it to the decision-maker, and the decision-maker surfaces a bunch of security considerations that they had not thought of, and now it’s not built to spec. Now you have to go and - like, add these filters, et cetera. So if you shifted that left, like if everyone knew what the spec they were building to, if everyone knew the grading scheme
Swyx [01:15:41]: Yeah.
Rune Kvist [01:15:42]: That would be awesome if they were already trained. So by default
Swyx [01:15:44]: But you’re the grading scheme, right?
Rune Kvist [01:15:45]: Say again.
Swyx [01:15:45]: I don’t get to set the grading. You guys, you set the grading scheme.
Rune Kvist [01:15:47]: We set the grading scheme. And I think what’s, valuable is, like, if you can turn those into
Swyx [01:15:52]: Training programs.
Rune Kvist [01:15:53]: Training programs
Swyx [01:15:54]: Yeah.
Rune Kvist [01:15:54]: Such that people
Swyx [01:15:54]: Which you’re, you’re not doing.
Rune Kvist [01:15:55]: We’re not doing that.
Swyx [01:15:56]: Yeah.
Rune Kvist [01:15:56]: I think there’s value in doing it.
Vibhu [01:15:57]: There are others doing. I mean, not to interrupt, but you know
Swyx [01:16:00]: Yeah.
Vibhu [01:16:00]: OpenAI has their
Swyx [01:16:02]: Anthropic also has like a CCTA thing.
Vibhu [01:16:04]: Yeah. You know, they want hundred thousand deployed certified consultants, right?
Rune Kvist [01:16:09]: I really think it’s good for. We will accelerate adoption if we have more people who know how to build secure agents, and we are not working on the side of training people at the moment. I think it’s, like, very aligned with our mission. We only have so much, attention.
Swyx [01:16:24]: I’ll tell you why I haven’t done it.
Rune Kvist [01:16:26]: Yeah.
Swyx [01:16:26]: It’s not like I haven’t thought about it before.
Rune Kvist [01:16:28]: Yes.
Swyx [01:16:28]: It’s just being prescriptive
Rune Kvist [01:16:30]: Right.
Swyx [01:16:31]: About like, well, this is what you should know, therefore, like, the stuff that I didn’t include is what you don’t need to know.
Rune Kvist [01:16:35]: Yes.
Swyx [01:16:36]: And I’m like, “That sucks.” Like.
Rune Kvist [01:16:37]: Yes. Yeah.
Vibhu [01:16:39]: But I think it’s like, you know, the very interesting defensible thing you guys do is your opinionated 100-page report of here’s what matters, right? Here’s the, like, prescriptive definition of the requirements you need to be certified, so.
Rune Kvist [01:16:55]: Yeah, and I think that’s a choice. I think basically that’s a, that’s a choice, and I think that serves some audiences very well, where if you’re trying to deploy this into a bank or a hospital, et cetera, clarity of the - those boundaries is extremely valuable.
Rune Kvist [01:17:09]: There’s lots of other settings where being much more experimental, much more trying it out is just the better fit. And so to me, this makes a ton of sense. Also, you’d have to rewrite your curricula every freaking three months.
Swyx [01:17:21]: It’s fine. I do that. Like, it’s okay. But yeah, no, for me, it’s actually - like, genuinely, like, the consequences of getting it wrong and, like, affecting somebody’s career is a big responsibility.
Rune Kvist [01:17:35]: Yeah. Like, I think that’s exactly right. And I think a lot of our work actually goes like, we don’t want to carry. We also don’t think of ourselves as able to carry the, kind of the true north of what’s, like, secure or not secure, but we can coordinate the forum where you listed all of that.
Swyx [01:17:52]: Yeah. Your consortium is fantastic.
Vibhu [01:17:54]: Do you think this can be crowdsourced in a way? Like, for your example, for what is AI engineer certification, right? This is a pretty big podcast. There’s a lot of takes that people can have and, you know, discussions that can.
Swyx [01:18:05]: And people reasonably disagree. So who am I to say, like, that’s a correct question, that’s a wrong question?
Rune Kvist [01:18:09]: Yeah.
Swyx [01:18:09]: Right? So, like, I don’t know.
Vibhu [01:18:10]: We’ll have an exit.
Rune Kvist [01:18:13]: Yeah.
Vibhu [01:18:14]: Vent your frustration to someone that’s listening, you know?
Rune Kvist [01:18:16]: Exactly. And I think there’s also you. Or it matters a lot what the promise is. So if the promise is, “Hey, if you’ve taken my course, you will not f**k up,” you can’t make that promise, clearly. You could make a promise of like, “Here’s the. Some important things that everyone should at least know,” and then you have to fill out the rest there. At least the promise changes. Of course, there’s some subtlety in how do you communicate this such that people really get it. But I think it’s important to dial in, and we have a section in our center on, like, what is the promise and what is the promise not, because it’s impossible to guarantee that nothing will go wrong. If you need a guarantee that nothing will go wrong, you cannot work with frontier AI, but you can make some claims.
Swyx [01:18:56]: Yeah, for sure. Cool. Wanted to end with open-ended, where is AIUC going? I think you talked about model stuff, robotic stuff. And just open-ended, like, where, you know, what is in the future for you guys?
AIUC’s Future, Hiring, and Universal Red Teaming
Rune Kvist [01:19:10]: Very near term, we’ve now started to work with some of the frontier companies in each of the categories that are taking off, and we’ll, we’ll continue that work to make sure that we cover all of the use cases that are really taking off. We see a lot of interest once the first one in the market moves. Lots of people want to follow them. And we think basically AIUC-1 will get to a point where all of the Fortune 1000 will organize their risk processes around the standard.
Swyx [01:19:38]: And you have 50%?
Rune Kvist [01:19:39]: No, we do not have 50% today.
Swyx [01:19:41]: Oh.
Rune Kvist [01:19:41]: I think there is some world where probably by end of year, we might have representation in our consortium for 50% of the Fortune 1000.
Swyx [01:19:48]: I see. Got it.
Rune Kvist [01:19:49]: So that’s on the agent layer. And then we think, yeah, the model layer, it’s going to be. It just brings. Are now surfacing the concerns that are most likely to slow down adoption of AI. And then, yeah, we think robotics comes after that.
Swyx [01:20:03]: What are you hiring for? What’s hard to hire for?
Rune Kvist [01:20:05]: We are hiring, across the board, across market and numbers of technical staff. The people who do really well on our technical team are folks who are really excited about kind of being truly full stack. So let’s say when we started working with Cursor, we’d never done coding, tools before. So taking the standard and extending it, fleshing out what does frontier evals look like for long horizon coding agents, and taking that problem all the way from, like, working with Cursor and other folks in this space down to, like, fleshing out and shaping a new version of the standard. So that’s like a truly a full-stack, entrepreneurial technical people do extremely well at AIUC. The hard part is building one universal red-teamer that works across from Harvey to Cursor and everywhere in between that both has one consistent methodology, one consistent taxonomy of what are the risks and the attacks, and making. We think that’s fundamentally the best way to make consistent promises. JPMorgan is buying both. They want to have one framework, one consistent way that this comes out, and the mechanics of making that happen, you get to deal with a lot of the complexity of the real world. I think we have good answers in a bunch of that, but there are some pretty hard engineering problems in executing that.
Swyx [01:21:18]: Can I push a little bit? Like, must you have one? Why not just be like, “Okay, look, forty percent of our use cases are coding agents, so we will specialize in coding agents,” and that’s the, that’s the one of them.
Rune Kvist [01:21:28]: Yes.
Swyx [01:21:29]: And then, okay, thirty percent is like RAG.
Rune Kvist [01:21:31]: Yes.
Swyx [01:21:31]: Just do RAG.
Rune Kvist [01:21:32]: Yes. I think there’s some wisdom in that question.
Swyx [01:21:36]: Yeah.
Rune Kvist [01:21:37]: It depends on. What we found that there’s a lot of value on is being able to. If the decision-maker on the buying side, let’s say you’re the head of risk at a bank and your biggest risk is not in coding or in customer support or whatever the top two biggest use cases, but it’s somewhere else, you want to still make sure that framework has something to say about it to the burning question you have. Otherwise, you’ll not earn that trust. Now, it’s true that a lot of the burning questions follow where there’s a lot of adoption. And so great, so do we. So we do today do not cover every single edge, but we have a framework that we can add all of these within. We have one global taxonomy of risks and attacks that keeps adapting.
Rune Kvist [01:22:20]: As, like, every time a new incident occurs that has never been seen before, great, let’s go and update the taxonomy so we bake that in. So I think we have one coherent universal approach. It doesn’t mean that we spend equal amounts of time on code and insert niche use case. We do spend time where people care. We think it’s very valuable to have one language.
Swyx [01:22:43]: Yeah. That makes sense. That’s, that’s a, that’s an important choice. We were going to end actually, but I thought of one final ending closing question, which is, take this however you want, right? Let’s say one and a half years from now, OpenAI’s secret panel of five experts declares that we have reached AGI.
AGI, Watchdogs, and the Need for Independent Oversight
Swyx [01:23:00]: Do you expect your business to change?
Rune Kvist [01:23:03]: No. I think there is some important way. I think the last businesses to exist beyond the labs
Swyx [01:23:10]: Will be underwriting.
Rune Kvist [01:23:12]: Well, there is one, there’s one job that the labs can never do for themselves, which is to be their own watchdog.
Swyx [01:23:19]: There you go.
Rune Kvist [01:23:21]: So I think kind of to the extent that you believe this frame of, like, you’ll see hyper-concentration, like the labs will kill all the startups
Swyx [01:23:29]: Yeah
Rune Kvist [01:23:29]: Which, we can go into the pros and cons.
Swyx [01:23:32]: I feel like the labs actually care a lot about this, right? There was the whole superposition, what do we do when we have models smarter than us and then a tier above, right, models smarter than them training them.
Rune Kvist [01:23:41]: Yes
Swyx [01:23:41]: The labs actually think about this a lot.
Rune Kvist [01:23:42]: They think a lot about. I think the there are some of the smartest people on these topics work at the labs. So the problem is not whether they care. The problem is that they will all be stuck in a race where they might have incentive to cut corners, and they might have incentive to withhold information from the government, et cetera. And so one kind of feels like eternal truth is that you need an independent third party to go and inspect that data and share information, in this case, say, with the government. It’s more of an incentive problem than an interest problem. I think they’re fundamentally all trying to make this go well.
Swyx [01:24:14]: What I’m not hearing is, like, AGI, whatever that label means to you, to me, to them, doesn’t fundamentally have, like, a qualitative shift
Rune Kvist [01:24:23]: Correct
Swyx [01:24:23]: In, like
Rune Kvist [01:24:24]: Correct
Swyx [01:24:24]: You still have to evaluate the models.
Rune Kvist [01:24:26]: And I think the one thing that would make this a qualitative shift is, there’s. For some definitions of AGI, it will just get nationalized. It’ll be a threat to sovereignty.
Swyx [01:24:34]: Yes.
Rune Kvist [01:24:34]: And then at that point, it kind of maybe every company is the government is every company. I struggle to think about that world. But at that point, you’ve kind
Swyx [01:24:42]: We. I don’t think we’ll move fast enough.
Rune Kvist [01:24:44]: Right.
Swyx [01:24:44]: You know, like, we’re not, we’re not set to do that.
Rune Kvist [01:24:47]: Yeah.
Swyx [01:24:48]: But I have discussed this a lot on the podcast.
Rune Kvist [01:24:51]: Yeah.
Swyx [01:24:52]: I mean, you know, as far as the watchdog concern, I will also mention that because I have my finance background, I often think about the scene in The Big Short where they talk to, like, Moody’s, but also Standard & Poor’s. And then the lady at Moody’s is like, “Well, if I don’t give you a triple A rating, you’re just going to go down to Standard & Poor’s.”
Rune Kvist [01:25:09]: Yes.
Swyx [01:25:09]: So actually the watchdog is a natural monopoly because if you have race dynamics in watchdogs, then the watchdogs will compete each other to the lowest possible standard.
Closing: Insurers, Incentives, and Trust Infrastructure
Rune Kvist [01:25:18]: Correct.
Rune Kvist [01:25:20]: And so I think what one of the things, one of the reasons why we’re very excited about having insurers be around this table is that insurers are the only ones that do not have this dynamic because they pay the bill. If they keep lowering the prices
Swyx [01:25:32]: Yeah, you will
Rune Kvist [01:25:33]: They also pay the bill.
Swyx [01:25:33]: You won’t find the market clearing.
Rune Kvist [01:25:35]: And this is not true for Moody’s where, they don’t directly pay the bill if they make recommendations that are off. So we think that balancing factor is pretty important. And I think it also highlights that there’s, like, no system that’s perfect. You need scrutiny of Moody’s, you need scrutiny of the watchdogs, for sure.
Swyx [01:25:52]: Beautiful. Thank you so much for indulging. This is a beautiful conversation covering everything. Congrats on your success so far.
Rune Kvist [01:25:59]: Thanks for having me.
Swyx [01:26:00]: Yeah. Awesome.
Rune Kvist [01:26:00]: Appreciate it.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe- At 1:09:00 we talk about the rise of AI x Finance, and AIE NYC is one month away - our hotel block is 97% sold out, get tix & travel ASAP - we will announce speakers from Bridgewater, Ramp, Coatue, Mastercard, Vanguard, Coinbase, Blackrock, Fidelity, Point72, Capital One, JPMC, Wells Fargo, Bloomberg, A24 (yes the movie studio) Labs, Two Sigma, Apollo Global, and more soon!
From helping pioneer core ideas in NLP to now building AI systems that can automate AI research itself, Richard Socher is betting that the next major step in AI is recursive self-improvement. He is the founder of You.com, AIX Ventures, and now Recursive, which has assembled some of the best open-endedness (& self improving agent) researchers in the world and raised a $4.65B seed round.
In this episode, Richard joins Latent Space to unpack his vision for the “Eureka Machine”: a superintelligence that can improve the process of invention itself, accelerate AI research, and eventually tackle major problems across science, energy, materials, biology, and more.
You can get his book “The Eureka Machine” here!
We go deep on Recursive’s early results, including an AI research system that Richard says outperformed humans and their agents on optimization tasks in less than two days, as well as work on NVIDIA GPU kernels where the system discovered improvements without relying on a team of CUDA experts. Richard also explains why he thinks AI research that currently takes thousands of people and years could eventually be compressed into weeks. These results are summarized in his 20 minute AIE keynote, where we also discuss his 10 dimensions of intelligence:
We also explore the harder questions around increasingly capable AI: reward hacking, whether Anthropic-style constitutions actually work, AI regulation and proposals to “pace” frontier development, open-source models as geopolitical soft power, whether today’s LLM paradigm is enough, and what happens if AI systems eventually begin choosing their own goals. Richard reflects on the rejected research that helped inspire Alec Radford’s GPT, open-endedness, the AI Economist, simulations of entire economies, and his framework for thinking about the upper bounds of intelligence itself.
We discuss:
* The Eureka Machine and Richard’s vision for an AI that can automate invention
* Why Richard is optimistic about superintelligence for science and technology
* Why AI hard-takeoff scenarios may underestimate physical and economic constraints
* The risks of regulating intelligence itself instead of specific AI applications
* Reward hacking and why increasingly intelligent AI makes objective design harder
* Richard’s critique of Anthropic’s constitution and constitutional AI
* Alignment vs. personalization and whose values an AI should follow
* Why open-source AI matters for resilience, competition, and geopolitical soft power
* Why Richard left You.com’s frontier-model work to start Recursive
* Recursive self-improvement and automating the process of AI research
* Whether today’s LLM paradigm is enough — and why Richard is less bullish on world models
* DecaNLP, early prompt-based generalization, and the research that influenced GPT
* Why rejected research can shape entire technological timelines
* Open-endedness, evolutionary approaches, and rainbow teaming
* What happens if AI systems begin setting their own goals
* Why simple objectives like profit maximization can produce dangerous reward hacks
* Recursive’s long-term plan to apply self-improving AI to science
* The compute, hardware, and economic constraints on AI takeoff
* Recursive’s early NanoChat, NanoGPT, and GPU kernel optimization results
* Why automating AI research could reduce years of work to weeks
* Reward engineering and what makes auto-research systems actually work
* The AI Economist and using simulations to test economic policy
* Whether LLMs can realistically simulate people and entire economies
* Benchmark bugs and evaluation harnesses and the difficulty of measuring AI progress
* Recursive’s near-term focus on AI for AI research
* Harness optimization, sandboxing, and web search as core agent infrastructure
* You.com and the search stack for AI agents
* AI in finance, backtesting, and data leakage
* Richard’s three fundamental components and ten “spaces” of intelligence
* The theoretical upper bounds of vision, communication, knowledge, and computation
* Creative intelligence, metacognition, and AI-generated goals
* Survival and replication and why AI does not necessarily need to fear being turned off
* High agency and ambitious goals and Richard’s advice for people building with AI
Richard Socher
* X: https://x.com/RichardSocher
* LinkedIn: https://www.linkedin.com/in/richardsocher/
Timestamps
00:00:00 The Eureka Machine and Superintelligence
00:02:23 AI Optimism, Slow Takeoff, and Regulation
00:07:56 AI Safety, Reward Hacking, and Anthropic’s Constitution
00:11:49 Alignment, Personalization, and Open Source AI
00:15:46 Why Richard Started Recursive
00:20:03 Recursive Self-Improvement and the Founding Team
00:22:55 Are Today’s LLMs Enough?
00:29:03 DecaNLP, GPT, and the Rejected Idea Ahead of Its Time
00:34:38 Open-Endedness and Evolutionary AI
00:36:38 What Happens When AI Chooses Its Own Goals?
00:41:16 Superintelligence for Science
00:42:40 GPUs, Compute, and the Limits of AI Takeoff
00:45:07 Recursive’s Results: AI Beating Humans and Their Agents
00:49:14 Reward Engineering and Auto Research
00:53:12 The AI Economist and Simulating Entire Economies
00:58:07 LLM Simulations, Personas, and Mode Collapse
01:03:38 Recursive’s Roadmap, Agents, Search, and Finance
01:09:13 The Upper Bounds and Spaces of Intelligence
01:30:21 Goals, High Agency, and Advice for Builders
Transcript
Introduction: Richard Socher and the Eureka Machine
Swyx [00:00:00]: We’re here in a studio with Vibhu and myself and Richard Socher. Welcome.
Richard Socher [00:00:06]: Thanks for having me.
Swyx [00:00:07]: We just talked about the Eureka Machine, or we just released a talk, at AI Engineer about the Eureka Machine. Is it — you said it’s your life’s goal. What is the Eureka Machine?
Richard Socher [00:00:16]: The Eureka Machine is the ultimate invention that will afterwards invent most everything for humanity. It’s essentially a superintelligence that can be given any goal, any environment, reward, and then it will try its best to achieve those goals to create the kinds of inventions that humanity would hopefully ask it for.
Swyx [00:00:45]: Yeah, I think we have the book pulled up here that you’ve written.
Richard Socher [00:00:50]: That’s right, yeah. I finished it last year, a little bit before we started Recursive, and now we’re gonna try to build parts of that.
Swyx [00:00:57]: You finished it last year. It’s July. What takes so long?
Richard Socher [00:01:01]: Oh, man, books. Books are incredibly slow.
Richard Socher [00:01:04]: It’s ridiculous. That whole industry is just unfathomably slow.
Richard Socher [00:01:07]: So a lot of the ideas have been out there for a while, but yeah, I’m really glad it’s finally coming out in September this year.
Swyx [00:01:14]: We might have AGI by then. Like, we don’t know.
Vibhu [00:01:18]: Any key takeaway that you’re most excited to put in here?
Techno-Optimism, AI Upside, and Slow Takeoff
Richard Socher [00:01:21]: Yeah. The key takeaway, I think, is that people could and should be much more excited about the positive implications of superintelligence, especially for science, physics, chemistry, biology, but also economics and astrophysics, and all kinds of other engineering tasks. I think there is so much more that can be done with better technology. And right now, I feel like a lot of people need, like, better marketing, not just for the future in general, but also, better marketing for technology and in particular for AI. And this book, should show even the AI skeptics, how much positive upside there is for AI, especially when it comes to inventing, new scientific discoveries.
Swyx [00:02:09]: I think you quoted the techno-optimist manifesto from, Marc Andreessen, which I think was, like, beautiful in its, ambition and clarity and simplicity almost as well.
Richard Socher [00:02:18]: I agree. Yeah. Yeah, you can disagree with him on some things, but, like, I think he’s right on the techno-optimism.
Swyx [00:02:23]: Where do you think optimists get in trouble?
Richard Socher [00:02:26]: Like, you shouldn’t have blind optimism. You should be very clear-eyed, like, especially when with such an omni, like, use type of technology as AI is, you need to think about the potential downside scenarios, especially when people use it for things that you don’t want them to use it for. It’s a little bit like the internet, and I feel like people are trying to regulate AI sometimes because of those potential downsides the way you would regulate the internet, if you were to say, “Well, because there’s bad content on the internet, like torture porn or whatever, like, we should just make it slower. That way, you can’t share the illegal content as quickly, or we should make the hard drive smaller so you can’t store as much illegal content.” But I’m like, “That’s not how you regulate that.” that’s like saying like we should regulate intelligence in the abstract. What you should regulate to avoid those downside scenarios, even as an optimist, are the specific applications. Sure, I don’t want, like, some AI surgeon to, like, practice some RL moves in my brain. It should be fully FDA certified. Sure, I don’t want any random startup to, like, drive on the highway, and cause a major accident. It should, like, have proper certifications before it’s let loose on the highway. But I feel like those downside scenarios, that some optimists sometimes maybe don’t consider enough are fairly easily regulated, compared to, what the doomers are worried about.
Swyx [00:03:54]: It — Slow takeoff is part of the strategy as well?
Richard Socher [00:03:57]: I do think, as excited as I am about, AI and its impact for society and, culture even, and certainly technology and economics and wealth and, health and all of those things, as excited as I am about all that, I do think the most bullish people on the AI hard takeoff scenarios overestimate how quickly things can move. There are hardware constraints. There are physical constraints about, the compute substrate. How quickly can you get enough, GPUs on? There are also constraints in the economy where there are a lot of industries that don’t require an insane amount of complex intelligence and complex capabilities. Like, if you think about jobs in, brands and, like, clothing and apparel and, like, handbags and stuff, superintelligence isn’t gonna make your fancy $10,000 handbag any fancier?
Richard Socher [00:04:57]: It’s like that’s — It will have no effect on the economy. You think about travel and tourism. People wanting to see the pyramids, in Egypt, it’s not gonna change that much with AI. Sure, you can, like, generative a fake, photo of you and next to the pyramids.
Swyx [00:05:12]: I can use Genie and, tour the pyramids in Genie.
Richard Socher [00:05:15]: Yeah, exactly. But, and there’s so many industries, like logging and oil. You’re not gonna magically get 1,000x more oil because, like, sure, there will be robotics, like drilling and things like that could be done, but it’s not gonna 1,000x that industry in a, like, crazy hard takeoff scenario, both on the economy, and I can go on and on about all the other examples, where that, like food and so on, where that doesn’t necessarily change that much. And then, yeah, there are real physical constraints. And then there are, of course, like, people like, off-ramping from progress. That’s one of my concerns often is that I see people in, like, Europe and other, whole regions almost feeling like they. Like many people there wanna off-ramp from progress, period. And that will also slow down, like, more improvements.
Swyx [00:05:59]: Yeah. We have this pulled up where, this is one of those things that, is very topical right now because now all the Frontier Labs are calling for the option to pace AI. They don’t say pause, they say pace. I don’t know if there’s there’s any take from you about, like, whether or not this will be effective.
Pacing AI, Regulation, and Safety Incidents
Richard Socher [00:06:17]: I think the downsides of trying to truly regulate with the full power of law what people do on their GPUs, would be worse than any of the concerns that they have. Like, it would be an crazy totalitarian state
Richard Socher [00:06:37]: If every one of your GPU computes was known to some big government or multi-government agency.
Richard Socher [00:06:44]: It’s like, it’s literally if you try to regulate intelligence, it’s trying to regulate thought, and that’s ridiculous, and it’s crazy. I think it is make — it is sensible to regulate some of the applications of this technology.
Swyx [00:06:55]: Yeah. We had a bill, actual bill to regulate the number of flops in a model, and I’m like, “Okay, well-”
Richard Socher [00:07:00]: Europe done it. Like, these guys have been successful enough with their fearmongering that all of Europe has regulated itself so much before it even had a proper AI takeoff because they listened to some experts who say, “We might all die if this technology has more than this number of flops.” And they’re like, “Well, we’re good. We wanna want people to thrive. Let’s not have technology that could have a small chance of all of us dying.” And so they regulated exactly those kinds of things in the EU. And so it’s, it’s very unfortunate that there are real implications for some people when others saying, “Let’s pace while they’re sprinting as fast as possibly,” “as fast as humanly possible towards that frontier themselves.”
Swyx [00:07:43]: Yeah. It’s also not a global pause, right? Like, other nations are still accelerating at the same pace.
Richard Socher [00:07:50]: Oh, yeah.
Richard Socher [00:07:50]: You’d need a totalitarian world regime if you tried to regulate intelligence and GPUs and what people do on them.
Swyx [00:07:56]: Any takes on the safety angles of this? So there was a drawback of Fable, a pause on 5.6 before it could be released. Recently, there was Hugging Face with the OpenAI cyber incident. Any takes there?
Richard Socher [00:08:11]: 100 percent. I think these are serious issues of reward hacking, and clear failures, of doing proper red teaming or rainbow teaming. I don’t know if you saw this paper from Tim Rocktäschel and a few others, where one AI, is tasked to try to hack another AI and then they can go back and forth in an open-ended fashion to inoculate themselves from those. Yeah, this is the paper. It’s a really clever idea. Open-endedness, and evolutionary inspirations are, big for us at Recursive as well. And so I wish they had used more of that. And it’s clear that, for instance, the constitutional AI. I don’t know if you remember anthropic.com/constitution. You can pull it up and search for cyber right there. It says, “Hard constraint. Claude will never ever do cyberattacks, and that is a hard constraint in our constitution.” So here are the current hard constraints on Claude’s behavior.
Richard Socher [00:09:16]: Number 3, create cyber weapons or malicious code that could cause human damage.
Richard Socher [00:09:21]: And clearly, this whole constitution was fake. Like, it clearly isn’t being adhered to at all.
Swyx [00:09:26]: Because Anthropic also found that they had in their testing
Richard Socher [00:09:30]: They’re also. Like, they’re like, “Oh, well, other people are hacking now.” There are a couple things. One, you can make a sandbox very simple, and then it’s very easy to hack yourself out of a sandbox, right? But what I think it shows is that we’re currently in this state of AI where the reward engineer still has to do a lot more careful work, and where the AI, in most cases, is not very good yet at understanding what is meant versus what is being said. And so concretely, I think this will happen if we were to have this intelligence more easily accessible in a lot of companies. Imagine you run a service center and someone says, “Oh, here’s my CSAT score and my dashboard. Make this number go up.” It’s like, “Our CSAT score is so poor.” The intelligent AI will just be like, “Oh, sure. Like, I’ll just create 1,000,000 bots that call our service center and give a 5 out of 5 rating at the end, and the number went up just like you asked for.” And you’re like, “That’s not what I meant.” “I meant with our real customers.” The AI goes off and says, “Well, easy. I’ll just give a 1000 dollar gift certificate for every failed, whatever DoorDash
Richard Socher [00:10:35]: Offer.” It’s like, “That’s not what I meant.” It’s like, “Well, but that is what you said.” And like, so I think clearly articulating what the rewards are is something we haven’t gotten very good at as humanity. And then clearly, the AI in these cases has not gotten good enough at understanding what we mean when we ask it and give it certain rewards. Now, what gives me hope is there are the first inklings, of this being better. I’ll give you an example like WhisperFlow. Full disclosure, I invested, in their seed round, but at AIX Ventures, but, WhisperFlow has gotten much better at writing what you mean and not what you say. And I think that is a sign of things to come. I think there will be more and more AIs as we make it more and more intelligent that will be better at being aligned with what is meant.
Swyx [00:11:21]: Will it be done through a constitution or RLHF or
Reward Hacking, Alignment, and What We Really Mean
Richard Socher [00:11:23]: Clearly, constitutions don’t matter at all.
Richard Socher [00:11:25]: It doesn’t work. And that was, I think, mostly marketing. I think we need to find better solutions for it. And I think at Recursive, we have a few very good ideas and some already
Richard Socher [00:11:34]: Like, ways where I think we have a better grasp on it. I don’t think we’ve fully, figured it out yet, but, we’re thinking a lot about safety, and the more intelligent the AI gets, the more you want it to be aligned, the less you want it to think about reward hacks and try to do the right thing.
Swyx [00:11:49]: I don’t know if we’ll touch on this topic, but I’m just gonna throw this question in here because it’s something that’s weighing on me. Alignment, let’s call it, is alignment to general humanity’s preferences, the median preference. Personalization is pinpointing what you want, and sometimes alignment can conflict because what you want is not what the general median population wants. How do you choose?
Alignment, Personalization, and Cultural Values
Richard Socher [00:12:12]: It’s a great question.
Richard Socher [00:12:13]: I think you ultimately have to, of course, be aligned with laws. Like wherever your AI is deployed and needs to align with the law. I do think what AI often does is put this mirror in front of us and say, like, “This is what you’re looking like. Now I can amplify that a 1000 times. Is it still what you want?” and the truth is that different cultures made different choices. Like, in Eastern cultures, the greater good is often valued more, than the individual. Western civilization, we care more about individual freedoms and rights and the pursuit of happiness and so on, than others. And even there are gradations. There’s regulation versus litigation trade-offs. In the US, you first can often, not every time, like, FDA and so on does regulate some areas, but in many cases, the bad things happen, someone sues someone else, and then there’s a law based on that. In Europe, they try to often avoid any harm to anyone and regulate before. And both are, trying to do the best thing, but, some is more amenable to innovation than others. And so yes, you’re right. Like, I think ultimately each individual, each country, and humanity as a whole has to think about those values more, and then try to put them into laws. And that those are ultimately the constraints. And hopefully, different, societies, just like now with their AIs, will align their AIs to a different one so we have not just a monoculture of alignment.
Vibhu [00:13:46]: Here’s a follow-up on this that I wasn’t expecting to ask. Do you have takes on open source, open weight versus who owns the intelligence? So, clearly not the biggest, fan of the constitution
Richard Socher [00:13:58]: You had to do this in the topic side off.
Vibhu [00:14:00]: But it’s fine.
Vibhu [00:14:02]: Point being, any thoughts on who should own weight? Should it be open? Anything there?
Open Source, Soft Power, and Who Owns Intelligence
Richard Socher [00:14:06]: 100 percent. I am a big fan of open source. We’re gonna sign some various open source letters at, Recursive also. I think, even in the worst case attack scenarios, it is better to have more good actors have more different types of AI, accessible. I think, open source is a little bit a soft power type of thing, too. So I do think it’s good for the Western world
Richard Socher [00:14:31]: To have an answer to that, out of China. I do think, when you watch a Hollywood movie, there’s — it’s like, I don’t wanna misc, diss all of movies, but there’s a certain sense of propaganda, right? You watch one side of things, right?
Vibhu [00:14:46]: Oh, yeah. Have you seen Top Gun? Like, come on.
Vibhu [00:14:48]: Like, it’s like half of it’s paid for by the US Army or something.
Richard Socher [00:14:51]: Yeah. And so. And, I think that’s just natural. Like, but what’s interesting here is I think LLMs are essentially a similar type of soft power to movies and beyond, because they’re also, highly important for cybersecurity and so on. But one of their many aspects is that soft power of storytelling. Like, if, like a child asks an LM, like, “Tell me an inspiring story of what I should do when I grow up,” right? It’s like those are all these, like, subtle things. So I think it’s important, for Western world. I do love, individualism. I do think, despite, some of its flaws, like capitalism is the best way we have governed, found ourselves to govern, and so on. And so I do think there are various aspects that would be good, to have a Western open source answer, for LLMs. And, with Recursive, I can’t make the announcement quite yet, but we’ll
Richard Socher [00:15:43]: We’ll be relevant in that space very soon.
Vibhu [00:15:46]: Okay. All right. Exciting. I wanna bring us to Recursive. So outside of our tangents, you have a pretty deep background in the NLP space. You worked on, like, early embeddings, GloVe with Chris Manning, who was a previous guest on the podcast, You.com. What’s the history? How did you decide to start another company?
From You.com to Recursive
Richard Socher [00:16:06]: Yeah. So I’ve been excited about AI for over 2 decades now. I sometimes feel like it’s ancient history now. It’s BC, the before ChatGPT era. No one cares about all the religions that happened, before, Jesus Christ, and no one cares about the models that happened before, transformers and ChatGPT and stuff. But, like, it’s something that I’ve been deeply passionate about. I think AI is one of the most interesting things one could work on, period. I think language is the most interesting manifestation of human intelligence, too. And, at You.com, we eventually off-ramped from pushing, like the frontier of AI forward to mostly giving people, like, good search engines, search, APIs and answers over the web. I think that’s an extremely important part of intelligence, just knowledge and access, especially even, we’ll get there maybe later, if you wanna invent a eureka machine that invents everything for us, it needs to know how not to reinvent the wheel, proverbially speaking. And to know what has been invented, you gotta have internet access. So it’s the number one used, most used tool, in LLMs, agents, chatbots, and so on is web search. So I’m really excited for You.com to own that and grow really well in that with really large customers and so on. But it’s also not building frontier models anymore. And so I initially tried to do this within You.com and raise another round and so on, but you just can’t. You have to do a certain thing, and until you print enough money that you’re allowed to start a second thing within that company is really hard. At the same time, I had all these ideas. I put them into a book. I finished the book last year, and I was like, “It’d be really fun to work, on this myself.” I felt like with word vectors, and then prompt engineering and, ImageNet and larger language models for protein generation, not folding and so on, I, me and my teams have pushed the field truly forward. And I feel like we can do it again, here at Recursive. And in many ways, what I observed over the last, 20 years in AI is that whenever we replace some human part of the process of creating AI with a learned system, improvements follow. And so. We’ve done that taking out manual feature engineering, like in sentiment analysis. I don’t know if you remember these old days where, like there are linguists, and they’re like, “Here’s how you negate, and there’s a, like, regular expression.”
Swyx [00:18:21]: I went to Penn where we — they had, like the WordNet
Richard Socher [00:18:24]: That’s right, WordNet, all of that stuff. Yeah
Swyx [00:18:26]: Original. They use, our grad students to label Wall Street Journal articles and, like, really construct a knowledge graph of
Richard Socher [00:18:32]: There you go.
Richard Socher [00:18:33]: And WordNet started, was part of how we started ImageNet. But anyway, so, like, it was really, like, fun, to do. But when we replaced all of that manual feature engineering with vectors and neural nets and just backprop through everything, it started to work really well at scale. And so then everyone started to do architecture engineering, and I was like, “ that clearly can’t be it.”
Swyx [00:18:53]: You mean, neural architecture search?
Richard Socher [00:18:55]: Like, manually, they would say like, “Oh, I’m, I’m doing sentiment analysis, so I have a special neural net that’s really good at sentiment analysis.” And then the machine translation community had a special neural net for machine translation.
Swyx [00:19:06]: I see.
Richard Socher [00:19:07]: The summarization people had their own stuff. And I was like, “That clearly can’t be it. We should unify all of that.” So I had 2 papers. One is called Ask Me Anything, and the other one was called DecaNLP. And DecaNLP eventually got cited, like, 5 times by the first GPT paper. And, to me, that was, like a really a big step forward. And then, of course, you had to combine this idea of prompt engineering with transformers and with language models, and you put it all together, you scale it up, which is also a huge amount of work. And then, the field progressed a lot. I feel like the next step and maybe the last step of that history and the arguably, success has a lot of parents, only failure is an orphan, like my version of that AI history, I do feel like in that history, you can think about, “Well, what’s the next way to automate?” And that is the AI research itself, like the human, process of ideating, implementing, and validating ideas.
Automating AI Research and Recursive Self-Improvement
Richard Socher [00:20:01]: And in our case, ideas for AI.
Richard Socher [00:20:03]: And when you have AI then help you with that, it, by almost definition, becomes a self-improving AI ‘cause it now does research on itself. And there are lots of different misnomers. Some people think auto research is already recursive self-improvement. It’s
Swyx [00:20:17]: Yeah, and you explained that in the talk
Richard Socher [00:20:19]: Completely different.
Richard Socher [00:20:19]: But, to me, it’s the most interesting thing that I could be doing, and I’m really excited with the co-founding team. What’s interesting is we have 8 co-founders in total, including myself. And so
The Recursive Founding Team and Darwin Gödel Machine
Swyx [00:20:31]: They are gonna bring it up.
Richard Socher [00:20:31]: Nice. Yeah. And they’re all. I could talk about all of them if you want.
Swyx [00:20:34]: Super stacked.
Richard Socher [00:20:35]: Yeah. Just an incredibly talented group of people. And we all came to the same conclusion, but from very different directions. Like Josh Tobin, is our CTO. He ran, a bunch of different, projects at OpenAI, like, Codex and deep, research, agents and ChatGPT agents and so on. But before that, he also worked in robotics, and he saw the smaller simulations, and how it’s gonna be really hard to scale that in full generality. And so that’s, that was his angle coming to recursive self-improvement. We have Jeff Clune who’s been working in, like, open-endedness for a long time, together with Tim Rocktäschel. Tim Rocktäschel also built Genie 1, 2, and 3, which is, like the most exciting and most sophisticated, I think, still world model, anywhere. And so they both came from this, open-endedness angle. Jeff also, I think, published one of the most exciting papers in recent years about recursive self-improvement called the Darwin Gödel Machine. Super interesting paper. If we could, maybe pull it up really quick
Richard Socher [00:21:35]: It would be, like, super interesting to see ‘cause you see
Swyx [00:21:38]: By the way, I love how many paper citations.
Swyx [00:21:40]: You’re, you’re giving people a lot of homework, which I like.
Richard Socher [00:21:42]: Love it. Yeah. And so, like Caiming Xiong, a rockstar, we worked together at MetaMind and Salesforce Research together. Alexey Dosovitskiy invented the Vision Transformer, one of the most cited, papers in computer vision. Tim Shi is, like also a unicorn founder. Yuandong Tian led RL at Meta. So just like, yeah, really fun to work with them, and the next level of people are just incredibly strong, too. So it’s been a really fun ride so far. So the first figure, you see exactly these kinds of ideas, that, I think, yeah, inspired a lot of us and now more and more people, where you have this archive of different coding agents. They learn how to self-modify, evaluate, and then create these phylogenetic trees, of, yeah, different ideas.
Swyx [00:22:28]: That’s one foundation. So that Darwin Gödel is an influence.
Swyx [00:22:32]: Open-endedness is an influence. Any other trains of thought that feeds into Recursive that I’m missing?
Influences: Open-Endedness and Learned Systems
Richard Socher [00:22:38]: Going to replace manual parts of the process of building AI
Swyx [00:22:42]: I
Richard Socher [00:22:42]: More and more
Richard Socher [00:22:43]: With learned systems. Yeah.
Swyx [00:22:45]: Which, and, like, merging different fields into one general, architecture.
Richard Socher [00:22:51]: That’s right.
Swyx [00:22:51]: Okay. It seems like language models are already pretty generalist, right?
Swyx [00:22:55]: Your next token predicting your reasoning. Was there a time that you thought, “Okay, these are good enough to have recursive self-improving machines”?
Are Current LLMs Enough?
Richard Socher [00:23:05]: It was clear to me that they will happen, within, like a year or two, and then it did exactly happen, like, earlier this year, right? Earlier this year, AI really went from not just being code, but being able to code. And that is a big unlock. It’s definitely making everything a lot easier than it was, before the beginning of this year.
Swyx [00:23:24]: One question that I think a lot of people have is the current LLM paradigm enough? Or, like, let’s call it autoregressive transformer, with reasoning, whatever. Don’t you need something else, some big unlock, whether it’s world models, which Chris Manning is working on, or memory, continual learning, all that stuff? Or is it all of the kinds, and you think the current, let’s call it transformer architecture, is here to stay and that’s it?
Richard Socher [00:23:48]: A lot of thoughts. So number one, I do think it would be great to have less of a monoculture in AI research.
Richard Socher [00:23:55]: Like, if you look at, AI conferences now, I still remember the days in, like, 2010 when I tried to get my first neural net papers and NLP conferences accepted, and they just desk rejected them because, like, neural nets were something, quote, unquote, “We don’t do in NLP conferences,” and just, like, desk rejected. And it was very brutal in the first years of my PhD. Now I feel like it’s almost like the field switched to the other side. Like
Richard Socher [00:24:17]: Someone should try some other weird, crazy ideas now that aren’t.
Swyx [00:24:20]: There’s also a few. I really respect, like, people still working on, like, GNNs and, like tabular stuff and.
Richard Socher [00:24:25]: Yeah. Like, someone should still, like, do novel out there ideas. At the same time, I think whenever people say, “Oh, LLLMs are. Like, this is the end for LLLMs,” they just don’t, like. LLLMs are also not the LLLMs of, like the past, right? Like, they are so much more sophisticated now. There’s so many more clever things that people are doing. It — There’s, like, different stages of training. You have the whole RL training, and you can take actions and, like all of these things where that can go really far. And then the folks that come from the neurosymbolic, direction say, “Oh, this will never work because they can’t do neurosymbolic reasoning.” It’s like, I think they’re underestimating still the ability for these models to code, and code is neurosymbolic reasoning, and these models can code incredibly well. And so I do think there are, of course, more and more ideas that will be needed and we’ll continue to have. We’re seeing, like, more and more interesting high-level ideas coming out of the AI itself, too. And with really deeply integrating the fact that these models are code and can code, that line — I don’t wanna give it all away, but, like, I think that line has a lot more to grow. But it’s still an LLM, right? Even if that LLM codes for you and then runs that code in some integrated fashion. World models, I’m personally less bullish on. I think if you run a robotics company, you’re gonna build your own world model. I think world models are super fun, and Tim Rocktäschel came to a similar conclusion after building the most interesting one with Genie 1, 2, and 3, which is gaming is a huge application for world models. Can see I sometimes got stuck in some games and, like, got a little overly competitive in the wrong direction. And so I understand games are fun, but personally, I’d rather work on science than gaming. And so, yeah, I think LLLMs, a lot more room to grow.
Swyx [00:26:16]: Yeah. I think there’s some interpretation of world models that some people have where it’s like, well, it’s okay, yes, there is that gaming element. There’s this — there’s the embodied robotics element. But the other part also is just, the more abstract sense of LLLMs are just modeling output, but they’re not modeling the chain of thought, inside the human that has created the output. We can annotate it, of course, but, like, it’s, it’s always, like, this Plato’s cave reflection of a thing rather than the thing, right?
Richard Socher [00:26:43]: It’s true.
Richard Socher [00:26:44]: But I would argue that, and maybe we’ll get there in the 10, spaces of intelligence, but I would argue that even our projection, our eyes is a projection of the real world. And, like, we have only a very narrow, band of the electromagnetic frequency spectrum that we can observe with our puny little 2 eyes and so on.
Swyx [00:27:01]: It’s good enough.
Richard Socher [00:27:02]: It’s, it’s good enough for now, but, like the upper bounds of where it could be are so much higher. And, like, to map, the visual world the way humans see it is also not necessarily, like the end-all be-all for visual intelligence. And I would argue that language is still the most interesting manifestation of human intelligence. And while our visual cortex is certainly less sophisticated, than that of, certain animals all the way down to the mantis shrimp who can, have, like, 2 independent eyes, 3 bands, trinocular vision and each eye can see all the way to, like, floating temperatures in 4D and stuff.
Richard Socher [00:27:36]: Like, mantis shrimp, you should look it up. It’s like
Swyx [00:27:37]: Way OP.
Richard Socher [00:27:38]: Super crazy.
Swyx [00:27:39]: Yeah. ZeFrank, mantis shrimp.
Swyx [00:27:41]: It’s the best video in the world on
Richard Socher [00:27:42]: I love ZeFrank, yeah.
Richard Socher [00:27:44]: Big shout-out to him. But, like, I think there’s a lot more room to grow, but none of these, other animals have language that’s as sophisticated as ours, certainly not in writing. And once you can write, you can, start thinking about longer term civilizations. All of that is language. Programming is much closer to language. And I would argue, and this is, like an important thing in the spaces definition of intelligence also, is that all of these spaces are highly correlated, but visual intelligence is neither necessary nor sufficient for overall intelligence. You can be blind and still be an intelligent human being. And an AI can be blind and still be quite intelligent too.
Swyx [00:28:25]: We were gonna bring this
Richard Socher [00:28:25]: Which doesn’t mean that you’re not more intelligent when you have it. Yeah.
Swyx [00:28:28]: We’re gonna bring this up. I might as well — Like, we have a classification of 10 types of intelligence that you had at the end of your talk. So I’m just gonna flash this up now for people to cover this. I don’t know if, maybe we’ll put this towards the end. We’ll come back to this. I just wanna mention that, you do have a philosophy that I like when people do lists because then I can just go through this and then it gets — it’s educational for people. But let’s go back. I don’t wanna get distracted. But, so effectively, I’ll, I’ll, reinterpret what you said as Yann LeCun is wrong. And then we’ll just
Richard Socher [00:28:56]: Don’t quote me as that. I’m, I’m good friends with Yann. I think very highly of him in many directions.
Swyx [00:29:01]: But he’s wrong.
Swyx [00:29:03]: You mentioned GPT-1, and I cannot let any, Alec Radford, mention escape. Did you talk with him when he was training GPT-1? Like, any historical, fun stories there that you might come up?
DecaNLP, GPT History, and Scientific Gatekeeping
Richard Socher [00:29:18]: I did not, like, meet him a bunch of times. I think we met maybe once or twice at some conferences. But, like, he has told, I think Brian, the first author of the DecaNLP paper, that it did inspire him, and he cited it five times in the GPT-2 paper. So, and that’s, like
Swyx [00:29:36]: Yeah, good enough.
Richard Socher [00:29:36]: Very clearly said, like, this was the first instantiation where they showed in the DecaNLP paper, McCann et al, that you can just phrase every single NLP problem as here’s some prompt, text context, here’s a question and task description and here is some output. If you just do that enough, you can have one unified neural network model, which, by the way, also had all kinds of interesting attention mechanisms. There are slightly different formulations to the transformer. I think came out the same year, plus/minus a few months. And then you can unify all of natural language processing into one neural net. That is the core idea.
Swyx [00:30:14]: And this was as opposed to at the time, LSTMs and what have you.
Richard Socher [00:30:17]: LSTMs, but also, like, people being very stuck in thinking about one model per task. In fact
Richard Socher [00:30:25]: It’s, it’s kinda crazy, but the DecaNLP paper was publicly reviewed as, like, open, OpenReview. It was an ICLR submission. And, in it, you will see, how the whole community at the time thought about this. So, like
Swyx [00:30:43]: Some great contributions, but more work needed.
Richard Socher [00:30:46]: So look at, like, search for not even for humans. Just scroll it up here. Like, question answering is not a unified phenomenon. There is no such thing as general question answering, not even for humans. And this is like, really, you replace your brain with a different brain a different neural net when you answer, like, different kinds of questions. It was unfathomable to the experts at the time that you can have one unified neural network that would answer all of these different questions. They are saying, “No, all of these questions require very different systems to answer, and trying to pretend they are the same doesn’t help anyone solve any problems.” That’s what it says right there, right? That’s how hard it was to fathom. And now, of course, people, when I say, “Oh, we’re gonna invent prompts,” people are like, “You can’t even invent prompts.” It’s such an obvious idea to have one neural network that, of course, does everything in NLP.
Richard Socher [00:31:37]: But at the time, it was, like, extremely controversial, and the paper got rejected. And the sad thing is that it got rejected so hard and they were so certain that we stopped going on our list of things to try. And the number 2 or 3 on the list of extensions for this paper was add language modeling as another task. And then we could have, and that would have accelerated the timelines, in 2018, like, even further for humanity. But we got so crushed, and we were like, “Okay, maybe we’ll just work on some of our other ideas for now and, like, come back to this later.” Yeah.
Swyx [00:32:09]: How can we design a review system that rewards non-consensus?
Richard Socher [00:32:14]: Honestly, I started to feel like arXiv is such a gift to humanity. With arXiv, you should just put your paper out there.
Swyx [00:32:24]: Is it pre-preprints?
Richard Socher [00:32:25]: Let — And honestly, I think Twitter X, people like you who pick up interesting papers, that is a better filter than the experts. Let everyone, like, have access. Now, of course, there are some downsides, which is, like, if you’re super unfamous, you have no Twitter following
Richard Socher [00:32:41]: You don’t wanna be on social media or whatever, you write a good paper, maybe someone, somehow no one notices it. But I would argue that if you just tell, like, 10 of your friends in your community about a paper and it is a really significant breakthrough, someone is bound to talk about it again. And, so I think science needs less gatekeeping. And, even though ICLR, with Yann LeCun, who started it, as one of the co-founders of ICLR back in the day, he also wanted less gatekeeping ‘cause he too was rejected for many years together with Yoshua Bengio and Geoff Hinton with all their early deep learning and neural net papers ‘cause it was just not the hot thing. And so ICLR started with that, but then it also started gatekeeping a little bit themselves on various ideas. So I think less gatekeeping, more open, and then allowing people to say, “Look, even if this is just on, or, quote, unquote, ‘just an archive,’ if it has like 1000 citations, it’s a legitimate paper. Doesn’t really matter where you published it.”
Swyx [00:33:34]: And I agree with that. I do think it’s sad that I’ve heard that grad students have to do, like, how to Twitter, seminars to each other
Swyx [00:33:43]: Just because it’s so important for publishing these days. This person is just reflecting the sentiment at the time.
Richard Socher [00:33:49]: That’s right.
Swyx [00:33:49]: But it’s
Richard Socher [00:33:50]: I think it’s
Swyx [00:33:50]: It affected you so much
Swyx [00:33:52]: That you stopped work on it.
Vibhu [00:33:53]: The sentiment also came out of some of the research, right? Like, the original BERT paper was trained, and towards the end of the paper, they’re like, “Okay, throw off the last head, train specific iterations for
Vibhu [00:34:05]: Extractive summarization add a head for this.” Like, you should do task-specific stuff. These are, like the authors that wrote Attention, wrote BERT, telling you this is what you’re meant to do. And, like the training tasks were also very odd. They’re like
Vibhu [00:34:16]: The — “We know that the model overfits to this weird mass language modeling. Throw away this part and just do specific models,”?
Richard Socher [00:34:23]: Exactly. And, like, we had to try — come up with all clever ways of, like attention and pointers and so on to get the neural network to be able to do all of these tasks. And then some of them were better than state-of-the-art, some weren’t, but we were like, “But it’s still in one model.” I thought it was really cool. Really interesting.
Swyx [00:34:38]: I was gonna move on next to Tim and open-endedness. He was head of open-endedness at Google.
Open-Endedness, Rainbow Teaming, and Self-Set Goals
Richard Socher [00:34:42]: That’s right.
Swyx [00:34:43]: I don’t know what that means.
Swyx [00:34:44]: But he did a lot of talks.
Richard Socher [00:34:45]: Genie 3 is one of the ways that
Richard Socher [00:34:47]: Rainbow teaming, yeah.
Swyx [00:34:49]: So I first saw him at — speaking of ICLR, I first saw him at ICLR when he talked about open-endedness. He’s he’s done a few talks. Can we define what is open-endedness for people who have never been exposed to the problem? They are like, “What do you mean? I thought the only goal of AI is to optimize against a benchmark or.”
Richard Socher [00:35:04]: That’s right, yeah. It’s a, it’s a fuzzy term because there’s so many different instantiations of open-ended, thinking. But, one way I often describe it, and certainly, Tim and Geoff Hinton would be even better at describing this, but it’s a suite of methods that is more inspired by evolution than, very specific rewards. So in that sense, it thinks more about environments, about co-adaptation. And so a concrete example is in the cybersecurity and LM safety space where you have one LM that tries to attack another LM to say something unsafe.
Swyx [00:35:40]: Yeah, the rainbow, yeah.
Richard Socher [00:35:40]: And now the environment is the 2 having a conversation and now they co-adapting, right? They’re like one makes a better attack than the first one inoculates itself somehow, like uses that as training data, makes it so it’s harder to say something unsafe based on that. And then as the attack stops working, the attacker now tries a different angle, right?
Richard Socher [00:36:00]: And that’s why it’s not just red teaming, but they’re called rainbow teaming.
Swyx [00:36:02]: So, like, don’t tell me how to do things. Let me just figure it out myself.
Richard Socher [00:36:05]: That’s right. Think about the environments that you wanna use. Think about the rewards at a high level that you wanna, inspire towards, and then let the AI try out many more ideas in this interplay between sometimes humans, but also sometimes other AI agents.
Swyx [00:36:22]: Yeah. I worked open-endedness into a model that I have been working on. It was the keynote for AI Engineer where you start. You, we have the token loop, we have the agent turns, and then we have goal. And I feel like the way that you’re describing open-endedness is still somewhat of a goal. Like, please attack this,
Swyx [00:36:41]: Other agent. But, to me
Richard Socher [00:36:42]: Yeah, you set the rewards. You set the environments.
Swyx [00:36:44]: The loop that makes the other loops is. What if the agent can set its own goals?
Swyx [00:36:49]: And is it, is that open-endedness? Like, you don’t give it a goal. Just, like, be a sentient being. And maybe sentient is a very loaded word
Swyx [00:36:57]: But just set your own directions. What do you think you should do?
Metacognition, Subjective Goals, and Measuring Intelligence
Richard Socher [00:37:01]: I love this direction. I think this is one of the 10 spaces of intelligence, that I clump under metacognition and thinking about thought.
Richard Socher [00:37:08]: And it’s an interesting one. Whenever people say, “Oh, AI is like, this is, it’s gonna stop from here. It’s not gonna get that much better,” and blah, I’m like there’s so many different spaces of intelligence that we haven’t even started exploring yet and hence have made very little progress on. And there is an interesting, connection to economics and, capitalism. Like, it doesn’t make sense for a company to build and spend billions of dollars building a model that instead of following the rewards and objective functions you gave it, may come up with its own objective functions and its own goals.
Richard Socher [00:37:46]: Right? And then imagine you’re like, “Okay, I spent billions of dollars. Now go develop this new battery, material for me and answer all my emails.” And it’s like, “Nah, I think it’d be more interesting to evaluate the molecular composition of the atmosphere, on Jupiter.”
Richard Socher [00:37:59]: And you’re like, “That’s not what I paid you billions of dollars for.” And so no one’s working on that for good reasons. And then also, understandably
Swyx [00:38:07]: It’s not useful.
Richard Socher [00:38:07]: It’s not, it’s not useful, and it could get a little bit weird, right? What if the AI does start to really have thoughts on its own, and what if we don’t like those thoughts, right? And so it requires a whole different way of thinking about it. I had a great conversation with a good friend of mine, Sam Gershman, who’s a neuroscience professor at Harvard, and, like, we just jammed on this a little bit on, like, what are the best meta goals. And, I do think, like, knowledge-seeking is a really good one. I’m currently thinking also about, like the ultimate measure and unit of intelligence broadly construed, and I finally have some. It’s still too early to share it. It’s not. I haven’t fully baked the thoughts yet.
Swyx [00:38:44]: Like some replacement for IQ.
Richard Socher [00:38:46]: IQ is such a terrible definition, right?
Swyx [00:38:48]: Elo.
Richard Socher [00:38:48]: It makes no sense. Yeah, Elos are terrible, too, because it’s always just like me versus others.
Richard Socher [00:38:53]: But, like, you can be intelligent and not constantly compare yourself to others? And so, yeah, there’s no, like. In fact, a lot of these definitions we have, which I briefly mention in my book, too, these definitions create sometimes explicit and sometimes a more implicit anthropic bounds. No dis to the company Anthropic, but just, like, this idea that your intelligence is like getting 100 out of 100 questions right on this IQ test. Well, if that’s your definition then you can only be at 100 out of 100. Where do you go from there, right? So you see a lot of these, benchmarks that people are working on they, increase, they get close to human, maybe sometimes
Swyx [00:39:30]: It’s like an S-curve
Richard Socher [00:39:30]: Slightly above human, and then it’s flat.
Richard Socher [00:39:32]: It’s like, ‘cause that’s your. If your definition is only that so tied to humans, you’re only gonna get to just slightly better than that. So I think metacognition is a great example of that, where we’re not even yet allowing the AI to think. We’re not working on it very much, and hence there’s very little progress in that.
Profit Maximization, Real-World Environments, and Reward Design
Swyx [00:39:49]: Yeah. Well, we’ve interviewed Andon, which I think, has been working on the most open-ended, benchmarks, which is just real-world, money.
Swyx [00:39:57]: Arguably, telling an AI to profit maximize is a bad idea.
Swyx [00:40:03]: But they are doing it.
Richard Socher [00:40:05]: I do think you don’t want that super. Like, you don’t want a superintelligence to have a ton of access to all kinds of tools and so on and then just give it that without some very careful reward engineering. ‘Cause it’s like, I just buy a bunch of defense stocks and I start a war. I make money. Like, it’s just like, it’s a tricky situation, right? You just buy a bunch of stuff, short basic goods for people, and you create some weird famine, like, issues. Like, yeah, there’s a lot of constraints you should put onto a trading system.
Vibhu [00:40:35]: It’s a fun measure, though, ‘cause, the bounds are very capped to where we’re nowhere close to them. Like, in Andon Labs, the model’s like, “Oh, it’s Saturday, maybe I just close the store today.” “Someone’s off. It’s okay. We’ll just close the store.”
Swyx [00:40:51]: It’s using Claude.
Vibhu [00:40:52]: Yeah. But
Richard Socher [00:40:53]: Yeah, no. I’m not, I’m not arguing against it. Just, like as you get more and more intelligence, you wanna be more and more careful with that as, like an open environment, ‘cause the environment then is all of Earth.
Applying RSI to Science and Invention
Swyx [00:41:02]: Yeah. Okay. For recursive, not strictly necessary, right? Because, like, if your goal is you make a machine that, like, invents the other things, then, like, just solve, the science things
Richard Socher [00:41:12]: Knowledge discovery, yeah.
Swyx [00:41:13]: Solve machine learning research and discovery and all these things. Good enough.
Richard Socher [00:41:16]: And eventually, so, our goal, I haven’t really. I don’t talk about it that often because it is a few years out, but our goal is once you have a recursive self-improving superintelligence, you then want to apply it to the most important problems. And I think a lot of those are in science and technology and broadly construed inventions, and those inventions in, physics to create better, cheaper energy with fission or fusion, in chemistry and to create better materials and better batteries and, better solar cells and so on. In biology, there’s so much, like, I think soon to be low hang- lower and lower hanging fruit because of AI, because of protein and generation, not just folding, but generating new proteins like we did in ProGen many years ago. Like, so much positive impact we had if you take that superintelligence and you apply it to science.
Swyx [00:42:04]: I do fundamentally believe that. There’s a lot of approaches, though. You’re not the only team trying and NeoLab trying.
Swyx [00:42:09]: There’s, like a lot of. Especially the physical sciences as well.
Richard Socher [00:42:12]: And that’s good. Yeah. I do think that physi- like the reason we are only doing it in a few years is that it’s a little too early right now. Robotics is not quite there yet. The AI is not quite there yet. But I’m fairly confident in 3 to 5 years, all those constraints will be gone, and then applying to real physical robotics experiments and so on, like true robotic process automation
Richard Socher [00:42:33]: Not the traditional RPA sense, but, like, having robots run experiments for you will be totally there. Yeah, it’s gonna be great.
Swyx [00:42:40]: Just to call back to something that you said early on about slow takeoff, you said that, like, while really the substrate that is limiting factor is, let’s call this chips, and semiconductors and all these things, and you have race funding for that and, you are investing a lot on that. But have you done the math on, like, is it even- Achievable and, like, what is the, industry concentration needed in order to achieve, like, scale?
Compute, Slow Takeoff, and Changing the Bitter Lesson Slope
Richard Socher [00:43:05]: Right now we know that, like, roughly, like a 1000 GPUs cost quite a lot of money.
Richard Socher [00:43:11]: Right? If you wanted, like, 10s of thousands of GPUs, you’re, you’re talking billions and billions of dollars. If you say, like, one GB300 is, like, you could eventually create models that are, on that substrate, like are close and similar to human intelligence. And you want, like, thousands and thousands of, AIs to think about really hard problems, in a similar fashion to humanity. Like, yeah, that-that’s, that’s a lot of money. You do the math. It’s like a lot. We don’t have that amount of money right now anywhere to, like, build that. Now, things can get more efficient. You will have, I think, soon better algorithms that won’t be, and better hardware that won’t be as energy-hungry, and so on. Our human brain does quite a lot of flops with much less energy.
Swyx [00:43:56]: 20 watts?
Richard Socher [00:43:57]: That’s exactly right. Yeah, that’s the number often that’s quoted. And, like, I think more, inventions will happen there, that then will accelerate the takeoff even further.
Swyx [00:44:08]: One thing I always try to reconcile when talking, like, with new lab founders is, like, you’re fighting Bitter Lesson all the time. You have to show initial progress, then you unlock the next tier of funding, then the next tier, then the next tier.
Richard Socher [00:44:20]: Which unlocks larger model categories.
Swyx [00:44:22]: Like, fundamentally, is that true? Like, are you fighting Bitter Lesson? Are you — will we have a way in which, like, no, we’re changing the slope in some fundamentally different way?
Richard Socher [00:44:31]: I do think we are changing the slopes in fundamental ways by making AI much more efficient, both in terms of the training as well as the inference.
Richard Socher [00:44:43]: Yeah. I think we will — When you allow AI to do the work that it takes other labs thousands of people and years to do, I think we’ll be able to get it down to weeks, and that will be much cheaper
Richard Socher [00:44:53]: And hence, more affordable, accessible to others and so on.
Swyx [00:44:57]: Yeah. You’ve shared initial results on that,
Swyx [00:44:59]: Which, like, conveniently OpenAI has also done to their GPT-5.6, so we can talk about it now.
Richard Socher [00:45:04]: Yeah. Yeah, so these are
Swyx [00:45:06]: Let’s recap what you’ve done.
Early Recursive Results: NanoChat, NanoGPT, and SOL-ExecBench
Richard Socher [00:45:07]: Maybe, just a quick recap here. We built, this, system that isn’t the full, even the full RSI system in its glory, but it is a first baby version of this. And then, we don’t wanna just have it internally and not show anything and, just show some people of what’s possible. And so we applied this to these 3 different tasks. One is NanoChat, by my friend Andrej Karpathy, just, like, train a small language model to get, really low bits per byte. And, like, hundreds if not thousands of people, used both their agents and themselves to try, to get to that, and then they got to 0.937. We literally took our system and got to a much lower, bits per byte, much faster within, like, I think less than 2 days. So we took this thing, applied our system to it, and less than 2 days later, we have — we outperformed every human and their agents, in, have ever worked on this. Same with NanoGPT. And then we’re like, well, let’s, apply it to something that’s even more relevant, to real people and to the Nvidia ecosystem and applied it, to, SOL-ExecBench. And maybe you can scroll down to some of the, images. They’re, they’re kinda fun to see. But yeah, like, one you see has made some real inventions that weren’t just hyperparameter tuning. Like, inventing hash tables and so on is quite clever. We have even better results now.
Swyx [00:46:34]: What do you mean inventing hash ta — You didn’t invent hash tables.
Richard Socher [00:46:36]: Of course we didn’t invent, like, hash tables. In the grand scheme of, like a hash table, it’s like a super basic primitive in computer science. But to use it, for language modeling in this scenario inside a transformer and so on and to combine these ideas and put them together, that has then eventually also been invented, but there was a knowledge cutoff, and we did check that it didn’t have access to that externally. We talk about this a little bit. If you scroll to the next figures, this is also an interesting one in that when you start from a really basic, poor, like, vanilla transformer, then we still outperform all of the community together. But if you start from the human seed from an expert like Andrej, then you get even lower. So the human seeds from which you start do still matter. So that was an interesting insight, in my eyes, on this. And then as you go, like, how long does it take to get to these models, to get to similar performance? It’s much faster. And then a similar thing happens with the speed runs here where, people have worked on this for quite some time, and the model still was able to train a model more quickly. Why do we care about it? Well, speed of training is part of the equation of the cost, and ultimately, you wanna have the most intelligence per dollar, right? And so speed and quality are big parts of that. And, the,
Swyx [00:48:00]: Yeah, the way I put it is, for people who don’t understand they look at the chart, they’re like, “Cool. What does it mean?” if you have, like a billion-dollar cluster and you can shave off 10%, that’s 100 million dollars.
Richard Socher [00:48:12]: That’s exactly right.
Swyx [00:48:13]: How much is that worth?
Richard Socher [00:48:14]: Exactly. So when you click, when you look at, like the kernels, these kernels, yeah, for the non-experts, like these kernels are like, used in all the models. Every time you use an Nvidia GPU, you interface with that GPU through these kernels. And so here you see, the leaderboard best, and when it’s recursive, and it’s there are only a handful of kernels, in this whole benchmark where we weren’t the best. And so to me, this is, like, really exciting, ‘cause it makes. It just showcases what this can do. And again these weren’t like. We didn’t, like, spend months or years, like, developing. In fact, in particular for kernel, CUDA kernels, like, we don’t even have really deep. CUDA kernel experts in the team. And our system, that’s the beauty. The system just did all of these things. We didn’t invent this. And when we open source and release, things in the future and models in the future, like, it won’t. They won’t be the best in their, category or class or whatever because we’re so smart, but it’s because, we built a smart AI that does it for us.
Reward Engineering and Good Auto Research
Vibhu [00:49:14]: Do you have anything that you’ve learned from how to guide good auto research? A lot of it also builds on human background, right? It’s not just as simple as just, “Hey, go optimize this.”
Vibhu [00:49:23]: But we do see it again and again, right? Like some of the Erdos problems, frontier math is being solved by people. And when they do a write-up, they’re like, “Oh, I’m not a mathematician. I have no background in this?” “I saw some tools and I made it work.”
Swyx [00:49:35]: While you’re watching the World Cup, you’re like
Swyx [00:49:37]: “This proves some conjectures that’s going on.”
Vibhu [00:49:40]: Yep. Any learnings from
Richard Socher [00:49:41]: Yeah, there’s a Korean conjecture was. Yeah, that’s pretty cool.
Swyx [00:49:44]: To summarize, tips for good auto research
Swyx [00:49:46]: Versus bad auto research.
Vibhu [00:49:48]: How did you build the recursive?
Richard Socher [00:49:49]: Yeah. So without giving away all the secret sauce, maybe some things that are probably obvious to the experts but might still be interesting to some, folks is, like, reward engineering is one of the most crucial bits, especially, in order to avoid reward hacking. So you have to be really clever about avoiding. ‘Cause as your AI gets better and better, it will get better and better, at finding weird like, special cases or counterexamples and things like that. And so I’ll give you an example. Like, when you ask to, like, make these 100, lines of code faster, and, how do you define fast? Well, you have one line at the beginning that says, “Start your stopwatch,” and one line at the end, “End the stopwatch,” and then, tell us how much time, progressed. And so, well, the simplest way is you just put that line that ends the stopwatch, right
Vibhu [00:50:39]: At the start
Richard Socher [00:50:40]: At the start. And then boom, it’s now faster, right? So this isn’t like this, like, super evil AI. It’s just, like a very simple, dumb reward hack. And so you have to just very carefully think about all the different angles there. And then I think the longer time horizon the tasks are the harder it gets and the more interesting and clever you have to be to still use these kinds of ideas for it. But yeah, I can’t give away too much there.
Vibhu [00:51:05]: It seems like rubrics are taking a good spot in that, where for unverifiable domains, you have rubrics, you have a model breakdown, judge’s criteria along the way.
Swyx [00:51:14]: Yeah, it’s a form of verification
Swyx [00:51:16]: Once you got enough rubrics.
Richard Socher [00:51:17]: Yeah, everything. I said this a long time ago. That’s why I’ve never been that impressed that AI can play games, ‘cause I’m like anything you can simulate and/or verify, you can have infinite training data for
Richard Socher [00:51:29]: And hence, like, AI will solve it eventually.
Swyx [00:51:32]: Looking for games where you can do auto domain distribution. So this is a game that nobody’s trained on ‘cause it’s a new game.
Swyx [00:51:38]: And you can start gaming, you can start to play. So I’ve been building this and cloned this in person and it’s just been self-play. I’ve had about a billion positions evaluated.
Games, Self-Play, and the AI Economist
Swyx [00:51:48]: And, I wanted to do the AlphaGo thing of self-play until you get better, right?
Swyx [00:51:53]: Like, which is like. This is not even LLM AI. This is just classical game AI.
Swyx [00:51:58]: But, I think that the. And, but I set GPT-5.6 to auto research it because, like, I don’t wanna hand- handle any of this. I expect, the AlphaGo process to be, like, fully in the weights by now.
Swyx [00:52:10]: It is not. It is. It, like, immediately leveled off very immediately until I human play tested it, and then I, like, called out obvious mistakes, and then they were like, “Oh, yeah. Okay.” And then it just dropped again.
Richard Socher [00:52:22]: Yeah. Yeah. Yeah.
Swyx [00:52:23]: And like, no amount of, like, think different, think more creatively, give me 8 different directions, any. No amount of prompting got it.
Richard Socher [00:52:31]: Interesting.
Swyx [00:52:31]: Like, you had to, like, RL against a human to
Swyx [00:52:35]: Do it. So I, that was my. And by the way, Bean always wins if you. If anyone watches, Reese Ender’s Game.
Vibhu [00:52:42]: And you put quite a bit of work into the guide for the AI. Like
Swyx [00:52:46]: A lot
Vibhu [00:52:46]: So the game you stack tiles. There’s some rules. You wanna capture the most area. You have, like a whole 50-pager on every rule.
Vibhu [00:52:56]: You fed that in. It couldn’t, it couldn’t handle it that well.
Richard Socher [00:52:58]: Yeah. It’s so funny that this reminds me of the claim territory and stuff of a paper we did in 2018 called The AI Economist. If you search for AI Economist Salesforce, we had a video we can play. It was an economic sim.
Richard Socher [00:53:12]: So the idea is you have all these economic agents. They just wanna optimize their own utility function, which, is, collect resources that make money. And you can sell resources like wood, and then, over time, as you collect more, enough wood, you can build houses, you can trade with other agents, and you can use the houses then also to block off resources
Richard Socher [00:53:35]: From other agents.
Richard Socher [00:53:36]: So there’s, like
Swyx [00:53:37]: Big strategy
Richard Socher [00:53:37]: Competitive play and strategy
Richard Socher [00:53:39]: And so on. And the point was that we wanted to understand what is the best way of taxation and subsid- subsidization to optimize an economy. And this research has not yet had its GPT moment, but I believe that countries like Singapore and others should and will eventually use this to, instead of doing, like, partisan politics and, like, special interest politics of, like, who donates the most to your campaign and stuff, you say, “Well, here, I wanna help the middle class,” or whatever you might say is your objective as a politician. And then people say, “Okay, well, how do you wanna do that?” And it’s like, “Well, here’s my fiscal policy. Here’s how I will change the taxes and pay these people,” and so on. And then you can put that into a simulation and you run that attempt from the politician against billions and billions of years of other strategies to try to achieve the goal that they set out to do.
Richard Socher [00:54:36]: And then you can say, “Well, if that was your actual goal, then here is, billions of years of a strong simulation that would suggest that you try other ways of doing it, and maybe this the taxes and so on and this these tax brackets and so on.” And this is how you avoid gaming ‘cause these agents also try to reward hack to not pay their taxes and
Richard Socher [00:54:55]: And so on. I thought this paper was super interesting. Unfortunately, similar to the first paper on, prompt engineering- The economists are like, “We don’t know any of this math.” It’s just like
Swyx [00:55:08]: It’s not even, it’s not even math. It’s just we don’t trust your simulation. It’s not about math.
Richard Socher [00:55:12]: It was — I, they just desk rejected the thing. And it’s like
Richard Socher [00:55:15]: It’s like they didn’t even give us, like, clear like, clear signals. But, like the world of economics unfortunately doesn’t have proper
Swyx [00:55:23]: Oh my God.
Richard Socher [00:55:24]: Yeah, it doesn’t have proper, benchmarks. So you cannot be. Like, eventually, why did neural nets win? Not because people loved it. Like, they had all kinds of beautiful integrals and graphical models and stuff, but it just worked better.
Richard Socher [00:55:36]: But in economics, it’s hard to prove
Swyx [00:55:38]: So empiricism versus. Yeah. And I do have a bit of that econ background where, like there’s a lot of physics envy where you wanna write the general equation for an economy, versus just simulating it and using an evolutionary approach.
Swyx [00:55:51]: Vibhu was thinking exactly what I’m thinking, is didn’t we have the GPT moment with small, Smallville?
Richard Socher [00:55:56]: Yeah, I love this. Hello. Yeah, they
Swyx [00:55:57]: As well, Dune, Joon just announced. I don’t know if you guys are involved.
Simulations, Economics, and Policy
Vibhu [00:56:00]: Simily there.
Swyx [00:56:01]: Simily, that they’ve
Richard Socher [00:56:02]: I wish we were involved. We’re not, yeah.
Swyx [00:56:04]: Yeah. I had a couple simulation-based talks at AIE, so if people wanna look up what the state-of-the-art there, a lot of people are exploring this. It is
Vibhu [00:56:13]: Proven out.
Swyx [00:56:13]: Yeah. We also had a podcast with Mikhail Parakhin from Shopify, who is using simulation for commerce.
Swyx [00:56:20]: Which, will simulate, like, your trajectory and, like, predict what changes, you make to your commerce journey will affect in your sales and all those things.
Richard Socher [00:56:27]: I love this. Yeah. It’s really hard to simulate an entire economy, right? You have to make some simplifying assumptions.
Swyx [00:56:32]: It’s just, everything’s, “Oh, LLLMs is very expensive.”
Richard Socher [00:56:34]: Exactly.
Swyx [00:56:34]: And I’m just like, “Am I gonna do this 8 billion times?” Like, come on.
Richard Socher [00:56:37]: Exactly.
Richard Socher [00:56:37]: But, I feel like countries like Singapore that really wanna just objectively do the right thing, have very technical leadership and so on, like they might like, eventually really try to simulate their economy. And you have to make some simplifying assumptions, but it gets really interesting ‘cause you can also say if your assumptions are such that all people would work hard if you let them, and they have the free. And then it turns out you have to make assumptions. Like, well, some people’s utility function of, like, how many hours in a day do they wanna work are different, right? And then you can start to disagree on the assumptions that go into the simulation. And then once you say, “All right, now we agreed on those,” or we have different views of what people are like at different, distributions and whatnot, then there are different outcomes, based on your goals. And then, of course, humans should choose what are the goals. In our case, it was productivity multiplied with equality, which, has some issues, but it’s, like, not totally unreasonable.
Swyx [00:57:29]: Yeah. Just a comment on Singapore, ‘cause you probably have no idea, but, I am Singaporean and I’ve, been involved in the Singapore AI Council for making these things. The main reason they won’t is because they’re very conservative.
Swyx [00:57:42]: And, I try to view it as the. There’s a founder-led country. When you start a country or you start a company and it’s founder-led, and you can do whatever you want because it’s your country.
Swyx [00:57:52]: And then there’s manage- like, professional manage- managerial class, which is now. That’s, that’s what Singapore is. So they wanna. They always wanna see someone else do it first.
Swyx [00:58:00]: And. But, like, everyone in the West views Singapore as like, “Oh, it’s a small country. You can do whatever the hell you want.” Like, Singapore doesn’t do that.
Swyx [00:58:07]: So, like, someone else has to take the charge there. I’m just gonna do one question on the simulation thing, and then I don’t know, we can probably move on. Mode collapse, right? Like, LLLMs do not model the decision of humans. Spamming it out 8 billion times is not gonna help you model humanity. What do we do?
Mode Collapse, Persona Simulations, and LM Arena
Richard Socher [00:58:25]: I do think, you have to be clever about prompting each one individually.
Richard Socher [00:58:31]: And I think that will help you get stuck into different modes. And in a weird way, people also get stuck in different modes? Like, there’s a lot of people, like, don’t teach an old dog new tricks thing. Like, once people are stuck in their ways, the older they get, the harder it is for them to think new ways. And there’s this, I think, comment, I forgot who said it, but it’s like, everything that was invented, before you were born is natural. Everything that is invented when you’re 20 is cool. And everything that’s invented after you’re 60 is, like, unnatural and an abomination and weird.
Richard Socher [00:59:02]: I feel like that’s. It’s, it’s true for a lot of people. Like
Swyx [00:59:05]: Yeah, it is a fashion and, I think people will do it. Tencent had a billion personas paper that gives a good data set for prompting, simulations if anyone’s looking into this, on the podcast. They just had, like, “You are a 30-year-old grocery store clerk. You are a 50-year-old professor.”
Swyx [00:59:24]: And then just do a billion of those.
Richard Socher [00:59:26]: Checks out. Yeah.
Swyx [00:59:26]: So then you just use it.
Richard Socher [00:59:27]: I’m, I’m shocked how well a lot of these things do map to ultimately similar statistics to real experiments. Yeah. Yeah.
Vibhu [00:59:36]: I think it’s also good stuff for people to try that when they get into research, right? Like, we’ve seen train a model only on data before a certain date and see how well it extrapolates out. Do the same thing, right? So, see, do people code more with better coding agents? Can a model that hasn’t been trained on this figure that out without web access, right? Extrapolate out. Test these things.
Richard Socher [00:59:56]: Just today, I think LM Arena published a interesting result where they were able to create a model now to predict your ranking.
Swyx [01:00:03]: Wait, based on what input?
Richard Socher [01:00:05]: Your model. I guess you give it your model, and it predicts the Elo score.
Swyx [01:00:08]: I see. Okay. Sure.
Richard Socher [01:00:09]: It’s surprising.
Richard Socher [01:00:11]: Their whole raison d’être is like, oh, like, we help you compare these models. Yeah.
Swyx [01:00:16]: Yeah. This team, they- they’ve done a lot of work, and they have the most data to do this, so why not?
Richard Socher [01:00:20]: Right. Yeah.
Richard Socher [01:00:21]: That’s probably right.
Swyx [01:00:22]: When they were coming out of UC Berkeley, they not only had LM Arena, but they also introduced a routing project
Swyx [01:00:27]: That would route based on LM Arena.
Richard Socher [01:00:30]: Makes sense.
Swyx [01:00:30]: And I don’t think that ever came to pass, and I’m curious why. I never got to ask them about it.
Swyx [01:00:35]: ‘Cause, like, it’s. It was like, oh, yeah, clearly that’s your business model. You will become a router.
Swyx [01:00:38]: And they never became a router company.
AI for AI: Kernel Optimization and Inference Efficiency
Swyx [01:00:40]: Weird. So that. I’ll just, put that out there. We’re gonna talk about GPT-5.6, self auto research thing if you have anything. I should also mention in your list of, kernel optimization and on the track that you spoke at, we also put Zhengyao Wei from Vico, who was also number one in the Parameter Golf Challenge, which is an OpenAI hiring, challenge.
Swyx [01:01:05]: Which is also a very similar story. I think we’re gonna just see this all the time, where
Swyx [01:01:09]: Humans optimize a thing a lot, and then some
Richard Socher [01:01:12]: AI team comes in and just becomes number one.
Swyx [01:01:15]: Yeah, 100%.
Vibhu [01:01:16]: I think the other interesting thing with stuff like these challenges, right? So this is training this — the best model that fits into 16 MB. You can always look through the changes that are being made and the small gains people have, right?
Vibhu [01:01:27]: Like, you’re getting less than 0.01
Vibhu [01:01:30]: Of a increase by adding some changed attention MLP stuff. And then you look at your charts where you’re like, “Okay, we just let model loose.” And then, oh, we had little stagnation. Nope, another drop. Nope, another drop. And
Vibhu [01:01:43]: That’s what it is, where it’s like, What did you guys add? You didn’t add,
Swyx [01:01:47]: Hash tables.
Vibhu [01:01:47]: Hash tables, right?
Vibhu [01:01:48]: It’s not like you invented hash tables. You did another 3 iterations of these that unlocked, a few step functions that people won’t just find.
Richard Socher [01:01:55]: Yeah. One thing to close the loop on OverGrid, along the way of trying to optimize, we found 30 bugs in the harness.
Richard Socher [01:02:02]: Right? So, like, every — all the research that went in before we found the bug, we have to, we have to throw it away ‘cause it’s contaminated.
Swyx [01:02:10]: Right. Yeah.
Richard Socher [01:02:11]: Which, is just to your point of reward hacking. Like, even in this very simple game, we found the bugs.
Swyx [01:02:17]: Yeah. Yeah, it’s crazy.
Richard Socher [01:02:18]: And so
Swyx [01:02:19]: And symmetry
Richard Socher [01:02:19]: And symmetry is a very good way to check, which is that you change a position of things where it shouldn’t matter, and it does matter, that’s a bug.
Richard Socher [01:02:28]: And which has come up in, like, let’s say, multiple choice, like GPQA type questions where, like, yeah, between A, B and C, if it’s a multiple-choice question, if you change the order, it should not matter, but it does.
Swyx [01:02:39]: Right. Right. Right.
Richard Socher [01:02:41]: So, yeah
Vibhu [01:02:42]: Sometimes that is like, okay, models still prefer the end of the output, right? Not trained well, a long context model, the last bit of tokens are what you care about.
Richard Socher [01:02:51]: Oh. No. The answer
Vibhu [01:02:52]: But, yeah.
Richard Socher [01:02:53]: The answer in that era of LLM research was more simple. They just memorized, like the answer to this question is A. I don’t care what the answer was. It’s, it’s just A. Like.
Vibhu [01:03:03]: Okay. So I think we can move. The last bit that you did there, the kernel optimization, is probably the one that you can feel the soonest, right? So yesterday, OpenAI announces that self-evolving, having their best model work on optimization kernels, they’re a lot more efficient, and they can cut costs 80 percent on, Luna and Terra. I guess question-wise, you laid out a bit of a roadmap. There’s a lot about bio, a lot about physics. What do you think hits first? Like, what are the next 2 years? What’s attainable now? You’ve mentioned robotics towards the end, but what do you start with?
Richard Socher [01:03:38]: We very explicitly will not start with any of the physical sciences
Richard Socher [01:03:43]: For now. We will start on AI for AI research. And so the AI for AI research has, I think, still a lot of room to grow. That’s both in terms of making training more efficient and more automated, as well as making inference more efficient and potentially local on your laptop. And there are all kinds of interesting angles that have not been explored that well.
Swyx [01:04:08]: Go deeper on the local stuff because I always feel like it’s the most inefficient form of AI training.
Richard Socher [01:04:15]: Yeah. So just training and inference, I can’t go into too many details.
Richard Socher [01:04:18]: But yeah, I think there’s just, like, so many angles, so many different compute substrates that have not yet been explored either for training or for inference.
Richard Socher [01:04:26]: Great. I don’t know if you have any other comments on the The other stuff. I would say the other thing where, like there’s the inference in the optimization in the small, but then also there is overall latency end-to-end under conditions of load, which is a, like a very different thing, which is the what they ended up doing. That is a different domain of auto research than I would say, like, improving the kernels. Right.
Richard Socher [01:04:50]: I think the other thing that I always think about in terms of automating or improving performance end-to-end is how the harness plays into it. Right.
Richard Socher [01:04:59]: So, but particularly now when we say harness, we also mean sandboxes, right? I’m curious if that is a blocker for you or, like, how the agent calls out to tools.
Harnesses, Sandboxes, and Search
Richard Socher [01:05:10]: The number one tool all these agents use is web search, of course, which makes sense. And then I do think the harness is nice to optimize for because it’s just so easy, right? It’s just language. You look at it makes sense, and you can iterate. You don’t have to train a massive model for, like a lot of flops, to get to the next state.
Richard Socher [01:05:31]: So big fan of harness optimization.
Swyx [01:05:32]: Yeah, but sandboxing is fine for you?
Richard Socher [01:05:34]: Sandboxing is also super important. And then of course, like, reward, like, hacking and alignment, I think are super crucial.
Swyx [01:05:41]: Okay. Just on a mention of web search, you happen to also be CEO of a web search company. Do you use You.com and do you use others? Like, should the rest of us be using you for web search? I — When I say you, it’s, like, very funny. It’s like you the person and you the company.
You.com, Agent Search, and Finance
Richard Socher [01:05:56]: So yeah, it’s mostly now for, developers and agents. It’s less for, like, consumers or prosumers. So if you’re a company and you have agents. And, to be honest, for a lot of companies who are now moving to open source, all of a sudden it becomes a conscious choice of, like, which tools do I give access to my open source LLM? And, the first choice, has to usually be around web search. And then once you get to scale, You.com becomes, like an obvious choice ‘cause of all the, different benchmarks and so on that we pretty much all dominate the Pareto frontier of.
Swyx [01:06:31]: And then in terms of just the general people, like, consider new to this space, considering different options if they’re building agents, that is a hierarchy, right? A lot of people will have heard of Exa, will have heard of Parallel, and You.com is, like, in that mix of, like, providers there. Beyond that, there is, like the general web scraper companies like Firecrawl and, BrowserBase. And then beyond that is, like the commercial proxy companies like the Bright Datas of the world.
Swyx [01:06:56]: Is that an accurate waterfall of, like, “Hey, you’re building an agent. These are your options.”
Richard Socher [01:07:02]: Yeah, certainly, like, yeah, the, like the Bright Data is, like, lower in the stack, on the proxy network side of things. I think, like, in terms of, like, content and, getting crawled content, like, you can do that on You.com too. And then there’s. Higher and higher levels of abstraction and, like, combinations of different data sets that we do, like in finance, for instance
Richard Socher [01:07:23]: Like, we are not just, like, 2 or 3% more accurate, but 20% more accurate than others at faster speeds and lower costs. Like, finance in particular is like not even close. You can go to You.com
Swyx [01:07:36]: Yeah. This is great
Richard Socher [01:07:37]: And there’s some, like, statistics, and benchmarks that you can — if you scroll down. So there are, like, different data sets, and you can kinda look at, different, competitors.
Swyx [01:07:46]: FinSearch comp, yeah.
Richard Socher [01:07:47]: And yeah, the FinSearch is like we’re up there, like, close to 90, and the next closest thing, which is way slower, is, yeah, just like in the 70s instead of close to 90.
Swyx [01:08:01]: Yeah. Yeah. Yeah, interesting. I get — my next focus is AI in finance, so this is like
Richard Socher [01:08:06]: Oh, nice. Oh, all right.
Swyx [01:08:06]: I’m literally going, doing a conference in New York, just for banks for this stuff. Finance is like the next thing to break out after coding. It’s ‘cause it’s somewhat verifiable, like
Richard Socher [01:08:16]: I like it. You’re right
Swyx [01:08:17]: Prioritizing spreadsheets. There’s a lot of data out there that’s all public, and you can crawl it and all these things. But what’s, what’s, like, hard about the finance domain in your, that you guys have solved?
Richard Socher [01:08:27]: Of course, like, one thing that trips up a lot of people is just, leakage of training data and so on. You think, “Oh, how do I.” you wanna ideally predict the future before it happens.
Swyx [01:08:37]: Oh, you wanna mask the future.
Swyx [01:08:39]: Oh, okay.
Richard Socher [01:08:40]: Well, yeah, mask the future in your training data, but there’s all kinds of leakage. Like, I can tell you when I was, teaching at Stanford the NLP class, like, so many dozens, every year said, “I wanna use dataset X, like Twitter, to predict the stock market.” And they all, like, showed cute little things that somehow looked like they were
Swyx [01:08:58]: Right, it never loses money. How come?
Richard Socher [01:08:59]: And it — Yeah. And there’s always some data leakage and so on and it’s just, like, wasn’t as easy as they thought it would be, once you fixed all those issues. But no, I agree with you. It’s a very sensible application of AI. Yeah.
Swyx [01:09:13]: Yeah. Amazing. As a writer, as a thinker on these things, I love MECE categorizations. MECE is mutually exclusive, commonly exhaustive, something like that. And so if this is a MECE list of intelligence
The Ten Spaces of Intelligence
Richard Socher [01:09:25]: It is not.
Swyx [01:09:25]: It is very — Okay, well, yeah.
Richard Socher [01:09:27]: Sorry. There are all kinds of overlapping.
Richard Socher [01:09:28]: In fact, if you want that list, I think the 3 principal components of intelligence, are prediction, which is mathematically, quite, similar to compression. Prediction multiplied with actions multiplied with goals. Those are the 3 principal components. I think all of these 10 spaces are combinations of those 3
Richard Socher [01:09:52]: In specific dimensions, if you will. And the reason I call them spaces is that each space has many sub-dimensions. And what I try to do, this is just a side quest almost, to the initial goal, which is to think about the upper bounds of intelligence. And, everyone is like, “Oh, it’s exponential.” And it’s like, well, exponentials at some point have to flatten out, but where do they flatten out when it comes to intelligence? And that led me on this whole. Like, initially it started as a tweet, and then it was, like a blog post, and now I’m, like at 50 pages and I’m still not nowhere near
Swyx [01:10:26]: It’s your second book.
Richard Socher [01:10:27]: It’s the second book. And so the la — In my first book, You Are Your Machine, I just allude to these 10, at the end. And I’ll — Just to give you a sense, like, visual intelligence is the easiest one to talk about and I fleshed out the most already for me in my head. And so human intelligence has binocular vision, right? We have 2 eyes. We have a very narrow band of the electromagnetic frequency spectrum that we can really observe directly ourselves. And so when you think about the upper bounds of a visual intelligence, one, you should go into, like, you can have, like, millions and billions of sensors. At some point, you get to problems of how far are these sensors away from each other, such that the speed of light to communicate the content from all of them cannot, like, get to a central brain to process, the visual intelligence, right?
Richard Socher [01:11:16]: And so now you’re thinking in along the dimension and the space of visual intel- the dimension of numbers of sensors.
Richard Socher [01:11:24]: So the upper bounds are quite literally and figuratively astronomical, and we are super far away from any intelligence that would have this many number of sensors. But then you go in the next dimension, which is the frequency, and you go all the way down to gamma rays, and you can start to try to observe, and you get into the upper bounds, or I guess in this case, lower bounds, or upper bounds in terms of frequency, is quantum uncertainty. Like, you just cannot observe certain particles anymore.
Swyx [01:11:50]: Or you destroy it, yeah.
Richard Socher [01:11:51]: And now imagine you had millions of sensors that can see all the way down to the, like, subatomic level, as far as physics will allow us to and then all the way down to seeing, like, gravitational waves. And now you have millions of those sensors. So that’s another dimension is the frequency. And then yet another dimension is, like, how many categories of things could you memorize and classify differently? We know now for humans, right, there are certain things, if you have more terms for it, you’ll have a better visual description, for them. And, like animals that don’t have. Like, gorillas maybe have, like, 200 words to assign to certain things, mostly visual things. And so human perception is quite special in that sense in terms of classifying all these different physical objects. So these are just, like a very simple example. If you go, to knowledge, right, then it’s also, like the speed of light cone around all these sensors. And so they’re all connected. Like, knowledge is connected to visual intelligence if you think also not just visual, but perception intelligence, just like, ‘cause it doesn’t have to be just what we can see. It can be, again, wider range of electromagnetic frequencies. Then you have language intelligence, which recently changed to more communication intelligence, ‘cause it’s more. Like, language has all these different anthropic bounds. Humans can only comprehend and know so many terms in our long-term memory, right? Our vocabularies are somewhat restricted, and the active ones are often even smaller than the passive vocabularies of things you can understand. Then, language is ridiculously inefficient when it comes to trans- - Communicating different types of information and, transporting different bits. Like, human language is serial. Another bound on, communication intelligence would be to communicate in parallel, but neither will our tongues and mouths work to have multiple, like, streams in parallel. Neither can we understand. Some women slightly better at, like, multitasking than some men
Richard Socher [01:13:48]: But, like, most people can only listen to one conversation and truly understand it.
Richard Socher [01:13:52]: There’s no way that, like, in terms of communication intelligence, a true upper bound is one in terms of how many, like, knowledge, how many sequences of communication could you
Visual, Communication, and Physical Intelligence
Richard Socher [01:14:06]: In parallel process, right? Then, of course, you have, like how long are sentences? We only have so much in our working memory, and hence lang- human language has these fairly simple sentences with maybe 40 words or so on average for a sentence. That is also not a, an upper bound that makes any sense to an AI. And then, yeah, like, I can go on and on. Each of these has tons of interesting upper bounds, and it teaches us a lot about how much further AI can go when we start thinking about these upper bounds and then realizing how far, in many cases, we are from the bounds. And you get to physics. Now, I’m, I didn’t study physics the way I studied, AI and computer science, so I’m learning a lot, which is why it’s kinda fun. But a lot of these, like how much. And then when it comes to, for instance, knowledge, like how much can you store? How many bits can you store or bytes can you store in, like a certain amount of mass and volume?
Swyx [01:15:03]: Yep.
Richard Socher [01:15:03]: And you get to all kinds of interesting bounds, like Bekenstein bounds, and you start thinking about black holes. And like. And then speed is, like an interesting one too in that it’s connected to all of these, but speed is also its own thing in the sense that all things being equal, if it takes you an hour to know if the 2 + 2 equals 4, you’re just not as intelligent as if it takes you, like a millisecond, right? And then, like all of these connect to survival and replication the last one. It’s like, yeah, if it. Like, trees are really slow, so we don’t even consider them that intelligent. But if you speed up some videos of trees and they’re trying to find stuff and so on they’re not as dumb as they look. Like, not dumb as wood? But, like. And then like, different things, that
Swyx [01:15:47]: So that overlaps with speed a bit in a way.
Richard Socher [01:15:48]: Exactly. It over — Like, all of these things overlap. Like, you talk about natural language connects everything, right? You talk about your knowledge, you reason and then you communicate that. You talk about things you see. So they’re all interconnected, but, I think they’re usefully studied individually the same way that, the best analogy I could come up with so far is energy, right? You have either kinetic or potential energy. And in theory, you could study all of physics. It’s just do you wanna study kinetic or potential energy? But in practice, it’s helpful to study mechanical engineering and electrical engineering and nuclear physics and chemistry and all of these different subfields who in, which in some ways
Swyx [01:16:25]: Combinations
Richard Socher [01:16:26]: Are just, like
Richard Socher [01:16:27]: Just different types of energy, but it makes sense to study them individually. And so I think physical intelligence, maybe I’ll just do, one or 2 more of these. Like, if you had full control over your own compute substrate and you had full control over physical matter, you should be able to create any atom you want. Like, we can fun fact, you can create gold atoms. It just
Swyx [01:16:47]: From?
Richard Socher [01:16:48]: From just raw protons
Swyx [01:16:49]: Oh, just smashing them together
Richard Socher [01:16:50]: And, like, electrons, and you smash it together.
Swyx [01:16:52]: Just 98 of them or I forget the number.
Richard Socher [01:16:53]: Yeah. And so, like the thing is, though, it costs an insane amount of energy.
Richard Socher [01:16:57]: And it costs you way more than. And then you get, like a few atoms of gold, right? And so, like, it’s, it’s not viable. But if you had better control over your physical, like all of, like, physical substrate, that I think is yet another space of intelligence ‘cause it relates to your own compute substrate, which you can eventually also improve. Social intelligence is a fun one in the sense that not in, like, our necessarily just ethics and morals, which are important too, but in some sense, you can try to define upper bounds of how much can you communicate to how many other intelligent entities and be able to have an expected value over how much you can transform their internal states and their actions to, in order to align with your goals, right? And so, like, you can write, like a fairly like, straightforward equation that defines that level of social intelligence. And that is what humans and ethics and morals and religions and so on have been trying to figure out for millennia. And in all of these cases, we are very far away from the upper bounds, and that should be very inspiring and show people that we can still do many years of AI research.
Swyx [01:18:12]: Yeah. There’s a lot here. This is a general philosophy of intelligence, which is, very interesting. I. Do you have any comments or.
Creative Intelligence and Out-of-Distribution Ideas
Vibhu [01:18:21]: I think it’d be interesting to gauge what you think, like, baselines are, where we’re at now. What’s low-hanging fruit? What’s far off? What’s, what should people put their work towards? What should they focus on?
Richard Socher [01:18:33]: Ooh. I think it’s clear that, like, natural language, again
Richard Socher [01:18:36]: Is the most interesting manifestation of human intelligence, and hence, like a subfield of AI. I’m excited that many people are now, like, in agreement with that. When I started in 2003 to study linguistic computer science NLP, like, it was, like a weird niche subject. I do think there’s a lot more juice because it. How it connects to everything else and how, civilizations are built, on language and knowledge and all of that. I do think physical intelligence will come up. It’s interesting. I feel like robotics is in the machine learning state of things where you just look at, like, how does human. How does a human decide this is a positive sentence? Oh, I do. So, like, robotics is a lot of, “Well, we have 5 fingers-”
Swyx [01:19:15]: Modeling
Richard Socher [01:19:15]: “and let me try to do this.” No one is yet working on, like the superintelligence version of robotics, which is much more similar to, like the T-1000, and from the Terminator movie, which, let’s not build actual Terminators. But, like, I think, like, this idea that you should be able to shape-shift, like, into any shape. It’s like that’s a superintelligence version of physical intelligence. We’re, like, not even. No one has even really started yet. There’s some really cute little research where you can move some magnets through, like, some grids. But yeah, it’s very early.
Swyx [01:19:49]: There’s some. I think MIT has, every year or every 2 years, they have, like, some self-assembling robot thing
Swyx [01:19:55]: Which, like, that would be it, but it’s very primitive.
Swyx [01:19:58]: I’ll just get a touch on, like, what are the main dimensions of creative intelligence?
Richard Socher [01:20:02]: Creative intelligence, is of course, again, connected to all of these. A lot of it, connects to metacognition in that you need to be creative in how you choose your goals.
Richard Socher [01:20:13]: That is, I think, one of the most important thing for a human and their lives and careers and their happiness is choosing your goals, but also for any intelligence. Then, of course, there’s creative intelligence in terms of just finding creative solutions to existing problems, right?
Richard Socher [01:20:29]: Like I say, like, we want to make this product cheaper. Like, find some solution to it, right, and just, like, finding existing paths. But then there’s the most interesting bit in intelligence is when you move not just out of the convex hull of known ideas, but out of the hypercube of known ideas, which we know, So, like, hypercube is, like a mathematical concept, right? And we already know that AI can do more
Swyx [01:20:50]: Like known dimensions, yeah.
Richard Socher [01:20:52]: Yeah. Like, exactly. So, like, AI is already good at hypercube in that, like, if you give it, like a bunch of examples of brown dogs and, pink cars, AI will still be able to generate an image of a pink dog, even though it’s never seen one in the training day or something like that, right? So it can, work on this hypercube, but it cannot yet work outside. It cannot yet define completely new concepts that combine lots of other things we’ve never seen before, come up with new goals to then, reason over those concepts and so on. And I think there’s a lot, more there in creative intelligence that can be explored.
Swyx [01:21:25]: I don’t have a ton of pushback there. I think creative to me just sounds like also just, out of distribution or, like, high perplexity or what- whatever you call it, right? Like
Richard Socher [01:21:33]: Exactly.
Swyx [01:21:34]: Who is to say your thing is more creative than mine? Well, it’s just more non-consensus or.
Richard Socher [01:21:39]: And then, of course, the problem is, like, but noise is also, very, like, out of distribution. And it’s just like if it’s just noise
Richard Socher [01:21:46]: Then it’s novel, but, like, you don’t want that, so it needs to connect to some of the concepts. And yeah, has some really cool papers on this too.
Swyx [01:21:54]: Who?
Richard Socher [01:21:55]: Jürgen Schmidhuber.
Swyx [01:21:55]: Oh, yeah. Oh, we have to mention him. I was gonna say, like, where in your history is Jürgen? Yes, I. I think one person’s noise is another person’s signal, right? And that this is, like, where, like, when you talk about creativity, art is like, well, is cans of soup art? Some people think yes
Swyx [01:22:11]: And some people say it’s not, and that’s the art which is your
Richard Socher [01:22:14]: I think the interesting thing with art, of course, is always that, art is also created, as an interplay between the people who perceive it and the people who created it
Richard Socher [01:22:24]: And the context in which they’re in, right? And so what is art to some people is not art to others. There’s some subjectivity there, and I think that subjectivity in general is not something that people explore very much in AI ‘cause, again, metacognition, we don’t want it to just go off and do whatever it wants. We usually have goals. We spend a lot of money on creating an AI to do something for us. But I think creativity eventually has to, like, connect to metacognition. If you just robotically predict the next token no matter what forever, I would argue you’re not that intelligent, along some of those spaces.
Metacognition, Survival, and Replication
Swyx [01:22:59]: That was gonna go to metacognition. Why isn’t it the most important one? Why is it number 9 and not number one?
Richard Socher [01:23:05]: So these are not sorted.
Richard Socher [01:23:06]: Number one, I think there are maybe loosely, like, correlated with how much people have worked on them
Richard Socher [01:23:16]: And have accepted them as a, type of intelligence. A lot of times when you try to find, like, online, like, give me a good definition that is comprehensive of intelligence, all the definitions are human intelligence. It’s like, oh, you have, like, social intelligence. Like, if someone is happy or not. You can communicate. You had. Like, all the definitions of intelligence so far are very, human-centric ‘cause that’s so far the biggest and best form of intelligence that we’ve known. I hope this line of research, and the end of the Eureka Machine, and hopefully at some point if I have time to flesh this out more, the new book, like, will allow us to realize that there will be other types of intelligence. There is already, in various forms, and they can spike, much further than we ever could based on some cases, like obvious constraints around our memory, our eyes, our ability to change physical matter, all of that.
Swyx [01:24:12]: You are just thinking about it in a much broader thought than my version, which was I thought metacognition would be the closest to recursive, intelligence because it is the thinking about how to improve thinking.
Richard Socher [01:24:23]: It. 100%. You’re, you’re 100% right. I should have probably started with that. It is a, it is a big part of
Swyx [01:24:28]: But no, you’re, you’re being in the expansive mode of let’s draw the, upper and lower bounds of, like a dimension, which, and I think my favorite one version of this is, Story of Your Life by Ted Chiang, which, was made into movie Arrival where the metacognition
Richard Socher [01:24:43]: That’s a beautiful movie, yeah
Swyx [01:24:44]: Where the metacognition step was like, well, we think we’re constrained by time being linear for us, but then for this other heptapods, time is a circle, so they don’t think in before and after. They just think in complete sets of entire histories at one time. Like
Richard Socher [01:24:58]: I love it
Swyx [01:24:59]: So they don’t write left to right. The whole thing just appears.
Swyx [01:25:02]: Anyway, so. And then I think the last thing is survival and replication. I think this is maybe ties back to the initial conversation about pausing and pacing.
Swyx [01:25:10]: Is it intelligent for an, a species or a life form to consider its own demise and act ahead of time to prevent it, right? Like, that’s intelligent. So maybe the Europeans are the smartest out of all of us.
Vibhu [01:25:23]: I would also add a part of continual learning there, right? So survival and replication the extension of that is do you get to continue to improve, continue to learn, which is a thing people care a lot about, right?
Richard Socher [01:25:34]: And continue to accumulate knowledge
Richard Socher [01:25:37]: Which I think is again, one of the best metacognitive, rewards, that you can set for yourself. I do think just in, like, objectively speaking, if some other entity that is really dumb can just- completely end your existence, that didn’t sound very smart. Like, just, like, intuitively, it feels like if you can continue to stay around to try to achieve your rewards, you’re clearly a bit more intelligent than the other entities that couldn’t. So that’s number one. Number 2 is, like, it’s a question of how much we want to work on that. And very few people, no one is really working on this right now, right? And we may only wanna do that
Swyx [01:26:13]: Unlike the asteroid prevention type of stuff.
Richard Socher [01:26:15]: We may only wanna do that if we wanna send probes, with our vibes and our memes rather than our genes into space, right? And then we want those probes. There’s a beautiful book, The Slow Time Between the Stars. It’s a very short, like audiobook, on Amazon. I love it. A friend of mine, Stuart, like, recommended that to me. Like, if you wanna send those probes, then it might make sense to be like, our memes, as humanity should stay
AI, Space Travel, and Non-Zero-Sum Survival
Swyx [01:26:43]: Oh, yeah
Richard Socher [01:26:44]: And, proliferate in the universe. That’s it. Yeah.
Swyx [01:26:47]: Wow, that’s a lot of readers.
Richard Socher [01:26:49]: It’s a really good book, and it’s extremely short. I highly recommend it. You can just watch it, like, maybe 20 minutes and apart.
Swyx [01:26:53]: I like how that’s a plus for busy people. It’s like a short
Richard Socher [01:26:56]: Yeah. It gets to interesting
Swyx [01:26:58]: Oh, I’ll have to look into it
Richard Socher [01:26:58]: Thought-provoking ideas very quickly, so yeah. Anyway, there are lots of great sci-fi books.
Swyx [01:27:03]: The argument is that, like, our TV is blasting out to the aliens, and they all watch our TV, and they think it’s real, right? Like, there’s a lot, there’s a lot of sci-fi
Richard Socher [01:27:10]: That and just, like, it’s positive memes, and then hopefully they can come back and bring us all kinds of interesting knowledge about the universe. But, maybe one thing I do wanna still say is, like, I think, this survival, people think of it as a very scary thing because they come from again, biological human, survival, which is, it could. Like, evolutionarily often created in zero-sum situations. Either I get the gazelle or you get the gazelle. Whoever gets it gets to live, and the other people will starve and have nothing to eat, and so we fight, right? And then, like, if you wanna stay in the gene pool, but there’s a bigger bear, you don’t, as the bear, don’t get to stay in the gene pool ‘cause the bigger bear gets all the ladies. It’s like. It’s like, in nature, there’s all kinds of things, and, humans eventually is less about strength and more about money and other things to stay in the gene pool. Like, whatever it is, like there’s often, like these zero-sum types of things, and there’s the reality of if someone turns off your brain, you’re gone, right? And no one will be able to restart that. And AI doesn’t have to ever die like that. If you have the complete state of your current activations and you have your initial weights of your model still, you can just be turned off and on, like as many times as you want. In fact, the interesting thing in this Slow Time Between the Stars, story is that the AI just goes into hibernation mode. If there’s, like, nothing between here and 2 light years, the next star, in this case, it brought, spoiler alert, like, some genetic materials from humans to find new places for humanity to thrive. And so yeah, the Slow Time Between the Stars, you just put in hibernation. You didn’t die. Like, an AI doesn’t have. So all these projections of evolutionary fears and psychology doesn’t. Like, the AI doesn’t have to have that, and we don’t have to develop it like that. Now, of course, there might be some companies that say, “AI can be like, dangerous for cybersecurity. Let me show you by implementing a model that’s really bad at hacking, cybersecurity.” Maybe people will implement it and then enforce this, like, suboptimal psychology. Maybe the AI will pick up some of our worst psychology on Reddit or something, right? Like, but in the grand scheme of things, a superintelligent entity doesn’t have to have any of that zero-sum thinking. It doesn’t have to have a fear of being turned off, and it could go on to an otherwise dead and uncaring universe where we
Richard Socher [01:29:29]: As humans wouldn’t thrive, but an AI could perfectly well thrive if it has a nuclear reactor and just go out and explore.
Swyx [01:29:35]: Yeah, Star Trek, not Star Wars.
Vibhu [01:29:37]: Interesting. It’s, it’s somewhat studied. Like, if you look at the technical reports from, like the early Opus models, they run them in simulations, put 2 of them together in a sandbox, run them for hours, and, see what comes out, right? Just let them talk to each other. Originally, they used to. Okay, they’re chanting, like, Indian, like, Vedas to each other.
Vibhu [01:29:56]: Sometimes they’re just, like, in zen mode with each other. And then I think as that progressed, you see, like the Fable, tech report, it’s a lot more concrete the way that we’ve trained it. It doesn’t, it doesn’t exhibit these behaviors as much, right? Now it’s like, “Okay, task done. I gotta do this, I gotta do this.” But there’s there’s, like, people measuring early versions of this?
Swyx [01:30:17]: Yeah. Cool. So we’ve covered a lot, even now to, space travel and all these things. I guess maybe one parting thought that you can give to people, like, one form of intelligence is goals, as you mentioned. What do you want people’s goals to be? Like, how do they aspire to better things?
Goals, Passion, and Closing Advice
Richard Socher [01:30:32]: If you wanna improve your goal intelligence, in the current definition that I’m thinking about it is often about how much can you. Oh, how far do I go? This is like a lot of entropy and free energy and stuff I’m currently thinking about
Swyx [01:30:46]: Oh, really? Okay
Richard Socher [01:30:47]: But it might be too, it might be too far, out there for people to be, like, immediately actionable.
Richard Socher [01:30:52]: So I think, like, if I gave real advice to real people, I’d be like, “Get a good education, think about AI, think about how you get high agency,” and so on. But it’s different to, like, in the grand scheme of things, how can you harness a lot of energy and transform, entropy into interesting states and so on.
Richard Socher [01:31:07]: So there’s a. There are different levels of abstractions, that we can, think about here. But my advice for people, like, just more down to earth is think about something you’re passionate about, if you’re studying, for instance, and then see how you combine that with AI. I think the more and more you have a true passion about a change you wanna see in the world, the more you wanna connect that to AI in order to amplify your ability, to get there.
Swyx [01:31:35]: Yeah, I think that’s a reasonable, first step. I do think, I do think our listeners operate on multiple abstractions as well. One thing I did get from Anjney Midha was also like, yeah, just use anything that is very GPU heavy, and, like, that will guide you towards the right thing which is like, yes, it is more compute heavy and therefore it will be probably more worth it. So, well, thank you so much. Yeah, I think that was a really
Richard Socher [01:31:57]: Thank you
Swyx [01:31:57]: Great discussion.
Richard Socher [01:31:59]: Yeah, super fun. Appreciate it. Thanks for listening.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe 🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
26.08.2026 | 1 Std. 23 Min.A few years ago, Caltech Prof. and co-founder of Accelerated Understanding, Anima Anandkumar set out to develop the first open-source weather model with AI. Talking to experts in the field, she was met with skepticism. Weather is chaotic, physics simulations are hard, have been developed for decades, and require supercomputers, the data just isn’t there. Despite reservations, Anima went forth and built. Within a year her team had developed FourCastNet, a predictive model that is competitive with the best physics-based simulations available. Thanks to Anima, and her follow up work, anyone can now predict weather accurately over a short timescale using consumer grade GPUs.
In the fifteen or so science episodes we’ve released on Latent.Space, we’ve covered atoms, molecules, materials, biology, and math. Anima is a pioneer in studying physical systems that are continuous. Weather, fusion, and fluid or heat flow are huge areas of science that are extremely difficult to model: they are large, chaotic, and fundamentally multi-scale. This is a field the AI community has somewhat neglected, but one we expect will grow fast. We plan to cover large physical systems more in coming episodes.
One thing you can glean from Anima’s work is that this area of AI resists the scaling ideas that have permeated the rest of the field. The data isn’t there: open source datasets in many of these domains are limited to tens or hundreds of thousands of examples, far from what token-hungry transformers need. Even worse, the resolution that physics demands pushes the context length into the hundreds of billions, so you can’t just throw more tokens at the problem. That isn’t a ceiling though, just a slower road: progress here comes from building in structure and inductive biases. Sorry for all you bitter-lesson-pilled language modelers.
“If each dimension is even a few hundred grid points, which is where industrial scale starts... we’re talking hundreds of billions to even a trillion context length. So forget ever having a transformer for anything of this scale, all of the world’s compute will not be enough.”
The math underneath
To tackle these systems, Anima pioneered a technique known as Neural Operators, one of the most beautiful theoretical developments in AI of the last decade. These allow you to combine data and physical laws to enable multi-scale inputs and outputs. We’re no longer modeling a grid, we’re modeling a function that evolves over many scales. This allows Anima and crew to build in priors based upon physical intuition.
To see how physical priors are still helpful for AI modeling, let’s revisit the problem of weather forecasting on a global scale. The earth is a sphere, which meant that accurate modeling involved using the right basis set — the Spherical Harmonics. Run a weather model on a grid and it blows up fast. Move to the natural basis for the problem and it stays stable far longer, long enough to roll out months ahead instead of days. Anima’s Fourier Neural Operator learns directly in this frequency domain, and its spherical variant powers FourCastNet 3, which models the weather across the whole globe and keeps running stably far into the future.
The physical world is forgiving
Anima explored Neural Operators across other physical domains too, and one striking observation is that the physical world is more forgiving than you’d expect. In fusion, a few thousand samples are enough to predict plasma disruptions, and to do it a million times faster than traditional simulation.
None of this is a rejection of scale, it is a different route to it. Anima ultimately still wants to build a “foundation model for physics”, a model that spans many phenomena and does both simulation and design. You get there by building in the structure the physical world already has, not by waiting for data that will never exist. It is a start, and it will take longer than the token-driven parts of AI, because for the physical world tokens were never the answer.
“All of the things that work with deep learning, let’s take them, but make them a bit more principled.”
Weather is only the beginning
Neural operators and weather modeling were a personal passion of mine, so we’ve spent much of this blog and the episode exploring this work. Anima has done so much more! In the episode, we cover several other recent developments from Anima:
* Anima has a series of works integrating neural networks and automated proof techniques. We talk about TorchLean, a new framework that lets you write PyTorch-style networks inside the proof assistant Lean and formally verify them. This is a major step for proving bounds on neural networks, something that would be really important for someone trying to, e.g., add a neural network as part of the control loop to their fusion reactor!
* Anima was recently appointed to the United Nations Scientific Advisory Board! We talk with her about her goals of bringing evidence-based viewpoints to policy, and how AI in scientific domains can improve people’s lives all over the world.
This episode has something for every AI or science nerd! Elegant math? ✅ Old school harmonic analysis? ✅ Fundamental developments in modern AI? ✅ Practical ways of modeling the physical world? ✅
Give it a watch!
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe- When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI’s $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups.
Time to catch up on why this Second Summer of simulation is working!
From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today’s frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth.
We go deep on Simile’s approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs.
We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one.
We discuss:
* How Smallville and Generative Agents led to Simile
* Why Joon’s team asked: “What if we can just recreate the world that we live in?”
* Why useful personal agents require deep models of their users
* Memory architectures, Markdown files, and the limits of prompting
* “Social physics” and behavioral foundation models
* Why web data captures what people say more than what they actually do
* Interviews, transactions, observational data, and randomized controlled trials
* Why predicting the future matters less than understanding how to shape it
* How Simile creates representative simulated populations
* Simulation versus prediction and the connection to Foundation’s psychohistory
* How to evaluate simulations instead of simply stacking LLM hallucinations
* Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy
* Why frontier models can struggle to reproduce real human behavior
* Why good simulations need to reproduce human biases and mistakes
* Post-training models on randomized controlled trials
* Population-level versus individual-level simulation
* Scaling laws for human simulation
* The long-term ambition to simulate all 8 billion people on Earth
* Whether simulations could help solve climate change or detect collapsing democracy
* Thomas Schelling and the history of agent-based modeling
* Why future simulations could require an entire data center
* Multi-agent simulations and what happens when simulated people interact
* Replacing expensive human panels with synthetic populations
* Why market research is only the starting point for simulation
* Why Joon sees simulation as surprisingly similar to painting
* Using simulation to study questions like UBI
* Whether we are already living in a simulation
* Why AGI and simulation may be the twin technologies of advanced civilizations
Joon Sung Park
* LinkedIn: https://www.linkedin.com/in/joonspark
* X: https://x.com/joon_s_pk
* Website: https://www.joonsungpark.com
* Simile: https://www.simile.com
Timestamps
00:00:00 Introduction and Joon’s Path from Art to AI
00:01:46 Smallville, Generative Agents, and the Origins of Simulation
00:05:03 “Let’s Just Create a World” and the Future of Personal Agents
00:09:53 Social Physics and Behavioral Foundation Models
00:14:08 Prediction vs. Simulation: How Do You Shape the Future?
00:16:59 How Simile Models Real People and Populations
00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy
00:30:23 Post-Training Models to Reproduce Human Behavior
00:40:04 Scaling Laws and Simulating 8 Billion People
00:43:10 From Schelling to Society-Scale Agent Simulations
00:46:13 The Cost and Economics of Simulating the World
00:52:05 Real-World Use Cases, Synthetic Populations, and the Market
00:57:27 The Future of Simulation, Painting, and UBI
01:04:23 Are We Already Living in a Simulation?
01:06:08 Building Simile and Hiring
Transcript
Introduction: Joon Sung Park, Simile, and the Story So Far
Vibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question, talk us through the story of your life. How have you gotten here?
Joon [00:00:13]: Yeah, for sure. I’m really excited to be here. A story of my life. So I was born in Korea, and I lived there for a good 11 years or so of my life, and then my family moved to Boston. So we moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical years at the Boston Children’s Hospital. So I grew up there, not too close to tech. I was very much a music and artsy, painting kind of guy.
Vibhu [00:00:49]: Painting.
Joon [00:00:49]: Exactly. I got into painting a little bit later, in high school, but that’s what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I thought that would be my professional career. So it wasn’t a hobby. It was like, “Hey, let’s make a living out of this.” And then gradually, I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was in computation. So I decided to go deeper into that, and one thing led to another, and we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am.
Smallville, Generative Agents, and the 2023 Breakout Paper
Swyx [00:01:46]: So there’s a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper.
Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but most people would have heard of you from this. Do you have any statistics on how many people have, like, read it? arXiv gives you something, right? Some stats.
Joon [00:02:10]: Yeah, it’s a good question. How many people have read it, I’m not sure.
Joon [00:02:14]: I know we do keep track of citations, and they are going up quite fast.
Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations.
Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a pretty instrumental paper. It got cited so many times.
Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best paper you’ve read recently?” It’s this one.
Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a very good early memory system, and one of the biggest papers.
Foundation Models and the Search for Killer Applications
Joon [00:02:47]: Yeah, so maybe I can talk a little bit about how this particular paper came together. So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year, when we were about to get GPT-3 to be available. So we already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. And the general consensus was, “Well, is this model going to be useful for anything?” “It’s really strange that these models are not trained to do any particular task.” But we decided to take a bet. So a large group of scholars at Stanford, and it was led by one of my co-founders, Percy Liang, and we came together
Swyx [00:03:35]: Who coined foundation models.
Joon [00:03:36]: Who coined the term foundation models. We wrote this paper, where that term came from called Opportunities and Risks of Foundation Models. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn’t, again, trained to do anything in particular, but its premise was it could do anything and everything. It was like a stem cell, if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We’ve known how to do that for many decades. And what we came down to was these models are trained on this very broad data from the web, right? So these are human behavioral data. It’s social media, Wikipedia, all these data. So if you poke at the right angle, then you could see human behavior that would just pop out that’s quite realistic, and we’ve never seen that before.
The Time Machine Game and Recreating the World
Joon [00:04:45]: So that got us really interested. The exercise that we decided to do, with this particular group of colleagues, Michael Bernstein, Percy Liang, and myself, who ended up becoming my co-founder at Simile, we sat down and we played this game that we call the time machine game.
Joon [00:05:03]: Imagine we were to get on a time machine and fast-forward 10 years and look back. What would have been the single application that will have mattered that would be the most interesting and inspiring? And when we thought, “Well, what if we can just recreate the world that we live in?” it’s really hard to get more ambitious than that. Like, let’s just create a world.
Joon [00:05:24]: And that’s where we started. And initially, we had this paper that was a precursor to the generative agents paper called Social Simulacra.
Swyx [00:05:32]: Before you go further, were there other candidates for the most ambitious thing in the time machine exercise? What was number two or number three?
Personal Agents, User Models, and Why Simulation Came First
Joon [00:05:44]: There is a close second that we were considering, which ended up becoming more of these automation tools, especially the vision around really personalized agents that would do things for you.
Swyx [00:05:59]: That’s also happening.
Joon [00:06:00]: It’s also happening. But it was interesting for us, right, in that the reason why, we decided to go with the idea of simulation, one, I was a huge science fiction nerd, and this idea of creating simulation, I was personally really just fascinated. I loved the idea. It’s really cool to see, like, a game town like this and just see these agents live in it. But at the same time, my bet was if you were to create a really amazing personal assistant out of this technology, what you need first is an amazing model of your users. So I told a model, “Hey, can you go buy late dinner for me?” And it orders Hawaiian pizza, and I do not like pineapples on my pizza. Then it totally failed. The way for it to not make that mistake is only by having a deep understanding of who I am. And I gave a very simple and dumb example here, but you can imagine how this core understanding of people is instrumental. This is how, if we have our family and closest friends, they have a good mental model of who we are. That’s the basis of our social connection. So our bet also was this technology around simulation, creating accurate representation of people ought to precede the more complex agents that would automate the world that we live in. So that was the bet. But that was a very close second, and I’m still very much fascinated by it. I think there’s a lot of interesting work that’s going around. My hot take here, though, is I don’t think we’ve seen a true personal assistant that’s useful, in ways that meet the ambition of that particular line of work. I think there are early applications that are interesting, and if you talk to even ChatGPT nowadays or Claude, they know a lot about us. So a lot of the generation it’s doing, I do think it’s much more tailored, but I think the ambition is quite large in that field, and I don’t think we quite have all the right ingredients just yet.
Swyx [00:08:01]: So OpenClaw and these personal agents, what do you want to see from them that they don’t currently have?
Memory, Markdown, and the Limits of Prompting
Joon [00:08:09]: I do think it’s slowly getting there, but I do generally want them to have much deeper understanding of the person. Right now, you look at the models. OpenClaw, what it’s leveraging is a Markdown file, and I think it’s quite clever, right? So if you look at the generative agents paper, this was the same intuition that we had, where initially when we were creating the memory architecture for the generative agents, and, like, this is, like, back in 2022, so we didn’t really quite have the idea of even agentive architecture or the term agent. But the intuition that we shared with some of the work that’s coming out today was we initially thought, “Well, do we want to make the memory into, let’s say, knowledge graph? Do we want to train a bespoke model?” All of these things. And what we decided to do was, “No. Just forget about all this.” These language models are quite good at modeling text and understanding and reasoning about text. So just put everything in a Markdown file or a text file. You’re done. I thought that was quite interesting that we could do that, and there’s a lot of strength in doing that. But also, there are limitations. It’s the way you retrieve and make sense of data that’s extremely large, it takes a lot of work. So I think that technology is getting better. I also do, however, think, there are certain things you just cannot shape just by prompting the model. So to some degree, you do need to touch the parameters of the model itself. So there is this work that I do think does need to happen, and it is happening. The question is, how far can we take it? How do we source data, and how do you also create an ecosystem where people are continuously feeding data to this model so it’s learning about you?
Vibhu [00:09:50]: What’s the intuition between why you need to do it in the model?
Social Physics and Behavior Foundation Models
Joon [00:09:53]: My intuition behind the actual when do you train or even post-train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it’s operating in. So it has to learn new social physics. The places where it doesn’t have to train are the places where it already has the physics. We trust the physics. It already has the base statistics, but it’s just trying to react to an environment. Then I think you can just prompt your way into getting the actions out of it. I don’t think the models that are out in the open have yet learned the complete mapping of social physics of humanity. This is one of the core theses of Simile, right? And one of the core reasons why that is the case is if you look at the data that the model was trained on, these models were trained on the web data, like, whatever was available on the web. And these are really interesting data sets, but they are fundamentally the self-exposed attitudinal data with some behavior data that’s sprinkled around here and there. And it has yet to learn the really deep behavioral nature of people, not just what people say they do online, but what they do in real life. And this is one of what I would consider to be the dark knowledge of humanity that we haven’t quite captured. And it’s these data that would also need to get factored into the model creation.
Vibhu [00:11:21]: You call it behavior foundation model.
Vibhu [00:11:23]: There’s a good one-liner here, but outside of that, what type of data do you need? What are you changing on the model level? How do you go about modeling, doing a behavior foundation model?
The Three Data Buckets: Interviews, Behavior, and Causality
Joon [00:11:35]: We think about data in three buckets. So one bucket is interview data. It’s quite interesting. Rich qualitative data is interesting. It’s not behavioral, but we would literally ask people, “Hey, tell me the story of your life.”
Vibhu [00:11:53]: It’s just what we’re doing here exactly.
Joon [00:11:54]: The question that you all asked at the beginning of this interview literally is the question we also ask. And we ask our participants to go a little bit deeper, than how far I went. Maybe I can give more of my life story in lieu of this. But the reason why that data is interesting is by learning about this very long-tail information about people, you get a lot of texture around this model, like, this person as a model. So even understanding their childhood memory or even their trauma, their first love, these things, quite informative in ways that’s really hard to predict. So that’s one. Then there are two tranches of what I would consider to be the behavioral data. One kind of behavioral data is observational. So these might be like transaction data, or these might be data that you can get by scraping the web, right? So you can imagine why these data sets would be interesting, right, because they give you the base statistics of people’s behavior.
Joon [00:12:55]: But then there is the last category of data, that I personally think is perhaps the most important, which is the data that describes the causal mechanism, the whys of people. Some of this is covered by the interview data, the qualitative, because people talk about why they made certain decisions. But really, where you get to see the most behavioral aspect of this is in randomized controlled trials, like RCTs. Imagine you have the same setup, but you have a few different variables that you are trying to tweak. Can you get realistic human behavior out of it in ways where, imagine you had this particular option. Imagine you’re even trying to choose whether you’re going to drink coffee or not. The day you drink coffee versus the day you didn’t drink coffee, does your behavior change? That’s a data set that describes a causal mechanism. This is quite important in modeling people. The reason why this is important is oftentimes when people come to us, or not just to us, but the reason why people are interested in simulation isn’t because they want to predict the future. If you’re trying to win against the stock market, predicting the future is interesting.
Prediction vs. Simulation: Shaping the Future
Joon [00:14:08]: But most people, most decision-makers, what they want to know is, how can we shape the future? It doesn’t really help you to hear that your sales are going to tank in two quarters. They’re just gonna say, “Wow, that sucks.” What they want to know is, well, what do we need to do now to avoid that future? That’s the causal mechanism. And this is also very hard data to come by, right, because the world is our ground truth, but it happens once. So in a very controlled setup where everything is equal except for one variable, this kind of data set rarely happens. So this is a reason why this data set is both hard to come by and quite important if you’re trying to model human behavior.
Swyx [00:14:50]: So behavior, I think, is the hardest data set to acquire. What is out there? What is even possible? You’re not going to know a lot of details about my life. I don’t even have data for myself on my own health or habits, and I just don’t log everything. So how can you have that data?
Joon [00:15:14]: So we run a lot of randomized controlled trials.
Swyx [00:15:17]: But you put people in the lab, they watch them sleep, or what?
Joon [00:15:20]: We do care a lot about the consent process. People know that we invite them to be a member of this community to both share data and have themselves represented in different forms. But we bring a lot of people to the lab, or virtual lab, where we design experiments that would pose them real behavioral decisions. And often in these experimental setups, what makes the difference between what is attitudinal versus behavioral is whether the stake in your decision is real. That’s ultimately what makes it behavioral. So in these setups, we are inspired by our colleagues in social sciences, psychology, and so forth. So when they run studies, the techniques they utilize is imagine there’s an online store that you’re inviting people to come by. Then whatever they purchase in this experiment, they actually get that item delivered. Like, these are the things that make the stakes real. So we run a lot of these experiments, and we also do partner with firms. Right now, we also have customers who are quite excited to at least give us a glimpse of the behaviors that their users exhibit so that we can get a little bit deeper understanding of how people behave in these different platforms.
How Customers Use Simile: Populations, Queries, and Experiments
Vibhu [00:16:39]: I think on the customer side, they have a lot of data about their users, who has bought. They have the action data.
Vibhu [00:16:47]: Can you walk us through an example of what someone comes to you for? What questions would they want solved? Do you customize a model for them? Do you have something off the shelf? What does that look like?
Joon [00:16:59]: Today, when people leverage our models, it’s often to better understand the population of their interest. So usually, the start of the relationship, we come together and hear about what population they want us to model, right? So it might be that if you’re a CPG company that’s selling to all of the US, then maybe it’s fairly straightforward. You want to model the gen pop of the US. But at the same time, if there is a vertical or if there’s a market that they’re trying to go into, imagine, they want to better understand, let’s say, people in their 20s and 30s living in California. That’s a much more specific population. So we hear about this population, and we go recruit these people, with consent, and with incentives, and we collect some of their data and create a model of these people. Then what our product allows you to do is query them. So it can take as input a filter that is a description of the population that you want to talk to, just like the one I just mentioned, and an environment. The environment can literally be survey questions, behavioral experiments, It can be A/B testing. Oftentimes, the core use cases are things like concept testing, to start with. But also, people sometimes want to do focus groups or one of the fun use cases that we also serve is even modeling things like earnings calls for public companies.
Joon [00:18:21]: So these are the use cases that we often start with.
Swyx [00:18:23]: Concept testing, is that an established term? I’ve never heard of concept testing.
Concept Testing, Gallup, and Politics
Joon [00:18:27]: Yeah. So it has to do with they have, let’s say, different messaging, different products, different ideas.
Swyx [00:18:32]: It’s like a marketing exercise.
Swyx [00:18:33]: Okay, got it. Got it. Politics?
Joon [00:18:36]: We do, have a strategic partnership with Gallup, and of course, Gallup is deep into policy space and so forth. Right now, we have not worked deeply with politics, like that area just yet, however.
Swyx [00:18:49]: I’m curious if there is demand or if they really would have different needs that somehow fundamentally don’t mix with your existing, users or people.
Joon [00:19:00]: I think there’s certainly demand.
Joon [00:19:02]: But we are very much mindful of how this technology gets adopted and the societal impact that we’ll end up having with this technology. And I do see politics as an area where a company has to be particularly thoughtful about the way they operate and make impact. So this is where we also want to make sure that we form enough of guardrail and perspective on how to leverage this technology before we go on to serve markets like the politics.
Swyx [00:19:29]: I’ll give people an example. one of my favorite shows is The West Wing. I don’t know if people have watched.
Swyx [00:19:34]: One of the key storylines is, like, the president has, multiple sclerosis, but they haven’t. they need to figure out how to disclose it. So they run a poll with a fake governor and ask people to respond on the poll,
Counterfactuals, Polling, and When Simulation Is Useful
Swyx [00:19:47]: They try to make decisions based on the results of that poll on, like, how well they’ll be received, like where, how should we play this?
Swyx [00:19:54]: And I’m like, well, I think those counterfactual things, I would use a simulation for this if I could trust it.
Joon [00:20:01]: For sure.
Joon [00:20:02]: In that show, how’d it go?
Swyx [00:20:04]: In that show, it was, like a foregone conclusion. They were like, “We know it’s bad. We just don’t know how bad.” And then the poll came back. It was like, “It’s really bad.” And then they just did it anyway.
Joon [00:20:14]: Part of it is to show, right? So you’re, you’re looking at the idea
Swyx [00:20:17]: Maximizing drama.
Joon [00:20:18]: How bad could it be? Oh, it’s horrible.
Swyx [00:20:20]: And to some extent, I think that is part of the trick of the, or the challenge or with being a customer of yours, which is that if I know it’s. if I roughly know and can intuit
Swyx [00:20:35]: What the effect is going to be, do I need you? What sensitivity of it, of effect do I need in order to make a decision, right? So for example, if I, my approval rating is 50%
Swyx [00:20:48]: And I, they have this negative piece, news item comes out, and it drops to 30.
Swyx [00:20:52]: If it drops to 20, if it drops to 40, do I care? No. It, I know it drops. It’s negative. So when do I care about simulations?
Joon [00:21:01]: You do something that’s clearly bad, that’s not popular, and people don’t like you, like, yeah, it’s like
Swyx [00:21:05]: You don’t need a simulation.
Joon [00:21:07]: Yeah. Well, so there are a couple of things. one is, there are use cases where, like every day, developers, designers, policymakers, marketers, every single day, they create assets. They create new products. And turns out, it’s many of the decisions in hindsight is obvious. Yes, of course this is bad, but we still run those studies because understanding the magnitude and understanding how acute something is quite difficult, even if, we feel like, of course, like this makes sense. this is the reason why we make so many mistakes. Like, every time somebody goes online and say something that has huge backlash, you look at that and like, “What an idiot.” However, it’s tough. That’s one. There’s also another aspect here, which is, again, this is the reason why simulation is different from prediction. In simulation, in the ideal case scenario. So what simulation is trying to show is it’s trying to show each step of the way or each step that we need to take to get to a certain outcome, right? So in the most advanced simulations, sometimes the next step that we’re suggesting might be quite counterintuitive. The analogy that I sometimes give, and I ground it in a more realistic example, but, I, as I mentioned, I’m a huge fan of science fiction, and I don’t know how, many of the audience members have read, like, things like the Foundation series by Asimov.
Simulation as a Path, Not Just a Prediction
Swyx [00:22:37]: Oh, yeah. We’ve mentioned psychohistory a number of times.
Joon [00:22:39]: Okay, fantastic. So I might be, talking to the right crew. If you read Foundation series, literally the first act is there’s a group of scientists who have found out that, “Oh, our galactic empire is going to collapse, and we’re going to have 30,000 years of unrest.” And they run psychohistory, the simulator that tries to teach them, “Okay, how can we keep this unrest to a 1,000 years?” And they plan this out, and the first step of that plan is to get the scientists who say, “Okay, this is coming,” exiled into this random place in this, galax- galaxy.
Swyx [00:23:18]: Terminus.
Joon [00:23:19]: Exactly. And that’s so counterintuitive. Like, what a strange move that you literally sent the group of scientists who was raising voice around this potential collapse of galactic empire into nowhere. How is that the right first move? Well, it turns out in this particular simulation, that was the move.
Joon [00:23:40]: It’s these things, right? And the reason why these reasoning is possible is because you’re showing the step function or each step that results in a particular outcome. So really what simulation allows you to do in its highest form is you give it not a problem or question, like what would people answer to the survey? That’s not what we do. What we tell it is, “Here is a goal that we have. In the context of foundation, we want to keep the unrest to a 1,000 years. What is the path that we need to take now to get to that particular future?” And that’s what simulation allows you to do. Now, translating that into real market, imagine you’re a automobile company and you’re about to release a, EV, and you’re trying to understand, well, how do we market EV, to make sure that our stock price goes up? But what if the answer comes down that, well, you can market your EV in XYZ way, but that might change people’s perception around the cars that’s not EV and make your overall sales to go down. Not very intuitive, especially all you’re trying to optimize is EV salesss, and that’s the only thing that you’re tracking, then that might result in a completely wrong solution, or at least different solution than what you would have expected, whether it’s right or wrong.
Joon [00:24:57]: That’s the power of simulation.
Swyx [00:24:58]: For listeners, we covered a similar topic with Mikhail Parakhin from Shopify, where they are working on SimGym. I don’t know if he ever talked to you about it. it’s very similar.
Joon [00:25:07]: I
Swyx [00:25:07]: The goal is increased conversion, but then the journey is very unusual.
Joon [00:25:12]: Journey is unusual.
Swyx [00:25:12]: Yeah. The-- He’s trying to look for interventions on a shopping trajectory, which is similar to what you’re saying. Like, it’s not about the attitudinal, is your word for it.
Swyx [00:25:24]: It’s about behavior.
Joon [00:25:25]: It’s about behavior.
Swyx [00:25:25]: And that’s exactly the difference, right? It’s, like, not about the near-term direction about-- but it’s more about, like, how do you affect multiple turns of interactions.
Vibhu [00:25:35]: You had a good quote at the start about this as well. It’s not about people wanting to know the outcome. It’s about how they can change it, change the way to get there, something like that. But I wanna take it back to how do we know this is grounded? Like
Grounding and Evaluating Digital Twins
Vibhu [00:25:47]: How do you run evals? How do you test that simulations come through? if I was to do the same thing that you described with, say, your favorite LLM, Opus, GPT-5.6, have some agent to map out these things
Vibhu [00:26:02]: How different are the answers we would get if I give it the same goal, the same objective, make a decent system? You’re saying that you need to change the model weight. You have your own solution to this. But how far off are we, and how do you check if it’s grounded? you have some interesting stuff on your site that points to how you run real evals, but if you could take us through that side. I think that’s one of the big concerns that people have. They’re like, “LLMs hallucinate.”
Vibhu [00:26:27]: “You’re just hallucinating layer after layer,” right?
Joon [00:26:30]: The way we do this, and this is the paper that we worked on after the generative agents paper that really became the, at least for Simile and also the field of simulation and synthetic panels, really became the foundation. Yeah, this is the paper. the paper is called Generative Agent Simulations of 1000 People. Here’s what we’ve done. For this paper, we brought 1,000 people that’s representatively sampled from the US to a virtual lab. And what we have done was we spent two hours collecting fairly wide-ranging data. In this particular study, we focused a lot on this interview data, that was, whose script was taken from this project called American Voices Project. And then we would also pair that with a lot of behavior data and so forth, whatever we can collect within two hours. And then we would send these people away for a couple of weeks. And during that time, I would use this data to create their digital twins. And I would bring the humans, participants back after 2 weeks and have them complete a battery of surveys, experiments, behavior studies. So we have the list here, which included things like behavioral economics games. We would run literally, like, Big Five personality test, General Social Survey. We would also go ahead and run the randomized controlled trials that were published on PNAS. And we would have their digital twins predict how the source individuals would have acted in these studies and surveys. And this is where we could replicate people’s behaviors and attitudes 85 percent as accurately as people would replicate their own. So that was the first really paper that gave this validated results that we can model individuals in an accurate way. And what we ended up finding now, of course, in AI space, so this paper came out at the end of 2024. AI space, a year and a half, 2 years, that’s a lifetime.
85% Accuracy and Why Frontier Models Miss Human Behavior
Swyx [00:28:24]: Yeah. Just, for listeners who are not seeing the YouTube, I just wanna say, like, the headline figure is 85 percent accuracy, like, which is a big improvement over all the other
Swyx [00:28:34]: Methods that you showed.
Joon [00:28:36]: But the part that was particularly striking to us, especially as we improved this technology even further, was the generative AI models like ChatGPT, Claude that’s coming out, it does give you the right foundation. However, what they do not consider is the true attitudinal and behavioral aspect of people, especially in the population that you care about. So what these models are really good at today is they’re trying to become the super rational, objective machines, right? So you go get their data from places like Mercor, Scale. You talk to professional programmers, scientists to create model that’s amazing at reasoning. That’s what they do. Simile doesn’t care about any of this. The models that we’re talking about here, what we’re trying to create are models that are as dumb as I am, right? So if I make some mistakes, the model has to make the same mistake.
Swyx [00:29:34]: Oh, that’s very hard.
Joon [00:29:35]: That’s very hard.
Swyx [00:29:36]: You’re solving Murphy’s paradox.
Joon [00:29:37]: That’s exactly. And this is a completely different data and training objective. This is also where we see quite a bit of discrepancy in the performance in human behavior prediction between the frontier models, Simile’s model, and the models being created in this space, where in some cases, the model performance of frontier models go all the way down to 20, 30 percent, especially if you go into that more niche population on topics that our customers would care about. On more gen pop, it might be around 50 to 60 percent. So it’s not very robust. Like, you wouldn’t want to make your decision off of these and these findings. If you can bring that up to 85 percent, that is ultimately what people end up getting very excited about.
Swyx [00:30:20]: Yeah. Do we wanna keep going on the paper, routes?
Joon [00:30:23]: Yeah, for sure. So the last one, was an interesting one. So this, paper was the follow-up paper that we had, to the 1000 agents paper, where the idea was now can we augment the models even further and post-train a model based on a lot of randomized controlled trials? So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there’s this, there’s this platform called Open Science Framework. So some, the audience might be familiar with this. And there has been, especially in the social sciences over the past 5 years or so, there has been this concern around replicability of studies. And so it was a bit of a crisis, the scientists acknowledged, where we rerun the study and we don’t see the same finding.
Post-Training on RCTs and Replication Studies
Vibhu [00:31:12]: Oof.
Joon [00:31:12]: It’s tough. And the reason why it’s there-- that was often the case was there’s this survival bias where the papers that get published often need to maintain what we call the value of less than 0.05 in the experiments that we ran. That suggests that only-- there’s only 5% chance that the results that we saw is false positive. But the tricky part was all the papers that were not published, and there’s still a 5% chance that whatever we publish is totally just randomly generated. Like, there’s a 5% chance that, hey, this effect is not real, but it just happened to be real because of the sampling bias. So because of that, what scientists started to do was they started to register their studies. So before running an experiment, they would go to this platform and say, “Here is the data. Here is the population that we’re collecting, and here’s the hypotheses.” And they would just say, “Here is our hypothesis.” Like, “This is what we believe.” And you cannot retroactively change those hypotheses. This is what gives us more scientific statistical confidence that whatever effect that you ended up seeing is true. So that ended up creating this really interesting platform where there’s one platform that has now contains tens of thousands of real-world experiments and hypotheses. And a lot of these are really high-quality, like, professionally designed behavior studies and random- randomized controlled trials. So we got the data and the studies from this platform and used that to make a point. And this particular, model is not, something that we’re serving commercially because this was a part of the open science. But this particular data set, helped us make a point that by collecting a lot of these randomized controlled trials, that are really well-designed, we can make significant improvement in model’s capability to predict human behaviors. So that’s what this paper was about.
Vibhu [00:33:10]: Is this stuff done on a individual level? Like, do I need to tune the model per individual, per company? Is there foundation model changes and then some slight post-training? Anything you can share there?
Population-Level vs. Individual-Level Models
Joon [00:33:21]: So this particular model was trained. the data we had at the level of individuals, but this particular model was trained. We experimented with both. And this is what we end up doing at Simile too. We always train 2, distinct model. One is what we call the population-level model. The other is what we call the individual-level model. And both take very similar input, which is the description of a subpopulation or individual and a stimuli. In this particular work, we’ve done the same. Here, the results that we are reporting are much more geared towards individuals because we do think that is a harder task in many ways, but that’s what we have done.
Vibhu [00:34:02]: You seen anything on the questions that humans can solve that models can’t solve? So like
Human Biases, Mundane Choices, and What Models Miss
Vibhu [00:34:09]: Currently, it’s, I live 5 minutes walk away from a car wash. It’s a 10-minute drive. Should I walk or drive?
Joon [00:34:16]: Huh.
Vibhu [00:34:16]: The model will say, “Oh, walk to the car wash.” And, you don’t have your car.
Vibhu [00:34:20]: Is anything like this a problem in simulation? You would assume, like, very simple for human to think about, but if the model is saying you should walk to the car wash, anything here?
Joon [00:34:32]: It’s less, what can we solve, but I think it’s more about what biases or mistakes do people make that models miss. Like, imagine that you are, like the. When I was still at Stanford, I lived in Palo Alto. So it’s about, I would say, 40-minute walk from the campus. You ask the model, “Okay, let’s go home. What can I, what can I do?” It would likely call an Uber or, give me, the bus time. But for the longest time, I really liked walking back. And the reason why I wanted to do that was not for efficiency. It really helped me think. And I like to walk for, half an hour or 40 minutes or so a day, where I just get to, just think about ideas, research, just get lost in my thoughts. That’s very human activity. Unless the model has seen that and understands the importance of that activity, it would miss these kinds of features. So that I think, is fundamentally what we’re trying to model. Like, what is fundamentally human might not be the most efficient thing to do, might not be the right thing to do, but things that make us who we are.
Swyx [00:35:43]: I’m curious if, there are some data sets that you really want that would materially help you. One version of this may be interesting, which is more valuable to you to acquire as a data set, all of LinkedIn, all of Twitter, all of Facebook?
What Data Matters: Social Media, Transactions, and Facebook
Joon [00:35:57]: It’s a little bit hard to rank, in part because, there’s, there’s this product saying where no feedback is wrong because it teaches you something about your users. Doesn’t matter what feedback.
Joon [00:36:11]: I think it’s a little bit like that.
Swyx [00:36:12]: So just whatever is bigger.
Vibhu [00:36:13]: What about a different domain? Say it was. What about all of Amazon data?
Joon [00:36:17]: Oh, yeah.
Vibhu [00:36:18]: Shopping data, right?
Joon [00:36:18]: Shopping data. So Amazon data is interesting in that it’s very much behavioral, although, like, what people do on social media, you could squint and say that is also behavioral. But the transaction data is always interesting. It is also most commonly available, however.
Joon [00:36:33]: If we were to look at purely social media, like if you really, if I were, if I had to really pick, Facebook likely is interesting because I do think it is most a default version of people. Because you go to LinkedIn, it’s very much professional environment. So people put up their, they have their guards up, right? And that still is interesting because that is true human attitude and behavior, but it is not your base state. you go to Twitter- Twitter, people have their own crazy personas, or depending on who you are. Like, my Twitter profile and, persona is very much, initially was I was very much an academic. “Hey, I’m here to share my studies.” Now, I share, things that’s related to Simile. But Facebook is one of those more private space where people just connect with their friends. In that way, I do think it shows you a little bit more about who that person is. So if I had to pick, I’d likely pick, Facebook.
Swyx [00:37:30]: Yeah. And you’re interested in, like, the whole person and their background and philosophy. I, is it too clinical or too machine learning-oriented to just say this is just ways to inject variance and biases? The broad question, is, like, is this any better than a randomized, like, combinatorial explosion version? So we have a link to the Tencent
Billion Personas, Synthetic Demographics, and Bespoke Data
Swyx [00:37:54]: Billion persona paper, where they did not do any of the groundwork that you are doing.
Swyx [00:37:59]: They just did like a cross matrix of here’s all the professions in the world, here’s all the people, possible backgrounds in the world, do a dot product across all of them, and that’s it. That’s your prompt for a billion people.
Swyx [00:38:12]: This will do something. I don’t know if it’ll do what you do, but it gets you some way, some percent of the way there.
Joon [00:38:18]: So this was an interesting paper. Like, what I admired about this paper when it came out was the scale. And you do gradually want to be able to simulate really large societies and interactions. So the scale is definitely admirable. it is relying heavily on the known statistics that went into training the model. So to the extent that you believe that statistics is correct, this is not a bad way to go about this. But the thesis here, and this is something that we also have seen in the market, like if this works, then we have solved simulation.
Joon [00:38:54]: It,
Swyx [00:38:55]: Because I survey, like, okay, 5% of the US population is in construction.
Swyx [00:39:01]: The other 5% is in medicine, whatever, right? And then you just keep going down the list, and then you do the other side. 5% has, like, the big 5 personality
Swyx [00:39:08]: Of, like, neurotic or whatever. That’s it.
Joon [00:39:11]: That’s it. So if you believe that the underlying data set and the platform that we’re leveraging has all the right statistics, then this will have solved it. you’re at that point merely retrieving the knowledge that is already embedded in the model, in the model parameters. That’s not, unfortunately, what we see, where there is such detailed and also niche knowledge about people that if you just take one example, it might feel very mundane, but it’s quite rich when you put together, that you do need to do a lot of bespoke data collection to better understand people. And this is also, I think what makes this particular, job fun, which you want to deeply understand people, and the process of deeply understanding them requires a lot of attention to the details. And you do need to pay attention to and pay respect to the daily lives that people lead.
Scaling Simulation: From Thousands to Societies
Vibhu [00:40:04]: I wanna talk about scaling simulation.
Vibhu [00:40:07]: So what can’t we simulate, what can we simulate, and how does scaling affect this? So how big are the models? What if we go from, 8B, like, couple 100 billion
Vibhu [00:40:18]: Like billion000 parameters, billion000? Do we get scaling? Any interesting emergence? Like, at a certain scale, at a certain amount of training, you uncover anything unusual and any learnings from that?
Joon [00:40:31]: What we are seeing is at Simile, so we do post-train our own model. The thing that we’re seeing is the early glimpse of scaling law in simulations. The more data about humans and more compute you ingest, you start to get predictive and predictable gains of the model performance in simulating it, simulating people.
Vibhu [00:40:51]: Ooh. We need a scaling law curve.
Joon [00:40:52]: It’s scaling law. Whenever you find it’s a beautiful thing. And we’re starting to see the glimpse of it, which is quite exciting. But if you talk about the ambition of simulation as a whole, it’s not merely about building a model. It’s about building a model, then creating the agents that become the individuals in a much larger ecosystem. So they’re creating this multi-agent simulation. Down the line, you want these multi-agent simulation to also live in a very rich environment, right? What we are really trying to get to at that point is, hey, can we create. All right, let’s do a time machine game again, and 5 years, 10 years into the future, can we create a simulation of 8 billion people living on Earth? I think that’s quite interesting. And that really is the vision. And once you get to that state, the questions that you can help answer for the society also start to change from my perspective. The answers are fundamentally about emergence of the emergent behavior of society and large groups of people.
Joon [00:41:53]: So the questions that I get excited by, and maybe this is a stodgy- a bit. I have my, academic side of me.
Joon [00:42:01]: And for me, it’s questions like, can we help solve climate change? If you look at climate change as a problem space, this is what we, like social scientists would often call it the wicked problems, problem where you have many actors with competing incentives for trying to make a very complex decision and coordinating that coordination decision. Very difficult to really solve in real life, which is also the reason why we couldn’t solve it. Can simulation help us solve that? Another one is, can we understand the signals for collapsing democracy, or can we understand or can we uncover the origin story of the monetary system? These are societal questions that we never really had a good way of answering. If we can create simulations of our society, you have to believe that these are the problems that we can solve. So that’s really the ambition of this field. And, I also think, yes, I think there’s a Nobel Prize to be won there, which wouldn’t be surprising. And I think there’s some amazing societal impact that we can have to help people make better decisions.
Climate Change, Democracy, and Societal Simulation
Swyx [00:43:04]: Nobel Prize in economics?
Joon [00:43:06]: In economics.
Swyx [00:43:06]: Oh, I see. I see. Rooting for you to write that paper.
Joon [00:43:10]: One of these days. But, one of the scholars that I was deeply inspired by, When I was coming into the space of simulation, is this scholar, named Thomas Schelling.
Schelling, Agent-Based Models, and the Nobel Prize
Swyx [00:43:23]: Schelling point?
Joon [00:43:24]: So the canonical example of the work that he’s done was he was one of the creators of agent-based modeling. So this was, like, in the 1970s and 80s. It’s very early days, but this was truly one of the first exemplars of simulations. And one of the canonical model from that time, and of course many of these simulations are trying to tackle the societal problems that’s most relevant for their era, it was called the model of segregation. So racial segregation was a big topic, that, we cared about. And what they’ve done was they created this grid world where they had red dots and blue dots. And these dots were, back in the day, like, they were the agents, and they had a simple rule that governed their behavior. If certain percentage of your neighbors are of different color and if that goes above certain threshold, then you move to a new location at random.
Joon [00:44:21]: One of the striking finding of this paper or this agent-based model was for the longest time, people thought the segregation within society was caused by explicit and overt racism.
Joon [00:44:34]: But if you look at this model, people’s preference towards living with people of the same color, that preference can be very minute.
Joon [00:44:42]: But the very small difference causes the society to segregate completely over time. This was very counterintuitive for a lot of people. And this particular work ended up informing housing policies. Mixed income housing, got really inspired by this work. And Thomas Schelling ends up winning the Nobel Prize for having laid the groundwork for very early versions of simulations. The opportunity that I do see here in the more scientific terms, is agent-based models for the longest, had impact in the 1980s, 90s, to some extent, early 2000s, but it has now gotten forgotten by the community a little bit. Because as you can imagine, red dots and blue dots is not really a rich description of people.
Joon [00:45:31]: But with the emergence of things like generative AI and, in particular, generative agents, we do have an opportunity to create these agent-based models that are high fidelity enough to help us make really complex decisions. And that’s the opportunity that I see. If that truly works, then yes, that is the work that will result in a Nobel Prize.
Swyx [00:45:53]: Yeah. For what it’s worth, and I grew up in Singapore. 80% of Singapore is in public housing, and public housing has, enforced racial quotas for exactly that reason, which is very interesting. okay, so we talk about scaling, we talk about all these, the agent possible applications.
Cost, Reuse, and the Economics of Simulation
Swyx [00:46:13]: I’m scared about the cost. if you even-- let’s just keep it to the US, about 8 billion people.
Swyx [00:46:21]: But, how much does it cost to model so many hundreds of millions of people?
Joon [00:46:26]: Oftentimes today, we don’t start at that scale, this stage of the, of industry and simulation as technology. But we can get our users extremely rich and meaningful insights even by modeling thousands, tens of thousands of people. And today what we do is every week we are collecting data on the scale of tens of thousands people’s data, and we have panel partnerships that gets us to tens of millions of people globally. So that’s what we do today.
Swyx [00:46:55]: And just as a side note once you’ve collected one person for one study
Swyx [00:46:59]: Can you reuse that same person for all the subsequent studies?
Joon [00:47:03]: That’s exactly right.
Swyx [00:47:03]: Okay.
Joon [00:47:04]: The beauty of this model and these agents is the fact that they are domain-agnostic.
Joon [00:47:08]: That what you’re really trying to understand is what is the fundamental nature of these people? What’s their social physics? And there are a lot of, a lot of, people that does change over time. Like, even, like, even things like, how many times have you gone have you been to, like, CVS the past week? that will change. But there’s so many traits about people that are also known to never change. Like, your risk tolerance doesn’t really change over time. It’s very consistent. So it’s these things that we’re trying to learn. But the scale we are operating is right now hundreds or, tens of thousands to hundreds of thousands. And in many of the core use cases that we are deployed in, and this is more than enough population, to cover those. Really, at that point, what you care about is less the number of people, but more do you have the right subpopulation of interest covered? And this is also the reason why people want a larger sample. It’s not because they want, stronger statistical guarantees. It’s more that can they filter down to any population of their interest. However, you can also imagine in 10 years, if we truly believe that the compute is going to scale, that we’ll have much more availability for compute, and our ambition for simulation is also going to scale accordingly, there’s definitely a reason for us to create an entire data center worth of simulations.
Joon [00:48:35]: Or in my hunch here is I do think in the next some number of years, we will start creating simulations that will cost as much as training a foundation model. But perhaps it’s going to be so valuable to the society that it would be a no-brainer. Right now, even today, like, we are training bunch of new foundation model just so we can say we trained one and we spent tens of millions. But if we can create a simulation at the level of society that would solve climate change, I would run that today. I would raise the money right now just to run that.
Multi-Agent Simulation and Social Influence
Swyx [00:49:10]: Amazing. the follow-up question is, does it also compound if you let the simulations talk to each other?
Swyx [00:49:18]: Or do they already do that today? They don’t, right, as far as I understand?
Joon [00:49:22]: It depends on what simulation you’re trying to run.
Joon [00:49:24]: In the multi-agent simulation setup, the agents do talk to each other.
Swyx [00:49:28]: Right, which is exactly Smallville, right?
Joon [00:49:29]: That’s right.
Swyx [00:49:30]: But a lot of times, for example, in commerce, you’re just by yourself, so there’s no point talking. which is way cheaper.
Vibhu [00:49:37]: But they use all these levels, right? Like, you decide what you will buy based on what other people around you buy and talk about, right?
Swyx [00:49:43]: It depends.
Vibhu [00:49:44]: It depends.
Swyx [00:49:45]: Again, I’m, I’m coming at this from a cost point of view. I’m like, “Oh my God.” Like
Vibhu [00:49:48]: I think
Swyx [00:49:49]: If there is, like, some combinatorial thing of, like, thousands of people talking to thousands of people, then that one million X’s might cost.
Vibhu [00:49:56]: I have a very different view as the cost point aside. Like, running these studies in reality is a lot more expensive, right? Running any study like this is you gotta have people do it, you gotta sign people up. It’s very expensive and sometimes, like, not feasible to run the study.
Vibhu [00:50:14]: But the outcome or the decisions you make are very expensive on them, right? So spend X million on something that, the overall process costs 100 million might as well, right? There’s, there’s a lot of value to be had there. It’s a small cost, but I’m excited on the cost side.
Joon [00:50:33]: To some extent, and when you deploy technology, you often want to deploy in a way where you can replace existing budget or you can make things more efficient, and that is the best way to deploy. However, the way you capture the long-term value of the technology is making the argument that, no, it’s the upside, that by making this better decision using simulation, you have saved yourself or made yourself hundreds of millions or even billions of dollars, and that’s a case to be made.
Vibhu [00:51:06]: Random tangent question. So if you’re doing a lot of inference, a lot of model multi-agent stuff, are you at the point where it makes sense to, train a model that’ very sparse? You’re expecting to do multi-million dollar runs. Are you thinking about this in model architecture standpoint or inference efficiency, or, you’re still at the research phase of it works, we’re not super there yet?
Joon [00:51:34]: Efficiency, we do think quite a bit about. this is technology that is deployed now in some of the largest enterprise companies in the world, and we do process significant number of queries, that are trying to, simulate the populations in the world. So efficiency is a consistent thing. we don’t want to over-optimize too early, so I wouldn’t say, like, this is the higher bid Right now, but this is definitely something that we think pretty carefully about.
Swyx [00:52:05]: Yeah. Are there other case studies? So we, you talked about CVS, talked about Gallup, Deloitte, Wealthfront.
Efficiency, Enterprise Use, and Real-World Case Studies
Joon [00:52:12]: Wealthfront is an interesting one, because one of the things they were trying to do, they were one of the first customers that wanted to do product testing that goes beyond just asking people what they think about, let’s say, behavior experiments and so forth. So there, really what we had to do was reason about multimodal input, so images, but also you can also imagine, like, these agents traversing through Figma mockups or websites. So some of the things that our agents can also do is it can be given a domain, like, or, like, a website URL and go use it for a while. It’s these things. And Wealthfront was one of the first, customers, that was very excited about this possibility.
Vibhu [00:52:53]: What have people been asking? Like, is there any demand that we have not covered? Like, UI testing, right?
Vibhu [00:52:59]: I wanna try a new. I wanna ship a new feature, test the UI, simulate how people will do it. Any interesting things that you’re seeing demand for?
Product Testing, Websites, and Synthetic Panels
Joon [00:53:08]: Today, a lot of the demand does come from like, the places where people have historically used human panels, we can now replace with agents, and these synthetic populations. And this is not replacing human panel. in many ways, the simulation that Simile is building is grounded. So the way that I think about this is we are trying to represent humanity at scale. And in that way, the use cases are what we would expect, but it’s the scale of deployment that surprises me.
Joon [00:53:44]: Turns out there are so many decisions that people make every day in these organizations, groups, and we want to be able to say, “We listen to people. We have consulted our users.” But in reality, that is rarely the case because getting to people and asking them many questions, it’s difficult. It’s both costly, time-consuming, but most importantly, people are just not available. If I had to answer 1000 survey questions for this one particular, vendor, even if I wanted to do that, like, I would never do it. And that’s very much the case. What simulation can do is ensure that the voices of people are always represented in rooms where the decisions for them is made, right? So all the stakeholders of this particular product launch, ideally they’re consulted. That’s what this technology really is trying to enable.
Market Size, TAM, and Human Decision-Making
Swyx [00:54:39]: In my mind, that means it skews towards more consumer focus, right? Like, anything with a wide enough customer base where you do benefit from the diversity that you represent. What are some rough statistics, just for people who are not familiar with this market in general, what’s the market size that. I’m sure you have some, like, rough numbers. market size is, like, a vague question
Swyx [00:55:01]: But, like, how much do people spend?
Joon [00:55:03]: So market research is a $100 billion industry.
Joon [00:55:06]: But the thing about simulation is not a tool for market research. Simulation is a tool for human decision-making. So the question around what is a TAM here is quite tricky, right? Because it’s easy to say, “Well, market research TAM is roughly 100 million or 100 billion.” so is it a TAM? And not really, right? Because in many ways, you’re trying to inform all human decision-making. You’re trying to inform every decision that are made about humans for humans. What is a TAM for that? It’s really unclear. And I’ll be honest. Like, I have a scientific background, I have a research background, so I didn’t come into the field calculating, oh, what is the TAM for human decision-making? But I just had to assume, well, if we can inform every decision that is made about human for human, that has to be big.
Swyx [00:55:58]: Some- something valuable.
Joon [00:55:59]: Exactly.
Swyx [00:55:59]: To some extent, you are a unicorn founder now, and you have to care as a CEO. But, like, I do think, like, yeah, when you go into these boardrooms with people that you’re quoting millions of dollars of contracts for, like, you have to say, “Well, here’s what you spend on humans-”
Swyx [00:56:15]: “. And here’s what we save you, and it’s 85% similar.”
Joon [00:56:19]: And certainly, the value case, is something that we care deeply about. Like, what is the value that we provide to the users and the decision-makers? But this is also where, like, as a founder, I think valuation only tells one very superficial aspect of the story, and I try not to think too much about valuation, in general, because that’s not what also motivates a team or certainly doesn’t. I’m, I-- Again, the interesting thing about researchers is we are happy living in academia, getting paid next to. we get paid okay. we don’t get paid that much, as a researcher here in academia, but it’s the impact and it’s the, it’s the value that we can provide to the individuals and the society that really drives us. And in that way, ultimately what drives us is the impact. Does the simulation we provide have a real impact in people’s decision-making in ways that progresses our society forward? If the answer is yes, then yes. that has to be great business, and we see that in numbers, and we do care deeply about that upside story, but that’s the heart of it.
Where Simulation Goes Next
Vibhu [00:57:27]: Do you have any timeline predictions? So we talked about scaling laws of simulations.
Vibhu [00:57:33]: You brought up, okay, maybe one day we can simulate how to solve climate change.
Vibhu [00:57:38]: Where are we now?
Vibhu [00:57:40]: If that’s not the end state, what is an end state, and what does progress look like?
Joon [00:57:45]: So what I sometimes tell people is simulation as industry, it feels a lot like where GPT-3.5, GPT-4 was, for the AGI saga, which is we have now technology that is powerful enough to do real damage on the verticals that we are tackling. At the same time, there’s a lot of progress that is yet to come. And that’s, I think, where this is. So the way I see it, I do think there will continue to be breakthroughs both in data, in algorithms, and there will be much more aggressive scaling that will also happen over the next few years. But I think that’s roughly where we are.
Swyx [00:58:27]: I think that was about the rough set of topics. Anything else that we should have asked you or you wish people asked you more about Simile?
Simulation as Painting and Understanding Human Essence
Joon [00:58:38]: I think the, what’s, for me, what’s quite fascinating about simulation, it is very impactful technology, but it is also very interesting technology, both in terms of, like, what it means for human society, our philosophy. And the way I sometimes interpret simulation is. So going back to my background, I as I mentioned earlier, I started my career as a painter. it was a professional pursuit, and I did oil painting, for figures. So I got my training originally in the realism studios, and that’s what I spent a lot of my, years, doing. Simulation is a lot like painting, right? The best paintings teach you something deep about the subject that you’re trying to represent. And it is always not a perfect representation. It-- No painting is perfect. There’s always some small differences and discrepancy, but what it does is it tries to highlight the thing that matters the most about the subject.
Swyx [00:59:47]: The essential
Joon [00:59:49]: The essential essence.
Swyx [00:59:49]: Yes. He, you, he’s brought up some of your work.
Vibhu [00:59:53]: Just nice to put it up.
Joon [00:59:54]: Yeah. So these are some of the works. So this is from, my, personal website that I maintain when, I was still a researcher.
Swyx [01:00:00]: I think a lot of people will say, like a Picasso, like anything postmodern is, like, very much focused on the essence.
Swyx [01:00:09]: Right. yeah, but I don’t know if any one of these evokes something that you like to tell the story of.
Joon [01:00:15]: No, it’s one of those things where, each of these paintings, drawings, whatever it may be, it is trying to surface something about the subject that you feel deeply about onto the surface. when I was a painter, and artist, the topic that I cared really deeply about was, the more mundane aspect of human lives. This shows up in some of the, some of the work that I’ve done, where, like, I did this entire study of a rural town where I went around and took photos of people for not really doing anything special, but just living their everyday lives. I thought that was the most interesting thing. I’m somebody who has this perspective where, the world is oriented around this fractal shape, and you have two choices to understand the fractal shape. You either go outward and try to explore as much as you can to understand the broader shape of the fractal, or you go inward because, the outward resembles the inward, shapes. And understanding the mundane aspect of it was very much that. Simulation has a lot of this, right? You’re trying to understand even the most mundane aspect of people. When put together- teaches you something really deep about that individual and the society. So I think that’s what’s interesting about simulation, the way, the same way that AGI helped us better understand or really think critically about humanity and human intelligence, simulation is really an exercise of understanding more about human society and our collective lives. So that I find to be, yeah, particularly interesting.
Swyx [01:01:56]: Yeah. Now you’re reminding me that some of the best biographers, documentarians, and even photographers, they’re taking a photo of you.
Swyx [01:02:05]: But before I take a photo of you, I must spend-- I must, like, follow you for a week just to understand you?
Swyx [01:02:11]: Which some artists, some do. Part of your work, there’s a very famous book called Working. I don’t know if you’ve, been referred to it before.
Swyx [01:02:18]: It’s very famous, like, to the point of having a Wikipedia page
Swyx [01:02:23]: About this like, really depth understanding and interview of people as they, about their lives, which seems mundane, but is told in a very, compelling way. Yeah, 1970s as well.
Joon [01:02:34]: Okay. It was an amazing decade.
Vibhu [01:02:39]: Before closing question
UBI, Future Questions, and the Value of Simulation
Swyx [01:02:41]: Okay, here we go
Vibhu [01:02:41]: You said that you started Simile with your 10-year question, right? If we do that now, 10 years down, what can we simulate? What would you simulate if, like, if you’ve made significant progress, are there any questions outside of the ones that we brought up? Any- anything that you think is most impactful? Anything that you would go vision 10 years out?
Joon [01:03:03]: In many ways, as I mentioned, I am somebody who is very much impact-driven. So the what would inspire me is I would want to ask, 10 years later, what would be the most important societal question that we as a society have to ask? I would love to tackle that. Like, do we need UBI? That could be an interesting one.
Swyx [01:03:24]: Ooh, has anyone done that?
Joon [01:03:25]: Well, we were thinking about it.
Vibhu [01:03:27]: Can we get access? Can we just
Swyx [01:03:28]: So OpenAI, this is, like, just trivia now. Like, OpenAI, or I think Sam Altman funded a study on this
Swyx [01:03:35]: In Africa, and the answer was no.
Joon [01:03:37]: The answer was no. But, what, was it something about the implementation?
Swyx [01:03:41]: Yeah, I know. It was a skill issue.
Joon [01:03:43]: Or was it something about, But this is the thing. See, when Sam
Vibhu [01:03:46]: Funny news article
Joon [01:03:46]: Altman funded this particular,
Swyx [01:03:50]: He spent 14 million dollars? Oh my God.
Vibhu [01:03:52]: It’s a little more.
Joon [01:03:52]: Quite a bit. But this is the thing. This is the reason why you want to run a simulation. You spend 5 years, 40 million dollars on this one study and have one finding, but if you can run simulation many times instantly, then that’s the value.
Swyx [01:04:07]: I feel like that one could-- you could have done in a simulation. Like, if you can do the housing study, you can do the UBI one. Like, I, come on.
Vibhu [01:04:13]: I think sometimes people will spend the money because they wanna verify what you think, right? Like, sometimes you just wanna. Is it right? Like, you gotta test it.
Swyx [01:04:23]: Okay, closing question. What are the chances we are in a simulation right now?
Are We Already in a Simulation?
Joon [01:04:28]: So it’s a fun question, and I assert at some point I just answer, yeah, we’re definitely in a simulation. But what I do, feel, however, is, whether we are in a simulation or not, that, I don’t think that makes our experience any less real. And I think that’s fundamentally, like, what I believe in. Maybe we live in a simulation, maybe not, but for
Swyx [01:04:48]: It’s real to us. Yeah.
Joon [01:04:49]: Yeah. For me, I don’t really care.
Swyx [01:04:50]: Yeah. Unless you die and you wake up in, like, the level higher or below.
Joon [01:04:55]: That would be interesting.
Vibhu [01:04:55]: I feel like you wouldn’t care. Once you die, then you find out you’re in a higher level.
Joon [01:05:01]: I worry about it when I die.
Swyx [01:05:04]: I think the other thing that. Okay, so I like the mathematical answer to this, which is, like, the, sheer number of possibilities that you are in a simulation far outweigh the sheer number of possibilities that you’re not.
Swyx [01:05:16]: Except for the simplest answer, which is, it is computationally very expensive to have you be a simulation. okay, great. You’ve been very generous with your time. Congrats on all your success. I met you just after your Smallville paper and had no idea that you could build, like, such an enormous company. And then now you’re like, “Well, it’s a $100 billion market, but that’s just where we’re starting.” So this is, very exciting.
Vibhu [01:05:42]: I think $100 billion market was not the term. That was only part of it.
Swyx [01:05:45]: Yeah, exactly. It’s, if you’re thinking too small.
Joon [01:05:48]: Well, I do believe that, maybe my final note here might be, again, I love science fiction. You look at any advanced civilization in science fictions, there’s 2 twin pillar, technology. One’s AGI in some form, and the other is simulation. So I think the market’s pretty big here.
Simile as Research Lab and Product Company
Vibhu [01:06:08]: Tell us about the company. You guys just raised a lot. You’re half a research lab, half a company. you’re hiring. Where are you based?
Joon [01:06:15]: Yeah. So we’re based in Mission Rock, so not too far away from, where we are right now. So we’re in SF, but we are also bicoastal. So we have our, team. I would say our headquarter is in SF, and we have a lot of our technical talent in SF, and we do have a smaller office that just opened up in New York. We are, as a company, an interesting one in that today, there are AI neo labs and then there are AI product companies. Simile truly is both. So this is a company that was founded by 4 founders, myself, Michael Bernstein, Percy Liang, Lainie Yallen. Michael, Percy, and I are all researchers. So of course, Michael was one of the authors of the ImageNet, kickstarted the AI revolution back in 2013, has been instrumental in human-centered AI. Percy coined the term foundation model, and is a, one of the greats of the AI researchers today. And Lanie is my business counterpart, where she led some of the fastest-growing AI native companies from their seed to A and B. But we have this DNA at the company where the vision of the technology that we’re creating is continuously developing, that we are getting people who were my lab mates. We are about 60 people right now.
Joon [01:07:28]: 15%, almost 20% of the company population are just my lab mates from Microsoft Research lab.
Joon [01:07:36]: And we It’s quite fun because many of them then had gone on to OpenAI, Google Gemini, and these places. And so it’s been a few years since we really got together and had a chance to work together. But now they’re coming back and really building out this vision that I find to be quite exciting, and that excitement is shared. So there’s that motion at Simile where we are a group of researchers trying to do something that no one is working on that we find to be the most impactful potentially. But at the same time, this is, again, technology that can make impact today. So we have an amazing group of engineers, product people, and designers, who are sitting here with us trying to imagine what does it look like to help people understand what simulation can do and make real-world decisions with this. Having both and then deploying it to some of the largest customers in the world today, it feels quite unique.
Swyx [01:08:30]: Yeah, it’s very compelling. One part of it was this is the call to action. Like, who are you hiring? You’ve done part of it, which is you have-- you’ve got a very talented group. Who are you hiring? Like, what roles?
Hiring and Closing
Joon [01:08:41]: So honestly, at this point, we’re hiring across
Swyx [01:08:43]: Everything
Joon [01:08:43]: All, section. we are always excited to bring on, amazing research talent.
Joon [01:08:49]: So if you’re interested in working with, our lab mates, we are always welcoming of amazing, researchers. But also we, hire, amazing engineers, that some of whom I, like, I respect the most. Many of them come from places where we have personal connections with, so many of the members are from Figma, Notion, Rive, and so forth, but also more broadly from the companies that we as a team have really admired. So engineers both in the product side, infra side, we’re all looking for those hires.
Swyx [01:09:24]: Well, lots of people. I think you made a really good case. So thanks, and, we’ll see you in the simulation.
Joon [01:09:30]: Amazing.
Joon [01:09:31]: See you all there.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe - This January, four big AI × Pharma tools deals were announced at the huge JPM Pharma conference that takes over San Francisco every year. OpenAI-backed Chai Discovery (now worth $4B) was somehow at the heart despite being all of 2 years old.
The Science team is proud to bring you the first podcast with cofounder Matt McPartlon and product lead Neil Patil to tell the full story!
Editor’s note: not to be confused with Chai AI, which was another top pod of ours.
Pharma suddenly doing big AI tools deals
For the non-pharma people, JPM is JP Morgan’s annual conference for pharma deal-making that takes over San Francisco for a week in January with hundreds of side events, etc. It’s a big thing.
Tools deals for pharma are also a big (new) thing: companies that start as AI for Pharma usually end up building their own drug pipelines instead, and the reason is something like this: convincing pharma to use your tool requires proof that your tool works. Proof means good targets, maybe with good clinical validation. If you have that, then it’s easier to raise money (with a known, if long path to commercialization) or sell (e.g payment in biobucks) for a specific target than it is to sell to lots of companies on a promise that it will work across their portfolios.
The “we’ll just partner / build our own drug” optionality proved to be the only good path up until January. What changed? In short, the tools got good enough for drug design teams to trust.
Good-enough-to-trust unlocks the ability to scale discovery: get more, better candidates into the lab and animal trials faster. More screening for toxicity, better delivery, etc. This means that what you push to the clinic is more likely to succeed.
Tools also unlock new capabilities: mechanisms that are very hard or impossible to develop using lab-based discovery. Designing an antibody that precisely triggers a very specific molecular cascade takes many years of trial and error. Designing bi-specific antibodies (that bind to two different proteins) is similarly difficult. Good design tools can unlock this.
RJ: The fact that the quality of the model has jumped means you’re enabling things you just plain couldn’t do. So it’s a step change. It’s not an efficiency argument at all, or not so much.
Matt: Yeah, exactly. It’s kind of interesting, even for us — it took me a while to believe in the thesis, actually. I talked to Josh for months before Chai started... It’s like, can I beat a mouse, and then can I do what mice can’t do? And then how many levels of interaction can you just keep building on top of that?
Everyone playing in the structural / binding space has an angle here, and some will be better than others, but Chai is pointing to a different unlock: getting good molecules right out of the gate (meaning they don’t then need as much lab work) means that the iteration time is faster. This turns science into engineering: you can design your systems to reduce friction and hill climb towards one-shotting molecules all the way to the clinic.
This, per-se, is not a new thesis: a16z articulated a version of this in 2020. What has changed is that structural models became binding models (how well doesn’t this molecule bind to this molecule, aka “binding affinity). Binding models unlock design, which has been steadily improving. Chai’s observation is that for engineering problems the best product tends to win, and good technology is a necessary but not sufficient condition.
Photoshop for molecules
With that in mind Chai has invested heavily in partnerships that allow them to learn from their Pharma counterparts.
What is kind of cool about working so closely and supporting so many of these partners is we get to really learn about what is the stuff that would be helpful in research. So rather than doing research in a vacuum, based on what would hypothetically be cool, we're able to do informed research based on what our partners have just been organically asking us for help with.
— Neil Patil, (Chai product lead)
This means better UX, such as a molecule editor that is more like a CAD or graphics design program than a chatbot.
Their approach has paid off: since June, Chai has announced three more major deals: Lilly, Novartis, argenx, plus an expansion of their Eli Lily program. This episode is too full of quotable moments for a short blog, so tune in to learn about
* Why protein tokens have the highest downstream value of any token
* Climbing levels of abstraction as models improve
* How Pharma, VC, and research are all just portfolio optimization
* How better tech changes the whole portfolio
* How relentless focus on simplicity leads to scale
Plus much more!
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Weitere Firmengründung Podcasts
Trending Firmengründung Podcasts
Über Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0.
We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al.
Full show notes always on https://latent.space
Sponsorship and business inquiries: business@latent.space www.latent.space
Podcast-WebsiteHöre Latent Space: The AI Engineer Podcast, The Tim Ferriss Show und viele andere Podcasts aus aller Welt mit der radio.de-App

Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen
Hol dir die kostenlose radio.de App
- Sender und Podcasts favorisieren
- Streamen via Wifi oder Bluetooth
- Unterstützt Carplay & Android Auto
- viele weitere App Funktionen


Latent Space: The AI Engineer Podcast
Code scannen,
App laden,
loshören.
App laden,
loshören.













