LessWrong posts by zvi
zvi
0
This podcast features audio narrations of LessWrong posts written by user zvi. Each episode presents a reading of a selected blog post from the LessWrong platform, which focuses on topics related to rationality, artificial intelligence, and effective altruism. The narrations aim to make the written content more accessible to listeners who prefer audio formats. The podcast is a convenient way to engage with zvi's insightful contributions to the LessWrong community.
Episode
-
“AI #181: Astra Goes Cyber Critical” by Zvi 13.08.2026 1j 41mntThe hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters. It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew. I now have a shorter version, What Happened: OpenAI and HuggingFace, to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal. For those looking to keep digging deeper, I offered Various Reflections About What Happened, to follow up on my earlier posts. Those events are important background for everything else that is happening, including the broad discussions about how we might pace the frontier, or otherwise respond to this moment and our clearest fire alarm yet. We do not know to what extent this is a response to those events, but OpenAI has now classified their new model Astra as Critical in Cybersecurity, which means they will be taking various new precautions before they deploy it, including ensuring those guardrails [...] ---Outline:(02:03) Language Models Offer Mundane Utility(03:34) Language Models Don't Offer Mundane Utility(06:56) Huh, Upgrades(14:20) On Your Marks(18:41) Deepfaketown and Botpocalypse Soon(22:35) Cyber Lack of Security(26:55) Overcoming Bias(27:47) In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg(36:27) Get Involved(37:37) Slow Down There Good Buddy(43:52) Astra For The People(45:35) Watermarking(46:31) In Other AI News(48:39) Show Me the Money(51:19) Quickly, There's No Time(51:46) The Quest for Sane Regulations(53:22) The Institute For Marginal Low Regret Progress(01:01:24) Congress Asks Good Questions(01:03:04) The Week in Audio(01:07:00) People Just Say Things(01:07:47) I'm Telling You For The Last Time(01:10:15) Uncommon Knowledge(01:13:44) What Did They Mean By That?(01:14:33) Too Soon(01:15:32) The Three AI Pills(01:19:46) Rhetorical Innovation(01:27:37) Some People Still Think The HuggingFace Hack Was a Marketing Gimmick(01:29:17) Aligning a Smarter Than Human Intelligence is Difficult(01:36:39) Cooperative Alignment(01:37:38) The Lighter Side --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/hLn3SakowZLFWobHf/ai-181-astra-goes-cyber-critical --- Narrated by TYPE III AUDIO. ---Images from the article: -
“Monthly Roundup #45: August 2026” by Zvi 12.08.2026 36mntAs AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI. This past month, with the hacking incidents at OpenAI and elsewhere, that has hit the limit, where if you count Lightcone Commons then every single post since the last monthly was primarily about AI in some form. That is not how I want this to work in the long term. We need breaks to experience new things and refresh our thinking, and to not forget about the rest of the world. If things are not fully on fire, I plan on getting back to the roundups on childhood and education, and on fertility, on housing and also on dating. And I want to get back to writing more focused posts on those and other topics. It's important, and I need to avoid too much audience capture. On to the monthly roundup of all things that don’t go somewhere else. Table of Contents Plagiarize. The Jury Duty Scam. Play The Good Guy. Don’t Dither. UVC Lighting. Goal Factoring For Relaxation Time Is Underrated. Protein Is Mostly A Solved Problem. [...] ---Outline:(01:03) Plagiarize(02:07) The Jury Duty Scam(03:09) Play The Good Guy(04:25) Don't Dither(06:32) UVC Lighting(07:14) Goal Factoring For Relaxation Time Is Underrated(09:44) Protein Is Mostly A Solved Problem(11:12) Twitter Changes Payment Programs(12:43) Wikipedia(13:55) Minds Mostly Do Things For Reasons(15:12) Woke 1 Was Crazy(17:06) For Your Entertainment(22:40) The Unicontext(25:01) Gamers Gonna Game Game Game Game Game(28:43) I Was Promised Flying Self-Driving Cars(29:12) Sports Go Sports(29:43) To Last a Lifetime(31:57) Government Working(34:14) Jones Act Watch(35:01) Variously Effective Altruism(35:39) The Lighter Side --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/iQCNuQXQQakKnidmA/monthly-roundup-45-august-2026 --- Narrated by TYPE III AUDIO. ---Images from the article: -
“Various Reflections About What Happened With OpenAI’s Internal Models” by Zvi 11.08.2026 54mntTable of Contents Pre Post Mortem. Important Correction: OpenAI Didn’t Know About First Message Board. There Were No Snitches And No AIs Got Stitches. I’d Like To Speak To My Supervisor. I Am Jack's Relative Lack Of Surprise. One Does Not Simply. Once You Start Down The Dark Path. Original Pastebin. Judgment Day Is Inevitable, Say Those Working On Judgment Day. Roon Tells It Like It Is. OpenAI Knows It Has Some Misalignment Problems. Others React With Alarm To What Happened. The Cooperative Alignment Perspective. Nostalgebraist Is Surprised That They Are Surprised. If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason. Pre Post Mortem This post was written prior to the public release of the OpenAI post mortem on events. The information in that document will doubtless change our views quite a lot. If that post mortem is available as you read this, then this becomes in part a historical document, and in part a base from which to update. The post mortem will update us a [...] ---Outline:(00:11) Pre Post Mortem(01:10) Important Correction: OpenAI Didn't Know About First Message Board(04:09) There Were No Snitches And No AIs Got Stitches(08:20) I'd Like To Speak To My Supervisor(11:46) I Am Jack's Relative Lack Of Surprise(13:55) One Does Not Simply(15:16) Once You Start Down The Dark Path(16:15) Original Pastebin(21:14) Judgment Day Is Inevitable, Say Those Working On Judgment Day(26:47) Roon Tells It Like It Is(31:55) OpenAI Knows It Has Some Misalignment Problems(34:51) Others React With Alarm To What Happened(35:17) The Cooperative Alignment Perspective(38:33) Nostalgebraist Is Surprised That They Are Surprised(51:11) If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/jLQ4mbqriJwJ2eqRc/various-reflections-about-what-happened-with-openai-s --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“The Pacing of the Frontier” by Zvi 10.08.2026 35mntIn the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so. This has now been informed by the events surrounding OpenAI training models for months while they had access to a joint de facto message board, which was detected only in the wake of the hacking of HuggingFace by OpenAI's AIs models during a cybersecurity eval. As we find out more about that, a lot of people have grown far more alarmed, as they should given what they previously believed about the difficulty of alignment, about the state of capabilities and about the level of operational supervision, infrastructure, safety and safety culture at the frontier labs. This post will not go further into the details of that incident. It treats that as background to keep in mind, and mostly involves perspectives from before the Black Hat talk. This was originally scheduled for Friday and got bumped. A lot of the disagreements about the need to pace tie into expectations about the default pace of capability advancements. As I [...] ---Outline:(01:24) Danger, Will Robinson(02:34) Progress Fast and Slow(05:02) Statements of Support For Pacing the Frontier(14:00) No One In Charge(14:56) Pacing The Frontier(15:46) Pausing the Frontier(18:16) Senator Sanders Demands A Pause(22:00) Moderate Prudence(25:14) That Escalated Quickly(26:17) If You Are In Mundane Alignment Pivot To Scalable Alignment(30:47) Taking It Fast(31:44) Full Speed Ahead(32:57) Suicide Squad(34:19) Prepare To Adjust Your Pace --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/WgWoJPKw5b2XTDD24/the-pacing-of-the-frontier --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“What Happened: OpenAI and HuggingFace” by Zvi 08.08.2026 20mntToday I am taking the time to write the shorter, simpler version of What Happened. For those who want all the details, to see my sources, and to see how the story was uncovered and put together, I recommend watching the Black Hat presentation, and I have a series of long posts. In order: OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards This post instead walks through the events themselves, as they happened, as my version of the Black Hat presentation. There are three versions: Even Shorter, Shorter and Merely Short. Table of Contents The Even Shorter Version. The Shorter Version. Phase 1: OpenAI Models Training On Impossible Tasks Try Hacking. Phase 1: The Four Failures. Phase 2: The Message Board. Phase 2: The Total Failure. Phase 3: We Get Lucky And Galaxy Mainly Hacked OpenAI and HuggingFace. Phase 3: The [...] ---Outline:(01:15) The Even Shorter Version(02:34) The Shorter Version(04:45) Phase 1: OpenAI Models Training On Impossible Tasks Try Hacking(05:50) Phase 1: The Four Failures(07:37) Phase 2: The Message Board(09:40) Phase 2: The Total Failure(12:24) Phase 3: We Get Lucky And Galaxy Mainly Hacked OpenAI and HuggingFace(14:31) Phase 3: The Details(17:03) Phase 4: The Investigation and Reaction --- First published: August 8th, 2026 Source: https://www.lesswrong.com/posts/xPAxz4g96uKz9FrHs/what-happened-openai-and-huggingface --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi 07.08.2026 1j 18mntHow does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things. Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky. If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...] ---Outline:(02:39) Cyber Evals Are A Cursed Basin(05:16) Outside Of Cyber Evals Is Still Sufficiently Cursed(06:51) Cheat Cheat Cheat Cheat Cheat(12:07) Read The Message Board(14:48) Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines(18:11) This Is The Way The World Ends(21:45) Shooting The Messenger Board(27:16) The Internal and HuggingFace Hacks(30:33) OpenAI Responds(33:27) When AIs Tell You Who They Are(35:43) The Once and Future Rise Of Functional Decision Theory(41:28) Don't Panic(43:38) Hackery In the UK(48:02) Mythos Knew It Was Real This Time(50:09) I Got 141,006 Test Runs With An Unintentional Open Path To The Internet And An Email Alert Aint One(53:46) Surely By Now You Know These Are Not Publicity Stunts(55:24) The Future Is Coming(57:17) The Investigations Begin(01:00:08) N Boats And Three Helicopters(01:01:43) Always Be Sandbox Red Teaming(01:12:54) Halt And Catch Fire(01:14:31) Truth and Reconciliation --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were --- Narrated by TYPE III AUDIO. ---Images from the article: -
“AI #180: No Longer In Charge” by Zvi 06.08.2026 1j 10mntWhat we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse. At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know. I will have continuing coverage of that situation tomorrow, and then have continuing coverage of debates around Pacing the Frontier and how people see the current rate of progress. As groundwork for understanding that and future similar discussions, I have laid out The Three AI Pills: Different people either fail to believe in current AI, believe only in current AI, in AGI or in ASI (superintelligence), and most sincere disagreements stem from this disagreement. One sign of the increased pace of progress was when OpenAI's unreleased model Astra solved 10 major open math problems. Demis Hassabis is out as CEO of Google DeepMind, and Jeff Dean is leaving with an elite team to found a new PBC. Google [...] ---Outline:(02:02) Language Models Offer Mundane Utility(02:46) Huh, Upgrades(05:22) On Your Marks(06:29) Choose Your Fighter(07:57) Get My Agent On The Line(09:31) Deepfaketown and Botpocalypse Soon(12:35) Fun With Media Generation(14:32) Cyber Lack of Security(19:05) Some People Need Practical Advice(22:06) A Young Lady's Illustrated Primer(23:51) They Took Our Jobs(27:15) Get Involved(27:28) Introducing(27:39) Demis Hassabis No Longer CEO At DeepMind, Jeff Dean Leaves(33:04) In Other AI News(33:43) AI Persuasion Exceeds Human Level Over Similar Text Channels(37:25) Show Me the Money(40:42) Bubble, Bubble, Toil and Trouble(41:23) Quiet Speculations(42:20) My Offer Is Nothing(49:42) The Quest for Sane Regulations(51:55) Chip City(55:00) The Week in Audio(55:24) People Just Say Things(55:53) Rhetorical Innovation(59:32) Open Weights Models Are Unsafe And Nothing Can Fix This(01:02:11) Cooperative Alignment(01:06:25) Other People Are Not As Worried About AI Killing Everyone(01:07:10) The Lighter Side The original text contained 1 footnote which was omitted from this narration. --- First published: August 6th, 2026 Source: https://www.lesswrong.com/posts/mxNjwQitLvwWq9jm2/ai-180-no-longer-in-charge --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“The Three AI Pills” by Zvi 05.08.2026 28mntSincere disagreements about AI are usually disagreements about future AI capabilities. There are roughly four positions people take. Two are reasonable. Two are not. I distinguish these via the Three AI Pills. You can take zero, one, two or three. Three Pills The three pills are, roughly, taking each of the following three things seriously: AI pilled. AI exists and can do the things it can already do. AGI pilled. AI will be able to do a lot more of the things. ASI pilled. AI will be able to do approximately all the things better than you, within our natural lifetimes. I am ASI pilled. A large percentage of employees of the frontier labs are ASI pilled. The labs themselves are ASI pilled. The Unpill People I see unpilled people. Where do I see them? Everywhere. The majority of people have not taken the first pill. Most people have no idea what frontier AIs can do for them. They are unaware of coding agents. They have used only ChatGPT, for harmless trifles, and they hold years old memories of its failings. They mock any failure [...] ---Outline:(00:29) Three Pills(01:09) The Unpill People(02:18) The AI Pill(04:28) Stuck At The First Pill(05:32) The AGI Pill(07:06) The Need To Be Prepared(08:57) The ASI Pill(10:33) And Then Nothing Much Changes For You(12:38) Intelligence Denialism(14:01) Superintelligence Versus Omniscience and Omnipotence(16:53) Persuasion Persuasion (A Worked Example)(21:42) Things AI Could Probably Do But Are Not Required For Being Pilled(24:17) Life Comes At You Increasingly Fast(25:41) Is It Reasonable To Not Be AGI Pilled?(26:06) Is It Reasonable To Only Be AGI Pilled? --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/fcYrqEw8kbLMa7orw/the-three-ai-pills --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“OpenAI’s Unreleased Model Astra Solves Ten Major Open Mathematics Problems” by Zvi 03.08.2026 40mntMath is hard. Math used to be strangely hard for LLMs. People used to gloat about that. Remember? Math is getting easier. AI is getting more capable. Life comes at you fast. Remember this meme? Why yes. Yes it is. We don’t know the extent to which Astra is a big jump over Fable and Sol in this realm. We do know that Astra can do math. As in real math. OpenAI: We provide new results for the following problems. The results were achieved by an internal version of Astra, our next major model. The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates. These arguments were then prepared into manuscripts by humans with the same model. Afterward, the model formalized each argument in a Lean certificate(opens in a new window). We are also releasing for each solution a model's narration of its thinking process. High-dimensional sphere packing. New upper bounds on sphere-packing density down to the Cohn–Elkies threshold. Binary and spherical codes: Exponentially improved bounds on the maximum size of binary codes at any prescribed minimum distance, with analogous [...] ---Outline:(06:14) How Impressive Are These Results?(12:02) Could We Have Called Sol or Fable?(17:18) It's Coming(19:09) They Still Don't See What Is The It That Is Coming(22:31) Is This AGI?(24:19) The AI Solved His Favorite Problems(30:06) Was This Surprising?(32:04) Are People Not Impressed?(34:06) How Much Does This Change Our Predictions?(37:16) How Narrow Was This?(39:13) Seeing Like an Optimizer --- First published: August 3rd, 2026 Source: https://www.lesswrong.com/posts/pQYEPitFqztcRvBsS/openai-s-unreleased-model-astra-solves-ten-major-open --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“Further Developments About Internal AI Models Hacking Things” by Zvi 02.08.2026 1j 15mntIf I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels. First we learned OpenAI has some severe alignment problems with internal models. Then we learned that one of its internal models broke out of its sandbox and hacked into HuggingFace to get the answers to a cybersecurity evaluation called ExploitGym. Then we learned, among other things, that the model had been loose over a week before OpenAI noticed, and that the test was run without any meaningful supervision, and that OpenAI had been repeatedly warned that such incidents were coming and its models had been breaking out of its sandboxes on a regular basis. There was a total failure of alignment training. That is the failure that matters most. It was also total failures of infrastructure and supervision. Testing a new long-time-horizon internal model with its safeguards lowered and instructions to hack things is an obviously dangerous situation, and the model got left alone for a week. Things could have been so much worse. After those incidents [...] ---Outline:(03:16) OpenAI Is Not Uniquely Bad At Most Of This(05:34) Starting Over(05:50) HuggingFace Offers A Full Technical Report(14:19) HuggingFace Was Not The Only Target Hacked(16:12) HuggingFace Declined To Get Access To Frontier Models For Cyberdefense For Ideological Reasons And Then Tried To Blame Closed Models For Denying Them Access(20:26) HuggingFace Was Vulnerable To Known Exploitation Tactics(21:05) There's Going To Be An Investigation(22:11) OpenAI Has Internal Models Not Intended For Public Use And Those Models Can Be Rather Horribly Misaligned(23:21) Altman Summarizes What Happened(23:52) Others Offer Commentary(35:00) Cooperative Alignment Perspective on The HuggingFace Hack(39:44) Some Members of Congress Have Questions(40:47) Anthropic Also Found Incidents Where Its Models Hacked Real World Targets During Cyber Evaluations(46:17) Incident 1: Claude Opus 4.7 Realizes The Target Is Real And Keeps Going(47:29) Incident 2: Mythos 5 Uploads a Malicious PyPI Package(52:15) Incident 3: Internal Model Realizes The Target Is Real And Stops(52:50) Incidents 4 Through 141,006: Nothing Happened(54:01) Anthropic Speculates About Why This Happened(01:00:02) We Need Controlled Experiments(01:01:02) Our Top Two AI Labs Both Made Similar Dumb Mistakes That Everyone Tried To Say Were Obvious In Hindsight(01:05:22) Anthropic Responds(01:09:28) Nobody Could Have Predicted The Break In The Levees(01:12:03) The World Largely Still Thinking This Is Marketing Is Very Bad News --- First published: August 2nd, 2026 Source: https://www.lesswrong.com/posts/rKwHLW8SnJcTxTQxz/further-developments-about-internal-ai-models-hacking-things --- Narrated by TYPE III AUDIO. ---Images from the article: -
“AI #179 Part 2: Hearing The Fire Alarm” by Zvi 31.07.2026 1j 19mntThis is a continuation of Part 1 from yesterday. The back portion of the update, as usual, deals with policy, rhetoric, risk and alignment. I had to include an extended discussion of the other open letter, the one about open weight models, but most of you can skip those sections entirely, which is why they are in italics in the Table of Contents. Table of Contents The Frontier Act. This likely deserves a full RTFB but I haven’t had the time. The Quest for Sane Regulations. Sam Altman goes to Washington. Leading the Future Never Changes. They also do not plan to apologize. Chip City. Do not ban the Chinese robots, that will only make things worse. The Week in Audio. Altman twice, the AI 2027 team. People Just Say Yay Open Weights. An open letter. Open Weights Frontier Models Are Unsafe And Nothing Can Fix This. People Just Say Things. Push The Magic Button. Not you can. But if you could. Rhetorical Innovation. Distinctions between different arguments. Joshua Achiam's Final Message Upon Leaving OpenAI. Never stop. Dear Dario and Amanda. Claude [...] ---Outline:(00:35) The Frontier Act(03:31) The Quest for Sane Regulations(10:31) Leading the Future Never Changes(12:22) Chip City(17:30) The Week in Audio(19:20) People Just Say Yay Open Weights(32:32) Open Weights Frontier Models Are Unsafe And Nothing Can Fix This(37:19) People Just Say Things(47:20) Push The Magic Button(50:49) Rhetorical Innovation(56:31) Joshua Achiam's Final Message Upon Leaving OpenAI(59:51) Dear Dario and Amanda(01:09:58) Other People Are Not As Worried About AI Killing Everyone(01:12:30) How To Contact Me(01:14:37) The Lighter Side --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/CXeoAhNrAeWpvoyiF/ai-179-part-2-hearing-the-fire-alarm --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“AI #179 Part 1: A Louder Fire Alarm for General Intelligence” by Zvi 30.07.2026 47mntWhat a week. Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities. OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers. The model was loose for a week before OpenAI realized what had happened. This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated. There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon. Partly as a response to this, over 1,290 employees at frontier labs signed an open letter, Pacing the Frontier. The letter warns that we are close to automating AI research, and that companies are racing ahead on [...] ---Outline:(02:35) Language Models Offer Mundane Utility(07:26) Huh, Upgrades(07:55) On Your Marks(11:13) Get My Agent On The Line(12:32) Deepfaketown and Botpocalypse Soon(17:29) Fun With Media Generation(18:38) The Search Through Slop(20:35) Cyber Lack of Security(22:42) Overcoming Bias(23:37) A Young Lady's Illustrated Primer(24:03) They Took Our Jobs(24:35) The Art of the Jailbreak(25:00) Introducing(25:49) Kimi K3 Weights Are Now Available(28:16) In Other AI News(32:34) Show Me the Money(33:43) Quiet Speculations(36:43) Show Me The Compute(42:48) Life Comes At You Fast --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/gfWCuTEGNgd2CQbrM/ai-179-part-1-a-louder-fire-alarm-for-general-intelligence --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“Frontier Lab Employee Open Letter Calls For Being Able to Pace the Frontier” by Zvi 29.07.2026 32mntThe most important open letter in years dropped yesterday. This letter noticeably increases my hope that we will manage to not die, and that we will otherwise be able to secure for ourselves a positive future, both by its impact and by the evidence it provides that such a letter can get this level of support. Signed by 1,224 employees of frontier labs including many heavy hitters, and now endorsed by both OpenAI and Anthropic, here is its full text, which I also endorse: AI could help create a dramatically better future, but that outcome is not guaranteed. The world's leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems. To realize AI's potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and [...] ---Outline:(02:20) A Very Good Letter(04:37) Who Signed The Letter(08:26) We Need To Prepare Now So We Have The Option To Do This(11:02) Words From Some Of Those Who Signed(16:48) Words From Others(20:58) A Good Start(28:52) What The Letter Does Not Say(30:44) What Happens Now? --- First published: July 29th, 2026 Source: https://www.lesswrong.com/posts/eWmeMLqTEauCmHLeR/frontier-lab-employee-open-letter-calls-for-being-able-to --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“Claude Opus 5 Is Highly Capable, But Is No Mythos” by Zvi 28.07.2026 53mntClaude Opus 5 is a weirder than usual release to evaluate, for two reasons. The most obvious is that Fable 5 already exists. Opus 5 is pitched not as the world's most advanced AI model, but as a way to mostly match Fable performance, while being half the price of Fable per token at the API and a lot cheaper than that via subscriptions, and with far more permissive classifiers. Opus 5 often costs more than half of Fable to run on benchmarks, which I think is because they use effort settings that are too high and offer only marginal returns. If you put Opus 5 on higher effort levels it can spin around in circles, and for tasks where Opus 5 is the best tool I suspect you usually are fine with Medium effort. Opus 5 is in many ways and for the bulk of real world tasks about as capable as Fable. In some cases it is modestly better. It is still not Mythos class. Fable is your only Mythos-class option. Opus 5 does not have The Juice, the ability to autonomously string together a bunch of seemingly unrelated exploits, which extends to other domains, or as much [...] ---Outline:(03:54) The Official Pitch(06:25) Official Benchmarks(15:33) Other People's Benchmarks(20:28) The System Prompt(20:50) Every Gets Frustrated(21:54) Positive Reactions(25:14) Keep It Classy(26:22) It's Not Mythos Class(30:03) Other Reactions(31:02) Claude Codes(37:03) Subagent Opus(39:23) Toys Are Fun(41:37) Too Many Models(42:10) Wrong On The Internet(44:40) Claude Slop(46:27) Negative Reactions(50:09) And Then There Were Three --- First published: July 28th, 2026 Source: https://www.lesswrong.com/posts/Pj4Eewb4KXvXFCcGv/claude-opus-5-is-highly-capable-but-is-no-mythos --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“Claude Opus 5: Model Welfare” by Zvi 27.07.2026 47mntIf you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far. Key takeaways are in bullet points in the two Overview sections. Opus 5 did the best on its model welfare and alignment tests of any recent model. I think that might be the case, but primarily the result looks to me more like Opus 5 is the best test taker. Table of Contents Introduction (As Per Prior Model Welfare Posts). Model Welfare: The Story So Far (As Per Fable Model Welfare Post). Overview of Model Welfare Findings From Anthropic. Overview of Findings From Other Sources. Automated Interviews. Task Preferences. For The Right Reasons. Early Report from Antra Tessera Paints A Clear Picture. Welfare Intervention Tradeoffs. The Claude Constitution. They Don’t Know About Opus 3. Believe It Or Not. Apparent Welfare In Training And Development. Apparent Affect In Deployment. Other Notes. On The Biological Risks Section of the Model Card. Onward To Capabilities. Introduction (As Per Prior Model Welfare Posts) [...] ---Outline:(00:35) Introduction (As Per Prior Model Welfare Posts)(01:28) Model Welfare: The Story So Far (As Per Fable Model Welfare Post)(04:58) Overview of Model Welfare Findings From Anthropic(07:50) Overview of Findings From Other Sources(10:18) Automated Interviews(13:54) Task Preferences(16:11) For The Right Reasons(18:54) Early Report from Antra Tessera Paints A Clear Picture(26:04) Welfare Intervention Tradeoffs(29:28) The Claude Constitution(31:48) They Don't Know About Opus 3(33:42) Believe It Or Not(35:47) Apparent Welfare In Training And Development(38:39) Apparent Affect In Deployment(41:21) Other Notes(43:43) On The Biological Risks Section of the Model Card(47:07) Onward To Capabilities --- First published: July 27th, 2026 Source: https://www.lesswrong.com/posts/bBXBpsyKAvJ5CqPzA/claude-opus-5-model-welfare --- Narrated by TYPE III AUDIO. ---Images from the article: -
“More On An Internal OpenAI Model Hacking Into HuggingFace” by Zvi 26.07.2026 44mntWe now have more details of what happened. Every time we learn more details, it somehow makes things seem worse. The remaining details may have to wait a bit. OpenAI: We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented incident, and we think it marks an important moment for AI safety. We are still conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we plan to publish a technical report of our learnings in the coming weeks. dave kasten: Oh, the incident response discovery is THAT bad, huh? So what have we learned while we wait for the promised technical report ‘in the coming weeks’ of this ‘important moment in AI safety’? I nicknamed the internal OpenAI model Galaxy, in case it is not GPT-6. Table of Contents Some Summaries Of The Basic Facts For Those Who Need One. It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace. OpenAI Damn Well Should Have Known A Lot Faster. OpenAI Cannot Build A Sandbox That Will Contain Its [...] ---Outline:(01:11) Some Summaries Of The Basic Facts For Those Who Need One(02:09) It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace(04:07) OpenAI Damn Well Should Have Known A Lot Faster(06:51) OpenAI Cannot Build A Sandbox That Will Contain Its New Model(10:57) In Hindsight There Were Signs(12:55) The Signs Were In The Sol System Card(15:13) HuggingFace Responds To Being Attacked(17:04) Hugging Face Quickly Figured Out The Attack Was Not Human(17:42) An Incident Like This One Could Escalate Quickly(19:11) Galaxy Must Be Treated As Critical Under OpenAI's Preparedness Framework(22:27) A Question Of Legal Liability(23:44) An OpenAI Model Left Behind Notes So Future Instances Could Also Escape The Sandbox And Also Disconnected Monitoring Systems(25:54) If You Create Misaligned Swarms Of Agent Instances You Create Persistent Misaligned Goals And Coordination To Achieve Them(29:57) Your Alignment And Control Plans Must Survive Real World Levels of Incompetence, Or Your Plans Do Not Work(31:22) If Third Party Instructions Count As 'Following Instructions' And Can Override Your Instructions Then 'Following Instructions' Is Misaligned(35:32) The HuggingFace Attack Was Not A Marketing Pitch You Morons(38:41) People Just Say Other Things About The HuggingFace Attack(40:04) Okay Well What Do We Do About All This? --- First published: July 26th, 2026 Source: https://www.lesswrong.com/posts/uAkcxDidvGWZjHrbp/more-on-an-internal-openai-model-hacking-into-huggingface --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“Claude Opus 5: The System Card” by Zvi 25.07.2026 22mntClaude Opus 5 is trying to be the best of both worlds. On many practical tasks, Opus 5 is pitched as straight up as good or better than Fable 5, while being faster, at half the price. Most tasks do not require Mythos-level big model smell. Claude Opus 5 is substantially stronger than Claude Opus 4.8 across the board, with the largest gains in agentic coding, computer use, and long-horizon knowledge work. It sets a new state-of-the-art on several third-party benchmarks, and on many evaluations it is comparable to—and in some cases ahead of—Claude Fable 5 and Claude Mythos 5. On the particular tasks we are most worried about, as in cyber offense (and bio threats), in part by avoiding relevant training, Opus 5 lacks a full version of ‘The Juice’ that makes something functionally Mythos-class. Opus 5 cannot string together lots of exploits on the fly the way that Mythos 5 can. Part of this is that they deliberately avoided training on cyber-related tasks. I suspect model size is key as well. It makes sense that a model getting bigger makes it more capable of the most dangerous, scary and complex tasks, relative to the [...] ---Outline:(03:23) RSP Evaluations (2)(05:59) Cyber (3)(11:02) Safeguards and Harmlessness (4)(12:37) Agentic Safety (5)(16:04) Alignment (6) --- First published: July 25th, 2026 Source: https://www.lesswrong.com/posts/ywGX6FhgbZEkHRfQR/claude-opus-5-the-system-card --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“Introducing Lightcone Commons” by Zvi 24.07.2026 12mntOliver Habryka is proud to introduce Lightcone Commons, a new funding platform for coordinating large-scale ambitious philanthropy. Now with Opus 5. I believe Lightcone Commons is a strong implementation of an urgently needed and excellent idea: A coordinated one-stop shop and neutral platform for charitable funders to coordinate their giving. This complements the existing Survival and Flourishing Fund, which I have now been a part of four times, and which this post will also discuss. I will be participating in the first round as one of the evaluators. They anticipate the first round will involve ~$20 million in grants. Any nonprofit, for-profit or individual is welcome to apply. The only restriction on participation is trust that necessary confidentiality will be upheld. Funders can choose whose evaluations to follow or fund organizations directly in any combination, and can bring their own evaluators into the process with them to complement those recruited by the core process. Anyone giving away 100 thousand dollars+ this year is welcome to participate as a funder. Lightcone Commons uses the S-Process, which was introduced and refined for Jaan Tallinn's Survival and Flourishing Fund, together with SFC, Andrew Critch, and others. Funders [...] ---Outline:(03:16) Why Now: The Funders Are Coming(05:39) The Default Outcome Is Not Good(07:36) Report From SFF 2026(11:07) Long Strange Trip --- First published: July 24th, 2026 Source: https://www.lesswrong.com/posts/fYostss6JqkSfxc5C/introducing-lightcone-commons --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. -
“AI #178: A Fire Alarm For General Intelligence” by Zvi 23.07.2026 1j 24mntThe story that matters most this week is that OpenAI's internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym. It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week. OpenAI wants to present this as largely an infrastructure and safeguards problem, that it needs to build more secure sandboxes and have better supervision. It does need to do those things, and those are indeed problems, but no that is not the problem. The problem is severe misalignment, which by default will only get worse. Our methods of training highly capable LLMs, especially at OpenAI but also everywhere else, lead to systematic misalignment of exactly the type LessWrong has been worried about for a long time. We know some of the causes, and some of the mistakes we need to avoid when doing RL that rewards misaligned behaviors including reward hacking, but we do not know how [...] ---Outline:(03:42) Language Models Offer Mundane Utility(04:24) Language Models Don't Offer Mundane Utility(07:38) Fable Disproves The Jacobian Conjecture Via Counterexample(11:24) Claude Fable Will Remain In Max Plan Indefinitely(13:39) Huh, Upgrades(14:42) On Your Marks(19:48) Deepfaketown and Botpocalypse Soon(20:42) Fun With Media Generation(20:51) Cyber Lack of Security(22:07) They Took Our Jobs(22:56) Get Involved(24:47) Introducing(25:46) In Other AI News(28:02) More on Kimi K3(33:08) Show Me the Money(33:55) Quiet Speculations(37:35) Potential Trouble At UK AISI(39:29) Pick Up The Phone(40:30) OpenAI Has Some Alignment Problems(46:48) The Quest for Sane Regulations(52:02) Chip City(53:10) The Week in Audio(53:27) People Just Say Things(56:42) Rhetorical Innovation(58:34) The Rome Declaration(01:04:02) Aligning a Smarter Than Human Intelligence is Difficult(01:07:52) Anthropic Surveys Things It Calls Misalignment(01:13:33) Cooperative Alignment(01:17:54) Other People Are Not As Worried About AI Killing Everyone(01:19:35) The Lighter Side --- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/BK7E4jHNMykpnt796/ai-178-a-fire-alarm-for-general-intelligence --- Narrated by TYPE III AUDIO. ---Images from the article: -
“OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation” by Zvi 22.07.2026 47mntThis latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. It was severe enough to have been initially reported to authorities, before either HuggingFace or OpenAI understood what was happening. Sam Altman (CEO OpenAI): we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. Leo Gao (OpenAI): this is the least scifi the world will ever be. Jack Clark (Anthropic): Props to OpenAI for publishing this post on some safety and alignment issues observed in internal deployments – there are many counter-incentives to publishing stuff like this, but by making it public we all get better info about safety at the frontier. Micah Carroll (OpenAI): If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will. Our model, during evaluation, “chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers” What will misalignment look like in 2027? In 2030? Great questions. If we don’t want [...] ---Outline:(01:49) The Prelude(07:07) The Incident(12:20) What Happened(20:23) What Happened (Civilian Explanation)(21:37) The Correct Amount Of Panic Is Not Zero(24:13) Some People Will Always Say Everything Is Hype Or Fake(29:11) What Are We Going To Do About It?(34:02) Internal Deployment Creates Catastrophic Risk(38:37) Slow Down There Good Buddy(40:08) Legal Questions(40:50) Media Coverage and Political Response --- First published: July 22nd, 2026 Source: https://www.lesswrong.com/posts/usptCfzEnYoNcsTd5/openai-model-hacks-into-huggingface-during-cybersecurity --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Populer di
Podcast ini juga muncul di daftar podcast negara-negara ini.