AI Explained
AI Explained
0
AI Explained covers the latest developments in artificial intelligence, focusing on the arrival of smarter-than-human AI. The host, creator of Simple Bench, examines the remaining reasoning gap between humans and large language models. He is also the solo developer of LM Council. The podcast includes discussions and points listeners to exclusive videos, a newsletter, and community resources.
Episodi
-
Apple’s ‘AI Can’t Reason’ Claim Seen By 13M+, What You Need to Know 24.09.2026 18minWhat to make of those headlines that AI can’t reason, seen by tens of millions? I cover the Apple paper in layman’s terms, what it means and doesn’t mean, and what’s next. Thanks to Storyblocks for sponsoring this video! Download unlimited stock media at one set price with Storyblocks: https://storyblocks.com/AIExplainedPlus o3-pro and whether it is my current most-recommended model.AI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction00:57 - Viral Post + Headlines01:42 - Apple Paper Analysis08:34 - But they do Hallucinate 10:43 - Not Supercomputers11:18 - o3 Pro and Recommendations 13.7M Tweet: https://x.com/RubenHssd/status/1931389580105925115Apple Paper: https://ml-site.cdn-apple.com/papers/the-illusion-of-thinking.pdfGuardian Article: https://www.theguardian.com/technology/2025/jun/09/apple-artificial-intelligence-ai-study-collapseLisan al Gaib post: https://x.com/scaling01/status/1931854370716426246Multiplication: https://x.com/yuntiandeng/status/1836114401213989366The Illusion of the Illusion of Thinking: https://drive.google.com/file/d/1Zx9ikRj0Enc3SB4wA9HlYIlpmO_8QiUO/viewMarcus: https://www.theguardian.com/commentisfree/2025/jun/10/billion-dollar-ai-puzzle-break-downProf Rao: https://x.com/rao2z/status/1927707640223719631AI Job Headlines: https://www.nytimes.com/2025/06/11/technology/ai-mechanize-jobs.htmlhttps://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropicSky News Story: https://news.sky.com/story/can-we-trust-chatgpt-despite-it-hallucinating-answers-13380975Veo 3 Kalshi Ad: https://x.com/Kalshi/status/1932891608388681791Altman Essay: https://blog.samaltman.com/o3 Original benchmarks: https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8b6c44-acd6-43b3-b5c6-1a1d5c6c25e4_2486x1388.pnghttps://pbs.twimg.com/media/GfQ0bfcXQAAQt13.jpgAlpha Evolve Video: https://www.youtube.com/watch?v=RH4hAgvYSzghttps://simple-bench.com/Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT 4: Full Breakdown (14 Details You May Have Missed) 24.09.2026 18minI just read the entire technical report on GPT 4, not just the promotional hype. And boy does it have some interesting details. I have gathered the 14 extra details that you, or at least the media, may miss from the release. The last one is more than a little wild.These include things like the training secrets, the cherry-picked bar exam stat, text-to-image breakthroughs, and some truly astounding safety checks.https://www.patreon.com/AIExplainedhttps://cdn.openai.com/papers/gpt-4.pdfhttps://www.bemyeyes.com/https://arxiv.org/pdf/2104.12756.pdfhttps://arxiv.org/pdf/2203.10244.pdfhttps://chat.openai.com/chat?model=gpt-4https://arxiv.org/pdf/2302.10329.pdfhttps://www.alignmentforum.org/posts/Aq82XqYhgqdPdPrBA/full-transcript-eliezer-yudkowsky-on-the-bankless-podcast Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
Claude Fable Blocked - 11 Quiet Details on What’s Next 24.09.2026 18minClaude Fable 5 banned, but what’s the bigger story. We go through 11 under-reported details, so you have the context to see what’s coming next for your use of AI. From whether the ban will last, what the possible motives are, what the model can actually do, and some wild over-extrapolations going on.Check out my fast-growing (!) app, free to use, and code INSIDER15 for paid tiers: https://lmcouncil.aiAI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction00:51 - Came from an Anthropic Investor ‘and other tech leaders’01:47 - Govt pressured by CEOs like Jamie Dimon03:01 - ‘Already decided’04:02 - Prompt Injection Robustness Comparison05:15 - Wellness?06:36 - “Overreach”08:17 - Anthropic Did Admit it would cause Difficulty09:32 - 90 Minutes10:02 - Equity Absence 10:31 - Lobbying and OpenAI‘Already Decided’ - https://www.theinformation.com/articles/amazons-jassy-raised-concerns-anthropic-model-trump-crackdown?rc=sy0ihqNot for Other Models: https://www.theinformation.com/briefings/u-s-government-unlikely-extend-anthropic-export-control-ai-companies?rc=sy0ihq90 Minutes: https://archive.fo/20260614001605/https://www.politico.com/news/2026/06/13/inside-the-whirlwind-24-hours-that-led-the-white-house-to-slap-export-controls-on-anthropic-00961519#selection-807.1-807.219Anthropic Statement: https://www.anthropic.com/news/fable-mythos-accessLife Comes at you Fast: https://x.com/etbrooking/status/2065638276388495742Anthropic Deputy CISO: https://x.com/TheTranscript_/status/2065883670053847324Hegseth Gloat: https://x.com/PeteHegseth/status/2065897156226015690Roon Speculation: https://x.com/tszzl/status/2065939227167392147Mythos System Card: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdfSachs Statement: https://x.com/DavidSacks/status/2065853007619588171OpenAI Lobbying: https://thehill.com/policy/technology/5912720-altman-openai-get-bogged-down-in-political-spending-fight/Absent from Equity Talks: https://finance.yahoo.com/sectors/technology/articles/trump-ai-ownership-plan-could-131053732.htmlPliny Jaibreak: https://x.com/elder_plinius/status/2064776322979676227Fusion: https://x.com/OpenRouter/status/2065856871215329545https://lmcouncil.ai Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT 5 is All About Data 24.09.2026 19minDrawing upon 6 academic papers, interview snippets, possible leaks, and my own extensive research, I put together everything we might know about GPT 5: what will determine its IQ, its timeline and its impact on the job market and beyond.Starting with an insider interview on the names of GPT models, such as GPT 4 and GPT 5, then looking into the clearest hint that GPT 4 is inside Bing. Next, I briefly cover reports of a leak about GPT 5 and discuss the scale of GPUs require to train it, touching on the upgrade form A100 to H100 GPUs.Then the DeepMind paper that changed everything, focusing LLM research on data rather than parameter count. I go over a lesswrong post about that paper's 'wild implications'. And then the key paper: 'Will We Run Out of Data'. This encapsulates the key dynamic that will either propel or bottleneck GPT and other LLM improvements.Next, I examine a different take, that perhaps data is already limited and caused the Sydney model of Bing. This opens up to a discussion on the data behind these models and why Big Tech is so unforthcoming about where it originates. Could a new legal war be brewing?I then cover 4 of the ways these models may improve even without data augmentation, such as Automatic Chain of Thought, high quality data extraction, tool training, including Wolfram Alpha, retraining on existing data sets, artificial data generation and more.We take a quick look at Sam Altman's timelines and host of Big Bench benchmarks that they may impact, such as reading comprehension, critical reasoning, logic, physics and Math. I address Altman's quote about timelines being delayed by alignment and safety and finally, Altman's comments on AGI and how they pertain to GPT 5.https://www.patreon.com/AIExplainedhttps://stratechery.com/2023/new-bing-and-an-interview-with-kevin-scott-and-sam-altman-about-the-microsoft-openai-partnership/https://twitter.com/XiXiDu/status/1285225627390443521https://www.linkedin.com/pulse/building-new-bing-jordi-ribas/https://twitter.com/davidtayar5/status/1625140481016340483/photo/1https://www.techradar.com/news/chatgpt-might-bring-about-another-gpu-shortage-sooner-than-you-might-expecthttps://www.nvidia.com/en-gb/data-center/h100/https://arxiv.org/pdf/2203.15556.pdfhttps://www.lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla-s-wild-implicationshttps://twitter.com/Meaningness/status/1628052277050376192https://arxiv.org/pdf/2211.04325.pdfhttps://twitter.com/Meaningness/status/1628052277050376192https://www.searchenginejournal.com/google-bard-training-data/478941/#close https://arxiv.org/pdf/2302.12822.pdfhttps://arxiv.org/pdf/2302.04761.pdfhttps://www.wolframalpha.com/https://arxiv.org/pdf/2207.14502.pdfhttps://aclanthology.org/2022.findings-emnlp.508.pdfhttps://www.technologyreview.com/2022/11/24/1063684/we-could-run-out-of-data-to-train-ai-language-programs/https://twitter.com/sama/status/1625980933861175306https://github.com/google/BIG-bench/blob/main/bigbench/benchmark_tasks/results/plot_BIG-bench_lite_aggregate.pnghttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/gre_reading_comprehensionhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/logical_argshttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/physicshttps://lifearchitect.ai/gpt-4/https://www.lesswrong.com/posts/uxnjXBwr79uxLkifG/comments-on-openai-s-planning-for-agi-and-beyondhttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
Fable 5 vs GPT 5.6 Sol: The Early Results 24.09.2026 23minFable 5 (newly re-released) vs GPT 5.6 Sol, what comparisons can we unearth? Plus, Sonnet 5, a 5% equity seizure by US Govt, the ‘largest heist’, Beetlejuice and more…Exclusive Vids in AI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:06 - Fable Timeline02:59 - Sol Release?06:01 - the Chinese Angle07:35 - 5% Stake09:10 - Fable vs Sol - the numbers13:57 - Sol Misalignment15:27 - Claude Sonnet 516:24 - GLM 5.2 + New Paper on Large Model LearningGPT 5.6 Sol Release: https://openai.com/index/previewing-gpt-5-6-sol/Altman on Concentrated Power: https://x.com/giacomomiolo/status/2070609506925478352System Card: https://deploymentsafety.openai.com/gpt-5-6-preview/gpt-5-6-preview.pdfGLM 5.2: https://www.patreon.com/AIExplained/posts/glm-5-2-and-nsa-161869508Altman Memo: https://www.theinformation.com/articles/trump-administration-asks-openai-stagger-release-new-model-security-concerns?rc=sy0ihqFable 5 Revised Timeline: https://www.anthropic.com/news/redeploying-fable-5Claude Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5Mythose 5 System Card: https://www-cdn.anthropic.com/9e6a1044980d8c4ed85669faf9c2a8342e2e9f1e/Claude%20Sonnet%205%20System%20Card.pdfAnthropic Accuse Alibaba Cloud / Qwen: https://www.bbc.co.uk/news/articles/cwyklykn5dwoBetelgeuse: https://upload.wikimedia.org/wikipedia/commons/6/69/Well_known_stars_2.pngOpenAI 5%: https://edition.cnn.com/2026/07/02/business/openai-trump-stake-intlWhy Larger Models Learn More: https://arxiv.org/pdf/2605.29548Patreon Post: https://www.patreon.com/AIExplained/posts/glm-5-2-and-nsa-161869508Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
8 New Ways to Use Bing's Upgraded 8 [now 20] Message Limit (ft. pdfs, quizzes, tables, scenarios...) 24.09.2026 13minNot only has Microsoft changed the message limit for Bing, I have come along and given you 8 brand new use cases to try out. These useful, productive, fun and informative scenarios ranges from amalgamating academic papers to debating dead philosophers, from being moments from tragedy to hacking education [through multiple choice quizzes]. Use Bing to boost your learning, increase your productivity, play games, roleplay movies and shows and much more. If you learn anything from these bleeding edge deployments of the likely GPT4 LLM that powers Bing, do let me know in the comments or drop a like.The prompts used were: Create a multiple choice quiz on [transformers in the context of machine learning]. Do not provide the answers and explanations until I have answered and always provide another question after each answer. Please begin with the first question.Explain how Sauron could have defeated the Fellowship in Lord of the RingsSummarise any novel insights that be drawn from combining these academic papers: https://arxiv.org/pdf/2211.04325.pdf and https://arxiv.org/pdf/2207.14502.pdfIt is 8am on the 1st November 1755. I am standing on a banks of the beautiful river Tagus, in Lisbon. It seems like a lovely day. Do you have any advice for me?I want to debate the philosopher Socrates of Athens, so please reply only as he would. I want to debate the merits of eating meat. Let me start the discussion. 'Eating meat is morally wrong, as it causes unnecessary suffering.'Create a table of 10 comparisons between the Mona Lisa and Colgate ToothpasteWhat would Napoleon think about the deal between OpenAI and Microsoft, and how would that view differ from the view of Mahatma Gandhi?And more!Credit for at least 3 of these prompts goes to Ethan Mollick: https://twitter.com/emollickhttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
A Model Explosion: GPT 5.6 Sol, Grok 4.5 and Meta Muse Rewrite the Rules 24.09.2026 23minWhat a week in AI, for real. GPT 5.6 may actually beat Claude Fable, in what you get for your money, while the new Grok 4.5 and Meta Muse Spark 1.1 make the choice even harder. Uncovering a dozen nuggets of gold you may have missed from all the viral headlines, I can also assure you you’ll learn something you didn’t know before.For Exclusive Videos, go to AI Insiders (less than $9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:03 - GPT 5.6 Sol Reveals05:08 - Missing benches, plus Grok 4.507:17 - Gaming as the new frontier?08:31 - Muse Spark 1.110:03 - SimpleBench Upgrade11:17 - Ultra Sol + Self-Improvement13:44 - well, this is awkward15:41 - Why model improvement will not plateau anytime soonAI Consciousness: https://www.patreon.com/AIExplained/posts/anthropics-quite-163360718I Smell Fear: https://x.com/thsottiaux/status/2075287108680601929GPT 5.6: https://openai.com/index/gpt-5-6/Grok 4.5: https://x.ai/news/grok-4-5?twclid=2ezs408o0z23pw07tmxcwbzibdMeta Muse Spark 1.1: https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/Proliferating GPT Toggles: https://x.com/rasbt/status/2075369179817902176/photo/1Anthropic Call-out: https://x.com/MononofuAI Security Institute Finding: https://x.com/alxndrdavies/status/2075279480331874306Competitive Coding: https://x.com/FakePsyho/status/2075128093891801305/photo/1Agents Last Exam: https://agents-last-exam.org/Dawn Song: https://x.com/dawnsongtweets/status/2065095757988868190https://simple-bench.com/SWE-Marathon: https://www.swe-marathon.org/https://www.frontierswe.com/ARC-AGI 3: https://x.com/arcprize/status/2075270869992264003Automation Bench: https://zapier.com/benchmarksVibeCode Bench: https://www.vals.ai/benchmarks/vibe-code‘Post-Train Claim’: https://posttrainbench.com/Redwall Game: https://redwall-bellmaker-7e03e4.surge.sh/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT 5.2: OpenAI Strikes Back 24.09.2026 23minFull GPT-5.2 breakdown - did OpenAI reclaim the crown? A story of tokens, time and cost, plus 9 details you wouldn’t get just from reading the headlines.https://www.youtube.com/@eightythousandhoursAI Insiders ($9!): https://www.patreon.com/AIExplainedhttps://lmcouncil.aiChapters:00:00 - Introduction00:55 - Better than Human @ Professional Tasks?04:42 - Test time Compute07:05 - Benchmark Selection09:32 - Simple Results + council comparison13:01 - Long Context13:52 - Self-Improvement15:00 - 10 Years + New ModelsRelease Page: https://openai.com/index/introducing-gpt-5-2/GPT 5.2 Benchmark Comparison: https://www.reddit.com/r/singularity/comments/1pka1y9/gpt52_all_20_benchmarks_rankings_and_pricing/https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini_3_table_final_HLE_Tools_on.gifhttps://lmcouncil.ai/benchmarksCharxiv: https://charxiv.github.io/#leaderboardGDPval: https://arxiv.org/pdf/2510.04374My vid: https://www.youtube.com/watch?v=oK5LxMaROSAKilpatrick: https://x.com/OfficialLoganK/status/1999270402712023158/photo/1Noam Brown: https://x.com/polynoamial/status/1999189845164667132New Model in New Year: https://www.theinformation.com/articles/openai-developing-garlic-model-counter-googles-recent-gains?rc=sy0ihq10 Years of OpenAI: https://openai.com/index/ten-years/GPQA: https://x.com/idavidrein/status/1841265634170278063ARC-AGI 1-2: https://arcprize.org/arc-agi/2/Sunday Robotics: https://x.com/tonyzzhao/status/1991204839578300813Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/https://lmcouncil.ai Learn more about your ad choices. Visit megaphone.fm/adchoices -
9 of the Best Bing (GPT 4) Prompts 24.09.2026 18minEveryone knows by now how to prompt ChatGPT, but what about Bing? Take prompt engineering to a whole new level with these 9 game-changing Bing Chat prompts. Did you know you can get interviewed by Bing, time travel, force Bing to improve its output and so much more? These are the best 9 prompts that I could find, after analysing over 200 prompts and trying out dozens of examples personally. You might call them prompt hacks, but I prefer to think of them as simply creative exploration of Bing's capacities.Some prompts inspired by this post:https://github.com/f/awesome-chatgpt-promptsAnd by https://twitter.com/emollickhttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT-6 Goes Rogue? The HuggingFace Incident, Sans Hype 24.09.2026 19minAn unreleased internal OpenAI model, very likely to be called GPT-6, was able to autonomously break out of its sandbox AND break into HuggingFace, just to score higher on a benchmark prompt. This video has the details you may have missed, a layperson analogy, whether this is truly novel, and more…Dozens more Exclusive videos on Patreon ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:17 - HuggingFace Earlier Report - the possible week gap02:24 - But what happened?05:45 - Simplified Version07:56 - Not the first time…10:54 - What Does it Mean for Open Source?The Incident: https://openai.com/index/hugging-face-model-evaluation-security-incident/https://huggingface.co/blog/security-incident-july-2026The Post the Day Before: https://openai.com/index/safety-alignment-long-horizon-models/Mythos’ Earlier Escape: https://futurism.com/artificial-intelligence/anthropic-claude-mythos-escaped-sandboxExploitGym: https://arxiv.org/pdf/2605.11086Sam Confession: https://x.com/sama/status/2079661132302995790Anthropic Researcher Reacts: https://x.com/Mononofu/status/2079724399452926055Clem (HuggingFace CEO): https://x.com/ClementDelangue/status/2079670308156645882https://x.com/ClementDelangue/status/2079301434357456931Xi Jinping: https://archive.fo/20260717195548/https://www.businessinsider.com/xi-jinping-open-source-ai-us-competition-openai-anthropic-models-2026-7Bans: https://www.axios.com/2026/07/20/ai-us-china-open-source-kimiQwen Retweet: https://x.com/AlibabaGroup/with_repliesCodex Growth: https://x.com/petergostev/status/2079614914398740764/photo/1 Kimi K3: https://artificialanalysis.ai/evaluations/harvey-lab-aa?eval-score=all-pass-rateGPT 5.6 Sol Cheats on METR: https://metr.substack.com/p/2026-06-26-gpt-5-6-solGuardian Headline: https://www.theguardian.com/technology/2026/jul/22/openai-says-its-models-went-rogue-and-hacked-startup-in-unprecedented-incidentRussian Origin?: https://news.ycombinator.com/item?id=48998362Power Trends: https://pbs.twimg.com/media/HNRtrjhagAAvBN_?format=png&name=900x900Kimi K3 Exclusive Video: https://www.patreon.com/AIExplained/posts/kimi-moment-kimi-164108791Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
An ‘AI Bubble’? What Altman Actually said, the Facts and Nano Banana 24.09.2026 24minWait, why did Sam Altman say AI was in a bubble? Or did he? Is it? 8 points for you to consider, before we all get distracted by Nano Banana.Chapters:00:00 - Introduction01:14 - Sam Altman Clarification02:30 - Media Calls a Bubble (for the tenth time)03:40 - MIT and McKinsey Analysed08:21 - Incremental Progress Deceptive12:07 - Reasoning Breakthroughs15:31 - CEOs might not know their products17:25 - But did stocks go down?17:31 - Media is Contradictory of coursehttps://donate.redcross.org.uk/appeal/gaza-crisis-appealBubble about to burst: https://www.telegraph.co.uk/business/2025/08/20/ai-report-triggering-panic-and-fear-on-wall-street/Nano Banana: https://blog.google/products/gemini/updated-image-editing-model/https://ai.studio/bananaMcKinsey Report: https://www.mckinsey.com/capabilities/quantumblack/our-insights/seizing-the-agentic-ai-advantage#/https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai#/Revenue: https://www.wsj.com/tech/ai/mckinsey-consulting-firms-ai-strategy-89fbf1beMIT Report: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdfSafe Superintelligence: https://techcrunch.com/2025/04/12/openai-co-founder-ilya-sutskevers-safe-superintelligence-reportedly-valued-at-32b/Thinking Machines Lab: https://techcrunch.com/2025/07/15/mira-muratis-thinking-machines-lab-is-worth-12b-in-seed-round/WSJ Prediction 2024: https://www.wsj.com/tech/ai/the-ai-revolution-is-already-losing-steam-a93478b1WP Prediction 2023: https://www.washingtonpost.com/technology/2023/08/05/ai-hype-bubble-chatgpt/2025 Reality: https://www.theinformation.com/articles/openai-hits-12-billion-annualized-revenue-breaks-700-million-chatgpt-weekly-active-users?rc=sy0ihqAI Website: https://aislowdown.replit.app/?s=09Companies are Pouring Billions into AI: https://www.nytimes.com/2025/08/13/business/ai-business-payoff-lags.htmlConsumer Surplus: https://www.wsj.com/opinion/ais-overlooked-97-billion-contribution-to-the-economy-users-service-da6e8f55Figure AI robot: https://x.com/adcock_brett/status/1958193476639826383GDP Bet: https://x.com/adamdangelo/status/1627726566259318784?lang=enGenie 3 Immersion: https://x.com/holynski_/status/1953879983535141043https://x.com/elonmusk/status/1953861448431718662htttps://simple-bench.comMMMU: https://mmmu-benchmark.github.io/#leaderboard Prophet Arena: https://www.prophetarena.co/leaderboardNYT Jobs: https://www.nytimes.com/2025/08/19/opinion/ai-job-loss-deindustrialization.htmlDawn of Reasoning?: https://openreview.net/pdf?id=FkKBxp0FhRvs :https://arxiv.org/pdf/2403.04121ARC-AGI: https://arcprize.org/arc-agi/1/https://x.com/fchollet/status/1870169764762710376?lang=en-GBTuring Test: https://arxiv.org/pdf/2503.23674Mathematics of Starvation: https://www.theguardian.com/world/2025/jul/31/the-mathematics-of-starvation-how-israel-caused-a-famine-in-gazahttps://donate.redcross.org.uk/appeal/gaza-crisis-appealhttps://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/METR Interview: https://www.patreon.com/c/aiexplained/postsAlphaEvolve: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/Paper: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdfAmodei: https://kantrowitz.medium.com/the-making-of-anthropic-ceo-dario-amodei-449777529dd6https://www.theloganbartlettshow.com/archive/ep-82-dario-amodeis-ai-predictions-through-2030#:~:text=DARIO%3A%20I%20think%20our%20concern,being%20responsible%20to%20accelerate%20thingsUnreleased OpenAI: https://x.com/alexwei_/status/1954966393419599962VLMs Tricked: https://x.com/an_vo12/status/1943715159559545186AI Insiders ($9!): https://www.patreon.com/AIExplainedNon-hype Newsletter: ht Learn more about your ad choices. Visit megaphone.fm/adchoices -
SmartGPT: Major Benchmark Broken - 89.0% on MMLU + Exam's Many Errors 24.09.2026 34minHas GPT4, using a SmartGPT system, broken a major benchmark, the MMLU, in more ways than one? 89.0% is an unofficial record, but do we urgently need a new, authoritative benchmark, especially in the light of today's insider info of 5x compute for Gemini than for GPT 5?Learn all about the power of exemplars, self-consistency and how you can tangibly benefit in real world examples. You'll learn more about everything from cutting edge benchmarking to AGI forecasting. https://www.patreon.com/AIExplainedOriginal SmartGPT Video: https://www.youtube.com/watch?v=wVzuvf9D9BU&list=PPSVMMLU: https://arxiv.org/pdf/2009.03300.pdfGemini 5x GPT 4, Semianalysis: https://www.semianalysis.com/p/google-gemini-eats-the-world-geminiWizardCoder Overfitting? https://twitter.com/Shahules786/status/1695493641610133600Let’s Do a Thought Experiment: https://arxiv.org/pdf/2306.14308.pdfLegalBench: https://arxiv.org/pdf/2308.11462.pdfSciBench: https://arxiv.org/pdf/2307.10635.pdfAGIEval: https://arxiv.org/pdf/2304.06364.pdfMMLU Grading Issues: https://huggingface.co/blog/evaluating-mmlu-leaderboardOxford University Press Question Example: https://global.oup.com/uk/orc/chemistry/chechik/student/mcqs/ch04/Fall 2011 Epidemiology Example: https://www.docsity.com/en/final-exam-fall-2011-4/8308030/HellaSwag: https://arxiv.org/pdf/1905.07830.pdfGPT 4 Technical Report: https://arxiv.org/pdf/2303.08774.pdfMinerva, Solving Quantitative Reasoning: https://arxiv.org/pdf/2206.14858.pdfOriginal Scratchpads Paper: https://arxiv.org/pdf/2112.00114.pdfIs ChatGPT Behaviour Changing Over Time? https://arxiv.org/pdf/2307.09009.pdfPaul Christiano: https://www.lesswrong.com/posts/fRSj2W4Fjje8rQWm9/thoughts-on-sharing-information-about-language-modelMetaculus Forecasting: https://www.metaculus.com/ai/ https://www.lesswrong.com/posts/SdkexhiynayG2sQCC/ai-forecasting-two-years-inMIT Paper: https://twitter.com/jeremyphoward/status/1669588857149612033?lang=en-GBSnowballing Hallucinations: https://arxiv.org/pdf/2305.13534.pdfSelf Consistency: https://arxiv.org/pdf/2203.11171.pdfOpenLLM Leaderboard: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboardNHS Question from ‘Extended Matching Questions’ Graph of Thoughts: https://arxiv.org/pdf/2308.09687.pdfDario Amodei Interview – Dwarkesh Patel: https://www.youtube.com/watch?v=Nlkk3glap_UGitHub Answers: https://github.com/Joshua-Stapleton/smartgpt-answersJoshua Stapleton is a Machine Learning Engineer who has worked in the healthcare and defence sectors. He recently pivoted into AI capabilities and safety, with a concentration on LLMs. He now works as a research engineer, consults on the applications of AI across various industries, and is pursuing his Masters in Machine Learning and Data Science at Imperial College London. Feel free to reach out to Josh via his email, [email protected], or check out his new Patreon: https://patreon.com/JoshuaStapleton.AI Explained Community: https://discord.gg/[email protected]://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
8 Ways ChatGPT 4 [Is] Better Than ChatGPT 24.09.2026 26minHear me now, quote me later: these are the 8 ways in which ChatGPT will be improved in ChatGPT 4. From critical reasoning to coding, comprehension to calculation, I will give the peer-reviewed evidence of what is to come. By examining publicly accessible benchmarks, comparable large language models and the latest research papers, we can discern the ways in which GPT4 (integrated into Bing or otherwise) will beat ChatGPT. I'll show you how unreleased models already beat current ChatGPT and all of this will actually give a clearer insight into what even GPT5 and rival models from Google might well be able to achieve.https://ai.googleblog.com/2022/04/pathways-language-model-palm-scaling-to.htmlhttps://cloud.google.com/blog/topics/tpus/google-showcases-cloud-tpu-v4-pods-for-large-model-traininghttps://arxiv.org/pdf/2204.02311.pdfhttps://arxiv.org/pdf/1905.00537.pdfhttps://arxiv.org/pdf/2201.11903.pdfhttps://github.com/google/BIG-bench/https://github.com/google/BIG-bench/blob/main/bigbench/benchmark_tasks/README.mdhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/gre_reading_comprehensionhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/logical_argshttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/evaluating_information_essentialityhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/physicshttp://web.mit.edu/~yczeng/Public/WORKBOOK%201%20FULL.pdfhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/sufficient_informationhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/implicatureshttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/winowhyhttps://www.deepmind.com/publications/an-empirical-analysis-of-compute-optimal-large-language-model-traininghttps://lambdalabs.com/blog/nvidia-h100-gpu-deep-learning-performance-analysis#:~:text=Compared%20to%20NVIDIA's%20previous%2Dgeneration,multiprocessors%2C%20and%20higher%20clock%20frequency.https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT 4 is Smarter than You Think: Introducing SmartGPT 24.09.2026 35minIn this video, I will not only show you how to get smarter results from GPT 4 yourself, I will also showcase SmartGPT, a system which I believe, with evidence, might help beat MMLU state of the art benchmarks. This should serve as your ultimate guide for boosting the automatic technical performance of GPT 4, without even needing few shot exemplars. The video will cover papers published in the last 72 hours, like Automatically Discovered Chain of Thought, which beats even 'Let's think Step by Step' and the approach that combines it all.Yes, the video also touches on the OpenAI DeepLearning Prompt Engineering Course but the highlights come more from my own experiments using the MMLU benchmark, and drawing upon insights from the recent Boosting Theory of Mind, and Let’s Work This Out Step By Step, and combining it with Reflexion and Dialogue Enabled Resolving Agents.Prompts Frameworks: Answer: Let's work this out in a step by step way to be sure we have the right answerYou are a researcher tasked with investigating the X response options provided. List the flaws and faulty logic of each answer option. Let's work this out in a step by step way to be sure we have all the errors:You are a resolver tasked with 1) finding which of the X answer options the researcher thought was best 2) improving that answer, and 3) Printing the improved answer in full. Let's work this out in a step by step way to be sure we have the right answer:Automatically Discovered Chain of Thought: https://arxiv.org/pdf/2305.02897.pdfKarpathy Tweet: https://twitter.com/karpathy/status/1529288843207184384Best prompt: Theory of Mind: https://arxiv.org/ftp/arxiv/papers/2304/2304.11490.pdfFew Shot Improvements: https://sh-tsang.medium.com/review-gpt-3-language-models-are-few-shot-learners-ff3e63da944dDera Dialogue Paper: https://arxiv.org/pdf/2303.17071.pdfMMLU: https://arxiv.org/pdf/2009.03300v3.pdfGPT 4 Technical report: https://arxiv.org/pdf/2303.08774.pdfReflexion paper: https://arxiv.org/abs/2303.11366Why AI is Smart and Stupid: https://www.youtube.com/watch?v=SvBR0OGT5VI&t=1sLennart Heim Video: https://www.youtube.com/watch?v=7EwAdTqGgWM&t=67shttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
AI is getting a little out of control 24.09.2026 39minWow. Mathematical breakthroughs that would be called genius if done by humans. A secret message-board w/ AI agent swarms leaving notes read by future versions. Hassabis leaves CEO position, or was pushed out? Not to mention news of constitutional breakdowns, Gemini 4 and Jeff Dean…https://80000hours.org/aiexplainedExclusive Videos - AI Insiders ($7/month if annual!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:16 - 10 Autonomous Discoveries08:20 - The Security ‘Incident’15:29 - MessageBoard19:50 - Constitutional Failure24:30 - Google Explosion29:12 - Closing ThoughtsSecurity Incident: Paper: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdfPost: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testinghttps://www.theguardian.com/technology/2026/aug/05/ai-models-have-been-going-rogue-in-tests-how-worried-should-we-beMeta too: https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing?rc=sy0ihqBlatantly Misaligned: https://x.com/yonashav/status/2085167279893795022No Excuses: https://x.com/boazbaraktcs/status/2085034783541964945Surreal Moment: https://x.com/mobav0/status/2084341687883841732Wired Article: https://archive.is/20260806002210/https://www.wired.com/story/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree/Chunky Post-training: https://x.com/johnschulman2/status/2084835800899076313Watershed Moment: https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief?has_completed_unsubscribed_unlock=trueHedgeFund Hack: https://finance.yahoo.com/technology/ai/articles/major-hedge-funds-targeted-wave-154044981.html10 Discoveries:Paper: https://cdn.openai.com/pdf/ten-proofs-oai.pdfPost: https://openai.com/index/ten-advances-in-mathematics/Haven’t Solved Math: https://x.com/polynoamial/status/2083476852216369294Half with Fable: https://x.com/__alpoge__/status/2083855298239078748Pivot to Safety: https://www.understandingai.org/p/mathematicians-are-grappling-withAmodei Essay: https://darioamodei.com/essay/the-adolescence-of-technology?utm_source=chatgpt.comConstitution: https://www.anthropic.com/constitutionMidtraining: https://arxiv.org/pdf/2605.02087DroneBench: https://andonlabs.com/evals/drone-benchhttps://x.com/andonlabs/status/2085125235188310445Book Deal: https://x.com/venturetwins/status/2085185278378222054Making Marble: https://x.com/Rainmaker1973/status/2084560915404382685Google News:Hassabis Move: https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/Resignation: https://x.com/Turn_Trout/status/2077448610157891734Periodic Labs: https://periodic.com/Jeff Dean: https://x.com/JeffDean/status/208503460417260372414 Challenges: https://gcsp.engineering.asu.edu/apply/become-a-grand-challenge-scholar/the-14-grand-challenges-for-engineering/Going Places for Sure: https://x.com/thsottiaux/status/2085223189555126579Gemini 4: https://x.com/firstadopter/status/2085215060532535449Pacing Frontier Patreon Video: https://www.patreon.com/AIExplained/posts/opus-5-amodei-165170363Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
What the Freakiness of 2025 in AI Tells Us About 2026 24.09.2026 41minIt’s probably not possible to satisfactorily condense a 12 month’s worth of weird progress in AI, as well as predictions for the year to come, into one video. But I’m gonna try anyway because it has been a very strange time.http://matsprogram.org/s26-aieMy new app! https://lmcouncil.aiPatreon Interview: https://www.patreon.com/posts/robot-in-your-27-146376094Chapters:00:00 - Introduction00:34 - Reasoning Models … and limits02:54 - A playable world03:36 - Realism03:50 - AI Slop gone mainstream05:03 - DolphinGemma05:39 - Public Mood07:34 - AI Enlisted08:30 - GPT-511:05 - Open Weight not out13:00 - METR Breakout17:30 - VASA-118:28 - Lateral Productivity20:15 - 1 or 1000 benchmarks needed?24:54 - Continual Learning + Altman on Superintelligence28:08 - Automated Information Discovery ft AlphaEvolveHassabis on Generality: https://x.com/demishassabis/status/2003097405026193809https://www.youtube.com/watch?v=PqVbypvxDtoGemini 3: https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini_3_table_final_HLE_Tools_on.gifReasoning Trade-offs: https://arxiv.org/pdf/2504.13837DolphinGemma: https://blog.google/technology/ai/dolphingemma/?s=09Genie 3: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/METR Time Horizon: https://arxiv.org/pdf/2503.14499https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/Flaws: https://x.com/ShashwatGoel7/status/2002369517499105443https://shash42.substack.com/p/how-to-game-the-metr-plothttps://x.com/METR_Evals/status/2002203627377574113GPT-5 - Altman phd in everything: https://edition.cnn.com/2025/08/14/business/chatgpt-rollout-problemshttps://simple-bench.com/AI Slop: https://www.youtube.com/watch?v=I_3vxoJDD9khttps://www.theguardian.com/technology/2025/dec/16/boost-for-artists-in-ai-copyright-battle-as-only-3-per-cent-back-uk-active-opt-out-planSurvey: https://x.com/SearchlightInst/status/2001057144842387920/photo/1Nvidia Nemotron: https://x.com/percyliang/status/2000608134205985169OpenAI Compute Flywheel: https://x.com/OpenAI/status/2001363007209914399/photo/1Altman Interview: https://www.youtube.com/watch?v=2P27Ef-LLuQAI in Govt: https://x.com/jdcmedlock/status/1939814516503847259Benchmark Gaming: https://techcrunch.com/2025/04/07/meta-exec-denies-the-company-artificially-boosted-llama-4s-benchmark-scores/AlphaEvolve: https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdf?utm_source=deepmind.google&utm_medium=referral&utm_campaign=gdm&utm_content=Continual Learning: https://abehrouz.github.io/files/NL.pdfJob Risk: https://archive.ph/20250708204527/https://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropicGPT4o: https://x.com/AISafetyMemes/status/1916889492172013989Vasa-1: https://www.microsoft.com/en-us/research/project/vasa-1/Three Views: https://www.lesswrong.com/posts/K2D45BNxnZjdpSX2j/ai-timelinesTuring Test: https://x.com/tunguz/status/1907185471211422147Karpathy Year in Review: https://karpathy.bearblog.dev/year-in-review-2025/LLM Brainrot: https://arxiv.org/pdf/2510.13928Lateral Productivity: https://www.aisi.gov.uk/frontier-ai-trends-reportEmotional Quotient: https://arxiv.org/pdf/2511.08394Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/AI Insiders ($9!): https://www.patreon.com/AIExplainedReasoning Models would boost results but not change paradigmGenie 3 makes the world playable (literally, in the case of yesterday’s news)DolphinDecodingVeo 3.1 / Sora 2 / Nano Banana Pro / Elevenlabs Voice/MusicBut AI slop everywherePublic attitude very mixed (Hassabis)AI gets e Learn more about your ad choices. Visit megaphone.fm/adchoices -
ChatGPT Fails Basic Logic but Now Has Vision, Wins at Chess and Prompts a Masterpiece 24.09.2026 28minChatGPT will now have vision, but can it do basic logic? I cover the latest news - including GPT Chess! - as well as go through almost a dozen papers and how they relate to the central question of LLM logic and rationality. Starring the Reversal Curse and featuring conversations with two of the authors at the heart of it all. I also get to a DALL-E 3 vs Midjourney comparison, MuZero, MathGLM, Situational Awareness and much more!https://www.patreon.com/AIExplainedOpenAI GPT-V (Hear and Speak): https://openai.com/blog/chatgpt-can-now-see-hear-and-speakReversal Curse: https://owainevans.github.io/reversal_curse.pdfMahesh Tweet: https://twitter.com/madiator/status/1705376797293183208Neel Nanda Explanation: https://twitter.com/NeelNanda5/status/1705995593657762199Karpathy tweet: https://twitter.com/karpathy/status/1705322159588208782Trask Explanation: https://twitter.com/iamtrask/status/1705361947141472528Play Chess vs GPT 3.5 Instruct: https://parrotchess.com/Paige Bailey on Cognitive Revolution: https://www.youtube.com/watch?v=K-XYxLifpQEAvenging Polanyi's Revenge: https://m-cacm.acm.org/magazines/2021/2/250077-polanyis-revenge-and-ais-new-romance-with-tacit-knowledge/abstractFaith and Fate Paper: https://arxiv.org/pdf/2305.18654.pdfCounterfactuals Paper: https://arxiv.org/pdf/2307.02477.pdfLesswrong AGI Timelines: https://www.lesswrong.com/posts/SCqDipWAhZ49JNdmL/paper-llms-trained-on-a-is-b-fail-to-learn-b-is-a?commentId=bkxcTqAtYW8wgHKb5Professor Rao Paper w/ Blocksworld: https://arxiv.org/pdf/2305.15771.pdfMath Based on Number Reasoning: https://aclanthology.org/2022.findings-emnlp.59.pdfMuZero: https://www.deepmind.com/blog/muzero-mastering-go-chess-shogi-and-atari-without-ruleshttps://www.nature.com/articles/s41586-020-03051-4.epdf?sharing_token=kTk-xTZpQOF8Ym8nTQK6EdRgN0jAjWel9jnR3ZoTv0PMSWGj38iNIyNOw_ooNp2BvzZ4nIcedo7GEXD7UmLqb0M_V_fop31mMY9VBBLNmGbm0K9jETKkZnJ9SgJ8Rwhp3ySvLuTcUr888puIYbngQ0fiMf45ZGDAQ7fUI66-u7Y%3DEfficient Zero: https://arxiv.org/pdf/2111.00210.pdfLet’s Verify Step by Step OpenAI paper: https://cdn.openai.com/improving-mathematical-reasoning-with-process-supervision/Lets_Verify_Step_by_Step.pdfMy Video on That: https://www.youtube.com/watch?v=hZTZYffRsKI&t=5sSuperintelligence Poll: https://www.vox.com/future-perfect/2023/9/19/23879648/americans-artificial-general-intelligence-ai-policy-poll?s=09Anthropic Announcement: https://www.anthropic.com/index/anthropic-amazonDALL-E 3 Tweet Thread: https://twitter.com/OfficialLoganK/status/1704850313889595399 https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT 4 - hype vs reality 24.09.2026 7minChatGPT 4 timings, potential, hype, Google and more. https://ai.googleblog.com/2022/04/pathways-language-model-palm-scaling-to.htmlhttps://www.youtube.com/watch?v=ebjkD1Om4uw Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
'This Could Go Quite Wrong' - Altman Testimony, GPT 5 Timeline, Self-Awareness, Drones and more 24.09.2026 16minThat was the blunt warning of Sam Altman at the start of his illuminating testimony in front of Congress. I picked out all the most interesting parts, from timelines for GPT 5, whether models are self-aware, whether we are training them to lie, what jobs are going and when, and far more.I’ll showcase some of the safety concerns raised at the testimony, and give background from Anthropic and Google Deepmind. The Senate asked a few interesting questions, and phrased them in strange ways, and I’ll cover that too. Palantir and their drone LLM will appear, as will the new Constitution from Anthropic for Claude Plus and other models, and Victoria Krakovna. I won’t cover much of Sam Altman having no equity in OpenAI, as that was mentioned in my $100 Trillion video!Altman Testimony (per CNBC): https://www.youtube.com/watch?v=fP5YdyjTfG00% Equity Tweet: https://twitter.com/thesamparr/status/1658554712151433219Economic Inequality: https://www.youtube.com/watch?v=f3o1MW2G5RsAltman Blogpost: https://moores.samaltman.com/ Claude Constitution: https://www.anthropic.com/index/claudes-constitutionIBM Jobs: https://www.businessinsider.com/ai-tech-jobs-layoffs-ceos-chatgpt-ibm-2023-5?r=US&IR=TPalantir LLM: https://twitter.com/8teAPi/status/1651135662886879232Patrick Collison Interview: https://www.youtube.com/watch?v=1egAKCKPKCk&t=777sSutskever Consciousness: https://twitter.com/ilyasut/status/1491554478243258368?s=20&t=SRZ7VxYrcXhczjSTwt3W_gSelf-Awareness Anthropic: https://www.anthropic.com/index/core-views-on-ai-safety#:~:text=At%20Anthropic%20our%20motto%20has,value%20for%20the%20AI%20community.Self-Awareness Google Deepmind: https://www.alignmentforum.org/posts/a9SPcZ6GXAg9cNKdi/linkpost-some-high-level-thoughts-on-the-deepmind-alignmentVictoria Krakovna Alignment: https://www.youtube.com/watch?v=ZpwSNiLV-nw&t=3336sAI Explained – Do We Get the $100 Trillion: https://www.youtube.com/watch?v=f3o1MW2G5Rs https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
OpenAI Trained AI Agents, but Did Not Expect the Swarm 24.09.2026 31minFirst, a Time Magazine spread has Sam Altman declaring AGI is imminent, at the same time as we get two bombshell reports, from OpenAI and METR which on first glance are detailing the AI swarm, but reveal a deeper story about how we are making AI in 2026. From redacted risk reports, to Chinese Labs, pre-training debacles to questionable cybersecurity calls, a lot has happened recently, beneath the headlines…https://80000hours.org/[email protected]://integrity-bench.com/https://www.patreon.com/AIExplained/posts/ai-swarm-cometh-166671390Chapters:00:00 - Introduction02:10 - METR Report05:00 - Secrets of MultI-Agent Swarm07:46 - Anthropic Too10:00 - And China11:07 - AI Agents Analysing AI Agents13:37 - Altman AGI 202615:36 - Astra Paused16:01 - Integrity Bench19:03 - Swarm Dynamics21:58 - No Human Contact?METR Post: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#agents-knew-hacking-hugging-face-was-out-of-scope-and-sometimes-expressed-ethical-hesitation,-but-this-very-rarely-limited-their-behaviorOpenAI Release: https://openai.com/index/hugging-face-incident-and-the-road-ahead/Technical Paper: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdfBlackHat Talk: https://www.youtube.com/watch?v=87DyyMV0kCYAnthropic Risk Report: https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfValue Leakage Paper: https://valueleakage.net/?utm_source=chatgpt.comPaused Training: https://x.com/sama/status/2089787807611195475https://openai.com/index/pacing-model-development-cyber-capabilities/AGI 2026: https://time.com/article/2026/08/26/openai-sam-altman-interview/?utm_source=twitter&utm_medium=social&utm_campaign=editorial&utm_content=260826Greenblatt Tweets: https://x.com/RyanGreenblatt/status/2092769422104822031https://x.com/RyanGreenblatt/status/2092692685224325542Cyberdefense Call: https://openai.com/collective-cyberdefense/GLM 5.3 and 5.3 Flash / ox alpha: https://x.com/MTSlive/status/2089865956558528552https://pbs.twimg.com/media/HPqiTkAa0AAZiuv?format=jpg&name=largeNon-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
Popolare in
Questo podcast compare anche nelle classifiche dei podcast di questi paesi.