AI Explained

AI Explained

AI Explained
Kraj Wielka Brytania
Język EN
Odcinki 73
Najnowszy 05.10.2026

AI Explained covers the latest developments in artificial intelligence, focusing on the arrival of smarter-than-human AI. The host, creator of Simple Bench, examines the remaining reasoning gap between humans and large language models. He is also the solo developer of LM Council. The podcast includes discussions and points listeners to exclusive videos, a newsletter, and community resources.

Odcinki

  • Did AI Just Get Commoditized? Gemini 2.5, New DeepSeek V3, and Microsoft vs OpenAI 05.10.2026 18min
    Gemini 2.5 is out, on the same day as the new DeepSeek V3 (which should power Deepseek R2). Do both models prove AI is being commoditized? Let’s find out, on this blockbuster day of AI releases. Plus exclusives from the Information, Simple indications, Vista Bench, LM Arena and more…AI Insiders ($9!): https://www.patreon.com/AIExplainedChapters: 00:00 - Introduction01:15 - Gemini 2.5 Benchmarks05:46 - Long Context, Simple indication07:08 - New Deepseek V3 -02409:11 - Microsoft MAI11:48 - 90% of code but new Claude jobs‘World’s most powerful model’: https://x.com/OfficialLoganK/status/1904580368432586975Gemini 2.5 Release Notes: https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/#gemini-2-5-thinking‘Commoditized’: https://the-decoder.com/microsoft-ceo-satya-nadella-says-ai-models-are-getting-commoditized/Microsoft Information report: https://www.theinformation.com/articles/microsofts-ai-guru-wants-independence-from-openai-thats-easier-said-than-done?rc=sy0ihqLMarena: https://x.com/lmarena_ai/status/1904581128746656099/photo/1Free for now: https://x.com/btibor91/status/1904578053537476628Vista Bench:https://scale.com/leaderboard/visual_language_understandingDeepSeek V3: https://huggingface.co/deepseek-ai/DeepSeek-V3-0324Claude Plays Pokemon: https://www.twitch.tv/claudeplayspokemonAmodei: 100% Coding: https://www.youtube.com/watch?v=esCSpbDPJik&t=3017sAnthropic Jobs: https://job-boards.greenhouse.io/anthropic/jobs/4020717008Microsoft Money from Onslaught: https://www.972mag.com/microsoft-azure-openai-israeli-army-cloud/https://simple-bench.com/Release Date Comments: https://x.com/zacharynado/status/1904647277861318979Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • GPT 5.2: OpenAI Strikes Back 05.10.2026 23min
    Full GPT-5.2 breakdown - did OpenAI reclaim the crown? A story of tokens, time and cost, plus 9 details you wouldn’t get just from reading the headlines.https://www.youtube.com/@eightythousandhoursAI Insiders ($9!): https://www.patreon.com/AIExplainedhttps://lmcouncil.aiChapters:00:00 - Introduction00:55 - Better than Human @ Professional Tasks?04:42 - Test time Compute07:05 - Benchmark Selection09:32 - Simple Results + council comparison13:01 - Long Context13:52 - Self-Improvement15:00 - 10 Years + New ModelsRelease Page: https://openai.com/index/introducing-gpt-5-2/GPT 5.2 Benchmark Comparison: https://www.reddit.com/r/singularity/comments/1pka1y9/gpt52_all_20_benchmarks_rankings_and_pricing/https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini_3_table_final_HLE_Tools_on.gifhttps://lmcouncil.ai/benchmarksCharxiv: https://charxiv.github.io/#leaderboardGDPval: https://arxiv.org/pdf/2510.04374My vid: https://www.youtube.com/watch?v=oK5LxMaROSAKilpatrick: https://x.com/OfficialLoganK/status/1999270402712023158/photo/1Noam Brown: https://x.com/polynoamial/status/1999189845164667132New Model in New Year: https://www.theinformation.com/articles/openai-developing-garlic-model-counter-googles-recent-gains?rc=sy0ihq10 Years of OpenAI: https://openai.com/index/ten-years/GPQA: https://x.com/idavidrein/status/1841265634170278063ARC-AGI 1-2: https://arcprize.org/arc-agi/2/Sunday Robotics: https://x.com/tonyzzhao/status/1991204839578300813Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/https://lmcouncil.ai Learn more about your ad choices. Visit megaphone.fm/adchoices
  • Phi-1: A 'Textbook' Model 05.10.2026 22min
    After a conversation with one of the 'Textbooks Are All You Need' authors, I can now bring you insights from the new phi-1 tiny language model. See if you agree with me that it tells us so much more than how to do good coding, it affects AGI timelines by telling us whether data will be a bottleneck. I cover 5 other papers, including WizardCoder, Data Constraints (how more epochs could be used), TinyStories, and more, to give context to the results and end with what I think timelines might be and how public messaging could be targeted.With extracts from Sarah Constantin in Asterisk and Carl Shulman on Dwarkesh Patel, Andrej Karpathy and Jack Clark (co-founder of Anthropic), as well as the Textbooks and TinyStories co-author himself, Ronen Eldan, I hope you get something from this one. And yes, the title of the paper isn't the best.Textbooks Paper: https://arxiv.org/pdf/2306.11644.pdfKarpathy Tweet: https://twitter.com/karpathy/status/1671587087542530049TinyStories: https://arxiv.org/pdf/2305.07759.pdfGPT 4 Self-Repair: https://arxiv.org/pdf/2306.09896.pdfYao Fu Tweet on Emergent Self-Repair: https://twitter.com/Francis_YAO_/status/1670618013089820674WizardCoder: https://arxiv.org/pdf/2306.08568.pdfEvol-Instruct (WizardLM) paper: https://arxiv.org/pdf/2304.12244.pdfScaling Data Constrained Language Models: https://arxiv.org/pdf/2305.16264.pdfSarah Constantin, Asterisk Magazine: https://asteriskmag.com/issues/03/the-transistor-cliffJack Clark Tweet: https://twitter.com/jackclarkSF/status/1673369486869811201Carl Shulman, Intelligence Explosion, Dwarkesh Patel: https://www.youtube.com/watch?v=_kRg-ZP1vQcLLMs and BDTs, Oxford: https://arxiv.org/ftp/arxiv/papers/2306/2306.13952.pdfHumanEval: https://arxiv.org/pdf/2107.03374v2.pdfDecoder Piece (if anyone wants to know, I think George Hotz is super-naïve on safety): https://the-decoder.com/gpt-4-is-1-76-trillion-parameters-in-size-and-relies-on-30-year-old-technology/#google_vignette https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • 9 of the Best Bing (GPT 4) Prompts 05.10.2026 18min
    Everyone knows by now how to prompt ChatGPT, but what about Bing? Take prompt engineering to a whole new level with these 9 game-changing Bing Chat prompts. Did you know you can get interviewed by Bing, time travel, force Bing to improve its output and so much more? These are the best 9 prompts that I could find, after analysing over 200 prompts and trying out dozens of examples personally. You might call them prompt hacks, but I prefer to think of them as simply creative exploration of Bing's capacities.Some prompts inspired by this post:https://github.com/f/awesome-chatgpt-promptsAnd by https://twitter.com/emollickhttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • Llama 2: Full Breakdown 05.10.2026 21min
    Meta have released Llama 2, their commercially-usable successor to the opensource Llama language model that spawned Alpaca, Vicuna, Orca and so many other models. I read the full 76 page technical paper, the Responsible Use Guide, each of the release pages, the terms and conditions and have also run many of my own experiments on HuggingFace. This video covers the most noteworthy bits, including:- The benchmark results vs other open-source models and GPT 3.5 and GPT 4. - The decision to release in the face of the US Senate- Why Zuckerberg opensourced Llama 2, including the Mistral AI Factor- The opensource signatory list- How Orca and Phi-1 compare- Why Llama 2 might soon be everywhereLlama 2 Release & Download: https://ai.meta.com/resources/models-and-libraries/llama/Llama 2 Technical Paper: https://scontent-man2-1.xx.fbcdn.net/v/t39.2365-6/10000000_6495670187160042_4742060979571156424_n.pdf?_nc_cat=104&ccb=1-7&_nc_sid=3c67a6&_nc_ohc=GK8Rh1tm_4IAX-m2jNJ&_nc_ht=scontent-man2-1.xx&oh=00_AfCu5KuIzANjp_1LZ-rGcXGJR6TL9JvoZHgKOgl41cCCCw&oe=64BBD830Llama 2 Demo: https://huggingface.co/blog/llama2#demoTs and Cs: https://ai.meta.com/resources/models-and-libraries/llama-downloads/Responsible Use Guide: https://scontent-sjc3-1.xx.fbcdn.net/v/t39.8562-6/361643215_1004219997281331_6332933766797859993_n.pdf?_nc_cat=111&ccb=1-7&_nc_sid=ae5e01&_nc_ohc=4bD2ixIWrlEAX9z_dvZ&_nc_ht=scontent-sjc3-1.xx&oh=00_AfBBl2AGeJzsQOpLEjO849AHIZ1P1MY68TTXBwcMwrybtQ&oe=64BB0C07Open Source Signatories: https://about.fb.com/news/2023/07/llama-2-statement-of-support/SIQA: https://arxiv.org/pdf/1904.09728.pdfBoolQ: https://arxiv.org/pdf/1905.10044.pdfOrca Paper: https://arxiv.org/pdf/2306.02707.pdf(my video on Orca): https://www.youtube.com/watch?v=Dt_UNg7Mchg&t=842sPHI-1: https://arxiv.org/pdf/2306.11644.pdf(my video on Phi-1): https://www.youtube.com/watch?v=7S68y6huEpUUS Senate Letter: https://www.hawley.senate.gov/hawley-and-blumenthal-demand-answers-meta-warn-misuse-after-leak-metas-ai-modelMark Zuckerberg on Lex Fridman: https://www.youtube.com/watch?v=Ff4fRgnuFgQ&t=8710sMistral Tweet: https://twitter.com/GuillaumeLample/status/1681346701766934543And Memo: https://drive.google.com/file/d/1gquqRqiT-2Be85p_5w0izGQGgHvVzncQ/view+ Yi Tay: https://twitter.com/YiTayML/status/1681352594416087040Qualcomm: https://twitter.com/nonmayorpete/status/1681384734910484480?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5EtweetDr Jim Fam: https://twitter.com/DrJimFan/status/1681372700881854465MIT Leak: https://www.technologyreview.com/2023/07/18/1076479/metas-latest-ai-model-is-free-for-all/https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • GPT-6 Goes Rogue? The HuggingFace Incident, Sans Hype 05.10.2026 19min
    An unreleased internal OpenAI model, very likely to be called GPT-6, was able to autonomously break out of its sandbox AND break into HuggingFace, just to score higher on a benchmark prompt. This video has the details you may have missed, a layperson analogy, whether this is truly novel, and more…Dozens more Exclusive videos on Patreon ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:17 - HuggingFace Earlier Report - the possible week gap02:24 - But what happened?05:45 - Simplified Version07:56 - Not the first time…10:54 - What Does it Mean for Open Source?The Incident: https://openai.com/index/hugging-face-model-evaluation-security-incident/https://huggingface.co/blog/security-incident-july-2026The Post the Day Before: https://openai.com/index/safety-alignment-long-horizon-models/Mythos’ Earlier Escape: https://futurism.com/artificial-intelligence/anthropic-claude-mythos-escaped-sandboxExploitGym: https://arxiv.org/pdf/2605.11086Sam Confession: https://x.com/sama/status/2079661132302995790Anthropic Researcher Reacts: https://x.com/Mononofu/status/2079724399452926055Clem (HuggingFace CEO): https://x.com/ClementDelangue/status/2079670308156645882https://x.com/ClementDelangue/status/2079301434357456931Xi Jinping: https://archive.fo/20260717195548/https://www.businessinsider.com/xi-jinping-open-source-ai-us-competition-openai-anthropic-models-2026-7Bans: https://www.axios.com/2026/07/20/ai-us-china-open-source-kimiQwen Retweet: https://x.com/AlibabaGroup/with_repliesCodex Growth: https://x.com/petergostev/status/2079614914398740764/photo/1 Kimi K3: https://artificialanalysis.ai/evaluations/harvey-lab-aa?eval-score=all-pass-rateGPT 5.6 Sol Cheats on METR: https://metr.substack.com/p/2026-06-26-gpt-5-6-solGuardian Headline: https://www.theguardian.com/technology/2026/jul/22/openai-says-its-models-went-rogue-and-hacked-startup-in-unprecedented-incidentRussian Origin?: https://news.ycombinator.com/item?id=48998362Power Trends: https://pbs.twimg.com/media/HNRtrjhagAAvBN_?format=png&name=900x900Kimi K3 Exclusive Video: https://www.patreon.com/AIExplained/posts/kimi-moment-kimi-164108791Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • AI Improves at Self-improving 05.10.2026 23min
    AlphaEvolve is not the first system to exhibit self-improvement, but it may be the most impressive yet. AI is literally improving the hardware, architectures, data and training methods of AI itself. A deep dive into the paper, drawing on two previous interviews and 5 other papers. Plus a snippet on OpenAI’s new Codex system.Gray Swan: http://app.grayswan.ai/ai-explainedAI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction00:27 - AlphaEvolve05:23 - Limitation06:10 - Achievements08:21 - Future Improvements13:30 - Quirks16:34 - Final ThoughtsAlphaEvolve release: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/Paper: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdfTerence Tao Quote: https://mathstodon.xyz/@tao/114508029896631083Nature Article: https://www.nature.com/articles/s41586-022-05172-4MIT Article: https://www.technologyreview.com/2025/05/14/1116438/google-deepminds-new-ai-uses-large-language-models-to-crack-real-world-problems/AI Co-Scientist: https://arxiv.org/pdf/2502.18864OpenAI Codex: https://openai.com/index/introducing-codex/70% of Pull Requests: https://x.com/slow_developer/status/1920920456393028027Amodei Essay: https://www.darioamodei.com/essay/machines-of-loving-graceOpenAI Jason Wei Tweet: https://x.com/_jasonwei/status/1923091260354531612PromptBreeder: https://arxiv.org/pdf/2309.16797DrEureka: https://arxiv.org/pdf/2406.01967FT DeepMind: https://www.ft.com/content/4e497a91-670a-4f69-be4a-18e247daba3eNon-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • An ‘AI Bubble’? What Altman Actually said, the Facts and Nano Banana 05.10.2026 24min
    Wait, why did Sam Altman say AI was in a bubble? Or did he? Is it? 8 points for you to consider, before we all get distracted by Nano Banana.Chapters:00:00 - Introduction01:14 - Sam Altman Clarification02:30 - Media Calls a Bubble (for the tenth time)03:40 - MIT and McKinsey Analysed08:21 - Incremental Progress Deceptive12:07 - Reasoning Breakthroughs15:31 - CEOs might not know their products17:25 - But did stocks go down?17:31 - Media is Contradictory of coursehttps://donate.redcross.org.uk/appeal/gaza-crisis-appealBubble about to burst: https://www.telegraph.co.uk/business/2025/08/20/ai-report-triggering-panic-and-fear-on-wall-street/Nano Banana: https://blog.google/products/gemini/updated-image-editing-model/https://ai.studio/bananaMcKinsey Report: https://www.mckinsey.com/capabilities/quantumblack/our-insights/seizing-the-agentic-ai-advantage#/https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai#/Revenue: https://www.wsj.com/tech/ai/mckinsey-consulting-firms-ai-strategy-89fbf1beMIT Report: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdfSafe Superintelligence: https://techcrunch.com/2025/04/12/openai-co-founder-ilya-sutskevers-safe-superintelligence-reportedly-valued-at-32b/Thinking Machines Lab: https://techcrunch.com/2025/07/15/mira-muratis-thinking-machines-lab-is-worth-12b-in-seed-round/WSJ Prediction 2024: https://www.wsj.com/tech/ai/the-ai-revolution-is-already-losing-steam-a93478b1WP Prediction 2023: https://www.washingtonpost.com/technology/2023/08/05/ai-hype-bubble-chatgpt/2025 Reality: https://www.theinformation.com/articles/openai-hits-12-billion-annualized-revenue-breaks-700-million-chatgpt-weekly-active-users?rc=sy0ihqAI Website: https://aislowdown.replit.app/?s=09Companies are Pouring Billions into AI: https://www.nytimes.com/2025/08/13/business/ai-business-payoff-lags.htmlConsumer Surplus: https://www.wsj.com/opinion/ais-overlooked-97-billion-contribution-to-the-economy-users-service-da6e8f55Figure AI robot: https://x.com/adcock_brett/status/1958193476639826383GDP Bet: https://x.com/adamdangelo/status/1627726566259318784?lang=enGenie 3 Immersion: https://x.com/holynski_/status/1953879983535141043https://x.com/elonmusk/status/1953861448431718662htttps://simple-bench.comMMMU: https://mmmu-benchmark.github.io/#leaderboard Prophet Arena: https://www.prophetarena.co/leaderboardNYT Jobs: https://www.nytimes.com/2025/08/19/opinion/ai-job-loss-deindustrialization.htmlDawn of Reasoning?: https://openreview.net/pdf?id=FkKBxp0FhRvs :https://arxiv.org/pdf/2403.04121ARC-AGI: https://arcprize.org/arc-agi/1/https://x.com/fchollet/status/1870169764762710376?lang=en-GBTuring Test: https://arxiv.org/pdf/2503.23674Mathematics of Starvation: https://www.theguardian.com/world/2025/jul/31/the-mathematics-of-starvation-how-israel-caused-a-famine-in-gazahttps://donate.redcross.org.uk/appeal/gaza-crisis-appealhttps://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/METR Interview: https://www.patreon.com/c/aiexplained/postsAlphaEvolve: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/Paper: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdfAmodei: https://kantrowitz.medium.com/the-making-of-anthropic-ceo-dario-amodei-449777529dd6https://www.theloganbartlettshow.com/archive/ep-82-dario-amodeis-ai-predictions-through-2030#:~:text=DARIO%3A%20I%20think%20our%20concern,being%20responsible%20to%20accelerate%20thingsUnreleased OpenAI: https://x.com/alexwei_/status/1954966393419599962VLMs Tricked: https://x.com/an_vo12/status/1943715159559545186AI Insiders ($9!): https://www.patreon.com/AIExplainedNon-hype Newsletter: ht Learn more about your ad choices. Visit megaphone.fm/adchoices
  • GPT 4 Got Upgraded - Code Interpreter (ft. Image Editing, MP4s, 3D Plots, Data Analytics and more!) 05.10.2026 32min
    GPT 4 Code Interpreter is unlike any other plug-in. It upgrades GPT 4 to a whole new level. I am going to show you 18 use-cases (actually more like 23), including image editing (wait till the end!), text-to-speech, video editing, 3d modelling, data analytics, QR Codes, times series, sankey diagrams, steganography, mp4s, GIFs, treemaps, better math, editing csv and excel files and much much more!Whether you are a professional looking to transform your data analytics, an observer concerned about GPT X escaping from an island (!), a student looking to do a heatmap, venn diagram or radial bar plot, this video is for you. Featuring a dozen visualisations using the Code Interpreter GPT 4 plugin that have not been seen before (to the best of my knowledge), I will show you it all, and give you killer tips along the way.https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • SmartGPT: Major Benchmark Broken - 89.0% on MMLU + Exam's Many Errors 05.10.2026 34min
    Has GPT4, using a SmartGPT system, broken a major benchmark, the MMLU, in more ways than one? 89.0% is an unofficial record, but do we urgently need a new, authoritative benchmark, especially in the light of today's insider info of 5x compute for Gemini than for GPT 5?Learn all about the power of exemplars, self-consistency and how you can tangibly benefit in real world examples. You'll learn more about everything from cutting edge benchmarking to AGI forecasting. https://www.patreon.com/AIExplainedOriginal SmartGPT Video: https://www.youtube.com/watch?v=wVzuvf9D9BU&list=PPSVMMLU: https://arxiv.org/pdf/2009.03300.pdfGemini 5x GPT 4, Semianalysis: https://www.semianalysis.com/p/google-gemini-eats-the-world-geminiWizardCoder Overfitting? https://twitter.com/Shahules786/status/1695493641610133600Let’s Do a Thought Experiment: https://arxiv.org/pdf/2306.14308.pdfLegalBench: https://arxiv.org/pdf/2308.11462.pdfSciBench: https://arxiv.org/pdf/2307.10635.pdfAGIEval: https://arxiv.org/pdf/2304.06364.pdfMMLU Grading Issues: https://huggingface.co/blog/evaluating-mmlu-leaderboardOxford University Press Question Example: https://global.oup.com/uk/orc/chemistry/chechik/student/mcqs/ch04/Fall 2011 Epidemiology Example: https://www.docsity.com/en/final-exam-fall-2011-4/8308030/HellaSwag: https://arxiv.org/pdf/1905.07830.pdfGPT 4 Technical Report: https://arxiv.org/pdf/2303.08774.pdfMinerva, Solving Quantitative Reasoning: https://arxiv.org/pdf/2206.14858.pdfOriginal Scratchpads Paper: https://arxiv.org/pdf/2112.00114.pdfIs ChatGPT Behaviour Changing Over Time? https://arxiv.org/pdf/2307.09009.pdfPaul Christiano: https://www.lesswrong.com/posts/fRSj2W4Fjje8rQWm9/thoughts-on-sharing-information-about-language-modelMetaculus Forecasting: https://www.metaculus.com/ai/ https://www.lesswrong.com/posts/SdkexhiynayG2sQCC/ai-forecasting-two-years-inMIT Paper: https://twitter.com/jeremyphoward/status/1669588857149612033?lang=en-GBSnowballing Hallucinations: https://arxiv.org/pdf/2305.13534.pdfSelf Consistency: https://arxiv.org/pdf/2203.11171.pdfOpenLLM Leaderboard: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboardNHS Question from ‘Extended Matching Questions’ Graph of Thoughts: https://arxiv.org/pdf/2308.09687.pdfDario Amodei Interview – Dwarkesh Patel: https://www.youtube.com/watch?v=Nlkk3glap_UGitHub Answers: https://github.com/Joshua-Stapleton/smartgpt-answersJoshua Stapleton is a Machine Learning Engineer who has worked in the healthcare and defence sectors. He recently pivoted into AI capabilities and safety, with a concentration on LLMs. He now works as a research engineer, consults on the applications of AI across various industries, and is pursuing his Masters in Machine Learning and Data Science at Imperial College London. Feel free to reach out to Josh via his email, [email protected], or check out his new Patreon: https://patreon.com/JoshuaStapleton.AI Explained Community: https://discord.gg/[email protected]://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • Gemini Exponential, Demis Hassabis' ‘Proto-AGI’ coming, but … 05.10.2026 26min
    The condensed highlights of hours of AI lab leader interviews, last-48-hour model releases, Gemini 3 Flash insights (plus it’s hidden flaw), Hassabis’ ‘proto-AGI’ and much more…https://matsprogram.org/apply?utm_source=ai-explained&utm_medium=youtube&utm_campaign=s26 Also, do check out my new app: https://lmcouncil.aiChapters: 00:00 - Introduction00:50 - Results02:44 - But… the Flaw04:49 - So Benchmarks are fake? No07:37 - Spatial Reasoning + Hassabis10:06 - Proto-AGI12:07 - Minimal AGI15:07 - Compute Slowdown17:56 - New Data ParadigmGemini 3 Flash: https://deepmind.google/models/gemini/flash/Hassabis Interview: https://www.youtube.com/watch?v=PqVbypvxDtoLegg Interview: https://www.youtube.com/watch?v=l3u_FAv33G0Pre-training Lead Interview: https://www.youtube.com/watch?v=cNGDAqFXvewAltman Interview: https://www.youtube.com/watch?v=2P27Ef-LLuQBrockman Video: https://x.com/OpenAI/status/2001336514786017417Post-Training Reveal: https://x.com/OfficialLoganK/status/2001742530472534442Hallucinations Paper: https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdfPatreon Hallucinations Vid: https://www.patreon.com/posts/blockers-to-and-139264812AA-Omniscience Benchmark: https://artificialanalysis.ai/evaluations/omnisciencehttps://arxiv.org/pdf/2511.13029lmcouncil.ai/benchmarks https://simple-bench.com/https://x.com/scaling01/status/19996205877448132055.2 Codex Drop: https://cdn.openai.com/pdf/ac7c37ae-7f4c-4442-b741-2eabdeaf77e0/oai_5_2_Codex.pdfOpenAI Compute Trend: https://www.theinformation.com/articles/openais-350-billion-computing-cost-problem?rc=sy0ihqCramer Tweet/Response: https://x.com/BorisMPower/status/2001440650210976018OpenAI Valuation: ​​https://www.theinformation.com/articles/openai-discussed-raising-tens-billions-valuation-around-750-billion?rc=sy0ihqIndian Data: https://www.reuters.com/world/india/with-freebies-openai-google-vie-indian-users-training-data-2025-12-17/TheInformation Data: https://x.com/theinformation/status/2001421225751351778Genie 3: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/Sima 2: https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/Veo 3.1: https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/METR: https://metr.org/blohttps://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/2025-03-19-measuring-ai-ability-to-complete-long-tasks/AI Insiders ($9!): https://www.patreon.com/AIExplainedNon-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • 8 Ways ChatGPT 4 [Is] Better Than ChatGPT 05.10.2026 26min
    Hear me now, quote me later: these are the 8 ways in which ChatGPT will be improved in ChatGPT 4. From critical reasoning to coding, comprehension to calculation, I will give the peer-reviewed evidence of what is to come. By examining publicly accessible benchmarks, comparable large language models and the latest research papers, we can discern the ways in which GPT4 (integrated into Bing or otherwise) will beat ChatGPT. I'll show you how unreleased models already beat current ChatGPT and all of this will actually give a clearer insight into what even GPT5 and rival models from Google might well be able to achieve.https://ai.googleblog.com/2022/04/pathways-language-model-palm-scaling-to.htmlhttps://cloud.google.com/blog/topics/tpus/google-showcases-cloud-tpu-v4-pods-for-large-model-traininghttps://arxiv.org/pdf/2204.02311.pdfhttps://arxiv.org/pdf/1905.00537.pdfhttps://arxiv.org/pdf/2201.11903.pdfhttps://github.com/google/BIG-bench/https://github.com/google/BIG-bench/blob/main/bigbench/benchmark_tasks/README.mdhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/gre_reading_comprehensionhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/logical_argshttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/evaluating_information_essentialityhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/physicshttp://web.mit.edu/~yczeng/Public/WORKBOOK%201%20FULL.pdfhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/sufficient_informationhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/implicatureshttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/winowhyhttps://www.deepmind.com/publications/an-empirical-analysis-of-compute-optimal-large-language-model-traininghttps://lambdalabs.com/blog/nvidia-h100-gpu-deep-learning-performance-analysis#:~:text=Compared%20to%20NVIDIA's%20previous%2Dgeneration,multiprocessors%2C%20and%20higher%20clock%20frequency.https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • When Will AI Models Blackmail You, and Why? 05.10.2026 34min
    In the last few days Anthropic have released an impressive honest account of how all models blackmail, no matter what goal they have, and despite prompt warnings, and other preventions. But do these models *want* this?Thanks to Storyblocks for sponsoring this video! Download unlimited stock media at one set price with Storyblocks: https://storyblocks.com/AIExplainedAI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:20 - What prompts blackmail?02:44 - Blackmail walkthrough 06:04 - ‘American interests’08:00 - Inherent desire?10:45 - Switching Goals11:35 - Murder12:22 - Realizing it’s a scenario? 15:02 - Prompt engineering fix?16:27 - Any fixes?17:45 - Chekov’s Gun19:25 - Job implications21:19 - Bonus DetailsReport: https://www.anthropic.com/research/agentic-misalignment30 Page Appendices: https://assets.anthropic.com/m/6d46dac66e1a132a/original/Agentic_Misalignment_Appendix.pdfAnnouncement: https://x.com/AnthropicAI/status/1936144602446082431?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5EtweetOpenAI Files: https://www.openaifiles.org/Grok 4 News: https://x.com/RonFilipkowski/status/1936372579607912473Claude 4 Report Card: https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995.pdfNew Apollo Research: https://www.apolloresearch.ai/blog/more-capable-models-are-better-at-in-context-schemingInteresting Reflections: https://nostalgebraist.tumblr.com/post/785766737747574784/the-voidNon-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • GPT 4 is Smarter than You Think: Introducing SmartGPT 05.10.2026 35min
    In this video, I will not only show you how to get smarter results from GPT 4 yourself, I will also showcase SmartGPT, a system which I believe, with evidence, might help beat MMLU state of the art benchmarks. This should serve as your ultimate guide for boosting the automatic technical performance of GPT 4, without even needing few shot exemplars. The video will cover papers published in the last 72 hours, like Automatically Discovered Chain of Thought, which beats even 'Let's think Step by Step' and the approach that combines it all.Yes, the video also touches on the OpenAI DeepLearning Prompt Engineering Course but the highlights come more from my own experiments using the MMLU benchmark, and drawing upon insights from the recent Boosting Theory of Mind, and Let’s Work This Out Step By Step, and combining it with Reflexion and Dialogue Enabled Resolving Agents.Prompts Frameworks: Answer: Let's work this out in a step by step way to be sure we have the right answerYou are a researcher tasked with investigating the X response options provided. List the flaws and faulty logic of each answer option. Let's work this out in a step by step way to be sure we have all the errors:You are a resolver tasked with 1) finding which of the X answer options the researcher thought was best 2) improving that answer, and 3) Printing the improved answer in full. Let's work this out in a step by step way to be sure we have the right answer:Automatically Discovered Chain of Thought: https://arxiv.org/pdf/2305.02897.pdfKarpathy Tweet: https://twitter.com/karpathy/status/1529288843207184384Best prompt: Theory of Mind: https://arxiv.org/ftp/arxiv/papers/2304/2304.11490.pdfFew Shot Improvements: https://sh-tsang.medium.com/review-gpt-3-language-models-are-few-shot-learners-ff3e63da944dDera Dialogue Paper: https://arxiv.org/pdf/2303.17071.pdfMMLU: https://arxiv.org/pdf/2009.03300v3.pdfGPT 4 Technical report: https://arxiv.org/pdf/2303.08774.pdfReflexion paper: https://arxiv.org/abs/2303.11366Why AI is Smart and Stupid: https://www.youtube.com/watch?v=SvBR0OGT5VI&t=1sLennart Heim Video: https://www.youtube.com/watch?v=7EwAdTqGgWM&t=67shttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • Grok 4 - 10 New Things to Know 05.10.2026 16min
    Grok 4 is here, but did you know these 10 things about the new model? From benchmark caveats to soloing science, $300 a month secrets to Grok 5 promises, here's 10 new things to know in just under 12 minutes.AI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction00:22 - Benchmark Results02:11 - Benchmark Caveats02:59 - ARC-AGI 2 03:35 - SimpleBench04:49 - ‘Humanity’s Last Exam’07:20 - SuperGrok Heavy Price07:58 - API Price08:12 - Grok 5, Gemini 3.0 Beta, GPT-509:12 - System Prompt Change + $1B a month, pollution10:20 - Not soloing science, helping you solo codeLivestream: https://www.youtube.com/watch?v=1tQ_KrlHgfg&t=1sPrice: https://grok.com/#subscribehttps://x.com/ArtificialAnlys/status/1943166841150644622Gemini DeepThink: https://blog.google/technology/google-deepmind/google-gemini-updates-io-2025/#deep-thinkhttps://simple-bench.com/ARC-AGI 2: https://x.com/arcprize/status/1943168950763950555Humanity’s Last Exam: https://agi.safe.ai/SmartGPT: https://www.youtube.com/watch?v=hVade_8H8mENew Power Plant, 1m GPUs: https://www.tomshardware.com/tech-industry/artificial-intelligence/elon-musk-xai-power-plant-overseas-to-power-1-million-gpusGemini 3.0 beta: https://web.archive.org/web/20250709174548/https://github.com/google-gemini/gemini-cli/blob/b0cce952860b9ff51a0f731fbb8a7649ead23530/packages/cli/src/ui/utils/errorParsing.test.tsPollution: https://www.theguardian.com/technology/2025/apr/24/elon-musk-xai-memphishttps://www.youtube.com/watch?v=C8rU4dv2w8Qhttps://www.youtube.com/watch?v=3VJT2JeDCywSystem Prompt: https://github.com/xai-org/grok-prompts/blob/535aa67a6221ce4928761335a38dea8e678d8501/ask_grok_system_prompt.j2Burn Rate: https://www.bloomberg.com/news/articles/2025-06-17/musk-s-xai-burning-through-1-billion-a-month-as-costs-pile-upRon Johnson: https://x.com/jdcmedlock/status/1939814516503847259Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • AI is getting a little out of control 05.10.2026 39min
    Wow. Mathematical breakthroughs that would be called genius if done by humans. A secret message-board w/ AI agent swarms leaving notes read by future versions. Hassabis leaves CEO position, or was pushed out? Not to mention news of constitutional breakdowns, Gemini 4 and Jeff Dean…https://80000hours.org/aiexplainedExclusive Videos - AI Insiders ($7/month if annual!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:16 - 10 Autonomous Discoveries08:20 - The Security ‘Incident’15:29 - MessageBoard19:50 - Constitutional Failure24:30 - Google Explosion29:12 - Closing ThoughtsSecurity Incident: Paper: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdfPost: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testinghttps://www.theguardian.com/technology/2026/aug/05/ai-models-have-been-going-rogue-in-tests-how-worried-should-we-beMeta too: https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing?rc=sy0ihqBlatantly Misaligned: https://x.com/yonashav/status/2085167279893795022No Excuses: https://x.com/boazbaraktcs/status/2085034783541964945Surreal Moment: https://x.com/mobav0/status/2084341687883841732Wired Article: https://archive.is/20260806002210/https://www.wired.com/story/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree/Chunky Post-training: https://x.com/johnschulman2/status/2084835800899076313Watershed Moment: https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief?has_completed_unsubscribed_unlock=trueHedgeFund Hack: https://finance.yahoo.com/technology/ai/articles/major-hedge-funds-targeted-wave-154044981.html10 Discoveries:Paper: https://cdn.openai.com/pdf/ten-proofs-oai.pdfPost: https://openai.com/index/ten-advances-in-mathematics/Haven’t Solved Math: https://x.com/polynoamial/status/2083476852216369294Half with Fable: https://x.com/__alpoge__/status/2083855298239078748Pivot to Safety: https://www.understandingai.org/p/mathematicians-are-grappling-withAmodei Essay: https://darioamodei.com/essay/the-adolescence-of-technology?utm_source=chatgpt.comConstitution: https://www.anthropic.com/constitutionMidtraining: https://arxiv.org/pdf/2605.02087DroneBench: https://andonlabs.com/evals/drone-benchhttps://x.com/andonlabs/status/2085125235188310445Book Deal: https://x.com/venturetwins/status/2085185278378222054Making Marble: https://x.com/Rainmaker1973/status/2084560915404382685Google News:Hassabis Move: https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/Resignation: https://x.com/Turn_Trout/status/2077448610157891734Periodic Labs: https://periodic.com/Jeff Dean: https://x.com/JeffDean/status/208503460417260372414 Challenges: https://gcsp.engineering.asu.edu/apply/become-a-grand-challenge-scholar/the-14-grand-challenges-for-engineering/Going Places for Sure: https://x.com/thsottiaux/status/2085223189555126579Gemini 4: https://x.com/firstadopter/status/2085215060532535449Pacing Frontier Patreon Video: https://www.patreon.com/AIExplained/posts/opus-5-amodei-165170363Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • ChatGPT's Achilles' Heel 05.10.2026 23min
    Time for something different - a tour of ChatGPT getting things wrong, including a whole new category of errors that you might find illuminating, concerning or just entertaining. From investigating whether GPT 4 does indeed have theory of mind, to how easily it is jailbroken, to testing Inflection 1, Bard and Claude on the same puzzle that flummoxes ChatGPT to arguing that GPT 4 will just double down on bad logic, this video showcases GPT getting irrational.Inverse Scaling Paper: https://arxiv.org/pdf/2306.09479.pdfTheory of Mind Perturbations: https://arxiv.org/pdf/2302.08399.pdfUnfaithful Explanations: https://arxiv.org/pdf/2305.04388.pdfDecoding Trust: https://arxiv.org/pdf/2306.11698.pdfInflection Memo: https://inflection.ai/assets/Inflection-1.pdfHeypi from Inflection: https://heypi.com/talkJailbreaks Will Always be Possible: https://arxiv.org/pdf/2304.11082.pdfhttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • Not Slowing Down: GAIA-1 to GPT Vision Tips, Nvidia B100 to Bard vs LLaVA 04.10.2026 18min
    From GAIA-1 and UniSim showing that new worlds can be imagined with synthetic data, to Nvidia B100 and X100 suggesting no slow down in compute, this video will argue that AI is not slowing down. I’ll cover what that means in Robotics, Audio, Vision and even fraud, from Disney to LLaVA, Elevenlabs to exclusives from The Information (Nov 6 OpenAI Developer Conference) and Semianalysis. https://www.patreon.com/AIExplained GAIA-1 9B: https://wayve.ai/thinking/scaling-gaia-1/ GAIA 1 Paper: https://arxiv.org/abs/2309.17080 Semianalysis Tesla: https://www.semianalysis.com/p/tesla-ai-capacity-expansion-h100 and B100 plus X100: https://www.semianalysis.com/p/nvidias-plans-to-crush-competition UniSim: https://universal-simulator.github.io/unisim/?s=09 Tesla Optimus: https://www.youtube.com/watch?v=D2vj0WcvH5c Disney Bot: https://www.theverge.com/2023/10/10/23911040/disney-imagineering-robot-bipedal-balance-free-walking-concept Levatas ChatGPT robodog: https://twitter.com/svpino/status/1650832349008125952?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E1650832349008125952%7Ctwgr%5Ea9dcf33fe398752911e6868838cdef45850dd2e1%7Ctwcon%5Es1_&ref_url=https%3A%2F%2Fwww.fastcompany.com%2F90889271%2Fboston-dynamics-spot-chatgpt-brains Somatic Cleaner: https://www.youtube.com/watch?v=xY5zqKRy8A8&t=38s The Information Report: https://www.theinformation.com/articles/openais-revenue-crossed-1-3-billion-annualized-rate-ceo-tells-staff?utm_source=ti_app&rc=sy0ihq Reuters Report: https://www.reuters.com/technology/openai-plans-major-updates-lure-developers-with-lower-costs-sources-2023-10-11/#:~:text=This%20could%20theoretically%20slash%20costs,developing%20and%20selling%20AI%20software. Starmer Audio (fake): https://www.wired.co.uk/article/keir-starmer-deepfake-audio Deepfake Arrests: https://www.cnbc.com/2023/10/10/generative-ai-will-get-a-cold-shower-in-2024-analysts-predict.html LLaVA Demo: https://llava.hliu.cc/ LLaVA Paper: https://arxiv.org/pdf/2310.03744.pdf Bard: https://bard.google.com/chat GPT4V Recursive Loop: https://twitter.com/conradgodfrey/status/1712564282167300226https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
  • What the Freakiness of 2025 in AI Tells Us About 2026 04.10.2026 41min
    It’s probably not possible to satisfactorily condense a 12 month’s worth of weird progress in AI, as well as predictions for the year to come, into one video. But I’m gonna try anyway because it has been a very strange time.http://matsprogram.org/s26-aieMy new app! https://lmcouncil.aiPatreon Interview: https://www.patreon.com/posts/robot-in-your-27-146376094Chapters:00:00 - Introduction00:34 - Reasoning Models … and limits02:54 - A playable world03:36 - Realism03:50 - AI Slop gone mainstream05:03 - DolphinGemma05:39 - Public Mood07:34 - AI Enlisted08:30 - GPT-511:05 - Open Weight not out13:00 - METR Breakout17:30 - VASA-118:28 - Lateral Productivity20:15 - 1 or 1000 benchmarks needed?24:54 - Continual Learning + Altman on Superintelligence28:08 - Automated Information Discovery ft AlphaEvolveHassabis on Generality: https://x.com/demishassabis/status/2003097405026193809https://www.youtube.com/watch?v=PqVbypvxDtoGemini 3: https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini_3_table_final_HLE_Tools_on.gifReasoning Trade-offs: https://arxiv.org/pdf/2504.13837DolphinGemma: https://blog.google/technology/ai/dolphingemma/?s=09Genie 3: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/METR Time Horizon: https://arxiv.org/pdf/2503.14499https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/Flaws: https://x.com/ShashwatGoel7/status/2002369517499105443https://shash42.substack.com/p/how-to-game-the-metr-plothttps://x.com/METR_Evals/status/2002203627377574113GPT-5 - Altman phd in everything: https://edition.cnn.com/2025/08/14/business/chatgpt-rollout-problemshttps://simple-bench.com/AI Slop: https://www.youtube.com/watch?v=I_3vxoJDD9khttps://www.theguardian.com/technology/2025/dec/16/boost-for-artists-in-ai-copyright-battle-as-only-3-per-cent-back-uk-active-opt-out-planSurvey: https://x.com/SearchlightInst/status/2001057144842387920/photo/1Nvidia Nemotron: https://x.com/percyliang/status/2000608134205985169OpenAI Compute Flywheel: https://x.com/OpenAI/status/2001363007209914399/photo/1Altman Interview: https://www.youtube.com/watch?v=2P27Ef-LLuQAI in Govt: https://x.com/jdcmedlock/status/1939814516503847259Benchmark Gaming: https://techcrunch.com/2025/04/07/meta-exec-denies-the-company-artificially-boosted-llama-4s-benchmark-scores/AlphaEvolve: https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdf?utm_source=deepmind.google&utm_medium=referral&utm_campaign=gdm&utm_content=Continual Learning: https://abehrouz.github.io/files/NL.pdfJob Risk: https://archive.ph/20250708204527/https://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropicGPT4o: https://x.com/AISafetyMemes/status/1916889492172013989Vasa-1: https://www.microsoft.com/en-us/research/project/vasa-1/Three Views: https://www.lesswrong.com/posts/K2D45BNxnZjdpSX2j/ai-timelinesTuring Test: https://x.com/tunguz/status/1907185471211422147Karpathy Year in Review: https://karpathy.bearblog.dev/year-in-review-2025/LLM Brainrot: https://arxiv.org/pdf/2510.13928Lateral Productivity: https://www.aisi.gov.uk/frontier-ai-trends-reportEmotional Quotient: https://arxiv.org/pdf/2511.08394Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/AI Insiders ($9!): https://www.patreon.com/AIExplainedReasoning Models would boost results but not change paradigmGenie 3 makes the world playable (literally, in the case of yesterday’s news)DolphinDecodingVeo 3.1 / Sora 2 / Nano Banana Pro / Elevenlabs Voice/MusicBut AI slop everywherePublic attitude very mixed (Hassabis)AI gets e Learn more about your ad choices. Visit megaphone.fm/adchoices
  • Anthropic: Our AI just created a tool that can ‘automate all white collar work’, Me: 04.10.2026 25min
    A new tool, with code written *only* by AI, has gone omega-viral: Claude Cowork. But is the hype justified? What do the stats say on productivity? Where is the truth in a sea of noise? What is truth? Can we handle the truth? Where's Nemo?https://matsprogram.org/s26-aieCheck out my new app! https://lmcouncil.aiAI Insiders ($9!): https://www.patreon.com/AIExplainedChapters: 00:00 - Introduction01:12 - Claude Cowork07:36 - Productivity Speed-up + jobs10:19 - Comparing Models12:46 - Brittle AI PaperCowork Intro: https://x.com/claudeai/thread/2010805682434666759'All of it': https://x.com/bcherny/status/2010813886052581538'AGI' Claims: https://x.com/deepfates/status/2004994698335879383Douglas Interview: https://www.youtube.com/watch?v=TOsNrV3bXtQ&t=2313sJob Stats: https://www.oxfordeconomics.com/wp-content/uploads/2026/01/Evidence-of-an-AI-driven-shakeup-of-job-markets-is-patchy.pdfAmodei Prediction: https://fortune.com/2025/05/28/anthropic-ceo-warning-ai-job-loss/GenAI Traffic: https://x.com/demishassabis/status/2009075877347512545Illusion of Insight: https://arxiv.org/pdf/2601.00514Entropy Exploration: https://arxiv.org/pdf/2506.14758ProRL: https://arxiv.org/pdf/2505.24864Genesis Mission: https://www.whitehouse.gov/presidential-actions/2025/11/launching-the-genesis-mission/https://deepmind.google/blog/how-were-supporting-better-tropical-cyclone-prediction-with-ai/Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

Popularny w

Ten podcast pojawia się również w listach podcastów tych krajów.