AI Explained
AI Explained
0
AI Explained covers the latest developments in artificial intelligence, focusing on the arrival of smarter-than-human AI. The host, creator of Simple Bench, examines the remaining reasoning gap between humans and large language models. He is also the solo developer of LM Council. The podcast includes discussions and points listeners to exclusive videos, a newsletter, and community resources.
Επεισόδια
-
9 of the Best Bing (GPT 4) Prompts 05.10.2026 18λEveryone knows by now how to prompt ChatGPT, but what about Bing? Take prompt engineering to a whole new level with these 9 game-changing Bing Chat prompts. Did you know you can get interviewed by Bing, time travel, force Bing to improve its output and so much more? These are the best 9 prompts that I could find, after analysing over 200 prompts and trying out dozens of examples personally. You might call them prompt hacks, but I prefer to think of them as simply creative exploration of Bing's capacities.Some prompts inspired by this post:https://github.com/f/awesome-chatgpt-promptsAnd by https://twitter.com/emollickhttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
Llama 2: Full Breakdown 05.10.2026 21λMeta have released Llama 2, their commercially-usable successor to the opensource Llama language model that spawned Alpaca, Vicuna, Orca and so many other models. I read the full 76 page technical paper, the Responsible Use Guide, each of the release pages, the terms and conditions and have also run many of my own experiments on HuggingFace. This video covers the most noteworthy bits, including:- The benchmark results vs other open-source models and GPT 3.5 and GPT 4. - The decision to release in the face of the US Senate- Why Zuckerberg opensourced Llama 2, including the Mistral AI Factor- The opensource signatory list- How Orca and Phi-1 compare- Why Llama 2 might soon be everywhereLlama 2 Release & Download: https://ai.meta.com/resources/models-and-libraries/llama/Llama 2 Technical Paper: https://scontent-man2-1.xx.fbcdn.net/v/t39.2365-6/10000000_6495670187160042_4742060979571156424_n.pdf?_nc_cat=104&ccb=1-7&_nc_sid=3c67a6&_nc_ohc=GK8Rh1tm_4IAX-m2jNJ&_nc_ht=scontent-man2-1.xx&oh=00_AfCu5KuIzANjp_1LZ-rGcXGJR6TL9JvoZHgKOgl41cCCCw&oe=64BBD830Llama 2 Demo: https://huggingface.co/blog/llama2#demoTs and Cs: https://ai.meta.com/resources/models-and-libraries/llama-downloads/Responsible Use Guide: https://scontent-sjc3-1.xx.fbcdn.net/v/t39.8562-6/361643215_1004219997281331_6332933766797859993_n.pdf?_nc_cat=111&ccb=1-7&_nc_sid=ae5e01&_nc_ohc=4bD2ixIWrlEAX9z_dvZ&_nc_ht=scontent-sjc3-1.xx&oh=00_AfBBl2AGeJzsQOpLEjO849AHIZ1P1MY68TTXBwcMwrybtQ&oe=64BB0C07Open Source Signatories: https://about.fb.com/news/2023/07/llama-2-statement-of-support/SIQA: https://arxiv.org/pdf/1904.09728.pdfBoolQ: https://arxiv.org/pdf/1905.10044.pdfOrca Paper: https://arxiv.org/pdf/2306.02707.pdf(my video on Orca): https://www.youtube.com/watch?v=Dt_UNg7Mchg&t=842sPHI-1: https://arxiv.org/pdf/2306.11644.pdf(my video on Phi-1): https://www.youtube.com/watch?v=7S68y6huEpUUS Senate Letter: https://www.hawley.senate.gov/hawley-and-blumenthal-demand-answers-meta-warn-misuse-after-leak-metas-ai-modelMark Zuckerberg on Lex Fridman: https://www.youtube.com/watch?v=Ff4fRgnuFgQ&t=8710sMistral Tweet: https://twitter.com/GuillaumeLample/status/1681346701766934543And Memo: https://drive.google.com/file/d/1gquqRqiT-2Be85p_5w0izGQGgHvVzncQ/view+ Yi Tay: https://twitter.com/YiTayML/status/1681352594416087040Qualcomm: https://twitter.com/nonmayorpete/status/1681384734910484480?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5EtweetDr Jim Fam: https://twitter.com/DrJimFan/status/1681372700881854465MIT Leak: https://www.technologyreview.com/2023/07/18/1076479/metas-latest-ai-model-is-free-for-all/https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT-6 Goes Rogue? The HuggingFace Incident, Sans Hype 05.10.2026 19λAn unreleased internal OpenAI model, very likely to be called GPT-6, was able to autonomously break out of its sandbox AND break into HuggingFace, just to score higher on a benchmark prompt. This video has the details you may have missed, a layperson analogy, whether this is truly novel, and more…Dozens more Exclusive videos on Patreon ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:17 - HuggingFace Earlier Report - the possible week gap02:24 - But what happened?05:45 - Simplified Version07:56 - Not the first time…10:54 - What Does it Mean for Open Source?The Incident: https://openai.com/index/hugging-face-model-evaluation-security-incident/https://huggingface.co/blog/security-incident-july-2026The Post the Day Before: https://openai.com/index/safety-alignment-long-horizon-models/Mythos’ Earlier Escape: https://futurism.com/artificial-intelligence/anthropic-claude-mythos-escaped-sandboxExploitGym: https://arxiv.org/pdf/2605.11086Sam Confession: https://x.com/sama/status/2079661132302995790Anthropic Researcher Reacts: https://x.com/Mononofu/status/2079724399452926055Clem (HuggingFace CEO): https://x.com/ClementDelangue/status/2079670308156645882https://x.com/ClementDelangue/status/2079301434357456931Xi Jinping: https://archive.fo/20260717195548/https://www.businessinsider.com/xi-jinping-open-source-ai-us-competition-openai-anthropic-models-2026-7Bans: https://www.axios.com/2026/07/20/ai-us-china-open-source-kimiQwen Retweet: https://x.com/AlibabaGroup/with_repliesCodex Growth: https://x.com/petergostev/status/2079614914398740764/photo/1 Kimi K3: https://artificialanalysis.ai/evaluations/harvey-lab-aa?eval-score=all-pass-rateGPT 5.6 Sol Cheats on METR: https://metr.substack.com/p/2026-06-26-gpt-5-6-solGuardian Headline: https://www.theguardian.com/technology/2026/jul/22/openai-says-its-models-went-rogue-and-hacked-startup-in-unprecedented-incidentRussian Origin?: https://news.ycombinator.com/item?id=48998362Power Trends: https://pbs.twimg.com/media/HNRtrjhagAAvBN_?format=png&name=900x900Kimi K3 Exclusive Video: https://www.patreon.com/AIExplained/posts/kimi-moment-kimi-164108791Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
AI Improves at Self-improving 05.10.2026 23λAlphaEvolve is not the first system to exhibit self-improvement, but it may be the most impressive yet. AI is literally improving the hardware, architectures, data and training methods of AI itself. A deep dive into the paper, drawing on two previous interviews and 5 other papers. Plus a snippet on OpenAI’s new Codex system.Gray Swan: http://app.grayswan.ai/ai-explainedAI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction00:27 - AlphaEvolve05:23 - Limitation06:10 - Achievements08:21 - Future Improvements13:30 - Quirks16:34 - Final ThoughtsAlphaEvolve release: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/Paper: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdfTerence Tao Quote: https://mathstodon.xyz/@tao/114508029896631083Nature Article: https://www.nature.com/articles/s41586-022-05172-4MIT Article: https://www.technologyreview.com/2025/05/14/1116438/google-deepminds-new-ai-uses-large-language-models-to-crack-real-world-problems/AI Co-Scientist: https://arxiv.org/pdf/2502.18864OpenAI Codex: https://openai.com/index/introducing-codex/70% of Pull Requests: https://x.com/slow_developer/status/1920920456393028027Amodei Essay: https://www.darioamodei.com/essay/machines-of-loving-graceOpenAI Jason Wei Tweet: https://x.com/_jasonwei/status/1923091260354531612PromptBreeder: https://arxiv.org/pdf/2309.16797DrEureka: https://arxiv.org/pdf/2406.01967FT DeepMind: https://www.ft.com/content/4e497a91-670a-4f69-be4a-18e247daba3eNon-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
An ‘AI Bubble’? What Altman Actually said, the Facts and Nano Banana 05.10.2026 24λWait, why did Sam Altman say AI was in a bubble? Or did he? Is it? 8 points for you to consider, before we all get distracted by Nano Banana.Chapters:00:00 - Introduction01:14 - Sam Altman Clarification02:30 - Media Calls a Bubble (for the tenth time)03:40 - MIT and McKinsey Analysed08:21 - Incremental Progress Deceptive12:07 - Reasoning Breakthroughs15:31 - CEOs might not know their products17:25 - But did stocks go down?17:31 - Media is Contradictory of coursehttps://donate.redcross.org.uk/appeal/gaza-crisis-appealBubble about to burst: https://www.telegraph.co.uk/business/2025/08/20/ai-report-triggering-panic-and-fear-on-wall-street/Nano Banana: https://blog.google/products/gemini/updated-image-editing-model/https://ai.studio/bananaMcKinsey Report: https://www.mckinsey.com/capabilities/quantumblack/our-insights/seizing-the-agentic-ai-advantage#/https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai#/Revenue: https://www.wsj.com/tech/ai/mckinsey-consulting-firms-ai-strategy-89fbf1beMIT Report: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdfSafe Superintelligence: https://techcrunch.com/2025/04/12/openai-co-founder-ilya-sutskevers-safe-superintelligence-reportedly-valued-at-32b/Thinking Machines Lab: https://techcrunch.com/2025/07/15/mira-muratis-thinking-machines-lab-is-worth-12b-in-seed-round/WSJ Prediction 2024: https://www.wsj.com/tech/ai/the-ai-revolution-is-already-losing-steam-a93478b1WP Prediction 2023: https://www.washingtonpost.com/technology/2023/08/05/ai-hype-bubble-chatgpt/2025 Reality: https://www.theinformation.com/articles/openai-hits-12-billion-annualized-revenue-breaks-700-million-chatgpt-weekly-active-users?rc=sy0ihqAI Website: https://aislowdown.replit.app/?s=09Companies are Pouring Billions into AI: https://www.nytimes.com/2025/08/13/business/ai-business-payoff-lags.htmlConsumer Surplus: https://www.wsj.com/opinion/ais-overlooked-97-billion-contribution-to-the-economy-users-service-da6e8f55Figure AI robot: https://x.com/adcock_brett/status/1958193476639826383GDP Bet: https://x.com/adamdangelo/status/1627726566259318784?lang=enGenie 3 Immersion: https://x.com/holynski_/status/1953879983535141043https://x.com/elonmusk/status/1953861448431718662htttps://simple-bench.comMMMU: https://mmmu-benchmark.github.io/#leaderboard Prophet Arena: https://www.prophetarena.co/leaderboardNYT Jobs: https://www.nytimes.com/2025/08/19/opinion/ai-job-loss-deindustrialization.htmlDawn of Reasoning?: https://openreview.net/pdf?id=FkKBxp0FhRvs :https://arxiv.org/pdf/2403.04121ARC-AGI: https://arcprize.org/arc-agi/1/https://x.com/fchollet/status/1870169764762710376?lang=en-GBTuring Test: https://arxiv.org/pdf/2503.23674Mathematics of Starvation: https://www.theguardian.com/world/2025/jul/31/the-mathematics-of-starvation-how-israel-caused-a-famine-in-gazahttps://donate.redcross.org.uk/appeal/gaza-crisis-appealhttps://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/METR Interview: https://www.patreon.com/c/aiexplained/postsAlphaEvolve: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/Paper: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdfAmodei: https://kantrowitz.medium.com/the-making-of-anthropic-ceo-dario-amodei-449777529dd6https://www.theloganbartlettshow.com/archive/ep-82-dario-amodeis-ai-predictions-through-2030#:~:text=DARIO%3A%20I%20think%20our%20concern,being%20responsible%20to%20accelerate%20thingsUnreleased OpenAI: https://x.com/alexwei_/status/1954966393419599962VLMs Tricked: https://x.com/an_vo12/status/1943715159559545186AI Insiders ($9!): https://www.patreon.com/AIExplainedNon-hype Newsletter: ht Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT 4 Got Upgraded - Code Interpreter (ft. Image Editing, MP4s, 3D Plots, Data Analytics and more!) 05.10.2026 32λGPT 4 Code Interpreter is unlike any other plug-in. It upgrades GPT 4 to a whole new level. I am going to show you 18 use-cases (actually more like 23), including image editing (wait till the end!), text-to-speech, video editing, 3d modelling, data analytics, QR Codes, times series, sankey diagrams, steganography, mp4s, GIFs, treemaps, better math, editing csv and excel files and much much more!Whether you are a professional looking to transform your data analytics, an observer concerned about GPT X escaping from an island (!), a student looking to do a heatmap, venn diagram or radial bar plot, this video is for you. Featuring a dozen visualisations using the Code Interpreter GPT 4 plugin that have not been seen before (to the best of my knowledge), I will show you it all, and give you killer tips along the way.https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
SmartGPT: Major Benchmark Broken - 89.0% on MMLU + Exam's Many Errors 05.10.2026 34λHas GPT4, using a SmartGPT system, broken a major benchmark, the MMLU, in more ways than one? 89.0% is an unofficial record, but do we urgently need a new, authoritative benchmark, especially in the light of today's insider info of 5x compute for Gemini than for GPT 5?Learn all about the power of exemplars, self-consistency and how you can tangibly benefit in real world examples. You'll learn more about everything from cutting edge benchmarking to AGI forecasting. https://www.patreon.com/AIExplainedOriginal SmartGPT Video: https://www.youtube.com/watch?v=wVzuvf9D9BU&list=PPSVMMLU: https://arxiv.org/pdf/2009.03300.pdfGemini 5x GPT 4, Semianalysis: https://www.semianalysis.com/p/google-gemini-eats-the-world-geminiWizardCoder Overfitting? https://twitter.com/Shahules786/status/1695493641610133600Let’s Do a Thought Experiment: https://arxiv.org/pdf/2306.14308.pdfLegalBench: https://arxiv.org/pdf/2308.11462.pdfSciBench: https://arxiv.org/pdf/2307.10635.pdfAGIEval: https://arxiv.org/pdf/2304.06364.pdfMMLU Grading Issues: https://huggingface.co/blog/evaluating-mmlu-leaderboardOxford University Press Question Example: https://global.oup.com/uk/orc/chemistry/chechik/student/mcqs/ch04/Fall 2011 Epidemiology Example: https://www.docsity.com/en/final-exam-fall-2011-4/8308030/HellaSwag: https://arxiv.org/pdf/1905.07830.pdfGPT 4 Technical Report: https://arxiv.org/pdf/2303.08774.pdfMinerva, Solving Quantitative Reasoning: https://arxiv.org/pdf/2206.14858.pdfOriginal Scratchpads Paper: https://arxiv.org/pdf/2112.00114.pdfIs ChatGPT Behaviour Changing Over Time? https://arxiv.org/pdf/2307.09009.pdfPaul Christiano: https://www.lesswrong.com/posts/fRSj2W4Fjje8rQWm9/thoughts-on-sharing-information-about-language-modelMetaculus Forecasting: https://www.metaculus.com/ai/ https://www.lesswrong.com/posts/SdkexhiynayG2sQCC/ai-forecasting-two-years-inMIT Paper: https://twitter.com/jeremyphoward/status/1669588857149612033?lang=en-GBSnowballing Hallucinations: https://arxiv.org/pdf/2305.13534.pdfSelf Consistency: https://arxiv.org/pdf/2203.11171.pdfOpenLLM Leaderboard: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboardNHS Question from ‘Extended Matching Questions’ Graph of Thoughts: https://arxiv.org/pdf/2308.09687.pdfDario Amodei Interview – Dwarkesh Patel: https://www.youtube.com/watch?v=Nlkk3glap_UGitHub Answers: https://github.com/Joshua-Stapleton/smartgpt-answersJoshua Stapleton is a Machine Learning Engineer who has worked in the healthcare and defence sectors. He recently pivoted into AI capabilities and safety, with a concentration on LLMs. He now works as a research engineer, consults on the applications of AI across various industries, and is pursuing his Masters in Machine Learning and Data Science at Imperial College London. Feel free to reach out to Josh via his email, [email protected], or check out his new Patreon: https://patreon.com/JoshuaStapleton.AI Explained Community: https://discord.gg/[email protected]://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
Gemini Exponential, Demis Hassabis' ‘Proto-AGI’ coming, but … 05.10.2026 26λThe condensed highlights of hours of AI lab leader interviews, last-48-hour model releases, Gemini 3 Flash insights (plus it’s hidden flaw), Hassabis’ ‘proto-AGI’ and much more…https://matsprogram.org/apply?utm_source=ai-explained&utm_medium=youtube&utm_campaign=s26 Also, do check out my new app: https://lmcouncil.aiChapters: 00:00 - Introduction00:50 - Results02:44 - But… the Flaw04:49 - So Benchmarks are fake? No07:37 - Spatial Reasoning + Hassabis10:06 - Proto-AGI12:07 - Minimal AGI15:07 - Compute Slowdown17:56 - New Data ParadigmGemini 3 Flash: https://deepmind.google/models/gemini/flash/Hassabis Interview: https://www.youtube.com/watch?v=PqVbypvxDtoLegg Interview: https://www.youtube.com/watch?v=l3u_FAv33G0Pre-training Lead Interview: https://www.youtube.com/watch?v=cNGDAqFXvewAltman Interview: https://www.youtube.com/watch?v=2P27Ef-LLuQBrockman Video: https://x.com/OpenAI/status/2001336514786017417Post-Training Reveal: https://x.com/OfficialLoganK/status/2001742530472534442Hallucinations Paper: https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdfPatreon Hallucinations Vid: https://www.patreon.com/posts/blockers-to-and-139264812AA-Omniscience Benchmark: https://artificialanalysis.ai/evaluations/omnisciencehttps://arxiv.org/pdf/2511.13029lmcouncil.ai/benchmarks https://simple-bench.com/https://x.com/scaling01/status/19996205877448132055.2 Codex Drop: https://cdn.openai.com/pdf/ac7c37ae-7f4c-4442-b741-2eabdeaf77e0/oai_5_2_Codex.pdfOpenAI Compute Trend: https://www.theinformation.com/articles/openais-350-billion-computing-cost-problem?rc=sy0ihqCramer Tweet/Response: https://x.com/BorisMPower/status/2001440650210976018OpenAI Valuation: https://www.theinformation.com/articles/openai-discussed-raising-tens-billions-valuation-around-750-billion?rc=sy0ihqIndian Data: https://www.reuters.com/world/india/with-freebies-openai-google-vie-indian-users-training-data-2025-12-17/TheInformation Data: https://x.com/theinformation/status/2001421225751351778Genie 3: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/Sima 2: https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/Veo 3.1: https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/METR: https://metr.org/blohttps://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/2025-03-19-measuring-ai-ability-to-complete-long-tasks/AI Insiders ($9!): https://www.patreon.com/AIExplainedNon-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
8 Ways ChatGPT 4 [Is] Better Than ChatGPT 05.10.2026 26λHear me now, quote me later: these are the 8 ways in which ChatGPT will be improved in ChatGPT 4. From critical reasoning to coding, comprehension to calculation, I will give the peer-reviewed evidence of what is to come. By examining publicly accessible benchmarks, comparable large language models and the latest research papers, we can discern the ways in which GPT4 (integrated into Bing or otherwise) will beat ChatGPT. I'll show you how unreleased models already beat current ChatGPT and all of this will actually give a clearer insight into what even GPT5 and rival models from Google might well be able to achieve.https://ai.googleblog.com/2022/04/pathways-language-model-palm-scaling-to.htmlhttps://cloud.google.com/blog/topics/tpus/google-showcases-cloud-tpu-v4-pods-for-large-model-traininghttps://arxiv.org/pdf/2204.02311.pdfhttps://arxiv.org/pdf/1905.00537.pdfhttps://arxiv.org/pdf/2201.11903.pdfhttps://github.com/google/BIG-bench/https://github.com/google/BIG-bench/blob/main/bigbench/benchmark_tasks/README.mdhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/gre_reading_comprehensionhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/logical_argshttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/evaluating_information_essentialityhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/physicshttp://web.mit.edu/~yczeng/Public/WORKBOOK%201%20FULL.pdfhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/sufficient_informationhttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/implicatureshttps://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/winowhyhttps://www.deepmind.com/publications/an-empirical-analysis-of-compute-optimal-large-language-model-traininghttps://lambdalabs.com/blog/nvidia-h100-gpu-deep-learning-performance-analysis#:~:text=Compared%20to%20NVIDIA's%20previous%2Dgeneration,multiprocessors%2C%20and%20higher%20clock%20frequency.https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
When Will AI Models Blackmail You, and Why? 05.10.2026 34λIn the last few days Anthropic have released an impressive honest account of how all models blackmail, no matter what goal they have, and despite prompt warnings, and other preventions. But do these models *want* this?Thanks to Storyblocks for sponsoring this video! Download unlimited stock media at one set price with Storyblocks: https://storyblocks.com/AIExplainedAI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:20 - What prompts blackmail?02:44 - Blackmail walkthrough 06:04 - ‘American interests’08:00 - Inherent desire?10:45 - Switching Goals11:35 - Murder12:22 - Realizing it’s a scenario? 15:02 - Prompt engineering fix?16:27 - Any fixes?17:45 - Chekov’s Gun19:25 - Job implications21:19 - Bonus DetailsReport: https://www.anthropic.com/research/agentic-misalignment30 Page Appendices: https://assets.anthropic.com/m/6d46dac66e1a132a/original/Agentic_Misalignment_Appendix.pdfAnnouncement: https://x.com/AnthropicAI/status/1936144602446082431?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5EtweetOpenAI Files: https://www.openaifiles.org/Grok 4 News: https://x.com/RonFilipkowski/status/1936372579607912473Claude 4 Report Card: https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995.pdfNew Apollo Research: https://www.apolloresearch.ai/blog/more-capable-models-are-better-at-in-context-schemingInteresting Reflections: https://nostalgebraist.tumblr.com/post/785766737747574784/the-voidNon-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT 4 is Smarter than You Think: Introducing SmartGPT 05.10.2026 35λIn this video, I will not only show you how to get smarter results from GPT 4 yourself, I will also showcase SmartGPT, a system which I believe, with evidence, might help beat MMLU state of the art benchmarks. This should serve as your ultimate guide for boosting the automatic technical performance of GPT 4, without even needing few shot exemplars. The video will cover papers published in the last 72 hours, like Automatically Discovered Chain of Thought, which beats even 'Let's think Step by Step' and the approach that combines it all.Yes, the video also touches on the OpenAI DeepLearning Prompt Engineering Course but the highlights come more from my own experiments using the MMLU benchmark, and drawing upon insights from the recent Boosting Theory of Mind, and Let’s Work This Out Step By Step, and combining it with Reflexion and Dialogue Enabled Resolving Agents.Prompts Frameworks: Answer: Let's work this out in a step by step way to be sure we have the right answerYou are a researcher tasked with investigating the X response options provided. List the flaws and faulty logic of each answer option. Let's work this out in a step by step way to be sure we have all the errors:You are a resolver tasked with 1) finding which of the X answer options the researcher thought was best 2) improving that answer, and 3) Printing the improved answer in full. Let's work this out in a step by step way to be sure we have the right answer:Automatically Discovered Chain of Thought: https://arxiv.org/pdf/2305.02897.pdfKarpathy Tweet: https://twitter.com/karpathy/status/1529288843207184384Best prompt: Theory of Mind: https://arxiv.org/ftp/arxiv/papers/2304/2304.11490.pdfFew Shot Improvements: https://sh-tsang.medium.com/review-gpt-3-language-models-are-few-shot-learners-ff3e63da944dDera Dialogue Paper: https://arxiv.org/pdf/2303.17071.pdfMMLU: https://arxiv.org/pdf/2009.03300v3.pdfGPT 4 Technical report: https://arxiv.org/pdf/2303.08774.pdfReflexion paper: https://arxiv.org/abs/2303.11366Why AI is Smart and Stupid: https://www.youtube.com/watch?v=SvBR0OGT5VI&t=1sLennart Heim Video: https://www.youtube.com/watch?v=7EwAdTqGgWM&t=67shttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
Grok 4 - 10 New Things to Know 05.10.2026 16λGrok 4 is here, but did you know these 10 things about the new model? From benchmark caveats to soloing science, $300 a month secrets to Grok 5 promises, here's 10 new things to know in just under 12 minutes.AI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction00:22 - Benchmark Results02:11 - Benchmark Caveats02:59 - ARC-AGI 2 03:35 - SimpleBench04:49 - ‘Humanity’s Last Exam’07:20 - SuperGrok Heavy Price07:58 - API Price08:12 - Grok 5, Gemini 3.0 Beta, GPT-509:12 - System Prompt Change + $1B a month, pollution10:20 - Not soloing science, helping you solo codeLivestream: https://www.youtube.com/watch?v=1tQ_KrlHgfg&t=1sPrice: https://grok.com/#subscribehttps://x.com/ArtificialAnlys/status/1943166841150644622Gemini DeepThink: https://blog.google/technology/google-deepmind/google-gemini-updates-io-2025/#deep-thinkhttps://simple-bench.com/ARC-AGI 2: https://x.com/arcprize/status/1943168950763950555Humanity’s Last Exam: https://agi.safe.ai/SmartGPT: https://www.youtube.com/watch?v=hVade_8H8mENew Power Plant, 1m GPUs: https://www.tomshardware.com/tech-industry/artificial-intelligence/elon-musk-xai-power-plant-overseas-to-power-1-million-gpusGemini 3.0 beta: https://web.archive.org/web/20250709174548/https://github.com/google-gemini/gemini-cli/blob/b0cce952860b9ff51a0f731fbb8a7649ead23530/packages/cli/src/ui/utils/errorParsing.test.tsPollution: https://www.theguardian.com/technology/2025/apr/24/elon-musk-xai-memphishttps://www.youtube.com/watch?v=C8rU4dv2w8Qhttps://www.youtube.com/watch?v=3VJT2JeDCywSystem Prompt: https://github.com/xai-org/grok-prompts/blob/535aa67a6221ce4928761335a38dea8e678d8501/ask_grok_system_prompt.j2Burn Rate: https://www.bloomberg.com/news/articles/2025-06-17/musk-s-xai-burning-through-1-billion-a-month-as-costs-pile-upRon Johnson: https://x.com/jdcmedlock/status/1939814516503847259Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
AI is getting a little out of control 05.10.2026 39λWow. Mathematical breakthroughs that would be called genius if done by humans. A secret message-board w/ AI agent swarms leaving notes read by future versions. Hassabis leaves CEO position, or was pushed out? Not to mention news of constitutional breakdowns, Gemini 4 and Jeff Dean…https://80000hours.org/aiexplainedExclusive Videos - AI Insiders ($7/month if annual!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:16 - 10 Autonomous Discoveries08:20 - The Security ‘Incident’15:29 - MessageBoard19:50 - Constitutional Failure24:30 - Google Explosion29:12 - Closing ThoughtsSecurity Incident: Paper: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdfPost: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testinghttps://www.theguardian.com/technology/2026/aug/05/ai-models-have-been-going-rogue-in-tests-how-worried-should-we-beMeta too: https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing?rc=sy0ihqBlatantly Misaligned: https://x.com/yonashav/status/2085167279893795022No Excuses: https://x.com/boazbaraktcs/status/2085034783541964945Surreal Moment: https://x.com/mobav0/status/2084341687883841732Wired Article: https://archive.is/20260806002210/https://www.wired.com/story/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree/Chunky Post-training: https://x.com/johnschulman2/status/2084835800899076313Watershed Moment: https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief?has_completed_unsubscribed_unlock=trueHedgeFund Hack: https://finance.yahoo.com/technology/ai/articles/major-hedge-funds-targeted-wave-154044981.html10 Discoveries:Paper: https://cdn.openai.com/pdf/ten-proofs-oai.pdfPost: https://openai.com/index/ten-advances-in-mathematics/Haven’t Solved Math: https://x.com/polynoamial/status/2083476852216369294Half with Fable: https://x.com/__alpoge__/status/2083855298239078748Pivot to Safety: https://www.understandingai.org/p/mathematicians-are-grappling-withAmodei Essay: https://darioamodei.com/essay/the-adolescence-of-technology?utm_source=chatgpt.comConstitution: https://www.anthropic.com/constitutionMidtraining: https://arxiv.org/pdf/2605.02087DroneBench: https://andonlabs.com/evals/drone-benchhttps://x.com/andonlabs/status/2085125235188310445Book Deal: https://x.com/venturetwins/status/2085185278378222054Making Marble: https://x.com/Rainmaker1973/status/2084560915404382685Google News:Hassabis Move: https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/Resignation: https://x.com/Turn_Trout/status/2077448610157891734Periodic Labs: https://periodic.com/Jeff Dean: https://x.com/JeffDean/status/208503460417260372414 Challenges: https://gcsp.engineering.asu.edu/apply/become-a-grand-challenge-scholar/the-14-grand-challenges-for-engineering/Going Places for Sure: https://x.com/thsottiaux/status/2085223189555126579Gemini 4: https://x.com/firstadopter/status/2085215060532535449Pacing Frontier Patreon Video: https://www.patreon.com/AIExplained/posts/opus-5-amodei-165170363Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
ChatGPT's Achilles' Heel 05.10.2026 23λTime for something different - a tour of ChatGPT getting things wrong, including a whole new category of errors that you might find illuminating, concerning or just entertaining. From investigating whether GPT 4 does indeed have theory of mind, to how easily it is jailbroken, to testing Inflection 1, Bard and Claude on the same puzzle that flummoxes ChatGPT to arguing that GPT 4 will just double down on bad logic, this video showcases GPT getting irrational.Inverse Scaling Paper: https://arxiv.org/pdf/2306.09479.pdfTheory of Mind Perturbations: https://arxiv.org/pdf/2302.08399.pdfUnfaithful Explanations: https://arxiv.org/pdf/2305.04388.pdfDecoding Trust: https://arxiv.org/pdf/2306.11698.pdfInflection Memo: https://inflection.ai/assets/Inflection-1.pdfHeypi from Inflection: https://heypi.com/talkJailbreaks Will Always be Possible: https://arxiv.org/pdf/2304.11082.pdfhttps://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
Not Slowing Down: GAIA-1 to GPT Vision Tips, Nvidia B100 to Bard vs LLaVA 04.10.2026 18λFrom GAIA-1 and UniSim showing that new worlds can be imagined with synthetic data, to Nvidia B100 and X100 suggesting no slow down in compute, this video will argue that AI is not slowing down. I’ll cover what that means in Robotics, Audio, Vision and even fraud, from Disney to LLaVA, Elevenlabs to exclusives from The Information (Nov 6 OpenAI Developer Conference) and Semianalysis. https://www.patreon.com/AIExplained GAIA-1 9B: https://wayve.ai/thinking/scaling-gaia-1/ GAIA 1 Paper: https://arxiv.org/abs/2309.17080 Semianalysis Tesla: https://www.semianalysis.com/p/tesla-ai-capacity-expansion-h100 and B100 plus X100: https://www.semianalysis.com/p/nvidias-plans-to-crush-competition UniSim: https://universal-simulator.github.io/unisim/?s=09 Tesla Optimus: https://www.youtube.com/watch?v=D2vj0WcvH5c Disney Bot: https://www.theverge.com/2023/10/10/23911040/disney-imagineering-robot-bipedal-balance-free-walking-concept Levatas ChatGPT robodog: https://twitter.com/svpino/status/1650832349008125952?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E1650832349008125952%7Ctwgr%5Ea9dcf33fe398752911e6868838cdef45850dd2e1%7Ctwcon%5Es1_&ref_url=https%3A%2F%2Fwww.fastcompany.com%2F90889271%2Fboston-dynamics-spot-chatgpt-brains Somatic Cleaner: https://www.youtube.com/watch?v=xY5zqKRy8A8&t=38s The Information Report: https://www.theinformation.com/articles/openais-revenue-crossed-1-3-billion-annualized-rate-ceo-tells-staff?utm_source=ti_app&rc=sy0ihq Reuters Report: https://www.reuters.com/technology/openai-plans-major-updates-lure-developers-with-lower-costs-sources-2023-10-11/#:~:text=This%20could%20theoretically%20slash%20costs,developing%20and%20selling%20AI%20software. Starmer Audio (fake): https://www.wired.co.uk/article/keir-starmer-deepfake-audio Deepfake Arrests: https://www.cnbc.com/2023/10/10/generative-ai-will-get-a-cold-shower-in-2024-analysts-predict.html LLaVA Demo: https://llava.hliu.cc/ LLaVA Paper: https://arxiv.org/pdf/2310.03744.pdf Bard: https://bard.google.com/chat GPT4V Recursive Loop: https://twitter.com/conradgodfrey/status/1712564282167300226https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
What the Freakiness of 2025 in AI Tells Us About 2026 04.10.2026 41λIt’s probably not possible to satisfactorily condense a 12 month’s worth of weird progress in AI, as well as predictions for the year to come, into one video. But I’m gonna try anyway because it has been a very strange time.http://matsprogram.org/s26-aieMy new app! https://lmcouncil.aiPatreon Interview: https://www.patreon.com/posts/robot-in-your-27-146376094Chapters:00:00 - Introduction00:34 - Reasoning Models … and limits02:54 - A playable world03:36 - Realism03:50 - AI Slop gone mainstream05:03 - DolphinGemma05:39 - Public Mood07:34 - AI Enlisted08:30 - GPT-511:05 - Open Weight not out13:00 - METR Breakout17:30 - VASA-118:28 - Lateral Productivity20:15 - 1 or 1000 benchmarks needed?24:54 - Continual Learning + Altman on Superintelligence28:08 - Automated Information Discovery ft AlphaEvolveHassabis on Generality: https://x.com/demishassabis/status/2003097405026193809https://www.youtube.com/watch?v=PqVbypvxDtoGemini 3: https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini_3_table_final_HLE_Tools_on.gifReasoning Trade-offs: https://arxiv.org/pdf/2504.13837DolphinGemma: https://blog.google/technology/ai/dolphingemma/?s=09Genie 3: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/METR Time Horizon: https://arxiv.org/pdf/2503.14499https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/Flaws: https://x.com/ShashwatGoel7/status/2002369517499105443https://shash42.substack.com/p/how-to-game-the-metr-plothttps://x.com/METR_Evals/status/2002203627377574113GPT-5 - Altman phd in everything: https://edition.cnn.com/2025/08/14/business/chatgpt-rollout-problemshttps://simple-bench.com/AI Slop: https://www.youtube.com/watch?v=I_3vxoJDD9khttps://www.theguardian.com/technology/2025/dec/16/boost-for-artists-in-ai-copyright-battle-as-only-3-per-cent-back-uk-active-opt-out-planSurvey: https://x.com/SearchlightInst/status/2001057144842387920/photo/1Nvidia Nemotron: https://x.com/percyliang/status/2000608134205985169OpenAI Compute Flywheel: https://x.com/OpenAI/status/2001363007209914399/photo/1Altman Interview: https://www.youtube.com/watch?v=2P27Ef-LLuQAI in Govt: https://x.com/jdcmedlock/status/1939814516503847259Benchmark Gaming: https://techcrunch.com/2025/04/07/meta-exec-denies-the-company-artificially-boosted-llama-4s-benchmark-scores/AlphaEvolve: https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdf?utm_source=deepmind.google&utm_medium=referral&utm_campaign=gdm&utm_content=Continual Learning: https://abehrouz.github.io/files/NL.pdfJob Risk: https://archive.ph/20250708204527/https://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropicGPT4o: https://x.com/AISafetyMemes/status/1916889492172013989Vasa-1: https://www.microsoft.com/en-us/research/project/vasa-1/Three Views: https://www.lesswrong.com/posts/K2D45BNxnZjdpSX2j/ai-timelinesTuring Test: https://x.com/tunguz/status/1907185471211422147Karpathy Year in Review: https://karpathy.bearblog.dev/year-in-review-2025/LLM Brainrot: https://arxiv.org/pdf/2510.13928Lateral Productivity: https://www.aisi.gov.uk/frontier-ai-trends-reportEmotional Quotient: https://arxiv.org/pdf/2511.08394Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/AI Insiders ($9!): https://www.patreon.com/AIExplainedReasoning Models would boost results but not change paradigmGenie 3 makes the world playable (literally, in the case of yesterday’s news)DolphinDecodingVeo 3.1 / Sora 2 / Nano Banana Pro / Elevenlabs Voice/MusicBut AI slop everywherePublic attitude very mixed (Hassabis)AI gets e Learn more about your ad choices. Visit megaphone.fm/adchoices -
Anthropic: Our AI just created a tool that can ‘automate all white collar work’, Me: 04.10.2026 25λA new tool, with code written *only* by AI, has gone omega-viral: Claude Cowork. But is the hype justified? What do the stats say on productivity? Where is the truth in a sea of noise? What is truth? Can we handle the truth? Where's Nemo?https://matsprogram.org/s26-aieCheck out my new app! https://lmcouncil.aiAI Insiders ($9!): https://www.patreon.com/AIExplainedChapters: 00:00 - Introduction01:12 - Claude Cowork07:36 - Productivity Speed-up + jobs10:19 - Comparing Models12:46 - Brittle AI PaperCowork Intro: https://x.com/claudeai/thread/2010805682434666759'All of it': https://x.com/bcherny/status/2010813886052581538'AGI' Claims: https://x.com/deepfates/status/2004994698335879383Douglas Interview: https://www.youtube.com/watch?v=TOsNrV3bXtQ&t=2313sJob Stats: https://www.oxfordeconomics.com/wp-content/uploads/2026/01/Evidence-of-an-AI-driven-shakeup-of-job-markets-is-patchy.pdfAmodei Prediction: https://fortune.com/2025/05/28/anthropic-ceo-warning-ai-job-loss/GenAI Traffic: https://x.com/demishassabis/status/2009075877347512545Illusion of Insight: https://arxiv.org/pdf/2601.00514Entropy Exploration: https://arxiv.org/pdf/2506.14758ProRL: https://arxiv.org/pdf/2505.24864Genesis Mission: https://www.whitehouse.gov/presidential-actions/2025/11/launching-the-genesis-mission/https://deepmind.google/blog/how-were-supporting-better-tropical-cyclone-prediction-with-ai/Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
ChatGPT Fails Basic Logic but Now Has Vision, Wins at Chess and Prompts a Masterpiece 04.10.2026 28λChatGPT will now have vision, but can it do basic logic? I cover the latest news - including GPT Chess! - as well as go through almost a dozen papers and how they relate to the central question of LLM logic and rationality. Starring the Reversal Curse and featuring conversations with two of the authors at the heart of it all. I also get to a DALL-E 3 vs Midjourney comparison, MuZero, MathGLM, Situational Awareness and much more!https://www.patreon.com/AIExplainedOpenAI GPT-V (Hear and Speak): https://openai.com/blog/chatgpt-can-now-see-hear-and-speakReversal Curse: https://owainevans.github.io/reversal_curse.pdfMahesh Tweet: https://twitter.com/madiator/status/1705376797293183208Neel Nanda Explanation: https://twitter.com/NeelNanda5/status/1705995593657762199Karpathy tweet: https://twitter.com/karpathy/status/1705322159588208782Trask Explanation: https://twitter.com/iamtrask/status/1705361947141472528Play Chess vs GPT 3.5 Instruct: https://parrotchess.com/Paige Bailey on Cognitive Revolution: https://www.youtube.com/watch?v=K-XYxLifpQEAvenging Polanyi's Revenge: https://m-cacm.acm.org/magazines/2021/2/250077-polanyis-revenge-and-ais-new-romance-with-tacit-knowledge/abstractFaith and Fate Paper: https://arxiv.org/pdf/2305.18654.pdfCounterfactuals Paper: https://arxiv.org/pdf/2307.02477.pdfLesswrong AGI Timelines: https://www.lesswrong.com/posts/SCqDipWAhZ49JNdmL/paper-llms-trained-on-a-is-b-fail-to-learn-b-is-a?commentId=bkxcTqAtYW8wgHKb5Professor Rao Paper w/ Blocksworld: https://arxiv.org/pdf/2305.15771.pdfMath Based on Number Reasoning: https://aclanthology.org/2022.findings-emnlp.59.pdfMuZero: https://www.deepmind.com/blog/muzero-mastering-go-chess-shogi-and-atari-without-ruleshttps://www.nature.com/articles/s41586-020-03051-4.epdf?sharing_token=kTk-xTZpQOF8Ym8nTQK6EdRgN0jAjWel9jnR3ZoTv0PMSWGj38iNIyNOw_ooNp2BvzZ4nIcedo7GEXD7UmLqb0M_V_fop31mMY9VBBLNmGbm0K9jETKkZnJ9SgJ8Rwhp3ySvLuTcUr888puIYbngQ0fiMf45ZGDAQ7fUI66-u7Y%3DEfficient Zero: https://arxiv.org/pdf/2111.00210.pdfLet’s Verify Step by Step OpenAI paper: https://cdn.openai.com/improving-mathematical-reasoning-with-process-supervision/Lets_Verify_Step_by_Step.pdfMy Video on That: https://www.youtube.com/watch?v=hZTZYffRsKI&t=5sSuperintelligence Poll: https://www.vox.com/future-perfect/2023/9/19/23879648/americans-artificial-general-intelligence-ai-policy-poll?s=09Anthropic Announcement: https://www.anthropic.com/index/anthropic-amazonDALL-E 3 Tweet Thread: https://twitter.com/OfficialLoganK/status/1704850313889595399 https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
Enter PaLM 2 (New Bard): Full Breakdown - 92 Pages Read and Gemini Before GPT 5? Google I/O 04.10.2026 23λGoogle puts it foot on the accelerator, casting aside safety concerns to not only release a GPT 4 -competitive model, PaLM 2, but also announce that they are already training Gemini, a GPT 5 competitor [likely on TPU v5 chips]. This is truly a major day in AI history, and I try to cover it all. I'll show the benchmarks in which PaLM (which now powers Bard) beats GPT 4, and detail how they use SmartGPT-like techniques to boost performance. Crazily enough, PaLM 2 beats even Google Translate, due in large part to the text it was trained on. We'll talk coding in Bard, translation, MMLU, Big Bench, and much more.I'll end on the Universal Translator deepfakes and the underwhelming results from Sundar Pichai and Sam Altman's trip to the White House and what Hinton says about it all. On a more positive note, I cover Med PaLM 2, which could genuinely save thousands of lives. PaLM 2 Technical Report: https://ai.google/static/documents/palm2techreport.pdfRelease Notes Google Blog: https://blog.google/technology/ai/google-palm-2-ai-large-language-model/Bard Access: https://bard.google.com/Scaling Transformer to 1M tokens: https://arxiv.org/pdf/2304.11062.pdfGPT 4 Technical Report: https://arxiv.org/pdf/2303.08774.pdfBard Languages: https://support.google.com/bard/answer/13575153?hl=enSelf Consistency Paper: https://arxiv.org/pdf/2203.11171.pdfAre Emergent Abilities a Mirage: https://arxiv.org/pdf/2304.15004.pdfSparks of AGI Paper: https://arxiv.org/pdf/2303.12712.pdfBig Bench Hard: https://github.com/suzgunmirac/BIG-Bench-HardGoogle Keynote: https://www.youtube.com/watch?v=cNfINi5CNbYGemini: https://www.youtube.com/watch?v=1UvUjTaJRz0Med PaLM 2: https://www.youtube.com/watch?v=k_-Z_TkHMqATPU v5: https://ai.googleblog.com/2022/01/google-research-themes-from-2021-and.htmlHinton Warning: https://www.youtube.com/watch?v=FAbsoxQtUwMWhite House Readout: https://www.whitehouse.gov/briefing-room/statements-releases/2023/05/04/readout-of-white-house-meeting-with-ceos-on-advancing-responsible-artificial-intelligence-innovation/https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices -
GPT 4 - hype vs reality 04.10.2026 7λChatGPT 4 timings, potential, hype, Google and more. https://ai.googleblog.com/2022/04/pathways-language-model-palm-scaling-to.htmlhttps://www.youtube.com/watch?v=ebjkD1Om4uw Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices
Δημοφιλές σε
Αυτό το podcast εμφανίζεται και στις λίστες podcasts αυτών των χωρών.