Search Off the Record

Search Off the Record

Google
Държава Съединени щати
Жанрове Технология
Език EN
Епизоди 100
Последен 30.07.2026

Search Off the Record takes you behind the scenes of Google Search and its inner workings. In each episode, the folks from the Search Relations team give background info on decision-making behind launches, feature prioritization in Search Console, and projects Google Search teams are working on. They share fun stories from conferences and their day-to-day working life at Google, and dive into trending conversations in the SEO community.

Епизоди

  • Should you block your Search result pages? - transcript 30.07.2026
  • Should you block your Search result pages? 30.07.2026 28мин
    Is your website's internal search feature secretly acting as an open invitation for crawling lots and lots? In this episode of Search off the Record, Martin Splitt and John Mueller pull back the curtain on how internal search results pages can turn into "infinite crawl spaces" that trap Googlebot, waste your crawl budget, and spike your database load. They break down the critical technical differences between blocking search pages via robots.txt versus noindex tags, and expose the massive security liabilities of leaving these pages indexable.  In this episode, you'll learn: The "Infinite Crawl Space" Concept: How Googlebot treats internal search functions as infinite crawl spaces that can generate an endless loop of new URLs. Server Strain & Performance: Why uncached internal search pages force constant database lookups and ranking calculations, slowing down your website for real users. Robots.txt vs. Noindex: The distinct technical trade-offs of using a broad robots.txt disallow rule versus a robots meta tag or HTTP header noindex. The Spam Vector Threat: How bad actors search for pharmaceutical, adult, or casino queries on your site to piggyback off your domain authority and display spammy contact info in Google's index. Why 500 Errors aren't great: Why serving a 500 server error code to stop bots on search URLs will backfire and reduce Googlebot's crawl rate across your entire website. Category Pages vs. Search Pages: How systems like Blogger use search parameters for tag landing pages and why you should treat valuable category pages differently. Key Takeaways for SEOs & Developers: Fix Crawling at the Source: Do not use the Google Search Console Removal Tool to handle infinite search URLs; it only filters search results temporarily and does not stop Googlebot from hammering your server. Broaden Your Robots Rules: Use one broad wildcard rule in your robots.txt (like /search?) to cover all query parameters, keeping your file maintainable and clean. Build Real Category Pages: Instead of using internal search parameters as makeshift categories, invest in clean, dedicated category pages to help search engines understand your site's hierarchy. Don't Depend on Auto-Systems: While Google's systems try to automatically recognize and deprioritize infinite spaces, it is slow and unreliable—proactive manual configuration is always safer.     Chapters 00:00 - Intro & Greetings 00:45 - Defining Internal Search Results Pages 01:26 - How Googlebot Discovers Search Features & Creates Infinite Spaces 04:15 - Crawl Budget, Server Load, and Database Performance Hurdles 07:32 - Solutions: Robots.txt Disallow vs. Meta Noindex 10:09 - The Fallacy of the Search Console Removal Tool & 404 Pages 12:35 - Why You Should Never Serve 500 Errors to Bots 13:56 -  CMS Nuances: Tag Landing Pages and Blog Categories 15:26 - Crafting Broad Robots.txt Patterns and Historical Guidelines 18:17 - When (and When Not) to Allow Indexed Search Pages 20:41 - The Spam Vector Threat: Hacked Content, Casino, & Pharma Exploits 25:06 - Taking Proactive Security Measures for Clients 27:04 - Lazy Search Redirection Hack, Outro & Subscribing Resources Mentioned: Google Search Console (Removal Tool)    Do you have a legitimate reason for letting search engines index your internal search results page? Let us know in the comments below, or find us on LinkedIn to share your thoughts! Don't forget to like and subscribe to the podcast on your favorite platform to catch every behind-the-scenes episode from the Search Relations team! Episode transcript →  https://goo.gle/sotr113-transcript Listen to more Search Off the Record → https://goo.gle/sotr-yt  Subscribe to Google Search Channel → https://goo.gle/SearchCentral Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch #SearchConsole #SEOTips #CrawlBudget #RobotsTXT #GoogleSearchConsole #WebPerformance #WebSecurity #SearchOffTheRecord  
  • How to read the Indexing Report - Transcript 16.07.2026
  • How to read the Indexing Report 16.07.2026 31мин
    Should you panic when your Search Console indexing report is showing pages that aren't indexed? Is a 404 error code always a sign of a broken website? In this episode of Search Off the Record, Martin Splitt and John Mueller from the Google Search Relations team dive deep into the Page Indexing report in Google Search Console. They unpack why treating this report as a static inventory checklist to "fix" things is the wrong response, how to spot massive SEO-ruining hosting or CDN traps, and why a healthy website doesn't actually need a 100% index rate. In this episode, you'll learn: The Indexing Report Shift: Insights from the Search Console team's Hillel on why you should look for trend lines and systemic patterns rather than treating the report as a giant list of errors. When 404s are Good: Why expected 404 errors are technically correct for deleted content, and how to survive the "boss panic" of numbers that won't go down. The Domain Property Advantage: How setting up a domain property handles canonical shifts, www vs. non-www tracking, and performance data much cleaner. The Site Query vs. Search Console: Why the site: query is an artificial tool that might show old domain moves or hreflang swaps for years, making Search Console your only true source of truth. Hosting & CDN Traps: How aggressive bot protections, hidden interstitials, and "Soft 200" error pages completely destroy your crawl data and lead to malicious canonicalization. Discovered vs. Crawled: What it actually means when pages sit in "Discovered/Crawled - currently not indexed," and how to recognize holistic site quality issues over technical bugs. Key Takeaways for SEOs & Developers: Patterns over Inventories: Use the report to verify that your intentional changes (like site migrations or page removals) are processing correctly over time. Forget the Ratio: There is no magic metric for indexed vs. non-indexed pages. Even Google's own developer documentation has a massive chunk of non-indexed content due to intentional choices. Watch Out for Soft Blocks: Ensure your security layers or CDNs aren't serving "Are you a bot?" challenge screens to Googlebot with a 200 success code. Computers Fail (And That's Fine): Minor server blips, failed DNS requests, or temporary 500 errors happen. Google's systems are resilient and will just try again later.   Chapters 0:00 - Introduction: The Search Central Live coverage report confusion. 1:45 - Shifting perspectives: Treating Search Console as a pattern tracker, not a checklist. 4:36 - Tracking site migrations and processing data delays. 6:25 - Why 404 errors can be a good thing  8:00 - Handling canonical shifts and the value of Domain Properties. 11:12 - Why the site: query isn't telling you what you think it does (Domain moves and hreflang bugs). 13:38 - CDN bot protection and the absolute nightmare of soft error pages. 17:49 - How to use "marked as fixed". 20:32 - Discovered vs. Crawled Not Indexed: Is it a technical or site quality issue? 25:31 - Debunking the indexed-to-non-indexed ratio myth. 27:48 - Final verdict: How to stop fearing your Indexing Report. Resources Mentioned: Google Search Central: https://developers.google.com/search Google Search Console: https://search.google.com/search-console Search Central Live Events: https://developers.google.com/search/events    Are you actively stressing over your non-indexed page counts, or are you tracking the big trend lines? Let us know in the comments! Episode transcript →  https://goo.gle/sotr112-transcript   Listen to more Search Off the Record → https://goo.gle/sotr-yt  Subscribe to Google Search Channel → https://goo.gle/SearchCentral   Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch #SearchConsole Speakers: Martin Splitt, John Mueller
  • Should I use markdown for my site? - Transcript 15.06.2026
  • Should I use markdown for my site? 15.06.2026 26мин
    Should you convert your website into Markdown to help Large Language Models (LLMs) understand your content better? Is "llms.txt" worth the effort for SEO? In this episode of Search Off the Record, Martin Splitt and John Mueller from the Google Search Relations team dive deep into the history of Markdown, its rise in the AI era, and whether it holds any real weight for search engine discovery. In this episode, you'll learn: The Origins of Markdown: From John Gruber and Aaron Swartz to its status as the "language of GitHub." Markdown vs. HTML: Why the "cleanliness" of Markdown is tempting for developers but potentially risky for site structure. LLMs & Markdown: Do AI crawlers actually prefer Markdown, or are they already experts at parsing HTML? The "Parallel Version" Trap: Why creating a separate text/Markdown version of your site for AI can lead to the same maintenance nightmares as dynamic rendering. Use Cases that Make Sense: When Markdown is actually superior (like developer documentation) and when it's totally unnecessary (like your shoe catalog). Key Takeaways for SEOs & Developers: Crawlers are built for the "messy" web: Google and other engines have decades of experience parsing HTML. Don't sacrifice discovery: Headers, footers, and sidebars in HTML provide critical context for site structure that a raw Markdown file might lack. Maintenance is king: Avoid the complexity of maintaining two versions of the same content. Chapters 0:00 - Introduction: Should we all be using Markdown? 3:45 - The history and purpose of Markdown. 7:15 - Why developers love it: Separation of style and content. 11:20 - Do crawlers need Markdown to understand your site? 14:50 - The danger of "parallel versions" and dynamic rendering lessons. 17:30 - Discussing the "llms.txt" proposal and AI agents. 21:00 - Where Markdown actually makes sense (Developer Docs). 24:00 - Final verdict: Stick to HTML for the web. Resources Mentioned: Google Search Central: https://developers.google.com/search Are you using Markdown for your site's frontend or just as a backend source? Let us know in the comments! Episode transcript →  https://goo.gle/sotr111-transcript Listen to more Search Off the Record → https://goo.gle/sotr-yt Subscribe to Google Search Channel → https://goo.gle/SearchCentral  Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch Speakers: Martin Splitt, John Mueller
  • Vibe Coding - yay or nay? - Transcript 07.05.2026
  • Vibe Coding - yay or nay? 07.05.2026 33мин
    In this episode of Search Off the Record, Martin Splitt and John Mueller from Google's Search Relations team dive deep into the world of AI-assisted development. They explore the reality of "Vibe Coding", the process of building apps and websites using natural language instead of manual syntax. Whether you're a developer looking to offload tedious setup tasks or an SEO expert trying to understand how AI-generated sites impact search, this conversation is for you. In this episode, you'll learn: * What is Vibe Coding? Understanding the shift from writing syntax to "talking" to your IDE. * The Developer's Trap: Why you still need technical knowledge (like linters, deployment scripts, and GitHub Actions) to prevent AI from breaking your project. * SEO & AI Architecture: Why you can't just "add SEO" at the end—and how to guide AI to build with canonicals and sitemaps from day one. * Tooling Breakdown: Martin and John share their experiences with AI Studio, Gemini CLI, Firebase, and GitHub. * Testing with AI Agents: How to use AI to remote control browsers (like Chromium) for automated testing. Chapters 00:00 – Intro: What exactly is "Vibe Coding"? 01:32 – Martin's experiment with AI Studio and client-side JS. 03:30 – The "English as a Programming Language" allure. 06:00 – Why the AI makes assumptions (and why that's dangerous). 08:51 – "Sprinkling SEO" vs. Building for SEO from the start. 12:40 – Can AI test itself? Using browser agents for QA. 20:27 – The technical debt of AI: Refactoring and maintainability. 25:42 – Moving to the terminal: Gemini CLI & Cloud Code. 31:34 – Using AI to skip the setup work. Resources Mentioned: * Google AI Studio * Firebase Hosting * Gemini CLI / Cloud Code * GitHub Actions for CI/CD What's your experience with Vibe Coding? Let us know in the comments! Episode transcript → https://goo.gle/sotr110-transcript  Listen to more Search Off the Record → https://goo.gle/sotr-yt Subscribe to Google Search Channel → https://goo.gle/SearchCentral  Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch Speakers: Martin Splitt, John Mueller
  • How AI Is Changing Google Search and SEO - transcript 01.05.2026
  • How AI Is Changing Google Search and SEO 01.05.2026 33мин
    In this episode of Search Off the Record, Martin speaks with Nikola Todorovic (director of Software Engineering at Google Search) about how AI is changing Google Search. They discuss the evolution from traditional search to AI Overviews and AI Mode, how Google tests and launches search changes, and why query behaviour is becoming more conversational and complex. Nikola also explains the role of machine learning in Search, how features are evaluated before launch, and what site owners and SEOs should focus on as AI becomes a bigger part of the search experience. If you work in SEO or web development, this episode offers a clear look at how Google approaches AI in Search and what it means for the future of search visibility. Episode transcript → https://goo.gle/sotr109-transcript Listen to more Search Off the Record → https://goo.gle/sotr-yt   Subscribe to Google Search Channel → https://goo.gle/SearchCentral  Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch Speakers: Martin Splitt, Nikola Todorovic
  • Analysing Robots.txt at scale with HTTP Archive and BigQuery - transcript 23.04.2026
  • Analysing Robots.txt at scale with HTTP Archive and BigQuery 23.04.2026 27мин
    In this episode of Search Off the Record, Martin and Gary turn a simple robots.txt question into a data‑driven deep dive using HTTP Archive, WebPageTest, custom JavaScript metrics, and BigQuery. They explore how millions of real robots.txt files are actually written in 2025–2026, which directives and user‑agents are most common, and what that means for modern crawling and AI bots. Perfect for beginner to mid‑level developers and SEOs, you'll learn how large‑scale web measurement works (HTTP Archive, Chrome UX Report, Web Almanac), and how to turn raw crawl data into actionable SEO insights. Subscribe for more candid conversations about crawling, indexing, and the data behind how Google Search and the web really work. Resources: Web Almanac →  https://almanac.httparchive.org/en/2025/ Robotstxt custom metric for the HTTP Archive →  https://github.com/HTTPArchive/custom-metrics/pull/191 robots.txt parser change → https://github.com/google/robotstxt/commit/4af32e54b715442bb04cd0470e99192f0ffb9792#commitcomment-178586774 Episode transcript → https://goo.gle/sotr108-transcript Listen to more Search Off the Record → https://goo.gle/sotr-yt   Subscribe to Google Search Channel → https://goo.gle/SearchCentral Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch Speakers: Martin Splitt, Gary Illyes
  • Are websites getting "fat"? Page weight, HTML size & Googlebot limits explained - transcript 30.03.2026
  • Are websites getting "fat"? Page weight, HTML size & Googlebot limits explained 30.03.2026 32мин
    In this episode of Search Off the Record, Gary and Martin dig into what "page size" and "page weight" actually mean for developers, users, and search engines. They discuss exploding web page sizes: median mobile homepages hit 2.3 MB in 2025 Web Almanac (up 3x from 2015), key insights for developers on page weight definitions, Googlebot's crawl limits, HTML bloat from structured data/images, and why size still hurts UX on slow connections despite faster networks. If you build or maintain websites, this conversation will help you rethink how much data your pages ship, where bloat really comes from, and why page weight still matters even as connections get faster. Resources: ​Web Almanac → https://almanac.httparchive.org/en/2025/ HTML living standard → https://html.spec.whatwg.org/multipage/ How page speed helps with conversions →  https://www.thinkwithgoogle.com/marketing-strategies/app-and-mobile/mobile-page-speed-data/  Episode transcript → https://goo.gle/sotr106-transcript Listen to more Search Off the Record → https://goo.gle/sotr-yt  Subscribe to Google Search Channel → https://goo.gle/SearchCentral Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch Speakers: Martin Splitt, Gary Illyes    
  • Google crawlers behind the scenes - transcript 12.03.2026
  • Google crawlers behind the scenes 12.03.2026 25мин
    Developers often talk about Googlebot as if it were a single program you could just run as "googlebot.exe", but that is not how Google's crawling actually works. In this episode of Search Off the Record, Martin and Gary from the Search Relations team unpack how Google's crawling infrastructure is really built and operated.​ They cover why "Googlebot" is a misnomer and how it relates to a central crawling software-as-a-service used by many Google products​, how crawl behavior is controlled centrally to avoid overwhelming sites (throttling, handling 503s, and "don't break the internet" safeguards)​ and more! If you build for the web, work on SEO, or just want a more accurate mental model of how Google crawls pages, this behind‑the‑scenes discussion is for you. Resources: ​Crawlers → https://developer.google.com/crawling  Episode transcript → https://goo.gle/sotr107-transcript  Listen to more Search Off the Record → https://goo.gle/sotr-yt   Subscribe to Google Search Channel → https://goo.gle/SearchCentral  Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch Speakers: Martin Splitt, Gary Illyes
  • How Browsers Really Parse HTML (and What That Means for SEO) - transcript 26.02.2026
  • How Browsers Really Parse HTML (and What That Means for SEO) 26.02.2026 32мин
    Martin and Gary unpack how HTML parsing really works, why the HTML standard is so lenient, and how messy markup can silently break key SEO signals like hreflang and rel=canonical. They revisit validators and cross‑browser hacks from the Netscape/IE days, and discuss whether semantic HTML and strict validity truly matter for search. You'll also hear when link hints like preload, prefetch, and DNS prefetch help performance (and indirectly SEO), and where meta and link tags really belong. ​ Resources: HTML Living Standard → https://html.spec.whatwg.org/ Episode transcript → https://goo.gle/sotr105-transcript Listen to more Search Off the Record → https://goo.gle/sotr-yt  Subscribe to Google Search Channel → https://goo.gle/SearchCentral Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch Speakers: Martin Splitt, Gary Illyes
  • Do You Still Need a Website in 2026? (Transcript) 12.02.2026
  • Do You Still Need a Website in 2026? 12.02.2026 28мин
    In this episode of Search Off the Record, Martin and Gary from the Google Search Relations team tackle a deceptively simple question: do you still need a website in 2026? Starting from the recurring industry claim that "the web is dead," they explore how the web has evolved through the rise of apps, AI chatbots, and social platforms, and why the answer almost always ends up being "it depends." Tune in for an engaging discussion on how websites remain relevant and what it means for content creation and discovery. Episode transcript → https://goo.gle/sotr103-transcript Listen to more Search Off the Record → https://goo.gle/sotr-yt  Subscribe to Google Search Channel → https://goo.gle/SearchCentral   Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.  #SOTRpodcast #SEO #GoogleSearch Speakers: Martin Splitt, Gary Illyes

Популярен в

Този подкаст се появява и в подкаст класациите на тези държави.