Should You Block Google's AI Crawler? The Math Is Different for a 20-Page Business Site.
Publishers are debating blocking Google's AI crawler. For a 20-page service business the math is completely different. Here is what each control does.
Kemal EsensoyĀ·Modified on September 25, 2026
"Saw that the BBC and the New York Times are blocking AI. We should probably do the same, right?"
That email came from a client who runs a five-truck HVAC company outside Phoenix. Twenty-two pages on the site. Roughly eleven inbound calls a month from organic search, which is most of his new business.
He wanted to copy the strategy of an organisation whose entire revenue model is advertising against pageviews. He is not that. Almost nobody reading this is that.
But the question deserves a real answer, not a brush-off, because the headlines are real and the controls are real and most of the advice floating around about them is confidently wrong. So here is what actually happens when you block Google's AI crawler, which of the four or five available controls does what, and how to work out which side of the line your business sits on.
Want the control reference in one place? Download the complete guide as a PDF ā 9 pages, every directive and what it actually does, no login required.
Google-Extended Does Not Control AI Overviews
Start here, because when someone says they want to block Google's AI crawler, this is usually the thing they mean, and it is the single most misreported detail in the topic.
Google-Extended is a robots.txt token. Google's crawler documentation says it lets publishers "manage whether content Google crawls from their sites may be used for training future generations of Gemini models" and for grounding in Gemini Apps and the Vertex AI API. The same page states plainly: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."
That last sentence is why people reach for it. It sounds like a free lunch. Block the AI, keep the rankings.
The part that gets left out: AI Overviews and AI Mode are surfaces inside Google Search, not separate products. They are built on the regular Search index, crawled by regular Googlebot. Google clarified this back in October 2023, shortly after introducing the token, when it confirmed Google-Extended did not apply to what was then called SGE. That has not changed. Add User-agent: Google-Extended and Disallow: / to your robots.txt and you have opted out of Gemini model training. You are still fully eligible to be summarised in an AI Overview tomorrow morning.
One more user agent, since it confuses people: Google-CloudVertexBot. Google's docs describe it as crawling "requested by the site owners' for building Vertex AI Agents" and say "it has no effect on Google Search or other products." Nobody crawls you with it unless a site owner asked. Not part of this debate.
The Real Opt-Out Now Lives in Search Console
For most of the last three years the honest answer to "can I leave AI Overviews without leaving Search?" was no. That changed in June 2026, and it changed because a regulator forced it.
On 30 September 2025 the UK's Competition and Markets Authority designated Google with Strategic Market Status in general search, having found Google handles more than 90 percent of UK general search queries. That designation unlocked a power under the Digital Markets, Competition and Consumers Act 2024: the CMA can impose binding conduct requirements without first proving wrongdoing in court. On 3 June 2026 it used it, ordering Google to give publishers a genuine opt-out from AI Overviews, AI Mode and related AI features that does not damage their ordinary search performance.
Google shipped it. There is now a property-level setting in Search Console, under Settings, covering AI Overviews, AI Mode and the generative AI features in Discover. Google began honouring the choice on 17 June 2026, with content dropping out within a day or two. The default for every property is "include," so nothing happens unless you deliberately switch it.
Three things it does not do, which matter:
- It does not cover AI training. That is still Google-Extended, a completely separate control.
- It does not cover the Gemini app.
- It is property-level, not page-level, at least for now. Google has nine months from the June order to comply fully, so finer-grained control could arrive as late as March 2027.
Availability is messy. It launched UK-first, started appearing outside the UK in July 2026, and the rollout is still partial. Some US properties have it, some do not, and as of early September it is not available in Germany, Austria or Switzerland. If you cannot find the setting, that is the rollout, not you.
And here is the caveat nobody wants to print: Google says the setting does not affect regular results, but AI Overviews and organic results are not cleanly separated in the live SERP. Top Stories, for instance, now appears inside the AI Overview on a meaningful share of trending news queries. What opting out costs you in practice is genuinely unmeasured. There is no public traffic dataset yet. Anyone telling you the number is making it up.
The Snippet Controls Are a Sledgehammer
Before the Search Console setting existed, the only lever pointed at AI features was the preview controls: nosnippet, max-snippet, data-nosnippet and noindex. Google's own AI features documentation still points site owners at exactly these to "limit the information shown from your pages in Search."
They work. They also cut both ways, which is the entire problem.
nosnippet as a meta robots tag removes your content from AI Overviews. It also removes your regular organic snippet. Your listing becomes a blue link with no description underneath, on every query, forever. For a business whose search listing is a sales pitch, that is a self-inflicted conversion problem far bigger than anything an AI Overview was doing to you. max-snippet:0 behaves the same way. max-snippet:[number] is a softer version of the same trade: you cap the text length, and the cap applies to your normal snippet too.
noindex needs no explanation. You are gone.
data-nosnippet is the only one of the four with any surgical value, and I will come back to it, because it is the foundation of the sane middle ground.
The short version: if the Search Console setting is available on your property, use that instead. Rolling out sitewide nosnippet in 2026 costs you more in ordinary search visibility than it saves in AI exposure.
Publishers and Plumbers Are Solving Different Equations
The blocking wave is real and it is large. A January 2026 Buzzstream analysis of roughly 100 of the biggest UK and US news sites found 79 percent block at least one AI training crawler, 71 percent block retrieval bots, and 34 percent of the top 50 block every bot they tested. The BBC, the New York Times, the Wall Street Journal, the Telegraph, the Daily Mail, NBC News and the New York Post all block everything. Fox News, Politico, the Independent and Substack allow everything.
Note what those publishers are blocking: training and retrieval crawlers from OpenAI, Anthropic, Common Crawl and others. They remain wide open to Googlebot. Blocking Google itself is newer and, as of July 2026 reporting, still talk. Reddit has discussed it internally. USA Today, Politico, Reuters, the Economist and People Inc are reported to be reassessing the relationship. No major publisher had actually pulled the plug on Google's crawlers.
Now the arithmetic. A publisher sells attention. A visitor who reads a summary instead of the article is revenue that did not happen, and there are millions of those a month. Withholding content is genuine leverage, because their content is what the summary is made of, and at their scale absence is noticeable.
Your HVAC company sells jobs. You need eleven phone calls, not eleven thousand visits. A prospect who reads an AI Overview saying "Company X in Chandler handles heat pump replacement and has 4.8 stars from 200 reviews" and then calls you is not a lost pageview. That is the entire outcome you were paying for. The pageview was only ever a means to it.
BrightLocal's 2026 Local Consumer Review Survey, a representative panel of 1,002 US adults, found 45 percent of consumers now use AI tools to find local businesses, up from 6 percent the year before, making AI the third-largest discovery channel behind Google and Facebook. Over the same period Google's own share of local discovery fell from 83 percent to 71 percent. Removing yourself from the channel that grew seven-fold in a year is a strange move for a business that needs to be found.
Whitespark's local study is the best illustration I have seen. They manually checked 540 queries across three US cities and six service industries. AI Overviews showed on 68 percent of local business queries on average, but the split by intent matters: 15 percent for pure local-intent queries like "plumbers near me," 92 percent for informational ones like "how much do plumbers charge," and 97 percent for hybrid queries. In their Houston plumber sample, 60 percent of AI citations went to third-party publishers like Yelp, Thumbtack, Reddit and HomeGuide, and only 40 percent to actual local businesses.
Read that last number again. On the queries where AI Overviews dominate, aggregators are already taking most of the citations. If you opt out, you do not remove the AI Overview. You just hand your 40 percent to Thumbtack. I wrote more about winning those informational queries in Local SEO for Contractors.
A Decision Rule You Can Actually Apply
Before you block Google's AI crawler anywhere, ask one question: does your site make money from the visit, or from what happens after it?
If revenue is a function of pageviews, ad impressions, subscription conversions on article pages, or affiliate clicks, then a summary that satisfies the reader is a direct loss and blocking is a coherent, if risky, negotiating position. That is publishers, large content sites, and comparison and review affiliates.
If revenue is a function of contacts, calls, bookings, quotes or purchases, then being described accurately inside an AI answer is the goal, not the injury. That is service businesses, local trades, B2B firms, agencies, SaaS, ecommerce and clinics. Blocking is self-harm.
The edge cases are real. A blog funded by display ads that also sells a service is genuinely split. If you are in the middle, the property-level toggle is a bad fit, which is why the next section exists.
The Middle Ground Is Selective, Not Total
This is where data-nosnippet earns its keep. It is an HTML attribute you put on a specific element, not a sitewide directive, and it tells Google not to use that particular block of text in any snippet or AI feature while the rest of the page stays fully available.
Wrap it around the things you genuinely do not want restated by a machine: an exact price table you renegotiate quarterly, a detailed process breakdown that is your actual differentiator, availability or lead-time text that goes stale. Leave everything else open. Your service descriptions, your service area, your credentials, your reviews and your contact details should be as machine-readable as you can make them, because those are what get you named in the answer.
Two rules I hold to on client sites. The money pages stay completely unrestricted: service pages, location pages, contact, and the pricing page if it is a real sales asset. And never use a sitewide directive when a block-level one exists. Blanket rules are how people delete their own snippets and then spend six weeks wondering why click-through fell off a cliff.
If you want the offensive version of this rather than the defensive one, How to Get Cited by ChatGPT covers what a 1.4 million prompt study found about which pages actually get pulled into answers, and controlling what AI says about your brand covers the reputation side, which is usually the thing clients are actually worried about underneath the blocking question.
Blocking GPTBot and ClaudeBot Is a Separate Argument
Worth saying clearly because the two conversations get mashed together: GPTBot, ClaudeBot, PerplexityBot, CCBot, Applebot-Extended and the rest have nothing to do with your Google rankings. They are other companies' crawlers. Disallowing every one of them tomorrow changes nothing about how Google sees you.
What it changes is whether ChatGPT, Claude and Perplexity can see and cite you. Those products distinguish training crawlers from retrieval crawlers, and the distinction is the whole decision. Blocking training bots costs you very little visibility. Blocking retrieval bots like OAI-SearchBot, ChatGPT-User and Perplexity-User makes you uncitable in the products where a growing share of buying research happens. Plenty of sites have done the second by accident while intending the first, which I went through in You Blocked the AI Bots. Now ChatGPT Can't Cite You. If you want a tidier way to state your preferences, llms.txt is worth a look, with the caveat that adoption is still thin.
Blocking on Bandwidth Grounds Is Fair, and It Is Not This
One legitimate reason to block a crawler has nothing to do with content policy: it is hammering your server. Cloudflare's own data suggests more than 50 percent of AI crawl traffic is spent re-fetching pages that have not changed. I have watched small sites on modest hosting get knocked over by crawlers that pull the same pages dozens of times an hour.
Rate-limit them. Block the badly behaved ones. That is an infrastructure decision and it needs no philosophical justification, only a log file. AI Bots Are Crawling Your Website to Death has the practical version.
The ground is shifting here too. From 15 September 2026 Cloudflare blocks mixed-use crawlers by default on ad-bearing pages for new customers, new sites on existing accounts and all free-tier accounts, splitting bots into Search, Agent and Training categories. If you are on a free Cloudflare plan and you have not looked at your bot settings recently, look now. You may find a policy applied to you that you never chose.
What I Told the HVAC Client
I told him not to block Google's AI crawler, and I told him why in about four sentences: the organisations he read about sell advertising, he sells heat pump installations, and an AI answer that names his company and lists his service area does his job for him rather than instead of him.
Then I did the thing that mattered, which was making his service and location pages easier for a machine to summarise correctly. The risk for a business his size is not being used by AI search. It is being described badly, or not found at all while an aggregator takes the slot.
I will be honest about the limits here. The Search Console opt-out is roughly three months old, the rollout is uneven, nobody has published credible before-and-after traffic data on using it, and the regulatory picture is still moving with Google's full compliance deadline running into 2027. If you are a publisher weighing this seriously, you are making a judgement call with incomplete information and you should say so out loud rather than pretend otherwise.
But for the twenty-page service business, the uncertainty does not really bite. The downside of staying visible is that some people read a summary instead of your homepage. The downside of blocking is that you disappear from the fastest-growing way your customers find businesses like yours, in exchange for leverage you do not have.
If you want a second pair of eyes on your robots.txt and your snippet directives before you change anything, that is the kind of thing I do at Wunderlandmedia. Bring your log files.
Find these posts useful? Mark Wunderlandmedia as a preferred source on Google ā my articles will then show up more often in your Search results, AI Overviews and AI Mode.
Set as preferred sourceAbout the Author
Kemal Esensoy
Kemal Esensoy, founder of Wunderlandmedia, started his journey as a freelance web developer and designer. He conducted web design courses with over 3,000 students. Today, he leads an award-winning full-stack agency specializing in web development, SEO, and digital marketing.