SEO

Publishers Are Threatening to Block Google

Blog Image

Rasit Cakir

Jul 27, 20265 min read

Publishers Are Threatening to Block Google

Twenty years of digital publishing ran on one assumption. You want Google to crawl you, because Google sends the traffic that pays for everything. That assumption is now being questioned out loud by some of the biggest names in media, and the reason is arithmetic rather than principle.

Who’s threatening what

The Wall Street Journal reported this month that Reddit has discussed cutting off Google’s access to its content for AI use, even though Google pays it a reported 60 million dollars a year for exactly that. Reddit isn’t alone. USA Today, Politico, Reuters, The Economist, and People Inc. have all been reported as weighing whether to block Google’s crawlers.

The numbers behind the anger come from Semrush, which measured organic Google search traffic from United States users between June of last year and June of this one. Business Insider was down more than 85%. USA Today lost close to half. CNN fell around 25%, Politico about 23%. Mike Reed, chief executive of the company that owns USA Today, said it’s time to take a stand and that enough is enough, and confirmed the company is prepared to block Google’s crawlers outright if the slide continues.

One thing to keep straight, though. Nobody has actually blocked anything yet. Read the coverage closely and this reads as a negotiating position, publishers using the threat to win better terms and paid access, rather than a coordinated walkout. The threat has value precisely because carrying it out would hurt.

The bundle is the whole problem

Why is turning off a crawler such a drastic step? Because Google’s crawler does two jobs with one visit. The same crawl that indexes you for regular search results also feeds your content into AI answers.

Google offers a control called Google-Extended that lets you opt out of having your content train its AI models. What it doesn’t cleanly do is keep you out of AI Overviews while leaving your blue-link rankings intact. To stay out of the AI summaries, you’re looking at tags like nosnippet, which hurt your normal search visibility too. Block the crawler outright and you vanish from both. Access is bundled, and the bundle is the leverage.

That arrangement may not last forever. A United Kingdom court ruled in June that Google must let publishers opt out of AI features without damaging their traditional search performance. Google has around nine months to comply, so the separation exists on paper and not yet in practice. Publishers are also skeptical of the existing controls, with executives telling trade press they expect using Google-Extended to cost them search visibility anyway.

The blocking nobody meant to do

While the media giants argue about whether to shut the door on purpose, a lot of ordinary sites have already shut it by accident.

A study published this month tested the top 1,000 websites and found that 40.9% of them are unreadable to GPTBot, the crawler behind ChatGPT. Nearly one in five, 18.4%, is completely dark to every AI crawler tested. The finding that should worry any marketer is the third one. About 17.6% of these sites allow GPTBot in their robots.txt file and then return a 403 error when it actually requests a page. Their stated policy says come in. Their firewall or content delivery network says no.

Those brands aren’t making a strategic choice about AI visibility. They don’t know they’re invisible. Somebody enabled bot protection, or a security default flipped on, and the AI systems stopped being able to read the site. No alert fires. Rankings look normal. The brand simply stops appearing in AI answers with no explanation.

Check before you strategize

There’s a useful order of operations in all this. Debating whether to appear in AI answers is a real strategic question, and it’s also completely academic if a crawler can’t reach your pages in the first place.

So verify before you theorize. Look at your robots.txt file and confirm which AI crawlers you allow. Then check your server logs or your content delivery network settings to see whether those same crawlers are getting 403 errors anyway, because the gap between stated policy and actual behavior is where most of the damage lives. Bot protection rules, firewall settings, and rate limiting all block AI crawlers by default in plenty of setups, and none of them announce it.

The publishers threatening to block Google have leverage most brands don’t. Reddit can negotiate because Google needs its forum data. USA Today can threaten because its brand carries weight. Everyone else has the opposite problem, which is getting seen at all in a system that cites only a handful of sources per answer. Being one of those sources comes from earned authority and the coverage that digital PR is built to produce. Make sure the crawlers can actually read the pages you worked to make citation-worthy, because right now a lot of them can’t, and almost nobody has checked.