Back in July, when Bay City News Foundation first began to track AI bot traffic on its news website, LocalNewsMatters.org, the referral rate of humans clicking on summaries from AI chatbots was as low as feared — just 0.7%. 

By comparison, the click-through rate on Google search appearances was 1.2%.

So if all the people who searched for news using Google switched to using chatbots — the so-called Google Zero scenario — clicks from user-led queries would drop by 40%. That’s a huge decline that would cut roughly 200,000 website visitors a year, presumably because the Artificial Intelligence summary was good enough to satisfy the user. 

The urge to do something is strong — block AI bots for instance, to dispute this unlicensed scraping and repurposing of content — but that could also cut off this new, emerging way to be visible to users.

The conundrum is real. If a single website blocks all AI bots, this emerging trickle of traffic will never grow into a river. Or worse, traffic could just flow elsewhere as other websites that do not block AI bots benefit from the attention.

It’s a prisoner’s dilemma and it’s playing out in real time. 

In the prisoner’s dilemma, the best strategy is for all accomplices to continue to keep their mouths shut, or in the case of media companies, for all of them to shut off the flow of news to these AI scrapers. 

If all publishers acted in coordination and shut off their sites to AI bots, chatbot summaries of the real world would become untethered to actual facts. Remember back in November 2022 when ChatGPT didn’t have any knowledge past September 2021? 

ChatGPT demonstrated the power of time travel, claiming that by September 2021, it was aware of an event that occurred in October 2021.

That was a news-to-chatbot breakdown, and led to the poor experience above. Hallucination can result, or even regurgitating the user’s own assumptions of what went on back to them. It’s the definition of fake news. 

Training Large Language Models only happens once every few months, so preventing scrapes could set real-world knowledge summaries back far enough to degrade the experience for users and force Frontier AI Labs to the negotiating table to license the very content they had been scraping for free.

Unfortunately the landscape is not so simple. 


There is a patchwork of news sites that block, and others that do not

When I sent Claude Code to examine 43 Bay Area news outlets, it found that 23 of them blocked scrapers: 14 were completely open, while four blocked some bots and not others. 

A summary of 43 San Francisco Bay Area news websites’ Robots.txt files by Claude https://claude.ai/artifact/F6snQYQ1NneCpjfhLHkf2u

Thankfully, the biggest AI players in the space, like Google, Anthropic, OpenAI, and Perplexity, seem to be respecting the most basic of controls — the Robots.txt file. 

This file declares to all who care to read it whether or not they’re allowed to scrape the site. News sites own the copyright to their published works, and those who violate this directive risk gigantic legal damages. It certainly helps that companies like The New York Times and MediaNews Group are suing OpenAI, maker of ChatGPT, to keep the AI industry on its toes. A Microsoft researcher even acknowledged the wrongness of OpenAI’s unlicensed scraping, calling it perhaps the “largest theft of labor in human history,” according to a recent legal brief in the case.

As one nice finding of compliance, Claude Code declined my directive to pretend to be a human being and evade detection. Way to be ethical — and legally safe — Claude!

My interaction with Claude Code, in which it declines my instructions to spoof its identity.

According to Known Agents, a freemium AI bot tracking service, compliance with Robots.txt happens over 90% of the time. There are some bad actors out there that are spoofing human visitors, and others that just don’t seem to be respecting the Robots.txt directive. But as a blocking mechanism, Robots.txt seems good enough for now. 

Has blocking degraded the chatbot experience? 

From what I can tell, blocking DOES make the chatbot experience worse. 

I live in San Jose, California, and so I’d expect the major newspaper in the area, The Mercury News, to be one of the biggest sources of local news. But its news didn’t appear in my Gemini query. I suspect that’s because the Merc’s Robots.txt is blocking most of the big AI bots, including one of Google’s, and its owner is MediaNews Group, one of the litigants against OpenAI. 

Gemini did not readily pull up sources of news from sources that blocked AI bots.

Gemini missed news of an 8-year-old girl’s death being ruled a homicide, or that California has run out of license plate numbers — both stories that were reported by the Merc prominently on the day I made the query. In my view, Gemini’s answer was limited because it was blocked.

Local news outlets could ironically be in a stronger position than national ones to demand payment from unlicensed scrapers, because there are fewer to coordinate to do the right thing and put up the barricade. 

What’s the value of AI to chatbot companies anyway?

Google is already making tens of billions of dollars annually, thanks to AI. 

AI Overview experiences drove 10% more queries globally and that benefit is increasing over time, CEO Sundar Pichai told investors on its Q2 2025 earnings call. Those queries monetize at about the same rate as other queries, said its Chief Revenue Officer, Phillipp Schindler, on the same call. Backing out the trailing 12 months of revenue from “Search and Other” of $243.3 billion from Q2 2026, a 10% lift is enormous. 


Proving the value that journalism brings

What fraction of that enormous benefit can the news industry claim came from its journalism? What experiments can we run to show the true value of real-time, human-verified facts and context? 

Those are the kinds of questions we’ll be seeking to answer at Bay City News Labs, a division of the foundation, over the next several months, thanks to support from the Maynard Institute’s Fire Up Entrepreneurship program. Stay tuned to this space for more. 

Interested in supporting our work? Contact us at labs@baycitynews.com


Ryan Nakashima is secretary of the board at Bay City News Foundation, technology advisor and participant in the Fire Up program, and held product roles at Hearst Newspapers and MediaNews Group. Reach him on LinkedIn.

Secretary of the Board, Bay City News Foundation