
Hacker News
July 20, 20264 min read
There’s a new boogeyman in the battles over AI: so-called “stealth crawlers.” We’ll admit it—the term “stealth crawlers” sounds quite nefarious. In reality, they’re anything but.
“Stealth crawlers” are simply automated tools to access and collect public web data—without disclosing the user’s identity. Private crawlers like these facilitate all kinds of important work that benefits the public, including investigative reporting, academic research, cybersecurity protection, and more.
Many publishers want to unmask crawlers anyways—and are pushing for new legislation that would give them new powers to do so. These legislative proposals threaten the open web, user privacy, and valuable research without directly addressing the problems they’re supposedly intending to solve.
Alarmingly, these harmful proposals are gaining traction. The New York state legislature has already passed such a bill, the NY Stealth Crawler Protection Act , which is now on Governor Hochul’s desk. We expect to see similar bills introduced in other states, and potentially in Congress. That’s a big problem for the open web—and the many benefits it provides.
Anonymous crawling enables some of the most publicly beneficial uses of the open web. Researchers, journalists, and other watchdog groups use unidentified automated tools to gather the information necessary to hold powerful institutions accountable and protect the public.
Anonymous crawling fuels important investigative journalism. For example, The Markup , a non-profit news site, used anonymous crawlers to investigate potentially anti-competitive practices by tech companies, such as Amazon’s tendency to prioritize Amazon brands and Amazon-exclusive products over competitors with higher ratings. The crawlers identified themselves as ordinary Firefox browsers to web servers, which allowed The Markup to understand how Amazon search results pages would appear to ordinary users. Similarly, ProPublica used an automated tool designed to simulate an ordinary Amazon customer to reveal that the site steered shoppers to more expensive products over cheaper alternatives.
Anonymous web scraping is also crucial for cybersecurity professionals , who use automated tools to monitor the web for information that helps them protect against malicious attackers. Privacy tools , including EFF’s own Privacy Badger , also crawl sites anonymously to identify trackers without compromising user privacy.
However, without the ability to scrape anonymously, these tools would likely be blocked. Sites can—and do—block crawlers operated by researchers, journalists, and activists who criticize them. For example, Facebook shut down accounts belonging to researchers who used automated tools to study misinformation on the platform and demanded that they take down published research. Many sites block automated access by anyone who hasn’t paid to crawl public webpages.
News publishers—and their allies in government— say that unmasking crawlers is necessary to protect news organizations from technological strain caused by AI-related crawling, and fears that AI could reduce news sites’ traffic and ad revenue. These are legitimate concerns.
But enacting broad, reactionary restrictions on automated access is not the answer. Legislation targeting anonymous crawling threatens the open web, user privacy, and valuable research without actually addressing these technological and potential economic harms of scraping.
The New York state legislature recently passed the NY Stealth Crawler Protection Act , a law that would make it illegal to crawl news websites without revealing who is operating the crawler and all possible future uses of the data collected by the crawler. The law would give websites the power to obtain court orders that unmask anyone using an unidentified crawler—without any evidence that they broke the law.
Laws like the New York bill sweep far beyond AI, and do not meaningfully address the technological or potential harms of AI-related web scraping. These policies would chill beneficial crawling by allowing publishers to veto lawful public access, giving them the power to block not just bad actors, but also security professionals, researchers, dissidents, or anyone who has not paid for a license to view public text. This needlessly undermines the free and open internet.
Digital news publishers—like most websites—face real technological challenges in the AI era. While web crawling has been around for decades, with the proliferation of AI, crawlers now collect far more public web data than they used to. This pushes servers closer to their maximum capacity, and if some bots collect information too aggressively, they may strain web servers to the point that it degrades site performance. The problem is not anonymity—so unmasking crawlers won’t solve it. The real problem is overaggressive crawling, which can be effectively addressed with technical measures that target harmful conduct without impeding anonymous access to information.
There are other, far less harmful ways to protect publishers from the harms these “stealth crawler” laws claim to target. Addressing the harms of AI-related crawling requires policies that narrowly target the causes of these issues–without undermining free expression and the open web. Policies that target crawlers and scrapers are anything but.
Read what's here, then head to the original whenever you're ready - never required.
Continue Reading on Hacker NewsMeasuring What Matters with Jules
Google Developers July 21, 2026