EFF Warns New Laws on 'Stealth Crawlers' Could Harm Open Web

3 min readSources: EFF

EFF says bills targeting 'stealth crawlers' mischaracterize them as threats.

Why it matters: As laws like New York's Stealth Crawler Prohibition Act limit unmarked web crawling, legal and tech professionals must understand implications for AI data access, privacy, and IP rights.

  • EFF states stealth crawlers—bots collecting public web data anonymously—aren't inherently harmful.
  • New York passed the Stealth Crawler Prohibition Act in June 2026 targeting deceptive bots on news sites.
  • Cloudflare starts blocking 'mixed-use' crawlers from ad-supported sites by default from Sept. 15, 2026.
  • Amnesty International raises privacy invasion concerns from unlawful web scraping powering generative AI.
  • 2.5 million sites have disallowed AI training via robots.txt as of May 2026.

The Electronic Frontier Foundation (EFF) criticized recent legislative efforts that label 'stealth crawlers' as threats, cautioning that bills like New York’s Stealth Crawler Prohibition Act could jeopardize the openness of the web. According to the EFF, stealth crawlers are automated tools that access public data without revealing user identities, and they are not inherently damaging to web ecosystems. (EFF analysis).

In June 2026, New York’s State Senate and Assembly passed the Stealth Crawler Prohibition Act, which forbids deceptive bots from accessing news websites without proper identification. This legislative move aims to shield publishers from excessive bot traffic. Danielle Coffey, CEO of the News/Media Alliance, applauded this as support for transparency and the information ecosystem. (News Media Alliance).

Further tightening web access, Cloudflare announced that starting September 15, 2026, it will block 'mixed-use' crawlers from accessing ad-supported pages by default, unless site owners choose to allow them. This policy pressures AI companies to negotiate access or pay publishers for content. (TechCrunch report).

Amnesty International spotlighted privacy risks tied to web scraping practices fueling major generative AI systems. These pipelines, they argue, are rooted in widespread invasions of privacy performed without consent, raising serious ethical and legal questions. Likhita Banerji, head of Amnesty’s Algorithmic Accountability Lab, stressed the scale and unlawfulness of such data harvesting. (Amnesty briefing).

Meanwhile, data reveals a growing number of websites—2.5 million as of May 2026—have explicitly disallowed AI training through robots.txt files, highlighting publisher resistance to non-consensual data use. This regulatory and technical pushback reflects a clash between AI development ambitions and the protection of online property and privacy rights.

By the numbers:

  • 2.5 million sites blocking AI training via robots.txt — as of May 31, 2026
  • September 15, 2026 — Cloudflare begins blocking 'mixed-use' crawlers by default on ad-supported pages
  • June 2026 — New York passes Stealth Crawler Prohibition Act

Yes, but: While the legislation targets deceptive bots, EFF warns it could inadvertently restrict benign crawlers and limit access to public data, impacting AI innovation and open web principles.

What's next: The effectiveness of New York’s law and Cloudflare’s policy in curbing harmful data scraping remains to be seen, with close attention expected from legislators, tech companies, and advocacy groups.