Perplexity Accused of Secretly Scraping Blocked Websites—Cloudflare Sounds the Alarm

Listen to this Post

Featured Image
Introduction: The Battle Over AI Web Crawling Just Got Personal

In a rapidly evolving AI landscape, tensions between tech platforms and AI companies are starting to boil over. Cloudflare, one of the internet’s most trusted content delivery and security companies, has publicly accused the rising AI player Perplexity of bypassing website restrictions in an underhanded attempt to harvest web content. These allegations strike at the heart of ongoing debates around AI training ethics, copyright law, and user consent. While many AI giants are striking deals with publishers or abiding by web protocols, Perplexity is now under scrutiny for allegedly doing the opposite—scraping data behind the scenes, despite clear barriers.

This controversy re-ignites the fundamental question: can AI startups continue scaling without respecting the digital boundaries set by content creators? And more importantly, are new technologies pushing the line between innovation and exploitation too far?

Original Cloudflare vs. Perplexity—The Hidden Crawling War

Cloudflare has accused Perplexity, a rising AI search engine startup, of violating website owners’ clear instructions to stay out. Websites commonly use a robots.txt file or firewall (WAF) rules to prevent bots from crawling their content. However, Cloudflare reports that Perplexity has been evading these protections by disguising itself as a regular user—mimicking a Chrome browser on a Mac and rotating its IP addresses to fly under the radar.

The accusations mirror prior claims from WIRED and Forbes, both of which also stated that Perplexity scraped their content despite blocks. Cloudflare’s investigation involved creating decoy domains with all known protections activated. Even then, Perplexity allegedly accessed content using stealth tactics, including rotating through different IP addresses and Autonomous System Numbers (ASNs) to hide its true identity.

Cloudflare says that not only did Perplexity access this restricted content—it then used it to answer queries through its AI platform, violating the explicit intent of site owners.

To combat this, Cloudflare has now upgraded its bot detection systems, allowing customers to block Perplexity’s disguised crawlers. This protection is included even for free-tier users. Cloudflare also contrasted Perplexity’s behavior with OpenAI, which it said respects website restrictions—although OpenAI is still facing lawsuits (including one from undercode’s parent company) over unauthorized content usage in AI training.

As more media outlets strike licensing deals with AI companies—including Gannett with Perplexity—the line between cooperation and infringement continues to blur. Still, Cloudflare insists Perplexity is breaking the rules, and the startup has yet to respond publicly to the allegations.

What Undercode Say: Perplexity’s Crawl Tactics Raise Bigger Industry Red Flags

Perplexity’s alleged stealth crawling isn’t just a technical issue—it’s a moral and legal landmine that could reshape how AI companies collect data in the future.

First, let’s consider the core accusation: Perplexity bypassed robots.txt directives and WAF firewalls by impersonating human browsers. This is akin to walking through a locked door by disguising yourself as the building’s janitor. The practice may be clever, but it clearly undermines the trust-based structure of the web. If true, it erodes the implied agreement between site operators and bot developers.

The use of rotating IPs and different ASNs is even more disturbing. These are techniques typically seen in botnets and scrapers trying to evade detection—not from companies that claim to operate ethically. If AI firms resort to such tactics, what separates them from traditional black-hat web scrapers?

From a broader industry lens, this situation underlines a key issue: content is now fuel for AI. But that fuel isn’t free. Companies like OpenAI, Anthropic, and Meta are increasingly paying for licensed access to data. Perplexity, by allegedly taking a shortcut, risks undercutting that entire licensing ecosystem. If some players cheat, why should others pay?

Cloudflare’s response is timely. By upgrading its bot detection and offering free protective rules, it’s empowering web owners to take back control. Their new “Pay Per Crawl” program is also a novel solution, allowing publishers to monetize access on their own terms. This could be a game-changer if adopted widely—replacing the unauthorized crawl economy with a sustainable one.

Yet, we must also recognize the legal gray zones. While bypassing robots.txt may be unethical, it’s not necessarily illegal. U.S. law has long debated whether data scraping is protected under fair use or if it violates terms of service. Until courts or lawmakers make a definitive ruling, we’ll continue to see these skirmishes escalate.

It’s also crucial to address perception. Perplexity has quickly gained a loyal user base and favorable media coverage for its fast, citation-based answers. But this positive press could unravel if the public sees the brand as a data thief rather than an innovator.

This case should serve as a warning across the tech world: transparency matters. As AI companies scramble to feed their models, those who sidestep ethical practices risk losing public trust—and facing massive lawsuits.

🔍 Fact Checker Results

✅ Cloudflare confirmed the stealth crawling behavior through direct testing using decoy domains.
✅ Previous reports from WIRED and Forbes independently accused Perplexity of similar actions.
❌ Perplexity has not issued a public denial or explanation as of August 2025.

📊 Prediction: Expect More Legal Showdowns Over AI Scraping

The Perplexity vs. Cloudflare conflict is just the beginning. In the coming 12 months, we can expect:

A rise in lawsuits against AI companies for unauthorized data use, especially from publishers and media companies.
More content delivery networks offering AI-blocking tools, possibly turning it into a standard website feature.
Regulatory scrutiny from governments worldwide, who are still playing catch-up on how to define and enforce digital consent in the age of AI.

Unless Perplexity responds with transparency and reforms, it risks both its legal standing and its public credibility. AI companies must choose: work with content creators—or risk being locked out for good.

🕵️‍📝✔️Let’s dive deep and fact‑check.

References:

Reported By: www.zdnet.com
Extra Source Hub:
https://www.reddit.com
Wikipedia
OpenAi & Undercode AI

Image Source:

Unsplash
Undercode AI DI v2

🔐JOIN OUR CYBER WORLD [ CVE News • HackMonitor • UndercodeNews ]

💬 Whatsapp | 💬 Telegram

📢 Follow UndercodeNews & Stay Tuned:

𝕏 formerly Twitter 🐦 | @ Threads | 🔗 Linkedin | 🦋BlueSky | 🐘Mastodon