Google built one of the world’s largest companies by scraping the entire web without asking permission first. This week a court told Google it cannot stop others from scraping Google.

On 20 July, a federal judge dismissed Google’s lawsuit against SerpApi. The firm scrapes Google search results and resells them as structured data. Google had sued in December under the Digital Millennium Copyright Act, arguing SerpApi circumvented its anti-bot defences to lift copyrighted material.

The problem, the court found, is that search results are not copyrighted material.

Why the case fell apart

The DMCA provision Google used, Section 1201, only protects technology that guards copyrighted works. Google’s barrier, called SearchGuard, guards its search results.

Plain results such as URLs, snippets and factual index data are public facts, not works, Judge Yvonne Gonzalez Rogers ruled. Those claims were thrown out with no chance to refile.

“Google has not alleged a plausible violation of the DMCA,” she wrote.

The judge was not soft on SerpApi’s methods. She agreed that spoofing browser fingerprints, rotating IP addresses and solving CAPTCHAs to get past SearchGuard counts as circumvention. It simply is not illegal unless the barrier protects copyrighted work with the owner’s permission.

The corner Google is now in

Google has one narrow way back. The court gave it 21 days to refile on a smaller claim, covering the licensed snippets that sometimes appear in its knowledge panels.

To do that, Google would have to argue those panels are full of copyrighted content it is authorised to protect. According to Meredith Rose, a DMCA specialist at the nonprofit Public Knowledge, that is a dangerous thing for Google to say out loud.

Google does not license everything in a knowledge panel. Suppose it argues that assembling those panels reproduces copyrighted material. It then invites the rights holders whose work fills its results, and its AI Overviews, to sue Google for the same scraping it is complaining about.

“They have talked themselves into a little bit of a corner,” Rose told Ars Technica. To win the small fight, in other words, Google may have to concede the big one.

Google says it will try anyway. Spokesperson José Castañeda said the company was “pleased to see that the Court rejected nearly all of SerpApi’s legal arguments” on standing. Google plans to file an amended complaint.

What Google was really protecting

Search data has never been more valuable. A chatbot cannot summarise the web if it cannot find it, and Google offers no official search API. That makes second-hand access through firms like SerpApi the main route to Google’s index. Its customers include Nvidia, Uber, Adobe and the AI search engine Perplexity.

SerpApi framed the dismissal as bigger than itself. Google and Reddit “do not own the Internet,” it said, accusing both of trying to “weaponise the DMCA to wall off the open Internet.” Its chief executive Julien Khaleghy noted that Google’s own damages maths, applied literally, would top the entire US economy.

The irony is not subtle, and critics have not let it pass. Google spent two decades scraping publishers to build a search empire, and is now using copyright law to stop a company doing something similar to Google.

Reddit is watching, and exposed

Google was not first here. Reddit filed a near-identical DMCA suit in October against SerpApi and Perplexity, over Reddit content that surfaces in Google results. Reddit faced its own dismissal hearing days after Google’s loss.

The Google ruling is awkward for it. Reddit is neither the copyright owner nor the exclusive licensee of content that appears in search results, which is the exact gap that sank Google’s case. SerpApi has warned that Reddit wants to act as a “toll collector,” charging for access to content its users wrote. Reddit has been tightening the taps on its data for a while. (Advance Publications, which owns Ars Technica’s parent company, is Reddit’s largest shareholder.)

The bigger fight over the open web

Rose put the case in a wider frame. A wave of aggressive AI scraping began around 2023. Publishers have raced to wall off their content ever since, a “re-enclosure” of a web that was mostly open.

That reaction does not only slow AI training. It also breaks anonymous crawling that research, archiving, journalism and public-health reporting quietly depend on. The tools built to stop the scrapers catch everyone.

The stakes are the same ones running through Google’s AI search overhaul and its fights with regulators over search data. Who gets to read the public web at scale, and who gets to charge for it, is being decided case by case. Google’s grip is already loosening in the AI era. It has 21 days to decide whether this fight is worth the risk of losing much more.

Get the TNW newsletter

Get the most important tech news in your inbox each week.