# AI Crawler Identity Is the ads.txt Moment for AI Search

*Publishers want AI bots to say who they are and why they came. The bill treats identity as a lock. The open web's history says identity is what makes a market.*

Published: 2026-09-29 | Read time: 5 min read | Author: Smalk AI Research | Source: https://www.smalk.ai/blog/ai-crawler-identity-ads-txt-moment-ai-search

---

AI crawler identity is about to be written into law and into internet standards, and publishers are treating it as a lock when it is really the foundation of a market. The open-web ad economy only scaled once every impression was described and every seller was verified. The agentic web has neither, and the Stealth Bot Prohibition Act would supply the first half. What publishers and brands build on that declaration decides whether AI Search gets a wall or an exchange.

- The Stealth Bot Prohibition Act (H.R. 9915) would force AI crawlers to declare their identity and purpose, and more than 300 publishing executives lobbied Congress for it in September 2026.
- The News/Media Alliance frames identity as step one toward blocking and licensing, yet licensing remains bespoke, opaque and small.
- Declared purpose is the missing field of an agentic bid request: an agent fetching a page to answer a user right now is a moment of influence, and therefore inventory.
- Brands need the same disclosure publishers do, because it is the only auditable record of which pages shape the AI answers about them.

## What AI crawler identity is, and what it is not

AI crawler identity is the verifiable declaration, by an automated agent, of who operates it and what it will do with the page it fetches: index it for search, train a model, or answer a user in real time. Today it is voluntary. GPTBot and ClaudeBot announce themselves; stealth scrapers impersonate browsers. Cloudflare already sorts crawlers into Search, Training and Agent, and the bill would make that declaration a legal obligation.

## From the verified impression to the verified agent

The open-web ad economy ran on declared identity. Every programmatic bid request described the visitor, the page and the context before money moved, and IAB Tech Lab's ads.txt in 2017 and sellers.json in 2019 made the supply side verifiable. Fraud never disappeared, but buyers could finally price what they could check.

The agentic web runs on undeclared visits. TollBit counted more than 22 billion AI bot scrapes in the first half of 2026 and warns that many scrapers are indistinguishable from human readers. Nobody can price a visitor they cannot name. Identity is not the product; it is the precondition for one.

## What PPC Land and Digiday reported, and what we verified

PPC Land reported on September 29, 2026 that more than 300 executives, organized by the News/Media Alliance and America's Newspapers, lobbied Congress for H.R. 9915. Introduced on July 23 by Representatives Laurel Lee, Valerie Foushee and Gus Bilirakis, the bill would be enforced by the FTC. The same month, Cloudflare conceded that robots.txt cannot identify who is crawling, and reported that 17% of its sites block AI training while fewer than 1% block search.

Digiday's recap of its September 2026 Publishing Summit supplied the commercial backdrop. The New York Times' Adam Greenberg called AI licensing marketplaces underdeveloped. USA Today Media's Kristin Roberts said Google's share of its referrals fell from over 70% to about 40%. Future's GEO product, Future Optic, has more than 30 clients measured on citations, with attribution still unsolved.

One figure needs context. Both outlets cite TollBit's AI bot to human ratio rising from 1 in 200 to 1 in 31, but that series ends in Q4 2025 and counts AI bots only, while the 'over 60% of traffic' figure is Cloudflare's measure of all automated traffic. In TollBit's first-half 2026 sample, 22 billion AI scrapes sit inside 987 billion visits, roughly 2%. The curve is steep and the share still small: that gap is the window to design the market rather than inherit it.

## Why identify, block, then license is the wrong sequence

### The strongest case for identity as a lock

Danielle Coffey, who leads the News/Media Alliance, is explicit about the plan: once stealth crawlers identify themselves, publishers can block them, control their content and negotiate licenses. The evidence is not trivial. Reuters has blocked bots by default since May 2026, and its bot traffic fell while monetizable traffic held steady. CNBC refuses every LLM deal until compensation matches the value of what is ingested.

> The bill is a means to an end.
>
> — Danielle Coffey, President and CEO, News/Media Alliance

The opposite camp makes a serious point too. CDT, EFF and the Internet Archive warn that compelled identification leads to a gated web of pay-per-scrape fortresses that only the largest AI labs can afford.

### Why a lock on a shrinking audience is not a revenue model

Reuters proves that blocking protects human monetization. It does not create machine monetization, and the human half is the one shrinking. Summit attendees agreed that search traffic is not coming back.

Licensing does not fill the gap. Greenberg says few deals have been signed, executives complained they cannot even compare terms, and TollBit found click-through from AI apps to publishers holding deals fell from 8.8% to 1.33% across 2025. CDT is right that identity plus walls produces a gated web. It is wrong that anonymity is the cure. The open web stayed free for humans because brands paid for access, and declared agents can be brought into the same bargain.

## What declared purpose means for brands and for publishers

### For CMOs, media buyers and agencies: ask for the agent log, not only the citation report

Future's Mike Peralta and Ziff Davis' Steve Horowitz both concede the AI visibility attribution loop is still open. A declared agent fetch is the closest thing the agentic web has to a served impression: an operator, a timestamp, a page and a stated purpose. Buyers should back bot identity standards the way they backed ads.txt, and require partners to show which pages AI engines actually fetched for their category.

### For publishers: license the archive, sell the live fetch

Treat the three declared purposes as three businesses. Search fetches keep you discoverable. Training fetches are an archive licensing question, negotiated rarely and in bulk. Agent fetches happen while a user waits for an answer, which makes them moments of influence brands will pay to be present in. Spend the identity dividend on pricing, not only on blocking.

## Three signals to watch before mid-2027

- Standards: the IETF Web Bot Auth drafts, written by engineers from Cloudflare, Amazon and Google, define a signed agent card listing identity, purpose and rate expectations, in effect a sellers.json for machine visitors.
- Deadlines: the UK CMA publisher conduct requirement takes effect on December 3, 2026, Microsoft targets robots.txt AI controls in early 2027, Google's page-level generative search controls arrive on March 3, 2027, and H.R. 9915 still needs committee and floor votes.
- Taxonomy: whether 'agent', as a declared purpose, gets a definition stable enough for buyers to transact on. That is the moment identity becomes inventory.

## Conclusion

Hold on to this: AI crawler identity is not a lock but the first line of an agentic bid request, and the open web has already shown what happens once every visitor can be named. Smalk AI is building the market on the other side of that declaration: a GEA ad network that places native ads for AI agents on the publisher pages AI engines rely on, giving brands measurable presence inside answers and publishers revenue from the machine audience. What to watch next: whether 'agent' becomes a declared purpose standardized enough to buy.

## FAQ

### What is the Stealth Bot Prohibition Act?

The Stealth Bot Prohibition Act (H.R. 9915) is a bipartisan US House bill introduced on July 23, 2026 that would require automated web crawlers to disclose their identity and purpose, and ban bots that impersonate humans for generative AI purposes. The FTC would enforce it. It must still pass committee and both chambers before becoming law.

### How is AI crawler identity different from robots.txt?

robots.txt is a preference a website publishes, while AI crawler identity is a declaration the crawler makes. Cloudflare acknowledged in September 2026 that robots.txt cannot identify who is crawling. Identity is what lets a site verify, enforce and eventually price access.

### Does blocking AI bots protect publisher revenue?

Blocking protects revenue from human readers, as Reuters showed after moving to block-by-default in May 2026. It does not generate revenue from the machine audience, which is the part of traffic that is growing. Blocking is a defense and needs a commercial model beside it.

### Why should advertisers care about AI bot identification?

Advertisers should care because the pages AI agents fetch shape the answers that recommend or ignore a brand. Verified, purpose-declared fetches are the only auditable record of that influence. They are the basis for measuring and buying AI Search visibility.

### Will forcing bots to identify themselves close the open web?

It can, if identity only feeds walls. Civil liberties groups including CDT and EFF warn of a gated web of pay-per-scrape fortresses. If identity feeds a brand-funded ad market instead, content can stay open to agents the way advertising kept it free for humans.
