AI crawler directory

Every AI crawler and user agent WriteWorks recognises

A reference to the AI crawlers that visit websites today: who runs each one, what its user agent looks like, and what WriteWorks classifies it as. 27 named signatures across 15 platforms, each mapped to citation, indexing, training or assistant intent.

7-day trial · Cancel anytime

The challenge

Why an AI crawler directory matters

Names multiply, purposes differ, and totals get muddy.

The list of AI user agents keeps growing

New AI crawlers appear and vendors split one bot into several. Keeping a hand-maintained regex of AI user agents current is a job nobody owns.

A name does not tell you the purpose

Bytespider, meta-externalagent and Amazonbot look alike in a log, but whether a bot collects training data or fetches pages for live answers changes what its visits mean.

Unknown bots muddy the totals

Without a registry, every self-declared crawler is counted the same way, and AI traffic is either overstated or lost in general bot noise.

Third parties hold the citation

Roundups, reviews and forums are quoted ahead of your own pages, and nothing tells you which source displaced you.

The solution

The AI crawlers WriteWorks recognises

Grouped by the engine behind them, with each bot's classification.

ChatGPT-User (AI Citation), OAI-SearchBot (AI Indexing), GPTBot (AI Training) and OAI-Operator (AI Assistant).

◆The AI crawlers WriteWorks recognisesLIVE
Customer story

How teams are winning AI search

From invisibility to category dominance across every major answer engine.

Share of voice, proven

“We retired three legacy tools and rebuilt our motion around WriteWorks. We can finally see our share of voice on the engines our buyers ask, and prove the lift on every run.”

JK
Jordan Kessler Head of Growth, GoTeachingJobs
Capabilities

How identification works

From raw user agent to engine and intent.

Longest-match identification

Applebot-Extended is never mistaken for Applebot.

Platform resolution

Raw user agents rolled up to the engine behind them.

Bot-type labels

Every named crawler mapped to one of four intents.

Per-platform view

Each engine's signatures with real per-agent visit counts.

Signed agent detection

Web Bot Auth Signature-Agent requests recognised.

Generic crawler capture

Unlisted self-declared bots logged under their own name.

Time savings

What changes with a maintained registry

No more hand-kept user-agent lists.

TaskBeforeWith WriteWorksTime saved
Keep an AI user-agent list currentHand-maintained regexMaintained registryOngoing upkeep
Know what each bot is forVendor docs, one by oneBot type on every rowResearch time
Included

Registry at a glance

Signatures, platforms and bot types.

27 named AI crawler signatures
15 platforms: ChatGPT, Claude, Perplexity, Google, Meta, Microsoft, Apple, Amazon, ByteDance, DeepSeek, xAI, You.com, Common Crawl, Mistral, Cohere
Four bot types: AI Citation, AI Indexing, AI Training, AI Assistant
Per-platform breakdown with per-agent visit counts
Generic capture for unlisted self-declared crawlers
Signed agent recognition via Signature-Agent
Built for

Who uses the directory

Anyone reading AI traffic in their logs.

Technical SEO

A working reference for every AI user agent in your logs.

Developers and DevOps

Know which crawlers are hitting your infrastructure and why.

Publishers

Separate training collection from search and live fetches before setting policy.

Featured story

“We closed the share-of-voice gap across every major engine. Our content team now ships citation-ready by default, and we have the proof to show it.”

Read the full story
10+ AI platforms
One closed loop
Citation gap
Diagnosed and closed
Precision
Not just presence
FAQ

Frequently asked questions

Everything teams typically ask before getting started.

What are AI crawlers?+
AI crawlers are bots run by AI companies to fetch web pages, either to collect training data, to build a search index an assistant can draw on, or to fetch a page live while answering a user. Each identifies itself with a user-agent string such as GPTBot or ClaudeBot.
What is Bytespider?+
Bytespider is the crawler operated by ByteDance, the company behind TikTok. WriteWorks resolves it to ByteDance and classifies it as AI Training.
What is CCBot?+
CCBot is Common Crawl's crawler. Common Crawl publishes an open archive of the web that many organisations use to build training datasets. WriteWorks resolves CCBot to Common Crawl and classifies it as AI Training.
What is meta-externalagent?+
meta-externalagent is a Meta crawler, and meta-externalfetcher is a related Meta agent. WriteWorks resolves both to Meta and classifies them as AI Training.
Can user agents be faked?+
Yes, any client can claim to be GPTBot. User-agent matching is the standard way to identify crawlers, so treat sudden spikes with care and verify against vendors' published IP ranges if a decision depends on it.
Explore more

Keep exploring

Related solutions, adjacent use cases, and platform features.

See which of these crawlers visit you

27 named AI crawler signatures across 15 platforms, classified by intent and counted against your own pages.