AI crawlers are now 21.6% of the identified-bot visits in our August 2026 data, against 3.8% in early 2024. Eight of the twenty busiest bots we see exist to feed AI systems. When we last counted, in March 2023, not one did.
That counts only the bots with a declared AI purpose. Googlebot, still the busiest crawler on the list, is not counted among them, even though Google uses the pages it fetches to help train Gemini. So the true share of crawling that feeds AI is higher than anything this article can measure.
The rest of the picture is familiar. HTTP client libraries still generate a mountain of automated traffic, only some of which is crawling. Almost nobody crawls with a tablet User-Agent. The AI tier is what moved, so we re-ran the numbers.
This list only includes bots that identify themselves. For the wider problem of undeclared and disguised automated traffic, see Introduction to Bot Traffic and Dark Traffic and Misrepresentation.
Contents
- Web crawlers list
- The rise of AI crawlers
- User-Agents of the most active web crawlers
- How to verify a crawler is genuine
- How to block AI crawlers
- Detecting web crawlers and bots
- Frequently asked questions
- Method
Web crawlers list
These are the most active self-identified bots in DeviceAtlas traffic for August 2026, classified by the DeviceAtlas Enterprise API and grouped by operator. The share is of identified-bot visits that month. Read it as an ordering, not as a measure of how much crawler traffic any one site gets.
| Web crawler | Purpose | Share |
|---|---|---|
| Googlebot and the wider Google fleet | Search indexing, ads, rendering, Read-Aloud, Inspection Tool. The crawled data also helps train Gemini | 18.4% |
| Meta-ExternalAgent, Meta-ExternalAds, meta-webindexer, facebookexternalhit | AI training, advertising, indexing and link previews | 14.2% |
| Headless Chrome | Chromium driven from a server, for rendering and extraction. Not one crawler but many | 4.5% |
| OAI-SearchBot, GPTBot, ChatGPT-User | OpenAI's search index, training crawler and assistant fetcher | 3.5% |
| AhrefsBot / SemrushBot / MJ12bot / DotBot | SEO and backlink audits | 3.4% |
| Bingbot / Bing Ads | Microsoft search and ads | 3.3% |
| Chrome-Lighthouse / Pingdom / UptimeRobot / WebPageTest / SpeedCurve | Performance and uptime monitoring | 2.8% |
| ClaudeBot, Claude-SearchBot, Claude-User, Claude-Code | Anthropic model training, AI search and assistant fetches | 2.5% |
| Bytespider | ByteDance's crawler. Documented for Toutiao search; widely classified as AI, though ByteDance states no purpose | 2.4% |
| Qwantbot | Qwant search (EU) | 2.1% |
| Baiduspider | Baidu search (China) | 1.7% |
| DuckDuckGo Bot | DuckDuckGo search — the token is DuckDuckBot, not the name above |
1.6% |
| Applebot | Apple search and Siri | 1.5% |
| Amazonbot | Amazon search, Alexa and AI answering | 1.1% |
| PetalBot | Huawei Petal search | 0.9% |
| Sogou Spider | Sogou search (China) | 0.7% |
| YandexBot | Yandex search | 0.7% |
Five of those rows crawl mainly or only for AI, and two of them are in the top six. In 2023, not one bot on this list was described that way.
The other change since 2023 is at the top. Google still leads, but Meta is now the second-busiest operator on the list: 14.2% of identified-bot visits against Google's 18.4%. Meta got there with a fleet rather than one bot. It documents five crawlers: meta-externalagent for model training and product indexing, meta-webindexer for Meta AI search, meta-externalads for advertising, facebookexternalhit for link previews, and meta-externalfetcher for pages a user asks about.
Google's block is striped for a reason that never shows up in a log. Google says it crawls a page once and may then use that data several times, and that content it has crawled may help train future Gemini models. So the opt-out it offers publishers is a robots.txt token with no crawler behind it, governing what happens to the data rather than the fetch. The AI agents Google runs in their own right, a Vertex AI crawler and an agent fetcher, are 0.7% of its traffic. Google is not the only operator in that position. Bytespider is the same question asked of a company that will not answer it. ByteDance points the crawler at two of its own documents: one calls it the Toutiao Search crawler, and the newer one, linked from the User-Agent itself, says only that it collects publicly available content and names no purpose at all.
The rise of AI crawlers
A whole class of bots built to feed AI systems arrived between the two versions of this article, and we can watch it happen. Over the past two and a half years the AI-crawler share of identified-bot visits rose more than fivefold, from 3.8% to 21.6%. Cloudflare, which fronts a large slice of the web, reported in 2025 that AI training had grown to drive close to 80% of the AI-bot activity it sees.
Meta's crawlers stopped looking like crawlers
Meta publishes a User-Agent for each of its crawlers, and every one is a bare token. In August 2026 only 13% of the traffic from its four meta-* bots arrived in that shape. The rest wraps the same token inside a full browser User-Agent.
Documented: what Meta publishes
meta-externalagent/1.1 |
meta-webindexer/1.1 |
meta-externalads/1.1 |
meta-externalfetcher/1.1 |
Observed: how meta-externalagent actually arrives
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36 (compatible; meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)) |
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36 (compatible; meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)) |
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36 (compatible; meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)) |
The whole fleet is moving, not one bot. meta-externalads changed in May 2025, and meta-externalagent and meta-webindexer both switched in July 2026.
Meta is still declaring itself. The token and the documentation link are both in the string, so this is not cloaking, and Meta could close the gap tomorrow by updating either the docs or the crawler.
That is the point worth keeping: a published User-Agent is a starting point for a crawler's identity, not a specification of it, and the gap between the two opens without notice. Keeping up with it is exactly what a detection service is for. DeviceAtlas classifies what actually arrives rather than what the documentation says should.
Three kinds of AI crawler, and why you might treat them differently
- Training crawlers collect text to train future models: GPTBot (OpenAI), ClaudeBot (Anthropic), Meta-ExternalAgent (Meta), CCBot (Common Crawl) and Bytespider (ByteDance). Two of those do more than train: Meta documents its crawler as indexing content for its products as well, and Bytespider also feeds ByteDance's own search.
- AI search crawlers build the indexes behind AI answer engines: Meta-WebIndexer, OAI-SearchBot, Claude-SearchBot and PerplexityBot. Blocking these is closer to blocking a search engine: you lose visibility in that product's answers.
- User-triggered fetchers, Google's term for the category, load one page in real time because a person asked an assistant about it: ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher and DuckDuckGo's DuckAssistBot. Each hit stands for one human request, so their volume tracks your visitors rather than a crawl schedule.
Does ChatGPT fetch my page when someone asks about it?
Yes, in real time, and it is a different bot from the one that trained the model. Two of the categories above are involved when an assistant answers from the web: the AI search crawler supplied the candidate pages weeks ago, and a user-triggered fetcher loads the page the assistant reads now. That live fetch is what lets the answer reflect a page as it stands today rather than as the model learned it in training. Every mainstream AI assistant now works this way. Answering from pages retrieved at question time, rather than from the model's memory alone, is called retrieval-augmented generation, usually shortened to RAG.
User-Agents of the most active web crawlers
AI crawlers: GPTBot, ClaudeBot, PerplexityBot and more
Most AI bots publish a distinct User-Agent token and still send it. Fewer send it on its own. Only 37% of AI-crawler visits carry the bare token, and the rest wrap it inside a full browser User-Agent. Meta's fleet is the biggest example, not the only one.
A representative set from our August 2026 data.
AI crawler User-Agent strings
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.4; +https://openai.com/gptbot) |
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com) |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) |
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler) |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) Chrome/119.0.6045.214 Safari/537.36 |
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com) |
Google generates the most crawler traffic, and it is not one bot but a fleet. Google now describes Googlebot as one client of a shared crawling platform rather than the crawler itself, and puts the consequence plainly: a Googlebot line in your log is Google Search and nothing else. Its other products crawl through the same infrastructure under their own names, which is where the ads bots, Google-Read-Aloud and the Inspection Tool used by Search Console come from. GoogleOther is the odd one out, a generic crawler that Google says is tied to no single product.
Google User-Agent strings
Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html) |
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/151.0.7922.137 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html) |
Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/138.0.0.0 Mobile Safari/537.36 (compatible; Google-Read-Aloud; +https://support.google.com/webmasters/answer/1061943) |
AdsBot-Google (+http://www.google.com/adsbot.html) |
GoogleOther |
The Chrome version in Google's mobile User-Agent tracks the current stable Chromium release rather than sitting frozen. Google's documentation writes it as W.X.Y.Z, and we saw Chrome/150 and 151 in the same month. The bot that does use Chrome's reduced Android 10; K User-Agent is Google-Read-Aloud.
Not every Google agent is browser-shaped. The image and video crawlers, the ads bot, Feedfetcher and several product fetchers send a bare token with no Mozilla/5.0 prefix at all, such as Googlebot-Image/1.0 or AdsBot-Google (+http://www.google.com/adsbot.html).
Bing
The Bing crawler, Bingbot, appears in both desktop and mobile forms.
Bingbot User-Agent strings
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/136.0.0.0 Safari/537.36 |
Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) |
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/136.0.0.0 Mobile Safari/537.36 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) |
Crawlers that claim to be a mobile device
Mobile-first indexing means Google's crawlers largely ask for the mobile version of your page, and they do it by borrowing a mobile device identity. It is not the rule everywhere. Meta's crawlers present as desktop, and across all identified bots desktop identities outnumber mobile ones. The device they claim is a fixture rather than a real handset: Googlebot Mobile still sends a Nexus 5X, a 2015 handset it has advertised since April 2016, Bytespider an Android 5.0, and Meta's ads crawler an iPhone on iOS 13. Only the testing tools carry anything current.
Mobile crawler User-Agent strings
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/150.0.7871.186 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html) |
Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/138.0.0.0 Mobile Safari/537.36 (compatible; Google-Read-Aloud; +https://support.google.com/webmasters/answer/1061943) |
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; https://bytedance.sg.larkoffice.com/docx/K5bxdypulop3IIxrJb0lOjLVgFe) |
Mozilla/5.0 (iPhone; CPU iPhone OS 13_2_3 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/13.0.3 Mobile/15E148 Safari/604.1 (compatible; meta-externalads/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)) |
Mozilla/5.0 (Linux; Android 7.0;) AppleWebKit/537.36 (HTML, like Gecko) Mobile Safari/537.36 (compatible; PetalBot;+https://webmaster.petalsearch.com/site/petalbot) |
Mozilla/5.0 (Linux; Android 11; moto g power (2022)) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/136.0.0.0 Mobile Safari/537.36 Chrome-Lighthouse |
One detail in there will break a naive pattern. About half of PetalBot's visits send (HTML, like Gecko) rather than KHTML, a typo in Huawei's own string, and most of the rest carry no such clause at all. A rule keyed on KHTML misses nearly all of it.
HTTP client libraries
These are not crawlers as such, but they are a large and shifting slice of non-human traffic. Anyone can send them, for any purpose, so a hit shows an automated client called your site but not why. The mix has changed since 2023, when OkHttp and cURL dominated. In August 2026, the most common library User-Agent is Dart, the language that powers Flutter apps, on 23% of library visits, then cURL at 19%. Python's requests and OkHttp follow on roughly 17% each, with Java, axios and Go behind them.
HTTP client library User-Agent strings
Dart/3.3 (dart:io) |
python-requests/2.33.1 |
curl/8.7.1 |
okhttp/4.12.0 |
Go-http-client/2.0 |
axios/1.13.5 |
SEO and monitoring crawlers
The audit and monitoring bots are steady fixtures: AhrefsBot, SemrushBot, MJ12bot and DotBot for backlinks and SEO; Chrome-Lighthouse, Pingdom, UptimeRobot and WebPageTest for performance and uptime. Facebook's facebookexternalhit still fetches pages to build link previews.
SEO and monitoring User-Agent strings
Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/) |
Mozilla/5.0 (compatible; SemrushBot/7~bl; +http://www.semrush.com/bot.html) |
facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php) |
Other web crawlers to note
Other User-Agent strings worth recognizing
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/148.0.0.0 Safari/537.36 |
Mozilla/5.0 (compatible; Applebot/0.1; http://www.apple.com/go/applebot) |
Mozilla/5.0 (compatible; Baiduspider-render/2.0; +http://www.baidu.com/search/spider.html) |
Mozilla/5.0 (compatible; YandexBot/3.0; http://yandex.com/bots) |
Mozilla/5.0 (compatible; Qwantbot/1.0; +https://help.qwant.com/bot/) |
Mozilla/5.0 (compatible; PetalBot;+https://webmaster.petalsearch.com/site/petalbot) |
Mozilla/5.0 (compatible; MJ12bot/v1.4.8; http://mj12bot.com/) |
Mozilla/5.0 (compatible; DotBot/1.2; https://opensiteexplorer.org/dotbot; help@moz.com) |
Pingdom.com_bot_version_1.4_(http://www.pingdom.com/) |
Puppeteer, Playwright and Selenium all drive Chromium, so they arrive as Headless Chrome and never under their own names. A block list that names the tool will not match any of them.
How to verify a crawler is genuine
Popularity attracts impersonation. A User-Agent that says GPTBot is trivial to fake, and plenty of requests claiming to be AI crawlers are not. Treat the User-Agent as a claim, not proof.
Two checks settle it. The first is a forward-confirmed reverse DNS lookup: resolve the requesting IP to a hostname, then resolve that hostname back and confirm it returns the same IP.
host 66.249.66.1 # -> crawl-66-249-66-1.googlebot.com host crawl-66-249-66-1.googlebot.com # -> must return 66.249.66.1
The second is an IP-range match. Most major operators publish their ranges as JSON, which is the more reliable check for bots that do not set reverse DNS:
- Google — googlebot.json and the other crawler ranges
- OpenAI — gptbot.json, searchbot.json and chatgpt-user.json
- Anthropic — published ClaudeBot ranges
- Bing — the Verify Bingbot tool
Anything you intend to act on (block, rate-limit, bill) should be verified this way rather than trusted from the User-Agent alone.
How to block AI crawlers
Because AI crawlers publish individual robots.txt tokens, you can make fine-grained choices without touching search. This keeps your content out of model training while leaving Googlebot, and your search ranking, untouched:
# Out of AI training, still in search User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: /
Disallowing OAI-SearchBot or PerplexityBot is the opposite trade: it pulls you out of those AI answer engines the way disallowing a search crawler pulls you out of that search engine. And neither block reliably stops a user-triggered fetch: OpenAI writes that for ChatGPT-User, "because these actions are initiated by a user, robots.txt rules may not apply", and Meta says Meta-ExternalFetcher "may bypass robots.txt rules" for the same reason. Anthropic, on the other hand, reserves no such right.
Neither choice is enforceable. robots.txt works only on bots that read it. Cloudflare reported in August 2025 that Perplexity kept fetching pages through undeclared crawlers after being blocked, and Bytespider has had that reputation for years. If you need enforcement rather than a polite request, that is a job for detection and blocking at the edge.
The pressure is real enough to have changed the defaults. In mid-2025, Cloudflare began blocking AI crawlers by default for new sites and opened a marketplace for publishers to charge AI companies per crawl.
Detecting web crawlers and bots
DeviceAtlas identifies non-human traffic in real time (robots, checkers, download agents, spam harvesters, filters, and feed readers) from the same request data shown throughout this article. Once a request is classified you can block unwanted bots, serve them a lighter response, or keep them out of your analytics so your human numbers stay accurate.
That works on the AI crawlers too, and unlike robots.txt it does not depend on the bot being honest.
Read more about bot detection and how device intelligence helps you limit ad fraud and protect your site. Device spoofing is the same problem in a different setting; see The Hidden Risk in Mobile Commerce.
For developers: paste any of the User-Agent strings above into the DeviceAtlas HTTP Headers Parser to see what properties DeviceAtlas extracts from it.
Frequently asked questions
Which web crawler is most active?
Google, spread across several bots: Googlebot, the ads bots, Google-Read-Aloud, the Inspection Tool and GoogleOther. Together they are a fifth of the identified-bot visits we see, ahead of Meta's fleet. Bing is the next-busiest single search engine.
Which AI crawler is most active?
Meta-ExternalAgent is the busiest in our August 2026 data, and Meta's AI-search crawler Meta-WebIndexer is second. Then come OpenAI's OAI-SearchBot and Bytespider, followed by ClaudeBot and Amazonbot.
What is the GPTBot User-Agent?
GPTBot User-Agent
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.4; +https://openai.com/gptbot) |
That is the form we saw most in August 2026 and the one OpenAI documents, but do not match it exactly. Thirty-five GPTBot variants arrived that month. They differ in version, with 1.2 and 1.3 both still live, and in where the brackets close: OpenAI's documentation closes one after Gecko where most of the traffic we see closes it at the end. What none of them did was wrap the token in a full browser User-Agent, which is the difference between GPTBot and Meta's crawlers. Match on the token GPTBot and nothing more.
How do I block GPTBot?
Add User-agent: GPTBot followed by Disallow: / to your robots.txt. That removes your content from OpenAI's model training without affecting Googlebot or your search ranking. It does not stop ChatGPT fetching your page live when a user asks about it; that is a separate bot, ChatGPT-User.
Can I block AI crawlers without hurting my Google ranking?
Yes. Disallowing training crawlers such as GPTBot, or the Google-Extended token, does not affect Googlebot or your search visibility. Blocking AI-search crawlers is a different trade-off, closer to opting out of a search engine.
Does blocking AI crawlers remove me from Google's AI Overviews?
No, and this is a common misunderstanding. Google-Extended controls whether your content trains Google's generative models. It is a robots.txt directive rather than a crawler, so it has no User-Agent and will never appear in your logs, which is also true of Applebot-Extended. AI Overviews are built from Google's ordinary search index, so pages that rank can still be summarized there.
Whether you appear in those features is a separate control. Google finished rolling out a Search Console toggle for generative AI Search features worldwide at the end of August 2026, and its documentation is explicit that the toggle does not affect training while Google-Extended does. The snippet controls, nosnippet and max-snippet, still limit what can be shown.
Method
Every figure is August 2026 DeviceAtlas traffic, classified with the DeviceAtlas Enterprise API. The counts are visits rather than requests. A crawler that fetches a page a thousand times counts the same as one that fetches it once, so this ranks reach across the sites the service sees rather than request volume.
Read the ranking as an order rather than a census. It reflects what reaches the service, so it supports no claim about how much crawler traffic any one site receives, and a national search engine ranks lower here than it would on sites serving its own market.
Every User-Agent shown is a real string, either observed in that traffic or published by the crawler's operator.