AI Crawler Activity Logs in 2026: How to Track GPTBot, Googlebot, and Bingbot Before They Eat Your Server Budget

Bots now outnumber humans on the open web. Here's how to read AI crawler logs, tell training bots from citation bots, and set up monitoring with Cloudflare, Screaming Frog, and crawler analytics in 2026.

Key takeaways

  • Automated requests passed human traffic for the first time on June 3, 2026: Cloudflare Radar put bots at 57.5% of HTML web traffic versus 42.5% human.
  • Not all crawlers deserve the same treatment. Training crawlers (GPTBot, ClaudeBot, Google-Extended) feed model training and send almost no traffic back. Search/citation crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot) decide whether your site gets cited in an answer.
  • Google-Extended is not a crawler, it's an opt-out flag. Googlebot does the actual fetching, and blocking Google-Extended does nothing to remove you from AI Overviews or AI Mode.
  • Copilot has no crawler of its own. It runs on Bingbot, so blocking Bingbot removes you from both Bing Search and Copilot at once.
  • Cloudflare's policy shift means that starting September 15, 2026, "mixed-use" crawlers get blocked by default from ad-carrying pages unless you opt back in, which makes log auditing more urgent, not less.

Why this suddenly matters

I'll be honest, a year ago "check your server logs" was advice mostly for enterprise SEOs with a Screaming Frog license and too much free time. That's not true anymore. On June 3, 2026, Cloudflare's own CEO Matthew Prince shared Radar data showing automated requests had crossed 57.5% of all HTML web traffic, pushing humans into the minority for the first time. Some of that is garden-variety bot noise, but a meaningful and fast-growing slice is AI crawlers specifically: Cloudflare put AI crawlers at 20.3% of verified bot traffic in May 2026, with AI-search bots adding another 6.5% on top. Add it up and you get roughly 26.7% of verified bot activity tied to AI in some form.

Screenshot of Cloudflare's blog post tracking AI bot crawl activity and referral traffic trends

Here's the part that should actually bother you as a site owner: Cloudflare attributed 51.8% of AI crawler requests in May 2026 to training purposes, and only 9.3% to search. Training crawlers chew through your bandwidth and CPU and send you essentially nothing back. No clicks, no citations, no referral traffic. If you're not looking at your logs, you have no idea what percentage of your infrastructure cost is going toward feeding someone else's model for free.

The three-lane mental model for AI crawlers

Before you open a log file, it helps to sort crawlers into three buckets. Treating all "AI bots" as one undifferentiated mass is the single most common mistake I see in this space.

Training crawlers exist to harvest content for model training. GPTBot, ClaudeBot (in its training capacity), Google-Extended, Applebot-Extended, and CCBot fall here. Blocking these does not remove your site from AI answers, because the model that's already trained doesn't need to re-crawl you for every response.

Search and citation crawlers are the ones that actually decide whether you show up in an answer. OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, and Bingbot live here. Block one of these and you disappear from that specific assistant's citations, full stop.

User-fetch crawlers retrieve a single page live because a person asked the assistant to read something specific. ChatGPT-User, Claude-User, and Perplexity-User are in this category. Block them and you break the "summarize this link for me" use case for your own pages.

The mistake that shows up again and again in audits: sites block the citation bot while leaving the training bot wide open. That configuration still feeds the model, it just loses you the citation. If you're going to be strict with anyone, be strict with the training lane and generous with the citation lane.

Google-Extended isn't what most people think it is

This one trips up a lot of people doing their own robots.txt audits. Google-Extended is not a crawler at all, it's a directive token. Regular Googlebot does the actual crawling of your pages; Google-Extended in your robots.txt simply tells Google not to use already-crawled content for Gemini training or AI Mode grounding. Blocking Google-Extended will not remove your site from Google AI Overviews or AI Mode, because those features run inside Googlebot's normal crawl, which is part of core Search and isn't opt-out-able the same way. If you actually want out of AI Overviews, you need the generative-AI toggle in Search Console or page-level nosnippet/max-snippet:0 tags, not a robots.txt entry.

Copilot runs on Bingbot, there's no separate agent

If you've been hunting your logs for a "Copilot crawler" user agent, stop. Microsoft Copilot's web results are powered entirely by Bingbot, the same crawler that's served Bing Search for years (UA string: bingbot/2.0; +http://www.bing.com/bingbot.htm). Blocking Bingbot removes you from both channels simultaneously. Microsoft publishes official Bingbot IP ranges in JSON format for verification, which matters because Bingbot's user-agent string is commonly spoofed. One Microsoft Q&A thread even documents a site owner who matched 124 supposedly "spamming" IPs against Microsoft's official ranges and confirmed every single one was legitimate Bingbot traffic, just a heavier crawl than expected.

There's a wrinkle worth knowing: Microsoft's newer agentic Copilot Actions feature inherits the end user's own browser user agent rather than a distinct bot UA, which makes it nearly impossible to isolate in standard log analysis.

What the crawler mix actually looks like right now

Pulling from Cloudflare's own May 2026 data, Googlebot remained the single largest AI-adjacent crawler at 27.26% of requests, with GPTBot at 11.48% and ClaudeBot at 9.73%. That's actually a reversal from April 2026, when ClaudeBot (11.69%) briefly overtook GPTBot (9.84%). Over the twelve months from July 2024 to July 2025, Googlebot's share of verified bot traffic climbed from 37.5% to 39%, GPTBot jumped from 4.7% to 11.7%, and Meta's crawler went from a rounding error at 0.9% to 7.5%. Bytespider, ByteDance's crawler, collapsed from 14.1% to 2.4% over the same window.

If you want a sense of how fast a single crawler can reshape this mix, Promptwatch's data on Meta-WebIndexer is a good illustration: that one crawler went from roughly 2.2% of all tracked AI crawler requests in mid-July 2026 to 37.8% by August 9, 2026, a roughly 17x jump in under a month. The surge reportedly ties to Meta building its own web index so Meta AI doesn't have to depend on Google, and it was heavy enough to trigger server load alerts on at least one independent site. See Promptwatch's breakdown of the Meta-WebIndexer surge for the daily numbers. If you're only watching for Meta-ExternalAgent or FacebookBot in your logs and ignoring Meta-WebIndexer specifically, you could be missing the fastest-growing line item in your whole bot budget.

Claude's citation crawler tells a similar, if smaller, story. Promptwatch's tracking shows it growing from about 30 visits a day in mid-December 2025 to several thousand a day by mid-April 2026, a hundred-fold increase in four months, though it's still only around 1% of tracked citation crawler traffic overall. Check the Claude citation crawler data if you want the inflection-point dates. It's small today. It won't stay small if the trend line holds.

Reading your own logs: what to pull and what it tells you

The fields that matter in a raw log line are the client IP, timestamp, requested URL, status code, user-agent string, and referrer. Cross-reference IP against the operator's published ranges whenever one is available (Microsoft publishes a JSON list for Bingbot; Anthropic, notably, does not publish fixed IP ranges at all, which makes IP-based blocking for Claude risky since a block can prevent the crawler from even reading your robots.txt).

A few patterns worth watching for:

  • Frequent 403s from a known AI user-agent string, which usually means a firewall rule or CDN bot-fight setting is silently blocking a crawler you meant to allow.
  • Zero recorded hits from a bot that should be visiting regularly, paired with declining AI citation visibility despite otherwise solid SEO, which often means you're unreachable for that provider and don't know it yet.
  • A single IP hitting a sequence of distinct pages (pricing, features, case studies) within a few minutes. That pattern usually signals a live user-fetch agent completing a task, not a bulk training crawler.
  • Almost no referrer headers on AI bot requests generally. Most AI bots fetch pages directly rather than arriving via a clicked link, so don't expect referrer data to help you here.

One genuinely useful audit heuristic: pull 30 days of logs and compute the ratio of training-bot requests to search/citation-bot requests. If training outnumbers citation-driving crawlers 5:1 or worse, you're effectively subsidizing someone else's model with no citation benefit in return. That's the moment to start blocking selectively rather than leaving everything open by default.

Tools for actually doing this

You have options at a few different layers: edge-level access control, log file analysis, and citation/visibility tracking. They're not interchangeable, and most sites end up using at least two.

ToolTracking layerFree tierStarting price
Cloudflare AI Crawl ControlNetwork edge, actual bot hitsYesFree; paid via Cloudflare plans
Screaming Frog Log File AnalyserServer log parsing with AI bot presets1,000 events free£99/year
PromptwatchCrawler logs + citation outcomes + content fixesExplore plan (10 prompts)$95/mo
Vercel BotIDApp-level invisible bot detectionYes (Basic mode)Included in Pro/Enterprise
TollBitScrape detection and licensingFree signupCustom

Cloudflare's AI Crawl Control is the easiest entry point if you're already on Cloudflare, it's available on every plan including free, with zero-config deployment. It monitors which AI services hit your content, tracks robots.txt compliance, and lets you set per-crawler policy (allow, charge via Pay Per Crawl, or block outright). The policy shift to watch: starting September 15, 2026, Cloudflare's default settings block "mixed-use" crawlers from any page carrying ads unless you explicitly opt back in. That applies to new customers, new sites on existing accounts, and all existing free-tier customers. Stack Overflow and reportedly Patreon have already partnered with Cloudflare on this model.

For server log analysis specifically, Screaming Frog's Log File Analyser ships with presets for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Perplexity-User, and CCBot. You drag in raw logs, filter to AI user agents, and optionally verify bot authenticity by IP lookup against published ranges. The free version caps out at 1,000 log events, the paid license is £99 a year.

Access control tells you whether bots can reach your pages. It doesn't tell you whether being crawled is actually translating into citations, traffic, or revenue. That's where a platform like Promptwatch fits in differently from the pure log tools above. It pulls AI crawler logs through your CDN (Cloudflare, Fastly, Vercel, Netlify, Akamai, AWS CloudFront, or a custom endpoint) and connects them to the citation side, showing you the crawl-to-citation path per page with an actual citation rate, not just a raw hit count. When a page gets crawled a lot but never cited, that's a signal worth acting on, and Promptwatch's Unified Actions and content agents can turn that signal directly into a fix rather than leaving you to interpret a spreadsheet of IPs.

Favicon of Promptwatch

Promptwatch

Track and optimize your brand visibility in AI search engines
View more
Screenshot of Promptwatch website

If you want to go deeper on the GEO software landscape beyond crawler logs specifically, the directory at bestgeosoftware.com covers the broader AI visibility category.

A practical robots.txt starting point

A lot of sites either block everything AI-related out of panic or allow everything out of neglect. Neither is great. A reasonable middle ground for a public-facing site that wants AI citations but not unpaid training data harvesting looks roughly like this:

# Allow AI search and retrieval bots
User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Bingbot
Allow: /

# Block AI training crawlers
User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

# Standard search engine crawlers
User-agent: Googlebot
Allow: /

User-agent: *
Allow: /

Sitemap: https://yourdomain.com/sitemap_index.xml

Treat this as a starting template, not gospel. Anthropic specifically recommends its non-standard Crawl-delay directive over outright IP blocking for rate control, since an IP block can prevent ClaudeBot from reading your robots.txt in the first place. And remember, this file only governs politeness, not enforcement. TollBit has reportedly detected more than 9 billion AI bot scrapes, with over 2.9 billion of them ignoring or bypassing robots.txt entirely. If you actually need enforcement, that's an edge-level job for something like Cloudflare, not a text file.

Putting it together

The honest takeaway here is that "checking your logs" used to be optional technical SEO hygiene. In a year where bots make up more than half of HTML traffic and training crawlers outnumber citation crawlers more than 5 to 1 on the open web, it's closer to a cost-control exercise. Pull 30 days of raw logs, separate training from citation crawlers, verify the ones you're unsure about by IP, and decide deliberately what you're willing to feed for free versus what you want showing up when someone asks ChatGPT, Gemini, or Copilot a question that your page actually answers. If you'd rather not do that manually every month, tools like Promptwatch can turn the log side and the citation side into one dashboard instead of two separate chores.

Share:

© 2026 Surferstack · Find the best Marketing tools for your GTM motion · RSS

Surferstack is an affiliate review site. When you click links to vendors or buy through links on our site, we may earn an affiliate commission at no extra cost to you.

Surferstack is a review website based on user reviews on Reddit and G2, and on publicly available information. We keep everything as up to date as possible, but pricing and features can change. Always confirm the details with the vendor before purchasing.