Track

AI agents and crawlers

See which AI assistants and crawlers read your site, and what they fetched.

AI assistants read the web for the people who use them. When someone asks ChatGPT, Claude or Perplexity a question, the assistant may fetch your pages on the spot; search indexes and training crawlers read them ahead of time. The AI agents page shows which agents visit, how often, what they fetched, and how many people then clicked through from an assistant.

Why your server reports them

Agents fetch pages without running JavaScript, so the browser script never sees them. Your server (or the edge in front of it) does. Add a few lines that report each request to /v1/agents. trackAgentVisit from @anyanalytics/analytics checks the User-Agent first and only sends requests from AI agents, so people's visits never leave your server. IP addresses are neither sent nor stored, and query strings are dropped.

Note

Crawlers that do render pages, like Googlebot and Bingbot, load the script and are picked up without any setup. Their pageviews still never count as visitors.

Set it up

Next.js

middleware.tsjs
import { trackAgentVisit } from "@anyanalytics/analytics";
import { type NextFetchEvent, type NextRequest, NextResponse } from "next/server";

export function middleware(request: NextRequest, event: NextFetchEvent) {
  // Only AI agents are sent; people's visits never leave your server.
  event.waitUntil(trackAgentVisit({ writeKey: "vxw_YOUR_WRITE_KEY", host: "https://api.anyanalytics.org" }, request));
  return NextResponse.next();
}

export const config = {
  matcher: ["/((?!_next/static|_next/image|favicon.ico).*)"],
};

Node / Express

server.jsjs
import { trackAgentVisit } from "@anyanalytics/analytics";

const agents = { writeKey: "vxw_YOUR_WRITE_KEY", host: "https://api.anyanalytics.org" };

app.use((req, res, next) => {
  // After the response, so the status code is known.
  res.on("finish", () => {
    trackAgentVisit(agents, {
      url: `${req.protocol}://${req.get("host")}${req.originalUrl}`,
      userAgent: req.get("user-agent") ?? "",
      signatureAgent: req.get("signature-agent"),
      status: res.statusCode,
    });
  });
  next();
});

Cloudflare Worker

worker.jsjs
import { trackAgentVisit } from "@anyanalytics/analytics";

// A Worker on your site's route, in front of your origin.
export default {
  async fetch(request, env, ctx) {
    const response = await fetch(request);
    ctx.waitUntil(trackAgentVisit({ writeKey: "vxw_YOUR_WRITE_KEY", host: "https://api.anyanalytics.org" }, request, response.status));
    return response;
  },
};

Any server

POST /v1/agentssh
# Send requests from AI agents (or all of them: others are ignored),
# up to 100 per call, e.g. from your access logs.
curl -X POST https://api.anyanalytics.org/v1/agents \
  -H "Authorization: Bearer vxw_YOUR_WRITE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "hits": [{
      "url": "https://example.com/pricing",
      "user_agent": "Mozilla/5.0 ... ChatGPT-User/1.0; +https://openai.com/bot",
      "status": 200,
      "timestamp": "2026-10-08T12:00:00Z"
    }]
  }'

Each request accepts up to 100 hits with url, user_agent and optional status, timestamp (up to a month back, so you can import access logs) and signature_agent (the Signature-Agent header that browsing agents like ChatGPT agent send). Hits from anything other than a known AI agent are ignored.

What the tabs mean

  • AI answers: fetched live because someone asked an assistant.
  • Indexing: building the search index assistants answer from.
  • Training: collecting pages to train models.

Agents we recognize

User-Agent tokenProductOperatorPurpose
ChatGPT-UserChatGPTOpenAIAI answers
Claude-UserClaudeAnthropicAI answers
Claude-WebClaudeAnthropicAI answers
Perplexity-UserPerplexityPerplexityAI answers
Gemini-Deep-ResearchGeminiGoogleAI answers
Google-NotebookLMNotebookLMGoogleAI answers
DuckAssistBotDuckDuckGoDuckDuckGoAI answers
MistralAI-UserMistralMistral AIAI answers
meta-externalfetcherMeta AIMetaAI answers
Amzn-UserAlexaAmazonAI answers
cohere-aiCohereCohereAI answers
OAI-SearchBotChatGPTOpenAIIndexing
Claude-SearchBotClaudeAnthropicIndexing
PerplexityBotPerplexityPerplexityIndexing
Google-CloudVertexBotVertex AIGoogleIndexing
GooglebotGoogleGoogleIndexing
bingbotBingMicrosoftIndexing
ApplebotAppleAppleIndexing
DuckDuckBotDuckDuckGoDuckDuckGoIndexing
Amzn-SearchBotAmazonAmazonIndexing
AmazonbotAmazonAmazonIndexing
YouBotYou.comYou.comIndexing
PetalBotPetalHuaweiIndexing
GPTBotOpenAIOpenAITraining
ClaudeBotAnthropicAnthropicTraining
anthropic-aiAnthropicAnthropicTraining
GoogleOtherGoogleGoogleTraining
meta-externalagentMetaMetaTraining
FacebookBotMetaMetaTraining
BytespiderByteDanceByteDanceTraining
CCBotCommon CrawlCommon CrawlTraining
cohere-training-data-crawlerCohereCohereTraining
Applebot-ExtendedAppleAppleTraining
AI2BotAi2Ai2Training
DiffbotDiffbotDiffbotTraining
TimpibotTimpiTimpiTraining
omgiliWebz.ioWebz.ioTraining
ImagesiftBotImageSiftHiveTraining
PanguBotPanguHuaweiTraining
Kangaroo BotKangarooKangaroo LLMTraining

Visitors sent by AI

The same page lists people who arrived from an assistant (chatgpt.com, perplexity.ai, claude.ai, gemini.google.com and others), from the referrer or utm_source the browser script records. That part needs no server setup.