Track
AI agents and crawlers
See which AI assistants and crawlers read your site, and what they fetched.
AI assistants read the web for the people who use them. When someone asks ChatGPT, Claude or Perplexity a question, the assistant may fetch your pages on the spot; search indexes and training crawlers read them ahead of time. The AI agents page shows which agents visit, how often, what they fetched, and how many people then clicked through from an assistant.
Why your server reports them
Agents fetch pages without running JavaScript, so the browser script never sees them. Your server (or the edge in front of it) does. Add a few lines that report each request to /v1/agents. trackAgentVisit from @anyanalytics/analytics checks the User-Agent first and only sends requests from AI agents, so people's visits never leave your server. IP addresses are neither sent nor stored, and query strings are dropped.
Note
Crawlers that do render pages, like Googlebot and Bingbot, load the script and are picked up without any setup. Their pageviews still never count as visitors.
Set it up
Next.js
import { trackAgentVisit } from "@anyanalytics/analytics";
import { type NextFetchEvent, type NextRequest, NextResponse } from "next/server";
export function middleware(request: NextRequest, event: NextFetchEvent) {
// Only AI agents are sent; people's visits never leave your server.
event.waitUntil(trackAgentVisit({ writeKey: "vxw_YOUR_WRITE_KEY", host: "https://api.anyanalytics.org" }, request));
return NextResponse.next();
}
export const config = {
matcher: ["/((?!_next/static|_next/image|favicon.ico).*)"],
};Node / Express
import { trackAgentVisit } from "@anyanalytics/analytics";
const agents = { writeKey: "vxw_YOUR_WRITE_KEY", host: "https://api.anyanalytics.org" };
app.use((req, res, next) => {
// After the response, so the status code is known.
res.on("finish", () => {
trackAgentVisit(agents, {
url: `${req.protocol}://${req.get("host")}${req.originalUrl}`,
userAgent: req.get("user-agent") ?? "",
signatureAgent: req.get("signature-agent"),
status: res.statusCode,
});
});
next();
});Cloudflare Worker
import { trackAgentVisit } from "@anyanalytics/analytics";
// A Worker on your site's route, in front of your origin.
export default {
async fetch(request, env, ctx) {
const response = await fetch(request);
ctx.waitUntil(trackAgentVisit({ writeKey: "vxw_YOUR_WRITE_KEY", host: "https://api.anyanalytics.org" }, request, response.status));
return response;
},
};Any server
# Send requests from AI agents (or all of them: others are ignored),
# up to 100 per call, e.g. from your access logs.
curl -X POST https://api.anyanalytics.org/v1/agents \
-H "Authorization: Bearer vxw_YOUR_WRITE_KEY" \
-H "Content-Type: application/json" \
-d '{
"hits": [{
"url": "https://example.com/pricing",
"user_agent": "Mozilla/5.0 ... ChatGPT-User/1.0; +https://openai.com/bot",
"status": 200,
"timestamp": "2026-10-08T12:00:00Z"
}]
}'Each request accepts up to 100 hits with url, user_agent and optional status, timestamp (up to a month back, so you can import access logs) and signature_agent (the Signature-Agent header that browsing agents like ChatGPT agent send). Hits from anything other than a known AI agent are ignored.
What the tabs mean
- AI answers: fetched live because someone asked an assistant.
- Indexing: building the search index assistants answer from.
- Training: collecting pages to train models.
Agents we recognize
| User-Agent token | Product | Operator | Purpose |
|---|---|---|---|
ChatGPT-User | ChatGPT | OpenAI | AI answers |
Claude-User | Claude | Anthropic | AI answers |
Claude-Web | Claude | Anthropic | AI answers |
Perplexity-User | Perplexity | Perplexity | AI answers |
Gemini-Deep-Research | Gemini | AI answers | |
Google-NotebookLM | NotebookLM | AI answers | |
DuckAssistBot | DuckDuckGo | DuckDuckGo | AI answers |
MistralAI-User | Mistral | Mistral AI | AI answers |
meta-externalfetcher | Meta AI | Meta | AI answers |
Amzn-User | Alexa | Amazon | AI answers |
cohere-ai | Cohere | Cohere | AI answers |
OAI-SearchBot | ChatGPT | OpenAI | Indexing |
Claude-SearchBot | Claude | Anthropic | Indexing |
PerplexityBot | Perplexity | Perplexity | Indexing |
Google-CloudVertexBot | Vertex AI | Indexing | |
Googlebot | Indexing | ||
bingbot | Bing | Microsoft | Indexing |
Applebot | Apple | Apple | Indexing |
DuckDuckBot | DuckDuckGo | DuckDuckGo | Indexing |
Amzn-SearchBot | Amazon | Amazon | Indexing |
Amazonbot | Amazon | Amazon | Indexing |
YouBot | You.com | You.com | Indexing |
PetalBot | Petal | Huawei | Indexing |
GPTBot | OpenAI | OpenAI | Training |
ClaudeBot | Anthropic | Anthropic | Training |
anthropic-ai | Anthropic | Anthropic | Training |
GoogleOther | Training | ||
meta-externalagent | Meta | Meta | Training |
FacebookBot | Meta | Meta | Training |
Bytespider | ByteDance | ByteDance | Training |
CCBot | Common Crawl | Common Crawl | Training |
cohere-training-data-crawler | Cohere | Cohere | Training |
Applebot-Extended | Apple | Apple | Training |
AI2Bot | Ai2 | Ai2 | Training |
Diffbot | Diffbot | Diffbot | Training |
Timpibot | Timpi | Timpi | Training |
omgili | Webz.io | Webz.io | Training |
ImagesiftBot | ImageSift | Hive | Training |
PanguBot | Pangu | Huawei | Training |
Kangaroo Bot | Kangaroo | Kangaroo LLM | Training |
Visitors sent by AI
The same page lists people who arrived from an assistant (chatgpt.com, perplexity.ai, claude.ai, gemini.google.com and others), from the referrer or utm_source the browser script records. That part needs no server setup.