# Crawlers

> RanqoBot, Ranqo's own crawler, and how to allow or block it, plus the AI crawlers Site Access tests and the bots Site Tracking recognises.

Source: https://ranqo.ai/docs/methodology/crawlers

Crawlers come up in Ranqo in two ways. Ranqo has its own crawler, RanqoBot, which fetches pages for some features. And two features work with the AI crawlers that visit your site: Site Access tests whether they can fetch it, and Site Tracking counts their real visits. Each of those two features keeps its own list, and this page says which is which.

## RanqoBot

RanqoBot is the crawler Ranqo uses whenever it follows links or fetches in volume. Its User-Agent contains the token `RanqoBot` and a link to [ranqo.ai/bot](https://ranqo.ai/bot), the page for site owners that describes exactly what it does. If you found RanqoBot in your access logs, that page is the full answer.

Each of its requests traces to someone using Ranqo:

- **Page discovery.** When a Ranqo customer audits a site, RanqoBot follows internal links on that one site to find pages worth auditing.
- **Outbound link checks.** When an audited page links to another site, RanqoBot checks the link still works, for up to 50 links per audit.
- **Citation sources.** When an AI engine cites a page, RanqoBot reads its title, or follows the engine's redirect link to the real address, so the publisher can be named.
- **Site Access.** When [Site Access](https://ranqo.ai/docs/guides/site-access) checks a Ranqo customer's site, RanqoBot reads its robots.txt, sitemap and llms.txt.
- **robots.txt.** Before any of the above, and before Site Access or the AI Crawler Inspector fetches a page, it reads the site's robots.txt.

### Allow or block RanqoBot

RanqoBot obeys robots.txt, including `Allow` rules, wildcards and patterns anchored to the end of a path. A group that names RanqoBot takes precedence over your `*` group. To block it entirely:

```text title="robots.txt"
User-agent: RanqoBot
Disallow: /
```

Ranqo keeps a copy of each site's robots.txt for 12 hours, so a change can take that long to apply. Refusing the User-Agent at your server or CDN takes effect at once. The [RanqoBot page](https://ranqo.ai/bot) has more examples and the full detail.

### Other ways Ranqo fetches a page

Not every request from Ranqo is RanqoBot, and a robots.txt rule for RanqoBot does not stop the others:

- **One page on request.** When someone pastes a URL into a page audit or most free tools, Ranqo fetches that single page with an ordinary browser User-Agent, follows no links, and queues nothing.
- **The free llms.txt generator** reads your homepage and sitemap under its own User-Agent, `RanqoLlmsTxtBot`.
- **Each crawler's own User-Agent.** Site Access and the free [AI Crawler Inspector](https://ranqo.ai/free-tools/ai-crawler-inspector) send GPTBot's, ClaudeBot's and the other crawlers' own User-Agents, to show how your server answers each one. These requests come from Ranqo, not from the AI vendors, and neither tool runs if your robots.txt disallows RanqoBot. xAI does not publish its crawler's User-Agent or robots.txt token, so the xAI check uses ones Ranqo assumes (`xAI`), and its result says how your server treats that string rather than xAI's real crawler.

The [RanqoBot page](https://ranqo.ai/bot) lists every User-Agent Ranqo sends.

### Ranqo's requests in Site Tracking

If your site runs [Site Tracking](https://ranqo.ai/docs/integrations/site-tracking), your install reports Ranqo's own requests like any other visit. Site Access checks your site on its own after your tracking runs, at most once every 7 days, so these requests arrive whether or not you open it.

- **Site Access and the AI Crawler Inspector** send each crawler's User-Agent from Ranqo's servers. On the [Traffic](https://ranqo.ai/docs/guides/traffic) page, those requests appear as visits from the crawlers they name, except DuckAssistBot and xAI, which Site Tracking's catalog does not list, so they count as people. Where Ranqo has a crawler's published IP ranges, the visit is marked unverified, because Ranqo's servers are not in them. The browser request each check compares against counts as a person.
- **RanqoBot and RanqoLlmsTxtBot** match no bot in the catalog either, so their requests, including RanqoBot's reads of robots.txt, sitemaps and llms.txt, count as people. So do the single-page fetches Ranqo makes with a browser User-Agent.

### If you block RanqoBot on your own site

If you are a Ranqo customer and your robots.txt disallows RanqoBot:

- **[Site Access](https://ranqo.ai/docs/guides/site-access)** reads your robots.txt and then stops before fetching any page, says so, and shows the verdict **Unknown**.
- **Page discovery** cannot follow links on your site to find pages to audit.
- **An audit of a single URL you enter** still runs, because that one fetch is not RanqoBot.

## Which crawler list belongs to which feature

| Feature | What it does with crawlers | The list it uses |
| --- | --- | --- |
| [Site Access](https://ranqo.ai/docs/guides/site-access), the free AI Crawler Inspector and the free Robots.txt Generator | Site Access and the Inspector fetch your site as each crawler and read what your robots.txt says to each; the Generator writes robots.txt rules for each | The AI crawler catalog below |
| [Site Tracking](https://ranqo.ai/docs/integrations/site-tracking), on the [Traffic](https://ranqo.ai/docs/guides/traffic) page | Classifies the real visits your server reports, by User-Agent | Site Tracking's own catalog |

The two lists overlap but are not the same, because the jobs differ. Site Access needs the crawlers that decide whether AI can read you. Site Tracking also has to recognise traffic that is not AI at all, so it is not counted as people. The lists are kept separately, so a crawler Site Access tests is not always one Site Tracking recognises: DuckAssistBot and xAI are in the first list only.

## The crawlers Site Access tests

A Site Access check fetches your homepage, and a few pages picked from your sitemap (up to 4 pages in all), as each of the 16 crawlers marked fetchable below, and as a regular browser to compare against. Some entries are permissions only: robots.txt tokens that a vendor honours but no crawler sends as a User-Agent. For those, Site Access reports what your robots.txt says, because that is the whole answer.

The User-Agent Site Access sends for each crawler is an approximation modelled on that crawler's own, and contains its token.

| Crawler | Vendor | Role | robots.txt token | Fetchable |
| --- | --- | --- | --- | --- |
| GPTBot | OpenAI | Training | GPTBot | Yes |
| ChatGPT-User | OpenAI | Assistant (live fetch) | ChatGPT-User | Yes |
| OAI-SearchBot | OpenAI | AI search index | OAI-SearchBot | Yes |
| ClaudeBot | Anthropic | Training | ClaudeBot | Yes |
| anthropic-ai | Anthropic | Training | anthropic-ai | Yes |
| Claude-User | Anthropic | Assistant (live fetch) | Claude-User | Yes |
| Google-Extended | Google | Training | Google-Extended | robots.txt token only |
| Googlebot | Google | Search engine | Googlebot | Yes |
| PerplexityBot | Perplexity | AI search index | PerplexityBot | Yes |
| Perplexity-User | Perplexity | Assistant (live fetch) | Perplexity-User | Yes |
| Applebot-Extended | Apple | Training | Applebot-Extended | robots.txt token only |
| Applebot | Apple | Search engine | Applebot | Yes |
| Bytespider | ByteDance | Training | Bytespider | Yes |
| CCBot | Common Crawl | Training | CCBot | Yes |
| Bingbot | Microsoft | Search engine | Bingbot | Yes |
| DuckAssistBot | DuckDuckGo | Assistant (live fetch) | DuckAssistBot | Yes |
| Meta-ExternalAgent | Meta | Training | Meta-ExternalAgent | Yes |
| xAI | xAI | Training | xAI | Yes |

The **Crawlers** tab in Site Access groups them by what they are for:

- **Live answers:** assistants that fetch a page while answering someone, and AI search indexes. A block here can take you out of AI answers.
- **Search:** traditional search engine crawlers.
- **Training:** crawlers that collect data to train future models. Blocking one is reported as a licensing choice, never as a blocker: AI answers cite pages through the live-answer and search crawlers.

See [Site Access](https://ranqo.ai/docs/guides/site-access) for what each result means.

## The bots Site Tracking recognises

Site Tracking checks the User-Agent of every visit your server reports against its own catalog, and puts each recognised bot in one of these groups on the Traffic page:

- **Training:** crawlers collecting data for AI model training.
- **Indexing:** AI search indexes that decide what an assistant can cite.
- **Agentic:** fetches by an AI assistant acting on someone's request.
- **Search Engine:** traditional search engine crawlers.
- **Social Preview:** link-preview generators from chat and social apps, and automated browsers such as performance audits and headless test browsers.

| Bot | Vendor | Category | IP verifiable |
| --- | --- | --- | --- |
| GPTBot | OpenAI | Training | Yes |
| OAI-SearchBot | OpenAI | Indexing | Yes |
| ChatGPT-User | OpenAI | Agentic | Yes |
| ClaudeBot | Anthropic | Training | Yes |
| Claude-User | Anthropic | Agentic | Yes |
| Claude-SearchBot | Anthropic | Indexing | Yes |
| PerplexityBot | Perplexity | Indexing | Yes |
| Perplexity-User | Perplexity | Agentic | Yes |
| Googlebot | Google | Search Engine | Yes |
| Google-Extended | Google | Training | Not applicable: robots.txt token only |
| GoogleOther | Google | Indexing | Yes |
| bingbot | Microsoft | Search Engine | Yes |
| BingPreview | Microsoft | Social Preview | Yes |
| Meta-ExternalAgent | Meta | Training | No |
| Meta-ExternalFetcher | Meta | Agentic | No |
| facebookexternalhit | Meta | Social Preview | No |
| Applebot | Apple | Search Engine | No |
| Applebot-Extended | Apple | Training | Not applicable: robots.txt token only |
| CCBot | Common Crawl | Training | No |
| Bytespider | ByteDance | Training | No |
| cohere-ai | Cohere | Training | No |
| Diffbot | Diffbot | Indexing | No |
| Amazonbot | Amazon | Search Engine | No |
| YouBot | You.com | Indexing | No |
| Slackbot | Slack | Social Preview | No |
| Discordbot | Discord | Social Preview | No |
| Twitterbot | X | Social Preview | No |
| LinkedInBot | LinkedIn | Social Preview | No |
| WhatsApp | Meta | Social Preview | No |
| TelegramBot | Telegram | Social Preview | No |
| Chrome-Lighthouse | Google | Social Preview | No |
| HeadlessChrome | Google | Social Preview | No |
| Playwright | Microsoft | Social Preview | No |
| Puppeteer | Google | Social Preview | No |
| Cypress | Cypress | Social Preview | No |

Rows marked **robots.txt token only** are permissions, not crawlers. No crawler sends them as a User-Agent (see the Site Access table above), so a visit that carries one is someone else using the name.

A visit whose User-Agent matches no bot is counted as a person: **Human (AI Referral)** when it arrived from an AI assistant's site, and **Human (Direct)** otherwise. That includes crawlers the catalog does not list, such as RanqoBot, DuckAssistBot, xAI and Googlebot-Image.

When a bot's vendor publishes the IP addresses its crawlers use, Ranqo checks each visit against them. On the Traffic page, the **Verified** column of **Agents detected** shows the share of a bot's checked visits that came from those addresses. A visit that claims to be a bot from anywhere else may be someone else using its name. Automated browsers cannot be verified this way, because anyone can run them.

See [Site Tracking](https://ranqo.ai/docs/integrations/site-tracking) for installing it and how visits are counted.
