# Site Access

> Check whether AI crawlers and search engines can actually fetch your site, what your server tells each one, and what to fix when one is blocked.

Source: https://ranqo.ai/docs/guides/site-access

Site Access answers one question: when an AI engine sends its crawler to your site, does it get the page? It is the one measurement in Ranqo that can see a rule on your server or CDN that turns crawlers away. Open it from **Site Access** in the Analyze section of the sidebar.

## What a check does

A check fetches your homepage once as a regular browser and once as each of the 16 crawlers in Ranqo's catalog, sending each crawler's own User-Agent. Then it compares what each crawler got with what the browser got.

- **The browser is the control.** It is what separates "this site blocks crawlers" from "this site was down". Without a browser answer, nothing is reported as a block.
- **Requests are spaced out, never sent at once**, and anything other than a 200 is asked again a few seconds later before it is recorded. Firing every request together would trip rate limits and report a block that Ranqo caused.
- **Deeper pages are checked too.** Up to 4 more pages, picked one per section from your sitemap, are fetched as every crawler. A refusal on the homepage that a deeper page contradicts is reported as unknown, not as a site-wide block.
- **The files that govern crawlers are read fresh:** robots.txt (and what it says to each crawler for your homepage), your sitemap (whether it is reachable and declared in robots.txt), and llms.txt (whether a real file is served, not an HTML page at that address).

A check takes about a minute. See [AI crawlers](https://ranqo.ai/docs/methodology/crawlers) for the full catalog and what each crawler is for.

**If your robots.txt disallows RanqoBot:**
Ranqo reads your robots.txt as RanqoBot before it fetches anything. If your rules disallow RanqoBot, the check stops there and says so, and no crawler is fetched. See [RanqoBot](https://ranqo.ai/bot) for what it fetches and how to allow it.

## The verdict: a state, not a score

Site Access deliberately has no 0 to 100 score. Measured across live sites, results cluster at the two ends, completely clear or hard blocked, and a score would blur the only distinction that matters. The banner at the top states one of four verdicts:

| Verdict | What it means |
|---|---|
| Clear | Every AI crawler and Googlebot can fetch your site. Nothing is in the way. |
| Worth a look | Nothing is certainly blocked, but something did not go the way it did for a browser, or a file needs attention. |
| Blocked | Something is stopping AI engines from reading your site. |
| Unknown | Ranqo could not load the site as a browser, or robots.txt disallows RanqoBot, so there is no verdict to give. |

Under the verdict, **Blockers** and **Worth a look** count the findings of each kind, and **Requests made** is how many fetches the check sent.

## Blockers, warnings and notices

Each finding has a severity. Only blockers make the verdict **Blocked**.

**Blockers** are anything that removes you from AI answers:

- A crawler that fetches pages to answer questions right now (a live-answer or AI search crawler) was refused by your server while a browser got the page. When Googlebot and Bingbot got in at the same time, the refusal is a rule about who is asking, usually a CDN or firewall bot filter.
- Your robots.txt disallows a live-answer, AI search or search engine crawler.

**Worth a look** covers results Ranqo cannot call a block with certainty:

- Your server refused every non-browser client, search engines included. Real crawlers arrive from IP ranges their vendors publish, and a filter that checks those may still let them in, while Ranqo sends each User-Agent from its own servers. Check your CDN's bot settings directly.
- Search engines were refused while a browser got the page (for the same reason).
- A training crawler was refused by your server. Worth confirming it is a decision and not a side effect.
- Your server answered 429 twice, which says more about request rate than policy.
- A request timed out or returned a server error, so that crawler is unmeasured, not clear.
- A crawler received a much thinner page than the browser did.
- robots.txt could not be read, or no sitemap was reachable.
- The check ran out of time before it reached every crawler, so the rest are unmeasured.
- Ranqo could not load your homepage as a browser (the verdict is then **Unknown**).

**Notices** are hygiene. They are listed but do not change the verdict: robots.txt disallowing training crawlers, a sitemap not declared in robots.txt, and a missing llms.txt.

**Training crawlers are a licensing choice:**
Crawlers that build training data are not how AI answers cite you today, and many publishers block them on purpose. Blocking one is reported, never counted as a blocker. What matters for citations is the live-answer and AI search crawlers.

## The six tabs

- **Overview.** Four cards (**Crawler fetch**, **robots.txt**, **llms.txt**, **Sitemap**), each opening its tab, then **What we found**: every finding with an explanation and a button to open the tab and crawler group it concerns. When your sitemap is not declared in robots.txt and Ranqo knows its address, the finding carries a **Copy fix** button with the exact `Sitemap:` line to add. On a clear site, the panel says so and points you to the Action Center, because the visibility gap is somewhere else.
- **Crawlers.** **Who can fetch your site**: one row per crawler with its status, response time and word count beside the browser's, filterable by **All**, **Live answers**, **Search** and **Training**. **Beyond the homepage** shows the deeper pages, and **Edge** uses response timing to say whether a refusal came from your CDN edge or your application, when the timing is clear enough to tell.
- **robots.txt.** Your file as fetched, annotated line by line, with **What it tells each kind of crawler** for your homepage.
- **llms.txt.** The file Ranqo found and the checks it ran on it. No engine has published that it reads llms.txt, so this is always a notice.
- **Sitemap.** Whether it is reachable and declared, what is in it, and the sample pages the check fetched.
- **History.** One row per completed check, with the verdict and a note on what changed since the previous one. Up to 26 checks are kept.

## When checks run

Site Access runs on its own after one of your brand's tracking runs, at most once every 7 days. To check again after a fix, press **Re-check** at the top of the page.

Re-checks have a cooldown of 60 minutes per brand, and only one check can run against a domain at a time. That is a limit on how often Ranqo hits your server, not a plan allowance: Site Access does not count toward any monthly limit.

When a check finds a blocker, it also appears in **This week's signals** on your Home page.

## What to do with a blocker

1. Open the finding's tab to see exactly which crawlers were refused and what your server answered.
2. If robots.txt is the cause, edit the rule for that crawler. If your server or CDN refused it, look for a bot-management or firewall rule that matches the crawler's User-Agent.
3. Press **Re-check** once the change is live, and confirm the finding is gone.

Clearing a blocker is a precondition for being read at all. Everything else on this page is hygiene, and Ranqo does not claim it will move your visibility.
