How to run a site audit with an AI agent

Have the agent start a crawl with run_site_audit, poll get_audit_status until it reports completed, then read the issues critical first with get_audit_issues. Group them by the fix they need, fix templates rather than single pages, and re-crawl to confirm. A coding agent such as Claude Code or Cursor can make the fixes in the same session.

Updated 27 September 2026

In short

  • A site audit runs in the background. Nothing the agent says about issues counts until the status is completed.
  • Triage by severity, then by cause. Forty pages missing a meta description are usually one template fix.
  • blocked-page means bot protection stopped the crawler, not that the page is broken. Allowlist the IndexZero-Audit user agent and run it again.
  • Lighthouse scores come from a lab run on a sample. Use them to find slow templates, not as a measure of what visitors experience.
  • Re-crawl after the fixes ship. The second audit is the evidence the fixes worked.

What the crawl checks

The crawler starts from your domain, reads robots.txt and obeys it for its own user agent (IndexZero-Audit/1.0), seeds itself from the sitemaps robots.txt lists, and stays on your site's origin. It checks every page it reaches against 27 issue types in three severities:

  • Critical (4): crawler was blocked, server error (5xx), broken internal link, missing title tag.
  • Warning (14): page returns an error (4xx), duplicate title, duplicate meta description, duplicate page content, missing meta description, missing H1 heading, multiple H1 headings, redirect chain, redirect loop, conflicting canonical signals, thin content, images missing alt text, orphan page, page has no outgoing links.
  • Info (9): title too long, title too short, meta description too long, meta description too short, heading levels skip, slow server response, page is noindex, canonicalized to another URL, page is deep in the site structure.

It also runs Lighthouse, on mobile and desktop, for up to 10 pages: the homepage plus one page per URL template. That is up to 20 paid runs, and each returns performance, accessibility, best-practices and SEO scores along with LCP, CLS, INP and TTFB.

Lighthouse is lab data: one page load under fixed network and device conditions. Real visitors bring every kind of device and connection, which is why lab and field data can differ. Read a poor lab score as a pointer to a slow template, then confirm it before you spend a week on it.

Each plan caps how many pages one audit may crawl: Free 50, Starter 1,000, Pro 5,000, Scale 10,000. A larger budget is trimmed to the cap rather than refused, and the response says so.

Before you start

Two things can leave an audit with nothing to read, and both are quicker to check by hand.

First, robots.txt. The crawler obeys it, so a Disallow: / in the group that applies to unknown bots leaves it with nothing to read. The free robots.txt tester fetches your live file and shows which rule decides each path.

Second, bot protection. If the site sits behind Cloudflare or a similar firewall, a bot challenge stops the crawler. Pages that answer with a challenge, a 403 or a 429 are reported as blocked-page rather than broken. On Cloudflare, a WAF custom rule that skips bot protection when the user agent contains IndexZero-Audit lets it through. Do this before a large run, or the audit will mostly measure your firewall.

Step 1: start the crawl

Check my credit balance, then run a site audit of example.com with the default page budget. Tell me the audit id and keep polling until it finishes. Don't report any issues until the status is completed.

Behind that prompt the agent calls whoami, then run_site_audit. The default budget is 50 pages. maxPages sets a different budget, from 10 up to your plan's cap (on Free, the default is already the cap). lighthouse: false skips the Lighthouse sample and its paid runs. startUrl starts somewhere other than the homepage, and redirects are followed first, so an apex domain that redirects to www is crawled on www.

The crawl runs on IndexZero's own crawler. What costs credits is the Lighthouse sample: up to 20 runs, however many pages you crawl, so with lighthouse: false the audit spends none. Reading its status, issues and pages afterwards is free, however many times the agent does it.

Step 2: wait for completed

run_site_audit returns straight away with an audit id. The crawl then moves through discovery, crawling, Lighthouse and finalizing, and get_audit_status reports the phase and how many pages it has crawled so far. Most crawls take minutes. A large site with a slow server takes longer.

Issues read mid-crawl are partial, and an agent that sees an empty list may tell you the site is clean. The tool descriptions warn against it, and the prompt above says it again, which is worth the extra sentence.

If the audit fails, get_audit_status returns an error code and a detail line. A timeout usually means the site responded too slowly, and a smaller page budget is the usual fix. The pages crawled before the failure, and the issues found on them, stay readable.

Step 3: triage by severity, then by cause

Read critical issues first by passing severity: "critical" to get_audit_issues, then warnings, then info. Each row carries the issue type, the URL and a detail field. The tool returns 200 rows by default and up to 1,000, so on a large site ask for one severity or one issue type at a time.

Read the critical issues from the latest audit, then the warnings. Group them by the fix they need rather than by page, with the number of pages each fix covers. For each group, tell me whether it looks like a template problem or a one-off.

Grouping by cause saves the most work. Forty missing-meta-description rows on blog posts usually mean the blog template has no description field. Ten duplicate-title rows on paginated category pages usually mean the page number isn't in the title template. Fix the template and the whole group goes.

Issues that need a second look

  • noindex-page and canonicalized-page are info, not errors. Login, thank-you and filtered listing pages are often meant to stay out of the index. Scan the list for a page that should rank.
  • orphan-page means no crawled page links to the URL, and the crawler found it only in the sitemap. Link to it from a relevant page, or take it out of the sitemap.
  • thin-content on a site that renders in the browser can mean the text isn't in the server-rendered HTML at all, so crawlers that read plain HTML see very little.
  • blocked-page is an access problem, not a page problem. Fix access and re-run before reading anything else.

get_audit_pages lists what the crawl reached, with status code, title, word count and H1 count. Compare it with the pages you expect. A whole section missing from the list is a finding in its own right.

Step 4: fix it in your codebase

IndexZero doesn't change your site. If the agent has your code open, as Claude Code, Cursor and Codex do, it can make the fixes there. Every issue type comes with an explanation and a documented fix (the full list is in the docs), and get_audit_issues returns only the type, URL and detail of each issue. Point the agent at that page, or paste the fix for the group you're working on.

Take the missing meta description group. Find the template that renders those pages in this repo, add a description field, and write descriptions for the ten pages with the most words. Show me the diff before committing.

Review the diffs as you would a colleague's. Titles and descriptions are copy, and copy should sound like you. The free meta tag checker shows how one page's title and description may look in Google once the change is live.

Keep the order the triage gave you: critical first, and within a severity, the fixes that cover the most pages.

Step 5: re-crawl and compare

Deploy, then run the audit again with the same budget. get_audit_status and get_audit_issues default to the latest audit, so keep the first audit's id, in the chat or in a saved report, to compare against:

Run a new audit of example.com with the same page budget as audit [previous audit id]. When it completes, compare the issue counts by type with that audit, reading each type separately so no count is cut off at the row limit, and list any issue type that is new.

An issue that survives the fix usually points at a second template or a cache. A new issue deserves a look before anything else, because fixes can break things.

For a quick check of one page between full audits, the free SEO checker runs the audit's on-page checks against a single URL with no account.

Make it a routine

Audits don't run on a schedule. You or the agent start them. A monthly re-crawl, or one after each large release, catches regressions while they are still small.

End each round with a saved report, so the record of what was found and fixed lives on the project rather than in a chat that scrolls away:

Save a report on the project for this audit: the audit id, issue counts by severity, what we fixed, and what's left for next time.

The audit is one of four workflows in SEO with AI agents, which covers connecting an agent, keeping its spending in check and checking its answers.

Frequently asked questions

How long does a site audit take?
Minutes for most sites. It depends on the page budget and how fast the server answers: the crawler slows down when pages are slow, blocked or erroring, and speeds up when the site responds well.
Does the crawler follow robots.txt?
Yes. It reads robots.txt for its user agent, IndexZero-Audit/1.0, obeys it, and starts from the sitemaps it lists.
What does an audit cost?
Only the Lighthouse sample costs credits: up to 20 runs, however many pages you crawl. The crawl itself runs on IndexZero's own crawler, so with lighthouse: false an audit spends no credits. Reading the results is free, and your plan caps the page budget.
Why are some pages reported as blocked?
Bot protection answered the crawler with a challenge, a 403 or a 429 instead of the page. The page may be fine for visitors. Allowlist the IndexZero-Audit user agent in your firewall and run the audit again.

Run it with your agent

Connect an agent to IndexZero

Every step in these guides is a tool call an agent can make. Connect Claude, ChatGPT, Cursor or Codex to IndexZero's MCP server and paste the prompts as they are. The free plan includes the MCP server and 2,000 credits a month.

Check one page first, free

More in this guide