Skip to content
App

Product guides

Website scanner

Configure crawl depth, jurisdictions, authenticated pages and scan schedules.

The scanner loads your pages in a real browser, records every network request and storage write, and replays the page with and without consent to see what fires early.

Scan options

  • Crawl depth — how many links deep from the entry URL, up to five.
  • Page cap — a hard upper bound so large sites stay within your plan quota.
  • Jurisdictions — GDPR, UK GDPR, CCPA/CPRA, LGPD, PIPEDA. Findings are labelled per regime.
  • Include subdomains — expand the crawl beyond the verified host.
  • Excluded paths — glob patterns for staging routes or endless pagination.

Scanning pages behind login

Provide a test account in the site settings. The scanner signs in once per scan and never stores the session beyond the run. Use a dedicated low-privilege account, never a real customer login.

Rate limits

Scans respect robots.txt and back off automatically. Aggressive WAF rules can still block the crawler — allowlist the scanner user agent if pages return 403.

Scan lifecycle

  1. 1Queued — the run is accepted and waiting for a browser worker.
  2. 2Crawling — pages are discovered and loaded in a real browser.
  3. 3Replaying — each page is re-loaded pre-consent and post-consent to compare behaviour.
  4. 4Analysing — requests, cookies and storage writes are matched against rule packs.
  5. 5Complete — a report is published and monitoring baselines update.

A typical 50-page scan finishes in two to four minutes. Deep crawls with authentication take longer because each page is loaded twice.

Was this page helpful?