Product guides
Website scanner
Configure crawl depth, jurisdictions, authenticated pages and scan schedules.
The scanner loads your pages in a real browser, records every network request and storage write, and replays the page with and without consent to see what fires early.
Scan options
- Crawl depth — how many links deep from the entry URL, up to five.
- Page cap — a hard upper bound so large sites stay within your plan quota.
- Jurisdictions — GDPR, UK GDPR, CCPA/CPRA, LGPD, PIPEDA. Findings are labelled per regime.
- Include subdomains — expand the crawl beyond the verified host.
- Excluded paths — glob patterns for staging routes or endless pagination.
Scanning pages behind login
Provide a test account in the site settings. The scanner signs in once per scan and never stores the session beyond the run. Use a dedicated low-privilege account, never a real customer login.
Rate limits
Scans respect robots.txt and back off automatically. Aggressive WAF rules can still block the crawler — allowlist the scanner user agent if pages return 403.
Scan lifecycle
- 1Queued — the run is accepted and waiting for a browser worker.
- 2Crawling — pages are discovered and loaded in a real browser.
- 3Replaying — each page is re-loaded pre-consent and post-consent to compare behaviour.
- 4Analysing — requests, cookies and storage writes are matched against rule packs.
- 5Complete — a report is published and monitoring baselines update.
A typical 50-page scan finishes in two to four minutes. Deep crawls with authentication take longer because each page is loaded twice.
Was this page helpful?
