POST /crawl
Start a durable, multi-page crawl of a website.
url. The crawl runs as a background job — use
the returned id to poll GET /crawl/{id} for status and results, or cancel
it with DELETE /crawl/{id}.
By default the crawl discovers URLs via sitemap first (same semantics as POST /map) and then
follows links found on each page, scoped to the starting URL’s path unless crawlEntireDomain or
allowSubdomains is set.
Request parameters (JSON body)
Required
url(string): Starting URL for the crawl.
Optional
email(string): Address to notify by email once the crawl completes. See Delivery.includePaths(array of strings): Only crawl URLs whose path matches one of these regex patterns.excludePaths(array of strings): Skip URLs whose path matches one of these regex patterns.maxDiscoveryDepth(integer): Maximum link-following depth from the starting URL. Unlimited when omitted.sitemap(string, default:include):include,skip,only— same semantics asPOST /map’ssitemapoption.ignoreQueryParameters(boolean, default:false): Treat URLs that differ only by query string as duplicates.limit(integer, default:10000, max:10000): Maximum number of pages to crawl.crawlEntireDomain(boolean, default:false): Follow links anywhere on the domain instead of only below the starting URL’s path.allowSubdomains(boolean, default:false): Follow links on subdomains of the starting URL’s registrable domain.output(object): Which fields to populate on each crawled document —markdown,html(raw),clean_html(boilerplate stripped),links,screenshot(all boolean, all defaultfalse). Defaults tomarkdownonly whenoutputis omitted entirely.webhook(object): Notify an endpoint on crawl lifecycle events in real time. See Delivery for the payload shape and event types.url(string, required): Webhook endpoint to notify. Must be HTTPS, point to a public host (no localhost or private IPs), and not contain credentials.headers(object): Extra headers to send with each webhook call.metadata(any): Arbitrary JSON echoed back in every webhook payload.events(array):started,page,completed,failed. Empty means all events.
scrape(object): Per-pagePOST /scrapeoptions merged into every page request (also accepts the aliasscrapeOptions).url,method, andbodyare ignored — the crawler sets these per page.scheduleToCloseTimeoutSecs(integer, default:86400): Overall time budget for the crawl job, in seconds, before it’s forced to fail.heartbeatTimeoutSecs(integer, default:30): Maximum time without progress, in seconds, before the crawl is considered stalled and retried.
Responses
- 202:
success,id(crawl job ID — use this to poll status or cancel), andurl(relative path to poll for status). - 422: Invalid crawl request.
- 503: Crawling is unavailable or unconfigured on this orchestrator instance.
Authorizations
Bearer authentication header of the form Bearer <token>, where <token> is your auth token.
Body
Starting URL for the crawl.
"https://example.com/docs"
Address to notify by email once the crawl completes. Must be a valid email address.
Only crawl URLs whose path matches one of these regex patterns.
Skip URLs whose path matches one of these regex patterns.
Maximum link-following depth from the starting URL. Unlimited when omitted.
Sitemap handling for initial URL discovery, same semantics as POST /map.
skip, include, only Whether URLs with query parameters should be treated as duplicates.
Maximum number of pages to crawl.
x <= 10000Follow links anywhere on the domain instead of only below the starting URL's path.
Follow links on subdomains of the starting URL's registrable domain.
Which fields to populate on each crawled document. Defaults to markdown only when omitted.
Optional webhook notified on crawl lifecycle events.
Per-page POST /scrape options merged into every page request (accepts alias scrapeOptions). url, method, and body are ignored — the crawler sets these per page.
Overall time budget for the crawl job, in seconds, before it is forced to fail.
Maximum time without progress, in seconds, before the crawl is considered stalled and retried.
Response
Crawl workflow started
Whether the crawl job was started.
Crawl job ID. Use this to poll status (GET /crawl/{id}) or cancel (DELETE /crawl/{id}).
"3f2b9e2e-6b7b-4a3a-9d0a-9b2a6e6f9c11"
Relative path to poll for this crawl's status.
"/crawl/3f2b9e2e-6b7b-4a3a-9d0a-9b2a6e6f9c11"
