POST
Start a durable, multi-page crawl beginning at url. The crawl runs as a background job — use the returned id to poll GET /crawl/{id} for status and results, or cancel it with DELETE /crawl/{id}. By default the crawl discovers URLs via sitemap first (same semantics as POST /map) and then follows links found on each page, scoped to the starting URL’s path unless crawlEntireDomain or allowSubdomains is set.

Request parameters (JSON body)

Required

  • url (string): Starting URL for the crawl.

Optional

  • email (string): Address to notify by email once the crawl completes. See Delivery.
  • includePaths (array of strings): Only crawl URLs whose path matches one of these regex patterns.
  • excludePaths (array of strings): Skip URLs whose path matches one of these regex patterns.
  • maxDiscoveryDepth (integer): Maximum link-following depth from the starting URL. Unlimited when omitted.
  • sitemap (string, default: include): include, skip, only — same semantics as POST /map’s sitemap option.
  • ignoreQueryParameters (boolean, default: false): Treat URLs that differ only by query string as duplicates.
  • limit (integer, default: 10000, max: 10000): Maximum number of pages to crawl.
  • crawlEntireDomain (boolean, default: false): Follow links anywhere on the domain instead of only below the starting URL’s path.
  • allowSubdomains (boolean, default: false): Follow links on subdomains of the starting URL’s registrable domain.
  • output (object): Which fields to populate on each crawled document — markdown, html (raw), clean_html (boilerplate stripped), links, screenshot (all boolean, all default false). Defaults to markdown only when output is omitted entirely.
  • webhook (object): Notify an endpoint on crawl lifecycle events in real time. See Delivery for the payload shape and event types.
    • url (string, required): Webhook endpoint to notify. Must be HTTPS, point to a public host (no localhost or private IPs), and not contain credentials.
    • headers (object): Extra headers to send with each webhook call.
    • metadata (any): Arbitrary JSON echoed back in every webhook payload.
    • events (array): started, page, completed, failed. Empty means all events.
  • scrape (object): Per-page POST /scrape options merged into every page request (also accepts the alias scrapeOptions). url, method, and body are ignored — the crawler sets these per page.
  • scheduleToCloseTimeoutSecs (integer, default: 86400): Overall time budget for the crawl job, in seconds, before it’s forced to fail.
  • heartbeatTimeoutSecs (integer, default: 30): Maximum time without progress, in seconds, before the crawl is considered stalled and retried.

Responses

  • 202: success, id (crawl job ID — use this to poll status or cancel), and url (relative path to poll for status).
  • 422: Invalid crawl request.
  • 503: Crawling is unavailable or unconfigured on this orchestrator instance.

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

application/json
url
string
required

Starting URL for the crawl.

Example:

"https://example.com/docs"

email
string | null

Address to notify by email once the crawl completes. Must be a valid email address.

includePaths
string[]

Only crawl URLs whose path matches one of these regex patterns.

excludePaths
string[]

Skip URLs whose path matches one of these regex patterns.

maxDiscoveryDepth
integer | null

Maximum link-following depth from the starting URL. Unlimited when omitted.

sitemap
enum<string>
default:include

Sitemap handling for initial URL discovery, same semantics as POST /map.

Available options:
skip,
include,
only
ignoreQueryParameters
boolean
default:false

Whether URLs with query parameters should be treated as duplicates.

limit
integer
default:10000

Maximum number of pages to crawl.

Required range: x <= 10000
crawlEntireDomain
boolean
default:false

Follow links anywhere on the domain instead of only below the starting URL's path.

allowSubdomains
boolean
default:false

Follow links on subdomains of the starting URL's registrable domain.

output
CrawlOutput · object | null

Which fields to populate on each crawled document. Defaults to markdown only when omitted.

webhook
CrawlWebhook · object | null

Optional webhook notified on crawl lifecycle events.

scrape
Scrape · object

Per-page POST /scrape options merged into every page request (accepts alias scrapeOptions). url, method, and body are ignored — the crawler sets these per page.

scheduleToCloseTimeoutSecs
integer
default:86400

Overall time budget for the crawl job, in seconds, before it is forced to fail.

heartbeatTimeoutSecs
integer
default:30

Maximum time without progress, in seconds, before the crawl is considered stalled and retried.

Response

Crawl workflow started

success
boolean
required

Whether the crawl job was started.

id
string
required

Crawl job ID. Use this to poll status (GET /crawl/{id}) or cancel (DELETE /crawl/{id}).

Example:

"3f2b9e2e-6b7b-4a3a-9d0a-9b2a6e6f9c11"

url
string
required

Relative path to poll for this crawl's status.

Example:

"/crawl/3f2b9e2e-6b7b-4a3a-9d0a-9b2a6e6f9c11"