Overview & Purpose

POST /crawl is the “get me the whole site” endpoint: give it a starting URL, and it discovers linked pages, follows them outward, and scrapes each one it visits — all as a single, durable background job. Use it when you want content from many pages and are happy to let pline discover them for you — a documentation site, a blog archive, a whole product catalog you don’t already have URLs for. If you only need the URL list, use Map — much cheaper. If you already have the exact URLs, use Batch scrape instead — it skips discovery entirely. Prerequisites: a valid API key. No special scope required.

Best practices

  • Unlike /scrape, a crawl doesn’t hand back a result immediately. POST /crawl returns a job id; discovery runs the same sitemap-first logic as Map, then scraping follows links found on each visited page, expanding outward. Poll GET /crawl/{id} for status until completed, or receive results via webhook/email instead — see Delivery.
  • Scope it deliberately. crawlEntireDomain and allowSubdomains can expand a crawl far beyond what you intended — start narrow (includePaths/excludePaths) and widen only if needed.
  • limit caps total pages crawled — set it explicitly for predictable cost on large sites.
  • Choose your output up front. output controls which fields (markdown, html, clean_html, links, screenshot) are populated per page — see Output. Requesting only what you need keeps jobs faster and cheaper.
  • Don’t poll aggressively. A crawl can run for minutes to hours depending on scope; poll on a reasonable interval, or better, use a webhook for the completed/failed events instead.

Practical Implementation Example

Scenario: crawl a documentation site for its Markdown content, and get notified by webhook instead of polling.

Error codes