Overview & Purpose
POST /crawl is the “get me the whole site” endpoint: give it a starting URL, and it discovers
linked pages, follows them outward, and scrapes each one it visits — all as a single, durable
background job.
Use it when you want content from many pages and are happy to let pline discover them for
you — a documentation site, a blog archive, a whole product catalog you don’t already have URLs
for. If you only need the URL list, use Map — much cheaper. If you
already have the exact URLs, use Batch scrape instead — it
skips discovery entirely.
Prerequisites: a valid API key. No special scope required.
Best practices
- Unlike
/scrape, a crawl doesn’t hand back a result immediately.POST /crawlreturns a jobid; discovery runs the same sitemap-first logic as Map, then scraping follows links found on each visited page, expanding outward. PollGET /crawl/{id}for status untilcompleted, or receive results via webhook/email instead — see Delivery. - Scope it deliberately.
crawlEntireDomainandallowSubdomainscan expand a crawl far beyond what you intended — start narrow (includePaths/excludePaths) and widen only if needed. limitcaps total pages crawled — set it explicitly for predictable cost on large sites.- Choose your output up front.
outputcontrols which fields (markdown,html,clean_html,links,screenshot) are populated per page — see Output. Requesting only what you need keeps jobs faster and cheaper. - Don’t poll aggressively. A crawl can run for minutes to hours depending on scope; poll on a
reasonable interval, or better, use a webhook for the
completed/failedevents instead.
