Overview & Purpose

POST /batch/scrape is for when you already know exactly which pages you want — a list of product URLs, a set of article links, output from Map — and want them all scraped with the same settings in one durable job. Use it when you’re scraping a fixed, known set of URLs. There’s no link discovery involved: you’re not asking pline to find pages, just to fetch a list you already have, consistently and at scale. If you don’t have the URL list yet, get it from Map first. If you want pline to discover and scrape pages by following links itself, use Crawl instead. Prerequisites: a valid API key. No special scope required.

Best practices

  • Same job lifecycle as Crawl, but list-driven. POST /batch/scrape returns a job id; every URL in urls is scraped with the identical options you supplied (output, proxyStrategy, actions, etc.). Poll GET /batch/scrape/{id} until status is completed, then use s3PresignedUrls to download the results.
  • limit also caps urls.length — the request is rejected if you supply more URLs than the limit allows.
  • JSON extraction (prompt/schema) works across the whole batch the same way it does on /scrape — see Output.
  • Large batches take time. Don’t poll aggressively; check status on a reasonable interval.

Practical Implementation Example

Scenario: you’ve already mapped a site’s product URLs; now scrape all of them for structured product data in one job.

Error codes