POST /batch/scrape
Start a durable batch scrape of many URLs with the same options.
id to poll GET /batch/scrape/{id} for
status and download links, or cancel it with
DELETE /batch/scrape/{id}.
All fields besides urls are the same per-page controls as POST /scrape (method, body, output,
geolocation, session, tag, actions, overrideParameter, prompt, schema), applied identically to
every URL in the batch.
Request parameters (JSON body)
Required
urls(array of strings): Pages to scrape.
Optional
method(string, default:GET):GET,POST,PUT,DELETE.body(object or string): Request body applied to every URL. Not supported withmethod: GET.jsRender(boolean, default:false): Use a headless browser to execute JavaScript on each page.onlyMainContent(boolean, default:true): Remove navigation, footers, ads, and other non-main layout content from cleaned outputs (clean_html,markdown,json), applied to every page.output(array, default:["html"]):html,clean_html,links,markdown,screenshot,screenshot_full_page, orjson(screenshot and screenshot_full_page are mutually exclusive). Each requested format is returned under its own key in every page’s result.geolocation(string): 2-letter ISO country code to route requests through.sessionId(string): Reuse cookies/headers across the batch (also accepts aliassession).tag(array of strings): Arbitrary labels echoed back for tracking.actions(array): Browser actions applied to every URL, same shape asPOST /scrape’sactions.overrideParameter(object): Per-request overrides, same shape asPOST /scrape’soverride_parameter.prompt(string): Natural language extraction instruction, applied to every page, used whenoutputincludesjson.promptorschemais required wheneveroutputincludesjson. Max 16,384 characters.schema(object): JSON schema for extracted data, applied to every page, used whenoutputincludesjson.promptorschemais required wheneveroutputincludesjson. Max 65,536 bytes when serialized.limit(integer, default:10000, max:10000): Maximum URLs this job may process; also capsurls.length.
Responses
- 202:
success,id(batch scrape job ID — use this to poll status or cancel), andurl(relative path to poll for status). - 422: Invalid batch scrape request.
- 503: Batch scraping is unavailable or unconfigured on this orchestrator instance.
Authorizations
Bearer authentication header of the form Bearer <token>, where <token> is your auth token.
Body
Pages to scrape.
HTTP method for every URL.
GET, POST, PUT, DELETE Request body (JSON object or raw string), applied to every URL. Not supported with method: GET.
Use a headless browser to execute JavaScript on each page.
Proxy cost tier to use for every URL. Omitting it escalates through the automatic set.
"auto"
Remove navigation, footers, ads, and other non-main layout content from cleaned outputs (clean_html, markdown, json), applied to every page.
html, clean_html, links, markdown, screenshot, screenshot_full_page, or json (screenshot and screenshot_full_page are mutually exclusive).
html, clean_html, links, markdown, screenshot, screenshot_full_page, json 2-letter ISO country code to route requests through.
Reuse cookies/headers across the batch (also accepts alias session).
Arbitrary labels echoed back for tracking.
Browser actions applied to every URL, same shape as POST /scrape's actions.
Per-page request timeout in milliseconds.
1000 <= x <= 60000CSS selector to wait for before returning a browser-rendered page.
Extra browser settle wait in milliseconds before returning.
0 <= x <= 60000Per-request overrides, same shape as POST /scrape's override_parameter.
Natural language extraction instruction, applied to every page, used when output includes json. prompt or schema is required when output includes json. prompt is limited to 16384 characters; schema is limited to 65536 bytes when serialized.
JSON schema for extracted data, applied to every page, used when output includes json. prompt or schema is required when output includes json. prompt is limited to 16384 characters; schema is limited to 65536 bytes when serialized.
Maximum URLs this job may process; also caps urls.length.
x <= 10000Response
Batch scrape workflow started
Whether the batch scrape job was started.
Batch scrape job ID. Use this to poll status (GET /batch/scrape/{id}) or cancel (DELETE /batch/scrape/{id}).
"8a1c9e2e-4b7b-4a3a-9d0a-9b2a6e6f9c22"
Relative path to poll for this job's status.
"/batch/scrape/8a1c9e2e-4b7b-4a3a-9d0a-9b2a6e6f9c22"
