POST
Start a durable job that scrapes many URLs with the same options. The job runs as a background job — use the returned id to poll GET /batch/scrape/{id} for status and download links, or cancel it with DELETE /batch/scrape/{id}. All fields besides urls are the same per-page controls as POST /scrape (method, body, output, geolocation, session, tag, actions, overrideParameter, prompt, schema), applied identically to every URL in the batch.

Request parameters (JSON body)

Required

  • urls (array of strings): Pages to scrape.

Optional

  • method (string, default: GET): GET, POST, PUT, DELETE.
  • body (object or string): Request body applied to every URL. Not supported with method: GET.
  • jsRender (boolean, default: false): Use a headless browser to execute JavaScript on each page.
  • onlyMainContent (boolean, default: true): Remove navigation, footers, ads, and other non-main layout content from cleaned outputs (clean_html, markdown, json), applied to every page.
  • output (array, default: ["html"]): html, clean_html, links, markdown, screenshot, screenshot_full_page, or json (screenshot and screenshot_full_page are mutually exclusive). Each requested format is returned under its own key in every page’s result.
  • geolocation (string): 2-letter ISO country code to route requests through.
  • sessionId (string): Reuse cookies/headers across the batch (also accepts alias session).
  • tag (array of strings): Arbitrary labels echoed back for tracking.
  • actions (array): Browser actions applied to every URL, same shape as POST /scrape’s actions.
  • overrideParameter (object): Per-request overrides, same shape as POST /scrape’s override_parameter.
  • prompt (string): Natural language extraction instruction, applied to every page, used when output includes json. prompt or schema is required whenever output includes json. Max 16,384 characters.
  • schema (object): JSON schema for extracted data, applied to every page, used when output includes json. prompt or schema is required whenever output includes json. Max 65,536 bytes when serialized.
  • limit (integer, default: 10000, max: 10000): Maximum URLs this job may process; also caps urls.length.

Responses

  • 202: success, id (batch scrape job ID — use this to poll status or cancel), and url (relative path to poll for status).
  • 422: Invalid batch scrape request.
  • 503: Batch scraping is unavailable or unconfigured on this orchestrator instance.

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

application/json
urls
string[]
required

Pages to scrape.

method
enum<string>
default:GET

HTTP method for every URL.

Available options:
GET,
POST,
PUT,
DELETE
body
any

Request body (JSON object or raw string), applied to every URL. Not supported with method: GET.

jsRender
boolean
default:false

Use a headless browser to execute JavaScript on each page.

proxyStrategy
string | null

Proxy cost tier to use for every URL. Omitting it escalates through the automatic set.

Example:

"auto"

onlyMainContent
boolean
default:true

Remove navigation, footers, ads, and other non-main layout content from cleaned outputs (clean_html, markdown, json), applied to every page.

output
enum<string>[]

html, clean_html, links, markdown, screenshot, screenshot_full_page, or json (screenshot and screenshot_full_page are mutually exclusive).

Available options:
html,
clean_html,
links,
markdown,
screenshot,
screenshot_full_page,
json
geolocation
string | null

2-letter ISO country code to route requests through.

sessionId
string | null

Reuse cookies/headers across the batch (also accepts alias session).

tag
string[]

Arbitrary labels echoed back for tracking.

actions
any[]

Browser actions applied to every URL, same shape as POST /scrape's actions.

timeout
integer | null

Per-page request timeout in milliseconds.

Required range: 1000 <= x <= 60000
waitSelector
string | null

CSS selector to wait for before returning a browser-rendered page.

waitMs
integer | null

Extra browser settle wait in milliseconds before returning.

Required range: 0 <= x <= 60000
overrideParameter
any

Per-request overrides, same shape as POST /scrape's override_parameter.

prompt
string | null

Natural language extraction instruction, applied to every page, used when output includes json. prompt or schema is required when output includes json. prompt is limited to 16384 characters; schema is limited to 65536 bytes when serialized.

schema
any

JSON schema for extracted data, applied to every page, used when output includes json. prompt or schema is required when output includes json. prompt is limited to 16384 characters; schema is limited to 65536 bytes when serialized.

limit
integer
default:10000

Maximum URLs this job may process; also caps urls.length.

Required range: x <= 10000

Response

Batch scrape workflow started

success
boolean
required

Whether the batch scrape job was started.

id
string
required

Batch scrape job ID. Use this to poll status (GET /batch/scrape/{id}) or cancel (DELETE /batch/scrape/{id}).

Example:

"8a1c9e2e-4b7b-4a3a-9d0a-9b2a6e6f9c22"

url
string
required

Relative path to poll for this job's status.

Example:

"/batch/scrape/8a1c9e2e-4b7b-4a3a-9d0a-9b2a6e6f9c22"