Overview & Purpose

POST /map answers one question fast: “what URLs exist on this site?” It returns URLs from the site’s sitemap — a flat, categorized list of discovered links, no page content, no HTML — and is the cheapest, quickest way to understand a site’s structure before deciding what to actually scrape. The sitemap parameter controls how the sitemap is used: include (the default) returns URLs found in the sitemap, skip skips the sitemap, and only returns URLs from the sitemap exclusively. Use it before Batch scrape when you need the URL list first, for quick SEO/content audits, or to check how many pages of a given type exist before committing to a full Crawl. Prerequisites: a valid API key. No special scope required.

Best practices

  • A well-formed sitemap means fast results. A site with a proper sitemap returns results almost instantly.
  • Filter by type in the response instead of guessing from the URL string — categorization is done for you.
  • Leave includeSubdomains off unless you specifically need links from subdomains too — it’s off by default to keep results scoped to the exact host you asked for.
  • Turn on ignoreQueryParameters when a site encodes the same page under many query strings (tracking params, filters, pagination) and you only want one entry per distinct page.
  • Set timeout if the default is too short for a large or slow site — otherwise leave it unset and let pline use its default.
Note: ignoreCache bypasses cached discovery results. Leave it off by default — repeated requests for the same site benefit from the cache.

Practical Implementation Example

Scenario: discover every product page on a site, then hand the filtered list to Batch scrape.

Error codes