Overview & Purpose
POST /map answers one question fast: “what URLs exist on this site?” It returns URLs from
the site’s sitemap — a flat, categorized list of discovered links, no page content, no HTML — and
is the cheapest, quickest way to understand a site’s structure before deciding what to actually
scrape.
The sitemap parameter controls how the sitemap is used: include (the default) returns URLs
found in the sitemap, skip skips the sitemap, and only returns URLs from the sitemap
exclusively.
Use it before Batch scrape when you need the URL list first, for
quick SEO/content audits, or to check how many pages of a given type exist before committing to a
full Crawl.
Prerequisites: a valid API key. No special scope required.
Best practices
- A well-formed sitemap means fast results. A site with a proper sitemap returns results almost instantly.
- Filter by
typein the response instead of guessing from the URL string — categorization is done for you. - Leave
includeSubdomainsoff unless you specifically need links from subdomains too — it’s off by default to keep results scoped to the exact host you asked for. - Turn on
ignoreQueryParameterswhen a site encodes the same page under many query strings (tracking params, filters, pagination) and you only want one entry per distinct page. - Set
timeoutif the default is too short for a large or slow site — otherwise leave it unset and let pline use its default.
Note: ignoreCache bypasses cached discovery results. Leave it off by default — repeated
requests for the same site benefit from the cache.
