POST
RAG pipeline
Turn a documentation site, help center, or blog archive into chunked, embeddable text for a RAG pipeline — without writing a crawler or an HTML parser.

Use case

  • Build or refresh a vector store from a docs site, knowledge base, or public wiki
  • Keep an internal assistant’s context current by re-running the crawl on a schedule
  • Skip boilerplate (nav, footers, ads) so embeddings aren’t diluted with layout noise

Step 1: Crawl the site into Markdown

Crawl discovers every linked page from a starting URL and scrapes each one — request markdown output so every page comes back ready to chunk, with navigation and footers already stripped (only_main_content default).

Step 2: Poll until the crawl completes

Python

Step 3: Chunk and embed each page

Download each page’s Markdown from s3PresignedUrls, split it into overlapping chunks (by heading or a fixed token window), and embed each chunk with your model of choice before upserting into your vector store.
Python

Request highlights

Tips

  • Re-run the crawl on a schedule (daily or weekly) and re-embed only pages whose content hash changed, instead of rebuilding the whole index every time.
  • Use webhook on the crawl request instead of polling if your pipeline runs as a background worker — see Delivery.
  • Keep tag set per knowledge source so you can trace crawl cost and history per source in Request history.