RAG pipeline
Crawl a knowledge base into clean Markdown and feed it into a retrieval-augmented generation pipeline.
POST
RAG pipeline
Turn a documentation site, help center, or blog archive into chunked, embeddable text for a RAG
pipeline — without writing a crawler or an HTML parser.
Use case
- Build or refresh a vector store from a docs site, knowledge base, or public wiki
- Keep an internal assistant’s context current by re-running the crawl on a schedule
- Skip boilerplate (nav, footers, ads) so embeddings aren’t diluted with layout noise
Step 1: Crawl the site into Markdown
Crawl discovers every linked page from a starting URL and scrapes each one — requestmarkdown output so every page comes back ready to chunk, with navigation and
footers already stripped (only_main_content default).
Step 2: Poll until the crawl completes
Python
Step 3: Chunk and embed each page
Download each page’s Markdown froms3PresignedUrls, split it into overlapping chunks (by
heading or a fixed token window), and embed each chunk with your model of choice before upserting
into your vector store.
Python
Request highlights
Tips
- Re-run the crawl on a schedule (daily or weekly) and re-embed only pages whose content hash changed, instead of rebuilding the whole index every time.
- Use
webhookon the crawl request instead of polling if your pipeline runs as a background worker — see Delivery. - Keep
tagset per knowledge source so you can trace crawl cost and history per source in Request history.
