RAG is a pipeline, not a single retrieval call
Retrieval-augmented generation gets demoed as one function call: embed the query, search an index, stuff the top results into a prompt. At scale -- millions of source documents, thousands of queries per minute, content that changes daily -- the retrieval call is the last five milliseconds of a pipeline that spans ingestion, chunking, embedding, indexing, and re-ranking, and every one of those upstream stages determines whether the final answer is actually grounded in something true.
This is a systems-design problem with the usual trade-offs: throughput versus freshness, recall versus latency, storage cost versus retrieval quality. This article works through the pipeline stage by stage, with the numbers that make each trade-off concrete.
Ingestion and chunking
Ingestion pulls source documents from wherever they live (object storage, a CMS, a database, a crawl) and normalizes them into text. The design decision that matters most here is chunking: how you split a long document into the units that actually get embedded and retrieved.
Fixed-size chunking (say, 512 tokens with a 10-15% overlap) is cheap and predictable, but it cuts across sentence and section boundaries indiscriminately, which fragments meaning right at chunk edges. Structure-aware chunking (split at headings, paragraphs, or semantic boundaries) preserves meaning better but produces uneven chunk sizes that complicate batching and can blow past a model's context budget if a section runs long. In practice, most production pipelines use a hybrid: split at structural boundaries first, then fall back to fixed-size splitting inside any section that's still too large.
Chunk size is a direct trade-off against retrieval precision. A 2000-token chunk gives the generator more surrounding context per retrieved item but dilutes the embedding -- a chunk about three different subtopics produces a smeared vector that matches all three queries poorly. A 200-token chunk embeds a focused idea well but forces the generator to stitch together many small fragments, increasing the chance a needed piece of context simply wasn't retrieved. 300-800 tokens with 10-15% overlap is the common production range; the right point within it depends on how internally coherent your source documents already are.
Embedding generation and the write path
Embedding is where ingestion throughput actually gets bottlenecked. A single-document, single-request embedding call wastes most of an embedding model's throughput -- production pipelines batch chunks (typically 32-256 per batch, tuned to the embedding provider's rate limits and your own compute if self-hosted) and run ingestion as an asynchronous job queue, not a synchronous step in the document's save path.
document arrives -> chunk -> enqueue chunks for embedding
|
embedding worker pool (batched, N chunks/call)
|
vector + metadata written to index
|
document marked "indexed" (async, eventually consistent)The write path into the index itself has its own throughput ceiling, and it's a different one than the embedding call's. Most vector indexes (see vector search infrastructure for the index-type trade-offs) are optimized for read (query) throughput, not write throughput -- inserting into an HNSW graph, for instance, gets more expensive as the graph grows, because each insert needs to find its neighbors in an increasingly large structure. High-ingestion-volume systems batch writes and often maintain a small, fast "staging" index for very recent content that gets periodically merged into the main index, rather than writing every chunk into the primary index synchronously.