The Context Window

Chunking Strategies and Retrieval Quality in RAG

Chunking matters as much as your embedding model—and most teams optimize the wrong layer.

Staff Writer · · 12 min read
Cover illustration for “Chunking Strategies and Retrieval Quality in RAG”
Context Delivery · September 21, 2026 · 12 min read · 2,637 words

Every RAG pipeline improvement conversation starts in the same place: swap the embedding model, upgrade to a newer LLM, rewrite the prompt. Chunking strategy rarely comes up, and that's backwards. No downstream component can retrieve information that got discarded or mangled when the document was first split into pieces. Feed a retriever garbage chunks and the best embedding model in the world just finds the garbage faster.

The evidence for this isn't anecdotal. A study published at NAACL 2025 (arXiv:2410.13070) by researchers at Vectara tested 25 chunking configurations against 48 embedding models and found that chunking configuration influenced retrieval quality as much as, or more than, the embedding model itself. That's a striking result, given how much attention embedding model selection gets in comparison. Teams spending weeks tuning prompts and swapping models while their retrieval quietly returns the wrong context on every third query are optimizing the wrong layer of the stack. This piece maps the major chunking approaches on the market, the trade-offs each one carries, and a decision path for choosing among them that holds up under scrutiny.

The two failure modes that bound every chunking decision

Chunking is a tug-of-war between two failure modes pulling in opposite directions, and every strategy on the market is a negotiated position between them. It's a tug-of-war between two failure modes pulling in opposite directions, and every strategy on the market is a negotiated position between them.

Chunks that are too small lose context. In the FloTorch 2026 benchmark, semantic chunking produced fragments averaging just 43 tokens, and those fragments scored only 54% accuracy on end-to-end question answering. The retrieval step found the right fragment. The LLM still couldn't answer the question, because 43 tokens rarely contains enough surrounding context to reason from. High recall doesn't save you here.

Chunks that are too large dilute relevance. A systematic analysis identified what it called a "context cliff" around 2,500 tokens, past which response quality drops off. Large chunks tend to cover multiple topics at once, and when you embed a chunk like that, the resulting vector is an average across everything in it, weakly relevant to any single query someone actually asks.

The accuracy spread between good and bad chunking choices is not subtle. Semantic and proposition-based chunking methods consistently post recall numbers in the 87 to 92% range, compared to 50 to 65% for a fixed-size baseline. Between the best and worst approaches, on the same model, the accuracy gap is substantial. That gap separates a system that works from one that doesn't.

Retrieval recall and end-to-end answer accuracy are not the same metric, and a strategy can win one while losing the other badly, a pattern that recurs throughout every comparison that follows. Semantic chunking's 91.9% retrieval recall against its 54% answer accuracy in the FloTorch 2026 benchmark is the cleanest illustration on record. Before trusting any chunking benchmark, check which of the two it's actually measuring.

Fixed-size and recursive splitting: the baseline every team starts with

Fixed-size chunking splits text at a uniform character or token count, with no regard for what the document actually looks like. It's trivial to build, deterministic, fast, and gives you a predictable index size, which explains why almost every RAG tutorial starts here. The weakness is equally plain: it splits sentences mid-thought, it can separate a table's header row from its data rows, and it has zero awareness of where one topic ends and another begins. It's fine for a prototype or for very uniform, well-structured text. It is not something to ship into production against real documents.

The Vectara study cited above cuts against the industry's default assumption. Semantic chunking, despite the added computational cost of embedding at chunk-boundary time, did not consistently beat fixed-size chunking across document retrieval, evidence retrieval, and answer generation tasks. Paying more didn't reliably buy better results.

Recursive character text splitting is the more practical default. It works through a hierarchy of separators, trying paragraph breaks first, falling back to sentence breaks, then word boundaries, then raw characters if nothing else works. It's the default in LangChain (LlamaIndex ships SentenceSplitter as its default instead), and it handles general prose noticeably better than fixed-size splitting for very little extra complexity. It's still blind to document-level structure like headings, tables, or code blocks, and the separator hierarchy needs to be configured per content type rather than tuning itself.

The benchmark result here is the one that should reset expectations industry-wide: in the largest real-document test run in 2026 (the FloTorch/premai.io benchmark guide), recursive splitting at 512 tokens with 50 to 100 tokens of overlap scored 69% end-to-end accuracy, beating every more expensive alternative tested. Microsoft's Azure documentation recommends 512 tokens with 25% overlap (128 tokens) as a starting point, measured in BERT tokens rather than raw character counts. Separately, Arize AI found that chunk sizes between 300 and 500 tokens combined with top-k retrieval of 4 chunks gave the best balance of speed and quality. A sensible starting range, before any tuning, is 256 to 512 tokens with 10 to 20% overlap.

None of this makes 69% a number worth celebrating on its own terms. It's a floor that every more expensive chunking method has to clear on end-to-end accuracy, not just on retrieval recall, and a lot of them don't.

Document-aware and semantic chunking: when content structure should drive boundaries

Document-aware, or structure-aware, chunking parses the actual structural elements of a document, headers, sections, paragraphs, tables, and code blocks, before it splits anything. The resulting boundaries mirror how a person would naturally break the document apart, keeping a table whole, leaving a code block uncut mid-function, and turning a heading boundary into a dependable split point rather than an accident of token counting. For content types that carry detectable structure, Markdown documentation, PDFs with a real layout, HTML pages, Jupyter notebooks, technical manuals, this approach delivers the best retrieval-effectiveness-per-dollar of anything covered here. It also costs about the same to run as recursive splitting, since it doesn't require additional model calls at chunking time.

The rules need to be domain-specific to work. Markdown should split on heading boundaries. Code should split on function or class boundaries. HTML tables should be chunked row by row, with the header row copied into every chunk so a retrieved row never loses its column labels.

Semantic chunking takes a different route: it uses embeddings to find where the topic actually shifts, rather than cutting wherever a token counter runs out. The appeal is obvious on paper, boundaries that track meaning instead of length, and Firecrawl's 2026 findings back it up in part, showing recall improvements of up to 9% over simpler chunking methods. Semantic chunking requires embedding every sentence at indexing time, which adds real API expense and latency at any meaningful scale. The Vectara peer-reviewed study reinforced this point, finding that semantic chunking did not consistently outperform fixed methods and carried meaningfully higher computational cost under real-world document conditions.

The recall-versus-accuracy trap mentioned earlier belongs squarely in this section. Semantic chunking's 91.9% retrieval recall in the Chroma benchmark, set against its 54% answer accuracy in FloTorch 2026, is the single clearest argument in this entire field for why retrieval-only benchmarks mislead people. Fragments averaging 43 tokens retrieve beautifully. They just don't give the LLM enough to work with once retrieved. Setting a minimum chunk-size floor isn't a nice-to-have with semantic methods; skip it and high recall turns into unusable answers.

When the document has real, parseable structure and indexing cost is a constraint, document-aware chunking gets most of semantic chunking's benefit for a fraction of the compute.

Hierarchical chunking: resolving the precision-context trade-off in production

Small chunks find things better. Large chunks explain things better. Hierarchical chunking is the strategy that stops trying to pick one and uses both: small chunks get indexed for retrieval precision, and when one of them matches a query, its larger parent chunk gets handed to the LLM so the model has the surrounding context it needs. This is the parent-child pattern, sometimes called "small-to-big" retrieval, and it's become the closest thing to a default production pattern across 2025 and 2026, because it resolves both failure modes from the opening section without requiring a single LLM call during indexing.

The benchmark numbers back up the adoption. Hierarchical chunking lifts retrieval quality by 18 to 25% over flat chunking methods according to HiCBench results, and in FinanceBench testing, 1,024-token chunks with 15% overlap produced a 22% jump in answer accuracy (per dasroot.net). LlamaIndex's HierarchicalNodeParser ships with three default levels, 2,048, 512, and 128 tokens, and its SentenceSplitter defaults to 1,024-token chunks with 20 tokens of overlap out of the box, both reasonable starting points for teams building this pattern themselves.

The enterprise fit is strong, particularly for corpora that mix document types and formats, since the different hierarchy levels naturally capture different granularities of meaning without extra engineering per source. The cost is moderate: no model calls at index time, but the index itself grows, since both parent and child chunks get stored.

Agentic and proposition chunking: when quality justifies the cost

Agentic chunking hands the job to an LLM directly: the model reads a document and proposes chunk boundaries based on what it understands the content to mean, not on a heuristic approximation of meaning. It can find where one claim ends and the next begins, where a procedure wraps up, where the context shifts, in a way no rule-based splitter can match. Figures from digitalapplied.com put the cost at a steep 10 to 50 times the indexing cost of fixed-size chunking.

Proposition chunking, the approach behind the Dense X Retrieval method, converts each passage into a set of self-contained, atomic factual statements and turns each one into its own chunk. Every resulting chunk answers exactly one question, cleanly and specifically. On paper, that sounds like a win. In the FloTorch 2026 benchmark, though, proposition-based and agentic approaches ranked among the worst performers on end-to-end accuracy, the exact same too-small-a-fragment problem that sank semantic chunking. Being self-contained doesn't help if the chunk still lacks the surrounding context the model needs to reason correctly. That's a design caution, not a reason to write off the method: proposition chunking needs careful minimum-length controls, or it walks straight into the 43-token failure mode described earlier.

The right home for agentic and proposition chunking is high-stakes, high-value corpora where retrieval accuracy actually matters and indexing cost is tolerable, legal contracts, regulatory filings, clinical guidelines, product specifications. It's the wrong choice for bulk indexing of millions of pages, or anywhere indexing speed and cost are hard constraints. The honest trade-off: the best retrieval quality on the table lives here, but only when the corpus is curated, relatively stable, and small enough that paying an LLM call per document doesn't break the budget.

Contextual retrieval and late chunking: the hybrid approach with the strongest failure-rate evidence

Two techniques, addressing two separate gaps, that work best combined. Contextual retrieval prepends a short, LLM-generated summary of where a chunk sits in the document before embedding it, solving the problem of a chunk that reads fine on its own but has lost track of its place in the larger argument. Late chunking flips the order of operations: it embeds the entire document first with a long-context model, then pools the resulting token vectors into chunks afterward, which preserves the cross-section dependencies that get severed the moment chunks are embedded independently.

Combined, the two techniques cut retrieval failure rates by 35 to 67% compared to naive fixed-size chunking, according to Anthropic's published benchmarks and Jina AI's paper on late chunking (as reported by aiworkflowlab.dev, 2026). That range is wide, but even the low end represents a meaningful jump over the baseline. Contextual retrieval adds an LLM call per chunk to generate its context summary, though far fewer calls than full agentic chunking requires, and late chunking needs a long-context embedding model but no generation calls at all, putting the cost of both in the middle of the pack. Enterprise fit is high on both counts. Contextual retrieval reads as low-risk with a meaningful accuracy payoff, applicable to nearly any pipeline, while late chunking suits long documents with real cross-section dependencies, research papers, contracts, technical manuals.

One live dispute belongs in this section rather than glossed over. Nearly every chunking guide in circulation recommends 10 to 20% overlap between chunks as a near-universal default. A systematic analysis published on arXiv found that overlap produced no measurable benefit. That's a genuine methodological disagreement that should be treated as such: run overlap as a variable to test against a specific corpus, rather than accepting the 10 to 20% figure on faith because it appears in every blog post on the topic.

Metadata-enriched retrieval: the layer that makes chunking decisions stick at enterprise scale

Diagram: The Two-Metric Trap: Recall vs. Answer Accuracy. Visualizes: Show the dangerous gap between retrieval recall and end-to-end answer accuracy using the FloTorch 2026 / Chroma benchmark data for semantic chunking.

Metadata enrichment attaches structured attributes to each chunk alongside its text: source URL, publication date, section heading, document type, author, a freshness signal, ownership, policy context. None of that touches how the chunk gets split. It changes what happens after.

Metadata lets a system filter before it ever runs semantic search, scoping retrieval down to a relevant subset of the corpus instead of searching everything every time. That's faster, and it's more precise, because the semantic search step is no longer competing against irrelevant documents that happen to share vocabulary. At real enterprise scale, this stops being a nice-to-have. Without freshness and lineage signals attached to chunks, RAG systems produce stale information, ungoverned content, or answers that quietly contradict each other, a failure mode that gets worse, not better, as the corpus grows (per atlan.com, 2026).

A production recipe from webscraft.org (2026): semantic chunking, 10 to 15% overlap, metadata covering source, section, and type, top-k retrieval of 3 to 5 chunks, and either MMR or a reranking step on top. On top-k specifically, 3 to 5 chunks covers most queries adequately; stuffing more chunks into the context window mainly adds noise and raises the risk of the "lost in the middle" problem, where the model loses track of relevant information buried in a long context. If recall looks weak, the fix is better chunking and better embeddings.

The DCD architecture, described in a paper by Kovalskii et al. (arXiv:2604.07590, April 2026), formalizes this pairing of metadata with hierarchy into a three-tier structure, enabling multi-stage routing that narrows both retrieval and generation scope in stages. It's a genuine blueprint for teams building governed enterprise RAG systems, not a toy concept. The trade-off is real too: metadata enrichment carries the highest enterprise fit of any approach in this piece, and also the highest implementation overhead. It's worth building for a corpus with real governance requirements. It's overkill for a weekend proof of concept.

How to measure whether your chunking strategy is working

Whether a chunking strategy is doing its job depends on two questions, and they require two different metrics to answer.

Context Recall asks whether all the relevant information needed to answer a query actually got retrieved. This measures the retrieval step in isolation, before the LLM ever sees the chunks, and it's the metric most chunking benchmarks report by default, which is why the Chroma evaluation numbers matter so much: semantic chunking looked excellent on recall (91.9%) while its end-to-end accuracy collapsed to 54%. Recall alone tells you whether the right needle got pulled from the haystack. It says nothing about whether the model, once handed that needle, could actually answer the question.

Any team evaluating a chunking strategy needs both numbers side by side, retrieval recall and end-to-end answer accuracy, before drawing a conclusion. A strategy that wins on one and loses badly on the other is a loss. It's a benchmark trap, and the FloTorch data makes clear that a lot of teams are already caught in it.

Sources

  1. Chunking Strategies for RAG: Methods, Trade-offs & Best Practices
  2. RAG Chunking Strategies: The 2026 Benchmark Guide
  3. Best Chunking Strategies for RAG (and LLMs) in 2026
  4. DCD: Domain-Oriented Design for Controlled Retrieval-Augmented Generation
  5. RAG Chunking Strategies: A 2026 Retrieval Playbook
  6. Chunking Strategies RAG 2026 : Best Ways to Split Data for Production
  7. RAG Chunking Strategies Guide (2026)
  8. dasroot.net
Filed underContext Delivery

More in Context Delivery