Choosing RAG Chunk Boundaries and Overlap to Preserve Section Context
This article explains how to choose effective chunk boundaries and overlap in retrieval-augmented generation (RAG) systems, comparing fixed-size and document-aware strategies. Using a hypothetical support note example, it demonstrates the risks of splitting across semantic units and shows how metadata and overlap can preserve context. The guidance is based on Microsoft Azure documentation for chunking in vector search and RAG workflows.
On this page
The short answer
To preserve section context in RAG, use document-aware chunking that respects semantic boundaries like headings, and include relevant metadata such as section titles in each chunk. Fixed-size chunks with overlap may split critical information, so supplement them with structural awareness. For example, a policy stating 'Returns: unopened items may be returned within 14 days' must keep both condition and duration together. When using fixed-size chunks, apply overlap (e.g., 10-25%) and repeat section headings as metadata. Always validate retrieval behavior near known boundaries and adjust one parameter at a time during tuning.
Choose the unit that preserves the answer
The primary goal of chunking is to ensure that each chunk contains enough context to independently answer potential queries. As stated in the Microsoft RAG guide, chunks that are too small and lack sufficient context lead to poor outcomes. In our hypothetical support note - 'Returns: unopened items may be returned within 14 days. Warranty: manufacturing defects are covered for 12 months.' - a query about return eligibility depends on both the condition (unopened) and time limit (14 days). If a chunk boundary splits 'within' and '14 days', retrieval might miss the full condition.
Choose a unit that keeps the relevant condition and its time window together while staying within the embedding model's actual input limit. Document structure and token length both matter: a long section may still need to be split. A complete short rule is a useful starting point, not a guarantee that the retriever will find it.
Keep headings with their passages
Section headings provide essential context for interpreting content. When a passage like 'unopened items may be returned within 14 days' appears without the 'Returns' header, its meaning becomes ambiguous. The Microsoft RAG documentation emphasizes preserving semantically relevant content, which includes associating headings with their respective text blocks.
Attach the section title to its passage. Storing a title as metadata alone does not ensure that it is embedded or shown to the language model. If the title supplies necessary context, include it in the text used for embedding and in the retrieved context passed to the answer generator. The dictionary below illustrates a stored chunk; it is not a complete Azure index configuration.
chunk = {
"text": "Returns: unopened items may be returned within 14 days.",
"metadata": {"section": "Returns"}
}Compare two illustrative boundaries
Consider two illustrative segmentations of the support note. First, suppose a character-based boundary cuts the note in the middle of a word:
Chunk 1: 'Returns: unopened items may be returned withi' Chunk 2: 'n 14 days. Warranty: manufacturing defects ' Chunk 3: 'are covered for 12 months.'
Here, the return policy is split between chunks, risking incomplete retrieval. Now compare a document-aware approach that uses 'Warranty:' as a boundary:
Chunk A: 'Returns: unopened items may be returned within 14 days.' Chunk B: 'Warranty: manufacturing defects are covered for 12 months.'
This version preserves both policies intact. While the first method relies solely on size, the second respects semantic structure - a key advantage highlighted in Azure's guidance on variable-sized and semantic chunking.
Choose an initial size and overlap
Azure AI Search recommends starting with 512 tokens (~2,000 characters) and 25% overlap (128 tokens) when using fixed-size chunking. This balances context continuity against redundancy. Overlap allows phrases split across chunks to appear fully in at least one result.
For our example, setting a chunk size of 60 characters with 15-character overlap might help bridge the gap between 'within' and '14 days'. However, overlap alone cannot guarantee preservation of intent if the logical unit spans more than one chunk. Thus, while overlap improves robustness, it does not replace structural awareness.
Account for repeated context
Overlap introduces duplicated text, increasing storage and indexing costs. As noted in the Microsoft RAG article, some approaches incur higher financial and temporal costs. Repeating section headers in multiple chunks also adds redundancy but improves interpretability.
In our case, repeating 'Returns:' at the start of every chunk under that section ensures clarity, even if it increases token usage. The trade-off favors accuracy over efficiency when answering user questions correctly depends on context.
Inspect answers near boundaries
After chunking, test queries that target information near likely split points. For example, ask 'How long can I return unopened items?' and verify whether the retrieved chunk includes both the subject and the time frame.
Since this is a hypothetical example, no actual retrieval system is tested. But in real implementations, inspecting results around structural transitions - such as after headings or mid-sentence splits - is crucial for validating chunk quality, as advised in the RAG chunking phase documentation.
Change one parameter at a time
When optimizing chunking, alter only one variable per iteration - size, overlap, or parsing method - to isolate its effect. For instance, first test fixed-size chunks at 500 vs 1000 characters, then adjust overlap from 10% to 25%, keeping other settings constant.
This systematic approach supports reliable evaluation, consistent with Microsoft’s recommendation to experiment with various chunking permutations and observe trade-offs before finalizing a strategy.
Know what a chunking check cannot prove
Visualizing chunks or checking overlap does not guarantee retrieval effectiveness. Semantic boundaries do not automatically imply relevance, and no static analysis proves that a chunk will be retrieved for a given query. As the Microsoft guide notes, your chunking approach is semipermanent and affects downstream processes, so assumptions must be validated empirically.
Keeping Returns with its policy makes the passage easier to interpret, but the actual retriever still has to select it and the answer has to use it correctly. Inspect retrieved evidence and review generated answers across representative queries. A good result on one query does not establish reliability; compare several boundary cases, including questions whose answers are absent from the document.
Things to check
- Does each chunk contain a complete semantic unit?
- Is section heading context preserved in every relevant chunk?
- Would a query about return window retrieve the full condition?
- Is overlap sufficient to cover mid-phrase splits?
- Has only one chunking parameter been changed during testing?
Where this applies
This analysis uses a hypothetical example and does not reflect actual retrieval performance. Tokenization effects, model-specific limits, and embedding quality are not evaluated. The recommendations assume access to document structure and depend on proper implementation of metadata attachment and overlap logic.