AI, Software Development

RAG chunking strategies: what to pick and how to test it

By James KillickSeptember 27, 2026

TL;DR: Start with recursive or structure-aware splitting at about 400 to 512 tokens and 10 to 20% overlap, counted in tokens, not characters. Treat that as a baseline, not a rule. Studies disagree on chunk size, and some found overlap adds cost with no gain. Build a small test set, change one setting at a time, and only pay for semantic, contextual or adaptive chunking when it beats your baseline on your own data.

Start with recursive or structure-aware splitting at about 400 to 512 tokens. Add 10 to 20% overlap. Then test it on your own questions before you change a thing.

That's the short answer on RAG chunking strategies. But every number in it is a place to start, not a rule. The research can't agree on the best chunk size. It can't even agree that overlap helps.

So here's what each method does, what the studies found, what it costs, and how to test one against another.

Where chunking sits

Chunking is the step where you split your documents into pieces before you embed them. Those pieces are what search finds and what the model reads.

We've covered the rest of the pipeline. Our guide to building a RAG pipeline walks through all four parts and gives the short version of chunking. New to the idea? Start with what RAG is. For the steps on each side, see embeddings explained simply and vector databases explained. AI-Led's walkthrough on how to build a RAG application shows the whole build, step by step.

This post stays on chunking.

The trade is simple. Small chunks are easy to find, but they can miss the context you need for a full answer. Big chunks hold the full answer, but they're harder to match to a question.

Pinecone's chunking guide has a handy test. If a chunk makes sense to a person with no text around it, it will make sense to the model too.

A 2025 paper points the same way. The HOPE study of chunk quality scored chunks across seven domains. Chunks that didn't lean on each other for meaning mattered most, with gains of up to 56.2% in factual correctness. Keeping one whole concept inside each chunk made little difference.

The main chunking methods compared

MethodHow it splitsBest forCost
Fixed-sizeCuts every set number of tokensPlain text with no clear structureLow. No model calls
SentenceCuts at the end of sentencesProse, as a fallbackLow. No model calls
RecursiveTries paragraphs, then lines, then wordsMost text. The usual defaultLow. No model calls
Structure-awareCuts at headings, sections or pagesManuals, contracts and PDFsLow to medium. Needs a parser
SemanticEmbeds sentences, cuts where the topic shiftsDocuments where topics overlap across sectionsHigh. Embeds every sentence
Contextual and late chunkingAdds document context to each chunkChunks that make no sense aloneExtra model calls or memory

The cost column is a rough guide. Microsoft's Azure guide to the chunking phase rates fixed-size and sentence splitting as low effort and low cost. It rates semantic chunking as high on both. It also says its own ratings are subjective.

Fixed-size is the blunt one. You pick a token count and cut. Pinecone says to start here, and to move on only when you've found it isn't enough.

Sentence splitting cuts at full stops. It's cheap. The catch? One sentence rarely holds a whole idea. Microsoft treats it as a fallback for when nothing else fits.

Recursive splitting is the usual default. LangChain's docs call its recursive splitter the one to use for generic text. It tries paragraph breaks first, then line breaks, then spaces. So a paragraph stays whole for as long as it fits.

One thing to watch. Chroma's chunking report found the default separators often made very short chunks. Their fix was to add sentence endings to the list.

Structure-aware splitting uses the shape of the document. Headings in Markdown. Tags in HTML. Pages in a PDF. In NVIDIA's chunking tests, page-level chunks had the best average accuracy at 0.648, and the steadiest results. It was close though. Fixed token sizes scored between 0.603 and 0.645.

AI Orchestrators makes the same point for business content. Their guide to a consultant knowledge base says to break long documents into focused chunks with one idea each.

Semantic chunking embeds each sentence, then cuts where the topic shifts. The chunks follow meaning, not length. You pay for that with an embedding call on every sentence.

Contextual retrieval and late chunking came later. Both fix a chunk that makes no sense on its own. With Anthropic's contextual retrieval, an LLM writes a short note that places each chunk in its document. You embed the note with the chunk. With late chunking, you run the whole document through the embedding model first. Then you cut it into chunks.

Chunk size: what the research found

Three studies. Three answers.

  • A 2025 study of chunk size found chunks of 64 to 128 tokens worked best when answers were short facts. Chunks of 512 to 1,024 tokens worked better when the answer needed wider context.
  • NVIDIA found the two ends of the range did worst. Chunks of 128 and 2,048 tokens mostly lost to the sizes in between. Fact questions did best at 256 to 512 tokens. Complex questions did better with bigger chunks or whole pages.
  • Chroma found LangChain's recursive splitter at 200 tokens, with no overlap, scored well on every measure it tracked.

So one study likes 128 tokens for facts, and another says 128 is too small. Both can be right. They used different data and different embedding models.

The embedding model matters too. In the chunk size study, one model did better with big chunks and another did better with small ones.

So what do you do? Match the chunk size to the questions your users ask. Short fact lookups? Go smaller. Questions that need a few linked points? Go bigger. Then test two or three sizes. Pinecone suggests 128 or 256 tokens at the small end, and 512 or 1,024 at the big end.

Overlap: a place to start, not a rule

Overlap repeats the end of one chunk at the start of the next. The goal is to stop an idea being cut in half at the join.

Most guides land on the same range. Firecrawl's guide to chunking strategies recommends recursive splitting at 400 to 512 tokens with 10 to 20% overlap as the default. For a 500-token chunk, that's 50 to 100 tokens. The Neural Base lesson on chunk overlap calls 10 to 20% the sweet spot for most RAG systems. NVIDIA tried 10, 15 and 20%, and 15% did best on one of its test sets. Its team says that wasn't a full search.

Microsoft goes higher. Its Azure AI Search chunking guide says to start at 512 tokens with 25% overlap, which is 128 tokens.

One problem though. Some tests found overlap didn't help at all.

Chroma's best recursive result at 400 tokens came with no overlap. It hit 89.5% recall, against 88.1% with a 200-token overlap. And a 2026 preprint on chunking for question answering found overlap gave no benefit it could measure, and it raised the indexing cost. That was one setup on one dataset. Firecrawl's guide flags the same study.

Overlap isn't free. Microsoft's guide splits one PDF e-book at 1,000 characters. With no overlap, it makes 172 chunks. With a 200-character overlap, it makes 216. That's 44 more chunks to embed, store and search. The Neural Base lesson warns that overlap of 50% or more doubles your storage cost.

So start at 10 to 20%. Then run the same test with no overlap. If the score holds, keep the smaller index.

Count in tokens, not characters

This one catches people out.

LangChain's recursive splitter measures chunk size in characters by default. The Neural Base lesson points out the overlap setting works the same way. But embedding models count in tokens. Microsoft puts a token at about four characters for common OpenAI models. So a chunk size of 1,000 in LangChain is roughly 250 tokens, not 1,000.

The fix is easy. Tell the splitter to count tokens, with the tokeniser that matches your embedding model. Microsoft's Azure guide also says to count tokens, not characters, for fixed-size chunks.

Then check your budget. Chunk size times the number of chunks you pull back, plus your prompt, has to fit the model's window. The LlamaIndex docs show the trade. When they halve the default chunk size from 1024 to 512, they double the chunks pulled back from 2 to 4.

Our post on AI context windows covers the limits. AI-Led's guide to context engineering covers what to put in the window.

If you swap embedding models, check it all again. Pinecone notes that different models can split the same text into tokens in different ways.

What each method costs you

Accuracy is half the picture. The other half is what you pay to build and run the index. There are four costs to watch.

  • Index time. Fixed, sentence and recursive splitting are plain text work. Semantic and LLM methods add model calls for every document.
  • Storage. Smaller chunks and more overlap both mean more vectors.
  • Memory. Late chunking holds the embedding of every token in memory while it builds the index.
  • Re-indexing. If your documents change each day, a slow method hurts each day.

An August 2026 preprint measured all of this. It ran eight chunking methods across two large sets of documents and three embedding models. It tracked indexing speed, query speed and memory next to the retrieval scores.

The result? The costly methods rarely gave steady gains over the simple ones. The authors call token-based chunking a strong default for most large systems. One cheap win stood out. Adding the document title to each chunk held up well across many settings.

LLM methods cost real money at index time. Anthropic puts contextual retrieval at US$1.02 per million document tokens, as a one-off, with prompt caching. That figure assumes 800-token chunks and 8,000-token documents.

Time counts too. Chroma's LLM chunker took up to tens of minutes to run.

And it's hard to undo. Microsoft calls your chunking approach a semi-permanent choice. Change it later and the work ripples through the rest of the pipeline.

Chunking is one part of the bill. Our post on token cost optimisation covers the rest, like pulling ten chunks when three would do.

How to test one strategy against another

Picking a chunking method with no test set is guessing. Here's a simple way to run it.

  1. Build a small test set. Twenty to thirty real questions is enough to start. Mix short fact questions with ones that need more context.
  2. Change one thing at a time. Keep the embedding model, the retriever and the prompt fixed. Swap only the chunking.
  3. Tune size and overlap first. Try two or three sizes and two overlap settings inside one method, before you compare methods.
  4. Score retrieval and answers. Did the right chunk come back? Did the answer stick to it? Our post on LLM evaluation covers the RAG metrics and how far to trust a judge.
  5. Score the cost. Count chunks, index time and embedding calls for every run.
  6. Read your chunks. Print ten and read them. If a chunk makes no sense to you, it won't help the model.
  7. Write down the settings. Splitter, size, overlap and embedding model, saved with every index build.

When the advanced methods are worth paying for

Sometimes. Less often than the hype suggests.

Semantic chunking. A NAACL 2025 Findings paper tested it on three retrieval tasks. Its verdict: the extra computing cost isn't justified by consistent gains. Chroma's numbers are mixed. Its own semantic chunker and its LLM chunker reached 91.3% and 91.9% recall. The recursive splitter reached 88.1% to 89.5% at 400 tokens. That's a gain, but a small one. And the semantic chunker that was later built into LangChain scored 83.6%, below the recursive splitter.

Adaptive chunking. This picks a method for each document. An LREC 2026 paper reports answer correctness of 72%, up from 62 to 64%, with no change to the models or prompts. The system answered 65 questions, against 49 for the baselines. It's a solid result from one study. Test it on your own documents before you count on it.

Contextual retrieval. Anthropic's own tests cut failed retrievals by 35%, from 5.7% to 3.7%. Adding the same context to keyword search took that to 49%. But the August 2026 preprint found contextual chunking rarely beat the simple baselines by a clear margin. So the evidence is split.

Late chunking. The paper reports better results across a range of retrieval tasks. The preprint found it was often weaker on bigger sets of documents.

Here's the thing. Every one of these results came from someone else's data. The only number that counts is the one from your test set. Pay for an advanced method when it beats your baseline by enough to cover the extra cost. Not before.

And check if you need chunking at all. Anthropic's 2024 post notes that a knowledge base under 200,000 tokens, about 500 pages, can go straight into the prompt.

Where Devwiz fits

Most teams don't need a smarter chunker. They need a baseline, a test set and the patience to change one thing at a time.

We build AI apps with chat trained on your own content, and search that finds the right answer. Chunking sits under both. If your content is your own program or method, our AI programs work turns it into a platform.

Teams swap chunkers because it feels like progress. Most of the time the fix is duller. Read your chunks. If they make no sense to you, they make no sense to the model.

*James*

Devwiz has shipped 200+ apps, including work for NSW Government (Justice and Corrective Services), Briometrix, Vivid and Huskee.

Building a RAG feature and not sure your chunks are the problem? Start with AI app development. Worth a chat.

Frequently asked questions

What is the best chunk size for RAG?

There isn't one. A 2025 study found 64 to 128 tokens worked best for short fact answers, and 512 to 1,024 tokens for answers that need wider context. NVIDIA's tests found 128 and 2,048 tokens both did poorly. Start around 400 to 512 tokens, then test two or three sizes on your own questions.

How much chunk overlap should I use?

Start at 10 to 20% of the chunk size, counted in tokens. Firecrawl and The Neural Base both suggest that range, and Microsoft suggests 25%. But Chroma's tests and a 2026 preprint found overlap gave little or no gain. So test with no overlap too, and keep the smaller index if the score holds.

Is semantic chunking better than fixed-size chunking?

Not reliably. A NAACL 2025 Findings paper found its extra computing cost isn't justified by consistent gains. In Chroma's tests, one semantic chunker beat recursive splitting by a few points of recall and another scored lower. Tune a recursive baseline first, then test semantic chunking against it.

What is recursive chunking?

It's a splitter that tries the biggest natural break first. LangChain's version tries paragraph breaks, then line breaks, then spaces, until each chunk fits your size limit. It keeps paragraphs whole where it can, needs no model calls, and is the usual default for general text.

Do I always need to chunk documents for RAG?

No. Anthropic noted in 2024 that a knowledge base under 200,000 tokens, about 500 pages, can go straight into the prompt. Chunking earns its place when your content is too big for the model's window, or when sending all of it on every call costs too much.

About James Killick

10+ years building digital products · 200+ apps shipped since 2015

James is a co-founder of Devwiz and an AI product specialist. Since 2015 he has helped ship 200+ apps for founders, businesses and government, including work for NSW Government, Briometrix and Huskee. He builds AI-first platforms and writes about turning a proven program into software. He also hosts the Up in the AI podcast.

More articles by James · James's personal site · LinkedIn · AI Orchestrators

Tags: AI, RAG, LLM, Data Pipelines

Browse all Devwiz articles·See our case studies