Skip to content

Generative AI

Text Chunking for RAG

O(n) time and space, where n is the number of words in the document. The overlap only makes each word get visited a small, fixed number of extra times.

The idea, in plain English

Imagine you slice a long loaf of bread into pieces small enough to eat. But instead of clean, unrelated slices, each piece overlaps a little with the next, so no bite loses its context. This is text chunking. Before an AI can search over or reason about a long document, you have to cut the document into small pieces called chunks that fit in the AI's working memory. Chunks usually overlap a little with their neighbors. That way, an idea sitting right on the cut line does not get sliced in half and lost.

How it works

  1. 1Split the document into words, or another small unit of text.
  2. 2Decide on a chunk size, which is how many words go in each chunk, and an overlap, which is how many of the last words in one chunk also start the next chunk.
  3. 3Walk through the document, taking a chunk of `chunk size` words at a time. But instead of jumping forward by the full chunk size each time, jump forward by (chunk size minus overlap) words. This smaller jump is what creates the overlap between chunks that follow each other. Stop once a chunk reaches the end of the document.

When you'd use it

Use this in Retrieval-Augmented Generation (RAG) systems. Before you can search or embed a long document, you first have to break it into chunks small enough to embed and hand to a language model. This preparation step happens before Vector Search / RAG Retrieval.

Common beginner mistakes

  • Don't chunk by a fixed number of characters instead of something that respects meaning, like whole words or sentences. This can slice a sentence, or even a word, right in half, making both halves harder to search and understand.
  • Don't use zero overlap. If the answer to a question spans the exact boundary between two chunks, neither chunk alone contains the full answer.
  • Don't use an overlap that is too large compared to the chunk size. You end up storing and searching almost the same text many times over. This wastes space and lets near-duplicate chunks compete with each other during retrieval.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Text: The quick brown fox jumps over the lazy dog while the cat watches quietly from the window
Total words: 17
Chunk size: 6  Overlap: 2

Chunk 1 (words 0-5): The quick brown fox jumps over
Chunk 2 (words 4-9): jumps over the lazy dog while
Chunk 3 (words 8-13): dog while the cat watches quietly
Chunk 4 (words 12-16): watches quietly from the window

Total chunks: 4

Not sure this is the right topic? See the learning paths → or where this leads →