AI & LLM Tools

RAG Text Chunker

A retrieval pipeline cannot embed an entire manual as one vector, so the text is cut into passages first. This tool does that cutting in your browser. Choose characters, words or estimated tokens, set the chunk size and the overlap, and it returns numbered chunks with their sizes, keeping whole words together where it can. Use it to experiment with sizes before writing code, to prepare a small sample corpus, or to see how overlap changes the number of chunks you would embed and pay for.

Your workspace

Runs locally in your browser

Result

Your input is processed locally in your browser and is not sent to Flutters servers. Inputs are not saved by this tool.

How to use this tool

Paste text, choose the unit, set a chunk size and an overlap smaller than the size, then choose Split text. Copy one chunk or all of them.

Example

Input

Text: a 1,200-word article
Unit: words
Chunk size: 200
Overlap: 30

Output

7 chunks, each up to 200 words, with 30 words shared between neighbours.

What does this tool do?

Retrieval-augmented generation stores a document as many short passages so a search step can fetch only the relevant ones. This chunker splits text by characters, words or estimated tokens, keeps an overlap between neighbouring chunks, avoids cutting words where practical and shows each chunk's size.

Common mistakes and limitations

There is no universally best chunk size; it depends on your content, embedding model and questions, so test retrieval quality on real queries. Token sizes are estimates, not a model tokenizer. This tool splits on whitespace only and does not understand headings, sentences, code blocks or tables, so structured documents may need a structure-aware splitter.

What chunking is for

In a retrieval-augmented setup, documents are split into passages, each passage is turned into an embedding and stored, and a question retrieves the closest passages to include in the prompt. Chunks that are too large blur several topics together and waste context; chunks that are too small lose the surrounding explanation. Overlap helps a sentence near a boundary appear intact in at least one chunk.

size 200, overlap 30:
chunk 1 -> units 1–200
chunk 2 -> units 171–370

Choosing a unit and testing the result

Characters are simple and predictable, words follow the text more naturally, and estimated tokens approximate what a model will see. No setting is best everywhere, so run your real questions against a few sizes and compare which chunks are retrieved. Structured sources such as Markdown headings, code and tables often work better when split along their own structure first.

Related: LLM Token Counter, Word and Character Counter

Chunk count drives cost

More overlap and smaller chunks mean more passages to embed, store and retrieve. Estimate token volume before processing a large corpus.

Related: AI API Cost Calculator

Frequently asked questions

Why use overlap?

Overlap repeats a little text at chunk edges so a sentence cut by a boundary still appears whole in one chunk. Too much overlap increases storage and duplicate matches.

Which chunk size should I use?

Start small enough to stay on one topic and large enough to carry context, then compare retrieval results. No single value suits every corpus.

Are token sizes exact?

No. They use the same estimate as the LLM Token Counter and can differ from your embedding model's tokenizer.

Related tools