Skip to content
CoronRing
All tools

Chunk Visualizer

Retrieval quality dies at chunk boundaries. Run seven real chunking strategies over your own document and see the cuts painted in place, with size distribution, the overlap you actually got, and every boundary that lands mid-sentence.

Runs entirely in this tab. No network requests, no logging, nothing leaves the page.

Document & In-Place Boundary Map
Samples:
Legend:■ Chunk■ Overlap⚠️ Cut
4 Chunks:
Splitter Strategy & Tuning
Budget Unit:

The default in most RAG stacks, and the right first choice. It tries to split on paragraphs, then lines, then sentences, then words, taking the largest unit that fits, so boundaries land at natural breaks unless the text leaves no option. Merging the pieces back up to the size budget is what keeps it from emitting a pile of one-line chunks.

Diagnostic Metrics & Sizing Curve
chunks
4
mean size
366c
median 359c · 88 tok
range
266–478c
index expansion
1.02×
chars stored ÷ source
overlap achieved
9c
asked 80c
sentence cuts
0
0 runts · 0 oversize
Chunk size distribution
0 charsbudget 500500 chars

About this tool

01 · In-Place Boundary Map

Tracks exact source offsets to paint cuts directly over document text, distinguishing overlapping regions from single-coverage zones.

Visual: Highlights mid-sentence breaks and fragment tails.

02 · Strategy Support

Seven standard splitter implementations: LangChain recursive-character, character, Markdown header, semantic topic boundary, code AST, sliding window, and sentences.

Real behavior: Descends separator tiers and merges up to target budget.

03 · Diagnostic Metrics

Measures achieved vs requested overlap, mid-sentence severed boundaries, index expansion factor, and size distribution clustering.

Sizing: Dynamic token-to-character ratio calibration.

Operational Limits & Assumptions

  • · Lexical Semantic Split: The semantic strategy uses term-vector shift client-side without remote model calls; treats synonym rephrasing as topic movement.
  • · Sentence Heuristics: Rule-based sentence segmentation may miss boundaries in OCR dumps, unpunctuated voice transcripts, or rare abbreviations.
  • · Index Overhead: Higher overlap increases total vector store storage and retrieval re-ranking costs.
  • · Eval Boundary: Visual inspection verifies splitter mechanics; end-to-end recall requires validation against your evaluation query set.

Found a case it gets wrong? Tell me about it .