Chunking revolution: AI retrieval gets smarter, cheaper, faster
The gist
AI retrieval just ditched clunky, fixed-size chunks—unlocking smarter, cheaper, and faster enterprise search by fusing multiscale context and structure-aware pipelines.
What to know
- By September 2026, operators like Pete Johnson were shipping contextualized chunking features that sidestep tedious trial-and-error tuning cycles.
- AI21 Labs’ Yuval Belfer revealed that indexing the same corpus at six chunk sizes and fusing results can recover up to 40% more retrieval quality lost by fixed chunks.
- Single-pass multimodal chunking now preserves document structure, slashes token and cost overhead, and boosts retrieval accuracy by as much as 14% across massive datasets.
Why Old Chunking Failed
Fixed-size document splits left critical context on the cutting room floor, forcing AI to guess at missing links between rules, exceptions, and explanations.
The baseline that operators were reacting against was well established: “By the early 2020s, retrieval-augmented generation (RAG) pipelines typically segmented documents into uniform chunks of a fixed size (e.g., 1,000 characters), embedded those chunks, and retrieved from that single set of vectors.” That design was cheap and simple, but it baked in recall loss because chunking often severed headings from explanations, tables from their interpretation, and related sections from each other, leaving retrieval to surface isolated fragments that forced the model to guess missing relationships.
What changed in September 2026 was that operators began describing practical fixes as deployable systems rather than research aspirations. On September 2, Pete Johnson said contextualized chunking could avoid the usual manual cycle of testing chunk sizes to balance storage cost against retrieval quality, and he tied that claim to an active release cadence: “We just released four of that in the last six weeks.” By mid-September, ByteByteGo was likewise framing chunking as an operational answer to concrete recall failures, such as when splitting a rule from its exception causes wrong answers despite the source text being present.
Fusing Multiple Chunk Views
Indexing documents at several chunk sizes and merging their results closes up to 40% of the retrieval gap left by single-size approaches, eliminating blind spots in enterprise search.
The mechanism matters now because it directly attacks the core retrieval failure in enterprise RAG: chunking freezes one view of a document even though queries need different levels of granularity. Yuval Belfer of AI21 Labs argues fixed-size chunking creates a structural retrieval “hole,” saying fixed-size pieces “silently” lose “between 20 and 40% of the retrieval quality you could have had” because chunking is “lossy compression” and “there is no right chunk size,” while freeCodeCamp.org notes naive chunks lose the surrounding context that makes them interpretable at retrieval time.
What changed is that operators now have a practical way to recover that lost signal without heavy new modeling: Belfer’s team “duplicated each dataset six times, with a different chunk size per copy — 50, 100, 200, 1,000 tokens, and so on,” then “query all six copies in parallel,” and fuse the rankings with reciprocal rank fusion. BigGo Finance reports “the fused multiscale method matched or beat the best fixed chunk size everywhere,” while contextual enrichment also helps suppress bad retrievals: “Entropic found that contextual retrieval reduces top 20 retrieval failures by 49% alone and 67% when combined with reranking.”
Preserving Structure With Markdown
A single multimodal LLM pass now rewrites messy PDFs and text into structured Markdown, rescuing vital tables and headings that old pipelines would irreversibly destroy.
Deterministic multimodal chunking starts by treating structure loss as the real ingestion failure, not a downstream retrieval bug. IT Brief New Zealand argues that flattening documents into raw text strips the headings, sections, and tables models rely on, while PDFs themselves are only presentation formats whose naive extraction can merge columns, inject headers and footers, and scramble reading order; D-RAC addresses that by first normalizing any input into PDF, then using a single multimodal LLM pass to convert rendered pages into retrieval-optimized Markdown instead of extracting brittle text streams.
That single-pass conversion matters because later stages cannot reconstruct structure once parsing has destroyed it. HackerNoon puts the failure bluntly: “None of those later stages can recover what the parser dropped,” so when “Tables become sentences… Flatten it, and you get a run of numbers,” even a simple question like “what was Q3 revenue in the EMEA segment” loses its column grounding; with “74.9%” of 20,000 scholarly PDFs lacking usable structure tags, D-RAC’s multimodal Markdown step preserves heading hierarchy, reading order, and rewrites tables into self-contained prose before deterministic chunking begins.
Token Budgets Drive Retrieval Design
Production systems trim and rerank chunks to fit strict token caps, turning retrieval optimization into a make-or-break factor for scaling enterprise AI without runaway costs.
In production, better chunking matters because it directly caps how much context an agent is allowed to consume, turning retrieval design into a cost-control mechanism rather than a tuning nicety. In the AI Agents case study, operators said, “We also have a limit of 100,000 tokens,” and enforced it by trimming excess chunks, while a staged pipeline retrieves the top 30 chunks and reranks to the top 5 so the agent processes only the most relevant material, cutting unnecessary token spend and making large document workflows more predictable under load.
That efficiency becomes economically decisive once enterprise corpora get big enough that every extra chunk compounds latency and inference cost across massive document estates. Pete Johnson says scale starts “in the neighborhood of 100,000 vectors,” the point where “milliseconds matter,” and he pairs that with retrieval gains of “as much as a 14% improvement,” asking whether that margin is “the difference between a hallucination and a correct answer”; with “the vast trove like the 90% of documents” sitting in SharePoint, Box, Dropbox, and S3, and “over 10 trillion plus pages” locked in files, those savings become the difference between pilots and viable operations.







