Retrieval is a query plan. You already know how to debug one.
Lightning Lesson
Dr. Aki Wijesundara · TAI Labs with Manu Jayawardana
24 minutes of content, 6 of questions. Open on the promise: if you understand indexes and query plans, you already own the mental model, and most RAG explainers waste your time re teaching you embeddings. Say there is no beginner recap coming. Flag the two chat moments, one now and one near the end.
01
The map
What RAG actually is
The model does not have your data. You fetch the rows, then you generate.
The model never searches. Retrieval does. Generation is just the last SELECT.
1 min. Name the two paths once. Index time is the write. Query time is the read. Point at Prompt plus chunks and say that is the JOIN. Then: we are not recapping embeddings. The rest of this session is the five failure stages on that retrieve box.
02
The map
Retrieval Is a Query Plan
Five stages. Every one has a database twin you already know.
Stage
You know this as
In retrieval
1
Row granularity
Chunking
2
Index selection
Dense vs sparse
3
WHERE clause
Metadata filter
4
ORDER BY
Reranking
5
Replica lag
Index freshness
Your retrieval is bad. One thing to check first. Type 1, 2, 3 or 4 in chat.1 chunking · 2 embeddings · 3 the filter · 4 stale index
Find the stage, not the symptom.
The spine of the session and the slide people screenshot. Run the poll properly, read three answers out loud by name. Costs 40 seconds, buys the room. Almost everyone says 1 or 2. Do not correct anyone, say "hold that thought" and advance. The reveal is the next slide.
03
The map
Where the Analogy Breaks
Most of you said chunking or embeddings. Those fail loudly. The two that cost you incidents fail silently.
01
There is no equals sign.
Ask for something not in your corpus and you still get confident nearest neighbours back.
02
WHERE does not compose.
A predicate that narrows a B tree can sever a graph traversal and collapse recall.
03
There are no transactions.
Your index is a stale replica with no lag metric and no alert.
A database tells you when it found nothing. Retrieval never does.
This slide earns the advanced framing. Slow down here. If you are behind later, take it from slide 5 or 7, never this one. The line to land is the callout, say it twice. Card 3 sets up slide 8, so flag that you will come back to it.
04
Stage 1 of 5
Chunking Is Schema Design
Chunk size is row granularity. Same tradeoff as deciding how far to normalise.
Three fixes, cheapest first
01
Split on document structure, not character counts.
Free, and it kills most boundary failures.
02
Contextual retrieval.
Prepend where the chunk sits in its parent, then embed. Anthropic measured 35 percent fewer retrieval failures.Anthropic
03
Index the small chunk, return the large one.
Match precisely, answer with full context.
Chunk boundaries are schema decisions. Never make them on character counts.
The diagram is the failure: a table splits across two chunks, numbers in one, headers in the other, both retrieve and neither answers. Everyone here has normalised a schema badly, so this lands fast. The tradeoff to narrate: small chunks match precisely but shred context, large chunks keep context but dilute the signal until the embedding averages toward nothing. Anthropic numbers are from their own engineering blog, state them flat. Costs about one dollar per million document tokens with prompt caching. If asked about late chunking: embed the whole document first, pool per chunk after, same goal, cheaper at query time.
05
Stage 2 of 5
You Would Never Run One Index
You would never put a single index on a table and call it done.
Index
Good at
Blind to
Dense
Paraphrase, intent, concepts
Exact strings, error codes, SKUs
BM25
Rare and exact terms
Anything phrased differently
Hybrid
Both, if you fuse by rank
Much less, at two queries plus fusion
Fusing by score fails, the scales differ. Reciprocal rank fusion ignores scores: each list gives 1 over (60 plus rank).
If a user can paste an identifier into your search box, dense only retrieval will fail and the score will still look healthy.
Say hybrid is blind to much less, at the cost of two queries and a fusion step. It is not blind to nothing: it still inherits both indexes’ recall ceilings and adds a fusion tuning problem, and someone in this room will push back if you say otherwise. The error code example is what lands. Ask them to picture a customer pasting ERR_CONN_REFUSED_5031 into support search. Cosine similarity sees a generic error shaped token sequence and returns generically error shaped documents. BM25 sees a rare term and goes straight to it. Contextual embeddings plus contextual BM25 took Anthropic to 49 percent fewer failures, neither alone gets close. Why 60: empirical, from the 2009 paper, it damps the top ranks so one confident list cannot dominate. This slide compresses if you are behind.
06
Stage 3 of 5
The WHERE Clause Trap
Scope to one customer or date range. In SQL that helps. Here your engine picks which failure you get.
Same query. Same data. The filter removed the path, not just the rows.
Filter shape
Recall
Most shapes
Above 97 percent
One broad value
90.8 percent
AND across two
39.7 percent
Qdrant's own benchmark on their own engine.
Benchmark recall with your filters applied, never without.
Click once to apply the filter: the right panel fades in and the path breaks after one hop. The two ways a filter fails: filter after searching discards non matches, so a selective filter returns fewer results than you asked for; filter before searching builds the eligible set first, and with too few eligible nodes the paths between them break and the walk dead ends early. Highest value slide for anyone running multi tenant retrieval and almost nobody measures it. The instinct to break: adding WHERE can only narrow, so it can only help. Those numbers are Qdrant benchmarking their own engine, say so. Yours will differ, the shape of the failure will not. Practical fix if asked: partition or namespace by tenant instead of filtering by tenant, so the constraint becomes which index you query. Some engines build filterable index structures and do neither naive approach, so the real instruction is to benchmark yours.
07
Stage 4 of 5
Cheap Recall, Then Expensive Precision
Same shape as a query planner. Fast scan for candidates, expensive work on a small set.
A bi encoder asks are these about the same subject. A cross encoder asks does this answer the question.
Recall is bought at retrieval and can never be recovered later. Precision is bought at rerank.
The funnel: retrieve wide with hybrid search, fuse, take 50 to 200 and optimise purely for recall. Rerank narrow with a cross encoder that reads query and document together, not precomputed vectors. Pass the top 20 to the model; Anthropic found 20 beat both 5 and 10. The distinction that matters: the bi encoder embedded that document months ago with no idea what would be asked, the cross encoder sees both at once. Reported cost is 100 to 150 ms on CPU for 50 candidates, 30 to 50 ms on GPU, and 5 to 10 nDCG points of gain. Say reported. Stack the whole pipeline for the big number: contextual retrieval plus hybrid plus rerank is 67 percent fewer retrieval failures, 5.7 down to 1.9. If behind schedule, this compresses to 90 seconds: just the two stage shape and the callout.
08
Stage 5 of 5
Your Index Is a Stale Replica
Vector similarity has no temporal dimension. A deprecated doc and the current one retrieve with equal confidence.
Three ways it goes stale
01
The source changed.
Track source timestamp against index timestamp. Trivial, and nobody does it.
02
You upgraded the embedding model.
New vectors, different geometry. Full reindex or nothing.
03
You changed chunking logic.
Same content, different vectors. A preprocessing change is an index migration.
Latency normal. Throughput normal. No errors. No alerts.
Set a shelf life per document type and monitor quality on a fixed query set, because nothing else will tell you.
This is card 3 from slide 3 cashing out, call that back. The line that gets a reaction: an embedding model upgrade is a full index migration and everyone treats it as a config change. One reported case saw recall drift 0.92 to 0.74 with nothing in the codebase changing, say reported. Practical: keep 50 real queries with known correct documents and rerun weekly. That is the whole monitoring story and it takes an afternoon.
09
Put it together
Diagnose This
Four real symptoms. Which stage broke? Type A, B, C or D in chat.
AFine on concepts. Falls apart on error codes and part numbers.
Stage 2. Dense only. Add BM25 and fuse.
BRight document, but the answer is missing a number from a table.
Stage 1. A chunk boundary split the values from their headers.
CScoped to one customer's docs and quality fell off a cliff.
Stage 3. The filter broke the traversal.
DGreat in March. Nobody touched the code. Bad now.
Stage 5. Freshness. "Nothing changed" is the tell.
All four are invisible in your logs. Latency and error rate look perfect through every one.
This is the payoff, give it the full three minutes, do not rush it to protect the closing slide. Reveal one at a time with the arrow key and read chat between each. D gets the biggest reaction because everyone has lived it: nothing in your code changed, the world did. If chat is quiet, call on someone by name, it is late and people need permission to type.
10
Take this with you
The Retrieval Triage Table
Symptom
Stage
Check first
Misses exact identifiers
2
Is BM25 in the mix
Right doc, incomplete answer
1
Split on structure
Worse when filtered
3
Benchmark recall with filters on
Relevant but badly ordered
4
Add a cross encoder
Was fine, now is not
5
Source timestamps, model version
Start measuring one thing tomorrow: recall at 20 on 50 real queries, with your filters applied.
This is week 3 of the AI Engineering Bootcamp with TAI Labs. Link in the chat.
Questions
Tell them explicitly to screenshot the table. One spoken mention of the bootcamp, then straight to questions, do not sell. Maven adds every signup to the waitlist and emails the recording in 48 hours, the work is done. Held for Q&A: GraphRAG, multi hop and agentic retrieval, ColBERT, quantisation, fine tuning embeddings.