TAI Labs

RAG for Engineers Who Already Know Databases

Retrieval is a query plan. You already know how to debug one.

Lightning Lesson

Dr. Aki Wijesundara · TAI Labs
with Manu Jayawardana

24 minutes of content, 6 of questions. Open on the promise: if you understand indexes and query plans, you already own the mental model, and most RAG explainers waste your time re teaching you embeddings. Say there is no beginner recap coming. Flag the two chat moments, one now and one near the end.
01
TAI Labs

What RAG actually is

The model does not have your data. You fetch the rows, then you generate.

Index time · write pathDocumentsChunkEmbedVector indexQuery time · read pathQuestionRetrievePrompt + chunksthe JOINLLMAnswerSame shape as a query: scan an index, fetch rows, then compute. Generation is the last SELECT.
The model never searches. Retrieval does. Generation is just the last SELECT.
1 min. Name the two paths once. Index time is the write. Query time is the read. Point at Prompt plus chunks and say that is the JOIN. Then: we are not recapping embeddings. The rest of this session is the five failure stages on that retrieve box.
02
TAI Labs

Retrieval Is a Query Plan

Five stages. Every one has a database twin you already know.

StageYou know this asIn retrieval
1Row granularityChunking
2Index selectionDense vs sparse
3WHERE clauseMetadata filter
4ORDER BYReranking
5Replica lagIndex freshness
Your retrieval is bad. One thing to check first. Type 1, 2, 3 or 4 in chat.1 chunking · 2 embeddings · 3 the filter · 4 stale index
Find the stage, not the symptom.
The spine of the session and the slide people screenshot. Run the poll properly, read three answers out loud by name. Costs 40 seconds, buys the room. Almost everyone says 1 or 2. Do not correct anyone, say "hold that thought" and advance. The reveal is the next slide.
03
TAI Labs

Where the Analogy Breaks

Most of you said chunking or embeddings. Those fail loudly. The two that cost you incidents fail silently.

01

There is no equals sign.

Ask for something not in your corpus and you still get confident nearest neighbours back.

02

WHERE does not compose.

A predicate that narrows a B tree can sever a graph traversal and collapse recall.

03

There are no transactions.

Your index is a stale replica with no lag metric and no alert.

A database tells you when it found nothing. Retrieval never does.
This slide earns the advanced framing. Slow down here. If you are behind later, take it from slide 5 or 7, never this one. The line to land is the callout, say it twice. Card 3 sets up slide 8, so flag that you will come back to it.
04
TAI Labs

Chunking Is Schema Design

Chunk size is row granularity. Same tradeoff as deciding how far to normalise.

Chunk 1 · headers, no numbersChunk 2 · numbers, no headersUnitsRevenueMargin1,20438.57.298631.06.81,51244.98.1chunk boundaryBoth retrieve.Neither answers.

Three fixes, cheapest first

01

Split on document structure, not character counts.

Free, and it kills most boundary failures.

02

Contextual retrieval.

Prepend where the chunk sits in its parent, then embed. Anthropic measured 35 percent fewer retrieval failures.Anthropic

03

Index the small chunk, return the large one.

Match precisely, answer with full context.

Chunk boundaries are schema decisions. Never make them on character counts.
The diagram is the failure: a table splits across two chunks, numbers in one, headers in the other, both retrieve and neither answers. Everyone here has normalised a schema badly, so this lands fast. The tradeoff to narrate: small chunks match precisely but shred context, large chunks keep context but dilute the signal until the embedding averages toward nothing. Anthropic numbers are from their own engineering blog, state them flat. Costs about one dollar per million document tokens with prompt caching. If asked about late chunking: embed the whole document first, pool per chunk after, same goal, cheaper at query time.
05
TAI Labs

You Would Never Run One Index

You would never put a single index on a table and call it done.

IndexGood atBlind to
DenseParaphrase, intent, conceptsExact strings, error codes, SKUs
BM25Rare and exact termsAnything phrased differently
HybridBoth, if you fuse by rankMuch less, at two queries plus fusion
Fusing by score fails, the scales differ. Reciprocal rank fusion ignores scores: each list gives 1 over (60 plus rank).
If a user can paste an identifier into your search box, dense only retrieval will fail and the score will still look healthy.
Say hybrid is blind to much less, at the cost of two queries and a fusion step. It is not blind to nothing: it still inherits both indexes’ recall ceilings and adds a fusion tuning problem, and someone in this room will push back if you say otherwise. The error code example is what lands. Ask them to picture a customer pasting ERR_CONN_REFUSED_5031 into support search. Cosine similarity sees a generic error shaped token sequence and returns generically error shaped documents. BM25 sees a rare term and goes straight to it. Contextual embeddings plus contextual BM25 took Anthropic to 49 percent fewer failures, neither alone gets close. Why 60: empirical, from the 2009 paper, it damps the top ranks so one confident list cannot dominate. This slide compresses if you are behind.
06
TAI Labs

The WHERE Clause Trap

Scope to one customer or date range. In SQL that helps. Here your engine picks which failure you get.

No filterentrytargetWith filterentrytarget

Same query. Same data. The filter removed the path, not just the rows.

Filter shapeRecall
Most shapesAbove 97 percent
One broad value90.8 percent
AND across two39.7 percent

Qdrant's own benchmark on their own engine.

Benchmark recall with your filters applied, never without.
Click once to apply the filter: the right panel fades in and the path breaks after one hop. The two ways a filter fails: filter after searching discards non matches, so a selective filter returns fewer results than you asked for; filter before searching builds the eligible set first, and with too few eligible nodes the paths between them break and the walk dead ends early. Highest value slide for anyone running multi tenant retrieval and almost nobody measures it. The instinct to break: adding WHERE can only narrow, so it can only help. Those numbers are Qdrant benchmarking their own engine, say so. Yours will differ, the shape of the failure will not. Practical fix if asked: partition or namespace by tenant instead of filtering by tenant, so the constraint becomes which index you query. Some engines build filterable index structures and do neither naive approach, so the real instruction is to benchmark yours.
07
TAI Labs

Cheap Recall, Then Expensive Precision

Same shape as a query planner. Fast scan for candidates, expensive work on a small set.

cheap · optimise for recallHybrid retrieve · 200 candidatesexpensive · optimise for precisionCross encoder rerank · score all 200Send to the model · top 20anything missed here can never be recovered
A bi encoder asks are these about the same subject. A cross encoder asks does this answer the question.
Recall is bought at retrieval and can never be recovered later. Precision is bought at rerank.
The funnel: retrieve wide with hybrid search, fuse, take 50 to 200 and optimise purely for recall. Rerank narrow with a cross encoder that reads query and document together, not precomputed vectors. Pass the top 20 to the model; Anthropic found 20 beat both 5 and 10. The distinction that matters: the bi encoder embedded that document months ago with no idea what would be asked, the cross encoder sees both at once. Reported cost is 100 to 150 ms on CPU for 50 candidates, 30 to 50 ms on GPU, and 5 to 10 nDCG points of gain. Say reported. Stack the whole pipeline for the big number: contextual retrieval plus hybrid plus rerank is 67 percent fewer retrieval failures, 5.7 down to 1.9. If behind schedule, this compresses to 90 seconds: just the two stage shape and the callout.
08
TAI Labs

Your Index Is a Stale Replica

Vector similarity has no temporal dimension. A deprecated doc and the current one retrieve with equal confidence.

Three ways it goes stale

01

The source changed.

Track source timestamp against index timestamp. Trivial, and nobody does it.

02

You upgraded the embedding model.

New vectors, different geometry. Full reindex or nothing.

03

You changed chunking logic.

Same content, different vectors. A preprocessing change is an index migration.

Latency normal. Throughput normal. No errors. No alerts.
Set a shelf life per document type and monitor quality on a fixed query set, because nothing else will tell you.
This is card 3 from slide 3 cashing out, call that back. The line that gets a reaction: an embedding model upgrade is a full index migration and everyone treats it as a config change. One reported case saw recall drift 0.92 to 0.74 with nothing in the codebase changing, say reported. Practical: keep 50 real queries with known correct documents and rerun weekly. That is the whole monitoring story and it takes an afternoon.
09
TAI Labs

Diagnose This

Four real symptoms. Which stage broke? Type A, B, C or D in chat.

AFine on concepts. Falls apart on error codes and part numbers.

Stage 2. Dense only. Add BM25 and fuse.

BRight document, but the answer is missing a number from a table.

Stage 1. A chunk boundary split the values from their headers.

CScoped to one customer's docs and quality fell off a cliff.

Stage 3. The filter broke the traversal.

DGreat in March. Nobody touched the code. Bad now.

Stage 5. Freshness. "Nothing changed" is the tell.

All four are invisible in your logs. Latency and error rate look perfect through every one.
This is the payoff, give it the full three minutes, do not rush it to protect the closing slide. Reveal one at a time with the arrow key and read chat between each. D gets the biggest reaction because everyone has lived it: nothing in your code changed, the world did. If chat is quiet, call on someone by name, it is late and people need permission to type.
10
TAI Labs

The Retrieval Triage Table

SymptomStageCheck first
Misses exact identifiers2Is BM25 in the mix
Right doc, incomplete answer1Split on structure
Worse when filtered3Benchmark recall with filters on
Relevant but badly ordered4Add a cross encoder
Was fine, now is not5Source timestamps, model version
Start measuring one thing tomorrow: recall at 20 on 50 real queries, with your filters applied.

This is week 3 of the AI Engineering Bootcamp with TAI Labs. Link in the chat.

Questions

Tell them explicitly to screenshot the table. One spoken mention of the bootcamp, then straight to questions, do not sell. Maven adds every signup to the waitlist and emails the recording in 48 hours, the work is done. Held for Q&A: GraphRAG, multi hop and agentic retrieval, ColBERT, quantisation, fine tuning embeddings.
11