1 / 22
Free Lesson

AI Engineering

Context Engineering

The skill that separates senior AI engineers from prompt hobbyists.

Dr. Aki WijesundaraDr. Aki Wijesundara The AI Internship

Agenda

1What is AI Engineering? — Where it sits in the landscape and why it's different.
2What is context engineering? — The definition, why it matters, and the guiding principle.
3The anatomy of good context — System prompts, tools, and examples done right.
4How agents get context — Pre-loaded vs just-in-time retrieval.
5Long-horizon tasks — Compaction, note-taking, and sub-agent architectures.

What is AI Engineering?

A new discipline

AI engineering is the practice of building production systems on top of foundation models.

1Not ML engineering. You don't train the models.
2Not pure software engineering. You're working with a non-deterministic component.
3Not prompt engineering. That's one skill inside it.
4The job: turn a probabilistic model into a reliable product.

The job is shifting

2023

Write the right prompt.

A discrete task. Optimise one string.

2026

Design the right context.

A system. Curate what the model sees on every call.

The AI engineer's stack: models, context, tools, evals, observability. Of these, context is where most teams lose. That's where we spend the rest of this session.

Context Engineering

What is context engineering?

Deciding what the model sees, when, and in what structure, on every single call.

Context is the set of tokens included when you sample from the model. Everything in the window: system prompt, tools, retrieved docs, message history, tool results, memory, MCP outputs.

Context engineering is the strategies for curating and maintaining the optimal set of those tokens during inference.

Source: Anthropic, "Effective context engineering for AI agents", Sep 2025

Prompt vs context engineering

Prompt engineering vs context engineering diagram

Source: Anthropic, "Effective context engineering for AI agents"

Context is finite

Two ideas to internalise before anything else.

Context rot

As tokens in the window grow, the model's ability to recall information from it degrades. Every model. Some gracefully, but the curve is universal.

Attention budget

Transformers create n² pairwise relationships. More tokens, thinner attention. Models were trained mostly on shorter sequences, so long-range reasoning is genuinely harder.

Bigger context windows do not solve this. They just delay the problem. Treat context as a precious, finite resource with diminishing returns.
Find the smallest possible set of high-signal tokens that maximise the likelihood of your desired outcome.

Everything else in this session is a tactic in service of this principle.

The anatomy of good context

Find the right altitude

Calibrating the system prompt spectrum diagram

Source: Anthropic, "Effective context engineering for AI agents"

Minimum viable set

Tools are the contract between the agent and the world. Bad tools poison context.

1Self-contained, robust to error, unambiguous.
2Token-efficient outputs. Don't return 8KB of JSON when 200 tokens of summary works.
3No overlapping functionality between tools.
If a human engineer can't tell which tool to use, an AI agent can't either.

Curated, not exhaustive

Few-shot prompting still works. But don't dump every edge case into the prompt.

Don't

Stuff a laundry list of every possible rule and edge case into the prompt.

Do

Pick a small set of diverse, canonical examples that show the shape of the desired behaviour.

For an LLM, examples are the pictures worth a thousand words.

How agents get context

Pre-loaded vs just-in-time

Pre-inference (RAG)

Embed everything ahead of time, retrieve top-K at query time, stuff into context.

Trade-off: fast, but you're guessing what's relevant before the agent has started thinking.

Just-in-time

Agent holds lightweight references (file paths, queries, IDs). Pulls data into context dynamically as it needs it.

Trade-off: slower, but mirrors how humans work. Used by Claude Code.

Progressive disclosure: the agent discovers context layer by layer. File names, folder structure, timestamps all become signals it can reason over. Working memory stays small.

Long-horizon tasks

When the task outgrows the window

Codebase migrations, deep research, multi-hour agents. Three techniques to know.

TechniqueWhat it doesBest for
Compaction Summarise the conversation, restart with the summary. Extensive back-and-forth.
Structured notes Agent writes notes outside the window, reads them back later. Iterative work with milestones.
Sub-agents Spawn focused sub-agents with clean contexts, return distilled summaries. Complex research, parallel exploration.

Compaction

When the conversation approaches the window limit, summarise it and start fresh with the summary.

Claude Code does this. It preserves architectural decisions, unresolved bugs, and key implementation details while dropping redundant tool outputs.

The art is in what you keep vs what you discard.
Lightest touch: clearing old tool results once they've been used. Tuning tip: maximise recall first (capture everything that matters), then tighten for precision.

Structured note-taking

The agent writes notes to a persistent store outside the context window and reads them back later. Simple patterns work: a TODO list, a NOTES.md file, a memory tool.

Claude playing Pokémon: tracks objectives across thousands of game steps. "For the last 1,234 steps I've been training my Pokémon in Route 1, Pikachu has gained 8 levels toward the target of 10." Maps of explored regions. Strategies. All in external notes. After a context reset, it reads its own notes and keeps going.

Sub-agent architectures

Instead of one agent holding all state, spawn focused sub-agents with clean context windows.

10,000+ tokens

Each sub-agent can burn this much exploring deep inside its own focused task.

1-2K tokens

What it returns to the lead agent: a condensed, distilled summary.

Clear separation of concerns. Detail stays inside sub-agents. The lead synthesises. Anthropic's multi-agent research system substantially outperformed single-agent setups on complex research tasks.
Find the smallest set of high-signal tokens that maximise the likelihood of your desired outcome.

RAG, agents, memory, MCP, compaction, sub-agents. Every one of them is a tactic in service of this principle.

Context as a finite, precious resource is not going away.

What's next?

Keep building with us

Two programmes — pick the one that fits where you are.

For engineers & builders

AI Engineering Bootcamp

4 weeks. Build RAG pipelines, LangChain apps, AI agents, and deploy production-ready systems from scratch. Capstone project + certificate.

Enroll — use code ENG35 →

For PMs & founders

AI Native Product Management Certificate

4 weeks. Ship a full-stack AI app, run AI evals like a senior PM, and deploy a live autonomous agent. No coding background required.

Enroll — use code PM35 →

Taking both? Get both for $800 — email us at admin@snapdrum.com

Alumni from

Google  ·  Meta  ·  Apple  ·  OpenAI  ·  Amazon  ·  NVIDIA  ·  McKinsey  ·  BCG

Questions?  tailabs.ai