Skip to content
The TAI Labs community is now on Skool
TAI Labs
All articles

12 August 2026 · 22 min read

Harness Engineering: Build the Car, Not Just the Engine

A coding agent is the model plus everything you build around it. From the ratchet habit and Week 1 API checklist through Ralph loops, HaaS, MCP trust, and where harnesses are going next.

By Aki Wijesundara, PhD, TAI Labs

From the archive. Originally published 2026-08-12. Tool details, examples and offers reflect that publication date. See our current guides for newer material.

The short version

  • Agent = Model + Harness. If you are not the model, you are the harness.
  • A decent model with a great harness beats a great model with a bad harness. Same weights, different scaffolding, different outcomes.
  • Default to skill issue, not model issue. Most failures are configuration, hooks, context, permissions, or evaluation.
  • The ratchet: every real mistake becomes a permanent rule. Earn each line in AGENTS.md, each hook, each schema check from a failure you observed.
  • Work backwards from behaviour. If you cannot name the behaviour a component delivers, it should not be there.
  • Start with the component harness: schema, timeout, retries, contract-preserving fallback, logging. Then grow into workflow control, permissions, evaluation, and state.
  • Harnesses do not shrink as models improve. They move. Drop dead scaffolding; add scaffolding for the new ceiling.
  • Build on harness APIs (HaaS), not raw completions. Iterate from a v0.1. Treat MCP tool text as trusted input.

Same model. Same prompt. Different outcome. In demos that looks like luck. In production it is almost always the system around the call.

We have spent two years arguing about models. Which one is smartest, which one hallucinates less, which one writes cleaner code. That debate is fine as far as it goes, and it misses the other half of the system. The model is one input into a running agent. The rest is the harness: prompts, tools, context policies, hooks, sandboxes, subagents, feedback loops, and recovery paths wrapped around the model so it can actually finish something.

The model is the engine. The harness is the car. You ship the car.

Viv Trivedy’s one-liner, popularised in Addy Osmani’s write-up of the discipline, does most of the work: Agent = Model + Harness. If you are not the model, you are the harness. A raw model is not an agent. It becomes one once a harness gives it state, tool execution, feedback loops, and enforceable constraints.

This post is how we teach that at TAI Labs: first as a production API component in Week 1 of the AI Engineering Bootcamp, then as a full agent loop in our Agent Harness lightning lesson. The framing draws on the emerging public discourse Osmani pulls together (Trivedy, HumanLayer, Anthropic, Dex Horthy, Birgitta Böckeler). The examples are the ratchet habit we want engineers to practise on their own codebases.

What is a harness, really?

A harness is every piece of code, configuration, and execution logic that is not the model itself. Concretely it includes:

  • System prompts, CLAUDE.md, AGENTS.md, skill files, and subagent prompts
  • Tools, skills, MCP servers, and their descriptions
  • Bundled infrastructure: filesystem, sandbox, browser
  • Orchestration: subagent spawning, handoffs, model routing
  • Hooks and middleware for deterministic execution: compaction, continuation, lint and typecheck
  • Observability: logs, traces, cost and latency metering
coding agent = AI model(s) + harness

That equation, articulated by Trivedy and echoed by HumanLayer, is where the work lives. The debate over the left-hand side is loud. Most of the actual leverage sits on the right. Simon Willison reduces the loop: an agent is a system that runs tools in a loop to achieve a goal. The skill is in the design of both the tools and the loop.

If that sounds like a lot of surface area, it is. And it is your surface area, not the model provider’s. Claude Code, Cursor, Codex, Aider, Cline: these are all harnesses. The model underneath is sometimes the same. The behaviour you experience is dominated by what the harness does.

Skill issue, not model issue

There is a pattern engineers fall into. The agent does something dumb, the engineer blames the model, and the blame gets filed under “wait for the next version.”

The harness-engineering mindset rejects that default. The failure is usually legible. The agent did not know a convention, so you add it to AGENTS.md. The agent ran a destructive command, so you add a hook that blocks it. The agent got lost in a forty-step task, so you split planner from executor. The agent kept “finishing” broken code, so you wire a typecheck back-pressure signal into the loop.

HumanLayer’s line is the right default: it is not a model problem, it is a configuration problem. Harness engineering is what happens when you take that seriously.

There is a striking data point that shows up in both Trivedy’s anatomy of a harness and HumanLayer’s writing, and that Osmani highlights: on Terminal Bench 2.0, the same Claude Opus 4.6 scores far differently depending on the harness around it. Viv’s team moved a coding agent from Top 30 to Top 5 by changing only the harness. Models get post-trained against the harness they ship with. Moving them into a harness with better tools for your codebase, a tighter prompt, and sharper back-pressure can unlock capability the original harness was leaving on the floor.

The gap between what today’s models can do and what you see them doing is largely a harness gap.

That is the opposite of the “just wait for GPT-6” narrative. In our lightning lesson we make the same point at agent altitude: almost every agent failure lives in workflow control, context, permissions, evaluation, or state, not in the prompt sentence you keep rewriting.

Three altitudes

Building with LLMs is three jobs. Teams that collapse them into “prompting” keep fixing the wrong layer.

01 PromptWhat you sayRole, task, constraints, output format. Precise, not clever.
02 ContextWhat the model can seeSystem brief, examples, retrieved docs, history, tools, and deliberate omissions.
03 HarnessWhat surrounds the modelLoop, validation, retries, timeouts, fallbacks, hooks, streaming. Turns an unreliable call into a component.

Two altitudes inside the harness itself:

  • Component harness: one model call behind an API. Schema, timeout, retries, safe fallback. This is Week 1.
  • Agent harness: a loop that plans, acts, observes, and continues. Workflow budgets, per-step context, permissions, evaluation, state. This is the lightning lesson.

A useful split for both: the prompt describes what good output looks like. The harness decides what happens when the output is not good.

“Always return valid JSON”Schema validation
“Never delete anything”Tool allowlist / pre-tool hook
“Remember what the user said”State + durable files
“Do not make things up”Retrieval + grounding check
“Stop when you have enough”Termination condition / step budget
“Be concise, use British spelling”Stays in the prompt

If you are writing “always” or “never” in a system prompt, you are describing a control, not an instruction. Every “never” is a guarantee you are hoping for instead of building. Failing prompts grow because each incident adds a sentence, the sentences compete, and past a point they degrade each other. The harness scales the other way: it is code, and code composes.

The ratchet: every mistake becomes a rule

The most important habit in harness engineering is treating agent mistakes as permanent signals. Not one-off stories to laugh about. Not “bad runs” to retry. Signals.

Osmani’s formulation is the discipline in one sentence: anytime you find an agent makes a mistake, you take the time to engineer a solution such that the agent never makes that mistake again. You only add constraints when you have seen a real failure. You only remove them when a capable model has made them redundant. Every line in a good AGENTS.md should be traceable back to a specific thing that went wrong.

This is also why harness engineering is a discipline rather than a framework you download. The right harness for your codebase is shaped by your failure history.

Worked ratchet examples

These are the kinds of slips we see when teams ship agents against real repos and APIs. Each one has a harness change, not a longer prompt.

Agent ships a PR with a commented-out or skipped test (.skip(, xit()Add to AGENTS.md: never skip tests; delete or fix. Pre-commit / pre-PR hook greps the diff and fails the loop. Reviewer subagent treats skips as blockers.
Agent runs rm -rf, force-pushes, or drops a tablePre-tool hook blocks destructive bash. Write tools require an allowlist per step. Planning steps cannot hold write tools.
Agent “finishes” with type errorsAfter every edit, run typecheck. Success is silent. Failure injects the error text back into the loop so attempt two is different from attempt one.
Endpoint returns plausible but invalid JSONSchema via Pydantic / response_format. Then rules (required fields, ranges). Model judgement last and only when schema and rules are not enough.
Agent declares done while the feature is brokenSeparate generator from evaluator. Agree a done-condition (sprint contract) before code. Self-grading skews positive; a second agent or a test suite does not.

HumanLayer’s principle for hooks is the one we want burned in: success is silent, failures are verbose. If typecheck passes, the agent hears nothing. If it fails, the error text enters the loop and the agent self-corrects. That makes the feedback loop almost free in the common case and directly actionable when something goes wrong.

# Pseudocode: after-edit hook
result = run("npm run type-check")
if result.exit_code != 0:
    # verbose failure enters the next model turn
    inject_into_context(result.stderr)
    continue_loop()
# success: say nothing

Work backwards from behaviour

The framing from Trivedy that is most useful when you are actually designing a harness: start from the behaviour you want and derive the harness piece that delivers it. If you cannot name the behaviour a component exists to deliver, it probably should not be there.

Work with real data durablyFilesystem + Git (workspace, offload, rollback, shared coordination)
Write and execute codeBash + code execution (general-purpose tools, not one gadget per action)
Safe execution and sane defaultsSandbox, allowlists, pre-installed runtimes, test CLIs, headless browser
Remember and update knowledgeMemory files (AGENTS.md), web search, MCP docs tools
Stay coherent as context fillsCompaction, tool-output offload, progressive skills, full session reset + hand-off
Finish long-horizon workStep budgets, plan files, continuation loops, planner ≠ evaluator

Walk the pieces. Each one exists because the model cannot deliver that behaviour on its own.

Filesystem and Git: durable state

The filesystem is the most foundational primitive, and it tends to be underrated because it is boring. Models can only directly operate on what fits in context. Without a filesystem you are copy-pasting into a chat window, and that is not a workflow.

Once you have a filesystem, the agent gets a workspace to read data, code, and docs; a place to offload intermediate work instead of holding it in context; and a surface where multiple agents and humans can coordinate through shared files. Adding Git on top gives you versioning for free, so the agent can track progress, roll back errors, and branch experiments. Most of the other harness primitives end up pointing at the filesystem for something.

Bash and code execution: the general-purpose tool

The main agent loop today is a ReAct loop: the model reasons, takes an action via a tool call, observes the result, and repeats. A harness can only execute the tools it has logic for. You can try to pre-build a tool for every possible action, or you can give the agent bash and let it build the tools it needs on the fly.

Willison’s take: agents already excel at shell commands; most tasks collapse to a few well-chosen CLI invocations. Harnesses still ship focused tools, but bash plus code execution has become the default general-purpose strategy. It is the difference between teaching someone to use a single kitchen gadget and handing them a kitchen.

Sandboxes and default tooling

Bash is only useful if it runs somewhere safe. Running agent-generated code on your laptop is risky, and a single local environment does not scale to many parallel agents.

Sandboxes give agents an isolated operating environment. The harness connects to a sandbox to run code, inspect files, install dependencies, and verify work. You can allow-list commands, enforce network isolation, spin up new environments on demand, and tear them down when the task is done.

A good sandbox ships with good defaults: pre-installed language runtimes and packages, Git and test CLIs, a headless browser for web interaction. Browsers, logs, screenshots, and test runners are what let the agent observe its own work and close the self-verification loop. The model does not configure its execution environment. Deciding where the agent runs, what is available, and how it verifies its output are all harness-level calls.

Memory and search: continual learning

Models have no additional knowledge beyond their weights and what is currently in context. Without the ability to edit weights, the only way to add knowledge is through context injection.

The filesystem is again the primitive. Harnesses support memory file standards like AGENTS.md that get injected on every start. As the agent edits that file, the harness reloads it, and knowledge from one session carries into the next. Crude, but effective continual learning.

For knowledge that did not exist at training time (new library versions, current docs, today’s data), web search and MCP tools like Context7 bridge the cutoff. Bake those primitives into the harness rather than leaving them to the user.

Claude Code, Cursor, Codex, Aider, Cline are all harnesses. The model underneath is sometimes the same. The behaviour you experience is dominated by what the harness does.

The component checklist (Week 1)

Before you grow into agents, ship a reliable reasoning component. A single model-backed endpoint with a weak harness fails in three boring ways: malformed output, timeouts and rate limits, and occasional hard failures. The harness we teach in Week 1 answers those with structured output, timeouts and retries, and graceful degradation that never returns a bare 500.

Each item below is a failure you have either already seen or will see on the first real traffic day.

1. Schema-validated structured output

Putting “respond in JSON” in the prompt is a hope. Schema enforcement at the API is a guarantee. Define a Pydantic model (or equivalent), pass it as response_format, and work with a typed object.

from pydantic import BaseModel

class Answer(BaseModel):
    answer: str
    sources: list[str]
    confidence: float  # 0.0 to 1.0

response = client.chat.completions.parse(
    model="gpt-5.4-mini",
    messages=[...],
    response_format=Answer,
)
result = response.choices[0].message.parsed

Downstream code should not check whether fields exist. They always do. That is what turns a demo into a component.

2. Client timeout set

Slow calls must fail cleanly, not hang forever. Set an explicit timeout. Twenty seconds is a common starting point for interactive APIs; tune to your SLA.

3. Retries on transient failures

Networks flap. Providers rate-limit. Occasional 5xx responses happen. Configure bounded retries with backoff. Do not retry forever, and do not retry non-idempotent side effects blindly.

client = OpenAI(
    api_key=os.getenv("OPENAI_API_KEY"),
    timeout=20.0,
    max_retries=3,
)

4. Safe fallback that preserves the response contract

Never crash the client with a bare 500. On the bad path, return the same shape you promised on success: a valid Answer with a clear message and confidence=0.0.

5. Errors logged, not swallowed

Graceful degradation is not silence. Log the real exception. Users get a calm response. On-call gets the truth.

6. Stream for humans; validate for machines

Streaming is core to chat UX. Machine consumers still need validated structured objects. Use the right mode for the consumer.

7. Optional upgrades

  • Secondary model fallback if the primary fails, same schema.
  • Same runtime everywhere via containers so local, staging, and production do not invent their own harness bugs.

Minimal code shape

The interesting part is not the happy path. It is that success and failure share one contract.

@app.post("/ask")
async def ask_question(question: str):
    try:
        response = client.chat.completions.parse(
            model="gpt-5.4-mini",
            messages=[{"role": "user", "content": question}],
            response_format=Answer,
        )
        return response.choices[0].message.parsed
    except Exception:
        # Log the real error. Still return HTTP 200 + valid Answer.
        return Answer(
            answer="Sorry, I couldn't answer that right now.",
            sources=[],
            confidence=0.0,
        )

Same Answer shape on success and failure. Clients never see a bare 500.

When the harness has to grow

A single call with schema, timeout, retry, and fallback is enough for a reliable reasoning component. Agents need more because they loop. Almost every agent failure lives in one of five layers.

  1. Workflow control: what runs, in what order, and when it stops. Max iterations, max tool calls, max wall-clock time. Not vibes. Continuation patterns (including Ralph-style “intercept exit, re-inject goal into a fresh window, read state from disk”) exist because early stopping and multi-window incoherence are harness problems.
  2. Context: assemble what the model needs for this step. Do not append forever. Rebuild per step with current state, retrieved facts, and the last relevant results.
  3. Permissions: allowlist tools per step. Read is not write. A planning step and an execution step should not share the same tools.
  4. Evaluation: Schema first (cheap), then rules, then model judgement last. Anthropic’s long-running harness work is explicit that separating generation from evaluation outperforms self-evaluation, because agents skew positive when grading their own work.
  5. State: attempts, failures, confirmed facts, workflow position. Without state, a retry is the same failing call again at three times the cost. Feeding state.log_failure into build_context is often the only reason attempt two can succeed.

Four out of five is not eighty percent of a harness. The missing layer is usually where the incident comes from.

Failure to fix map

Agent looped foreverWorkflow controlStep budget
Answered from stale dataContextRebuild context per step
Deleted the wrong recordPermissionsPer-step allowlist
Plausible but wrong answerEvaluationRules tier after schema
Forgot a confirmed fact mid-taskStatePersist confirmed facts
Retry failed identicallyState + contextFeed the error back into context

Hooks: “I told it” vs “the system enforces it”

Hooks are what separate a polite system prompt from an enforceable system. A hook runs at a lifecycle point: before a tool call, after a file edit, before commit, on session start. That is the right place for things the agent should never forget but often does: typecheck after edit, block destructive bash, require approval before opening a PR, auto-format on write so the agent does not waste tokens on whitespace.

Keep AGENTS.md short. HumanLayer keeps theirs under sixty lines for a reason: every line competes for attention. Pilot’s checklist, not style guide. Earn each line from a failure. The same discipline applies to tools: ten focused tools beat fifty overlapping ones, and every MCP description is trusted text in the prompt.

Battling context rot

Context rot is the observation that models get worse as the window fills. Harnesses are delivery mechanisms for good context engineering. Three techniques show up repeatedly in production harnesses, with a fourth for the longest jobs:

  • Compaction: summarise and offload older context before the API errors.
  • Tool-call offloading: keep head and tail of huge tool outputs; put the full blob on disk for on-demand reads.
  • Skills with progressive disclosure: do not load every tool and MCP at startup; reveal instructions when the task needs them.
  • Full context resets: Anthropic’s long-running harness work is explicit that compaction alone was not enough for some long tasks; tear the session down and rebuild from a compact hand-off file.

We go deeper on usable vs advertised windows, agents-by-accumulation, and application-layer controls in Your context window is smaller than you think.

Long-horizon execution: Ralph loops, planning, verification

Autonomous long-horizon work is the hardest thing to get right. Today’s models suffer from early stopping, poor decomposition of complex problems, and incoherence as work stretches across multiple context windows. The harness has to design around all three.

Ralph-style continuation

Osmani restates the Ralph Loop in harness terms: a hook intercepts the model’s attempt to exit and re-injects the original prompt into a fresh context window, forcing the agent to continue against a completion goal. Each iteration starts clean but reads state from the previous one through the filesystem. It is a simple trick for turning a single-session agent into a multi-session one, and it is the kind of primitive you would never derive from “just use a smarter model.”

# Pseudocode: continuation hook
if model_tries_to_exit() and not goal_complete(state_on_disk):
    reset_context()
    inject(original_goal)
    inject(read_progress_from_filesystem())
    continue_loop()

Planning on disk

Planning is when the model decomposes a goal into a sequence of steps, usually into a plan file on disk. The harness supports this with prompting and reminders about how to use the plan file. After each step, the agent checks its work via self-verification: hooks run a pre-defined test suite and loop failures back with the error text, or the model reviews its output against explicit criteria.

Planner, generator, evaluator, and the sprint contract

Anthropic’s long-running harness work is explicit that separating generation from evaluation into distinct agents outperforms self-evaluation, because agents reliably skew positive when grading their own work. Osmani calls the related pattern the sprint contract: the generator and evaluator negotiate what “done” actually means before code gets written. Writing down the done-condition before starting catches more scope drift than most prompt changes.

Agree done before you generate. Self-grading is not a substitute for a second agent or a test suite.

AGENTS.md, tool choice, and MCP trust

The flat markdown rulebook at the root of your repo is still the single highest-leverage configuration point, because it lands in the system prompt every turn. Conventions go here: package manager, test framework, formatting, “never touch /legacy,” “always use our logger.”

Two hard-won lessons from HumanLayer, which Osmani amplifies:

  • Keep it short. Under about sixty lines. Every line competes for attention. More rules make each rule matter less. Pilot’s checklist, not style guide.
  • Earn each line. Rules should trace to a specific past failure or a hard external constraint. If they do not, they are noise. Ratchet; do not brainstorm.

Same discipline applies to tools. Each tool’s name, description, and schema gets stamped into the prompt every request. Ten focused tools outperform fifty overlapping ones because the model can hold the menu in its head.

There is also a real security concern: tool descriptions populate the prompt, so any MCP server you install is trusted text the model will read. A sloppy or malicious MCP can prompt-inject your agent before you have typed anything. Treat MCP installation like dependency installation with code execution rights.

What this looks like in production

Osmani points to Fareed Khan’s estimated breakdown of Claude Code’s architecture as the clearest public picture of a mature harness. You do not need the diagram to use the layer map. Almost every concept above shows up as a named component in a shipping product.

InputUI, session manager, permission gateWho can act; what needs approval
KnowledgeSkill registry, context compressor, task graph, memory storeContext injection and compaction
IntegrationMCP runtime and external serversTrusted tool text and remote capabilities
ExecutionTool dispatch, streaming runtime, prompt cacheWhere bash and MCP both plug in
OutputVerified task resultsDone means verified, not claimed
ObservabilityEvent bus, background executorTraces, cost, latency
Multi-agentSubagent spawn, mailboxes, FSM, worktree isolatorContext firewalls between agents

The master agent loop sits at the centre. Destructive-action hooks sit behind the permission gate. Subagent context firewalls are the multi-agent layer. Khan’s argument, via Osmani, is the same as Trivedy’s: Claude Code’s trajectory is about the harness at least as much as about the model underneath it.

Harnesses do not shrink, they move

The naive story is that better models make harnesses obsolete. If the model can plan, delete the planner. If the model is coherent at long horizons, delete context resets.

Anthropic’s observation, which Osmani foregrounds, is sharper: as models improve, the space of interesting harness combinations does not shrink. It moves. Opus 4.6 largely killed the context-anxiety failure mode that made Sonnet 4.5 wrap up work early as it approached what it thought was its context limit. A whole class of anxiety-mitigation scaffolding becomes dead code.

But the ceiling moved with the model. Tasks that were unreachable are in play, and they bring new failure modes: multi-day memory policy, multi-agent coordination, evaluators for design quality in generated UIs. Anthropic’s line: every component in a harness encodes an assumption about what the model cannot do on its own. When the model gets better at something, that component becomes load-bearing for nothing and should come out. When the model unlocks something new, new scaffolding is needed to reach the new ceiling.

The model-harness training loop

Trivedy names the feedback loop explicitly. A useful primitive is discovered in a harness, standardised into the product, used when training the next generation of models, and the next model improves at using that primitive. Cycle repeats.

Today’s agent products are post-trained with harnesses in the loop. The model gets specifically better at the actions the harness designers think it should be good at: filesystem operations, bash, planning, subagent dispatch. That is why Opus 4.6 feels different inside Claude Code than inside someone else’s harness, and why changing a tool’s logic sometimes causes strange regressions. A genuinely general model would not care whether you used apply_patch or str_replace, but co-training creates overfitting.

Two practical implications. A harness is a living system, not a config file you set once. And the “best” harness is not necessarily the one the model was trained inside; it is the one designed for your task. Viv’s Top 30 to Top 5 Terminal Bench jump remains the clearest proof point Osmani cites.

Harness-as-a-Service

Trivedy’s other contribution is the HaaS framing: Harness-as-a-Service. We are moving from building on LLM APIs (which give you a completion) to building on harness APIs (which give you a runtime). The Claude Agent SDK, the Codex SDK, and the OpenAI Agents SDK all point the same way. You get the loop, the tools, the context management, the hooks, and the sandbox primitives out of the box, and you customise them.

The default path used to be: build your own loop, wire up your own tool-calling, handle your own conversation state, invent your own approval flow. The default path now is: pick a harness framework, configure it along the four pillars (system prompt, tools, context, subagents), and put the rest of your effort into domain-specific prompt and tool design.

That is what makes “skill issue” tractable. You are not rebuilding an agent from scratch every time something goes wrong. You are tuning a configuration surface that is already well-factored.

Good agent building is an exercise in iteration. You cannot do iterations if you do not have a v0.1.

Week 1’s FastAPI component is that v0.1 for many of our students: a real loop surface (even if the “loop” is one call) with schema, timeout, retry, and fallback. Then you ratchet.

Where this is going

Look at the top coding agents side by side (Claude Code, Cursor, Codex, Aider, Cline) and they look more like each other than their underlying models do. The models are different. The harness patterns are converging. That is not an accident. It is the industry finding the load-bearing scaffolding that turns a generative model into something that can ship.

Trivedy’s open problems, which Osmani flags as the exciting frontier:

  • Parallel agents on a shared codebase: orchestration, worktree isolation, and conflict handling as first-class harness concerns.
  • Self-repairing harnesses: agents that analyse their own traces to identify and fix harness-level failure modes, not just task failures.
  • Just-in-time assembly: harnesses that dynamically assemble the right tools and context for a given task instead of being pre-configured at startup. That last one is where harnesses stop being static config and start becoming something closer to a compiler.
# Component harness (one call)
# ✓ validated output shape
# ✓ timeout set
# ✓ retries on transient failures
# ✓ safe fallback on the bad path
# ✓ errors logged not swallowed
# ✓ stream for humans; validate for machines
#
# The ratchet habit
# ✓ every real failure becomes a permanent rule
# ✓ hooks: success silent, failures verbose
# ✓ AGENTS.md: short, earned, traceable to incidents
#
# Agent harness (the loop)
# ✓ workflow control · per-step context · permissions
# ✓ evaluation (schema → rules → judgement) · state
# ✓ planner ≠ evaluator · done-condition before code
# ✓ Ralph-style continuation when goals span windows
#
# Living system
# ✓ drop scaffolding the model no longer needs
# ✓ add scaffolding for the new ceiling
# ✓ prefer harness SDKs (HaaS); iterate from v0.1
# ✓ treat MCP descriptions as trusted prompt text

Prompt tuning will keep improving. It still will not replace the code that surrounds the call. Build the harness first. Then argue about temperature.

Want to build this properly?

We teach the component-level harness in Week 1 of the AI Engineering Bootcamp, and the agent-level version in a dedicated lightning lesson.

AI Engineering Bootcamp → Reliable APIs, RAG, agents, evals, and memory

The Agent Harness lightning lesson → Why prompts fail in production, and what to build instead

Sources

  • Addy Osmani: Agent Harness Engineering (April 2026). Primary synthesis this essay engages; Terminal Bench / Top 30→Top 5 harness point, Ralph Loop, HaaS, and production-layer framing as reported there.
  • Viv Trivedy: Anatomy of an Agent Harness (Agent = Model + Harness; behaviour → harness derivation; model-harness training loop; HaaS; open problems).
  • HumanLayer: configuration / “skill issue” framing; short AGENTS.md; success-silent hooks; MCP tool-description trust.
  • Anthropic engineering: long-running agent harness design; compaction vs full context resets; generator/evaluator separation; “harnesses move” as models improve; context-anxiety failure modes.
  • Dex Horthy: tracking the harness-engineering pattern as it emerges (credited in Osmani).
  • Birgitta Böckeler: harness from the user’s side (credited in Osmani).
  • Fareed Khan: estimated Claude Code architecture breakdown (via Osmani).
  • Simon Willison: agent as tools in a loop to achieve a goal.
  • TAI Labs: Your context window is smaller than you think
  • TAI Labs: AI Engineering Bootcamp Week 1: Reliable Reasoning Components
  • TAI Labs: The Agent Harness lightning lesson