TAI Labs
AI Engineering · 30-minute course session

Graph &
Loop
Engineering

Structure your Agent's work. Make the workflow explicit.
Topic
Agent architecture
Format
30 minutes · live teaching
Speaker
Dr. Akshika Wijesundara
The promise of this session

Don't just prompt an agent. Engineer its work.

By the end of this session, you should be able to look at an agentic task and decompose it into nodes, state, decisions and loops — then explain why each part exists.

The mental shift: an LLM call generates an answer. An engineered agent manages a workflow.
01 · Start with the problem

What happens when we simply tell an LLM to “do the job”?

The naive mental model

Prompt → LLM → Answer

The model decides what to do next, what tools to call, what information to retain, and when it is finished.

That can work for a demo. It becomes difficult to reason about when the task contains branching, failures, retries or quality checks.

The engineering mental model

Work → Graph → State → Control

The workflow is designed explicitly. The model still makes decisions, but those decisions happen inside boundaries we can inspect, test and operate.

Question: “Why did the agent take this step?” should have an observable answer.

Core idea: prompting controls behaviour locally; graph engineering controls the shape of work globally.
02 · The mental model

An agent is a graph of work.

NODEa unit of work EDGEwhat happens next STATEwhat work carries forward CONTROLconditions · retries · termination
Remember: nodes perform work, state carries work, edges route work, and control logic decides whether work continues.
03 · Nodes

A node should have one understandable job.

01

Research

Find and collect evidence. The node should not also decide the final answer.

Work unit
02

Generate

Turn available context into a candidate output with an explicit contract.

Work unit
03

Validate

Check the candidate against a definition of good. It can pass, fail or request another iteration.

Decision input
Design test: if a node needs a paragraph to explain everything it does, it probably contains multiple responsibilities.
04 · Edges

Edges are the agent's routing rules.

Research

Gather evidence

Generate

Produce candidate

Validate

Evaluate quality

Finish

Return accepted result

Linear edge

Always continue

Research → Generate

Useful when the workflow is deterministic.

Conditional edge

Route based on state

Validate → Finish if quality passes.

Validate → Research if evidence is insufficient.

05 · State

State is the shared memory of the workflow.

state = { "task": "Research a topic and produce a reliable report", "research": [], "draft": "", "feedback": "", "iteration": 0, "status": "running" }
Why explicit state matters: it makes intermediate work inspectable. A failed agent run should leave evidence of what it knew, what it produced and which decision moved it forward.
06 · Put it together

The first complete agent graph.

STARTtask Researchcollect evidenceupdate state Generatedraft outputupdate state Validatequality gatechoose next edge PASS → ENDaccepted output FAIL → RESEARCHiterate the graph becomes iterative when validation can route work backward
07 · Why linear chains break

Real work does not look like A → B → C → D.

Real tasks contain uncertainty. Information may be missing. A tool may fail. The answer may be below the quality bar. A human may need to intervene.

Engineering question: where does each uncertainty go in the graph?
08 · Conditional edges

Let state determine the next step.

Evaluate evidenceupdate state Sufficient?a decision over state YES NO Synthesize Search again
Good conditional edges are explicit: define the condition, define the destination, and make the state that triggered the decision observable.
09 · Loop engineering

Loops turn workflows into agents.

PLANdecide what to attempt EXECUTEuse tools / models OBSERVEcollect result / errors EVALUATEis it good enough? FAIL → improve and try again PASS → DONE
A loop is not “run forever.” It is a controlled feedback mechanism: attempt → observe → judge → improve.
10 · The basic agent loop

The loop is useful because the agent can learn from its own work.

01

Plan

Choose the next action from the current state. The plan can be revised when new information arrives.

02

Execute + observe

Call tools, retrieve information, generate output, and capture both success and failure signals.

03

Evaluate

Compare the result to a quality contract. The result is not “finished” merely because the model returned text.

Key distinction: generation produces a candidate. Evaluation determines whether the candidate earns the right to leave the graph.
11 · Loop control

Every loop needs an exit strategy.

Without termination conditions, an agent can waste time, tokens and money while repeatedly “improving” an answer that was already good enough.

Production rule: a loop must be bounded by quality, iteration count, failure handling, and resource budget.
12 · Four loop controls

Control the loop from four directions.

01

Success condition

A measurable quality threshold. Example: quality_score ≥ 0.90.

Stop when good
02

Maximum iterations

A hard ceiling such as max_iterations = 5. The system must be able to say “enough”.

Bound effort
03

Failure condition

Named terminal or recovery states: tool failure, invalid output, missing evidence, policy violation.

Fail safely
04

Resource budget

Token, time, tool-call and monetary budgets. Reliability without operational control is not production engineering.

Control cost
13 · Loop implementation

The loop should be visible in the code.

while state["status"] != "done": plan(state) execute(state) observe(state) evaluate(state) if state["quality"] >= 0.90: state["status"] = "done" elif state["iteration"] >= 5: state["status"] = "stopped" elif state["budget_exceeded"]: state["status"] = "stopped"
Important note: don't hide termination inside a vague “agent.run()”. Make the stopping logic inspectable and testable.
14 · Reusable graph patterns

Five patterns cover most agent workflows.

01
Sequential · A → B → C. Use when every step is required and ordered.
02
Branching · classify → route. Use when different inputs require different workflows.
03
Loop · generate → evaluate → improve. Use when quality emerges through iteration.
04
Fan-out / fan-in · parallel work → synthesis. Use when independent evidence can be gathered concurrently.
05
Supervisor / worker · coordinator → specialists → coordinator. Use when tasks need specialized capabilities.
15 · Pattern 1 + 2

Sequential when order matters. Branching when context matters.

Sequential

Extract → Transform → Store

Every record must pass through the same stages. The graph is simple because the business process is simple.

Extract→Transform→Store
Branching

Classify → route

The classifier chooses which specialized path is appropriate.

Request→Classify→SQL / Web / Code
Rule: don't introduce a loop just because you can. Use the simplest graph that accurately represents the work.
16 · Pattern 3 + 4

Loops create refinement. Fan-out creates breadth.

LOOP · DEPTH Generate Critique improve until pass FAN-OUT / FAN-IN · BREADTH Task Search A Search B Search C SYNTH
17 · Pattern 5

Supervisor / worker: coordinate without making one agent do everything.

SUPERVISORplan · delegate · integrate Research workerevidence gathering Analysis workerspecialized reasoning Validation workerquality / policy checks
Use sparingly: multi-agent graphs add coordination overhead. Split work when specialization or isolation genuinely improves the system.
18 · Worked example

Let's engineer a research agent.

Task: “Research a topic and produce a reliable report.” We will evolve the design one engineering decision at a time.

Watch the transformation: prompt → workflow → graph → loop → controlled system.
19 · Research agent · Step 1

Before the agent searches, make it write questions.

Naive approach

“Go research X and write a report.”

The model jumps straight into searching and writing. You cannot see what it decided to research, or why.

If the report is weak, you do not know whether the plan was bad or the research was bad.

Engineered approach

Step 1 = PLAN only

First node turns the task into a short list of research questions and saves them in state.

  • What is the definition?
  • What are the benefits / limits?
  • What evidence exists?
  • What needs a primary source?

It does not search. It does not write the report. Later nodes do that, using the plan.

Why this matters: you can inspect the plan, change it, or test it — before any searching starts.
20 · Research agent · Step 2

Research can fan out, then fan in.

PLANresearch questions DISPATCHparallel evidence tasks Source research Counter-evidence Current state SYNTHESIZEevidence → draft
Fan-out is not automatically “more agents.” The key property is independent work that can be performed concurrently and merged deterministically.
21 · Research agent · Step 3

Validation turns a generated report into an engineered output.

PASS

Ship the report

  • Claims have supporting evidence.
  • Important claims use stronger sources.
  • Contradictions are resolved or disclosed.
  • Required sections are complete.
FAIL

Route back to research

  • Missing evidence → search.
  • Conflicting sources → investigate.
  • Weak claim → find stronger source.
  • Incomplete section → targeted research.
Important: validation should produce structured feedback, not just “this looks bad.” Feedback becomes the next iteration's state.
22 · Complete architecture

The research agent is now a controlled graph.

START PLANquestions RESEARCHevidence SYNTHESIZEcandidate VALIDATEquality gate DONE RESEARCH AGAIN
The loop is now meaningful: validation failure identifies a specific information gap, and the graph routes the agent to address that gap instead of blindly regenerating the report.
23 · State through the graph

State changes as work moves.

01
After planningstate.questions[] is populated; iteration = 0.
02
After researchstate.sources[] and evidence[] contain collected material; missing[] may still exist.
03
After synthesisstate.draft contains the candidate report and claim/source mappings.
04
After validationstate.quality, feedback[] and next_action tell the graph where to go next.
Observability payoff: a trace can reconstruct not just what the model said, but how the workflow evolved.
24 · Engineering reality

The hardest question is not “what did the agent output?” It is “why did it stop?”

Production agents need traces that show node execution, state changes, edge decisions, tool failures, retries, iteration count and termination reason.

Debugging principle: if the graph cannot explain the run, the architecture is too opaque.
25 · Production considerations

A good graph is designed for operation, not only execution.

01

Observability

Trace nodes, inputs, outputs, tool calls, state changes, edge decisions and termination reasons.

02

Reliability

Define retries, fallbacks, refusal paths and human checkpoints before the first production incident.

03

Cost + latency

Loops and parallel branches multiply model/tool usage. Budget them explicitly.

04

Evaluation

Test individual nodes and full graph trajectories. A workflow can fail even when each node looks reasonable in isolation.

26 · Engineering principles

Five rules to take into your next agent build.

01
One responsibility per node. Small work units are easier to test and replace.
02
Make state explicit. Intermediate work should be inspectable and serializable.
03
Make transitions explicit. Conditions belong in the graph, not only in a prompt.
04
Every loop needs termination. Quality, iterations, failures and budgets all matter.
05
Observe the graph. You should be able to explain every important trajectory.
27 · The big picture

Four concepts turn an LLM call into an engineered agent.

01

Graph

Defines the topology of work: nodes and edges.

02

State

Carries the context and intermediate artefacts between steps.

03

Decisions

Conditions determine which path the workflow takes.

04 · Loops provide controlled iteration so the system can respond to feedback instead of assuming the first attempt is correct.
28 · Takeaway

Graph the work. Loop the uncertainty.

Nodes make work modular. State makes it visible. Edges make decisions explicit. Loops make improvement possible. Controls make the whole thing safe to operate.

Final question for the room: “If your agent fails tomorrow, can you point to the node, state and edge that explain why?”
29 · Quick exercise

Before you build your next agent, draw this first.

Five questions
  • What are the nodes?
  • What state moves between them?
  • What are the conditions?
  • Where is the loop?
  • What stops it?
One-minute design test

Can another engineer understand the workflow without reading the prompt?

If yes, your architecture is becoming explicit.

If no, move more of the reasoning into the graph, state model and control logic.

Thank you · Questions

Build agents as systems of work.

Dr. Akshika Wijesundara

Graph & Loop Engineering · Structure Your Agent's Work

01 / 32