
By the end of this session, you should be able to look at an agentic task and decompose it into nodes, state, decisions and loops — then explain why each part exists.

The model decides what to do next, what tools to call, what information to retain, and when it is finished.
That can work for a demo. It becomes difficult to reason about when the task contains branching, failures, retries or quality checks.
The workflow is designed explicitly. The model still makes decisions, but those decisions happen inside boundaries we can inspect, test and operate.
Question: “Why did the agent take this step?” should have an observable answer.


Find and collect evidence. The node should not also decide the final answer.
Work unitTurn available context into a candidate output with an explicit contract.
Work unitCheck the candidate against a definition of good. It can pass, fail or request another iteration.
Decision input
Gather evidence
Produce candidate
Evaluate quality
Return accepted result
Research → Generate
Useful when the workflow is deterministic.
Validate → Finish if quality passes.
Validate → Research if evidence is insufficient.


Real tasks contain uncertainty. Information may be missing. A tool may fail. The answer may be below the quality bar. A human may need to intervene.



Choose the next action from the current state. The plan can be revised when new information arrives.
Call tools, retrieve information, generate output, and capture both success and failure signals.
Compare the result to a quality contract. The result is not “finished” merely because the model returned text.
Without termination conditions, an agent can waste time, tokens and money while repeatedly “improving” an answer that was already good enough.

A measurable quality threshold. Example: quality_score ≥ 0.90.
Stop when goodA hard ceiling such as max_iterations = 5. The system must be able to say “enough”.
Bound effortNamed terminal or recovery states: tool failure, invalid output, missing evidence, policy violation.
Fail safelyToken, time, tool-call and monetary budgets. Reliability without operational control is not production engineering.
Control cost


Every record must pass through the same stages. The graph is simple because the business process is simple.
The classifier chooses which specialized path is appropriate.


Task: “Research a topic and produce a reliable report.” We will evolve the design one engineering decision at a time.

The model jumps straight into searching and writing. You cannot see what it decided to research, or why.
If the report is weak, you do not know whether the plan was bad or the research was bad.
First node turns the task into a short list of research questions and saves them in state.
It does not search. It does not write the report. Later nodes do that, using the plan.




Production agents need traces that show node execution, state changes, edge decisions, tool failures, retries, iteration count and termination reason.

Trace nodes, inputs, outputs, tool calls, state changes, edge decisions and termination reasons.
Define retries, fallbacks, refusal paths and human checkpoints before the first production incident.
Loops and parallel branches multiply model/tool usage. Budget them explicitly.
Test individual nodes and full graph trajectories. A workflow can fail even when each node looks reasonable in isolation.


Defines the topology of work: nodes and edges.
Carries the context and intermediate artefacts between steps.
Conditions determine which path the workflow takes.
Nodes make work modular. State makes it visible. Edges make decisions explicit. Loops make improvement possible. Controls make the whole thing safe to operate.

If yes, your architecture is becoming explicit.
If no, move more of the reasoning into the graph, state model and control logic.
Dr. Akshika Wijesundara
Graph & Loop Engineering · Structure Your Agent's Work