Graph Engineering

Confidence 0.75 · 7 sources · last confirmed 2026-09-18

The practice of structuring agentic work as an explicit workflow graph — nodes the engineer defines, edges the engineer specifies, state passed from node to node — rather than as a single agent looping until it satisfies a goal. Its stated payoff is predictability, debuggability and control; its precondition is that you already know the workflow before you build it.

The corpus’s primary source is [[2026-09-03-thurium-wang-google-cloud-graph-engineering-101|Google Cloud’s Graph Engineering 101]] (3 Sep 2026), the third video in an AI Builder Essentials arc that also supplied the wiki’s harness boundary and its loop-failure taxonomy.

The three-layer nesting

The most useful thing this vocabulary does is stack cleanly, which the wiki’s three Google Cloud sources establish one layer at a time:

LayerWhat it isSource
Harness”everything around the model including its tools, memory and guardrails”Baugues & Thurium, Jul 2026
Loop”the cycle of the agent running inside that harness” — reason, decide, pick a tool, re-pick, until the goal is metThurium & Wang, Aug 2026
Graph”the organization chart” — a workflow of agent nodes and function nodes (“which can be a deterministic logic”)Thurium & Wang, Sep 2026

A loop can be a node in a graph — “you can put loop as part of the graph” — so these are not competing architectures but nested scopes. See agent-harness.

The mechanism that makes traversal meaningful is shared state: “when you traverse the graph, you’re passing memory or information down the graph to the next node.” In ADK this exists both across agents in a multi-agent system and across nodes in a workflow.

The three patterns

Named in the primary source’s PR-review walkthrough, and worth holding as the working vocabulary:

  • Fan-out — split one task into n independent sub-tasks and run them simultaneously. “Much faster than running them one after another and waiting for them all to finish.”
  • Join — a synchronisation node that “waits for the slowest processing” and then synthesises the parallel results into something downstream nodes can consume.
  • Router — a conditional edge: take the input, dispatch to a different sub-agent, specialist or sub-workflow. The example’s condition is risk-asymmetric — failure routes to a specialised fixer agent; success routes to human approval.

Wang’s own summary of what these are: “we’re taking basic principles of control flow and applying them to AI engineering.”

What it is not — three disambiguations

Not a knowledge graph. They share a word and nothing else: “Knowledge Graph emphasises on the data, and Graph Engineering emphasises on the behavior — basically what goes in, in what order, and then what happens next.” GraphRAG sits on the data side too. See knowledge-graphs.

Not loop engineering. A loop is one running cycle to a goal; a graph is nodes and edges. The stated threshold is a complexity heuristic rather than a principle — a one-paragraph summary is loop work, “a 50 page PDF with a bunch of graphics” is graph work — and it is the same example the loop-failure taxonomy used for its complexity overflow mode.

Not an agent swarm — and this is the load-bearing contrast:

Graph engineeringAgent swarm
Who specifies the structurethe engineer, node by node, including “how the data looks”nobody: “each agent just gets its own personality and that’s it”
Node’s relationship to history”an agent that’s at a certain node doesn’t need to know what happened before”agents hold the problem, not a position in it
Payoff”really good predictability, debuggability and control”flexibility
Fits”problems that can be really strictly defined""not all problems are easily defined like that”

Against the wiki’s earlier resolution

On 2026-09-03 agent-fleet-management folded squad, fleet, orchestra and graph into a single primitive — role-specialised parallelism — and instructed future ingests not to proliferate sub-concepts. Three of those four still belong together; graph does not, and the correction comes from the vendor that supplied the wiki’s use of the term.

The distinction is not pedantic. Parallelism is one pattern inside a graph (fan-out), while what makes a graph a graph is engineer-specified control flow over history-free nodes. A squad of role-specialised agents passing a task between them is much closer to the swarm end of the primary source’s own contrast — the thing graph engineering is defined against. Resolve squad / fleet / orchestra to agent-fleet-management; resolve graph here.

When it is actually justified

The vendor’s answer is a precondition plus a size heuristic: use a graph “if you already know the workflow ahead of time,” and when the output is large and structured. The corpus’s evidence suggests the precondition is the real criterion and the size heuristic is decoration:

  • Tran & Kiela find single agents beat multi-agent decompositions at equal token budgets, so decomposition is not free and must be paid for by something other than parallel wall-clock.
  • MAST locates a whole failure category in inter-agent misalignment — the seam between nodes is where multi-agent systems break.
  • Google Cloud’s own loop taxonomy is notable for a vendor arguing against reaching for the more elaborate architecture by default: escalate only when a single loop stops coping.

The convergent reading: a graph is justified when the decomposition is knowable in advance, not when the task is big. Nobody has measured where that threshold sits.

The deflationary reading, from the vendor

The primary source closes with its own presenters undercutting the vocabulary:

Thurium: “Are we just reinventing data structures and algorithms for the agentic age?” — Wang: “Yeah, maybe.” — Thurium: “Maybe hash table engineering or stack engineering.” — Wang: “Abstraction-maxxing.”

Held next to GitHub’s developer advocates the day before and Kokane’s argument that most harness engineering is mature systems design on a new substrate, this is the first-party version of the same claim. The terms are real and first-class in vendor materials; they are also, by their own authors’ admission, control flow with new nouns. Weight them accordingly.

agent-harness (the layer this sits inside) · ai-agents (agent nodes) · agent-fleet-management (the adjacent primitive it is not) · multi-agent-failure-modes (what breaks at the seams) · knowledge-graphs (the homonym) · agentic-pull-requests (the worked example’s domain) · agent-oversight-and-delegation (the router’s human-approval branch)

Debates and supersession

  • Determinism versus capability is an untested trade. The claim that history-free nodes buy debuggability is plausible and matches isolation prescriptions elsewhere in the corpus. But BFCL finds memory and long-horizon state to be the unsolved half of agentic capability, and MAST finds the seams between agents to be where systems fail. Discarding history has a cost that no source here quantifies.
  • The swarm side of the contrast is unsourced. The wiki holds no primary material on agent swarms as a named architecture — the category enters the corpus only as the thing graph engineering is defined against. Treat swarm as a placeholder until a real source lands.
  • Every claim on this page is vendor-asserted. Predictability, debuggability and control are stated, never demonstrated: no latency numbers, no cost comparison, no evidence that a graph-structured pipeline outperforms a single loop on the same job. Confidence is capped at 0.75 for that reason, per the vendor-source rule.
  • Open: does the harness/loop/graph layering survive contact with other vendors’ vocabularies, or is it Google Cloud’s local taxonomy? GitHub’s hosts say “there’s no standard again — no standards in AI”, which is at least consistent with the latter.

The loop layer, in running code (added 2026-09-15)

The Agent Factory demonstrates what sits inside a graph’s nodes, which this page has described but never shown. Billy builds three harnesses by hand:

PatternShapeWhen
Linearinspect → output → exit, same flow every time”great if you need some level of determinism”
Closed loopapply edit → run tests → capture why it failed → feed the error into memory → repeat (capped ~5)code editing; run until resolved
ADK with guardrailsmemory compaction plus a block_destructive_commands check before executioncustomisation without writing the plumbing

The detail worth carrying is in the closed loop: the payload fed back is the failure text, not a boolean — “we’re going to see why it failed and feed that error back into memory.” A graph’s function nodes carry deterministic logic; its agent nodes carry loops of this shape, and the linear/closed distinction is the same determinism-versus-adaptivity trade this page draws between graph and swarm, one level down.

The framing is also a decent argument for the page’s existence: “you can just use the framework, but a great engineer will really understand the framework. Building a harness yourself is how you understand what’s happening under the hood when things break.”

Three planning levels, and a contract at each (added 2026-09-18)

Rashad (PyData, September 2026) grades a harness’s planning into three levels: a fixed workflow the engineer writes (by hand or in a graph library), bounded planning where the tasks are known and the model manages the queries routed to them, and a fully dynamic task graph built at runtime, which coding agents and open-ended research need. It is the same escalation as Thurium and Wang’s loop-to-graph path. He adds one design point: each level needs a contract, and at the dynamic end the contract governs which graph operations the model may perform. Versioned plans let a user steer mid-run without restarting; a fixed workflow cannot be edited in flight.