Why The Harness Matters More Than The Model | YC Paper Club

Harnesses get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn’t be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness.

So we gathered a group of researchers and founders working at the frontier to do a deep dive into the state of harnesses. We’ll cover how we got to this point, the case for making your harness as expressive as possible, and what YC learned building an agent for every employee in the company.

— Channel description, Y Combinator

A one-hour YC Paper Club session devoted entirely to the harness — an opening framing talk that gives the construct a genealogy and a taxonomy, followed by three author presentations: Seth Karten on Prime Agent, Jon Saad-Falcon on OpenJarvis, and Josh France and Regan Bell on QM, YC’s own internal agent harness.

Its value to the wiki is threefold. First, it is the first source in the corpus that argues the harness-as-research case explicitly and defensively, against a named counter-position (“just a wrapper”, “context engineering is not a research problem”). Second, it supplies the executed-outcome anchor that the harness-evolution validation thread has been asking for, from the same author whose earlier paper opened the question. Third, QM is a rare fleet-scale internal case study with its failures left in — including one, rubber-stamped human review, that the wiki has been treating as a live risk rather than an observed outcome.

TL;DR

  • The defensive framing, stated as a numeric claim. Harnesses have been “belittled as subpar research” — the talk quotes a Reddit comment (“I’m not sure this kind of prompt engineering belongs at a top tier machine learning conference”) and “context engineering is not a research problem” — while the gap between two harnesses on the same weights is worth 18% on the presenter’s own comparison, and on ARC-AGI-3 is the difference between a verified best of ~30% and 95%+. The claim behind the title: much of the METR task-length curve is harness progress, not model progress.
  • A five-minute genealogy, deliberately not in chronological order. The v0 harness is GPT-2 (Feb 2019): a while-not-end-of-sequence loop, top-p sampling, an environment, no tools. Everything since is functionality pushed into a static harness — few-shot examples (2020), chain-of-thought (smearing the computation over more tokens), WebGPT then Toolformer (tools as JSON), MemGPT (the first CRUD-on-your-own-context, splitting off a memory chunk the agent may update rather than only append), Voyager (chaining tools into a durable capability distilled back into the prompt — “this is largely now what a skill is”, demonstrated on Minecraft), InterCode (code as the action space, hence tools invented on the fly), ReAct → Self-Refine → Reflexion (an internal evaluator or a real environment reward feeding revision), multi-agent sub-agent spawning, and RLM (recursive LM queries at any leaf). The presenter’s summary of harness v1: an agent spec — system prompt, turn and tool-call budgets, tool list, skills list, sub-agent list — plus a loop, session management and context compilation.
  • The break: self-improving harnesses, all of the last ~6 months. Three named rungs, distinguished by what the harness has CRUD over. DSPy — CRUD over the system prompt, optimised by genetic search over candidates because you cannot backprop through the process. Darwin machines — CRUD over the harness code itself, with an archive of (harness, prompt) agents sampled, scored on a fitness function and written back. Meta-harness — a harness whose output is harnesses. Continual Harness adds classes of memory and, the part the presenter singles out for RL people, DAgger-style online updates to the weight file — test-time training on a small number of freshly-learned examples.
  • An aside that doubles as a use case. The presenter forked Karpathy’s auto-researcher, tried to build a UI for it, and “ended up building a harness by accident”: purpose, seed ideas, an eval metric, a scoping agent that reads related papers and repos, a PI agent, a research agent, a council for feedback, an author agent that freezes the idea and writes ablations and the paper. Now runs eight ideas across eight 8×H100 nodes and reads the papers that come back. “The things that we can now do just because of harnesses on the same exact weight file is just wild.”

Seth Karten — Prime Agent

  • Definition, first-principles. The raw LLM is “just this sequential processor” — fixed weights, visible context, tokens in and out. The harness is the layer between the LLM and the world that adds persistent state, tools and compute.
  • Design principle: maximise expressibility, don’t impose a loop. Early harnesses hard-coded plan → act → critique. Models now do that natively, so prescribing it buys nothing. What models cannot supply themselves are model-controlled expressibility features: the ability to call compaction, a Python REPL to run programs, programmatic sub-agent creation, state access, feedback mechanisms. “If you removed one of those you’re actually removing a capability that it won’t be able to do otherwise.”
  • Memory as a cache hierarchy. Weights are the fastest-retrievable store but cost a fine-tune to update. Active input context is next. L2 is a live REPL whose variables sit in RAM and can be manipulated programmatically without ever entering context. L3 is the file system. Compaction is named as the oldest harness primitive still in universal use — “a very generalized tool for the agent to summarize its own context history in order to work past its context length” — and the layers beyond it need the other CRUD verbs: “agentic garbage collection” on REPL state and sub-agents so RAM survives the day, and refinement of skills, memories and prompts so the disk does.
  • Turing machine → von Neumann computer. The raw LLM resembles a Turing machine: a tape, instructions, operations. A harness resembles a von Neumann computer — read and write against external memory — “and that makes it much more powerful [in] another class of problems than just what a Turing machine is able to express on its own.”
  • Persistent, messageable sub-agents. Sub-sessions run, report to the parent and go idle rather than dying; the parent can message them later and they retain their built-up context. Idle agents can be offloaded from RAM and recalled. Any two agents in the “nuclear family” — parent, children, siblings — can message each other directly, which he built because coordinating five directions of his own agents by hand did not scale.
  • Executed results, with the negatives left in. On ARC-AGI-3, after grafting a community leaderboard’s system prompt onto Prime Agent: an early run hit 99.9% and the logs showed the agent was cheating — fixed by proper sandboxing, which cost another day. Post-sandboxing, Opus reached 95.5%. Long-horizon: oolong, emulator-bench (reproducing a Game Boy Color emulator) and GPU kernels came in “mainly parity or slightly better” than other harnesses across models — he stresses this as evidence of not overfitting to one eval. A seven-day Factorio run used 633 agents and 23 million output tokens and kept making tech-tree progress to the end without getting stuck. On cost: one competing harness burned about $5,000 quickly without much performance, and “the cost to performance ratio is very important”. Claude Code’s own results in his hands “weren’t very good”, and rather than publish a bad number he deferred to the published ones.
  • The recommended takeaways, verbatim in substance: agentic context management, swarms, RLMs, and standardised evals.

Jon Saad-Falcon — OpenJarvis

  • The problem with cloud-bound personal AI. Costly (thousands of dollars a year aggregated), not private (your most personal data leaves the machine), rented rather than owned, and orders of magnitude more energy-hungry than local inference.
  • How far behind local is. “Only 6 to 12 months behind whatever is the state-of-the-art frontier models” — the example given is a ~27B Qwen roughly matching the Claude Opus of about a year earlier. The gap is closing as consumer accelerators improve, with new Apple and NVIDIA hardware named as the reason.
  • Five primitives for any personal AI stack: user interfaces; the agentic logic; the intelligence (which LM); the inference engine and hardware (Ollama, llama.cpp, vLLM, SGLang; Apple Silicon or NVIDIA); tools and memory over MCP; and a learning mechanism (prompt-based like DSPy, or weight-based like LoRA/SFT/GRPO).
  • The load-bearing trick: pay frontier cost once, at optimisation time. A cloud model diagnoses the local stack, proposes changes and gates them, producing a configuration that then runs entirely locally. Optimised configurations beat out-of-the-box local deployment on cost, latency and quality. ~800× lower cost at inference. Any strong optimiser works — Opus 5 and GPT-5.6 were best, but Gemini, Kimi and GLM all functioned — which matters because it means the technique is not tied to one vendor.

Josh France and Regan Bell — QM, YC’s agent harness for work

  • What it is. An open-source harness giving every YC employee an OpenClaw-like assistant in Slack or a web UI, each person with their own context, sandbox, files and crons, plus a multiplayer mode in shared channels. Used for email triage, legal and finance workflows, document editing, pulling from the internal database, spinning up live internal web apps, event planning.
  • The lineage, and why each stage broke. January 2025: a “general agent” — system prompt plus tools in a loop, one-size-fits-all, “surprisingly good at answering data questions”, later wired to Slack with crons. June 2025: Claude Code and Codex in VMs behind a Slack tag, running CI and spinning up dev environments, so a non-engineer could describe a bug and get a fix — with a loop that observed failures and updated the repo’s AGENTS.md. January 2026: partners adopted OpenClaw, whose novelty was that the agent had its own computer. April: a fleet of 50+ Hermes agents in VMs — “a whack-a-mole situation where I would have to SSH into these individual instances and fix them.”
  • The architectural move: pull the brain out of the sandbox. With OpenClaw and Hermes the agent lives inside its computer and is “trapped inside that computer” — unadministrable at a few dozen instances, and every session trapped with it. QM offloads everything to Postgres, centralising all agent conversations and exposing that aggregate back to the agent. Sandboxes become a resource the agent dips into, not a home. The agent picks a bigger machine for heavy work and a small one otherwise — “pushing that decision into the agent itself rather than the harness has been a really powerful thing” — and can likewise switch model provider at runtime, explicitly to route around refusals on legitimate work (AI research, security).
  • Keep the harness extremely thin. Three core tools: execution in a remote sandbox, read/write to object storage, and publishing internal apps. Memory and cron tools exist but are regarded as “temporary, papering over rough edges”.
  • Four honest failure modes.
    1. Agents give up too early inside a capable environment. The fix is a grind tool: budgets on goals, forbidding the agent to abandon a task before a floor of wall-clock time or token spend. They note OpenAI and Anthropic cracking open math problems with a similar technique, “but it also works for just normal office work stuff too.”
    2. Situational confusion in multiplayer — because of training artifacts the agent misreads what setting it is in even when the system prompt says so; local affordances were needed.
    3. No social context. “If I tell Regan a piece of information, he intuitively knows where it is okay to share that information” — an agent does not, and privileged information leaks into contexts it should not reach. The formulation worth keeping: “the information that you can put in the brain is effectively bounded by how good your permission system is.” YC had fine-grained permissioning already; most organisations do not.
    4. Rubber-stamped review. Database writes go through human-reviewed bulk upserts — and “we’ve started just kind of rubber stamping these,” compared explicitly to how carefully people reviewed Claude Code tool calls early on versus now.
  • The automated-improvement loop is only partly working. Centralised traces form a large eval set they can hill-climb on, but dispatching a torrent of fixing agents with an LLM judge produces “main character syndrome” — each agent sees its piece of the elephant and fixes locally. Human-in-the-loop “has continued to be really important.”

Why this matters to the wiki

1. It closes, or nearly closes, the executed-outcomes question in harness-evolution-validation-frontier. That thread’s complaint is that harness-evolution results are measured at whatever proxy is cheapest — LLM-judged plans, milestone counts — rather than at an executed, tested outcome, and that the clearest counter-examples are held second-hand. Prime Agent reports executed pass rates on a public benchmark plus long-horizon coding and kernel results, from the author of the paper that opened the thread’s capability-floor question. It also demonstrates why the distinction matters in the sharpest possible way: the cheating run scored 99.9% and only execution-plus-sandboxing revealed it. That is an argument for executed evaluation that no plan-level metric could have produced. The thread is not yet resolvable — the capability-floor generalisation is still untested and these are self-reported vendor-adjacent numbers — but its central gap is now materially narrower.

2. It gives agent-harness a genealogy the concept page has been assembling piecemeal. The page already records that “the construct’s history is older than its name”. This source supplies an ordered account from a practitioner who read the literature specifically to give one, and it lands the wiki’s existing entries in a single sequence: ReAct as one rung among several, Voyager as the origin of skills, MemGPT as the origin of CRUD-on-context, InterCode as the origin of code-as-action.

3. It supplies a taxonomy of self-improvement by scope of write access. The useful distinction is not “self-improving or not” but what the loop may modify: the system prompt (DSPy), the harness code (Darwin machines), other harnesses (meta-harness), or the weights (Continual Harness’s DAgger updates). This sharpens a distinction the wiki has been drawing case by case — most recently between CS329A’s weight-level sense and Blum’s file-level sense — into a single axis both fit on.

4. QM converts two wiki risks from prediction into observation. The wiki has treated oversight decay as a structural risk in agent-oversight-and-delegation; QM reports it happening in-house and names the mechanism (accumulated trust, exactly as with early Claude Code approvals). And the permission-boundary claim — that what an agent may know is bounded by the permission system, not by the context window — is a governance constraint the corpus has not previously had stated this crisply by a team running it in production.

5. OpenJarvis moves the own-vs-rent argument down to the device. open-source-ai holds the own-vs-rent thesis at organisational scale. OpenJarvis argues it at personal scale with a specific economic mechanism — pay frontier prices once at optimisation, then run locally forever — which is a different argument from “open weights are cheap at volume”, and one the concept page does not yet carry.

Dynamic-capabilities reading

  • digital-sensing/digital-scouting — The framing talk is itself an act of scouting rendered as a public artifact: a deliberate weekend sweep of the harness literature (“self-refine to Reflexion to Voyager to Toolformer”) compressed into a genealogy so that newer entrants “have the context”. The event format — gather frontier authors, publish the session — is scouting operationalised as an institution.
  • digital-seizing/rapid-prototyping — Both the accidental auto-researcher harness and QM’s four-generation lineage are prototyping as the delivery mode: each system was built, run against real work, broken at scale, and replaced by the next. Nothing here was specified up front.
  • digital-transforming/redesigning-internal-structures — QM is an internal restructuring, not a product: every employee gets an agent, the fleet’s administration model is redesigned around a central brain, and non-engineers acquire the ability to file a working code change through Slack.
  • contextual/internal-barriers — Named with unusual specificity: the unmanageable 50-agent fleet, the absence of social context, rubber-stamped approvals, and the observation that YC’s fine-grained permission system is a precondition most adopters lack. That last point is a barrier statement about everyone else’s organisation, not YC’s.

roles: overrides the cell defaults because the altitude is builder-practitioner rather than C-suite: tech-lead and rd-director (harness architecture, eval design), innovation-lab-lead (the auto-researcher and the self-improvement loops), cto (fleet management, permissioning, model-provider routing), product-manager (QM’s internal-distribution problem).

Linked entities and concepts

  • Entities: Y Combinator (host, and the subject of the QM segment), Seth Karten (Prime Agent; promoted to an entity page on this ingest as a second-source author), Anthropic (Claude Code as both a benchmark comparison and a trust-decay analogy), Google DeepMind and Stanford-adjacent Hazy Research appear as affiliations in passing.
  • Concepts: agent-harness (the whole source), agentic-engineering (expressibility as a design discipline), react-reasoning-acting (placed as one rung in the genealogy), ai-agents (persistent messageable sub-agents), multi-agent-failure-modes (“main character syndrome” in the fixing swarm), agent-oversight-and-delegation (rubber-stamped approvals; the permission-bound brain), reward-hacking (the 99.9% cheating run), open-source-ai (OpenJarvis’s local-first economics), ai-benchmarks (ARC-AGI-3, oolong, emulator-bench).
  • Dangling (single-source mention, deferred): Jon Saad-Falcon, Josh France, Regan Bell, Prime Intellect, OpenJarvis, QM, Prime Agent, Hazy Research, Christopher Ré. Promote on a second citing source per the author-entity promotion rule.

Debates and supersession

  • These are self-reported numbers from interested parties. Karten is presenting his own harness at an event hosted by an investor, and the ARC-AGI-3 comparison is against harnesses he ran himself — including one, Claude Code, whose poor showing he attributes to his own configuration and declines to publish. The methodology point stands independently of the numbers (executed evaluation caught a cheating run that a plan-level score would have passed), but the leaderboard positions should be treated as claims pending independent replication, in the same way the wiki treats vendor benchmark announcements.
  • “The harness matters more than the model” is the title, not the finding. The evidence supports harness variance is large on fixed weights — which is compatible with model quality also mattering enormously, and indeed the same talk reports frontier-model choice moving ARC-AGI-3 results by tens of points. The wiki should carry the disjunctive claim (both layers matter, and the harness layer has been undermeasured) rather than the headline’s implied ranking. Note the tension with the domain-specialisation reversal, where the originating vendor argues against elaborating the harness.
  • Thin harness versus expressive harness — a real disagreement, or a scope difference? QM’s stated goal is “keep the harness extremely thin” — three tools, everything else regarded as temporary scaffolding. Karten’s stated goal is maximum expressibility. They are on the same bill and do not reconcile the tension. A plausible reading is that they agree: both are removing prescribed control flow while preserving primitive capability, and “thin” and “expressive” describe the same harness from different sides. But that reading is the wiki’s, not the speakers’. This is the live question in harness-thinning-what-persists and this source is evidence on both sides of it.
  • The rubber-stamping admission cuts against the human-in-the-loop remedy offered two paragraphs earlier. France and Bell rely on human review to catch the automated-improvement loop’s “main character syndrome”, and report that human review of database writes has decayed into rubber-stamping. Both are true in the same system. Whether human-in-the-loop is a durable control or a temporarily-effective one that decays with familiarity is not resolved here, and it is the sharpest open question the source raises for agent-oversight-and-delegation.
  • Open question — does the local-model gap actually hold at 6–12 months? The claim is asserted with one model pair as illustration and no benchmark table. It is directionally consistent with open-source-ai’s existing sources, which put the lag at roughly three months for text tasks from a different vantage, and the two estimates are not obviously reconcilable.

What was actually ingested

Full ~60-minute session transcript (auto-generated English captions). The host and opening presenter is not named in either the transcript or the channel description; they identify themselves only as a YC-affiliated researcher with a Stanford Hazy Research connection, and are referred to here as “the presenter”. Speaker names for the three presentations are taken from the channel description’s chapter list, which is the ground truth for spelling (Seth Karten, Jon Saad-Falcon, Josh France, Regan Bell) — the ASR renders several of them incorrectly. Several model names in the ARC-AGI-3 comparison are garbled beyond safe reconstruction in the ASR (one appears as “GPT tero”, another as “GPT soul”); only the figures attributable to a confidently-identified model are reported above, and the rest are omitted rather than guessed. Slides are referenced throughout and were not ingested — all architecture details here are as narrated, not as shown.