Runkle & Lovell — Building a Harness with Jev

TL;DR

A LangChain integration post announcing langchain-typesafe, which wraps Jev, a model released by TypeSafe AI in September 2026. The authors are Sydney Runkle (product manager, LangChain open-source) and Hunter Lovell. Jev is not a text generator. TypeSafe calls it a “System One model”: you send it a state (text, structured data or LangChain messages) and a set of typed questions, and it returns typed answers with probabilities. The post’s argument is that many decisions an agent currently makes with a full LLM call are really classification tasks. Those can be moved off the main model into a cheaper, faster one without changing what the agent can do.

The headline numbers (“up to 200x faster inference and 400x lower cost than comparable LLMs on classification tasks”) are TypeSafe’s claims, attributed as such. This post does not measure them. The companion eval post does measure latency and cost, and gets much smaller ratios (see §Scope and reliability).

What Jev is, per the post

  • Three question types. Choice picks one of several options and returns a probability per option plus overall confidence. Score rates against ordered levels (e.g. low/medium/high) and returns a continuous score, the distribution and a confidence. Noul (TypeSafe’s term for a yes/no question) returns the probability that a statement is true.
  • Parallel questions against one state. “System One models evaluate every question in a request in parallel. Adding questions barely changes the response time.” The post contrasts this with the sequential, token-by-token way an LLM reaches a judgement.
  • Training objective. Jev was trained with “reinforcement learning for calibrated decisions (RLCD)”, according to TypeSafe. The post gives no model size, architecture or benchmark.
  • Where it sits. “Jev isn’t a drop-in replacement for an LLM… use an LLM for open-ended reasoning and generation, and Jev for fast, structured decisions along the way.” It is a component inside the loop, not a model for the loop.

The post places this as the third step in making LLMs usable in software. Tool calling gave models structured requests and structured outputs gave them typed results, but “the agent loop is still slow and costly: every decision requires another model call.”

The two harness components it ships

Both are shipped as middleware under langchain_typesafe.experimental.middleware, the extension point that agent-harness records as the Constraints layer.

1. ModelRouterMiddleware — model routing. You declare model choices with plain-language criteria, e.g. fast for “direct lookups, extraction, and localized changes” and powerful for “architecture and high-stakes decisions”. Jev then picks one from the latest user message, with the instruction “Choose the least costly model that can complete the task.” The choice is made once per run: “The router selects a model from the latest user message and uses it throughout the run.” The probabilities stay in agent state. This puts in code the routing idea on small-language-models: send each request to the cheapest model that can serve it. It routes per run, not per step. For the per-model economics, see Belcak et al. (NVIDIA) and Google Cloud’s loop-count argument.

2. AutoModeMiddleware — tool-risk gating. Jev checks each tool call (the example scopes it to bash) and blocks risky ones before the tool runs. The post frames it as open-sourcing something closed:

“Coding harnesses like claude, codex, cursor have shipped some kind of way to classify dangerous actions before they’re taken which has slowly helped to build trust in agents. Up until now, this classifier step has been locked away in the closed source parts of the harness. Now that a cheap and performant classifier model exists, we can take the same pattern and adopt it to all agents!”

The closed version it refers to is Claude Code’s auto mode, whose design Lydia Hallie described on The Agent Factory. The post’s stated reason for gating is trust: “Agents are still inherently untrustworthy. An agent can receive bad instructions (either naturally or from a motivated enough attacker).” See agent-oversight-and-delegation.

The harness framing matches LangChain’s own vocabulary (Trivedy, The Anatomy of an Agent Harness): both components are harness middleware, not changes to the model.

Dynamic-capabilities reading

  • digital-sensing/digital-scouting — the post is a scouting artefact: a vendor spots a new model class in the week it appears and shows where it fits in an existing agent stack. That is the identifying new technology options activity, from the builder side.

What was actually ingested

The full post body: prose, both JSON examples from TypeSafe’s quickstart, and the three Python snippets. One figure (a three-panel illustration of the question types) is an image; the prose next to it covers the same content. The page’s structured metadata names only a placeholder author (“LangChain Accounts”). The visible byline, Sydney Runkle, Hunter Lovell, September 17, 2026, is used here.

Linked entities and concepts

Scope and reliability

This is an integration announcement from a company that ships the integration. It contains no measurements. The speed and cost multipliers are TypeSafe’s own figures for classification tasks, quoted without method. Three things the reader cannot check from this post:

  • The multipliers. LangChain’s own measurement four days later (Shea & Roche) found Jev at $0.00035 per call against GPT-5.6 Luna’s $0.00039, about 1.1× cheaper, and 0.44 s against 2.16–2.83 s, about 5–6× faster. Jev was ~80× cheaper than Claude Sonnet 4.6. The “400×” figure depends on which LLM you compare against, and the post does not say which.
  • The gate’s error rates. AutoModeMiddleware blocks calls Jev scores as risky. No false-block or false-approve rate is reported, and no threshold is documented in the post.
  • The model. Size, architecture and training data are not disclosed. Both middleware components are in an experimental namespace.

The social-proof paragraph (named builders using Jev for browser agents, a trading agent and email triage) links to X posts, not to any results.