Runkle / LangChain — Building a Harness with Jev (video)

Learn all about Jev, a new System One model from TypeSafe AI, and how you can use it in your agent harness. Jev is up to 200x faster and 400x cheaper than LLMs on classification style tasks, which makes it great for applications like model routing and “jev as a judge” online evals.

TL;DR

A nine-minute slide-and-demo explainer from LangChain, presented by Sydney Runkle, product manager on LangChain’s open-source team. It is the video version of Runkle & Lovell’s blog post: what Jev from TypeSafe AI is, its three question types, the langchain-typesafe integration, and three uses (model routing, auto mode, Jev-as-judge). Most of it repeats the post. Four things are only here:

1. The name comes from Kahneman. “System One” is borrowed from Daniel Kahneman’s Thinking, Fast and Slow: “fast and cheap, System 1 intuition, versus slow and expensive, System 2 reasoning.” TypeSafe positions Jev as the System 1 half, “returning almost instant typed decisions instead of generating text”. Standard LLMs “behave more like System 2, slower and pricier, but adhering to more of this step-by-step reasoning over open-ended problems.” This is the vendor’s branding, not a claim about cognition. It still gives a compact picture of how the two model types split the work in a harness.

2. A side-by-side demo. The question “is there PII in this text?” goes to an LLM with structured output (about five seconds) and to Jev (“almost immediately, chance that there’s PII is 98%”). One example, timed by eye.

3. The speed and cost claim is stated as a range. “20 to 200 times faster and 40 to 400 times cheaper on classification style tasks compared to an LLM.” The blog post and the video description quote only the top of the range.

4. A first-person account of turning oversight off because it was slow. This is the most useful minute in the video:

“Anecdotally, I actually turned off auto mode in my coding agent recently, because the classification step of whether a given tool call was risky was too slow for my coding agent to feel productive. But it’s back on now that Jev can make these decisions so quickly.”

So a safety gate was switched off because of its latency, not because it was wrong. It came back on when the gate got faster. This is the same class of cost as the permission fatigue that Lydia Hallie described for Claude Code’s auto mode. Hallie’s gate is ignored because it asks too often; Runkle’s was dropped because it answered too slowly. See agent-oversight-and-delegation.

On routing, Runkle adds an internal use: LangChain wants “our internal coding agents to switch between fast and cheap models and more expensive and powerful models depending on the complexity of a given question or coding task”. That is the heterogeneous-model argument on small-language-models, as a stated intention, not a reported result. On evals she summarises the Jev-as-a-Judge experiment as “much cheaper and much faster, but also… much more reliable and consistent”. That goes beyond what its five-case result supports (see that page’s scope section). She places it as “an evolution of LLM-as-a-judge online evals” on the Test/Monitor stages of agent-development-lifecycle.

Dynamic-capabilities reading

  • digital-sensing/digital-scouting — a vendor introducing a new model class to its developer audience days after release, and mapping where it fits in the agent-harness. This is technology scouting done in public.

What was actually ingested

The full transcript from the creator-uploaded (manual) English caption track, human-curated, so no ASR cleanup was needed. The slides themselves were not captured; the example state (“Hi, I’ve been trying to connect my Stripe account for three days…”) and the scores quoted for it (billing 0.84 with confidence 0.596; frustration score 1.035; urgency 0.999) are read aloud in the transcript.

Linked entities and concepts

Scope and reliability

Vendor explainer, no measurements. The demo is one prompt. The multipliers are TypeSafe’s. The auto-mode anecdote is one person’s experience, useful as a mechanism (latency as a reason to disable a gate) and not as evidence of how often it happens. The evals claim overstates its own source.