AI-Wiki

Tag: online-evals

2 items with this tag.

  • Sep 22, 2026

    Jev-as-a-Judge for Agent Evals

    • langchain
    • langsmith
    • deep-agents
    • typesafe-ai
    • jev
    • system-one-models
    • daniel-shea
    • sean-roche
    • llm-as-judge
    • agent-evaluation
    • online-evals
    • evaluator-variance
    • repeatability
    • eval-cost
    • gpt-5-6
    • claude-sonnet-4-6
    • signal-value
    • vendor-experiment
    • small-n
    • type/source
    • kind/article
  • Sep 22, 2026

    Building a Harness with Jev

    • langchain
    • typesafe-ai
    • jev
    • system-one-models
    • sydney-runkle
    • kahneman
    • thinking-fast-and-slow
    • agent-harness
    • model-routing
    • auto-mode
    • llm-as-judge
    • online-evals
    • pii-detection
    • langchain-typesafe
    • science-technology
    • vendor-explainer
    • type/source
    • kind/video

Created with Quartz v0.1.0 © 2026

307 sources · 192 entities · 53 concepts · 9 threads · 6 syntheses

  • GitHub