AI-Wiki
Search
Search
Dark mode
Light mode
Explorer
Tag: ai-evaluation
5 items with this tag.
Aug 31, 2026
durable-skills
durable-skills
21st-century-skills
future-ready-skills
ai-deskilling
scalable-assessment
psychometrics
ai-evaluation
hiring-criteria
leadership-skills
type/concept
Aug 30, 2026
ai-benchmarks
ai-benchmarks
ai-evaluation
foundation-models
capability-reliability-gap
scar-fragmentation
type/concept
Aug 30, 2026
METR
metr
ai-evaluation
ai-benchmarks
ai-safety
reward-hacking
rct
developer-productivity
task-horizons
re-bench
hcast
type/entity
kind/organization
Aug 12, 2026
Evals 101 — Doug Guthrie, Braintrust
evals
ai-evaluation
llm-as-judge
observability
agent-development-lifecycle
agent-harness
braintrust
ai-engineer-worlds-fair
flywheel-effect
human-in-the-loop
online-scoring
offline-evals
type/source
kind/video
Aug 12, 2026
AI Evaluations Clearly Explained in 50 Minutes (Real Example) — Hamel Husain on Peter Yang's podcast
evals
ai-evaluation
llm-as-judge
observability
binary-pass-fail
agreement-metric
true-positive-rate
true-negative-rate
error-analysis
axial-coding
spreadsheet-evals
hamel-husain
peter-yang
agent-development-lifecycle
agent-harness
ai-product-management
type/source
kind/video