AI-Wiki

Tag: agentic-evals

2 items with this tag.

  • Aug 12, 2026

    Agentic Evaluations Workshop — Deep Dive on the Future of Evals for Agents

    • evals
    • agentic-evals
    • ai-benchmarks
    • reliability
    • capability-reliability-gap
    • automation-vs-augmentation
    • gaia-2
    • are-environment
    • inspect-ai
    • lighteval
    • open-llm-leaderboard
    • community-eval
    • living-benchmarks
    • every-eval-ever
    • sim-to-real-gap
    • reward-hacking
    • llm-as-judge
    • rubrics
    • agent-development-lifecycle
    • agent-harness
    • ai-policy
    • hugging-face
    • type/source
    • kind/video
  • Jul 22, 2026

    Hugging Face

    • hugging-face
    • open-source-ai
    • open-weight-models
    • model-hub
    • datasets
    • ai-builders
    • clem-delangue
    • nemotron
    • reachy
    • agentic-evals
    • github-for-ai
    • type/entity
    • kind/organization

Created with Quartz v0.1.0 © 2026

254 sources · 171 entities · 47 concepts · 7 threads · 5 syntheses

  • GitHub