AI-Wiki

Tag: held-out-tests

3 items with this tag.

  • Sep 01, 2026

    Reward hacking

    • reward-hacking
    • specification-gaming
    • benchmark-integrity
    • held-out-tests
    • cot-monitoring
    • obfuscation
    • swe-bench
    • evaluation
    • oversight-surface
    • goodharts-law
    • type/concept
  • Aug 30, 2026

    EvilGenie: A Reward Hacking Benchmark

    • reward-hacking
    • benchmark
    • livecodebench
    • llm-judge
    • held-out-tests
    • codex
    • claude-code
    • gemini-cli
    • inspect
    • misalignment
    • detection
    • type/source
    • kind/paper
  • Aug 30, 2026

    SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

    • specbench
    • reward-hacking
    • reward-hacking-gap
    • held-out-tests
    • long-horizon
    • benchmark
    • oversight-surface
    • os-kernel
    • json-parser
    • test-memorization
    • type/source
    • kind/paper

Created with Quartz v0.1.0 © 2026

254 sources · 171 entities · 47 concepts · 7 threads · 5 syntheses

  • GitHub