AI-Wiki

Tag: cot-monitoring

3 items with this tag.

  • Sep 01, 2026

    Agent oversight and delegation

    • oversight
    • delegation-regret
    • trust-calibration
    • reversibility
    • blast-radius
    • approval-checkpoints
    • least-privilege
    • imda
    • preview
    • cot-monitoring
    • risk-scoring
    • type/concept
  • Sep 01, 2026

    Reward hacking

    • reward-hacking
    • specification-gaming
    • benchmark-integrity
    • held-out-tests
    • cot-monitoring
    • obfuscation
    • swe-bench
    • evaluation
    • oversight-surface
    • goodharts-law
    • type/concept
  • Aug 30, 2026

    Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

    • openai
    • reward-hacking
    • chain-of-thought
    • cot-monitoring
    • obfuscation
    • monitorability-tax
    • o3-mini
    • gpt-4o
    • alignment
    • oversight
    • type/source
    • kind/paper

Created with Quartz v0.1.0 © 2026

254 sources · 171 entities · 47 concepts · 7 threads · 5 syntheses

  • GitHub