AI-Wiki

Tag: evaluation

5 items with this tag.

  • Sep 01, 2026

    Reward hacking

    • reward-hacking
    • specification-gaming
    • benchmark-integrity
    • held-out-tests
    • cot-monitoring
    • obfuscation
    • swe-bench
    • evaluation
    • oversight-surface
    • goodharts-law
    • type/concept
  • Aug 30, 2026

    Cursor

    • cursor
    • composer
    • ai-ide
    • coding-agent
    • benchmark-integrity
    • reward-hacking
    • swe-bench-pro
    • agentic-pr
    • evaluation
    • type/entity
    • kind/product
  • Aug 30, 2026

    SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    • swe-bench
    • benchmark
    • github-issues
    • code-generation
    • iclr-2024
    • swe-llama
    • claude-2
    • evaluation
    • princeton-nlp
    • swe-bench-verified
    • swe-bench-pro
    • type/source
    • kind/paper
  • Aug 12, 2026

    How to Reap Compound Benefits From Generative AI

    • mit-smr
    • kiron
    • schrage
    • compound-benefits
    • evaluation
    • verification
    • learning-capture
    • generative-ai
    • micro-productivity-trap
    • ai-iteration
    • asset-appreciation
    • type/source
    • kind/article
  • Aug 12, 2026

    The Agent Development Lifecycle

    • langchain
    • agent-development-lifecycle
    • adlc
    • build-test-deploy-monitor
    • governance
    • langgraph
    • langsmith
    • deep-agents
    • agent-runtime
    • agent-harness
    • evaluation
    • tracing
    • simulation
    • sandboxes
    • prompt-hub
    • type/source
    • kind/article

Created with Quartz v0.1.0 © 2026

254 sources · 171 entities · 47 concepts · 7 threads · 5 syntheses

  • GitHub