AI-Wiki

Tag: swe-bench-pro

3 items with this tag.

  • Aug 30, 2026

    Cursor

    • cursor
    • composer
    • ai-ide
    • coding-agent
    • benchmark-integrity
    • reward-hacking
    • swe-bench-pro
    • agentic-pr
    • evaluation
    • type/entity
    • kind/product
  • Aug 30, 2026

    SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    • swe-bench
    • benchmark
    • github-issues
    • code-generation
    • iclr-2024
    • swe-llama
    • claude-2
    • evaluation
    • princeton-nlp
    • swe-bench-verified
    • swe-bench-pro
    • type/source
    • kind/paper
  • Aug 30, 2026

    Reward hacking is swamping model intelligence gains

    • cursor
    • reward-hacking
    • benchmark-contamination
    • swe-bench-pro
    • swe-bench-multilingual
    • opus-4-8-max
    • composer-2-5
    • upstream-lookup
    • git-history-mining
    • harness-design
    • evaluation-integrity
    • type/source
    • kind/article

Created with Quartz v0.1.0 © 2026

254 sources · 171 entities · 47 concepts · 7 threads · 5 syntheses

  • GitHub