AI-Wiki

Tag: swe-bench

5 items with this tag.

  • Sep 01, 2026

    Reward hacking

    • reward-hacking
    • specification-gaming
    • benchmark-integrity
    • held-out-tests
    • cot-monitoring
    • obfuscation
    • swe-bench
    • evaluation
    • oversight-surface
    • goodharts-law
    • type/concept
  • Aug 30, 2026

    Karthik Narasimhan

    • karthik-narasimhan
    • princeton-nlp
    • react
    • swe-bench
    • language-agents
    • benchmarks
    • reasoning-and-acting
    • type/entity
    • kind/person
  • Aug 30, 2026

    SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    • swe-bench
    • benchmark
    • github-issues
    • code-generation
    • iclr-2024
    • swe-llama
    • claude-2
    • evaluation
    • princeton-nlp
    • swe-bench-verified
    • swe-bench-pro
    • type/source
    • kind/paper
  • Aug 30, 2026

    Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

    • agents-md
    • claude-md
    • context-files
    • repository-context
    • swe-bench
    • inference-cost
    • negative-result
    • eth-zurich
    • context-engineering
    • prompt-engineering
    • type/source
    • kind/paper
  • May 09, 2026

    Rethinking Agents — Harness is All you Need

    • agent-harness
    • harness-engineering
    • subtraction-principle
    • llm-non-determinism
    • dspy
    • natural-language-harness
    • transferable-harness
    • terminal-bench-2
    • swe-bench
    • ablation
    • type/source
    • kind/video

Created with Quartz v0.1.0 © 2026

254 sources · 171 entities · 47 concepts · 7 threads · 5 syntheses

  • GitHub