AI-Wiki
Search
Search
Dark mode
Light mode
Explorer
Tag: swe-bench
5 items with this tag.
Sep 01, 2026
Reward hacking
reward-hacking
specification-gaming
benchmark-integrity
held-out-tests
cot-monitoring
obfuscation
swe-bench
evaluation
oversight-surface
goodharts-law
type/concept
Aug 30, 2026
Karthik Narasimhan
karthik-narasimhan
princeton-nlp
react
swe-bench
language-agents
benchmarks
reasoning-and-acting
type/entity
kind/person
Aug 30, 2026
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
swe-bench
benchmark
github-issues
code-generation
iclr-2024
swe-llama
claude-2
evaluation
princeton-nlp
swe-bench-verified
swe-bench-pro
type/source
kind/paper
Aug 30, 2026
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
agents-md
claude-md
context-files
repository-context
swe-bench
inference-cost
negative-result
eth-zurich
context-engineering
prompt-engineering
type/source
kind/paper
May 09, 2026
Rethinking Agents — Harness is All you Need
agent-harness
harness-engineering
subtraction-principle
llm-non-determinism
dspy
natural-language-harness
transferable-harness
terminal-bench-2
swe-bench
ablation
type/source
kind/video