Li, Zhang & Hassan — AIDev: Studying AI Coding Agents on GitHub

TL;DR

The population dataset underneath the agentic-PR literature, and the reason the 2026 findings in this ingest are population statements rather than case studies.

Agent-authored pull requests (“Agentic-PRs”)932,791
Repositories116,211
Developers involved72,189
Agents covered5 — OpenAI Codex, Devin, GitHub Copilot, Cursor, Claude Code
Curated subset33,596 PRs from 2,807 repos with >100 stars, enriched with comments, reviews, commits and related issues

Why the numbers themselves are the finding. Nearly a million agent-authored pull requests across 116,211 repositories is not an emerging practice — it is an established one, measured in early 2026. The involvement of 72,189 developers is the other half: these PRs are landing in front of humans who have to do something about them. That is the volume that makes review capacity, not authoring capacity, the binding constraint — the premise Merge Mommy and LGTM both start from.

The curated subset is the methodological contribution. 33,596 PRs from repositories with >100 stars, with comments, reviews, commits and linked issues attached, is what makes qualitative work on agent failure possible at all — it is the sampling frame Abujadallah et al. draw their 306 hand-coded rejections from.

Note that the five agents named here are exactly the five that recur across this ingest — Codex, Devin, Copilot, Cursor, Claude Code. The same products appear in Carson’s stack, in EvilGenie’s reward-hacking evaluation, and in Liu et al.’s technical-debt study. The corpus is converging on one small set of tools, which is convenient for cross-study comparison and a reason to be careful about generalising beyond them.

Dynamic-capabilities reading

  • digital-sensing/digital-scouting — the dataset’s entire function is making an emerging practice observable; it is scouting infrastructure for the research community.

Linked entities and concepts

Scope and reliability

Abstract only. A dataset paper — it makes no causal claims and should not be cited for any. Coverage limits worth holding: public GitHub only (so no enterprise or private-repo behaviour), five agents (so no coverage of in-house or less popular tooling), and attribution depends on agents identifying themselves in PR metadata, which will undercount agent involvement where a human takes credit. The >100-star curated subset skews toward visible, well-maintained projects — which is precisely the population LGTM finds least willing to auto-merge, so findings from the subset will understate laissez-faire practice in the long tail.