Cui, Demirer, Jaffe, Musolff, Peng & Salz — The Effects of Generative AI on High-Skilled Work
TL;DR
Three randomised field experiments, at Microsoft, Accenture, and an anonymous Fortune 100 electronics manufacturer, measuring what happens when developers are granted access to GitHub Copilot. 4,867 developers. Pre-registered as AEARCTR-0014530; published in Management Science.
Headline: a 26.08% increase (standard error 10.3%) in completed tasks for developers granted access.
Two things about that number deserve emphasis over the number itself.
1. The standard error is large. 10.3% on a 26.08% point estimate means the 95% interval runs roughly from +6% to +46%. The finding is that the effect is positive and probably substantial — not that it is 26%. Citing “26%” as a settled figure overstates what three experiments with this much variance can establish.
2. The heterogeneity is the real finding. Less experienced developers adopted the tool at higher rates and gained more from it. That single result is what makes this study and METR’s RCT compatible rather than merely contradictory: METR studied 16 maintainers with five years of context on their own repositories — the far tail of the experience distribution, where this paper’s own gradient predicts the smallest gain. The two studies together support a gradient claim: assistance is worth most where the marginal work is acquiring context, and worth least — possibly negative — where the developer already holds it.
3. Note what the outcome variable is. Completed tasks, not shipped value and not defect-free code. The measure is a volume measure. 2026-03-30-liu-debt-behind-the-ai-boom measures what a comparable volume of AI-authored change costs downstream (22.7% of introduced issues still alive at the latest revision), and DORA finds throughput and stability moving in opposite directions. None of these studies is wrong; they are measuring different points in the pipeline, and only the first one is flattering.
Dynamic-capabilities reading
digital-transforming/improving-digital-maturity— three large firms measuring a tool rollout with randomisation rather than a satisfaction survey is itself the maturity behaviour the corpus keeps prescribing.contextual/internal-enablers— the adoption gradient by seniority is an enabling-condition finding: who takes up the tool determines where the return lands.
Linked entities and concepts
- Entities: Microsoft, GitHub
- Concepts: ai-coding-productivity-evidence, automation-vs-augmentation, ai-employment-effects, enterprise-ai-adoption
- Dangling (single-source mention, deferred): Kevin Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, Tobias Salz
Scope and reliability
Publisher abstract and headline estimates only. The journal article is paywalled and the working-paper PDF could not be converted in this environment (no local PDF toolchain), so the per-experiment breakdowns, the secondary outcomes (commits, pull requests, builds), and the robustness checks were not read. Two of the three sites are the tool vendor (Microsoft, which owns GitHub) and a large systems integrator; the vendor-site issue is mitigated by randomisation and pre-registration but not eliminated. Before quoting the 26.08% figure in print, read the full paper — in particular to confirm whether the headline pools all three experiments and how the Microsoft site alone behaves.