Dell’Acqua et al. — The Cybernetic Teammate (Organization Science, 2026)

The strongest source in the corpus on what generative AI does to teams — a preregistered 2×2 field experiment with 791 professionals at Procter & Gamble, published in Organization Science and licensed CC BY. Eleven authors across Harvard Business School’s AI Institute, ESSEC, Warwick, Wharton and P&G itself, with Fabrizio Dell’Acqua first, Ethan Mollick and Karim Lakhani among them.

It is the sequel to the BCG jagged-frontier study the wiki already holds, and it answers a different question: not how well does an individual do with AI, but can AI stand in for a teammate.

TL;DR

“Individuals with AI matched the performance of teams without AI, suggesting that AI can effectively replicate certain benefits of human collaboration.” — Abstract

The design is unusually clean for organisational research, and the wiki should say so plainly: randomised, preregistered, real business problems from participants’ own units, a full working day of effort, professional stakes (the best proposals went to business-unit leaders, and the best ideas entered P&G’s actual innovation pipeline), and evaluation by multiple independent assessors under a protocol validated by P&G managers. Four arms: individual alone, human pair, individual + GenAI, human pair + GenAI. Every pair was one commercial professional and one R&D professional, so cross-functional collaboration was structural rather than incidental.

The four results

1. AI substitutes for a teammate — and the second human adds almost nothing on top.

ConditionEffect vs. individual working alone
Team, no AI+0.24 SD (~6.3%), p<0.05
Individual + AI+0.37 SD (~9.6%), p<0.01
Team + AI+0.39 SD (~10.2%), p<0.01 — not significantly different from individual + AI

The authors read this as diminishing marginal returns to team expansion once AI is present: going individual → dyad → AI-augmented individual → AI-augmented dyad is a sequential head-count expansion, and the fourth step buys nothing on the average. Their own gloss is worth keeping: AI’s effect “appears to stem more from its capacity to bolster individual cognitive capabilities than affecting human-to-human collaboration.”

Two readings the corpus should hold alongside this. Cui et al. find the largest gains among juniors, where the marginal work is acquiring context — and what AI substitutes for here is precisely context the other professional would have supplied, so the two results are the same mechanism read at the seniority and the cross-functional level. And the silo result below is the expert-generalist thesis produced mechanically rather than cultivated.

2. AI dissolves functional silos — the finding the authors call their most noteworthy. Without AI, the silos are visible in the output: commercial professionals submitted predominantly commercial solutions, R&D professionals predominantly technical ones. Pairing a commercial with an R&D professional produced balanced solutions — that is what cross-functional teams are for. Individuals using AI achieved the same balance alone. AI acts “not just as an information provider but as an effective boundary-spanning mechanism.”

3. The emotional result, which is the one most likely to be dismissed and probably shouldn’t be. Positive emotions rose +0.457 SD for individuals with AI and +0.635 SD for teams with AI (both p<0.01); negative emotions fell. The mechanism proposed is the language interface itself — an LLM “trained on human language and often act[ing] more like a person than a machine” fills part of the social and motivational role of a human teammate. The wiki should hold this as self-reported affect, which is what it is, but note that it is measured under randomisation.

4. Breakthroughs need the pair and the AI. On a binary Top 10% Solutions measure, Team + AI were 9.2 percentage points more likely to land in the top decile against a control mean of 5.8% — roughly three times the rate. Individuals with AI showed a small, statistically insignificant effect. So the average-performance story (where the second human is redundant) and the tail story (where the second human is decisive) point in opposite directions. If an organisation cares about exceptional outcomes rather than mean quality, the pair still earns its cost.

The perception gap, inverted

The single most important thing this paper does for this wiki. Participants using AI were 9.2 percentage points less likely to expect their solution to place in the top 10% (p<0.05) — while objectively performing better.

The corpus’s standing claim, from METR, is that developers slowed 19% believed they had been sped up 20%, and that the survey instrument most organisations use is therefore invalid. That finding stands. But the inference many readers draw from it — self-report is inflated by AI use — does not survive this paper. Here the bias runs the other way, in a larger, preregistered sample. The transferable claim is narrower than the wiki has been stating it: self-assessment is decoupled from performance under AI, in a direction that is not predictable from the technology alone. Recorded on ai-coding-productivity-evidence.

Where the human stays

The decomposition analysis splits the innovation process and finds AI “primarily enhances the quality of generated ideas, shifting the distribution of creative output upward, whereas human judgment retains value in evaluative selection.” Generation improves; selection does not. That is the evidentiary basis for the claim Ethan Mollick makes without numbers in his Sinek interview — evaluation, not generation, is the bottleneck.

A second detail worth keeping: the AI-content retention analysis is polarised, with a non-trivial share of participants retaining zero AI-generated sentences. Two distinct usage styles — AI as ghostwriter, and AI as sounding board — both inside the treatment arm. Averaged treatment effects hide that.

Dynamic capabilities (Warner & Wäger)

  • digital-transforming/redesigning-internal-structures — the paper’s own stated implication: “organizations may need to reevaluate optimal team sizes and compositions.” It is a measurement of whether a structural unit (the cross-functional dyad) is still load-bearing, which is exactly a redesign question.
  • digital-seizing/rapid-prototyping — the experimental task is early-stage new product development, run as a one-day flash-team sprint that mirrors P&G’s real ideation routine; the finding is about how fast an organisation can generate and screen viable product concepts.
  • strategic-renewal/collaborative-approach — the silo-dissolution result is a claim about how expertise circulates across functional boundaries, which is the collaborative-approach cell read at the knowledge level rather than the partnership level.

Linked entities and concepts

Scope and reliability

The highest evidential tier the wiki holds: peer-reviewed, preregistered, randomised, N=791, real tasks with real stakes, independent evaluation, open access. Cite it for magnitudes, not just framing.

The authors’ own limits, stated rather than buried — and they cut both ways:

  • The effects are probably a lower bound. Participants were “relatively inexperienced with AI prompting techniques”, and the tools were not built for collaborative work.
  • One day, one company, one industry, one model. Virtual collaboration between largely unfamiliar participants — the authors explicitly call these flash teams, not established teams with embedded relationships. Extended coordination and iterative rework cycles are absent by construction. Consumer packaged goods, early-stage NPD only.
  • Cross-functional pairs only. Same-expertise pairs and larger teams may behave differently, and the paper does not test them.
  • A single AI model at a point in time.

The practical consequence: the “AI replaces a teammate” headline is an artifact of a one-day flash team as much as of AI. What a human teammate contributes over weeks — relationship, memory, rework, disagreement that survives a night’s sleep — is not what this experiment could measure.

Debates and supersession

  • The perception gap is no longer a one-directional finding. See above; this page and METR are both correct and point opposite ways, so the wiki now holds decoupling rather than inflation.
  • Average versus tail is an unresolved design question. Individual + AI matches Team + AI on the mean and loses badly on the top decile (insignificant vs. 3×). An organisation optimising for throughput and one optimising for breakthroughs should read this paper differently, and the paper does not adjudicate.
  • Does boundary-spanning build expertise or only rent it? The authors ask this themselves: “Does AI-enabled boundary spanning foster genuine knowledge growth, or merely facilitate temporary access to existing expertise?” Unanswered here, and directly load-bearing for ai-deskilling — a commercial professional producing technically balanced proposals has not thereby learned engineering.
  • Open: every result is one-day. Nothing in the corpus measures what happens to team performance, silo structure, or affect after months of AI-augmented work.