Pochampally, An & Chen — Assistant or Actor? Delegation Regret
TL;DR
The paper that names the failure mode nobody had a word for.
Delegation regret: “a pattern in which users regret not that the agent erred, but that it acted beyond what they would have authorized.”
That distinction is the contribution. Every existing framework for agent quality measures whether the agent was right. This measures whether the agent was authorised — and the two come apart. Finding 3 is the sharp version: “delegation regret appeared consistently when the agent executed actions without preview, even when the output was rated as successful.” A correct outcome does not repair an unauthorised action.
Design. 20 university students, five common daily tasks with an agent (OpenClaw), tasks chosen to vary in privacy, stakes and reversibility. Trust, perceived control, transparency, supervision burden and approval preference measured on 5-point Likert scales, plus thematic coding of free-text reflections.
Three findings:
-
Trust is calibrated per task, not per agent. Participants granted wide autonomy for advisory and low-stakes work and demanded confirmation for irreversible, externally visible actions. There is no such thing as “how much do you trust this agent” — the question is malformed. This is the empirical basis for per-task autonomy policies, and it is why a single global permission setting will always be wrong in both directions.
-
Irreversibility × external visibility drives trust withdrawal — not stakes. The moderate-stakes email task produced the sharpest trust drop (M = 3.10) and the highest approval demand (M = 4.65), while a high-stakes but verifiable task did not. Sending an email is not high-stakes; it is unrecallable and seen by someone else. That combination, not consequence magnitude, is what people actually guard.
-
Preview is the mechanism, not permission. Regret appeared whenever the agent acted without preview. The design implication is specific: it is not enough to have granted the agent permission in advance — users want to see what it is about to do.
Why this belongs next to the software-engineering material. Merge Mommy scores reversibility and blast radius among its six dimensions and escalates on them; Carson keeps production credentials out of agent hands and approves sensitive actions by hand; his Land PR loop gates on watching a video of what the agent did before merging — which is preview, exactly as finding 3 prescribes. Practitioners converged on these controls commercially; this study explains why they feel necessary, and Singapore’s regulator arrives at the same place from policy.
Terminology note. The paper frames trust calibration and per-task autonomy. It does not use the phrase “span of control”; any manuscript attributing that phrase to this paper should be corrected.
Dynamic-capabilities reading
contextual/internal-barriers— delegation regret is an adoption barrier that persists even when the agent works, so it cannot be engineered away by improving quality.digital-transforming/redesigning-internal-structures— the design implications (expose action boundaries, per-task autonomy policies, separate advisory output from agentic execution) are structural prescriptions for how agent authority is organised.
Linked entities and concepts
- Concepts: agent-oversight-and-delegation, agent-fleet-management, responsible-ai, ai-agents
- Dangling (single-source mention, deferred): Shiva Pochampally, Shengwei An, Yan Chen
Scope and reliability
Abstract only. N = 20 university students, one agent, five tasks — this is a mechanism-naming qualitative study, not a population estimate, and the Likert means (M = 3.10, M = 4.65) should be read as descriptive of this sample rather than as effect sizes. The student population is a real limit: professional users with accountability for outcomes may calibrate differently, plausibly more conservatively. The concept — delegation regret, and the irreversibility × visibility trigger — is the durable contribution and transfers well beyond the study’s scope; the numbers do not.
Date discrepancy, recorded rather than resolved: the arXiv identifier is 2607.* (July 2026) while the listing page states a v1 submission of Thu, 14 May 2026. This page uses the stated submission date; the identifier is the citable handle.