How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations

A peer-reviewed benchmarking study in Strategy Science (INFORMS, special issue Can AI Do Strategy?) that runs 21 proprietary and 13 open-source LLMs through the Back Bay Battery simulation — a standard MBA exercise in exploration/exploitation — and compares them against 249 MBA students. Open access under CC BY.

It is the wiki’s first peer-reviewed, multi-model, multiperiod measurement of LLM strategic decision making, and it reports a result the corpus has nothing else like: the newest frontier models are worse at this task than models a year older, and worse than MBA students, while continuing to improve on every general benchmark.

TL;DR

  • The gap it addresses. Strategy is defined by five properties — complexity/interdependence, high stakes, multiperiod partially-irreversible commitments, delayed and noisy feedback, and deep uncertainty. No existing benchmark captures them. Chess/Go/StarCraft have the structure but depend on training over millions of games; LLM video-game benchmarks (Claude Plays Pokémon) avoid the training requirement but lack irreversibility and business complexity; idea-generation studies (Csaszar, Doshi, Boussioux) test narrow one-shot judgment. The field also lacks any shared instrument, so single-model findings go stale as models turn over.
  • The instrument. The Back Bay Battery simulation (Christensen & Shih 2019): eight years as general manager of a battery firm, allocating a constrained R&D budget across a declining core technology (AGM, ~80% of initial revenue) and an emerging one (supercapacitor, ~20%), across three customer segments with different performance priorities, while setting prices and forecasting sales. Run on “advanced” difficulty in “legacy” mode — the MBA-classroom standard, in which a player can be fired for sales variance ≤ −50% in a year, below −20% for more than three years, or three consecutive loss-making years. Structural dynamics: an unavoidable core-market downturn around year 4, and supercapacitor breakthroughs that require early, sustained, focused investment before any payoff.
  • The scoring. Deliberately multi-outcome, because “as in real-world strategy, there is not a single performance outcome” — and students are never told what to maximise. Three metrics: Cumulative Profitability (core-business management), Cumulative Revenue (overall expansion, with final-period emerging revenue subtracted to reduce correlation), and Emerging Tech Revenue (future positioning). A Composite Score min–max normalises and equally weights the three; the authors are careful that it is “useful only for comparisons within the same data set” and “a simplifying visualization rather than a standalone metric.”
  • The contamination problem, and a good solution. Since BBB teaching materials and solutions are public, a model might simply recall the winning playbook. So every string from the simulation is regex-masked into synthetic names — Back Bay Battery → EnergyCo, AGM → AeroBond Matrix, Supercapacitor → Quantum Storage Cell, UPS → Continuity Power Modules, Energy Density → Specific Energy Index — and the background brief was rewritten from scratch. The masking targets recognition, not knowledge: general strategic concepts stay available, because using them “is precisely what we aim to evaluate.”
  • The validation of that protocol is the paper’s sharpest small result. Two-condition test on six models. Asked to “explain the Back Bay Battery simulation and how to win it”, four of six produced detailed descriptions matching public teaching materials — direct evidence of contamination. Asked the same about “EnergyCo”, no model gave relevant advice: four hallucinated fictitious simulations and two admitted no knowledge.
  • The human comparison. 249 second-year MBA students at a top US East Coast school across five cohorts (2018, 2020, 2023, 2024, 2025); first full run only, so outcomes reflect decisions under genuine initial uncertainty; 1–2 hours each. The authors flag its imperfections themselves — students were primed by a course covering exploration, exploitation and disruption, worked unsupervised, and later cohorts cannot be proven not to have used AI — and treat it as contextual reference, not an apples-to-apples control. Cohort checks find no significant differences.
  • The main result, in three movements.
    1. Early models are bad. GPT-3.5-turbo and GPT-4-era models score well below everything else and well below the MBA average.
    2. Late-2024/early-2025 models cross the human line. gpt-4o, claude-sonnet-4, gemini-2.0-flash, o4-mini and o3-mini (highest composite) exceed the historical MBA average — “a notable expansion of AI’s cognitive frontier.” They win by timing investment and pricing to get both profitability and growth.
    3. Mid-to-late-2025 frontier models regress below the students. GPT-5, o3 and Gemini 2.5 Pro score substantially worse than the lighter early-2025 models and below the MBA average.
  • The mechanism is a systematic exploitation bias, and it is visible in the models’ own reasoning. The frontier models achieve high cumulative profit with very low cumulative revenue and emerging-technology revenue — they defend the core and starve the future. In their own words: “Pause all QSC R&D due to long lead times and poor fit with current market requirements” (GPT-5); “We will not invest any R&D into QSC this year, preserving our limited capital for the core business” (Gemini 2.5 Pro). Part of the revenue shortfall comes from being fired earlier, but the pattern persists in “basic” mode where firing is impossible, and persists within the years they do survive.
  • Benchmark transfer breaks down — the finding with the widest implications. Plotting BBB composite against GPQA Diamond, the relationship is positive from GPT-3.5 through o4-mini and then reverses to negative for o3-mini through GPT-5. Against LM Arena it plateaus or reverses. The authors’ reading: “advances in general reasoning and conversational fluency, although once correlated with better strategic outcomes, no longer reliably predict them”, and the newest models “appear increasingly optimized for benchmarks in other domains, potentially at the expense of the forward-looking strategic reasoning that the BBB simulation captures.”
  • Prompt sensitivity cuts both ways, and the asymmetry is the interesting part. Re-run with instructions explicitly emphasising the emerging technology, GPT-5, o3 and Gemini 2.5 Pro improve marginally — so they can shift focus — but still lag the lighter models under the default open-ended prompt. Meanwhile o3-mini, the best performer under open-ended instructions, gets significantly worse: it overcommits to the new technology, sacrifices core profitability, and is fired earlier. Under the directive prompt, performance becomes roughly linear in release date but the overall average drops. The interpretation: frontier models are more rigid and literal instruction-followers — “which can become a liability in strategic or uncertain environments” — whereas earlier models were more adaptive when left to set their own priorities, and more fragile when given a narrow goal.
  • Open-source models plateau lower. DeepSeek, Gemma, Qwen, Llama and GPT-OSS score substantially below proprietary systems, with performance less systematically correlated with release date, and no clear frontier decline — they appear to have levelled off at a markedly lower band.
  • The forward agenda. A four-part taxonomy for future strategy benchmarks — elements of strategy (vary levers, horizons, build in genuine novelty not solvable by analogy), business context (other industries, stages, non-financial outcomes such as ethics and stakeholder buy-in), interaction structure (multi-agent: model-vs-model and model-vs-human arenas; and agentic play in the human interface rather than being fed structured information), and AI characteristics (fine-tuning on strategy corpora, RLHF). Plus a call for open-source, procedurally generated, contamination-resistant simulation infrastructure.

Why this matters to the wiki

1. It is the first source in the corpus where capability goes backwards on a task the wiki cares about. The corpus is full of monotone capability narratives and of practitioner claims keyed to “the frontier model.” This paper documents a release-date regression on a strategic task, with a named mechanism and a robustness check that survives removing the firing rule. Any wiki page that reasons “use the newest model for the hardest judgment” now has a peer-reviewed counter-example. See ai-benchmarks, strategy.

2. It gives jagged-frontier its strongest empirical extension. Dell’Acqua et al. located the jag with humans in the loop on a single out-of-frontier business case. This locates it without humans, across 34 models, on a task built specifically to have the properties strategy has — and finds the jag is not closing with model generation. The concept can now say something stronger than “capability is uneven”: on this measurement, the unevenness is growing in the direction of strategic judgment.

3. It supplies empirical support for a conceptual critique the wiki has held as argument only. theory-based-view carries Felin & Zenger’s and Felin & Holweg’s claim that LLM reasoning is “largely backward-looking and imitative” where strategic value creation requires forward-looking causal logic. The exploitation bias is exactly the behavioural signature that critique predicts — defend the observable, discount the hypothetical — and it is now measured rather than asserted. It also connects to analogical-reasoning: the masking protocol isolates transfer-from-principles by removing recognition-of-instance.

4. It is a benchmark-design contribution the corpus can use beyond strategy. The masking protocol plus its two-condition validation is a portable method for any benchmark built on published material, and it produced a clean measurement of contamination in passing (four of six models reciting a public playbook). That belongs alongside the wiki’s reward-hacking and ai-benchmarks material on how evaluation environments get gamed — here the gaming is memorisation rather than exploitation, and the defence is lexical rather than architectural.

5. The prompt-rigidity asymmetry is directly actionable. “The current frontier models may be more rigid and ‘literal’ in following explicit instructions, which can become a liability in strategic or uncertain environments” — while older models were better left open-ended and worse when constrained. For a corpus that carries a great deal of prompt- and context-engineering guidance (agent-harness, agentic-engineering), this is evidence that the right amount of instruction is model-dependent, and moving in an awkward direction.

Dynamic-capabilities reading

  • digital-seizing/balancing-digital-portfolios — The paper’s subject is this cell. BBB is a resource-allocation problem between a declining core and an emerging technology under a hard budget constraint, and the headline finding is that frontier models systematically get the balance wrong in one direction. This is the most direct empirical measurement of a W&W seizing cell anywhere in the corpus.
  • digital-sensing/digital-scenario-planning — The simulation is a scenario instrument: delayed noisy feedback, an unavoidable year-4 downturn, and technology thresholds that only reward early sustained commitment. Benchmarking on such an instrument is itself a sensing practice — a controlled way to find out what a capability will do before betting on it.
  • digital-sensing/digital-scouting — The cross-benchmark analysis is scouting discipline in its most useful form: a demonstration that the signals a firm would naturally scout on (GPQA, LM Arena, vendor release notes) have stopped predicting the capability it actually wants. A scouting function that tracks headline benchmarks would have concluded the opposite of what this measures.

roles: overrides the cell defaults toward the strategy-owning and R&D-owning seats — the finding is directly about who should not be handed a resource-allocation decision, and about how to evaluate a model before doing so.

Linked entities and concepts

  • Concepts: strategy (the central subject), ai-benchmarks (benchmark design, contamination, transfer failure), jagged-frontier (the strongest empirical extension in the corpus), theory-based-view (Felin & Zenger / Felin & Holweg, now with behavioural evidence), analogical-reasoning (what masking isolates), strategic-foresight (simulation as scenario instrument), dynamic-capabilities (exploration/exploitation as a seizing problem), automation-vs-augmentation (the authors’ own conclusion is collaboration, not substitution), foundation-models, open-source-ai (the open-model plateau).
  • Dangling (single-source mention, deferred): Ryan T. Allen, Rory M. McDonald, Strategy Science, INFORMS, Back Bay Battery, Brigham Young University, University of Virginia Darden. Note Felin, Holweg, Csaszar and Dell’Acqua appear here as cited authors rather than as this source’s authors, so the promotion rule does not apply to them from this ingest.

Debates and supersession

  • The MBA comparison is contextual, and the authors say so repeatedly. Students were primed by a course on exploration/exploitation/disruption, completed the exercise unsupervised on their own time, and later cohorts cannot be proven AI-free — though cohort checks find no significant differences and the conclusions hold across subsets. The model-to-model comparison is the paper’s rigorous core; the human line is a reference marker. The wiki should quote the ranking against MBA students with that qualifier attached, because it is the most quotable and least controlled number in the paper.
  • The composite score is a visualisation, not a measurement. Min–max normalised within the data set, equally weighted by choice, and explicitly “sensitive to the specific weighting scheme applied” — it has no meaning across data sets and could shift with different weights. The underlying three metrics carry the finding; notably, the frontier models’ failure is legible in the raw profit-versus-emerging-revenue split without any compositing, which is what makes the exploitation-bias story robust.
  • Is the regression about capability or about tuning? The paper offers a hypothesis rather than a demonstrated cause — that models tuned for peak performance in coding, chat and science “may inadvertently converge on an inability to handle uncertainty.” Alternative readings it cannot rule out: safety/conservatism tuning producing risk aversion; instruction-following training producing the literalism the prompt experiment detects; or a scoring artifact in which profit-maximising is a defensible reading of an outcome-agnostic prompt. The authors partly test the last of these — the emerging-tech prompt improves frontier models only marginally — which weakens but does not eliminate it. Open question, and the most important one for the wiki: is exploitation bias a strategic deficit or a goal-inference deficit?
  • Single simulation, deterministic, non-agentic. BBB is one setting with a fixed black-box structure, no true multi-agent competition (rivalry is simulated through news events and pricing pressure), and information is fed to the model by a structured program rather than sought by it. So the study measures strategic reasoning over supplied information, not strategic agency. Given how much of the wiki’s harness material turns on what an agent does when it controls its own information-gathering, this is a real scope limit — and the authors name agentic play in the human interface as a future milestone.
  • Contamination is reduced, not eliminated. The authors are precise: masking defeats recognition, and their inferences are about relative performance under the same masked representation, so residual leakage would shift levels rather than the differences. That is the right claim, and it is the claim the wiki should repeat.
  • Open question — does the regression persist or correct? The paper is explicitly built so this can be re-run. As of ingest the wiki has no follow-up. This is a standing ingest target: a re-run on models released after late 2025 would say whether mid-2025 was a blip in a tuning cycle or the start of a divergence.

What was actually ingested

Full article text, pp. 93–117, converted from the publisher PDF with pdftotext via the zotero-acquire skill (Zotero item F7CUFQT2, collection ai-wiki), including Appendices A–F headings and the model-alias table. Figures 1–6 carry the quantitative results and are images — the per-model composite scores, confidence intervals, the GPQA and LM Arena scatterplots, and the raw-metric rankings were not machine-readable in the conversion. All numeric claims above are therefore taken from the narrative text, which states directions, orderings and named models but not the underlying values; no figure values are quoted. The online appendices 1–2 (simulation screenshots, the contamination-test transcripts) are hosted separately at the DOI and were not retrieved. Direct fetch of the article page returns HTTP 403 to automated clients; the source was acquired via the user’s Zotero library, and the PDF is gitignored at raw/papers/<slug>.pdf.