The AI shift most companies didn’t see coming

What does it take to move the needle with AI? Top AI voice Allie K. Miller joins WorkLab to share why some of the biggest advances aren’t coming from the usual places—or people. She explains how organizations can unlock new value by giving unconventional thinkers room to experiment, why speed and iteration matter more than perfection, and how leaders can turn edge ideas into real business impact. Plus: how agentic workflows, voice interfaces, and cross-disciplinary teams are reshaping the future of work.

TL;DR

A ~41-minute live WorkLab recording from the Copilot Summit on the Microsoft channel, published 10 August 2026 — host Molly Wood with Allie K. Miller, founder and CEO of Open Machine, who advises Fortune 500 C-suites on AI transformation. Load-bearing claims, following the episode’s ten chapters:

  1. The thesis, stated in the cold open: “AI is not a tool.” “If you think of AI as a tool, you are going to budget as if it is SaaS or traditional procurement, roll it out as if it is a traditional SaaS platform, measure its impact solely based on productivity — and you’re not going to think of brand new business lines, org restructuring, process reinvention. None of that really happens if you’re just rolling out a productivity tool.”
  2. The third shift, and why it landed as a cost shock. Miller’s periodisation: (1) late 2022–2023, general chatbot adoption; (2) end of 2024, big reasoning models; (3) the last several weeks, multi-agent — “this idea that I can actually work autonomously on many tasks at a decent reliability level for an hour or more.” Her read is that “most leaders are still operating under 2024, 2025 assumptions.” What makes the third shift bite is spend: with frontier models and agents “running amok 24/7 and not 9 to 5, then your costs, instead of being thousands per head per year — there are companies spending thousands per head per day, right? Mostly in the engineering space.” She also cites a statistic that 74% of CEOs are worried they will be fired in the next two years because they handled AI wrong.
  3. The frontier unit — her structural answer to the cost shock. Rather than choose between company-wide adoption and siloed pilots, run both paths. The whole organisation keeps upskilling, experimenting within departments, and sharing in Teams channels and town halls. Separately, a unit of roughly 60–120 people gets a materially higher per-head budget, tests every new model “the second a new model comes out,” works “in teams of 2 to 8 to be able to actually get production prototypes done,” and hands winners back to the wider organisation. “You don’t really want 80,000 people spending thousands of dollars a day while you’re in the experimentation phase.” Later in the episode she adds the design constraints: cross-discipline, not all engineers (“I want to see the weirdos in marketing, in sales, in legal, in finance”) so prototypes do not die on arrival at a department; “almost no layers of management”; and ownership that varies by company — R&D/innovation, or a single bought-in C-suite executive, often a non-technical CFO or CMO, but “always a passionate weirdo.”
  4. The leadership gap is literal. “I think the average C-suite has not built a single agent, and that has really, really bad implications for the rest of the business.” She reports doing more one-on-one C-suite workshops and coaching in the last six months than in the previous three and a half years, with CEOs “who kind of whispered to me in the hushed tone of, like, I haven’t used any of this stuff — can you help me?” Against that she cites a Microsoft study of 1,800 employees globally: where managers actively model AI use themselves, employees report a 22-point lift in critical thinking about their AI work and a 30-point lift in trust of agentic AI. Her summary: “strategy lives or dies at the middle layer. Adoption lives and dies at the employee in the cubicle.”
  5. Measure what people use AI for, not how many use it. “I hope all of you know your monthly active, weekly active, daily active AI users inside your org. But what I would love people to look at is what are they actually using AI for? Because if my whole company is using AI to rewrite emails… that is not meaningful AI adoption. You might have people who believe that they’re amazing AI users, and who are using AI every single day, and they’re using it at 5% of the power.”
  6. “Gremlin mode” — a three-level ladder of usage depth, with a worked example. Level 1 is the failure mode: one prompt, disappointment, “you think that AI is horrible, you think everything is overhyped, you don’t invest in it.” Her anecdote is a CMO in a board meeting who said the company does not use AI for writing “because it’s not good enough… I just feel like it doesn’t have heart” — whose prompt had been to ask it to be funny and heartfelt, without iterating. Level 2, giving examples and iterating, she still rates “10% level.” Full Gremlin mode, for a healthcare blog: take your ten top-performing posts, put each individually into AI and have it write a brief for the post, open a fresh thread and have it write from that brief, “literally have it loop eight times to be able to improve that blog post,” then save the final as the gold standard and reuse the brief-generation process for the next topic. “But that process was six hours, and you’re using AI’s goal features and looping features to be able to get to that. We’re not in this single-chat-thread land where you’re giving it one sentence.”
  7. Three hiring skills for the AI age. (a) High agency — “are they going to let the AI dictate their day? And if you didn’t get what you want, are you just going to walk away and complain and still do things manually, or are you going to go in bulldozer mode and try it from 17 different angles?” (b) A strong sense of wonder — “this has nothing to do with age, nothing to do with department… some people are two years from retirement who are the most curious people I’ve ever met.” (c) Systems thinking — “that’s been the really big change the last year, where we’re managing multi-agent systems.” Her structural point about who this disadvantages: an IC “who has never been a people manager might actually have a pretty difficult time managing a mini org, a temporary digital workforce of AI agents, because they just haven’t been exposed to that,” which puts early-career people “in a really tough position on that systems thinking side, very strong on the AI side.”
  8. How to book it, and how to measure it. On budget category she is deflationary: it will keep landing in employee salary plus tool-and-tech procurement under the office of the CIO, as cloud did, and “it might take years for that to trickle into enterprises.” What should change is measurement — pull the spend out to measure it differently, and look past employee productivity to “engineering velocity, net new ideas that have been tested out, how trustworthy your customers feel about you, or how fulfilled your employees feel. But growth is the one that I think it should be measured on, not just productivity.
  9. “Cannibalize your business lines” — and the analyst who made millions in a month. The worked example: a company with a single business type where “one analyst, in his free time, decided to use generative AI to build out a brand new auditing product for the energy usage for their clients. They saved their clients 1 to 3%. The reward, the business model, was that they took 50%. And this one person’s idea has made millions in its first month.” Her implication for allocation: “you might have to move some of your analysts away from whatever they’re doing now to be able to reallocate toward net new business creation.”
  10. The employee-side tension, quantified from Microsoft’s own index. “The recent Microsoft Work Trend Index… 65% of AI users say they are afraid of falling behind if they don’t move quickly, but 45% say it feels safer to focus on current goals rather than to redesign their work with AI.” Her prescription starts by ruling out the obvious lever: leading by fear “never works. You’ll have mutiny” — she cites a company that told staff to figure it out within three months or be out, after which “every single employee[‘s] productivity plummeted.” Instead: (a) find out what is actually blocking people — time, a VP who said “AI sucks, I hate it” over a beer, reorg anxiety, or simply the absence of a clear data-privacy protocol; (b) own it personally — hiring chief AI officers in 2023 was “the worst idea ever… you are offloading the responsibility of your entire company strategy and reinvention to one person who, by the way, is probably a brand new hire and does not have the background of your company”; (c) treat speed of iteration as “the moat of 2026 and beyond” — small experiments, small teams, fewer layers, lower bureaucracy, and a budget teams know they can operate freely within. Her competitive framing for why: the lean AI leaderboard of sub-50-person companies racing on revenue per employee toward “$10 million-plus, $20 million-plus… to be able to build a very small company that is valued at a billion-plus.”
  11. The agent org chart. Miller runs 34 agents structured like a company: a chief of staff, Simon, who has his own memory-and-documentation assistant, Toby; six direct reports named after Friends characters (Chandler on marketing, Rachel on client work, Phoebe as the wildcard); each with three to seven sub-agents, which can in turn spin up temporary agents she calls “civilians.” She mostly talks only to Simon. She reports a founder running the identical structure through a chief-of-staff agent called Maya over live voice while walking 18,000 steps a day, and a Silicon Valley startup where “every single person had a $60 microphone at their desk, and they’re dictating all day long in an open office setting.” Dictation, she claims, is “four times faster than writing.” Agents sit in Teams and Slack, and “they’re having agents teach other agents new skills.” Her conclusion loops back to the thesis: “this is why you have to believe that AI is not a tool. Because if you thought that it was a tool, would you think that you should set up a self-learning flywheel and have 60 agents kick off a thousand? Absolutely not. We’re having to treat it as this like operating system.”
  12. The AI layer: make the company queryable. “You have to make your whole company queryable in the AI age to be able to have this second brain and take action off it.” Her example is an executive who dictates five to forty minutes at the end of every day — what key decisions were made, what she is happy about, what she is stressed about, and the human nuances that “actually drive 80% of our decision making… maybe you’re recording meetings, but you’re not capturing [that] the general counsel is staring at their phone the whole time.” After months of accumulation she can ask AI to “look at my notes for the last three months, am I getting more stress, less stress, map out my energy” and adjust the following weeks accordingly.
  13. Find the weirdos. “Find these people who are operating at the edge… they might be engineers, they might be in HR, they might be gamers in their free time, they might be Gen X, Gen Z — it is in every single corner. But you have to figure out leadership ways to unearth these people. Maybe it’s a hackathon, maybe it’s having people submit different ideas, maybe it’s just looking at who’s sharing things in Teams or town hall.”
  14. The 1:5 budget ratio, and the least obvious dollar in it. On McKinsey’s guidance of $5 on people for every $1 on AI tools, Miller says it still holds. The $1 is “token maxing” — managing big models, small models, architecture. The $5 obviously covers upskilling, changed hiring processes and retraining people managers; “the less obvious dollar that is being spent is on incentives.” Her mechanism: pay real money as the prize in an agent-system hackathon — “that is still very much a people reward, but it is not just going on upskilling. You sneak in the upskilling into the hackathon, and in doing so you have upskilled your whole org.” Two pre-genAI cautionary tales support the people-side weighting: a retailer that spent millions on “the greatest prediction system ever” for inventory and got zero adoption, only afterwards asking location owners what they wanted; and a computer-vision project for manufacturing facilities designed to be entirely voice-enabled, where “you can’t do that when the machines are crazy loud.”
  15. The exponential, in one benchmark series. “On coding benchmarks, the recent frontier models… they are hitting 94. A year ago, we were hitting 72. And a year before that, we were at… seventeen.” Her use of it is procurement-facing: “if you’re making a multi-year procurement strategy, if you signed a three-year tech deal — that is where we are two years later.” The diagnosis behind it: leaders “feel the nausea, like they feel the acid reflux, but they don’t quite see the exponential.”
  16. The 30-day close. Start with the context layer: “crack the context puzzle. Are you connecting it into tools, SharePoint, Outlook? Are you having your meetings recorded and transcribed even if you’re in a regulatory environment? Are you having your executives dictate these things? Are you allowing for your employees to build up context docs? I have my AI interview me to build up context docs. So context, context, context. We are in the era of context engineering.” Second, make AI proactive: “you should have AI prompting you for things… whether it’s scheduled or trigger based or time based. If you are the only thing that is kicking off AI workflows, you are 3x already, because you’re only working eight hours and there’s 24 hours in a day.” And the thing to stop: “I would stop typing… I want people to move into crazy multimodal, gremlin-level communication — giving images, giving dictation… so that AI is functioning as an operating system.” She flags the rhetorical inflation herself: “I say this because then I think you’re going to reduce it by 10%. Like I have to set the bar high.”

What was actually ingested

The full human-curated English caption track (kind: manual, not ASR) — the cleanest transcript in this batch. All ten chapters present and consistent with duration: 40:39 / length_seconds: 2439. Fetch-pipeline note: two transcript panels were mounted on the page and every caption was scraped once per panel with different line-wrapping, so the raw capture contained 712 segments that collapse to 356 unique ones; the duplication was removed at acquire time and the fetch tooling fixed. The Copilot Summit venue framing, audience show-of-hands moments and the closing show trail are excluded from the substantive summary above.

Dynamic-capabilities tagging

  • digital-seizing/rapid-prototyping — the frontier unit is a rapid-prototyping structure with explicit parameters: teams of two to eight, “almost no layers of management,” a mandate to test each new model on release, and the goal of getting “production prototypes done” that roll out to a wider department. Gremlin mode is the same discipline at the individual level — an eight-loop iteration cycle treated as a six-hour investment rather than a single prompt.
  • digital-seizing/balancing-digital-portfolios — the double-path allocation is a portfolio decision made explicit: the 80,000 minus 120 get upskilling and departmental experimentation, the 120 get a materially higher per-head token budget, and winners are selected from the frontier unit and propagated. The McKinsey 1:5 ratio ($5 on people per $1 on tools) is the second allocation rule, with incentives named as the underweighted line item within the $5.
  • digital-transforming/redesigning-internal-structures — the 34-agent org chart (chief of staff, six named direct reports, three to seven sub-agents each, spun-up temporary “civilians”) is an internal structure in the literal sense, and Miller’s hiring criteria follow from it: systems thinking matters because ICs who have never managed people must now run a digital mini-org. The frontier unit’s cross-discipline composition and near-flat management are structural prescriptions, not tooling ones.
  • digital-sensing/digital-mindset-crafting — “find the weirdos” is a sensing mechanism aimed at internal edge practice (hackathons, idea submissions, watching who shares in Teams and town halls), and the manager-modelling finding (22-point critical-thinking lift, 30-point trust lift where managers use AI themselves) is the mindset-crafting mechanism. The exponential-benchmark exhibit (17 → 72 → 94) is used precisely to shift leaders’ internal model of the rate of change.
  • strategic-renewal/business-model — “cannibalize your business lines” is a business-model renewal prescription, with the energy-audit product built by one analyst in his free time — sold on a 50% share of client savings — as the worked example, and the reallocation implication that analysts should be moved off existing work “toward net new business creation.” The measurement argument follows: growth, not productivity, as the metric the investment should be judged on.

Linked entities and concepts

  • Microsoft — publisher (WorkLab, recorded live at the Copilot Summit); also the source of the 1,800-employee manager-modelling study and the Work Trend Index figures cited. Updated in this ingest.
  • McKinsey & Company — origin of the 1:5 people-to-tools budget ratio Miller endorses.
  • Argenti — the same optimisation-versus-transformation lever from the CIO chair; see this source’s relationships:.
  • Netflix — systems thinking as the rising hiring criterion, reached three weeks earlier from an operator vantage.
  • YC — make-the-company-queryable, at startup scale rather than enterprise scale.
  • Wood — same show, the historical form of the same bolt-on argument.
  • BBC AI Decoded — the Gremlin-mode CMO as a worked instance of the adoption gap.
  • micro-productivity-trap — “AI is not a tool” as a budgeting-and-measurement account of the trap; the growth-not-productivity metric; the 65%/45% Work Trend Index split as the employee-side mechanism.
  • enterprise-ai-adoption — the frontier unit, the double-path allocation, token maxing, the 1:5 ratio, and the depth-of-use measurement critique.
  • systems-thinking — named as the third and most-changed hiring skill, with agent-fleet management as the reason.
  • durable-skills — high agency, sense of wonder and systems thinking as the three interview criteria.
  • ai-agents — the 34-agent org chart, agents teaching other agents, and proactive trigger-based agent invocation.
  • agentic-engineering — Gremlin mode as an iteration discipline, and the context layer as the prerequisite (“we are in the era of context engineering”).
  • ai-employment-effects — the claim that early-career ICs are structurally disadvantaged at managing agent fleets because they have never managed people.
  • dynamic-capabilities — the five cells tagged above.

Dangling (single-source mention, deferred per author-entity promotion): Allie K. Miller, Open Machine, Molly Wood.

Source quality note

Human-curated caption track, so transcription quality is high; a handful of residual errors were corrected at acquire time (Allie K. Miller, Work Trend Index, 22-point lift, agentic AI — see the raw file’s notes:).

This is vendor-published content recorded at the vendor’s own customer conference. WorkLab is Microsoft’s show, the Copilot Summit is a Microsoft event, and the audience is Microsoft customers — the incentive runs toward depth-of-adoption arguments. Miller is an independent advisor rather than a Microsoft employee, but the two Microsoft-sourced statistics she cites (the 1,800-employee manager-modelling study, the Work Trend Index 65%/45% split) are first-party research presented on the publisher’s own channel and were not independently checked in this ingest.

Nearly every other number is unsourced or unverifiable as stated: the 74%-of-CEOs-fear-firing figure has no attribution; the analyst-made-millions-in-a-month case is anonymised with no denominator; “four times faster than writing” for dictation is asserted; and the coding-benchmark series (17 → 72 → 94) names neither the benchmark nor the models, which makes the exponential exhibit rhetorically strong and analytically unusable. Miller is candid that some of her prescriptions are deliberately overstated to move behaviour (“I have to set the bar high to then let’s shoot for the stars, settle on the moon”). Treat the structural contributions — the frontier unit, the Gremlin-mode ladder, the agent org chart, the context-layer prescription — as the durable content.