How Agentic AI Is Reshaping Finance Workflows, Processes, and Governance

How is agentic AI changing investment management? In this CFA Institute roundtable, Rhodri Preece, CFA, Senior Head of Research at CFA Institute, is joined by Brian Pisaneschi, CFA, and James Tate to explore how agentic AI is transforming investment workflows, financial data analysis, governance, and decision-making.

— Channel description, CFA Institute

A ~27-minute roundtable from the CFA Institute Research and Policy Center, structured as three questions put by Rhodri Preece to Brian Pisaneschi (director, Applied Investment Practice and Tools) and James Tate (data science researcher): what tools are actually available, how to handle incomplete financial data, and what governance and ethics look like when the tool has judgment.

Its value to the wiki is that it is the corpus’s clearest regulated-profession vantage on agentic AI — a professional body with a code of ethics and a fiduciary frame reasoning in public about the same primitives the vendor and practitioner sources describe. The governance segment supplies something the wiki does not otherwise have: an argument that LLM biases mirror the specific behavioural biases the profession already has a literature on, which cuts directly against the assumption that delegating to a machine launders human bias out of a decision.

TL;DR

  • Skills plus MCP have displaced hand-built agent frameworks — in one year. Pisaneschi: “a year ago [I was] building these agentic AI solutions within Python… I had to learn these different Python libraries.” Now, skills — “just these markdown files… semi-structured text that outline a very specific workflow” — plus MCP servers, which let the model “almost natively interact with outside data APIs and the like.” A skill can carry resources with it: an Excel valuation model to pre-populate, or a connector that pulls and manipulates outside data.
  • The de-skilling of agent-building, stated as an access claim. “If you have some sort of repetitive task, you can come up with a skill file and you don’t need to know how to program… you don’t need to be a highly technical person.” With the honest qualifier: “However, if you do have those abilities, you can make them far more robust” — custom MCP servers, more capable workflows.
  • The interface inversion. Preece names it explicitly: traditionally a practitioner moves between Excel, a Bloomberg terminal and several platforms, extracting and then analysing. Now the AI platform “sits above all of these” and is the primary point of interface. Pisaneschi confirms — once sources are connected, the agent decides whether an incoming question matches a defined task, pulls the relevant skill and starts that workflow.
  • The compounding claim, and it is the most quotable thing in the source. Iterating a skill when it fails — “you screwed up there, don’t do that again” — is individual. The collective version is the point: the same task performed by many analysts in a firm, or across the industry, with everyone improving the same skill file, means “even if the models don’t get any better at all… we can really create incredibly robust workflows just by our own knowledge and iteration.” Capability growth decoupled from model progress.
  • Three tiers for missing financial data (Tate). Simple gaps in a return series: linear interpolation. A panel with macroeconomic variables and company information: k-nearest-neighbours imputation. The interesting tier: synthetic data from a trained generative model — train on historical returns jointly with macro variables (inflation, yields, GDP, employment), let the model learn the relationships, then simulate. Preece draws out the methodological claim and Tate confirms it: “you don’t have to prescribe a functional form on the data or assume a given distribution” — no linearity assumption, the model is data-driven. The application is scenario simulation, portfolio stress testing and backtesting, and the argument for it is regime change: bootstrapping from the empirical return distribution assumes the past distribution still holds; a generative model conditioned on macro variables can be asked what happens if it does not.
  • Open versus proprietary — a three-part answer. (1) Confidentiality forces the question for many firms: “plenty of people in our industry are just focused only on using open source because they can’t send any of their information and data outside.” (2) The lag is real but modest — open models trail by roughly three months on Pisaneschi’s estimate. (3) The differentiator has moved to the harness. “It’s almost turned to a lot of the value being… they call it harness, but it’s essentially how the architecture underlying the agentic framework is.” His test case: take an open model, “[throw] it into the open source harness, it is not going to quite… be as good as Claude Code” — because the closed vendors have built real parallelisation and optimisation into theirs. His read of the market: “it’s a winner take all scenario.”
  • What open source is actually for: cost optimisation. “When you’re dealing with a lot of data, if you just send API calls to OpenAI every single time, that’s going to be a lot of money.” Instead, train a small language model on your dataset — “garbage in, garbage out” on the training set, but “if you can get a good enough model… you can run it locally and you’ve saved tons of money… it’s practically free. So it becomes a cost optimisation problem with open source rather than the actual capabilities itself.”
  • The scalability limit, from someone who hit it (Tate). Methodologically he did the right thing: use the proprietary model as the baseline to beat, run open alternatives through the identical process, compare. Result: “with a much smaller large language model that was open source, I was able to replicate more or less what the proprietary OpenAI model was able to do for the task.” The failure was throughput — a ~36B model “still took a very, very long time,” against a proprietary batch API taking 50,000 requests and returning within 24 hours. “It’s very very difficult to replicate that locally unless you invest a lot in the infrastructure.”
  • The governance section’s central finding: model bias mirrors investor bias. Tate names two layers. Pre-training skew — “a very heavy skew towards Western data sets, of American stocks, tech stocks”, surfacing as unexplained default recommendations. And the sharper one: “a lot of research has shown that the biases in these large language models basically mirror investor biases. You know, they’re trained on human generated data. So naturally they mirror human biases.” The worked example: LLMs have been shown to exhibit loss aversion, so given a portfolio with gains and losses “they’re more likely to hold on to the losses when potentially they should… delegate those funds to a better investment.” He is careful about status — “that’s something that’s still being actively explored in the research as to what are the actual implications” — but the practical instruction is unambiguous: a manager using these models “needs to be aware that they have these limitations.”
  • Preece draws the inference that matters. The behavioural finance literature establishes that humans have biases; “there is often a misplaced assumption that if I can delegate some tasks to a machine tool… that would be a way to negate some of the human biases.” With LLMs, “that’s not the case.” Bias must now be managed on both sides of the delegation.
  • The mitigation: skills as bias controls, red-teamed. Pisaneschi’s contrast is instructive: rules-based tactics — he uses a stop-loss order as the minimal example — mitigate human emotional bias precisely because they remove discretion. LLMs reintroduce “decision making authority that is subject to the biases.” The remedy is to specify what the model must not decide implicitly: give it explicit screening criteria rather than let it “implicitly choose one of those ratios”; red-team the model for its biases — “what the cybersecurity industry does to try and essentially hack their own system” — and encode the mitigations in a skill file. “If that control is predictable, then you can absolutely reduce the human bias that we have, as well as reduce the bias that the large language model has.”
  • Trust as the object of governance, and evals as the mechanism. “Everything about governance, everything about ethics within the actual technical AI side is really just about gaining trust.” One expert judging one output is “just like kind of an employee” — trust accrues over repeated tasks. The scaling question is the real one: “how do you gain trust… at scale for an entire firm? You give it lots and lots of tasks. You run evaluations. You create the distributions of outputs… and then you gain trust by reviewing the outputs and the evaluations.”

Why this matters to the wiki

1. It is the corpus’s first substantive statement of the harness argument from a regulated profession’s research arm. agent-harness is dense with vendor, practitioner and academic sources; this adds a professional-body vantage whose obligations are fiduciary rather than commercial. Notably, Pisaneschi arrives at the harness as the locus of competitive advantage independently and reluctantly — he is describing why he cannot simply substitute an open model, not selling a harness.

2. The bias-mirroring finding gives responsible-ai a mechanism it lacks. Most of the wiki’s responsible-AI material concerns representational bias, transparency and accountability. This supplies a domain-specific, testable claim — that LLM decision biases recapitulate the documented biases of human decision-makers in the same domain, with loss aversion as the named instance — and the corollary that automating a judgment does not launder the bias out of it. For a corpus that carries automation-vs-augmentation and agent-oversight-and-delegation, that is a substantive addition: it means the case for human oversight cannot rest on the human being the biased party.

3. “Robust workflows even if the models never improve” is the strongest form of a claim the wiki holds in weaker versions. agent-harness records that harness variance is large on fixed weights. This source states the organisational consequence: capability can compound through shared, iterated procedure independent of model releases — which is a claim about where a firm’s AI investment should go, and one a professional body is well placed to make because its members share tasks across competing firms.

4. The synthetic-data segment is new territory for the corpus. Nothing else in the wiki addresses generative models as a data-generating instrument for scenario analysis, or the argument that regime change is what makes empirical bootstrapping inadequate. It connects to strategic-foresight from an unexpected direction — scenario planning implemented as a trained conditional distribution.

Dynamic-capabilities reading

  • digital-sensing/digital-scouting — The roundtable format is the sensing function of a professional body: a research centre scanning what is newly available (skills, MCP, small local models, generative data synthesis) and reporting it to a membership that cannot each scan for itself. Pisaneschi’s “the thing I’m talking about most every day now is skills” is a scouting report.
  • digital-seizing/rapid-prototyping — The skill-file workflow is explicitly iterative: build, let it fail, correct, encode the correction. Tate’s open-versus-proprietary comparison is prototyping as method — baseline with the proprietary model, run the alternatives through the same pipeline, compare.
  • digital-transforming/improving-digital-maturity — The access claim is a maturity intervention aimed at a whole profession: analysts who cannot program can now build agentic workflows, and the compounding argument depends on many of them doing so.
  • digital-transforming/redesigning-internal-structures — The interface inversion is a structural change to how investment work is done: the AI platform sits above the terminal and the spreadsheet rather than beside them, which relocates where analysis happens and what the analyst’s day looks like.
  • contextual/internal-barriers — Named concretely: data that cannot leave the firm; open-model throughput that cannot match a batch API without infrastructure investment; and the governance barrier of fiduciary duty and client confidentiality, which Preece raises as the thing organisations “are struggling with.”

roles: overrides the cell defaults toward the technical and risk-owning C-suite plus the research function, because the source’s audience is investment professionals and their governance leadership rather than a general executive audience.

Linked entities and concepts

  • Entities: Anthropic (skills as the named standard; Claude Code as the harness benchmark), OpenAI (batch API as the throughput baseline). CFA Institute itself remains deferred — it is the author: on this single source, and the promotion rule waits for a second.
  • Concepts: agent-harness (the harness as the durable advantage, from a reluctant witness), responsible-ai (bias mirroring, red-teaming, evals as trust machinery), open-source-ai (confidentiality, lag, cost optimisation, the throughput limit), enterprise-ai-adoption (non-programmers building workflows; the interface inversion), agent-oversight-and-delegation (delegation does not remove bias), automation-vs-augmentation (rules-based tactics remove discretion; LLMs reintroduce it), ai-benchmarks and strategic-foresight (evals; generative scenario simulation), document-intelligence (skills carrying spreadsheet artifacts).
  • Dangling (single-source mention, deferred): Rhodri Preece, Brian Pisaneschi, James Tate, CFA Institute and its Research and Policy Center, Bloomberg (as terminal), DeepSeek. Promote on a second citing source per the author-entity promotion rule.

Debates and supersession

  • The “leaked their entire code base” claim should not be repeated as fact. Pisaneschi says of Anthropic that “ironically, they actually leaked their entire code base and the open source community has taken that and run with it and created really great open source alternatives.” This is stated in passing, unsourced, and the wiki has no corroborating source for it. It is recorded here as something the speaker said, not as something the wiki asserts. The surrounding argument — that open harnesses exist and still underperform the vendor’s own — does not depend on it.
  • Open-model lag: three months here, six-to-twelve months elsewhere. Pisaneschi puts the gap at “probably around, I would say like three months”; Saad-Falcon at YC Paper Club says 6–12 months for locally runnable models. These are probably not in conflict — the frontier of open weights is not the frontier of what fits on a laptop — but neither speaker distinguishes the two, and the wiki should keep the distinction that they elide.
  • The winner-take-all reading conflicts with the sovereign-AI reading, and this is a live disagreement. Pisaneschi: if you don’t have the best model “you’re missing something,” and the harness advantage compounds it. Huang, three weeks earlier, argues the opposite — that open weights are now malleable enough to exceed frontier performance in-domain. Both speak from good vantage points. The disagreement likely turns on scope (general capability versus a narrow domain with proprietary data) and on whether you have the post-training capability to exploit malleability, but neither addresses the other and the wiki should not resolve it by fiat. Recorded as a contradicts edge.
  • The bias research is cited, not identified. Tate refers to “a lot of research” showing LLM loss aversion and bias-mirroring without naming papers, and flags the implications as unresolved. This is an ingest target: the underlying behavioural-finance-meets-LLM literature would let the wiki hold the claim at first hand rather than second.
  • Open question — does red-teaming for behavioural bias actually work? The proposal is coherent (find the bias, encode a mitigation in a skill, verify over a distribution of outputs) but no results are offered, and it inherits the general problem that a mitigation written into a prompt is a soft constraint. Whether skill-file bias controls survive contact with the model’s priors is untested here.

What was actually ingested

Full ~26.5-minute roundtable transcript (auto-generated English captions), after repairing a scrape defect at acquire time. The Playwright scrape returned 955 segments running to 36:22 against a declared runtime of 26:35; the final 124 segments belonged to a different video (on missing time-series data and imputation) whose transcript panel had also mounted on the page. The genuine talk ends at 26:19 with the host’s sign-off. The foreign tail was removed and the incident recorded in the raw file’s notes:. This is a new variant of the two-panel bug in the transcript skill’s failure modes — not exact duplication, which the skill’s dedupe catches, but a second distinct video’s segments appended; the tell was the last timestamp exceeding length_seconds. Speaker-name spellings follow the channel description (Rhodri Preece, Brian Pisaneschi, James Tate), which differs from the ASR rendering in at least one case. No slides or on-screen material were ingested.