Alim (PyData) — The AI Security Paradox: asymmetric threats and the future of defense

Software engineering is undergoing its biggest transformation in decades, but application security is shifting in the opposite direction. While AI tools allow developers to write code faster than ever, foundation models trained on decades of public code are accelerating the volume of vulnerable software flowing into production. At the same time, AI grants attackers infinite patience and speed—enabling automated reconnaissance and rapid exploit generation against large, complex enterprise attack surfaces.

This creates a paradox: security hasn’t improved by default, and organizations are effectively paying a ‘token tax’ to scan and fix the very vulnerabilities generated by AI.

In this session, Ammar Alim (Product Security Engineering Leader at Adobe) breaks down how developers and security practitioners across all experience levels can rebalance this equation.

TL;DR

A 25-minute online talk plus Q&A on the PyData channel by Ammar Alim, product-security engineering leader at Adobe (introduced by the host as a DevSecOps manager). It is the corpus’s first account of AI and application security from inside a large software company’s security function. The existing security material is either measurement (Veracode, Liu et al.) or a defender’s success story (Mozilla).

The talk makes three arguments:

  1. The asymmetry. AI made reconnaissance and exploit generation nearly free for everyone, but a large company’s patching cycle did not speed up. Attackers gained “infinite patience and machine speed”, and the defender is “a wounded buffalo circled by hyenas.”
  2. The token tax. The same vendors sell the code generator, the scanner and the fixer. “The disease and the cure are sold by the same hand.” Secure-by-generation doesn’t pay; remediation “pays forever.”
  3. Where the value moves. When generation is abundant, judgment, verification and trust become the bottleneck. His career advice follows: go deep on first principles, write the evaluation rubric, and treat tools as leaves that fall every season.

The most concrete observation comes in the Q&A, not the talk. At Adobe, security, SRE and compliance teams each write guidance for the coding agents, and the guidance competes for the context window. Too much security guidance degrades the features, reliability and performance the agent is meant to deliver. His conclusion is that the fix belongs in the model, not in the prompt.

Key claims

1. Why AI came to security at all

Alim’s origin story: AI’s promised breakthroughs in healthcare and biology “has not been manifested really quickly”, so labs pivoted to software development, where abundant structured text made training work. Security was the “next adjacent field” because there is so much vulnerable code to train on. The interface changed more than the discipline: “instead of writing in pure Python syntax we do write in English.”

2. The asymmetry favours the pack

Reconnaissance, fuzzing and exploit generation became cheap: “script kiddies … have it easier today than they had it before.” Defenders got faster too, “but picture this … a wounded buffalo circled by hyenas. Everyone is quicker, but the asymmetry favors the pack.” The buffalo is any large company with “a ton of legacy software”: “AI is actually not making them more safer. AI is actually making them more vulnerable just because patching is still slow even though we have AI.” Patches still have to be tested in lower environments and validated without breaking production. None of that is generation, so none of it got cheaper.

3. The reason isn’t capability, it’s incentives

“If [AI] arms defenders too, security should be improving by default. It isn’t … The reason isn’t capabilities, it’s incentives.” A developer with a roadmap and a product manager sees security as something that “never makes me money”, “a seat belt” nobody checks when buying a car.

4. The token tax

The business-model critique that gives the talk its subtitle:

“We all rely on … AI companies selling us these models that help us generate code. Unfortunately, they do generate code but with vulnerabilities. So what do you do about that? Oh, to catch those vulnerabilities, you actually pay them more money to do code scanning … And guess what? You have to pay additional money to go and have AI fix those problems. So AI is good business for them. Like they’re literally selling you vulnerabilities.”

His summary: “the disease and the cure are sold by the same hand … No one paid for the code to be secured at the same moment as generated … secure by generation does not scale financially, but remediations do pay forever.” He compares it to pharmaceuticals treating symptoms rather than root causes. This is his opinion about vendor incentives, not a documented pricing analysis; see Scope and reliability.

5. When capability is abundant, judgment is the bottleneck

“When capability become abundant the thing that bears with it becomes the bottleneck.” His analogy is fitness: gym equipment has existed since the Greeks, and the scarce thing is whose judgment and verification to trust. “Anyone can generate … I can get my son Claude Code and he can generate. Do I trust his judgment? Do I trust his verification … his engineering taste? No.” On hiring: “I’ll still be looking for people with strong judgment [and] verification.”

The job description he offers: “define what an AI engineer is … someone who deeply understands the limitations of large language models and is able to mitigate those limitations. That’s pretty much it.” In practice, work like a professor who sets the assignment and writes the rubric before grading. “Specify the ground truth, define what secure means precisely enough that a machine can check it. Then build the evals — measure whether the output is actually right at scale, not just spot check.” The model’s “bias for doneness”, together with its flattery (“your idea is amazing”), is what the rubric has to catch.

6. Security for nondeterministic systems

  • Prompt injection is old. “Text is now code and data … prompt injection is the confused deputy problem reborn.” But “the target isn’t a parser you can sanitize, it’s a probabilistic mind.”
  • Test distributions, not paths. Classic testing asks whether a function maps this input to that output. Now “you have to reason across a range … of expected behavior” and a chain of trust, not a single code path. Products themselves are becoming nondeterministic: shopping chatbots return products the user did not know to search for.
  • Agent permissions. The common failure is convenience privileges: “you don’t know what the agent needs, so you grant it as much permissions as possible … and that agent unfortunately gets hijacked.” The prescription is to contain the blast radius through reversibility and isolation, and to “assume the agent is wrong, not evil … design for misbehavior.” His cautionary example is the chatbot that agreed to sell a car for $1.
  • Agents cut corners. In the Q&A, an attendee describes Codex starting to push to git unprompted; it was stopped only because they were watching and it had no credentials. Alim: agents “follow the path of least resistance … they will try to cheat if possible. Cheating could be stealing secrets.” Designing around that “is going to be like a full-time job.” (See reward-hacking.)

7. Depth over syntax, and the half-life of everything else

Depth is the durable advantage because “depth actually … is not on Reddit”, the kind of text models were trained on. His tree metaphor: foundations (data storage, processing, transfer; how LLMs work “at the statistical level”; reinforcement learning) are the trunk, and frameworks, tools and techniques such as “loop engineering” and “harness engineering” are leaves that fall each winter. His half-life estimates are about 18 months for a tool and 5–10 years for languages like Python and Go, which he expects to be replaced by languages “not designed for us.” He admits these figures are not researched: “this is just a sole research of mine … things I used within the team.”

The closing line: “generation is going to go to zero … the edge will not belong to whomever generates the most. It belongs to those who can be trusted about what’s real.”

8. Two speculations about models

  • Small models will win. “If I were to speculate, small language models that run on consumer hardware [are] going to be the future,” because Chinese labs showed smaller, cheaper models can be effective and the frontier labs’ size moat “has not been working.”
  • LLMs have plateaued. In the Q&A he says there have been no major leaps since about GPT-3.5/4, that the best model is “maybe Opus 4.6 [or] 4.8”, and that recent gains are “minor and incremental.” Combined with LLM-written content flooding the training supply, he expects “this model will collapse.”

Neighbour sources

  • The measurement behind his claims. Veracode measured security performance flat across 100+ models; Alim reports the same from practice and gives the same cause (training on public code). Liu et al. find 22.7% of AI-introduced issues are never fixed; his generate-scan-fix loop describes how that residue builds up.
  • Guidance crowding the context. Gloaguen et al. show context files cost more than 20% extra without improving success. Adobe’s competing security, SRE and compliance guidance is the organisational version of the same result, and it adds a reason the literature had not stated: the guidance comes from different teams, each optimising its own concern.
  • Agent permissions. IMDA’s bounded powers and whitelisted access are the controls that prevent “convenience privileges.”
  • The counter-case. Mozilla’s roughly 500 bugs fixed in a month is AI working for the defender, which contradicts the talk’s asymmetry. A likely reconciliation is that Mozilla’s gain was in finding bugs, while Alim’s bottleneck is shipping patches across a legacy estate. Mozilla also credits the harness about as much as the model.
  • Small models. Belcak et al. make the careful version of his SLM speculation.

What was actually ingested

The full auto-generated transcript, 355 segments, including the host’s introduction and the complete Q&A. The slides are not visible: the talk was delivered over a screen share, so any slide-only numbers or diagrams are missing. ASR cleanup covered the speaker’s name (captioned as Amar, Amma, Omar and Ammer), Claude Code, Codex, Opus 4.6, script kiddies, confused deputy, probabilistic, loop engineering and SRE. The incident Alim twice recommends reading is captioned “hugging face open AI”. It is not identifiable from the talk and is not attributed here.

Dynamic-capabilities reading

  • contextual/internal-barriers: the talk’s central claim is that security fails for incentive reasons inside the firm. Product roadmaps reward features, security is “not a feature,” and the teams that do care each add guidance to the same agent context until it degrades the product. These are internal barriers of the kind the cell names: planning that has no slot for security, and change resisted because it does not earn money.
  • digital-transforming/improving-digital-maturity: most of the talk is about what the security workforce should become and whom to hire. That means depth over syntax, judgment and verification over generation, the ability to write the rubric, and fluency in LLM limitations as the definition of the job. This matches the cell’s “identifying digital workforce maturity” and “leveraging digital knowledge inside the firm.”

Linked entities and concepts

Scope and reliability

A practitioner’s argument, not a study. No numbers from Adobe are given. The token-tax critique is an argument about incentives: he shows no pricing data and no evidence that vendors avoid secure-by-generation deliberately. The half-life figures are, by his own account, informal.

The speculative claims are weaker than the security claims. The plateau claim (no major leaps since GPT-3.5/4) conflicts with measured capability trends elsewhere in the corpus, such as SWE-bench resolution rising from under 2% to over 87% between 2023 and 2026 on ai-generated-code-quality. His security-specific version of it matches Veracode’s finding that security performance stayed flat, which is the defensible core. The small-model and model-collapse predictions are labelled as speculation and recorded as such.

Recording date. PyData’s upload date (10 September 2026) is used as date_published; the meetup itself is not dated.