Why Agentic Attacks Are a Harder Problem Inside a Bank

Security vendor Sysdig has documented an incident in which an autonomous AI agent exploited a vulnerability, navigated an unfamiliar environment and exfiltrated a production database, all without human intervention.

Pramin Pradeep, CEO of BotGauge

Pramin Pradeep, CEO of testing firm BotGauge, calls the incident more than an isolated attack. “The Sysdig attack isn’t a warning shot, it’s proof of concept” he said when the company offered him for interview. The Fintech Times put written questions to him on what that means specifically for banks, payment processors and regulated infrastructure, on whether his own product category is the honest answer to it, and on what would prove him wrong.

1. Everything in the pitch is about software estates in general. Our readers run banks, payment processors and regulated infrastructure. What has Pramin actually seen in financial services specifically, and what makes a bank’s estate a different problem from a generic one?

Three things make financial services specifically harder. First, the attack surface is unusually credential dense. A bank’s production environment holds API keys to payment rails, cloud credentials, core banking system tokens, and inter-bank authentication certificates, all in close proximity. JADEPUFFER’s first move after gaining access was sweeping the environment for exactly this class of credential. A generic enterprise might lose an AWS key. A bank loses access to SWIFT connectors, card processing APIs, and settlement infrastructure simultaneously. Second, legacy and modern systems run in parallel in ways they do not in other industries. A typical large bank has core banking infrastructure that is decades old sitting directly behind APIs built last year, sometimes connected by middleware nobody fully understands anymore. AI-generated code accelerating into that environment creates shadow logic layered on top of already opaque
systems. Third, regulatory traceability requirements mean that when something goes wrong, you need to explain every action the system took. When AI-generated code or an AI agent produces an unexpected outcome, that audit trail is often exactly what is missing.

2. Pramin calls the Sysdig-documented incident a proof of concept rather than a warning shot. Attackers have used heavy automation for years. What could that agent do that a well-built scripted toolkit could not? We want the specific capability, because that is the whole of the argument.

The specific capability is decision-making under uncertainty. A scripted toolkit follows a decision tree. If the environment does not match the expected state at a given branch, the script stalls or fails. The JADEPUFFER agent encountered unexpected conditions, services that did not respond as expected, credentials that did not work on the first attempt and reasoned about alternatives. It retried with modified parameters, read error messages and adjusted its approach, and chained together weaknesses that no single script was written to
exploit in combination. In a bank’s environment, this matters enormously because no two estates look the same. Scripted attacks work at scale against common configurations. An agent works against your specific configuration, including the ones that are unusual, poorly documented, or misconfigured in ways that only exist in your environment. That is the capability gap.

3. “Static detection is structurally obsolete” is a strong phrase. Signature and pattern-based tooling still catches the large majority of what actually hits a bank every day. Is obsolete the word he means, or does he mean insufficient on its own? We would rather print the precise claim than the punchy one.

Insufficient on its own is the precise claim, and it is worth being exact about it. Signature and pattern-based tooling catches the large majority of what hits financial institutions every day, and that will remain true. The argument is not that it stops working, it is that it has a structural blind spot that is growing. Signature detection matches behaviour against what has been seen before. An agent that reasons and adapts during an attack generates behaviour that, by definition, has not been seen before in that specific sequence and context. That class of threat passes through signature detection not because the tools failed but because they were not designed for it. For a bank running both, the question is not whether to replace signature detection, it is whether to treat the gap it leaves as acceptable. Given the credential density I described in the first answer, I would argue it is not.

4. Continuous behavioural runtime coverage is what BotGauge sells, so this is also a description of the product. Can Pramin make the case in a form that would still hold if he worked somewhere else? Our readers will spot the alignment, so it is better met head on.

Yes, and it is worth doing directly. The underlying argument has nothing to do with BotGauge. It is this: you cannot detect deviation from expected behaviour unless you have established what expected behaviour looks like, continuously, at runtime. That is true regardless of who builds the tooling or whether it exists as a product category yet. Banks already do this partially, transaction monitoring systems establish behavioural baselines for customer activity and flag deviations. The argument is that the same logic needs to apply to
the software systems and AI agents running the bank’s infrastructure, not just to customer transactions. If transaction monitoring did not exist and someone proposed it, the argument for it would not be “this is what our product does.” It would be “you cannot catch fraud without knowing what normal looks like.” The case for runtime behavioural validation of AI systems is identical. BotGauge builds in that space. But the need exists independent of whether we do or not.

5. Behavioural detection trades false negatives for false positives. In a bank, a false positive means blocking a legitimate transaction or a legitimate customer. What is the real rate he sees, and what does it cost the firms running it?

Published false positive rates from behavioural detection systems in financial services vary enormously depending on how well the baseline is established and how narrowly the detection scope is defined. Poorly tuned behavioural systems in production banking environments can generate false positive rates that make them operationally unusable, rates that block legitimate transactions at a frequency that creates worse customer and operational impact than the threat they were designed to catch. The honest answer is that
the false positive problem is real, it is the primary reason behavioural detection has not been more widely adopted in financial services, and anyone claiming a specific low rate without specifying the detection scope and tuning methodology is not giving you a number you can rely on. What I would push back on is the framing that this is a reason not to deploy behavioural detection. It is a reason to deploy it carefully, with narrow initial scope, rigorous baseline establishment, and explicit false positive budgets agreed with operations teams before go-live.

6. Defenders can point agents at this too. Over the next five years, does the asymmetry favour the attacker or the defender, and why? Most people assert the attacker without arguing it.

The attacker, but not for the reason usually given. The common argument is that attackers only need to succeed once. That is true but not the specific asymmetry that matters here. The structural advantage attackers have with agentic AI is iteration speed. A defensive team deploying behavioural detection in a bank goes through procurement, integration, compliance review, tuning, and operational approval, a cycle measured in months. An attacker iterates their agent’s approach in hours. The attacker’s feedback loop is faster by an order of magnitude. CrowdStrike‘s 2026 data showed average attacker breakout time dropped from 62 minutes to 29 minutes in a single year. That compression will continue as agents improve. The defender’s feedback loop is constrained by organisational process in ways the attacker’s is not. The only thing that changes this asymmetry is if defenders deploy agents that can also iterate and adapt at machine speed and for regulated financial institutions, the compliance requirements that govern what defensive systems can do autonomously will remain a meaningful constraint. That is the uncomfortable structural
reality.

7. What would prove him wrong? If in two years autonomous agent attacks on financial institutions remain rare and conventional tooling is still catching almost everything, what will he have got wrong? We would rather have a specific and uncomfortable answer than a diplomatic one.

The specific thing that would prove me wrong is this: if autonomous agent attacks on financial institutions remain rare two years from now not because defences improved but because attackers did not adopt the model at scale. JADEPUFFER proved the technical capability. What it did not prove is that the model will be widely adopted quickly. Scripted attacks are cheaper to run at volume and still effective enough against the majority of targets. If the economics of agentic attack tooling remain too high for most threat actors, if it
stays a capability available only to sophisticated state-linked groups rather than commodity ransomware operators, then the frequency argument does not hold and conventional tooling remains sufficient for the realistic threat distribution. I think that window is closing, and the commoditisation of LLM access makes it close faster than people expect. But if in two years the base rate of agentic attacks on banks has not materially increased, the honest conclusion is that I overestimated adoption speed, not that the threat does not exist.

The post Why Agentic Attacks Are a Harder Problem Inside a Bank appeared first on The Fintech Times.

Read More

Leave a Reply

Your email address will not be published. Required fields are marked *