Financial services firms are moving from simple chatbots to autonomous AI agents, and their AI bills are rising even as the price of individual model calls falls. At the same time, the EU’s Digital Omnibus amendments have pushed back some of the most demanding obligations in the AI Act.

Vikas Krishan, chief digital business officer and head of UK and EMEA at Altimetrik, spoke to The Fintech Times about why a rising token bill says little about value, what firms should measure instead, and why he sees the AI Act timeline as a question for the board rather than a breathing space.
Krishan’s starting point is that token consumption is evidence of activity rather than of return. “A rising token bill tells you that AI systems are doing more work, but it doesn’t tell you whether that work is useful,” he said.
He argues the distinction matters more in financial services than elsewhere because the workloads are token-heavy by nature. “Detailed compliance records, regulatory filings and transaction data all create large inputs before the model has even begun reasoning,” he said. “On top of that, financial institutions need more validation, audit logging and human review than many other sectors.”
“Higher consumption could mean successful adoption. It could equally mean inefficient retrieval, unnecessary model calls or routine tasks being sent to models that are far more powerful and expensive than they need to be,” Krishan said.
The question for leadership teams, in his view, is not how many tokens the business is consuming but what it is getting back for that consumption. “Rising usage alongside rising output and measurable business value can be healthy. Rising usage while outcomes remain flat is a more expensive way of standing still.”
Cheaper tokens, bigger bills
Falling model prices have not brought bills down. “The economics of AI have changed faster than the budgeting models businesses use to manage it,” Krishan said. “The cost of individual tokens has fallen dramatically, but the number of tokens organisations consume has exploded.”
He pointed to data from US fintech and financial operations platform Ramp, which he said found that token usage among businesses with connected AI grew by approximately 1,001 per cent between January 2025 and April 2026. Falling per-token prices did not offset that increase: by the same measure, total AI spend grew by 497 per cent over the period.
Much of that growth comes from the shift from chatbots to agents. “A chatbot responds when somebody asks it a question. An agent can plan, retrieve information, call tools, check results and retry tasks autonomously,” he said. “Agentic workflows consume substantially more tokens because a single task can involve repeated reasoning, document retrieval, tool calls and retries.”
Budgeting habits have not caught up either. The other problem, Krishan said, is that “companies are still mentally budgeting for AI like conventional SaaS. Seat-based software produces relatively predictable costs. AI consumption is variable and can continue in the background without a person actively initiating every call.” Falling model prices, he warned, “should not create false comfort: the unit is getting cheaper, but businesses are using vastly more units.”
Measuring outcomes, not consumption
Krishan is not arguing that tokens should go unmeasured. “Token consumption still needs to be measured, but it needs an outcome for it to align to,” he said. Depending on the use case, that might be the cost per resolved customer query, per document processed, per fraud case screened or per task completed, alongside completion rates, escalation rates from cheaper to more capable models, output quality against a human or pre-AI baseline, and variance against forecast.
Attribution is the second requirement. “Every AI call should ideally be attributable to a team, feature, workflow, use case and, for international businesses, market,” he said. “Without that, the business knows what its total AI bill is but not what is producing it.”
He added: “The model fintech businesses need is one with clear ownership, where questions can be asked at workload level: what did this work cost, and what did it return?”
On controlling costs, his advice is to start with visibility rather than optimisation. “If you cannot attribute usage to individual workflows, teams and features, you do not yet know what you should be optimising,” he said. Once that instrumentation is in place, he would look first at stable, high-volume workloads, routing tasks to cheaper models where they perform reliably and caching the lengthy, relatively stable compliance and policy instructions that are sent repeatedly.
Every optimisation, he said, needs a quality gate. “If a change cuts token consumption by 30 per cent and passes the same accuracy, validation and compliance tests, you have found a saving. If it cuts consumption by 30 per cent and nobody reruns those tests, you have created an unmeasured risk.”
The AI Act timeline
Krishan does not read the Digital Omnibus delay as a reprieve, because “the dates have moved, but the underlying governance problem has not.”
“The risk-based structure of the AI Act remains intact. Banned practices remain banned, general-purpose AI requirements remain, and several transparency obligations still apply this year. GDPR has not moved at all,” he said. “The extension gives organisations more implementation time; it does not remove the obligation to implement.”
He also sees a strategic question for organisations weighing whether to accelerate deployments before later obligations become applicable. “That is not a technical or product-management choice,” he said. “It changes the organisation’s regulatory exposure and potentially its future remediation burden, which makes it a decision for the board and senior risk leadership.”
“The danger with a reprieve mentality is that organisations postpone the difficult work: establishing what AI they have, where it is being used, how it should be classified and who owns the associated risks,” Krishan said. “A deadline moving does not make that work disappear. It just changes how much runway you have to do it properly.”
For fintech leadership teams, he said, the first job is discovery: a complete inventory of the AI systems in use across the organisation, “including systems bought centrally, embedded into third-party products and adopted informally within individual teams.” Those systems then need to be classified, matched to the obligations and deadlines that apply to each, and given an accountable owner and governance that can change as the technology and regulation change.
“The firms treating the change as permission to wait will probably look no different for a while. The difference becomes visible when the deadline approaches,” he said. “One group will already understand its AI estate and have operating governance around it; the other will still be trying to work out what systems it has.”
An early warning for other sectors
Krishan describes financial services as an early warning because it combines “large document-heavy workloads, high volumes, sensitive data, strict accuracy requirements, regulatory scrutiny, auditability and significant consequences when a model gets something wrong.” None of those issues, he said, is unique to the sector.
He also argues that cost and governance are more closely linked than organisations tend to assume. “Some of the easiest ways to make an AI system cheaper, such as reducing context, removing validation steps or routing everything to the cheapest available model, are the same changes that can make it less reliable,” he said.
“The wider lesson for fintechs is to treat AI the way they would treat any other piece of operational infrastructure,” Krishan said. “Know who owns each workload, what it costs, what it produces and which controls cannot be traded away when you optimise it.”
The post Altimetrik’s Vikas Krishan on AI Token Costs and the AI Act appeared first on The Fintech Times.