GenAI productivity metrics banking: IBM partner proposes a throughput test
An IBM consulting partner says banks using GenAI should add business-output measures to speed and stability metrics, though the model lacks validation.
By Rafael Ortiz · Fintech Correspondent
· 3 min read
GenAI productivity metrics in banking should move beyond story points and code output, according to Srini Bala, an IBM consulting partner who has proposed a throughput measure for financial-services software teams. In an external, unedited opinion published by Finextra, Bala argues that faster AI-assisted coding can inflate conventional measures without showing whether banks are releasing more usable services or features.
The proposal comes as financial institutions seek evidence that spending on generative AI improves delivery, while retaining the security, compliance and review controls required for regulated systems. Bala’s framework is an opinion rather than an established industry standard, and the material provided includes no independent bank study or case evidence validating it as a measure of GenAI return on investment.
How should banks measure GenAI software productivity?
Bala’s argument is that story points, which are intended to estimate a task’s complexity, risk and manual effort, become less meaningful when an AI assistant can generate routine code quickly. A team may complete the same number of estimated points with less human effort, producing an apparent jump in productivity without a comparable increase in business outcomes.
He also says that common delivery measures do not resolve that question on their own. Deployment frequency and lead time to production show the pace of software delivery. Under Bala’s proposed model, change failure rate would track stability. He would add a third measure, which he calls “Payload Size per Deployment”, to count the volume of functional work delivered in a consistent release window.
Rather than count lines of code or commits, Bala suggests defining a standard set of business artifacts and tracking how many reach production. Examples include accepted features or epics, deployed microservices, and validated, deployed end-to-end user journeys. The categories would need to be consistently defined within each organisation for comparisons across releases to be meaningful.
Why could faster code generation still slow delivery?
Bala argues that AI can move the constraint from writing code to checking and releasing it. Architectural reviews, security and vulnerability scans, compliance controls and human pull-request approvals may not accelerate at the same rate as code generation. He describes that accumulation of downstream checks as a “verification tax”.
In that situation, larger development output could increase lead time and reduce release frequency if review capacity does not keep pace, Bala writes. His suggested response is to examine output alongside speed and stability, and consider investments in automated testing, pull-request summaries and AI-assisted security review.
His illustration is conditional: if a team raised the number of major features from five to nine in comparable biweekly releases without worsening change failure rate, it could point to a measurable improvement in throughput. It is not evidence of a result achieved by a particular bank.
For a cautious comparison, teams would need a baseline before introducing GenAI and pre-agreed rules for each business unit and release window. Change failure rate also requires a fixed definition: vendor guidance from Swarmia defines it as the proportion of production changes that cause degraded service and subsequently need remediation, while cautioning that inconsistent failure definitions can distort results.
The unresolved issue is whether different features, microservices and user journeys can be weighted or standardised sufficiently to compare teams. Bala’s proposal identifies a way to test whether coding gains are reaching production, but it does not yet establish a universal measure of business value.
This story draws on original reporting from Finextra Research.