KYC Hub co-founder says RL is a poor fit for AML alert queues
Farnoush Mirmoeini argues AML alerts lack the verified feedback that reinforcement learning systems need to improve decisions safely.
By Rafael Ortiz · Fintech Correspondent
· 3 min read
Reinforcement learning is being pitched as a way to cut anti-money laundering alert volumes and compliance costs, but KYC Hub co-founder Farnoush Mirmoeini says the technology lacks the decisive feedback needed for the job. In an external opinion published by Finextra, Mirmoeini argues that AML alert queues do not produce reliable ground truth on whether a cleared transaction was in fact clean.
The argument turns on how reinforcement learning improves. A system tests actions, receives a reward or penalty, and adjusts future behaviour toward choices that score better. Mirmoeini writes that this mechanism depends on a prompt and unambiguous signal that a decision was correct.
She contrasts AML with areas where artificial intelligence has advanced through verifiable rewards, including maths and code. In those fields, a proof can be checked or a program can be run, giving the model a low-cost and objective assessment. Mirmoeini also cites recent remarks by Anthropic chief executive Dario Amodei, who she says has described unverifiable tasks as a constraint on AI progress even while expressing strong confidence in the broader technology.
AML alert handling, in her view, breaks that feedback loop. Firms must decide whether to clear an alert, hold a transaction, ask for more information or file a suspicious activity report. Those decisions have commercial and regulatory consequences, which makes them appear suitable for reinforcement learning. The difficulty, Mirmoeini says, is that an analyst’s clearance shows only that the investigation satisfied the analyst at the time, rather than proving the funds were lawful.
That distinction matters because money laundering is adversarial. According to Mirmoeini, a launderer’s objective is to present plausible explanations and documentation that withstand review. If an illicit payment is cleared, the compliance record may still treat the case as a false positive, creating training data that marks a bad outcome as a good one.
Mirmoeini criticises research claims that reinforcement learning can sharply reduce false positives when those findings rely on synthetic datasets with transactions already labelled clean or suspicious. She argues that such labels are precisely what real AML operations usually lack. A benchmark can reward correct classifications, while a live alert queue often cannot confirm the underlying truth.
She also rejects three proposed ways to create a reward signal for AML decisions. Rubrics, she says, can assess whether an investigation was thorough, but cannot prove whether the money was illicit. Formal rules can define compliance procedures, but cannot reveal the hidden parties or purpose behind a payment. Enforcement actions can produce factual findings, but they are rare, slow and usually arrive through regulators or prosecutors long after the original decision.
The practical risk, according to Mirmoeini, is not only insufficient data but misleading data. A reinforcement learner trained on cleared alerts could become more confident in clearing cases that should have been escalated if the historical record wrongly coded missed laundering as successful decisions.
Mirmoeini says reinforcement learning may still be useful in financial crime settings where outcomes are observable. She points to card fraud, where chargebacks provide a delayed but real signal, and public blockchains, where flows can be traced and some criminal addresses later confirmed. Her central claim is narrower: AML alert queues do not routinely tell firms whether individual decisions were right, limiting the usefulness of reinforcement learning for that specific workflow.
This story draws on original reporting from Finextra Research.