Operating-model quality determines AI quality
How to Prove the AI Is Right, Not Just Agreeing With You
Adyen, a Dutch Payments Platform, holds a slice of live traffic where its models never act, runs every new model in ghost mode before it influences anything, and promotes it only when it beats the incumbent with statistical confidence. Most AI pilots are decided on an agreement rate, which stops describing the model the moment a reviewer can see its answer.
By Christopher Hughes
18 August 2026
The choice of whether to give an AI system live volume is generally based on how frequently it has agreed with the people currently doing the job, but that figure does not include the cases where the model and the reviewer were both wrong. It also stops describing the model as soon as a reviewer can see its answer before deciding.
The Rollout Decision Rests on a Number Nobody Interrogates
In a go-live conversation the board pack in front of the steering group might report that the model reached the same answer as the current process 97% of the time. Every executive in the room knows agreement is not accuracy, but the decision gets made on it anyway, because nobody produced a better number.
The pack cannot show the difference, because a reviewer tends to start agreeing with an answer already in front of them, and the rate then measures the interface rather than the model.
The FCA's second AI Live Testing cohort started in April 2026, with Barclays, Lloyds Banking Group, and UBS among the eight firms. [1] The regulator describes itself as exploring output-driven validation rather than publishing a measurement checklist, and neither cohort has published findings, [2] so the evidence standard for the next two years is being set inside firms, by whoever designs the next comparison. Agreement wins that design because it is the cheapest number to compute.
Blind the Comparison, Then Count What You Agreed On
A number that is not an agreement rate needs both systems deciding the same cases and both scored against something outside either of them, which is only possible while the candidate runs without deciding anything. A bounded parallel run buys that window, and it carries four decisions. The four are how to blind the comparison, what both sides are scored against, how much of the log gets adjudicated and by whom, and what promotes the system at the end.
The first decision is how to blind the comparison, and blinding means arranging the run so that one side's answer cannot influence the other. It decides whether the run produces evidence at all, and it takes two forms.
Reviewer blinding has the human decide first, with the model's answer going to a log the reviewer never sees, which keeps the human judgement clean. Population blinding holds out a fixed slice of live traffic the model never acts on, which keeps a comparison population alive after go-live. The two answer different questions, and most teams run neither.
Each form has a price the sponsor should hear before the run rather than halfway through it. Reviewer blinding buys several weeks of dual-running with no efficiency benefit. Population blinding leaves a slice of volume permanently on the old process, which somebody will eventually ask the owner to justify.
The second decision is what both sides are scored against, and agreement with the incumbent decision is the cheap answer that is structurally incapable of showing the model beating the incumbent. The real outcome is the only target that can show it, and it usually arrives after the promotion decision is due, which leaves the operative choice as naming the early proxy and writing down what it costs in confidence. An early slice of the true label, such as a 30-day dispute rate, is objective but truncated towards fast-moving cases, while a senior adjudicator's blind re-review is defensible, slower, and inherits the institution's own bias.
The size of that cost depends on how late the label lands. Visa and Mastercard allow cardholders up to around 120 days to file a dispute, [3] credit teams treat 12 months as the regulatory baseline and 24 as the point a vintage curve has settled, [4] and in anti-money-laundering work the label never really arrives, because a suspicious activity report is an internal judgement and the UK Financial Intelligence Unit publishes sanitised, thematic feedback rather than case-level outcomes. [5]
The incumbent's decision is the agreement target wearing a proxy label, and a team choosing it owes the decision-maker one sentence saying so.
The third decision is how much of the log gets adjudicated, and by whom. At low volume every disagreement gets a verdict. Above that, a stratified sample weighted to the highest-consequence segments carries a fixed weekly quota, a named senior adjudicator, and three buckets on each case (model wrong, incumbent wrong, both defensible). The decision almost nobody makes is what share of the agreed cases enters the same quota, because cases where the model and the human reached the same wrong answer generate no disagreement and appear in no log.
The fourth decision pre-commits the exit rule, the date, the signer, and what survives promotion. Promotion moves either by scope, one segment or value band at a time, or by autonomy, keeping full scope while the system goes from advising to acting on a defined class of case. The sponsor writes the bar before the first case is scored, so the run cannot be re-argued once the numbers are in.
Adyen Published the Method, Not Just the Result
Adyen, a payments platform, is useful because an outsider can read how the firm knows its model won, in a design Andreu Mora, then its SVP and Global Head of Engineering Data, set out in January 2025. [6][7] A unified control group is "a split of traffic that is consistent across all elements of a transaction and where models do not act". New models run first in ghost mode, logging telemetry without influencing outcomes, then as a challenger on a traffic split, and a challenger becomes the principal only when it beats the principal with statistical confidence. [6]
Adyen's baseline is deliberately hard, comparing against what a proficient payment provider would already offer, which it says produces "a lower, but more honest uplift calculation". [6] Mora also names the funnel problem on his own product, writing that once a decision is applied to a transaction the firm does not have access to the outcome had it not acted. [6]
Adyen reported customer conversion up 0.9 percentage points on average by the end of the first half of 2026, attributed jointly to Adyen Uplift and Dynamic Identification, in its own results release. [8] Both that figure and the method account are Adyen writing about Adyen, and neither is independently audited.
The unified control group is population blinding rather than reviewer blinding, so it proves the untouched-comparison half of the parallel run and says nothing about what a human does when shown a model's answer, which is measured where blinded reading has been studied directly, in clinical double reading. Ghost mode and challenger testing are established practice in payments, so Adyen's contribution is the disclosure rather than the design, at a level of detail an outsider can check. Stripe, on the same domain, reported an 80% decrease in successful card-testing attacks in January 2025 and published no comparison method at all. [9]
Where a Parallel Run Quietly Stops Being Evidence
When reviewers can see the model's answer, the agreement rate becomes an artefact of the interface. The signal is agreement climbing week over week while adjudicated accuracy stays flat, and the climb is fastest in the teams with the most model visibility. A second signal is handling time on agreed cases falling faster than on disagreed ones, which is a reviewer accepting rather than deciding.
The Rubber-Stamp Problem is usually a production failure, where an approval step degrades into a keystroke and the control is fake. [10] The same behaviour inside an evaluation window corrupts the evidence instead, and the manufactured number then authorises the rollout the control was meant to catch.
Blinding is what puts a number on that gap, and in a 2021 breast-screening cohort of 1,119,191 women, second readers agreed with a recall recommendation 74.7% of the time unblinded against 69.8% blinded, with blinded reading producing higher positive predictive value and specificity. [11] In a study of 28 pathologists, roughly 7% of AI-assisted decisions accepted an incorrect suggestion that contradicted the reviewer's own prior correct call. [12]
Nobody counts the cases where the model and the reviewer were both wrong. The signal is a disagreement log with entries and an agreed-case sample with none, alongside a measured error rate suspiciously close to zero in exactly the segments where the incumbent process is weakest.
Almost nobody budgets adjudicator time for matching answers, which is precisely why almost nobody has the number. No named financial institution has published a joint error rate, and in the clinical double-reading literature the agreed-and-wrong case appears to go unmeasured rather than solved. Some firms may hold the number internally and never publish it, which changes nothing for a team that still has to produce its own from its own run.
The stream under test is not the real stream, because every case in the log entered through the incumbent's own funnel, so the model has only been tested on what the existing filter already surfaced. The signal is precision that looks excellent alongside a recall figure nobody in the room can produce. In credit-scoring experiments, retraining on approved-only outcomes holds accuracy steady while recall collapses, and standard metrics reward the strategies that amplify the bias, so the fix is deliberately approving 2 to 5% of rejected applications to observe what happens. [13]
Key Takeaways
- Blind the comparison for the length of the evidence window, and keep a slice of live volume the model never touches after it. Once a reviewer can see the model's answer before deciding, the agreement rate measures the interface rather than the model, and the tell is agreement climbing while adjudicated accuracy stays flat.
- Budget adjudicator time for the cases where the model and the human agreed. Joint errors generate no disagreement and appear in no log, so a run without an agreed-case sample cannot produce a joint error rate, and no named financial institution appears to have published one.
- The outcome label almost always arrives after the decision is due, so name the early proxy and write down what it costs in confidence before the run starts. An undeclared proxy quietly returns the measurement to agreement with the incumbent.
Sources
[1] Financial Conduct Authority. "FCA announces second cohort for AI Live Testing." FCA, 2026. https://www.fca.org.uk/news/press-releases/fca-announces-second-cohort-ai-live-testing
[2] Financial Conduct Authority. "FS25/5: AI Live Testing." FCA, 2025. https://www.fca.org.uk/publications/feedback-statements/fs25-5-ai-live-testing
[3] Chargeflow. "Chargeback Time Limit by Card Network." Accessed 18 August 2026. https://www.chargeflow.io/blog/chargeback-time-limit (Industry explainer; card-scheme rulebooks are member-only, and windows vary by scheme, reason code and jurisdiction.)
[4] GARP. "A Review on the Probability of Default for IFRS 9." January 2022. https://www.garp.org/hubfs/Whitepapers/a2r5d000003s7K0AAI_RiskIntel.WP.IFRS9.PD.Jan22.pdf (The 12-month horizon is regulatory convention; the 24-month point is industry convention from vintage-curve practice, not a single sourced figure.)
[5] National Crime Agency. "UK Financial Intelligence Unit." https://www.nationalcrimeagency.gov.uk/what-we-do/crime-threats/money-laundering-and-illicit-finance/ukfiu
[6] Mora, Andreu. "The AI behind Uplift." Adyen Tech, 17 January 2025. https://www.adyen.com/knowledge-hub/the-ai-behind-uplift (Adyen's own account of its own product design.)
[7] Retail Technology Innovation Hub. "Andreu Mora announces departure from fintech Adyen to build an AI company from scratch." 24 February 2026. https://retailtechinnovationhub.com/home/2026/2/23/andreu-mora-announces-departure-from-fintech-adyen-to-build-an-ai-company-from-scratch
[8] Adyen. "Adyen publishes H1 2026 financial results." Adyen press release, 13 August 2026. https://www.adyen.com/press-and-media/adyen-publishes-h1-2026-financial-results-3wjne (Self-reported; the company's own attribution of a conversion metric to its own products.)
[9] Meltzer, Jacob and Chadalapaka, Viswanath. "How Stripe Radar responded to a new wave of card testing." Stripe, 15 January 2025. https://stripe.com/blog/how-stripe-radar-responded-to-a-new-wave-of-card-testing (Self-reported; no comparison method published.)
[10] Hughes, Christopher. "The Rubber-Stamp Problem: Why Human-in-the-Loop Is a False Promise, and What Should Replace It." cgh.dev, 2026. https://cgh.dev/thinking/rubber-stamp-problem/
[11] Cooper, J.A. et al. "Optimising breast cancer screening reading: blinding the second reader to the first reader's decisions." European Radiology 32(1):602-612, 12 June 2021. DOI 10.1007/s00330-021-07965-z
[12] Rosbach, E. et al. "Stuck on Suggestions: Automation Bias, the Anchoring Effect, and the Factors That Shape Them in Computational Pathology." arXiv, 13 March 2026. https://arxiv.org/html/2603.11821v2 (Study of 28 pathologists.)
[13] Scarone, B. and Baeza-Yates, R. "The Illusion of Improvement: Reject Inference Strategies in Credit Scoring." arXiv, 16 June 2026 (rev. 12 August 2026). https://arxiv.org/html/2606.18479
Take it further