Newer AI models missed more payment fraud in Coinbase’s benchmark
A fixed historical replay found weaker fraud coverage across three model upgrades, while GPT’s precision improved. The post Newer AI models missed more payment fraud in Coinbase’s benchmark appeared first on CryptoSlate.
Coinbase reported Oct. 7 that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service, despite an unchanged decision policy. The findings challenge the assumption that upgrading a model improves an existing payment screener.
The company’s evaluation replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions. The cohort covered nine weeks before its risk agent rolled out, retaining all matured fraud cases while sampling legitimate traffic.
Each candidate reviewed recent transaction behavior under fixed guidance and the same policy for turning risk classifications into decisions. This isolated the decision model’s behavior within that setup, rather than comparing redesigned screening systems.
Related ReadingCoinbase compared Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Every newer version had lower recall, a lower combined precision-and-recall score called F1, and lower dollar-weighted recall. Recall measures the share of fraud cases a model catches; dollar-weighted recall measures how much of the total fraud value it catches.
Sonnet’s recall fell 22.2 percentage points and its dollar-weighted recall dropped 22.9 points. Opus’s recall declined 0.8 points. Both newer models also had lower precision, meaning a smaller share of transactions they classified as fraud were actually fraudulent.
GPT showed why one improving score can be misleading. Its precision rose 11.5 percentage points, but recall fell 20.7 points and dollar-weighted recall fell 21.8 points. Its fraud flags were more accurate, while more fraud cases and value escaped detection in the replay.
The replay does not establish customer losses from deploying those versions. Coinbase also said it could identify the regressions without establishing their cause.