Coinbase reported on Oct. 7 that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service, despite an unchanged decision policy. The findings challenge the common assumption that upgrading a model automatically improves an existing payment screener.
The company’s evaluation replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions. The cohort covered nine weeks before its risk agent rolled out, retaining all matured fraud cases while sampling legitimate traffic.
Each candidate reviewed recent transaction behavior under fixed guidance and the same policy for turning risk classifications into decisions. This isolated the decision model’s behavior within that setup, rather than comparing redesigned screening systems.
Results from a fixed historical replay
Coinbase compared Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Every newer version had lower recall, a lower combined precision-and-recall score called F1, and lower dollar-weighted recall. Recall measures the share of fraud cases a model catches; dollar-weighted recall measures how much of the total fraud value it catches.
Sonnet’s recall fell 22.2 percentage points and its dollar-weighted recall dropped 22.9 points. Opus’s recall declined 0.8 points. Both newer models also had lower precision, meaning a smaller share of transactions they classified as fraud were actually fraudulent.
GPT showed why one improving score can be misleading. Its precision rose 11.5 percentage points, but recall fell 20.7 points and dollar-weighted recall fell 21.8 points. Its fraud flags were more accurate, while more fraud cases and value escaped detection in the replay.
The replay does not establish customer losses from deploying those versions. Coinbase also said it could identify the regressions without establishing their cause.
Coinbase’s earlier online experiment compared adding selective LLM review with the existing models and rules alone. That agent-enabled flow recorded 30% fewer fraudulent transactions and 22% less fraud value; it did not compare newer model versions.
In their limitations, the SR-Fraud researchers say the proprietary dataset cannot be released, restricting independent replication and generalization. Their related payment-fraud study first appeared Sept. 23 and was revised Sept. 30, before the October blogs.
A separate case for a custom model
In its Oct. 8 disclosure, Coinbase reported that a post-trained Qwen3.5-9B model exceeded Opus 4.5 across four fraud-detection metrics. F1 improved 9.6 percentage points and dollar-weighted recall rose 35.4 points. The company specialized it using historical fraud outcomes and deterministic rewards balancing fraudulent and legitimate examples.
Separately, production measurements put median end-to-end LLM-request latency at 0.683 seconds versus 1.515 seconds for Opus 4.5, a 55% relative reduction. Faster inference and stronger benchmark detection came from different evaluations.
For payment providers, the upgrade question is whether a candidate improves fraud coverage under their actual decision setup. Coinbase recommends testing that configuration first, then evaluating changed prompts or thresholds separately, with latency, reliability and cost alongside detection quality.
Also Read
- BTCPay Docker Users Must Enable Tor Manually to Retain Onion Access
- CFTC Issues Dual Rules Defining Prediction Market Contracts as Swaps and Excluding Traditional Casino Gambling
- block at beginning, then rewritten article HTML body text. Must remove promotional lines, “SEE ALSO” links, footer junk. Also must keep all media elements (images, iframe, etc.) unchanged and in original positions. Need to rewrite headline to unique, fresh, professional. Then rewrite the news story body text into natural, high-quality professional English. We need to keep HTML tags, but also ensure media tags remain exactly. Let’s identify media elements: There’s a with an and a . The img class includes many srcset and width/height attributes. Must keep exactly. There’s also a <img alt=”” loading=”lazy”… class=”shrink-0 rounded-full max-sm:h-[22px] max-sm:w-[22px]” style=”color:transparent” src=”https://coin-images.coingecko.com/coins/images/1/large/bitcoin.png?1696501400″/> (a small bitcoin icon). Must keep. There are some lines with etc, but they are part of UI elements like price, market data. Those seem like promotional or UI, not part of article body? The instructions: Remove promotional lines, “SEE ALSO” links, and footer junk. The price/value and market projections might be promotional UI. So we likely need to remove those elements that are not part of news story. Also the Daily Debrief Newsletter and its paragraph likely are promotional newsletter sign-up, thus remove. Also the with Bitcoin label etc maybe part of UI; likely promotional. So we should remove those. We need to keep the article body: The news story comprises paragraphs with tags describing the committee actions. Also list items () describing amendments. We need to ensure we preserve the with three list items (the amendments). Also the figure with image remains. Also need to keep the paragraph after the figure (the final paragraph) that continues article: “The Assembly adopted another crypto-inclusive tax…” that’s part of article? It describes another amendment, still news. That seems relevant. Also the final paragraph about floor debate dates. That also part of article. We must rewrite headline: Original: “French Committee Backs Stablecoin Swap Tax and Crypto Exit Tax, Then Rejects the Budget”. Need unique, fresh, professional. Maybe: “French Finance Committee Approves Taxes on Crypto Swaps and High‑Value Holdings, Rejects Budget”. But need professional and fresh. Let’s craft: “French Finance Committee Endorses Taxes on Crypto Swaps and Wealthy Holdings, Rejects Budget”. That is okay. But we need to output [TITLE]… then the headline? Actually [TITLE] block should contain the rewritten headline. The instruction: Output ONLY the requested [TITLE] block at the absolute beginning, followed immediately by the rewritten content. So [TITLE] block contains the headline (rewritten). Then immediately after that, the cleaned, rewritten article HTML body text. Thus we need: [TITLE]French Finance Committee Endorses Taxes on Crypto Swaps and Wealthy Holdings, Rejects Budget
- Reserve Bank of India Introduces Special Dollar Window for Public Sector OMCs to Curb Rupee Depreciation


