If Your Next Exam Pulled 100 Random Calls, Would Your QA Sample Have Reviewed Any of Them?

The thought experiment that should keep BSA officers up at night

Picture the next exam. The examiner sits down, opens the call recording system, and pulls 100 random conversations from the last quarter. Disclosures, complaint handling, fee disputes, account closures, fraud claims. A clean random sample, no filtering.

Now ask the harder question. Of those 100 calls, how many would your QA team have already reviewed?

Most compliance leaders, asked this honestly, land somewhere between "a couple" and "I'd have to check." That answer is the entire problem. It means the evidence chain your program rests on, the sampled reviews, the scorecards, the calibration notes, has almost no overlap with the conversations the examiner is actually looking at.

The math, worked out

Take a community bank running 10,000 customer calls a month across deposit operations, loan servicing, and the fraud queue. A typical QA program reviews 2 percent of those calls. That is 200 calls reviewed against 9,800 that no one inside the bank ever listened to again.

When the examiner pulls 100 random calls from that same month, the expected overlap between the QA sample and the exam sample is straightforward. Each call the examiner picks has a 2 percent chance of being one your QA already reviewed. Across 100 picks, the expected overlap is 2 calls.

Two. Out of one hundred.

The probability that none of the examiner's calls overlap with your QA sample at all is roughly 13 percent. The probability that fewer than 5 overlap is over 95 percent. Push the QA sample rate down to 1 percent, which is closer to what many programs actually run when you net out monitoring backlogs and staff turnover, and the expected overlap drops to 1 call. The probability of zero overlap rises to about 37 percent.

This is not an attack on your QA team. They are doing exactly what they were asked to do. The math is just what it is. Random sampling at 2 percent does not produce evidence about the other 98 percent, and the examiner is, by design, looking at the other 98 percent.

What this means in the exam room

The examiner asks about a specific call. A wire that should have triggered a CTR. A complaint that escalated to a fee reversal. A Reg E claim that closed without a provisional credit. The conversation is in front of them. The reasoning is not.

You have a few options in that moment. You can pull the recording and listen to it live, which signals that you did not review it before they did. You can describe the program in general terms, training, policies, scorecards, hoping the examiner accepts process evidence in place of conversation evidence. You can promise a written response after the exam window.

None of these is the answer the examiner wants. The examiner wants to know what happened in that call, whether the required steps were followed, and whether your program already knew. When the answer is "we have not reviewed that specific call," what they hear is "we cannot tell you whether our team is compliant on the conversations we did not happen to pick."

That gap, between what was reviewed and what was reviewable, is where MRAs get written.

Why this got worse, not better

Five years ago, the math was uncomfortable but at least the volume was stable. A QA team reviewing 200 calls a month had a known relationship to the pool of human-handled conversations.

That relationship has broken. Volume keeps climbing as banks push more first-touch interactions into voice channels, chat, and now AI-handled flows that hand off to humans only when something goes wrong. The conversations that reach your team are denser, more complex, and disproportionately the ones with compliance exposure, because the easy stuff was already handled upstream. At the same time, QA headcount has not grown. Most programs are still reviewing the same 100 to 300 calls a month they were reviewing in 2021.

Meanwhile, AI agents now handle a meaningful slice of customer-facing volume in deposit operations, payments support, and basic loan servicing. Those conversations carry the same regulatory weight as human calls. Reg E timing, UDAAP exposure, fair lending language, BSA red flags, none of those obligations care whether the speaker on your side was a person or a model. But the sampling rate did not change to account for the new surface area. In practice, AI-handled calls are reviewed even less frequently than human ones, because most QA workflows were built around agent scorecards and have no clear owner for the AI traffic.

The sample rate stayed flat. The risk surface grew. The exam expectations did not soften.

What evidence-based review actually changes

The way out is not to sample more aggressively. Pushing 2 percent to 5 percent does not close the gap, it just moves the math from "almost no overlap" to "still almost no overlap." A 5 percent sample still leaves a 60 percent chance that fewer than 4 of the examiner's 100 calls were touched by QA.

The shift that matters is moving from sampled judgment to recorded evidence on every conversation. Not scoring every call with the same checklist your QA team uses today, that is just scaling the old system. Evaluating every call against the specific obligations that apply to it. Did the Reg E timing window get communicated correctly. Was the required disclosure read in full, in the language the customer was speaking. Did the agent verify identity before discussing account specifics. Was a complaint actually treated as a complaint, or quietly routed back to a service request.

When every call has a structured record of what was said, what was required, and whether the two matched, the examiner's 100-call pull stops being a stress test. You already know what is in each of those calls, because every call had its review the first time it happened. The QA team's job shifts from "find the bad calls in a 2 percent sample" to "validate the patterns the system surfaced, and focus deep review on the calls that look genuinely ambiguous."

The compliance program stops resting on probability. The evidence chain stops having gaps the size of 98 percent of the call volume.

The question worth asking before the next exam

Run the thought experiment with your team this week. Pull 100 random calls from last month. Cross-reference against your QA log. Count the overlap.

Whatever the number is, that is the size of your current evidence base relative to what the examiner sees. If it is small, that is not a QA team failure. It is a sampling model failure. The model assumes that 2 percent of conversations can speak for the other 98, and the regulators are increasingly making clear that it cannot.

If you want to see what coverage looks like when every call has a record before the examiner asks, Chordia's Compass team runs a no-commitment compliance scan on a sample of your recent calls. You get back what we found, where the risk lives, and what would have been visible to your program if every conversation had been evaluated the first time. No deck, no demo theater. Just the evidence, on your own calls.

The Question

If a regulatory examiner pulled 100 random calls, would my QA sample have already reviewed any of them?

If you sample 2 percent of calls, the expected overlap with a random 100-call exam pull is only 2 calls, and there is about a 13 percent chance of zero overlap. Sampled QA does not produce evidence about the calls examiners actually review.

Discover how Compass gives you a full understanding of your customers.

Conversation Intelligence Terminology

Read more from Insights