Most bank compliance programs still review somewhere between 1% and 3% of customer-facing calls. The sample exists for a sensible reason. When listening was manual, full coverage was impossible. A small sample was the only sample available, so QA was built around making the most of it: weighted selection, rotating reviewers, a rubric that produced consistent scoring across analysts.
The program does what it was designed to do. It tells you, on the calls you reviewed, whether the agent handled the interaction within policy. It does not tell you anything about the other 97%.
For a long time, this was acceptable. Examiners understood the constraint. The expectation was that the sample was representative, the reviewers were calibrated, and remediation closed the loop on what you found. That math has not changed. How examiners are pulling has.
When OCC, FDIC, state DFI, or CFPB examiners ask for call evidence during a consumer compliance or BSA exam, they pull on their own criteria. A complaint cohort. A product tied to a fair lending or UDAAP focus. A date range matched to a marketing campaign. A list of accounts that triggered a Reg E dispute or a Reg Z disclosure event. The request references calls you have never reviewed.
Risk-based examination playbooks now drive specific pulls into specific call populations, and the selection criteria are almost never aligned with how your QA team built its sample. Your program covers a small, deliberately representative slice. The examiner reaches into the 97% your QA never touched and asks what you can tell them about it.
The moment a compliance officer answers an examiner with "we will need to pull and review that call before we can respond" is the moment the exam shifts tone. The honesty of the answer is not the issue. What the examiner has just learned is that the institution does not have a current view of the conversation in question, and that the only evidence chain available is the one being built reactively, after the request.
This is usually written up as a process or oversight observation. In sharper exams, it lands as a matter requiring attention. The underlying point is the same: the monitoring program did not produce evidence on a call that mattered to the examination.
Reactive review carries its own problem. The agent may no longer be on the team. The customer record has moved through several states. The context that made the call accurate to review the week it happened has decayed. By the time the review reaches the examiner, it is a reconstruction, not a record.
The sampling approach was a workaround for a capacity limit. It was never a regulatory standard. Examiners accepted it because there was no alternative. The expectation was always that the bank would monitor for compliance risk across its consumer-facing operations. Sampling was how that monitoring got done given the available tooling.
Once the tooling changes, the expectation changes with it. If a bank can review every call for the disclosures, fair-lending language, complaint handling, and UDAAP red flags relevant to its products, then "we sampled 2%" stops looking like a methodology and starts looking like a choice not to look. None of this requires a new rule. Examiner expectations rise to match what is operationally possible.
Many of the recent AI-for-QA pitches are scorecard automation. A model reads a transcript, applies a rubric, and produces a number. Category coverage is wider than a human reviewer could sustain, and the consistency is higher, but the output is still a scored judgment about the call. Useful for coaching. Not what an exam needs.
An examiner asking about a specific call is not asking for a score. They are asking what happened. Was the Reg E error resolution timeframe communicated correctly? Did the agent acknowledge a complaint and route it through the complaint process, or treat it as a service issue? Was the customer told something during a collection call that crosses into a UDAAP concern? Was a required disclosure delivered before the action it had to precede?
These are evidence questions. They have answers in the transcript, anchored to specific moments. The exam wants the moment, the quote, the timestamp, and the structure that lets the institution show it has been monitoring this category of risk across all relevant calls, not just the ones the sample caught.
Scoring is judgment. Evidence is truth. The bank that walks into an exam with evidence walks into a different conversation than the bank that walks in with scores.
Continuous, evidence-based analysis on 100% of customer-facing calls produces something specific. For each call, you have a structured record of which compliance-relevant conditions were present, what was said, when it was said, and how it relates to the policy that governs it. Disclosure delivery is captured as a moment in the call, not a checkbox the agent self-attested to. Complaint signals are flagged the day the call happens, with the customer language that triggered the flag. Reg E dispute mentions are tied to the timeline of the institution's response.
When an examiner asks about a specific call, the institution can answer in the room. When the examiner asks about a population, the institution can describe what the monitoring program saw across all of them and produce the evidence behind that description. The conversation stays inside the bank's evidence chain instead of moving onto the examiner's.
The benefit shows up between exams too. A disclosure that drifts after a script change is caught the week it begins, not the quarter the examiner notices.
Before the next OCC, FDIC, or state DFI visit, the useful exercise is not "are our QA scores good?" It is "if the examiner pulls any call from the last twelve months in a specific product or complaint cohort, can we describe what happened on that call with evidence today, or do we have to go listen to it after the request?"
If the answer is the second one, the exposure is structural. The sample was never going to cover the call the examiner picked. That is the difference between a program that monitors a slice and one that monitors the population.
Banks that have moved to evidence on every call describe the change in modest terms. Exams get quieter. Requests get answered in the room. Findings related to monitoring coverage stop appearing. The work between exams shifts from preparing defenses to running a steady, current view of the conversations the bank is actually having.
If you want to see what your last twelve months of calls would look like under continuous analysis before the next exam window, Compass runs a scoped review on a sample of your own recordings and surfaces the evidence the QA sample did not cover. The same view an examiner would build, run on your data first.
Examiners do not pull calls from the bank's QA sample. They select on their own criteria, complaint cohorts, specific products, fair lending or UDAAP focus areas, which means most call requests reference calls QA never reviewed. Continuous evidence-based analysis on 100% of customer-facing calls produces a structured record before the exam request, instead of a reactive reconstruction after it.