Why Coaching Needs Evidence

Your QA team and your supervisors are grading two different shows.

QA scores a small sample against a rubric. Supervisors coach from memory, escalations, and gut. The two never quite line up, and your reps know it. Compass gives both jobs the same evidence base, on every conversation, across every site and program.

The problem

You run customer operations. Maybe you sit over an in-house team of 80 reps and six supervisors. Maybe you run a QA function covering 34 supervisors across four sites. Maybe you are a VP at a BPO holding calibration across five client programs with five rubrics. The titles differ. The structural problem is the same. Coaching is the work your supervisors are evaluated on, and it is the work that slips first when the floor catches fire.

So you built or inherited a QA program to fill the gap. It pulls two to five calls per rep per month, scores them against a rubric, and feeds the scores into one-on-ones. The intent is honest. The execution rarely holds. A supervisor with 15 to 22 reps cannot listen to enough calls to coach with confidence, and a QA team auditing two to five percent of conversations cannot produce a picture anyone trusts. You have two parallel grading systems running on the same agents with different inputs. When they disagree, and they always disagree, the rep learns to game the one that affects pay and treat the other as a tax.

The math is unforgiving. A supervisor with 18 reps coaching four calls per rep per month is making decisions on one to two percent of each rep's work. Your new hires are taking 40 to 60 calls a day from week one and developing habits no one is watching. By the time a sampled call catches a problem, the habit has been there for weeks.

Then peak hits, a client launches a new program, or you open a second site. Coaching calendars collapse. Newer reps take harder calls because everyone is on the phone. By the time the monthly review shows CSAT or conversion drift, you are not coaching. You are explaining.

What sampling misses

A two to five percent sample tells you almost nothing about your bottom-quartile reps and even less about your top reps. The signal you need, the difference between a rep who recovers a frustrated customer and one who escalates the same situation, lives in the 95 percent of conversations no one is listening to. Random sampling will not find it. Targeted sampling finds whatever the sampler was already looking for.

Sampling misses the strong-middle rep. Top performers get watched because you brag about them. Bottom performers get watched because they are on a plan. The eight to twelve reps in the middle, where every point of variance is worth more than coaching the tails, are invisible. They are the reason your team average moves, and you usually find them by accident when one quits.

Sampling misses drift. Reps regress one habit at a time. The discovery question that used to land gets skipped. The empathy phrase gets shorter. The verification step gets dropped when the queue is deep. A tenured rep has slowly stopped doing one of the behaviors that correlated with their early CSAT, and no one will know until the trendline turns. It also misses the moments that should have become coaching and never did. The save attempt that almost worked. The objection one rep handles in 40 seconds and four reps fumble for three minutes. All of it shows up only when you can see every conversation, grouped by rep, situation, and outcome.

What 100% understanding surfaces

  • Variance with a reason attached. Compass evaluates every conversation across more than 130 behavioral signals (the specific things reps do that have measurable impact: opens, verification, objection responses, recap, empathy markers, hold etiquette). You stop seeing "Rep A is a 92, Rep B is an 84" and start seeing "Rep B drops the value reinforcement step on most retention attempts."
  • Ramp curves week by week, per cohort. Every new hire has a tracked curve against team average on the behaviors that move outcomes. You see at week three which reps are tracking and on which signals, and intervene then instead of at week ten. That matters more when ramp has a deadline: a peak season, a client go-live, a contract SLA.
  • The strong-middle rep, finally visible. Compass ranks reps by where one coaching moment would move the most, not by where the scorecard is loudest. You spend coaching time where it changes the team average.
  • Drift detection across tenure. The trendline on a tenured rep's key signals shows the inflection point, anchored by the calls where the change appears. You catch the regression while it is still a coaching conversation, not a performance plan.
  • Coachable moments instead of coachable calls. Listening to a full call to find a 90-second teachable moment is the wrong unit of work. Compass surfaces the moment with the customer's reaction. The session is 12 minutes instead of 40.
  • Outcome Lift on the behaviors you debate. Compass attaches Outcome Lift, the difficulty-adjusted measure of how much a behavior shifted the outcome, to each signal on your data. You see what the empathy phrase is worth on the call types your team actually handles.
  • Supervisor and site calibration drift. Two supervisors looking at the same rep reach different conclusions. Two sites coach the same script to different standards. Compass makes the inter-supervisor and inter-site variance visible at the signal level, so calibration sessions become working meetings.
  • Repeat-contact root cause. When a customer calls back, Compass connects the second call to the first and shows which resolution patterns produce repeats. You stop coaching reps for problems your process created.

Why Coaching Needs Evidence

Coaching without evidence is opinion delivered with authority. It can be right. It is often right. But when a tenured rep disagrees with a supervisor's read, the conversation goes one of two ways. Either the rep accepts feedback they do not believe and changes nothing, or the supervisor escalates and damages the relationship. Either way, the coaching did not work.

Coaching with evidence is different. The supervisor and the rep listen to the same moment together, look at the signal Compass flagged, and compare it to the rep's calls last week where the signal landed. The rep is not being told they are doing something wrong. They are being shown what happened. That produces a behavior change in less time than the argument.

The structural case is simpler. Every operation we talk to has at least three definitions of "good" running at once: the QA rubric, the supervisor's mental model, and whatever the rep took from training. The rep lives in the gap. The fix is not more calibration meetings. The fix is one evidence base that QA, supervisors, training, and operations leadership all work from. When a supervisor and a QA analyst disagree, they are no longer disagreeing about what happened on the call. They are disagreeing about what to do next. That is the conversation you want them having.

The capacity case matters too. Most supervisors carry 15 to 22 reps and 40 other responsibilities, and most of the hour per week per rep gets spent finding the call, not coaching. Compass inverts that. The call is pulled, the signal is identified, the framing is drafted. The supervisor walks in with a one-pager and walks out with one committed behavior change. That is a better coach-to-rep ratio than the one you have today, without widening span or adding QA headcount.

How Compass works

Compass takes every conversation your team has, on every channel, and turns it into structured understanding. The pipeline is four steps. Conditions describe what was happening in the call: customer intent, complexity, sentiment, product mix, channel. Signals describe what the rep did, scored at the moment they occurred. Outcome Lift attaches difficulty-adjusted impact to each signal against the outcomes you care about: conversion, rebook, FCR, CSAT, repeat contacts, retention. Guidance turns the picture into the specific moments a supervisor coaches to that week, ranked by impact.

The primary pillar is Conversation Coaching, supported by Conversation Quality, Conversation Insights, and, where scripted or disclosure-bound work exists, Conversation Compliance. The intent is not to add a fifth definition of good, but to replace the disagreement with one shared evidence base.

  • Conversation Coaching. Every supervisor gets a weekly view of their reps with prioritized coaching moments. Each moment links to the transcript segment, the signal it surfaced, the Outcome Lift it carries, and a suggested framing. One-on-ones get prepared in 15 minutes instead of skipped.
  • Conversation Quality. The Conditions, Signals, Outcome Lift, Guidance loop replaces the traditional scorecard fight. Keep a simplified rubric for reporting if you need to. The underlying evaluation runs on evidence, and calibration drift between supervisors, sites, or client programs becomes visible and fixable.
  • Conversation Insights. Floor-wide and program-wide patterns. Where ramp is slowing for the new cohort. Which signals have drifted in the last 30 days. Which call types are getting harder. Which sites coach differently from each other.
  • Conversation Compliance. For operations that carry script, disclosure, or licensing requirements, adherence is tracked at the step level across every conversation. The output is a clean audit trail without turning coaching into a compliance review.

Common questions

Q: We already have a QA team. What changes for them? A: The job changes, the headcount usually does not. QA stops pulling and scoring samples and becomes calibration owners: auditing Compass output, tuning signal definitions for your call types, owning the relationship between Outcome Lift and your KPIs, and partnering with supervisors on the high-impact moments. Most operations leaders describe this as the most senior work their QA team has ever done.

Q: Will it replace our scorecard? A: It can. Most operations keep a simplified scorecard for reporting with the underlying evaluation done by Compass. The scorecard stops being the source of truth and becomes a view into it. Reps see one score, derived consistently, that matches what their supervisor coaches them on.

Q: We run multiple programs or sites with different rubrics. Does that break the model? A: No. Signal libraries and Outcome Lift weights can be tuned per program, per client, or per site, so a BPO running five client programs or an in-house operation running four locations does not have to pretend everything is the same. Reporting rolls up to the level you need: rep, supervisor, team, site, program, or enterprise.

Q: What about our recording platform and CRM? A: Compass ingests from the common recording and telephony platforms and pulls metadata from your CRM. You do not change your phone system, case tool, or scheduling stack. If you are on something less common, we will tell you in the first call whether the integration is straightforward.

Q: Our recordings are governed by client MSAs or customer agreements. How do you handle that? A: We sign standard paperwork: NDA, MSA, DPA, and a BAA where PHI is in scope. No customer data is used to train models that serve other customers. SOC 2 is in progress, and vendor security documentation, including our subprocessor list, is available on request during your review. We work through your security process with you, and we do not need recordings before that paperwork is signed.

Q: Will my reps feel surveilled? A: They are already recorded. The question is whether the recording produces useful coaching or sits in a vault. Compass is built around behaviors and outcomes, not gotcha clipping. Most operations roll it out by showing reps their own dashboards first, so the tooling reads as development.

Q: How does this connect to performance management and HR? A: Coaching evidence, signal trends, and outcome impact export into the artifacts your HR process needs: PIP documentation, promotion packets, performance review prep, calibration records. We do not write into your HRIS, but we produce the source documents your supervisors and HR partners currently assemble by hand.

Q: How is this different from Gong, NICE, Verint, Observe.AI, CallMiner, or the analytics in our recording platform? A: Recording-platform analytics summarize calls and surface keywords, which is not coaching. Conversation intelligence built for sales pipelines was not built to replace QA. Quality management built around scoring was not built to drive coaching. Compass was designed end to end around one loop, Conditions, Signals, Outcome Lift, Guidance, and the same evidence base feeds QA, coaching, and the weekly operations review.

Q: How long until supervisors get value? A: Most operations see usable per-rep coaching queues within 30 days of ingest. Deeper Outcome Lift modeling takes longer because it needs enough conversation and outcome data to calibrate against your KPIs. We tell you that on the first call, not after you sign.

Your QA team and your supervisors are grading two different shows.

Book a 30-minute working session, not a slide deck. We can walk the platform against a shared sandbox, or, once NDA and any required BAA are signed, against a handful of your own calls: two from a top performer, two from a struggling new hire, and one each from two supervisors who do not always agree. You will leave with a clearer view of where your team's variance actually lives, whether or not we end up working together.