Call quality assurance is the practice of reviewing customer conversations against a defined standard, so that service quality can be measured, coached and kept consistent. In most contact centres it means a supervisor listening to a handful of recordings per agent per month and filling in a scorecard.
That description sounds reasonable until you look at the arithmetic. A team handling 20,000 calls a month, reviewing five per agent across forty agents, is sampling 200 calls — one percent. Every conclusion about quality is drawn from that one percent, and the sample was not random.
This guide covers what call QA actually measures, how the sampling problem distorts it, and what changes when coverage goes from one percent to all of it.
What call QA measures
A QA scorecard breaks a conversation into observable behaviours. Most scorecards group them into categories along these lines:
| Category | What it checks | Example |
|---|---|---|
| Call opening | Greeting, identification, delay before first response | Did the agent introduce themselves? |
| Personalisation | Use of the customer's name and history | Was the customer addressed by name? |
| Communication | Clarity, grammar, jargon, filler words | Was the explanation understandable? |
| Listening | Interruptions, acknowledgement, probing | Did the agent let the customer finish? |
| Tone | Friendliness, professionalism, patience | Was the tone appropriate throughout? |
| Hold procedure | Permission, reason, thanks on return | Was a reason for the hold given? |
| Problem solving | Diagnosis, ownership, resolution | Was the issue actually resolved? |
| Closing | Further assistance, verification, sign-off | Was next-step confirmation given? |
The categories matter less than the fact that they are written down. An unwritten standard is not a standard — it is each supervisor's preference, which is why two reviewers scoring the same call often disagree by twenty points.
The sampling problem
Manual QA cannot review everything, so it samples. Two things go wrong with that sample.
First, it is small. One to two percent coverage is typical. A behaviour that occurs on one call in twenty will show up in a five-call sample about a quarter of the time — so whether an agent gets flagged is substantially luck.
Second, it is rarely random. Reviewers gravitate to long calls, escalated calls, or the ones a complaint was attached to. Those are the least representative conversations you own.
| Manual QA | Automated QA | |
|---|---|---|
| Calls reviewed | 1–2% sample | 100% |
| Selection | Often escalations and long calls | Every call, no selection bias |
| Consistency | Varies by reviewer and by day | Same criteria applied identically |
| Time to feedback | Days to weeks after the call | Minutes |
| Cost driver | Analyst hours | Minutes of audio analysed |
What changes at 100% coverage
Full coverage is not simply more of the same. It changes what questions you can ask.
- Systemic problems become visible. If one agent fails a parameter, that is coaching. If ninety percent of the floor fails it, the problem is the script, the training or the parameter itself. A 2% sample cannot distinguish these.
- Coaching stops being anecdotal. "You interrupted the customer on this call" is arguable. "You interrupted on 34 of 210 calls this month, against a team average of 9" is not.
- Rare but serious events surface. Abusive language, a compliance breach, an agent dropping a call — these are exactly the events too rare for a small sample to catch reliably, and the most expensive to miss.
- Disputes get evidence. When a customer escalates, the scorecard and the recording are already there.
Fatal incidents: the calls that override the score
Most QA programmes eventually add a zero-tolerance category. A call can be technically compliant on every parameter and still be unacceptable — the agent was dismissive, sarcastic, refused to help, or hung up on the customer.
These need separate treatment because averaging hides them. An agent with a 92% average and one abusive call does not have a 92% problem. Xperia flags these as fatal incidents and reports them as a count, not a percentage, and it distinguishes an agent hanging up from a customer hanging up — a distinction that generates false accusations when missed.
Who owns call QA
In smaller teams, QA is part of the team lead's week. Past roughly fifty agents it usually becomes a dedicated function, and that is when the calibration question appears: if three analysts score the same call differently, the scores are not comparable across the teams they cover.
Calibration sessions — everyone scores the same call, then the group reconciles the differences — are the standard fix. They work, and they are also the point at which most teams discover their scorecard is ambiguous.
Getting started without buying anything
- Write down your scorecard. Whatever your reviewers are doing implicitly, make it explicit. Ten to fifteen parameters is enough to start.
- Weight it. Not every parameter matters equally. Sales floors weight probing and upselling; support desks weight listening and resolution.
- Run one calibration session. Three reviewers, one call, compare. The spread will tell you which parameters are ambiguous.
- Fix the ambiguous ones. This is the highest-value hour in the whole exercise.
- Only then think about coverage. A precise standard applied to 2% beats a vague standard applied to everything.
Call quality assurance is not really about scoring calls. It is about having a written, agreed standard for what a good conversation sounds like, and enough coverage that the numbers describe reality rather than a biased sample.
Most teams have the first without the second. Fixing coverage is a tooling question. Fixing the standard is not — and it has to come first.
