Call QA

QA Sampling: How Many Calls Should You Actually Score?

N

Najoomi Press

Blog Main Image

Five calls per agent per month is the most common QA sample size in the industry. It is also, for most of the conclusions drawn from it, close to meaningless — not because five is a small number, but because of what people then do with it.

This post is about what a sample of a given size can and cannot support, and how to choose one on purpose rather than by convention.

What five calls can and cannot tell you

Suppose an agent fails a particular behaviour on one call in ten. Across a five-call sample, the chance of catching it at least once is about 41%. Which means it is more likely than not that a real, recurring problem is invisible this month.

Run the same maths the other way and it gets worse: whether an agent appears to have a problem is substantially determined by which five calls were pulled.

Sample per agentChance of catching a 1-in-10 behaviourChance of catching a 1-in-50 behaviour
5 calls41%10%
10 calls65%18%
25 calls93%40%
100 calls99.9%87%
All calls100%100%

The right column matters more than it looks. Rare behaviours are where the expensive risk lives — abusive language, compliance breaches, an agent ending a call improperly. Those are exactly the events small samples cannot see.

The bigger problem is that the sample is not random

Sample size is arithmetic and at least easy to reason about. Selection bias is worse because it is invisible.

Left to human choice, reviewers tend to pull:

  • Long calls, because something presumably happened
  • Escalated calls, because they are already flagged
  • Calls attached to a complaint, because someone asked
  • Calls from the start of the month, because that is when the task was scheduled

Each is defensible individually. Together they produce a sample that is systematically unrepresentative of ordinary work — and ordinary work is what most of your customers actually experience.

Sampling deliberately

If you cannot review everything, decide the rule in advance and write it down. Three approaches that work:

  1. Random, stratified by agent. Everyone gets the same number, selected without human choice. Fair, and the baseline everything else is measured against.
  2. Random plus targeted. A random base for everyone, extra volume for new hires, agents under review, or a queue you are worried about.
  3. Risk-weighted. Higher sampling on calls matching conditions you care about — long duration, specific queues, particular dispositions — while keeping a random base so you retain a representative picture.

Xperia supports the second and third directly: a daily processing percentage, hard daily caps and per-agent sampling, plus filter rules on your own CRM metadata so only the calls matching your criteria are analysed.

What changes when the answer is "all of them"

Full coverage removes both problems at once, and it changes the questions you can ask rather than just improving precision.

QuestionAt 2% sampleAt 100%
Is this agent improving?Weak signal, high noiseDirect comparison month to month
Does the whole floor fail this parameter?Cannot distinguish from chanceImmediately visible
How often do fatal incidents occur?Unreliable — too rareExact count
Which work code has the worst quality?Sample too thin to splitEvery code, every month

The third row is the one that changes behaviour. Most teams discover their fatal-incident rate is not what they assumed, in one direction or the other.

Cost is now the real constraint

When QA was analyst hours, coverage was limited by headcount. When it is per-minute analysis, coverage is limited by budget — a different problem with different levers.

Which is why sampling controls still matter at 100% capability: a daily cap and a processing percentage let you choose coverage deliberately instead of discovering it on an invoice. See how per-minute pricing works for the arithmetic.

Five calls a month was never a statistical decision; it was what a supervisor could fit around their other work. That is a legitimate constraint, but it should be named rather than mistaken for a methodology.

Whatever coverage you choose, choose it deliberately, make the selection random unless you have a reason otherwise, and be honest about which conclusions your sample can carry.

Najoomi Technologies

Learn more about us

Contact

Delaware, United States
hello@najoomi.ai

Follow Us

hello

© 2026 Najoomi Technologies — All Rights Reserved