Contact centres measure a lot and act on little of it. Part of the reason is that the standard metrics each have a well-known failure mode, and reporting them individually invites exactly the behaviour that breaks them.
What each actually measures, and what it does when you make it a target.
The standard set
| Metric | Measures | Failure mode when targeted |
|---|---|---|
| AHT (average handle time) | Call duration plus wrap-up | Agents rush, transfer more, resolve less |
| FCR (first contact resolution) | Issues closed without a callback | Calls marked resolved that were not |
| CSAT | Post-call satisfaction rating | Low response rate, skewed to extremes |
| NPS | Willingness to recommend | Measures the brand more than the call |
| QA score | Adherence to your standard | Improves without service improving, if the scorecard is weak |
| Abandonment rate | Callers who gave up | Solved by staffing, not by agents |
AHT is the most misused number in the industry
Handle time is a cost metric wearing a quality metric's clothing. Shorter is cheaper; shorter is not better. Push AHT down as a target and you reliably get more transfers, more callbacks and lower resolution — the work moves rather than disappearing.
AHT is useful as a diagnostic. A specific queue with rising AHT is worth investigating. AHT as an agent target produces the behaviour you did not want.
CSAT has a sampling problem you already know about
Typical post-call survey response rates sit in the single digits to low teens, and the people who respond are disproportionately delighted or furious. Everything in between is missing.
So CSAT is a real signal about a self-selected minority. Movement of a point or two is usually noise, and comparing agents on CSAT with a handful of responses each is not meaningful.
Watch these together
Individually each metric can be gamed. In pairs it is much harder:
- AHT with FCR. Falling AHT and falling FCR means work is being deferred, not done.
- QA score with CSAT. Rising QA and flat CSAT suggests your scorecard measures compliance rather than experience.
- FCR with repeat contacts. Verifies FCR against behaviour rather than disposition codes.
- QA score with fatal incidents. Averages hide rare events. A high average with a rising incident count is a serious signal.
The metric most teams are missing
Almost nobody tracks coverage — what proportion of calls anyone has actually assessed. It belongs on the dashboard because it qualifies everything else. A QA score drawn from 2% of calls and one drawn from 100% are not the same kind of number, and reporting both as "QA score" hides the difference.
A workable dashboard
Four numbers and one count, reviewed weekly:
- QA score, with coverage stated alongside it
- FCR, verified against repeat contacts rather than disposition codes
- AHT as a diagnostic only, never as an agent target
- CSAT with the response rate always shown next to it
- Fatal incidents as an absolute count, never averaged
Xperia reports the first, third and fifth directly, with category and parameter drill-downs so a movement can be traced to a specific behaviour rather than left as a number that went down.
Every one of these metrics is gameable alone and informative in combination. The single most common mistake is turning a diagnostic into an agent target — AHT being the textbook case.
And state your coverage next to your QA score. Without it, the score is a claim about a sample presented as a claim about the operation.
