Most QA scorecards are inherited rather than designed. Someone adapted a template years ago, parameters accumulated, and nobody has asked since whether the thing measures what the business cares about.
This is a method for building one deliberately. It takes an afternoon and it is worth considerably more than the tooling you put around it — a precise scorecard applied to a small sample beats a vague scorecard applied to every call.
Step 1: Write down what a good call sounds like
Before parameters, prose. In plain language, describe a call you would be happy for a customer to receive. Three or four sentences. Do it separately with two or three colleagues and compare.
The disagreements are the useful output. If one person's ideal call is efficient and another's is thorough, you have found a tension that will otherwise show up as inconsistent scoring forever.
Step 2: Turn behaviours into parameters
A parameter has to be observable. "Was the agent professional?" is not a parameter; it is an opinion. "Did the agent interrupt the customer?" is.
| Weak parameter | Why it fails | Better |
|---|---|---|
| Was the agent professional? | Unobservable, reviewer-dependent | Did the agent avoid slang and filler? |
| Was the call handled well? | Restates the whole score | Was the stated issue resolved on this call? |
| Was the customer satisfied? | Not observable from the call | Did the agent confirm the customer's issue was addressed? |
| Good tone | No threshold | Did the agent stay calm when the customer raised their voice? |
Aim for ten to fifteen parameters to start. Teams that begin with forty abandon the programme; teams that begin with twelve extend it.
Step 3: Group them, then weight the groups
Individual parameters are hard to reason about in bulk, so group them — opening, listening, communication, problem solving, closing, and whatever else your business runs on. Then assign each group a share of 100 points.
The weighting is where strategy enters. A collections team and a technical support desk should not have the same distribution:
| Category | Support desk | Outbound sales |
|---|---|---|
| Listening | 20 | 10 |
| Problem solving | 25 | 10 |
| Communication | 15 | 15 |
| Probing | 10 | 25 |
| Upselling | 0 | 20 |
| Closing | 15 | 15 |
| Opening | 15 | 5 |
If one scorecard has to serve both, one of those teams is being measured on the wrong things. Xperia scopes scoring models per department for this reason — same account, different standards.
Step 4: Decide what a failure costs
Binary pass/fail on every parameter is blunt. An agent who greeted the customer but forgot to give their name has not failed the opening as completely as one who launched straight into questions.
Partial credit handles this: a missed parameter scores a defined fraction of its points rather than zero. Decide that fraction deliberately, because it changes behaviour — a floor of zero makes agents defensive about edge cases.
Step 5: Separate fatal breaches from scoring
Some things should not be a deduction. Abusive language, refusing to help, hanging up on a customer, or a compliance breach are categorically different from a weak closing.
Keep them outside the score, as a flag and a count. An agent averaging 92% with one abusive call does not have a 92% problem, and any averaging that hides it is working against you.
Step 6: Calibrate before you roll out
Three reviewers, one call, score independently, compare. Then do it twice more with different calls.
Expect a wide spread the first time. Every disagreement points at a parameter whose wording is ambiguous, and fixing those is the highest-value hour in this whole process. A scorecard that two reviewers apply differently is not a standard.
Step 7: Review it quarterly, not never
Products change, scripts change, and parameters that made sense two years ago start measuring nothing. Two checks each quarter: which parameters does everyone always pass, and which does everyone always fail?
- Always passed — either the behaviour is now automatic, or the parameter is too easy. Either way it is not telling you anything.
- Always failed — almost never an agent problem at that scale. Look at the script, the training, or the parameter's wording.
A good scorecard is short, observable, weighted to your actual priorities, and applied consistently enough that two reviewers reach the same number. Nothing about that requires software.
Coverage is the part that does. Once the standard is right, the question becomes how much of your call volume it is applied to — and one to two percent is where most programmes stall.
