Multilingual

Multilingual Call QA: Scoring Urdu, Arabic and Code-Switched Calls

N

Najoomi Press

Blog Main Image

If your agents handle calls in more than one language, most call QA tooling will quietly mislead you. It will produce scores. The scores will look plausible. And they will have been generated by translating the conversation into English first and then judging the translation.

That works acceptably for factual content and badly for everything QA actually measures. Politeness, deference, warmth and irritation live in word choice, register and delivery — precisely the things translation flattens.

What translation-first QA loses

Consider a support call in Urdu. The agent uses the formal second person throughout, then switches to the familiar form when the customer becomes agitated. To a native speaker that shift is significant — it either de-escalated the call or patronised the caller. Translated to English, both forms become "you" and the entire signal disappears.

The same applies to Arabic honorifics, Hindi register, and the difference between a brisk and a curt reply in almost any language. A scorecard that measures tone cannot measure tone through a translation layer.

SignalSurvives translation?Why
Facts and intentMostly yesContent transfers reasonably well
Politeness registerNoFormal/familiar distinctions collapse into English "you"
HonorificsNoUsually dropped entirely
Tone and deliveryNoNot present in text at all
Interruptions and pacingNoLost when audio becomes a transcript

Code-switching is the part that actually breaks things

In multilingual markets people do not pick a language and stay in it. A single support call routinely opens with an English greeting, moves into Urdu for the problem description, returns to English for technical terms and product names, and closes in Urdu again.

Systems that detect a language once at the start of the call and commit to it degrade from the first switch onward. Everything after that point is being interpreted through the wrong model, and the resulting score is noise dressed as data.

Ask a vendor "do you support Arabic?" and you will get a yes. Ask "what happens when the caller switches between Arabic and English four times?" and you will learn something.

What Xperia does differently

Xperia analyses the call audio directly rather than translating it into English and scoring the text. Over 50 languages are supported — including English, Urdu, Arabic, Hindi, French, Spanish, Bengali and Punjabi — and language changes within a single conversation are followed rather than treated as one fixed language.

Because the audio is what gets evaluated, tone and delivery are assessed in the language actually spoken. The same scoring parameters and department weightings apply regardless of language, so a multilingual floor is measured against one consistent standard instead of a different scorecard per language.

The detail most tools get wrong: right-to-left reports

Scoring the call is half the job. The other half is handing a supervisor a report they can read. Arabic and Urdu are right-to-left scripts, and most PDF generators either reverse the character order or disconnect the letter forms, producing output that is technically present and practically illegible.

Xperia reshapes RTL text correctly in exported PDF scorecards, so a report on an Urdu call reads as Urdu. It is an unglamorous detail that decides whether the output is usable.

How to test a vendor properly

Do not evaluate multilingual capability from a feature list. Take four of your own recordings and check the output yourself:

  1. A call entirely in your second language. Does the score look defensible to a native speaker on your team?
  2. A call that switches language at least twice. Does quality hold after the first switch, or does the transcript drift?
  3. A call where the agent is polite but firm. Does the tone score reflect that, or does it read as either warm or hostile?
  4. Export the scorecard as a PDF. If the script renders backwards, the integration is not finished.

That last test takes thirty seconds and eliminates a surprising number of vendors.

Why this matters commercially

Contact centres in South Asia, the Gulf and much of Africa operate multilingually by default, and they are routinely sold QA tooling designed for monolingual English floors. The result is a programme that scores the English calls credibly and everything else approximately — while reporting a single average that hides the difference.

If most of your volume is not in English, multilingual handling is not a feature to compare late in the process. It determines whether any of the numbers mean anything.

Multilingual call QA is not about the length of a supported-languages list. It is about whether the system evaluates the conversation that happened, in the language it happened in, including the parts where the language changed.

Test it with your own calls before you believe anyone, including us. Xperia includes free starting credit so you can score real recordings from your own floor and judge the output yourself.

Najoomi Technologies

Learn more about us

Contact

Delaware, United States
hello@najoomi.ai

Follow Us

hello

© 2026 Najoomi Technologies — All Rights Reserved