Aller au contenu
Voix9 septembre 202614 min de lecture

Voice agents in Arabic dialects, and where they break

Modern Standard Arabic is a written language. Callers speak Levantine, Gulf, Egyptian and Maghrebi, switch to English mid-sentence and read numbers in two directions. Here is where the systems fail, and what a bank or a ministry should demand before it lets one answer the phone.

Par Altuon

Modern Standard Arabic is the language of the news bulletin, the contract and the ministry circular. Nobody uses it to ask why a card was declined. A caller in Amman, Riyadh, Cairo or Casablanca speaks the Arabic of the street they grew up on, drops into English or French for the words the bank taught them, reads an account number the way their mother taught them to count, and expects to be addressed with the respect their name carries. A voice agent that was demonstrated in a quiet room on a scripted Standard Arabic dialogue has been tested against none of this.

This article is written for the executive who has to sign: the chief operating officer of a bank, the director general of a ministry, the head of a telecoms contact centre. It sets out where Arabic voice agents fail, in the order a call exposes the failures, and what to demand of a vendor before one is allowed to answer a customer. The position is simple. The technology is ready for production in Arabic, but only for the institution that tests it on its own calls, deploys it where its regulator can see it, and gives it a graceful way to say that it does not know.

The language on the demo is not the language on the line

Arabic is one written language and several spoken ones. Modern Standard Arabic is learned at school and used for writing, broadcasting and formal speech; it is nobody's mother tongue. The spoken varieties differ from it and from each other in vocabulary, grammar, pronunciation and rhythm. Four groups matter for a service line in the region: Levantine (Jordan, Syria, Lebanon and Palestine), Gulf (the Arabian peninsula), Egyptian (the most widely understood, after decades of film and television), and Maghrebi (Morocco, Algeria and Tunisia, the furthest from the rest and the most heavily mixed with French).

The differences are not accents. They are different words for the most frequent functions in a sentence.

EnglishLevantineGulfEgyptianMaghrebiStandard
I wantbiddiabi, abghaʿayizbghiturīd
whatshuwesh, shuehash, shnumādhā
nowhallaʾal-hīndilwaʾtidābaal-ān
finemnīḥzēnkwayyismezyānjayyid

Pronunciation moves too. The letter qāf is a glottal stop in Cairo and in urban Levantine speech, a hard g in most of the Gulf and in rural Jordan, and close to the textbook sound in the cities of the Maghreb. The Egyptian jīm is a hard g. Negation, question formation and the future tense are each built differently. A speech model trained on Standard Arabic and one prestige dialect will transcribe a Cairo caller adequately and a Casablanca caller badly, and it will do both with equal confidence.

The consequence for a buyer is that "supports Arabic" is not a claim that can be evaluated. The questions are which dialects, tested how, on whose recordings, and with what result per dialect rather than on average.

Callers switch languages, and numbers run in two directions

Code-switching is the norm

A Jordanian or Gulf caller moves between Arabic and English inside a single sentence, and does so most often for exactly the words a bank or a telecoms company cares about: transfer, account, card, statement, OTP, app, branch. A Lebanese caller may add French. A Moroccan or Tunisian caller will conduct the whole transactional part of the call in French borrowings: virement, compte, carte, prélèvement. This is not a failure of the caller's Arabic. It is how the language is spoken by the people who hold accounts.

Recognisers handle this in one of three ways. Some are locked to one language per call and transcribe the English words as Arabic nonsense. Some detect the switch and swap models, losing the first syllable after every switch. A few are trained on mixed speech and treat it as ordinary. Only the third is fit for a service line. The test is a recording of a real customer saying "biddi aʿmal transfer la account tani" and a transcript that gets every word.

Numbers

Spoken Arabic reads compound numbers units first: twenty-five is "five and twenty", and a hundred and twenty-five is "a hundred and five and twenty". A caller reading an account or reference number may read it digit by digit, in pairs, in hundreds, or entirely in English, and may change method halfway through. Written Arabic runs from right to left but the digits inside it run from left to right, which matters for every transcript an agent shows to a person and for every read-back it generates. Egypt, the Levant and the Gulf write numbers with the Eastern Arabic digits; the Maghreb uses the same Western digits as Europe. A system that normalises numbers correctly for one market will corrupt them for another, and a corrupted account number is not a transcription error. It is a payment to the wrong person.

The design answer is procedural, not linguistic. Every critical value is read back to the caller in the caller's own convention and confirmed before anything is done with it. A vendor who cannot demonstrate the read-back is not ready.

Names and honorifics

Arabic names arrive in several forms. The name on the account may be a chain of given name, father, grandfather and family; the caller may give their kunya, Abu Khalid or Umm Sara, instead of their given name; a family name may carry the definite article or drop it; the same name is transliterated in the bank's systems as Mohammed, Muhammad or Mohamed. Matching the spoken name to the record is a data problem as much as a speech problem, and it has to be solved before verification can rely on it.

Honorifics are the register in which the institution speaks to its customers. Ustaz, Doktor, Sheikh, Hajj, Bash-muhandis, the vocative ya, the plural forms of respect: a caller expects them, and an agent that uses the wrong one, or none, has told the customer what the institution thinks of them. Arabic also inflects the verb and the adjective for the gender of the person addressed. The agent must know whether it is speaking to a man or a woman, or be written so that it never has to guess. Neither is automatic in a system designed in English.

Where the recogniser fails

Speech recognition fails in three layers, and each has to be tested on its own.

The first is dialect, described above. The second is the acoustic environment: a caller on a mobile phone in a car, in a souk, in a kitchen with a television on, on a speakerphone, with a child on the line. Arabic has a set of consonants distinguished by fine acoustic cues, including the emphatic and pharyngeal sounds that have no equivalent in European languages, and noise removes exactly those cues first.

The third is the telephone itself. Most calls still travel over narrowband channels sampled at eight kilohertz, with a passband that ends well below the frequencies that separate s from and t from . A model that scored well on studio audio or on smartphone app recordings has not been tested on the channel your customers use. The test set must be made from telephone recordings, not from recordings of telephones.

Two further behaviours deserve a named test. Large speech models sometimes produce fluent text during silence, hold music or noise, a behaviour usually called hallucination; on a service line that text becomes an intent and then an action. And Arabic transcripts are written without the short vowels, so a transcript can be correct as text and still ambiguous as meaning. What the model passes downstream matters as much as what it heard.

Recognising words is not understanding intent

A perfect transcript can still lead to the wrong action. Arabic conversation, in every dialect, tends to open with greeting and blessing formulas, place the request at the end of a long preamble, and soften it with indirectness. The caller says a great deal about the card they received last month, the branch, the queue and the previous phone call before the sentence that contains what they want, and that sentence may contain two requests. A language model tuned on English support tickets looks for the request in the first line.

The second failure is translation in the middle. Some systems transcribe Arabic, translate it to English for the reasoning step, and translate the answer back. The dialect idiom, the honorific and the caller's exact wording of a number are lost twice. The agent must reason in the caller's language, or be tested to prove that the round trip loses nothing that matters.

The third is over-confidence. An agent that has not understood should say so, once, in the caller's dialect, and ask again. An agent that guesses at intent and acts on the guess is worse than the menu it replaced, because the menu at least made the caller press the wrong button themselves.

An agent that has not understood should say so, once, in the caller's dialect, and ask again. An agent that guesses and acts is worse than the menu it replaced.

Every call has a legal opening before it has a conversational one. Jordan's Personal Data Protection Law, the GDPR for any European customer or branch, the revised Swiss FADP, the data-protection laws of the Gulf states and the consumer-protection rules of the region's telecoms and banking regulators each, in their own terms, require that a caller be told that the call is recorded and why, and that the recording be kept only as long as its purpose requires. The EU AI Act adds, where it applies, an obligation to tell a person that they are dealing with a machine unless that is obvious from the context. The Central Bank of Jordan and its peers in the region expect a supervised institution to know where its customer data is, who processes it, and how a complaint about the channel is received and resolved.

The notice has to be in the caller's language, spoken before recording begins, and it has to have an answer for refusal. A caller who declines is not a failed call; they are routed to a person on an unrecorded line, where that is what the law and the institution's policy require. Consent, where it is the lawful basis, has to be captured in a form that can be produced later, per call.

Voice itself is a further matter. A voiceprint used for verification is biometric data, a special category under the GDPR and sensitive data under several of the region's laws. Voice verification therefore needs explicit consent, a documented case for necessity and a plan for deletion, and it should never be the only factor for a transaction the policy marks as sensitive.

Where the audio lives

For a bank, an insurer or a ministry in the region, the architecture question comes before the accuracy question. A hosted speech model behind an API means that customer audio leaves the country every time the phone rings. The data-protection regimes of Saudi Arabia, the United Arab Emirates, Jordan and Egypt each constrain the cross-border transfer of personal data, by different mechanisms; banking regulators add outsourcing rules that require the institution to know, approve and be able to audit every party that processes customer data; and a sovereign programme may require that the models themselves run on national infrastructure.

The workable design places the speech recogniser, the language model and the transcript store inside the institution's own data plane, on its premises or in a national cloud, and keeps only policy, configuration and aggregate metrics in any control plane the vendor operates. Ask the vendor three questions. Does "region" mean that audio is stored in the region, or that it is also processed there? Is the model that runs on-premises the same model that was demonstrated, or a smaller one? And what is in the sub-processor list, by name, for the path a recording takes from the carrier to the transcript?

The trade-off is honest and should be stated in the proposal. On-premises deployment requires infrastructure the institution provides and operates, limits the choice of models to those that fit that footprint, and carries a cost line of its own. The alternative is a residency promise in a contract, which no auditor can inspect.

How to test a vendor

Do not accept a demonstration. Build the test, and make it from your own calls.

Take recordings from your existing contact centre, under the lawful basis and the anonymisation your privacy function approves, and build a set that is stratified by dialect, by gender, by age group, by acoustic condition, by channel and by intent. Include the calls your agents found hard. Include silence, hold music and a caller who says nothing for ten seconds. Include the caller who asks for a person in the first sentence and the caller who asks a question the agent is not permitted to answer. Then measure what matters to the business rather than what is easy to compute: whether the account number was captured correctly, whether the intent was understood, whether the agent handed over when it should have and did not when it should not have, and how often it said "I don't know" when that was the true answer.

AskA serious answerAn answer that ends the conversation
Which dialects were in the training data?Named dialects, with per-dialect results on our own test set"Arabic", or a single average figure
What happens on a request outside policy?A demonstrated refusal in the caller's dialect, then a handover"The model is very capable"
What does the agent do when it is not sure?Asks once, then hands over; the threshold is ours to setConfidence is not exposed
Where is audio processed, as distinct from stored?A named data plane and named sub-processorsA region name with no distinction
Who owns the evaluation set and the conversation design?The client, as deliverables on paymentRetained as vendor property
Can we run the test set ourselves before every release?Yes, with the harness handed overTesting is done on the vendor's side

Refusal behaviour deserves particular attention. An agent that cannot say "I don't know" will eventually say something worse, on a recorded line, to a customer. Test it with questions that have no answer in its policy, with a caller who insists, and with a caller who tries to talk it into a transaction it is not permitted to make. The agent that declines politely, in dialect, every time, is the one that can be defended to a regulator.

What to do on Monday

  1. Pull a sample of recordings from your existing service line, under a basis your privacy officer signs, and have a native speaker tag each one by dialect, language mix, noise and intent. This is the beginning of your test set, and it belongs to you.
  2. Write down the intents the agent may resolve alone, the ones it may resolve after verification, and the ones it must always hand to a person. Have compliance sign it before any vendor sees it.
  3. Ask your legal function for the recording notice, the consent capture and the retention period for each jurisdiction you serve, and decide what the agent does when a caller declines.
  4. Decide the data plane before the vendor shortlist: on-premises, national cloud or hosted, and what your banking or sector regulator requires you to be able to show for each.
  5. Put the six questions in the table above into the request for proposal, and score the written answers before any demonstration is scheduled.
  6. Plan the first slice of live traffic as one intent, one dialect and one segment, with a person reviewing every call, and set the numbers that have to hold before the slice widens.

Dites-nous ce qui ne peut pas échouer.

Une demande de proposition compte sept étapes courtes et est lue par le responsable de mission qui conduirait le travail. Des références de missions réelles sont fournies sous confidentialité, selon votre secteur et votre région.