For twenty years, Arabic-language assessment has run on a quiet compromise. The learner writes in Arabic; the system translates; the marker reads English; the feedback travels back through the same wash cycle. Somewhere in that loop the evidence dies: the phrase the learner actually wrote, the idiom that carried the argument, the quote a moderator would need to check the grade. Four Gulf states have now decided, in policy and in procurement, that the compromise is over. Their assessment AI reads Arabic first.
Saudi Arabia, the UAE, Qatar and Kuwait all run Sokros as the Arabic-first layer in their qualification and skills programmes. What that means in practice is four things: legibility, extraction, training and marking, each done in the language the learner used. This is what each involves, and who is running it.
Legibility: reading what is actually there
Arabic script is unforgiving to systems trained on Latin text. Diacritics change meaning. Handwriting collapses letterforms. A professional qualification script mixes right-to-left prose with left-to-right numerals, English technical terms and the occasional Latin abbreviation, sometimes in the same sentence. Generic OCR pipelines lose the diacritics, guess at the handwriting and silently flip the mixed-direction lines.
The Sokros extraction stack was built with Arabic as a first-class marking language, developed with MBZUAI and hardened on real submissions from Gulf centres. The legibility numbers below are measured on held-out scripts from live programmes, not on clean test sets.
Arabic-first extraction · measured on live programme scripts
Character-level legibility on typed Arabic submissions
98.9%
vs 91.2% for translate-then-read pipelines
Legibility on handwritten Arabic, including mixed-direction lines
96.7%
+5.4 pts since January
Dialect and register variants covered, Gulf, Levantine and Modern Standard
41
Quotes translated before marking: evidence stays in the learner's own words
0
Extraction with the evidence trail intact
Verbatim quoting is the hardest thing to keep when the quote is right-to-left inside a left-to-right marksheet. Most systems solve it by translating, which is to say they do not solve it: a moderator cannot check a grade against a quote the system paraphrased into another language. Sokros quotes in the original script, pins the quote to the page and line it came from, and carries that pin through marking, moderation and appeal. The evidence chain holds in both scripts because it never leaves either of them.
Two ways to handle an Arabic script
Arabic-first, as deployed
- Script read natively: diacritics, handwriting and mixed-direction lines preserved
- Every extracted field quoted in the learner's own Arabic, pinned to page and line
- Mark scheme applied in Arabic and English against one descriptor set
- Moderator re-reads the evidence trail in the script's own language
The translate-then-read pipeline
- Arabic flattened into English before anyone or anything reads it
- Quotes paraphrased by the translation step, unpinned from the page
- Idiom and register lost: the strongest evidence weakened by the wash cycle
- Moderation checks a translation of a judgement about a translation
Four states, four programmes
The deployments differ in shape but not in principle. Saudi Arabia runs the largest: Arabic-first marking inside the national skills programmes aligned with SDAIA's frameworks, with the Sokros layer operating inside HUMAIN's sovereign stack. The UAE pairs delivery with research: KHDA-regulated providers mark in Arabic and English against one scheme, while the MBZUAI collaboration keeps pushing legibility forward. Qatar's programme runs through its education and higher-education institutions, and Kuwait's through its applied training colleges, both anchored on the same requirement: the record must stand in front of the national regulator, in Arabic, without a translator in the room.
Arabic-language submissions marked · trailing 12 months
Training, in-region, in Arabic
Arabic-first is also a training commitment. Calibration sets for Gulf programmes are built from Arabic scripts, marked by Arabic-speaking assessors, and the models in the marking path are tuned in-region on infrastructure the state controls. Learner work does not leave the jurisdiction to make the system better. The calibration arithmetic does not change with the language: dozens of human-marked scripts per unit, not thousands, and every release replays the Arabic record with the same byte-for-byte gate as the English one.
A qualification marked in Arabic should be auditable in Arabic. Anything less is a compromise the learner pays for.
Where this goes next
The next step is already in motion: Arabic-first moderation, where the IQA sample, the EQA visit and the appeal all run in the script's own language end to end, and the first cross-state frameworks where a unit marked in Jeddah and a unit marked in Doha land on one comparable record. The language of the learner is not a localisation problem. It is the ground the whole system stands on.