Kingdom Summit does not suffer from a shortage of ambition. Walk the hall and every third stand promises an agent that runs the whole workflow: intake, extraction, decision, done. The demos are polished. The coffee queue, though, kept asking a plainer question. When the workflow runs itself, who answers for what it did?
We took that question on from the HUMAIN stand, where Sokros runs as the judgement layer inside HUMAIN's sovereign AI stack, and again in a packed afternoon slot on the main stage. Our answer is short enough to fit on one slide, and it took three years of production marking to earn: an agent can run the whole workflow, provided the workflow is written down before it runs, every step it takes quotes its evidence, and the whole run can be replayed afterwards, byte for byte. Autonomy first. Answerability always.
What autonomous actually means here
The word "agent" covers a lot of theatre at the moment, so it is worth being exact about ours. A Sokros agent is not a model improvising its way through a process. It is a runner executing a written procedure. A learner's brief arrives; the extraction stack reads it and quotes every field it fills; the centre's own rules check the evidence, present or not; the judgement is recorded with the quotes that support it; and the run is stored so it can be replayed later, exactly, whoever asks.
The agent decides how the work gets done: which document to read first, when a scan is too poor to use and needs re-requesting, which rule to check next. It never decides what the rules are. Those stay where they have always lived, with the awarding body and the centre, in writing.
One workflow, held to the record
- 01 · Intakebrief, scripts, enrolment
- 02 · Extractionevery field quoted
- 03 · Ruleswritten criteria, checked
- 04 · Judgementbanded, reasoning recorded
- 05 · Replaybyte for byte, on demand
The Sokros layer inside HUMAIN
HUMAIN builds the Kingdom's sovereign AI capacity: the models, the compute, the national programmes that run on both. What a sovereign stack still needs, before a ministry lets it near a consequential decision, is a way to make that decision answerable. That is the piece Sokros contributes. Inside HUMAIN's education and government programmes, our layer sits between the models and the record: HUMAIN's systems read and reason, Sokros rules judge, and the joint record replays in front of whoever regulates the outcome.
The partnership matters because of what it rules out. No learner work leaves the Kingdom's boundary. No grade is a foreign API call. And no minister has to take a vendor's word for how a decision was made, because the decision carries its own evidence. Sovereignty stopped being a hosting question and became a governance question. This is the governance layer.
The summit put numbers on the programme's reach for the first time.
Sokros × Gulf sovereign AI programmes · 2026
National AI programmes running the Sokros judgement layer, KSA and UAE
4
2 added in 2026
Ministries and authorities with Sokros-marked qualifications in delivery or pilot
11
+5 year on year
In-region deployments, Riyadh and Abu Dhabi, zero third-party AI in the marking path
2
Of recorded decisions replayable byte for byte on regulator request
100%
The job description is the governance
The part of the demo that drew the most questions was not the speed. It was the boundary. Every autonomous workflow we deploy carries a written job description: the list of things the agent may finish on its own, and the shorter list of things it must hand back to a person, with the evidence attached. Nobody signs off on "the AI handles assessment". They sign off on two columns.
The boundary, as deployed
The agent finishes alone
- Chases a missing signature page before an assessor ever opens the script
- Re-requests an unreadable scan, and says which page failed
- Applies the marking rules to quoted evidence, criterion by criterion
- Queues borderline scripts for review with the quotes already attached
It hands back to a person
- Any integrity signal: a flag is evidence for a caseworker, never a penalty
- Any grade sitting on a band boundary
- Any appeal, resubmission or reasonable-adjustment case
- Anything the written rules do not cover
Why this landed in Riyadh
The Gulf conversation has moved past whether AI should touch regulated assessment. The question now is accountability, and it is being asked by people with the power to make it stick: ministries, qualification authorities, the bodies standing up national skills frameworks. An autonomous workflow that cannot show its working is a non-starter in that room, whatever it scored on a benchmark. It is why the region's biggest programmes chose a layer that replays over a model that merely performs.
An agent can run the whole workflow. Written rules, quoted evidence and a replay are what keep it answerable.
What happens next
The follow-ups from the summit are already booked: scoped programmes with education authorities and providers across the Kingdom and the wider Gulf, each starting the same way every engagement starts. One workload, judged in front of the people who own it. Written rules, quoted evidence, a replay. Then, and only then, the agent gets the keys.
Autonomy is the easy part to demo and the hard part to defend. We would rather do it the other way round.
