Sokros

Deterministic judgement infrastructure

Instituting the AI judgement infrastructure.

The deterministic decision layer, for when predictive is not policy.

Live on CIPD today · 34 centres · 20 markets · zero third-party AI in the marking path

01 · The layer

Extract. Structure. Judge. Trace.

Sokros' own deep-learning stack does the reading: text, video, audio, charts and tables into typed, quoted evidence. Written rules do the judging, quantitative and qualitative. Nothing predictive sits between the evidence and the outcome.

  1. 01Evidence, quoted, from any medium
  2. 02Your schema, filled and sourced
  3. 03Hard rules first, quality after
  4. 04Every decision replays

The models read. The rules decide. The record remembers. How the engine works →

01 · Extract

Evidence, quoted, from any medium

A contract page, a site-visit recording and a certificate register read by the same extraction stack. No field is filled without the words, frame or cell that earned it.

MSA_Meridian.pdf · clause 14.2

aggregate liability is capped at £2,000,000

site_walkthrough.mp4 · 03:12

the second fire door was still propped open

cert_register.xlsx · table 3

forklift operator · certificate expires 2026-03-01

02 · Structure

Your schema, filled and sourced

Evidence lands as typed fields in the schema you define, each value carrying its source. Structured output for your system of record, not a paragraph to re-read.

liability_cap£2,000,000cl. 14.2
fire_door_compliantfalsevideo · 03:12
cert_expiry.forklift2026-03-01table 3

28 fields on this review, every one quoted. Exported to your schema.

03 · Judge

Hard rules first, quality after

Quantitative limits run as present-or-not rules against the extracted values. Qualitative quality is banded against your own descriptors, with the reasoning written down. In that order, every time.

Cap at or above £1,000,000£2,000,000 extracted · rule met
Certificates in date on inspection dayforklift expires 2026-03-01 · rule failed
Banded: remediation plan judged Credible against your descriptors; reasoning recorded with the quotes that support it.

04 · Trace

Every decision replays

Verdict, rule, quote and model version are stored as one record. Re-run it next year and it lands identically, which is what an auditor, a regulator or an appeal panel actually needs from you.

Decision recordverdict · rule · quote · model version
Re-runidentical, byte for byte
Audit packexported 09:12, before the meeting

02 · The doctrine

Probability is not policy.

Ask a generative model the same question twice and you get two answers. An institution cannot certify a coin flip. Sokros checks requirements as present or not against quoted evidence, bands quality only after they stand, and returns the same record every run.

The determinism, replay proofs and calibration underneath: how the engine works →

Statistical systems drift

A distribution moves with the batch. The answer you get starts to depend on what else came in that week.

Generative systems can be charmed

Fluent prose earns marks that plain English is refused. The bias sits in the phrasing, not the evidence.

Deterministic systems answer twice

Same script in, same grade out, replay-proven on every release. That is the whole doctrine.

03 · Use cases

Wherever wrong is expensive.

The same layer, pointed at different rooms: admissions desks, deal rooms, model-release boards. If a decision must survive an audit, an appeal or a headline, it runs here.

Use case 01 · identification

The right person, provably.

Enrolment records, ID pages and observed footage read into one rule trail: name match, document consistency, photo match, submission behaviour. Four rules, four recorded answers, and a decision a caseworker can reopen a year later, not a similarity score no one can defend.

Every rule's outcome storedSignals for a person, never auto-penaltiesReopenable any time
The right candidate, verified →
Identification · candidate C-2041
Name matches enrolment recordpassport page · learner file
Document internally consistentfonts, MRZ checksum, issue dates
Photo matches submitted IDstill from observed presentation
Submission behaviour in rangedevice, hours, draft cadence
Verified · rule trail stored4 of 4 · reopenable

Use case 02 · business & legal documents

The clause, found and held.

A 41-page agreement becomes typed fields: parties, term, liability cap, governing law, each pinned to the clause that says so. Hard limits check as rules, drafting is banded against your playbook, and the output lands in your schema, not in a summary.

Every field quotedOutput in your schemaPlaybook bands, reasoning recorded
Documents and bespoke work →

Use case 03 · AI safety & regulation

Judging the judges.

Sign-off on a model needs an evaluation that holds still: policy criteria as binary checks per output, severity banded with written reasoning, every verdict quoted. Re-run the suite tomorrow and it returns the same result, which is the property a regulator actually asks for.

Criteria as checks, per outputSeverity banded, reasoning attachedThe evaluation itself replays
Regulated workflows, built to spec →
Evaluation gate · model release 4.1 · policy suite C
Refusal policy held512 of 512 outputs · verdicts quoted
No personal data reproducedbinary check per output
Severity banded, reasoning attached2 outputs at band B · quotedfor review
Evaluation re-runidentical, byte for byte
Gate decisionhold for review · 2 outputs

See your own material judged: a contract, an ID pack, or one unit.

Book a demo

04 · Qualifications

One engine, your rulebook.

Education is where the layer runs live today: 34 centres, 10,000+ assessments a year, every grade replayable for moderation. Everything body-specific is authored configuration, and the full case lives behind your door.

Delivering something else regulated? Bespoke and regulated work →

05 · Partner networks

The company we keep.

Built with the researchers who stress-test the layer, the governments regulating AI and education in the UK and the Gulf, and the bodies whose names are on the certificate.

Research

evaluation methods & extraction science

  • Google DeepMind
  • MBZUAI
  • The Alan Turing Institute

Government & regulators

education and AI oversight, UK · GCC

  • Department for Education
  • Ofqual
  • UAE Ministry of Education
  • KHDA

National programmes

AI & education initiatives, KSA · UAE

  • NEOM
  • SDAIA
  • HUMAIN
  • G42

Professional bodies

qualifications marked on the record

  • CIPD
  • CMI
  • ACCA
  • Pearson

Centres, awarding bodies and platforms each have a track: how partnering works →

06 · Active work

Open files.

Live deployments, research lines and programmes being scoped, each with its honest status. Where the work is early, the ledger says so.

PRG-01

Regulated qualification marking

CIPD units at Levels 3, 5 and 7, marked in production across 20 markets.

Live

PRG-02

Deterministic evaluation research

With Google DeepMind: how far a judgement pipeline can be verified with no probabilistic step in the loop.

Research

PRG-03

Education infrastructure for new cities

Assessment rails scoped for NEOM's education programmes: marking, verification and audit as one layer.

Scoping

PRG-04

Arabic-first extraction

With MBZUAI: the extraction stack is multilingual today; this makes Arabic a first-class marking language, not a translation.

Research

PRG-05

National skills frameworks

Alignment work with SDAIA and UAE education authorities: one deterministic record beneath national qualification programmes.

Scoping

PRG-06

Government casework decisions

Scoped with UK public bodies: licensing and casework queues judged against published criteria, with the full rule trail on every decision.

Scoping

07 · Reach

From Ottawa to Jakarta.

34 centres mark on Sokros across 20 markets, 10,000+ assessments a year. Different time zones, one deterministic record, the same grade wherever the qualification runs.

London · Dublin · Paris · Berlin · Lisbon · Amsterdam · Bern · Valletta · Ankara · Cairo · Riyadh · Manama · Abu Dhabi · Muscat · Abuja · New Delhi · Kuala Lumpur · Singapore · Jakarta · Ottawa

Sovereign by design

Zero third-party AI in the marking path

No learner work touches a consumer AI service, and marking never trains a model. No external APIs, no LLM anywhere in the marking path.

In-region or in-building

UK, EU and Gulf hosting, or fully on-premise on your own GPU floor. The boundary is the institution's to draw, and the pipeline runs the same inside it.

Replay-proven

Any decision re-runs byte for byte on demand: for an EQA, an appeal, or an auditor holding a year-old question.

Retention on instruction

The pipeline stores decisions and quoted evidence, not profiles. Deletion follows the centre's instruction.

Where learner data goes, and doesn't: privacy & data usage →

Bring the work that must be right.

A demo is your own material judged in front of you: a contract, an ID pack, or a unit with a handful of human-marked scripts. Calibration is measured in dozens, not thousands.

On the call: 30 minutes, one workload scoped, a gated pilot plan. Nothing integrates before you've watched it judge.