Sokros

Sokros 3.0 — the model family

Marking you can rerun and get the same answer.

Ask a chatbot to mark twice and you'll get two answers. Sokros 3.0 is a family of proprietary models resolved by written policy — pinned versions, explicit gates, and releases replay-tested byte for byte.

Four parts, one verdict.

No single model decides a grade. Each part does the one job it's built for, and a deterministic layer settles the outcome.

Reader

Structured extraction

Reads a submission criterion by criterion into typed fields, every claim pinned to a verbatim quote. Schema-driven per unit, multilingual-capable including Arabic.

Judge

Rubric-anchored evaluation

Grades each criterion against the unit's own rubric and a compiled grade-guidance framework, producing a band with written reasoning — runnable multiple times with a median taken.

Scorers

Statistical corroboration

An ensemble of engineered-feature models, entailment checks, cross-encoder relevance and reference-credibility classification — signals that corroborate, never unexplained verdicts.

Policy

Deterministic resolution

Written rules — band boundaries, gate caps, dispatch — resolve every judgement into the final grade. The strongest model output never wins by default; the policy does.

The logic, opened up

Why a grade can't be charmed.

Generative graders reward confident prose and drift with their context. Sokros 3.0 is built so neither is possible — four properties, each inspectable on any script.

This is what makes the system defensible on Ofqual-regulated qualifications: fair to the learner whose English is plain but whose answer is complete, and immovable for the one whose answer is polished but empty.

Extraction substantiates everything

Each criterion is read into typed fields, every claim pinned to a verbatim quote from the learner's own words. A criterion cannot be credited if no extraction exists for its subareas — there is nothing for fluent-but-empty writing to score against, and nothing for plainly-worded, complete answers to lose.

Verbatim quotesTyped fieldsMultilingual, incl. Arabic
Core requirements are binary

The brief's hard rules — counts, named elements, case application, exclusions — run as present-or-not checks against the extracted evidence. Only once they stand is any quality judgement allowed; the same order a professional marker works in, enforced by architecture.

Present or notEvidence attached
Judgement is scoped, then resolved by policy

The evaluation model bands only quality-after-requirements against the unit's own rubric, with written reasoning, runnable multiple times with a median taken. Written policy — band boundaries, gate caps, dispatch rules — settles the final grade. The strongest model output never wins by default.

Median of runsWritten reasoning
Confidence is mathematics, not prediction

Flags come from measurable margins: gate outcomes, judging spread, corroborating statistical signals. A flag means a number was thin and a human should look — never a model guessing at its own reliability.

Deterministic flagsHuman routing
Evidence modes — AC 2.4 · presentation task
Video — recorded presentationtranscript extracted · evidence found at 04:1212:40
Chart — engagement survey (fig. 3)native chart data parsed · trend claim verifiedparsed
Table — appendix Brows read as typed fields · counts checkedread
Image — process diagramstages detected against the criterion's required listread

…as our engagement survey shows, absence fell once…04:12

Same gates, same rubric, same feedback sheet

One pipeline, every mode

A presentation, a chart and an essay walk into the same gate.

Where other systems stall on anything but text — or on a layout they weren't trained on — Sokros 3.0 extracts and classifies every evidence mode through the same schema-driven reading: transcripts from video, natively-parsed charts, tables as typed fields. The gates and rubric that decide a written answer decide these too.

See the modes

Ready before your data arrives.

Sokros is not an empty model waiting for your cohort to teach it. Units ship calibrated — the system is fully capable from the first assessment.

Sharp from script one

No thousand-assessment warm-up and no training on your queue. An outcome is decided identically whether it is the first script marked or the five-hundredth.

Any centre's layout

Templated or untemplated documents, tables, appendices, embedded charts — parsed as centres actually receive them, and output in each centre's own feedback sheet.

Every evidence mode

Video, audio, chart, image and table evidence extracted and gated exactly like text — one pipeline, preconfigured per criterion.

Integrity built in

Reference credibility, AI-usage analysis, plagiarism and within-cohort collusion indicators — evidence-attached signals inside the same run.

The same script, marked twice, months apart.

Determinism isn't a promise — it's an architecture. Model versions are pinned, gates are declarative, and every release replays past scripts and compares the output byte for byte before it ships.

Zero tolerance: a release is blocked if any calibration script flips between Refer and a clear Pass

Every behavioural change is a numbered amendment with written rationale

Past brief years stay runnable and auditable forever

Replay check — release 3.0.14
Script 12 of 38 — 5CO01, attempt 1Identical
Script 13 of 38 — 5CO01, attempt 1Identical
Script 14 of 38 — 5CO01, resubmissionReviewing diff

Feedback wording changed under amendment #214 — grade unchanged, diff attached

Byte-level comparison37 identical · 1 explained · 0 flips

Third generation, same principle.

Every generation kept the rule that matters: rules where rules exist, judgement where judgement is needed — and a written reason for everything.

1.0

Proved a single unit could be marked with evidence attached to every decision.

2.0

Made the engine unit-agnostic: new qualifications and brief years became configuration, not code.

3.0

The current generation: median-of-runs judging, the statistical scoring ensemble, resubmission awareness and byte-level replay testing on every release.

Watch 3.0 mark one of yours.

Bring one human-marked script. We'll run it through the full pipeline and walk you through every decision it made — and the evidence behind each one.