Sokros 3.0 — the model family
Marking you can rerun and get the same answer.
Ask a chatbot to mark twice and you'll get two answers. Sokros 3.0 is a family of proprietary models resolved by written policy — pinned versions, explicit gates, and releases replay-tested byte for byte.
Four parts, one verdict.
No single model decides a grade. Each part does the one job it's built for, and a deterministic layer settles the outcome.
Reader
Structured extractionReads a submission criterion by criterion into typed fields, every claim pinned to a verbatim quote. Schema-driven per unit, multilingual-capable including Arabic.
Judge
Rubric-anchored evaluationGrades each criterion against the unit's own rubric and a compiled grade-guidance framework, producing a band with written reasoning — runnable multiple times with a median taken.
Scorers
Statistical corroborationAn ensemble of engineered-feature models, entailment checks, cross-encoder relevance and reference-credibility classification — signals that corroborate, never unexplained verdicts.
Policy
Deterministic resolutionWritten rules — band boundaries, gate caps, dispatch — resolve every judgement into the final grade. The strongest model output never wins by default; the policy does.
The logic, opened up
Why a grade can't be charmed.
Generative graders reward confident prose and drift with their context. Sokros 3.0 is built so neither is possible — four properties, each inspectable on any script.
This is what makes the system defensible on Ofqual-regulated qualifications: fair to the learner whose English is plain but whose answer is complete, and immovable for the one whose answer is polished but empty.
01Extraction substantiates everything
Each criterion is read into typed fields, every claim pinned to a verbatim quote from the learner's own words. A criterion cannot be credited if no extraction exists for its subareas — there is nothing for fluent-but-empty writing to score against, and nothing for plainly-worded, complete answers to lose.
02Core requirements are binary
The brief's hard rules — counts, named elements, case application, exclusions — run as present-or-not checks against the extracted evidence. Only once they stand is any quality judgement allowed; the same order a professional marker works in, enforced by architecture.
03Judgement is scoped, then resolved by policy
The evaluation model bands only quality-after-requirements against the unit's own rubric, with written reasoning, runnable multiple times with a median taken. Written policy — band boundaries, gate caps, dispatch rules — settles the final grade. The strongest model output never wins by default.
04Confidence is mathematics, not prediction
Flags come from measurable margins: gate outcomes, judging spread, corroborating statistical signals. A flag means a number was thin and a human should look — never a model guessing at its own reliability.
“…as our engagement survey shows, absence fell once…04:12
One pipeline, every mode
A presentation, a chart and an essay walk into the same gate.
Where other systems stall on anything but text — or on a layout they weren't trained on — Sokros 3.0 extracts and classifies every evidence mode through the same schema-driven reading: transcripts from video, natively-parsed charts, tables as typed fields. The gates and rubric that decide a written answer decide these too.
See the modesReady before your data arrives.
Sokros is not an empty model waiting for your cohort to teach it. Units ship calibrated — the system is fully capable from the first assessment.
Sharp from script one
No thousand-assessment warm-up and no training on your queue. An outcome is decided identically whether it is the first script marked or the five-hundredth.
Any centre's layout
Templated or untemplated documents, tables, appendices, embedded charts — parsed as centres actually receive them, and output in each centre's own feedback sheet.
Every evidence mode
Video, audio, chart, image and table evidence extracted and gated exactly like text — one pipeline, preconfigured per criterion.
Integrity built in
Reference credibility, AI-usage analysis, plagiarism and within-cohort collusion indicators — evidence-attached signals inside the same run.
The same script, marked twice, months apart.
Determinism isn't a promise — it's an architecture. Model versions are pinned, gates are declarative, and every release replays past scripts and compares the output byte for byte before it ships.
Zero tolerance: a release is blocked if any calibration script flips between Refer and a clear Pass
Every behavioural change is a numbered amendment with written rationale
Past brief years stay runnable and auditable forever
Feedback wording changed under amendment #214 — grade unchanged, diff attached
Third generation, same principle.
Every generation kept the rule that matters: rules where rules exist, judgement where judgement is needed — and a written reason for everything.
1.0
Proved a single unit could be marked with evidence attached to every decision.
2.0
Made the engine unit-agnostic: new qualifications and brief years became configuration, not code.
3.0
The current generation: median-of-runs judging, the statistical scoring ensemble, resubmission awareness and byte-level replay testing on every release.
Watch 3.0 mark one of yours.
Bring one human-marked script. We'll run it through the full pipeline and walk you through every decision it made — and the evidence behind each one.