Sokros

Solutions · Grading Engine

Judgement that answers twice.

Ask a chatbot to mark the same script twice and you get two answers. Sokros gives you one, and can prove it.

Zero third-party AI in the marking path

Replay check · release 3.0.14
Script 12 of 38 · 5CO01, attempt 1Identical
Script 13 of 38 · 5CO01, attempt 1Identical
Script 14 of 38 · 5CO01, resubmissionReviewing diff

Feedback wording changed under amendment #214. Grade unchanged, diff attached.

Byte-level comparison37 identical · 1 explained · 0 flips

01 · Four parts, one verdict

No single model decides a grade.

Each part does the one job it is built for, and a deterministic layer settles the outcome.

01 · Structured extraction

Reader

Reads a submission criterion by criterion into typed fields, every claim pinned to a verbatim quote from the learner's own words. Schema-driven per unit, multilingual-capable including Arabic.

02 · Rubric-anchored evaluation

Judge

Grades each criterion against the unit's own rubric and a compiled grade-guidance framework, producing a band with written reasoning. Runnable multiple times with a median taken.

03 · Statistical corroboration

Scorers

An ensemble of engineered-feature models, entailment checks, cross-encoder relevance and reference-credibility classification. Signals that corroborate, never unexplained verdicts.

04 · Deterministic resolution

Policy

Written rules for band boundaries, gate caps and dispatch resolve every judgement into the final grade. The strongest model output never wins by default; the policy does.

02 · The logic, opened up

Why a grade can't be charmed.

Generative graders reward confident prose and drift with their context. Sokros is built so neither is possible: four properties, each inspectable on any script.

This is what makes the system defensible on Ofqual-regulated qualifications: fair to the learner whose English is plain but whose answer is complete, and immovable for the one whose answer is polished but empty.

Extraction substantiates everything

Each criterion is read into typed fields, every claim pinned to a verbatim quote from the learner's own words. A criterion cannot be credited if no extraction exists for its subareas: there is nothing for fluent but empty writing to score against, and nothing for a plainly worded, complete answer to lose.

Verbatim quotesTyped fieldsMultilingual, incl. Arabic
Core requirements are binary

The brief's hard rules (counts, named elements, case application, exclusions) run as present-or-not checks against the extracted evidence. Only once they stand is any quality judgement allowed. The same order a professional marker works in, enforced by architecture.

Present or notEvidence attached
Judgement is scoped, then resolved by policy

The evaluation model bands only quality-after-requirements against the unit's own rubric, with written reasoning, runnable multiple times with a median taken. Written policy settles the final grade: band boundaries, gate caps, dispatch rules. The strongest model output never wins by default.

Median of runsWritten reasoning
Confidence is mathematics, not prediction

Flags come from measurable margins: gate outcomes, judging spread, corroborating statistical signals. A flag means a number was thin and a human should look, never a model guessing at its own reliability.

Deterministic flagsHuman routing

03 · Marked twice, months apart

Same script in, same grade out.

Determinism is not a promise here; it is an architecture. Model versions are pinned, gates are declarative, and every release replays past scripts and compares the output byte for byte before it ships.

A general-purpose model keeps no record of why either answer happened. This engine is built the other way round: a 5CO01 script marked in October and re-marked in March returns the same grades, with the reasoning on record both times.

Zero flips tolerated

A release is blocked if a single calibration script flips between Refer and a clear Pass.

Every change numbered

Each behavioural change ships as a numbered amendment with written rationale, replayed against past scripts before release.

Brief years stay runnable

A new cohort year is a new configuration version. Nothing resets, nothing retrains, and both years stay auditable side by side.

04 · Every evidence mode

Whatever the unit submits.

Evidence modes · AC 2.4 · presentation task
Video · recorded presentationtranscript extracted · evidence found at 04:1212:40
Chart · engagement survey (fig. 3)native chart data parsed · trend claim verifiedparsed
Table · appendix Brows read as typed fields · counts checkedread
Image · process diagramstages detected against the criterion's required listread

…as our engagement survey shows, absence fell once…04:12

Same gates, same rubric, same feedback sheet

One pipeline, every mode

A presentation, a chart and an essay walk into the same gate.

Where other systems stall on anything but text, or on a layout they were not trained on, Sokros extracts and classifies every evidence mode through the same schema-driven reading: transcripts from video, natively parsed charts, tables as typed fields. The gates and rubric that decide a written answer decide these too.

See the modes

05 · From day one

Ready before your data arrives.

Sokros is not an empty model waiting for your cohort to teach it. Units ship calibrated on dozens, not thousands, of your marked scripts, fully capable from the first assessment.

Sharp from script one

No thousand-assessment warm-up and no training on your queue. An outcome is decided identically whether it is the first script marked or the five-hundredth.

Any centre's layout

Templated or untemplated documents, tables, appendices, embedded charts: parsed as centres actually receive them, and output in each centre's own feedback sheet.

Every mode, one gate

Video, audio, chart, image and table evidence extracted and gated exactly like text. One pipeline, preconfigured per criterion.

Integrity built in

Reference credibility, AI-usage analysis, plagiarism and within-cohort collusion indicators: evidence-attached signals inside the same run.

How this stacks up against a marking panel and a generic chatbot:the honest comparison in full →

06 · Generations

Third generation, same principle.

Every generation kept the rule that matters: rules where rules exist, judgement where judgement is needed, and a written reason for everything.

1.0

Proved a single unit could be marked with evidence attached to every decision.

2.0

Made the engine unit-agnostic: new qualifications and brief years became configuration, not code.

3.0

The current generation: median-of-runs judging, the statistical scoring ensemble, resubmission awareness and byte-level replay testing on every release.

The commitments in full: what we hold ourselves to →

Watch it mark one of yours.

Bring one human-marked script. We run it through the full pipeline and walk you through every decision it made, and the evidence behind each one.

A 30-minute briefing: one unit scoped, your own script marked in front of you before anything integrates.