Solutions · Grading Engine
Judgement that answers twice.
Ask a chatbot to mark the same script twice and you get two answers. Sokros gives you one, and can prove it.
Zero third-party AI in the marking path
Feedback wording changed under amendment #214. Grade unchanged, diff attached.
01 · Four parts, one verdict
No single model decides a grade.
Each part does the one job it is built for, and a deterministic layer settles the outcome.
01 · Structured extraction
Reader
Reads a submission criterion by criterion into typed fields, every claim pinned to a verbatim quote from the learner's own words. Schema-driven per unit, multilingual-capable including Arabic.
02 · Rubric-anchored evaluation
Judge
Grades each criterion against the unit's own rubric and a compiled grade-guidance framework, producing a band with written reasoning. Runnable multiple times with a median taken.
03 · Statistical corroboration
Scorers
An ensemble of engineered-feature models, entailment checks, cross-encoder relevance and reference-credibility classification. Signals that corroborate, never unexplained verdicts.
04 · Deterministic resolution
Policy
Written rules for band boundaries, gate caps and dispatch resolve every judgement into the final grade. The strongest model output never wins by default; the policy does.
02 · The logic, opened up
Why a grade can't be charmed.
Generative graders reward confident prose and drift with their context. Sokros is built so neither is possible: four properties, each inspectable on any script.
This is what makes the system defensible on Ofqual-regulated qualifications: fair to the learner whose English is plain but whose answer is complete, and immovable for the one whose answer is polished but empty.
01Extraction substantiates everything
Each criterion is read into typed fields, every claim pinned to a verbatim quote from the learner's own words. A criterion cannot be credited if no extraction exists for its subareas: there is nothing for fluent but empty writing to score against, and nothing for a plainly worded, complete answer to lose.
02Core requirements are binary
The brief's hard rules (counts, named elements, case application, exclusions) run as present-or-not checks against the extracted evidence. Only once they stand is any quality judgement allowed. The same order a professional marker works in, enforced by architecture.
03Judgement is scoped, then resolved by policy
The evaluation model bands only quality-after-requirements against the unit's own rubric, with written reasoning, runnable multiple times with a median taken. Written policy settles the final grade: band boundaries, gate caps, dispatch rules. The strongest model output never wins by default.
04Confidence is mathematics, not prediction
Flags come from measurable margins: gate outcomes, judging spread, corroborating statistical signals. A flag means a number was thin and a human should look, never a model guessing at its own reliability.
03 · Marked twice, months apart
Same script in, same grade out.
Determinism is not a promise here; it is an architecture. Model versions are pinned, gates are declarative, and every release replays past scripts and compares the output byte for byte before it ships.
A general-purpose model keeps no record of why either answer happened. This engine is built the other way round: a 5CO01 script marked in October and re-marked in March returns the same grades, with the reasoning on record both times.
Zero flips tolerated
A release is blocked if a single calibration script flips between Refer and a clear Pass.
Every change numbered
Each behavioural change ships as a numbered amendment with written rationale, replayed against past scripts before release.
Brief years stay runnable
A new cohort year is a new configuration version. Nothing resets, nothing retrains, and both years stay auditable side by side.
04 · Every evidence mode
Whatever the unit submits.
“…as our engagement survey shows, absence fell once…04:12
One pipeline, every mode
A presentation, a chart and an essay walk into the same gate.
Where other systems stall on anything but text, or on a layout they were not trained on, Sokros extracts and classifies every evidence mode through the same schema-driven reading: transcripts from video, natively parsed charts, tables as typed fields. The gates and rubric that decide a written answer decide these too.
See the modes05 · From day one
Ready before your data arrives.
Sokros is not an empty model waiting for your cohort to teach it. Units ship calibrated on dozens, not thousands, of your marked scripts, fully capable from the first assessment.
Sharp from script one
No thousand-assessment warm-up and no training on your queue. An outcome is decided identically whether it is the first script marked or the five-hundredth.
Any centre's layout
Templated or untemplated documents, tables, appendices, embedded charts: parsed as centres actually receive them, and output in each centre's own feedback sheet.
Every mode, one gate
Video, audio, chart, image and table evidence extracted and gated exactly like text. One pipeline, preconfigured per criterion.
Integrity built in
Reference credibility, AI-usage analysis, plagiarism and within-cohort collusion indicators: evidence-attached signals inside the same run.
How this stacks up against a marking panel and a generic chatbot:the honest comparison in full →
06 · Generations
Third generation, same principle.
Every generation kept the rule that matters: rules where rules exist, judgement where judgement is needed, and a written reason for everything.
1.0
Proved a single unit could be marked with evidence attached to every decision.
2.0
Made the engine unit-agnostic: new qualifications and brief years became configuration, not code.
3.0
The current generation: median-of-runs judging, the statistical scoring ensemble, resubmission awareness and byte-level replay testing on every release.
The commitments in full: what we hold ourselves to →
Watch it mark one of yours.
Bring one human-marked script. We run it through the full pipeline and walk you through every decision it made, and the evidence behind each one.
A 30-minute briefing: one unit scoped, your own script marked in front of you before anything integrates.