Three ways to mark. One survives an audit.
Marking panels drift and drown. Chatbots improvise. Here's the honest comparison your quality assurer would draw.
| What matters | Marking panel | Generic AI chat | Sokros |
|---|---|---|---|
| Same script, same grade | ✕ | ✕ | ✓ |
| Verbatim evidence per decision | ± | ✕ | ✓ |
| Brief requirements as explicit checks | ± | ✕ | ✓ |
| Output in the official template | ✓ | ✕ | ✓ |
| Complete, replayable audit trail | ✕ | ✕ | ✓ |
| Learner data stays in your boundary | ✓ | ✕ | ✓ |
| Consistent at end-of-term volume | ✕ | ✓ | ✓ |
| Resubmissions marked against the referral | ± | ✕ | ✓ |
± — achievable sometimes, dependent on individual diligence and workload.
The chatbot problem, in one picture.
Ask a general-purpose model to mark the same essay twice and you get two different answers — and no record of why either happened. Sokros is engineered the other way: pinned model versions, declarative gates, a written policy layer, and byte-for-byte replay tests on every release.
A release is blocked if a single calibration script flips between Refer and a clear Pass. That's the bar an awarding body deserves.
And your assessors?
They stop being throughput and start being judgement. Sokros does the reading, checking and drafting; your team samples, reviews and signs off — with the evidence already on the table.