Three ways to mark. One survives an audit.
Marking panels drift and drown. Chatbots improvise. The comparison below is the one your quality assurer would draw.
| What matters | Marking panel | Generic AI chat | Sokros |
|---|---|---|---|
| Same script, same grade | โ | โ | โ |
| Verbatim evidence per decision | ยฑ | โ | โ |
| Brief requirements as explicit checks | ยฑ | โ | โ |
| Output in the official template | โ | โ | โ |
| Complete, replayable audit trail | โ | โ | โ |
| Learner data stays in your boundary | โ | โ | โ |
| Consistent at end-of-term volume | โ | โ | โ |
| Resubmissions marked against the referral | ยฑ | โ | โ |
ยฑ: achievable sometimes, dependent on individual diligence and workload.
The chatbot problem, in one picture.
Ask a general-purpose model to mark the same essay twice and you get two different answers, and no record of why either happened. Sokros is engineered the other way: pinned model versions, declarative gates, a written policy layer, and byte-for-byte replay tests on every release.
A release is blocked if a single calibration script flips between Refer and a clear Pass. That's the bar an awarding body deserves.
And your assessors?
They stop being throughput and start being judgement. Sokros does the reading, checking and drafting; your team samples, reviews and signs off, with the evidence already on the table.