Why we publish benchmarks
Claims about AI accuracy in education are easy to make and notoriously hard to verify. We believe the only way to build trust with teachers is to benchmark our models against experienced human markers and publish the exact methodology and numbers.
Benchmark methodology
Test dataset
- 400 student responses spanning short-answer questions (1 to 4 marks) and extended responses (6 marks or more).
- Eight core topics from the Singapore-Cambridge GCE O-Level Chemistry syllabus.
Marking evaluation
Each student submission was scored independently by two parties:
- Human markers. Experienced O-Level and GCSE chemistry teachers marking against the official exam rubric.
- Ren's AI engine. Ingesting the same official answer key and scoring criteria.
We treated the consensus teacher scores as our baseline and measured Ren's agreement against them. This setup measures alignment with human grading conventions. When a discrepancy occurs, it can stem from AI error, human variance, or ambiguities in the underlying rubric.
Overall results
| Metric | Ren score |
|---|---|
| Exact mark agreement | 82.4% |
| Within 1 mark | 96.1% |
| Cohen's Kappa (inter-rater agreement) | 0.79 |
| Mean absolute error | 0.31 marks |
For comparison, two independent human markers typically achieve exact agreement between 80% and 85% on these same papers. Ren's score suggestions fall directly within the range of normal human inter-marker variation.
Performance by question type
| Question type | Exact agreement | Within 1 mark |
|---|---|---|
| Short answer (1 to 2 marks) | 91.2% | 99.1% |
| Medium response (3 to 4 marks) | 79.8% | 95.3% |
| Extended response (6 marks or more) | 68.4% | 89.2% |
Agreement is highest on structured questions with clear factual criteria. As responses become longer and more open-ended, agreement drops for both the AI and human markers.
Key findings
Where the AI excels
- Factual recall verification. The engine reliably detects whether specific chemical nomenclature, state symbols, and equations are present.
- Structural compliance. The model consistently tracks whether multi-part responses follow required sequences, such as "state, explain, and illustrate" formats.
- Fatigue-free consistency. The model applies identical rigor to the 500th script as it does to the first, avoiding the late-night grading fatigue that affects human marking panels.
Where human review is essential
- Borderline judgments. Scripts that sit precisely between two grade bands require a teacher's pedagogical context.
- Unconventional valid reasoning. Students occasionally solve a problem correctly using an alternative method not explicitly listed in the answer key. A subject teacher must verify that the reasoning is chemically sound.
- Messy scans and illegible handwriting. On physical paper scans, heavily smudged pencil marks or ambiguous chemical formulas still require human clarification.
Iterating on our models
We re-run these benchmark suites every quarter across updated model weights and revised prompt architectures. Our engineering focus centers on:
- Expanding our evaluation datasets with real student work from partner schools.
- Refining how the model handles strict negative marking constraints.
- Working directly with curriculum leads to handle ambiguous edge cases.
Running a benchmark for your department
Before deploying in a new school, we run a calibration benchmark using your department's past papers and rubrics to prove performance on your specific exams.
If you would like to run a benchmark trial with your school's past papers, contact our team.