!IS1108 Ethics in Computing cohort report
Grading at university scale usually forces an uncomfortable compromise between depth and turnaround time. You can write thorough, individualized feedback for every submission, or you can return grades to a 400-student cohort within a week. Large courses almost never manage both.
Last semester, we partnered with NUS School of Computing to run Ren across IS1108 (Ethics in Computing), marking individual reports for all 426 enrolled students.
Grounding feedback in course frameworks
Instead of using a generic grading prompt, we built a dedicated context layer tailored to the IS1108 curriculum. We ingested the course's FISh framework, official rubrics, and specific ethical reasoning guidelines so the AI's suggestions reflected the exact concepts taught in lectures.
Every student received specific margin commentary referencing their ethical arguments and structural choices. The teaching assistants reviewed and adjusted these suggestions directly in the platform before releasing them.
!Teaching assistants mark IS1108 reports in Ren
Teaching assistants worked in Ren's split-view workspace, reading the student's submission side by side with AI-generated notes, tagged curriculum concepts, and rubric criteria. They could edit feedback or override marks in a click before publishing results.
Why grading consistency matters
Marking at scale is not just about turnaround time, it is about fairness across a large grading panel.
When nineteen different teaching assistants grade hundreds of open-ended papers, individual interpretations of the rubric inevitably diverge. In our calibration baseline, scores awarded by 19 independent markers evaluating the exact same report varied by up to 4.5 marks on a 20-mark scale.
Because Ren evaluated every paper against a unified rubric baseline, the system reduced that variation dramatically. Across 390 audited reports, Ren's initial score suggestions differed from final TA-approved grades by an average of just 0.40 marks.
The numbers
We benchmarked Ren's suggested scores against the final grades approved by teaching assistants on a 20-mark scale:
| Metric | Result |
|---|---|
| Mean absolute error | 0.40 marks |
| Within 1 mark | 95% |
| Within 2 marks | 100% |
| Human variance on same paper (19 markers) | 4.5 marks |
Turnaround speed improved just as significantly. Grading an ethical analysis report manually took teaching assistants roughly 33 minutes per paper. Reviewing, refining, and approving Ren's generated feedback took 3.3 minutes per paper. That 10× speedup saved roughly 250 hours of TA grading time across the cohort.
Uncovering cohort-wide misconceptions
After completing the batch, we delivered cohort-level analytics to Prof Boon Kee Lee and the instructional team. The report highlighted common misconceptions in ethical reasoning frameworks, identified specific case studies where students struggled to substantiate claims, and tracked class-wide distribution curves.
In traditional paper-and-pen marking or disconnected spreadsheets, these diagnostic insights stay scattered across individual grader notes and are rarely aggregated.
Looking ahead
This pilot showed what is possible when software handles the heavy lifting of initial drafting while instructional staff maintain pedagogical control.
We want to thank Prof Boon Kee Lee and the 19 teaching assistants who partnered with us throughout the semester.
We are expanding university pilots for upcoming academic terms. If you want to pilot Ren in your course, reach out to our academic team.