Can generative AI grade university essays accurately? Not reliably enough to replace human judgement, according to a new Cardiff University-led study. When researchers asked two versions of ChatGPT to mark 50 undergraduate bioscience essays, the scores varied substantially and often compressed very different pieces of work toward the middle of the scale.
What did the researchers test?
The team, involving researchers at Cardiff University and the University of Melbourne, used 50 real undergraduate bioscience essays. Each essay was assessed against seven criteria under four prompting conditions, allowing the researchers to compare AI-generated scores with marks awarded by human assessors.
This matters because a system can appear accurate when its average result is close to a human average while still being unreliable for individual students. The study looked beyond headline averages to examine how marks changed essay by essay and criterion by criterion.
What were the main findings?
- AI marks varied considerably across models and prompting approaches.
- In all but one comparison, the AI awarded a higher average mark than human assessors.
- The largest difference between mean marks was 16.1 percentage points.
- For an individual essay, the largest gap reached 40 percentage points.
- High-scoring work tended to be marked down, while lower-scoring work tended to be marked up.
That last pattern is especially important. It suggests the models were not simply “generous” or “strict”. Instead, their scores gravitated toward the centre. In assessment terms, that can flatten meaningful distinctions between excellent, adequate and weak work.
Why can an average score be misleading?
Imagine a marking tool gives one student 20 points too many and another 20 points too few. The average error is zero, but both students receive the wrong outcome. A similar problem arises when an AI marker produces a class average close to a human marker while making large errors on particular submissions.
Universities therefore need evidence about agreement at the individual level, not just similar overall averages. Decisions about progression, awards and feedback affect individual students, so consistency across the cohort is not enough.
Does this mean AI has no role in assessment?
No. The study addresses automated grading, not every possible educational use of AI. Generative systems may help staff draft example questions, organise feedback themes or identify passages that deserve closer attention. They may also give students low-stakes opportunities to practise, provided the output is clearly presented as provisional rather than authoritative.
But the evidence does argue against treating a general-purpose chatbot as a dependable substitute for an academic marker. Prompt wording, model version, subject context and the structure of a rubric can all affect the result. A tool that changes its judgement when those conditions change is difficult to defend in a high-stakes process.
What should universities do now?
Institutions considering AI-assisted marking should validate systems on their own subjects and assessment types, publish how the tools are used, keep a qualified human responsible for final decisions and provide a route for students to challenge errors. They should also examine whether performance differs across writing styles or student groups rather than assuming a single accuracy figure applies to everyone.
The Cardiff-led study used student work with consent, an important detail for a sector also navigating questions about privacy, intellectual property and the reuse of submitted assignments. Any deployment would need safeguards for those issues as well as evidence of marking quality.
What is the practical takeaway?
Generative AI can produce a plausible-looking mark, but plausibility is not the same as dependable assessment. The new results show why institutions should test individual outcomes, look for score compression and retain meaningful human oversight before AI-generated grades affect a student’s record.
Source note: This article is based on Cardiff University’s report of the peer-reviewed research, published on 24 August 2026. Read the Cardiff University summary.

