Sunday, September 6, 2026

News

AI Essay Detectors Fail One in Five Times as Universities Lose Control Over Grading

ResearchPatryk Raba
AI Essay Detectors Fail One in Five Times as Universities Lose Control Over Grading
Fot. Alison wood, Wikimedia Commons (CC BY 3.0)

A Daily Maverick analysis finds that AI-text detection tools are wrong at massive scale, prompting universities worldwide to abandon them. A growing chorus argues that a degree no longer confirms a graduate's actual knowledge.

Contents
  1. Detectors that get it wrong
  2. False accusations as a side effect
  3. Radical ideas to break the impasse
  4. Implications for Poland

None of the fourteen independently tested tools for detecting AI-written text achieved accuracy above 80 percent. That is the conclusion of an analysis published on September 6, 2026 by Daily Maverick, which describes how generative AI is dismantling the grading system at universities worldwide.

The author of the Daily Maverick analysis, part of the Crossed Wires series, describes the situation as a fundamental crisis of trust in the university degree. For two hundred years, a diploma meant an institution had verified a graduate's knowledge. Today, with a significant share of written work produced with the help of chatbots, universities can no longer say with certainty whether they are grading a student's skills or a language model's capabilities.

Detectors that get it wrong

The problem goes beyond low accuracy. AI detection systems don't even agree with each other - at the University of Cape Town, two different tools flagged completely different sections of the same text as suspicious. One student was warned that 30 percent of his paper may have been generated by AI, even though he had written it entirely himself.

Researchers cited in the analysis, McKenna and Kramm, explain that detection tools don't identify any digital fingerprint proving AI use. Instead, they look for stylistic patterns, which can falsely flag a student who simply writes precise, well organized prose, or who overuses dashes and words like 'crucial'.

Coursework may no longer be something that can be reliably attributed to a specific student - Quality Assurance Agency (QAA), United Kingdom

False accusations as a side effect

The scale of false positives has led one university after another to simply abandon automated detection. University of Cape Town's rationale for its decision spoke directly of an atmosphere in which students live in constant fear of being wrongly accused of academic dishonesty. Access to alternative tools can also be costly - the competing system GPTZero is priced at 19 dollars per check, which adds up to a real cost for institutions running thousands of papers a year.

Meanwhile, other research undermines a one-sided fix to the problem, namely handing grading over to AI itself. Cambridge researchers who had leading generative models grade hundreds of undergraduate papers found that AI matched the human grade classification in only about half of cases, systematically undervaluing the best papers and overvaluing the weakest ones. Another study, cited by Inside Higher Ed, found that models tend to award essays higher grades than human examiners, favoring writing style over substantive accuracy.

Radical ideas to break the impasse

The Crossed Wires author also describes the most radical proposal circulating among higher-education researchers: reducing the degree to a mere 'certificate of attendance,' and shifting actual knowledge verification to professional bodies - medical boards, bar exams for lawyers, accounting internships or job interviews. The author himself calls the idea a 'tragic capitulation' to the noble tradition of university education, while acknowledging that universities urgently need some kind of solution.

Among the alternatives currently being tested are oral exams, which can verify genuine understanding of material but are costly and don't scale for large cohorts, and a return to multiple-choice tests, which only work well for a narrow range of subjects. The pace of academic research remains a problem too - data from two or three years ago is already practically useless, since language model capabilities are advancing faster than academic publications can keep up.

Implications for Poland

For Polish universities and employers, the problem isn't abstract. National surveys show that more than 90 percent of Polish programmers already use AI tools, while distrust of the quality of generated code is growing at the same time - a similar mechanism of distrust is starting to spread to degrees and diplomas as well. If employers stop treating a thesis grade as a reliable signal of competence, the burden of verifying skills will shift to hiring processes and professional exams, much as the most radical scenario described by Daily Maverick suggests.

Poland's Ministry of National Education (Ministerstwo Edukacji Narodowej) is only beginning to put artificial intelligence at the center of education policy, with decisions on primary and secondary school students coming ahead of regulations for higher education. The experience of universities in Australia, the United Kingdom and South Africa shows that simply banning AI or relying on automated detection doesn't solve the problem - what's needed is a redesign of how universities test knowledge in the first place.

Share: