Take-home assessments have taken a beating in recent years. Students now have AI tools that can draft an essay in seconds, and a lot of that work reaches the teacher’s desk looking passable.
A 2026 study by UC Berkeley researcher Igor Chirikov shows how strongly AI may be affecting assessment results. After analyzing more than 500,000 grades at a research university in Texas, he found that the share of A grades climbed much faster in courses with substantial writing and coding work than in less AI-exposed courses, with the effect concentrated where homework carried the most weight.
To save the take-home format, many institutions have turned to AI detectors like Turnitin, GPTZero or others, hoping an AI-writing score can identify work generated with tools such as ChatGPT. The question is whether the method actually holds up:
- Is the detection good enough in the first place, or can students simply work around it?
- Even when something is flagged, how do you handle every detected case? Are you prepared for the yes-no disputes, the potential misconduct hearings and the appeals that follow?
- And what if you actually want to allow AI use, but still need to control how it is used?
This article looks at whether using AI detection tools like Turnitin is a sustainable way to protect your assessments and introduces an alternative way without having to go back to pen and paper exams.
How accurate are AI detectors, and can students bypass them
On raw, unedited output, AI detectors do reasonably well. One 2025 evaluation of GPTZero found that the large majority of purely AI-generated essays were correctly identified, often at high confidence. If a student pastes straight from ChatGPT and submits it, a detector will usually catch it.
The picture changes once the text is edited. A peer-reviewed 2026 study from Vrije Universiteit Brussel tested GPTZero, Turnitin, Copyleaks, and Pangram, on 160 academic papers. On humanized AI text, most tools struggled, with Turnitin catching around half and Copyleaks a quarter, while Pangram detected the great majority. In a separate test against paraphrased and humanized output, GPTZero's detection rate on humanized text dropped to roughly 52 percent. Each of these editing techniques can materially reduce detection accuracy, though the impact varies by detector, model and method. The honest reading is not that detection never works, but that it is uneven and that a determined student can easily weaken it.
There is a second error that lands on honest students: detectors sometimes flag human writing as AI. A 2023 study in the journal Patterns found that AI detectors misclassified more than 61 percent of essays by non-native English speakers as AI, while classifying native writing almost perfectly. This risk is serious enough that a score should never be treated as proof on its own.
Many institutions have reached the same conclusion. Vanderbilt University disabled Turnitin's AI detector in 2023, noting that a 1 percent false-positive rate across its submissions could still mean hundreds of students wrongly flagged. Johns Hopkins turned it off over accuracy concerns, and Curtin University switched its AI detection off from January 2026, stating that a detector cannot be the sole or primary basis for a misconduct allegation.
What an AI detector can never tell you: who wrote it
There is one limitation that no improvement in AI detection accuracy will fix. A detector can only ask whether text looks like AI. It cannot tell you who wrote it.
A student who pays someone else to write their assignment, a contract cheater, sets off no AI alarm at all, because the work is genuinely human. Contract cheating is a commercial industry that grew sharply during the pandemic and now bundles generative AI into its services, sold to students online with the same guarantees and convenience as any other e-commerce product. Even a perfect AI detector is blind to it, and it is equally blind to unauthorized collaboration or help from someone else. For anything your institution has to certify, this is the question that matters most, and detection does not touch it.
The hidden cost of acting on every flagged case
Suppose the AI detector is right. The flag is accurate and the student did use AI. The workload of having to act on it is often underestimated.
Because an AI detection score from a detector is not proof. To act on it, you have to build a case, and that means gathering evidence, documenting the process, giving the student a chance to respond, and often taking it to an examination board. Every step takes staff time. Every contested flag can turn into a back-and-forth, then an appeal. And because false positives are real, every accusation carries the risk of being wrong about an honest student.
This is a time-consuming process. It is manageable for a single flagged paper and unworkable across hundreds. An AI detector is cheap to run, but expensive to enforce.
The deeper issue is that all of this happens after the fact. You are reviewing work that has already been submitted, trying to reconstruct what a student did behind a screen you never controlled. You are reacting to cheating rather than preventing it, which means the student always gets to try and you will always have to be suspicious.
A more reliable alternative: controlled in-person assessment without going back to pen and paper
Many institutions are placing less weight on take-home assessments and increasingly want to test the same skills in person, where the conditions can be controlled.
The written exam is often the component that counts heaviest in a course, and it has traditionally been theoretical instead of practical. Many institutions are now using the same setup to assess competencies that used to be demonstrated at home. Students are asked to produce actual work in the exam room, not just answer questions about it.
A writing assignment can be done in the classroom with access to a word processor. The task looks much like the take-home version, except you know the conditions it was produced under.
That only works if you can shape those conditions precisely. Students get access to exactly what the assessment needs, and nothing beyond it:
- Websites, limited to a whitelist you define
- Files, such as a supplied PDF, without enabling access to the student's personal files
- Applications, like Word for an essay, without access to personal files, AI or Grammarly
- AI, deliberately included if that is your pedagogical choice, while ensuring that the handed in work is made by the student and not a contract cheater
Tools such as Schoolyear's Safe Exam Workspace are used to enforce this. It puts a device into a controlled exam mode and enforces the restrictions at system level, so you decide which websites, files and desktop applications are available, while access to AI, personal files, and other potential cheating routes are restricted.
Compared with running AI detection over submitted work, this changes the process of protecting academic integrity:
- Prevention instead of detection: You define what is available during the exam rather than reconstructing afterwards what was used.
- No after-the-fact review: There is no queue of flagged submissions to work through after the assessments are handed in.
- No disputes over a score: There are no discussions about whether a detection score is correct.
- No false positives: Honest students are not put under suspicion by a tool that misreads their writing.

Does this mean take-home assessments will disappear? Almost certainly not, particularly for larger pieces of work that simply can’t be done in a matter of hours. However, institutions are increasingly building in an in-person component to gain more certainty about what a student can actually do.
Conclusion
AI detection can provide a useful signal, particularly on unedited AI output, but not a verdict. This is why detector vendors and a growing number of institutions now say a score should never stand alone.
Taking the assessment into a controlled environment is a more reliable and efficient path for grades that carry weight. Schoolyear’s Safe Exam Workspace lets you run digital, in-person assessments where you decide what resources students can access, whether that is a locked-down word processor or a curated set of tools with AI included. You get prevention rather than detection, and you keep the benefits of digital assessment without betting your academic integrity on an AI detection tool.

.jpg)


