By
Daniel Hickey
|
|
|
I was invited to deliver the keynote address at the
annual Faculty Assessment and Quality Improvement Retreat at Saint Mary’s of
the Woods in May 2026. SMWC is a highly regarded private liberal arts college
on the Indiana-Illinois border in Terre Haute. It was an honor to follow in the
footsteps of Indiana University Bloomington’s Associate English
Professor-turned AI guru Justin Hodgeson, who delivered the 2025 Keynote. Zach
Selby (Assistant Professor of Theology) and colleagues organized a highly
productive retreat, and my two talks were well received. As an assessment
scholar and innovator, I would love to see all colleges and universities pay
such careful attention to assessment.
I participated in an afternoon breakout session where faculty discussed how the college was using Turnitin’s controversial Originality AI detector. This post describes how this session and new research have led me to question my prior opposition to using AI detectors to discourage students' unauthorized use of GenAI in submitted work and take-home assessments.
The bottom line is that I still think we should focus on helping students use GenAI transparently and productively. I also worry about bias and unintended side effects of policing student work. But I came away convinced that judicious use by thoughtful faculty may have positive consequences that critics and skeptics should be aware of
My Experience with Detectors
One of my presentations included the findings of a
report I previously drafted for UNESCO entitled Higher
Education Assessment in the Era of AI and Digital Technologies: Practices and
Recommendations. I presented my findings
at the Assessment Institute Conference at Indiana University Indianapolis in
2025. Despite the 7:00 AM “rise and shine” presentation, the session drew a
large and enthusiastic crowd that included one of the retreat organizers. I am glad the report has had some impact;
because of the mayhem caused by our recent withdrawal from UNESCO (again), it
still has not been published.
Of course, academic integrity featured prominently in
my report. One of the Trends I researched and made a
Recommendation for concerned using remote proctors to help
thwart GenAI-driven cheating (I
recommended “use remote proctors judiciously”). Likewise, one of the Challenges
I studied and made a Recommendation for concerned GenAI detection
programs like Originality and GPTZero.
All detectors have well-documented false positive rates (e.g., Webber-Wulff, et al.,
2023);
some have been shown to be biased against non-native English writers (Liang et al., 2023).
As such, I recommended “Use AI detector judiciously (if at all).” More
importantly, following widespread advice, I recommended never using evidence
from detectors as the sole basis for concluding unauthorized use. Rather,
corroborating evidence (such as from an interview) is necessary to require a
redo or reduce a grade—or worse. Unless conducted immediately following
submission, an interview likely can’t rule out unauthorized use. But interviews
likely can confirm students’ knowledge of their submitted content,
regardless of unauthorized use. The real question for me is whether the risk
of suspicion and a follow-up interview discourage the most unproductive
forms of unauthorized use. This is now
widely acknowledged “cognitive
offloading” (Gerlich, 2025), where students submit GenAI
content that they have not bothered to read or learn from.
Meanwhile, of course, the rise of GenAI means that
most of us need to think very differently about our assignments and
assessments. While I reported extensively on this topic, this post is not about
that, other than a few word at the end about a new project.
What I Learned in the AI Detector Breakout
Session
The breakout session I attended was titled From
Policy to Pedagogy: Navigating the SMWC AI Use Guidelines. It was
skillfully hosted by Shalini Persaud (Assistant
Professor of Science and Math) and Sara Amstutz (Lecturer of
Business and Leadership). Participating instructors and faculty
spoke knowledgeably and sensitively about how they used Originality. Several
explained why they chose not to use it; instead, they encouraged unfettered,
fully acknowledged use. I was impressed by (a) the explanations other
instructors provided for using Originality, (b) how those explanations
interacted with their particular discipline and course, and (c) how they
followed up on suspected unauthorized use.
I came away convinced of four things:
·
The risk of being flagged and interviewed
was reducing unauthorized and unproductive GenAI use.
·
The school's careful and planful rollout
helped create an effective community of practice around the tool and how to
best use it.
·
The rollout and the resulting
community of practice were having an overall positive impact on student learning
of course content
·
The positive impact seemed particularly
strong in fully online asynchronous courses.
However, I was certainly not convinced that
there were no unintended negative consequences. More on that later below.
How this Compares with Indiana University
Bloomington (and many other campuses)
The situation above is very different than the
current situation at my own IU Bloomington. I have been following detectors closely since
I helped draft the Generative AI Task
Force Report for the Indiana University System in
2024. That report reminded IU faculty that they are forbidden from using AI
detectors on student work. But the prohibition is “indirect”: instructors are
forbidden from uploading student work to any unapproved third-party
platform. This prohibition is written into DM-02 Disclosing
Institutional Information to Third Parties.
IU’s
Center for Innovative Teaching and Learning offers some great advice on alternative
responses that they update regularly, and our University
Information Technology Services has stated they will
revisit the decision to not enable Originality. But it appears quite unlikely
that any detection software will ever be approved (or possibly even seriously
considered)
I just stepped down from a two-year term as Co-Chair
of the Bloomington Faculty Council Technology Policy Committee. We tried (and
failed) to advance a more explicit prohibition policy that explained detector
shortcomings and suggested alternative responses to integrity concerns. In our
committee discussions, it became clear that many faculty were using AI detectors, despite the prohibition. But they were
apparently not discussing their use with colleagues or admitting their use to
students, presumably because of the policy prohibition and corresponding sanctions
for violating student privacy.
In the midst of our
discussions, I received an email from a faculty member whose undergraduate son
had his grade on a group project marked down substantially “for obvious
unauthorized AI use.” In a massive
violation of university policy, the instructor did not file a formal academic
misconduct charge. Such allegations are mandatory whenever a student’s grade is
impacted by alleged misconduct. Misconduct hearings let students see and
respond to the evidence of misconduct. Ultimately, the instructor removed the
assertion from the online gradebook and instead asserted that the submitted
work was unsatisfactory—a terrible outcome.
Widely viewed social media
posts on Reddit
and media reports confirm that most IU students know that their instructors are
prohibited from using AI detectors. A growing body of research shows that a
great deal of unauthorized AI use is motivated by the same factors uncovered in
research reviewed in James Lang's excellent 2013 book Cheating
Lessons: Learning from Academic Dishonesty. Multiple anonymous survey
studies of academic misconduct (which presumably under-report such behavior)
show that most undergraduates will cheat if (a) they believe they are being
unfairly disadvantaged by classmate cheating and (b) they are confident they
will not be caught.
What Recent Research Says
I like to think of such students as the “honest
cheaters.” And I worry a lot about them. Undergraduates are busy. Between
heavy course loads, sports, social obligations, and home life, students have many
competing demands on their time. A growing body of new research is confirming
that these factors are driving a surge in unproductive unauthorized GenAI use.
Giray et al. (2026)
found that students who admitted using generative AI to cheat often framed the
behavior as an academic survival strategy, while peer networks helped normalize
the practice and shape expectations about acceptable conduct. Similarly, Fu et
al. (2026) reported that students’ immediate peer groups could create informal
norms of AI use that competed with official institutional rules, particularly
under pressure from deadlines, grades, and examinations (see also Johnston et
al, 2024).
I was unable to locate any causal research showing
that institutional rollout of AI detectors reduces cognitive offloading and
unproductive GenAI use. Here is what I was able to uncover (with some help from
ChatGPT EDU Pro)
Causal evidence is
limited. The systematic review by Salamah (2026)
characterized the evidence that AI-text detection is an effective, mature
response to generative-AI misuse as “limited.” It also found little direct
comparative testing of punitive or surveillance-based approaches. The review’s
search ended in December 2024.
Perceived risk is
associated with lower reported GenAI use. Ortiz-Bonnin and
Blahopoulou (2025) surveyed 468 undergraduates at a Spanish
university. Perceived risks surrounding ChatGPT were associated with lower use
frequency, p = -21, and lower intention to
use it in the future, p = -33. However, this was a
cross-sectional survey at one university; it measured general ChatGPT use
rather than verified misuse and did not study institutional deployment of a
detector. The authors explicitly caution against causal interpretation.
Fear of accusation
appears to put some students off AI. The
UK Higher Education Policy Institute (2025) surveyed over a
thousand UK undergraduates. They found that 53% said that the possibility of
being accused of cheating discouraged them from using AI, and 76% believed
their institution would be able to spot AI use in assessed work. At the same
time, 88% reported using generative AI to help with assessments. Thus,
perceived enforcement pressure exists, but the study does not isolate the
effect of detectors or distinguish productive assistance from
learning-replacing use
The most
detector-specific survey found both deterrence and evasion.
A survey by Copyleaks
(2025) of approximately 1,100 US students reported that 36%
used AI less because of detection concerns. But 37% said they edited AI outputs
to make them less detectable, and 62% had actively tried to avoid detection at
least once. This is vendor-published, self-reported survey evidence rather than
a peer-reviewed or causal evaluation. It suggests that detectors may change behavior,
but some of the changes are concealment rather than reduced misuse.
In summary, I find that current evidence supports a
plausible deterrence association: perceived risk of detection or accusation may
discourage some unproductive unauthorized GenAI use. However, no robust causal
evidence currently demonstrates that institution-wide AI-writing detectors
selectively reduce unauthorized or learning-substituting generative-AI use.
Published surveys also indicate that detector awareness can encourage non-disclosure,
paraphrasing, and active evasion. I find non-disclosure and evasion
particularly worrisome because it works against helping students learn to use
AI productively.
While I am certainly not as pro-detector as Derek
Newton’s prolific Substack The
Cheat Sheet, I certainly am not as
opposed to them as I was a few months ago. I also strongly agree with the last
(of five) propositions from the excellent report from Jason Lodge and
colleagues at the Australian Tertiary Education Quality and Standards Agency
(TEQSA, 2023, p. 6). The report acknowledged that “in many disciplines, there
may be a need to understand and evidence what students are capable of without
AI.”
A convincing evaluation of detectors’ deterrent
effects would need a controlled or staggered institutional rollout,
measurements taken before and after implementation, an appropriate comparison
group, and independent outcomes such as writing-process evidence, oral
verification and unaided learning assessments, and not just the detectors’ own
scores. Such research seems particularly called for in some high-stakes fully
online settings where in-person examinations are not possible and digital
proctoring is not feasible.
I will first close by pointing out that I strongly
encourage acknowledged AI use in my own (graduate-level) courses, and I include
optional GenAI elements in most of my assignments. A previous post showed how students
could use an early version of ChatGPT to write
a decent literature review for my class; I also concluded
that it would have been difficult to do so without still learning the content. That
is definitely not the case with current pro-level models, which can
draft entire literature reviews with accurate references and quotations from a
single prompt.
I will finally
close by acknowledging that online colleagues and I are currently exploring a
“completion-based” response where each assignment includes a significant number
of carefully aligned (a) reflections, (b) open-ended formative assessments, and
(c) summative multiple-choice items. Students must complete them in order to
progress through the course. Our idea is
to “make it easier to learn than to cheat,” but without significantly
increasing instructor workload (the assessments and feedback are automatically
administered by the LMS). Right now, we have more questions than answers, but
it seems promising. Fortunately, GenAI is remarkably proficient at generating
item sets and feedback for instructors and designers to choose from. I hope I
have more to say about that a few months from now.
References
Copyleaks,
Inc. (2025). AI in action: Normalized AI in the classroom. Author. https://copyleaks.com/knowledge-base/form?resource=ai-in-education-ai-in-the-classroom
Elkhatat,
A.M., Elsaid, K. & Almeer, S. Evaluating the efficacy of AI content
detection tools in differentiating between human and AI-generated text. Int
J Educ Integr 19, 17 (2023). https://doi.org/10.1007/s40979-023-00140-5
Gerlich,
M. (2025). AI Tools in society. Impacts on cognitive offloading and the future
of critical thinking. Societies, 15 (1) https://www.mdpi.com/2075-4698/15/1/6
Higher
Education Policy Institute (2025). Student generative AI survey 2025. Author. https://www.hepi.ac.uk/reports/student-generative-ai-survey-2025/
Johnston,
H., Wells, R. F., Shanks, E. M., Boey, T., & Parsons, B. N. (2024). Student
perspectives on the use of generative artificial intelligence technologies in
higher education. International Journal for Educational Integrity, 20,
Article 2. https://doi.org/10.1007/s40979-024-00149-4
Liang,
W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are
biased against non-native English writers. Patterns, 4 (7),
Article 100779. https://doi.org/10.1016/j.patter.2023.100779
Ortiz-Bonnin,
S., & Blahopoulou, J. (2025). Chat or cheat? Academic dishonesty, risk
perceptions, and ChatGPT usage in higher education students. Social
Psychology of Education, 28, Article 13. https://doi.org/10.1007/s11218-025-10080-2
Giray,
L., Jacob, J., Encanto, V., & Mansilungan, C. J. (2026). Cheating writing
with generative AI: Exploring student motivations using the theory of planned
behavior. Journal of Academic Ethics, 24(1), Article 19. https://doi.org/10.1007/s10805-025-09695-z
Fu,
Y., Lin, Y., Wang, J., Tran, S., & Hiniker, A. (2026). “Everyone’s using
it, but no one is allowed to talk about it”: College students’ experiences
navigating the higher education environment in a generative AI world
[Preprint]. arXiv. https://doi.org/10.48550/arXiv.2602.1772
Salamah,
W. (2026). The role of artificial intelligence in detecting and preventing
academic dishonesty in higher education: A systematic review. Frontiers in
Education, 11, https://doi.org/10.3389/feduc.2026.1880283
Tertiary
Education Quality and Standards Agency (TEQSA, 2023, November). Assessment
reform for the age of artificial intelligence. [Report]https://www.teqsa.gov.au/guides-resources/resources/corporate-publications/assessment-reform-age-artificial-intelligence
Weber-Wulff,
D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J.,
Popoola, O., Å igut, P., & Waddington, L. (2023). Testing of detection tools
for AI-generated text. International Journal for Educational Integrity, 19(1),
Article 26. https://doi.org/10.1007/s40979-023-00146-z

No comments:
Post a Comment