Ending the Arms Race and Embracing an Integrity Infrastructure

Share via:

Rethinking High-Stakes Exams in an Age of Technological Assistance

With assessments submitted and marking and ratification processes about to begin, discussions about cheating and integrity will again take centre stage. 

A recent paper produced by Cornell and UC Berkeley provided large-scale evidence on the use of Generative AI to cheat in assessments, adding that it was now part of the fabric of student work. 

Whilst the paper strengthened the evidence, it offered nothing new, in fact, had a group of us gathered around a whiteboard I suspect that we would have reached the same conclusions within thirty minutes. 

However, the scale of the research, almost 100,000 students, puts substance behind the anecdotes. For example: 

  • Around 37% of students use generative AI to aid their work at least monthly, rising to more than half in subjects like mathematics and business, and 62% in computer science. 
  • At least 9% of students admit to using AI to cheat. Daily AI users are three times more likely to cheat as those who use it less frequently. 
  • Misuse is not evenly distributed, AI-assisted cheating rates are more than three times higher in disciplines such as economics and journalism than in biology, reinforcing that both disciplines and assessment methods matter as much as the tools themselves.

The conclusion of the report was simple: generative AI poses a problem for assessment validity and, by extension, the credibility of what we do, and “assessment reform is necessary and urgent.”

That conclusion is difficult to argue with. 

And, if we accept that there is an infrastructure for cheating, the only credible long-term answer is to build a parallel infrastructure for integrity.

A Challenge as Old as Time 

 Long before large language models, candidates were devising ingenious ways of cheating.  

I often cite the example of the civil service examinations in Imperial China. 

These were gruelling multi-day, multi-disciplinary assessments, beyond anything that we deliver today, and, for centuries, they determined the futures of those who sat them. 

Whilst most candidates embarked on months of rigorous study and perhaps the occasional prayer to Kui Xing, the primary God of Exams, other candidates and officials engaged in an incredible range of tactics including smuggling, bribery and impersonation, in an attempt to cheat.  

One of the most inventive cheats was the creation of miniature books of Confucian classics, small enough to fit into a matchbox or be sewn into the lining of clothing, with characters so small that a single grain of rice could cover several of them. 

Authorities were not oblivious to what was going on.  

Candidates were searched before entering the exam hall, had their food and drink examined, their writing materials inspected. 

As cheating evolved, exams were often invigilated by armed imperial guards, and in some era’s officials who failed to prevent cheating faced exile or even execution. 

Despite these enforcement policies, the practice continued, and the methods became more and more inventive. 

Whenever we create a high-stakes gateway to opportunity, people will cheat.

We respond with controls.

Candidates adapt.

So, we introduce new controls.

And the cycle continues.

It is an arms race that we will never win.

The Ambient Cheating Infrastructure 

Over the last few months, I have spoken with assessment providers across the world to gain their insights into the challenges that we face today. 

What has changed is not the existence of cheating but its accessibility.

Candidates no longer need to invent sophisticated schemes for themselves, they now operate inside an ambient cheating infrastructure: a combination of devices, commercial services and digital tools that make assistance constantly available.

Instead of isolated acts, we are now dealing with a multi-layered, partly commercial ecosystem that surrounds high-stakes assessment.

Understanding the Cheating Ecosystem 

Three layers combine to create the current climate:

The Device Layer: Invisible Assistance

Smart glasses, watches, wearable tech and micro-devices offer discreet channels to move information in and out of exam halls. 

Candidates can easily use camera glasses to capture exam questions, transmit them to an accomplice and receive answers via smartwatches or hidden earpieces.  

AI-enabled glasses allow candidates to query AI platforms and display answers within the lens; in one reported case, students using them achieved average scores of 92.5 compared to a class average of 72. 

It is also now cheap to buy a “cheat watch” which stores documents and displays notes in plain text, where the screen looks blank until viewed through special lenses. 

Regulators report that mobile phones and smart devices account for around 44% of reported exam malpractice, prompting some assessment providers to introduce blanket bans on all watches. 

The Service Layer: Commercial Cheating  

Beyond the tech sits a global market in contract cheating, where candidates outsource assignments or exam tasks to third parties, from bespoke essay writing to “ghost-sitting”.  

Several recent studies have concluded that commercial cheating is increasing. 

What is less well known is that these studies also reveal a significant rise in associated harms: the sharing of personal data, blackmail, and long-term reputation damage to professionals where qualifications are obtained without the necessary underlying skills. 

The Social Layer: Sharing Answers at Scale

Multiple platforms now exist which allow assessment information, worked solutions and even live exam content to be captured, stored and traded.

Incidents of assessment papers being leaked via social media platform, are also increasing, with high-stakes assessments recently leaked in Asia, UK and mainland Europe. The Indian Government recently imposed a temporary restriction on Telegram ahead of the NEET re-examination, arguing that the platform had become a vehicle for organised exam fraud and paper leaks.

A Surveillance State is Not the Solution 

We have built exam halls. 

Deployed invigilators. 

Introduced proctoring.

Installed AI detectors.

Yet the fundamental pattern has never changed.

Why?

Because cheating is not a technology problem.

It is a systems problem.

Technology did not create this dilemma. It simply accelerated and expanded an ecosystem that already existed.

The line between legitimate collaboration and organised cheating has blurred, particularly when assessment designs have not evolved.  

During my recent conversations, I have heard one response ring loud:  

 “We need to build more controls; we need to get tougher”. 

Some suggest that means stricter device bans, more robust proctoring, more sophisticated plagiarism and AI-detection and the introduction of tougher penalties.   

Some of these measures are appropriate; there are exams where independent performance must be verified, and where controlled conditions are entirely justified.

But, when the stakes are highest, or where there is a misalignment between assessed tasks and real-world competence, there will always be an incentive to cheat, so we risk remaining trapped in that arms race where we have to run faster and further to stand still. 

There is another way. 

Returning briefly to imperial China, despite the harshest of penalties and the ultimate in micromanaged procedures, cheating persisted for centuries, especially among candidates who had the resources or the connections.

Today, technology has amplified the problem, but a thousand years of assessment history suggests that enforcement alone has never been enough.

Not only will motivated candidates find a way to beat the system, false positives and restrictive regimes are also harmful, often adding additional barriers, to the wider candidate population. 

A sustainable response must focus less on catching malpractice and more on a genuine, principled redesign of our assessment delivery.  

Integrity – A Shared Responsibility

Integrity is not something that we can impose on students. It never has been. 

It is something that students, staff and institutions collectively create.

Integrity works best when it is treated as a shared responsibility – not a disciplinary process.

Rather than framing assessment integrity as a list of forbidden behaviours, it should be framed as a fundamental part of what it means to belong to a discipline or profession. 

If we only discuss integrity when someone breaks the rules, it becomes a compliance exercise. When we embed it into our teaching, assessment design and professional expectations, it becomes part of the culture.

To achieve that takes much more than asking a student to sign a declaration. 

It means a fundamental change to the culture within which candidates and staff operate. 

It means discussing integrity openly throughout the learning journey. 

It means revisiting expectations before major assessments. 

It means creating deliberate moments where students actively affirm their commitment to academic standards.

Research consistently shows that students are influenced by their peers. 

When they believe that cheating is widespread, even tolerated, misconduct becomes easier to justify.

Conversely, when integrity is visible, expectations are clear and staff actively model the behaviours they expect, integrity becomes the norm and misconduct reduces. 

It is impossible to over emphasise the importance of culture in this debate, because people rarely make ethical decisions in isolation

The Design Dimension 

Design also plays a critical role in the development of an integrity infrastructure. 

This means that it is time to move away from single, high-stakes, recall-heavy exams towards more authentic, higher-order tasks that are harder to outsource. For example: 

  • Design tasks that require explanation, critique, personal or context-specific application as opposed to simple content reproduction. 
  • Structure programmes around sequences of lower-stakes assessments with feedback, reducing the perceived need to cheat the system in “all-or-nothing” exams.
  • Incorporate process evidence like learning journals, code repositories, vivas, or videos, where we focus as heavily on how the candidate worked as we do on the final deliverable. 
  • Develop assessment content that explicitly includes AI, asking students to document how they used it, to critique AI-generated outputs, or justify why they chose not to use them. 

This transforms AI from a hidden accomplice to a visible part of the learning process, which candidates can engage with critically and ethically.

Turning Theory into Cultural Practice 

Of course, recognising the importance of integrity culture is one thing. Building it is another.

If integrity is to become something that is actively cultivated rather than simply enforced, organisations need to translate these principles into day-to-day practice.

  • Articulate the purpose: Provide clear, explicit guidance at both a programme and module level. What is the assessment for? What kinds of knowledge, judgement and behaviours need to be evidenced? Where is independent performance essential?
  • Manage controlled environments: There will always be a place for controlled, proctored, device-free environments, but use them only where they are essential, when the construct requires that candidates can perform independently. 
  • Design “Assistance-Aware” assessments:  Accept that, in the real world, work will be conducted with some form of digital tool. Assume that candidates have access to AI, notes, peers and the internet, and design assessments that generate meaningful evidence under those conditions. 
  • Co-Design integrity commitments: Involve students in designing and maintaining integrity and AI policies. Provide clear examples of acceptable collaboration, and discuss the detail, as part of class activity. 
  • Support and encourage staff:  This is often overlooked. Change requires that we  develop new skillsets, and to deliver it effectively we must invest in staff development, whilst providing support and governance that allows them to embrace the change confidently. 

A Final Thought 

Every generation believes it faces the cheating threat that will destroy assessment.

For Imperial China it was hidden manuscripts. For universities it was essay mills.

Today it is generative AI.

Yet over thousands of years of examinations, the story has remained consistent.

Perhaps the problem is not that our controls are too weak, but that we have spent decades, centuries, trying to solve the wrong problem.

The technologies change, but the arms race continues. 

Candidates discover a new advantage – Institutions develop a new control.

Candidates adapt. 

And the cycle repeats.

AI is not the first challenge we have faced as assessors, and it will not be the last. 

But it has taken us to a junction where we have a choice about what happens next. 

We can spend the next decade accelerating the arms race, or we can change the battlefield.

Integrity will not be secured through increasingly complex layers of surveillance.

It will be secured by creating the strongest integrity infrastructure: assessments that remain meaningful in the presence of technology, cultures that value professional responsibility, and systems designed around trust, authenticity and defensible evidence.

The future of assessment integrity will not be secured by trying to catch every attempt to cheat.

It will be secured by making cheating a less attractive, less effective and less necessary strategy in the first place.

Share via:
Topics
Picture of Robert Burns
Robert Burns
Robbie works with assessment organisations on the practical realities of delivering high-stakes assessments, with a particular focus on operational resilience, assessment integrity and candidate experience in digital delivery models.
Would you like to receive Cirrus news directly in your inbox?
More posts in Better Assessments
Better Assessments

Choosing the Right Proctoring Model for Your Assessment

Everyone wants to know the best way to proctor an exam. There isn’t one. Every qualification carries its own purpose, stakes, candidates and constraints, so the useful question is not which model is best, but which model fits, and whether you can explain why you chose it.

Read More »
Better Assessments

How Secure Is Secure Enough?

There is no universal definition of secure enough. The answer depends on what the task exposes, what is at stake if the result is wrong, and the scale you are working at. The third article in The Proctoring Question sets out how to match the control to the risk, and what it costs when you get it wrong.

Read More »

Know exactly where your assessment operation stands

Your free personalised report in 4 minutes

Answer 12 questions across strategy, delivery, design and data and get a clear, personalised breakdown of where your assessment operation is strong and where to focus next.