A reviewer opens clip 214 of the day. Same exam, same alert: candidate looked away from the screen. At nine this morning they would have watched it twice and written a careful note. It is now late afternoon, the queue is still growing, and they dismiss it in eight seconds.
Nothing malicious has happened. Nobody has cut a corner deliberately. And yet a decision that could one day sit in front of an appeals panel, or a regulator, has just been made by someone running on empty.
In the first article of The Proctoring Question we asked whether candidates get a fair deal across delivery channels. This second article turns the camera around, onto the people watching the footage. Because for all the attention we give to AI monitoring, live invigilation and hybrid models, the integrity of remote proctoring is decided somewhere far less visible: the human review queue.
Where integrity is determined
Behind every flagged incident, every suspicious eye movement or reflection in a mirror, sits a human reviewer. Our teams navigate hours of video footage and analytical data, often working under extreme time pressure and balancing a multitude of operational demands.
For the system to maintain genuine integrity, every single clip demands focus, consistency and fairness.
Yet when hundreds of AI-generated “possible incidents” pile up, and the queue for human reviews gets longer, the workload can quickly become overwhelming.
This creates a challenge long recognised in other high stakes environments like aviation and healthcare: cognitive fatigue.
When we begin a review cycle, attention is high and judgement is sharp. Over time, as deadlines loom and pressure builds, consistency begins to erode. Subtle indicators that we might have flagged four hours ago can be missed or dismissed.
This is not a question of capability or diligence. It is simply our human response to sustained cognitive load.
When volume undermines judgement
In proctoring, this phenomenon has a name: flag fatigue. It hits us when automated systems generate more alerts than reviewers can realistically process with attention and care, or when resourcing does not match demand.
When this happens, the risk moves beyond operational inefficiency and towards inconsistency in decision-making.
Similar incidents may be judged differently depending on when they are reviewed and by whom.
For organisations operating in regulated environments, this introduces a serious concern around the defensibility of outcomes. If decisions cannot be shown to be consistent, transparent and repeatable, the integrity of the assessment process itself can be called into question.
There is an Alanis Morissette-like irony here: cutting edge technology designed to enhance fairness risks creating inconsistency, not from technical failures or human negligence, but simply from overwork and cognitive exhaustion.
From operational pressure to institutional risk
This challenge is not purely operational. It reflects broader organisational choices about culture, accountability and decision-making.
I have seen some outstanding proctoring processes, carefully and thoughtfully designed, where risks are well understood and senior leadership takes personal responsibility for the outcomes.
I have also witnessed less mature processes where motions are gone through and boxes are ticked, but where neither the processes nor the oversight are sufficient to guarantee integrity.
It can be easy for us to celebrate automation while underestimating the human labour that sits beneath it, or to celebrate resource savings while not fully understanding the pressure points that we are pushing deep into our processes.
At the same time, compliance requirements, audit expectations and reputational concerns continue to drive broader monitoring. Yet resourcing for post-assessment review often fails to keep pace.
Reviewers working long hours, or balancing multiple tasks, are less likely to escalate genuine issues and less likely to record detailed rationales, and as a result we will have less confidence in their decisions.
All of which means that the processes we put in place to safeguard integrity actually risk compromising it.
Designing for the human layer: five things to get right
None of this is an argument against the technology. Proctoring software has opened up assessment delivery that few of us would have imagined a decade ago, and it keeps getting better.
But better algorithms and glossier dashboards will not fix a review process that exhausts the people running it. That takes deliberate design: processes where integrity is built into every step, and where senior leaders understand, and own, what happens after the flag is raised.
The good news is that none of these fixes are exotic or expensive. In my experience they come down to five things:
- Resource for reality: The volume and complexity of flagged events should directly inform how review teams are structured, whether those teams are in-house or external. Under-resourcing this stage does not reduce cost. It simply transfers risk further down the line.
- Triage intelligently:AI models should be adjusted to suit our specific assessments and our unique environment. They should be there to prioritise patterns of genuine concern rather than flagging every potential deviation. Reducing the volume of false positives allows our teams to focus their attention where it is most needed.
- Set review limits: Air traffic controllers are not permitted to work more than a set number of hours, to reduce the risk of error. Proctoring reviewers should work on the same principles, with clear limits that acknowledge the fact that we all have boundaries beyond which our concentration diminishes.
- Check in on wellbeing: Wellness is a word that is often overused today, and yet I do not think I have ever heard it mentioned in this context. As senior leaders we need to understand what is involved in this process, and appreciate that burnout not only poses health risks to the individuals concerned but directly impacts the reliability of our output. Proactive wellness policies therefore reinforce both ethical practice and operational accuracy.
- Train and calibrate: Whether you are using in-house or external resource teams, consistency always improves when everyone shares the same set of clear and concise definitions. Organising calibration activities, as we do for markers, can be a hugely beneficial practice to ensure consistent reviewer judgement.
It is also important for us to recognise that reviewer workload is not created in isolation. Poorly calibrated AI models, high false-positive rates and assessment designs that generate unnecessary flags all contribute to the volume of review required.
Managing workload effectively requires alignment between technology configuration, assessment design and operational capacity.
Towards smarter integrity management
For remote proctoring to retain the trust of our candidates, our stakeholders and our regulators, it must evolve from a purely technical solution to one where human capacity and technological capability are well balanced and complement each other.
The integrity of our outcomes is not determined by how many behaviours can be detected, but by how reliably those detections are interpreted and how consistently judgement is applied.
Technology will continue to advance. But we cannot depend on technology alone. The integrity of what we do will always rely on the human layer: on the capacity, consistency and wellbeing of those making the final decisions.
If we are serious about protecting the integrity of our assessments, it is imperative that we design our processes around that human layer, for their performance, their workload and their wellbeing.
This is the second article in The Proctoring Question, our series on the tensions shaping remote proctoring. Missed the first? Read When choice becomes risk: fairness in multi-channel exam delivery. Next in the series: Security in an age of organised cheating.


