If AI Can Pass Your Assessment, Your Assessment Is Broken

Share via:

I spent a few days last week at the Learning Technologies conference in London, speaking with employers, assessors and learning leaders about how their assessment models are evolving.

When I set out for London town, it was with the intention not to speak about AI, which seemed to be the tagline that illuminated every seminar and every booth. I arrived determined to focus on the distinction between learning platforms and assessment platforms, and how they differ in terms of measuring the capability of employees. 

But, in the end, almost every conversation that I had, both during and after the conference, pulled strongly towards the same concern:

“How do we trust assessment outcomes in a world where AI can generate answers instantly?”

AI hasn’t broken assessment; it has simply exposed that many of our assessments were never measuring the right things in the first place. If an assessment relies heavily on recall, predictable responses, or generic questions, AI will perform well. It can mask gaps in understanding, generate credible answers, and create the illusion of competence.

But the opposite is also true.

Well-designed assessments don’t allow AI to hide those gaps in understanding, in fact, they expose them. Because where an assessment is well designed, the differentiator is no longer who can produce an answer… it is who can interpret, evaluate, adapt, and apply it.

Not a Level Battlefield   

Within the sector, most of the conversations that I hear still centre on protecting integrity. Protecting reputation. Protecting the level playing field. Preventing misconduct. Preventing learners from using the tools that they will use in the real world. 

For someone who has spent almost my entire working life passionately defending the integrity of both assessments and reputations all of that is understandable. 

But. 

If our strategy is simply to make existing assessments harder to game, to implement stricter controls and develop additional monitoring techniques, we are stepping into a battle we can never win. 

AI will continue to improve. It will become more embedded in everyday workflows. It will generate better, more human responses, and it will do so even faster than it does today. And our learners will continue to find new ways to use it. 

We cannot outpace this level of change simply by making weak questions more difficult. 

What is AI Really Exposing 

In reality, the world has always required something quite different. 

The real world demands the ability to analyse and interpret information, to understand context, to make decisions under constraints and to apply knowledge in unfamiliar situations. 

AI hasn’t changed that.

If we are still relying heavily on recall tasks, predictable essays, flat objective tests or generic questions, AI is merely highlighting how weak those formats are when being used to gather evidence of competence. 

From Answers to Outcomes 

We need to reframe the questions and instead of asking whether a learner can reproduce information simply using their powers of recall, we need to ask whether they can produce the right outcome in a realistic environment where tools like AI exist. 

Because, in most workplaces, AI has not been banned, Google has not been disabled and employees are not judged on memory. They are judged on the quality of their decisions, the strength of their reasoning and their ability to analyse and apply information ethically and effectively. 

Integrity AND Relevance 

There is a tendency within the sector to frame this as a defensive challenge: How do we protect the integrity of our assessments?

But there is a much more powerful opportunity here.

We can protect integrity and improve the relevance of assessment at the same time.

We can design assessments that reflect real world conditions, that prepare learners more effectively for work. Assessments that provide more meaningful evidence of capability and which increase confidence for employers 

This isn’t about lowering standards, it is about raising them, by aligning assessment more closely with the performance outcomes that we actually need to see. 

AI As an Amplifier

It came up again and again at Learning Technologies: AI does not affect all learners equally.

In many cases, it amplifies underlying capability.

Learners with a shallow understanding tend to rely heavily on AI output, but they typically struggle to explain, adapt, or defend that output when challenged. 

Learners with stronger understanding use AI differently, they apply judgement, refine outputs, challenge assumptions, and adapt responses to fit the context. 

As a marker, that difference becomes visible quickly, not in the answer itself, but in what the learner can do with it. This doesn’t mean that AI automatically reveals ability, but it does mean that well-designed assessments can use AI to make differences in capability more visible, not less. 

This is Something We Can’t Police Our Way Out Of 

Remote proctoring, detection tools, and tighter controls and policies all have a role to play. But they don’t solve the underlying issue. If an assessment can be easily completed with AI, a determined candidate will always find a way.

So, we need to stop thinking about how we can stop learners from using AI and start thinking about how we design assessments where using AI well requires genuine competence 

That’s a fundamentally different challenge, but a far more valuable one to solve.

A best-in-class model must combine thoughtful design, validation, trust, contextualisation and deep analysis to ensure that evidence of competence is generated across the full assessment suite rather than inferred from a single response. 

A Moment of Opportunity

AI is already embedded in many workplaces. Within a very short time, it will be almost impossible to find roles where it does not play an integral role. 

Assessment models that ignore this reality risk becoming less authentic, less credible and less relevant. 

But for organisations willing to reimagine their approach, this moment offers something more positive, an opportunity to redesign assessment so that it measures what actually matters.

Not just knowledge, but judgement, critical thinking, interpretation and the ability to perform effectively in a modern, technologically enabled environment. 

A Final Thought 

At the end of a busy week, as I left the sunshine of London behind to head home to grey skies and rain, one question kept returning:

What would a better assessment look like if we designed it today, knowing that AI exists?

At first, it felt like a difficult question.

But the more I thought about it, the more exciting it became, because it grants us permission to move beyond defensive thinking and towards something much more ambitious.

It offers us an opportunity to rethink what we value, what counts as evidence and how closely our assessments reflect the real world 

If we take that opportunity, AI will not become the force that undermines assessment – It will become the catalyst that pushes it towards something more meaningful, more demanding, and, ultimately, more useful.

In Part II, we will explore what this means in practice, and how assessment design can evolve to ensure that capability remains visible, even in an AI-enabled world.

Share via:
Topics
Picture of Robert Burns
Robert Burns
Robbie works with assessment organisations on the practical realities of delivering high-stakes assessments, with a particular focus on operational resilience, assessment integrity and candidate experience in digital delivery models.
Would you like to receive Cirrus news directly in your inbox?
More posts in Better Assessments
Better Assessments

Choosing the Right Proctoring Model for Your Assessment

Everyone wants to know the best way to proctor an exam. There isn’t one. Every qualification carries its own purpose, stakes, candidates and constraints, so the useful question is not which model is best, but which model fits, and whether you can explain why you chose it.

Read More »
Better Assessments

How Secure Is Secure Enough?

There is no universal definition of secure enough. The answer depends on what the task exposes, what is at stake if the result is wrong, and the scale you are working at. The third article in The Proctoring Question sets out how to match the control to the risk, and what it costs when you get it wrong.

Read More »

Know exactly where your assessment operation stands

Your free personalised report in 4 minutes

Answer 12 questions across strategy, delivery, design and data and get a clear, personalised breakdown of where your assessment operation is strong and where to focus next.