When AI Thinks for Us: Why Assessing Critical Thinking Has Never Been More Urgent

Share via:

The moment a trainee accountant consults an AI tool to interpret a complex set of financial statements, instead of working through the numbers themselves, something essential shifts in their professional development. Rather than building the analytical mindset and resilience needed for high-stakes decision making in the finance sector, they may be outsourcing the mental processes that underpin sound professional judgement.

This very tension surfaced in our recent webinar roundtable on “The Financial Testing Challenge”, where assessment leaders and industry experts discussed how AI is now woven into every aspect of both learning and examination in accountancy and finance. As artificial intelligence embeds itself into professional training and increasingly into assessment environments, we face a pressing challenge: how do we make sure certified finance professionals can still think critically and independently, especially when AI-generated answers might be incomplete, flawed or even misleading?

The Cognitive Offloading Reality

Recent work by Professor Michael Gerlich at SBS Swiss Business School shows a clear pattern: heavy users of AI tools score lower on critical-thinking tests, and the drop is sharpest among younger professionals. Psychologists call this ‘cognitive offloading’, outsourcing mental effort to a machine rather than doing the reasoning yourself.

Fresh evidence makes the risk hard to ignore. A Microsoft study of 319 knowledge-workers found that the more confidence people place in GenAI, the less they challenge its output. An EEG study at MIT recorded lower brain engagement when participants wrote essays with ChatGPT compared with Google search or no tool at all, and their performance declined over time. In the legal world, judges have already fined solicitors for submitting filings packed with invented case law generated by ChatGPT, see the Avianca Airlines brief and the growing list of more than 120 similar incidents tracked this year alone. Finance is no safer: a peer-reviewed study of 21 real-life money-management scenarios showed ChatGPT offering advice that missed basic tax wrappers, made arithmetic errors and overlooked legal duties (Journal of Risk and Financial Management).

The research base is still young and many samples are modest, yet the findings keep pointing in the same direction: AI can assist, but it also invites complacency. Maintaining rigorous critical-thinking habits is therefore not a nice-to-have; it is the safety net that catches the mistakes a fluent chatbot will not.

The Reliability Gap

The urgency of this challenge becomes apparent when we consider AI’s reliability limitations. Recent evaluations of leading AI models show that even advanced systems like OpenAI-o1, Claude 3.7 Sonnet, and DeepSeek R1 failed more than 70% of the time when presented with novel reasoning tasks. For professional contexts where accuracy is paramount, medical diagnosis, financial analysis, maritime navigation, this error rate is unacceptable.

This reliability gap has serious implications across sectors. In healthcare, large language models achieved only 61% accuracy on medical examinations, a failure rate that would be catastrophic in clinical practice. Similarly, in finance, AI systems struggle with novel regulatory scenarios that require contextual judgment.

The reliability concerns extend beyond simple accuracy. Research on AI-driven assessments highlights how AI systems can perpetuate biases embedded in training data, leading to discriminatory outcomes in professional evaluation. Without robust critical thinking skills, professionals may accept these outputs uncritically, potentially making decisions that harm both individuals and organisations.

Given these reliability concerns, professional bodies face a critical question: how can certification processes adapt to ensure competence in an AI-enhanced world?

Professional Certification in the AI Era

Professional bodies are responding by weaving AI literacy into their qualifications. For example, some accountancy programmes now require future accountants to understand not only how to use AI tools, but also when and why to challenge their results and identify possible shortcomings. This adaptation is essential: as AI automates more technical tasks, the comparative value of human critical thinking and independent judgment only increases.

However, these advances also raise a vital question for assessment: how can we validate that professionals retain the ability to think and decide independently, even as their learning and daily practice become increasingly AI-enabled? 

This issue is especially pronounced in safety-critical fields. In maritime and aviation, for instance, AI-driven training and decision support systems are now standard, helping practitioners interpret data and make preliminary decisions. Yet both industries (and regulators like the FAA) stress that the final responsibility still rests with the human professional, who must be ready to override or question AI guidance, particularly in unpredictable situations.

Similarly, in medicine, healthcare certification bodies and specialist programmes focus not just on AI operations, but on empowering clinicians to critically appraise and, when necessary, countermand algorithmic recommendations for patient care.

The Assessment Imperative

Traditional professional exams have a problem. Most of them test exactly what AI does best: recalling information, applying standard procedures and generating polished responses. When a finance candidate can ask ChatGPT to analyse a balance sheet or an accounting trainee can get AI to explain complex regulatory requirements, what exactly are we measuring anymore?

This isn’t just a theoretical concern. Consider a typical CPA case study about financial risk assessment. A candidate might spend hours researching market conditions, calculating ratios and crafting their analysis. But with AI, that same analysis could be generated in minutes. The traditional assessment approach suddenly becomes meaningless because it’s testing skills that are now automated.

The solution isn’t to ban AI or pretend it doesn’t exist. Instead, we need assessments that focus on what humans do that AI still can’t: making judgements in messy, real-world situations where there isn’t a clear right answer.

Recent research from universities redesigning their assessments shows this shift is already happening. Rather than asking candidates to produce analysis from scratch, these new approaches present candidates with AI-generated outputs and ask them to critique, improve, or adapt them for specific contexts. The assessment becomes about evaluation and judgment, not just production.

Think about what this looks like in practice. Instead of asking an accounting candidate to calculate financial ratios, you present them with AI-generated ratio analysis and ask them to evaluate whether the conclusions are appropriate for a family-owned business considering succession planning. Suddenly, the test isn’t about mathematical ability but about professional judgement, contextual understanding and critical thinking.

Evidence from professional testing organisations supports this approach. Situational Judgement Tests, which present realistic workplace scenarios requiring professional decision-making, have proven effective at measuring competencies that predict job performance better than traditional knowledge tests.

The key insight is this: we need to stop testing what candidates know and start testing how they think. This means assessments that require candidates to navigate competing priorities, weigh ethical considerations, and justify their reasoning process. These are the skills that matter when qualified professionals face real challenges where AI might provide information, but human judgment determines the outcome.

The Path Forward

Shifting our assessments from product to process is not a one-off tweak but a long-term project that calls for four concrete moves.

  1. Put the reasoning on show: Tasks need to expose a candidate’s thinking, not just the final figure or paragraph. Situational Judgement Tests already do this well by asking professionals to work through realistic dilemmas and explain their choices. They have outperformed traditional knowledge tests in predicting on-the-job decision making (OPM guidance). Building similar “show-your-work” steps into finance case studies or clinical scenarios lets examiners see how applicants weigh evidence, spot gaps and challenge AI outputs.
  2. Test judgment in context: Real competence shows when the rules run out. Scenario-rich frameworks such as the MAGE model encourage assessors to present messy, open-ended problems and insist candidates adapt their approach to the specifics (Zaphir et al., 2024). In practice, this might mean giving an AI-generated audit report that looks plausible but contains subtle errors, then asking the candidate to decide whether it is fit for purpose and justify any amendments.
  3. Measure metacognition, not just cognition: AI makes it easy to look smart; metacognition shows whether the professional knows why an answer is trustworthy. Simple prompts such as “Where could this output fail?” or “Which extra data would you check?” help examiners capture the planning, monitoring and self-correction skills that underpin sound judgement (Times Higher Education, 2024).
  4. Keep ethics on the table: Every field now carries AI-specific obligations around privacy, bias and accountability. UNESCO’s global recommendation frames these as part of everyday professional duty (UNESCO, 2021). Embedding short ethics vignettes into exams, for example, asking an accountant how they would handle a model that advantages certain clients, ensures that candidates can recognise and act on ethical red flags.

Taken together, these steps turn assessment into a live demonstration of critical thinking in an AI-saturated workplace. They confirm that a pass mark still signals trustworthy professional judgement, even when the first draft came from a machine.

Conclusion

The evidence is clear: AI tools can enhance professional capability when used appropriately, but they can also undermine the development of critical thinking skills that form the foundation of professional competence. The challenge for professional bodies and educational institutions is not to resist AI adoption but to ensure that certification processes continue to validate genuine professional capability in an AI-enhanced world.

The stakes are high. In healthcare, finance, maritime operations and other critical fields, professional judgment can mean the difference between success and catastrophe. If we allow cognitive offloading to erode the analytical skills that underpin professional competence, we risk creating a generation of certified professionals who cannot think critically when AI systems fail, provide unreliable information, or encounter scenarios beyond their training data.

The future of professional certification depends on our ability to harness AI’s benefits while maintaining the human judgment that remains irreplaceable in complex, high-stakes professional contexts. This requires not just new assessment methods but a fundamental rethinking of what professional competence means in an age where artificial intelligence is everywhere.

This is the first post in our four-part series “Assessing Critical Thinking in the Age of AI.” Subscribe to our newsletter to receive the next three posts directly in your inbox, plus insights on the future of professional assessment.

Share via:
Topics
Picture of Dani van Weert
Dani van Weert
Cirrus' Marketing Manager Dani is interested in how we can make technological advances work for us, to improve education and make it more accessible.
Would you like to receive Cirrus news directly in your inbox?
More posts in Better Assessments
Better Assessments

Choosing the Right Proctoring Model for Your Assessment

Everyone wants to know the best way to proctor an exam. There isn’t one. Every qualification carries its own purpose, stakes, candidates and constraints, so the useful question is not which model is best, but which model fits, and whether you can explain why you chose it.

Read More »
Better Assessments

How Secure Is Secure Enough?

There is no universal definition of secure enough. The answer depends on what the task exposes, what is at stake if the result is wrong, and the scale you are working at. The third article in The Proctoring Question sets out how to match the control to the risk, and what it costs when you get it wrong.

Read More »

Know exactly where your assessment operation stands

Your free personalised report in 4 minutes

Answer 12 questions across strategy, delivery, design and data and get a clear, personalised breakdown of where your assessment operation is strong and where to focus next.