AI Essay Marking Today: Progress and Solutions to Ongoing Challenges

Share via:

Essays are invaluable for assessing students’ higher-order thinking skills, problem-solving, and ability to apply knowledge. Yet they come with a high cost: marking essays is time-consuming, expensive, and prone to inconsistencies. AI-powered essay marking offers a promising alternative, combining speed, consistency, and personalised feedback. Blees AI, experts in AI-driven educational solutions, are pioneering advancements in this field, leveraging their experience to tackle the most pressing challenges of automated marking.

In recent trials, Blees AI’s models reached accuracy rates of up to 88.7% – well exceeding the 80% consistency threshold benchmark maintained by the testing organisation. They also deliver results significantly faster than human markers, with marking speeds up to 38 times quicker.

At Future-Proofing Education in Dublin, Sojin Lee from Blees AI shared insights into the potential of AI marking, based on their recent research. These results highlight the current benefits of AI-driven essay marking and point to even greater possibilities as the field advances. While the technology isn’t flawless yet, ongoing research is helping to tackle challenges and make AI marking faster, fairer, and more adaptable for educational institutions.

Here’s a closer look at the current challenges in AI essay marking – and the innovative approaches being used to address them:

1. The need for a large data set

Marking essays with AI isn’t as simple as flipping a switch. To deliver accurate results, AI essay-marking models typically require extensive training data – anywhere from 200 to 400 scripts on familiar topics, and potentially more for obscure topic areas. This demand can be a significant barrier for many institutions, particularly those without large datasets.

To address this, Blees AI is exploring few-shot and zero-shot learning in combination with large language models (LLMs), allowing AI to function effectively with minimal labelled examples. By finetuning a pre-trained large language model and using prompt engineering rather than extensive annotations, they are significantly reducing the time and data needed to automate essay scoring. This approach not only requires limited data to get started, making it practical for organisations with smaller cohorts, but it also saves considerable time that would otherwise go to manual annotations. Additionally, it offers greater scalability, as it removes the need to build individual AI models for each subject area. 

In this approach, prompt engineering makes a real difference. Using the same prompt, Blees AI found that breaking instructions down – giving them one at a time – produced better results than providing all rubric criteria in one go. With this “discrete prompting” approach, the AI zeroes in on each grading aspect individually, leading to more accurate and reliable assessments.

2. Delivering feedback that promotes learning

Delivering feedback that genuinely supports learning is one of the biggest challenges in AI essay marking. Traditional AI feedback often relies on pre-set comments, which can limit personalization.

Blees AI’s approach aims to address this. By using Large Language Models (LLMs), their feedback process is designed to be more personalised and constructive. Instead of relying on standard remarks, the AI identifies specific strengths and areas for improvement in a student’s work, offering targeted suggestions. This feedback isn’t just accurate – it’s also encouraging, giving students a clear sense of what they did well, what they can improve, and how to move forward.

With this individualised, constructive feedback, students are more likely to engage with their results, seeing feedback as part of their learning journey rather than as judgment. This approach helps promote understanding and builds confidence, making feedback a more meaningful part of the educational process.

3. Tackling bias & transparency in AI grading

Bias and transparency are critical issues in AI grading. Bias can enter AI systems through the data they process, often reflecting broader social or linguistic patterns that may disadvantage certain groups. In education, even subtle biases can impact grading fairness, making this a key area of concern. Transparency is equally important, as understanding how AI reaches its decisions is crucial in high-stakes contexts like assessment. Yet with complex large language models (LLMs), achieving clear, explainable outcomes can be challenging.

The need to address these issues has led to entire fields of research, with some organisations focusing exclusively on bias and transparency in AI, and regulatory bodies beginning to push for clearer standards. Blees AI is tackling this through ongoing research in both areas. Over the coming months, they are exploring methods to track AI decision-making, identify where bias might emerge, and develop techniques to improve clarity in grading outcomes.

To address bias specifically, Blees AI collaborates with domain experts to identify high-risk areas for bias, such as regional language variations or cultural differences. They use use prompt engineering (chain of thoughts – Chain of Thought (CoT)) to refine AI outputs and aim to reduce subtle biases that might affect different groups disproportionately. For transparency, Blees AI is exploring LIME (Local Interpretable Model-agnostic Explanations), a technique that helps explain how specific input features—like certain phrases or structures in an essay—contribute to the final grading decision. This allows educators and stakeholders to better understand the factors influencing each score, adding a layer of accountability to AI-driven assessments.

4. Addressing fears of replacement and managing change

“Will AI make my job redundant?” This is a common worry among educators as AI technology, particularly in grading, becomes more prevalent. Many fear that AI might replace parts of their work, creating uncertainty about how it will impact daily responsibilities. This concern is understandable, especially with complex technology that can feel distant or inaccessible. Beyond the fear of replacement, introducing AI into established workflows requires training, adjustment, and adaptation—new tools can disrupt familiar processes, and resistance to change is natural.

Blees AI suggests the following steps for testing organisations implementing AI:

  • Communicate AI’s role as supportive, not substitutive: AI should be positioned as a support for human expertise, not a replacement. By handling repetitive, high-volume tasks, AI can free educators to spend more time on personalised feedback and direct engagement with students. Reinforcing AI as a tool that reduces workload, rather than replacing roles, can help ease concerns.
  • Provide training and resources: Educators need clarity on both the capabilities and limits of AI. Blees AI recommends creating training materials that outline AI’s functions, boundaries, and benefits, helping educators see where human judgment remains essential. Clear training fosters confidence, making the transition to AI-assisted grading smoother.
  • Offer support systems for career growth and transition: For those impacted by AI adoption, career coaching and job transition assistance can be crucial. Providing resources for skill development, career planning, or even support for role changes demonstrates a commitment to educators’ professional growth and long-term value within the organisation.
  • Involve educators early in the process: Early involvement in AI adoption allows educators to feel more in control of the changes. Engaging staff from the planning stage, inviting their input, and incorporating feedback into implementation can create a sense of ownership and reduce resistance.
  • Provide ongoing support: Managing change is not a one-off effort. Continuous support, regular check-ins, and feedback opportunities ensure that educators feel supported long after AI is introduced. This approach addresses immediate concerns while fostering long-term confidence in AI as a valuable educational tool.

AI-automated essay marking offers a compelling solution to longstanding challenges in education – speeding up grading, enhancing consistency, and delivering personalised feedback at scale. While there are still hurdles to overcome, ongoing research is making steady progress. Blees AI’s work in this field highlights how these advancements are bringing us closer to fair, efficient, and adaptable AI assessment tools that support both educators and students.

Interested in learning more about AI essay marking with Cirrus and Blees AI? Get in touch to stay updated on the latest developments.

Share via:
Topics
Picture of Dani van Weert
Dani van Weert
Cirrus' Marketing Manager Dani is interested in how we can make technological advances work for us, to improve education and make it more accessible.
Would you like to receive Cirrus news directly in your inbox?
More posts in Better Assessments
Better Assessments

Choosing the Right Proctoring Model for Your Assessment

Everyone wants to know the best way to proctor an exam. There isn’t one. Every qualification carries its own purpose, stakes, candidates and constraints, so the useful question is not which model is best, but which model fits, and whether you can explain why you chose it.

Read More »
Better Assessments

How Secure Is Secure Enough?

There is no universal definition of secure enough. The answer depends on what the task exposes, what is at stake if the result is wrong, and the scale you are working at. The third article in The Proctoring Question sets out how to match the control to the risk, and what it costs when you get it wrong.

Read More »

Know exactly where your assessment operation stands

Your free personalised report in 4 minutes

Answer 12 questions across strategy, delivery, design and data and get a clear, personalised breakdown of where your assessment operation is strong and where to focus next.