The Most Dangerous Number on Your Dashboard

Share via:

One of the more interesting debates to emerge from this year’s World Cup isn’t about VAR, water breaks or weather interruptions. 

It’s about statistics, or “stat-padding” as it has become known. 

Specifically, whether every goal tells the same story.

For those who have not yet become deeply engrossed in this debate, critics argue that modern tournaments make it easier for elite players to accumulate impressive statistics against significantly weaker opposition. 

Every goal counts. But should every goal count equally when we compare careers?

From a personal perspective, I can see both sides. As a footballer you can only beat the team that is in front of you, and any player who is seeking personal records cannot be blamed for taking the chance to attack those records against weaker teams. 

But I can also see the argument that a player who has scored 10 world cup goals against elite teams has worked harder and performed more consistently, than a player who has scored 10 goals against less powerful footballing nations. 

The more I thought about the debate, the more I realised that it was not really about football. 

What fascinated me wasn’t whether some goals should count for more.

It was what the discussion revealed about data itself.

Numbers rarely settle an argument.

They start one.

That should sound familiar to all of us who operate in the assessment space. 

A pass rate, average score, completion rate, or item performance can look compelling at first glance, but it tells only part of the story if it is not interpreted in context. 

Numbers have an extraordinary ability to shut down curiosity.

Once a figure appears on a dashboard, we stop asking questions.

We mistake measurement for understanding.

Data does not remove the need for judgement. It demands better judgement.

The Seduction of The Headline Stats 

Headline statistics are powerful because they are simple, memorable and easy to compare. They can often reduce complex systems to a single number, and that is what makes them attractive.

In football, goals, assists, and appearances create instant narratives. In assessment, the equivalents are pass rates, score distributions, question performance, and outcomes by cohort. These metrics are useful, but their simplicity is exactly what makes them risky.

A large number can create the illusion of certainty. It can imply strength where there is merely volume, consistency where there is only repetition, or quality where there has simply been favourable context. 

That is why the World Cup stat-padding debate persists: Everyone instinctively understands that all goals count, but some, when they look deeper, realise that not all goals are necessarily equal. 

The same applies in assessment. 

A strong pass rate may signal effective learning, a more accessible paper, a better-prepared cohort, generous standard setting, or perhaps a combination of all four. Without context, the metric describes an outcome but it does not explain it.

Today, the danger is no longer that we have too little data.

The danger is believing the data has already done the thinking for us.

Signals, Noise and False Confidence 

Good analytics is fundamentally an exercise in separating signals from noise. 

In football, that means distinguishing a genuine pattern in performance from a flattering spike produced by games against weaker opposition or a small sample size. 

In assessment, it means resisting the urge to over-interpret one sitting, one cohort, one paper, or one delivery mode as though it represents the whole system.

And this is where over-confidence, perhaps even complacency can creep in. 

We mistake visibility for insight.

A metric appears on a dashboard, highlighted in bold colours, and we instinctively assume that it explains something.

In reality, it may simply describe what happened.

Explanation requires analysis, because dashboards don’t make decisions. People do.

And the real value of analytics lies less in reporting yesterday’s number than in helping leadership teams understand the data, and judge whether that number represents a durable pattern, a contextual anomaly, or an early warning signal.

All Data Thrives on Strong Opposition

A great player does not only impress against the minnows, they raise their game when they face the toughest opponents, and when the stakes are highest. 

Data also needs strong opposition. 

Performance only becomes meaningful when it meets resistance and comparison.

In football, strong opposition may come from the quality of the team faced, the stage of the tournament, the tactical conditions, or the stakes in the game. 

In assessment, opposition takes different forms: item difficulty, cohort characteristics, prior attainment, language demands, delivery conditions, assessor behaviour, or the comparative difficulty of one form versus another. 

The data point matters, but the conditions surrounding it matter just as much.

If we are taking our data and our insights seriously, the goal should not simply be to count and record events, but to challenge the meaning of those events.

A score, on its own, is just a number.

A score interpreted against comparable conditions becomes evidence.

That is where insight begins. 

Fair Comparison Requires Normalisation

Fair comparison is impossible without context.

In sport, the very best analysts don’t simply collate data, they analyse it over time. They adjust it to reflect the quality of the opposition, the importance of the fixture, the sample size involved. 

In assessment, the same discipline is essential if results are to support valid conclusions about candidate performance, item functioning, or organisational quality.

That does not mean turning every dataset into an abstract psychometric exercise. It means asking more challenging questions. 

Was this result strong relative to difficulty? Is this trend visible across multiple cohorts? Did performance change because learning improved, because the assessment changed, or because the population changed? 

The quality of our conclusion depends on the quality of the comparison beneath it.

With such a wealth of data now available to us, we have a greater opportunity to use it, and to learn from it than we have ever had before. 

More data can be useful. 

Better data is valuable. 

Better framing is more valuable than both.  

Analytics should not be about presenting glossy reports to our boards and panels to reassure ourselves that everything is going well, we should be using it to make detailed comparisons, distinguish trends from blips, and separate cause from coincidence. 

Good analytics moves us from reporting, to judgement. 

A Final Thought 

I am not suggesting that raw statistics are useless, they are often the starting point, but too often we can be guilty of mistaking the starting point for the conclusion.

That is why the World Cup stat-padding debate is about more than football. 

The real value of analytics is not that it gives us more numbers.

It gives us better questions.

Every dashboard should provoke curiosity before confidence. Every statistic should be challenged before it is trusted.

Good analytics is not about finding numbers that support our assumptions. It is about finding evidence strong enough to challenge them.

Because better decisions come not from having more data, but from understanding what that data had to overcome.

Better data, better judgement

If this piece raised questions about your own dashboards, that’s a good sign. Answering them takes two things: data that’s built for comparison, and the expertise to challenge what it appears to say.

That’s exactly where we can help. Cirrus Data & Insights gives you the context behind the headline numbers, trends across cohorts, sittings and forms, not just single figures in bold colours. And our consulting team, led by Robbie Burns, works with awarding bodies and assessment providers to turn those numbers into evidence, and evidence into decisions.

[Explore Data & Insights] [Talk to our consulting team]


Header image generated using AI. The stats shown are illustrative, which, fittingly, is rather the point.

Share via:
Topics
Picture of Robert Burns
Robert Burns
Robbie works with assessment organisations on the practical realities of delivering high-stakes assessments, with a particular focus on operational resilience, assessment integrity and candidate experience in digital delivery models.
Would you like to receive Cirrus news directly in your inbox?
More posts in Better Assessments
Better Assessments

Choosing the Right Proctoring Model for Your Assessment

Everyone wants to know the best way to proctor an exam. There isn’t one. Every qualification carries its own purpose, stakes, candidates and constraints, so the useful question is not which model is best, but which model fits, and whether you can explain why you chose it.

Read More »
Better Assessments

How Secure Is Secure Enough?

There is no universal definition of secure enough. The answer depends on what the task exposes, what is at stake if the result is wrong, and the scale you are working at. The third article in The Proctoring Question sets out how to match the control to the risk, and what it costs when you get it wrong.

Read More »

Know exactly where your assessment operation stands

Your free personalised report in 4 minutes

Answer 12 questions across strategy, delivery, design and data and get a clear, personalised breakdown of where your assessment operation is strong and where to focus next.