This is a guest blog post by James Benson*, a South Australian secondary teacher with nearly 20 years classroom and leadership experience across a range of education sectors and settings.
My father always warned me to be wary of people who began sentences with two words: “Yeah, but”. He believed that whatever followed was usually going to be an excuse, a qualification, or an attempt to make an inconvenient fact a little less inconvenient.
I have been thinking about those two words a lot lately, particularly in relation to the way we use data in education.
We certainly do not suffer from a shortage of data. We have NAPLAN, PISA, PIRLS, TIMSS, PAT, DIBELS and an ever-growing collection of school and system assessments. We have dashboards, spreadsheets, data walls, data dives, improvement cycles and meetings devoted entirely to examining results. Being “data-informed” and “evidence-informed” has become part of the language of education.
Yet our commitment to data can be remarkably conditional. When the results accord with what we already believe, they are readily incorporated into our thinking. When they do not, we can become very good at finding reasons why they are not quite as important as they first appeared.
Of course, data needs to be interrogated. That is the point of having it. But interrogation should mean trying to understand what the data is telling us, not searching for reasons to make an uncomfortable result disappear. Different assessments illuminate different parts of student learning, and interpreting them well requires knowledge and judgement. The problem starts when scrutiny becomes a convenient form of dismissal.
Anyone who has spent time around school data will recognise the pattern. A disappointing result arrives, and the explanations (aka rationalisations) are rarely far behind. This was a difficult cohort. Attendance was poor. Students were anxious. The test did not suit them. The assessment failed to capture what teachers knew students could really do. The cohort had particular needs. The list goes on.
Some of these explanations may be correct. Context matters. But context should help us make better sense of data, not automatically provide a reason to disregard it. There is a difference between explaining why a problem exists and explaining the problem away.
I suspect there is another issue here that receives considerably less attention. For all the data we collect, I am not convinced we have invested nearly as much in developing the knowledge required to interpret it.
Data collection and data literacy are not the same thing. A school can produce an impressive dashboard without necessarily understanding what the numbers on it mean. Colour-coding students according to benchmarks, calculating percentages and producing graphs can create the appearance of analytical sophistication, but none of those things guarantees an understanding of what an assessment was designed to measure, what constitutes meaningful growth or what conclusions the evidence will actually sustain.
When that knowledge is weak, analysis can become fairly shallow. Small movements are given significance they may not deserve. Percentages are discussed without considering the size of the cohort behind them. Correlation slides quietly into causation. Measures designed for different purposes are compared as though they are interchangeable, and averages are discussed as if they describe every child in the group.
It is also very easy in these circumstances to reach for explanations that fit the educational mood of the moment. Poor achievement might be attributed to engagement, motivation, wellbeing, disadvantage, test anxiety, technology, COVID, socioeconomic circumstances or flaws in the assessment itself. Any of these could be relevant and several may operate at once. But a plausible explanation is not necessarily a demonstrated explanation.
If the data tells us students are struggling, we need to examine what they were taught, whether the curriculum was coherent and sufficiently ambitious, how effectively that curriculum was taught and whether students actually acquired the knowledge and skills they were supposed to acquire. Examination of context should not absolve us from examining the things schools can change.
Reading comprehension provides a useful example. If assessment data suggests students are struggling to comprehend what they read, describing them as disengaged readers does not get us very far. We need to know whether they can decode accurately and fluently, whether they possess sufficient vocabulary and background knowledge, whether they can navigate increasingly complex syntax and, importantly, what they have actually been taught and read.
The same applies in mathematics. Student anxiety may need to be considered, but it does not remove the need to establish what mathematics students know, what they have been taught and where the gaps lie.
An assessment result is not a diagnosis. It is the beginning of an investigation.
This is the broader context in which I have been watching the response to the latest PISA results. PISA is only one source of educational data, but the discussion surrounding it provides a useful illustration of our sometimes-complicated relationship with evidence.
The latest results provide plenty to think about. Across the OECD, achievement in reading and mathematics has fallen substantially over the past decade. Australia remains above the OECD average, but that should not obscure the longer-term deterioration in our own performance. These trends deserve serious attention.
What interests me just as much as the results themselves is the way international data has been used over time.
For years, Finland's (2000) strong performance on reading was treated as evidence worth examining. Its success generated books, conferences, international study tours and extensive commentary about what other education systems might learn from it. Professor Pasi Sahlberg became an influential advocate for Finnish education, while prominent education thinkers including Professor Andy Hargreaves drew attention to Finland as an alternative to some of the reform approaches being pursued elsewhere, e.g., England.
There was nothing unreasonable about this. When a country achieves unusually strong educational outcomes, we should be curious. We should look at its curriculum, teacher preparation, policies, demographics, culture and classroom practice and consider whether there are useful lessons.
What becomes harder to defend is changing our view of the measure when the story it tells changes.
Finland's performance subsequently declined, while England's relative performance strengthened. Over the same broad period, England pursued a substantial program of education reform: a stronger emphasis on a knowledge-rich curriculum, systematic synthetic phonics, mathematics mastery and more explicit approaches to instruction, alongside significant reforms to behaviour management, discipline and school culture. These were not minor adjustments at the margins. They reflected a quite different view of curriculum, teaching and the conditions required for learning, and many of them ran directly against approaches that had been fashionable in education for decades.
England's results do not establish that these reforms caused its international performance to lift. National education systems are far too complex for a claim of that kind. They do, however, give us a reason to be interested in what England has been doing.
Instead, some responses have concentrated on reasons commentary on England's performance should be qualified. Sampling, demographics, participation rates and the difficulties inherent in comparing education systems have all featured in the discussion. These are legitimate considerations. They are also considerations that apply, to varying degrees, across international assessment data more broadly. No international comparison takes place under perfectly controlled conditions, and no national sample is a miniature laboratory in which every contextual difference has conveniently disappeared.
The issue, then, is not whether we should scrutinise samples or examine the characteristics of the students who participated. Of course we should. The issue is whether we apply that scrutiny consistently.
If differences in sampling, demographics or context are sufficient to substantially discount England's performance, then the same standard needs to be applied when interpreting the performance of Finland, Singapore, Estonia, Canada, Australia or any other education system we happen to be discussing. We cannot treat contextual differences as background noise when the results support a preferred educational narrative and then elevate them to the central explanation when they do not.
Methodological caution cannot become something we remember only when a result is inconvenient.
If international assessment was sufficiently informative to support a global conversation about lessons from Finland, then it remains sufficiently informative to provoke curiosity when a system following a different policy direction performs strongly. PISA cannot be an illuminating source of evidence when it supports an educational philosophy we favour and suddenly become much less interesting when it produces a result we find uncomfortable.
That's the “Yeah, but”, right there.
The same habit appears with domestic and school-level data. When NAPLAN repeatedly identifies weakness in writing, saying that NAPLAN does not measure everything about a child is both true and beside the point. Nobody sensible believes that it does. When DIBELS identifies significant problems with fluency, pointing out that DIBELS does not measure every aspect of reading does not resolve the fluency problem. When PAT indicates weak growth, understanding the limitations of PAT should inform the investigation rather than end it.
No assessment gives us the whole picture. That is precisely why we use multiple sources of evidence, that must be interpreted with care.
A single result may be anomalous. When different measures, designed for different purposes and administered at different points in schooling, begin to tell a similar story, however, the collective evidence becomes increasingly difficult to explain away. The task is not to find the perfect assessment. It is to determine what the weight of the evidence is telling us.
There is also a deeper issue sitting underneath much of this discussion. Our arguments about what counts as meaningful educational data may partly reflect disagreement about what we think schools are actually for.
I should declare my bias here.
I believe the central educational purpose of school is the acquisition of domain knowledge. Schools should teach young people things they do not already know, including things they may have little opportunity to learn outside school. They should open access to mathematics, science, history, geography, literature, the arts and the wider cultural inheritance, while developing the literacy and numeracy that allow students to participate fully in those domains.
This does not mean schools should be cold or joyless places where children's wellbeing is irrelevant. Schools should be safe and inclusive. Children should be known, respected and cared for. Relationships matter. Behaviour matters. Belonging matters.
I regard these things, however, as important conditions for successful schooling rather than replacements for its educational purpose. A school may be warm, inclusive and caring, but if students leave unable to read demanding texts, write coherently or work confidently with mathematics, something fundamental has been missed.
There is nothing historically unusual about debating the purpose of schooling. Mass education has always carried several expectations. Schools have transmitted knowledge and culture, prepared young people for work and citizenship, contributed to social cohesion and played an important role in socialisation. The balance between these purposes has shifted over time.
What does seem increasingly apparent is how much more we now ask schools to do. Their remit has expanded to encompass wellbeing, resilience, identity, agency, engagement, creativity, collaboration, social and emotional development and an expanding collection of dispositions said to prepare young people for an uncertain future.
Many of these are worthwhile aims. The difficulty comes when they begin to compete with, rather than support, the core business of acquiring knowledge and skills, or when they become alternative indicators of success whenever academic outcomes disappoint.
Belonging, motivation and engagement can provide useful information about students' experiences, but they tell us different things from measures of achievement. Self-reported engagement does not establish whether a student can comprehend a difficult text. A sense of belonging does not tell us whether that student understands fractions. A measure of growth mindset does not tell us how much background knowledge a student possesses.
We should be particularly cautious when these softer measures become substitutes for learning itself.
If academic learning is regarded as one outcome among a very large collection of equally important goals, sustained declines in achievement can be accommodated relatively easily. Attention can move towards wellbeing, engagement, creativity or collaboration, followed by the familiar rhetoric that conventional assessments fail to measure everything that matters.
Of course they do.
The relevant issue is whether they measure something that matters.
If the acquisition of knowledge remains a central responsibility of schooling, sustained deterioration in reading, mathematics or science requires a serious response. Those results do not tell us everything about education, but they tell us enough that we should pay attention.
There is an equity dimension to this as well. Families with significant cultural and economic resources can often compensate for what schools do not provide. They can buy books, tutoring, experiences, travel, music lessons and academic support. They can fill gaps in knowledge and navigate the education system on behalf of their children.
Many children do not have that luxury.
For them, school may be the only institution capable of systematically providing access to the knowledge that others acquire partly through circumstance. This is why I find the argument that schools should be less concerned with knowledge particularly difficult to reconcile with claims about educational equity. Knowledge is not an optional extra for disadvantaged children. It is one of the most powerful things schools can distribute more fairly.
And that brings us back to data.
Sometimes a cohort genuinely is problematic. Sometimes a movement in results means very little. Sometimes attendance, disadvantage or disruption explains much of what we see.
And sometimes students simply have not learned enough of what we were charged with teaching them
An evidence-informed profession needs to be capable of entertaining that possibility without immediately reaching for an explanation that allows us to preserve what we already believe. That requires better assessment literacy, better data literacy and a willingness to follow evidence into places that may sit awkwardly with the prevailing educational zeitgeist.
Data needs interpretation, but interpretation should help us get closer to the problem rather than provide increasingly sophisticated ways of avoiding it. If evidence only counts when it can be used to confirm what we already believe, we are not really allowing it to inform us at all.
My father was probably right to be suspicious of “Yeah, but”. The phrase becomes most seductive at precisely the point when the evidence is telling us something we would rather not hear.
That is probably when we need to listen most carefully.
*James Benson is a pseudonym. The author is a real teacher, who wishes to preserve anonymity, for professional reasons.
(c) Pamela Snow & "James Benson" (2026)