A student takes a national exam and scores 75 out of 100. The number appears on their certificate, determines their university options, and follows them for years. It looks precise. It looks final. It looks like a fact about the student, in the same way that their height or weight is a fact.
It isn’t. And understanding why it isn’t – where the imprecision comes from, how big it is, and what test developers do to control it – is one of the most useful things anyone who takes, gives, or interprets exams can learn.
This post is about three connected ideas: what an exam score actually represents, where measurement error comes from, and how psychometricians make scores comparable when different students take different versions of a test. That last topic – equating – is one of the most important and least understood operations in all of assessment.
The Score Is Not the Ability
Classical test theory begins with a deceptively simple equation: X = T + E. The observed score (X) equals the true score (T) plus measurement error (E).
The observed score is the 75 the student actually got. The true score is a theoretical quantity – the score the student would average if they could take infinitely many parallel versions of the test, under all possible conditions, with all the randomness washed out. And the error is the gap between the two on this particular occasion.
We never observe the true score. We only ever see one draw from a distribution of possible scores. If the same student took a parallel exam tomorrow, they might score 72. Or 78. Their ability hasn’t changed overnight – but the specific items, the specific day, the specific circumstances all shift the observed result around the underlying truth.
This is the first mental shift: an exam score is an estimate of ability, not a measurement of it in the physical sense. It comes with an error margin, whether or not that margin is printed on the certificate.
Where Error Comes From
So what creates the gap between observed and true scores? In short: everything. Every element of the testing situation contributes something to the error. It helps to group the sources into three categories.
The items themselves. A question may be ambiguously worded, so that a knowledgeable student reads it differently than intended. A distractor may accidentally contain a clue. And there’s the deeper issue of content sampling: any exam is a sample of items drawn from a much larger domain. The specific 40 questions on this year’s mathematics exam happen to cover certain topics and skip others. A student who happened to study exactly those topics gets a boost; an equally able student whose strengths lie in the skipped topics gets penalized. Neither score difference reflects a real ability difference – it reflects luck in the sampling of items.
The testing conditions. Noise from the corridor. A hot room with a broken air conditioner. An uncomfortable chair, poor lighting, an exam scheduled at 8 a.m. for a student who functions best in the afternoon. Small disruptions accumulate, and they push observed scores away from true scores in ways that have nothing to do with what the exam is measuring.
The students themselves. A test taker might be fatigued, anxious, or coming down with an illness. They might have a momentary lapse and misread a question they knew the answer to. They might guess on three items and get lucky on two, or unlucky on all three. Human performance fluctuates from day to day and hour to hour, and the exam captures a single snapshot of that fluctuation.
None of these sources can be eliminated entirely. But they can be reduced through better item writing, standardized administration conditions, and longer tests that average out the noise. This is much of what professional test development actually consists of: the systematic hunting down and shrinking of error sources.
Why Different Students Get Different Booklets
Now to the part of exam design that puzzles people the most: why do different students often take different versions of the same exam?
There are two main reasons, corresponding to two very different assessment contexts.
In national high-stakes exams, the reason is largely logistical. When an entire country’s graduating cohort must be examined, it is extremely difficult to have every student write at the same moment. The organizational preparations for a single simultaneous administration at that scale are daunting – venues, proctors, materials, security. So exams run in multiple sessions, and each session must use a different variant of the test. If the same items were reused across sessions, students in later sessions would have an obvious advantage: the content would leak.
In international large-scale assessments like PISA, TIMSS, and PIRLS, the reason is entirely different: content coverage. The goal of these studies is to measure a whole domain – reading literacy, mathematical literacy – with high accuracy at the country level. Covering a domain that broad requires a very large pool of items, far more than any single student could answer in a reasonable testing session. So the item pool is distributed across many different booklets, each student takes only one booklet, and the full domain coverage is achieved across the sample rather than within each individual.
Both designs are sensible. Both are necessary. And both create the same new problem: the booklets are never exactly equally difficult. However carefully test developers try to balance the variants, one booklet will always turn out slightly harder than another. Which means the existence of multiple booklets is itself a source of error – a student’s score now depends partly on which variant they happened to receive.
This is where equating comes in.
Equating Without Common Items: The Equipercentile Method
In national exams, security requirements typically mean that no items are shared between sessions. Booklet A and booklet B have zero questions in common. How can scores from two completely different tests be placed on the same scale?
The answer used in Georgia’s national exams – and in many similar systems worldwide – is equipercentile equating. The logic is elegant.
Put simply, we calculate what score a student would have received if they had written a different variant. We then compare their raw score with their estimated scores on the other variants, and take the maximum – that maximum is their equated score. Put another way: a student’s equated score is the score they would have received if they had happened to write the easiest variant of the test.
How can we possibly know what someone would have scored on a test they never took? Through percentile ranks. Because each booklet is written by a very large number of students – roughly nine to ten thousand per variant in the Georgian national exams – we have a precise picture of the score distribution in each session. The method works like this: we determine what percentage of test takers a student outperformed within their own variant. Then we find, in each other variant, the theoretical score whose holder would have outperformed exactly the same percentage of test takers in that variant.
If you beat 80% of the students who took your booklet, your equated equivalent on another booklet is whatever score beats 80% of that booklet’s takers. Same percentile, different raw score – but the same standing relative to the population.
The natural objection: what if one variant simply attracted better-prepared students than another? Then the percentile comparison would be distorted. But this is practically ruled out by design, because of the principle booklets are assigned to students randomly. With random assignment and thousands of students per variant, the groups writing each booklet are statistically equivalent, which gives high statistical confidence that any difference in score distributions between variants reflects differences in booklet difficulty – not differences in the students. That is exactly the condition equipercentile equating requires.
The elegance of the “take the maximum” rule is also worth noting: it means no student can ever be disadvantaged by having received a harder variant. The equating always resolves in the student’s favor, aligning everyone to the easiest booklet’s scale.
Equating With Common Items: Anchors in ILSAs
International assessments solve the comparability problem differently, using anchor items – items that appear in more than one booklet.
In PISA, TIMSS, and PIRLS, the many booklets are deliberately constructed so that they overlap: each booklet shares blocks of items with several other booklets. These shared items are the anchors. Because every booklet is connected to the others through common items, all items, and all students, can be calibrated onto a single common scale using Item Response Theory. No student answers every item, but the network of overlaps ties everything together into one coherent measurement system.
Anchor items do something else that equipercentile equating cannot: they link assessments across time. A subset of items is carried over from one assessment cycle to the next – from PISA 2018 to PISA 2022, for example. These trend anchors are what make it possible to say that a country’s reading performance rose or fell between cycles. Without common items spanning the years, each cycle would be its own isolated measurement, and cross-year comparisons would be meaningless.
This is one of the quiet miracles of modern assessment methodology: hundreds of thousands of students, in dozens of countries, across multiple years, taking different booklets in different languages – all placed on one comparable scale, held together by a carefully designed web of anchor items.
Does Equating Remove the Error?
Equating dramatically reduces the unfairness that booklet differences would otherwise create. But it’s important to be honest: it does not remove error entirely. Equating is itself a statistical procedure performed on finite samples, and it introduces its own uncertainty – known, fittingly, as equating error.
The equated score conversion is estimated from data, and like any estimate, it has a margin. This is one reason why large samples per booklet matter so much, and why the design of the equating study – random assignment, adequate anchor coverage, quality control of anchor item behavior – receives so much technical attention. Serious assessment programs don’t ignore this residual uncertainty; ILSAs explicitly include equating error as a component of the total error budget when reporting results and trends.
So the honest summary is: multiple booklets introduce error, equating removes most of it, and a small, quantifiable amount remains. That remaining amount is the price of running exams at scale, and knowing its size is far better than pretending it doesn’t exist.
What This Means for You
If you take one idea away from this post, let it be this: exam scores are not exact measurements like a person’s height or weight. Every score contains error, and the size of that error varies.
This is not a reason to distrust exams. It’s a reason to understand them. A score of 75 means “somewhere around 75” – and how wide that “somewhere around” is depends entirely on the quality of the test. Good tests, written by skilled item writers, administered under standardized conditions, and equated with proper methodology, have small error. Poor tests have large error, and their scores can mislead badly while looking just as precise on paper.
The whole discipline of psychometrics exists in that gap – in the systematic effort to make the error smaller, to know how big it is, and to be honest about what a number on a certificate can and cannot tell us about a human being.
—
Giorgi Tchumburidze
July, 2026