How IQ tests are built and normed
A score only means something because of work done long before you sat down. Understanding that pipeline is the fastest way to tell a real test from a quiz wearing its clothes.
On this page
1. The blueprint
Before an item is written, a test needs a specification: what it claims to measure, and how that claim maps onto content.
For cognitive ability the usual reference is the Cattell-Horn-Carroll model, built from a survey of more than 460 factor-analytic datasets. It arranges abilities in three strata: general ability at the top, around a dozen broad abilities beneath it, and many narrow ones below those.8 A blueprint says which broad abilities will be sampled, with how many items each, and why.
Skipping this step is what produces tests that are really just a pile of puzzles. If nobody can say what construct an item was chosen to measure, the total has no defined meaning.
2. Writing and piloting items
Items are drafted against the blueprint, reviewed for ambiguity and cultural loading, then piloted - given to a sample purely to find out how they behave. Piloting is where most items die. Typical failure modes:
- Two defensible answers. A second reading of the item works and is not the key.
- A giveaway distractor. One option is obviously odd, so the item becomes a three-way guess.
- Negative discrimination. High scorers get it wrong more often than low scorers - a reliable sign the item is measuring something other than intended.
- Ceiling or floor. Almost everyone passes or almost everyone fails, so the item separates nobody.
An automatic generation approach changes this stage considerably: if items are produced from explicit rules, correctness can be proved rather than piloted, and difficulty can be predicted from structure.7 It does not remove the need for empirical data - it removes the need to discover that an item is broken.
3. Item statistics
Every surviving item gets numbers attached. Classical test theory gives two:3
- Difficulty (p) - the proportion answering correctly.
- Discrimination - how well the item separates people who scored high overall from those who scored low.
Modern tests use item response theory instead, which models the probability of a correct answer as a function of ability.4 The three-parameter logistic model gives each item:
| Parameter | Meaning |
|---|---|
| b - difficulty | The ability level at which the item becomes an even bet |
| a - discrimination | How sharply the item separates people around that level |
| c - guessing | The floor set by chance on a multiple-choice item |
The payoff is that a hard item passed counts as stronger evidence than an easy one, and the model reports how precisely ability has been estimated at each point on the scale - which is where the confidence interval on a score comes from.
4. The standardisation sample
This is the step that actually creates the IQ number, and the step cheap tests skip.
A large sample is recruited to match a target population on age, sex, education, region and other characteristics. The WAIS-IV standardisation sample, for instance, comprised 2,200 adults stratified against United States census figures.2 Their raw scores are then transformed so the mean becomes 100 and the standard deviation 15.
Everything an IQ score means comes from this sample. The number is a rank against a specific reference group, not a measurement of a quantity. Change the reference group and the same performance yields a different score. A test with no standardisation sample is not producing an IQ in any meaningful sense, whatever number it prints.
Norms are usually banded by age, because raw performance varies systematically across the lifespan. A 65-year-old and a 25-year-old with identical raw scores receive different IQs, because each is ranked against their own age group.
5. Reliability and validity
Two different questions, routinely confused.
Reliability asks whether the test measures consistently. Internal consistency (coefficient alpha, or preferably omega) checks whether items agree with each other; test-retest checks stability over time. The WAIS-IV reports full-scale reliability around 0.98, which is exceptionally high and reflects its length.2
Validity asks whether it measures the right thing - and is not a single number but an accumulating argument built from several kinds of evidence: does the internal structure match the blueprint, does it correlate with established measures, does it predict relevant outcomes.51
A test can be highly reliable and invalid. A bathroom scale that always reads four kilos heavy is perfectly reliable. Reliability is necessary and nowhere near sufficient.
6. Norms decay
A standardisation sample is a photograph of a population at a moment. Raw performance drifted upward through the twentieth century at roughly 0.28 IQ points per year.6
So a test normed in 2000 and still scored against those norms will, by 2026, be handing out scores several points too generous. This is why publishers renorm every decade or two, and why the date of the norms is a material fact about any score. See what is an average IQ for how renorming keeps the mean pinned at 100.
Where this test sits
Honestly, against the pipeline above:
| Stage | This test |
|---|---|
| Blueprint | Yes. Four broad CHC abilities, item counts and rationale published in the methodology. |
| Item correctness | Proved, not piloted. Rule-generated items, verified by two independent implementations on every build. |
| Item parameters | Rational, not empirical. Difficulty assigned from structural complexity, anchored to published data for comparable item families. |
| Scoring model | Yes. Three-parameter logistic IRT with expected a posteriori estimation. |
| Standardisation sample | No. This is the honest gap. |
| Reliability | Estimated from the model, not from a retest study. |
| Validity evidence | Argued from item design, not demonstrated against a criterion. |
The machinery downstream of norming is genuine. The norms themselves are provisional, which is why the score is reported as an interval and capped at the extremes rather than presented as a precise figure. A site that quietly skipped step 4 and printed a confident number would be doing something worse than being imprecise - it would be hiding which part is guesswork.
Common questions
What is a standardisation sample and why does it matter?
It is the group whose performance defines the scale. Their scores are transformed so the mean is 100 and the standard deviation 15, and everyone else is ranked against them. Without one, an IQ number has no defined reference point - it is a raw score wearing a costume.
What is the difference between reliability and validity?
Reliability is consistency; validity is measuring the right thing. A scale that always reads four kilos heavy is perfectly reliable and completely invalid. Reliability is necessary but nowhere near sufficient, and a test reporting only reliability is telling you the easier half of the story.
Why do IQ tests need renorming?
Because raw performance drifted upward at roughly 0.28 points a year through the twentieth century. Norms collected in 2000 would by now produce scores several points too generous, so publishers restandardise every decade or two to pull the mean back to 100.
Does this test have proper norms?
No, and we say so throughout. The item parameters are rational - assigned from structural complexity and anchored to published difficulty data for comparable item families - rather than calibrated on a standardisation sample. The scoring model is real IRT; the norms it operates on are provisional.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. AERA.
- Wechsler, D. (2008). Wechsler Adult Intelligence Scale - Fourth Edition: Technical and Interpretive Manual. Pearson.
- Lord, F. M., & Novick, M. R. (1968). Statistical Theories of Mental Test Scores. Addison-Wesley.
- Embretson, S. E., & Reise, S. P. (2000). Item Response Theory for Psychologists. Lawrence Erlbaum.
- Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302.
- Pietschnig, J., & Voracek, M. (2015). One century of global IQ gains: A formal meta-analysis of the Flynn effect (1909-2013). Perspectives on Psychological Science, 10(3), 282-306. doi:10.1177/1745691615577701
- Condon, D. M., & Revelle, W. (2014). The International Cognitive Ability Resource: Development and initial validation of a public-domain measure. Intelligence, 43, 52-64. doi:10.1016/j.intell.2014.01.004
- Carroll, J. B. (1993). Human Cognitive Abilities: A Survey of Factor-Analytic Studies. Cambridge University Press.
Related reading
See the pipeline in action
Thirty items, an IRT-scored result, and a methodology page that shows its working.
Start the test30 questions · about 30 minutes · no sign-up