The main types of IQ test, and what each is for

"IQ test" covers instruments that differ enormously in what they sample, who may administer them, and what their scores can support. Knowing which kind you are looking at answers most questions about it.

Individual clinical batteries

These are the instruments people mean by "a real IQ test": administered one-to-one by a qualified psychologist, taking one to two hours, and producing a profile rather than a single number.

The Wechsler Adult Intelligence Scale is the most widely used. It reports a Full Scale IQ built from index scores covering verbal comprehension, perceptual reasoning, working memory and processing speed. Its standardisation sample comprised 2,200 adults stratified against census figures, and full-scale reliability is around 0.98.1 Parallel versions exist for children.

The Stanford-Binet descends from the original 1905 Binet-Simon scale, built to identify children needing additional support in school - the task that created the field.7 The current edition covers five factors in both verbal and non-verbal formats.2

What you are paying for with these is not harder puzzles. It is:

Non-verbal and matrix tests

These strip out language entirely. Raven's Progressive Matrices is the archetype: grids of figures with a missing cell, no words, no arithmetic.3

They exist for good reasons. They can be given across languages with only the instructions translated; they suit people with language impairments or limited schooling; and they load heavily on fluid reasoning, the ability most associated with the general factor.8

The trade is narrowness. A matrices score says little about verbal comprehension, acquired knowledge or processing speed. It is a deep measurement of one broad ability rather than a survey of several - which is exactly right for research and incomplete for clinical assessment. See how matrices work.

Group and screening tests

Group tests are administered to many people at once, on paper or by computer, without individual observation. Military selection batteries and educational screening instruments are the main examples, and the format dates to the need to sort very large numbers of recruits quickly.

They are cheap per head and reasonably reliable, but they lose the administrator’s observation - nobody notices that you misheard the instruction, or worked carefully rather than quickly. Brief screeners exist for the same reason: a fifteen-minute estimate that flags whether a full assessment is warranted, without pretending to replace one.

Online tests

The category covers everything from serious research instruments to entertainment.

At the serious end sit publicly documented research measures. The International Cognitive Ability Resource was developed and validated on 96,958 participants across 199 countries, with published reliability figures and item statistics.5 Its 16-item short form reached an internal consistency of 0.81, and a later study reported a correlation of about 0.81 with WAIS-IV Full Scale IQ in a university sample. A short online battery can carry real signal.

At the other end are tests that produce a number by a formula nobody publishes, frequently behind a payment, often with no norms at all.

Four questions separate them, and all four are answerable before you start:

  1. Is the scoring method published?
  2. Is a margin of error reported with the score?
  3. Is the norm group described - and if there is not one, does the site admit it?
  4. Are you asked to pay to see your result?

A "no" to the first three or a "yes" to the fourth is enough to disregard the number. The longer version of this checklist is in are online IQ tests accurate.

Side by side

Clinical batteryMatricesGroup testThis test
Time60-120 min20-45 min30-90 min~30 min
SupervisedYes, one-to-oneUsuallyGroup settingNo
Abilities sampledFour or moreOne (fluid)VariesFour
NormsLarge, stratified, age-bandedPublishedPopulation-specificProvisional
Reports uncertaintyYesYesUsuallyYes
Diagnostic standingYesLimitedNoNone
CostOften substantialVariesVariesFree

The row that matters most is norms. Everything downstream of it - what the number means, whether it can be compared with anyone else’s - depends on the reference group, and it is the row where this test is weakest and says so.

Which one do you need

This site is firmly in the last category, and the methodology page sets out exactly what that supports: a real item response model, rule-generated items verified on every build, an honest confidence interval - and provisional norms rather than a standardisation sample.

Common questions

What is the most accurate IQ test?

An individually administered clinical battery such as the WAIS, given by a qualified psychologist. Full-scale reliability is around 0.98, the norms come from a large stratified sample, and an administrator observes how you work rather than only what you answer.

What is the difference between the WAIS and Raven's Matrices?

Breadth. The WAIS samples four broad domains and reports a Full Scale IQ with an index profile. Raven's measures fluid reasoning only, using no language at all - deeper on one ability, silent on the others.

Can a free online test replace a professional assessment?

No. Unsupervised tests cannot verify identity or conditions, usually lack a proper standardisation sample, and carry no diagnostic standing. A well-built one is informative about how you reason; it is not evidence for any decision that matters.

Why do different tests give me different scores?

Because they sample different abilities, use different norm groups collected in different years, and carry their own measurement error. A gap of several points between two properly built tests is entirely ordinary, which is why scores should be read as intervals.

References

  1. Wechsler, D. (2008). Wechsler Adult Intelligence Scale - Fourth Edition: Technical and Interpretive Manual. Pearson.
  2. Roid, G. H. (2003). Stanford-Binet Intelligence Scales, Fifth Edition: Technical Manual. Riverside Publishing.
  3. Raven, J. (2000). The Raven Progressive Matrices: Change and stability over culture and time. Cognitive Psychology, 41(1), 1-48. doi:10.1006/cogp.1999.0735
  4. Carroll, J. B. (1993). Human Cognitive Abilities: A Survey of Factor-Analytic Studies. Cambridge University Press.
  5. Condon, D. M., & Revelle, W. (2014). The International Cognitive Ability Resource: Development and initial validation of a public-domain measure. Intelligence, 43, 52-64. doi:10.1016/j.intell.2014.01.004
  6. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. AERA.
  7. Binet, A., & Simon, T. (1905). Methodes nouvelles pour le diagnostic du niveau intellectuel des anormaux. L'Annee Psychologique, 11, 191-244.
  8. Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment. Journal of Educational Psychology, 54(1), 1-22.

Related reading

Find out where you land

Thirty questions, about 30 minutes, and a score with the uncertainty attached.

Start the test

30 questions · about 30 minutes · no sign-up