Are IQ tests biased?
"Biased" means something specific in psychometrics, and it is not what most arguments about it assume. Separating the technical question from the political one makes both easier to think about.
On this page
Three different things called bias
Arguments about test bias usually go badly because the participants are using the word to mean three unrelated things. The testing standards used across the field distinguish them carefully.1
| Sense | The question it asks | How it is tested |
|---|---|---|
| Predictive bias | Does the test systematically over- or under-predict a real outcome for one group? | Compare regression slopes and intercepts across groups |
| Measurement bias | Does the test measure the same thing, on the same scale, in each group? | Measurement invariance testing; differential item functioning |
| Content or cultural loading | Does answering require knowledge that is unevenly distributed for reasons unrelated to reasoning? | Item review, cross-cultural comparison |
These can come apart. A test can be free of predictive bias and still be culturally loaded. It can measure the same construct in two groups while those groups have had wildly unequal opportunities to develop it. Establishing one thing tells you very little about the others - which is precisely why the argument goes in circles.
A score gap is not itself bias
This is the technical point that most often gets lost, and it cuts both ways.
If two groups differ in average score, that difference alone tells you nothing about whether the instrument is faulty. A thermometer is not biased because it reads lower in Reykjavik than in Cairo. Whether a test is measuring badly is a separate question from whether the thing being measured differs, and separate again from why it differs.
The corollary is the part people skip: showing a test is technically unbiased does not show that the scores reflect equal opportunity. If a group has had systematically less schooling, an unbiased test will faithfully record the consequence. Faithful measurement of an unequal situation is still measurement of an unequal situation.
That matters here because education has a demonstrated causal effect on measured intelligence - roughly one to five IQ points per additional year, established using compulsory-schooling reforms as natural experiments.7 Unequal schooling is therefore sufficient to produce score differences without any defect in the test at all.
Major expert reviews have consistently reported that well-constructed tests show little evidence of predictive bias, while stressing that the causes of group differences in scores remain unresolved and are not settled by the psychometric evidence.23 This site takes no position on those causes; a test score is not the kind of evidence that could settle them.
Cultural loading is real
Some item types are obviously culture-bound. A vocabulary question asks whether you have met a word, and which words you have met depends on your language, schooling and reading. A general knowledge question is worse still. These items can be excellent predictors within a culture and close to meaningless across cultures.
Figural matrices were designed to reduce this. They use no words, require no facts, and ask only that you infer a rule from a pattern - which is why they travel better than verbal tests and why they dominate cross-cultural work.6
"Culture-reduced" is not "culture-free", though. Matrix items still assume:
- familiarity with reading a grid left-to-right and top-to-bottom;
- comfort with abstract, decontextualised puzzles as a legitimate activity;
- experience of formal testing, and of the implicit rule that you should keep going when stuck;
- schooling in the analytic style of thinking these items reward.
None of that is innate. All of it is unevenly distributed. It is one reason cross-national score comparisons are treated sceptically by measurement specialists - such datasets are frequently built from small, unrepresentative and non-comparable samples, and rarely satisfy the invariance requirements that would make the comparison meaningful in the first place.8
Stereotype threat and what it explains
A well-known line of research showed that making a negative stereotype salient before a test depressed the performance of the stereotyped group.4 The finding shaped decades of discussion about testing.
A later analysis raised a sharp technical objection. If stereotype threat lowers scores, it should show up as a violation of measurement invariance - the test would be functioning differently under threat. Re-analysing the data, the authors found the evidence for this was weaker and more inconsistent than the standard account implied.5
The honest summary is that situational factors clearly affect test performance, that stereotype threat is one candidate mechanism, and that its size and generality are actively contested.
What is not contested is that motivation moves scores substantially. A meta-analysis of studies offering material incentives found average gains of around 0.64 standard deviations, concentrated among lower scorers.9 Any unsupervised test - including this one - is measuring effort as well as ability, and cannot distinguish the two.
Where this test is weakest
Applying all of the above to this site specifically, rather than in the abstract:
| Concern | Status here |
|---|---|
| Language dependence | Real. Seven of 30 items are verbal and assume fluent English, including two vocabulary-in-context items. A fluent non-native speaker will be underestimated. |
| Cultural content | Partly mitigated. The other 23 items are figural, numeric or spatial and require no cultural knowledge - but they still assume familiarity with formal testing. |
| Predictive bias | Unknown. Testing it requires an external criterion and group data we do not have and do not collect. |
| Measurement invariance | Untested. This needs a large sample with demographic data. We hold no such data by design. |
| Motivation and conditions | Uncontrolled. Nobody is supervising you. Effort, interruptions and fatigue all enter the score. |
Two of those say "unknown", and that is not evasion - it follows directly from a deliberate design decision. This test stores nothing and asks for no demographic information, which means the data needed to test for bias does not exist. We think that is the right trade for a free public test, but it has a cost, and the cost is that we cannot make the reassuring claim.
The practical implication: if English is not your first language, treat the verbal domain score as an underestimate and weight the figural, numeric and spatial domains more heavily. The methodology page sets out the rest of what the test can and cannot support.
Common questions
Does a difference in average scores between groups prove the test is biased?
No. Bias in the technical sense means the test measures differently or predicts differently across groups. A difference in scores is compatible with an unbiased test faithfully recording differences in the environments and opportunities that produced them. Equally, showing a test is unbiased does not show those environments were equal.
Are non-verbal tests like matrices culture-free?
Culture-reduced, not culture-free. They avoid vocabulary and factual knowledge, which is why they travel better across languages. But they still assume familiarity with grids, with abstract puzzles as a worthwhile activity, and with formal testing itself - all of which are products of schooling.
Will this test underestimate me if English is not my first language?
Probably, yes. Seven of the 30 items are verbal and two turn on knowing specific English words. Your verbal domain score is the one to discount; the figural, numeric and spatial domains are far less language-dependent.
Has this test been checked for bias?
No, and it cannot be with the current design. Testing for predictive bias or measurement invariance requires demographic data and an external criterion. This test stores nothing about you and asks no demographic questions, so that data does not exist. We would rather say this plainly than imply a check we have not done.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. AERA.
- Neisser, U., Boodoo, G., Bouchard, T. J., et al. (1996). Intelligence: Knowns and unknowns. American Psychologist, 51(2), 77-101.
- Nisbett, R. E., Aronson, J., Blair, C., Dickens, W., Flynn, J., Halpern, D. F., & Turkheimer, E. (2012). Intelligence: New findings and theoretical developments. American Psychologist, 67(2), 130-159. doi:10.1037/a0026699
- Steele, C. M., & Aronson, J. (1995). Stereotype threat and the intellectual test performance of African Americans. Journal of Personality and Social Psychology, 69(5), 797-811.
- Wicherts, J. M., Dolan, C. V., & Hessen, D. J. (2005). Stereotype threat and group differences in test performance: A question of measurement invariance. Journal of Personality and Social Psychology, 89(5), 696-716. doi:10.1037/0022-3514.89.5.696
- Raven, J. (2000). The Raven Progressive Matrices: Change and stability over culture and time. Cognitive Psychology, 41(1), 1-48. doi:10.1006/cogp.1999.0735
- Ritchie, S. J., & Tucker-Drob, E. M. (2018). How much does education improve intelligence? A meta-analysis. Psychological Science, 29(8), 1358-1369. doi:10.1177/0956797618774253
- Wicherts, J. M., Borsboom, D., & Dolan, C. V. (2010). Why national IQs do not support evolutionary theories of intelligence. Personality and Individual Differences, 48(2), 91-96. doi:10.1016/j.paid.2009.05.028
- Duckworth, A. L., Quinn, P. D., Lynam, D. R., Loeber, R., & Stouthamer-Loeber, M. (2011). Role of test motivation in intelligence testing. PNAS, 108(19), 7716-7720. doi:10.1073/pnas.1018601108
Related reading
Take it with the caveats in mind
Thirty questions, a score with the uncertainty attached, and every limitation stated openly.
Start the test30 questions · about 30 minutes · no sign-up