Key takeaways: IQ testing began in early 1900s France as a practical tool to identify students needing extra educational support, not as a way to rank general intelligence. It evolved through several major instruments — Binet’s original scale, the Stanford-Binet, and the Wechsler scales — each refining both the format and the statistical methods used to calculate a score.
The Original Purpose: Identifying Students Who Needed Help
The story of IQ testing starts with French psychologist Alfred Binet, who in the early 1900s was commissioned by the French government to develop a way to identify schoolchildren who needed additional educational support, rather than being placed in mainstream classrooms without extra help. Working with Théodore Simon, Binet developed what became known as the Binet-Simon scale — a series of tasks of increasing difficulty designed to estimate a child’s “mental age” relative to their actual chronological age.
This origin matters for understanding what the concept was originally built to do: it was a practical, targeted diagnostic tool for a specific educational purpose, not a general ranking system for adult intelligence, and Binet himself reportedly expressed reservations about the scale being used to label a person’s fixed, permanent intellectual worth — a concern that turned out to be prescient given how the concept was later applied and sometimes misused.
The Original “Quotient” Calculation
The term “intelligence quotient” itself comes from German psychologist William Stern, who proposed expressing the relationship between mental age and chronological age as a ratio: mental age divided by chronological age, multiplied by 100. A ten-year-old performing at the level of a twelve-year-old would score 120 under this method; a ten-year-old performing at the level of an eight-year-old would score 80.
This ratio-based approach worked reasonably well for children, where mental age develops in a roughly comparable way to chronological age, but it broke down for adults, since cognitive development doesn’t continue increasing linearly with age the way this simple ratio assumes. This limitation is part of why the field eventually moved away from the ratio method entirely.
The Stanford-Binet
Binet’s original scale was brought to the United States and substantially revised by Lewis Terman at Stanford University in 1916, becoming the Stanford-Binet Intelligence Scale — a name that persists today through several subsequent revisions (the current edition is the fifth). Terman’s version expanded the scale’s scope considerably and, notably, extended its intended use beyond Binet’s original educational-support purpose toward broader claims about measuring general intelligence across the population, a shift that set the direction much of the field would follow for decades.
Terman also led the well-known longitudinal study following gifted children (identified via IQ testing) over the course of their lives — a project that produced valuable data but also revealed the limits of IQ as a predictor of extraordinary lifetime achievement, since several children with more moderate scores in the study cohort went on to more distinguished careers than some of the study’s highest scorers.
The Deviation IQ and the Modern Scale
The ratio-based “mental age / chronological age” method was eventually replaced by what’s called the deviation IQ — the approach still used today, and the one described throughout this site. Instead of a ratio, your score is calculated by comparing your performance to a reference population and expressing the result as a standardized score with a fixed mean (100) and standard deviation (commonly 15). David Wechsler was instrumental in popularizing this shift, in part because it solved the adult-scoring problem the ratio method struggled with, and because it made scores directly comparable using well-established statistical methods. See standard deviation and the bell curve for the full mechanics of how this scale works.
The Wechsler Scales
David Wechsler introduced his own test — initially the Wechsler-Bellevue Intelligence Scale in 1939, later evolving into the Wechsler Adult Intelligence Scale (WAIS) and the Wechsler Intelligence Scale for Children (WISC) — built around a structurally different approach from the Stanford-Binet. Rather than one continuous scale, Wechsler’s tests were built from multiple subtests grouped into broader categories (verbal comprehension, perceptual/spatial reasoning, working memory, and processing speed, with the specific groupings evolving across editions), combining into a Full Scale IQ score.
This subtest-based structure — measuring several distinct cognitive domains rather than one blended task type — is the direct conceptual ancestor of the category-breakdown approach used across most modern tests, including the four-category structure (spatial, numerical, logical, applied) used in this site’s own free IQ test and explained further in what does an IQ test measure.
Addressing Bias: Culture-Fair Testing
As IQ testing became more widespread through the mid-20th century, researchers increasingly recognized that many tests relied heavily on language and culturally specific knowledge, which could disadvantage test-takers from different linguistic or educational backgrounds in ways unrelated to their actual reasoning ability. This concern led to the development of culture-fair and nonverbal tests — Raven’s Progressive Matrices being among the most well-known and widely used examples — designed to minimize reliance on language and rely primarily on abstract visual pattern reasoning instead. See professional and official IQ tests explained for more on how these instruments are used today.
The Rise of Online Testing
The most recent chapter in this history is the proliferation of self-administered online IQ tests over the past two decades, made possible by widespread internet access and the relatively low cost of building a scored digital assessment compared to the resources required for a fully validated clinical instrument. This has made informal cognitive testing radically more accessible than at any prior point in the field’s history — though, as covered in how accurate are free online IQ tests, this accessibility comes with real tradeoffs in supervision and validation rigor compared to a professionally administered assessment.
What This History Tells Us About Using IQ Tests Today
Understanding this trajectory — from a targeted educational diagnostic tool, through several major statistical and structural revisions, to today’s mix of clinical and informal instruments — helps put any single IQ score in better context. The concept has always been evolving, has always had known limitations, and was never intended (even by its original creator) to serve as a complete, fixed judgment of a person’s overall intellectual worth. That context is worth keeping in mind whether you’re looking at a score from a century-old test or one you just got from taking the free test on this site.
Frequently Asked Questions
Who invented the IQ test? Alfred Binet, working with Théodore Simon, developed the first practical intelligence test in the early 1900s in France, originally to identify students needing additional educational support.
What does “deviation IQ” mean? It’s the modern scoring method that replaced the original ratio-based “mental age divided by chronological age” calculation, instead comparing your performance to a reference population and expressing the result on a standardized scale (mean 100, SD 15).
What’s the difference between the Stanford-Binet and Wechsler tests? The Stanford-Binet evolved from Binet’s original single-scale approach, while the Wechsler scales (like the WAIS) use a structurally different subtest-based design, combining several distinct cognitive domain scores into a Full Scale IQ.
Why were culture-fair tests developed? To reduce reliance on language and culturally specific knowledge that could disadvantage test-takers from different backgrounds, relying instead primarily on abstract visual reasoning.
Did Alfred Binet believe IQ tests should rank general intelligence? Historical accounts suggest Binet had reservations about his scale being used to label a person’s fixed, permanent intellectual capacity — his original intent was a targeted diagnostic tool for identifying students needing support, not a general intelligence ranking system.