Intelligence Testing and Achievement Study Pack

Kibin's free study pack on Intelligence Testing and Achievement includes a 6-section study guide, 25 quiz questions, 30 flashcards, and 5 open-ended Explain review questions. Sign up free to track your progress toward mastery, plus upload your own notes and recordings to create personalized study packs organized by course.

Last updated May 28, 2026

Topic mastery0%

Intelligence Testing and Achievement Study Guide

Unpack the key debates in intelligence testing — from Spearman's g factor to Gardner's and Sternberg's multiple intelligences — and master how IQ scores, the Flynn Effect, and achievement vs. aptitude distinctions apply to real-world outcomes.

Key Takeaways

  • Intelligence is most commonly defined as the capacity to learn from experience, solve novel problems, and adapt to new environments, though researchers disagree about whether it is a single general ability or a collection of distinct mental faculties.
  • Charles Spearman proposed a general intelligence factor (g) derived from statistical correlations across cognitive tasks, while Howard Gardner and Robert Sternberg independently argued for multiple, qualitatively different forms of intelligence.
  • The most widely used individual intelligence tests — the Wechsler scales and the Stanford-Binet — yield a standardized IQ score with a mean of 100 and a standard deviation of 15, placing most individuals within a predictable bell-curve distribution.
  • Intelligence scores show high test-retest reliability and predict real-world outcomes such as academic achievement, job performance, and health outcomes better than most other single psychological measures.
  • Both genetic and environmental factors shape measured intelligence; heritability estimates from twin studies range roughly from 50–80% in adults, yet Flynn Effect data show IQ scores rising across generations, confirming that environment substantially influences test performance.
  • Achievement tests measure mastery of a specific body of knowledge or skills already taught, whereas intelligence (aptitude) tests estimate a person's capacity to learn or perform in the future — a distinction central to how each type of test is used in schools and clinical settings.
  • Stereotype threat, test bias, and socioeconomic disparities in access to enriched environments are documented sources of group differences in test scores, meaning observed score gaps do not reflect fixed differences in underlying cognitive potential.

Defining Intelligence: Competing Theoretical Frameworks

Before any test can measure intelligence, researchers must decide what intelligence actually is — and that question has generated significant theoretical debate over more than a century of study.

Spearman's General Factor (g)

  • Charles Spearman noticed in the early 1900s that people who scored high on one cognitive test tended to score high on others, suggesting a shared underlying ability he called g, or general intelligence.
  • Spearman used factor analysis — a statistical technique that identifies clusters of correlated variables — to extract g as the common variance across verbal, spatial, and numerical tasks.
  • According to Spearman's model, specific abilities (called s factors) also exist for individual task types, but g accounts for the largest portion of performance differences between people.

Cattell-Horn Fluid and Crystallized Intelligence

  • Raymond Cattell and John Horn split g into two broad components: fluid intelligence (Gf), the ability to reason through novel problems without relying on prior knowledge, and crystallized intelligence (Gc), the accumulated knowledge and verbal skills gained through experience and education.
  • Fluid intelligence peaks in early adulthood and declines with age; crystallized intelligence remains stable or grows across the lifespan, which explains why older adults often outperform younger adults on vocabulary tasks but lag on timed reasoning tasks.

Gardner's Theory of Multiple Intelligences

  • Howard Gardner proposed in 1983 that intelligence is not a single capacity but at least eight distinct intelligences: linguistic, logical-mathematical, spatial, musical, bodily-kinesthetic, interpersonal, intrapersonal, and naturalistic.
  • According to Gardner's theory, each intelligence has its own developmental trajectory and neural basis, and traditional IQ tests only capture two of the eight (linguistic and logical-mathematical).
  • Gardner's framework is influential in education but remains contested among psychometric researchers, who argue that the proposed intelligences are not truly independent and overlap substantially with g.

Sternberg's Triarchic Theory

  • Robert Sternberg's triarchic theory identifies three forms of intelligence: analytical (academic problem-solving), creative (generating novel ideas and solutions), and practical (applying knowledge effectively in real-world contexts, sometimes called 'street smarts').
  • Sternberg argued that standard IQ tests assess only analytical intelligence, leaving creative and practical abilities unmeasured despite their importance for success outside formal schooling.

Major Intelligence Tests: Design, Scoring, and Interpretation

Intelligence tests translate theoretical constructs into standardized, scored instruments, and understanding their design is essential to interpreting what scores actually mean.

Standardization and the IQ Scale

  • A test is standardized by administering it to a large, representative normative sample and then scaling scores so that the population mean equals 100 with a standard deviation of 15.
  • This means approximately 68% of the population scores between 85 and 115, and roughly 95% falls between 70 and 130 — creating the familiar bell-curve distribution of IQ scores.
  • Standardized IQ scores are called deviation IQ scores because they express how far an individual's performance deviates from the age-matched average, replacing older ratio-based calculations.

The Wechsler Scales

  • David Wechsler developed a family of individually administered tests — the WAIS (adults), WISC (children 6–16), and WPPSI (preschool children) — that are among the most widely used intelligence instruments worldwide.
  • Each Wechsler scale yields a Full Scale IQ and four index scores: Verbal Comprehension, Perceptual Reasoning, Working Memory, and Processing Speed, giving clinicians a detailed cognitive profile rather than a single number.
  • The Wechsler scales are frequently used in clinical neuropsychological assessments to detect learning disabilities, ADHD, traumatic brain injury effects, and intellectual disability.

The Stanford-Binet Intelligence Scale

  • The Stanford-Binet originated with Alfred Binet's 1905 test designed to identify French schoolchildren who needed additional academic support — one of the earliest practical intelligence tests ever created.
  • Lewis Terman at Stanford adapted and extended Binet's work, introducing the intelligence quotient (mental age ÷ chronological age × 100) and promoting the test's use in the United States.
  • The current fifth edition of the Stanford-Binet measures five cognitive factors — fluid reasoning, knowledge, quantitative reasoning, visual-spatial processing, and working memory — across both verbal and nonverbal domains.

Reliability and Validity of IQ Tests

  • Intelligence tests demonstrate high test-retest reliability, meaning scores remain relatively stable when the same individual is tested again weeks or months later.
  • They show strong predictive validity: IQ scores correlate with academic achievement (r ≈ 0.5), job performance in complex occupations, and even health and longevity outcomes.
  • However, a high correlation between IQ and an outcome does not mean IQ is the sole cause; socioeconomic status, educational opportunity, and motivation all contribute independently.

Achievement Testing: Measuring Learned Knowledge and Skills

Achievement tests serve a fundamentally different purpose from intelligence tests — they evaluate what a person has already learned in a specific domain rather than estimating general cognitive potential.

Distinguishing Achievement from Aptitude

  • An achievement test assesses mastery of a defined body of content — for example, a final exam in algebra, an Advanced Placement exam in U.S. history, or a licensing exam in nursing — and scores reflect what instruction and study have produced.
  • An aptitude test (including most IQ tests) attempts to estimate future learning capacity or potential performance, not current knowledge, making it a tool for prediction rather than certification of learning.
  • In practice, the boundary between the two types blurs: the SAT, originally designed as a pure aptitude measure, was redesigned in 2016 to align more explicitly with high school curriculum content, moving it closer to an achievement test.

Standardized Achievement Testing in Education

  • Large-scale standardized achievement tests such as the National Assessment of Educational Progress (NAEP), state proficiency exams, and international assessments like PISA allow comparisons of student performance across schools, districts, and countries.
  • These tests use item response theory (IRT) to calibrate item difficulty so that scores are comparable across different versions of a test administered in different years.
  • High-stakes achievement testing — where scores determine grade promotion, graduation, or school funding — has generated policy debate about whether test pressure narrows curriculum and disadvantages students from under-resourced schools.

Criterion-Referenced vs. Norm-Referenced Achievement Tests

  • A norm-referenced achievement test ranks a student relative to peers in the normative sample, producing percentile ranks rather than a judgment of mastery.
  • A criterion-referenced achievement test compares a student's performance to a fixed performance standard (e.g., 'demonstrates proficiency at grade-level reading'), regardless of how other students performed.
  • Most state accountability exams are criterion-referenced because policymakers want to know what percentage of students meet a defined proficiency threshold, not merely who scores above or below average.

Unlock the rest of this study guide

  • Access the full study pack
  • Track your mastery and be test-day ready
  • Upload your own notes to build personalized study guides, quizzes, flashcards, and more
Sign up free →

About this Study Pack

Created by Kibin to help students review key concepts, prepare for exams, and study more effectively. This Study Pack was checked for accuracy and curriculum alignment using authoritative educational sources. See sources below.

Sources

More in Psychology 101

See all topics →

Browse other courses

See all courses →