- Introduction
- Chapter 1 The Birth of the Test: Carl Brigham and the Early SAT
- Chapter 2 The Logic of Similarity: Why Analogies Were Chosen
- Chapter 3 Building the Vocabulary Engine: How the Questions Were Written
- Chapter 4 Inside the Psychometrician's Lab: The Science of Difficulty
- Chapter 5 RUNNER is to MARATHON: Deciphering the Bridge Sentence
- Chapter 6 The Golden Age of the Analogy: 1950–1980
- Chapter 7 Coaching the Uncoachable: The Rise of the Test Prep Industry
- Chapter 8 Stanley Kaplan’s Revenge: Cracking the SAT Code
- Chapter 9 Class, Race, and Cognition: The Cultural Bias Debate
- Chapter 10 The Oarsman and the Regatta: The Question That Sparked a Controversy
- Chapter 11 Standardized America: How Analogies Shaped College Admissions
- Chapter 12 The Linguistics of the SAT: Why Rote Memorization Failed
- Chapter 13 Flashcards and Word Roots: The Subculture of Prep
- Chapter 14 The Psychometric Backlash: Critics Challenge the Analogies
- Chapter 15 Gender Disparities and the Verbal Section
- Chapter 16 The University of California Threat: Richard Atkinson's Ultimatum
- Chapter 17 Inside the College Board: The Great Debate Over Reform
- Chapter 18 Redesigning the Rite of Passage: The Decision to Eliminate
- Chapter 19 The Class of 2005: Facing the New SAT
- Chapter 20 Post-Mortem of a Question Type: What Replaced the Analogy
- Chapter 21 The Legacy of Verbal Aptitude: Did Admissions Change?
- Chapter 22 The Nostalgia of the Analogy: A Generation’s Shared Trauma
- Chapter 23 Lost Art of Verbal Reasoning: Is Vocabulary Instruction Dying?
- Chapter 24 The Digital SAT: The Evolution of Assessment in the 21st Century
- Chapter 25 A is to B: The Lasting Imprint of the Analogy on the American Mind
The Rise and Fall of SAT Analogies
Table of Contents
Introduction
For nearly eight decades, a stark, minimal string of text served as an unmistakable rite of passage for millions of American teenagers: RUNNER : MARATHON. Printed in crisp black ink across the newsprint pages of the SAT, these paired words were more than a test item. They were a cultural icon, a psychological instrument, and for generations of college-bound students, a source of profound anxiety. To solve an analogy was to perform a specific kind of intellectual acrobatics—to discern the invisible relationship between two concepts, abstract that logic, and map it onto an entirely new set of terms. For middle-class America in the twentieth century, mastering this formula was seen as the key to unlocking the gates of higher education and securing a place in the meritocracy.
The analogy was the pure, distilled essence of what early test designers called "scholastic aptitude." Unlike history or biology, it did not explicitly test what a student had been taught in a high school classroom. Instead, it purported to measure raw mental agility, verbal reasoning, and the capacity for abstract thought. For decades, the College Board and the Educational Testing Service (ETS) defended analogies as the ultimate leveler—a tool that could spot brilliant minds in underfunded rural schools just as easily as in elite New England prep academies. In the mid-twentieth century, as college admissions transformed from an informal network of social privilege into a massive, standardized sorting machine, the analogy became the machine’s most precise gear.
Yet, behind the promise of pure measurement lay a complex and increasingly volatile history. The very characteristics that made analogies psychometrically elegant also rendered them deeply controversial. What test developers viewed as neutral measures of relational thinking, critics increasingly identified as mirrors of cultural and socioeconomic privilege. A question hinges not just on logic, but on the precise, nuanced definitions of words—words that were far more likely to be spoken in wealthy suburban households than in working-class communities. When a single test item like OARSMAN : REGATTA could penalize students simply because they had never encountered elite collegiate sports, the illusion of pure aptitude began to shatter.
This book is the story of how a simple question format rose to define American meritocracy, ignited a battle over civil rights and educational philosophy, and ultimately met its demise. It traces the journey of the analogy from its origins in the early twentieth-century intelligence tests of Carl Brigham, through the golden age of post-war higher education, to its explosive role in the test-prep arms race led by figures like Stanley Kaplan. It explores the boardroom debates, psychometric breakthroughs, and public controversies that culminated in 2005, when the College Board finally excised analogies from the SAT altogether under immense pressure from university leaders and reformers.
The Rise and Fall of SAT Analogies is not merely a nostalgic look back at a retired question type, nor is it just a history of standardized testing. It is an examination of how America attempted to quantify intelligence and engineer a fair society through words. By examining the life cycle of the analogy, we gain a clear window into our evolving ideas about talent, privilege, race, gender, and the purpose of education itself. Whether you remember analogies with fond nostalgia, lingering resentment, or simple curiosity, this story reveals how two pairs of words separated by colons helped shape the modern American mind.
CHAPTER ONE: The Birth of the Test: Carl Brigham and the Early SAT
In June of 1926, precisely 8,040 high school students sat down in drafty classrooms across the United States to take an experimental, three-hour exam. They were handed test booklets, sharpened pencils, and a set of instructions unlike anything most American teenagers had ever encountered. The exam was called the Scholastic Aptitude Test. It was the brainchild of a young, intense Princeton University psychology professor named Carl Campbell Brigham, and its arrival marked the beginning of a quiet revolution in American education.
To understand why Brigham created the SAT, one must first look at the chaotic landscape of college admissions at the turn of the twentieth century. In the late 1800s and early 1900s, applying to college was an unpredictable and largely regional affair. Elite Ivy League institutions, such as Harvard, Yale, and Princeton, maintained their own proprietary entrance exams. These tests were heavily focused on traditional, classical curricula: translation of Latin and Greek passages, knowledge of ancient history, geometry, and English literature.
This system served a distinct social purpose. It catered almost exclusively to elite private boarding schools—institutions like Phillips Exeter, Andover, and Groton—whose entire curricula were tailored to match the specific entrance criteria of elite universities. A bright student attending a rural public high school in Ohio or a municipal school in Iowa had virtually no chance of passing these exams, simply because their schools did not offer the necessary years of Greek verbs or classical rhetoric. Higher education was not a sorting mechanism for national talent; it was a finishing school for the American establishment.
By 1900, secondary schools and colleges were desperate for standardization. That year, representatives from twelve prominent colleges met in New York to form the College Entrance Examination Board. The idea was simple: instead of every college setting its own test, the Board would create a uniform set of essay examinations given simultaneously across the country. A student could take one set of tests, and the results could be submitted to any participating university.
While the essay-based College Board exams brought order to the admissions process, they did not alter its fundamentally exclusionary nature. The tests still measured achievement in specific, high-level academic subjects. They rewarded students whose families could afford private prep schools or rigorous classical tutoring. Furthermore, grading tens of thousands of essay hand-written exams was slow, expensive, and notoriously subjective. Readers grading essay exams on hot summer afternoons in Manhattan were prone to wild inconsistencies. A paper that received an "A" from one professor might easily receive a "C" from another.
Enter Carl Brigham and the burgeoning field of psychological measurement. Born in Massachusetts in 1890, Brigham was a brilliant and ambitious scholar who came of age during the dawn of experimental psychology. He earned his doctorate at Princeton in 1916, immersing himself in the new science of psychometrics—the quantitative measurement of mental traits, capacities, and processes.
When the United States entered World War I in 1917, the American Psychological Association saw an unprecedented opportunity to demonstrate the practical utility of their young discipline. Led by Robert Yerkes, a team of prominent psychologists, including Brigham, gathered at Vineland Training School in New Jersey to construct standardized intelligence tests for the U.S. Army.
The military faced a massive logistical challenge: it needed to rapidly evaluate, classify, and assign positions to nearly two million recruits. The psychologists developed two primary instruments: the Army Alpha, a written test for literate recruits, and the Army Beta, a pictorial test for illiterate or non-English-speaking soldiers. The Army Alpha included a variety of subtests, ranging from numerical problems and logical sequences to vocabulary items and primitive verbal relationships.
The wartime testing program was a staggering logistical success. Millions of soldiers were categorized in a matter of months. For the psychologists involved, the army data represented the largest collection of human intelligence measurements in history. Brigham immersed himself in analyzing this mountain of data, drawing conclusions that would shape the rest of his career—and eventually cause him immense regret.
In 1923, Brigham published a controversial book titled A Study of American Intelligence, based on his analysis of the Army Alpha data. In it, he argued that the test results proved that native-born, white Americans possessed higher intellectual capacity than recent immigrants from Southern and Eastern Europe, as well as African Americans. Brigham’s work was quickly embraced by the eugenics movement and cited in congressional debates that led to the restrictive Immigration Act of 1924.
However, Brigham’s thinking was about to undergo a profound evolution. As he continued to analyze psychometric data and refine his testing methodologies through the mid-1920s, he realized that the Army Alpha test had not measured innate, biological intelligence at all. Instead, it had measured familiarity with American culture, length of residence in the United States, and access to formal education. Brigham possessed a rare academic trait: when confronted with undeniable evidence of his own error, he publicly recanted his previous theories. By the end of the decade, he repudiated A Study of American Intelligence, disavowing the racial conclusions he had drawn and acknowledging the profound influence of environment and schooling on test performance.
What remained unchanged, however, was Brigham’s belief in the power of standardized psychometric tools to measure abstract cognitive abilities. He envisioned a new kind of test—one that would bypass the superficial advantages of elite prep school curricula and identify raw scholastic talent wherever it existed.
In 1924, Princeton began using an experimental psychometric test developed by Brigham for its incoming freshman classes. Impressed by the test's ability to predict academic performance, the College Entrance Examination Board invited Brigham to lead a committee to design a national test. The goal was to create an exam that could supplement the existing essay-based achievement tests, providing colleges with a uniform, objective measure of a candidate's underlying mental aptitude.
Brigham and his committee set to work constructing the first Scholastic Aptitude Test. From its inception, the test was designed to be radically different from traditional classroom exams. It was not a test you could cram for by memorizing historical dates or chemical formulas. Instead, it aimed to measure general mental capacity—what Brigham called "scholastic aptitude"—through speeded, highly structured logical and verbal exercises.
The original 1926 SAT consisted of nine distinct subtests, designed to measure various facets of cognitive functioning:
- Definitions
- Arithmetical Problems
- Classification
- Artificial Language
- Antonyms
- Number Series
- Analogies
- Logical Inference
- Paragraph Reading
The exam was heavily weighted toward verbal reasoning, reflecting Brigham’s conviction that mastery of language and the ability to manipulate abstract verbal concepts were the primary ingredients of academic success.
To modern eyes, the 1926 exam looks like an eclectic mix of logic puzzles and linguistic experiments. The "Artificial Language" section, for instance, required students to learn a fictional set of grammatical rules and vocabulary within minutes, and then translate sentences into the imaginary tongue. The "Number Series" subtest asked students to complete numerical patterns, a forerunner of modern IQ test items.
Yet tucked away as Subtest 7 was a format that would outlive almost all the others: Analogies.
Brigham did not invent the verbal analogy. The format had been used in laboratory psychology experiments in Europe during the late nineteenth century to study mental association. It had also appeared in early intelligence scales, including the Binet-Simon scale in France and the Army Alpha test in America. But Brigham recognized that analogies possessed a unique set of psychometric properties that made them exceptionally suited for high-stakes testing.
In its 1926 iteration, the analogy subtest looked somewhat different from the streamlined format that would later become famous. The items often required students to identify relationships across multiple words or select missing components from a broader array of choices. But the core mechanics were already present: students were asked to hold two concepts in their mind, identify the precise structural relationship between them, and recognize an equivalent relationship in another pair of words.
The rationale behind including analogies in the 1926 SAT was rooted in the psychological theories of the era, particularly the work of British psychologist Charles Spearman. Spearman had argued that human intelligence consisted of a general factor—g—which represented the fundamental capacity to perceive relationships and deduce correlates. Analogies were viewed as the purest possible operationalization of this capacity. They stripped away extraneous context and forced the brain to perform pure, relational logic using words as the medium.
The initial administration of the SAT on June 23, 1926, was viewed as a grand scientific experiment. Test booklets were shipped to 318 test centers across the country and overseas. Students were given strict time limits for each subtest—often as little as ten or twelve minutes per section—creating a high-pressure environment where speed and accuracy were paramount.
When the scores were tallied, Brigham and his team analyzed the results with rigorous mathematical tools. They evaluated each subtest not only on how well it predicted a student’s future college grades, but also on how internally consistent it was. If high-performing students consistently missed a particular question while low-performing students got it right, that question was discarded as psychometrically flawed.
Through this process of statistical Darwinism, the nine subtests began to compete for survival. Some sections proved too cumbersome to score, while others failed to correlate reliably with academic performance. The "Artificial Language" section, while novel, was deemed too artificial to offer meaningful long-term predictive value. "Classification" and "Logical Inference" subtests often suffered from subtle ambiguities that frustrated statistical modeling.
Analogies, however, performed exceptionally well. They showed high internal reliability: a student who excelled at one analogy was very likely to excel at others. More importantly, analogy scores correlated strongly with college performance across a wide range of disciplines, from humanities to hard sciences. They provided a clean, elegant spread of scores, effectively separating students along a bell curve without requiring complex, multi-page reading passages.
In the years following the 1926 launch, Brigham worked tirelessly to refine and streamline the SAT. Operating out of his laboratory at Princeton, which would eventually evolve into the Educational Testing Service's research infrastructure, Brigham transformed test development from an informal academic exercise into an exacting science.
By the late 1920s and early 1930s, the structure of the SAT was shifting. Brigham recognized that testing nine separate subtests in a single sitting was inefficient and exhausting for candidates. The test began to converge into two primary domains: mathematical aptitude and verbal aptitude.
As the verbal section was condensed, the item types were winnowed down to the most efficient and reliable formats. By 1930, the test was restructured into a more cohesive battery, and by the mid-1930s, the core components of the verbal SAT had crystallized. Among them, analogies emerged as a primary pillar of the exam, taking their place alongside antonyms, sentence completions, and reading comprehension.
Brigham’s vision for the SAT was deeply idealistic, despite his problematic early work on intelligence. He truly believed that a standardized test of scholastic aptitude could serve as an instrument of social mobility. In his view, American high schools varied wildly in quality, grading standards, and local resources. An "A" grade from a rural school in the Midwest did not mean the same thing as an "A" from an elite preparatory school in Massachusetts.
By offering a test that measured abstract reasoning rather than specific content mastery, Brigham hoped to create a national, objective metric. A brilliant young mind working on a farm in Kansas or in a working-class neighborhood in Philadelphia could sit for the SAT, score in the 90th percentile, and catch the eye of an admissions director at Harvard or Columbia.
In this vision, the analogy was the ultimate meritocratic tool. It did not matter if a student’s high school lacked a Latin teacher, a advanced physics laboratory, or a library stocked with European classics. If that student possessed the mental agility to discern that a CONDUCTOR was to an ORCHESTRA as a DIRECTOR was to a CAST, they possessed the fundamental aptitude required for higher education.
However, Carl Brigham himself remained deeply skeptical of over-relying on the very test he created. As the SAT gained popularity throughout the 1930s, Brigham frequently warned the College Board against treating test scores as absolute measures of human worth or innate intellect. He resisted the idea that a single number could fully capture a mind, and he strongly opposed converting SAT scores into rigid cutoffs for admissions decisions. He viewed the test as an indicator of developed aptitude—a snapshot of cognitive skills shaped by both nature and nurture—rather than a static measurement of fixed intelligence.
Brigham died in 1943 at the age of fifty-two, leaving behind an intellectual legacy that would fundamentally reshape American higher education. He did not live to see the post-World War II college boom, the passage of the G.I. Bill, or the massive expansion of the SAT from an experimental tool for a few thousand students into a national rite of passage for millions.
Nor did he live to see how the simple verbal analogy, which he helped pluck from psychological laboratories and embed into the 1926 exam, would become the defining feature of the test he designed. What began as Subtest 7 in an experimental battery was about to become the premier instrument for measuring the American mind.
This is a sample preview. The complete book contains 27 sections.