Jaimee T. Mendillo
Standardized testing is a powerful and controversial force in American public education. Although often regarded as an objective way to measure student achievement, these assessments have a complex and troubling history intertwined with scientific racism, structural inequality, and the development of public schooling. To fully understand the role of standardized testing today, educators must examine its historical context. This essay provides background for the unit Challenging the Standard: Testing & the Fight for Educational Justice by tracing the development of public education, the institutionalization of standardized testing, and both historic and contemporary critiques of its impact on students, schools, and society.
The Origins and Evolution of Public Education
Before public school systems were formally established, education in the United States was highly localized, informal, and unequal. Access depended heavily on race, class, and geography, with instruction provided through church-supported schools, charity schools, dame schools, and home tutoring.1 Many children, especially Black, Indigenous, and poor white children, were excluded altogether. Education was neither compulsory nor universally funded.
Some of the Founding Fathers, including Thomas Jefferson and John Adams, viewed education as vital to sustaining American democracy. In 1785, John Adams wrote: “The education of a nation, instead of being … for the instruction of the few, must become the national care and expense, for the information of the many.”2 He advocated for free, public education to cultivate informed citizens capable of protecting their rights and participating in civic life. Federal land grants in the 1780s further laid the financial groundwork for school development in new states, reinforcing the principle that education should serve the public good.3
Horace Mann’s common school movement of the 1830s advanced this vision, promoting universal, nonsectarian, publicly-funded education to foster literate, moral, and productive citizens. Reformers argued that education could uplift the poor and unify a diverse population. Yet in practice, access remained limited. Many advocates excluded non-white children in their visions of universal education, and public schools often reinforced rather than disrupted existing social hierarchies.4
The expansion of public schools was uneven. In 1830, only 55% of children aged 5 to 14 were enrolled in school; by 1870, that number rose to 78%. High school attendance remained low, with only 14% of adults having attended by 1910.5 Despite growing enrollment, access continued to be shaped by race, gender, class, and ability. Southern states prohibited the education of enslaved people and later enforced segregation through Jim Crow laws. Immigrant, Latinx, Asian American, and female students faced systemic barriers nationwide, and children with disabilities were often placed in separate programs or denied meaningful schooling altogether.6
Despite persistent inequities, public schools became spaces of civic education and community life. They also emerged as critical battlegrounds in the struggle for equity. Landmark moments like Brown v. Board of Education (1954) and the Elementary and Secondary Education Act (1965) placed public education at the center of federal efforts to redress inequality and promote opportunity. Still, the decentralized structure of U.S. education – with overlapping federal, state, and local control – often complicates these efforts.7
Today, public school districts are legally obligated to educate all students “regardless of income, race, ethnicity, academic level, disability, immigration status, language proficiency, or other characteristics.”8 However, in practice, many students do not receive the education they need.
This history reveals a persistent contradiction: schools have been imagined as instruments of democracy and equal opportunity, yet they often uphold systems of social and racial hierarchy. One key mechanism sustaining these hierarchies has been standardized testing, particularly assessments claiming to measure intelligence. Far from being neutral, these tests emerged from a broader eugenic effort to classify and control populations under the guise of science.
The Eugenic Roots of Intelligence Measurement
In the late 19th and early 20th centuries, intelligence testing in the United States developed alongside the rise of the eugenics movement. Historian Erik L. Peterson, defines eugenics as “scientific ideas that some [people] are inherently healthy, robust, smart, charismatic (the ‘fit’) – and others sickly, weak, mentally impaired, lazy, alcoholics (the ‘unfit’).”9 Eugenics aimed to improve the human population by promoting reproduction among the “fit” and discouraging, or even preventing, reproduction of the “unfit,” continuing an “American tradition of those with power controlling the bodies of others with less power.”10 This ideology appealed to social reformers, scientists, and policymakers who believed psychological measurement could provide a scientific basis for identifying, sorting, and controlling individuals according to perceived genetic worth.
In 1904, Alfred Binet and Théodore Simon developed the first widely known intelligence assessment in France to identify schoolchildren needing support. It was never meant to measure innate ability; Binet explicitly warned that intelligence was not fixed and that his test measured only a limited range of abilities.11 However, American psychologist Henry Goddard “convinced his medical colleagues to redefine mental deficiency in terms of intelligence” and repurposed the test to reinforce hereditarian views.12 Psychologist Lewis Terman modified the Binet-Simon into the Stanford-Binet IQ test. Robert Yerkes, president of the American Psychological Association and chair of the Eugenics Research Association’s Committee on Inheritance of Mental Traits, collaborated with Terman, Goddard, and others to promote the idea that intelligence was hereditary and quantifiable. Together, they adapted intelligence tests for mass administration and IQ testing was soon embraced by school officials, policymakers, and military leaders as a “scientific” method for sorting and categorizing people.13
Millions of immigrants, soldiers, and students were tested. Results were used to justify racial discrimination, forced sterilizations, and immigration restrictions. Tests placed disproportionate numbers of Black and immigrant children in “special education” programs and lower academic tracks, as schools began “organizing classrooms according to students’ mental – rather than chronological – ages.”14 IQ tests also supported laws enacted by at least thirty states to permit some level of forced or involuntary sterilization. The Supreme Court supported these sterilization policies in Buck v. Bell (1927), when Justice Oliver Wendell Holmes Jr. infamously wrote: “It is better for the world, if … society can prevent those who are manifestly unfit from continuing their kind.”15 IQ tests also informed the 1924 Immigration Act, which restricted Southern and Eastern European immigration based on perceived intellectual inferiority.16
While explicit eugenic rhetoric has faded, the underlying assumptions persist. Ideas about who is considered “gifted,” who is deemed “at risk,” and whose intelligence is worth cultivating still shape policies and practices. Understanding this history exposes intelligence testing as a tool of exclusion masquerading as meritocracy.
The Emergence of Standardized Testing in American Public Schools
Although advocates promoted standardized testing as a means to foster meritocracy, in practice, these assessments have reinforced social inequalities. Early IQ tests reflected the cultural knowledge and educational experiences of white, middle-class Americans, rather than innate intelligence. Marginalized groups predictably scored lower. These results were often misinterpreted as proof of inherent biological inferiority rather than evidence of systemic barriers to educational and economic opportunity.17
As Emily Merchant explains, “Intelligence testing became a big business after World War I, with universities and employers adopting intelligence tests as gatekeeping mechanisms for entering the middle class, making correlation among IQ, educational attainment, and socioeconomic status a self-fulfilling prophecy.”18 Standardized testing expanded beyond IQ tests to large-scale achievement exams promising objective and efficient measures of student performance. In reality, these tests replicated disparities in access to quality education, culturally relevant curricula, and economic resources.19
By the late 20th century, standardized testing had become central to school accountability. Consequences included curricula narrowed to match test content, increased levels of student stress, and high-stakes decisions based on limited and often irrelevant scores. These effects disproportionally harmed students of color, multilingual learners, and students with disabilities. As Terry Meier writes, standardized testing, rooted in white, middle-class norms, is “deeply at variance with the values and strengths of many minority communities.”20 Marginalized students were often labeled low-achieving and denied access to advanced coursework, gifted programs, and higher education. Standardized tests privilege specific forms of language, expression, knowledge, and values while penalizing others.
Critics like Meier have long questioned the validity and purpose of standardized testing, arguing that they fail to reflect the diverse ways students learn and demonstrate understanding.21 Cindy Long documents educators’ growing resistance to high-stakes testing policies that harm students in under-resourced schools.22 Although testing tools have changed, the legacy of eugenics, with its focus on ranking and sorting, still shapes definitions of intelligence and achievement.
Even Edward Thorndike, a pioneer of educational psychology, admitted in 1918 that “there is no known unit of intelligence,” and therefore, measuring it is “inordinately tricky.”23 In 1923, psychologist Edwin G. Boring unironically declared intelligence is “what intelligence tests test.”24 Despite there being no universally accepted definition of intelligence nor a clear method to measure it, standardized achievement tests continue to be treated as definitive indicators of student ability and remain central to educational policy and decision-making.
Although modern assessments are often framed as tools to promote equity and accountability, they still operate within framework rooted in early 20th-century eugenic thinking. Understanding this history contradicts the idea that standardized testing is a neutral or fair measure of potential. Instead, it reveals how these systems uphold dominant cultural norms and justify unequal outcomes, reaffirming the belief that merit is both measurable and inherently unequal.
High-Stakes Testing in the 21st Century
High-stakes testing continues to shape 21st-century educational policy, driving decisions about curriculum, teacher evaluation, school funding, and closures. Operating under the pretense of accountability, these tests assess school and student “performance” with profound consequences for those falling below arbitrary benchmarks. Despite evidence that “standardized tests are inaccurate, inequitable, and often ineffective at gauging what students actually know,” they continue to be treated as objective indicators of merit.25
Terry Meier describes the tragedy of defining excellence through such narrow means as a standardized test just as educational psychology embraces broader and more complex understandings of intelligence.26 Though modern assessments no longer claim biological inferiority, they replicate racial, economic, and linguistic hierarchies.
Initiatives like the Massachusetts Consortium for Innovative Education Assessment (MCIEA), a partnership of eight public school districts, counter this trend by building more comprehensive accountability systems. MCIEA co-founder Jack Schneider notes that test data mainly “correlates to race, income, and family educational achievement,” revealing more about advantage than achievement. 27 Yet schools still adjust curricula based on this data, often cutting enrichment programs in favor of test prep.28 The result is a narrowed curriculum that undermines equity and engagement. Additional consequences include denied student promotion, penalized teachers, and closed schools.29
Policies like No Child Left Behind (2001) and Race to the Top (2009) embedded testing into federal accountability measures. Though intended to close achievement gaps, they often deepened inequities by holding marginalized communities accountable for systemic conditions without addressing root causes. Meaningful change must come from grassroots movements that challenge standardized testing and reimagine assessments grounded in skill development and growth.
Resistance and Critique
A 2018 panel during Black Lives Matter in Schools week at the Institute for Collaborative Education in New York City, titled A Standardized Test is a Poor Substitute for Justice exemplifies the grassroots advocacy reshaping the national conversation around assessment. Organized by parents, the event featured educators and a student reflecting on the injustices of standardized testing and the validity of alternative assessments. Educator Kristin Taylor shared how federal mandates reshaped her curriculum and expectations for student proficiency. She expressed frustration that high-stakes standardized testing contradicts her school’s motto: “We meet the child where they are.”30
The student panelist, Chris Lopez, spoke about his inability to afford expensive SAT test-preparation programs that helped prepare his more economically-advantaged peers. He reflected on how his SAT scores diminished his confidence in the college application process and his relief that there are test-optional colleges.31 Another way to challenge the detrimental standardized tests is for parents to opt their children out of them. Zipporiah Mills, a retired principal, emphasized the power of parental action in opposing such harmful practices: “By not standing up against something, you’re supporting it.”32
In 2012, Wayne Au and Melissa Bollow Tempel edited a collection of essays titled Pencils Down: Rethinking High-Stakes Testing and Accountability in Public Schools. In their introduction, they criticized the intensified role of high-stakes testing under No Child Left Behind: despite rhetoric supporting multiple measures to assess learning and teaching, “high-stakes test scores are being used to quantify, rank, and judge everything in public schools.”33 The essays in Pencils Down “deconstruct the damage” caused by standardized tests, “highlight their inaccuracy as tools of measurement, and offer visionary forms of assessment that are more authentic, democratic, fair, and accurate.”34
Educators nationwide are implementing alternative assessments that prioritize growth, reflection, and authentic demonstrations of knowledge and skills in context and over time. These models include portfolios, performance-based assessments, exhibitions, and collaborative projects. These alternatives aim to foster learning and re-center assessment around equity, engagement, and deeper understanding. National organizations like FairTest support this shift by working to end the misuse of standardized tests and promote fair, educationally sound evaluations systems.
Reimagining Assessment for Equity and Justice
As communities confront the harms of high-stakes testing, they also envision new assessment measures rooted in growth, context, and student voice. MCIEA educators Alissa Holland and Dan Cote describe how performance-based assessment shifts the focus of learning. Holland explains, “Students can show much more of their knowledge over time, in a form they choose, rather than on one day, on one test question, where only one answer is right.” Cote adds that the process “acknowledges that learning happens in different stages, where we can give feedback and build skills during assessment.”35
At the Institute for Collaborative Education in New York City, Jehan Senai Worthy’s students develop real-world skills through performance-based assessment tasks. Her students become historians by conducting research, constructing arguments grounded in evidence, and presenting their knowledge in writing and through panel presentations.36 Chris Lopez credits his experience with such research projects and panel presentations as the reason he excelled in his college interviews and earned a full-ride scholarship to his chosen college.37
These performance-based approaches reflect the vision of assessment advanced by educators and scholars in Pencils Down. The book’s final section, “Beyond High-Stakes Standardized Testing,” advocates for authentic forms of evaluation. Effective methods of assessment exist, though they are “messier, more intensive, more democratic, and less punitive.”38 These methods prioritize meaningful growth and understanding over ranking and affirm the knowledge and dignity that students bring into the classroom.
In his 1963 “A Talk to Teachers,” James Baldwin reminded educators that “the purpose of education…is to create in a person the ability to look at the world for himself.”39 Reimagining assessment is part of that purpose. When students help shape how their learning is evaluated, assessment becomes a practice of liberation. By confronting the past and designing new models for the future, educators and students can create systems that honor learning in all its complexity.