A Globusz Books discovery
Word Embeddings Quantify 100 Years of Gender and Ethnic Stereotypes
Nikhil Garg et al. · English
Ever wonder how deep stereotypes run in the language we use every day? Or how much the words in old newspapers or books might reveal about society’s biases over the last century? Turns out, you can actually measure prejudice and progress with math—if you’re willing to let machines crunch the words for you.
Source-grounded summary
What the book is about
Let’s face it: language isn’t just words tossed around to get a point across. It’s a mirror reflecting what we think, believe, and sometimes what we’re too polite to say out loud. The 2018 study by Nikhil Garg and colleagues dives headfirst into this mess by using word embeddings—fancy machine learning tools that turn words into points in a high-dimensional space—to track how gender and ethnic stereotypes have morphed in American English over the last hundred years.
Word embeddings aren’t some new magic trick. They’re mathematical maps that capture how words relate to each other based on context. If words like “nurse” and “woman” hang out close together in this space, the model suggests a stereotypical association. Garg and his team took these embeddings trained on massive text collections spanning from the early 1900s to the 2010s and asked: what do these relationships tell us about societal biases over time?
Here’s where it gets interesting. They didn’t just eyeball word clouds or count word frequencies. They married these embeddings with actual demographic and occupational data from the U.S. Census. This means they could see, for example, whether adjectives like “emotional” or “ambitious” shifted in their closeness to “women” or “men” as women’s roles in society evolved. Or how jobs traditionally linked to certain ethnic groups changed in the cultural imagination.
The results are a bit like social history told by a machine. During the 1960s and 70s, as women pushed for equality, the embeddings showed noticeable shifts in how language described gender. Words associated with women began to reflect more diverse and complex roles, moving away from purely domestic or passive traits. Similarly, the study tracks the changing stereotypes around Asian Americans, especially as immigration patterns and social attitudes shifted.
What’s clever about this approach is that it quantifies something notoriously slippery: bias. Instead of relying solely on surveys or historical analysis, the embeddings provide a continuous, data-driven measure of how stereotypes wax and wane. This gives social scientists a powerful new tool to study prejudice and progress.
But it’s not all sunshine and roses. The study’s biggest limitation is its dependence on the texts it analyzes. If the source material skews toward certain voices—say, mainstream newspapers or literature—it might miss or misrepresent minority perspectives. Also, focusing solely on the U.S. means these findings don’t necessarily translate elsewhere, where language and social dynamics differ wildly.
Another catch: the embeddings show correlation, not causation. They reveal that stereotypes changed, but not why or how exactly those changes happened. Was it activism? Economic forces? Shifts in media? The method can’t answer that. It’s a powerful lens, but not a crystal ball.
Still, this work lands at a pivotal moment in AI and social science. It sounds a warning: if machines learn from our language, they also learn our biases. Understanding how these biases evolved helps us spot them in algorithms today. It’s a reminder that technology isn’t neutral; it’s a reflection of the messy, biased world we live in.
In short, Garg et al. offer a fresh, data-heavy way to track the slow crawl of social change through the words we use. It’s a tool for historians, linguists, and tech folks alike—if you’re willing to look beneath the surface of everyday language and accept that progress is neither linear nor perfect.
Beyond the plot
What might this book awaken in you?
This study is a smart reminder that the words we use aren’t just harmless chatter—they’re loaded with history, prejudice, and progress all tangled up. Machines can help us spot these patterns, but they can’t fix the mess. Understanding the past encoded in language is useful, but it’s just one piece of the puzzle in tackling bias today.
Before you commit
Why you might read this
Ever wonder how deep stereotypes run in the language we use every day? Or how much the words in old newspapers or books might reveal about society’s biases over the last century? Turns out, you can actually measure prejudice and progress with math—if you’re willing to let machines crunch the words for you.
Themes worth noticing
Language as social mirror
Language reflects collective attitudes, prejudices, and values, making it a powerful tool to study social change.
Bias in technology
Machine learning models inherit biases from their training data, raising ethical concerns and the need for critical scrutiny.
Interdisciplinarity
Combining computational methods with social sciences enriches understanding and opens new research frontiers.
Historical change and continuity
Stereotypes evolve slowly, with periods of progress and regression, reflecting complex social dynamics.
Questions to carry with you
- How much of what I say or think is shaped by hidden stereotypes in language?
- Can machines ever truly understand the social context behind words, or just mimic patterns?
- What voices are missing when we analyze history through preserved texts?
- How do we balance the power of data-driven insights with the complexity of human experience?
- In what ways does our current technology reflect—and reinforce—the biases of our past?
Continue the journey
Read the original when you are ready.
Globusz Books helps you decide whether a book deserves your time. This public-domain work can also be read free at Project Gutenberg.
