
Big Data, New Data, and What the Internet Can Tell Us About Who We Really Are
Seth Stephens-Davidowitz · 2017 · Psychology
Original summary · AI-drafted, human-published · added by Library
Everybody Lies argues that people routinely mislead pollsters, surveys, and even themselves about sex, prejudice, and other sensitive behavior, but their anonymous searches on Google and other platforms reveal what they actually think and do. Drawing on billions of search queries, Stephens-Davidowitz treats this data as a new kind of confession booth, using it to measure hidden racism, closeted sexuality, and parental bias, while also showing where big data's predictive power breaks down without careful causal reasoning.
Pick a finish date and Genius lays out the days — the plan shows today's target and keeps you honest.
Start a circle and share the code — everyone sees everyone's honest place in the book. Accountability, not leaderboards.
- Marketers and researchers who want a rigorous alternative to unreliable survey self-report. - Curious readers interested in what large-scale data reveals about prejudice, sexuality, and parenting bias. - Data scientists and economists looking for a readable case for pairing big data with causal experimentation.
Self-reported survey data is systematically unreliable because people misrepresent sensitive behavior to pollsters and to themselves.
Anonymous Google searches function as a more honest data source than surveys because searchers have no audience to perform for.
Geographic patterns in racist search terms predicted anti-Obama voting better than any survey measure of racial attitudes.
Search and behavioral data suggest a meaningfully larger closeted gay population than surveys capture, and its size varies with local tolerance.
Parents' private search behavior reveals gender bias about their children's intelligence and appearance that they would not admit on a survey.
Other digital platforms beyond search, including forums, adult sites, and dating services, provide behavioral data that Google alone cannot supply.
Big data's apparent predictive power collapses without theoretical grounding, as shown by Google Flu Trends' well-documented failure.
Raw correlational patterns in big data are not enough for decision-making; they require experiments or natural experiments to establish cause and effect.
Behavioral datasets can anticipate harmful or economically significant events earlier than official statistics, when built and tested carefully.
Seth Stephens-Davidowitz is an economist trained at Harvard, a former Google data scientist, and a contributing writer to The New York Times. He built his reputation mining search-engine and other digital exhaust for evidence about attitudes people hide from pollsters, and he consults widely on applying these methods to business and social science questions.