One of the things I remember most about my grade-school math classes was the teacher’s insistence on showing my work. The answer at the bottom of the page mattered less than how I produced it, because a right answer reached by an unrepeatable method could just be a lucky guess, and a lucky guess is worthless on the next problem.
Hold that standard in mind, because this year 865 researchers applied it to the social sciences… and the social sciences bombed disastrously.
The effort, called SCORE and funded by the Defense Advanced Research Projects Agency (DARPA), spent seven years examining claims from roughly 3,900 papers published between 2009 and 2018 in 62 journals spanning economics, psychology, education, sociology, political science, business, and more. When the results landed in Nature this spring (on April Fools’ Day, hilariously enough), the headline was shocking: only about half of social-science studies could be replicated.
When independent teams reran the studies with new participants and new data, the original findings held up about as often as a coin lands heads. That number is obviously bad. But a quieter companion paper tells an even worse story, one that received a fraction of the above headline’s attention.

Replication asks whether a finding survives a brand-new study (with new data not used in the original). Reproducibility asks something far more modest: if I use your data and your method, will I get the same answer you asserted? No new experiment, no new sample, and no interpretive judgment call. The same numbers go through the same analysis and get checked against the published result. It is the scholarly equivalent of my teacher checking how I arrived at the answer on my math worksheet or, in other words, the lowest bar a quantitative claim can be asked to clear.
The reproducibility study drew a random sample of 600 papers. The first obstacle was simply obtaining the evidence. Authors, it turns out, had made their data available for only about a quarter of them. Three in four published findings, in other words, could not even be checked, because the people who made the claims would not or could not produce the data. In education research, data were available for only 2.9 percent of papers!
Among the papers that could be tested, only about half reproduced precisely (meaning the reanalysis matched what had been published). Economics and political science fared best, at roughly two-thirds. Across the remaining fields, precise reproduction fell to 38 percent. Education (the discipline whose findings shape how tens of millions of American children spend their days) produced the most arresting result. Of the handful of education papers the team could test, half could not be reproduced at all, and the fewest could be reproduced precisely.
I should note that a failed reproduction does not prove a finding false: code gets lost, documentation is thin, and researchers attempting reproduction make mistakes. Yet consider what that concession implies. If a published result cannot be verified from the author’s own materials, then nobody knows whether it is true, often including the author.
Here lies the heart of the matter. Science is (when followed correctly) a method for earning the right to hold a conclusion. Its authority comes from exposure to challenge: I tell you what I did, you do it yourself, and reality adjudicates between us. If you remove the possibility of checking, what remains is authority without accountability. A field that cannot reproduce its own work is not doing science. It’s purveying academic propaganda.
Why do so many credentialed people tolerate such a low standard? The answer is less mysterious than it seems. Academic careers are built on publication, and publication rewards the novel, the striking, and the sympathetic. A bold finding about the latest pedagogical fashion, or about some fresh injustice in need of a government program, earns citations, grants, and press releases. A careful null result earns a desk drawer and barely a media mention. Who wants the latter? Certainly no aspiring academic aiming to climb the ivory tower’s ladder.
Peer review—the process meant to catch errors—is, of course, performed by peers: often drawn from the same small circles, sharing the same assumptions, and rarely asking to see the data. The SCORE team found that as of 2018, just 6.8 percent of journals outside economics and political science required authors to share data, share code, or submit to a reproducibility check. Reviewers were effectively grading others’ homework without ever asking to see the work. Some “peer review,” that.
Then there’s replication. In the studies SCORE reran (i.e., with new data), the typical effect shrank to less than half of what the originals reported. Any policy built on the original estimate was built on a finding inflated more than twofold. This scandal effectively creates a double-hit for taxpayers who had to pay once to fund the shoddy research, and pay again to live under the policies and programs these findings were used to justify.
Scandalous as this is, it doesn’t require a conspiracy. It requires only ordinary human nature operating inside a system that rewards confidence without auditing it. We are social creatures. We defer to credentials, to consensus, to the solemn phrase “peer-reviewed”—and those who traffic in those phrases know it. Propaganda has always worked best when it borrows the vocabulary of truth. John Oliver covered this pretty well a decade ago:
Yet the same research that exposes the rot also points to the remedy, which is almost embarrassingly simple. Papers published in journals that required data sharing were precisely reproduced about 70 percent of the time; papers in journals with no such policy, about 40 percent. (70 percent is no “A,” but hey, it’s an improvement…) And journals are catching on: by mid-2025, roughly 94 percent of economics and political science journals in the sample had adopted at least one transparency policy, along with 43 percent of journals in other fields.
Incentives shaped the problem, and incentives can reshape it. So change them. Any study funded by taxpayers, directly or indirectly, whether in whole or in part, should publish its data and code as a condition of the grant. Any finding cited in legislation or regulation should be reproducible by an independent party before it is used to justify a new policy or program. And any journalist tempted to write “a new study finds…” should note whether the study has been reproduced, let alone replicated, by anyone else.
For the rest of us, there’s a simple question we need to utilize far more frequently: how do you know? Ask it of any expert asserting something as fact. Teach children to ask it, early and often, to dismantle their deference and strengthen their critical thinking capacity. A citizenry that demands evidence is harder to govern by shocking headlines and propaganda press releases.
Eight-year-old Connor had to learn the rule: show your work. It is long past time we held the people who presume to instruct today’s children, and the rest of us, to the same standard.





