Exposome Perspectives Blog

Do Geneticists Dream of Cloned Sheep? The Case for Replication Science

In this blog post, Dr. Wright highlights how a lack of incentives has driven a scientific replication crisis. To rebuild public trust, he advocates for establishing a dedicated discipline of "replication science" to standardize how research is understood.

Exposome Perspectives Blog by Robert O. Wright, MD, MPH

Reality denied comes back to haunt.

Philip K. Dick (from Flow My Tears, the Policeman Said, Doubleday, 1974.)

The novel Do Androids Dream of Electric Sheep? is the source material for the Blade Runner movies. It describes AI robots called “replicants” that look and act so much like humans that we have difficulty distinguishing them from us—and, as it turns out, they may not be able to distinguish them from us either. The author, Philip K. Dick (PKD), was fascinated by the idea of replicating human consciousness. You may not have read his work, but you likely saw a movie based on it, from Blade Runner to Minority Report to Screamers.

His themes include “What is the nature of truth and reality?” and “What has consciousness and is therefore alive?” His replicants are near perfect imitations of humans, manufactured using bioengineering and corporate patents. In this alternate world, mankind carefully considered human replication with artificial intelligence (AI). We tried to reproduce exactly how intelligence and consciousness are created. The replicants live a human life over many years. Unlike AI or the robots of today, which take monumental shortcuts, replicants learn about life by living, i.e., by having occupations, family, and co-workers. They experience love, hatred, envy, and friendships. PKD introduced a time-dependent fly into the ointment of AI, consciousness, and philosophy. However, this post is about replication, not AI.

Replication is a problem across all science

I had a couple of alternative titles for this post including: “Do Epigeneticists Dream of Elongated Giraffe Necks?” and “Do Climate Scientists Dream of Butterfly Wings Flapping?” I hoped to convey that scientific replication is not limited to genomics. The lay press seems to focus on social science not replicating its work while giving molecular biology a pass. The molecular scientific fields have the exact same problem, exposomics included. The largest issue is that the incentives to replicate a study are nonexistent, especially when considering how science operates. Therefore, capital “S” Science hasn’t created a set of best practice guidelines for when and how replication research should be done—it’s all ad hoc. There are myriad reasons studies don’t replicate, and it’s not just because of “junk” science or malfeasance. True findings in studies sometimes won’t replicate, and until we take replication seriously, the public (and scientists) will remain confused about what the truth actually is.

First, we shouldn’t expect that true findings will always replicate because they won’t. Context, in other words, matters. A study in Moscow on air pollution and diet probably won’t replicate in upstate Michigan even if all the variables are measured similarly. The same is true for a genomic study in Moscow vs. Michigan. Genomics is not immune from confounding and context. There are many interesting questions embedded in replication research, particularly the role of variable environmental and genetic context which we presently ignore.

Twenty years ago, John Ioannidis famously wrote a provocative article in the style of a scientific proof entitled: “Why Most Published Research Findings are False.” He listed small sample sizes, multiple comparisons with omics assays, small effect sizes in large observational studies, contextual dependencies, information/selection/confounding bias, chance, financial interests, and “trendy” research topics all contributing to publishing false results. Fraud, while present, is likely only a small part of the problem. Perhaps the rise of omic research, which generates mostly false positive results, together with biotech commercialization of research and its propensity to hype findings prematurely, helped to create the replication crisis, which in turn is contributing to the erosion of public trust in scientific findings. “Science” bemoans misinformation generated by non-scientists, but we need to get our own house in order to reply with authority. The replication crisis is a form of misinformation, and this didn’t happen overnight.

This issue is bigger than we care to admit

Replication research is rare because the research funding incentives heavily favor the status quo of big data, commercially motivated hype, and never replicating a published finding—especially if you got a press release. In the current age of omic science, grant reviewers are encouraged to focus on “innovation” as a review criterion. Imagine submitting a grant proposal repeating someone else’s prior work and competing against a multi-omic proposal with millions of genetic, epigenetic and proteomic measures, and sophisticated technology. The latter will seem sexy and modern compared with the staid, simple former. As long as “multi-omics” is the king of science, it will be better to propose big and “innovative” new technology studies over replicating a previously published finding. We value hypothesis-free discovery science over hypothesis-testing replication science—which means we don’t value hypothesis testing at all. Genomics itself was not the 21st century’s revolution in science; discovery research was; but discovery (i.e., omics) is a means to an end, and not an end itself. We’ve forgotten that the goal of the revolutionary “omic” discovery/replication paradigm was to eventually test a hypothesis.

Nothing lasts forever—even research findings

Biology is contextual and time-varying. The world of 2026 is wildly different from 1986. It may even be impossible to replicate a study conducted in 1976 because the conditions no longer exist—even if the people still do. The demographics, environment, average age, economics, industries, entertainment options, etc., are different, and those differences may be affecting results. Even at the individual level, the passage of time changes everything. We may have the same DNA sequence at age 6 as at age 80, but our bodies and minds are wildly different and react to our environment differently. Cross-sectionally, the parameters of a study conducted in Salt Lake City may not translate to Philadelphia. In today’s world, no one can plausibly make a career out of replication research. However, the public needs results that are trustworthy to influence health care, policy, and everyday lifestyle choices. This is not a trivial issue.

Where are the incentives?

Not only is replication hard to do, we don’t seem to want to do it. I already mentioned the higher weight given to discovery research over replication research in terms of perceived innovation. NIH guides its reviewers to prefer something new over something “old” like replication. (It’s not clear if they prefer something borrowed over something blue though).

Here are two more reasons the replication crisis exists:

1) Publication bias: Because primarily positive studies get published, we publish a disproportionate number of false positive studies that shouldn’t replicate. This is compounded by the dominance of “omic” science (including exposomics), which measures thousands to millions of “things” (DNA, RNA, proteins, metabolites, chemicals etc), creating a Bayesian nightmare of false positive results that cannot be fixed by statistical corrections alone. Replication can’t keep up because it requires an independent population or experiment, but that runs into problems of budget and lack of perceived innovation.

2) Money and prestige: A “discovery” research finding that ultimately turns out to be false—particularly one that may lead to drug development—is often touted as a “game-changing disruption” that will cure the incurable and fix societal ills. Reporters are easy prey to scientists with multiple degrees prone to hyperbole. Try selling a replication study to a major media outlet.

Practice makes perfect, or at least a little bit better

Replication doesn’t have to be repetition. We can make the replication study better, like a PKD replicant. If we understand replication science, perhaps we can even improve on an important finding (larger sample size, better measurements, targeted population, etc.). For this to be operationalized, there need to be NIH review criteria that favor replication. We don’t have such incentives. Reviewers may even demand a replication study be done in the application, even though there are no results yet. Maybe it’s not needed? We need solutions on “how to replicate research” so we can separate the true from the false. There are lots of reasons that a study won’t replicate—being a false finding is just one of them.

True findings also may not replicate, and right now we can’t always tell the difference unless it’s obvious.

We should direct resources towards organizing a field that creates a set of principles and systems to address the ideal ways to replicate research—factoring in the highly contextual nature of science, the rise of false positive results from omic science, publication bias, and the all-too-human characteristic of self-promotion. At present, it is haphazard, unevenly conducted, and not at all valued. We need to incentivize replication.

Failed replication is complicated and multifactorial—confounding bias, unmeasured gene–environment interactions, mismatched covariates between discovery and replication cohorts, cohort effects across time, population stratification, insufficient statistical power, measurement error, and even chance play a role. There is also the influence of money. Corporate interests can back studies that seem designed to fail, and we need criteria to objectively evaluate the quality of replication. Being funded by industry doesn’t guarantee bias, but industry would be better served if their replication studies were held to a rigorous standard.

Today, we lack a systematic scientific framework dedicated to studying replication itself. What conditions are needed to justify funding a replication study (i.e., importance, actionability, etc.)? Some “negative” findings may even need replication—but what are the criteria? What constitutes a good replication study vs. a poor one? Can we compare the bias in the first vs. the replication study, i.e., studies can be set up to succeed or set up to fail—we need a set of criteria to evaluate if the replication study is relatively unbiased. We also need to address the question of when we translate findings into interventions or policies. We can’t have endless replication either. I’m a little tired of reading, “More research is needed.” When is it enough?

Replication science can address the problem directly

These gaps suggest the need for a new discipline: Replication science—a field devoted to understanding when, why, and under what conditions scientific findings reproduce, partially reproduce, or fail to reproduce. Replication science would create experts in conducting and evaluating replication research and address three key questions: 1) if replicated, should we believe results; 2) if not replicated, should we believe results; and 3) what factors drove the success or failure of the replication results? The results and methods of all similar studies won’t just be compared; they can even be pooled or included in a meta-analysis to see how variable the effect is. We already know a lot about how to do this, but we lack an organizational framework for replication.

Replication science would lead to new tools to aid replication research—like calculating the impact of different levels of confounders or effect modifiers between populations. We would have experts who could help us discern false-positive, false-negative, and biased research. In the age of misinformation, social media, and corporate interests that sow distrust, we desperately need a field that can distinguish truth from chance and biased untruth. Replication science, if done well, can help the media and, in turn, the public discern good vs. bad vs. chance vs. intentionally false. Right now, we don’t take it seriously. We also don’t see that there are many interesting questions embedded in replication.

There are even precedents for creating replication science. Translational science emerged out of the observation that biomedical discoveries were not making it into real-world patient care. Replication is part of the same bench-to-bedside/policy continuum. Let’s give it the respect it deserves. Let’s create a field of replication science and study how to do it.