Translating animal studies to humans is harder to measure than assumed
Promising animal results that fail in humans: it is one of the biggest frustrations in biomedical science. But how do you actually measure whether that translation has succeeded? It turns out to be far more difficult than assumed.
Researchers conducted a simulation study to evaluate nine widely used metrics for assessing translation success. The study, published in eLife, used parameters from a meta-analysis on amino acid supplementation in pregnant women as a starting point. A total of 648 scenarios were then simulated, varying effect sizes, heterogeneity between studies, and sample sizes.
The findings are sobering. No single method performed consistently well across all conditions. Most metrics controlled false-positive rates adequately only when there was little variability between studies. As heterogeneity increased, too many false-positive conclusions slipped through. One metric, based on meta-analysis, too readily indicated translation success even when strong evidence came from only one of the two species.
Small animal studies undermine translation power
A second finding is directly relevant to longevity research. So-called translation power, the probability of a true positive outcome, was strongly limited by the weakest link. A small, underpowered animal study dragged down the entire chain, even when the human study was well designed. This partly explains why longevity interventions that work in mice so rarely succeed in humans.
The researchers recommend using multiple metrics simultaneously and carefully examining the assumptions underlying each one. The skeptical p-value that controls overall type-one error and the weighted version of Edgington’s method performed most consistently, but even these are not universally applicable.
A methodological problem with practical consequences
This may sound abstract, but the implications are concrete. If we cannot reliably measure whether animal results are translatable, pharmaceutical companies and research institutions invest in clinical trials that are destined to fail. For longevity interventions, where human evidence is already scarce, this is an urgent problem that deserves more attention.
Want to research this yourself?
Search for example:
- translational validity animal studies
- translation power clinical research
- skeptical p-value replication