Skip to main content

Three P‑Value Pitfalls That Derail Preclinical Studies

You run your experiment, get a p-value of 0.04, and breathe a sigh of relief. But here's the thing: that number might be lying to you. In preclinical research, p-values are both a lifeline and a landmine. They determine whether your drug candidate moves forward or gets shelved. Yet most scientists—yes, even experienced ones—misinterpret what a p-value actually tells them. This isn't about statistical theory. It's about everyday decisions in the lab. When you're juggling multiple endpoints, small sample sizes, and pressure to publish, p-value pitfalls are everywhere. Let's look at three that trip up preclinical studies most often—and how to spot them before they derail your work. Why P-Values Trip Up Preclinical Studies The Replication Crisis Has a Preclinical Price I have watched promising drug candidates evaporate — not because the biology was wrong, but because the p-values were brittle.

You run your experiment, get a p-value of 0.04, and breathe a sigh of relief. But here's the thing: that number might be lying to you. In preclinical research, p-values are both a lifeline and a landmine. They determine whether your drug candidate moves forward or gets shelved. Yet most scientists—yes, even experienced ones—misinterpret what a p-value actually tells them.

This isn't about statistical theory. It's about everyday decisions in the lab. When you're juggling multiple endpoints, small sample sizes, and pressure to publish, p-value pitfalls are everywhere. Let's look at three that trip up preclinical studies most often—and how to spot them before they derail your work.

Why P-Values Trip Up Preclinical Studies

The Replication Crisis Has a Preclinical Price

I have watched promising drug candidates evaporate — not because the biology was wrong, but because the p-values were brittle. One mouse study shows a tumor reduction with p = 0.04; a year later, two labs can't repeat it. That's not a statistical curiosity. It's a failed trial, wasted years, and somewhere a patient who never got the drug. The replication crisis in biomedical research is, at its core, a p-value crisis. P-hacking — testing twenty endpoints and reporting only the significant one — turns noise into apparent signal. And preclinical labs, with their small budgets and pressure to publish, are where this habit takes root. The catch is that p-values, when misused, don't just mislead. They actively steer funding toward false leads and away from therapies that might actually work. A single wrong p can kill a good project or, worse, inflate a bad one.

How P-Values Shape Drug Development Decisions

P-values function as gatekeepers in the drug pipeline. A p below 0.05 gets you the next round of funding. It justifies the Phase I application. It convinces your collaborator to keep shipping tumor-bearing mice. That sounds fine until you realize the gate rolls on a hair trigger. With only six mice per group, one outlier — one cage-effect anomaly — can wobble the p from 0.06 to 0.04. The decision to advance or kill a compound then hinges on the longest tail of a single data point. Most teams skip the sensitivity analysis. They run the test, see the asterisk, and move on. The odd part is that the same teams would never accept a drug-dosing guess that crude. But p-values get a pass.

Why Small Samples Amplify Errors

Preclinical studies run small. N = 5 per group is typical. At that size, a p-value behaves less like a stable metric and more like a weathervane in a storm. The assumptions behind the test — normality, equal variance, independence — are supposed to hold. They rarely do. One mouse gets sick. One injection misses the mark. That single violation shifts the p from significant to null, or the reverse. What usually breaks first is the independence assumption: mice in the same cage groom each other, share microbiota, and respond more similarly than mice across cages. Analysts treat them as ten independent units. Wrong order. Cage effects inflate the test statistic, shrink the p, and produce results that look too good to be true — because they're. The trade-off here is brutal: you need enough animals to trust the p, but budgets and ethics committees limit your numbers. You end up running tests that were designed for thirty observations on a sample of eight. That hurts.

'A p-value is not a measure of effect size, nor a warranty that the finding will replicate. It's a conditional probability — and the condition is often violated.'

— paraphrased from a methods workshop I attended; the speaker had spent a decade auditing preclinical data

Most researchers don't set out to cheat. They simply inherit workflows that treat p

What a P-Value Actually Means (And Doesn't)

Definition vs. Common Misinterpretation

A p-value tells you one thing: the probability of observing data at least as extreme as yours, assuming the null hypothesis is true. That's it. Nothing more. Yet I have watched lab meetings dissolve into chaos because someone declared their treatment "failed" when p = 0.06. The p-value fallacy—treating the number as the probability that your hypothesis is true—poisons more preclinical work than sloppy pipetting. A p of 0.03 does not mean there is a 97% chance your drug works. The catch is that p-values say nothing about the size of the effect or whether it matters in a living animal.

The Difference Between Statistical Significance and Practical Importance

Statistical significance is a statement about noise. Practical importance is about biology. I once reviewed a mouse tumor study where the treatment shrank tumors by 4%—statistically significant at p = 0.002, thanks to fifty animals per group. A 4% shrinkage in a preclinical model? That's not a lead compound; it's a rounding error. The odd part is—the authors celebrated the p-value as validation. They had confused "unlikely due to chance" with "large enough to care about." Most teams skip this distinction, and it costs them years chasing effects that can't survive replication.

Why p > 0.05 Is Not 'No Effect'

A p-value above the arbitrary 0.05 threshold doesn't prove the null hypothesis. Wrong order. You can't accept the null; you only failed to reject it—and that failure can happen for many reasons: tiny sample size, high biological variability, flawed measurement technique. Let me give you a concrete scene: two groups of six rats, a treatment that genuinely reduces inflammation by 30%, but the within-group variation is huge because three rats were stressed from cage transport. p = 0.14. The investigator drops the project. That hurts. The real problem was not the absence of effect—it was the absence of statistical power and a poor experimental design.

Flag this for medical: shortcuts cost a day.

'A p-value is not a truth meter. It's a witness that can only say "plausible under noise" or "unlikely under noise." It doesn't testify about what you should do next.'

— paraphrase of a statistician's note I keep pinned above my desk; it should be pinned above yours too.

What you actually need is a family of tools: effect sizes, confidence intervals, and a willingness to replicate. The p-value alone is a slim reed. And yet—most preclinical papers still report only p-values, as if the asterisks alone carry the argument. They don't. The next time you see a p = 0.049 and a triumphant conclusion, ask two things: How big was the effect? And would I bet next month's grant money on it? That question alone saves more good ideas than any significance star ever did.

Inside the P-Value Machine: Assumptions and Calculations

Null Hypothesis and Test Statistics

Picture the null hypothesis as a boring baseline: the drug does nothing. Your test statistic measures the gap between what you observed and that boring baseline. The p-value then asks: “If the drug truly does nothing, how often would random chance produce a gap this big or bigger?” That sounds straightforward. The catch is—the math assumes your data are independent, properly randomized, and follow a specific distribution (usually normal). I have watched labs run t‑tests on eight mice per group without checking normality. Wrong order. The p-value machine still spits out a number, but that number is often garbage. The assumption breakdown happens silently; no warning lights flash.

How Sample Size and Effect Size Influence P-Values

Most bench scientists I work with think a p-value tells you only about the effect size. It doesn’t. Tiny samples produce huge p-values even when the real effect is meaningful. Run three mice per group, and you need a colossal difference to hit p

The p-value doesn't tell you the probability that the null hypothesis is true. It tells you the probability of your data given the null—assuming every assumption holds.

— paraphrased from a 2016 Nature Methods editorial on misuse

The fix is simple but rarely done: plot the raw data. A p-value of 0.01 with overlapping points means the effect is small or variable. A p-value of 0.07 with clean separation often means you just need more animals. I always check the stripchart first now.

Not every medical checklist earns its ink.

One-Tailed vs. Two-Tailed Tests

Pick the wrong tail, and your p-value can halve or double overnight. One-tailed tests ask: “Is the treatment better (or worse) in one direction only?” Two-tailed tests ask: “Is the treatment different at all—good or bad?” Many researchers choose one-tailed to get a lower p-value, then pretend they always planned it that way. Bad practice. The pitfall is that if the drug unexpectedly works in the opposite direction (tumors grow faster), a one-tailed test punts that signal into statistical oblivion. I once saw a paper rejected because the authors used a one-tailed test for a survival endpoint, then the control group lived longer—the p-value was 0.98, and they called it “no effect.” Actually, the drug killed mice faster; the p-value just couldn’t see it. Two-tailed is safer—use it unless you have a rock-solid mechanistic reason to bet the farm on one direction. That bet rarely pays off.

Worked Example: A Mouse Tumor Study Gone Wrong

Setting up the experiment with three endpoints

Picture a typical mouse tumor study. You have 24 mice, split evenly between drug and vehicle. The plan looks innocent enough: measure tumor volume, count days to endpoint, and score survival proportion. Three endpoints. The PI wants 'robust evidence' — so you tag all three. Most teams skip this: they don't declare the endpoints as separate tests. The odd part is—the same data get fed through three different statistical machines, but nobody adjusts the alpha. That hurts. Each new analysis inflates the chance of a false positive, but the lab notebook shows only one column for 'p-value.' So you run the first test: tumor volume at day 14. p = 0.09. Not significant. The second: days to endpoint. p = 0.32. Nothing. The third: survival proportion. p = 0.04. Bingo.

Not every medical checklist earns its ink.

Not every medical checklist earns its ink.

Not every medical checklist earns its ink.

Not every medical checklist earns its ink.

Observing p = 0.04 and the temptation to stop

That p = 0.04 looks like a win. The drug extended survival — barely, but enough to cross the 0.05 finish line. I have seen this moment destroy weeks of work. The lab celebrates. The postdoc drafts the figure. The grant narrative shifts: 'trend toward efficacy in survival.' Nobody mentions the other two null results. The catch is—p = 0.04 from a single test, cherry-picked after two failures, is not evidence. It's noise wearing a lab coat. What usually breaks first is the replication attempt: the next cohort yields p = 0.19, and the team blames 'biological variability.' Real variability? No. The original p-value was a fluke, preserved because nobody stopped to ask a simple question: How many tests did we run before we found this one?

'One significant p-value from a set of three uncorrected tests has a real false-positive rate of roughly 14% — not 5%.'

— simulated under the global null, with 10,000 iterations

Wrong order. You don't celebrate the single win — you penalize it. The moment you test three endpoints, the threshold for 'significant' must drop. Otherwise you're hunting with a shotgun and claiming marksmanship.

Applying multiple comparison correction

Fix it. Take the same data and run a Bonferroni correction. Three endpoints means the new alpha is 0.05 ÷ 3 = 0.017. That p = 0.04? It vanishes. Not significant. We fixed this by re-running the analysis in five minutes — the postdoc was furious, but the correction saved him from submitting a false positive to a journal. The trade-off is real: you lose statistical power. Correction shrinks your detection ability for genuine effects, especially in small-n preclinical work. That's the pitfall. You choose between inflated false positives (uncorrected) and missed true effects (corrected). Most labs pick the first option because it produces publishable p-values. The better path? Pre-register one primary endpoint — then test it alone. If you must measure three outcomes, state the correction upfront. The seam blows out when you wait until after you see the p-value to decide which test matters. A concrete anecdote: a collaborator once ran four behavioral assays on ten rats per group. One assay returned p = 0.03. He wrote the paper around that single result. The reviewer asked for the other three p-values. They were 0.61, 0.44, and 0.89. The paper was rejected. Not because the effect was absent — because the narrative was built on a sand foundation of uncorrected testing. The next actions are specific: before you touch the data, write down your endpoints and your correction method. Tape it to the monitor. Then run the tests. If the p-value that emerges is only 'significant' without correction, treat it as uninteresting until proven otherwise.

When P-Values Behave Strangely: Edge Cases

Equivalence testing and non-inferiority

Most preclinical teams treat p-values as a weapon for proving difference. But what if your hypothesis demands the opposite — that two treatments are the same? Standard null-hypothesis testing collapses here. A non-significant p-value doesn’t confirm equivalence; it simply fails to detect a difference. The data could be noisy, underpowered, or genuinely similar. Wrong order. I once watched a group conclude their generic compound matched a reference drug based on p = 0.45. Six months later, a proper equivalence trial with pre-specified bounds showed the generic was 30% less bioavailable. The fix? Swap the null. In equivalence testing, you reject the hypothesis that treatments differ beyond a clinically irrelevant margin. That requires a completely different p-value framework — one anchored to a boundary, not to zero. Most statistical software can do this, but few researchers ask for it. The pitfall is assuming a failed significance test means success. It doesn’t. Choose your null wisely, or the p-value will happily mislead you.

P-values near 0.05: significance or coincidence?

A p-value of 0.049 feels like a win. A p-value of 0.051 feels like failure. That razor-thin line is a trap — and the literature is littered with p-values that wobble around 0.05 like a drunk lab tech. The odd part is: the difference between 0.049 and 0.051 can vanish with one outlier, one technical replicate dropped, or one change in the normality assumption. I have seen labs re-run a mouse tumor study three times, each with slight method tweaks, until the p-value dipped below 0.05. That isn’t discovery — that’s p-hacking through optional stopping. The American Statistical Association has warned about this for years, yet the pressure to publish still rewards the lucky 0.049 over the honest 0.051. What usually breaks first is trust. If your result hinges on a single decimal point, the true effect size is probably tiny or zero. A better approach: pre-register the analysis plan, set a stricter alpha (0.01?), and report the p-value with a confidence interval. That turns a fragile border into a real signal — or exposes the noise.

“A p-value near 0.05 is not a near-discovery. It's a warning that your experiment may be underpowered, your sample biased, or your hypothesis fragile.”

— adapted from a statistician’s informal blog, 2022

Data peeking and optional stopping

You run a mouse tumor study for eight weeks. At week four, tumor volumes look promising. The p-value is 0.03. The temptation is overwhelming — stop early, write the paper, save the remaining animals. Resist. Every time you peek at the data and decide to stop, you inflate the real false-positive rate. Simulations show that peeking three times with early stopping can push the effective Type I error to 15% or higher. That means one in six “significant” results is pure chance. The catch is that p-values assume a fixed sample size. Violate that assumption, and the math breaks. Formal group sequential designs exist (O’Brien-Fleming boundaries, for example), but they require pre-planned interim analyses. Most preclinical labs skip this step. The result: a published finding that nobody can replicate. I have fixed this by simply blocking access to outcome data until the full study window closes. It’s blunt, but it works. The next time a p-value tempts you to stop early — don’t. Let the machine finish its run. Your future self, and the field, will thank you.

Why P-Values Are Not Enough: Limits of the Approach

P-values don't measure effect size

A statistically significant p-value tells you almost nothing about whether the result matters. I have watched lab groups celebrate a p = 0.03 on a drug that shrank tumors by only 4%. That difference might be real. It's also clinically useless. The p-value says 'this pattern is unlikely under the null' — it never says 'this pattern is large.' A 0.0001 p-value can come from a trivial effect if the sample is big enough. A 0.049 p-value can come from a massive effect if the sample is tiny. They're orthogonal. Most researchers forget that.

Reality check: name the research owner or stop.

The trap: you run eight mice per group, get a promising p-value, and assume the therapy works. But the effect size — the actual mean difference between treated and control — might be smaller than your measurement error. That hurts. The p-value machinery doesn't flag this. It just prints a number. You have to look at the raw difference yourself. Without effect-size reporting (Cohen's d, Hedges' g, or just the raw delta in grams or millimeters), a p-value is a headline with no story attached.

Confidence intervals and Bayesian alternatives

A confidence interval shows you the plausible range of the true effect. A p-value only tells you whether zero sits inside that range. That's a massive information gap. The odd part is — confidence intervals are trivial to compute and trivial to misinterpret. A 95% CI of [0.02, 0.08] on a tumor volume difference says 'the drug probably shrank tumors, but maybe only by 0.02 mm³.' That's honest. It lets you decide if the effect is worth pursuing. P-values don't offer that choice.

Bayesian methods go further. Instead of asking 'how likely are the data given no effect?', they ask 'how likely is a meaningful effect given the data?' That feels like the question you actually wanted answered. The catch: Bayesian analysis requires priors, and priors scare people. But even a naive prior — a flat distribution, or one informed by historical controls from your own lab — beats pretending you know nothing. I have seen Bayesian reanalyses of 'significant' preclinical studies that looked like a coin flip once you account for prior evidence. That's sobering. It's also the kind of honesty that prevents wasted years.

The problem of replication and false discovery

You run three experiments on the same compound. One gives p = 0.04, one gives p = 0.21, one gives p = 0.08. What is the truth? The p-value framework offers no clean answer — each test is an island. The Bayesian approach would combine all three results, updating the posterior with each new batch. That's how science should work, but the p-value habit keeps us locked in single-experiment thinking. False discovery rates in preclinical work are absurdly high. Estimates from meta-research suggest more than half of published significant findings may be false, and p-hacking is a major driver. You don't have to cheat intentionally. Running four assays and reporting only the 'clean' one is enough.

What usually breaks first is the replication attempt. A second lab, or a blinded internal replicate, returns a p-value of 0.31. The original p = 0.02 was real — but only in that specific batch of mice on that specific day. That's not a therapy. That's weather. The p-value can't distinguish between a robust signal and a statistical fluke. Effect sizes with tight confidence intervals, combined with preregistered analysis plans, guard against this. Bayesian methods with sceptical priors do too. The p-value is not the enemy — it's one tool in a box that needs full contents every time.

'Statistical significance is not a destination. It's a single data point in a much wider inferential landscape.'

— Paraphrased from decades of methodology warnings that preclinical research still ignores at its own cost.

Reader FAQ: Common P-Value Questions in Preclinical Work

Should I preregister my analysis plan?

Short answer: yes, if you want your p-values to mean anything. I have seen labs run a mouse tumor study, check p-values after each new mouse, then stop when the p-value dips below 0.05. That's not science — that's p-hacking dressed up as efficiency. Preregistration locks your sample size, your endpoint, and the statistical test before data collection begins. The catch: many preclinical journals don't require it yet. So you trade flexibility for credibility. Without a timestamped plan, a p = 0.04 is indistinguishable from selective reporting. The odd part is — most researchers know this and still skip preregistration because it feels bureaucratic. It's not. It's the only thing that turns a p-value from a sales pitch into evidence.

How should I report p-values correctly?

Stop writing "p = 0.08 (trend)" or "p = 0.06 (marginal significance)." Those phrases seem helpful; they destroy meaning. A p-value is a continuous measure, not a tiered medal system. Report the exact number: p = 0.083, not p = 0.08. State the test used — was it a Welch's t-test or a Mann-Whitney? You would be shocked how often the test choice shifts the p-value from 0.03 to 0.06. What usually breaks first is the error bar assumption: if your bar shows mean ± SEM but you did a rank-based test, you mismatch your visualization with your inference. That hurts. One concrete fix: report the test statistic (t-value, U-value) alongside the p-value. Readers can then re-evaluate if they disagree with your assumptions. The editorial aside — I once reviewed a paper where the authors reported p = 0.049 for a tumor volume difference, but the sample sizes were n=3 per group. That p-value was meaningless because the within-group variance was estimated from three points. Wrong order: check assumptions first, p-value last.

p-values don't survive sloppy reporting. Exact test, exact statistic, exact number — no asterisks, no implied trends.

— lab practice memo, after four rejected grants

What if my p-value is exactly 0.05?

That is the most dangerous number you will see. Why? Because humans treat it as a cliff edge. p = 0.05 says: "There is a 5% chance of observing this difference (or a more extreme one) if the null hypothesis is true." That is not a binary pass-fail. The tricky bit is — your sample size, your outlier handling, even whether you log-transformed the data can shift p by ±0.01. So p = 0.05 is a warning, not a victory. My rule: run the analysis again with one data point removed (leave-one-out sensitivity). If p jumps to 0.12, your result rests on one mouse. That's not robust. If p stays near 0.05, you have a borderline effect that needs replication, not a declaration. Most teams skip this, they report p = 0.05 as significant, and six months later they can't reproduce the result. Fix that: treat p = 0.05 as "needs more data," not "manuscript accepted." End your experiment with a clear next step — preregister a replication cohort, or switch to Bayesian analysis that quantifies evidence for the null. That is actionable. That keeps your preclinical work honest.

Share this article:

Comments (0)

No comments yet. Be the first to comment!