Skip to main content

Morphium’s Guide to Avoiding Sample Size Mistakes in Early Trials

You’re designing an early-phase trial. The protocol is half-written, the endpoint is solid, but there’s a number that keeps you up at night: How many patients do I need? Too few, and a promising drug looks like a dud. Too many, and you burn cash you don’t have. Sample size isn’t just a statistic; it’s a strategic bet. And in early trials — where uncertainty is high and data is thin — that bet can break your program. This isn’t a math lecture. It’s a practical guide for people who have to decide, by next week , what number to put in the protocol. We’ll cover what’s out there, what to look for, and what to do after you choose. Who Decides — and When the Clock Starts Ticking The decision-makers: PI, sponsor, biostatistician The number sits between three chairs — and somebody has to own it.

You’re designing an early-phase trial. The protocol is half-written, the endpoint is solid, but there’s a number that keeps you up at night: How many patients do I need? Too few, and a promising drug looks like a dud. Too many, and you burn cash you don’t have. Sample size isn’t just a statistic; it’s a strategic bet. And in early trials — where uncertainty is high and data is thin — that bet can break your program.

This isn’t a math lecture. It’s a practical guide for people who have to decide, by next week, what number to put in the protocol. We’ll cover what’s out there, what to look for, and what to do after you choose.

Who Decides — and When the Clock Starts Ticking

The decision-makers: PI, sponsor, biostatistician

The number sits between three chairs — and somebody has to own it. In early trials I have watched the principal investigator assume sample size is “just a stats thing,” the sponsor assume it's baked into the protocol template, and the biostatistician assume someone already picked one. Wrong order. The PI brings clinical sense: how big a difference would matter to patients? The sponsor brings portfolio pressure: can we afford 40 subjects or only 20? The biostatistician brings math: given the anticipated variability, what number keeps the false-negative rate below 20%? The catch is — these three rarely talk to each other until the clock runs out.

Timeline pressure: protocol submission deadlines

Sample size is not an infinite coffee-break thought experiment. It lands on an actual line-item in the protocol, which lands on a regulatory submission date. That deadline is usually pinned before anyone asked “How many patients do we need for a reliable signal?” Most teams skip this: they treat sample size as a late-stage fill-in-the-blank. But the ethics committee wants a justification, the IRB wants a statistical plan, and the site coordinator wants a recruitment target. If you fix the number the week before submission, you're fixing it wrong. I have seen a phase II protocol go to IRB with a sample size copied from a completely different disease — embarrassment, delay, 14 lost days.

“You don't pick a sample size. You inherit one from the calendar.”

— observed by a regulatory consultant after a third protocol revision

Consequences of delaying the choice

That hurts. When the decision is postponed, the number gets made by default — the budget cap, the chief’s gut feeling, the number that “looks right” from the last study. The odd part is: those defaults are rarely challenged. Yet a wrong number early wastes money, masks a real effect, or buries a promising compound. The flip side is also real — picking too early without variability estimates forces a formula based on a wild guess. So when does the clock actually start ticking? The moment you define your primary endpoint. Not before. Not after. That's the narrow window where the decision lives — and if you miss it, the number picks itself.

Three Routes to a Number: Formula, Simulation, Adaptive

Classic formula-based approach (power analysis)

Most teams start here. You plug an expected effect size, a desired power (usually 0.80), and an alpha into a closed-form equation — et voilà, a sample size appears. The math is settled: Cohen’s d for continuous endpoints, Fleiss for proportions, Schoenfeld for survival data. I have used these in five-minute coffee-break calculations. They work — until they don’t.

The catch is hidden in the inputs. Formula-based power analysis assumes your data will behave: normal distributions, equal variances, no dropouts, no weird correlation structures. One early-phase trial I consulted on used a textbook formula for a binary outcome; the nominal 80% power collapsed to 54% after the first interim look because the control event rate was half what the literature suggested. That hurts. The formula gave a number, not a guarantee.

Trade-off summary: fast, reproducible, cheap to compute — but brittle. A single wrong guess about standard deviation or effect size and you're flying blind. Use formulas when you have solid prior data from the same clinic, same patient population, same endpoint definition. Anything less and the number is a placeholder, not a plan.

Simulation-based methods (Monte Carlo, bootstrapping)

You write a model instead. Generate ten thousand virtual trials under plausible scenarios — varying dropout rates, effect sizes, seasonal recruitment dips — then count how often your test rejects the null. This is how we fixed that 54% disaster I mentioned. We simulated 20,000 trials with a distribution of control rates drawn from three historical studies. The required n jumped from 84 to 134 per arm.

The tricky bit is building the simulation. You need to specify data-generating mechanisms: a logistic regression for a binary outcome, a mixed model for repeated measures, maybe a frailty term for center effects. Most teams skip this: they reuse a script from a previous trial without checking whether the correlation structure matches. Wrong order. A bootstrap that samples from the wrong marginal distribution inflates power estimates the same way a broken speedometer shows 60 while you're stuck in traffic.

Trade-off: far more flexible than formulas — you can model non-compliance, staggered enrollment, even competing risks. But you pay in time and transparency. A colleague spent three weeks debugging a simulation that silently assumed a constant hazard when the actual failure rate accelerated after month six. That said, once validated, simulation gives you a power curve, not a single point. You see the trade-off between sample size and detectable effect — a luxury formulas rarely offer.

Adaptive designs with sample size re-estimation

You don't need to guess perfectly on day one — you can let the data adjust the number mid-flight.

— Design principle from group-sequential trials, applied to early-phase dose-finding

An internal pilot study. At a pre-planned interim look (say after 40% of patients are enrolled), you re-estimate the nuisance parameters — variance, control event rate — and recalculate the required total sample size. The final n adapts to what the trial actually sees, not what a literature review guessed six months ago. The odd part is: many researchers think this is cheating. It's not. The type-I error rate stays controlled if you pre-specify the re-estimation rule and use a blinded variance estimate.

Flag this for medical: shortcuts cost a day.

The pitfall? Adaptive designs demand tighter logistics. You need a data monitoring committee on standby, an unblinded statistician, and software that can run the re-estimation within days — not weeks. One team I know missed their interim window by two weeks because the CRO’s data entry lagged. By the time they had clean data, enrollment had overshot the original target. The adaptation was useless. A rhetorical question: would you rather have the wrong number with elegance, or the right number with operational chaos?

Trade-off: highest efficiency — you waste fewer patients and preserve power when initial guesses are off. But the complexity of implementation, regulatory scrutiny, and the risk of delayed analysis make this unsuitable for tiny, fast-moving Phase I/II trials. Reserve adaptive re-estimation for studies where recruitment timelines are predictable and a single wrong variance estimate would cost you months of re-randomization.

What to Look For: Criteria That Actually Matter

Statistical Power: 80 % vs 90 % — Which One Hurts Less?

The 80 % threshold is a habit, not a law. Too many teams grab it out of tradition—then watch a promising candidate fail because it never had enough shots on goal. The catch is simple: going from 80 % to 90 % can double your sample size. That sounds fine until your recruitment pipeline runs dry at month seven. I have seen a phase II trial limp along with 62 patients because nobody checked whether 90 % power would push enrollment past what the three local clinics could supply. The odd part is—80 % still works for low-risk repurposing studies. The real question is: what are you willing to miss? A 20 % false-negative rate means one in five real effects walks past you. For a first-in-human therapy where the animal data looked shaky, that risk might be acceptable. For a pivotal early trial that investors will read, 90 % often saves more time than it costs.

Precision of the Effect Size Estimate — Not Just a Number Game

Power tells you whether you’ll detect something. Precision tells you whether you’ll know how much. Most teams skip this: they pick a sample size based on a single guessed effect—say a 0.5 standard-deviation improvement—and never ask what the confidence interval around that estimate will look like. Wrong order. A trial sized for 80 % power might still produce a 95 % confidence interval so wide that the result is useless for planning the next phase. You get a p-value of 0.04 and a plausible range from “tiny” to “huge.” That hurts. Precision depends on variance, which you often overestimate from published data. The fix? Run three effect-size scenarios—optimistic, moderate, and pessimistic—and check how the interval width behaves. If the pessimistic scenario gives you a confidence interval that spans zero, your sample size is too small even if the power looks fine. One concrete anecdote: a colleague once fixed this by increasing enrollment by 18 % and halved the interval width. That small bump turned a borderline result into a clear go or no-go signal.

Feasibility Constraints — Budget, Recruitment, Timeline

“The perfect sample size is useless if you can’t recruit the patients.”

— Common remark at the end of failed trial reviews

What usually breaks first is not the statistics—it’s the calendar. A simulation might say you need 140 patients, but your principal investigator can only find 12 eligible subjects per month across four sites. That's 11 months of enrollment before you even get to follow-up. Most teams ignore the timeline constraint until the grant report is due. The trick: overlay your recruitment curve on the sample size plan. If the curve flattens after month six, you either lower the power, accept wider precision, or change the endpoint to something with less variance. Budget matters too—each extra patient adds monitoring costs, lab fees, and data-entry hours. I have watched a startup burn through its cash reserve chasing 90 % power when 80 % with a cheaper biomarker endpoint would have answered the same question. Feasibility is not a second-class criterion. It's the one that keeps your trial alive. Start with the hard numbers—recruitment rate per site, per-month cost, dropout history—then let those numbers pick your approach, not the other way around.

Trade-Offs at a Glance: Formula vs. Simulation vs. Adaptive

Cost and complexity

Formulas are cheap. A textbook, a calculator, maybe an afternoon — you're done. That sounds fine until you realize the formula assumes a world that barely exists: equal groups, no dropouts, effect sizes you pulled from a dream. The catch is hidden cost. You save time upfront but pay later in wrong numbers, underpowered trials, wasted doses.

Simulation costs more — a few days of coding, some compute credits, a person who actually knows R or Python. The odd part is: most teams skip this because it feels like overkill. But simulation lets you watch the trial fail before it runs. That's cheap compared to a phase II flop. Adaptive designs? They're the most expensive upfront — statisticians, software, protocol flexibility — and they demand pre-specified decision rules. You can't wing it mid-trial. However, if dropouts are high or effect sizes uncertain, adaptive approaches often save net time by stopping futile arms early.

What usually breaks first is not the math — it's the budget for the statistician. I have seen teams invest in a pricey platform but refuse to pay for the person to set it up. Wrong order. You need the brain, then the tool.

Accuracy and flexibility

Formulas give you one number. That's their strength — and their trap. They're accurate only when your assumptions are perfect. When are your assumptions perfect? Never. Simulation, by contrast, lets you stress-test: what if the control event rate is 30% instead of 25%? What if 15% of patients drop out? You can run 10,000 virtual trials and see the probability that your sample size still works. That's a different kind of accuracy — not precision on paper, but robustness in reality.

Adaptive designs trade the neatness of a fixed plan for flexibility. You can increase sample size if the effect is smaller than expected, or drop a dose that's obviously toxic. That flexibility requires iron-clad pre-planning; you can't peek at the data and then decide. The pitfall is over-flexing — adjusting too many things and losing interpretability. A trial that changes its primary endpoint, sample size, and treatment arms mid-stream is no longer a trial. It's a fishing expedition.

‘A formula gives you the answer to a question you probably got wrong. Simulation lets you ask better questions.’

— Researcher running a virtual phase II, after three failed power analyses

Not every medical checklist earns its ink.

Not every medical checklist earns its ink.

Not every medical checklist earns its ink.

Not every medical checklist earns its ink.

Not every medical checklist earns its ink.

Regulatory acceptance

Regulators accept all three — but not equally. The formula route is the oldest, most documented, and safest for standard designs. If your trial looks like a thousand before it, regulators nod. The rub: they also expect you to justify that similarity. Blindly plugging in a formula from a 1998 paper will get a request for more detail.

Simulation is increasingly accepted by FDA and EMA, especially for complex designs like platform trials or dose-finding. The key: show your code, your assumptions, and a sensitivity analysis. Hiding the mess behind a chart is a mistake. I have seen a regulator ask: “Why did you only simulate 500 trials?” — and the team had no answer. If you simulate, be ready to defend every knob you turned.

Adaptive designs face the highest scrutiny. Regulators want to see the pre-planned decision rules, the alpha-spending functions, and an independent data monitoring committee. A sponsor who says “we will look at interim data and see what makes sense” will be rejected. The good news? Once accepted, adaptive designs often get faster approvals — because you built the flexibility into the protocol, not patched it in later. That hurts to build but pays off.

After You Pick a Number: Implementing Your Choice

Lock It in the Protocol — or Lose It

You have a number. Now what? Most teams skip this: writing the why behind that number into the protocol. I have seen IRB submissions arrive with a single line — "Sample size = 60 per arm" — and no trace of how we got there. That's a time bomb. The rationale should sit in a dedicated subsection: expected effect size, assumed standard deviation, dropout rate, power target. Even a one-paragraph note forces you to check your own assumptions. The odd part is — you will catch half your mistakes just by typing them out. Don't rely on memory. Six months later, when the manuscript lands on a reviewer’s desk, "we ran a simulation in R" without archived code reads as fiction.

Documentation is not bureaucratic theatre. It's the brake pedal. Without it, you can't defend a mid-trial adjustment or explain a failed recruitment target. A concrete example: we once fixed a sample size based on a 15% dropout assumption. We wrote it down. When dropout hit 22% at interim, the stored rationale let us recalculate without re-litigating old arguments. That document saved two weeks of e‑mail ping-pong.

Dropouts, Missing Data, and the Hole in Your Plan

Planned sample size and evaluable sample size are rarely the same. The gap kills power faster than a wrong effect size. You need a contingency — not a vague "we will handle missing data with MICE."
What usually breaks first is the dropout budget. Most protocols inflate the target by 10–15% and call it done. That works until the dropout pattern is non‑random — sicker patients leaving, or placebo dropouts clustering. The fix? Make the inflation explicit: "We anticipate 12% attrition, so we will recruit 112 patients to yield 100 completers." Then add a trigger: if actual attrition exceeds 18%, pause enrolment and reassess. That sounds fine until your site is racing to meet a quarterly target. The catch is — pausing feels like failure. It's not. Proceeding blind feels worse when your primary endpoint ends up underpowered by 20 patients.

A second layer: define your primary analysis population upfront. Intention‑to‑treat, per‑protocol, or modified ITT? Each shifts the effective sample size. Pick one, state the fallback, and don't let a data‑driven decision at unblinding rewrite the rulebook.

Interim Looks: When to Peek and When to Stop

Adaptive designs are trendy. Implementing them sloppily is a regulatory black eye. If you plan one interim analysis for futility, write two things: the boundary (e.g., conditional power

“A stopping rule is not a suggestion. It's a contract with your own future bias.”

— paraphrased from a biostatistics director I worked with, after watching a promising drug fail replication

Set the looks in the protocol calendar. Tie each look to a specific data cut. And for heaven’s sake — assign one statistician to run the unblinded analysis and a separate team to interpret the results. Chinese walls are not paranoia; they're the difference between a clean Phase II result and a conversation with the FDA about inflation of Type I error.

One last concrete step: archive the simulation code or formula reference in the trial master file. Not as a PDF screenshot. As runnable code or a clearly labelled spreadsheet. Future you — or a skeptical reviewer — will thank you.

When the Number Is Wrong: Risks of Skipping Proper Planning

Underpowered Study: False Negatives That Haunt You Later

A trial that's too small doesn't just miss significance—it buries a potentially viable treatment. I have watched teams spend eighteen months recruiting, treating, and analyzing, only to produce a p-value of 0.09. The drug looked dead. Management killed the program. Six months later a competitor ran the same experiment with thirty more patients and got p = 0.03. That hurts. The first team had been right about the biology; they were just wrong about the math. False negatives masquerade as failures, and in early-stage research those failures often stop development cold. You lose the therapy, the investment, and years of work—all because no one checked whether the sample could actually detect the effect size you cared about.

Reality check: name the research owner or stop.

The catch is that small studies also inflate your effect estimate when they do turn up significant. This is the winner's curse: the few positive results that squeak through an underpowered trial are systematically larger than the truth. That sounds fine until a Phase II replication fails spectacularly. Then you face a harder question: was your original finding real but weak, or was it noise dressed up as a breakthrough? Nobody budgets time to untangle that mess.

Overpowered Study: Wasted Resources and Ethical Blind Spots

Bigger isn't automatically better. An overpowered trial—say, 400 patients when 120 would suffice—burns cash and investigator goodwill. The odd part is that many sponsors treat extra patients as cheap insurance. It's not. Every extra body means more monitoring visits, more data queries, more serum samples that freeze and thaw and degrade. The administrative drag alone can delay database lock by weeks. Worse, you expose more people to an experimental intervention than necessary. That's an ethical failure, plain as day. Institutional review boards catch this sometimes, but not always. Most teams skip this: they run a power calculation with one effect size, then double the sample because they're scared of missing a signal. Wrong order. You should determine the minimum viable sample to answer your question, then add a fixed buffer for dropout—not double the whole thing out of vague anxiety.

Regulatory scrutiny sharpens when a sample looks inflated. I have seen FDA reviewers ask, "Why 200 patients if your pre-specified analysis required 80?" If you can't defend the number with a concrete rationale—not just "we wanted to be safe"—the agency flags the trial as poorly planned. That triggers questions about data integrity. Did you stop early? Did you change endpoints midstream? Did you fish for subgroups because the overall result was borderline? Proper planning inoculates you against those queries. Skipping it invites them. And regulators have long memories: one sketchy sample justification stains your entire development program.

“An underpowered trial is a waste of everyone's time. An overpowered trial is a waste of someone's trust.”

— paraphrased from a biostatistician after a Data Safety Monitoring Board meeting, 2023

Regulatory Scrutiny and the Data Integrity Trap

The worst risk is subtle: when sample size is wrong, the entire trial's credibility fractures. Underpowered studies produce unstable variance estimates. Overpowered ones produce statistically significant but clinically trivial results. Either way, your data set looks suspicious. Regulators will ask for the original power calculation, the assumptions behind it, and the audit trail for every deviation. If those assumptions were plucked from thin air— "we used Cohen's d = 0.5 because that's standard" —your submission loses all persuasive force. One concrete anecdote: a team I worked with submitted a Phase Ib that showed a 40% response rate. The agency asked for the sample size justification. The statistician had used a two-sample t-test for a single-arm study. That mistake killed six months of review time and forced a costly addendum. Fix the number upfront, or fix it later under a microscope—your choice.

Mini-FAQ: Sample Size Questions You’re Afraid to Ask

Why not always use the largest sample possible?

Because “more data” isn’t free — it’s a tax on time, money, and ethics. I have seen teams balloon a Phase I trial to 60 subjects simply because “it looks better for the FDA.” What they didn’t plan for: the recruitment pipeline dried up at month seven, the placebo arm started showing weird adverse events, and the whole study had to pause. Bigger isn’t stronger; bigger is just slower. The catch is statistical power follows a diminishing curve — after a certain N, you gain decimal points of precision while burning real calendar months. Worse, an oversized early trial can mask a signal you need to see. Small effects wash out in noise, sure, but tiny samples show you the shape of variability. That shape tells you where to invest next. So no — don't default to “max N.” Default to “enough to see the effect move, not to prove it.”

How do I account for dropout rates without guessing?

You guess. But you guess with data, not fear. Most researchers inflate by 15% and call it done. That's lazy — and it usually fails in one direction: too conservative, wasting slots, or too optimistic, leaving you underpowered. What actually works? Pull attrition rates from your own clinic’s last three similar trials. Did patients drop because of side effects? Travel distance? Follow-up frequency? Each driver has a different curve. Side-effect dropouts cluster in week one. Travel dropouts bleed steady across the trial. So instead of a flat 15% buffer, I model two scenarios: worst-case (30% attrition, back-loaded) and realistic (12%, front-loaded). Then I pick the sample size that survives the realistic case and still hits power. That sounds fine until your funder says “no budget for the worst-case.” Then you negotiate: shorten the follow-up window or add a mid-trial enrollment top-up clause. This is not math — it’s logistics with a calculator.

One more thing: track your assumed dropout rate as a protocol amendment risk. If actual dropout hits 20% by month three, you need a pre-planned response. Not a panic. A rule. “If dropout exceeds X, we trigger a blinded sample-size review.” Put that in writing before the first patient consents.

Can I change sample size mid-study?

Yes — but only if you wrote the rules before you saw the data. This is where adaptive designs earn their keep. A well-planned interim analysis lets you increase N based on blinded variance estimates, or even re-estimate power using pooled effect sizes without unblinding. The pitfall? Doing it because the results “look promising.” That's p-hacking dressed in good intentions. A legitimate adaptive approach requires a pre-specified decision rule — for example: if the conditional power falls between 50% and 80% at the interim, we boost N by 30%. If it’s above 80%, stop early for efficacy. Below 50%, stop for futility. Those rules must be locked in the SAP before the first injection. Otherwise, regulators see a fishing expedition, and you lose credibility fast.

“The worst time to decide sample size is when you already love the data. Love blinds. Plan before you peek.”

— paraphrased from a clinical operations director I worked with, after she watched a promising trial collapse under post-hoc adjustments

The practical take: if you anticipate needing flexibility, build a simple adaptive plan into your protocol. Simulation tools (see section 2) make this dead simple — run 1,000 virtual trials, test different stopping rules, and print the decision boundaries. Then your mid-study change is not a change; it’s an execution of the plan you already had. That's how you stay credible while staying agile.

Our Take: Start Simple, Simulate If You Can, But Never Skip Feasibility

When to use formula vs. simulation

Formulas win when your endpoint is simple—a single mean, a proportion, a clean time-to-event with no competing risks. I have designed over a dozen early trials where a closed-form calculation took twenty minutes and worked fine. The catch: formulas assume perfect data. No dropouts. No delayed enrollment. No covariance structure you forgot to specify. Simulation shines when reality gets messy—unequal allocation, longitudinal biomarkers, a planned interim look that bends the operating characteristics. The odd part is—teams often simulate too late. They build a model after the protocol draft is locked. Swap the order: simulate first, then confirm with a formula. That catches hidden inflation before you commit to a number you can't defend.

‘A formula gives you the sample size you deserve; simulation gives you the sample size you actually need.’

— paraphrased from a biostatistician I worked with at a Phase Ib start-up

The one thing you must do before submitting

Feasibility. Not sample size. Feasibility. Most teams skip this: they calculate n=24, submit to ethics, and discover the site can recruit two patients per month. That hurts. Before you write a single line of code, ask: how many eligible patients have passed through this clinic in the last year? What is the screen-failure rate for similar protocols? If the answer is fuzzy, do a mini chart review—five hours of work can save you six months of slow enrollment. I have seen a perfectly powered trial collapse because the PI overestimated referral volume by a factor of three. The sample size was correct. The assumption was wrong. Feasibility is not a soft skill; it's the hard constraint that makes your careful calculation either meaningful or moot.

A final checklist

Run through this before you finalise any number:
- Did you account for at least one source of real-world noise (dropouts, missing visits, protocol deviations)?
- Does your simulation (if you built one) assume the treatment effect is fixed? Try a range—optimistic, realistic, pessimistic.
- Can you recruit the required n within the trial timeline, given historical rates?
- Have you shown your assumptions to a clinician who disagrees with them?
Wrong order. Not yet. That last bullet catches more errors than any statistical test. Find a sceptic. Buy them coffee. Listen to what breaks first—then fix the sample size, not the objection.

Share this article:

Comments (0)

No comments yet. Be the first to comment!