Guides  /  experimental design

Experimental design: a practical guide

An experiment earns the right to a causal claim through its design, not its analysis — and no statistical technique can recover what a flawed design gave away. This guide covers the choices that determine what your study can conclude, the threats each one guards against, and the decisions that must be made before a single participant is recruited.

Hafiz Ahmad Tariq Written and reviewed by Hafiz Ahmad Tariq, Senior Biostatistician
Updated 17 August 202618 min read
What is experimental design?

Experimental design is the set of choices about how a study is structured so that observed differences in the outcome can be attributed to the manipulated variable rather than to something else. The core elements are a manipulated independent variable, a measured dependent variable, random allocation of participants to conditions, and control of extraneous variables.

Definition

What makes a study an experiment

An experiment manipulates something and observes the effect, with everything else held as constant as the design allows. Three features are required, and a study missing any of them is quasi-experimental or observational.

Manipulation — the researcher sets the level of the independent variable rather than observing it
Random allocation — participants are assigned to conditions by chance
Control — extraneous variables are held constant or accounted for
DesignManipulationRandom allocationCausal claim
True experimentYesYesSupported
Quasi-experimentYesNo — pre-existing groupsWeak; confounding likely
CorrelationalNoNoNot supported
Natural experimentNo — but allocation is as-if randomApproximatelySometimes defensible

The distinction between an independent and a dependent variable is worth restating precisely, because the terms get used loosely. The independent variable is the one you set: it is independent of the participant, because you decided its value. The dependent variable is the one you measure, and its value depends on what you did. If you did not set the value of a variable — if you merely recorded whether participants happened to smoke, or which school they attended — it is not an independent variable in the experimental sense, however your software labels it, and the study is observational.

Causation comes from the design

This is worth being blunt about, because it is the single most common overreach in student work. Regression with twenty covariates on observational data does not establish causation. A simple randomised comparison with two groups does. The difference is in how participants came to be in their conditions, not in the sophistication of the analysis.

Structure

Between, within and mixed designs

BETWEEN-SUBJECTS — different people in each condition P1–P20 → AP21–P40 → B no carryover; needs more participants WITHIN-SUBJECTS — everyone does every condition P1–P20 → AP1–P20 → B more powerful; risks order and practice effects MIXED — one factor between, one within group 1: time 1 → time 2 → time 3 group 2: time 1 → time 2 → time 3 the common trial shape: treatment between, time within The choice fixes your analysis: independent tests, paired tests, or a mixed model.
The three basic shapes. The choice determines your analysis before you collect anything.
Between-subjectsWithin-subjects
ParticipantsDifferent people per conditionSame people in all conditions
Individual differencesAdd noise between conditionsControlled — each person is their own control
PowerLowerSubstantially higher
Participants neededMoreFewer
Order effectsNoneA real risk
AnalysisIndependent t-test, ANOVAPaired t-test, repeated measures ANOVA

Within-subjects designs are considerably more efficient, because the largest source of noise in most behavioural research is variation between people, and using each participant as their own control removes it entirely. A within-subjects design can often detect the same effect with a third of the participants.

Factorial designs extend the same logic to more than one manipulated variable. A 2 × 2 design crosses two factors with two levels each, producing four conditions, and it delivers three results for roughly the cost of one: two main effects and an interaction. The interaction is usually the reason to run it, since it answers whether one factor's effect depends on the other — a question a pair of separate experiments cannot address at all. The cost is that participant requirements grow with the number of cells, and interactions need appreciably more power to detect than main effects, typically around four times as much for the same effect size.

When a within-subjects design is impossible

If the manipulation changes the participant permanently — teaching them something, revealing a deception, altering how they approach the task — they cannot meaningfully do the other condition afterwards. Any manipulation that involves learning is usually between-subjects for this reason.

Allocation

Randomisation, and what it buys you

Random allocation is what licenses the causal claim. Because participants are assigned by chance, the groups differ only by chance at the outset — on everything, including variables you never measured and never thought of.

That last point is the whole value of it. Statistical adjustment can only control for variables you measured; randomisation balances the ones you did not.

MethodHowUse when
SimpleRandom number per participantLarge samples
BlockRandomise within blocks of fixed sizeEnsures groups stay balanced in size
StratifiedRandomise within levels of a key variableA known strong prognostic factor
ClusterRandomise groups, not individualsClassrooms, wards, practices
MinimisationAllocate to balance covariates as you goSmall trials with several key covariates

One point that surprises people: randomisation does not guarantee balanced groups in any single study, and checking baseline balance with significance tests is actively discouraged in trial reporting. If allocation was genuinely random, any imbalance is by definition due to chance, so a significance test asks a question whose answer is already known. Report the baseline characteristics in a table so readers can judge whether any imbalance is clinically or substantively important, and adjust for a variable in the analysis only if that was pre-specified.

Allocation concealment is separate from randomisation

Randomisation decides the sequence; concealment stops whoever recruits participants knowing what comes next. Without concealment a recruiter can consciously or unconsciously steer a participant towards a condition, which reintroduces exactly the selection bias randomisation exists to prevent. Use sealed opaque envelopes or a remote system.

Note that simple randomisation with small samples will not necessarily produce balanced groups — with 20 participants, an 8 to 12 split is entirely possible. Block randomisation prevents this, and stratifying on a variable known to predict the outcome strongly is worth doing when the sample is small.

Comparison

Control conditions

What you compare against determines what your result means. A treatment that beats nothing has demonstrated something much weaker than one that beats a credible alternative.

ControlWhat it isolatesLimitation
No treatmentEverything about the interventionCannot separate the active ingredient from attention
WaitlistAs above, with an ethical routeParticipants know they are waiting
Placebo / shamThe specific active componentNot always constructible
Active controlWhether it beats current practiceThe clinically relevant question
Attention controlThe effect beyond contact timeDemanding to design well

Ethics constrain this choice more than any other part of the design. Withholding an effective treatment to create a no-treatment arm is rarely acceptable where one exists, which is why active controls dominate in clinical research. A waitlist design is the usual compromise in psychological and educational work: everyone eventually receives the intervention, and the comparison is made before the waiting group starts. Note that this limits follow-up, since after crossover there is no longer an untreated group to compare against.

The comparison decides the claim

“Better than nothing” and “better than what is currently done” are different findings, and only the second usually matters for practice. Choose the comparator that answers the question a reader will actually have, and state explicitly what your control did and did not receive.

Within-subjects

Order effects and counterbalancing

In a within-subjects design, whatever participants do first affects what they do next. Three distinct effects are worth separating.

EffectWhat happens
PracticePerformance improves through familiarity with the task
FatiguePerformance declines through tiredness or boredom
CarryoverThe first condition changes how the second is experienced

Counterbalancing

Practice and fatigue are handled by varying the order across participants, so that the effects cancel out across the sample rather than accumulating in one condition.

Complete counterbalancing — every possible order used equally. With 3 conditions that is 6 orders; with 4 it is 24, which quickly becomes impractical
Latin square — each condition appears in each position once. Practical for 4 or more conditions
Random order per participant — acceptable with a reasonably large sample
Balanced Latin square — also balances which condition precedes which, controlling simple carryover

A practical note on how many orders you need. Complete counterbalancing requires the number of participants to be a multiple of the number of orders, so with three conditions and six orders you want a multiple of six — 24, 30, 36. Planning a sample of 25 leaves one order over-represented, which reintroduces exactly the imbalance the counterbalancing was meant to remove. Decide the design first and let it constrain the target sample size, rather than the reverse.

Counterbalancing does not fix carryover

If the first condition genuinely changes the participant — they learn a strategy, or a drug is still active — reversing the order does not cancel it out; it just distributes the contamination. The remedies are a washout period long enough for the effect to dissipate, or a between-subjects design.

Get the design settled before you collect data

Send your research question and proposed design. A named statistician confirms the structure, the allocation method and the sample size — while every one of them is still changeable.

See sample size and power

Bias control

Blinding

LevelWho is unawarePrevents
Single-blindParticipantsExpectancy and demand characteristics
Double-blindParticipants and researchersThe above, plus experimenter effects
Triple-blindPlus the analystAnalytic decisions favouring a hypothesis

Blinding matters because expectation affects outcomes on both sides. Participants who know they received the intervention may report improvement they would not otherwise notice; researchers who know may unconsciously alter how they administer a task or rate an outcome.

Record who was blinded, at what stage, and how the blinding was maintained, since a claim of double-blinding with no procedural detail is routinely queried.

When blinding is impossible, say so

Many interventions cannot be blinded — a participant knows whether they attended a workshop. Where that is the case, blind whoever assesses the outcome instead, which is usually feasible and addresses much of the risk. State plainly what was and was not blinded rather than omitting the question.

Validity

Threats to validity

validity shown as a trade-off between controlled laboratory conditions and realistic field conditions"> tight lab experiment high internal validity low external validity field study low internal validity high external validity pragmatic trial Internal validity: can the effect be attributed to the manipulation? External validity: does it hold outside these conditions? Controlling more raises internal validity and usually lowers external validity. Decide which your question needs.
Internal and external validity trade off. Tightening control buys attribution at the cost of generalisability.
ThreatWhat it isGuard
SelectionGroups differ at the outsetRandomise; check baseline balance
HistorySomething happens during the studyConcurrent control group
MaturationParticipants change naturally over timeControl group; shorter follow-up
TestingThe pre-test affects the post-testSolomon four-group design, or no pre-test
InstrumentationThe measure changes between occasionsStandardise; train and check raters
Regression to the meanExtreme scorers drift towards averageDo not select on an extreme score
AttritionDropout differs by conditionIntention-to-treat analysis; report by arm
Demand characteristicsParticipants guess and complyBlinding; cover story where ethical

Attrition deserves particular attention because it can undo randomisation entirely. If participants who find the intervention difficult drop out of that arm while the control arm loses nobody, the groups you analyse are no longer the groups you randomised, and the comparison has quietly become observational. Analysing by intention to treat — keeping everyone in the arm they were assigned to, whatever they actually received — preserves the randomisation, at the cost of diluting the estimated effect. Report both that and the per-protocol analysis where the two diverge.

Regression to the mean is the quiet one

Select the lowest-scoring students for an intervention and they will improve at retest whether or not it works, simply because their first score was partly bad luck. Without a control group selected the same way, this effect is indistinguishable from a treatment effect — and it has produced a great many spurious findings.

Power

Sample size and power

Sample size is a design decision, not something to determine afterwards. A power analysis needs three inputs: the smallest effect worth detecting, the significance level, and the power you want.

DesignEffect (d)Powern per group
Two independent groups0.5 (medium).8064
Two independent groups0.8 (large).8026
Two independent groups0.3 (small).80176
Paired / within-subjects0.5.8034 total

The first and last rows show the efficiency of within-subjects designs: the same effect at the same power needs 128 participants between-subjects and 34 within-subjects. That difference is frequently decisive for a student project.

The effect size to plug in should be the smallest one worth detecting, not the one you hope to find or the one a previous study reported. Published effect sizes are systematically inflated by publication bias, so powering a study on them routinely produces one that is underpowered for reality. Asking “how small a difference would still change what anyone does?” gives a defensible figure and is the question examiners expect you to have asked.

Post-hoc power is not informative

Calculating power after a non-significant result, using the effect size you observed, is a direct transformation of the p-value you already have. It cannot tell you whether the study was adequately powered to detect the effect you cared about. Report the confidence interval instead — it answers the question honestly.

Work out the sample size your design needs

Our power analysis calculator gives the number required for the effect you want to detect, at your chosen power and significance level.

Use the calculator

Pitfalls

Six mistakes that undermine a design

1. No control group

A pre-post comparison cannot separate the intervention from maturation, history or regression to the mean.

2. Allocating by convenience

Assigning by class, session or willingness is not randomisation, and the groups will differ systematically.

3. No counterbalancing in a within-subjects design

Order effects then load entirely onto one condition and are indistinguishable from the effect you are testing.

4. Deciding the sample size afterwards

Stopping when a result reaches significance inflates the false positive rate substantially.

5. Selecting participants on extreme scores

Regression to the mean will produce apparent improvement with no intervention at all.

6. Ignoring differential attrition

If dropout differs between arms, randomisation is broken. Report attrition per arm and analyse by intention to treat.

Reporting

Reporting your design

Name the design precisely — “2 × 3 mixed design, with condition between-subjects and time within”
State how randomisation was performed and how allocation was concealed
Describe the control condition in full
Report counterbalancing and the method used
State what was blinded, and what could not be
Give the power analysis, including the effect size assumed and its source
Report attrition by arm with reasons
Follow CONSORT for a trial, including the flow diagram
A sentence that earns marks

“A 2 × 3 mixed design was used, with condition (intervention, active control) between-subjects and time (baseline, post, 3-month follow-up) within-subjects. Participants were randomised in blocks of four using a computer-generated sequence held by a researcher independent of recruitment. An a priori power analysis indicated 64 per group would detect d = 0.5 at 80% power with α = .05.”

Answers

Frequently asked questions

What is the difference between a true experiment and a quasi-experiment?

A true experiment randomly allocates participants to conditions; a quasi-experiment uses pre-existing groups. Random allocation is what balances unmeasured variables between groups and licenses a causal claim. Without it, differences between groups may reflect whatever made them different in the first place.

Should I use a between-subjects or within-subjects design?

Within-subjects wherever the manipulation permits, because each participant acts as their own control, removing individual differences and substantially increasing power. Use between-subjects when the manipulation changes participants permanently — anything involving learning, revelation of a deception, or lasting effects.

What is counterbalancing and when do I need it?

Varying the order of conditions across participants so that practice and fatigue effects cancel out rather than loading onto one condition. It is required for any within-subjects design. Use complete counterbalancing for two or three conditions, and a Latin square for four or more.

Why is random allocation so important?

Because it balances the groups on everything at the outset, including variables you never measured. Statistical adjustment can only control for what you recorded; randomisation handles unknown confounders as well, which is why it is the basis of any causal claim from an experiment.

What is regression to the mean?

The tendency for extreme scores to move towards the average on retesting, because part of an extreme score is chance. If you select the lowest scorers for an intervention, they will improve regardless of whether it works. A control group selected the same way is the only reliable guard.

How many participants do I need for an experiment?

It depends on the effect size you want to detect. For two independent groups at 80% power and alpha .05, detecting a medium effect (d = 0.5) requires about 64 per group; a small effect (d = 0.3) requires about 176. A within-subjects design detecting d = 0.5 needs about 34 in total.

What is the difference between internal and external validity?

Internal validity is whether the observed effect can be attributed to your manipulation. External validity is whether it generalises beyond your study's conditions. They trade off: tighter control raises internal validity and usually lowers external validity, so the balance should follow from your question.

Do I need a control group?

For a causal claim, yes. A pre-post comparison without a control cannot separate your intervention from maturation, external events, testing effects or regression to the mean. Which control you choose — no treatment, placebo, or active comparison — determines what your finding actually claims.

Send the data. Get a fixed quote.

Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.