Guides / experimental design
Experimental design: a practical guide
An experiment earns the right to a causal claim through its design, not its analysis — and no statistical technique can recover what a flawed design gave away. This guide covers the choices that determine what your study can conclude, the threats each one guards against, and the decisions that must be made before a single participant is recruited.
Experimental design is the set of choices about how a study is structured so that observed differences in the outcome can be attributed to the manipulated variable rather than to something else. The core elements are a manipulated independent variable, a measured dependent variable, random allocation of participants to conditions, and control of extraneous variables.
Definition
What makes a study an experiment
An experiment manipulates something and observes the effect, with everything else held as constant as the design allows. Three features are required, and a study missing any of them is quasi-experimental or observational.
| Design | Manipulation | Random allocation | Causal claim |
|---|---|---|---|
| True experiment | Yes | Yes | Supported |
| Quasi-experiment | Yes | No — pre-existing groups | Weak; confounding likely |
| Correlational | No | No | Not supported |
| Natural experiment | No — but allocation is as-if random | Approximately | Sometimes defensible |
The distinction between an independent and a dependent variable is worth restating precisely, because the terms get used loosely. The independent variable is the one you set: it is independent of the participant, because you decided its value. The dependent variable is the one you measure, and its value depends on what you did. If you did not set the value of a variable — if you merely recorded whether participants happened to smoke, or which school they attended — it is not an independent variable in the experimental sense, however your software labels it, and the study is observational.
This is worth being blunt about, because it is the single most common overreach in student work. Regression with twenty covariates on observational data does not establish causation. A simple randomised comparison with two groups does. The difference is in how participants came to be in their conditions, not in the sophistication of the analysis.
Structure
Between, within and mixed designs
| Between-subjects | Within-subjects | |
|---|---|---|
| Participants | Different people per condition | Same people in all conditions |
| Individual differences | Add noise between conditions | Controlled — each person is their own control |
| Power | Lower | Substantially higher |
| Participants needed | More | Fewer |
| Order effects | None | A real risk |
| Analysis | Independent t-test, ANOVA | Paired t-test, repeated measures ANOVA |
Within-subjects designs are considerably more efficient, because the largest source of noise in most behavioural research is variation between people, and using each participant as their own control removes it entirely. A within-subjects design can often detect the same effect with a third of the participants.
Factorial designs extend the same logic to more than one manipulated variable. A 2 × 2 design crosses two factors with two levels each, producing four conditions, and it delivers three results for roughly the cost of one: two main effects and an interaction. The interaction is usually the reason to run it, since it answers whether one factor's effect depends on the other — a question a pair of separate experiments cannot address at all. The cost is that participant requirements grow with the number of cells, and interactions need appreciably more power to detect than main effects, typically around four times as much for the same effect size.
If the manipulation changes the participant permanently — teaching them something, revealing a deception, altering how they approach the task — they cannot meaningfully do the other condition afterwards. Any manipulation that involves learning is usually between-subjects for this reason.
Allocation
Randomisation, and what it buys you
Random allocation is what licenses the causal claim. Because participants are assigned by chance, the groups differ only by chance at the outset — on everything, including variables you never measured and never thought of.
That last point is the whole value of it. Statistical adjustment can only control for variables you measured; randomisation balances the ones you did not.
| Method | How | Use when |
|---|---|---|
| Simple | Random number per participant | Large samples |
| Block | Randomise within blocks of fixed size | Ensures groups stay balanced in size |
| Stratified | Randomise within levels of a key variable | A known strong prognostic factor |
| Cluster | Randomise groups, not individuals | Classrooms, wards, practices |
| Minimisation | Allocate to balance covariates as you go | Small trials with several key covariates |
One point that surprises people: randomisation does not guarantee balanced groups in any single study, and checking baseline balance with significance tests is actively discouraged in trial reporting. If allocation was genuinely random, any imbalance is by definition due to chance, so a significance test asks a question whose answer is already known. Report the baseline characteristics in a table so readers can judge whether any imbalance is clinically or substantively important, and adjust for a variable in the analysis only if that was pre-specified.
Randomisation decides the sequence; concealment stops whoever recruits participants knowing what comes next. Without concealment a recruiter can consciously or unconsciously steer a participant towards a condition, which reintroduces exactly the selection bias randomisation exists to prevent. Use sealed opaque envelopes or a remote system.
Note that simple randomisation with small samples will not necessarily produce balanced groups — with 20 participants, an 8 to 12 split is entirely possible. Block randomisation prevents this, and stratifying on a variable known to predict the outcome strongly is worth doing when the sample is small.
Comparison
Control conditions
What you compare against determines what your result means. A treatment that beats nothing has demonstrated something much weaker than one that beats a credible alternative.
| Control | What it isolates | Limitation |
|---|---|---|
| No treatment | Everything about the intervention | Cannot separate the active ingredient from attention |
| Waitlist | As above, with an ethical route | Participants know they are waiting |
| Placebo / sham | The specific active component | Not always constructible |
| Active control | Whether it beats current practice | The clinically relevant question |
| Attention control | The effect beyond contact time | Demanding to design well |
Ethics constrain this choice more than any other part of the design. Withholding an effective treatment to create a no-treatment arm is rarely acceptable where one exists, which is why active controls dominate in clinical research. A waitlist design is the usual compromise in psychological and educational work: everyone eventually receives the intervention, and the comparison is made before the waiting group starts. Note that this limits follow-up, since after crossover there is no longer an untreated group to compare against.
“Better than nothing” and “better than what is currently done” are different findings, and only the second usually matters for practice. Choose the comparator that answers the question a reader will actually have, and state explicitly what your control did and did not receive.
Within-subjects
Order effects and counterbalancing
In a within-subjects design, whatever participants do first affects what they do next. Three distinct effects are worth separating.
| Effect | What happens |
|---|---|
| Practice | Performance improves through familiarity with the task |
| Fatigue | Performance declines through tiredness or boredom |
| Carryover | The first condition changes how the second is experienced |
Counterbalancing
Practice and fatigue are handled by varying the order across participants, so that the effects cancel out across the sample rather than accumulating in one condition.
A practical note on how many orders you need. Complete counterbalancing requires the number of participants to be a multiple of the number of orders, so with three conditions and six orders you want a multiple of six — 24, 30, 36. Planning a sample of 25 leaves one order over-represented, which reintroduces exactly the imbalance the counterbalancing was meant to remove. Decide the design first and let it constrain the target sample size, rather than the reverse.
If the first condition genuinely changes the participant — they learn a strategy, or a drug is still active — reversing the order does not cancel it out; it just distributes the contamination. The remedies are a washout period long enough for the effect to dissipate, or a between-subjects design.
Send your research question and proposed design. A named statistician confirms the structure, the allocation method and the sample size — while every one of them is still changeable.
See sample size and powerBias control
Blinding
| Level | Who is unaware | Prevents |
|---|---|---|
| Single-blind | Participants | Expectancy and demand characteristics |
| Double-blind | Participants and researchers | The above, plus experimenter effects |
| Triple-blind | Plus the analyst | Analytic decisions favouring a hypothesis |
Blinding matters because expectation affects outcomes on both sides. Participants who know they received the intervention may report improvement they would not otherwise notice; researchers who know may unconsciously alter how they administer a task or rate an outcome.
Record who was blinded, at what stage, and how the blinding was maintained, since a claim of double-blinding with no procedural detail is routinely queried.
Many interventions cannot be blinded — a participant knows whether they attended a workshop. Where that is the case, blind whoever assesses the outcome instead, which is usually feasible and addresses much of the risk. State plainly what was and was not blinded rather than omitting the question.
Validity
Threats to validity
| Threat | What it is | Guard |
|---|---|---|
| Selection | Groups differ at the outset | Randomise; check baseline balance |
| History | Something happens during the study | Concurrent control group |
| Maturation | Participants change naturally over time | Control group; shorter follow-up |
| Testing | The pre-test affects the post-test | Solomon four-group design, or no pre-test |
| Instrumentation | The measure changes between occasions | Standardise; train and check raters |
| Regression to the mean | Extreme scorers drift towards average | Do not select on an extreme score |
| Attrition | Dropout differs by condition | Intention-to-treat analysis; report by arm |
| Demand characteristics | Participants guess and comply | Blinding; cover story where ethical |
Attrition deserves particular attention because it can undo randomisation entirely. If participants who find the intervention difficult drop out of that arm while the control arm loses nobody, the groups you analyse are no longer the groups you randomised, and the comparison has quietly become observational. Analysing by intention to treat — keeping everyone in the arm they were assigned to, whatever they actually received — preserves the randomisation, at the cost of diluting the estimated effect. Report both that and the per-protocol analysis where the two diverge.
Select the lowest-scoring students for an intervention and they will improve at retest whether or not it works, simply because their first score was partly bad luck. Without a control group selected the same way, this effect is indistinguishable from a treatment effect — and it has produced a great many spurious findings.
Power
Sample size and power
Sample size is a design decision, not something to determine afterwards. A power analysis needs three inputs: the smallest effect worth detecting, the significance level, and the power you want.
| Design | Effect (d) | Power | n per group |
|---|---|---|---|
| Two independent groups | 0.5 (medium) | .80 | 64 |
| Two independent groups | 0.8 (large) | .80 | 26 |
| Two independent groups | 0.3 (small) | .80 | 176 |
| Paired / within-subjects | 0.5 | .80 | 34 total |
The first and last rows show the efficiency of within-subjects designs: the same effect at the same power needs 128 participants between-subjects and 34 within-subjects. That difference is frequently decisive for a student project.
The effect size to plug in should be the smallest one worth detecting, not the one you hope to find or the one a previous study reported. Published effect sizes are systematically inflated by publication bias, so powering a study on them routinely produces one that is underpowered for reality. Asking “how small a difference would still change what anyone does?” gives a defensible figure and is the question examiners expect you to have asked.
Calculating power after a non-significant result, using the effect size you observed, is a direct transformation of the p-value you already have. It cannot tell you whether the study was adequately powered to detect the effect you cared about. Report the confidence interval instead — it answers the question honestly.
Our power analysis calculator gives the number required for the effect you want to detect, at your chosen power and significance level.
Use the calculatorPitfalls
Six mistakes that undermine a design
1. No control group
A pre-post comparison cannot separate the intervention from maturation, history or regression to the mean.
2. Allocating by convenience
Assigning by class, session or willingness is not randomisation, and the groups will differ systematically.
3. No counterbalancing in a within-subjects design
Order effects then load entirely onto one condition and are indistinguishable from the effect you are testing.
4. Deciding the sample size afterwards
Stopping when a result reaches significance inflates the false positive rate substantially.
5. Selecting participants on extreme scores
Regression to the mean will produce apparent improvement with no intervention at all.
6. Ignoring differential attrition
If dropout differs between arms, randomisation is broken. Report attrition per arm and analyse by intention to treat.
Reporting
Reporting your design
“A 2 × 3 mixed design was used, with condition (intervention, active control) between-subjects and time (baseline, post, 3-month follow-up) within-subjects. Participants were randomised in blocks of four using a computer-generated sequence held by a researcher independent of recruitment. An a priori power analysis indicated 64 per group would detect d = 0.5 at 80% power with α = .05.”
Answers
Frequently asked questions
What is the difference between a true experiment and a quasi-experiment?
A true experiment randomly allocates participants to conditions; a quasi-experiment uses pre-existing groups. Random allocation is what balances unmeasured variables between groups and licenses a causal claim. Without it, differences between groups may reflect whatever made them different in the first place.
Should I use a between-subjects or within-subjects design?
Within-subjects wherever the manipulation permits, because each participant acts as their own control, removing individual differences and substantially increasing power. Use between-subjects when the manipulation changes participants permanently — anything involving learning, revelation of a deception, or lasting effects.
What is counterbalancing and when do I need it?
Varying the order of conditions across participants so that practice and fatigue effects cancel out rather than loading onto one condition. It is required for any within-subjects design. Use complete counterbalancing for two or three conditions, and a Latin square for four or more.
Why is random allocation so important?
Because it balances the groups on everything at the outset, including variables you never measured. Statistical adjustment can only control for what you recorded; randomisation handles unknown confounders as well, which is why it is the basis of any causal claim from an experiment.
What is regression to the mean?
The tendency for extreme scores to move towards the average on retesting, because part of an extreme score is chance. If you select the lowest scorers for an intervention, they will improve regardless of whether it works. A control group selected the same way is the only reliable guard.
How many participants do I need for an experiment?
It depends on the effect size you want to detect. For two independent groups at 80% power and alpha .05, detecting a medium effect (d = 0.5) requires about 64 per group; a small effect (d = 0.3) requires about 176. A within-subjects design detecting d = 0.5 needs about 34 in total.
What is the difference between internal and external validity?
Internal validity is whether the observed effect can be attributed to your manipulation. External validity is whether it generalises beyond your study's conditions. They trade off: tighter control raises internal validity and usually lowers external validity, so the balance should follow from your question.
Do I need a control group?
For a causal claim, yes. A pre-post comparison without a control cannot separate your intervention from maturation, external events, testing effects or regression to the mean. Which control you choose — no treatment, placebo, or active comparison — determines what your finding actually claims.
Send the data. Get a fixed quote.
Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.