Guides / reliability and validity
Reliability and validity: a complete guide
Reliability asks whether a measure gives consistent results; validity asks whether it measures what you claim. They are commonly treated as a single box to tick, and the distinction between them decides whether your findings mean anything at all. This guide covers every type of each, the statistic to report for it, and what to do when your instrument falls short.
Reliability is consistency: whether a measure gives the same result on repeated application, across items, or between raters. Validity is accuracy: whether it measures the construct it claims to measure. A measure can be highly reliable and completely invalid, producing the same wrong answer every time. Reliability is necessary for validity but does not guarantee it.
Definition
The difference, and why it matters
Reliability is consistency. Validity is accuracy. A bathroom scale that reads three kilograms heavy every time is perfectly reliable and entirely invalid; one that fluctuates by five kilograms between weighings is neither.
The relationship between them is asymmetric and worth stating precisely. Reliability is a necessary but not sufficient condition for validity. A measure swamped by random error cannot be measuring anything consistently, so it cannot be valid. But a measure can be beautifully consistent and still capture something other than what you intended.
There is a further reason to keep the two separate in your write-up. Reliability sets a ceiling on validity that can be quantified: a measure's correlation with any criterion cannot exceed the square root of its reliability. A scale with an alpha of .64 can correlate at most about .80 with anything, however well designed. This is why a weak reliability figure is not merely untidy — it caps what the instrument can possibly demonstrate, and no amount of validity argument will lift that ceiling.
Reliability produces a single number your software will calculate in seconds. Validity requires an argument, usually assembled from several sources of evidence, and there is no button for it. That asymmetry is why a great many dissertations report Cronbach's alpha and stop — and why examiners ask about validity.
Reliability
Types of reliability
| Type | Question it answers | When to use |
|---|---|---|
| Internal consistency | Do the items measure the same thing? | Any multi-item scale |
| Test-retest | Is it stable over time? | Traits expected to be stable |
| Inter-rater | Do different raters agree? | Observation, coding, clinical judgement |
| Intra-rater | Does one rater agree with themselves? | Repeated coding by the same person |
| Parallel forms | Do two versions give the same result? | Where practice effects rule out retesting |
Test-retest needs a defensible interval
Too short and participants remember their answers, inflating the correlation; too long and genuine change gets counted as unreliability. Two to four weeks is conventional for stable traits. The right interval depends on how quickly the construct itself is expected to move, and you should justify the choice rather than default to it.
Split-half reliability is largely of historical interest now but explains where alpha came from. The idea was to divide a scale in two, correlate the halves, and adjust upwards for the fact that each half is shorter than the whole. The difficulty is that the result depends on how you split it — odd against even items gives a different figure from first half against second. Cronbach's alpha resolves this by being, in effect, the average of every possible split-half coefficient, which is why it superseded the older method.
Mood, fatigue and situational anxiety are supposed to fluctuate. A low test-retest correlation for a state measure is evidence that it is doing its job, not that it is unreliable. Use internal consistency instead, and say why.
Statistics
Which reliability statistic to use
| Statistic | Use for | Acceptable |
|---|---|---|
| Cronbach's α | Internal consistency, continuous items | ≥ .70; ≥ .80 preferred |
| McDonald's ω | Internal consistency, fewer assumptions | ≥ .70; increasingly preferred to α |
| KR-20 | Internal consistency, binary items | ≥ .70 |
| Cohen's κ | Two raters, categorical | ≥ .61 substantial; ≥ .81 almost perfect |
| Weighted κ | Two raters, ordered categories | As above, credits near misses |
| Fleiss' κ | Three or more raters, categorical | As above |
| ICC | Continuous ratings, or test-retest | ≥ .75 good; ≥ .90 excellent |
Kappa has a known weakness worth anticipating, since a reviewer may raise it. When one category dominates heavily, kappa can be low even though the coders agree on almost everything — the so-called kappa paradox. The correction for chance becomes so large that little room is left for the statistic to move. Where this occurs, report kappa alongside the raw agreement and the marginal distributions, and note the imbalance rather than presenting the low kappa as though the coders had performed badly.
Why kappa rather than percentage agreement
Percentage agreement ignores the agreement expected by chance. If two coders each assign 90% of extracts to the same dominant category, they will agree about 82% of the time by chance alone, so 85% agreement is close to worthless. Running the numbers makes this concrete: with expected agreement of .82, an observed 85% gives κ = (.85 − .82) / (1 − .82) = .17 — slight agreement, not the substantial figure the raw percentage suggests. Kappa corrects for this, which is why it is the reportable statistic.
The ICC has several forms, and software will ask you to choose between them. The distinctions concern whether raters were selected at random or are the only ones of interest, and whether you want agreement in absolute terms or merely consistency in ranking. Report which form you used, because a two-way mixed model assessing absolute agreement and a one-way random model assessing consistency can give noticeably different values on the same data.
It depends on the sample as well as the items. The same questionnaire returns a lower alpha in a homogeneous group, because restricted variation reduces inter-item correlation. Report alpha for your sample rather than citing the figure from the original validation paper, and note the number of items alongside it — alpha rises with length regardless of item quality.
Validity
Types of validity
| Type | Question | How to evidence it |
|---|---|---|
| Face | Does it look right? | Judgement; the weakest form |
| Content | Does it cover the whole construct? | Expert panel, content validity index |
| Construct | Does it measure the underlying concept? | Factor analysis, hypothesis testing |
| Convergent | Does it correlate with related measures? | Correlation with established instruments |
| Discriminant | Is it distinct from unrelated ones? | Low correlation where expected |
| Criterion — concurrent | Does it match a standard now? | Correlation with a gold standard |
| Criterion — predictive | Does it predict a future outcome? | Correlation with a later criterion |
| Ecological | Does it hold outside the study setting? | Replication in real conditions |
Convergent and discriminant validity are usually assessed together, because either alone proves little. A new measure of anxiety should correlate substantially with an established anxiety scale and only modestly with a measure of, say, extraversion. High correlations with everything indicate a general response tendency rather than a specific construct.
Construct validity deserves particular attention because it is the type examiners probe hardest and the one most often asserted without evidence. It is not established by a single analysis; the modern view treats it as the overarching concern that the others contribute to. Factor analysis showing the expected structure is one piece of evidence. So is the instrument behaving as theory predicts — scores rising after an intervention designed to raise them, or differing between groups known to differ on the construct. Each finding that matches prediction strengthens the case; each that does not weakens it, and should be reported rather than omitted.
Once the instrument is administered it is too late to discover that an important facet of the construct was never asked about. Have subject experts review the item pool against a definition of the construct, and record their judgements — a content validity index is straightforward to compute and easy to defend.
Practice
Establishing them in your own study
If you use an existing validated instrument
If you develop your own
This is a substantial undertaking and is usually a study in itself rather than a preliminary step. The minimum credible sequence is: define the construct; generate an item pool from the literature; have experts review it for content validity; pilot on a small sample and revise; administer to a larger sample; examine the factor structure; compute internal consistency; and test convergent and discriminant validity against established measures.
Translating an instrument counts as modification, and a translated scale requires revalidation in the new language. The accepted procedure is forward translation by two independent translators, reconciliation, back-translation by a third who has not seen the original, and comparison — followed by a pilot to check that items read naturally and mean what was intended. Reporting that a scale was “translated into Urdu” without describing this process invites a straightforward challenge in a viva.
Common guidance is at least 10 participants per item, with a floor of about 200. A 20-item scale therefore needs roughly 200 respondents before a factor structure can be examined credibly. If you cannot reach that, use an existing instrument — a bespoke scale validated on 60 people is weaker evidence than a borrowed one.
Send your draft items or your chosen instrument. A named researcher reviews the wording, structure and scoring, and advises on what reliability and validity evidence you will be able to claim.
Get a fixed quoteWorked example
Worked example: assessing a new scale
A 10-item scale measuring study confidence is administered to 240 undergraduates, alongside an established self-efficacy measure and a measure of extraversion. Two weeks later, 60 participants complete it again.
| Evidence | Result | Verdict |
|---|---|---|
| Cronbach's α | .88 | Good internal consistency |
| McDonald's ω | .89 | Confirms α is not inflated by item count |
| Test-retest ICC (2 weeks) | .81 | Good stability |
| Factor analysis | One factor, 54% variance | Consistent with a single construct |
| Correlation with self-efficacy | r = .62 | Convergent validity supported |
| Correlation with extraversion | r = .11 | Discriminant validity supported |
| Predicts end-of-year marks | r = .34 | Modest predictive validity |
Read as a set, this is a credible case. The scale is internally consistent and stable, behaves as one factor, relates strongly to a conceptually similar measure and weakly to an unrelated one, and predicts a real outcome. No single row would establish much; the pattern across them is the argument.
Note also what is missing from that table and would strengthen it further: evidence that the scale behaves as theory predicts under intervention, and evidence that it performs comparably in a second, different sample. Validation is cumulative rather than a single study, and an honest write-up says which pieces of the argument are still outstanding.
A correlation of .34 with marks means the scale accounts for about 12% of the variance in attainment. That is a genuine and reportable relationship, and it is not a basis for using the scale to make decisions about individual students. Predictive validity at the group level and usefulness for individual prediction are different claims.
Qualitative research
The qualitative equivalents
Reliability and validity are concepts from measurement, and applying them unchanged to qualitative work misrepresents what that work claims. The established alternative framework is trustworthiness.
| Quantitative | Qualitative equivalent | Established by |
|---|---|---|
| Internal validity | Credibility | Member checking, triangulation, prolonged engagement |
| External validity | Transferability | Thick description so readers can judge fit |
| Reliability | Dependability | Audit trail, documented decisions |
| Objectivity | Confirmability | Reflexivity, evidence traceable to data |
Note that inter-rater reliability still has a place in qualitative work, but a contested one. In content analysis, where the aim is systematic classification, kappa is appropriate and expected. In reflexive thematic analysis, where interpretation by the researcher is the instrument rather than a source of error, reporting a kappa misrepresents the method — and Braun and Clarke are explicit about this.
Triangulation deserves a note, since it is the trustworthiness criterion most often invoked and least often specified. There are several distinct kinds: data triangulation across sources or time points, investigator triangulation across analysts, theoretical triangulation across frameworks, and methodological triangulation across methods. Saying “triangulation was used” conveys almost nothing. Name which kind, describe what it consisted of, and say what it revealed — including where sources disagreed, which is usually the more interesting result.
Reporting Cronbach's alpha for interview data, or claiming inter-rater reliability for a reflexive thematic analysis, signals that the framework was applied without regard to what the method actually claims. Use trustworthiness criteria, name them, and evidence each one.
Pitfalls
Six mistakes that cost marks
1. Treating reliability as evidence of validity
A consistent measure can consistently measure the wrong thing. They are separate claims requiring separate evidence.
2. Citing the original study's alpha instead of your own
Alpha depends on your sample. Compute and report it for your data.
3. Reporting percentage agreement rather than kappa
Percentage agreement ignores chance agreement and overstates reliability substantially.
4. Modifying a validated instrument silently
Changing wording, response options or item order voids the original validation. Say what you changed and why.
5. Reporting alpha above .95 as excellent
Very high alpha usually indicates redundant items asking the same question repeatedly, not a superior scale.
6. Applying quantitative criteria to qualitative work
Use credibility, transferability, dependability and confirmability, and evidence each.
Reporting
Reporting reliability and validity
| Element | How to write it |
|---|---|
| Internal consistency | The 10 items showed good internal consistency (α = .88, ω = .89) |
| Test-retest | ICC = .81, 95% CI [.70, .88], n = 60, two-week interval |
| Inter-rater | κ = .74, indicating substantial agreement across 20% double-coded |
| Convergent validity | r = .62 with the GSE, p < .001 |
| Existing instrument | α in the present sample was .84 (originally reported as .89) |
Send your instrument and dataset. A named researcher checks the reliability analysis, examines the factor structure, and advises what validity evidence your design supports.
See SPSS data analysisAnswers
Frequently asked questions
What is the difference between reliability and validity?
Reliability is consistency — whether a measure gives the same result on repeated application, across items or between raters. Validity is accuracy — whether it measures the construct it claims to. A measure can be perfectly reliable and completely invalid, giving the same wrong answer every time.
Can a measure be reliable but not valid?
Yes, and it is the most dangerous combination. A scale that reads three kilograms heavy every time is perfectly consistent and entirely wrong. Reliability is necessary for validity but nowhere near sufficient, which is why reporting alpha alone does not establish that an instrument works.
What is an acceptable Cronbach's alpha?
Above .70 is conventionally acceptable and above .80 good. Below .60 the items should not be combined into a single score. Be cautious about values above .95, which usually indicate redundant items. Alpha also rises with the number of items, so report the item count alongside it.
What is the difference between Cronbach's alpha and McDonald's omega?
Alpha assumes every item contributes equally to the construct, which is rarely true. Omega relaxes that assumption and is generally the more accurate estimate, so it is increasingly preferred. Report omega where your software provides it, and both if you can.
Should I report percentage agreement or kappa for inter-rater reliability?
Kappa. Percentage agreement ignores the agreement expected by chance, which can be very high when one category dominates — two coders using the same dominant category 90% of the time will agree about 82% of the time by chance alone. Kappa corrects for this.
How do I establish validity for a new questionnaire?
Build an argument from several sources: expert review for content validity, factor analysis for construct validity, correlation with an established measure of the same construct for convergent validity, low correlation with an unrelated measure for discriminant validity, and where possible correlation with a real outcome for predictive validity.
What are the qualitative equivalents of reliability and validity?
Credibility, transferability, dependability and confirmability, collectively called trustworthiness. They are established through triangulation, member checking, thick description, an audit trail of decisions, and reflexivity, rather than through statistics.
Do I need inter-rater reliability for thematic analysis?
It depends on the variant. For content analysis and codebook approaches, where systematic classification is the aim, kappa is appropriate and expected. For reflexive thematic analysis, the researcher's interpretation is the analytic instrument rather than a source of error, and Braun and Clarke argue explicitly against reporting a reliability coefficient.
Send the data. Get a fixed quote.
Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.