Guides  /  thematic analysis

Thematic analysis: a complete guide, with a worked example

Thematic analysis is the most widely used qualitative method in the social and health sciences, and the most widely misreported. This guide covers what it is, the six phases as Braun and Clarke actually defined them, a coding example worked through from extract to theme, and the mistakes that cost marks in a viva.

Elaine Halliburton Written and reviewed by Elaine Halliburton, Professor of Applied Statistics
Updated 14 August 202618 min read
What is thematic analysis?

Thematic analysis is a qualitative method for identifying, analysing and reporting patterns of meaning across a dataset. It involves coding extracts of data, grouping those codes, and developing them into themes that answer a research question. Unlike content analysis it does not count occurrences; it interprets what the patterns mean.

Definition

What thematic analysis is

Thematic analysis is a method for identifying, analysing and reporting patterns of meaning — themes — across a qualitative dataset. You read the data closely, label segments of it with codes, group those codes, and develop the groups into themes that answer your research question.

It was formalised for the social sciences by Virginia Braun and Victoria Clarke in 2006, in a paper that has since become one of the most cited in qualitative research. Their central argument was that thematic analysis was already being used everywhere but rarely described properly, so it looked like an absence of method rather than a method in its own right.

Three properties distinguish it from the methods it is most often confused with.

MethodWhat it producesHow it differs
Thematic analysisInterpretative themes across a datasetNot tied to a theoretical framework; flexible across epistemologies
Content analysisFrequencies and categoriesCounts occurrences; usually more descriptive and often quantitative
Grounded theoryA theory grounded in the dataAims to generate theory; requires theoretical sampling and constant comparison
IPADetailed accounts of lived experiencePhenomenological, idiographic, small samples, double hermeneutic
Framework analysisA case-by-theme matrixPartly a priori; built for applied policy research and cross-case comparison
The most consequential sentence in your methods chapter

“Themes emerged from the data” is the phrase examiners most reliably challenge. Themes do not emerge. An analyst constructs them by making decisions, and describing those decisions is what makes the analysis assessable rather than merely asserted.

Choosing the method

When to use it — and when not to

Thematic analysis is flexible, which is its strength and its trap. Because it can be applied to almost any qualitative dataset, it is often chosen by default rather than because it fits the question.

Use thematic analysis when

  • You want to identify patterns of meaning across participants
  • The research question is about what people think, feel or experience
  • You have a reasonably sized dataset — interviews, focus groups, open-text responses
  • You need a method that is accessible to a mixed-methods audience
  • Your question is exploratory or descriptive rather than theory-generating

Use something else when

  • You need a theory as the output → grounded theory
  • You need the detailed texture of one person's experience → IPA
  • You need cross-case comparison in a matrix → framework analysis
  • You genuinely need frequencies → content analysis
  • Your interest is in how language is used rather than what is said → discourse analysis

The honest test: if you cannot say what thematic analysis gives you that another method would not, you have probably chosen it because it is familiar. That is a defensible reason to change method at the design stage and an uncomfortable one to defend at a viva.

Braun and Clarke

The six phases, in practice

The six phases are widely quoted and widely misunderstood. They are not a linear pipeline: phases four and five routinely send you back to phase two, and Braun and Clarke are explicit that the process is recursive.

PHASE 1 Familiarisation PHASE 2 Generating codes PHASE 3 Constructing themes PHASE 4 Reviewing themes PHASE 5 Defining and naming PHASE 6 Writing up Dashed line = the return loop Phase 4 routinely sends you back to phase 2. Analysis that ran straight through once has almost certainly not reviewed its themes properly.
The six phases, with the return loop that most published accounts leave out.

Phase 1 — Familiarisation

Read the entire dataset before coding anything. If you transcribed the interviews yourself, you have already begun; if you did not, read each transcript at least once without a highlighter in your hand. Take notes on early impressions, but resist turning them into codes — the point of this phase is to know the shape of the whole dataset before you start fragmenting it.

Practical marker: you should be able to describe what each interview was broadly about from memory before you code.

Phase 2 — Generating initial codes

Work systematically through the dataset, labelling segments that are relevant to your research question. Codes are short and descriptive: felt dismissed, waiting without information, relied on other parents. Code generously at this stage — it is far easier to merge codes later than to notice a pattern you never labelled.

Two rules save the most time. First, code the extract with enough surrounding text that it still makes sense when read in isolation, because in phase four you will read hundreds of extracts stripped of context. Second, allow the same extract to carry more than one code; real speech rarely does one thing at a time.

Phase 3 — Constructing candidate themes

Group codes that seem to speak to the same underlying idea. A theme is not a bucket of similar codes — it is a claim about what those codes mean taken together. If you can only describe a theme by listing its codes, it is a category, not yet a theme.

Phase 4 — Reviewing themes

This is the phase most often skipped, and the one that separates a defensible analysis from a plausible one. Review at two levels: first against the coded extracts, then against the entire dataset.

Does each theme hold together internally, or is it two ideas wearing one name?
Are the themes distinct from each other, or do two of them overlap so heavily they should merge?
Does the set of themes fairly represent the dataset, or only the most articulate participants?
Is there data that contradicts a theme? If so, it belongs in the analysis, not outside it.
Would a theme survive if you removed the two strongest extracts supporting it?
The two-extract test

If a theme collapses when its two best quotes are removed, it is a quotation with ambitions. Examiners find these quickly, because the same extract tends to appear every time the theme is discussed.

Phase 5 — Defining and naming

Write a short definition of each theme — two or three sentences saying what it captures and, importantly, what it excludes. If you cannot write the boundary, the theme is not yet defined. Names should be informative rather than clever; a reader should know what a theme is about from its name alone.

Phase 6 — Writing up

The write-up is analysis, not reporting. Extracts illustrate a claim you are making; they do not make it for you. A results section that is mostly quotations with connecting sentences is a common and heavily penalised pattern.

Stuck between phase 3 and phase 4?

This is where most projects stall — you have codes, you have candidate themes, and you cannot tell whether they hold. A statistician who works in qualitative methods can review your framework against your data and tell you what is a theme and what is a category.

See qualitative coding support

Worked example

A worked example: from extract to theme

Consider a study interviewing parents about accessing a support service. Below are six extracts, the codes applied to them, and the two candidate themes they were grouped into. The dataset in the real study was 22 interviews; this is a fragment.

EXTRACT CODE CANDIDATE THEME “I didn’t know who to ask” “Nobody explained the form” “I found out from another parent” “They talked over my head” “I felt like a nuisance” “You have to push to be heard” no obvious route in unexplained process informal information talked down to feeling a burden having to advocate Navigating in the dark 3 codes · 14 extracts Having to fight to matter 3 codes · 19 extracts Codes are descriptive. Themes are interpretative. A theme is not a summary of its codes — it is a claim about what they mean together.
Six extracts, six codes, two candidate themes. Note that the theme names make claims — they are not restatements of the codes.

Why these codes and not others

Take the extract “I found out from another parent”. It could have been coded word of mouth, which is descriptive and accurate. It was coded informal information instead, because the analytic interest was in how information reached parents at all, not in the specific channel. That decision is defensible either way — what matters is that it was made deliberately and recorded.

This is what a codebook entry looks like for one of these codes:

FieldEntry
Code nameinformal information
DefinitionInstances where a participant obtained information about the service other than from the service itself
IncludeOther parents, community groups, social media, chance conversations
ExcludeInformation obtained from the service but poorly explained (use unexplained process)
Anchor extract“I found out from another parent at the school gate, nobody had told me it existed”
Created / revisedPhase 2, revised phase 4 (merged with heard from friend)

From codes to a theme

The three codes no obvious route in, unexplained process and informal information were grouped as Navigating in the dark. The name makes a claim: that these parents were not merely uninformed but were required to find their own way through a system that assumed knowledge they did not have.

That claim is what makes it a theme. Had it been named “information” it would have been a category — a place to put things rather than something the analysis says.

What phase 4 changed

On review against the full dataset, a third candidate theme — “waiting” — was dissolved. Its codes distributed between the two themes above: waiting without information belonged to Navigating in the dark, and waiting while being treated as a nuisance belonged to Having to fight to matter. The waiting itself turned out not to be what participants were talking about.

This is the part to write down

Dissolving a theme is a finding about your data, and describing it in your methods chapter demonstrates that phase 4 actually happened. Most write-ups present the final themes as if they arrived fully formed.

Coding approach

Inductive, deductive, or both

A separate decision from which variant of thematic analysis you use is where your codes come from. Most methods chapters claim one and describe the other.

Inductive coding

  • Codes derived from the data itself
  • No pre-existing framework imposed
  • Suits exploratory questions and under-researched topics
  • Slower, and the codebook grows unpredictably
  • Risk: reinventing a framework that already exists in your literature

Deductive coding

  • Codes derived from theory or a prior framework
  • Applied to the data and refined
  • Suits evaluation against defined questions or an established model
  • Faster, and comparable with other studies using the same framework
  • Risk: seeing only what the framework anticipates

In practice almost every real analysis is hybrid. You begin with sensitising concepts from your literature, code openly, and find that some prior categories survive while others do not. That is a perfectly defensible approach — but it needs saying, because a chapter claiming purely inductive coding while using a framework lifted from a published model is an easy criticism to make.

A test for which you actually did

Look at your first ten codes. If you could have written them before reading a single transcript, your coding was more deductive than your methods chapter probably admits.

Where a study is deductive, name the framework, cite it, and say what you did with codes that did not fit it. Data that falls outside an a priori framework is often the most interesting finding in the dataset, and discarding it silently is the single biggest weakness of deductive coding done badly.

Which TA?

The three variants, and why the difference matters

Braun and Clarke have since been explicit that “thematic analysis” names a family of methods with genuinely different assumptions. Saying which one you used is now expected, and mixing them is the most common methodological criticism in review.

Reflexive TACodebook TACoding reliability TA
Codes areAnalyst's interpretationsA mix, structured in advanceDomain summaries, fixed early
CodebookDevelops throughoutDeveloped early, appliedFixed before main coding
Multiple codersOptional; for richness, not agreementCommonRequired
Kappa / IRRNot appropriateSometimesCentral
Researcher subjectivityA resourceManagedA problem to control
Best forExploratory, interpretative questionsApplied research with deadlinesTeam research needing consistency
The mismatch reviewers catch most often

Citing Braun and Clarke for reflexive thematic analysis and then reporting Cohen's kappa. Reflexive TA treats coding as interpretation, so an agreement statistic is not just unnecessary — it contradicts the epistemology you have just claimed.

Common problems

Seven mistakes that cost marks

MistakeWhy it costs youWhat to do instead
“Themes emerged from the data”Implies no analytic decisions were madeDescribe how themes were constructed and revised
Themes that are topics, not claims“Communication” tells the reader nothingName the finding: “Navigating in the dark”
Themes mirroring interview questionsSuggests the schedule was summarised, not analysedLook for patterns that cut across questions
Quotation-led results sectionsExtracts are illustration, not argumentMake the claim first, then evidence it
No negative casesReads as cherry-pickingReport data that complicates the theme
Reporting kappa with reflexive TAContradicts the stated epistemologyChoose the variant deliberately and be consistent
No account of phase 4The most important phase is invisibleSay what changed on review, and why

Six of these seven are write-up problems rather than analysis problems, which is worth noticing: most people do more analytic work than their methods chapter gives them credit for.

Methods and results

How to write it up

What belongs in the methods chapter

Which variant of thematic analysis, with a citation, and why it suits your question
Whether coding was inductive, deductive or both
Your epistemological position, in a sentence
How coding proceeded, including software
What happened at phase 4 — themes merged, split or dissolved
Whether anyone else coded, and what role they played (not necessarily agreement)
How many participants and extracts sit behind the final themes

What belongs in the results

One subsection per theme is conventional. Open each with the claim the theme makes, develop it across two or three paragraphs, and use extracts to evidence specific points rather than to carry the argument. Attribute extracts consistently (P07, or a pseudonym) and keep them short — a half-page quotation is almost always doing less work than three lines would.

A thematic map is optional but often earns its place, particularly where themes relate to each other rather than sitting in parallel. If you include one, it must match the themes as written; a map showing four themes above a results section describing three is a preventable inconsistency.

Worked example, continued

What a written-up theme actually looks like

Advice to “make the claim, then evidence it” is easier to give than to follow. Below is the opening of the Navigating in the dark theme from the worked example, written the way it would appear in a results chapter, with the moves annotated.

The written version

“Participants consistently described the service as something they had to find rather than something offered to them. This was not simply a matter of missing information: the process assumed a level of prior knowledge that most parents did not have, and the gap was filled informally or not at all. As P07 put it, ‘I found out from another parent at the school gate, nobody had told me it existed.’ Where information did arrive through official channels it was often unusable — P12 described a form that ‘might as well have been in another language’. The result was a period, sometimes months long, in which parents were nominally eligible for support but practically unable to reach it. Two participants who had navigated the system successfully both attributed this to prior professional experience of similar services rather than to anything the service itself had done.”

What that paragraph is doing

MoveWhereWhy it matters
States the claim firstSentence 1The reader knows what the theme argues before meeting any evidence
Refines the claimSentence 2Distinguishes this theme from a simpler ‘lack of information’ reading
Evidence, attributedP07Short, specific, and illustrating the claim already made
Second evidence, different angleP12Shows the pattern is not one participant's experience
ConsequenceSentence 5Moves from description to what it meant for participants
Negative / complicating caseFinal sentenceReports the participants who did succeed, and why — which strengthens rather than weakens the theme
The final sentence is the one examiners notice

Reporting the cases that ran counter to your theme, and explaining them, is the clearest signal that you analysed the dataset rather than assembled a case. Its absence is the most common reason a results chapter reads as advocacy.

How much extract to include

The two quotations above total nineteen words. That is deliberate. Long extracts feel like evidence but usually contain one usable clause surrounded by context the reader has to do the work of discarding. If an extract needs more than about thirty words, consider whether you are quoting because it proves the point or because you found it striking.

A practical ratio: across a results chapter, extracts should occupy perhaps a quarter of the words. Much more than that and the chapter is a collection; much less and the claims are unevidenced.

Self-assessment

A quality checklist before you submit

Braun and Clarke published a fifteen-point checklist alongside the original paper. The version below is adapted for what UK examiners and reviewers actually query, grouped by phase.

PhaseCheckCommon failure
TranscriptionTranscripts checked against recordingsErrors carried silently into coded extracts
CodingEvery data item given equal attentionThe first three interviews coded richly, the rest skimmed
CodingThemes derived from thorough coding, not anecdoteA theme built on one memorable participant
CodingAll relevant extracts collated for each themeExtracts found once and never revisited
ReviewThemes checked against the coded extracts and the full datasetOnly the first check performed
ReviewThemes internally coherent and mutually distinctTwo themes that are one idea under two names
NamingEach theme has a written definition and a boundaryThemes named but never defined
Write-upExtracts illustrate claims rather than replacing themQuotation-led results section
Write-upThe analytic approach is stated and matched to the epistemologyReflexive TA cited alongside a kappa statistic
Write-upNegative or contradictory cases reportedOnly confirming data presented
ThroughoutAnalytic decisions recorded as they were madeThe audit trail reconstructed from memory afterwards
Want a second opinion before it goes to your supervisor?

A results chapter review checks that the analysis suits the design, that the framework is defensible, and that no claim over-reaches the data — before the impression forms with the person marking it.

See results chapter review

Tools

Software: NVivo, and the alternatives

Thematic analysis does not require software. It requires organisation, and software provides it once a dataset passes roughly fifteen interviews.

OptionBest forWatch out for
NVivoMost UK universities; large datasets; coding comparisonLicence cost outside a university; the Mac version differs
MAXQDAMixed methods; visual toolsSmaller UK institutional presence
ATLAS.tiNetwork views; multimedia dataSteeper learning curve
DedooseTeam projects, browser-basedSubscription; data hosted externally — check with your ethics committee
Word or ExcelUnder ~10 interviewsNo audit trail; retrieval becomes unmanageable quickly

Whatever you use, the requirement is the same: at the end you should be able to retrieve every extract behind every theme in seconds. If you cannot, you will not be able to answer the question an examiner is most likely to ask.

Answers

Frequently asked questions

How many themes should I have?

Most studies report three to six. Fewer than three often means the themes are too broad to say anything; more than six usually means categories have been mistaken for themes, or that themes overlap and should be merged. There is no rule, and the number should follow from the data rather than be decided in advance.

How many interviews do I need for thematic analysis?

Between 15 and 30 is typical for a reflexive thematic analysis on a reasonably homogeneous group. A more heterogeneous sample needs more. What matters at examination is that you can justify the number in terms of your question and your data — not that it matches a published range.

Is thematic analysis the same as content analysis?

No. Content analysis counts occurrences and is often quantitative; thematic analysis interprets patterns of meaning and does not rely on frequency. A code appearing twice can carry a theme if what it shows is important; a code appearing forty times may be trivial.

Do I need two coders?

Not for reflexive thematic analysis, where coding is treated as interpretation rather than measurement. Coding reliability TA does require it, and codebook TA often uses it. Decide which variant you are using first, then the answer follows.

Can thematic analysis be deductive?

Yes. Codes can be derived from existing theory or a prior framework, applied to the data, and refined. Most real analyses are partly both — say which, rather than claiming purely inductive coding you did not do.

What is the difference between a code and a theme?

A code labels a segment of data descriptively. A theme is an interpretative claim about what a group of codes means together. If your theme can be fully described by listing its codes, it is still a category.

Should I report saturation?

Only if you can demonstrate it. Saturation means new data stopped generating new codes, which requires you to have tracked new codes per interview as you went. Asserting it retrospectively is a claim you cannot evidence, and reviewers increasingly ask.

Can I use thematic analysis on open-text survey responses?

Yes, and it is common. The main difference is that survey responses are short and lack the follow-up an interview provides, so codes tend to stay closer to the surface of the data. Say so in your limitations rather than writing up shallow themes as if they had interview depth behind them.

Can thematic analysis be used in mixed methods research?

Yes. It is one of the most common qualitative components in mixed methods designs, partly because it does not commit you to a theoretical framework that might conflict with the quantitative strand. State how the strands relate — whether the qualitative work explains, expands or triangulates the quantitative findings.

How long does thematic analysis take?

For 20 to 30 interviews, expect three to five weeks of concentrated work if you are experienced, and considerably longer if it is your first project. Familiarisation and phase 4 take far more time than people budget for; initial coding usually takes less than expected.

What is a thematic map, and do I need one?

A thematic map is a diagram showing the themes and how they relate. It is optional. It earns its place where themes are connected rather than parallel, and it must match the themes as written — a map showing four themes above a results section describing three is an avoidable inconsistency.

Do I need to report how many participants mentioned each theme?

Not in reflexive thematic analysis, where prevalence is not the measure of importance. If you do give counts, be consistent and explain why they matter, because mixing an interpretative method with frequency claims invites the question of why you did not simply run a content analysis.

Send the data. Get a fixed quote.

Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.