Guides / thematic analysis
Thematic analysis: a complete guide, with a worked example
Thematic analysis is the most widely used qualitative method in the social and health sciences, and the most widely misreported. This guide covers what it is, the six phases as Braun and Clarke actually defined them, a coding example worked through from extract to theme, and the mistakes that cost marks in a viva.
Thematic analysis is a qualitative method for identifying, analysing and reporting patterns of meaning across a dataset. It involves coding extracts of data, grouping those codes, and developing them into themes that answer a research question. Unlike content analysis it does not count occurrences; it interprets what the patterns mean.
Definition
What thematic analysis is
Thematic analysis is a method for identifying, analysing and reporting patterns of meaning — themes — across a qualitative dataset. You read the data closely, label segments of it with codes, group those codes, and develop the groups into themes that answer your research question.
It was formalised for the social sciences by Virginia Braun and Victoria Clarke in 2006, in a paper that has since become one of the most cited in qualitative research. Their central argument was that thematic analysis was already being used everywhere but rarely described properly, so it looked like an absence of method rather than a method in its own right.
Three properties distinguish it from the methods it is most often confused with.
| Method | What it produces | How it differs |
|---|---|---|
| Thematic analysis | Interpretative themes across a dataset | Not tied to a theoretical framework; flexible across epistemologies |
| Content analysis | Frequencies and categories | Counts occurrences; usually more descriptive and often quantitative |
| Grounded theory | A theory grounded in the data | Aims to generate theory; requires theoretical sampling and constant comparison |
| IPA | Detailed accounts of lived experience | Phenomenological, idiographic, small samples, double hermeneutic |
| Framework analysis | A case-by-theme matrix | Partly a priori; built for applied policy research and cross-case comparison |
“Themes emerged from the data” is the phrase examiners most reliably challenge. Themes do not emerge. An analyst constructs them by making decisions, and describing those decisions is what makes the analysis assessable rather than merely asserted.
Choosing the method
When to use it — and when not to
Thematic analysis is flexible, which is its strength and its trap. Because it can be applied to almost any qualitative dataset, it is often chosen by default rather than because it fits the question.
Use thematic analysis when
- You want to identify patterns of meaning across participants
- The research question is about what people think, feel or experience
- You have a reasonably sized dataset — interviews, focus groups, open-text responses
- You need a method that is accessible to a mixed-methods audience
- Your question is exploratory or descriptive rather than theory-generating
Use something else when
- You need a theory as the output → grounded theory
- You need the detailed texture of one person's experience → IPA
- You need cross-case comparison in a matrix → framework analysis
- You genuinely need frequencies → content analysis
- Your interest is in how language is used rather than what is said → discourse analysis
The honest test: if you cannot say what thematic analysis gives you that another method would not, you have probably chosen it because it is familiar. That is a defensible reason to change method at the design stage and an uncomfortable one to defend at a viva.
Braun and Clarke
The six phases, in practice
The six phases are widely quoted and widely misunderstood. They are not a linear pipeline: phases four and five routinely send you back to phase two, and Braun and Clarke are explicit that the process is recursive.
Phase 1 — Familiarisation
Read the entire dataset before coding anything. If you transcribed the interviews yourself, you have already begun; if you did not, read each transcript at least once without a highlighter in your hand. Take notes on early impressions, but resist turning them into codes — the point of this phase is to know the shape of the whole dataset before you start fragmenting it.
Practical marker: you should be able to describe what each interview was broadly about from memory before you code.
Phase 2 — Generating initial codes
Work systematically through the dataset, labelling segments that are relevant to your research question. Codes are short and descriptive: felt dismissed, waiting without information, relied on other parents. Code generously at this stage — it is far easier to merge codes later than to notice a pattern you never labelled.
Two rules save the most time. First, code the extract with enough surrounding text that it still makes sense when read in isolation, because in phase four you will read hundreds of extracts stripped of context. Second, allow the same extract to carry more than one code; real speech rarely does one thing at a time.
Phase 3 — Constructing candidate themes
Group codes that seem to speak to the same underlying idea. A theme is not a bucket of similar codes — it is a claim about what those codes mean taken together. If you can only describe a theme by listing its codes, it is a category, not yet a theme.
Phase 4 — Reviewing themes
This is the phase most often skipped, and the one that separates a defensible analysis from a plausible one. Review at two levels: first against the coded extracts, then against the entire dataset.
If a theme collapses when its two best quotes are removed, it is a quotation with ambitions. Examiners find these quickly, because the same extract tends to appear every time the theme is discussed.
Phase 5 — Defining and naming
Write a short definition of each theme — two or three sentences saying what it captures and, importantly, what it excludes. If you cannot write the boundary, the theme is not yet defined. Names should be informative rather than clever; a reader should know what a theme is about from its name alone.
Phase 6 — Writing up
The write-up is analysis, not reporting. Extracts illustrate a claim you are making; they do not make it for you. A results section that is mostly quotations with connecting sentences is a common and heavily penalised pattern.
This is where most projects stall — you have codes, you have candidate themes, and you cannot tell whether they hold. A statistician who works in qualitative methods can review your framework against your data and tell you what is a theme and what is a category.
See qualitative coding supportWorked example
A worked example: from extract to theme
Consider a study interviewing parents about accessing a support service. Below are six extracts, the codes applied to them, and the two candidate themes they were grouped into. The dataset in the real study was 22 interviews; this is a fragment.
Why these codes and not others
Take the extract “I found out from another parent”. It could have been coded word of mouth, which is descriptive and accurate. It was coded informal information instead, because the analytic interest was in how information reached parents at all, not in the specific channel. That decision is defensible either way — what matters is that it was made deliberately and recorded.
This is what a codebook entry looks like for one of these codes:
| Field | Entry |
|---|---|
| Code name | informal information |
| Definition | Instances where a participant obtained information about the service other than from the service itself |
| Include | Other parents, community groups, social media, chance conversations |
| Exclude | Information obtained from the service but poorly explained (use unexplained process) |
| Anchor extract | “I found out from another parent at the school gate, nobody had told me it existed” |
| Created / revised | Phase 2, revised phase 4 (merged with heard from friend) |
From codes to a theme
The three codes no obvious route in, unexplained process and informal information were grouped as Navigating in the dark. The name makes a claim: that these parents were not merely uninformed but were required to find their own way through a system that assumed knowledge they did not have.
That claim is what makes it a theme. Had it been named “information” it would have been a category — a place to put things rather than something the analysis says.
What phase 4 changed
On review against the full dataset, a third candidate theme — “waiting” — was dissolved. Its codes distributed between the two themes above: waiting without information belonged to Navigating in the dark, and waiting while being treated as a nuisance belonged to Having to fight to matter. The waiting itself turned out not to be what participants were talking about.
Dissolving a theme is a finding about your data, and describing it in your methods chapter demonstrates that phase 4 actually happened. Most write-ups present the final themes as if they arrived fully formed.
Coding approach
Inductive, deductive, or both
A separate decision from which variant of thematic analysis you use is where your codes come from. Most methods chapters claim one and describe the other.
Inductive coding
- Codes derived from the data itself
- No pre-existing framework imposed
- Suits exploratory questions and under-researched topics
- Slower, and the codebook grows unpredictably
- Risk: reinventing a framework that already exists in your literature
Deductive coding
- Codes derived from theory or a prior framework
- Applied to the data and refined
- Suits evaluation against defined questions or an established model
- Faster, and comparable with other studies using the same framework
- Risk: seeing only what the framework anticipates
In practice almost every real analysis is hybrid. You begin with sensitising concepts from your literature, code openly, and find that some prior categories survive while others do not. That is a perfectly defensible approach — but it needs saying, because a chapter claiming purely inductive coding while using a framework lifted from a published model is an easy criticism to make.
Look at your first ten codes. If you could have written them before reading a single transcript, your coding was more deductive than your methods chapter probably admits.
Where a study is deductive, name the framework, cite it, and say what you did with codes that did not fit it. Data that falls outside an a priori framework is often the most interesting finding in the dataset, and discarding it silently is the single biggest weakness of deductive coding done badly.
Which TA?
The three variants, and why the difference matters
Braun and Clarke have since been explicit that “thematic analysis” names a family of methods with genuinely different assumptions. Saying which one you used is now expected, and mixing them is the most common methodological criticism in review.
| Reflexive TA | Codebook TA | Coding reliability TA | |
|---|---|---|---|
| Codes are | Analyst's interpretations | A mix, structured in advance | Domain summaries, fixed early |
| Codebook | Develops throughout | Developed early, applied | Fixed before main coding |
| Multiple coders | Optional; for richness, not agreement | Common | Required |
| Kappa / IRR | Not appropriate | Sometimes | Central |
| Researcher subjectivity | A resource | Managed | A problem to control |
| Best for | Exploratory, interpretative questions | Applied research with deadlines | Team research needing consistency |
Citing Braun and Clarke for reflexive thematic analysis and then reporting Cohen's kappa. Reflexive TA treats coding as interpretation, so an agreement statistic is not just unnecessary — it contradicts the epistemology you have just claimed.
Common problems
Seven mistakes that cost marks
| Mistake | Why it costs you | What to do instead |
|---|---|---|
| “Themes emerged from the data” | Implies no analytic decisions were made | Describe how themes were constructed and revised |
| Themes that are topics, not claims | “Communication” tells the reader nothing | Name the finding: “Navigating in the dark” |
| Themes mirroring interview questions | Suggests the schedule was summarised, not analysed | Look for patterns that cut across questions |
| Quotation-led results sections | Extracts are illustration, not argument | Make the claim first, then evidence it |
| No negative cases | Reads as cherry-picking | Report data that complicates the theme |
| Reporting kappa with reflexive TA | Contradicts the stated epistemology | Choose the variant deliberately and be consistent |
| No account of phase 4 | The most important phase is invisible | Say what changed on review, and why |
Six of these seven are write-up problems rather than analysis problems, which is worth noticing: most people do more analytic work than their methods chapter gives them credit for.
Methods and results
How to write it up
What belongs in the methods chapter
What belongs in the results
One subsection per theme is conventional. Open each with the claim the theme makes, develop it across two or three paragraphs, and use extracts to evidence specific points rather than to carry the argument. Attribute extracts consistently (P07, or a pseudonym) and keep them short — a half-page quotation is almost always doing less work than three lines would.
A thematic map is optional but often earns its place, particularly where themes relate to each other rather than sitting in parallel. If you include one, it must match the themes as written; a map showing four themes above a results section describing three is a preventable inconsistency.
Worked example, continued
What a written-up theme actually looks like
Advice to “make the claim, then evidence it” is easier to give than to follow. Below is the opening of the Navigating in the dark theme from the worked example, written the way it would appear in a results chapter, with the moves annotated.
The written version
“Participants consistently described the service as something they had to find rather than something offered to them. This was not simply a matter of missing information: the process assumed a level of prior knowledge that most parents did not have, and the gap was filled informally or not at all. As P07 put it, ‘I found out from another parent at the school gate, nobody had told me it existed.’ Where information did arrive through official channels it was often unusable — P12 described a form that ‘might as well have been in another language’. The result was a period, sometimes months long, in which parents were nominally eligible for support but practically unable to reach it. Two participants who had navigated the system successfully both attributed this to prior professional experience of similar services rather than to anything the service itself had done.”
What that paragraph is doing
| Move | Where | Why it matters |
|---|---|---|
| States the claim first | Sentence 1 | The reader knows what the theme argues before meeting any evidence |
| Refines the claim | Sentence 2 | Distinguishes this theme from a simpler ‘lack of information’ reading |
| Evidence, attributed | P07 | Short, specific, and illustrating the claim already made |
| Second evidence, different angle | P12 | Shows the pattern is not one participant's experience |
| Consequence | Sentence 5 | Moves from description to what it meant for participants |
| Negative / complicating case | Final sentence | Reports the participants who did succeed, and why — which strengthens rather than weakens the theme |
Reporting the cases that ran counter to your theme, and explaining them, is the clearest signal that you analysed the dataset rather than assembled a case. Its absence is the most common reason a results chapter reads as advocacy.
How much extract to include
The two quotations above total nineteen words. That is deliberate. Long extracts feel like evidence but usually contain one usable clause surrounded by context the reader has to do the work of discarding. If an extract needs more than about thirty words, consider whether you are quoting because it proves the point or because you found it striking.
A practical ratio: across a results chapter, extracts should occupy perhaps a quarter of the words. Much more than that and the chapter is a collection; much less and the claims are unevidenced.
Self-assessment
A quality checklist before you submit
Braun and Clarke published a fifteen-point checklist alongside the original paper. The version below is adapted for what UK examiners and reviewers actually query, grouped by phase.
| Phase | Check | Common failure |
|---|---|---|
| Transcription | Transcripts checked against recordings | Errors carried silently into coded extracts |
| Coding | Every data item given equal attention | The first three interviews coded richly, the rest skimmed |
| Coding | Themes derived from thorough coding, not anecdote | A theme built on one memorable participant |
| Coding | All relevant extracts collated for each theme | Extracts found once and never revisited |
| Review | Themes checked against the coded extracts and the full dataset | Only the first check performed |
| Review | Themes internally coherent and mutually distinct | Two themes that are one idea under two names |
| Naming | Each theme has a written definition and a boundary | Themes named but never defined |
| Write-up | Extracts illustrate claims rather than replacing them | Quotation-led results section |
| Write-up | The analytic approach is stated and matched to the epistemology | Reflexive TA cited alongside a kappa statistic |
| Write-up | Negative or contradictory cases reported | Only confirming data presented |
| Throughout | Analytic decisions recorded as they were made | The audit trail reconstructed from memory afterwards |
A results chapter review checks that the analysis suits the design, that the framework is defensible, and that no claim over-reaches the data — before the impression forms with the person marking it.
See results chapter reviewTools
Software: NVivo, and the alternatives
Thematic analysis does not require software. It requires organisation, and software provides it once a dataset passes roughly fifteen interviews.
| Option | Best for | Watch out for |
|---|---|---|
| NVivo | Most UK universities; large datasets; coding comparison | Licence cost outside a university; the Mac version differs |
| MAXQDA | Mixed methods; visual tools | Smaller UK institutional presence |
| ATLAS.ti | Network views; multimedia data | Steeper learning curve |
| Dedoose | Team projects, browser-based | Subscription; data hosted externally — check with your ethics committee |
| Word or Excel | Under ~10 interviews | No audit trail; retrieval becomes unmanageable quickly |
Whatever you use, the requirement is the same: at the end you should be able to retrieve every extract behind every theme in seconds. If you cannot, you will not be able to answer the question an examiner is most likely to ask.
Answers
Frequently asked questions
How many themes should I have?
Most studies report three to six. Fewer than three often means the themes are too broad to say anything; more than six usually means categories have been mistaken for themes, or that themes overlap and should be merged. There is no rule, and the number should follow from the data rather than be decided in advance.
How many interviews do I need for thematic analysis?
Between 15 and 30 is typical for a reflexive thematic analysis on a reasonably homogeneous group. A more heterogeneous sample needs more. What matters at examination is that you can justify the number in terms of your question and your data — not that it matches a published range.
Is thematic analysis the same as content analysis?
No. Content analysis counts occurrences and is often quantitative; thematic analysis interprets patterns of meaning and does not rely on frequency. A code appearing twice can carry a theme if what it shows is important; a code appearing forty times may be trivial.
Do I need two coders?
Not for reflexive thematic analysis, where coding is treated as interpretation rather than measurement. Coding reliability TA does require it, and codebook TA often uses it. Decide which variant you are using first, then the answer follows.
Can thematic analysis be deductive?
Yes. Codes can be derived from existing theory or a prior framework, applied to the data, and refined. Most real analyses are partly both — say which, rather than claiming purely inductive coding you did not do.
What is the difference between a code and a theme?
A code labels a segment of data descriptively. A theme is an interpretative claim about what a group of codes means together. If your theme can be fully described by listing its codes, it is still a category.
Should I report saturation?
Only if you can demonstrate it. Saturation means new data stopped generating new codes, which requires you to have tracked new codes per interview as you went. Asserting it retrospectively is a claim you cannot evidence, and reviewers increasingly ask.
Can I use thematic analysis on open-text survey responses?
Yes, and it is common. The main difference is that survey responses are short and lack the follow-up an interview provides, so codes tend to stay closer to the surface of the data. Say so in your limitations rather than writing up shallow themes as if they had interview depth behind them.
Can thematic analysis be used in mixed methods research?
Yes. It is one of the most common qualitative components in mixed methods designs, partly because it does not commit you to a theoretical framework that might conflict with the quantitative strand. State how the strands relate — whether the qualitative work explains, expands or triangulates the quantitative findings.
How long does thematic analysis take?
For 20 to 30 interviews, expect three to five weeks of concentrated work if you are experienced, and considerably longer if it is your first project. Familiarisation and phase 4 take far more time than people budget for; initial coding usually takes less than expected.
What is a thematic map, and do I need one?
A thematic map is a diagram showing the themes and how they relate. It is optional. It earns its place where themes are connected rather than parallel, and it must match the themes as written — a map showing four themes above a results section describing three is an avoidable inconsistency.
Do I need to report how many participants mentioned each theme?
Not in reflexive thematic analysis, where prevalence is not the measure of importance. If you do give counts, be consistent and explain why they matter, because mixing an interpretative method with frequency claims invites the question of why you did not simply run a content analysis.
Keep reading
Related guides and services
Content analysis
The other way to analyse qualitative material — and when frequency is the right evidence.
ServiceQualitative coding and thematic analysis
Codebooks, systematic coding and themes evidenced back to the extracts.
Case studyReflexive thematic analysis of 45 interviews
How a coding framework was built across 700 pages of transcripts.
Send the data. Get a fixed quote.
Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.