Guides  /  nvivo for thematic analysis

How to use NVivo for thematic analysis

NVivo does not do thematic analysis, but it does make phases three and four considerably easier — and phase four is the one most projects skip. This maps the six phases onto what you actually click, and flags where the software can push you the wrong way.

Written and reviewed by Elaine Halliburton, Professor of Applied Statistics
Updated 2 September 20267 min read
Can you do thematic analysis in NVivo?

Yes. NVivo supports thematic analysis by holding all transcripts in one project, letting you code passages to nodes, group those nodes into candidate themes, and retrieve every extract at a theme for review. It does not identify themes for you — the analytic judgement at every phase remains the researcher's.

Overview

The six phases in NVivo

Braun & Clarke phase What that is in NVivo 1. FamiliarisationRead sources. Write memos. Do not code. 2. Generating codesCode to nodes. Flat list, no hierarchy yet. 3. Searching for themesGroup nodes into parents. Merge duplicates. 4. Reviewing themesOpen each node, read every reference back. 5. Defining and namingWrite the node description. That IS the definition. 6. Producing the reportCoding query to pull extracts for writing. Phase 4 is where NVivo genuinely earns its place — and where most projects skip straight past.
Braun and Clarke's phases and their NVivo equivalents. Phase 4 is where the software pays for itself.
Name your variant

Reflexive thematic analysis, codebook and coding reliability approaches make different assumptions and expect different things of you. Braun and Clarke are explicit that reflexive TA does not use inter-rater reliability. State which variant you followed and cite it — examiners ask.

Reading

Phase 1: familiarisation

Import your sources and read them. Do not code anything yet.

The temptation with NVivo open is to start tagging on the first pass, and it produces codes that describe individual sentences rather than meaning. Read each transcript through, then read it again.

Use annotations for margin notes on specific passages — select text, Annotate
Use memos for thinking that spans a whole transcript — Create → Memo, linked to the source
Do not create nodes yet

Coding

Phase 2: generating codes

Now code. Select passages and code to new nodes, keeping the list flat — no parents, no structure.

Name codes for what is happening rather than the topic. “Support” is a topic; “seeking support without appearing to need it” names a process and carries the tension that makes it analytically useful.

Expect the list to grow to fifty or eighty codes across the first several transcripts. That is normal and not a sign of disorganisation.

Write the node description as you create it

Two sentences on what counts and what does not. This is your codebook, it is what keeps coding consistent at transcript thirty, and it becomes phase five's theme definition almost verbatim.

The core

Phases 3 and 4: themes and review

Phase 3 — searching for themes

Group nodes into parents by what they share analytically. Dragging one node onto another merges them; dragging into a parent nests them. Merging is safe — references are combined, nothing is lost.

A theme is not a summary of a topic. It should say something: a shared tension, a pattern of action, a way participants handle a situation.

Phase 4 — reviewing themes

This is where NVivo genuinely earns its place, and where most projects skip straight past.

Open each candidate theme node and read every reference coded there, in sequence. You are checking two things: that the extracts cohere with each other, and that the theme holds across the dataset rather than resting on two articulate participants.

Do the extracts belong together? If half of them are about something else, the theme needs splitting
How many participants contribute? A node with 30 references from two people is not a dataset-wide theme
Check the source count — NVivo shows references and sources separately for exactly this reason
Re-read the uncoded material for anything the theme structure is missing
References versus sources

NVivo displays both. A theme with 47 references across 3 sources and one with 47 references across 18 sources are entirely different findings. The second is a pattern; the first is two people talking at length.

Want the theme structure tested before you write it up?

Send your transcripts and node structure. A named researcher reviews whether the themes hold against the extracts and where they need splitting or merging.

See qualitative coding

Finishing

Phases 5 and 6: defining and writing

Phase 5 — defining and naming

If you wrote node descriptions as you went, this phase is largely done. Each theme needs a definition that says what it is about and what its boundaries are, plus a name that is informative rather than a single word.

Phase 6 — producing the report

Use a coding query to pull every extract at a theme, filtered by case attribute if you are comparing groups. Export it and write from that rather than scrolling the project.

Extracts need context

An extract that made sense while you were coding may not on the page. Include enough surrounding turns for the reader to follow, and give the participant identifier consistently.

Cautions

Where NVivo pushes you wrong

The featureThe risk
Word frequency and word cloudsAttractive and nearly meaningless. Common words in interviews are fillers. Orientation only, never evidence
Auto-coding by sentimentAssigns positive or negative labels without the interpretive judgement TA requires
AI-assisted codingUseful for structural splitting; not a substitute for analytic coding. Say what it did and what you did
The hierarchy viewInvites building structure early, before the data have earned it
Reference countsTempting to treat as prevalence. Frequency is not importance in reflexive TA
Counting is not thematic analysis

It is easy to write “this theme had 47 references” because NVivo displays the number. In reflexive TA, prevalence is not what makes a theme significant — a pattern mentioned by four participants may matter more than one mentioned in passing by twenty. If you report counts, say why they are relevant.

Reporting

What to write in your methods

Name the variant of TA and cite it — reflexive, codebook, or coding reliability
State that NVivo was used for data management and coding, not for analysis
Give the version — NVivo 14, NVivo 12 and so on
Describe how codes became themes, which is the step examiners probe hardest
Mention memos if you kept them, as evidence of the analytic trail
Only report kappa if your variant calls for it — reflexive TA does not
A sentence that earns marks

“Transcripts were managed and coded in NVivo 14. Following Braun and Clarke's reflexive thematic analysis, initial codes were generated inductively across all 22 transcripts before any grouping was attempted. Candidate themes were reviewed by reading every extract coded at each node and checking distribution across sources; two candidate themes were collapsed at this stage as the extracts did not cohere. Analytic memos were kept throughout.”

Answers

Frequently asked questions

Can you do thematic analysis in NVivo?

Yes, and it suits it well — particularly phases three and four, where you group codes into candidate themes and then read every extract back to check the theme holds. NVivo manages and retrieves; it does not identify themes or judge which are analytically significant.

Do I need NVivo for thematic analysis?

No. Braun and Clarke are explicit that software is not required, and thematic analysis is done well in word processors and spreadsheets. NVivo earns its place at roughly twenty transcripts or more, or when you need to compare systematically across participant groups.

How do Braun and Clarke's six phases map onto NVivo?

Familiarisation is reading and memoing without coding. Generating codes is coding to a flat node list. Searching for themes is grouping nodes into parents. Reviewing themes is opening each node and reading every reference. Defining and naming is writing the node description. Producing the report is running a coding query to pull extracts.

Should I report inter-rater reliability for thematic analysis in NVivo?

It depends on the variant. Codebook and coding reliability approaches expect it, and NVivo computes kappa through Explore > Coding Comparison. Reflexive thematic analysis does not — Braun and Clarke argue the researcher's interpretation is the analytic instrument rather than a source of error.

Does the number of references show how important a theme is?

Not in reflexive thematic analysis. Prevalence is not significance: a pattern raised by four participants may matter more than one mentioned in passing by twenty. Check the source count as well as the reference count, and if you report frequencies, explain why they are relevant.

What should I write in my methods section about NVivo?

Name the variant of thematic analysis and cite it, state the NVivo version, say the software was used for data management and coding rather than analysis, and describe how codes became themes. That last step is what examiners probe hardest.