Guides  /  grounded theory

Grounded theory: a practical guide

Grounded theory builds a theory from data rather than testing one against it, and it is the qualitative method most often claimed and least often actually carried out. This guide sets out the coding stages, explains the two features that distinguish real grounded theory from thematic analysis with different labels, and covers what an examiner will look for.

Elaine Halliburton Written and reviewed by Elaine Halliburton, Professor of Applied Statistics
Updated 17 August 202618 min read
What is grounded theory?

Grounded theory is a qualitative research method in which theory is developed inductively from data rather than tested against a pre-existing framework. Data collection and analysis proceed together, with each round of analysis determining what is collected next, and coding moves from many descriptive codes to a small number of categories and finally to a core category that explains the central process. It ends when further data produce no new insight.

Definition

What grounded theory is

Grounded theory develops a theory from the data rather than testing a theory against them. You begin without a hypothesis, code closely, and build upwards until you have an explanatory account of the process you were studying.

Its output is distinctive. Thematic analysis produces themes describing what participants said; grounded theory produces a theory — an account of how something works, what conditions bring it about, and what consequences follow. That difference in output is what should drive the choice of method.

The method originated in the 1960s as a deliberate corrective. Sociology at the time was dominated by grand theorising tested through survey work, and Glaser and Strauss argued that this produced theory disconnected from the settings it claimed to describe. Their alternative was to derive theory from close, systematic engagement with data — hence “grounded”. Understanding that origin helps explain features that otherwise look arbitrary, particularly the suspicion of an early literature review, which was aimed at preventing existing frameworks from being imposed on data that might not fit them.

Choose it for the question, not the label

Grounded theory suits questions about process: how something unfolds, how people manage a situation over time, how a practice is sustained. If your question is “what do participants think about X?”, thematic analysis is the honest answer and will be easier to carry out well.

Variants

The three schools, and choosing one

Grounded theory split into distinct traditions after its founders disagreed, and examiners expect you to know which one you followed.

SchoolKey figuresStanceCoding
ClassicGlaserTheory emerges; the literature is read lateOpen, selective, theoretical
StraussianStrauss & CorbinMore structured; a coding paradigm guides analysisOpen, axial, selective
ConstructivistCharmazTheory is co-constructed; the researcher is present in itInitial, focused, theoretical

The Straussian version is the most prescriptive and therefore the easiest to defend in a viva, because each step has a stated procedure. The constructivist version is the most widely used in contemporary social science and is the most comfortable fit for a researcher who cannot plausibly claim to have approached the topic with no prior knowledge.

Name your school and stick to it

Mixing Glaser's rejection of a prior literature review with Corbin's coding paradigm produces a method that belongs to nobody, and it is the kind of inconsistency an examiner will press on. State which tradition you followed, cite it, and apply it consistently.

The distinction

What makes it grounded theory and not thematic analysis

collectcodememo saturationstop theoretical sampling — what to collect next is decided by what the analysis so far needs Collection and analysis run together. You do not gather all the data and then start coding. This is the feature that most distinguishes grounded theory, and the one most often dropped.
Data collection and analysis alternate. What you collect next is determined by what the emerging analysis requires.

Two features distinguish genuine grounded theory, and a study lacking either is thematic analysis under a different name.

FeatureGrounded theoryThematic analysis
Data collectionIterative — interleaved with analysisAll collected, then analysed
SamplingTheoretical — driven by the emerging theoryPurposive, decided in advance
Stopping ruleTheoretical saturationPlanned sample size
OutputAn explanatory theoryDescriptive or interpretive themes
Literature reviewLate, or used cautiouslyNormally up front
The commonest failure in dissertations

Interviewing twenty people, then coding all twenty transcripts, then reporting categories — and calling it grounded theory. Without iterative collection and theoretical sampling, the two defining procedures are absent. Examiners check for this specifically, and the safer route is to describe accurately what you did rather than to claim a method you did not follow.

Not sure which qualitative method your question calls for?

Send your research question and the data you have or plan to collect. A named researcher advises which method fits and what it will require of you.

Get a fixed quote

Coding

The three stages of coding

Coding narrows as the analysis proceeds OPEN CODING — many codes, close to the data, often in participants’ own words category Acategory Bcategory C AXIAL CODING — codes grouped into categories, and relationships between them identified CORE CATEGORY SELECTIVE CODING — one core category that explains the most variation, with the others related to it The output is a theory grounded in the data, not a description of it.
Coding narrows from many descriptive codes through categories to a single core category.

Open coding

Work line by line or incident by incident, naming what is happening in the data. Codes stay close to the material and are often gerunds — “managing uncertainty”, “deferring to expertise” — because grounded theory is interested in action and process rather than in topics. Expect a large number of codes at this stage; several hundred from a handful of transcripts is normal.

Axial coding

Group codes into categories and, crucially, work out how the categories relate. The Straussian coding paradigm offers a scaffold for this: for each category, identify the conditions that give rise to it, the actions and interactions involved, and the consequences that follow. Not every analysis needs the paradigm applied mechanically, but the relational question — how does this connect to that? — is what separates axial coding from sorting.

Selective coding

Identify the core category: the one that accounts for the most variation and to which the others can be related. Everything else is then integrated around it, and the resulting account is the theory. If no core category emerges, the analysis is not finished, and reporting a list of parallel categories instead is the point at which many grounded theory studies quietly become thematic analyses.

Software helps with the mechanics but does not do the analysis. NVivo, MAXQDA and Dedoose all manage codes, retrieve extracts and store memos efficiently, which matters once you have several hundred codes across twenty transcripts. None of them identifies a core category for you, and a coding tree that looks tidy in software can still be a description rather than a theory.

Codes should be actions, not topics

“Support” is a topic. “Seeking support without appearing to need it” is a process, and it carries the tension that makes a theory worth building. Naming codes as gerunds is a practical discipline that keeps the analysis oriented towards process.

Procedure

Theoretical sampling and constant comparison

These are the two engines of the method, and both are procedural rather than conceptual — you either did them or you did not.

Theoretical sampling

After the first few interviews you analyse, notice a gap or an ambiguity in the emerging categories, and choose the next participants specifically to address it. If a category about “managing uncertainty” is developing but every participant so far has been experienced, you deliberately recruit novices next to see whether the category holds.

This means your sampling strategy cannot be fully specified in advance, which has implications for ethics applications. Write the protocol to allow it: state that sampling will be theoretical, that later participants will be selected on the basis of emerging analysis, and give the broad pool you will draw from.

In practice, theoretical sampling rarely means recruiting entirely new participants each round. It may mean returning to earlier participants with sharper questions, seeking documents that bear on a developing category, or observing a setting you had not planned to visit. What defines it is not the source but the reasoning: the choice follows from what the analysis needs, and that reasoning is recorded.

Constant comparison

Every new piece of data is compared against the codes and categories already developed, and every code against the others. The comparisons are what force categories to become precise: encountering a case that does not fit obliges you either to redefine the category or to identify the condition under which it does not apply.

Keep an audit trail

Record why each sampling decision was taken and what the analysis at that point required. This trail is the main evidence that theoretical sampling actually occurred, and reconstructing it afterwards is both difficult and obvious to a reader.

Saturation

Saturation, and how to justify stopping

Theoretical saturation is reached when further data no longer add to the properties of your categories. It is a property of the categories, not of the number of interviews, and the distinction matters when you come to defend it.

New data yield no new codes within the categories that matter
Categories are well developed in their properties and dimensions
Relationships between categories are established and stable
The core category accounts for the variation you have seen

Sample sizes in published grounded theory studies commonly fall between 20 and 40 participants, but citing a number is not a justification. What defends the decision is describing the point at which categories stopped developing and showing the evidence — typically that the last several interviews contributed nothing new to any core category.

Saturation is frequently overclaimed

“Saturation was reached after 12 interviews” with no supporting detail is a claim examiners routinely challenge. Say which categories saturated, at approximately what point, and what the final interviews added. If you stopped for practical reasons — time, access, funding — say so and treat it as a limitation. That is far stronger than an unevidenced claim.

Memos

Memo writing

Memos are analytic notes written throughout, and in grounded theory they are not optional. They are where the theory is actually developed; the codes are only its raw material.

Memo typeContents
Code memoWhat a code means, its boundaries, an example extract
Theoretical memoHunches about relationships between categories
Operational memoSampling and procedural decisions, with reasons
Integrative memoHow categories fit together as the core emerges

Write them immediately, date them, and do not edit them for style. Their value lies in capturing thinking as it develops, including the ideas later abandoned — which is precisely what lets you reconstruct and defend the analytic path in a viva.

There is no minimum, but a study producing a handful of memos across several months has probably not been doing the analytic work the method requires. Many researchers find they write more memo text than they eventually write findings text, and that is a reasonable sign the analysis is developing rather than merely accumulating.

Memos are the answer to 'how did you get from codes to theory?'

This is the question grounded theory candidates find hardest, because the step is genuinely interpretive. A dated sequence of memos showing the reasoning is the only convincing answer, and it cannot be produced retrospectively.

Get your coding framework and analysis reviewed

Send your transcripts and coding so far. A named researcher reviews the codebook, tests the categories against the extracts, and advises on developing the core category.

See qualitative coding

Worked example

Worked example: from extract to category

A study explores how newly qualified practitioners manage situations beyond their competence. Three extracts from different participants, and the codes assigned during initial coding.

ExtractInitial codes
“I'd ask, but I'd phrase it like I was checking rather than like I didn't know.”disguising the request; protecting credibility
“You learn quite fast who you can look stupid in front of and who you can't.”mapping safe others; assessing risk of exposure
“Sometimes I'd just look it up afterwards rather than ask at the time.”deferring the question; avoiding visible uncertainty

Moving to a category

Every code involves a tension between needing information and managing how the need appears to others. Grouping them produces a candidate category: managing visible uncertainty. Note that it is stated as a process, not a topic — “uncertainty” alone would describe a subject area, while “managing visible uncertainty” names something participants are doing.

Developing its properties

PropertyDimensions observed
Who the audience isSenior colleague ↔ peer ↔ patient or client
Cost of exposureTrivial ↔ career-relevant
Strategy usedDisguising ↔ deferring ↔ asking openly
TimingIn the moment ↔ afterwards

What theoretical sampling does next

Every participant so far has been newly qualified. To test whether the category is specific to inexperience or persists, the next round deliberately recruits practitioners with ten or more years in post. If they describe the same management of visible uncertainty, the category is broader than first assumed and the theory must account for what sustains it. If they do not, the conditions under which it operates become part of the theory.

This is the step that makes it grounded theory

The decision about who to recruit next was made because of what the analysis showed, and could not have been specified in the original protocol. Recording that decision and its reasoning is what evidences theoretical sampling to an examiner.

Pitfalls

Six mistakes examiners look for

1. Collecting all the data before coding any of it

This removes theoretical sampling entirely, and with it the defining procedure of the method.

2. Reporting categories with no core category

A list of parallel categories is a thematic analysis. Grounded theory requires integration around a core.

3. Mixing schools

Name your tradition and apply it consistently. Combining Glaser and Corbin invites exactly the question you cannot answer.

4. Claiming saturation without evidence

Say which categories saturated and what the last interviews added.

5. Codes that are topics rather than processes

Topic codes produce a description. Process codes produce a theory.

6. No memos

Without them there is no record of how the theory was built, and no defence of it.

Reporting

Writing it up

Name the school and cite its methodological source
Describe the iterative process, including how many rounds of collection and analysis
Explain theoretical sampling with concrete examples of decisions taken
Evidence saturation rather than asserting it
Present the core category first, then the categories related to it
Use extracts to evidence each category, with participant identifiers
Include a diagram of the theory — readers expect one
State your own position, particularly in a constructivist study

Structure the findings around the theory rather than around the interview schedule. A results chapter organised by question order reveals that the analysis stayed descriptive; one organised around the core category and its relationships demonstrates that a theory was actually built.

A sentence that earns marks

“Following Charmaz's constructivist approach, data collection and analysis were interleaved across three rounds. After the first six interviews, initial coding identified ‘managing visible uncertainty’ as a developing category; participants for round two were then selected specifically for contrasting levels of experience to test its boundaries. Saturation of this category was reached by interview 19, with the final four interviews contributing no new properties.”

Answers

Frequently asked questions

What is grounded theory in simple terms?

A qualitative method that builds a theory out of data rather than testing an existing one. You collect and analyse together, let categories emerge from close coding, choose later participants on the basis of what the analysis needs, and stop when new data add nothing further.

What is the difference between grounded theory and thematic analysis?

Two procedural features. Grounded theory interleaves data collection with analysis and uses theoretical sampling, so what you collect next depends on the emerging analysis. It also aims at an explanatory theory with a core category, whereas thematic analysis produces themes describing the data.

Which version of grounded theory should I use?

Straussian is the most prescriptive and therefore the easiest to defend procedurally. Constructivist, following Charmaz, is the most widely used in contemporary social science and is the honest choice for a researcher with existing knowledge of the field. Name whichever you use and apply it consistently.

What is theoretical sampling?

Selecting later participants or data sources on the basis of what the developing analysis requires, rather than deciding the whole sample in advance. If a category needs testing against a contrasting case, you recruit that case next. It is one of the two defining procedures of the method.

How many participants do you need for grounded theory?

There is no fixed number; published studies commonly fall between 20 and 40. What matters is theoretical saturation — the point at which further data add nothing to the properties of your categories. Justify stopping by describing what the final interviews contributed, not by citing a number.

What is the core category?

The category that accounts for the most variation in the data and to which the other categories can be related. Identifying it is the goal of selective coding, and integrating the analysis around it is what turns a set of categories into a theory.

Do I need a literature review before starting?

It depends on the school. Classic grounded theory delays it to avoid imposing existing concepts; Straussian and constructivist approaches accept an early review provided you remain reflexive about its influence. Most degree programmes require a review up front, so state how you managed its influence on your coding.

What are memos and why do they matter?

Analytic notes written throughout the study recording what codes mean, how categories might relate, and why sampling decisions were taken. They are where the theory is actually developed, and a dated sequence of them is the only convincing answer to a viva question about how you moved from codes to theory.

Send the data. Get a fixed quote.

Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.