Guides / grounded theory
Grounded theory: a practical guide
Grounded theory builds a theory from data rather than testing one against it, and it is the qualitative method most often claimed and least often actually carried out. This guide sets out the coding stages, explains the two features that distinguish real grounded theory from thematic analysis with different labels, and covers what an examiner will look for.
Grounded theory is a qualitative research method in which theory is developed inductively from data rather than tested against a pre-existing framework. Data collection and analysis proceed together, with each round of analysis determining what is collected next, and coding moves from many descriptive codes to a small number of categories and finally to a core category that explains the central process. It ends when further data produce no new insight.
Definition
What grounded theory is
Grounded theory develops a theory from the data rather than testing a theory against them. You begin without a hypothesis, code closely, and build upwards until you have an explanatory account of the process you were studying.
Its output is distinctive. Thematic analysis produces themes describing what participants said; grounded theory produces a theory — an account of how something works, what conditions bring it about, and what consequences follow. That difference in output is what should drive the choice of method.
The method originated in the 1960s as a deliberate corrective. Sociology at the time was dominated by grand theorising tested through survey work, and Glaser and Strauss argued that this produced theory disconnected from the settings it claimed to describe. Their alternative was to derive theory from close, systematic engagement with data — hence “grounded”. Understanding that origin helps explain features that otherwise look arbitrary, particularly the suspicion of an early literature review, which was aimed at preventing existing frameworks from being imposed on data that might not fit them.
Grounded theory suits questions about process: how something unfolds, how people manage a situation over time, how a practice is sustained. If your question is “what do participants think about X?”, thematic analysis is the honest answer and will be easier to carry out well.
Variants
The three schools, and choosing one
Grounded theory split into distinct traditions after its founders disagreed, and examiners expect you to know which one you followed.
| School | Key figures | Stance | Coding |
|---|---|---|---|
| Classic | Glaser | Theory emerges; the literature is read late | Open, selective, theoretical |
| Straussian | Strauss & Corbin | More structured; a coding paradigm guides analysis | Open, axial, selective |
| Constructivist | Charmaz | Theory is co-constructed; the researcher is present in it | Initial, focused, theoretical |
The Straussian version is the most prescriptive and therefore the easiest to defend in a viva, because each step has a stated procedure. The constructivist version is the most widely used in contemporary social science and is the most comfortable fit for a researcher who cannot plausibly claim to have approached the topic with no prior knowledge.
Mixing Glaser's rejection of a prior literature review with Corbin's coding paradigm produces a method that belongs to nobody, and it is the kind of inconsistency an examiner will press on. State which tradition you followed, cite it, and apply it consistently.
The distinction
What makes it grounded theory and not thematic analysis
Two features distinguish genuine grounded theory, and a study lacking either is thematic analysis under a different name.
| Feature | Grounded theory | Thematic analysis |
|---|---|---|
| Data collection | Iterative — interleaved with analysis | All collected, then analysed |
| Sampling | Theoretical — driven by the emerging theory | Purposive, decided in advance |
| Stopping rule | Theoretical saturation | Planned sample size |
| Output | An explanatory theory | Descriptive or interpretive themes |
| Literature review | Late, or used cautiously | Normally up front |
Interviewing twenty people, then coding all twenty transcripts, then reporting categories — and calling it grounded theory. Without iterative collection and theoretical sampling, the two defining procedures are absent. Examiners check for this specifically, and the safer route is to describe accurately what you did rather than to claim a method you did not follow.
Send your research question and the data you have or plan to collect. A named researcher advises which method fits and what it will require of you.
Get a fixed quoteCoding
The three stages of coding
Open coding
Work line by line or incident by incident, naming what is happening in the data. Codes stay close to the material and are often gerunds — “managing uncertainty”, “deferring to expertise” — because grounded theory is interested in action and process rather than in topics. Expect a large number of codes at this stage; several hundred from a handful of transcripts is normal.
Axial coding
Group codes into categories and, crucially, work out how the categories relate. The Straussian coding paradigm offers a scaffold for this: for each category, identify the conditions that give rise to it, the actions and interactions involved, and the consequences that follow. Not every analysis needs the paradigm applied mechanically, but the relational question — how does this connect to that? — is what separates axial coding from sorting.
Selective coding
Identify the core category: the one that accounts for the most variation and to which the others can be related. Everything else is then integrated around it, and the resulting account is the theory. If no core category emerges, the analysis is not finished, and reporting a list of parallel categories instead is the point at which many grounded theory studies quietly become thematic analyses.
Software helps with the mechanics but does not do the analysis. NVivo, MAXQDA and Dedoose all manage codes, retrieve extracts and store memos efficiently, which matters once you have several hundred codes across twenty transcripts. None of them identifies a core category for you, and a coding tree that looks tidy in software can still be a description rather than a theory.
“Support” is a topic. “Seeking support without appearing to need it” is a process, and it carries the tension that makes a theory worth building. Naming codes as gerunds is a practical discipline that keeps the analysis oriented towards process.
Procedure
Theoretical sampling and constant comparison
These are the two engines of the method, and both are procedural rather than conceptual — you either did them or you did not.
Theoretical sampling
After the first few interviews you analyse, notice a gap or an ambiguity in the emerging categories, and choose the next participants specifically to address it. If a category about “managing uncertainty” is developing but every participant so far has been experienced, you deliberately recruit novices next to see whether the category holds.
This means your sampling strategy cannot be fully specified in advance, which has implications for ethics applications. Write the protocol to allow it: state that sampling will be theoretical, that later participants will be selected on the basis of emerging analysis, and give the broad pool you will draw from.
In practice, theoretical sampling rarely means recruiting entirely new participants each round. It may mean returning to earlier participants with sharper questions, seeking documents that bear on a developing category, or observing a setting you had not planned to visit. What defines it is not the source but the reasoning: the choice follows from what the analysis needs, and that reasoning is recorded.
Constant comparison
Every new piece of data is compared against the codes and categories already developed, and every code against the others. The comparisons are what force categories to become precise: encountering a case that does not fit obliges you either to redefine the category or to identify the condition under which it does not apply.
Record why each sampling decision was taken and what the analysis at that point required. This trail is the main evidence that theoretical sampling actually occurred, and reconstructing it afterwards is both difficult and obvious to a reader.
Saturation
Saturation, and how to justify stopping
Theoretical saturation is reached when further data no longer add to the properties of your categories. It is a property of the categories, not of the number of interviews, and the distinction matters when you come to defend it.
Sample sizes in published grounded theory studies commonly fall between 20 and 40 participants, but citing a number is not a justification. What defends the decision is describing the point at which categories stopped developing and showing the evidence — typically that the last several interviews contributed nothing new to any core category.
“Saturation was reached after 12 interviews” with no supporting detail is a claim examiners routinely challenge. Say which categories saturated, at approximately what point, and what the final interviews added. If you stopped for practical reasons — time, access, funding — say so and treat it as a limitation. That is far stronger than an unevidenced claim.
Memos
Memo writing
Memos are analytic notes written throughout, and in grounded theory they are not optional. They are where the theory is actually developed; the codes are only its raw material.
| Memo type | Contents |
|---|---|
| Code memo | What a code means, its boundaries, an example extract |
| Theoretical memo | Hunches about relationships between categories |
| Operational memo | Sampling and procedural decisions, with reasons |
| Integrative memo | How categories fit together as the core emerges |
Write them immediately, date them, and do not edit them for style. Their value lies in capturing thinking as it develops, including the ideas later abandoned — which is precisely what lets you reconstruct and defend the analytic path in a viva.
There is no minimum, but a study producing a handful of memos across several months has probably not been doing the analytic work the method requires. Many researchers find they write more memo text than they eventually write findings text, and that is a reasonable sign the analysis is developing rather than merely accumulating.
This is the question grounded theory candidates find hardest, because the step is genuinely interpretive. A dated sequence of memos showing the reasoning is the only convincing answer, and it cannot be produced retrospectively.
Send your transcripts and coding so far. A named researcher reviews the codebook, tests the categories against the extracts, and advises on developing the core category.
See qualitative codingWorked example
Worked example: from extract to category
A study explores how newly qualified practitioners manage situations beyond their competence. Three extracts from different participants, and the codes assigned during initial coding.
| Extract | Initial codes |
|---|---|
| “I'd ask, but I'd phrase it like I was checking rather than like I didn't know.” | disguising the request; protecting credibility |
| “You learn quite fast who you can look stupid in front of and who you can't.” | mapping safe others; assessing risk of exposure |
| “Sometimes I'd just look it up afterwards rather than ask at the time.” | deferring the question; avoiding visible uncertainty |
Moving to a category
Every code involves a tension between needing information and managing how the need appears to others. Grouping them produces a candidate category: managing visible uncertainty. Note that it is stated as a process, not a topic — “uncertainty” alone would describe a subject area, while “managing visible uncertainty” names something participants are doing.
Developing its properties
| Property | Dimensions observed |
|---|---|
| Who the audience is | Senior colleague ↔ peer ↔ patient or client |
| Cost of exposure | Trivial ↔ career-relevant |
| Strategy used | Disguising ↔ deferring ↔ asking openly |
| Timing | In the moment ↔ afterwards |
What theoretical sampling does next
Every participant so far has been newly qualified. To test whether the category is specific to inexperience or persists, the next round deliberately recruits practitioners with ten or more years in post. If they describe the same management of visible uncertainty, the category is broader than first assumed and the theory must account for what sustains it. If they do not, the conditions under which it operates become part of the theory.
The decision about who to recruit next was made because of what the analysis showed, and could not have been specified in the original protocol. Recording that decision and its reasoning is what evidences theoretical sampling to an examiner.
Pitfalls
Six mistakes examiners look for
1. Collecting all the data before coding any of it
This removes theoretical sampling entirely, and with it the defining procedure of the method.
2. Reporting categories with no core category
A list of parallel categories is a thematic analysis. Grounded theory requires integration around a core.
3. Mixing schools
Name your tradition and apply it consistently. Combining Glaser and Corbin invites exactly the question you cannot answer.
4. Claiming saturation without evidence
Say which categories saturated and what the last interviews added.
5. Codes that are topics rather than processes
Topic codes produce a description. Process codes produce a theory.
6. No memos
Without them there is no record of how the theory was built, and no defence of it.
Reporting
Writing it up
Structure the findings around the theory rather than around the interview schedule. A results chapter organised by question order reveals that the analysis stayed descriptive; one organised around the core category and its relationships demonstrates that a theory was actually built.
“Following Charmaz's constructivist approach, data collection and analysis were interleaved across three rounds. After the first six interviews, initial coding identified ‘managing visible uncertainty’ as a developing category; participants for round two were then selected specifically for contrasting levels of experience to test its boundaries. Saturation of this category was reached by interview 19, with the final four interviews contributing no new properties.”
Answers
Frequently asked questions
What is grounded theory in simple terms?
A qualitative method that builds a theory out of data rather than testing an existing one. You collect and analyse together, let categories emerge from close coding, choose later participants on the basis of what the analysis needs, and stop when new data add nothing further.
What is the difference between grounded theory and thematic analysis?
Two procedural features. Grounded theory interleaves data collection with analysis and uses theoretical sampling, so what you collect next depends on the emerging analysis. It also aims at an explanatory theory with a core category, whereas thematic analysis produces themes describing the data.
Which version of grounded theory should I use?
Straussian is the most prescriptive and therefore the easiest to defend procedurally. Constructivist, following Charmaz, is the most widely used in contemporary social science and is the honest choice for a researcher with existing knowledge of the field. Name whichever you use and apply it consistently.
What is theoretical sampling?
Selecting later participants or data sources on the basis of what the developing analysis requires, rather than deciding the whole sample in advance. If a category needs testing against a contrasting case, you recruit that case next. It is one of the two defining procedures of the method.
How many participants do you need for grounded theory?
There is no fixed number; published studies commonly fall between 20 and 40. What matters is theoretical saturation — the point at which further data add nothing to the properties of your categories. Justify stopping by describing what the final interviews contributed, not by citing a number.
What is the core category?
The category that accounts for the most variation in the data and to which the other categories can be related. Identifying it is the goal of selective coding, and integrating the analysis around it is what turns a set of categories into a theory.
Do I need a literature review before starting?
It depends on the school. Classic grounded theory delays it to avoid imposing existing concepts; Straussian and constructivist approaches accept an early review provided you remain reflexive about its influence. Most degree programmes require a review up front, so state how you managed its influence on your coding.
What are memos and why do they matter?
Analytic notes written throughout the study recording what codes mean, how categories might relate, and why sampling decisions were taken. They are where the theory is actually developed, and a dated sequence of them is the only convincing answer to a viva question about how you moved from codes to theory.
Keep reading
Related guides and services
Thematic analysis
The method grounded theory is most often confused with, and when it is the honest choice.
GuideReliability and validity
The trustworthiness criteria that apply to qualitative work.
ServiceQualitative coding
Codebooks, systematic coding and categories evidenced back to the extracts.
Send the data. Get a fixed quote.
Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.