Guides / content analysis
Content analysis: a complete guide, with a worked example
Content analysis is the systematic categorisation of communication — text, transcripts, documents, media — in order to describe it. This guide covers the manifest and latent distinction, how to build and test a coding frame, a worked example with real counts, and the question that decides whether content analysis or thematic analysis is the right method for your study.
Content analysis is a research method for systematically categorising the content of communication — such as interview transcripts, documents, or media — and describing what is there. It can be quantitative, counting how often categories occur, or qualitative, describing categories in context. Unlike thematic analysis it treats frequency as meaningful evidence.
Definition
What content analysis is
Content analysis is a method for systematically categorising communication and describing what it contains. The material can be interview transcripts, open-text survey responses, policy documents, newspaper coverage, social media posts, clinical notes or broadcast media — anything recorded that can be read.
Its defining feature is systematic categorisation. Every unit of the material is examined against the same set of categories, applied by the same rules, so the description that results is a property of the material rather than of whoever happened to read it. That is why content analysis is often the method of choice where a finding has to be reproducible.
It has a longer history than most qualitative methods, originating in communication research in the first half of the twentieth century, and it retains the quantitative instincts of that origin: content analysis is comfortable counting, in a way that most interpretative methods are not.
The core distinction
Manifest and latent content
Every content analysis sits somewhere on a spectrum between describing what is literally there and interpreting what it means.
Manifest content
The surface of the material: the words actually used, the topics explicitly raised, the categories a reader could identify without inference. Coding manifest content is closer to measurement — two competent coders should agree most of the time, and where they do not, the coding frame is usually at fault.
Example: coding whether each policy document mentions equality impact assessment. Either the phrase and its concept appear, or they do not.
Latent content
The underlying meaning: what is implied, what tone is being struck, what is conspicuously absent. Coding latent content requires interpretation, which means agreement between coders is harder to achieve and reliability statistics need to be read more carefully.
Example: coding whether each policy document treats equality as a compliance obligation or as a purpose. No single phrase settles it; the coder is making a judgement about the document as a whole.
A study that codes latent content but reports reliability as though it were manifest is making a claim it cannot support. Stating the level you worked at, in one sentence, pre-empts the most common criticism of published content analyses.
Choosing
Content analysis or thematic analysis?
This is the most common question we are asked about qualitative method selection, and the honest answer is short: does frequency answer your question?
| Content analysis | Thematic analysis | |
|---|---|---|
| Output | Categories, often with counts | Interpretative themes |
| Frequency | Meaningful evidence | Not the measure of importance |
| Coding frame | Usually fixed before main coding | Develops through the analysis |
| Multiple coders | Common; reliability often reported | Optional, and not for agreement in reflexive TA |
| Researcher role | Minimised, treated as a source of error | A resource, made explicit |
| Dataset size | Scales to hundreds of documents | Practically limited by close reading |
| Best when | You need to describe what is there, comparably | You need to interpret what it means |
A concrete test. If your finding would be reported as “equality was mentioned in 34 of 51 documents, most often as a compliance requirement”, you are doing content analysis. If it would be reported as “equality functioned as a language of reassurance rather than a commitment”, you are doing thematic analysis. Both are defensible; only one answers each question.
Choosing the wrong qualitative method is expensive, because you usually discover it at the analysis stage rather than the design stage. A statistician who works in qualitative methods can look at your research question and data and tell you which method fits before you commit.
See qualitative coding supportTypes
Three approaches to qualitative content analysis
Hsieh and Shannon’s 2005 typology is the one most UK examiners expect to see cited, and naming which of the three you used is now close to mandatory in health and social research.
| Approach | Where categories come from | Use when | Reported as |
|---|---|---|---|
| Conventional | Derived from the data itself | Existing theory on the topic is limited or fragmented | Categories with descriptions and illustrative extracts |
| Directed | Derived from existing theory or prior research | A framework exists and you are extending or testing it | Categories, plus data that did not fit the framework |
| Summative | Specific words or content counted, then interpreted | Interest is in usage of particular terms and what that usage implies | Counts, followed by latent interpretation |
The distinction matters most for directed content analysis, because its characteristic weakness is confirmation. If your categories come from a framework, the analysis will find the framework. The safeguard is to report explicitly what did not fit — material that falls outside an a priori frame is frequently the most interesting result in the study, and the studies that discard it silently are the ones reviewers criticise.
It begins with counting particular terms, then asks what the pattern of use means — which is latent interpretation. A study that stops at the frequency table has done half the method and should not claim the label.
Design
Sampling the material
Content analysis is often applied to material that already exists, which makes sampling a design decision rather than a recruitment problem. It is also the decision most often left undescribed.
Where the corpus is large, a random or systematic sample is defensible and should be described as such. Where it is small enough to analyse whole, say that you analysed the population rather than a sample — it is a stronger claim and it costs nothing to state.
The most common sampling weakness in published content analyses is convenience framed as completeness: analysing the documents that happened to be online, and reporting percentages as though they described all documents of that type. If your corpus is what you could find, say so.
Method
Building a coding frame
The coding frame is the instrument. In content analysis it does the work that a questionnaire does in survey research, and it deserves the same care.
Where categories come from
Deductive frames are built from theory, prior research or a policy framework before the data is read. Inductive frames are built from a subset of the data and then applied to the rest. Most real studies are hybrid: a prior frame, revised after a pilot on 10–20% of the material.
What each category needs
Categories must also be exhaustive and, in most designs, mutually exclusive. If a unit can fall into two categories at once, either the categories overlap and should be merged, or the design should allow multiple codes per unit — which is fine, but it changes how you report frequencies and must be stated.
The unit of analysis
Decide before coding whether your unit is the word, the sentence, the paragraph, the speaking turn, or the whole document. This decision determines what your counts mean, and changing it midway invalidates everything coded before the change.
Apply the draft frame to roughly 10% of the material with a second coder. Almost every frame needs revision at this point, and revising after full coding means recoding everything.
Worked example
A worked example, with counts
A study examined how 51 local authority strategy documents discussed digital exclusion. The unit of analysis was the paragraph; a paragraph could receive more than one code.
| Category | Definition | Documents | Paragraphs |
|---|---|---|---|
| Access framing | Digital exclusion presented as lack of devices or connectivity | 44 (86%) | 312 |
| Skills framing | Presented as lack of confidence, literacy or capability | 31 (61%) | 148 |
| Design framing | Presented as services being hard to use | 9 (18%) | 37 |
| Choice framing | Presented as people choosing not to engage | 7 (14%) | 19 |
| Named responsibility | A specific body given responsibility for addressing it | 12 (24%) | 28 |
| Measurable target | A stated, quantified objective | 5 (10%) | 9 |
What the counts show
Access framing dominates: it appears in 86% of documents and accounts for more than half of all coded paragraphs. Design framing — the possibility that services themselves create exclusion — appears in fewer than one document in five. Only 10% of strategies contain a measurable target.
That pattern is the finding, and it is a finding frequency can carry. The contrast between how often the problem is described as the citizen’s (access, skills, choice: 479 paragraphs) and how often it is described as the service’s (design: 37 paragraphs) is a quantitative claim about a body of policy text, and it would be considerably weaker stated as an impression.
What the counts do not show
They do not show that access framing is wrong, or that authorities using design framing perform better. Content analysis describes the material; it does not evaluate it. Studies that slide from describing what documents say to judging what authorities do are making a claim their design does not support.
Reliability
Inter-coder reliability
Because content analysis claims replicability, it usually has to demonstrate it. A subset of the material — commonly 10–20% — is coded independently by a second coder, and agreement is calculated.
| Statistic | Use when | Interpretation |
|---|---|---|
| Cohen’s kappa | Two coders, nominal categories | <.40 poor · .41–.60 moderate · .61–.80 substantial · >.80 almost perfect |
| Fleiss’ kappa | Three or more coders | Same conventional bands |
| Krippendorff’s alpha | Any number of coders; handles missing data and ordinal categories | ≥.80 accepted · .67–.79 tentative conclusions only |
| Percentage agreement | Reporting alongside a chance-corrected figure | Overstates reliability; never report alone |
Percentage agreement is the figure most often reported and the least informative. Two coders assigning a category that occurs 90% of the time will agree 90% of the time by chance alone. Chance-corrected statistics exist precisely for this, and reviewers know it.
The instinct is to retrain the second coder. Usually the category definitions are ambiguous, the boundary rules are missing, or the unit of analysis is too large. Fix the frame, then recode.
Common problems
Six mistakes that undermine a content analysis
| Mistake | Consequence | Instead |
|---|---|---|
| Coding frame built after coding begins | Early and late material coded to different rules | Pilot, revise, fix, then code |
| Unit of analysis never stated | Counts cannot be interpreted or compared | State it, and keep it constant |
| Percentage agreement reported alone | Overstates reliability, often substantially | Report kappa or Krippendorff’s alpha |
| Counting without context | Frequency presented as importance | Report counts and illustrate with material |
| Categories that overlap | Units forced into arbitrary bins | Merge, or allow multiple codes and say so |
| Describing, then evaluating | A claim the design cannot support | Keep description and inference separate |
Reporting
How to write it up
Methods
Results
Report counts in a table rather than in prose — six categories described in sentences is unreadable, while the same six in a table is scannable. Give both the number of documents and the number of coded units where they differ, because a category appearing many times in one document tells a different story from one appearing once in many.
Illustrate categories with short extracts even in a quantitative content analysis. Counts establish the pattern; extracts establish what the category actually looks like in the material, and without them the reader has to take the frame on trust.
We build coding frames, code the material, calculate reliability and hand back the coded dataset — or review a frame you have already built and tell you where it will not hold.
See qualitative coding supportAnswers
Frequently asked questions
Is content analysis qualitative or quantitative?
It can be either, and often both. Quantitative content analysis counts occurrences of predefined categories; qualitative content analysis describes categories in context with less emphasis on frequency. Most applied studies sit between the two and should say where.
What is the difference between content analysis and thematic analysis?
Content analysis systematically categorises material and treats frequency as evidence. Thematic analysis constructs interpretative themes and does not use frequency as the measure of importance. If your finding would be reported as a count, you are doing content analysis.
How much material do I need?
There is no minimum. Content analysis scales well, so the practical question is whether you have enough material for the categories to be populated. Six documents will not support percentage claims; sixty will.
Do I need two coders?
If you are claiming replicability — which most content analyses do — yes, at least for a subset. Single-coder content analysis is possible but should be acknowledged as a limitation, since the method's central claim is that another researcher would reach similar results.
What software should I use?
NVivo, MAXQDA and ATLAS.ti all support content analysis and calculate coding comparison. For purely quantitative frames on large corpora, R packages such as quanteda are faster. Under about twenty documents, a spreadsheet is adequate.
What is a good kappa value?
Above .80 is conventionally almost perfect and .61 to .80 substantial. Krippendorff's alpha of .80 or above is the usual threshold, with .67 to .79 supporting tentative conclusions only. These are conventions rather than rules, and latent coding reasonably produces lower values than manifest coding.
Can I use content analysis on interview data?
Yes, though consider whether it answers your question. Interviews are usually collected to explore meaning, which is thematic analysis territory. Content analysis suits interview data where you need comparable description across many participants rather than depth.
Can content analysis be used with a small sample?
Qualitative content analysis can. Quantitative content analysis with percentage claims cannot — reporting that 67% of six documents did something is arithmetically true and rhetorically misleading.
What is the difference between conventional, directed and summative content analysis?
Conventional derives categories from the data, directed derives them from existing theory, and summative counts specific content and then interprets what the pattern of usage means. Name which you used; examiners in health and social research increasingly expect the Hsieh and Shannon citation.
How do I choose a unit of analysis?
Pick the smallest unit that carries the meaning you are coding. Words suit summative analysis of terminology; sentences and paragraphs suit most category coding; whole documents suit presence-or-absence questions. Whatever you choose, keep it constant — changing it midway invalidates everything coded before.
Should I report percentages or raw counts?
Both. Raw counts let the reader judge the base; percentages let them compare. Percentages alone are misleading with small corpora, which is why reporting that a category appeared in 67% of a nine-document sample invites the obvious question.
Is content analysis suitable for a dissertation?
Yes, and it is often a good fit. It scales to material you can obtain without recruiting participants, the method is explicit enough to describe fully in a methods chapter, and reliability can be demonstrated — which is easier to defend at a viva than an interpretative method you are learning for the first time.
Keep reading
Related guides and services
Thematic analysis
When your question is about meaning rather than frequency.
ServiceQualitative coding and thematic analysis
Coding frames built, applied and tested for reliability.
Case studyFramework analysis across stakeholder groups
Comparing perspectives without flattening what made each distinct.
Send the data. Get a fixed quote.
Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.