Guides  /  content analysis

Content analysis: a complete guide, with a worked example

Content analysis is the systematic categorisation of communication — text, transcripts, documents, media — in order to describe it. This guide covers the manifest and latent distinction, how to build and test a coding frame, a worked example with real counts, and the question that decides whether content analysis or thematic analysis is the right method for your study.

Elaine Halliburton Written and reviewed by Elaine Halliburton, Professor of Applied Statistics
Updated 14 August 202616 min read
What is content analysis?

Content analysis is a research method for systematically categorising the content of communication — such as interview transcripts, documents, or media — and describing what is there. It can be quantitative, counting how often categories occur, or qualitative, describing categories in context. Unlike thematic analysis it treats frequency as meaningful evidence.

Definition

What content analysis is

Content analysis is a method for systematically categorising communication and describing what it contains. The material can be interview transcripts, open-text survey responses, policy documents, newspaper coverage, social media posts, clinical notes or broadcast media — anything recorded that can be read.

Its defining feature is systematic categorisation. Every unit of the material is examined against the same set of categories, applied by the same rules, so the description that results is a property of the material rather than of whoever happened to read it. That is why content analysis is often the method of choice where a finding has to be reproducible.

It has a longer history than most qualitative methods, originating in communication research in the first half of the twentieth century, and it retains the quantitative instincts of that origin: content analysis is comfortable counting, in a way that most interpretative methods are not.

It is systematic — the same rules applied to every unit of data
It is replicable — another researcher applying the frame should reach similar results
It can be quantitative or qualitative, or both in the same study
It works on found data as readily as on data you collected
It scales to material far larger than an interpretative method could handle

The core distinction

Manifest and latent content

Every content analysis sits somewhere on a spectrum between describing what is literally there and interpreting what it means.

thematic analysis"> MANIFEST · what was said LATENT · what it means Counting Quantitative CA Qualitative CA Thematic analysis Frequencies ofpredefined words Categories counted,coding frame fixed Categories described,some interpretation Interpretativethemes, no counting The question is not which is better. It is whether your research question is answered by how often something appears, or by what it means when it does.
Content analysis is not one method but a range. Where you sit on it should follow from your research question, and you should say where that is.

Manifest content

The surface of the material: the words actually used, the topics explicitly raised, the categories a reader could identify without inference. Coding manifest content is closer to measurement — two competent coders should agree most of the time, and where they do not, the coding frame is usually at fault.

Example: coding whether each policy document mentions equality impact assessment. Either the phrase and its concept appear, or they do not.

Latent content

The underlying meaning: what is implied, what tone is being struck, what is conspicuously absent. Coding latent content requires interpretation, which means agreement between coders is harder to achieve and reliability statistics need to be read more carefully.

Example: coding whether each policy document treats equality as a compliance obligation or as a purpose. No single phrase settles it; the coder is making a judgement about the document as a whole.

Say which you did

A study that codes latent content but reports reliability as though it were manifest is making a claim it cannot support. Stating the level you worked at, in one sentence, pre-empts the most common criticism of published content analyses.

Choosing

Content analysis or thematic analysis?

This is the most common question we are asked about qualitative method selection, and the honest answer is short: does frequency answer your question?

Content analysisThematic analysis
OutputCategories, often with countsInterpretative themes
FrequencyMeaningful evidenceNot the measure of importance
Coding frameUsually fixed before main codingDevelops through the analysis
Multiple codersCommon; reliability often reportedOptional, and not for agreement in reflexive TA
Researcher roleMinimised, treated as a source of errorA resource, made explicit
Dataset sizeScales to hundreds of documentsPractically limited by close reading
Best whenYou need to describe what is there, comparablyYou need to interpret what it means

A concrete test. If your finding would be reported as “equality was mentioned in 34 of 51 documents, most often as a compliance requirement”, you are doing content analysis. If it would be reported as “equality functioned as a language of reassurance rather than a commitment”, you are doing thematic analysis. Both are defensible; only one answers each question.

Not sure which your question needs?

Choosing the wrong qualitative method is expensive, because you usually discover it at the analysis stage rather than the design stage. A statistician who works in qualitative methods can look at your research question and data and tell you which method fits before you commit.

See qualitative coding support

Types

Three approaches to qualitative content analysis

Hsieh and Shannon’s 2005 typology is the one most UK examiners expect to see cited, and naming which of the three you used is now close to mandatory in health and social research.

ApproachWhere categories come fromUse whenReported as
ConventionalDerived from the data itselfExisting theory on the topic is limited or fragmentedCategories with descriptions and illustrative extracts
DirectedDerived from existing theory or prior researchA framework exists and you are extending or testing itCategories, plus data that did not fit the framework
SummativeSpecific words or content counted, then interpretedInterest is in usage of particular terms and what that usage impliesCounts, followed by latent interpretation

The distinction matters most for directed content analysis, because its characteristic weakness is confirmation. If your categories come from a framework, the analysis will find the framework. The safeguard is to report explicitly what did not fit — material that falls outside an a priori frame is frequently the most interesting result in the study, and the studies that discard it silently are the ones reviewers criticise.

Summative content analysis is not word counting

It begins with counting particular terms, then asks what the pattern of use means — which is latent interpretation. A study that stops at the frequency table has done half the method and should not claim the label.

Design

Sampling the material

Content analysis is often applied to material that already exists, which makes sampling a design decision rather than a recruitment problem. It is also the decision most often left undescribed.

Define the population of material — all documents of a type, in a date range, from a defined set of sources
State the inclusion and exclusion criteria, in the same way a systematic review would
Say how you searched, including the databases, archives or sites used and the terms
Record what you excluded and why, with numbers
Justify the time window — policy language shifts, and a five-year span may straddle a change that explains your findings
Note what you could not obtain, because unavailable material is rarely missing at random

Where the corpus is large, a random or systematic sample is defensible and should be described as such. Where it is small enough to analyse whole, say that you analysed the population rather than a sample — it is a stronger claim and it costs nothing to state.

The most common sampling weakness in published content analyses is convenience framed as completeness: analysing the documents that happened to be online, and reporting percentages as though they described all documents of that type. If your corpus is what you could find, say so.

Method

Building a coding frame

The coding frame is the instrument. In content analysis it does the work that a questionnaire does in survey research, and it deserves the same care.

Where categories come from

Deductive frames are built from theory, prior research or a policy framework before the data is read. Inductive frames are built from a subset of the data and then applied to the rest. Most real studies are hybrid: a prior frame, revised after a pilot on 10–20% of the material.

What each category needs

A name that a second coder would interpret the same way
A definition saying what the category captures
An inclusion rule — what qualifies
An exclusion rule — what looks similar but does not qualify
At least one anchor example taken from the actual material
A note on boundary cases and how they were resolved

Categories must also be exhaustive and, in most designs, mutually exclusive. If a unit can fall into two categories at once, either the categories overlap and should be merged, or the design should allow multiple codes per unit — which is fine, but it changes how you report frequencies and must be stated.

The unit of analysis

Decide before coding whether your unit is the word, the sentence, the paragraph, the speaking turn, or the whole document. This decision determines what your counts mean, and changing it midway invalidates everything coded before the change.

Pilot the frame before you commit

Apply the draft frame to roughly 10% of the material with a second coder. Almost every frame needs revision at this point, and revising after full coding means recoding everything.

Worked example

A worked example, with counts

A study examined how 51 local authority strategy documents discussed digital exclusion. The unit of analysis was the paragraph; a paragraph could receive more than one code.

CategoryDefinitionDocumentsParagraphs
Access framingDigital exclusion presented as lack of devices or connectivity44 (86%)312
Skills framingPresented as lack of confidence, literacy or capability31 (61%)148
Design framingPresented as services being hard to use9 (18%)37
Choice framingPresented as people choosing not to engage7 (14%)19
Named responsibilityA specific body given responsibility for addressing it12 (24%)28
Measurable targetA stated, quantified objective5 (10%)9

What the counts show

Access framing dominates: it appears in 86% of documents and accounts for more than half of all coded paragraphs. Design framing — the possibility that services themselves create exclusion — appears in fewer than one document in five. Only 10% of strategies contain a measurable target.

That pattern is the finding, and it is a finding frequency can carry. The contrast between how often the problem is described as the citizen’s (access, skills, choice: 479 paragraphs) and how often it is described as the service’s (design: 37 paragraphs) is a quantitative claim about a body of policy text, and it would be considerably weaker stated as an impression.

What the counts do not show

They do not show that access framing is wrong, or that authorities using design framing perform better. Content analysis describes the material; it does not evaluate it. Studies that slide from describing what documents say to judging what authorities do are making a claim their design does not support.

Reliability

Inter-coder reliability

Because content analysis claims replicability, it usually has to demonstrate it. A subset of the material — commonly 10–20% — is coded independently by a second coder, and agreement is calculated.

StatisticUse whenInterpretation
Cohen’s kappaTwo coders, nominal categories<.40 poor · .41–.60 moderate · .61–.80 substantial · >.80 almost perfect
Fleiss’ kappaThree or more codersSame conventional bands
Krippendorff’s alphaAny number of coders; handles missing data and ordinal categories≥.80 accepted · .67–.79 tentative conclusions only
Percentage agreementReporting alongside a chance-corrected figureOverstates reliability; never report alone

Percentage agreement is the figure most often reported and the least informative. Two coders assigning a category that occurs 90% of the time will agree 90% of the time by chance alone. Chance-corrected statistics exist precisely for this, and reviewers know it.

Low agreement is a frame problem, not a coder problem

The instinct is to retrain the second coder. Usually the category definitions are ambiguous, the boundary rules are missing, or the unit of analysis is too large. Fix the frame, then recode.

Common problems

Six mistakes that undermine a content analysis

MistakeConsequenceInstead
Coding frame built after coding beginsEarly and late material coded to different rulesPilot, revise, fix, then code
Unit of analysis never statedCounts cannot be interpreted or comparedState it, and keep it constant
Percentage agreement reported aloneOverstates reliability, often substantiallyReport kappa or Krippendorff’s alpha
Counting without contextFrequency presented as importanceReport counts and illustrate with material
Categories that overlapUnits forced into arbitrary binsMerge, or allow multiple codes and say so
Describing, then evaluatingA claim the design cannot supportKeep description and inference separate

Reporting

How to write it up

Methods

The material, how it was selected, and how much of it there is
Whether coding was manifest, latent, or both
The unit of analysis
How the coding frame was developed, and whether deductively or inductively
Whether categories were mutually exclusive
How many coders, how much was double-coded, and which reliability statistic
The software used, if any

Results

Report counts in a table rather than in prose — six categories described in sentences is unreadable, while the same six in a table is scannable. Give both the number of documents and the number of coded units where they differ, because a category appearing many times in one document tells a different story from one appearing once in many.

Illustrate categories with short extracts even in a quantitative content analysis. Counts establish the pattern; extracts establish what the category actually looks like in the material, and without them the reader has to take the frame on trust.

Need the coding done, or checked?

We build coding frames, code the material, calculate reliability and hand back the coded dataset — or review a frame you have already built and tell you where it will not hold.

See qualitative coding support

Answers

Frequently asked questions

Is content analysis qualitative or quantitative?

It can be either, and often both. Quantitative content analysis counts occurrences of predefined categories; qualitative content analysis describes categories in context with less emphasis on frequency. Most applied studies sit between the two and should say where.

What is the difference between content analysis and thematic analysis?

Content analysis systematically categorises material and treats frequency as evidence. Thematic analysis constructs interpretative themes and does not use frequency as the measure of importance. If your finding would be reported as a count, you are doing content analysis.

How much material do I need?

There is no minimum. Content analysis scales well, so the practical question is whether you have enough material for the categories to be populated. Six documents will not support percentage claims; sixty will.

Do I need two coders?

If you are claiming replicability — which most content analyses do — yes, at least for a subset. Single-coder content analysis is possible but should be acknowledged as a limitation, since the method's central claim is that another researcher would reach similar results.

What software should I use?

NVivo, MAXQDA and ATLAS.ti all support content analysis and calculate coding comparison. For purely quantitative frames on large corpora, R packages such as quanteda are faster. Under about twenty documents, a spreadsheet is adequate.

What is a good kappa value?

Above .80 is conventionally almost perfect and .61 to .80 substantial. Krippendorff's alpha of .80 or above is the usual threshold, with .67 to .79 supporting tentative conclusions only. These are conventions rather than rules, and latent coding reasonably produces lower values than manifest coding.

Can I use content analysis on interview data?

Yes, though consider whether it answers your question. Interviews are usually collected to explore meaning, which is thematic analysis territory. Content analysis suits interview data where you need comparable description across many participants rather than depth.

Can content analysis be used with a small sample?

Qualitative content analysis can. Quantitative content analysis with percentage claims cannot — reporting that 67% of six documents did something is arithmetically true and rhetorically misleading.

What is the difference between conventional, directed and summative content analysis?

Conventional derives categories from the data, directed derives them from existing theory, and summative counts specific content and then interprets what the pattern of usage means. Name which you used; examiners in health and social research increasingly expect the Hsieh and Shannon citation.

How do I choose a unit of analysis?

Pick the smallest unit that carries the meaning you are coding. Words suit summative analysis of terminology; sentences and paragraphs suit most category coding; whole documents suit presence-or-absence questions. Whatever you choose, keep it constant — changing it midway invalidates everything coded before.

Should I report percentages or raw counts?

Both. Raw counts let the reader judge the base; percentages let them compare. Percentages alone are misleading with small corpora, which is why reporting that a category appeared in 67% of a nine-document sample invites the obvious question.

Is content analysis suitable for a dissertation?

Yes, and it is often a good fit. It scales to material you can obtain without recruiting participants, the method is explicit enough to describe fully in a methods chapter, and reliability can be demonstrated — which is easier to defend at a viva than an interpretative method you are learning for the first time.

Send the data. Get a fixed quote.

Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.