Case studies / Third sector
Exploratory Factor Analysis and Ordinal Regression on a National Survey
Reducing 75 overlapping questionnaire items to interpretable factors, then modelling an ordinal outcome correctly.
The challenge
What the organisation came with
A national organisation had run a large survey and could not identify which factors influenced outcomes. They had the data; what they lacked was a way through it.
Two problems sat underneath that. First, many of the 75 questionnaire items measured overlapping concepts. Items written to capture the same underlying idea from slightly different angles are highly correlated with each other, and entering them together into a regression produces severe multicollinearity — unstable coefficients that swing wildly depending on which items are included.
Second, the primary outcome was ordinal. Responses had a meaningful order but no guarantee that the distance between adjacent categories was equal. Treating that as a continuous variable assumes an interval scale the data does not have.
The approach
How it was analysed, and why that method
Reducing the item set before modelling
Exploratory factor analysis was run first, to identify the smaller number of underlying constructs the 75 items were collectively measuring. This solves the multicollinearity at source: instead of entering dozens of correlated items, the model uses a handful of interpretable factors, each representing a coherent concept the survey was actually measuring.
Factor solutions were assessed for interpretability as well as fit. A statistically defensible solution that cannot be given a meaningful name is of no use to an organisation that has to act on it.
Modelling an ordinal outcome as ordinal
Ordinal logistic regression was used for the outcome model. Linear regression was rejected because the dependent variable was ordinal — using it would have assumed equal spacing between response categories and produced predictions outside the range of the scale.
Treating an ordinal outcome as continuous because it is coded 1 to 5. The coding is a convenience; it is not evidence that the gap between 1 and 2 equals the gap between 4 and 5.
Delivered
What the client received
The value
What changed as a result
The findings informed programme improvements and future strategic planning.
Because the analysis produced a small number of named factors rather than 75 coefficients, the results could be presented to a board and acted on — which is the difference between an analysis that is correct and one that is used.
Written from the assigned statistician's own project notes. The client is not named and no identifying detail, data or figures are published. Scale and timeframe are approximate.
Related services
The services behind this work
Statistical consultancy
Design, analysis and reporting for charities, universities, the NHS and business.
From £950Qualitative coding
Codebooks, systematic coding and themes evidenced back to the extracts.
Fixed quoteData cleaning & preparation
Messy data made analysis-ready, with a documented log of every change.
Fixed quoteSend the data. Get a fixed quote.
Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.