Case studies  /  Third sector

Exploratory Factor Analysis and Ordinal Regression on a National Survey

Reducing 75 overlapping questionnaire items to interpretable factors, then modelling an ordinal outcome correctly.

Third sectorSector
EFA and ordinal logistic regressionMethod
SPSS and RSoftware
~2,200 respondents, 75 variablesScale
About three weeksDuration

The challenge

What the organisation came with

A national organisation had run a large survey and could not identify which factors influenced outcomes. They had the data; what they lacked was a way through it.

Two problems sat underneath that. First, many of the 75 questionnaire items measured overlapping concepts. Items written to capture the same underlying idea from slightly different angles are highly correlated with each other, and entering them together into a regression produces severe multicollinearity — unstable coefficients that swing wildly depending on which items are included.

Second, the primary outcome was ordinal. Responses had a meaningful order but no guarantee that the distance between adjacent categories was equal. Treating that as a continuous variable assumes an interval scale the data does not have.

The approach

How it was analysed, and why that method

Reducing the item set before modelling

Exploratory factor analysis was run first, to identify the smaller number of underlying constructs the 75 items were collectively measuring. This solves the multicollinearity at source: instead of entering dozens of correlated items, the model uses a handful of interpretable factors, each representing a coherent concept the survey was actually measuring.

Factor solutions were assessed for interpretability as well as fit. A statistically defensible solution that cannot be given a meaningful name is of no use to an organisation that has to act on it.

Modelling an ordinal outcome as ordinal

Ordinal logistic regression was used for the outcome model. Linear regression was rejected because the dependent variable was ordinal — using it would have assumed equal spacing between response categories and produced predictions outside the range of the scale.

The most common error we see in survey work

Treating an ordinal outcome as continuous because it is coded 1 to 5. The coding is a convenience; it is not evidence that the gap between 1 and 2 equals the gap between 4 and 5.

Delivered

What the client received

Statistical report with the factor structure and its interpretation
Predictive models for the primary outcome
Presentation materials for a non-technical audience

The value

What changed as a result

The findings informed programme improvements and future strategic planning.

Because the analysis produced a small number of named factors rather than 75 coefficients, the results could be presented to a board and acted on — which is the difference between an analysis that is correct and one that is used.

About this case study

Written from the assigned statistician's own project notes. The client is not named and no identifying detail, data or figures are published. Scale and timeframe are approximate.

Send the data. Get a fixed quote.

Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.