Case studies  /  Manufacturing

Generalised Linear Mixed Models to Identify the Drivers of Production Defects

A quarter of a million production records across six lines and two years, where univariate analysis could not separate interacting process variables.

ManufacturingSector
Generalised linear mixed models, AIC selectionMethod
RSoftware
~250,000 records, six lines, two yearsScale
About four weeksDuration

The challenge

What the organisation came with

A manufacturing organisation needed to identify the factors driving defects in its production process, across roughly 250,000 records from six production lines over two years.

Defect rates varied between production lines and between shifts, and multiple process variables interacted simultaneously. Standard univariate analyses — testing one variable at a time — could not identify combined effects, and would attribute to a single variable what was actually the product of several acting together.

The clustering was the second problem. Records from the same production line share equipment, calibration and staffing. Treating 250,000 records as independent observations would have overstated the precision of every estimate.

The approach

How it was analysed, and why that method

Why a mixed model, and why generalised

Generalised linear mixed models were used. Generalised because the outcome was a defect rate rather than a continuous measurement, which requires an appropriate error distribution and link function rather than a normal-errors model. Mixed because random effects for production line and shift account for the clustering.

Simple logistic regression was rejected because it ignores clustering by production line and shift. At this scale that is not a marginal concern: it produces confident-looking coefficients whose standard errors are wrong.

Scale does not solve dependence

A quarter of a million records clustered in six lines carries far less independent information than the raw count suggests. More data does not fix a model that misrepresents the structure of the data.

Model selection

Candidate models were compared using the Akaike Information Criterion, which balances fit against complexity and guards against a model that explains the historical data well while generalising poorly. The selected model was then used to identify the process variables with the largest effect on defect rates, singly and in combination.

Delivered

What the client received

Statistical report identifying the principal drivers of defects
Predictive model of defect rates
Quality-control dashboard
Reproducible R scripts

The value

What changed as a result

The organisation introduced targeted process improvements that reduced defect rates and improved production consistency.

The dashboard turned the model from a one-off finding into an operational tool, monitoring the identified drivers as production continued.

About this case study

Written from the assigned statistician's own project notes. The client is not named and no identifying detail, data or figures are published. Scale and timeframe are approximate.

Send the data. Get a fixed quote.

Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.