Confirmatory factor analysis: testing a test's proposed structure
Confirmatory factor analysis tests a structure specified in advance against real data. Here is how IQ test publishers use it to justify their index scores.
Dr. Russell T. WarneChief Scientist
Share
"Confirmatory factor analysis," usually shortened to CFA, states in advance which underlying abilities a test is supposed to measure and which subtests should load on each of them, then measures how badly that stated model fails to reproduce the correlations observed in a sample. It underlies almost every claim a test publisher makes about its index scores. When the WISC-V manual asserts that the battery measures five distinct factors, CFA is the evidence offered.
This page covers how a model is specified, what the fit indices mean and where their cut-offs came from, how competing models are compared, and what happened when independent researchers reran the analyses behind current commercial batteries. For what a factor is, and for the exploratory side of the method, see our companion article on factor analysis.
Specifying the model before you look
The defining feature of CFA is that the structure is a hypothesis written down first. The analyst names the latent factors, assigns each subtest to one or more of them, fixes the loadings theory says should be zero at exactly zero, and leaves the rest free. The software finds the parameter values that best reproduce the observed correlation matrix and reports how far off it still is.
That last point is usually got backwards. A CFA does not ask whether a structure is true. It asks how much worse the world looks than the model says it should. The oldest measure of this is the chi-square test of exact fit, which asks whether the gap between observed and model-implied correlations exceeds what sampling error would explain. With the sample sizes used for cognitive test standardization, often two thousand cases or more, that test rejects essentially every model, including good ones. The apparatus of approximate fit indices exists because of this.
Some choices matter more than they look. A model in which one subtest loads on three factors at once abandons the simple structure that makes factors interpretable. A "higher-order" model places a general factor above the group factors so that it influences subtests only indirectly, while a "bifactor" model lets a general factor and the group factors act on subtests directly and independently. Those are different claims about how ability is organised, and they can fit the same data very differently. Our page on the g factor covers what they argue about.
Models also fail mechanically. If an estimated factor variance comes out negative, an impossibility called a Heywood case, the solution cannot be interpreted and its fit statistics are meaningless. That happened to the WISC-V.
The fit indices, and where the cut-offs actually came from
Four approximate fit indices dominate reporting, and their conventional thresholds have a specific, often misquoted history.
• RMSEA, the root mean square error of approximation: Misfit per degree of freedom, so lower is better and zero is perfect. Steiger and Lind introduced the statistic in 1980. The familiar guidance came later: Browne and Cudeck (1993) recommended treating values below .05 as close fit, .05 to .08 as fair, and above .10 as poor, and MacCallum, Browne and Sugawara (1996) added .08 to .10 as mediocre while insisting the numbers "are intended as aids for interpretation of a value that lies on a continuous scale and not as absolute thresholds."
• CFI, the comparative fit index: Bentler (1990) proposed it as an improvement on earlier incremental indices, scaled so that 1.0 means the model reproduces the data as well as a saturated model, and measured against a baseline in which all variables are uncorrelated.
• TLI, the Tucker-Lewis index: Older than CFI, from Tucker and Lewis (1973), similar but penalised for model complexity, so it often sits slightly below CFI for the same model.
• SRMR, the standardized root mean square residual: The average difference between observed and model-implied correlations, read directly in correlation units.
The ".90 is acceptable" rule that still appears in textbooks traces to Bentler and Bonett (1980). The stricter numbers most journals now expect come from Hu and Bentler (1999), whose simulations concluded that cut-offs "close to .95" for TLI and CFI, "close to .08 for SRMR" and "close to .06 for RMSEA" produced lower Type II error rates.
Three details about that paper are routinely dropped. Its actual recommendation was a two-index rule rather than a checklist: CFI or TLI near .95 in combination with SRMR, and in the combinational analysis the recommended SRMR value was close to .09, with .96 and SRMR above .09 giving the lowest combined error rate. Their conclusions were conditional on sample size, with warnings that below 250 cases the trade-off between error types changes. Marsh, Hau and Wen (2004) then showed that treating these as golden rules produces perverse results, including cases where the probability of correctly rejecting a misspecified model falls as the sample grows. McNeish and Wolf (2023) argued that cut-offs have to be simulated for the specific model and sample at hand, since a fixed threshold means different things for a 10-indicator model and a 40-indicator one.
Comparing models against each other
Because no single model passes or fails cleanly, and because these thresholds are conventions rather than properties of good measurement, CFA in test development is mostly comparative. Two models are "nested" when one is a constrained version of the other, and the difference in their chi-square values against the difference in degrees of freedom formally tests whether the extra parameters bought anything. For non-nested comparisons, information criteria such as the AIC penalise complexity and allow ranking.
Rules of thumb are also used. Analysts commonly treat a CFI change of at least .01 together with an RMSEA change of at least .015 as evidence that one model fits meaningfully better. For measurement invariance, the question of whether a structure holds the same way across groups, Cheung and Rensvold (2002) proposed a CFI change of no more than .01. These conventions carry the same caveat as the absolute cut-offs.
What happened when publishers' models were rerun
Internal-structure evidence in a technical manual is a requirement rather than a courtesy. Standard 1.13 of the AERA, APA and NCME Standards for Educational and Psychological Testing requires it wherever a score interpretation depends on premises about the relationships among parts of a test, and Standard 1.14 adds that when composite scores are developed, "the basis and rationale for arriving at the composites should be given."
The WISC-V is the best-documented case. Its Technical and Interpretive Manual rests its structural claims entirely on CFA and settles on a five-factor higher-order model in which Arithmetic cross-loads on three group factors and the standardized path between the general factor and Fluid Reasoning is 1.00, a value implying the two are redundant. Canivez, Watkins and Dombrowski (2017) reran every model in that manual on the same standardization sample of 2,200 children using maximum likelihood estimation. Every higher-order model with five group factors produced negative variance for Fluid Reasoning and so could not be interpreted at all. Among the models that estimated properly, a one-factor model fitted poorly, with CFI .843, TLI .819, SRMR .065 and RMSEA .103. A four-factor higher-order model fitted well, at CFI .969, TLI .963, SRMR .032 and RMSEA .047, and the best fit of the sixteen models tested was a four-group-factor bifactor model, at CFI .986, TLI .980, SRMR .023 and RMSEA .034.
The substantive conclusion mattered more than the fit numbers. In the bifactor solutions the general factor dominated, and the reliability of the group factors once it was removed was low enough that the authors judged most index scores of questionable interpretive value on their own. Our comparison of alpha and omega reliability coefficients explains the omega-hierarchical statistic that argument turns on.
The pattern repeats. DiStefano and Dombrowski (2006) analysed the Stanford-Binet Fifth Edition standardization data and found one or two factors adequate where the manual specifies five. Dombrowski, McGill and Canivez (2018) recovered seven factors from the Woodcock-Johnson IV full battery across the school-age range, with subtests migrating away from their assigned Cattell-Horn-Carroll factors and Fluid Reasoning and Quantitative Reasoning collapsing into one. In both cases the published and the replicated structures clear conventional fit thresholds, which is the problem with treating those thresholds as a verdict.
Reading a CFA claim critically
A fit index describes a model's distance from one sample's covariances. It cannot tell you the factors are the abilities their labels name, and acceptable fit for a five-factor model does not rule out a four-factor model that fits as well. Internal structure is also only one source of validity evidence the Standards describe, which is why our page on construct validity treats it as part of a larger argument.
Four questions are worth asking of any published CFA. Were rival models tested, or only the preferred one? Were the estimates proper, with no negative variances? Was the winning model's advantage large enough to matter? And were the scores the structure justifies shown to be reliable once the general factor is accounted for? A publisher who reports enough detail for those questions to be answered is doing something the field does not always do. Readers who want index scores from an online IQ test built by psychometricians can take the Reasoning and Intelligence Online Test.
Frequently asked questions
How is confirmatory factor analysis different from exploratory factor analysis?
Exploratory analysis asks the data how many factors there are and which subtests belong to each. Confirmatory analysis begins with that structure specified and estimates how well it reproduces the observed correlations. Fixing loadings at zero in advance is the mechanical difference.
What counts as good fit in CFA?
By the most-cited convention, CFI and TLI at or above about .95 with RMSEA at or below about .06 and SRMR at or below about .08. Those figures come from Hu and Bentler's 1999 simulations and are conditional on the model type and sample size studied there, making them a starting point for judgement rather than a pass mark.
Why do large samples make every model look bad?
The chi-square test of exact fit depends on sample size, so with thousands of cases even trivial discrepancies become statistically significant. Approximate fit indices measure how large the misfit is rather than whether any exists.
What is a bifactor model?
A model in which a general factor and several group factors all influence the subtests directly, rather than the general factor acting through the group factors. It makes the variance attributable to the general factor visible, and for several Wechsler batteries it has fitted better than the higher-order models the manuals report.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Hu, L., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling, 6(1), 1-55. doi.org
3. MacCallum, R. C., Browne, M. W., & Sugawara, H. M. (1996). Power analysis and determination of sample size for covariance structure modeling. Psychological Methods, 1(2), 130-149. statpower.net
4. Bentler, P. M. (1990). Comparative fit indexes in structural models. Psychological Bulletin, 107(2), 238-246. doi.org
5. Tucker, L. R., & Lewis, C. (1973). A reliability coefficient for maximum likelihood factor analysis. Psychometrika, 38(1), 1-10. doi.org
6. Marsh, H. W., Hau, K.-T., & Wen, Z. (2004). In search of golden rules: Comment on hypothesis-testing approaches to setting cutoff values for fit indexes. Structural Equation Modeling, 11(3), 320-341. ora.ox.ac.uk
7. McNeish, D., & Wolf, M. G. (2023). Dynamic fit index cutoffs for confirmatory factor analysis models. Psychological Methods, 28(1), 61-88. doi.org
8. Cheung, G. W., & Rensvold, R. B. (2002). Evaluating goodness-of-fit indexes for testing measurement invariance. Structural Equation Modeling, 9(2), 233-255. doi.org
9. Canivez, G. L., Watkins, M. W., & Dombrowski, S. C. (2017). Structural validity of the Wechsler Intelligence Scale for Children-Fifth Edition: Confirmatory factor analyses with the 16 primary and secondary subtests. Psychological Assessment, 29(4), 458-472. doi.org
10. DiStefano, C., & Dombrowski, S. C. (2006). Investigating the theoretical structure of the Stanford-Binet Fifth Edition. Journal of Psychoeducational Assessment, 24(2), 123-136. eric.ed.gov
11. Dombrowski, S. C., McGill, R. J., & Canivez, G. L. (2018). Hierarchical exploratory factor analyses of the Woodcock-Johnson IV full test battery. School Psychology Quarterly, 33(2), 235-250. doi.org
Hero image: curtain wall under construction at the Amsterdam Public Library, by Fons Heijnsbroek, released under CC0 1.0 (creativecommons.org/publicdomain/zero/1.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.