🇺🇸The official website of Riot IQ
Log in
  • Home
  • About

Measure your
intelligence online.

Google

Assessments

  • All IQ Tests
  • Basic IQ Test
  • Full IQ Test
  • Custom IQ Test
  • Free IQ Test

Our Socials

  • X
  • YouTube
  • Facebook
  • LinkedIn

Other IQ Tests

  • WAIS-V
  • SB-5
  • Raven's 2
  • RIAS-2
  • CogAT 9
  • WISC-V

Community

  • Join Subreddit
  • Join Discord

Other Pages

  • Test Manual
  • Administer IQ Tests
  • About Us
  • Articles
  • Data
  • FAQ

Research

  • What do polygenic scores really predict?
  • Working speed and ability on the RIOT

Intelligence Journals & Organizations

  • Human Intelligence Research & Education (HIRE) Foundation
  • International Society for Intelligence Research (ISIR)
  • Intelligence & Cognitive Abilities Journal (ICA)
  • Intelligence Journal
  • Mensa Foundation

Contact

  • Email
  • Support

News & Press

  • International Society for Intelligence Research
  • American Thinker
  • Mensa Foundation (1/2)
  • Mensa Northern New Jersey
  • The University of Western Australia
  • Prolific
  • Quillette
  • Brainz

Our Articles

  • Gmatclub (1/2)
  • ApolloTechnical
  • LessWrong
  • Psychreg
  • Study in Switzerland
  • SuccessConsciousness
  • Creative Organizational Design (1/2)
  • ABNewsWire
  • Vanderbilt University

Our Articles

  • The Globe and Mail
  • Barchart
  • Journal
  • Mensa Foundation (2/2)
  • Psychologs
  • Creative Organizational Design (2/2)
  • AZBigMedia
  • Thoughts on Life and Love
  • Anxiety and Depression Association of America

Our Articles

  • Before It's News
  • Siglo XXI
  • TechBullion
  • Medium
  • Gmatclub (2/2)
  • MSN
  • National Review
  • Minding the Campus
  • Launching Next

Our Articles

  • Comparing Cronbach’s Alpha and McDonald’s Omega Reliability
  • Breaking the Intelligence & IQ Taboo
  • What is the Flynn Effect?
  • A Comprehensive History of IQ Tests
  • The 15 Subtests of the RIOT

Our Articles

  • How to Take an IQ Test
  • How to Calculate IQ
  • What is the RIOT IQ Test?
  • The Pro-Human Aspects of Intelligence Research
  • What is an IQ Test? A Beginner's Guide.

Our Articles

  • 5 Best IQ Tests in 2025
  • Cognitive Profiles on the RIOT IQ Test Results
  • 6 Cognitive Abilities of the RIOT
  • Are There Any Professional and Real Online IQ Tests?

Our Articles

  • Resources to Learn About IQ and Intelligence
  • Studying IQ Matters
  • The Search for Albert Einstein's IQ
  • Do Non-g Gains from the Flynn Effect Matter?

Riot IQ © 2026

  • Terms of Service
  • Privacy Policy
  • BAA Agreement
  • Test Administrator Terms
  • Terms of Service
  • •Privacy Policy
  • •BAA Agreement
  • •Test Administrator Terms

Table of Contents

  • Specifying the model before you look
  • The fit indices, and where the cut-offs actually came from
  • Comparing models against each other
  • What happened when publishers' models were rerun
  • Reading a CFA claim critically
  • Frequently asked questions
  • How is confirmatory factor analysis different from exploratory factor analysis?
  • What counts as good fit in CFA?
  • Why do large samples make every model look bad?
  • What is a bifactor model?
  • References
Sep 27, 2026·Advanced Topics & Research

Confirmatory factor analysis: testing a test's proposed structure

Confirmatory factor analysis tests a structure specified in advance against real data. Here is how IQ test publishers use it to justify their index scores.

Dr. Russell T. WarneChief Scientist
Share
Confirmatory factor analysis: testing a test's proposed structure
"Confirmatory factor analysis," usually shortened to CFA, states in advance which underlying abilities a test is supposed to measure and which subtests should load on each of them, then measures how badly that stated model fails to reproduce the correlations observed in a sample. It underlies almost every claim a test publisher makes about its index scores. When the WISC-V manual asserts that the battery measures five distinct factors, CFA is the evidence offered.

This page covers how a model is specified, what the fit indices mean and where their cut-offs came from, how competing models are compared, and what happened when independent researchers reran the analyses behind current commercial batteries. For what a factor is, and for the exploratory side of the method, see our companion article on factor analysis.


Specifying the model before you look

The defining feature of CFA is that the structure is a hypothesis written down first. The analyst names the latent factors, assigns each subtest to one or more of them, fixes the loadings theory says should be zero at exactly zero, and leaves the rest free. The software finds the parameter values that best reproduce the observed correlation matrix and reports how far off it still is.

That last point is usually got backwards. A CFA does not ask whether a structure is true. It asks how much worse the world looks than the model says it should. The oldest measure of this is the chi-square test of exact fit, which asks whether the gap between observed and model-implied correlations exceeds what sampling error would explain. With the sample sizes used for cognitive test standardization, often two thousand cases or more, that test rejects essentially every model, including good ones. The apparatus of approximate fit indices exists because of this.

Some choices matter more than they look. A model in which one subtest loads on three factors at once abandons the simple structure that makes factors interpretable. A "higher-order" model places a general factor above the group factors so that it influences subtests only indirectly, while a "bifactor" model lets a general factor and the group factors act on subtests directly and independently. Those are different claims about how ability is organised, and they can fit the same data very differently. Our page on the g factor covers what they argue about.

Models also fail mechanically. If an estimated factor variance comes out negative, an impossibility called a Heywood case, the solution cannot be interpreted and its fit statistics are meaningless. That happened to the WISC-V.


The fit indices, and where the cut-offs actually came from

Four approximate fit indices dominate reporting, and their conventional thresholds have a specific, often misquoted history.

• RMSEA, the root mean square error of approximation: Misfit per degree of freedom, so lower is better and zero is perfect. Steiger and Lind introduced the statistic in 1980. The familiar guidance came later: Browne and Cudeck (1993) recommended treating values below .05 as close fit, .05 to .08 as fair, and above .10 as poor, and MacCallum, Browne and Sugawara (1996) added .08 to .10 as mediocre while insisting the numbers "are intended as aids for interpretation of a value that lies on a continuous scale and not as absolute thresholds."

• CFI, the comparative fit index: Bentler (1990) proposed it as an improvement on earlier incremental indices, scaled so that 1.0 means the model reproduces the data as well as a saturated model, and measured against a baseline in which all variables are uncorrelated.

• TLI, the Tucker-Lewis index: Older than CFI, from Tucker and Lewis (1973), similar but penalised for model complexity, so it often sits slightly below CFI for the same model.

• SRMR, the standardized root mean square residual: The average difference between observed and model-implied correlations, read directly in correlation units.

The ".90 is acceptable" rule that still appears in textbooks traces to Bentler and Bonett (1980). The stricter numbers most journals now expect come from Hu and Bentler (1999), whose simulations concluded that cut-offs "close to .95" for TLI and CFI, "close to .08 for SRMR" and "close to .06 for RMSEA" produced lower Type II error rates.

Three details about that paper are routinely dropped. Its actual recommendation was a two-index rule rather than a checklist: CFI or TLI near .95 in combination with SRMR, and in the combinational analysis the recommended SRMR value was close to .09, with .96 and SRMR above .09 giving the lowest combined error rate. Their conclusions were conditional on sample size, with warnings that below 250 cases the trade-off between error types changes. Marsh, Hau and Wen (2004) then showed that treating these as golden rules produces perverse results, including cases where the probability of correctly rejecting a misspecified model falls as the sample grows. McNeish and Wolf (2023) argued that cut-offs have to be simulated for the specific model and sample at hand, since a fixed threshold means different things for a 10-indicator model and a 40-indicator one.


Comparing models against each other

Because no single model passes or fails cleanly, and because these thresholds are conventions rather than properties of good measurement, CFA in test development is mostly comparative. Two models are "nested" when one is a constrained version of the other, and the difference in their chi-square values against the difference in degrees of freedom formally tests whether the extra parameters bought anything. For non-nested comparisons, information criteria such as the AIC penalise complexity and allow ranking.

Rules of thumb are also used. Analysts commonly treat a CFI change of at least .01 together with an RMSEA change of at least .015 as evidence that one model fits meaningfully better. For measurement invariance, the question of whether a structure holds the same way across groups, Cheung and Rensvold (2002) proposed a CFI change of no more than .01. These conventions carry the same caveat as the absolute cut-offs.


What happened when publishers' models were rerun

Internal-structure evidence in a technical manual is a requirement rather than a courtesy. Standard 1.13 of the AERA, APA and NCME Standards for Educational and Psychological Testing requires it wherever a score interpretation depends on premises about the relationships among parts of a test, and Standard 1.14 adds that when composite scores are developed, "the basis and rationale for arriving at the composites should be given."

The WISC-V is the best-documented case. Its Technical and Interpretive Manual rests its structural claims entirely on CFA and settles on a five-factor higher-order model in which Arithmetic cross-loads on three group factors and the standardized path between the general factor and Fluid Reasoning is 1.00, a value implying the two are redundant. Canivez, Watkins and Dombrowski (2017) reran every model in that manual on the same standardization sample of 2,200 children using maximum likelihood estimation. Every higher-order model with five group factors produced negative variance for Fluid Reasoning and so could not be interpreted at all. Among the models that estimated properly, a one-factor model fitted poorly, with CFI .843, TLI .819, SRMR .065 and RMSEA .103. A four-factor higher-order model fitted well, at CFI .969, TLI .963, SRMR .032 and RMSEA .047, and the best fit of the sixteen models tested was a four-group-factor bifactor model, at CFI .986, TLI .980, SRMR .023 and RMSEA .034.

The substantive conclusion mattered more than the fit numbers. In the bifactor solutions the general factor dominated, and the reliability of the group factors once it was removed was low enough that the authors judged most index scores of questionable interpretive value on their own. Our comparison of alpha and omega reliability coefficients explains the omega-hierarchical statistic that argument turns on.

The pattern repeats. DiStefano and Dombrowski (2006) analysed the Stanford-Binet Fifth Edition standardization data and found one or two factors adequate where the manual specifies five. Dombrowski, McGill and Canivez (2018) recovered seven factors from the Woodcock-Johnson IV full battery across the school-age range, with subtests migrating away from their assigned Cattell-Horn-Carroll factors and Fluid Reasoning and Quantitative Reasoning collapsing into one. In both cases the published and the replicated structures clear conventional fit thresholds, which is the problem with treating those thresholds as a verdict.


Reading a CFA claim critically

A fit index describes a model's distance from one sample's covariances. It cannot tell you the factors are the abilities their labels name, and acceptable fit for a five-factor model does not rule out a four-factor model that fits as well. Internal structure is also only one source of validity evidence the Standards describe, which is why our page on construct validity treats it as part of a larger argument.

Four questions are worth asking of any published CFA. Were rival models tested, or only the preferred one? Were the estimates proper, with no negative variances? Was the winning model's advantage large enough to matter? And were the scores the structure justifies shown to be reliable once the general factor is accounted for? A publisher who reports enough detail for those questions to be answered is doing something the field does not always do. Readers who want index scores from an online IQ test built by psychometricians can take the Reasoning and Intelligence Online Test.


Frequently asked questions

How is confirmatory factor analysis different from exploratory factor analysis?

Exploratory analysis asks the data how many factors there are and which subtests belong to each. Confirmatory analysis begins with that structure specified and estimates how well it reproduces the observed correlations. Fixing loadings at zero in advance is the mechanical difference.

What counts as good fit in CFA?

By the most-cited convention, CFI and TLI at or above about .95 with RMSEA at or below about .06 and SRMR at or below about .08. Those figures come from Hu and Bentler's 1999 simulations and are conditional on the model type and sample size studied there, making them a starting point for judgement rather than a pass mark.

Why do large samples make every model look bad?

The chi-square test of exact fit depends on sample size, so with thousands of cases even trivial discrepancies become statistically significant. Approximate fit indices measure how large the misfit is rather than whether any exists.

What is a bifactor model?

A model in which a general factor and several group factors all influence the subtests directly, rather than the general factor acting through the group factors. It makes the variance attributable to the general factor visible, and for several Wechsler batteries it has fitted better than the higher-order models the manuals report.


References

1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net

2. Hu, L., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling, 6(1), 1-55. doi.org

3. MacCallum, R. C., Browne, M. W., & Sugawara, H. M. (1996). Power analysis and determination of sample size for covariance structure modeling. Psychological Methods, 1(2), 130-149. statpower.net

4. Bentler, P. M. (1990). Comparative fit indexes in structural models. Psychological Bulletin, 107(2), 238-246. doi.org

5. Tucker, L. R., & Lewis, C. (1973). A reliability coefficient for maximum likelihood factor analysis. Psychometrika, 38(1), 1-10. doi.org

6. Marsh, H. W., Hau, K.-T., & Wen, Z. (2004). In search of golden rules: Comment on hypothesis-testing approaches to setting cutoff values for fit indexes. Structural Equation Modeling, 11(3), 320-341. ora.ox.ac.uk

7. McNeish, D., & Wolf, M. G. (2023). Dynamic fit index cutoffs for confirmatory factor analysis models. Psychological Methods, 28(1), 61-88. doi.org

8. Cheung, G. W., & Rensvold, R. B. (2002). Evaluating goodness-of-fit indexes for testing measurement invariance. Structural Equation Modeling, 9(2), 233-255. doi.org

9. Canivez, G. L., Watkins, M. W., & Dombrowski, S. C. (2017). Structural validity of the Wechsler Intelligence Scale for Children-Fifth Edition: Confirmatory factor analyses with the 16 primary and secondary subtests. Psychological Assessment, 29(4), 458-472. doi.org

10. DiStefano, C., & Dombrowski, S. C. (2006). Investigating the theoretical structure of the Stanford-Binet Fifth Edition. Journal of Psychoeducational Assessment, 24(2), 123-136. eric.ed.gov

11. Dombrowski, S. C., McGill, R. J., & Canivez, G. L. (2018). Hierarchical exploratory factor analyses of the Woodcock-Johnson IV full test battery. School Psychology Quarterly, 33(2), 235-250. doi.org

Hero image: curtain wall under construction at the Amsterdam Public Library, by Fons Heijnsbroek, released under CC0 1.0 (creativecommons.org/publicdomain/zero/1.0). Via Wikimedia Commons.

Take our professional IQ test

Want to know your IQ? Try the first ever professional online IQ test.

Try our IQ test
Author
Dr. Russell T. WarneChief Scientist

Contact

Table of Contents

  • Specifying the model before you look
  • The fit indices, and where the cut-offs actually came from
  • Comparing models against each other
  • What happened when publishers' models were rerun
  • Reading a CFA claim critically
  • Frequently asked questions
  • How is confirmatory factor analysis different from exploratory factor analysis?
  • What counts as good fit in CFA?
  • Why do large samples make every model look bad?
  • What is a bifactor model?
  • References
Article Categories
All ArticlesUnderstanding IQ ScoresTaking an IQ TestRIOT-Specific InformationGeneral IQ & IntelligenceAdvanced Topics & ResearchIQ Scores & InterpretationMensa & High-IQ SocietiesOnline IQ Tests IQ Test Basics & FundamentalsAverage IQ & DemographicsFamous People & IQHistory & Origins Of IQ TestingAccuracy, Reliability & CriticismSpecial Population & Related ConditionsImproving IQ / PreparationSpecific IQ Tests & FormatsIQ Testing for HR & RecruitmentSkills Assessment
Related Articles
What is an independent educational evaluation? The rules in 34 CFR §300.502Response to intervention: how schools decide a student needs more helpComputer adaptive testing: how a test that rebuilds itself as you go worksItem response theory: how a test models one item at a timeClassical test theory: the model behind most published test scoresConfirmatory factor analysis: testing a test's proposed structureFactor analysis: how the structure of an IQ test is discoveredPsychometrics: the science of building and evaluating mental testsWhat Is the Difference Between Intellectual Disability and Learning Disability? What Jobs Need High Spatial Ability?Is IQ Correlated With Dementia?Does High IQ Actually Correlate With Higher Salary After Age 30?High IQ vs. High EQ: Which One Predicts Long-Term Relationship Happiness?Are You Left-Brained or Right-Brained? What Neuroscience Actually SaysWhat Does Too Much Screen Time Do to Children's Brains?Types of IQ: The Quotients Explained (IQ, EQ, SQ, AQ, CQ)What Is the Dunning-Kruger Effect? What the Research Actually ShowsSame Test, Different Patterns: How ADHD and Autism Show Up Differently Across IQ SubtestsGeneral Intelligence vs. Multiple Intelligences: What Each Theory Gets RightNeuroplasticity in Your 30s and 40s: What the Science Actually SaysThe G-Factor vs. Gardner's Multiple Intelligences: What the Evidence Actually ShowsCan Hyperlexia Make You Seem Smarter Than You Actually Are?Is There a Correlation Between IQ and Reaction Time?Can Exercise Affect Your IQ Score?How Does IQ Change as a Person Ages?What Is the Flynn Effect and Why Are IQ Scores Rising?What Part of the Brain Controls IQ and Cognitive Function?Unlocking Your Potential: The Role of Online IQ TestingLogical Reasoning on an IQ Test: How It's Defined, Measured, and Why It Predicts So MuchThe Science Behind IQ Tests: Understanding Intelligence AssessmentExploring the Controversies Surrounding IQ TestsThe Future of IQ Testing: Trends and InnovationsVisual & Spatial Reasoning: What It Is and Why It Shows Up on an IQ TestFluid vs. Crystallized Intelligence: What the Difference Actually MeansIs Everyone About as Smart as I Am???Does Intelligence Research Undermine the Fight against Inequality?Does Intelligence Research Lead to Negative Social Policies?Do Past Controversies Taint Modern Research on Intelligence?Should Controversial or Unpopular Ideas Be Held to a Higher Standard of Evidence?Does Stereotype Threat Explain Score Gaps among Demographic Groups?Do Unique Influences Operate on One Group’s Intelligence Test Scores?Are Racial/Ethnic Group IQ Differences Completely Environmental in Origin?Do Males and Females Have the Same Distribution of IQ Scores?Is Emotional Intelligence a Real Ability that Is Helpful in Life?Is Very High Intelligence More Beneficial than Moderately High Intelligence?Are Intelligence Tests Designed to Create or Perpetuate a False Meritocracy?Is Intelligence Important in the Workplace?Do IQ Scores Just Measure How Good Someone is at Taking Tests?Are Admissions Tests A Barrier to College for Underrepresented Students? Do Non-Cognitive Variables Have Powerful Effects on Academic Achievement?Can Effective Schools Make Every Child Academically Proficient?Is Every Child Gifted?Does Improvability of IQ Mean Intelligence Can Be Equalized?Can Braining-Training Programs Raise IQ?Can Social Interventions Drastically Raise IQ?Are Genes Important for Determining Intelligence?Is Raising IQ Possible?Does IQ Reflect A Person’s Socioeconomic Status?Are Intelligence Tests Biased Against Diverse Populations?Is Practical Intelligence a Real Ability Separate from General Intelligence?Is Intelligence Just A Western Concept?Does IQ Correspond to Brain Anatomy or Functioning?Measuring Cognitive Aging with Memory and Processing Speed TasksComparing Cronbach’s Alpha and McDonald’s Omega ReliabilityWhat is the Flynn Effect? Is the World’s Collective IQ Increasing or Decreasing?How Do You Test Cognitive Functions?Does High IQ Correlate with Success?Creating an IQ TestIncreasing Your IQCulture-Fair Intelligence TestsFlynn EffectCognitive DevelopmentIQ Test QualityDo Non-g Gains from the Flynn Effect Matter?
Take our IQ tests

Basic IQ Test

5 subtests + 5 cognitive abilities

Take the IQ test

Features

  • ~13 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±5.6 IQ margin of error

5/15 Subtests

Learn more
Vocabulary
Matrix Reasoning
SToVeS
Visual Reversal
Symbol Search
Most comprehensive

Full IQ Test

15 subtests + all cognitive abilities

Take the IQ test

Features

  • ~52 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±3.7 IQ margin of error

15/15 Subtests

Learn more
Vocabulary, Information, Analogies
Matrix Reasoning, Visual Puzzles, Figure Weights
Object Rotation, SToVeS, Spatial Orientation
Computation Span, Exposure Memory, Visual Reversal
Symbol Search, Abstract Matching
Simple Reaction Time, Choice Reaction Time
Compare all tests