Factor analysis

Exploratory Factor Analysis (EFA)

Works out from the data how many latent dimensions sit behind a set of items, and which items move together on each of them.

Method summary

Exploratory factor analysis takes a pool of items and reduces the variance they share to a small number of latent dimensions. Its raw material is the correlation matrix between items: items that rise and fall together are read as indicators of the same dimension. In the panel the items are standardised, a Pearson correlation matrix is built, and that matrix is decomposed into eigenvalues and eigenvectors. If you do not state a number of factors, the count of eigenvalues greater than one is used, which is the Kaiser rule. Loadings are formed from the eigenvectors along the principal axis route and varimax rotation is applied by default, so that each item stands out on a single dimension. The word exploratory carries weight here: the method describes the structure visible in the data instead of testing a structure you brought with you. Testing a structure specified in advance is the job of confirmatory factor analysis.

Which research questions does it answer?

  • Does the 18-item scale I am developing measure one thing, or does it break into several subscales?
  • How many distinct clusters do brand image statements form in respondents' minds?
  • Can a long survey battery be reduced to fewer dimensions without losing much information?
  • Which items drift away from the dimension they were written for and are candidates for removal?

When should you use it?

  • When there is no binding theoretical expectation about the dimensional structure and the structure is to be derived from the data.
  • When a scale is newly written or adapted to another culture and the dimensions have to be settled before reliability is computed.
  • When a large set of highly correlated variables needs to be condensed into a few dimensions for use in later analyses.
  • When at least three items are available; the engine refuses to run with fewer than three variables.
  • When all the items come from the same respondents and share the same response format.

Required variable types

  • Every column entering the analysis must be numeric: Likert items, scores or scale totals. Setting the measurement level to 'scale' on the Variable tab lets the method picker recognise them.
  • There is no split into dependent and independent variables. The method treats a set of items as a whole and no outcome variable is chosen.
  • At least three items are required. Reverse-worded items must already be recoded, otherwise their loadings appear with the opposite sign.
  • One row per respondent. A row with a missing value on any of the selected items drops out of the analysis entirely.

Key assumptions

Meaningful association among the items
If the correlation matrix is close to an identity matrix there is no shared variance and therefore nothing to extract. Put differently, if the items are mutually unrelated the method has nothing to condense.
Sampling adequacy
The variance each item shares with the others should be sizeable relative to the partial correlations. Where adequacy is low, the extracted factors remain specific to the sample and do not reappear in another one.
Ratio of observations to items
Correlation coefficients are unstable in small samples, and a factor structure drawn from an unstable matrix is unstable in turn. The literature often cites 10 observations per item and a total of 200 or more respondents as a rule of thumb rather than a hard threshold.
Linearity and continuous measurement
Because the associations are summarised with Pearson correlations, relationships are assumed linear and items are assumed to be measured close to interval level. With binary or markedly skewed items the coefficients understate the association.
Absence of influential respondents
A single extreme response pattern can visibly change the correlation between two items and shift the factor structure.

How YouReply checks these assumptions

  • Meaningful association among the items: Bartlett's test of sphericity is computed on every run, and its chi-square value and p value appear on a separate line of the result. A large p value means the correlation matrix is not suitable for factoring.
  • Sampling adequacy: The Kaiser-Meyer-Olkin measure is computed and shown with a verbal label, on a ladder running from the top band at 0.90 and above down to unacceptable below 0.50. The labels follow Kaiser's published classification.
  • Ratio of observations to items: The engine enforces a floor of 5 complete observations per item: below that the analysis does not run and the error states how many rows were found. Clearing that floor does not make a sample adequate, it only makes the computation possible.
  • Linearity and continuous measurement: The panel builds the correlation matrix from Pearson coefficients only; there is no polychoric or tetrachoric option. To inspect the item distributions you need to run descriptive statistics or the normality tests separately.
  • Absence of influential respondents: The panel does not check this assumption automatically; the researcher evaluates it.

How the analysis is run

  1. 1Drag the file holding your item pool (CSV or XLSX) onto the panel, arranged with respondents in rows and items in columns.
  2. 2In the editable grid on the Data tab, check that reverse-worded items are already recoded; define a derived column on the Variable tab if they are not.
  3. 3On the Variable tab set the items to the 'scale' measurement level, declare missing value codes such as 99 as missing, and enter value labels.
  4. 4Open the method picker on the Analysis tab and choose this method from the factor analysis category. With the 'Only ones that fit my data' switch on, unsuitable methods are dimmed with a plain-language reason rather than hidden, and the data requirements box shows which columns can serve as items.
  5. 5Tick the items in the parameter form. Leaving the factor count empty applies the Kaiser rule; the rotation field arrives set to varimax and the extraction field to the principal axis route.
  6. 6Start the run. The scree plot sits at the top of the results, with the sampling adequacy block, the loading matrix, the eigenvalues, the explained variance and the communalities listed below it in collapsible sections.
  7. 7Download the tables as Excel and the charts as PNG. The computation credits card names the library call and its version, and the history tab keeps your earlier runs. The free plan allows 50 runs per month.

Statistics and tables produced

Sampling adequacy (KMO) with its label
A single KMO value together with a verbal label for the band it falls into. Item-level KMO values are not produced.
Bartlett's test of sphericity
The chi-square statistic and p value testing whether the correlation matrix differs from an identity matrix. It is derived from the determinant and reported on one line.
Loading matrix
The loading of every item on every retained factor. Where a rotation has been applied the table holds rotated loadings, and the result states which rotation was used.
Eigenvalues
The eigenvalues of the retained factors, plus the full list of eigenvalues of the correlation matrix. The full list is what the scree plot and the decision on factor count rest on.
Explained variance
The percentage of variance accounted for by each retained factor and the cumulative total. The percentages come from each eigenvalue as a share of total variance.
Communalities
For each item, the proportion of its variance accounted for by the retained factors. Items with low communalities contribute little to the solution.
Details of the solution
Number of factors, number of items, number of complete observations used, and the names of the rotation and extraction applied. The observation count is what remains after rows with missing values are dropped.
Scree plot and loading radar
A line chart of explained variance by factor order is drawn on every run. When there are at most 15 items and at most 5 factors, a radar chart of the loadings is added.

Effect size and confidence intervals

Cumulative explained variance
How much of the total variance in the item pool the solution gathers. A cumulative figure around 50% is often quoted as a lower reference point in the social sciences; it is a point of comparison rather than a threshold, and the percentage climbs on its own as more factors are retained.
Communality
The share of an item's variance explained by the solution. Values below 0.30 suggest the item is weakly tied to the common structure, and removal decisions weigh this against the item's contribution to content coverage.
Item loading
The strength of an item's link to a factor. Cut-offs between 0.30 and 0.40 are common in practice, and items loading similarly on more than one factor (cross-loadings) warrant a closer look.

No output of exploratory factor analysis comes with a confidence interval. Loadings, eigenvalues, communalities and explained variance percentages are reported as point estimates, and no standard errors accompany them either. There is therefore no test of whether a loading differs from zero or whether two loadings differ from one another. A researcher who wants to see whether the structure replicates can split the sample and fit a confirmatory factor analysis on the second half.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
How many dimensions do the 18 items written for an organisational trust scale form, and which items are candidates for removal? (Illustrative example, not real data.)
Variables
18 items on a five-point Likert format, all numeric and flagged as scale level · 420 respondents with complete answers (23 observations per item) · Factor count left empty, varimax rotation
Example result
Sampling adequacy was KMO = 0.87 and Bartlett's test of sphericity was significant, chi-square(153) = 2841.6, p < 0.001. The Kaiser rule retained three factors with eigenvalues of 5.65, 2.30 and 1.48. Explained variance was 31.4%, 12.8% and 8.2% respectively, 52.4% cumulative. Rotated loadings on the intended factors ranged from 0.58 to 0.81, with cross-loadings under 0.30. Communalities ran from 0.34 to 0.71; only item 7 stood apart, with a communality of 0.21 and two close loadings of 0.34 and 0.31.
Interpretation
The adequacy measure and the sphericity test both support using this correlation matrix for factoring. The three-factor solution gathers a little over half of the total variance and most items stand out cleanly on one dimension. What the three dimensions mean has to be read off the content of the items; the numbers do not name them. Item 7 is a candidate for removal on two counts, its low communality and its cross-loading, and the decision should also weigh what the item contributes to content coverage, with the analysis re-run once it is dropped. This solution is a structure worth testing, not evidence that the structure is correct.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

Variable examined

Scale development data (synthetic)

A twenty item five point Likert scale answered by three hundred people. The items come from two latent constructs: the first ten measure one construct, the last ten the other, and the two are moderately correlated. The factor and reliability methods run on this file.

Rows
300
Columns
respondent_id, item_1, item_2, item_3, item_4, item_5, item_6, item_7, item_8, item_9, item_10, item_11, item_12, item_13, item_14, item_15, item_16, item_17, item_18, item_19, item_20
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: Do the twenty items separate into the two expected factors, and which item loads on which factor?

Analyzed variables
item_1,item_2,item_3,item_4,item_5,item_6,item_7,item_8,item_9,item_10,item_11,item_12,item_13,item_14,item_15,item_16,item_17,item_18,item_19,item_20
Requested number of factors
2
Rotation method
varimax
Extraction method
principal_axis
KMO measure of sampling adequacy
0.914
Bartlett's test of sphericity (chi-square)
2,109.90
Bartlett's test of sphericity (p value)
0
Number of factors
2
Number of variables
20
Valid observations
300
rotation
varimax
extraction
principal_axis

Loading matrix

Loading matrix
RowFactor 1Factor 2
item_1-0.734-0.046
item_2-0.688-0.052
item_3-0.701-0.072
item_4-0.727-0.065
item_5-0.722-0.092
item_6-0.701-0.082
item_7-0.678-0.057
item_8-0.686-0.031
item_9-0.707-0.082
item_10-0.736-0.048
item_11-0.103-0.372
item_120.006-0.393
item_13-0.035-0.421
item_14-0.064-0.410
item_15-0.149-0.415
item_16-0.229-0.404
item_17-0.199-0.401
item_18-0.114-0.439
item_19-0.069-0.430
item_20-0.155-0.439

Eigenvalues

Factor 1
5.70
Factor 2
3.79

Variance explained (percent)

Factor 1
28.52
Factor 2
18.93

Cumulative variance explained

Factor 1
28.52
Factor 2
47.45

Communalities

item_1
0.541
item_2
0.477
item_3
0.496
item_4
0.533
item_5
0.530
item_6
0.498
item_7
0.462
item_8
0.472
item_9
0.507
item_10
0.543
item_11
0.149
item_12
0.154
item_13
0.179
item_14
0.172
item_15
0.194
item_16
0.216
item_17
0.201
item_18
0.206
item_19
0.190
item_20
0.217

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260915

How to report the result

The item pool was submitted to exploratory factor analysis with varimax rotation (n = 420). The correlation matrix was suitable for factoring, KMO = .87, Bartlett's chi-square(153) = 2841.6, p < .001. Three factors were retained under the eigenvalue-greater-than-one rule, accounting for 52.4% of the total variance, with loadings on the intended factors ranging from .58 to .81.

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • It does not test a dimensional structure drawn from theory. The question of whether items fall into the three dimensions you expected calls for a test of fit, and that test belongs to confirmatory factor analysis.
  • The 'ml' label in the extraction field does not correspond to a separate computation; it runs the same principal axis route. Genuine maximum likelihood estimation, the fit chi-square that goes with it, and a likelihood ratio test for the number of factors are not part of this method.
  • Among the rotation options, quartimax returns the varimax solution. Promax is applied as a power transformation of the varimax solution, and even when an oblique rotation is chosen no matrix of inter-factor correlations is reported, so the output cannot tell you how far the dimensions separate.
  • No standard errors or confidence intervals are computed for any output, so differences between loadings are judged by eye.
  • The correlation matrix rests on Pearson coefficients. With binary items, or scales showing floor and ceiling effects, these coefficients can understate associations and give rise to artificial extra factors.
  • A row missing a value on any selected item is dropped in full. Where missingness is widespread, the analysis runs on a progressively narrower and self-selected subset of the data.

What to use when the assumptions are not met

  • Confirmatory Factor Analysis (CFA)Why: When an expectation about the structure is already in place; item-to-factor assignments are declared up front and the model's fit to the data is tested.
  • Parallel AnalysisWhy: When the factor count should not rest on the Kaiser rule alone; it compares observed eigenvalues against those from random data.
  • Principal Component Analysis (PCA)Why: When the aim is not to interpret latent dimensions but to reduce the number of variables while preserving variance.
  • Cronbach's AlphaWhy: When the structure is already settled and the question is the internal consistency of a subscale.

Frequently asked questions

How many factors should I keep?
Leave the factor count empty and eigenvalues greater than one are counted, which is the Kaiser rule. On its own that rule tends to retain more factors than warranted, so the elbow in the scree plot, the interpretability of the dimensions and the theoretical sense of the subscales all feed into the decision. Parallel analysis is offered as a separate method and supports the choice by comparing your eigenvalues with those from random data. Once you have decided, you can enter the factor count by hand and run the analysis again.
Does this prove my dimensional structure, or do I still need a confirmatory analysis?
An exploratory solution proposes a structure; it offers no test that the structure is right. Testing it requires confirmatory factor analysis, where item-to-factor assignments are declared in advance and fit indices are computed. The preferred route is to split the sample or collect a second one rather than stacking both analyses on the same data: confirming a structure on the very data that produced it yields misleadingly good fit.
What sample size do I need?
The engine requires at least 5 complete observations per item and will not run below that, reporting how many rows it found. That is only the floor at which the computation goes through. A widely cited rule of thumb in the literature is 10 observations per item with a total of 200 or more respondents; where inter-item correlations are strong and communalities high, stable solutions are known to emerge from smaller samples.
Which rotation should I choose?
The default varimax rotation treats factors as independent and aims to make each item stand out on a single dimension, which usually yields the most readable solution. Oblique rotation is preferred where dimensions are theoretically related, but in this method the quartimax option amounts to the varimax solution and promax is applied as a power transformation of it, with no inter-factor correlations reported. If the relationship between the dimensions is itself central to your question, confirmatory factor analysis is the better address.

References

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.