Structural equation

Mediation Analysis

Estimates how much of a predictor's effect on an outcome travels through an intervening third variable and tests that share with a Sobel test.

Method summary

Mediation analysis pursues the question of how: through what mechanism the link between two variables operates. The variable proposed as that mechanism is called the mediator, and the model splits into three parts, the link from predictor to mediator, the link from mediator to outcome, and the direct link left from predictor to outcome. The indirect effect is the product of the first two links, and interest usually rests on that product because the theoretical claim is usually about the mechanism. YouReply Analyze estimates the three paths in one model, computes the indirect and total effects, gives the share of the total that runs through the mediator, and reports a Sobel test for the indirect effect. The model is built around a single mediator, and control variables can be added if you want them.

Which research questions does it answer?

  • Does the effect of training hours on job performance travel through perceived self-efficacy?
  • What part does perceived trust play in the link between brand familiarity and purchase intention?
  • Does the intensity of remote work reduce commitment through feelings of isolation?
  • Does perceived risk mediate the effect of price sensitivity on channel preference?

When should you use it?

  • When the theoretical claim is about a mechanism, that is, when the question is why the predictor affects the outcome.
  • When the mediator is a process that can be argued to occur after the predictor and before the outcome; the ordering is established by design, not by the data.
  • When all three variables are measured continuously and the mediator is represented by one variable.
  • When seeing the magnitude of the indirect share in a screening pass is enough; a firm claim of significance calls for a tool that computes a bootstrapped interval.

Required variable types

  • Predictor: a continuous numeric column.
  • Mediator: a continuous numeric column, playing both the outcome and the predictor role within the model.
  • Outcome: a continuous numeric column.
  • Control variables: optional, one or more continuous numeric columns. When added, all three paths are estimated with their share held in the model.

Key assumptions

The causal ordering is established by design
A mediation claim presumes a sequence in time: predictor first, mediator next, outcome last. When all three are measured at the same moment of the same questionnaire, that sequence is absent from the data and comes from theory alone. Swapping the mediator and the outcome in the same dataset can produce equally significant results.
Adequate sample size
Because the indirect effect is a product of two estimates, its uncertainty exceeds that of a single coefficient, and in small samples the Sobel test tends to miss a real indirect effect.
Linear relationships
All three paths are treated as linear. If the link between predictor and mediator, or between mediator and outcome, curves, the product forming the indirect effect comes out biased.
No unmeasured common cause
If an unmeasured variable influences both the mediator and the outcome, the second path is biased and the indirect effect looks larger than it is. Since the mediator's value is never randomly assigned, this risk is always present in cross-sectional designs.
The mediator is measured reliably
A mediator carrying measurement error understates both of its paths and lowers the indirect effect. Mediators measured with a single item are fragile in this respect.

How YouReply checks these assumptions

  • The causal ordering is established by design: The panel does not check this assumption automatically; the researcher evaluates it.
  • Adequate sample size: The engine requires at least twenty complete observations and stops below that. The number is a technical floor rather than a criterion of adequacy; the guidance in the literature for indirect effects is far higher.
  • Linear relationships: The form of the relationships is not tested. Examine any pair you have doubts about with the correlation or regression methods separately.
  • No unmeasured common cause: The panel cannot detect the presence of such a variable. Columns added to the optional control variables field hold the share of those columns only.
  • The mediator is measured reliably: The model is built from observed variables, so measurement error does not enter it. For a multi-item mediator scale you should evaluate internal consistency separately.

How the analysis is run

  1. 1Upload the file holding the three variables and any control variables.
  2. 2On the Data tab confirm that the columns resolved to numbers and that reverse-coded items have been turned round, since a reverse-coded mediator flips the sign of a path.
  3. 3On the Variable tab set the measurement level of the three columns to scale, declare your missing-value codes and review the value labels.
  4. 4Choose mediation analysis on the Analysis tab; the data requirements box shows which columns meet the continuous condition.
  5. 5In the parameter form pick the predictor, the mediator and the outcome separately, adding control variables to the optional field if needed. There is no canvas on which to draw the model; the three roles are set through the form fields.
  6. 6Run the analysis. The a and b paths, the direct path, the indirect and total effects, the share running through the mediator and the Sobel test open as collapsible sections.
  7. 7Read the full or partial mediation label in the output as a note rather than a verdict, and reach the verdict yourself from the reported coefficients.
  8. 8No method-specific chart is drawn; the tables export to Excel. The computation credits card names the call and version, and the method requires the semopy library on the server.

Statistics and tables produced

The a path
The estimate from predictor to mediator with its standard error. This is where you see whether the mediator relates to the predictor.
The b path
The estimate from mediator to outcome with its standard error. It is estimated with the predictor's share held in the model.
The direct path
The link that remains from predictor to outcome. It comes back as an estimate ONLY, with no p value or standard error beside it.
Indirect effect
The product of the a and b paths, which is the share reaching the outcome from the predictor by way of the mediator.
Total effect and the mediated share
The sum of the direct and indirect effects, plus the ratio of the indirect effect to that sum. The ratio expresses how much of the relationship runs through the mediator.
Sobel test
The test of the indirect effect: a z value, a p value and a significance flag. It is the only instrument in this output that tests the indirect effect.
Mediation type note
A statement of full or partial mediation. The label is produced by a simplified rule inspecting the size of the direct coefficient and is not a statistical verdict.

Effect size and confidence intervals

Indirect effect
The product of a and b says how much the outcome moves with a one-unit change in the predictor by way of the mediator. The figure is unstandardised, so comparing it directly with indirect effects from other samples misleads; convert the three variables to standard scores before the analysis if you need that comparison.
Share running through the mediator
The ratio of the indirect effect to the total effect offers an intuitive measure of the mechanism's weight. When the direct and indirect effects carry opposite signs, the ratio can exceed one or lose meaning altogether, so report it together with the signs of both parts.
Magnitude of the a and b paths
The indirect effect does not reveal which path is the weak one: 0.40 times 0.10 and 0.20 times 0.20 give the same result. Reading the two paths separately shows at which end of the mechanism the fragility sits.

This method returns no confidence interval for any figure, and no bootstrapped interval is computed for the indirect effect. The only information you have about its uncertainty is the z and p value of the Sobel test. That test assumes the product is normally distributed, whereas the product of two estimates is usually skewed, which is why it tends to miss a real indirect effect in small and moderate samples. Current practice prefers a bias-corrected bootstrap interval. Treat the result in this output as a SCREENING test and describe it as such in your write-up: state that the indirect effect was tested with a Sobel test and that no bootstrapped interval was computed.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
Representative example: an organisation asks whether the effect of training hours on performance ratings travels through perceived self-efficacy. All three variables were measured in the same period, 248 complete observations were used and no control variables were added.
Variables
Predictor: training hours received during the year (continuous) · Mediator: perceived self-efficacy (mean score on a 1-7 scale, continuous) · Outcome: manager-assigned performance rating (0-100, continuous)
Example result
The a path was 0.021 (SE = 0.004), the b path 6.42 (SE = 1.18) and the direct path 0.043. The indirect effect came to 0.021 times 6.42 = 0.135, the total effect to 0.178, and the indirect share to 0.76 of the total. The Sobel test gave z = 3.94, p < .001 with a positive significance flag, and the output attached a partial mediation note to this pattern.
Interpretation
Both paths are strong in their own right: training hours raise self-efficacy, and self-efficacy relates clearly to the performance rating. The indirect share is about three quarters of the total effect, so most of the association between training and performance appears to operate through self-efficacy. The Sobel test supports that share, but it is a screening result: with no bootstrapped interval, the range around the indirect effect is unknown. A direct path of 0.043 is not by itself enough to decide anything, because no p value accompanies it, and the same caution applies to the partial mediation note, which comes from a simple rule looking at the size of that coefficient. Since the three variables were measured at once, the ordering is theoretical: an account in which higher self-efficacy leads people to seek out more training fits the same data.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

Variable examined

General customer survey (synthetic)

A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.

Rows
300
Columns
respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: Does the effect of service perception on loyalty work through satisfaction?

Independent variable
service_score
Mediator (M)
satisfaction
Dependent variable
loyalty
Direct effect
-0.126
Indirect effect
0.134
Total effect
0.008
Ratio of the indirect to the direct effect
17.34
Valid observations
300
independent_var
service_score
mediator
satisfaction
dependent_var
loyalty

Paths

Paths
RowEstimateStandard error
a_path0.4620.048
b_path0.2910.116
c_prime-0.126-

Sobel test

z statistic
2.42
p value
0.015
Significant
Yes

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914

How to report the result

The indirect effect of training hours on performance ratings through perceived self-efficacy was estimated at 0.135 (a = 0.021, SE = 0.004; b = 6.42, SE = 1.18) and tested with a Sobel test, z = 3.94, p < .001; the direct effect was 0.043 and the total effect 0.178, with the indirect part accounting for 0.76 of the total. No bootstrapped confidence interval was computed for the indirect effect.

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • The indirect effect is tested with a Sobel test alone and no bootstrapped confidence interval is computed; since current practice prefers an interval, your write-up has to state that the result is a screening one.
  • No p value or standard error is returned for the direct path, so this output cannot tell you whether the direct effect is significant.
  • The full or partial mediation label comes from a simplified rule inspecting the size of the direct coefficient rather than its significance, so base the verdict on the reported coefficients instead of the label.
  • Only a single mediator is supported; parallel mediators and serial mediators following one another cannot be estimated here.
  • No confidence interval is given for any figure.
  • With cross-sectional data mediation establishes no causal claim, and because the mediator is never assigned, an unmeasured variable influencing both mediator and outcome can inflate the indirect effect.

What to use when the assumptions are not met

  • Path AnalysisWhy: When there are several mediators or several outcomes; it estimates all the paths in one model and returns fit indices.
  • Moderation AnalysisWhy: When the question is about conditions rather than mechanism, that is, in which subgroups the effect strengthens.
  • Linear Regression (OLS)Why: When the share of the predictors in the outcome is the question and no mediation claim is being made; coefficients are reported with their intervals.
  • Structural Equation Modeling (SEM)Why: When the variables are measured with multi-item scales; latent constructs bring measurement error inside the model.

Frequently asked questions

Why is a bootstrapped interval preferred for an indirect effect?
The indirect effect is the product of two estimates, and the product of two normally distributed numbers is not itself normal: its distribution is skewed. The Sobel test rests on precisely that assumption of normality for the product. As a result it tends to miss a genuine indirect effect in small and moderate samples, which is to say it lacks power. Bootstrapping instead draws thousands of resamples from the data, computes the product in each, and reads an interval straight off the resulting distribution, so the skewness is accommodated without being assumed away. This output performs no bootstrapping, so the Sobel result is the only test you have. A significant result is a supportive sign for the indirect effect, while a non-significant one may reflect the low power of the test rather than the absence of an effect.
Full or partial mediation, how should I decide?
Do not decide from the note in the output. The label is produced by a simplified rule that looks at the size of the direct coefficient, whereas the distinction in the literature turns on whether the direct effect is significant, and no p value for the direct path appears here. Treat the note as a summary and ground your verdict in the reported coefficients instead, weighing the magnitudes of a, b and the direct path together with the mediated share of the total and the Sobel result. There is a broader point as well: the current literature leans less and less on the full-versus-partial dichotomy, because the classification is sensitive to sample size. Conveying the magnitude of the indirect effect and its uncertainty is more informative than picking one of the two labels.
All three variables came from the same questionnaire. Can I still speak of mediation?
You can estimate the indirect effect, but cross-sectional data carries no claim about a causal mechanism. The model presumes a sequence in time, and measuring everything at once leaves that sequence out of the data. Rebuilding the model with the mediator and the outcome swapped often yields an equally significant indirect effect from the same three columns, and the output says nothing about which ordering is right. When presenting the finding, state that the ordering comes from theory and that the data is silent about direction. Strengthening a claim about mechanism calls for manipulating the predictor experimentally, or at least measuring mediator and outcome at different times.
If one of the a or b paths is weak, should I still report the indirect effect?
Because the indirect effect is a product, it shrinks as either path approaches zero, yet its magnitude hides which path is the weak one. Report it alongside the a and b values rather than on its own. The two cases are worth separating: if a is strong and b is weak, the predictor moves the mediator but the mediator does not relate to the outcome, meaning the mechanism you proposed does not explain the outcome. If b is strong and a is weak, the mediator relates to the outcome but the predictor is not moving it. Those two situations carry different theoretical implications, and giving only the product keeps the distinction from the reader.

References

  • Hayes, A. F. (2022). Introduction to Mediation, Moderation, and Conditional Process Analysis
  • Baron, R. M., & Kenny, D. A. (1986). The Moderator-Mediator Variable Distinction in Social Psychological Research
  • Preacher, K. J., & Hayes, A. F. (2008). Asymptotic and Resampling Strategies for Assessing and Comparing Indirect Effects in Multiple Mediator Models
  • semopy documentation

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.