Group comparison

One-Sample t-Test

Compares the mean of one continuous variable against a fixed reference value supplied from outside the data, and reports a confidence interval for the difference.

Method summary

There are no two groups in a one-sample t-test. On one side is the mean of the column you collected; on the other, a single number that comes from theory, from published norms, from a contract, or from an earlier wave. The test divides the gap between them by the standard error of the mean, with n - 1 degrees of freedom. Because the reference value is not estimated from the data, everything rests on how defensible that number is. YouReply Analyze returns the mean, standard deviation, standard error, sample size, t value, degrees of freedom, p value, the mean difference and a 95% confidence interval around that difference, which makes this the only method in the whole panel that produces an interval. Leave the reference value empty in the parameter form and it is treated as 0.

Which research questions does it answer?

  • Does our sample's mean job satisfaction differ from the norm published for the population?
  • Is the average call handling time below the 120 seconds promised in the service agreement?
  • On a five-point satisfaction scale, does the mean sit away from the scale midpoint of 3?
  • Does average weekly usage in this panel differ from the figure reported a year ago?

When should you use it?

  • There is no second group and no second measurement, only a single continuous column.
  • The value you are testing against comes from outside the study: a published norm, a legal threshold, a target, a scale midpoint.
  • The question reads 'is the mean different from this value', rather than 'which group scores higher'.
  • You need to report the size of the difference together with an interval, which no other method in the panel provides.

Required variable types

  • One continuous column: numeric type, with the 'scale' measurement level set on the Variable tab.
  • The reference value is not a column. It is a single number typed into the parameter form, expressed in the same units as the variable.
  • No grouping column is used; this method has no independent variable.

Key assumptions

The reference value comes from outside the data
If the comparison value is derived from the same sample, say by rounding its own mean, the test is checking the data against itself and the p value means nothing. State where the number came from when you report it.
Independence of observations
Each row should come from a different respondent, and one respondent's value must not influence another's. Repeat responses from the same person, or clustered sampling, make the standard error too small.
Approximate normality of the variable
Normality is expected of the measured variable itself. As n grows (roughly 30 and above) the central limit theorem pulls the sampling distribution of the mean towards normal and the test becomes more forgiving.
No single observation dragging the mean
Both the t value and the interval are built from the mean and the standard deviation, so one extreme value can move the whole result.

How YouReply checks these assumptions

  • The reference value comes from outside the data: The reference value is typed into the parameter form and defaults to 0 when left empty. The panel does not judge whether the number is a defensible benchmark, and on most scale totals 0 is not one.
  • Independence of observations: The panel does not check this assumption automatically; the researcher evaluates it.
  • Approximate normality of the variable: This method runs no normality test. Normality tests are available as a separate method, which is worth running first on small samples.
  • No single observation dragging the mean: The panel does not screen for outliers, and this method returns no median to compare against the mean. Review the column in the Data tab grid, and declare missing-value codes such as 99 or -1 on the Variable tab, because undeclared codes are averaged in as if they were real scores.

How the analysis is run

  1. 1Drop your CSV or XLSX file onto the panel; a single column holding the variable is all this method needs.
  2. 2On the Variable tab make sure the column is numeric with the right measurement level, and declare missing-value codes so they stay out of the mean.
  3. 3Decide on the reference value before you run anything, and write down where it comes from. Adjusting it after seeing the output turns the test into a search.
  4. 4Open the method picker on the Analysis tab. The data requirements box states that this method needs one numeric column and shows which of yours qualify.
  5. 5In the parameter form select the column, type the reference value (empty means 0) and switch to a one-tailed test if your hypothesis is directional.
  6. 6Run it, then read the mean difference together with its 95% interval. Export the tables to XLSX and the chart to PNG; the history tab keeps earlier runs, so runs against different reference values stay comparable.

Statistics and tables produced

Mean, standard deviation, standard error and sample size
The sample descriptives. Because the standard error is reported separately, you can see where the width of the interval comes from.
t statistic and degrees of freedom
Degrees of freedom are n - 1, and the sign of t tells you whether the mean sits above or below the reference value.
Significance (p) value and a significance flag
The p value for the direction you chose, plus a field stating whether the result counts as significant at the .05 level.
Mean difference
The gap between the sample mean and the reference value, in the units of the variable. This is what the interpretation rests on.
95% confidence interval for the mean difference
Lower and upper bounds in separate fields. The level is fixed at 95% and cannot be changed from the parameter form.
Cohen's d
The standardised effect size, obtained by dividing the mean difference by the standard deviation.
Mean chart
A bar chart of the mean is drawn above the tables and can be saved as a PNG.

Effect size and confidence intervals

Cohen's d
How far the mean sits from the reference value, measured in standard deviations: d = 0.50 means half a standard deviation away. The 0.20, 0.50 and 0.80 landmarks are disciplinary habit rather than law. The sign carries the direction, so a negative d means the mean falls below the reference value.

This method returns the lower and upper bound of a 95% confidence interval for the mean difference, and it is the only method in the panel that returns an interval at all. The level is hardcoded: there is no 90% or 99% option. The interval is built around the difference rather than around the mean itself, so to get an interval for the mean you simply add the reference value to both bounds. In a two-tailed test, an interval excluding zero and a p value below .05 lead to the same decision, but the interval says more: it shows which differences in size the data are compatible with. No interval is returned for the effect size.

Example research question and example result

The numbers below are a representative example, not data from a real study or a real user.

Research question
Illustrative example: a study of teachers measured job satisfaction (0-100) and compared the result with a previously published norm of 65.0. Does the sample depart from the norm?
Variables
Variable: job satisfaction total score (continuous, 0-100) · Reference value: 65.0, the mean reported in last year's national report
Example result
n = 210, mean 61.8 (SD = 14.2, SE = 0.98). The mean difference is -3.2: t(209) = -3.27, p = .001, 95% CI [-5.13, -1.27], d = -0.23.
Interpretation
The sample mean falls 3.2 points below the norm, and since the interval excludes zero the gap is not explained by sampling variation. The width matters as much as the sign: the data are compatible with a true difference of roughly 1.3 to 5.1 points, so the gap is established but its magnitude is not pinned down. With d = -0.23 the effect is small. The report should also note that the norm was collected in a different year, most likely from a different sampling frame, so this comparison on its own is not evidence of a decline.

Real output on a sample dataset

The results below were produced by the analysis engine from this data file. Changing the variable changes the research question as well; every run was computed in advance, so the page sends no request to the engine.

Variable examined

General customer survey (synthetic)

A wide survey of three hundred respondents: two and three category grouping variables, continuous measures, a five point ordinal scale, a binary purchase outcome, a four category brand choice, three repeated measurements, paired binary questions, three raters, four price questions and deliberately empty cells.

Rows
300
Columns
respondent_id, gender, education, region, age, income, satisfaction, service_score, price_score, quality_score, loyalty, nps_score, purchased, brand_choice, satisfaction_level, pre_score, post_score, measure_1, measure_2, measure_3, use_before, use_after, use_followup, rater_1, rater_2, rater_3, price_too_cheap, price_cheap, price_expensive, price_too_expensive, feedback_score, followup_rating
Download the dataset as CSV

The data is synthetic: it comes from a fixed random seed, not from a real study. The values below were produced by the analysis engine from this file, so uploading the same file to the panel gives the same results.

Research question: Does mean satisfaction differ from 4, the midpoint of the seven point scale?

Analyzed column
satisfaction
Hypothesized value
4
column
satisfaction
test_value
4
Mean
5.24
Standard deviation
0.960
Standard error
0.055
Valid observations
300
t statistic
22.44
Degrees of freedom
299
p value
< 0.001
Mean difference
1.24
Confidence interval lower
1.13
Confidence interval upper
1.35
Cohen's d
1.30
Significant
Yes

Computation credits: scipy 1.18.0 · statsmodels 0.14.6 · scikit-learn 1.9.0 · numpy 2.5.1 · pandas 3.0.5 · semopy 2.3.11 · 89fc29a · Data seed: 20260914

How to report the result

Teachers' job satisfaction (M = 61.8, SD = 14.2) was significantly below the norm value of 65.0, t(209) = -3.27, p = .001, mean difference = -3.2, 95% CI [-5.13, -1.27], d = -0.23.

An example sentence close to APA style; the numbers are representative.

When you should not use it

  • Not a way to compare two groups; once your data contain a grouping column the question becomes a between-groups one and calls for the independent samples t-test.
  • Not for two measurements of the same people, where the paired t-test works on the difference scores instead.
  • When the reference value is taken from the same data the result is circular; the test only means something with an externally supplied number.
  • An empty reference value is read as 0, and on scale totals that comparison is almost always trivially significant and tells you nothing.
  • With a small sample and a clearly skewed distribution the interval looks narrower than it should, and a rank-based test is the better option.
  • On a single Likert item the mean and standard deviation are contested quantities, which makes the comparison against a benchmark equally contested.

What to use when the assumptions are not met

  • Wilcoxon Signed-Rank TestWhy: When the distribution is skewed or the variable is ordinal; it compares a single measurement against a reference value through the median.
  • Independent Samples t-TestWhy: When a two-category grouping column exists and the question turns into a comparison between groups.
  • Paired Samples t-TestWhy: When the comparison is against a second measurement of the same people rather than a fixed number.
  • Descriptive StatisticsWhy: When there is no benchmark to test and the goal is simply to characterise the distribution without a hypothesis test.

Frequently asked questions

What should I enter as the reference value, and what if I leave it empty?
An empty field is treated as 0. That can be a sensible benchmark for change scores, but not for a score on a 0 to 100 scale: you already know the mean is not zero, the test comes out significant almost by construction, and nothing is learned. Fix the value beforehand and make it defensible: a published norm, a contractual threshold, a scale midpoint, or the mean of a previous wave.
Can I ask for a 99% interval?
No. The interval is computed at 95% and that level is fixed in the engine, not offered as an option in the parameter form. If you need a different level, the mean difference and the standard error in the result table let you work it out yourself.
What if the normality assumption does not hold?
The panel does not test normality in this method, so the judgement is yours. With a few hundred observations, moderate skew does little harm to the sampling distribution of the mean. On small samples, run the normality tests method first; if the skew is pronounced or there are extreme values, move to the Wilcoxon signed-rank test. Note that doing so also costs you the interval, since this is the only method that reports one.
Do the p value and the interval tell me the same thing?
They lead to the same decision but carry different information. In a two-tailed test, an interval that excludes zero corresponds to p below .05. Beyond that, the interval makes the size of the difference discussable: if both bounds sit inside a range you would call practically irrelevant, a significant p value still describes a difference nobody will act on. Report both.

References

  • Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics
  • Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences
  • Cumming, G. (2012). Understanding The New Statistics: Effect Sizes, Confidence Intervals, and Meta-Analysis
  • scipy.stats.ttest_1samp documentation

Try it with your own data

The free plan includes 50 analysis runs a month and needs no card.