t-test calculator
Paste raw data or enter n, mean and standard deviation. See t, the degrees of freedom, the p-value, the confidence interval of the mean difference and Cohen's d for a one-sample, independent two-sample (Welch or pooled) or paired t-test, with a plot of the distribution and the critical region.
Separate values with new lines, spaces, tabs or semicolons (a column copied from Excel works). Use a point or a comma as the decimal mark; do not type thousands separators. Parts that are not numbers are skipped.
The calculation runs in your browser; the data you paste is not sent anywhere.
Usually 0 (no difference).
Paste data or load the sample.
Would you like to make sure your experiments and A/B comparisons are designed with the right sample and the right test? Let us look at it together.
Request a call01
How to use
A
Choose the type of test (one sample, two independent groups or paired), then paste raw data or enter n, mean and standard deviation.
B
Choose the alternative hypothesis (two-sided or one-sided), the level α and, for two groups, Welch or pooled variance.
C
Read t, the degrees of freedom, the p-value, the confidence interval of the mean difference and Cohen's d; in the chart see where the observed t falls relative to the critical region, and copy the result.
02
Which t-test should you choose?
A one-sample t-test compares the mean of one group with a known or target value: is the average fill of a filling machine equal to 500 g? An independent two-sample t-test compares the means in two unrelated groups: is the cycle time of line A different from that of line B?
A paired t-test is for two measurements of the same units (energy use before and after maintenance, the time the same users take on two interfaces). Here the difference of each pair is taken and the question is whether the mean of those differences differs from zero. Applying an independent-samples test to paired data wastes the power that comes from the pairing; applying a paired test to independent data gives a wrong result.
03
Welch and pooled t-test
The pooled (Student) test assumes the two groups have equal variance and estimates one common variance; its degrees of freedom are n₁ + n₂ − 2. The Welch test does not make that assumption: each group's variance is divided by its own sample size and the degrees of freedom come out as a fractional number from the Welch–Satterthwaite formula.
When variances are equal Welch gives almost the same result; when they are not (especially if group sizes differ too) the pooled test can distort the type I error rate. That is why Welch is the default. The pooled option is there to reproduce textbook examples or to compare with an older report.
04
p-value, confidence interval and effect size
The p-value is the probability of seeing a t as extreme as the observed one, or more, if the null hypothesis were true. A small p says the data are inconsistent with H₀; it does not say the difference is large or important. The confidence interval of the mean difference lets you look at the size and the uncertainty of the difference directly: if the interval does not contain zero, the two-sided test rejects H₀ at the same α.
Cohen's d expresses the difference in units of standard deviation; it is unit-free and comparable across studies. Common rough thresholds: 0.2 small, 0.5 medium, 0.8 large; these vary by field. In small samples d is slightly inflated, and Hedges' g corrects this.
FAQ
- Should I use the Welch or the pooled t-test?
- Use Welch by default. When variances are equal it gives a result very close to the pooled test, and when they are not it stays correct where the pooled test breaks down. Choose the pooled test only if you have a reason to know the equal variances in advance or want to reproduce a source.
- Should I run a one-sided or a two-sided test?
- Use one-sided only if a difference in one direction alone matters to you and you decided that before looking at the data. Choosing the direction after seeing the result makes the p-value artificially small. If unsure, use two-sided; the one-sided p-value is half the two-sided one when the result lies in the stated direction.
- What if the sample is small or the data are not normal?
- With a small sample the t-test relies on the data being approximately normal; skewness and outliers distort the p-value and the interval. As the sample grows (roughly beyond 30 per group, more if the skew is strong) the distribution of the mean approaches normal and the test becomes more robust. With very small samples treat the result cautiously, look at the plot and the data, and use a transformation or a rank-based test if needed.
- If p ≥ 0.05, can I say there is no difference?
- No. Failing to reject does not show that there is no difference; the sample may not be large enough to detect it. Look at the confidence interval: if it contains both zero and practically important differences, the data are inconclusive. Showing equivalence takes a separate method (TOST).
- How large must Cohen's d be to count as a large difference?
- A rough guide: 0.2 small, 0.5 medium, 0.8 large. These thresholds are Cohen's suggestion and vary by field; in production even a very small d can be large in money terms. Also judge the effect in your own unit (the interval).
- Can I enter summary values instead of raw data?
- Yes. In the "Summary values" mode enter n, the mean and the sample standard deviation (with n−1) for each group; for the paired test enter the mean and standard deviation of the differences. With summary input the skewness check cannot be done.
Let us test your comparisons properly
Let us design together how to show whether two conditions in your production, sales or field data really differ, with which sample and which test.