Maths Guide

Hypothesis Testing: Steps, P-Values and Examples

The five steps of a hypothesis test, what a p-value really means, how to choose the right test, Type I and Type II errors, and a fully worked t-test example.

Updated October 2026 · 7 min read

Hypothesis testing is a method for deciding whether sample data give enough evidence against a default claim, called the null hypothesis. You compute a test statistic, find its p-value, and compare that with a significance level chosen in advance.

This guide covers the steps of hypothesis testing, what p-values mean and do not mean, how to choose a test, errors and power, and a worked example with code. It ends with how to report results and how STEM Donkey can help.

The Five Steps of a Hypothesis Test

  1. State the null hypothesis H₀ and the alternative hypothesis H₁.
  2. Choose the significance level α, usually 0.05, before looking at the result.
  3. Pick a test that fits your data type and design, and check its assumptions.
  4. Calculate the test statistic and its p-value.
  5. Decide: if p ≤ α, reject H₀; otherwise, do not reject it. Then interpret the result in context, with an effect size.

Writing the Null and Alternative Hypotheses

The null hypothesis is the "no effect" or "no difference" claim. The alternative is what you are looking for evidence of.

Research questionH₀H₁
Does a fertiliser change mean plant height?μtreated = μcontrolμtreated ≠ μcontrol (two-tailed)
Does a new alloy have higher mean strength than 250 MPa?μ = 250 MPaμ > 250 MPa (one-tailed)
Is smoking status associated with disease?The variables are independentThe variables are associated

Hypotheses are about population parameters such as μ, not sample statistics such as x̄. Choose one-tailed or two-tailed before collecting data, and justify a one-tailed test from the research question, not from the result.

What a P-Value Means

The p-value is the probability of getting a test statistic at least as extreme as the one observed, assuming the null hypothesis is true. A small p-value means the data would be unusual if H₀ were true.

It helps to be clear about what a p-value is not:

  • It is not the probability that H₀ is true.
  • It is not the probability that the result happened "by chance".
  • It does not measure the size or importance of an effect.
  • A p-value above 0.05 does not prove H₀; it means the evidence was not strong enough to reject it.

With very large samples, tiny and unimportant differences can give small p-values. That is why every result should be reported with an effect size and, ideally, a confidence interval.

Choosing the Right Test

The right test depends on the type of outcome, the number of groups and whether measurements are paired.

SituationParametric testNon-parametric alternative
One sample mean against a valueOne-sample t-testWilcoxon signed-rank test
Two independent groups, continuous outcomeIndependent-samples t-test (Welch's version if variances differ)Mann-Whitney U test
Same subjects measured twicePaired t-testWilcoxon signed-rank test
Three or more independent groupsOne-way ANOVAKruskal-Wallis test
Two categorical variablesChi-square test of independenceFisher's exact test for small expected counts
Relationship between two continuous variablesPearson correlationSpearman rank correlation

Parametric tests assume conditions such as approximately normal data or residuals and, for some tests, equal variances. Check them with plots and, where useful, formal tests, and say what you found.

A Worked Example: One-Sample T-Test

Question. A water supplier states that mean nitrate concentration is 5.0 mg/L. A student takes 16 samples and finds a mean of 5.4 mg/L with a standard deviation of 0.8 mg/L. Is there evidence, at α = 0.05, that the true mean differs from 5.0 mg/L?

Step 1. H₀: μ = 5.0 mg/L. H₁: μ ≠ 5.0 mg/L (two-tailed).

Step 2. α = 0.05.

Step 3. One-sample t-test, assuming the samples are independent and the concentrations roughly normal.

Step 4. Standard error = s/√n = 0.8/√16 = 0.2 mg/L. Test statistic t = (x̄ − μ₀)/SE = (5.4 − 5.0)/0.2 = 2.0, with df = n − 1 = 15. The two-tailed critical value is t₀.₀₂₅,₁₅ ≈ 2.131, and the p-value is about 0.064.

Step 5. Since 0.064 > 0.05 (equivalently, 2.0 < 2.131), do not reject H₀. The data do not give sufficient evidence at the 5 percent level that mean nitrate differs from 5.0 mg/L.

Confidence interval. 95% CI = 5.4 ± 2.131 × 0.2 = 4.97 to 5.83 mg/L. It contains 5.0, which agrees with the test.

Effect size. Cohen's d = (5.4 − 5.0)/0.8 = 0.5, a medium effect by common rules of thumb. A larger sample might detect it, which is worth saying in the discussion.

The same test takes one line in software. In Python, with your measurements in a list:

from scipy import stats

result = stats.ttest_1samp(nitrate_mg_per_l, popmean=5.0)
print(result.statistic, result.pvalue)

In R, the equivalent is t.test(nitrate, mu = 5), which also prints the 95 percent confidence interval.

Type I Errors, Type II Errors and Power

H₀ is trueH₀ is false
Reject H₀Type I error (false positive), probability αCorrect decision, probability 1 − β (power)
Do not reject H₀Correct decision, probability 1 − αType II error (false negative), probability β

Power is the probability of detecting a real effect. It rises with larger samples, larger effects, less variable data and a larger α. Many studies aim for a power of 0.8, and a power analysis before data collection tells you the sample size needed.

Lowering α reduces Type I errors but, with everything else fixed, increases Type II errors. The choice should reflect which mistake is more costly in your context.

Multiple Comparisons

Each test at α = 0.05 carries a 5 percent chance of a false positive when H₀ is true. Run 20 independent tests on data with no real effects and you should expect about one "significant" result.

  • After a significant ANOVA, use a post hoc test such as Tukey's HSD rather than many separate t-tests.
  • The Bonferroni correction divides α by the number of tests; it is simple but conservative.
  • Decide your main comparisons in advance and report every test you ran.

How to Report a Hypothesis Test

A good report gives the test, the statistic with its degrees of freedom, the p-value, an effect size and a plain-language conclusion.

Example report sentence. A one-sample t-test found no significant difference between the mean nitrate concentration (M = 5.4 mg/L, SD = 0.8) and the stated value of 5.0 mg/L, t(15) = 2.00, p = .064, d = 0.50, 95% CI [4.97, 5.83].

This example follows APA conventions, which drop the leading zero for p-values. Other styles keep it, so check your course guide. Give exact p-values, such as p = .064, rather than only "p > .05", unless p is below .001.

Common Hypothesis Testing Mistakes

MistakeFix
Saying "we accept H₀"Say "we do not reject H₀"; absence of evidence is not proof
Writing hypotheses about sample meansWrite them about population parameters such as μ
Choosing a one-tailed test after seeing the dataDecide the direction before analysis
Independent t-test on paired dataUse a paired t-test for before and after measurements on the same subjects
Ignoring assumptionsCheck normality, independence and equal variances, or use an alternative test
Reporting p onlyAdd an effect size and a confidence interval

Hypothesis testing rewards a steady, step-by-step approach. Follow the same five steps every time and most errors never appear.

How STEM Donkey Helps with Hypothesis Testing

Send your data, research question, course notes and the software you use, such as SPSS, R, Python or Excel. A writer prepares a custom analysis with clearly stated hypotheses, assumption checks, the test result, effect size, confidence interval and a written interpretation, plus commented code or output.

Your data is used exactly as provided and never altered to reach significance. Free revisions within the original scope are included for 14 days.

Want Your Hypothesis Test Done and Explained?

Send your data, research question and software. You get a custom analysis with hypotheses, assumption checks, the test result, effect size and a plain-language interpretation.

Order Your Statistics Assignment

Free revisions within scope for 14 days · Full refund if late · Written from scratch for your order

Frequently Asked Questions

What is a null hypothesis?

The default claim of no effect, no difference or no association, which the test assesses evidence against.

What does p < 0.05 mean?

If the null hypothesis were true, a result at least this extreme would occur less than 5 percent of the time. It does not mean there is a 95 percent chance the alternative is true.

What is the difference between a one-tailed and a two-tailed test?

A two-tailed test looks for a difference in either direction. A one-tailed test looks in one direction only, and must be justified before seeing the data.

What is statistical power?

The probability that a test detects a real effect of a given size. It depends on sample size, effect size, variability and α.

When should I use a non-parametric test?

When assumptions such as normality clearly fail and cannot be fixed, especially with small samples or ordinal data.

Is a significant result always important?

No. Statistical significance says the effect is unlikely to be zero; the effect size and context tell you whether it matters.

What is the link between confidence intervals and hypothesis tests?

For a two-tailed test at α = 0.05, H₀ is rejected exactly when the 95 percent confidence interval excludes the null value.

Can you run the analysis in SPSS or R?

Yes. State your software in the brief and you receive the output, commented syntax or code, and the written interpretation.