Hypothesis testing is a method for deciding whether sample data give enough evidence against a default claim, called the null hypothesis. You compute a test statistic, find its p-value, and compare that with a significance level chosen in advance.
This guide covers the steps of hypothesis testing, what p-values mean and do not mean, how to choose a test, errors and power, and a worked example with code. It ends with how to report results and how STEM Donkey can help.
The Five Steps of a Hypothesis Test
- State the null hypothesis H₀ and the alternative hypothesis H₁.
- Choose the significance level α, usually 0.05, before looking at the result.
- Pick a test that fits your data type and design, and check its assumptions.
- Calculate the test statistic and its p-value.
- Decide: if p ≤ α, reject H₀; otherwise, do not reject it. Then interpret the result in context, with an effect size.
Writing the Null and Alternative Hypotheses
The null hypothesis is the "no effect" or "no difference" claim. The alternative is what you are looking for evidence of.
| Research question | H₀ | H₁ |
|---|---|---|
| Does a fertiliser change mean plant height? | μtreated = μcontrol | μtreated ≠ μcontrol (two-tailed) |
| Does a new alloy have higher mean strength than 250 MPa? | μ = 250 MPa | μ > 250 MPa (one-tailed) |
| Is smoking status associated with disease? | The variables are independent | The variables are associated |
Hypotheses are about population parameters such as μ, not sample statistics such as x̄. Choose one-tailed or two-tailed before collecting data, and justify a one-tailed test from the research question, not from the result.
What a P-Value Means
The p-value is the probability of getting a test statistic at least as extreme as the one observed, assuming the null hypothesis is true. A small p-value means the data would be unusual if H₀ were true.
It helps to be clear about what a p-value is not:
- It is not the probability that H₀ is true.
- It is not the probability that the result happened "by chance".
- It does not measure the size or importance of an effect.
- A p-value above 0.05 does not prove H₀; it means the evidence was not strong enough to reject it.
With very large samples, tiny and unimportant differences can give small p-values. That is why every result should be reported with an effect size and, ideally, a confidence interval.
Choosing the Right Test
The right test depends on the type of outcome, the number of groups and whether measurements are paired.
| Situation | Parametric test | Non-parametric alternative |
|---|---|---|
| One sample mean against a value | One-sample t-test | Wilcoxon signed-rank test |
| Two independent groups, continuous outcome | Independent-samples t-test (Welch's version if variances differ) | Mann-Whitney U test |
| Same subjects measured twice | Paired t-test | Wilcoxon signed-rank test |
| Three or more independent groups | One-way ANOVA | Kruskal-Wallis test |
| Two categorical variables | Chi-square test of independence | Fisher's exact test for small expected counts |
| Relationship between two continuous variables | Pearson correlation | Spearman rank correlation |
Parametric tests assume conditions such as approximately normal data or residuals and, for some tests, equal variances. Check them with plots and, where useful, formal tests, and say what you found.
A Worked Example: One-Sample T-Test
Question. A water supplier states that mean nitrate concentration is 5.0 mg/L. A student takes 16 samples and finds a mean of 5.4 mg/L with a standard deviation of 0.8 mg/L. Is there evidence, at α = 0.05, that the true mean differs from 5.0 mg/L?
Step 1. H₀: μ = 5.0 mg/L. H₁: μ ≠ 5.0 mg/L (two-tailed).
Step 2. α = 0.05.
Step 3. One-sample t-test, assuming the samples are independent and the concentrations roughly normal.
Step 4. Standard error = s/√n = 0.8/√16 = 0.2 mg/L. Test statistic t = (x̄ − μ₀)/SE = (5.4 − 5.0)/0.2 = 2.0, with df = n − 1 = 15. The two-tailed critical value is t₀.₀₂₅,₁₅ ≈ 2.131, and the p-value is about 0.064.
Step 5. Since 0.064 > 0.05 (equivalently, 2.0 < 2.131), do not reject H₀. The data do not give sufficient evidence at the 5 percent level that mean nitrate differs from 5.0 mg/L.
Confidence interval. 95% CI = 5.4 ± 2.131 × 0.2 = 4.97 to 5.83 mg/L. It contains 5.0, which agrees with the test.
Effect size. Cohen's d = (5.4 − 5.0)/0.8 = 0.5, a medium effect by common rules of thumb. A larger sample might detect it, which is worth saying in the discussion.
The same test takes one line in software. In Python, with your measurements in a list:
from scipy import stats
result = stats.ttest_1samp(nitrate_mg_per_l, popmean=5.0)
print(result.statistic, result.pvalue)
In R, the equivalent is t.test(nitrate, mu = 5), which also prints the 95 percent confidence interval.
Type I Errors, Type II Errors and Power
| H₀ is true | H₀ is false | |
|---|---|---|
| Reject H₀ | Type I error (false positive), probability α | Correct decision, probability 1 − β (power) |
| Do not reject H₀ | Correct decision, probability 1 − α | Type II error (false negative), probability β |
Power is the probability of detecting a real effect. It rises with larger samples, larger effects, less variable data and a larger α. Many studies aim for a power of 0.8, and a power analysis before data collection tells you the sample size needed.
Lowering α reduces Type I errors but, with everything else fixed, increases Type II errors. The choice should reflect which mistake is more costly in your context.
Multiple Comparisons
Each test at α = 0.05 carries a 5 percent chance of a false positive when H₀ is true. Run 20 independent tests on data with no real effects and you should expect about one "significant" result.
- After a significant ANOVA, use a post hoc test such as Tukey's HSD rather than many separate t-tests.
- The Bonferroni correction divides α by the number of tests; it is simple but conservative.
- Decide your main comparisons in advance and report every test you ran.
How to Report a Hypothesis Test
A good report gives the test, the statistic with its degrees of freedom, the p-value, an effect size and a plain-language conclusion.
Example report sentence. A one-sample t-test found no significant difference between the mean nitrate concentration (M = 5.4 mg/L, SD = 0.8) and the stated value of 5.0 mg/L, t(15) = 2.00, p = .064, d = 0.50, 95% CI [4.97, 5.83].
This example follows APA conventions, which drop the leading zero for p-values. Other styles keep it, so check your course guide. Give exact p-values, such as p = .064, rather than only "p > .05", unless p is below .001.
Common Hypothesis Testing Mistakes
| Mistake | Fix |
|---|---|
| Saying "we accept H₀" | Say "we do not reject H₀"; absence of evidence is not proof |
| Writing hypotheses about sample means | Write them about population parameters such as μ |
| Choosing a one-tailed test after seeing the data | Decide the direction before analysis |
| Independent t-test on paired data | Use a paired t-test for before and after measurements on the same subjects |
| Ignoring assumptions | Check normality, independence and equal variances, or use an alternative test |
| Reporting p only | Add an effect size and a confidence interval |
Hypothesis testing rewards a steady, step-by-step approach. Follow the same five steps every time and most errors never appear.
How STEM Donkey Helps with Hypothesis Testing
Send your data, research question, course notes and the software you use, such as SPSS, R, Python or Excel. A writer prepares a custom analysis with clearly stated hypotheses, assumption checks, the test result, effect size, confidence interval and a written interpretation, plus commented code or output.
Your data is used exactly as provided and never altered to reach significance. Free revisions within the original scope are included for 14 days.
Want Your Hypothesis Test Done and Explained?
Send your data, research question and software. You get a custom analysis with hypotheses, assumption checks, the test result, effect size and a plain-language interpretation.
Order Your Statistics AssignmentFree revisions within scope for 14 days · Full refund if late · Written from scratch for your order
Frequently Asked Questions
The default claim of no effect, no difference or no association, which the test assesses evidence against.
If the null hypothesis were true, a result at least this extreme would occur less than 5 percent of the time. It does not mean there is a 95 percent chance the alternative is true.
A two-tailed test looks for a difference in either direction. A one-tailed test looks in one direction only, and must be justified before seeing the data.
The probability that a test detects a real effect of a given size. It depends on sample size, effect size, variability and α.
When assumptions such as normality clearly fail and cannot be fixed, especially with small samples or ordinal data.
No. Statistical significance says the effect is unlikely to be zero; the effect size and context tell you whether it matters.
For a two-tailed test at α = 0.05, H₀ is rejected exactly when the 95 percent confidence interval excludes the null value.
Yes. State your software in the brief and you receive the output, commented syntax or code, and the written interpretation.