ANOVA Calculator
Run one-way ANOVA on raw data or on n, mean, and SD per group. Get the full ANOVA table, step-by-step working, effect sizes, a Levene/Brown-Forsythe variance check, optional Welch ANOVA, and Tukey HSD when the overall test is significant.
Balanced follow-up for every pairwise difference: Tukey HSD calculator. Two factors at once: two-way ANOVA calculator. Nonparametric alternative: Kruskal-Wallis test.
Numbers separated by commas, spaces, or new lines (at least 2 per group).
Numbers separated by commas, spaces, or new lines (at least 2 per group).
Numbers separated by commas, spaces, or new lines (at least 2 per group).
Related Calculators
F-Table (F-Distribution Critical Values)
Look up F critical values by numerator and denominator degrees of freedom for α = 0.10, 0.05, 0.025, and 0.01.
T-Test Calculator
Compare sample means with common t-test workflows and interpretable outputs.
Chi-Square Calculator
Test observed counts against expected counts with a chi-square goodness-of-fit test.
Learn More
One-Way ANOVA Explained
How ANOVA turns variances into a verdict about means: the full SS/df/MS/F bookkeeping on nine data points, the F-table decision, and what a significant F does not say.
How to Read an F Table
Look up the F critical value from the numerator and denominator degrees of freedom, with worked ANOVA examples, why the order matters and software equivalents.
What a one-way ANOVA tests
A one-way analysis of variance (ANOVA) tests whether the means of three or more independent groups are all equal. The null hypothesis is H₀: μ₁ = μ₂ = … = μₖ, and the alternative is that at least one mean is different. The test compares two estimates of the same variance: the spread of the group means around the grand mean (the between-group mean square, MSB) and the spread of the observations around their own group mean (the within-group mean square, MSW). If all groups share one mean the two are similar and F = MSB / MSW is close to 1. A large F means the group means lie further apart than the noise inside the groups can explain.
Use it for one numeric outcome and one factor with several levels, for example plant weight under a control and two treatments. With two groups it gives the same p-value as the pooled two-sample t-test (F = t²), so ANOVA is the natural extension of the t-test to more than two groups. Two factors at once need the two-way ANOVA calculator. A significant F only says that some means differ; the post-hoc tests below say which. New to the idea? Read ANOVA explained.
How to use the ANOVA calculator
- Choose the input type: raw data (one list per group) or summary statistics (n, mean and sample SD per group). Add another group gives you up to 10 groups.
- Enter each group's observations separated by commas, spaces or new lines, at least 2 per group. The groups may have different sizes.
- Pick the method, classical (assumes equal variances) or Welch (does not), and the significance level α, usually 0.05.
- Press Calculate ANOVA. Read the F-value and p-value, the ANOVA table, the effect sizes, the variance check and the step-by-step working. When the classical test is significant, the Tukey HSD comparisons follow.
- Load example enters the PlantGrowth data used below, and Copy link to this calculation shares your exact analysis.
ANOVA formulas
Group mean: x̄ᵢ = Σⱼ xᵢⱼ / nᵢ, grand mean: x̄ = Σᵢ Σⱼ xᵢⱼ / N
SST = Σᵢ Σⱼ (xᵢⱼ − x̄)²
SSB = Σᵢ nᵢ (x̄ᵢ − x̄)²
SSW = Σᵢ Σⱼ (xᵢⱼ − x̄ᵢ)² = Σᵢ (nᵢ − 1) sᵢ²
SST = SSB + SSW; df: between k − 1, within N − k, total N − 1
MSB = SSB / (k − 1), MSW = SSW / (N − k)
F = MSB / MSW, p = P(F(k − 1, N − k) ≥ F)
η² = SSB / SST, ω² = (SSB − (k − 1)·MSW) / (SST + MSW), f = √(η² / (1 − η²))
Welch: wᵢ = nᵢ / sᵢ², W = Σ wᵢ, x̄w = Σ wᵢ x̄ᵢ / W
Welch: F = [Σ wᵢ (x̄ᵢ − x̄w)² / (k − 1)] / [1 + 2 (k − 2) Λ], Λ = Σ (1 − wᵢ / W)² / (nᵢ − 1) / (k² − 1)
Welch: denominator df = 1 / (3 Λ)
The classical F and p-value are the ones printed by scipy.stats.f_oneway and by R's aov; the Welch F, degrees of freedom and p-value are those of R's oneway.test and of statsmodels' anova_oneway with unequal variances.
How to read the ANOVA table
- SS, df and MS. Each source of variation has a sum of squares (SS), degrees of freedom (df) and a mean square (MS = SS / df). Between groups measures how far the group means lie from the grand mean; within groups is the scatter of the observations around their own group mean, the error term.
- F and p. F = MSB / MSW. The p-value is the chance of an F at least this large if all population means were equal. Only a large F counts against H₀, so the test is right-tailed. Reject H₀ when p ≤ α, which is the same as F exceeding the critical value F(α; k − 1, N − k); read it in an F table or with the F table calculator.
- Effect size. η² is the share of the total variance that goes with group membership, ω² is a less biased estimate and Cohen's f = √(η² / (1 − η²)). Cohen (1988) calls η² = 0.01, 0.06 and 0.14 small, medium and large, and f = 0.10, 0.25 and 0.40. ω² can come out slightly negative when F is below 1; read a negative value as 0. See effect size explained.
- Variance check. The Brown-Forsythe (median Levene) test under the result compares the group variances. A small p, or a largest-to-smallest SD ratio above 2, is a reason to choose Welch ANOVA.
Worked example by hand: three groups of three
Three fertilizers are tried on three plots each. The yields are A: 4, 5, 6; B: 6, 7, 8; C: 8, 9, 10. Type them into the calculator to reproduce every number below.
| Group | n | Mean | Σ(x − x̄ᵢ)² | nᵢ(x̄ᵢ − x̄)² |
|---|---|---|---|---|
| A | 3 | 5 | 2 | 12 |
| B | 3 | 7 | 2 | 0 |
| C | 3 | 9 | 2 | 12 |
Working
Grand mean x̄ = (5 + 7 + 9) / 3 = 7, N = 9 and k = 3.
SSB = 12 + 0 + 12 = 24 and SSW = 2 + 2 + 2 = 6, so SST = 30.
df between = 2 and df within = 6, so MSB = 24 / 2 = 12 and MSW = 6 / 6 = 1.
F = 12 / 1 = 12 and p = P(F(2, 6) ≥ 12) = 0.008. The critical value F(0.05; 2, 6) is 5.1433, so H₀ is rejected.
η² = 24 / 30 = 0.8: 80% of the variation in yield goes with the fertilizer. Tukey HSD separates A and C (p = 0.0065); A against B and B against C are not significant on their own (p = 0.1089 each).
Worked example with real data: PlantGrowth
The PlantGrowth data set that ships with R (Dobson, 1983) holds the dry weight of 30 plants: ten controls and ten plants for each of two treatments. Press Load example to enter the weights.
| Group | Weights |
|---|---|
| Group 1 (control) | 4.17, 5.58, 5.18, 6.11, 4.50, 4.61, 5.17, 4.53, 5.33, 5.14 |
| Group 2 (treatment 1) | 4.81, 4.17, 4.41, 3.59, 5.87, 3.83, 6.03, 4.89, 4.32, 4.69 |
| Group 3 (treatment 2) | 6.31, 5.12, 5.54, 5.50, 5.37, 5.29, 4.92, 6.15, 5.80, 5.26 |
| Source | SS | df | MS | F | p |
|---|---|---|---|---|---|
| Between groups | 3.7663 | 2 | 1.8832 | 4.8461 | 0.0159 |
| Within groups | 10.4921 | 27 | 0.3886 | n/a | n/a |
| Total | 14.2584 | 29 | n/a | n/a | n/a |
F(2, 27) = 4.846 and p = 0.0159, so at α = 0.05 the mean weights differ. The critical value is F(0.05; 2, 27) = 3.3541. The effect sizes are η² = 0.2641, ω² = 0.2041 and Cohen's f = 0.5991: about 26% of the variation in weight goes with the condition, a large effect by Cohen's benchmarks. The Brown-Forsythe test finds no sign of unequal variances (W = 1.1192, p = 0.3412), so the classical test is appropriate; Welch's F is 5.181 with 2 and 17.128 degrees of freedom (p = 0.0174).
| Pair | Difference | Adj. p | 95% CI |
|---|---|---|---|
| G1 − G2 | 0.371 | 0.390871 | (-0.3202, 1.0622) |
| G1 − G3 | -0.494 | 0.197996 | (-1.1852, 0.1972) |
| G2 − G3 | -0.865 | 0.012006 | (-1.5562, -0.1738) |
Only treatment 2 against treatment 1 is significant after Tukey's correction (adjusted p = 0.012). Written up: F(2, 27) = 4.85, p = .016, η² = .26. R prints the same numbers: summary(aov(weight ~ group, data = PlantGrowth)) gives F = 4.846 and p = 0.0159, TukeyHSD gives the pair table above and oneway.test the Welch result.
When the group variances differ: classical versus Welch
The classical F test pools the variances of all groups. When the group sizes are unequal and the variances differ, its error rates drift away from α, and Welch's F test, which weights each group by n / s² and adjusts the denominator degrees of freedom, is the safer choice. Check the SD ratio and the Brown-Forsythe test under the result, or use the Levene test calculator. Example with A: 12.1, 11.8, 12.4, 12.0; B: 14.2, 15.9, 13.1, 16.4, 12.8, 15.0; C: 10.0, 13.5, 8.9, 15.2, 9.4, 14.8, 11.1 (group SDs 0.25, 1.4652 and 2.6248):
| Method | Statistic | df | p |
|---|---|---|---|
| Classical ANOVA | F = 3.6412 | 2, 14 | 0.0533 |
| Welch ANOVA | F = 7.7025 | 2, 7.6935 | 0.0146 |
| Brown-Forsythe variance check | W = 5.059 | 2, 14 | 0.0222 |
The classical test misses the difference at α = 0.05 (p = 0.0533), Welch's test finds it (p = 0.0146) and the variance check confirms that the spreads differ. Decide on the method from the design and the variance check, not from whichever p-value is smaller.
Assumptions of one-way ANOVA
- Independent observations. Each unit belongs to one group and units do not influence each other. No test can settle this; it comes from the design. Repeated measurements on the same subjects need a different model.
- Roughly normal values within each group. The F test tolerates moderate departures, especially when the groups are similar in size and not tiny. Check with a histogram or the Shapiro-Wilk test (see normality tests explained).
- Equal variances (classical F only). Use the Brown-Forsythe result and the SD ratio, or switch to Welch ANOVA.
- No extreme outliers. They inflate the within-group variance and pull the means. Screen the data with the outlier calculator.
- A numeric outcome. For ranks or strongly skewed data use the Kruskal-Wallis test (compare the two in parametric vs nonparametric tests).
After a significant ANOVA: which groups differ?
A significant F says that at least one mean differs, not which. Testing every pair with separate t-tests is no answer: k groups have k(k − 1)/2 pairs, and at α = 0.05 separate t-tests without correction give about 14.26% probability of at least one false positive for 3 groups (3 pairs) and 40.13% for 5 groups (10 pairs).
- Tukey HSD compares every pair while holding the family-wise error rate at α. It appears under the result when the classical test is significant; the Tukey HSD calculator also accepts group means with the MSE of any ANOVA.
- Bonferroni or Holm adjust the p-values of a smaller, planned set of comparisons; use the Bonferroni correction calculator.
- Games-Howell is the usual choice when the variances are unequal. This site does not offer it; SPSS lists it under Post Hoc and R has it in the rstatix package.
The reasoning behind these rules is in multiple comparisons explained.
One-way ANOVA in Excel, R, Python and SPSS
| Software | How to run it |
|---|---|
| Excel | Data > Data Analysis > Anova: Single Factor (Analysis ToolPak add-in), groups in columns; the output lists SS, df, MS, F, P-value and F crit. Without the ToolPak: =F.DIST.RT(F, df1, df2) for the p-value and =F.INV.RT(alpha, df1, df2) for the critical value |
| R | summary(aov(y ~ g, data = d)); Welch: oneway.test(y ~ g, data = d); post hoc: TukeyHSD(aov(y ~ g, data = d)) |
| Python | scipy.stats.f_oneway(a, b, c); Tukey: scipy.stats.tukey_hsd(a, b, c); Welch: statsmodels.stats.oneway.anova_oneway([a, b, c], use_var="unequal") |
| SPSS | Analyze > Compare Means and Proportions (Compare Means before version 29) > One-Way ANOVA; Options: Descriptive, Homogeneity of variance test, Welch; Post Hoc: Tukey, Games-Howell |
Frequently Asked Questions
What is a one-way ANOVA?
A test of whether the means of three or more independent groups are equal. It compares the variation between the group means with the variation inside the groups; a large ratio (F) is evidence that at least one mean differs.
How do I interpret the F-value and the p-value?
F is the between-group mean square divided by the within-group mean square. The p-value is the probability of an F at least that large if all population means were equal. If p is at most your α (usually 0.05), reject the hypothesis that all means are equal.
How do I calculate a one-way ANOVA by hand?
Find the group means and the grand mean, then SSB = Σ n(group mean − grand mean)² and SSW = Σ (x − group mean)². Divide by k − 1 and N − k degrees of freedom to get MSB and MSW, and F = MSB / MSW. The calculator shows every step for your data, and the worked example above does it for three groups of three.
What is the difference between ANOVA and a t-test?
A t-test compares two means, ANOVA compares two or more. With two groups they are the same test: F equals t squared and the p-values agree. With more groups, repeated t-tests inflate the chance of a false positive, which is why ANOVA is used first.
Why not run t-tests on every pair of groups?
With k groups there are k(k−1)/2 pairs; at α = 0.05 the chance of at least one spurious rejection is about 14.26% for k = 3 (3 pairs) and 40.13% for k = 5 (10 pairs). ANOVA tests all means in one step, and Tukey HSD then compares the pairs with the error rate controlled.
ANOVA is significant. Which groups differ?
The F test only says at least one mean differs. Use Tukey HSD (equal variances), Games-Howell (unequal variances) or Bonferroni-corrected pairs to locate the differences. The calculator lists the Tukey comparisons when the classical test is significant.
When should I use Welch ANOVA instead of classical ANOVA?
When the group variances differ (a small Brown-Forsythe p-value or a largest-to-smallest SD ratio above 2) or the group sizes are unequal. Welch weights each group by n / s² and adjusts the denominator degrees of freedom, so it does not assume equal variances.
What are the assumptions of one-way ANOVA?
Independent observations, roughly normal values within each group, equal variances across groups (classical F only) and a numeric outcome. Moderate non-normality is tolerable, unequal variances are handled by Welch ANOVA, and dependent observations need a different model.
Can I enter means and standard deviations instead of raw data?
Yes. Choose Summary mode and enter n, mean and sample SD per group. The classical and Welch results are the same as for the raw data, but the variance check, the plot and the Tukey table need the raw values.
What is eta squared (η²) and what counts as a large effect?
η² is SSB divided by SST, the share of the total variance that goes with group membership. Cohen's benchmarks are 0.01 (small), 0.06 (medium) and 0.14 (large); the equivalent Cohen's f values are 0.10, 0.25 and 0.40. Treat them as rough guides, not thresholds.
How does this map to the ANOVA output of Excel?
Excel's Anova: Single Factor lists Between Groups and Within Groups with SS, df, MS, F, P-value and F crit, the same quantities as the table here. This calculator adds the effect sizes, the variance check, the Welch option and the Tukey comparisons.
What if my data are not normal?
Moderate non-normality with similar group sizes is usually acceptable. For strong skew, outliers or ordinal data use the Kruskal-Wallis test or transform the values first.
Is ANOVA a one-tailed or a two-tailed test?
The F test is right-tailed: only a large F counts against the null hypothesis, and the p-value is the upper-tail area of the F distribution. It has no two-tailed version, but it is not directional about which mean is larger.
How many observations do I need per group?
The calculator needs at least 2 values per group. More is better: small groups give little power, and non-normal data or unequal variances are safer with larger, equal-sized groups. Plan the size before collecting data if you can.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.