Correlation Coefficient Calculator

Calculate Pearson's correlation coefficient r for paired X and Y data. The calculator returns r and R², the population and sample covariance, the t statistic, degrees of freedom and two-sided p-value of the test that the true correlation is zero, a Fisher z confidence interval for ρ, the working steps and a scatter plot.

Related guides: correlation vs causation, covariance vs correlation and how to interpret R².

Enter numbers separated by commas, spaces or new lines

One Y for every X, in the same order

A decimal such as 0.95: sets the interval and the significance level 1 − level

What the correlation coefficient measures

Pearson's correlation coefficient r summarizes how closely paired data follow a straight line. It runs from −1 (a perfect downhill line) through 0 (no linear pattern) to +1 (a perfect uphill line), and it does not change when either variable is rescaled or shifted, so heights in centimeters and inches give the same r. Squaring it gives R², the share of the variation in Y that a straight line on X explains.

r describes a linear pattern only. A curve such as a perfect U can give r near 0, and a single extreme point can create or hide a correlation, so always look at the scatter plot. To predict Y from X, use the line of best fit calculator or the linear regression calculator; for a rank-based measure that tolerates outliers and curved monotonic trends, use the Spearman correlation calculator.

When the two variables are items of one questionnaire scale, the correlations between all pairs of items are what Cronbach's alpha summarizes: the Cronbach's alpha calculator reports alpha for the whole scale together with the correlation of each item with the rest.

Formula

r = Sxy / √( Sxx · Syy )

Sxy = Σ(x − x̄)(y − ȳ), Sxx = Σ(x − x̄)², Syy = Σ(y − ȳ)²

Covariance: Sxy / n (population), Sxy / (n − 1) (sample); r = covariance / (s_x·s_y) with the matching standard deviations

t = r·√( (n − 2) / (1 − r²) ) on n − 2 degrees of freedom

Fisher z: z′ = atanh(r), SE = 1 / √(n − 3), CI for ρ = tanh( z′ ± z*·SE )

The calculator adds up deviations from the means with compensated summation and evaluates the t and normal tail probabilities directly, so results stay accurate for data far from zero, such as years, and for r very close to ±1.

How to interpret r

|r|Common description
0.00 – 0.19Very weak or none
0.20 – 0.39Weak
0.40 – 0.59Moderate
0.60 – 0.79Strong
0.80 – 1.00Very strong

These bands are conventions, not tests. Cohen (1988) suggested 0.1, 0.3 and 0.5 as small, medium and large effects for behavioral research, and physics or engineering data routinely reach 0.99. The sign gives the direction: positive means Y tends to rise with X, negative means it tends to fall.

Is the correlation statistically significant?

The p-value answers a narrower question than the size of r: if the true correlation ρ were 0, how often would a sample of this size give an |r| at least this large? It depends heavily on n, as the smallest significant |r| shows:

Pairs (n)Degrees of freedom|r| needed at 5%|r| needed at 1%
530.87830.9587
860.70670.8343
1080.63190.7646
15130.5140.6411
20180.44380.5614
30280.3610.4629
50480.27870.361
100980.19660.2565

Values are two-sided critical values of |r| = t*/√(df + t*²). With only a handful of pairs, even a large r can be a chance result, and with thousands of pairs a trivial r becomes significant, so report r with its interval rather than the p-value alone.

Worked example: study hours and quiz scores

Five students study X = 1, 2, 3, 4, 5 hours and earn quiz scores Y = 2, 4, 5, 4, 5. Load example fills in these numbers.

  1. Means and sums of squares: x̄ = 3, ȳ = 4, Sxx = 10, Syy = 6, Sxy = 6.
  2. Correlation: r = 6 ÷ √(10 × 6) = 6 ÷ 7.7460 = 0.7746, so R² = 0.6.
  3. Covariance: 6 ÷ 5 = 1.2 for the population and 6 ÷ 4 = 1.5 for a sample.
  4. Test: t = 0.7746 × √(3 ÷ 0.4) = 2.1213 on 3 degrees of freedom, so p = 0.124027.
  5. Interval: z′ = atanh(0.7746) = 1.0317, SE = 1 ÷ √2 = 0.7071, z* = 1.96, so the 95% interval for ρ is tanh(1.0317 ± 1.3859) = (−0.3401, 0.9842).

Interpretation: r = 0.77 is a fairly strong positive linear pattern, and 60% of the variation in scores goes with study time. But with five students p = 0.124 is above 0.05 and the interval includes 0, so the data do not rule out no correlation at all; the table above shows that n = 5 needs |r| ≥ 0.8783.

Assumptions and cautions

  • Two numeric variables measured on the same units. Each X must be paired with the Y of the same student, day or item; the lists need the same length.
  • A linear pattern. Check the scatter plot for curves before relying on r.
  • Few outliers. Pearson's r uses every distance from the mean, so one extreme point can move it a long way; the Spearman correlation is more resistant.
  • Roughly bivariate normal data for the p-value and the Fisher z interval when n is small. The coefficient itself needs no normality.
  • Restricted range. Looking at only part of the range of X, such as top students only, makes r smaller than in the full population.
  • No causal claim. A strong r can come from a third variable or from chance (see correlation vs causation).

Correlation in other software

ToolCommand
TI-84STAT > CALC > 4:LinReg(ax+b) shows r and r² after 2nd > CATALOG > DiagnosticOn; STAT > TESTS > F:LinRegTTest tests ρ
Excel / Google Sheets=CORREL(x, y) or =PEARSON(x, y); =RSQ(x, y); p-value: =T.DIST.2T(ABS(r)*SQRT((n-2)/(1-r^2)), n-2)
Rcor(x, y); cor.test(x, y) adds the t test and the Fisher z interval
Pythonscipy.stats.pearsonr(x, y) returns r and the p-value; numpy.corrcoef(x, y)

For the covariance on its own, with a choice of sample or population, use the covariance calculator; to split R² into SSR, SSE and SST, use the R-squared calculator.

Frequently Asked Questions

What is a good correlation coefficient?

It depends on the field and the purpose. Textbooks often call |r| of 0.2 to 0.4 weak, 0.4 to 0.6 moderate, 0.6 to 0.8 strong and above 0.8 very strong, while Cohen's benchmarks for behavioral research are 0.1, 0.3 and 0.5. Physical measurements often exceed 0.99. Report r with its confidence interval and look at the scatter plot rather than relying on a label.

Does a high correlation prove that one variable causes the other?

No. Correlation measures association, not causation. Ice cream sales and drownings correlate because both rise in summer. Establishing causation needs a controlled experiment or careful causal inference, not just a large r.

What is the difference between r and R²?

r is the correlation coefficient, between −1 and 1, with a sign for the direction. R² is its square, between 0 and 1: the proportion of the variance in Y explained by a straight-line fit on X. An r of 0.7746 gives R² = 0.6, meaning 60% of the variation is explained.

When should I use Spearman instead of Pearson?

Use Spearman's rank correlation when the data are ordinal, when there are outliers, or when the relationship is monotonic but curved. Pearson measures how well a straight line fits the raw values and is sensitive to outliers; Spearman applies the same formula to the ranks.

What does r close to 0 mean?

It means there is little or no linear relationship. It does not rule out a strong nonlinear one: a perfect U-shaped pattern can give r near 0. Always inspect the scatter plot before concluding that the variables are unrelated.

How many pairs do I need for a significant correlation?

It depends on the size of r. At the 5% level a correlation must be at least 0.8783 in absolute value with 5 pairs, 0.6319 with 10, 0.4438 with 20, 0.2787 with 50 and 0.1966 with 100 (two-sided). Small samples also make the confidence interval for ρ very wide.

How is the confidence interval for the correlation calculated?

The calculator applies the Fisher z transformation z′ = atanh(r), whose sampling distribution is approximately normal with standard error 1 / √(n − 3), builds the interval z′ ± z*·SE, and transforms the limits back with tanh. It needs at least 4 pairs and assumes roughly bivariate normal data.

Why do my X and Y lists need the same length?

Correlation is computed over paired observations: each X value must belong to a Y value measured on the same unit, such as the same student or day. Lists of different lengths leave some observations without a pair, so the calculator reports the two counts instead of guessing.

Embed This Calculator

Add this free calculator to your course page or LMS.

Adjust the height value to fit your page.