Linear Regression Calculator
Fit a least-squares line to paired X and Y data. The calculator returns the regression equation, slope, intercept, R² and adjusted R², the standard error, t statistic, p-value and confidence interval of both coefficients, the ANOVA table, a scatter plot with the fitted line, every residual and, for any X you enter, the predicted Y with its confidence and prediction intervals.
Related guides: linear regression explained, how to interpret R² and correlation vs causation. For curved trends, use the quadratic regression calculator or the exponential regression calculator, or see LOWESS smoothing. To score predictions against observed values, use the mean squared error calculator; to judge the fit with a residual plot, use the residual calculator.
Enter numbers separated by commas, spaces or new lines
One Y for every X, in the same order
A decimal such as 0.95, for the intervals
Related Calculators
Line of Best Fit Calculator
Find the least-squares line of best fit for X–Y data with the equation, slope, intercept, correlation, working table, residuals and a scatter plot.
Correlation Coefficient Calculator
Calculate Pearson's r with its p-value, Fisher z confidence interval, R², covariance, worked steps and a scatter plot.
R-Squared Calculator
Get R², adjusted R², r, and the SSR/SSE/SST breakdown from X–Y data or observed vs predicted values.
What a linear regression calculator does
Simple linear regression describes how a response Y changes with one predictor X by the straight line ŷ = b₀ + b₁x that makes the sum of squared vertical distances between the points and the line as small as possible. Enter the paired values above and the calculator gives the line, how well it fits, and how much uncertainty surrounds it: the standard errors, t tests and confidence intervals of the slope and the intercept, the analysis of variance table, and confidence and prediction intervals at any X you type in.
If you only need the equation and a plot, the line of best fit calculator shows the hand calculation step by step. To test the slope the way a TI-84 does, use the LinRegTTest calculator; for the strength of the relationship alone, use the correlation calculator.
Formulas
Slope: b₁ = Sxy / Sxx, Sxy = Σ(x − x̄)(y − ȳ), Sxx = Σ(x − x̄)²
Intercept: b₀ = ȳ − b₁·x̄
R² = 1 − SSE / SST = SSR / SST = r², SSE = Σ(y − ŷ)², SST = Σ(y − ȳ)²
Adjusted R² = 1 − [SSE / (n − 2)] / [SST / (n − 1)]
Standard error of the estimate: s = √( SSE / (n − 2) )
SE(b₁) = s / √Sxx, SE(b₀) = s·√( 1/n + x̄² / Sxx )
t = b / SE(b) on n − 2 degrees of freedom; CI: b ± t*·SE(b)
F = MSR / MSE = SSR / s², and F = t² for the slope
Mean of Y at x₀: ŷ ± t*·s·√( 1/n + (x₀ − x̄)² / Sxx )
New Y at x₀: ŷ ± t*·s·√( 1 + 1/n + (x₀ − x̄)² / Sxx )
The sums are taken over deviations from the means, so data with a large offset, such as years or prices in the millions, keep their precision. t and F tail probabilities come from direct evaluations of the t and F distributions, which is why very small p-values are reported exactly instead of being rounded to 0.
How to read the results
| Output | What it tells you |
|---|---|
| Slope b₁ | The average change in Y for a one-unit increase in X. Its p-value tests the null hypothesis that the true slope is 0. |
| Intercept b₀ | The predicted Y when X is 0. It is only meaningful when X = 0 is inside, or close to, the range of your data. |
| R² Value | The share of the variation in Y that the line explains, from 0 to 1. In simple regression it equals r², the square of the correlation. |
| Adjusted R² | R² corrected for the degrees of freedom used by the fit. It is never above R² (and is negative for a very poor fit), and matters most when comparing models of different sizes. |
| Residual standard error (s) | The typical distance of a point from the line, in the units of Y. |
| Confidence interval for mean Y | A range for the average Y of all observations at that X. It shrinks as n grows. |
| Prediction interval for new Y | A range for one new observation at that X. It is always wider, because it adds the scatter of a single point around the line. |
| F and its p-value | The overall test of the regression. With one predictor it is the square of the slope's t statistic and gives the same p-value. |
Worked example: ad spend and sign-ups
A startup records weekly ad spend (in $100s), X = 1, 2, 3, 4, 5, and new sign-ups (in dozens), Y = 2, 4, 5, 4, 5. Load example fills in these numbers and asks for a prediction at x = 6.
- Means and sums of squares: x̄ = 3, ȳ = 4, Sxx = 10, Sxy = 6, Syy = 6.
- Slope: b₁ = 6 ÷ 10 = 0.6. Intercept: b₀ = 4 − 0.6 × 3 = 2.2. The line is ŷ = 0.6x + 2.2.
- Fit: SSR = 0.6 × 6 = 3.6 and SSE = 6 − 3.6 = 2.4, so R² = 3.6 ÷ 6 = 0.6, adjusted R² = 1 − (2.4 ÷ 3) ÷ (6 ÷ 4) = 0.4667, and s = √(2.4 ÷ 3) = 0.8944.
- Slope test: SE(b₁) = 0.8944 ÷ √10 = 0.2828, t = 0.6 ÷ 0.2828 = 2.1213 on 3 degrees of freedom, p = 0.1240. With t* = 3.1824 the 95% interval for the slope is 0.6 ± 0.9001 = (−0.3001, 1.5001).
- Prediction at x = 6: ŷ = 0.6 × 6 + 2.2 = 5.8. The 95% confidence interval for the mean sign-ups at that spend is (2.8146, 8.7854); the 95% prediction interval for one new week is (1.6751, 9.9249).
Interpretation: each extra $100 of weekly ad spend is associated with 0.6 dozen (about 7) more sign-ups. Yet R² = 0.6 with only five weeks is not enough evidence: the slope's p-value of 0.124 is above 0.05 and its interval includes 0. And x = 6 lies outside the observed range 1 to 5, so the calculator flags that prediction as an extrapolation.
Assumptions and pitfalls
- Linear: the average of Y changes in a straight line with X. A curved pattern in the residuals means a line is the wrong model (the residual calculator draws the residual plot); try a smoother or transform the data.
- Independent observations: repeated measurements of the same subject, or data ordered in time, break the standard errors and p-values.
- Constant spread: the residuals should scatter evenly around the line. A funnel shape means the prediction intervals are too narrow for large X and too wide for small X.
- Roughly normal residuals: needed for the t tests and intervals when n is small; with large n they are robust to moderate skew.
- Outliers: one extreme point can tilt the line. Look at the plot and at the largest residuals, and check each pair's leverage and Cook's distance in the residual calculator, before trusting the slope.
- Extrapolation and causation: a fitted line describes the range of X that was observed, and an association does not show that changing X changes Y (see correlation vs causation).
Regressing Y on X is not the same as regressing X on Y: the line minimizes vertical distances, so swapping the variables gives a different line unless the points lie exactly on one. R² and the p-value of the slope stay the same.
The same regression in other software
| Tool | Command |
|---|---|
| TI-84 | STAT > CALC > 4:LinReg(ax+b) gives the slope a, the intercept b, and r² and r with DiagnosticOn; STAT > TESTS > F:LinRegTTest adds t, p and s |
| Excel / Google Sheets | =SLOPE(y, x), =INTERCEPT(y, x), =RSQ(y, x), =STEYX(y, x), =FORECAST(x0, y, x), =LINEST(y, x, TRUE, TRUE) |
| Excel Data Analysis | Data > Data Analysis > Regression: coefficient table, ANOVA, confidence intervals and residuals |
| R | fit <- lm(y ~ x); summary(fit); confint(fit); predict(fit, data.frame(x = 6), interval = "prediction") |
| Python | scipy.stats.linregress(x, y); statsmodels.api.OLS(y, sm.add_constant(x)).fit().summary() |
All of them use the same least-squares formulas, so their results agree with this calculator apart from rounding in the last digits. To break R² into SSR, SSE and SST, or to score predictions against observed values, use the R-squared calculator.
Frequently Asked Questions
How do I interpret the slope and the intercept?
The slope is the expected change in Y for a one-unit increase in X: a slope of 0.6 means each extra unit of X goes with 0.6 more units of Y on average. The intercept is the predicted Y when X is 0, which is only meaningful if X = 0 is a realistic value near your data. Use the coefficient table to see whether each estimate is distinguishable from 0 (p-value) and how precisely it is known (confidence interval).
What does R-squared tell me, and can it be high while the slope is not significant?
R² is the proportion of the variance in Y explained by the line, from 0 to 1; 0.6 means 60% of the variation is captured. It does not say the model is correct or the slope is significant. With few points a large R² can come from chance: five points with R² = 0.6 give a slope p-value of about 0.124. Judge the fit with the p-value, the interval for the slope and the residual plot together.
What is the difference between a confidence interval and a prediction interval?
The confidence interval bounds the average Y of all observations at a given X, so it narrows as the sample grows. The prediction interval bounds one new observation at that X and adds the scatter of a single point around the line, so it is always wider and does not shrink to zero even with a huge sample. Use the prediction interval to say where a future value is likely to fall.
Can I use the line to predict outside my data range?
Extrapolation is risky. The relationship has only been observed over the range of your X values; beyond it the true relationship may bend, flatten or break, and the intervals assume the line still holds. The calculator flags a prediction outside the observed range so that you treat it with caution.
What is the difference between correlation and regression?
Correlation is a single symmetric number for the strength and direction of a linear relationship. Regression fits an asymmetric equation for predicting Y from X, with a slope, an intercept and predictions. They are linked: in simple regression R² equals the squared correlation coefficient, and the slope equals r times the ratio of the standard deviations of Y and X.
Why does the calculator say the tests and intervals are unavailable?
There are three cases with no error to estimate: only 2 pairs (the line passes through both points exactly), all Y values equal (a horizontal line, so R² is undefined), and points that lie exactly on a line (zero residual variance, so the standard errors are 0 and the t and F statistics are infinite). In each case the equation and predictions are still shown.
Does my data have to be normally distributed?
No. The least-squares line needs no distributional assumption. The t tests and intervals assume the residuals (the vertical distances from the line) are roughly normal with constant spread, which matters mainly for small samples. Check the residual table and the plot for curvature, funnel shapes and outliers.
How many data points do I need?
The line needs at least 2 points with different X values and the tests need at least 3, but with fewer than about 10 pairs the intervals are very wide and the assumptions cannot be checked. A common rule of thumb for a reliable fit is 10 to 20 observations per predictor. The calculator accepts up to 2000 pairs.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.