Skip to content

Correlation and simple regression calculator

Paste two columns (x and y). See the Pearson correlation, the least-squares line, the slope's standard error, 95% confidence intervals and the p-value, and get a prediction of y for any x.

Free tool · Data analytics

One (x, y) pair per line. The separator can be a tab, ; , | or a space (columns copied from Excel work); decimal comma and point are detected automatically. A header row is skipped; the first two columns are used.

The calculation runs in your browser; the data you paste is not sent anywhere.

Let us look together at whether the relationships in your data really help and how they can become decision support.

Request a call

01

How to use

  1. A

    Paste the x and y columns (you can copy them from Excel) or load the sample data.

  2. B

    Read the correlation, R², slope and intercept, standard errors, 95% intervals and the p-value; check in the chart how well the points follow the line.

  3. C

    Enter an x value to get the predicted y and its intervals; copy the result.

02

What do Pearson r and R² tell you?

The Pearson correlation coefficient r shows the direction and strength of the linear relationship between two variables, from −1 to +1. A value near zero says there is no linear relationship, but another kind of relationship (a U shape, for example) may exist. A rough guide for describing strength in words: |r| < 0.1 negligible, < 0.3 weak, < 0.5 moderate, < 0.8 strong, above that very strong; these thresholds vary by field.

R² = r² is the share of the variability in y explained by the linear model. R² = 0.60 means 60% of the variability is explained by x and 40% is not. A high R² does not mean the model is right or that the relationship is causal.

03

Slope, standard error, confidence interval and p-value

The least-squares line minimises the sum of squared vertical deviations. The slope b tells how many units y changes on average when x rises by one unit. The standard error shows how much the slope could vary in another sample; the 95% confidence interval is b ± t·SE, where t comes from the Student t distribution with n − 2 degrees of freedom.

The p-value is the probability of seeing a slope this extreme or more if the true slope were zero. A small p says the relationship is hard to explain by chance alone; it says nothing about the size or practical importance of the relationship. With a very large sample even a tiny slope turns out "significant".

Two intervals are given for the prediction: the confidence interval for the mean response (how uncertain the line itself is) and the prediction interval for a single new observation (the scatter of a single measurement is added). The prediction interval is always wider.

04

Correlation is not causation

Two variables moving together does not prove that one causes the other. A common third factor (air temperature raises both ice-cream sales and electricity use), reverse causation or plain chance can produce the same result. Two series that both grow over time almost always show high correlation.

A causal claim needs an experiment, a well-controlled comparison or additional knowledge (a physical mechanism). Watch for outliers too: a single extreme point can change r greatly; do not trust the number without looking at the chart.

FAQ

How large must a correlation coefficient be to count as strong?
There is no fixed threshold. A rough guide: |r| < 0.3 weak, 0.3–0.5 moderate, 0.5–0.8 strong, above that very strong. In physics and engineering 0.9 is expected, while in human-behaviour data even 0.3 can be meaningful. Always look at the chart too.
What is the difference between R² and r?
R² is the square of r and gives the share of the variability in y explained by the linear model. r also carries the direction (plus or minus), R² does not. r = −0.8 and r = +0.8 give the same R² (0.64).
If the p-value is below 0.05, is the relationship strong?
No. The p-value concerns how likely the relationship is to be a chance result, not its strength. With a large sample even a very weak relationship gives p < 0.05; for the size of the relationship look at r, R² and the slope interval.
How many data points do I need?
The tool calculates with 3 pairs, but that is very few. For a sound assessment at least 20–30 pairs are recommended; with few pairs a single outlier can change the result entirely.
Can I predict outside the range of the data with the regression line?
Technically yes, but it is unreliable. Nobody knows whether the linear relationship continues in a region that was not measured. The tool shows a warning for an x outside the data range.
In what format should I paste the data?
Two columns: x and y on each line. The separator can be a tab (copied from Excel), ;, | or a space. Decimal comma or point is detected automatically; a header row is skipped and if there are more than two columns the first two are used.

Turn the relationships in your data into decisions

Let us look together at which variables in your production, sales or field data really affect each other and how you can turn that into decision support.