Skip to content

Correlation matrix

Paste a multi-column table with a header row. See the Pearson or Spearman correlation of every pair of columns as a heat table, with missing values handled by pairwise deletion and the n of each pair; read the strongest pairs and multicollinearity warnings.

Free tool · Data analytics

The first row must hold the column names. You can copy from Excel or upload a CSV; separator and decimal format are detected automatically. Non-numeric columns are skipped; empty cells count as missing. At most 30 columns and 20,000 rows.

The data is read and calculated in your browser; it is not sent anywhere.

Type of correlation

Let us look together at which relationships in your data really help and what to watch for when building a model.

Request a call

01

How to use

  1. A

    Paste the multi-column table with a header row or upload a CSV; you can also try the sample data.

  2. B

    Choose Pearson (linear) or Spearman (rank) correlation; missing values are deleted pairwise.

  3. C

    Read the heat table, the 10 strongest pairs and the |r| > 0.9 multicollinearity warning; download the matrix as CSV.

02

What does a correlation matrix show?

The matrix collects the correlation coefficient of every pair of columns in one table: the diagonal is 1 (each column with itself) and the matrix is symmetric about it. Values lie between −1 and +1; the sign shows the direction (rising together, or one rising while the other falls) and the absolute value the strength of the relationship. With many variables it is a quick way to see at a glance which ones move together.

In the heat table the cell background darkens with |r| and the value itself is written in the cell. Because intensity and a number are used instead of colour, the table stays readable in black-and-white print and for colour-blind readers.

03

Pearson or Spearman?

Pearson measures the linear relationship of the values; it is sensitive to outliers and can understate a curved relationship. Spearman replaces the values with their ranks and computes Pearson on those: only the ordering matters, so it catches every monotonic relationship (rising or falling, even if not linear) and resists outliers. Equal values (ties) get the average rank.

If the two results differ a lot there is usually curvature or an outlier; look at the scatter plot of that pair.

04

Missing values and multicollinearity

For each pair only the rows where both are filled are used (pairwise deletion); the n of each cell may differ, and the tool shows it. With many missing values different pairs come from different subsets, so be careful comparing cells with one another. The alternative is to reduce all rows to one common set (listwise deletion), which is more consistent but loses data.

Pairs with |r| > 0.9 carry almost the same information (multicollinearity). If you feed such columns together into a regression or machine-learning model the coefficients become unstable; dropping one, averaging them or moving to principal components are common remedies.

FAQ

Why were my columns left out?
Non-numeric columns (text, dates) are skipped and the skipped ones are listed. A column counts as text if fewer than 80% of its values are numbers. Clean formats that keep numbers as text (a unit suffix such as "12 kg", for example).
What is the small number in a cell?
It is n, the number of rows used for that pair of columns. Without missing values it is the same everywhere and the tool shows a single n; with missing values it can differ from cell to cell.
How many rows are enough?
The calculation works with 3 rows, but that is very few. At least 30 rows are recommended for a reliable reading; with a small n, |r| can come out high by chance alone. The p-value in the strongest-pairs table shows this uncertainty.
Does a high correlation mean cause and effect?
No. Two variables moving together does not show that one causes the other; a common third factor, the reverse direction or chance may be behind it. Two series that both grow over time almost always show a high correlation.
Can I trust the p-values when there are many pairs?
Be careful. With 30 columns 435 pairs are tested; even with no real relationship about 5% of them give p < 0.05. This tool applies no multiple-comparison correction; read the p-value as a ranking aid.
Why does a constant column give an error?
A column whose values are all the same has zero variance; the correlation formula divides by zero and is undefined. Such a column carries no relationship with other variables anyway; remove it from the matrix.

Turn relationships into the right model

Let us find which variables in your production, sales or field data really influence each other; we can build forecasting and decision-support models together while managing multicollinearity.