Skip to content

Correlation & Covariance

Correlation and covariance are fundamental concepts in multivariate statistics. PCA is directly based on covariance or correlation matrices, while many other multivariate methods use related ideas of similarity, distance, and shared variation. Understanding these concepts will make the methods that follow much easier to interpret.

Covariance

Covariance measures how two variables change together:

\text{cov}(X, Y) = \frac{\sum_{i=1}^{n}(x_i - \bar{x})(y_i - \bar{y})}{n-1}
  • Positive: X and Y tend to increase together
  • Negative: X increases as Y decreases
  • Zero: no linear association
height <- c(160, 165, 170, 175, 180)
weight <- c(55, 60, 65, 70, 75)

cov(height, weight)
# [1] 62.5

For multiple variables, covariances are organised in a covariance matrix. The diagonal contains variances; the off-diagonal entries are pairwise covariances.

data <- data.frame(height, weight, age = c(20, 22, 25, 28, 30))
cov(data)
#        height weight  age
# height   62.5   62.5 32.5
# weight   62.5   62.5 32.5
# age      32.5   32.5 17.0

The Problem with Covariance

Covariance depends on the units of measurement. Expressing height in metres instead of centimetres changes the covariance 100-fold, even though the relationship is identical. This makes covariances difficult to compare across variables.


Correlation

Correlation standardises covariance to always fall between -1 and +1:

\text{cor}(X, Y) = \frac{\text{cov}(X, Y)}{\sigma_X \cdot \sigma_Y}

It is unitless and directly interpretable:

Value Interpretation
+1 Perfect positive linear relationship
0 No linear relationship
-1 Perfect negative linear relationship
Between -1 and +1 Strength depends on context

The closer |r| is to 1, the stronger the linear association. But there are no universal cutoffs for what counts as "weak" or "strong".

data(iris)
cor_matrix <- cor(iris[, 1:4])
round(cor_matrix, 2)
#              Sepal.Length Sepal.Width Petal.Length Petal.Width
# Sepal.Length         1.00       -0.12         0.87        0.82
# Sepal.Width         -0.12        1.00        -0.43       -0.37
# Petal.Length         0.87       -0.43         1.00        0.96
# Petal.Width          0.82       -0.37         0.96        1.00

Which Method to Use

Three correlation coefficients are commonly used. The choice depends on your data:

Method Measures Useful when
Pearson Linear association Continuous variables and approximately linear relationships
Spearman Monotonic association Ranked/ordinal data, or relationships that are monotonic but not necessarily linear
Kendall Rank concordance Ordinal/ranked data, especially when ties or small samples matter
x <- c(1, 2, 3, 4,  5,  6,  7,  8,  9, 100)  # note the outlier
y <- c(2, 4, 6, 8, 10, 12, 14, 16, 18, 20)

cor(x, y, method = "pearson")   # > 0.59
cor(x, y, method = "spearman")  # = 1.00
cor(x, y, method = "kendall")   # = 1.00

Notice that the relationship remains perfectly monotonic, but is no longer perfectly linear. Pearson therefore changes substantially, while the rank-based measures do not.

Correlation measures association, not just any relationship

Pearson correlation measures linear association, while Spearman and Kendall measure monotonic association. A zero Pearson correlation therefore signifies the absence of a linear relationship rather than the absence of any relationship. Therefore, always plot your data.

x <- seq(-5, 5, 0.1)
y <- x^2           # perfect nonlinear relationship
plot(x, y)
cor(x, y)          # small but not zero

Visualising Correlation Structure

A correlation heatmap is usually the most informative first step:

library(corrplot)

corrplot(cor(iris[, 1:4]),
         method = "circle",
         type = "upper",
         order = "hclust",
         addCoef.col = "lightbluelibrary(GGally)",
         tl.col = "black",
         tl.srt = 45)

For a fuller overview including scatterplots and distributions:

library(GGally)

ggpairs(iris,
        columns = 1:4,
        aes(color = Species, alpha = 0.5))

Covariance vs. Correlation: When to Use Which

Correlation is best for interpretation, visualization, and comparing variables on different scales. Covariance is preferable when scale carries meaning, particularly when running PCA on variables with the same unit to maintain their natural variance.

PCA and scaling

PCA on the correlation matrix gives each variable equal variance after standardisation. PCA on the covariance matrix retains the original scale of the variables, so variables with larger variance contribute more strongly. When in doubt, scale your data (scale. = TRUE in prcomp()).


Exercise

Using the mtcars dataset:

  1. Compute and visualise the correlation matrix
  2. Identify the three strongest pairwise correlations, regardless of direction
  3. Test whether Pearson and Spearman give meaningfully different results
Solution
library(corrplot)

cor_p <- cor(mtcars, method = "pearson")
cor_s <- cor(mtcars, method = "spearman")

corrplot(cor_p, method = "circle", type = "upper",
         order = "hclust", tl.col = "black")

source("https://www.gdc-docs.ethz.ch/UniBS/MultivariateStatistics/examples/top_cor_pairs.R")
top_cor_pairs(cor_p, n = 3, use_abs = TRUE)
# Strongest correlations (Pearson)
# cyl  & disp:  0.90
# cyl  & hp  :  0.89
# disp & wt  :  0.87

# Compare Pearson vs Spearman
round(cor_p - cor_s, 2)
# Large differences indicate outliers or non-linearity