Correlation & Covariance
Correlation and covariance are fundamental concepts in multivariate statistics. PCA is directly based on covariance or correlation matrices, while many other multivariate methods use related ideas of similarity, distance, and shared variation. Understanding these concepts will make the methods that follow much easier to interpret.
Covariance
Covariance measures how two variables change together:
- Positive: X and Y tend to increase together
- Negative: X increases as Y decreases
- Zero: no linear association
height <- c(160, 165, 170, 175, 180)
weight <- c(55, 60, 65, 70, 75)
cov(height, weight)
# [1] 62.5
For multiple variables, covariances are organised in a covariance matrix. The diagonal contains variances; the off-diagonal entries are pairwise covariances.
data <- data.frame(height, weight, age = c(20, 22, 25, 28, 30))
cov(data)
# height weight age
# height 62.5 62.5 32.5
# weight 62.5 62.5 32.5
# age 32.5 32.5 17.0
The Problem with Covariance
Covariance depends on the units of measurement. Expressing height in metres instead of centimetres changes the covariance 100-fold, even though the relationship is identical. This makes covariances difficult to compare across variables.
Correlation
Correlation standardises covariance to always fall between -1 and +1:
It is unitless and directly interpretable:
| Value | Interpretation |
|---|---|
| +1 | Perfect positive linear relationship |
| 0 | No linear relationship |
| -1 | Perfect negative linear relationship |
| Between -1 and +1 | Strength depends on context |
The closer |r| is to 1, the stronger the linear association. But there are no universal cutoffs for what counts as "weak" or "strong".
data(iris)
cor_matrix <- cor(iris[, 1:4])
round(cor_matrix, 2)
# Sepal.Length Sepal.Width Petal.Length Petal.Width
# Sepal.Length 1.00 -0.12 0.87 0.82
# Sepal.Width -0.12 1.00 -0.43 -0.37
# Petal.Length 0.87 -0.43 1.00 0.96
# Petal.Width 0.82 -0.37 0.96 1.00
Which Method to Use
Three correlation coefficients are commonly used. The choice depends on your data:
| Method | Measures | Useful when |
|---|---|---|
| Pearson | Linear association | Continuous variables and approximately linear relationships |
| Spearman | Monotonic association | Ranked/ordinal data, or relationships that are monotonic but not necessarily linear |
| Kendall | Rank concordance | Ordinal/ranked data, especially when ties or small samples matter |
x <- c(1, 2, 3, 4, 5, 6, 7, 8, 9, 100) # note the outlier
y <- c(2, 4, 6, 8, 10, 12, 14, 16, 18, 20)
cor(x, y, method = "pearson") # > 0.59
cor(x, y, method = "spearman") # = 1.00
cor(x, y, method = "kendall") # = 1.00
Notice that the relationship remains perfectly monotonic, but is no longer perfectly linear. Pearson therefore changes substantially, while the rank-based measures do not.
Correlation measures association, not just any relationship
Pearson correlation measures linear association, while Spearman and Kendall measure monotonic association. A zero Pearson correlation therefore signifies the absence of a linear relationship rather than the absence of any relationship. Therefore, always plot your data.
x <- seq(-5, 5, 0.1)
y <- x^2 # perfect nonlinear relationship
plot(x, y)
cor(x, y) # small but not zero
Visualising Correlation Structure
A correlation heatmap is usually the most informative first step:
library(corrplot)
corrplot(cor(iris[, 1:4]),
method = "circle",
type = "upper",
order = "hclust",
addCoef.col = "lightbluelibrary(GGally)",
tl.col = "black",
tl.srt = 45)
For a fuller overview including scatterplots and distributions:
library(GGally)
ggpairs(iris,
columns = 1:4,
aes(color = Species, alpha = 0.5))
Covariance vs. Correlation: When to Use Which
Correlation is best for interpretation, visualization, and comparing variables on different scales. Covariance is preferable when scale carries meaning, particularly when running PCA on variables with the same unit to maintain their natural variance.
PCA and scaling
PCA on the correlation matrix gives each variable equal variance after standardisation. PCA on the covariance matrix retains the original scale of the variables, so variables with larger variance contribute more strongly. When in doubt, scale your data (scale. = TRUE in prcomp()).
Exercise
Using the mtcars dataset:
- Compute and visualise the correlation matrix
- Identify the three strongest pairwise correlations, regardless of direction
- Test whether Pearson and Spearman give meaningfully different results
Solution
library(corrplot)
cor_p <- cor(mtcars, method = "pearson")
cor_s <- cor(mtcars, method = "spearman")
corrplot(cor_p, method = "circle", type = "upper",
order = "hclust", tl.col = "black")
source("https://www.gdc-docs.ethz.ch/UniBS/MultivariateStatistics/examples/top_cor_pairs.R")
top_cor_pairs(cor_p, n = 3, use_abs = TRUE)
# Strongest correlations (Pearson)
# cyl & disp: 0.90
# cyl & hp : 0.89
# disp & wt : 0.87
# Compare Pearson vs Spearman
round(cor_p - cor_s, 2)
# Large differences indicate outliers or non-linearity