Linear Regression and Correlation
Modelling
Linear Regression and Correlation
Syllabus tag: KASNEB CPA | Advanced Level | CA34S1 Business Data Analytics
1. Correlation
Correlation measures the strength and direction of the linear relationship between two quantitative variables.
The Pearson correlation coefficient (r) ranges from –1 to +1:
- r = +1: perfect positive linear relationship
- r = 0: no linear relationship
- r = –1: perfect negative linear relationship
Interpretation: |r| < 0.3 (weak), 0.3–0.7 (moderate), > 0.7 (strong).
Important: correlation does not imply causation. Two variables may be correlated because of a common cause (spurious correlation) or coincidence.
2. Simple linear regression
Simple linear regression models the relationship between one independent variable (x) and one dependent variable (y):
y = a + bx
Where: a is the y-intercept (value of y when x = 0) and b is the slope (change in y for each unit increase in x).
3. Calculating regression coefficients
Given n pairs of (x, y) values:
b = [n·Σxy – Σx·Σy] / [n·Σx² – (Σx)²]
a = ȳ – b·x̄
Where x̄ and ȳ are the means of x and y respectively.
4. The coefficient of determination (R²)
R² measures the proportion of variation in y explained by the regression model:
R² = r² (for simple linear regression)
R² ranges from 0 to 1 (or 0% to 100%). R² = 0.75 means 75% of the variation in y is explained by x; the remaining 25% is unexplained (residual variation).
5. Residuals and model fit
A residual is the difference between the observed value of y and the predicted value: e = y – ŷ. A good regression model has residuals that are: small (low prediction error); randomly distributed (no pattern); normally distributed with mean zero. Heteroscedasticity (residuals increasing with x) and autocorrelation (patterns in residuals) indicate model violations.
6. Multiple linear regression
When more than one independent variable is included:
y = a + b₁x₁ + b₂x₂ + ... + bₙxₙ
Multicollinearity — a problem when independent variables are highly correlated with each other, making it difficult to estimate individual coefficients reliably.
7. Practical applications in business
Sales forecasting: regress sales on advertising spend, economic indicators, or seasonality variables. Cost analysis: regress total cost on output volume to separate fixed and variable costs (high-low method is a simple form). Revenue modelling: regress revenue on store size, staff numbers, or location characteristics. Extrapolation caution: predicting beyond the range of observed data is less reliable and should be clearly caveated.
