Skip to content
SmartStudy

Linear Regression and Correlation

Modelling

Linear Regression and Correlation

Syllabus tag: KASNEB CPA | Advanced Level | CA34S1 Business Data Analytics

1. Correlation

Correlation measures the strength and direction of the linear relationship between two quantitative variables.

The Pearson correlation coefficient (r) ranges from –1 to +1:

  • r = +1: perfect positive linear relationship
  • r = 0: no linear relationship
  • r = –1: perfect negative linear relationship

Interpretation: |r| < 0.3 (weak), 0.3–0.7 (moderate), > 0.7 (strong).

Important: correlation does not imply causation. Two variables may be correlated because of a common cause (spurious correlation) or coincidence.

2. Simple linear regression

Simple linear regression models the relationship between one independent variable (x) and one dependent variable (y):

y = a + bx

Where: a is the y-intercept (value of y when x = 0) and b is the slope (change in y for each unit increase in x).

3. Calculating regression coefficients

Given n pairs of (x, y) values:

b = [n·Σxy – Σx·Σy] / [n·Σx² – (Σx)²]
a = ȳ – b·x̄

Where x̄ and ȳ are the means of x and y respectively.

4. The coefficient of determination (R²)

R² measures the proportion of variation in y explained by the regression model:

R² = r²     (for simple linear regression)

R² ranges from 0 to 1 (or 0% to 100%). R² = 0.75 means 75% of the variation in y is explained by x; the remaining 25% is unexplained (residual variation).

5. Residuals and model fit

A residual is the difference between the observed value of y and the predicted value: e = y – ŷ. A good regression model has residuals that are: small (low prediction error); randomly distributed (no pattern); normally distributed with mean zero. Heteroscedasticity (residuals increasing with x) and autocorrelation (patterns in residuals) indicate model violations.

6. Multiple linear regression

When more than one independent variable is included:

y = a + b₁x₁ + b₂x₂ + ... + bₙxₙ

Multicollinearity — a problem when independent variables are highly correlated with each other, making it difficult to estimate individual coefficients reliably.

7. Practical applications in business

Sales forecasting: regress sales on advertising spend, economic indicators, or seasonality variables. Cost analysis: regress total cost on output volume to separate fixed and variable costs (high-low method is a simple form). Revenue modelling: regress revenue on store size, staff numbers, or location characteristics. Extrapolation caution: predicting beyond the range of observed data is less reliable and should be clearly caveated.

Next in Business Data AnalyticsIndex Numbers and Time Series Analysis →