Descriptive Statistics — Central Tendency and Dispersion
Statistics
Descriptive Statistics — Central Tendency and Dispersion
Syllabus tag: KASNEB CPA | Advanced Level | CA34S1 Business Data Analytics
1. Measures of central tendency
Measures of central tendency describe the "centre" of a data distribution.
Arithmetic mean (average): sum of all values divided by the number of values.
Mean (x̄) = Σx / n
Sensitive to outliers — a single extreme value significantly affects the mean.
Median: the middle value when data is arranged in ascending or descending order. For an even number of values, the median is the average of the two middle values. Less affected by outliers than the mean.
Mode: the most frequently occurring value. A dataset can have no mode, one mode (unimodal), or multiple modes (bimodal, multimodal). Used for categorical data.
2. Measures of dispersion
Measures of dispersion describe how spread out the data values are.
Range: maximum value – minimum value. Simple but ignores all values except the two extremes.
Variance (population):
σ² = Σ(x – x̄)² / n
The average of the squared deviations from the mean. Squaring eliminates negatives and emphasises large deviations.
Standard deviation (population):
σ = √[Σ(x – x̄)² / n]
The square root of variance — expressed in the same units as the data. Measures the average distance of data points from the mean.
Sample standard deviation: uses (n–1) in the denominator (Bessel's correction) to provide an unbiased estimate of the population standard deviation:
s = √[Σ(x – x̄)² / (n–1)]
3. Coefficient of variation (CV)
The CV measures relative variability — useful for comparing variability between datasets with different means or units:
CV = (σ / x̄) × 100%
A higher CV indicates greater relative variability. Example: comparing the variability of sales in two branches with very different average sales volumes.
4. Skewness and kurtosis
Positive skew (right skew): the tail extends to the right; the mean > median > mode. Income distributions are typically positively skewed.
Negative skew (left skew): the tail extends to the left; the mean < median < mode.
Symmetric distribution: mean = median = mode (e.g. the normal distribution).
Kurtosis: measures the "peakedness" of the distribution. High kurtosis (leptokurtic) = fat tails and a sharp peak; low kurtosis (platykurtic) = thin tails and a flat peak.
5. Practical computation steps
For a dataset: 48, 55, 62, 49, 56, 60, 53, 57:
n = 8
Sum = 440
Mean = 440 / 8 = 55.0
Deviations: -7, 0, 7, -6, 1, 5, -2, 2
Squared deviations: 49, 0, 49, 36, 1, 25, 4, 4
Sum of squared deviations = 168
Population variance σ² = 168/8 = 21.0
Population SD σ = √21 = 4.58
CV = (4.58/55.0) × 100% = 8.33%
