Probability and Normal Distribution
Statistics
Probability and Normal Distribution
Syllabus tag: KASNEB CPA | Advanced Level | CA34S1 Business Data Analytics
1. Basic probability concepts
Probability is the measure of the likelihood that an event will occur, expressed as a value between 0 (impossible) and 1 (certain).
Classical probability: P(event) = number of favourable outcomes / total equally likely outcomes.
Relative frequency probability: based on observed frequency over many trials — P(event) = frequency of event / total observations.
Conditional probability: P(A|B) = P(A and B) / P(B) — the probability of A given that B has occurred.
Mutually exclusive events: cannot occur simultaneously — P(A or B) = P(A) + P(B).
Independent events: the occurrence of one does not affect the probability of the other — P(A and B) = P(A) × P(B).
2. The normal distribution
The normal distribution (Gaussian distribution) is the most important probability distribution in statistics. It is completely described by two parameters: mean (μ) and standard deviation (σ). Properties:
- Bell-shaped and symmetric about the mean
- Mean = Median = Mode
- The total area under the curve = 1
- Empirical rule: approximately 68.3% of values lie within 1σ of the mean; 95.4% within 2σ; 99.7% within 3σ
3. The standard normal distribution (z-distribution)
The standard normal distribution has μ = 0 and σ = 1. Any normal distribution can be standardised using the z-score:
z = (x – μ) / σ
The z-score tells us how many standard deviations a value is above or below the mean.
4. Using standard normal tables
Standard normal tables give the area under the curve from z = 0 to a specified z value (one-tailed area). To find P(X < x):
- If z > 0: P(X < x) = 0.5 + table area for z
- If z < 0: P(X < x) = 0.5 – table area for |z|
To find P(X > x) = 1 – P(X < x)
Worked example: X ~ N(50, 10²). Find P(X < 58).
z = (58 – 50) / 10 = 0.8
Table area for z=0.80 = 0.2881
P(X < 58) = 0.5 + 0.2881 = 0.7881
P(X > 58) = 1 – 0.7881 = 0.2119
5. Business applications of the normal distribution
Quality control: specifying acceptable ranges (control limits) for product measurements. Financial modelling: modelling returns on investments (though real returns often have fat tails). Risk management: Value at Risk (VaR) calculations assume normally distributed returns. HR: performance ratings, aptitude test scores often approximately normally distributed. Inventory management: demand forecasting with safety stock calculations.
6. The Central Limit Theorem (CLT)
The CLT states that the distribution of sample means approaches a normal distribution as the sample size increases, regardless of the shape of the underlying population distribution. This is fundamental to statistical inference — it justifies using normal distribution tables for hypothesis testing and confidence interval construction even when the population is not normal (provided n ≥ 30).
