8 December 2026
Inference is only as credible as both the sampling process and the analysis.
| Population parameter | Sample statistic |
|---|---|
| mean \mu | mean \bar{x} |
| standard deviation \sigma | standard deviation s |
Greek letters describe the population; Roman letters describe the observed sample.
\hat p=\frac{X}{n}.
A proportion must lie between 0 and 1.
Bias is a systematic tendency to over- or under-represent the target quantity.
Common sources include:
Increasing n reduces random variation, not systematic bias.
An app asks active users, “How satisfied are you with our excellent new design?”
Identify:
A smartwatch company emails only customers who completed a workout yesterday and asks whether its product increases exercise.
If we repeatedly take random samples of the same size, we get different values of \bar{x} and \hat p.
A sampling distribution describes those possible statistic values—not the distribution of individual observations.
This sampling variation is the uncertainty quantified by standard errors and confidence intervals.
As a random sample grows, the sample mean tends to settle closer to the population mean.
For sufficiently large random samples, the sampling distribution of the mean is approximately Normal:
\bar X\approx N\left(\mu,\frac{\sigma}{\sqrt n}\right).
The approximation concerns sample means, even when individual observations are not Normal.
The standard deviation of the sampling distribution is the standard error:
SE(\bar X)=\frac{\sigma}{\sqrt n}\quad\text{or estimated by}\quad\frac{s}{\sqrt n}.
As n quadruples, the standard error halves. Precision improves with the square root of sample size.
Suppose \sigma=20.
The population variability did not change; uncertainty in the sample mean did.
Assume transaction values have \sigma=48.
For a sufficiently large random sample:
\hat p\approx N\left(\pi,\sqrt{\frac{\pi(1-\pi)}{n}}\right).
For estimation, check using the observed proportion:
n\hat p\ge5,\qquad n(1-\hat p)\ge5.
Every confidence interval has the same architecture:
\text{estimate}\pm\text{critical value}\times\text{standard error}.
For a population mean:
\bar{x}\pm z^*\frac{\sigma}{\sqrt n}\qquad(\sigma\text{ known}),
\bar{x}\pm t^*\frac{s}{\sqrt n}\qquad(\sigma\text{ unknown},\ df=n-1).
In business applications, \sigma is usually unknown, so the t interval is common.
For confidence level 1-\alpha:
=NORM.S.INV(1-alpha/2)=T.INV(1-alpha/2,n-1)For 95% confidence, z^*\approx1.96. A t^* value is larger for small samples and approaches 1.96 as n grows.
An audit sample has n=80, \bar{x}=86.05, and known \sigma=22.38.
In Excel:
=NORM.S.INV(0.975);=22.38/SQRT(80);Interpret the interval for the population mean balance.
Use n=80, \bar{x}=86.05, and s=24.02.
A 95% confidence method captures the fixed population parameter in about 95% of repeated random samples.
For the interval already calculated, say:
We are 95% confident that the population mean lies between the calculated bounds.
Do not say 95% of individual observations lie inside the interval.
Repeat the audit interval at 90% and 99% confidence.
When the large-count condition is satisfied:
\hat p\pm z^*\sqrt{\frac{\hat p(1-\hat p)}{n}}.
The standard error uses \hat p because the unknown \pi is the quantity being estimated.
Of 600 surveyed employers, 126 used a recruitment service recently.
Open week3-proportions-data.xlsx, sheet Marketing.
Separate confidence intervals can describe subgroup proportions, but visual overlap is not a formal test of a difference.
Before comparing groups, ask:
For female and male respondents separately:
Create labelled input cells for:
Keep inputs visually separate from formulas. Add an assumption check and a plain-language conclusion beneath the calculations.
Using your template, hold the estimate fixed and compare:
Write two general rules a manager could use when planning a study.
Can you: