ETX1100/ETX5900 Business Statistics

Week 3: Sampling and estimation

8 December 2026

Today’s journey

  1. Connect population questions to random samples.
  2. Recognise bias and sampling variation.
  3. Use the law of large numbers and central limit theorem.
  4. Build confidence intervals for means and proportions.
  5. Explain precision, confidence, and limitations in business language.

Samples and sampling

Retrieval check

  • What is the difference between a parameter and a statistic?
  • What does P(A\mid B) do to the denominator?
  • When is a t distribution useful?
  • Why does a large dataset not automatically represent its target population?

Population to sample and back

  1. Define the target population.
  2. Select a sample using a defensible process.
  3. Calculate a sample statistic.
  4. Use the statistic to estimate a population parameter.
  5. Quantify uncertainty and communicate limitations.

Inference is only as credible as both the sampling process and the analysis.

Notation for means

Population parameter Sample statistic
mean \mu mean \bar{x}
standard deviation \sigma standard deviation s

Greek letters describe the population; Roman letters describe the observed sample.

Notation for proportions

  • Population proportion: \pi.
  • Sample count with the characteristic: X.
  • Sample size: n.
  • Sample proportion:

\hat p=\frac{X}{n}.

A proportion must lie between 0 and 1.

Sampling bias

Bias is a systematic tendency to over- or under-represent the target quantity.

Common sources include:

  • undercoverage;
  • voluntary response;
  • non-response;
  • convenience sampling; and
  • leading or ambiguous measurement.

Increasing n reduces random variation, not systematic bias.

Watch me: diagnose a survey

An app asks active users, “How satisfied are you with our excellent new design?”

Identify:

  • target population;
  • sampling frame;
  • likely missing users;
  • wording bias; and
  • a more credible design.

Now you try: redesign the evidence

A smartwatch company emails only customers who completed a workout yesterday and asks whether its product increases exercise.

  1. Identify two sources of bias.
  2. Explain why a very large response does not fix them.
  3. Propose a sampling and measurement plan.
  4. State the population to which your conclusion would apply.

Sampling distributions

Statistics vary from sample to sample

If we repeatedly take random samples of the same size, we get different values of \bar{x} and \hat p.

A sampling distribution describes those possible statistic values—not the distribution of individual observations.

This sampling variation is the uncertainty quantified by standard errors and confidence intervals.

Law of large numbers

As a random sample grows, the sample mean tends to settle closer to the population mean.

  • It does not say every larger sample is closer.
  • It does not repair biased sampling.
  • It describes long-run behaviour under repeated random sampling.

Central limit theorem

For sufficiently large random samples, the sampling distribution of the mean is approximately Normal:

\bar X\approx N\left(\mu,\frac{\sigma}{\sqrt n}\right).

The approximation concerns sample means, even when individual observations are not Normal.

Standard error of a mean

The standard deviation of the sampling distribution is the standard error:

SE(\bar X)=\frac{\sigma}{\sqrt n}\quad\text{or estimated by}\quad\frac{s}{\sqrt n}.

As n quadruples, the standard error halves. Precision improves with the square root of sample size.

Watch me: compare sample sizes

Suppose \sigma=20.

  • At n=25, SE=20/\sqrt{25}=4.
  • At n=100, SE=20/\sqrt{100}=2.

The population variability did not change; uncertainty in the sample mean did.

Now you try: sampling variation

Assume transaction values have \sigma=48.

  1. Calculate SE(\bar X) for n=36 and n=144.
  2. Which sample mean is more precise?
  3. How much must n change to halve the standard error?
  4. Would either calculation remove non-response bias?

Sampling distribution of a proportion

For a sufficiently large random sample:

\hat p\approx N\left(\pi,\sqrt{\frac{\pi(1-\pi)}{n}}\right).

For estimation, check using the observed proportion:

n\hat p\ge5,\qquad n(1-\hat p)\ge5.

Confidence intervals for means

Point estimate plus margin of error

Every confidence interval has the same architecture:

\text{estimate}\pm\text{critical value}\times\text{standard error}.

  • The point estimate identifies the centre.
  • The margin of error quantifies sampling uncertainty.
  • The confidence level controls the long-run coverage rate.

Known versus unknown \sigma

For a population mean:

\bar{x}\pm z^*\frac{\sigma}{\sqrt n}\qquad(\sigma\text{ known}),

\bar{x}\pm t^*\frac{s}{\sqrt n}\qquad(\sigma\text{ unknown},\ df=n-1).

In business applications, \sigma is usually unknown, so the t interval is common.

Critical values in Excel

For confidence level 1-\alpha:

  • z^*: =NORM.S.INV(1-alpha/2)
  • t^*: =T.INV(1-alpha/2,n-1)

For 95% confidence, z^*\approx1.96. A t^* value is larger for small samples and approaches 1.96 as n grows.

Watch me: audit balances, known \sigma

An audit sample has n=80, \bar{x}=86.05, and known \sigma=22.38.

In Excel:

  1. compute =NORM.S.INV(0.975);
  2. compute =22.38/SQRT(80);
  3. multiply for the margin of error; and
  4. calculate lower and upper bounds.

Interpret the interval for the population mean balance.

Now you try: audit balances, unknown \sigma

Use n=80, \bar{x}=86.05, and s=24.02.

  1. Construct a 95% t interval.
  2. Identify df and the Excel critical-value formula.
  3. State the margin of error.
  4. Interpret the interval in dollars and in context.

Interpreting confidence correctly

A 95% confidence method captures the fixed population parameter in about 95% of repeated random samples.

For the interval already calculated, say:

We are 95% confident that the population mean lies between the calculated bounds.

Do not say 95% of individual observations lie inside the interval.

Now you try: confidence versus precision

Repeat the audit interval at 90% and 99% confidence.

  1. Rank the three intervals by width.
  2. Explain the confidence–precision trade-off.
  3. Predict what happens if n decreases.
  4. Identify one aspect of data quality that interval width does not capture.

Confidence intervals for proportions

Interval for a population proportion

When the large-count condition is satisfied:

\hat p\pm z^*\sqrt{\frac{\hat p(1-\hat p)}{n}}.

The standard error uses \hat p because the unknown \pi is the quantity being estimated.

Watch me: recruitment-service use

Of 600 surveyed employers, 126 used a recruitment service recently.

  1. \hat p=126/600=0.21.
  2. Check 600(0.21) and 600(0.79).
  3. Use z^*=1.96 for 95% confidence.
  4. Calculate the standard error and bounds.
  5. Format the result as percentages.

Now you try: product trial proportion

Open week3-proportions-data.xlsx, sheet Marketing.

  1. Find n, the count who have tried the product, and \hat p.
  2. Check the large-count condition.
  3. Construct a 95% confidence interval for \pi.
  4. Does the interval suggest that more than half the market has tried it?
  5. State a conclusion and one sampling limitation.

Compare groups carefully

Separate confidence intervals can describe subgroup proportions, but visual overlap is not a formal test of a difference.

Before comparing groups, ask:

  • Were both groups sampled credibly?
  • Are sample sizes adequate?
  • Are categories defined consistently?
  • Is the practical difference meaningful?

Now you try: subgroup intervals

For female and male respondents separately:

  1. calculate n, X, and \hat p for product trial;
  2. construct a 95% interval for each population proportion;
  3. compare centres and widths; and
  4. write a cautious conclusion without claiming causation.

Reusable estimation workflow

Watch me: build an Excel template

Create labelled input cells for:

  • estimate, standard deviation or count, n, and confidence level;
  • critical value and standard error;
  • margin of error; and
  • lower and upper bounds.

Keep inputs visually separate from formulas. Add an assumption check and a plain-language conclusion beneath the calculations.

Now you try: scenario analysis

Using your template, hold the estimate fixed and compare:

  1. 90%, 95%, and 99% confidence;
  2. sample sizes n=50, 200, and 800; and
  3. mean versus proportion standard errors.

Write two general rules a manager could use when planning a study.

Common traps

  • Confusing the distribution of observations with the sampling distribution.
  • Assuming large n removes selection or measurement bias.
  • Using z when \sigma is unknown without justification.
  • Forgetting the square root of n.
  • Interpreting a confidence interval as containing most individual values.
  • Reporting bounds without units, population, or business meaning.

Closing check

Can you:

  • distinguish bias from random sampling variation?
  • explain the law of large numbers and central limit theorem?
  • calculate and interpret a standard error?
  • choose a mean or proportion interval and its critical value?
  • explain how confidence level and sample size affect width?
  • communicate an interval as evidence, not certainty?