ETX1100/ETX5900 Business Statistics

Week 4: Testing business claims

15 December 2026

Today’s journey

  1. Translate a business claim into hypotheses.
  2. Use one five-step testing process.
  3. Test population means and proportions in Excel.
  4. Interpret p-values, decisions, and possible errors.
  5. Separate statistical evidence from practical importance.

From claims to hypotheses

Retrieval check

  • What does a confidence interval estimate?
  • What is a standard error?
  • When do we use t rather than z for a mean?
  • Why must hypotheses refer to population parameters rather than sample statistics?

What hypothesis testing does

Hypothesis testing asks whether sample evidence is difficult to reconcile with a stated population benchmark.

  • Begin with a business claim.
  • Assume a null benchmark temporarily.
  • Measure how far the sample result is from that benchmark.
  • Quantify extremeness with a p-value.
  • Make a decision with a pre-selected significance level.

The five-step workflow

  1. State H_0 and H_1.
  2. Calculate the test statistic.
  3. Calculate the p-value in the correct tail or tails.
  4. Compare the p-value with \alpha.
  5. State a contextual conclusion about H_1.

Null and alternative hypotheses

  • H_0 contains equality and represents the benchmark.
  • H_1 expresses the claim being investigated.
  • Right-sided: H_1:\theta>\theta_0.
  • Left-sided: H_1:\theta<\theta_0.
  • Two-sided: H_1:\theta\ne\theta_0.

Choose direction from the question before seeing the result.

Watch me: translate a click-through claim

An analyst claims the population mean click-through rate is below 12 clicks per minute.

H_0:\mu=12,\qquad H_1:\mu<12.

The parameter is \mu, not the observed \bar{x}. The word below makes this a left-sided test.

Now you try: write the hypotheses

Write hypotheses and identify the tail for each claim:

  1. Average delivery time exceeds 30 minutes.
  2. Mean transaction value differs from AUD 85.
  3. Fewer than 4% of invoices contain an error.
  4. More than half of customers use the loyalty app.

Tests for a population mean

Test statistic for a mean

When \sigma is unknown:

t=\frac{\bar{x}-\mu_0}{s/\sqrt n},\qquad df=n-1.

The statistic is the observed difference measured in standard errors. Its sign indicates direction; its magnitude indicates distance from H_0.

P-values and tails in Excel

Alternative Excel p-value
H_1:\mu<\mu_0 T.DIST(t,df,TRUE)
H_1:\mu>\mu_0 1-T.DIST(t,df,TRUE)
H_1:\mu\ne\mu_0 T.DIST.2T(ABS(t),df)

Using ABS with T.DIST.2T avoids separate positive/negative formulas.

What a p-value means

The p-value is the probability—assuming H_0 is true—of obtaining a test statistic at least as extreme as the one observed, in the direction specified by H_1.

It is not:

  • the probability that H_0 is true;
  • the size or importance of an effect; or
  • the probability that the result occurred “by chance”.

Watch me: test click-through rate

Given n=45, \bar{x}=10.8, s=3.2, \mu_0=12, and \alpha=0.05:

  1. t=(10.8-12)/(3.2/\sqrt{45});
  2. use =T.DIST(t,44,TRUE);
  3. compare the p-value with 0.05; and
  4. conclude about the population mean click-through rate.

Now you try: discretionary spending

A random sample of 40 students has \bar{x}=283 and s=17.85. Test at \alpha=0.05 whether population mean monthly discretionary spending exceeds AUD 275.

Complete all five steps, show the Excel p-value formula, and state the conclusion without saying “accept H_0”.

One-sided versus two-sided

A two-sided test considers evidence in both directions and uses both tails. A one-sided test concentrates the rejection region in one pre-specified direction.

Do not choose one-sided after observing the sample. That inflates the chance of finding a convenient result.

Now you try: change only the claim

Using the same spending sample, test whether the mean differs from AUD 275.

  1. Which steps change?
  2. Does the test statistic change?
  3. Which p-value formula changes?
  4. Compare the one- and two-sided conclusions.

Tests for a population proportion

Large-sample proportion test

For H_0:\pi=\pi_0:

z=\frac{\hat p-\pi_0}{\sqrt{\pi_0(1-\pi_0)/n}}.

Check:

n\pi_0\ge5,\qquad n(1-\pi_0)\ge5.

The null value \pi_0 appears in the test standard error.

Proportion p-values in Excel

Alternative Excel p-value
H_1:\pi<\pi_0 NORM.S.DIST(z,TRUE)
H_1:\pi>\pi_0 1-NORM.S.DIST(z,TRUE)
H_1:\pi\ne\pi_0 2*(1-NORM.S.DIST(ABS(z),TRUE))

The tail follows the alternative hypothesis, not merely the sign of z.

Watch me: transaction success rate

A bank benchmarks error-free processing at 95%. In a random sample of 250 transactions, 229 succeed.

  1. \hat p=229/250=0.916.
  2. Test H_0:\pi=0.95 against H_1:\pi<0.95.
  3. Verify the large-count condition using \pi_0.
  4. Calculate z and its left-tail p-value.
  5. Conclude at \alpha=0.10.

Now you try: invoice accuracy

A company claims at least 98% of invoices are accurate. A random audit finds 288 accurate invoices out of 300.

Test at \alpha=0.05 whether the population accuracy rate is below 98%.

Include the assumption check, statistic, Excel p-value, decision, and business conclusion.

Confidence intervals and tests agree

For a two-sided test at significance \alpha:

  • reject H_0:\theta=\theta_0 when the corresponding (1-\alpha) confidence interval excludes \theta_0;
  • fail to reject when it includes \theta_0.

The interval also communicates plausible effect sizes, which the p-value alone does not.

Now you try: interval cross-check

For the invoice sample:

  1. construct a 95% confidence interval for \pi;
  2. compare it with the 98% benchmark;
  3. explain why this is a two-sided cross-check, while the original claim was left-sided; and
  4. identify what the interval adds to the decision.

Decisions, errors, and importance

Decision language

If p-value <\alpha:

Reject H_0; there is sufficient evidence that …

If p-value \ge\alpha:

Fail to reject H_0; there is insufficient evidence that …

“Fail to reject” is not proof that the null is true.

Type I and Type II errors

Decision H_0 true H_0 false
Reject H_0 Type I error correct
Fail to reject H_0 correct Type II error
  • P(\text{Type I error})=\alpha.
  • Reducing \alpha makes rejection harder and can increase Type II risk if other factors stay fixed.

Watch me: errors in context

For H_1:\mu>275:

  • Type I: conclude mean spending exceeds AUD 275 when it does not.
  • Type II: fail to find evidence that mean spending exceeds AUD 275 when it actually does.

Name the real-world cost of each error before choosing \alpha.

Now you try: business consequences

For the invoice-accuracy test:

  1. describe Type I and Type II errors in context;
  2. identify who bears the cost of each;
  3. recommend whether 10%, 5%, or 1% significance is most defensible; and
  4. justify the recommendation without calculating a new test.

Statistical versus practical significance

  • Statistical significance asks whether the result is difficult to explain by sampling variation under H_0.
  • Practical significance asks whether the effect is large enough to matter.
  • Large samples can make tiny effects statistically significant.
  • Report an effect estimate and confidence interval alongside the p-value.

Now you try: challenge the headline

A report says, “Waiting time fell by 0.2 minutes, p<0.001, so the redesign was a major success.”

Identify:

  1. what the p-value supports;
  2. what it does not support;
  3. additional quantities needed; and
  4. a more defensible conclusion.

Integrated evidence brief

Watch me: audit a testing worksheet

Check that the spreadsheet separates:

  • inputs and units;
  • hypotheses and chosen \alpha;
  • assumption checks;
  • statistic and p-value;
  • decision; and
  • contextual conclusion.

Trace formulas to ensure the sample statistic, benchmark, and tail are consistent.

Now you try: complete the evidence brief

Choose either a mean claim from week4-testing-data.xlsx or a product-trial proportion from week4-proportions-data.xlsx.

Produce a one-page brief containing the five test steps, an interval or effect estimate, one possible error, and a two-sentence recommendation.

Common traps

  • Writing hypotheses with \bar{x} or \hat p.
  • Putting equality in H_1.
  • Choosing the tail after seeing the data.
  • Using \hat p rather than \pi_0 in a test standard error.
  • Saying a p-value is the probability that H_0 is true.
  • Saying “accept H_0”.
  • Treating statistical significance as business importance.

Closing check

Can you:

  • translate a claim into directional hypotheses?
  • complete the five-step process for a mean or proportion?
  • select the correct Excel tail?
  • state a decision and conclusion precisely?
  • describe Type I and Type II errors in context?
  • distinguish evidence strength from effect importance?