ETX1100/ETX5900 Business Statistics

Week 8: Integrated business case and revision

26 January 2027

Today’s journey

  1. Select methods from a business question and variable types.
  2. Analyse ETF’s supermarket staffing and profit case.
  3. Turn output into recommendations and limitations.
  4. Correct common statistical misconceptions.
  5. Consolidate the six learning outcomes for the examination.

Choosing an analysis

The four purposes

Purpose Core question Example method
Describe What happened in this sample? tables, charts, summaries
Infer What does the sample suggest about a population? intervals, tests
Predict What outcome is expected for given inputs? regression
Optimise What controllable decision is best? Data Tables, Goal Seek, Solver

Watch me: unpack a business question

For “Should stores introduce manager-retention incentives to improve profit?” identify:

  1. outcome and candidate explanatory variables;
  2. variable types and observational unit;
  3. descriptive and relationship methods;
  4. whether causal language is justified; and
  5. information needed before a decision.

Now you try: method selection

Choose a purpose and method for each question:

  1. What percentage of stores open 24 hours?
  2. Is population mean service quality above 80?
  3. Which staffing factors predict profit?
  4. What advertising allocation maximises profit under a budget?

Justify each choice from the variables and claim.

Supermarket case: understand the data

Business context

A supermarket chain is considering incentives intended to improve store performance.

Open week8-case-data.xlsx.

Each row represents a store. Variables cover:

  • profit and operating context;
  • manager and crew tenure;
  • manager and crew skill; and
  • service quality.

Data dictionary highlights

Variable Meaning
Profit current-year store profit
MTenure, CTenure average manager and crew tenure in months
MgrSkill, CrewSkill staffing skill measures
ServQual service-quality score
Pop, Comp, Visibility, PedCount, Res, Hours24 store context

Confirm units and coding in the workbook before analysing.

Watch me: audit before modelling

  1. Check row count, headers, and missing values.
  2. Classify each variable.
  3. Inspect ranges and units.
  4. Create descriptive statistics for key numerical variables.
  5. Plot Profit and ServQual for unusual values.

Document—not silently delete—anything suspicious.

Now you try: data-quality brief

Create a compact audit containing:

  1. observational unit and target population;
  2. roles and types for five variables;
  3. two useful descriptive statistics for Profit;
  4. one appropriate visualisation; and
  5. two data-quality or scope limitations.

Explore relationships

Staffing questions

The ETF case asks:

  • Are manager and crew skills associated?
  • Do crews stay longer where managers stay longer?
  • Are skill measures associated with service quality?
  • Is service quality associated with profit?

These are observational association questions—not experiments.

Watch me: correlation matrix

Use Data → Data Analysis → Correlation for selected numerical variables.

  1. Include labels.
  2. Confirm the output is symmetric with 1s on the diagonal.
  3. Pair strong correlations with scatterplots.
  4. Describe direction, strength, form, and unusual points.
  5. Avoid ranking variables by correlation alone.

Now you try: prioritise relationships

From the correlation matrix:

  1. identify the strongest positive and negative relationships;
  2. create two scatterplots;
  3. flag one possible multicollinearity concern;
  4. identify one weak relationship that may still matter in a multiple model; and
  5. write two cautious business interpretations.

High service-quality stores

ETF defines excellent service as ServQual >= 90.

Comparing skill across this threshold can be useful descriptively, but dichotomising a numerical variable discards information. Use it to complement—not replace—the continuous analysis.

Now you try: compare service groups

Create an ExcellentService flag and compare CrewSkill and MgrSkill across groups.

Use appropriate centre, spread, and plots. State what the comparison suggests and why it does not establish that skill caused service quality.

Build and refine models

Model service quality

Begin with a focused explanatory question:

ServQual_i=\beta_0+\beta_1MTenure_i+\beta_2CTenure_i+\beta_3MgrSkill_i+\beta_4CrewSkill_i+\varepsilon_i.

The chosen variables should reflect a rationale, not merely every available column.

Watch me: service-quality model

  1. Fit the model in Excel.
  2. Write the fitted equation.
  3. Interpret two partial slopes.
  4. Evaluate adjusted R^2 and Significance F.
  5. Inspect coefficient p-values and residual plots.
  6. Record limitations before considering reduction.

Now you try: refine the service model

Create one defensible reduced model.

Compare it with the full model using coefficient stability, adjusted R^2, standard error, individual tests, overall evidence, residuals, and interpretability. Recommend one model.

Model profit in context

Profit may depend on staffing and store environment. Candidate predictors include tenure, skills, service quality, population, competition, visibility, pedestrian count, residential context, and 24-hour operation.

Adding controls can change staffing coefficients because store context is not evenly distributed.

Watch me: profit model

Build a full model with a clearly stated rationale.

  • Identify dummy-variable coding.
  • Interpret staffing slopes conditionally.
  • Distinguish statistical from practical significance.
  • Check overall fit and residuals.
  • Consider whether correlated predictors make individual coefficients unstable.

Now you try: decision model

Develop a reduced profit model that supports the incentive question.

Your output must include:

  1. fitted equation and variable coding;
  2. three coefficient interpretations;
  3. coefficient and overall tests;
  4. adjusted R^2 and diagnostics;
  5. a prediction for one observed store; and
  6. a comparison with its actual profit.

From analysis to recommendation

Evidence chain

A defensible recommendation connects:

question → data → method → result → limitation → action

Weak reports jump directly from a significant coefficient to a policy. Strong reports explain the comparison, uncertainty, assumptions, feasible action, and what should be tested next.

ETF case implications

Possible evidence may suggest that tenure and skill are positively associated with performance, but:

  • tenure develops over time;
  • incentives may affect retention rather than tenure directly;
  • omitted store characteristics may confound relationships;
  • current data are observational; and
  • an incentive programme has costs not represented in the regression.

Now you try: executive recommendation

Write a five-sentence executive recommendation:

  1. decision and strength of recommendation;
  2. most relevant quantitative evidence;
  3. practical meaning;
  4. principal limitation; and
  5. a pilot or additional evidence that would reduce uncertainty.

Statistical misconception clinic

Misconception 1: conditional probability

“Because 70% of app users are highly satisfied, 70% of highly satisfied customers use the app.”

The statement reverses the condition:

P(Satisfied\mid App)\ne P(App\mid Satisfied).

The denominators differ.

Now you try: correct the conditional claim

Rewrite the claim correctly and describe the additional contingency-table information needed to calculate the reversed conditional probability.

Misconception 2: confidence intervals

“There is a 95% probability that this fixed population mean is inside our calculated interval.”

Frequentist confidence attaches 95% to the long-run method, not a probability assigned to the fixed parameter after observing the interval.

Now you try: repair the interval statement

Write a correct contextual interpretation and explain why the interval does not describe where 95% of individual customers lie.

Misconception 3: regression coefficients

“The manager-skill coefficient is the raw difference in profit between high- and low-skill stores.”

In multiple regression it is a conditional comparison, holding other included predictors constant. Its unit and coding must be stated.

Now you try: repair the coefficient claim

Choose one coefficient from your profit model and provide:

  1. a correct interpretation;
  2. the base category if relevant; and
  3. one reason it should not be described as causal.

Misconception 4: evidence and importance

“Because p<0.05, the effect is large, useful, and the null hypothesis has a less than 5% chance of being true.”

A p-value addresses compatibility with H_0 under assumptions. Effect size, uncertainty, costs, and practical thresholds require separate evidence.

Now you try: repair the evidence claim

Rewrite the claim using the estimate, confidence interval, and business threshold. State what decision would follow if the effect is statistically clear but too small to matter.

Examination consolidation

Learning-outcome map

  1. Classify and visualise data.
  2. Calculate and interpret probability.
  3. Sample, estimate, and test claims.
  4. Analyse relationships.
  5. Build simple and multiple regression in Excel.
  6. Formulate and solve spreadsheet business models.

The examination can ask you to select, execute, interpret, or critique—not merely recall a formula.

Watch me: decode an exam question

Underline:

  • the population or decision target;
  • variable types and units;
  • claim direction;
  • supplied sample information;
  • requested method and output; and
  • required contextual conclusion.

Then sketch the workflow before calculating.

Now you try: rapid method map

For eight instructor-provided mini-scenarios, record only:

  1. parameter, outcome, or objective;
  2. appropriate method;
  3. key assumption or constraint;
  4. Excel tool or function; and
  5. conclusion template.

Compare answers before doing any arithmetic.

Final checklist

  • Define symbols and retain units.
  • Show the formula or Excel method used.
  • Check assumptions before interpreting output.
  • Use the correct tail, denominator, or base category.
  • Distinguish sample evidence from population claims.
  • Avoid causal wording without a causal design.
  • State limitations and answer the business question.

Closing reflection

Can you move independently through:

business question → method → Excel → validation → conclusion → limitation → recommendation?

That complete reasoning chain—not an isolated calculation—is the central skill of Business Statistics.