1 December 2026
Have Tried categorical or numerical?Income categorical or numerical?| Business wording | Notation | Meaning |
|---|---|---|
| A and B | A\cap B | both occur |
| A or B | A\cup B | at least one occurs |
| not A | A^c | complement of A |
P(A^c)=1-P(A)
P(A\cup B)=P(A)+P(B)-P(A\cap B)
For an online order, define:
Translate:
Let C mean a customer clicks an advertisement and P mean they purchase.
A contingency table cross-classifies observations by two categorical variables.
Open week2-categorical-data.xlsx and select Marketing.
Create a PivotTable:
Income in Rows;Have Tried in Columns;Person ID in Values; andCheck that all cells sum to 100%.
Always state the denominator:
P(A)=\frac{\text{observations satisfying }A}{\text{all observations}}.
Using the PivotTable, calculate and label:
For each answer, identify whether it is marginal, joint, or union probability.
Conditional probability changes the denominator:
P(A\mid B)=\frac{P(A\cap B)}{P(B)}.
Read P(A\mid B) as “the probability of A, given B”. The event after the vertical bar defines the group being considered.
To calculate
P(\text{income}>50{,}000\mid\text{not tried}),
either:
Have Tried is in columns.Verify that each conditioning column sums to 100%.
Calculate both:
Explain why these are different questions. Which denominator is used in each calculation?
Events A and B are independent when any equivalent condition holds:
P(A\mid B)=P(A),
P(B\mid A)=P(B),
P(A\cap B)=P(A)P(B).
A small difference may arise from sampling variation; it is not proof of causation.
Create a second PivotTable for Gender and Have Tried.
Compare:
P(\text{tried}\mid\text{female})\quad\text{with}\quad P(\text{tried}).
Narrate the conclusion as an association in this sample, not a causal effect.
Use Live Alone and Have Tried.
If X\sim N(\mu,\sigma), the distribution is:
Approximately 68%, 95%, and 99.7% lie within one, two, and three standard deviations of the mean.
| Question | Excel function |
|---|---|
| P(X\le x) | NORM.DIST(x, mean, sd, TRUE) |
| P(X>x) | 1-NORM.DIST(x, mean, sd, TRUE) |
| percentile with lower-tail probability p | NORM.INV(p, mean, sd) |
| standard Normal percentile | NORM.S.INV(p) |
For an interval, subtract two cumulative probabilities.
Assume adult systolic blood pressure is Normal with mean 124 and standard deviation 10.
=NORM.DIST(117,124,10,TRUE)=1-NORM.DIST(140,124,10,TRUE)=NORM.INV(0.85,124,10)Sketch and shade the requested area before typing the formula.
Assume service time is Normal with mean 8 minutes and standard deviation 1.5 minutes.
The standard score
z=\frac{x-\mu}{\sigma}
counts how many standard deviations x is above or below the mean.
A store’s weekly sales are 1.4 standard deviations above the chain mean. Its profit is 0.6 standard deviations above the chain mean.
Ask in order:
The method follows the information available—not the desired answer.
Choose Normal, standard Normal, t, or “not enough information”:
For “Are high-income respondents more likely to have tried the product?”:
The conclusion should answer the business question, not narrate menu clicks.
Using the Marketing worksheet, choose one demographic variable and answer:
Prepare a three-sentence briefing and one clearly labelled table.
NORM.DIST without checking the tail or interval requested.Can you: