5 January 2027
| Explanatory variable | Outcome variable | Useful tools |
|---|---|---|
| categorical | categorical | contingency table, conditional probability |
| numerical | numerical | scatterplot, covariance, correlation, regression |
| categorical | numerical | group summaries, boxplots |
Method selection begins with the business question and variable types.
Two categorical variables are associated when the outcome distribution changes across explanatory-variable groups.
Compare conditional probabilities such as:
P(\text{tried}\mid\text{high income})
with the corresponding probabilities in other income groups or with the marginal probability.
Using a contingency table, ask:
A loyalty-app sample finds 72% high satisfaction among users and 61% among non-users.
A scatterplot reveals:
Put the explanatory variable X on the horizontal axis and response Y on the vertical axis.
Sample covariance measures joint movement:
s_{XY}=\frac{\sum_{i=1}^{n}(x_i-\bar{x})(y_i-\bar{y})}{n-1}.
Correlation standardises covariance:
r=\frac{s_{XY}}{s_Xs_Y},\qquad -1\le r\le1.
Excel: =CORREL(x_range,y_range).
Open week5-regression-data.xlsx, sheet IncomeSpend.
Spend against Income.=CORREL(Income,Spend).On HouseSize:
For each statement, identify the problem:
Population model:
Y_i=\beta_0+\beta_1X_i+\varepsilon_i.
Fitted sample model:
\hat Y_i=\hat\beta_0+\hat\beta_1X_i.
Residual:
e_i=Y_i-\hat Y_i.
Least squares chooses the line that minimises:
SSE=\sum_{i=1}^{n}e_i^2.
Squaring prevents positive and negative residuals cancelling and penalises larger errors more strongly.
The fitted line describes the mean response predicted at each X.
Include both variables, units, direction, and “on average”. Question whether X=0 is observed or meaningful.
For Spend as Y and Income as X:
Data → Data Analysis → Regression;From the Excel output:
To test for a linear relationship:
H_0:\beta_1=0,\qquad H_1:\beta_1\ne0.
Excel reports the two-sided p-value for the slope. A small p-value supports a population linear association, subject to model assumptions and study design.
R^2=\frac{\text{variation explained by the model}}{\text{total variation in }Y}.
Interpret R^2 as the percentage of sample variation in the response explained by its linear relationship with X.
It is not the percentage of observations predicted correctly.
For the fitted model
\widehat{Spend}=8254.575+0.296(Income),
substitute an income within the observed range. Keep units consistent and distinguish:
Regress market price (AUD thousands) on house age.
A useful linear model has residuals that are:
A pattern in residuals is information the model failed to capture.
Create predicted values and residuals for the age model.
Audit the IncomeSpend model:
Prediction outside the observed X range is extrapolation.
Before predicting spending for a household income of AUD 500,000:
A concise brief contains:
Using HouseSize, compare separate simple regressions of market price on:
Recommend the more useful single predictor using plots, correlation, R^2, slope evidence, residuals, and business reasoning. Explain why this is not yet a multiple-regression conclusion.
Can you: