9. Linear Model (Regression)

One predictor

Author
Affiliation

Johnny van Doorn

University of Amsterdam

Published

18 September 2026

In this lecture we discuss:

  • Linear model
    • Theory
    • Assumptions
    • In JASP
  • Predictions and model fit

Reading: Chapter 8 (§8.1–8.8)

Last week

Last week: correlation & NHST

  • \(COV_{xy}\): association between two variables, in the original measurement units
  • \(r\): completely standardized: \(r_{xy} = \frac{COV_{xy}}{s_x s_y}\)
  • Get \(p\) by converting to \(t\): \(t_r = \frac{r \sqrt{N-2}}{\sqrt{1 - r^2}}\)
  • Partial correlation: the same association, holding a third variable constant

Note on one-sided vs. two-sided testing

A choice in the analysis before seeing the data (i.e., informed by the research question, not the result).

  • Two-sided
    • Safer choice if unsure about the direction
    • “Spending \(\alpha\)” on two tails needs a more extreme test statistic to reach significance
  • One-sided
    • Blind to an effect in the opposite direction
    • Same \(\alpha\), but the \(\alpha \%\) rejection area is in 1 tail
    • Higher power, relevant when \(N\) is small (expensive studies, rare cases)

Additional rambling here

The data

Predict album sales (x 1,000 copies) based on the advertising budget (x £1,000).


The data

Two different units

A deviation of 100 means something different in each variable

Standardize

\[z_i = \frac{x_i - \bar{x}}{sd_x}\]

Mean = 0 and sd = 1, now the deviations are comparable.

Correlation is the covariance of z-scores

\[z_i = \frac{x_i - \bar{x}}{s_x} \qquad\qquad r_{xy} = \frac{COV_{xy}}{s_x s_y} = COV_{z_x z_y}\]

z.adverts <- (adverts - mean(adverts)) / sd(adverts)
z.sales   <- (sales   - mean(sales))   / sd(sales)

cov(z.adverts, z.sales)
[1] 0.5784877
cor(adverts, sales)
[1] 0.5784877

If we standardize both variables first, the covariance is the correlation.

Three levels of standardization

cov(adverts, sales)                               # original scale
[1] 22672.02
cov(adverts, sales) / (sd(adverts) * sd(sales))   # standardize x and y -> the correlation
[1] 0.5784877
cov(adverts, sales) / var(adverts)                # only x -> today's slope
[1] 0.09612449

Standardizing makes a statistic universally interpretable, but you lose information

Today

  • The regression slope \(b_1\): \(b_1 = r_{xy} \frac{s_y}{s_x}\)
  • Same trick for significance: coefficient \(b\) \(\rightarrow t \rightarrow p\)
  • Making predictions based on a model
  • One predictor today, multiple predictors in block 3

Regression
(one predictor)

Regression

\[\LARGE{\text{outcome} = \text{model prediction} + \text{error}}\]

In statistics, linear regression is a linear approach for modeling the relationship between a scalar dependent variable y and one or more explanatory variables denoted X. The case of one explanatory variable is called simple linear regression.

\[\LARGE{Y_i = \beta_0 + \beta_1 X_i + \epsilon_i}\]

In linear regression, the relationships are modeled using linear predictor functions whose unknown model parameters are estimated from the data.

Source: wikipedia

Assumptions of the Linear Model

A selection from Field (Section 8.4.1):

  • Sensitivity (Section 6.4)
    • Are there outliers?
  • Normality (Section 6.8)
    • Are the model errors non-normally distributed?
  • Linearity (Section 6.6)
    • Is the association non-linear?
  • Homoscedasticity (Section 6.7)
    • Is there systematic model error?

The greater the assumption violation, the less reliable your results are


Sensitivity

Outliers

  • Extreme residuals (e.g., > 3)
  • Cook’s distance (e.g., > 1)
  • Check Q-Q, residuals plots, casewise diagnostics

Sensitivity

Normality

  • Residuals should be normally distributed
  • Look at Q-Q plot standardized residuals. Most points should be on the diagonal
  • After the analysis is complete because it’s based on the residuals
  • See figure 6.25

figure 6.25

When does normality matter?

  • Normality doesn’t matter in large samples (thank you CLT)
  • Parameter estimation itself does not require normality
  • Small samples: check the Q-Q plot of the residuals

Field, Section 6.8.3

Linearity

  • The relationship between predictors and the outcome should be linear
  • Inspect using plots:
    • Scatter plots (or partial plots) of predictors vs. outcome
    • Residuals vs. predicted
  • See also Figure 6.28

Homoscedasticity

  • Variance of residuals should be equal across all expected values
  • Look at scatterplot of residuals vs. predicted values. Ideally it’s an uncorrelated, round cloud
  • For examples of violations, see here
  • After the analysis is complete because it’s based on the residuals

Calculate regression parameters

\[{sales}_i = b_0 + b_1 {adverts}_i + \epsilon_i\]

adverts  <- data$Adverts
sales    <- data$Sales

Calculate \(b_1\)

\[b_1 = r_{xy} \frac{s_y}{s_x}\]

# Calculate b1

cor.sales.adverts <- cor(sales,adverts)
sd.sales          <- sd(sales)
sd.adverts        <- sd(adverts)

b1 <- cor.sales.adverts * ( sd.sales / sd.adverts )
b1
[1] 0.09612449

Calculate \(b_0\)

\[b_0 = \bar{y} - b_1 \bar{x}\]

mean.sales    <- mean(sales)
mean.adverts  <- mean(adverts)

b0 <- mean.sales - b1 * mean.adverts
b0
[1] 134.1399

The slope


The slope

The slope - zoomed in

For every additional £1,000 spent on advertising, we predict sales to increase by 0.096 (x 1,000 copies)

Define regression equation

\[\widehat{sales} = {\text{model prediction}} = b_0 + b_1 {adverts}\] \[\widehat{sales} = {\text{model prediction}} = 134.14 + 0.096\times {adverts}\]

So now we can add the expected sales based on this model

prediction <- b0 + b1 * adverts
data$prediction <- round(prediction, 2)

Predicted values

Let’s have a look

\(y\) vs \(\hat{y}\)

And lets have a look at this relation between model prediction and observed

Error

The error (residual) is the difference between the model predictions and observed values

error <- sales - prediction
data$error <- round(error, 2)

Normality of Residuals

Are the residuals normally distributed?

Model fit

  • The fit of the model can be viewed in terms of the correlation (\(r\)) between the predictions and the observed values: if the predictions are perfect, the correlation will be 1.
  • For simple regression, this is equal to the correlation between adverts and sales. For multiple regression (block 3), these will differ.
r <- cor(prediction, sales)
r
[1] 0.5784877

Explained variance

Squaring this correlation gives the proportion of variance in sales that is explained by adverts:

r^2
[1] 0.3346481

Explained variance visually (\(n = 10\))

\(r^2\) is the proportion of blue to orange, while \(1 - r^2\) is the proportion of red to orange

Calculate t-values for b’s for hypothesis testing

We can also convert the slope \(b_1\) to a \(t\)-statistic, since that has a known sampling distribution:

\[\begin{aligned} t_{n-p-1} &= \frac{b_1 - \mu_{b_1}}{{SE}_{b_1}} \\ df &= n - p - 1 \\ \end{aligned}\]

Where \(b_1\) is the beta coefficient, \({SE}\) is the standard error of the beta coefficient, \(n\) is the number of subjects and \(p\) the number of predictors. \(\mu_{b_1}\) is the null-hypothesized value for \(b_1\) - usually set to 0.

Converting b to \(t\)

# Get Standard error's for b (bonus)
se.b1 <- sqrt((n/(n-2)) * mean(error^2) / (var(adverts) * (n-1))); se.b1
[1] 0.009632366
# Calculate t for b1
mu.b1 <- 0
t.b1  <- (b1 - mu.b1) / se.b1; t.b1
[1] 9.979322
n     <- nrow(data) # number of rows
p     <- 1          # number of predictors
df.b1 <- n - p - 1

\(p\)-value for \(b_1\)

Locate in \(t\)-distribution

\[P(|t| \geq 9.98 \mid H_0) < .001\]

So how many @!&#$ ways do we have for assessing an association?!

# the correlation between x and y, standardized (between -1, 1)
cor(sales, adverts)    # .58: moderate and positive, whatever the units were
[1] 0.5784877
# the covariance between x and y, unstandardized 
cov(sales, adverts)    # same association but original units
[1] 22672.02
# regression coefficient in linear regression, unstandardized 
# generalizes easily to settings with multiple predictors
b1 # how much does y-prediction increase, if we increase x by 1 unit?
[1] 0.09612449
# t-statistic: standardized difference between b1 and 0
t.b1 # used for testing the null hypothesis that b1 = 0
[1] 9.979322
     # 9.98: b1 sits about ten standard errors above zero

# The metrics below are more indicative of an overall model's performance

# the correlation between y and model prediction, standardized (between -1, 1)
cor(sales, prediction) # can be squared to get proportion explained variance
[1] 0.5784877
                       # same as original cor: one predictor *is* the model
                       # squared -> .33, a third of the variance in sales
                       # is explained by advertising

Move \(r\), \(s_x\) and \(s_y\) around and watch the covariance, the slope and \(r\) respond

Closing

Exam note

Alternative hypothesis:

  • The question context tells you which test/alternative hypothesis to use
  • Linear regression in JASP only reports two-sided \(p\), so those stay two-sided

Assumptions:

  • Only assessing the assumptions, not yet how to handle violations
  • Only Sensitivity (scatterplot, casewise diagnostics), Normality (Q-Q plot), Linearity (scatterplot, residuals vs. predicted)
  • Don’t need to study ch6; lecture and application in ch8 suffice

See trial exam questions for examples of both

Next Week

  • Exam!
    • Mix of using JASP/interpreting output/conceptual understanding
    • WA will have trial exam with old exam questions
    • Field book available (like in the trial exam), including glossary

Contact

CC BY-NC-SA 4.0