Chapter 2: Simple Linear Regression
Sections ΒΆ Summary ΒΆ 2.1 Overview ΒΆ Simple linear regression model:
h ( x i ) = w 0 + w 1 x i h(x_i) = w_0 + w_1 x_i h ( x i β ) = w 0 β + w 1 β x i β w 0 w_0 w 0 β is the intercept; w 1 w_1 w 1 β is the slope.
Mean squared error for the simple linear regression model is a bowl-shaped function of w 0 w_0 w 0 β and w 1 w_1 w 1 β :
R sq ( w 0 , w 1 ) = 1 n β i = 1 n ( y i β ( w 0 + w 1 x i ) ) 2 R_\text{sq}(w_0, w_1) = \frac{1}{n}\sum_{i=1}^n \big(y_i - (w_0 + w_1 x_i)\big)^2 R sq β ( w 0 β , w 1 β ) = n 1 β i = 1 β n β ( y i β β ( w 0 β + w 1 β x i β ) ) 2 2.2 Partial Derivatives ΒΆ Suppose f ( x , y , z , . . . ) f(x, y, z, ...) f ( x , y , z , ... ) is a function of multiple input variables. β f β x \frac{\partial f}{\partial x} β x β f β β the partial derivative of f f f with respect to x x x β is the derivative of f f f with respect to x x x , treating all other inputs as constants. Itβs still a function of all inputs.
Critical points: all partial derivatives are 0 at once. E.g. f ( x , y ) = x 2 + y 2 9 f(x, y) = \frac{x^2 + y^2}{9} f ( x , y ) = 9 x 2 + y 2 β has partials 2 x 9 , 2 y 9 \frac{2x}{9}, \frac{2y}{9} 9 2 x β , 9 2 y β ; minimum at ( 0 , 0 ) (0, 0) ( 0 , 0 ) .
2.3 Finding Optimal Parameters ΒΆ β R sq β w 0 = β 2 n β i = 1 n ( y i β ( w 0 + w 1 x i ) ) \frac{\partial R_\text{sq}}{\partial w_0} = -\frac{2}{n}\sum_{i=1}^n \big(y_i - (w_0 + w_1 x_i)\big) β w 0 β β R sq β β = β n 2 β i = 1 β n β ( y i β β ( w 0 β + w 1 β x i β ) ) β R sq β w 1 = β 2 n β i = 1 n x i ( y i β ( w 0 + w 1 x i ) ) \frac{\partial R_\text{sq}}{\partial w_1} = -\frac{2}{n}\sum_{i=1}^n x_i \big(y_i - (w_0 + w_1 x_i)\big) β w 1 β β R sq β β = β n 2 β i = 1 β n β x i β ( y i β β ( w 0 β + w 1 β x i β ) ) Setting both partials to 0 and solving for w 0 w_0 w 0 β and w 1 w_1 w 1 β gives:
w 1 β = β i = 1 n ( x i β x Λ ) ( y i β y Λ ) β i = 1 n ( x i β x Λ ) 2 w 0 β = y Λ β w 1 β x Λ w_1^* = \frac{\sum_{i=1}^n (x_i - \bar x)(y_i - \bar y)}{\sum_{i=1}^n (x_i - \bar x)^2} \qquad\quad w_0^* = \bar y - w_1^* \bar x w 1 β β = β i = 1 n β ( x i β β x Λ ) 2 β i = 1 n β ( x i β β x Λ ) ( y i β β y Λ β ) β w 0 β β = y Λ β β w 1 β β x Λ Since β i = 1 n ( x i β x Λ ) = 0 \sum_{i=1}^n (x_i - \bar x) = 0 β i = 1 n β ( x i β β x Λ ) = 0 , also w 1 β = β i = 1 n ( x i β x Λ ) β y i β i = 1 n ( x i β x Λ ) 2 w_1^* = \frac{\sum_{i=1}^n (x_i - \bar x)\, y_i}{\sum_{i=1}^n (x_i - \bar x)^2} w 1 β β = β i = 1 n β ( x i β β x Λ ) 2 β i = 1 n β ( x i β β x Λ ) y i β β , and several other useful forms.
The errors of the fit line satisfy β i = 1 n ( y i β h ( x i ) ) = 0 \sum_{i=1}^n (y_i - h(x_i)) = 0 β i = 1 n β ( y i β β h ( x i β )) = 0 .
The fit line passes through ( x Λ , y Λ ) (\bar x, \bar y) ( x Λ , y Λ β ) .
2.4 Correlation ΒΆ Correlation coefficient:
r = 1 n β i = 1 n ( x i β x Λ Ο x ) ( y i β y Λ Ο y ) r = \frac{1}{n}\sum_{i=1}^n \left(\frac{x_i - \bar x}{\sigma_x}\right)\left(\frac{y_i - \bar y}{\sigma_y}\right) r = n 1 β i = 1 β n β ( Ο x β x i β β x Λ β ) ( Ο y β y i β β y Λ β β ) where
Ο x = 1 n β i = 1 n ( x i β x Λ ) 2 \sigma_x = \sqrt{\frac{1}{n}\sum_{i=1}^n (x_i - \bar x)^2} Ο x β = n 1 β i = 1 β n β ( x i β β x Λ ) 2 β r r r is the mean product of z-scores. β 1 β€ r β€ 1 -1 \le r \le 1 β 1 β€ r β€ 1 ; β£ r β£ |r| β£ r β£ is the strength and the sign the direction; it is unitless, and r ( x , y ) = r ( y , x ) r(x, y) = r(y, x) r ( x , y ) = r ( y , x ) . It measures linear association only: r β 0 r \approx 0 r β 0 can hide a strong curve.
x i β a x i + b x_i \to a x_i + b x i β β a x i β + b , y i β c y i + d y_i \to c y_i + d y i β β c y i β + d (a , c β 0 a, c \ne 0 a , c ξ = 0 ) leaves β£ r β£ |r| β£ r β£ unchanged; r r r flips sign if exactly one of a , c a, c a , c is negative.
The optimal slope can be written in terms of r r r :
w 1 β = r β Ο y Ο x w 0 β = y Λ β w 1 β x Λ w_1^* = r\,\frac{\sigma_y}{\sigma_x} \qquad\quad w_0^* = \bar y - w_1^* \bar x w 1 β β = r Ο x β Ο y β β w 0 β β = y Λ β β w 1 β β x Λ So the line depends only on x Λ , y Λ , Ο x , Ο y , r \bar x, \bar y, \sigma_x, \sigma_y, r x Λ , y Λ β , Ο x β , Ο y β , r .
2.5 Least Squares ΒΆ Model Optimal parameters Minimum MSE h ( x i ) = w h(x_i) = w h ( x i β ) = w w β = y Λ w^* = \bar y w β = y Λ β Ο y 2 \sigma_y^2 Ο y 2 β h ( x i ) = w 0 + w 1 x i h(x_i) = w_0 + w_1 x_i h ( x i β ) = w 0 β + w 1 β x i β w 1 β = r Ο y Ο x w_1^* = r\frac{\sigma_y}{\sigma_x} w 1 β β = r Ο x β Ο y β β , β
β w 0 β = y Λ β w 1 β x Λ \; w_0^* = \bar y - w_1^* \bar x w 0 β β = y Λ β β w 1 β β x Λ Ο y 2 ( 1 β r 2 ) β€ Ο y 2 \sigma_y^2 (1 - r^2) \le \sigma_y^2 Ο y 2 β ( 1 β r 2 ) β€ Ο y 2 β