Data Science
Linear Regression
It takes O(n) time to compute the sums needed for slope and intercept — a single pass over the data.
The idea, in plain English
Linear regression finds the single straightest line through a scatter of dots. Think of stretching a rubber band across a cloud of points until it settles right through the middle of them. Once you have that line, you can predict a value you have not seen yet. Just read it off the line.
How it works
- 1Take your (x, y) data points — for example, hours studied (x) versus test score (y).
- 2Use the least-squares formula to compute the line's slope (how steep it is) and intercept (where it crosses the y-axis). This formula picks the line that keeps the total squared distance from every point to the line as small as possible.
- 3To predict a new y for any x, plug x into the line's equation: y = slope * x + intercept.
When you'd use it
Use this whenever you want to predict a numeric outcome from a numeric input, and the relationship looks roughly like a straight line. Examples include predicting house prices from square footage, or sales from ad spend.
Common beginner mistakes
- Do not trust the line far outside the range of your original data (this is called 'extrapolating'). The real relationship might curve or break down out there.
- Do not use linear regression on data that is not actually linear. Check with a scatter plot first, or the line will fit poorly no matter what.
Try it — edit and run
Click the code to edit · press ⌘/Ctrl+↵ to run
Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.
Points: (1,3) (2,5) (3,7) (4,9)
Slope: 2.00
Intercept: 1.00
Prediction at x=5: 11.00Not sure this is the right topic? See the learning paths → or where this leads →