01Home 02Work With Me 03Speaking 04Research 05Teaching 06Students 07Products 08News 09Writing 10About 11Platform New 12Contact 13Curriculum Vitae
MBI806BLesson

Three ways to predict things

Linear regression, decision trees and random forests, explained with no maths. Move a line until it fits, grow a tree until it cheats, watch nine trees outvote the best one — then run real Python in the page.

Business Data Analytics with AI and ML16 min readFree, no login
MBI806B · Three ways to predict things The first screen of Three ways to predict things

This is the written version of Three ways to predict things, taken from the lesson itself. The simulations, drag-and-drop activities and quizzes only work in the interactive lesson.

Linear regression, decision trees and random forests. These three do most of the work in business analytics. You will use all three without writing a single formula.

Reading about 25 minutes Assumes nothing Needs a browser

By the end of this lesson you can

  • Explain all three models to a friend, in your own words
  • Pick the right one for a job
  • Run all three in Python

FirstWhat is a model?

Two minutes on the one idea all three share. Then we get to the models.

Normally, a person writes the rule. Someone at a bank decides: approve the loan if income is over $50,000. They typed that in. They can change it tomorrow.

A model works the other way round. You give it examples. It works out the rule itself.

Same three boxes, and the coloured one has swapped ends. That is the difference.

Why bother? Because some rules are too hard to write. Nobody can explain what makes a squiggle a 7 and not a 1. You just know. Show a computer a million squiggles and it works it out.

Two words you will keep seeing. The facts you have — size, visits, age — are the inputs. The thing you want to know is the answer. That is it.

Worth knowing

A model copies the past. If your old decisions were unfair, it learns to be unfair too — just faster. Nothing in this lesson fixes that.

Model 1 of 3Linear regression.

Use it when the answer is a number. How much will this flat rent for? How many will we sell?

Put your data on a chart. Draw one straight line through the dots. That is the whole model.

The line is the rule. Give it a new flat size, read the rent off the line.

Your turn

Twelve flats. Move the two sliders and get the line as close to all of them as you can.

Move the line

Get it as close to all twelve dots as you can.

How wrong the line is, in total $1,367

The best possible is $164. You are $1203 away.

Each dotted line is how wrong the line is about that flat.

That was hard work. The computer does the same thing in a fraction of a second: it tries thousands of lines and keeps the one that misses least.

And look at what it found. About $6.70 for every extra square metre. That is not just a prediction — it is a sentence you can say in a meeting. That is why people still use this model.

It goes wrong when the dots make a curve

Ice cream sales climb with heat, then drop when it is too hot to go out. A straight line cannot do that.

And when you ask about something it has never seen

These flats are 28 to 98 square metres. Ask about a 400 square metre penthouse and it will still give you a number. It should not be trusted.

Where the name comes from: an 1886 study of how tall children turn out. It tells you nothing about what the model does. Everyone just calls it regression.

Model 2 of 3Decision tree

Use it when the answer is a choice. Will this customer cancel — yes or no?

Twist your ankle badly and you used to get an X-ray, almost automatically. Then some doctors in Ottawa wrote a short flowchart. Can you walk on it? Does this bone hurt? Answer those and you know if the X-ray is worth doing.

Hospitals that use it do far fewer X-rays and still catch the breaks. It is still on the wall today.

A decision tree is that flowchart, except the computer writes it.

How it picks the questions. It tries all of them and keeps whichever one splits the group best. Then it does the same again on each half.

A gym wants to ring people before they cancel. Here are fourteen members it already knows about. Which question should the tree ask first?

Which question should the tree ask first?

Filled square = cancelled · hollow circle = stayed

Good: you can print it

Hand it to a nurse or a loan officer and every decision has a reason you can read out loud. In medicine and lending that is often the law, not a nice-to-have.

Bad: it memorises

You just saw this. Let it keep asking questions and it ends up with one branch per person.

The fix is one word

Stop it early. Three questions deep is plenty. In Python that is max_depth=3.

Model 3 of 3Random forest

Many small trees. Each one votes. The answer with the most votes wins.

One decision tree can be wrong.

So make lots of small trees. Ask every one of them. Go with the answer most of them give. That is a random forest.

Grow lots of small trees

Each tree asks its own question.

Ask every tree

Each tree says quit or stay.

Count the votes

Most votes wins: quit.

Your turn: ask the trees

The same gym. We want to know who will quit. Here are five tiny trees — each asks just one question, so you can read them. Pick a member and watch each tree vote.

Step 1 · Pick a gym member

Fay comes 3× a month joined 7 months ago

Step 2 · Each tree answers its one question

  • Fay comes 3× a month → no
  • Fay comes 3× a month → yes
  • Fay joined 7 months ago → no
  • Fay joined 7 months ago → yes

Step 3 · Count the votes

The forest says Fay will quit

Now check all 20 at once

We already know who really quit. So we can mark every tree, and the forest, on every member.

All 20 members · ✓ = the tree got them right · ✗ = wrong

MemberTree 1Tree 2Tree 3Tree 4Tree 5Forest vote
Amy
Ben
Cal
Dan
Eva
Fay
Gus
Hana
Ivy
Jay
Kim
Leo
Mia
Ned
Oli
Pia
Raj
Sam
Tia
Zoe
Got right13/2015/2015/2015/2016/2019/20

Look down a column

Every tree has crosses. The best single tree gets 16 out of 20. The worst gets 13.

Now look across a row

The crosses land on different people. Most rows have more ticks than crosses, so the ticks win the vote. The forest gets 19 out of 20.

Find Dan’s row. 3 of the 5 trees got Dan wrong, so the vote went wrong too. When most trees are fooled, the forest is fooled.

The big idea: each tree makes mistakes, but on different people. When they vote, the right answers outnumber the wrong ones.

Why is it called a random forest?

Forest

Because it is lots of trees.

Random

Because each tree learns from a random handful of the data. So every tree comes out a little different — and makes different mistakes.

Above, we wrote the five questions by hand so you could read them. A real random forest does it for you:

Each tree learns from a random 12 of the 20 members

Filled dot = this tree gets to see that member · empty = it never sees them

Good: it is right more often

One tree is easy to fool. It is much harder to fool most of them at the same time.

Not so good: it is hard to explain

You can read one tree. You cannot read 300. If you have to tell someone why, use one decision tree instead.

In Python you only choose how many trees

n_estimators=300 just means 300 trees. The computer grows them and counts the votes.

All three, side by sideWhich one do I use?

Usually decided by what you have to explain afterwards, not by which one scores highest.

Linear regression

One straight line through your data.

Can you explain it Yes — in one sentence

How accurate Fine, if the dots are roughly straight

Watch out for Curves. And questions outside your data.

Use it when The answer is a number and you want to know what drives it.

A flowchart of yes/no questions.

Can you explain it Yes — print the chart

How accurate Fine, if you stop it early

Watch out for Memorising your data instead of learning from it.

Use it when Someone will ask you to justify every decision.

Hundreds of trees. They vote.

What it answers A choice, or a number

Can you explain it No — 300 trees is not a chart

How accurate Usually the best of the three

Watch out for You cannot show anyone your working.

Use it when It mostly just has to be right.

Same twelve dots in all three pictures. Only the shape of the rule changes. “Usually the best” means exactly that — on a problem where the answer really is a straight line, the line wins.

Always try the simple one first. A line takes ten minutes and tells you whether there is anything in your data at all. If there is not, no forest will save you.

The codeSix lines of Python

You do not need to be a programmer. Every model is the same few lines, and we will read them together.

Here is a whole, real program. It trains a random forest on the 20 gym members you just met, then asks about two new people.

Press Next line to walk through it. Watch the picture underneath change as you go.

gym.py · a real program
Copy
Open Colab ↗
  1. 1 from sklearn.ensemble import RandomForestClassifier
  2. 2
  3. 3 # Inputs: [visits a month, months a member]
  4. 4 X = [[1, 2], [1, 9], [2, 3], [2, 18], [3, 1],
  5. 5 [3, 7], [3, 26], [4, 4], [4, 11], [4, 30],
  6. 6 [5, 2], [5, 13], [6, 5], [6, 20], [7, 3],
  7. 7 [8, 10], [9, 2], [9, 22], [11, 6], [12, 15]]
  8. 8
  9. 9 # Answers: 1 = quit, 0 = stayed
  10. 10 y = [1, 1, 1, 0, 1, 1, 0, 1, 1, 0, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0]
  11. 11
  12. 12 model = RandomForestClassifier(n_estimators=300)
  13. 13 model.fit(X, y)
  14. 14
  15. 15 new_members = [[2, 4], [10, 20]]
  16. 16 print(model.predict(new_members))

Step 1 of 7Bring in the tool

scikit-learn is a free box of ready-made models. This line takes one model out of the box so the program can use it.

Word by word

  • from sklearn.ensemble from the part of scikit-learn that keeps the forests
  • import bring in
  • RandomForestClassifier the random forest. Classifier means it picks between answers, like quit or stay.

Want to run it? Press Copy, then open a free Colab notebook, paste it into the box and press ▶. Colab already has scikit-learn, so there is nothing to install.

Swap one word, get a different model

Pick a model. Only the highlighted words change. The last two lines are the same for all three.

  • from sklearn.ensemble import RandomForestClassifier changes
  • model = RandomForestClassifier(n_estimators=300) changes
  • model.fit(X, y) never changes
  • model.predict(new) never changes

Use it when

The answer is a choice, and being right matters more than explaining.

The setting n_estimators=300

Grow 300 trees. More trees, steadier votes.

  1. 1 import the model
  2. 2 make it
  3. 3 fit — learn from examples
  4. 4 predict — ask about someone new

Look inside: run it right here

scikit-learn hides the work inside fit. The three boxes below do that work by hand, in plain Python, so you can see it. They run in your browser — nothing to install.

Read each one line by line first. Then switch to Edit and run, change a number, and press Run.

The only Python you need for them

  • # a note A note for people. Python skips everything after the #.
  • size = 52 Store a value under a name. Read = as “is now”.
  • [28, 35, 41] A list: several values in order, inside square brackets.
  • size[0] Pick one item from a list. Counting starts at 0, so this is the first.
  • print("hello") Show something on the screen.
  • f"Rent is {rent}" Text with a value dropped in. The f in front lets you put a name inside { }.
  • for i in range(12): Repeat the lines under it 12 times, with i = 0, 1, 2 … 11.
  • if visits < 4: Only run the lines under it when this is true. else: is what happens otherwise.
  • def total_miss(…): Make your own command. return hands the answer back.
  • (four spaces) The spaces at the start of a line matter. They show which lines belong together.

1 — finding the best line

  1. # Twelve flats: how big they are (m²) and their rent ($ a month).
  2. size = [28, 35, 41, 46, 52, 58, 63, 70, 76, 84, 91, 98] A list of 12 flat sizes, in square metres.
  3. rent = [410, 430, 505, 520, 585, 600, 665, 690, 760, 780, 870, 880] Their rents, in the same order. The first flat is 28 m² and rents for $410.
  4. # How far does a line miss? Add up the gap for every flat.
  5. def total_miss(start, per_m2): Our own command. Give it a line — a starting price and a price per m² — and it adds up how far that line misses all 12 flats.
  6. miss = 0 Start the total at zero.
  7. for i in range(12): Go through the flats one by one: i is 0, then 1, … up to 11.
  8. guess = start + per_m2 * size[i] What the line says this flat should rent for. * means “times”.
  9. miss = miss + abs(rent[i] - guess) How far off was that? abs drops any minus sign, so too high and too low both count.
  10. return miss Hand the total back.
  11. # 1. Try a line yourself. Change these two numbers.
  12. my_start = 200 Your line: start at $200 …
  13. my_per_m2 = 7 … and add $7 for every square metre. Change these two, press Run, and try to get the miss down.
  14. print(f"My line misses by ${total_miss(my_start, my_per_m2):.0f}") Show how far your line misses. The first $ is just a dollar sign; {…} drops the number in;:.0f means “no decimals”.
  15. # 2. Let the computer try 60,000 lines and keep the best one.
  16. best_miss = total_miss(my_start, my_per_m2) The best line so far. We start with yours.
  17. best_start = my_start
  18. best_per_m2 = my_per_m2
  19. for start in range(150, 300): Try every starting price from $150 to $299 …
  20. for cents in range(500, 900): … and for each one, every price per m² from $5.00 to $8.99. 150 × 400 = 60,000 lines.
  21. per_m2 = cents / 100 range only counts in whole numbers, so we count in cents and divide by 100.
  22. miss = total_miss(start, per_m2)
  23. if miss < best_miss: Does this line miss by less than the best so far?
  24. best_miss = miss Then it is the new best. Remember it.
  25. best_start = start
  26. best_per_m2 = per_m2
  27. print(f"Best line: ${best_start} + ${best_per_m2} per m²") This is what LinearRegression finds. It just uses maths to get there instead of trying every line.
  28. print(f"It misses by ${best_miss:.0f} in total")
  29. # 3. Use the best line to predict a new flat.
  30. print(f"A 65 m² flat -> about ${best_start + best_per_m2 * 65:.0f} a month") Predicting is just using the line: start + price per m² × 65.

Runs in your browser. Nothing is installed and nothing is sent anywhere.

Try this: change my_start and my_per_m2 and get your miss as low as you can. Can you beat the computer's $164?

2 — a tree picking its first question

  1. # Fourteen gym members.
  2. # How often each one comes (visits a month)...
  3. visits = [1, 2, 2, 3, 3, 4, 4, 5, 6, 7, 8, 9, 11, 12] A list of 14 numbers, one per member. The first member comes once a month.
  4. #... and what really happened: 1 = cancelled, 0 = stayed.
  5. cancelled = [1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0] The answers, in the same order. The first six cancelled.
  6. # Ask "visits less than cut_off?" and count the mistakes.
  7. def count_wrong(cut_off): Our own command. Give it a cut-off and it tells you how many members that question gets wrong.
  8. wrong = 0 Start the mistake counter at zero.
  9. for i in range(14): Go through the members one at a time: i is 0 for the first, up to 13 for the last.
  10. if visits[i] < cut_off: The tree’s question. visits[i] is how often member number i comes.
  11. guess = 1 # yes -> guess "cancels" Yes side of the tree: we guess they cancel.
  12. else:
  13. guess = 0 # no -> guess "stays" No side: we guess they stay.
  14. if guess!= cancelled[i]:!= means “is not equal to”. Our guess does not match what really happened …
  15. wrong = wrong + 1 … so add one mistake.
  16. return wrong Hand the count back.
  17. # Try every cut-off. Keep the one with the fewest mistakes.
  18. best = 2 The best cut-off so far. We start with the first one we will try.
  19. for cut_off in range(2, 10): Try cut-offs 2, 3, 4 … 9. range stops just before 10.
  20. print(f"visits < {cut_off} -> {count_wrong(cut_off)} of 14 wrong") Show the score for this cut-off. The f lets us drop numbers into the text with { }.
  21. if count_wrong(cut_off) < count_wrong(best): Fewer mistakes than the best so far? Then this one is the new best.
  22. best = cut_off
  23. print(f"The tree's first question: visits less than {best}?") This is what DecisionTreeClassifier does for its first question — for every column, very fast.

Try this: does the winning cut-off match the question you picked in the tree widget above?

3 — five tiny trees, then a vote

  1. # The same 20 gym members as the widget above.
  2. visits = [1, 1, 2, 2, 3, 3, 3, 4, 4, 4, 5, 5, 6, 6, 7, 8, 9, 9, 11, 12] Three lists, all in the same order. The first number in each belongs to Amy.
  3. months = [2, 9, 3, 18, 1, 7, 26, 4, 11, 30, 2, 13, 5, 20, 3, 10, 2, 22, 6, 15]
  4. answers = [1, 1, 1, 0, 1, 1, 0, 1, 1, 0, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0] # 1 = quit What really happened: 1 = quit, 0 = stayed.
  5. # Five tiny trees. Each asks ONE question. True means "will quit".
  6. def tree1(i): return visits[i] < 3 A tiny tree. For member number i it answers True (will quit) or False (will stay).
  7. def tree2(i): return visits[i] < 5
  8. def tree3(i): return visits[i] < 7
  9. def tree4(i): return months[i] < 6
  10. def tree5(i): return months[i] < 12 Five trees, five different questions — the same five as Tree 1 to Tree 5 above.
  11. trees = [tree1, tree2, tree3, tree4, tree5] Put all five trees in one list, so we can go through them.
  12. # The forest: ask all five trees and count the "quit" votes.
  13. def forest(i): The forest is a command too. It answers for member number i.
  14. votes = 0 Start the quit votes at zero.
  15. for tree in trees: Ask each tree in turn.
  16. if tree(i): Does this tree say quit?
  17. votes = votes + 1 Then add one vote.
  18. return votes >= 3 # 3 or more of 5 say quit -> quit Most votes wins: 3 or more of the 5 means quit.
  19. # How many of the 20 does a model get right?
  20. def score(model): Give it any model — one tree, or the whole forest — and it counts how many of the 20 it gets right.
  21. right = 0
  22. for i in range(20):
  23. if model(i) == answers[i]: == asks “are these equal?”. Did the model guess what really happened?
  24. right = right + 1
  25. return right
  26. for n in range(5): Score each tree on its own …
  27. print(f"Tree {n + 1} on its own: {score(trees[n])} / 20 right")
  28. print(f"All five voting: {score(forest)} / 20 right") … then the forest. Compare the numbers.

Try this: change a cut-off in one of the trees, or change votes >= 3 to votes >= 2. The vote is hard to make worse — that is the useful part.

Stuck? Press Reset to get the original back. And while True: will freeze the page, because Python is running in this tab — just reload.

Quick quizThree questions

Nothing is saved and nobody sees it. Read why the wrong ones are wrong — that is the useful bit.

You want to predict how much a flat will rent for. Which model?

What this lesson saidFive things to remember

  • 01 A model finds the rule for you. You give it examples. It gives you back a rule. Nobody types the rule in.
  • 02 Linear regression draws one line. Use it when the answer is a number. It also tells you things like "about $6.70 per square metre".
  • 03 A decision tree is a flowchart. Use it when the answer is a choice, and when you have to explain yourself. Stop it early or it just memorises.
  • 04 A random forest is lots of trees voting. More accurate, but you cannot print it. Good when being right matters more than explaining.
  • 05 Try the simple one first. A line takes ten minutes and tells you if there is anything there at all.

Sources: Galton's height study (1886); the Ottawa ankle rules (Stiell and colleagues, early 1990s).

Now try the real thing.

Everything above is on one page so you can read it anywhere. The lesson itself runs in your browser: no login, no install.