Normally, programming means writing the rules yourself: if this, do that. Machine learning flips it around. You hand the computer a pile of examples with the right answers attached, and it works out the rules on its own. This page builds that idea from scratch with a deliberately tiny problem, telling apples from oranges, so you can see every moving part: training data, features, labels, models, predictions, and the trap called overfitting.
Say you want a program that looks at a fruit and says "apple" or "orange". The classic way is to sit down and write the rule by hand:
if colour_score >= 5:
answer = "orange"
else:
answer = "apple"
That works for fruit, because you already know what makes an orange an orange. Now try writing the rules for "is this email spam?" or "is there a cat in this photo?". You'd need thousands of ifs, you'd miss cases, and the rules would go stale the moment spammers change tactics. Nobody can write those rules by hand.
Machine learning takes the other route. Instead of rules, you collect examples: lots of emails, each one already marked "spam" or "not spam" by a person. A learning algorithm (a fixed recipe of steps) then searches for a rule that gets those examples right. The rule it ends up with is called a model.
Machine learning has a handful of words you'll see everywhere. Here they are on our toy problem. Imagine we weighed some fruit and gave each one a colour score from 0 (green or red) to 10 (fully orange):
| fruit # | weight (g) | colour score | label |
|---|---|---|---|
| 1 | 180 | 2 | apple |
| 2 | 150 | 3 | apple |
| 3 | 140 | 9 | orange |
| 4 | 160 | 8 | orange |
This style, where every training example comes with its label, is called supervised learning, and "pick one of a few categories" problems like ours are called classification. (Predicting a number instead, like a house price, is called regression.) Most of the machine learning you meet day to day, from spam filters to photo tagging, is supervised.
Start as small as possible: one feature, colour score, and a rule of the shape if colour >= threshold then orange. The only thing to decide is the threshold, the cut-off number. Drag the slider to choose one yourself, and watch how many training fruits your rule gets wrong (they get a white ring).
Rule writer · one feature
What the button does is, honestly, all "training" means here: try every possible threshold, count the mistakes on the training data for each, and keep the one with the fewest. Real algorithms search far more cleverly, and their models have thousands or billions of adjustable numbers instead of one, but the idea is identical: adjust the model until it fits the examples.
Notice also that no threshold gets every fruit right. One apple is unusually orange-looking and one orange is unusually pale. Real data is like that: messy, overlapping, sometimes simply mislabelled. A model that is right 100% of the time on real data is rare, and as you're about to see, often a warning sign.
Threshold rules are one kind of model. Here's another that is just as easy to understand and is used for real: k-nearest neighbours, or k-NN. To classify a new fruit:
k is a number you choose, like 3 or 7.That's the whole thing. The "learning" part is barely there: the model simply remembers all the examples. The colour of the background in the chart below shows what the model would predict for a fruit at each spot: reddish means "apple", orange means "orange". The edge between the two areas is called the decision boundary.
Tap Mystery fruit and then tap the chart to see a vote happen. Switch to Add apple or Add orange to put your own examples in and watch the boundary bend around them.
Nearest-neighbour playground · two features
A detail real systems care about: weight goes up to hundreds of grams, colour only to 10. If you measured distance with raw numbers, weight would drown colour out completely. So features are usually rescaled to similar ranges first. The chart does this for you by measuring distance as drawn, where both axes have the same length.
You've probably noticed the two scores under the chart. This is the most important habit in all of machine learning.
The filled shapes are training data: the examples the model gets to see. The hollow shapes are test data: fruits that were set aside before training and that the model never looks at. They have labels too, but we keep those labels secret and only use them for marking the model's work afterwards.
Why bother? Because scoring a model on the examples it learned from is like a teacher handing out the exam questions, with answers, the week before the exam. A student could memorise them all, score 100%, and still understand nothing. What we actually care about is how the model does on new fruit, the kind it will meet in real use. The test set is our stand-in for "new fruit". This ability to do well on unseen examples has a name: generalisation.
Now go back to the playground and drag k all the way to 1, then all the way to 39. Keep an eye on both scores.
The tell-tale sign of overfitting is a big gap between the two scores: great on training data, noticeably worse on test data. Every kind of model can overfit, from tiny ones like this to the huge neural networks behind modern AI, and a lot of practical machine learning is the craft of avoiding it: more data, simpler models, and always, always checking on data the model hasn't seen.
Settings like k, which you choose rather than the model learning them, are called hyperparameters. Picking them by peeking at the test score over and over quietly turns the test set into training data, so in real projects people set aside a third slice, the validation set, for tuning, and look at the test set only once at the very end.
k-nearest neighbours is short enough to write in a few lines of plain Python, no libraries needed. You don't need to follow every line yet; just notice there's no fruit rule written anywhere. The answers come from the examples.
examples = [
# weight (g), colour (0 = green, 10 = orange), label
(180, 2, "apple"),
(150, 3, "apple"),
(200, 1, "apple"),
(140, 9, "orange"),
(160, 8, "orange"),
(210, 9, "orange"),
]
def distance(a, b):
# grams are big numbers and colour is 0-10, so shrink
# the weight difference to keep the two features fair
return ((a[0] - b[0]) / 20) ** 2 + (a[1] - b[1]) ** 2
def predict(weight, colour, k=3):
new_fruit = (weight, colour)
by_closeness = sorted(examples, key=lambda e: distance(e, new_fruit))
nearest = by_closeness[:k]
labels = [e[2] for e in nearest]
return max(labels, key=labels.count)
print(predict(170, 7))
print(predict(190, 2))
orange
apple
sorted(..., key=...) lines the examples up from closest to furthest, [:k] keeps the first k, and max(labels, key=labels.count) picks the label that appears most often: the vote. Add more rows to examples and the program gets better at fruit without you changing a single rule. That's machine learning in miniature. In practice you'd reach for a library such as scikit-learn, which has k-NN and dozens of other models ready to use.
In a spam filter, what is the label of each training email?
The label is the correct answer supplied with the example. The number of links is a feature (a measurement the model can use), and the filter's guess is a prediction.
A model scores 99% on its training data and 61% on test data. What's the most likely story?
A big gap between training and test scores is the classic sign of overfitting: it has memorised its examples, quirks included, and doesn't generalise. An underfitting model does poorly on both.
Why keep test data hidden from the model during training?
Test data does have labels; we just don't let the model learn from them. Then the test score is an honest estimate of how the model will do on examples it has never seen.
In k-nearest neighbours, what usually happens if you make k as large as the whole training set?
If every example gets a vote, the vote comes out the same wherever the new point is, so the model always says the majority label. That's underfitting, the opposite end from k = 1.
Every AI system you hear about, from a spam filter to a chatbot, is a scaled-up version of what you just did with fruit: lots of examples, a model with knobs to adjust, and a test it hasn't seen.