Three ways to learn

Part 1 of this series looked at what machine learning is in general: a system that finds patterns in data and uses them to make predictions or decisions, instead of following rules a person wrote by hand. That description is deliberately broad, because machine learning is not one single technique. It is a family of approaches that differ in one important way: what kind of feedback the learning algorithm gets while it is learning.

That single question - what feedback does the algorithm have access to - is enough to split almost all of machine learning into three broad categories: supervised learning, unsupervised learning, and reinforcement learning. Each one fits a different kind of problem, needs a different kind of data, and answers a different kind of question. Once you can recognise which category a problem belongs to, you already know a lot about how to approach it, long before you pick a specific algorithm.

This article builds intuition for each of the three, gives a precise definition, works through a small example you can follow by hand, and lists where each one shows up in practice and where people commonly get confused. It does not yet explain how a model is actually trained, how we measure error, or how we check that a model generalises to new data - those come in later parts of this path. For now, the goal is simpler: given a problem, can you tell which of the three kinds of learning it is?

Supervised learning: learning from answered examples

The intuition

Imagine you are learning a new language by studying flashcards. Each card shows a word on one side and its translation on the other. You look at the word, guess the translation, flip the card, and see whether you were right. After enough cards, you start to notice patterns - certain endings, certain roots - and you can guess the translation of a word you have never seen before, because it resembles ones you have studied.

Supervised learning works the same way. You give the algorithm a set of examples where you already know the correct answer, and it searches for a pattern that connects the input to that answer. Once it has found a pattern that works well on the examples it was shown, you can use it to guess the answer for new examples where you do not know the answer yet.

The precise definition

In supervised learning, every example in the training data comes as a pair: an input and a known correct output. The known output is usually called a label. The job of the learning algorithm is to find a function - a rule, however complicated - that maps inputs to outputs in a way that matches the labels it was given as closely as possible, and that also works reasonably well on new inputs it has never seen.

Supervised learning problems come in two main flavours, depending on what kind of thing the label is:

  • Regression: the label is a number, such as a price, a temperature, or a duration. The model predicts a quantity.
  • Classification: the label is a category, such as spam or not spam, cat or dog, or the species of a flower. The model predicts which category an input belongs to.

A worked example: predicting house prices

Suppose you have records of five houses that were recently sold, along with their size in square feet and the price they sold for, in thousands of dollars:

  • 800 sq ft sold for 150
  • 1000 sq ft sold for 175
  • 1200 sq ft sold for 200
  • 1500 sq ft sold for 240
  • 2000 sq ft sold for 300

Look at the numbers for a moment. As size goes up by 200, price goes up by roughly 24 to 25. That suggests a rough rule: price is about 0.12 times the size, plus 60. Check it: 0.12 times 800 is 96, plus 60 is 156, close to the real 150. For 2000 square feet, 0.12 times 2000 is 240, plus 60 is 300, exactly matching the real price. This rule, learned by looking for a straight-line relationship that fits the examples, is a supervised learning model. The inputs are sizes, the labels are prices, and the model is the rule connecting them.

A real learning algorithm does something similar but far more systematically: it searches over many possible rules (not just straight lines, often far more complex shapes) and picks the one that fits the training examples best, using a precise way of measuring how good a fit is. That precise measurement is called a loss function, and it is the subject of Part 5 of this path. For now, the important idea is simpler: the algorithm is guided the whole time by labels it is told are correct, which is exactly what makes this supervised learning.

A classification example works the same way but the label is a category rather than a number. Imagine emails described by two numbers: how many exclamation marks they contain, and how many times the word free appears. You are given a handful of emails already marked spam or not spam. The learning algorithm looks for a boundary - for example, more than two exclamation marks and the word free appearing at all - that separates the spam examples from the non-spam examples as cleanly as possible. A new email is then checked against that boundary and classified accordingly.

Where supervised learning is used

  • Predicting a continuous number: house prices, delivery times, next month's sales, a patient's blood pressure.
  • Sorting things into known categories: spam detection, diagnosing a disease from test results, recognising handwritten digits, deciding whether a transaction is fraudulent.
  • Any situation where you have historical examples with known outcomes and you want to predict the outcome for new, similar cases.

Common mistakes with supervised learning

The most common mistake is assuming that having lots of data is the same as having lots of labelled data. A million photographs of handwriting are only useful for supervised learning once each one is tagged with which digit it shows. Collecting the labels is frequently the slowest and most expensive part of a supervised learning project, not the modelling itself.

A second mistake is trusting a model that fits its training examples perfectly without checking whether it also works on examples it has not seen. A rule that memorises five house prices exactly is not automatically useful for the sixth house. This question - whether a model generalises beyond the data it was trained on - is central enough that an entire later part of this path, Part 6, is devoted to it. For now, just keep in mind that fitting the examples you have is necessary but not sufficient.

Unsupervised learning: finding structure without answers

The intuition

Now imagine a different situation. You are handed a box of mixed buttons - different sizes, colours, and numbers of holes - with no instructions at all. Nobody has sorted them or labelled them. You naturally start grouping them: these look alike, those look alike, this small pile is different again. Nobody told you the groups existed beforehand or what to call them; you found the structure yourself by noticing similarities and differences.

That is the spirit of unsupervised learning. There are no labels telling the algorithm the right answer. Instead, the algorithm looks at the data on its own and finds patterns, groupings, or simpler ways of describing it, based purely on how the examples relate to one another.

The precise definition

In unsupervised learning, the training data consists only of inputs, with no accompanying correct answer. The learning algorithm's job is to discover some useful structure in that data - groups of similar examples, an underlying simpler representation, or patterns in how variables relate to each other - without being told in advance what that structure should look like.

Two common tasks inside unsupervised learning are:

  • Clustering: grouping examples into clusters of similar items, without being told in advance how many groups there are or what defines them.
  • Dimensionality reduction: compressing data that has many measured variables into a smaller number of new variables that still capture most of what matters, making the data easier to visualise or work with.

A worked example: grouping customers

Suppose an online shop records, for six customers, how many times they visited last month and how much they spent on average per visit, in dollars:

  • Customer A: 2 visits, 80 dollars average spend
  • Customer B: 3 visits, 75 dollars average spend
  • Customer C: 15 visits, 10 dollars average spend
  • Customer D: 18 visits, 8 dollars average spend
  • Customer E: 1 visit, 90 dollars average spend
  • Customer F: 14 visits, 12 dollars average spend

Nobody has told us these customers belong to named groups. But looking at the numbers, two clear clusters jump out: customers A, B and E visit rarely but spend a lot per visit, while customers C, D and F visit often but spend little each time. A clustering algorithm does this same kind of grouping automatically, by measuring how close examples are to each other in terms of their numeric features (here, visits and average spend) and gathering nearby examples into the same group. It does not know these groups mean anything like infrequent big spenders versus frequent small spenders - it only finds that the six points naturally fall into two clumps. Giving the clumps meaningful names and deciding what to do with them is still a human job.

Dimensionality reduction solves a related but different problem. Imagine instead of two features you had fifty - pages viewed, time on site, device type, number of items in cart, and so on. Fifty numbers per customer is hard to plot or reason about directly. A dimensionality reduction technique looks for a much smaller number of new variables, perhaps just two or three, that still capture most of the differences between customers, so you can visualise them on a chart or feed a simpler summary into another step of analysis. The technique has no labels to check against either; it is purely finding structure in how the fifty original numbers vary together.

Where unsupervised learning is used

  • Customer segmentation in marketing, grouping shoppers or users by behaviour without predefined categories.
  • Anomaly detection, where most data forms a dense cluster and unusual points stand out as outliers, useful for fraud or fault detection.
  • Organising large collections of documents or images by similarity when no one has manually tagged them.
  • Compressing high-dimensional data for visualisation or as a preparation step before other analysis.

Common mistakes with unsupervised learning

A frequent misunderstanding is treating the groups an algorithm finds as objectively true categories. The clustering of customers above is only as meaningful as the features you chose to measure. If you had used visits and favourite colour instead of visits and spend, you would get entirely different, probably far less useful, clusters. Unsupervised learning finds structure in the numbers you give it, not some hidden truth about the world; choosing the right features to describe your data still requires human judgement, which is exactly the topic of Part 3 of this path.

A second mistake is assuming there is one correct number of clusters. In the example above, two groups made intuitive sense, but with different data, or a different clustering method, you could just as reasonably end up with three or four groups. Unsupervised learning often involves a judgement call about how finely to split the data, and different reasonable choices can all be defensible.

Reinforcement learning: learning from consequences

The intuition

Think about teaching a dog to sit. You do not hand the dog a labelled list of situations and correct responses, the way you would in supervised learning. You also are not just looking for patterns in a pile of unlabelled data, the way you would in unsupervised learning. Instead, the dog tries something, and you react: a treat for sitting, nothing for standing around, maybe a firm no for jumping on the furniture. Over many repetitions, the dog adjusts its behaviour to get more treats and fewer nos. Nobody told the dog the rule directly; it worked out a good strategy through trial, error, and consequences.

Reinforcement learning is this same idea applied to a computer program. A system, usually called an agent, takes actions in some environment, and after each action it receives a signal, called a reward, saying roughly how good or bad the outcome was. The agent does not know in advance which actions are good. It has to try things, observe the rewards, and gradually learn a strategy - called a policy - that tends to produce good outcomes over time.

The precise definition and its vocabulary

Reinforcement learning is usually described with a small set of terms that are worth learning precisely, because you will see them again whenever reinforcement learning comes up:

  • Agent: the learner or decision-maker, the thing being trained.
  • Environment: everything the agent interacts with and that reacts to its actions.
  • State: a description of the current situation the agent finds itself in.
  • Action: a choice the agent can make from its current state.
  • Reward: a number given to the agent after an action, indicating how good or bad that action turned out to be, in that moment.
  • Policy: the strategy the agent has learned, mapping states to the actions it tends to choose.
  • Episode: one full run of the agent interacting with the environment, from a starting state to some ending point.

Unlike supervised learning, the agent is never told the single correct action for a given state. Unlike unsupervised learning, it is not simply looking for structure in a fixed pile of data with nothing guiding it. It is guided by rewards, but those rewards only say how good an outcome was, not what the best possible action would have been. The agent has to explore different actions, notice which ones tend to lead to better rewards over time, and build up a policy from that experience.

A worked example: a robot vacuum learning a room

Picture a simple robot vacuum in a small room represented as a grid. At each step it can move forward, turn left, or turn right. The room has some dirt patches and some walls. Suppose the rewards are set up like this: plus 1 for moving onto a square with dirt and picking it up, minus 1 for bumping into a wall, and 0 for moving onto an already-clean, empty square.

At the very start, the robot knows nothing about the room. It might move forward and bump a wall, getting a reward of minus 1. It tries turning instead, moves onto a dirty square, and gets a reward of plus 1. Over many episodes of wandering around the room, it starts to notice which sequences of actions, from which positions, tend to lead to more plus-1 rewards and fewer minus-1 rewards. Slowly, a policy forms: in this corner, turn right rather than go forward, because going forward here usually hits a wall; near this patch, keep moving forward, because dirt tends to follow dirt in this room's layout.

Notice what the robot was never given: nobody handed it a labelled example saying this exact grid position means turn left, the way a supervised learning example would. It had to discover that by acting, observing rewards, and gradually adjusting its behaviour. It also was not simply finding structure in a static pile of recorded data; the data it learns from (which squares it visits, which rewards it gets) depends on its own earlier actions, which keep changing as it learns. That back-and-forth between acting and learning from the consequences of acting is the hallmark of reinforcement learning.

Where reinforcement learning is used

placeholder