Starting with a question instead of a definition
Before defining machine learning, it helps to notice a problem that ordinary programming struggles with. Imagine you want to write a program that decides whether an email is spam. You could sit down and write rules: if the subject contains the word "free", mark it as spam. If it has more than three exclamation marks, mark it as spam. If the sender is unknown and the email mentions money, mark it as spam.
This works for a while. Then spammers change their wording, legitimate marketing emails start tripping your rules, and you find yourself adding exceptions to your exceptions. The list of rules grows every week and never quite catches up. The underlying difficulty is that nobody, including you, can write down a complete and stable set of rules that separates spam from non-spam. The pattern is real, but it is too complex, too shifting, and too full of edge cases to describe by hand.
Machine learning offers a different approach: instead of writing the rules yourself, you show the computer many examples of emails that are already labelled spam or not spam, and you let an algorithm work out the pattern from the data. The program that results was not typed in line by line. It was produced by a training process that adjusted itself based on examples.
Two ways to get from input to output
It is worth being very concrete about the difference, because this is the idea that everything else in this learning path builds on.
In traditional programming, a person writes the logic, and the computer applies that logic to data to produce an answer.
rules + data --> program --> answers
In machine learning, the logic itself is what gets produced. A person provides data together with the answers that were observed for that data, and a training algorithm searches for a rule that fits those answers well. Once trained, that rule, called a model, can be applied to new data it has never seen.
data + answers --> training algorithm --> model (the rule)
model + new data --> answers for new data
This flip, from writing rules to learning rules, is the single idea that defines the field. Everything else, the different algorithms, the maths, the evaluation techniques, exists to make that flip work reliably.
A more precise definition
A commonly used way to pin this down, going back to early textbooks on the subject, is to describe a learning system in terms of three things: a task, an experience, and a performance measure.
- The task is what you want the system to do, such as classifying emails as spam or not spam, or predicting the price of a house.
- The experience is the data the system is given to learn from, such as a collection of past emails that have already been labelled.
- The performance measure is how you judge whether the system is any good, such as the percentage of emails it labels correctly.
A system is said to be learning if, as it is given more or better experience (data), its performance on the task, measured in a way you have defined, improves. That last part matters: learning is not just storing examples, it is improving at a task you can measure. If a system memorises the exact training emails but cannot judge a new email it has never seen, it has not learned anything useful, it has just copied.
This framing also tells you what you need to have in place before you can do machine learning at all: a clearly stated task, a source of experience (data) relevant to that task, and a way to measure performance that you trust. Later articles in this path spend entire lessons on each of these three pieces, because getting any one of them wrong quietly ruins a project.
A worked example with numbers you can check by hand
Suppose you are given the sizes and sale prices of five small houses in the same neighbourhood.
- House A: 50 square metres, sold for 150,000
- House B: 60 square metres, sold for 170,000
- House C: 70 square metres, sold for 190,000
- House D: 80 square metres, sold for 210,000
- House E: 90 square metres, sold for 230,000
Look at the pattern. Each extra 10 square metres adds 20,000 to the price. You can describe this with a simple rule: price equals 50,000 plus 2,000 times the size in square metres. Check it against House C: 50,000 plus 2,000 times 70 equals 50,000 plus 140,000, which is 190,000. It matches exactly.
You, a person, spotted that pattern by eye because the data was small and tidy. A machine learning algorithm does the same thing mechanically, except it can do it with thousands of houses, dozens of features such as number of bedrooms, distance to the city centre, and age of the building, and noisy, imperfect prices that do not line up so neatly. The algorithm does not guess the rule by inspection. It starts with some initial, usually wrong, guess for the numbers 50,000 and 2,000 (or whatever the equivalent numbers are called for the problem at hand), measures how far off its predictions are from the real sale prices, and nudges those numbers, repeatedly, in the direction that reduces the error. After many such nudges it settles on numbers close to the true pattern.
Two things in that last sentence deserve attention, because they reappear constantly later in this path. First, "measures how far off its predictions are" is the idea behind a loss function, which a later article covers in depth. Second, "settles on numbers close to the true pattern" only happens if the data genuinely contains that pattern and there is enough of it; if the five houses had wildly inconsistent prices with no relationship to size, no algorithm could invent a pattern that is not there.
It is also worth noticing what happens when you ask this learned rule about a house it has never seen, say one of 65 square metres. It would predict 50,000 plus 2,000 times 65, which is 180,000. Whether that prediction is any good for a real house depends on whether the pattern learned from five houses actually holds for a sixth, different house. That question, whether a model's learned pattern holds up on new data, is called generalization, and it is arguably the central concern of the entire field. A model that fits its training examples perfectly but falls apart on anything new has not solved the task, it has just described the past.
Where this actually gets used
Once you see the pattern, examples are easy to find, because the same shape of problem, learn a rule from labelled past examples and apply it to new cases, shows up across very different domains.
- Email and messaging systems learn to separate spam from legitimate mail, following roughly the process described above.
- Recommendation systems on shopping or streaming sites learn, from what people have watched, bought or rated before, which items a given person is likely to want next.
- Banks and lenders use models trained on past loan outcomes to estimate the risk that a new applicant will default.
- Medical imaging tools are trained on large sets of scans that have been labelled by specialists, so they can flag patterns that may warrant a closer look by a doctor.
- Voice assistants rely on models trained on huge amounts of recorded speech paired with its correct transcription, to turn new audio into text.
- Manufacturing lines use models trained on past sensor readings and known failures to flag equipment that may be about to break down.
Notice the common ingredient in every example: there is a large set of past cases where both the input and the correct outcome are known, and the goal is to produce a model that handles new cases well. Where that ingredient is missing, for instance when there is no reliable record of correct outcomes, or when the task keeps changing its definition, machine learning struggles or needs a different approach, which the next article in this path discusses when it introduces supervised, unsupervised and reinforcement learning.
What machine learning is not
The term gets used loosely in everyday conversation, and that looseness causes real confusion for people starting out. It is worth being precise about a few boundaries.
It is not the same thing as artificial intelligence
Artificial intelligence is the broader, older goal of getting computers to perform tasks that normally require human intelligence, by whatever method works. Machine learning is one strategy for pursuing that goal, specifically the strategy of learning rules from data rather than writing them by hand. Other strategies exist too, including hand-written rule systems (sometimes called expert systems) and search-based methods that explore possible moves in a game without learning anything from past data. Every machine learning system is a form of artificial intelligence, but not every artificial intelligence system uses machine learning, and plenty of software that is marketed as "AI" is, underneath, ordinary rule-based logic with no learned component at all.
It is not a system that understands or reasons the way a person does
A trained spam filter has not understood the concept of unwanted mail the way a person understands it. It has found numerical patterns, such as certain words, sender patterns and formatting choices, that correlated with the label "spam" in its training data. This distinction matters in practice: such a system can be confidently wrong on an email that a human would immediately recognise as an unusual but legitimate case, precisely because it is pattern-matching against its training examples rather than reasoning about intent. Treating a model's output as understanding rather than as a statistical estimate is one of the most common sources of misplaced trust in these systems.
It is not guaranteed to work just because you throw data at it
Training a model only produces something useful if the data genuinely contains a learnable pattern relevant to the task, and if there is enough of it to distinguish that pattern from noise. Feeding an algorithm data that has no real relationship to the outcome you want to predict, or data riddled with errors and inconsistent labels, does not produce a working model, it produces a model that has learned the noise, which will perform badly the moment it meets anything new. This is often summarised as "garbage in, garbage out", and it is the reason later lessons in this path spend so much time on the quality of features and labels before they spend any time on algorithms.
It is not the same as writing more detailed rules by hand
It can be tempting, once a rule-based system starts failing, to assume that machine learning is simply a more sophisticated version of the same rule-writing exercise. It is not. In the house price example above, nobody told the algorithm that size should be multiplied by some number and added to a constant in that specific way for this specific data; the algorithm estimated those particular numbers from the examples it was given. Swap in a different neighbourhood with different prices, and the same algorithm, with no code changes, learns different numbers. A hand-written rule system has no equivalent way to adapt itself; a person has to go back and edit it.
It does not establish cause and effect by itself
A model can learn that houses nearer a particular school tend to sell for more, without that meaning proximity to the school causes the higher price. It might be that both are driven by some third factor, such as the general desirability of the neighbourhood. Machine learning, as typically practised, finds patterns of association in data; it does not, by itself, tell you which direction the causal arrow points, or whether there is a causal arrow at all. Drawing causal conclusions from a model's predictions requires additional reasoning, and sometimes additional experiments, beyond what the model provides.
It does not sit outside of statistics
Machine learning overlaps heavily with statistics; the house price example above is, mathematically, a form of linear regression that has been taught in statistics courses for a long time. What distinguishes the way the field of machine learning tends to use these ideas is less about brand-new mathematics and more about emphasis: a strong focus on prediction accuracy on new, unseen data, a willingness to use very flexible models with many adjustable numbers, and a reliance on large datasets and computing power to fit them. It is more accurate to think of machine learning as a particular style of applying statistical and computational ideas to prediction problems, rather than as a wholly separate subject.
Common mistakes people make when they are starting out
A few misunderstandings come up repeatedly in the early stages of learning this subject, and it helps to name them early.
- Assuming a model that fits the training examples well will automatically do well on new data. As the house price example showed, fitting the past is necessary but not sufficient; generalization to new cases is a separate question, covered in depth later in this path.
- Assuming more data always helps regardless of its quality. A thousand mislabelled examples teach a model a thousand wrong lessons just as efficiently as correct ones would teach it right ones.
- Assuming the choice of algorithm matters more than the quality of the task definition and the data. In practice, a clearly defined task with clean, relevant data, handled by a simple method, regularly beats a sophisticated method applied to a vague task with messy data.
- Assuming a working model explains why something happens. As discussed above, a model's predictions describe a pattern of association; explaining the underlying mechanism is a separate and often harder question.
- Confusing the learning process (training), which happens once or periodically using past data, with the use of the resulting model (prediction), which happens repeatedly on new data afterwards. These are distinct phases with different costs, different risks, and different things that can go wrong.
Keeping these distinctions straight from the beginning saves a lot of confusion later, because almost every later topic in this path, from how datasets are split, to how performance is measured, to why simple baselines matter, is really an answer to one of these common mistakes.
Summary and what comes next
Machine learning is the practice of producing a model, a rule for turning inputs into outputs, by learning from labelled examples rather than by writing the rule by hand. A system can be said to be learning when, given more or better experience in the form of data, its measured performance on a clearly defined task improves. It is a widely useful approach precisely because many real patterns, from what counts as spam to what a house is worth, are too complex and too shifting for anyone to write down as an explicit set of rules. At the same time, it is not artificial intelligence in general, not a form of human-style understanding, not a guarantee of good results regardless of data quality, not a substitute for careful rule design, not a source of causal explanations on its own, and not a field wholly separate from statistics.
The next article in this path, Supervised, Unsupervised and Reinforcement Learning, builds directly on the idea of experience introduced here, by looking at the different shapes that experience can take: cases where every example comes with a known correct answer, cases where it does not, and cases where the system learns by taking actions and observing the consequences over time.
Comments (0)
No comments yet. Be the first to share your thoughts.