From Learning Styles to Learning Material

Part 2 drew the big map: supervised learning learns from examples that already come with the right answer, unsupervised learning looks for structure in examples that have no answer attached, and reinforcement learning learns from the consequences of actions. This article zooms into the supervised case and asks a more concrete question: what does one of those labeled examples actually look like, and what are the pieces it is made of? By the end you should be able to look at almost any table of data used for machine learning and point confidently at which parts are the features, which part is the label, and which direction one row runs.

What a Dataset Actually Is

Most people already have an intuitive picture of a dataset without realising it: a spreadsheet. Imagine a sheet listing students, with one column for hours of sleep last night and another for their score on a test the next morning. That spreadsheet, in its simplest form, is a dataset. Formally, a dataset is a collection of examples that all share the same structure, meaning every example has a value recorded for the same set of properties. A hospital's table of patient visits, a shop's table of past orders, a weather station's table of daily readings, and a biologist's table of measured leaves are all datasets in exactly this sense.

Not every dataset is literally a spreadsheet. Photographs, audio recordings, and blocks of text are datasets too, and later parts of this path and other paths on MLHub will deal with them directly. But even those messier forms are eventually turned into rows of numbers before a model touches them, so the vocabulary in this article, rows, columns, features and labels, applies underneath almost any kind of machine learning, not just the tabular kind.

Rows: Examples, Instances, Samples, Records

Each row of a dataset is one unit that the model will learn from or make a prediction about. In the student spreadsheet, one row is one student on one particular morning. In a house price dataset, one row is one house that was sold. You will see this row called an example, an instance, a sample, an observation, or a record, depending on which book, course, or software library you are reading. Statisticians tend to say observation, machine learning papers and libraries like scikit-learn tend to say sample, and many introductory texts just say example. They all mean the same thing: one unit of data, described by a fixed set of properties, usually with one known outcome attached when the learning is supervised.

Columns (Minus the Target): Features

A feature is a measured or recorded property of an example that the model is allowed to use when making its prediction. You will also hear features called attributes, variables, predictors, or inputs. For a house, features might include its floor area in square meters, the number of bedrooms, the city it is in, the year it was built, and the distance to the nearest school. None of these is the thing we are trying to predict, each one is a clue the model is given to work with.

Numerical features

Numerical features are features whose values are numbers that support ordinary arithmetic. Continuous numerical features can take any value within a range, such as a floor area of 50.0 square meters or 73.4 square meters, where the difference between any two values is meaningful and a value exactly between two others makes sense. Discrete numerical features are whole-number counts, such as the number of bedrooms being 1, 2, or 3. The key test for whether a feature is genuinely numerical is whether subtraction and comparison mean something real: a house with 90 square meters really is 40 square meters larger than one with 50, and a house with 3 bedrooms really does have one more bedroom than a house with 2.

Categorical features

Categorical features record which group or category an example belongs to, rather than a quantity. There are two important kinds. Nominal categorical features have categories with no natural order, such as city (Leeds, Bristol, Oxford) or payment method (card, cash, transfer). It would be wrong to secretly encode these as Leeds equals 1, Bristol equals 2, Oxford equals 3 and feed that straight into a model, because nothing says Oxford is three times anything or sits further along a line than Leeds. Ordinal categorical features do have a meaningful order, such as education level (high school, bachelor's, master's, doctorate) or a customer satisfaction rating (poor, fair, good, excellent). Order matters here, a doctorate really does come after a bachelor's degree, but the gaps between categories are not necessarily equal in size, the jump from poor to fair might not represent the same change in satisfaction as the jump from good to excellent.

Binary features

A binary feature is a special, very common case of a categorical feature that takes exactly two values, such as has_garden being yes or no, or is_weekend being true or false. Binary features are usually stored as 0 and 1, which is convenient because, unlike nominal categories with three or more values, there is no ordering problem to worry about: 0 and 1 are just labels for two states, and most models handle them without any special treatment.

Text, images and other raw material

Raw text, raw images, and raw audio are not features by themselves, they are material that features get built from. The number of characters in an email, the count of exclamation marks it contains, or the average brightness of a photograph are numeric features derived from that raw material. Turning raw text and images directly into model-ready numbers is a substantial topic on its own and belongs to later paths on MLHub rather than this one; for now it is enough to recognise that somewhere between the raw record and the tidy dataset, someone decided which measurable properties were worth extracting.

The Label: What You Are Trying to Predict

The label is the answer, the outcome, the thing the model is trying to learn to produce. It goes by several other names depending on the source: target, outcome, dependent variable, response, ground truth, or simply y. As part 2 explained, a label is only present in supervised learning, it is precisely what separates a supervised problem from an unsupervised one. If every house in your dataset has a recorded sale price, you can train a model to predict price from features, that is supervised learning. If no such recorded outcome exists and you are only looking for groups or patterns among the houses, that is unsupervised learning, with no label in sight.

The data type of the label decides what kind of supervised problem you have. If the label is a category, such as passed_exam being yes or no, or email being spam or not spam, the task is classification. If the label is a number, such as exam_score ranging from 0 to 100, or house price in thousands of pounds, the task is regression. The same raw situation can sometimes be framed either way: whether a student passed is classification, the exact score they got is regression, and choosing between the two framings is a decision you make deliberately, based on what question you actually need answered.

Putting It Together: the Feature Matrix X and the Label Vector y

Almost every machine learning text and library uses the same shorthand. The capital letter X refers to the feature matrix, a grid with one row per example and one column per feature. The lowercase letter y refers to the label vector, one entry per example, lined up so that the label in position 3 of y belongs to the example in row 3 of X. If a dataset has n examples and p features, X has n rows and p columns, and y has n entries, one for each row of X.

Here is a tiny, concrete version of this idea. Imagine five houses that have recently sold, with three features recorded for each, floor area in square meters, number of bedrooms, and city, and one label, the sale price in thousands of pounds.

🐍Python
# Assumes Python 3, no external libraries required

data = [
    {"size_m2": 50, "bedrooms": 1, "city": "Leeds",    "price_k": 120},
    {"size_m2": 65, "bedrooms": 2, "city": "Leeds",    "price_k": 150},
    {"size_m2": 80, "bedrooms": 3, "city": "Bristol",  "price_k": 210},
    {"size_m2": 45, "bedrooms": 1, "city": "Bristol",  "price_k": 110},
    {"size_m2": 95, "bedrooms": 3, "city": "Leeds",    "price_k": 240},
]

feature_names = ["size_m2", "bedrooms", "city"]
label_name = "price_k"

X = [[row[f] for f in feature_names] for row in data]
y = [row[label_name] for row in data]

for features, label in zip(X, y):
    print(features, "->", label)

Running this prints five lines, each showing one row of X next to its matching entry in y, for example [50, 1, 'Leeds'] -> 120 and [95, 3, 'Leeds'] -> 240. Nothing clever has happened here, we have simply separated every column we would hand to a model, size_m2, bedrooms, city, from the one column we want the model to predict, price_k. When a model is trained in part 11, it will search for some function f such that f applied to each row of X comes out close to the matching value in y. Every supervised learning algorithm you will meet in this path, however sophisticated it sounds, is ultimately doing some version of that search.

Where Datasets and Features Come From

Datasets are not born as tidy tables, they are assembled from somewhere: a hospital's record system logging each patient visit, a web application's server logs recording every click, a field surveyor filling in a form, a weather station streaming sensor readings, or a camera capturing images. Features are extracted or computed from whatever that source happens to record.

Labels deserve separate attention because of how they are obtained. Sometimes a label is recorded automatically by the system itself, for instance whether a customer actually returned an item is simply logged by the shop's own database the moment it happens, no human judgement required. Other times a label requires a person to look at the example and decide, for instance a radiologist marking whether a chest X-ray shows a tumor, or a moderator marking whether a post counts as spam. This second kind of labeling is slow and expensive. Labeling ten thousand chest X-rays, even at a brisk few seconds of attention per image, still adds up to many hours of a specialist's time, which is one reason hand-labeled datasets in fields like medicine are often far smaller than the mountains of unlabeled images that exist. This scarcity of labeled data, rather than a scarcity of raw data, is frequently the real bottleneck in a supervised learning project.

A Feature Is Only as Good as Its Measurement

Raw recorded data rarely arrives already shaped into useful features. Suppose a database stores the date a house was listed for sale as a text string like 2024-03-01. That string is not directly useful to most models. A more useful feature is days_on_market, computed as the number of days between the listing date and the date it sold, a single meaningful number built from two raw dates. The process of turning raw recorded fields into columns that actually help a model is called feature engineering, and it occupies a large part of real machine learning work, far more than most beginners expect.

Not every raw field deserves to become a feature, either. A database might also store the buyer's first and last name, but there is no sensible reason to believe the letters in someone's name predict the price of the house they bought, so a careful practitioner leaves that field out entirely. Deciding what counts as a useful feature always depends on the task at hand, the same raw table can yield very different feature sets depending on what question you are trying to answer.

For now, the main thing to take away is that the clean feature columns you see in a tidy dataset are usually the product of earlier decisions and earlier work, not the rawest possible form the data could take. Later, hands-on parts of this path will walk through some of that engineering directly.

Common Mistakes with Features and Labels

A surprising share of problems in supervised learning trace back to confusion over what belongs in X and what belongs in y, long before any algorithm is chosen. The following mistakes are worth watching for deliberately.

  • Leakage, the label in disguise: including a feature that was itself computed directly from the label, for example using total_price as a feature to predict price_per_sqm when total_price was calculated as price_per_sqm times size in the first place. The model appears to perform brilliantly during development and then fails on genuinely new houses, because it was secretly allowed to see the answer.
  • Using information that would not exist yet at prediction time: for example using cancellation_reason as a feature to predict whether a booking will be cancelled. You only find out the cancellation_reason after the cancellation has already happened, so a real system trying to predict the future could never have that column available.
  • Treating identifiers as features: columns like customer_id, row_number, or order_number carry no genuine signal about the outcome and can quietly mislead a model, particularly if the identifiers happen to correlate with time, such as customers who signed up earlier having systematically lower id numbers.
  • Encoding nominal categories as plain numbers: assigning Leeds equals 1, Bristol equals 2, Oxford equals 3 and feeding that column straight into a model that treats numbers as having order and distance. This silently tells the model that Oxford is numerically three times Leeds, which means nothing. Proper ways to handle nominal categories, such as one-hot encoding, are covered in the hands-on part of this path.
  • Noisy or inconsistent labels: if two people reviewing the same set of posts disagree about which ones count as spam, the label column itself contains errors, and no model can be more accurate than the labels it was trained to match.
  • Mixing up which column plays which role: the same dataset can serve more than one question. Price can be the label when the task is predicting price, but price can equally be a feature when the task is predicting something else, such as how long a listing stays on the market. Features and labels are not fixed properties of a dataset, they are defined by the question you choose to ask of it.

Worked Example End to End

Consider eight students, each described by two features, hours studied for an exam and percentage of classes attended, together with two possible labels recorded for each student, their exact exam score out of 100 and whether they passed.

  • Student A: 2 hours studied, 40 percent attendance, score 38, passed: no
  • Student B: 5 hours studied, 60 percent attendance, score 58, passed: no
  • Student C: 7 hours studied, 75 percent attendance, score 66, passed: yes
  • Student D: 1 hour studied, 50 percent attendance, score 41, passed: no
  • Student E: 9 hours studied, 90 percent attendance, score 81, passed: yes
  • Student F: 6 hours studied, 55 percent attendance, score 60, passed: yes
  • Student G: 3 hours studied, 45 percent attendance, score 44, passed: no
  • Student H: 8 hours studied, 85 percent attendance, score 77, passed: yes

The two feature columns, hours studied and attendance, never change. What changes is which label you line them up against. If the question is will this student pass, y is the passed column, a category with two possible values, and the task is classification. If the question is what score will this student get, y is the score column, a number, and the task is regression. Notice that the pass threshold used here looks roughly like 60 or above, but the dataset itself does not state a rule, it only records outcomes, and it is the model's job during training to discover a rule that fits the pattern in the data, not to be handed one.

This small example also makes the earlier leakage warning concrete. It would be a mistake to add a feature called passing_grade_estimate that some earlier spreadsheet formula had already computed from the score column, because a model given that feature would just be reading off an already-computed version of the answer rather than learning anything from hours studied and attendance. The safest check, when you are unsure whether a column belongs in X or should be left out entirely, is to ask whether that value would realistically be available at the moment you need to make the prediction, before the true outcome is known. If the answer is no, it does not belong among the features.

Summary and What Comes Next

A dataset is a collection of examples sharing the same structure, where each row, also called an example, instance, sample, or observation, describes one unit the model will learn from or predict about. Each column other than the target is a feature, a measured property the model is allowed to use, and features come in several flavors: continuous and discrete numerical features that support arithmetic, nominal categorical features with no natural order, ordinal categorical features with an order but uneven spacing, and binary features as a simple two-valued special case. The label, present only in supervised learning, is the outcome the model is trying to predict, and whether it is a category or a number determines whether the task is classification or regression. Together, the feature columns form the matrix X and the label column forms the vector y, lined up row by row. Good features avoid leaking the answer, avoid encoding categories as if they were ordered numbers, and reflect information genuinely available at prediction time, while labels are only as trustworthy as the process, automatic or human, that produced them.

With a clear picture of what rows and columns actually represent, the next natural question is what to do with all of them before training anything. Part 4 covers how to split a dataset like the ones built here into training, validation, and test portions, so that you can honestly check whether a model has learned a genuine pattern or merely memorised the examples it was shown.