Why the first question is always "what kind of data is this"
Imagine you are handed a spreadsheet with a column called "rating" that contains the numbers 1, 2, 3, 4 and 5. Someone asks you for the average rating. You compute it in two seconds: sum the column, divide by the count, done. But here is the uncomfortable question: does that average mean anything? If the numbers are customer satisfaction scores where 1 means "very unhappy" and 5 means "very happy", the gap between 1 and 2 might not feel the same to a customer as the gap between 4 and 5. Averaging them treats every step as equally sized, which may or may not be true. If instead the numbers are counts of items a customer bought, the average behaves exactly as you would expect, because counts really do add and divide in a meaningful way.
This is the point of this article. The first article in this path argued that statistics comes before machine learning because you need to understand your data before you can model it responsibly. Knowing the type of each column is the very first step in that understanding, and it comes before you even look at a mean, a chart or a model. Get it wrong and every summary statistic, every chart choice and every encoding decision downstream can be quietly misleading, even though the code runs without a single error.
The big split: qualitative and quantitative
Every column of data you will meet in practice falls, broadly, into one of two camps. Qualitative data (also called categorical data) describes a quality or a category: a colour, a country, a yes-or-no answer, a product type. Quantitative data (also called numerical data) describes an amount: a count, a weight, a duration, a temperature. The split matters because qualitative data answers the question "which group?" while quantitative data answers the question "how much?", and those two questions call for different tools.
- Qualitative examples: eye colour, marital status, the department an employee works in, whether an email is spam or not
- Quantitative examples: number of emails received per day, weight of a package in kilograms, time to load a webpage in milliseconds
Within each of these two camps there is a further, finer split, and this is where the four types in the title come from. Qualitative data splits into nominal (no natural order) and ordinal (has a natural order). Quantitative data splits into discrete (countable, often whole numbers) and continuous (measurable, can fall anywhere on a scale). The rest of this article works through each of the four in turn, with small examples you can check by hand.
Categorical data (nominal)
Nominal data is the simplest type to picture. It puts each observation into one of several named buckets, and there is no sense in which one bucket is "more" or "less" than another. "Country of residence" is nominal: France is not greater than or less than Japan, they are simply different labels. "Blood type" is nominal: A, B, AB and O are categories with no ranking between them. "Payment method" (cash, card, transfer) is nominal for the same reason.
The defining test for nominal data is this: if you shuffled the label names, would anything meaningful be lost? For "country", swapping the word "France" for the word "Zeta" and "Japan" for "Alpha" changes nothing about the structure of the data, you have just renamed the categories. That is the signature of a nominal variable: the labels are identifiers, not measurements.
- Customer ID, product SKU, zip code, phone area code
- Gender categories as typically recorded in a form
- Type of vehicle: car, van, motorbike, bicycle
- Marketing channel: email, social media, search, referral
A common and costly mistake is to store nominal data as numbers and then forget it is nominal. Zip codes look like numbers, and a spreadsheet will happily let you compute their "average", but the average of two zip codes is not a meaningful location, it is a number with no interpretation at all. The same trap catches phone numbers, customer IDs and product codes. If adding, subtracting or averaging the values produces nonsense, the column is nominal, no matter how numeric it looks on the screen. The safe habit is to ask, before doing any arithmetic on a numeric-looking column: does this number represent an amount, or is it just a label that happens to be written with digits?
Ordinal data
Ordinal data is categorical data with a meaningful order, but without a guarantee that the gaps between categories are equal in size. "Education level" (no formal schooling, primary, secondary, bachelor's, master's, doctorate) is ordinal: a doctorate clearly represents more formal education than a bachelor's degree, so the categories can be ranked, but nobody can say that the step from "primary" to "secondary" is the same size, in any meaningful sense, as the step from "bachelor's" to "master's". "Customer satisfaction" on a scale such as very dissatisfied, dissatisfied, neutral, satisfied, very satisfied is ordinal for the same reason: satisfied is clearly above neutral, but the emotional distance between those two levels is not guaranteed to equal the distance between dissatisfied and neutral.
Here is a small worked example that shows why this matters in practice, not just in theory. Suppose five customers rate a product on a 1-to-5 scale, where 1 is very dissatisfied and 5 is very satisfied. Four customers say 1 (very dissatisfied) and one customer says 5 (very satisfied). The sum of the ratings is 1 + 1 + 1 + 1 + 5 = 9, and the mean is 9 divided by 5, which is 1.8. Reported on its own, a mean of 1.8 out of 5 sounds like "mostly dissatisfied, leaning towards very dissatisfied", which happens to match the true picture reasonably well here. But change the numbers slightly: three customers say 2 (dissatisfied) and two say 4 (satisfied). The sum is 2 + 2 + 2 + 4 + 4 = 14, and the mean is 2.8, close to the "neutral" midpoint of 3. Yet not a single customer actually chose neutral. The mean has invented a middle ground that nobody reported, because it treated the numeric labels 1 to 5 as if they were measured quantities with equal spacing, rather than ranked categories.
This does not mean you should never summarise ordinal data with numbers. It means you should be honest about what the number represents. Reporting the count or percentage of customers in each category (how many said "satisfied", how many said "neutral") preserves the real information. Reporting the median or the mode is usually safer than reporting the mean for ordinal data, because the median only asks "what is the middle category when everyone is lined up in order" and does not assume anything about equal spacing. The mean, by contrast, performs arithmetic that implicitly assumes those gaps are equal, which for ordinal data is an assumption, not a fact. Other common ordinal examples include clothing sizes (small, medium, large, extra large), movie age ratings, Likert-scale survey questions, and star ratings on shopping sites, even though those star ratings are often averaged in practice as a convenient approximation.
Discrete numerical data
Discrete data is quantitative, meaning it represents an amount rather than a label, and it takes values you can count, typically whole numbers (though not always, as the next section will clarify). The number of children in a household, the number of items in a shopping basket, the number of support tickets a team closes in a week, and the number of goals scored in a football match are all discrete. You cannot have 2.5 children or 3.7 goals; the value jumps from one whole number to the next with nothing meaningful in between.
- Number of emails received per day: 0, 1, 2, 3, and so on
- Number of defective items in a batch of 100
- Number of visits to a website in an hour
- Number of bedrooms in a house
The key difference between discrete numerical data and ordinal categorical data is that arithmetic on discrete data is genuinely meaningful. If one household has 2 children and another has 4, it is correct and useful to say the second household has twice as many children, and that the difference is 2 children. Compare this with ordinal satisfaction scores, where saying a rating of 4 is "twice as satisfied" as a rating of 2 is not something you can defend, because the numbers 1 to 5 there were just ordered labels, not counted quantities. The surface appearance of both types can be identical, a column of small whole numbers, but the meaning underneath is completely different, and that meaning is what should drive your choice of summary statistics and charts, not the appearance of the column.
Continuous numerical data
Continuous data is also quantitative, but instead of being counted, it is measured, and in principle it can take any value within a range, including fractions and decimals that go on as finely as your measuring instrument allows. Height, weight, temperature, time taken to complete a task, and the price of a product are all continuous. Between a height of 170 centimetres and 171 centimetres there are infinitely many possible heights (170.1, 170.23, 170.238, and so on), limited in practice only by how precisely your ruler, scale or sensor can measure.
- Height in centimetres, weight in kilograms
- Temperature in degrees Celsius
- Time to load a webpage, measured in milliseconds
- Average basket value of a customer's purchases, in euros
In real datasets, continuous data is almost always stored with some rounding, because every measuring instrument has a limit of precision. A kitchen scale might record weight to the nearest gram, a stopwatch to the nearest hundredth of a second. This rounding does not turn continuous data into discrete data; the underlying quantity (an actual weight, an actual duration) still varies smoothly, we have simply chosen to record it with a certain number of decimal places. A useful rule of thumb: if, in principle, you could always imagine a value falling between any two recorded values given a more precise instrument, the data is continuous, even if in practice your spreadsheet only shows whole numbers or two decimal places.
What operations make sense for each type
The reason to go through this classification carefully is that each type permits a different set of honest operations, and using an operation the data does not support is how misleading statistics get produced even from correct arithmetic. Here is a practical summary, moving from the least to the most information-rich type.
- Nominal data supports: counting how many observations fall in each category, computing the mode (the most common category), and computing proportions or percentages. It does not support ordering, averaging, or subtraction.
- Ordinal data supports everything nominal data supports, plus ranking (this is higher than that), the median, and percentile-type summaries based on position in the order. Averaging is possible to compute but needs to be interpreted with caution, because equal spacing between categories is usually an assumption rather than a verified fact.
- Discrete numerical data supports everything ordinal data supports, plus addition, subtraction, the mean, the variance, and ratios (twice as many, half as much), because the numbers represent real, countable quantities with genuine distances between them.
- Continuous numerical data supports all of the above, plus it can be divided into arbitrarily fine intervals, which is why it is the natural fit for tools such as histograms with many narrow bins, and for the probability distributions covered later in this path.
A more precise framework: four levels of measurement
Statisticians often describe the same idea using four levels of measurement, which refine the categorical-versus-numerical split a little further and are worth knowing because you will see the terms "interval" and "ratio" used in textbooks and documentation. The first two levels match what has already been covered: nominal (labelled categories with no order, like country or blood type) and ordinal (ordered categories with unknown or unequal spacing, like satisfaction ratings or education level).
The remaining two levels both fall under what this article has been calling numerical data, but they are distinguished by one specific detail: whether zero means "none of this quantity" or is just an arbitrary reference point. Interval data has equally spaced values, so differences between numbers are meaningful, but it lacks a true zero. The classic example is temperature measured in degrees Celsius. The 10-degree gap between 20°C and 30°C is the same size as the 10-degree gap between 5°C and 15°C, so subtraction is meaningful. But 0°C does not mean "no temperature", it is simply the freezing point of water, a human-chosen reference. Because zero is arbitrary, multiplication and ratios break down: it is not correct to say that 20°C is "twice as hot" as 10°C, even though 20 is twice 10 as plain numbers. Calendar years work the same way: the year 2000 is not "twice as much time" as the year 1000, because year 0 does not mean the absence of time, it is just a reference point chosen by a calendar system.
Ratio data has equally spaced values and a true, meaningful zero, which means both differences and ratios are valid. Weight, height, age, income, and counts of things are all ratio data, because zero genuinely means "none": zero kilograms means no weight at all, zero euros means no money at all. This is why it is completely correct to say a 20-kilogram suitcase is twice as heavy as a 10-kilogram suitcase, in a way that simply does not hold for 20°C versus 10°C. In practice, for everyday data work, the nominal-ordinal-discrete-continuous split used in the rest of this article (and across this path) is usually the more directly useful distinction, while interval-versus-ratio is a detail worth recognising mainly so that you do not compute ratios of temperatures, dates or other interval-scaled numbers and present them as meaningful percentages.
Seeing this in a dataset
It helps to look at a tiny dataset and classify every column before touching a single statistic. Imagine eight rows describing customers of a small online shop: a customer ID, the country they ordered from, how satisfied they said they were on a 1-to-5 scale, how many orders they have placed, and the average amount spent per order in euros. The code below builds exactly that table using pandas (version 2.x; the ideas apply to any recent version) and prints what pandas infers about each column's storage type, which is a useful starting point, though not the final word, for thinking about data type.
import pandas as pd
data = {
"customer_id": ["C001", "C002", "C003", "C004", "C005", "C006", "C007", "C008"],
"country": ["France", "Japan", "France", "Brazil", "Japan", "Brazil", "France", "Kenya"],
"satisfaction": [4, 2, 5, 3, 4, 1, 5, 3],
"num_orders": [3, 1, 7, 2, 5, 1, 9, 4],
"avg_order_value_eur": [42.50, 19.99, 88.10, 25.00, 61.75, 15.40, 102.30, 37.60]
}
df = pd.DataFrame(data)
print(df)
print()
print(df.dtypes)
Running this prints the eight rows followed by a list of pandas dtypes: customer_id and country will show as object (pandas' default for text), satisfaction and num_orders will show as int64, and avg_order_value_eur will show as float64. Notice what this tells you, and what it does not. Pandas correctly separates text columns from numeric columns, but it has no idea that satisfaction is ordinal while num_orders is discrete numerical. Both are stored as int64, yet they mean very different things: calling df["satisfaction"].mean() will run without error and return a number, but as the earlier worked example showed, that number can misrepresent what customers actually reported, while df["num_orders"].mean() returns a genuinely meaningful average number of orders per customer. The dtype pandas assigns is about storage, not about meaning. Knowing the four data types from this article is what lets you decide, column by column, which pandas operations are trustworthy and which need a second thought, something the hands-on article later in this path will return to with a larger, real dataset.
Where this choice shows up in real data work
Chart choice follows directly from data type. Bar charts, which show a separate bar for each category with no implied order between bars, suit nominal data, such as a bar for each country showing how many customers come from there. When the categories have a natural order, as with ordinal data, bars are still appropriate but should be drawn in the natural order (very dissatisfied through very satisfied, left to right) rather than alphabetically, so the order is visible. Histograms, which group continuous (or sometimes discrete) numerical data into bins and show a bar per bin, suit continuous data such as order value, where there is no natural small number of categories but a smooth range of possible amounts.
Encoding for machine learning models is another place this distinction has direct, practical consequences. A nominal column like country should typically be one-hot encoded, turned into a separate 0/1 column per country, because feeding the model a single numeric column where France is 1, Japan is 2 and Brazil is 3 would invite the model to treat Japan as being "between" France and Brazil in some numeric sense, which is meaningless for a nominal variable. An ordinal column like education level or satisfaction, by contrast, can reasonably be encoded as a single increasing numeric column (1, 2, 3, 4, 5), because there the order is real, even if the exact spacing is uncertain; some modelling approaches even have dedicated ordinal encoders for exactly this reason. Discrete and continuous numerical columns are usually left as plain numbers, sometimes after scaling, since arithmetic on them is already meaningful.
The choice also quietly decides which statistical tests and summaries are appropriate later in this path. Tests and summaries built around means and variances, which later articles in this path cover in detail, assume the arithmetic they perform is meaningful, which is why they are natural for discrete and continuous data and need extra care for ordinal data. Tests built around ranks or category counts are the safer default for ordinal and nominal data respectively. You do not need the names of those tests yet, only the habit of asking what type of data you have before choosing a tool, which is exactly the habit this article is building.
Common mistakes to avoid
- Averaging ordinal survey scales (1 to 5 ratings, education levels, pain scales) and reporting the result as if it were a precise measurement, when the categories may not be evenly spaced.
- Treating identifier-like numbers (zip codes, customer IDs, phone numbers, flight numbers) as quantities to be summed or averaged, when they are really nominal labels written in digits.
- Assuming any column stored as int64 or float64 in pandas, or as a "number" type in a spreadsheet, must be discrete or continuous data, without checking whether it actually represents an amount.
- Alphabetising ordinal categories in a chart (dissatisfied, neutral, satisfied, very dissatisfied, very satisfied) instead of plotting them in their natural order, which hides the pattern the chart was meant to reveal.
- Computing a ratio or percentage change on interval data such as temperatures in Celsius or calendar years, where the zero point is arbitrary and ratios are not meaningful.
- One-hot encoding an ordinal variable and discarding its order, or numerically encoding a nominal variable and accidentally introducing a false order, when building features for a model.
Summary and what comes next
Every column of data is either qualitative, describing a category, or quantitative, describing an amount. Qualitative data splits into nominal, where categories have no order (country, blood type, payment method), and ordinal, where categories are ordered but the spacing between them is not guaranteed to be equal (satisfaction ratings, education level, clothing size). Quantitative data splits into discrete, where values are counted and usually whole (number of orders, number of children), and continuous, where values are measured and can fall anywhere within a range, limited only by instrument precision (weight, height, time, price). A more formal framework adds the interval-versus-ratio distinction within numerical data, turning on whether zero means "none of this quantity" (ratio: weight, age, counts) or is just an arbitrary reference point (interval: temperature in Celsius, calendar years). Getting this classification right, before computing a single statistic, is what lets you choose summary statistics, charts, encodings and tests that actually mean what they claim to mean.
With the four types in hand, the next article in this path puts them to work directly: it looks at the mean, the median and the mode, three different ways of describing the "centre" of a dataset, and shows concretely, with small worked examples, why the right choice of centre depends on exactly the kind of data distinction built in this article.
Comments (0)
No comments yet. Be the first to share your thoughts.