The problem machine learning has to solve
Strip away the buzzwords and almost every machine learning model is doing one thing: turning a description of something into a prediction about it. The description might be the pixels of a photograph, the words in a sentence, the price history of a stock, or the size and age of a house. The prediction might be a category, a number, or a probability. In between the description and the prediction sits a function, and that function is built, almost without exception, out of additions and multiplications of numbers arranged in rows and columns. That arrangement, and the rules for combining it, is exactly what linear algebra studies.
This is not a coincidence or a historical accident. Linear algebra is the branch of mathematics that describes how to combine many numbers at once in a consistent, efficient way, and machine learning is fundamentally a discipline of combining many numbers at once. Once you see this connection clearly, formulas that look intimidating in a textbook or a paper start to read as ordinary bookkeeping: lists of numbers, grouped into tables, combined by a small number of repeated operations.
Data is already a list of numbers
Consider a single house you want to put a price on. You might describe it with three numbers: its size in hundreds of square feet, the number of bedrooms, and its age in years. A house with 1500 square feet, 3 bedrooms and 10 years of age becomes the list [15, 3, 10]. That ordered list of numbers is a vector. Nothing mystical has happened yet; a vector at this stage is just a row of measurements about one example, written in a fixed order so that the first number always means size, the second always means bedrooms, and so on.
The next two articles in this series spend real time on what a vector means geometrically, as a direction and a length in space, and why that geometric picture matters for machine learning. For now, the simpler reading is enough: a vector is one example's worth of numbers, kept in order.
Stacking data: a dataset is a matrix
Real datasets rarely contain one example. Suppose you have three houses instead of one.
- House A: 1500 sqft, 3 bedrooms, 10 years old → [15, 3, 10]
- House B: 2000 sqft, 4 bedrooms, 5 years old → [20, 4, 5]
- House C: 1000 sqft, 2 bedrooms, 20 years old → [10, 2, 20]
Stack these three vectors as rows and you get a table with three rows and three columns. That table is a matrix. Rows usually correspond to examples (here, individual houses) and columns correspond to features (size, bedrooms, age). This is the same shape you would get by opening a spreadsheet of house listings: one row per listing, one column per attribute. A matrix, at this stage, is simply a dataset written as a grid of numbers instead of a spreadsheet with labels.
A prediction is a matrix-vector product
Now suppose you already have a simple pricing rule: multiply size by 3, bedrooms by 10, age by minus 2 (older houses lose value), add them up, then add a base price of 50, all in thousands of dollars. For House A that is 15 times 3, plus 3 times 10, plus 10 times minus 2, plus 50, which comes out to 45 + 30 - 20 + 50 = 105, so a predicted price of $105,000. For House B you get 60 + 40 - 10 + 50 = 140, or $140,000. For House C you get 30 + 20 - 40 + 50 = 60, or $60,000.
Notice what you just did for each house: you multiplied each feature by a fixed weight, summed the results, and added a constant. That operation — multiply matching entries, sum them, repeat for every row — is precisely matrix-vector multiplication. The weights [3, 10, -2] form a vector, call it w. The house data forms a matrix, call it X. The predictions come from computing X times w, then adding the constant b for every row. Written as code, this is a single line.
import numpy as np
# three houses, three features each: size (hundreds of sqft), bedrooms, age
X = np.array([
[15, 3, 10],
[20, 4, 5],
[10, 2, 20]
])
w = np.array([3, 10, -2]) # weight per feature
b = 50 # base price, in thousands of dollars
prices = X @ w + b
print(prices)
# expected output: [105 140 60]
The @ operator in NumPy (version 1.x and later) performs matrix multiplication. X @ w takes each row of X, multiplies it entry by entry with w, and sums the result, producing one number per row. Adding b then shifts every prediction by the same base price. The printed array [105, 140, 60] matches the hand calculation exactly. This is linear regression, one of the oldest and simplest models in machine learning, and it is nothing more than a matrix-vector product plus a constant. Every time you read the formula y = Xw + b in a textbook, this is what it is describing.
Neural network layers are the same idea, repeated
It is tempting to think that once you move from plain linear regression to neural networks, the linear algebra disappears and something more exotic takes over. It does not. A single layer of a neural network takes an input vector, multiplies it by a weight matrix, adds a bias vector, and then applies a nonlinear function to the result. The matrix multiplication part is identical in spirit to the house price example; it is just repeated many times and the output of one layer becomes the input to the next.
Take a tiny example: an input vector with three numbers, x = [1, 2, 3], and a layer that produces two outputs instead of one. The weights for this layer form a 2-by-3 matrix, because you need three weights (one per input) for each of the two outputs.
import numpy as np
x = np.array([1, 2, 3])
W = np.array([
[0.5, -1.0, 0.2],
[1.0, 0.0, 0.5]
])
b = np.array([0.1, -0.2])
z = W @ x + b
print(z)
# expected output: [-0.8 2.3]
# apply a ReLU activation: replace negative values with 0
a = np.maximum(z, 0)
print(a)
# expected output: [0. 2.3]
Walking through the first output by hand: 0.5 times 1, plus minus 1 times 2, plus 0.2 times 3, gives 0.5 - 2 + 0.6 = -0.9, and adding the bias 0.1 gives -0.8. The second output: 1 times 1, plus 0 times 2, plus 0.5 times 3, gives 1 + 0 + 1.5 = 2.5, and adding the bias -0.2 gives 2.3. So z = [-0.8, 2.3]. The ReLU activation function then zeroes out negative values, giving [0, 2.3]. Stack several of these layers, each with its own weight matrix and bias vector, feeding one into the next, and you have the computational skeleton of a neural network. The nonlinearity (ReLU here) is what lets the network represent more than straight lines and flat planes, but the heavy lifting of combining inputs is still matrix multiplication, layer after layer.
Comparing things: a preview of dot products
A great deal of machine learning also involves measuring how similar two things are: two documents, two user profiles, two images encoded as vectors of numbers. The basic tool for this, the dot product, is also linear algebra, and it will get a full article to itself later in this path. For now, just notice the shape of the idea. Take two vectors of the same length, multiply their matching entries, and sum the results. For u = [1, 0, 1] and v = [1, 1, 0], the dot product is 1 times 1, plus 0 times 1, plus 1 times 0, which is 1. This single number ends up being the mathematical backbone of recommendation systems, search engines and the attention mechanism inside transformer language models, all of which repeatedly ask some version of the question: how aligned are these two vectors? You will see the precise geometric meaning of that question, including the connection to angles and cosine similarity, in the next two articles.
Why not just write loops
Everything shown so far could, in principle, be written with ordinary Python loops: for each row, for each feature, multiply and add. So why does machine learning code almost never look like that? Two reasons, one about speed and one about how hardware works.
The speed reason is concrete. A Python for-loop processes one number at a time, with a fair amount of bookkeeping overhead for every single step. A library like NumPy, or a deep learning framework like PyTorch or TensorFlow, hands the same arithmetic down to code written in C or Fortran and compiled ahead of time, often using routines from a library called BLAS (Basic Linear Algebra Subprograms) that has been tuned over decades specifically for multiplying matrices quickly. For a dataset of three rows this difference is invisible. For a dataset of three million rows, or a neural network layer with millions of weights, a loop-based version can be tens to hundreds of times slower than the equivalent vectorized operation, simply because of how much overhead accumulates per step.
The hardware reason goes further. Graphics processing units, GPUs, were originally built to update millions of pixels on a screen at once, which meant they were designed from the ground up to do many independent, simple arithmetic operations in parallel. Matrix multiplication is, structurally, exactly that kind of workload: every output entry is computed independently from every other, using the same simple recipe of multiply-and-sum. This is why the same hardware that renders video games turned out to be extremely well suited to training large neural networks, and why frameworks like PyTorch express an entire model as a sequence of matrix and tensor operations rather than as loops over numbers. Writing computation in the language of linear algebra is not just mathematically convenient; it is literally the format that lets modern hardware do the work in parallel instead of one number at a time.
A geometric preview
So far every vector here has been described purely as a list of numbers and every matrix purely as a table. That is enough to follow the arithmetic, but it hides half the reason linear algebra is such a good fit for machine learning. A vector of numbers can also be read as a point, or an arrow from the origin to that point, in space. A vector with three entries is a point in three-dimensional space; a vector with 784 entries, the kind you get from flattening a small grayscale image, is a point in a 784-dimensional space that is impossible to draw but perfectly well defined mathematically. Two data points that are similar tend to sit close together in this space, and the dot product from the previous section is closely tied to the angle between two such points.
Matrices, read this way, are not just tables of weights; they describe transformations of that space, stretching it, rotating it, or flattening it in certain directions. When a neural network layer computes Wx + b, it is not only doing arithmetic, it is moving the point x somewhere else in space, and training the network is the process of finding transformations that move similarly labeled points close together and differently labeled points far apart. This geometric reading is where the real intuition for linear algebra in machine learning lives, and it is the subject of the next several articles in this path: vectors as direction and length, matrices as transformations, and what happens when you chain transformations together through matrix multiplication.
Common misconceptions worth clearing up now
- Linear algebra is not only for "linear models". Linear regression is linear algebra plus nothing else, but neural networks are linear algebra (the matrix multiplications) alternated with nonlinear activation functions. The nonlinearity is essential, but it sits on top of, not instead of, the linear algebra.
- A vector is not defined only by being "a list of numbers". The order of the entries matters, because it fixes which number means what (size, then bedrooms, then age, in the house example). Swapping the order silently without updating your weights produces nonsense predictions, and this is a common source of real bugs.
- Shape mismatches are the single most common practical error. A matrix with 3 rows and 4 columns cannot be multiplied by a vector with 3 entries the way the house example was; the inner dimensions have to match. Later articles in this path, especially the ones on matrix multiplication and on NumPy in practice, spend time on exactly this kind of bookkeeping, because it is where beginners lose the most time.
- You do not need the geometric picture to write correct code, but you do need it to debug and improve a model with any confidence. Treating vectors and matrices as pure arithmetic works until something goes wrong, and then the geometric intuition (closeness, direction, transformation) is usually what tells you why.
- Using a library like NumPy does not mean you can skip understanding the underlying operation. Knowing that X @ w computes a weighted sum per row is what lets you read almost any machine learning formula you encounter, long before you memorize any particular model's equations.
Summary and what comes next
Machine learning models take descriptions of things, written as vectors, and combine many of them at once, written as matrices, using a small set of repeated operations: multiply and sum, stack rows, chain transformations. Linear regression is a direct matrix-vector product. A neural network layer is the same product with a bias added and a nonlinear function applied, repeated layer after layer. Measuring similarity between two vectors uses a dot product. None of this is incidental; it is the reason libraries are built around matrix operations and the reason GPUs, designed for parallel arithmetic, turned out to be the right hardware for training models at scale.
The rest of this path builds up the pieces introduced here one at a time and in more depth. The next article looks closely at what a vector actually is, as direction and length, and how a single data point becomes a point in space. After that comes the dot product and the idea of angle and similarity that was only sketched here, followed by matrices as transformations, matrix multiplication as composition, and eventually the decompositions, like eigenvalues and singular value decomposition, that explain why certain techniques such as PCA work at all. Each article keeps returning to the same two questions this one started with: what does this operation mean geometrically, and where exactly does a real model use it.
Comments (0)
No comments yet. Be the first to share your thoughts.