A camera does not send a picture, it sends numbers

When you look at a photograph, you see shapes, colours and faces. A computer never sees any of that directly. What it receives is a grid of numbers, one number (or a small group of numbers) for every tiny square on that grid. Each square is a pixel, short for picture element, and the number attached to it says how bright or how coloured that tiny square is. A digital camera sensor is physically a grid of light-sensitive cells arranged in rows and columns. Each cell measures how much light hit it and converts that measurement into a number. Line those numbers up in the same rows and columns as the sensor cells, and you get the image as the computer actually stores it.

This is the idea you need before anything else in computer vision makes sense: an image is data in the shape of a grid, and a tensor is simply the general name for a grid of numbers that can have more than two dimensions. A single number is a tensor with zero dimensions (a scalar). A list of numbers is a tensor with one dimension (a vector). A table of numbers is a tensor with two dimensions (a matrix). Once you add more axes on top of rows and columns, you call it a tensor rather than a matrix, mostly out of convention. Images, as you will see, usually need three or four axes, so the word tensor fits naturally and you will see it everywhere in the deep learning libraries you use for computer vision, such as PyTorch and TensorFlow.

The simplest case: a grayscale image is a matrix

Start with a grayscale image, meaning an image with no colour information, only brightness. Each pixel is a single number describing how dark or light that spot is. By a very common convention, that number is an 8-bit integer, which means it can take any value from 0 to 255: 0 means pure black, 255 means pure white, and everything in between is a shade of grey. An 8-bit integer type is usually written as uint8 in code (unsigned, meaning no negative values, 8-bit wide).

A grayscale image that is, say, 4 pixels tall and 4 pixels wide is therefore nothing more than a 4 by 4 matrix of numbers between 0 and 255. Here is one, built directly in numpy so you can see the shape and the values together.

🐍Python
import numpy as np

# numpy >= 1.21

# A tiny 4x4 grayscale image, values from 0 (black) to 255 (white)
gray = np.array([
    [ 10,  20,  30,  40],
    [ 50,  60,  70,  80],
    [ 90, 100, 110, 120],
    [130, 140, 150, 160]
], dtype=np.uint8)

print(gray.shape)   # (4, 4)  -> height=4, width=4
print(gray.dtype)   # uint8
print(gray[0, 0])   # 10  -> top-left pixel is nearly black
print(gray[3, 3])   # 160 -> bottom-right pixel is a mid grey

Two things are worth slowing down on. First, gray.shape prints (4, 4), and by convention the first number is the height (number of rows) and the second is the width (number of columns). This height-first order trips up a lot of newcomers because we normally say "a 4 by 4 image" meaning width by height in everyday speech, but numpy arrays index rows before columns, so shape always reports height first. Second, gray[0, 0] gives you the pixel in the top row, left column, with value 10, which is close to black, while gray[3, 3] gives 160, a medium grey in the bottom-right corner. Indexing a tensor is just indexing into this grid, row then column.

Color images add a third axis

A colour image needs more than one number per pixel, because colour is usually described using three components: how much red, how much green and how much blue light is present. This is the RGB model. So instead of one number per pixel you now have three, and the natural way to store that is to add a third axis to the grid. A colour image of height H and width W becomes a tensor of shape (H, W, 3): for every one of the H times W spatial positions, there are 3 numbers stacked along a new axis called the channel axis.

🐍Python
import numpy as np

# A tiny 2x2 colour image, channel order is (R, G, B)
color = np.array([
    [[255,   0,   0], [0, 255,   0]],   # top row: pure red, pure green
    [[0,     0, 255], [255, 255, 0]],   # bottom row: pure blue, yellow
], dtype=np.uint8)

print(color.shape)        # (2, 2, 3) -> height=2, width=2, channels=3
print(color[0, 0])        # [255   0   0] -> the red pixel
print(color[1, 1])        # [255 255   0] -> the yellow pixel

# Pulling out just the red channel gives back a plain 2D matrix
red_channel = color[:, :, 0]
print(red_channel.shape)  # (2, 2)
print(red_channel)

The shape (2, 2, 3) reads as height 2, width 2, channels 3. color[0, 0] gives the three numbers [255, 0, 0], which is as red as red gets and no green or blue at all. color[1, 1] gives [255, 255, 0], which mixes full red and full green with no blue, producing yellow. When you slice out color[:, :, 0] you take only the red number from every pixel and get back a plain two-dimensional matrix, exactly like the grayscale example earlier. This is a useful trick to remember: a single channel of a colour image is a grayscale image in its own right, and many simple image processing operations (brightness adjustment, thresholding) are first understood on a single channel before being applied to all three.

HWC or CHW: the order matters

So far the channel axis came last: (height, width, channels), usually abbreviated HWC. This is how most image libraries that load files from disk will hand you the array, including Pillow and OpenCV (though OpenCV loads colour channels in BGR order rather than RGB, which is a classic source of bugs worth remembering). It is also the order that matplotlib expects when you ask it to display an array as an image.

PyTorch, on the other hand, expects the channel axis first for its convolutional layers: (channels, height, width), abbreviated CHW. This is not a matter of one being more correct than the other, it is purely a convention baked into how each library's internal operations are written. Because of this, converting an image into a tensor for a PyTorch model almost always involves moving the channel axis from last to first. The torchvision library's ToTensor transform does exactly this conversion and, at the same time, rescales pixel values and changes their type, which brings us to the next topic.

🐍Python
import numpy as np
import torch
from torchvision import transforms

# torch >= 2.0, torchvision >= 0.15

# Same 2x2 colour image as before, shape (H, W, C), dtype uint8, range 0-255
color = np.array([
    [[255,   0,   0], [0, 255,   0]],
    [[0,     0, 255], [255, 255, 0]],
], dtype=np.uint8)

to_tensor = transforms.ToTensor()
tensor_img = to_tensor(color)

print(tensor_img.shape)   # torch.Size([3, 2, 2])  -> channels moved to the front
print(tensor_img.dtype)   # torch.float32
print(tensor_img.max())   # tensor(1.)  -> values rescaled from 0-255 to 0.0-1.0

Notice the shape flips from (2, 2, 3) to (3, 2, 2): channels, then height, then width. This single swap is one of the most common sources of shape errors when people move between a library that loaded or displayed an image in HWC order and a model that expects CHW. Whenever a tensor shape error mentions mismatched dimensions in computer vision code, checking whether the channel axis ended up in the wrong place is a good first move.

Pixel values: integers on a 0-255 scale, or floats

Raw image files almost always store pixel intensities as 8-bit integers from 0 to 255, because that is enough precision for human vision and it is compact to store. Neural networks, however, work much better with small floating point numbers centred near zero, because the arithmetic inside the network (sums of weighted inputs, gradient updates) stays numerically well behaved in that range. Large, strictly positive integers like 255 can make training slower and less stable. So before an image tensor is fed into a model, it is standard to convert it from uint8 in the range 0 to 255 into float32 in a much smaller range, typically 0.0 to 1.0.

The conversion itself is simple division: take every pixel value and divide by 255.0. A pixel that was 200 becomes 200 divided by 255, which is about 0.784. A pixel that was 0 stays 0.0, and a pixel that was 255 becomes exactly 1.0. Some pretrained models go a step further and also subtract a fixed mean and divide by a fixed standard deviation, so that the resulting numbers are centred around zero rather than only between 0 and 1; this extra step is called normalization and you will meet it again when the path covers transfer learning, because the exact mean and standard deviation values must match whatever the original model was trained with.

🐍Python
import numpy as np

# The grayscale image from earlier, dtype uint8, values 0-255
gray = np.array([
    [ 10,  20,  30,  40],
    [ 50,  60,  70,  80],
    [ 90, 100, 110, 120],
    [130, 140, 150, 160]
], dtype=np.uint8)

# Convert to float32 and rescale to the 0.0-1.0 range
gray_float = gray.astype(np.float32) / 255.0

print(gray_float.dtype)      # float32
print(gray_float[0, 0])      # 0.039215688 -> 10 / 255
print(gray_float[3, 3])      # 0.627451    -> 160 / 255
print(gray_float.min(), gray_float.max())  # 0.039... 0.627...

A detail that matters in practice: the division has to happen after converting to a floating point type. If you divide the uint8 array directly, many libraries will either raise an error or silently produce wrong, truncated results, because uint8 arithmetic cannot represent fractions. Calling astype(np.float32) first, then dividing, avoids this trap.

Stacking images into batches: the fourth axis

Training a neural network one image at a time would be extremely slow, because modern hardware (GPUs in particular) is designed to do the same arithmetic on many pieces of data at once. So in practice, images are grouped into batches, and a batch of images is itself a tensor with one more axis than a single image: (batch size, height, width, channels) in HWC-first libraries, or (batch size, channels, height, width) in PyTorch's CHW convention. This fourth axis is usually written as N, short for the number of samples in the batch.

🐍Python
import numpy as np

# Three tiny 2x2 colour images, each shape (2, 2, 3), dtype uint8
img1 = np.zeros((2, 2, 3), dtype=np.uint8)
img1[:, :, 0] = 255               # all red

img2 = np.zeros((2, 2, 3), dtype=np.uint8)
img2[:, :, 1] = 255               # all green

img3 = np.zeros((2, 2, 3), dtype=np.uint8)
img3[:, :, 2] = 255               # all blue

# Stack them along a new first axis to form a batch
batch = np.stack([img1, img2, img3], axis=0)

print(batch.shape)   # (3, 2, 2, 3) -> batch=3, height=2, width=2, channels=3
print(batch[0, 0, 0])  # [255 0 0] -> top-left pixel of the first image

The shape (3, 2, 2, 3) is easy to misread at first glance because two of the numbers happen to be the same, but working through it step by step it says: 3 images in the batch, each 2 pixels tall, 2 pixels wide, with 3 colour channels. A very common beginner mistake is to feed a single image, shaped (H, W, C), straight into a model that expects a batch axis, shaped (N, H, W, C). The usual fix is to add a batch dimension of size 1 using a function such as numpy's expand_dims or PyTorch's unsqueeze, so a single image of shape (28, 28, 1) becomes (1, 28, 28, 1) before it is passed to the model.

Why this representation matters for the rest of the path

Everything a convolutional network does to an image is, underneath, an operation on this grid of numbers. The next article in this path introduces the convolution operation itself, which slides a small grid of weights (a filter, also called a kernel) across the height and width axes of the input tensor, combining nearby pixel values into new ones, while the channel axis is used to mix information across colour channels or across the feature maps produced by earlier layers. None of that will make sense without first being comfortable with the idea that an image is a grid with a height axis, a width axis, a channel axis and, during training, a batch axis. Pooling layers, also covered next, likewise act on the height and width axes to shrink the spatial grid while leaving the channel axis alone. Data augmentation, covered later in the path, is in the end a set of operations (flipping, cropping, rotating) applied directly to these same height and width axes before the tensor ever reaches the network. Keeping a clear mental picture of which axis is which will save you a great deal of debugging time throughout the rest of this path.

Common mistakes

  • Mixing up HWC and CHW: loading an image with Pillow or OpenCV gives HWC, but PyTorch models expect CHW, so the channel axis needs to be moved, not just reinterpreted.
  • Forgetting the channel axis for grayscale images: a grayscale image is often stored as a plain (H, W) matrix, but most model code expects an explicit channel axis, so it needs to become (H, W, 1) or (1, H, W).
  • Dividing uint8 pixel values by 255 without first converting to a floating point type, which silently truncates the result instead of producing the expected fraction.
  • Assuming every float image is already scaled to 0.0-1.0: some pipelines leave floats on the original 0-255 scale, and feeding that into a model expecting 0-1 will badly hurt training or predictions.
  • Passing a single image where a batch is expected, missing the leading batch axis of size 1.
  • Confusing RGB and BGR channel order, since OpenCV reads and writes images in BGR by default while most other tools use RGB; a swapped red and blue channel is a very common silent bug, not a crash.
  • Reading shape tuples in the wrong order and assuming the first number is width rather than height, which leads to accidentally swapped height and width when resizing or cropping.

A short checklist before you feed an image to a model

  • What is the shape, and in what order are the axes: HWC or CHW, and is there a batch axis?
  • What is the dtype: uint8 integers from 0 to 255, or float32 numbers, and if float, what range are they actually in?
  • Does the number of channels match what the model expects: 1 for grayscale, 3 for RGB, occasionally 4 if an alpha (transparency) channel was not dropped?
  • If the model was pretrained elsewhere, does the normalization (division and any mean and standard deviation subtraction) match what that model was originally trained with?

Summary and what's next

An image, as far as a machine is concerned, is a grid of numbers: one number per pixel for grayscale, three for colour, arranged along height, width and channel axes, with an extra batch axis added when many images are processed together. Pixel values typically start as 8-bit integers from 0 to 255 and are converted to small floating point numbers before training, and different libraries disagree on whether the channel axis comes first or last, which is one of the most common sources of shape bugs in computer vision code. None of this is specific to any one model or architecture; it is the shared language that every technique in this path is built on. The next article, Convolution, Padding, Stride and Pooling, takes this grid representation and shows precisely how a convolutional filter slides across it, what padding and stride change about the output size, and how pooling shrinks the spatial grid while keeping the information that matters.