Where we left off

Part 1 got Python installed, a virtual environment working, and a notebook or editor open. This part is about the actual language: how Python stores a value under a name, what kinds of values exist, and how you combine them into expressions. Everything in this path, from cleaning a CSV to calling an API, is built on these basics, so it is worth slowing down here even if some of it looks familiar from other languages.

What a variable really is

In many languages a variable is described as a box that holds a value. In Python it is closer to a name tag stuck onto an object that already exists somewhere in memory. When you write price = 19.99, Python first creates a float object holding 19.99, then attaches the label price to it. If you later write price = 24.99, Python does not change the original object; it creates a new float object and moves the label to it. The old object, if nothing else refers to it, is eventually cleaned up.

This distinction matters because more than one name can point at the same object, and what happens next depends on whether that object can be changed in place (mutable) or not (immutable). Numbers, strings and booleans are immutable: once created, their value never changes, only the labels pointing at them move. Lists and dictionaries, which you will meet in part 3, are mutable, and that is where this labels-not-boxes picture starts to really matter. For now, a small example with numbers makes the idea concrete.

🐍Python
x = 10
y = x      # y now points at the same object as x
y = 20     # y is moved to a new object; x is untouched

print(x)   # 10
print(y)   # 20
print(id(x), id(y))  # different memory addresses

id() returns a number that identifies where an object lives in memory (the exact number is not something you should rely on, it is just useful for teaching). The point of the example is that reassigning y never touched x, because integers are immutable and assignment just moves labels around.

Naming variables

Python names can contain letters, digits and underscores, but cannot start with a digit, and are case-sensitive, so total_sales and Total_Sales are two different names. The near-universal convention in Python, written down in the official style guide called PEP 8, is snake_case for ordinary variables: words in lower case separated by underscores, for example customer_id, order_total, is_active. Constants that should not change during a run are usually written in all capitals, such as TAX_RATE or MAX_ROWS, purely as a signal to readers; Python does not actually stop you from reassigning them.

A handful of words are reserved by the language itself and cannot be used as variable names, including if, for, while, import, class, return and a few others. You can list them at any time if you are unsure.

🐍Python
import keyword
print(keyword.kwlist)

Beyond the hard rules, naming is a habit worth building early. In data work you will have many similar-looking values in the same cell or function, so df, x, tmp and val are a poor choice once there is more than one of them around. Prefer names that say what the value is: revenue_2023, customer_count, avg_order_value. It costs a few extra keystrokes and saves real time when you, or someone else, reread the code a month later.

  • Do not start a name with a digit: 2024_sales is invalid, sales_2024 is fine
  • Avoid shadowing built-in names like list, str, sum or type, since reusing them as variables hides the real function for the rest of that scope
  • Avoid single letters except for short-lived loop counters you will meet in part 4, where i or j is conventional

Python's type system: dynamic but not loose

Python is dynamically typed, which means a name is not locked to one type forever; you can point count at an integer now and at a string later, and Python will not complain at the point of assignment.

🐍Python
count = 5
print(type(count))   # <class 'int'>

count = "five"
print(type(count))   # <class 'str'>

This is different from statically typed languages, where a variable's type is fixed when it is declared. Dynamic typing gives flexibility, but it also means Python will not catch a type mistake until the line actually runs, which is one reason good variable names and small, testable pieces of code matter as your programs grow (you will see this formalised with type hints in part 11 and with tests in part 12).

It is important to be precise about what is dynamic: it is the name that is flexible, not the object. Every object in Python has one fixed type for its entire life. The integer 5 is always an int; it is the name count that can be reassigned to point at a different kind of object. You check an object's type with the built-in type() function, and you check whether an object belongs to a type (or one of its relatives) with isinstance(), which is the more robust choice inside conditions you will write from part 4 onwards.

🐍Python
value = 42
print(type(value))              # <class 'int'>
print(isinstance(value, int))   # True
print(isinstance(value, float)) # False

Numbers: int and float

Python has two everyday numeric types for data work: int for whole numbers with no size limit other than available memory, and float for numbers with a decimal point, stored in a fixed-precision binary format (the same double-precision format used by most other languages, typically 64 bits). A literal without a decimal point, like 7, is an int. A literal with one, like 7.0, is a float, even though mathematically it is a whole number.

🐍Python
orders = 120          # int
average_order = 48.5  # float

print(type(orders))         # <class 'int'>
print(type(average_order))  # <class 'float'>

The arithmetic operators are +, -, *, / for addition, subtraction, multiplication and ordinary division, plus three that are less familiar if you are coming from spreadsheets: // for floor (integer) division, % for the remainder (modulo), and ** for exponentiation.

🐍Python
total_rows = 103
batch_size = 20

full_batches = total_rows // batch_size   # 5, whole batches that fit
leftover_rows = total_rows % batch_size   # 3, rows in the final partial batch

print(full_batches, leftover_rows)

print(2 ** 10)       # 1024, exponentiation
print(10 / 4)        # 2.5, ordinary division, always returns a float
print(10 // 4)       # 2, floor division, drops the remainder

Notice that / always returns a float in Python 3, even when the result is a whole number, which is a common surprise if you are used to languages where integer division truncates by default. If you specifically want the whole-number quotient, use //. This pair, // and %, is genuinely useful in data work for splitting a dataset into fixed-size batches, computing page numbers, or working out how many full weeks are in a span of days.

Floating point numbers deserve one warning that trips up almost everyone the first time they see it: because floats are stored in binary, numbers that look simple in decimal, like 0.1, cannot always be represented exactly, so arithmetic on them can produce tiny rounding errors.

🐍Python
print(0.1 + 0.2)          # 0.30000000000000004
print(0.1 + 0.2 == 0.3)   # False

This is not a Python bug; it happens in essentially every language that uses the same binary floating point standard. The practical lesson is: never compare two floats with == when they come from a calculation. Instead compare their difference against a small tolerance, or round before comparing.

🐍Python
a = 0.1 + 0.2
b = 0.3

print(abs(a - b) < 1e-9)   # True, this is the safe way to compare floats

Strings: working with text data

A string is a sequence of characters, written between single quotes, double quotes, or triple quotes for text that spans multiple lines. Python does not distinguish between single and double quotes functionally; the common convention is to pick one and stay consistent, switching to the other only when the text itself contains a quote mark.

🐍Python
product = 'USB cable'
note = "customer said: it's great"   # double quotes avoid clashing with the apostrophe
description = """A multi-line
string, useful for longer
text blocks."""

print(product)
print(note)

Strings are sequences, so you can index into them to get a single character, and slice them to get a substring. Indexing starts at 0, and negative indices count from the end.

🐍Python
customer_id = "CUST00123"

print(customer_id[0])      # 'C', first character
print(customer_id[-1])     # '3', last character
print(customer_id[4:])     # '00123', everything from index 4 onward
print(customer_id[:4])     # 'CUST', everything before index 4
print(customer_id[4:7])    # '001', characters at index 4, 5, 6

Slicing follows a start:stop pattern where start is included and stop is excluded, which is why customer_id[4:7] gives three characters, not four. This same slicing syntax will reappear with lists in part 3, so it is worth getting comfortable with it now on something as simple as a string.

Strings are immutable: you cannot change a character in place, you can only build a new string. A handful of string methods cover most of the cleaning work you will do on messy real-world text, which part 13 of this path will use heavily on an actual CSV file.

🐍Python
raw = "  Customer Name \n"

print(raw.strip())        # 'Customer Name', removes leading/trailing whitespace
print(raw.strip().lower())  # 'customer name'
print(raw.strip().upper())  # 'CUSTOMER NAME'
print(raw.strip().replace("Name", "Surname"))  # 'Customer Surname'

row = "2024-01-15,48.50,completed"
fields = row.split(",")
print(fields)              # ['2024-01-15', '48.50', 'completed']

rebuilt = "-".join(fields)
print(rebuilt)             # '2024-01-15-48.50-completed'

split() breaks a string into a list of pieces wherever a given separator appears, which is exactly what you need when reading a line from a CSV file by hand. join() does the reverse: it stitches a list of strings back together with a given separator between each piece. These two show up constantly once you start reading real files in part 7.

For building strings that mix text and values, f-strings (formatted string literals) are the clearest and most common approach in modern Python, available from Python 3.6 onward. You write an f right before the opening quote, and put any expression inside curly braces.

🐍Python
customer = "Priya"
order_total = 128.4
item_count = 3

message = f"{customer} ordered {item_count} items for a total of ${order_total:.2f}"
print(message)
# Priya ordered 3 items for a total of $128.40

The :.2f inside the braces is a format specification telling Python to show the float with exactly two decimal places, which is the kind of formatting you will want constantly when printing prices, percentages or summary statistics. Before f-strings became standard, the same thing was done with .format() or the % operator; you will still see both in older code, but f-strings are the recommended style for new code.

Booleans, None, and truthiness

The bool type has exactly two values, True and False, written with capital letters. Booleans are what comparison operators produce, and what conditions in part 4 will be built from.

🐍Python
price = 19.99
is_expensive = price > 15

print(is_expensive)       # True
print(type(is_expensive)) # <class 'bool'>

None is a special, unique value that represents the absence of a value, roughly similar to NULL in a database or NaN for missing data in a spreadsheet, though not identical to either. It is its own type, NoneType, and it is commonly used as a placeholder for a result or a field that has not been set yet.

🐍Python
middle_name = None
print(middle_name)       # None
print(type(middle_name)) # <class 'NoneType'>

# the recommended way to check for None is 'is', not '=='
print(middle_name is None)   # True

You check for None with is rather than == because None is meant to represent identity (there is exactly one None object in the whole program), and is checks that two names point at the very same object rather than just an equal-looking one. In practice this distinction rarely bites you with None specifically, but is None is the idiomatic, expected style, and == None will even trigger a style warning in most linters.

Every value in Python has a truthiness, meaning it behaves like True or False when used somewhere a boolean is expected, such as inside an if statement you will meet in part 4. The rule is short: zero, empty collections, empty strings and None are all falsy; almost everything else is truthy.

🐍Python
print(bool(0))        # False
print(bool(1))        # True
print(bool(""))       # False, empty string
print(bool("no"))     # True, any non-empty string is truthy, even the text 'no'
print(bool(None))     # False
print(bool([]))       # False, empty list (lists are covered in part 3)

That last line is worth staring at: the string "no" is truthy, because truthiness is about emptiness, not about the word's meaning. This is a common source of bugs when people expect a text field containing the word 'false' to behave like the boolean False; it will not, since it is a non-empty string.

Converting between types

Data rarely arrives already in the type you want. A CSV file, for example, gives you every field as text, even the ones that are clearly numbers, so converting between types is something you will do constantly. Python's built-in functions int(), float(), str() and bool() each try to construct a value of that type from whatever you give them.

🐍Python
raw_price = "19.99"
raw_quantity = "3"

price = float(raw_price)
quantity = int(raw_quantity)

total = price * quantity
print(total)              # 59.97
print(type(total))        # <class 'float'>

print(str(total))          # '59.97', back to text, e.g. for printing or writing to a file
print(str(total) + " USD") # '59.97 USD'

A few conversions that look like they should work do not, and it is worth knowing the exact failure rather than guessing. int() cannot parse a string that contains a decimal point, even though the number is a whole number once you ignore the decimal part.

🐍Python
# int("19.99") would raise:
# ValueError: invalid literal for int() with base 10: '19.99'

# the correct two-step conversion, when you really want a whole number:
price_text = "19.99"
price_as_int = int(float(price_text))
print(price_as_int)   # 19, the decimal part is simply dropped, not rounded

Converting via float first and then int works, but notice that int() truncates rather than rounds, so int(19.99) gives 19, not 20. If you want proper rounding, use the built-in round() function instead, which is covered along with its rounding rules in more depth once you reach numeric cleaning in part 13. Attempting to convert text that is not a valid number at all, like float("abc"), raises a ValueError immediately; handling that kind of failure gracefully is the subject of part 8 on errors and exceptions, so for now it is enough to know that bad conversions fail loudly rather than silently producing a wrong number, which is actually the safer behaviour.

Expressions, operator precedence and comparisons

An expression is anything Python can evaluate down to a single value: a literal number, a variable, or a combination of these joined by operators. 2 + 3 * 4 is an expression; so is price * quantity; so is a whole chain like (price * quantity) - discount + shipping_fee.

When an expression mixes several operators, Python follows a fixed order of precedence, the same general idea as the order of operations taught in school arithmetic. Exponentiation (**) binds tightest, followed by multiplication, division, floor division and modulo (*, /, //, %), followed by addition and subtraction (+, -). Parentheses override all of this and should be used liberally whenever the order is not obvious at a glance, both for correctness and for readability.

🐍Python
result = 2 + 3 * 4
print(result)        # 14, not 20, because * runs before +

result_clear = 2 + (3 * 4)
print(result_clear)  # 14, same answer, but the intent is unmistakable

subtotal = 100
tax_rate = 0.08
discount = 10

# written without parentheses, this is easy to misread
total = subtotal + subtotal * tax_rate - discount
print(total)   # 98.0

# the same calculation, with intent made explicit
total_clear = subtotal + (subtotal * tax_rate) - discount
print(total_clear)  # 98.0, identical result, clearer to the next reader

Comparison operators (==, !=, <, >, <=, >=) take two values and produce a bool. A detail worth knowing early: == checks whether two values are equal, while = is the assignment operator that gives a name a value. Confusing the two is one of the most common early mistakes, and in most contexts Python will actually raise a syntax error if you write = where == was required, which at least stops the mistake from running silently.

🐍Python
budget = 500
spent = 480

print(spent == budget)   # False, checking equality
print(spent < budget)    # True
print(spent <= budget)   # True

# a Python-specific convenience: chained comparisons read naturally
age = 34
print(18 <= age < 65)    # True, equivalent to (18 <= age) and (age < 65)

Logical operators combine booleans: and requires both sides to be True, or requires at least one side to be True, and not flips a boolean. Python evaluates and and or with short-circuiting, meaning it stops as soon as the overall result is already decided, which also means the right-hand side is sometimes skipped entirely.

🐍Python
has_discount_code = True
order_total = 45

qualifies_for_free_shipping = has_discount_code and order_total > 30
print(qualifies_for_free_shipping)  # True

is_eligible = (order_total > 100) or (has_discount_code and order_total > 20)
print(is_eligible)  # True

# 'not' flips a boolean
out_of_stock = False
print(not out_of_stock)  # True, meaning it IS in stock

These comparison and logical expressions are exactly what you will place inside if statements in part 4; this part is giving you the pieces, the next one gives you the control flow that reacts to them.

Common mistakes worth knowing now

A short list of errors that come up constantly for people new to Python, each with the reason behind it, so you recognise them instantly instead of guessing.

  • Using = instead of == inside a condition. Python usually catches this as a syntax error, but it is worth knowing the rule by heart rather than relying on the error message.
  • Comparing floats with == after arithmetic, which can fail due to binary rounding as shown earlier with 0.1 + 0.2. Compare with a tolerance instead.
  • Mixing incompatible types in an operation, such as "3" + 4, which raises a TypeError because Python will not silently guess whether you meant to add numbers or concatenate text. Convert explicitly with int(), float() or str() first.
  • Assuming int() rounds when it actually truncates, so int(2.9) gives 2, not 3. Use round(2.9) if you want proper rounding.
  • Shadowing a built-in name, for example writing list = [1, 2, 3] or type = "customer", which replaces the built-in function or type for the rest of that scope and causes confusing errors later when you try to use list(...) or type(...) normally.
  • Forgetting that strings and numbers are immutable, then being surprised that a method like .strip() or .upper() does not change the original variable unless you reassign it, for example name = name.strip() rather than just name.strip().
🐍Python
name = "  Alicia  "
name.strip()          # this creates a new string but throws it away
print(repr(name))     # '  Alicia  ', unchanged, because nothing was reassigned

name = name.strip()   # now the cleaned string actually replaces the old one
print(repr(name))     # 'Alicia'

This last mistake deserves emphasis because it will reappear, in a different shape, with lists and dictionaries in part 3: most string and number methods return a new value rather than changing the original, since strings and numbers are immutable. Always check whether a method's documentation says it returns a new object or modifies in place, and when in doubt, reassign the result back to the variable.

Putting it together: a small worked example

To see these pieces working as a group, here is a short script that takes a single raw order line of the kind you might read from a file (that proper file reading is the subject of part 7), converts its fields to the right types, and produces a formatted summary.

🐍Python
raw_line = "ORD-2048,3,19.99,false"

order_id, quantity_text, price_text, is_paid_text = raw_line.split(",")

quantity = int(quantity_text)
price = float(price_text)
is_paid = is_paid_text.strip().lower() == "true"   # text 'false' is not the boolean False

subtotal = quantity * price
tax = subtotal * 0.08
total = subtotal + tax

summary = (
    f"Order {order_id}: {quantity} units at ${price:.2f} each, "
    f"subtotal ${subtotal:.2f}, tax ${tax:.2f}, total ${total:.2f}, paid: {is_paid}"
)

print(summary)
# Order ORD-2048: 3 units at $19.99 each, subtotal $59.97, tax $4.80, total $64.77, paid: False

Notice the line computing is_paid: the raw field is the text "false", which is truthy as a string (it is non-empty), so bool(is_paid_text) alone would wrongly give True. The fix is to compare the cleaned, lower-cased text against the literal string "true" and let that comparison produce the correct boolean. This single line packs in string cleaning (strip, lower), a comparison operator, and the truthiness distinction covered earlier, which is exactly the kind of small, deliberate type handling that shows up throughout real data cleaning work.

Summary and what comes next

A variable in Python is a name bound to an object, and that binding can move to a different object, including one of a different type, at any time. Every object, however, has one fixed type for its whole life. The built-in types covered here, int, float, str, bool and None, along with the operators that combine them, cover a large share of everyday data work: arithmetic on numbers, cleaning and formatting text, checking conditions, and converting between the text a file gives you and the numbers or booleans you actually need to compute with. The recurring themes to carry forward are: Python will not silently fix a type mismatch for you, floats need careful comparison, and conversions fail loudly rather than guessing, which is a feature, not an inconvenience.

Part 3 builds directly on this: lists, tuples, sets and dictionaries are Python's built-in ways of holding many of these values together, and the indexing and slicing you practised on strings here will reappear, almost unchanged, when you start working with lists of numbers, rows of data, and lookup tables keyed by name.