Why this first step matters
Before you can clean a spreadsheet, train a model or call an API, you need a Python that actually works on your machine, a way to keep your projects from interfering with each other, and a place to write and run code. None of this is glamorous. It is also the single biggest source of frustration for people starting out in data work, because the errors you get from a broken setup look nothing like the errors you get from a mistake in your logic, and they are much harder to search for. This article is about getting that foundation right once, so that every later article in this path can simply say "open your environment and run this" without you wondering why it does not work.
We will cover three things: installing Python itself, creating and using virtual environments, and the difference between writing code in plain scripts versus notebooks. By the end you should have a working setup and, more importantly, understand why each piece exists, so that when something goes wrong later you know which part to check.
Getting Python onto your machine
Which Python, and why the version matters
Python is a programming language, but "Python" on your computer is also a specific program, an interpreter, that reads your code and runs it. There have been two major lines of Python: Python 2, which is no longer maintained, and Python 3, which is what everyone uses today. Within Python 3 there are minor versions, written as 3.9, 3.10, 3.11 and so on. Each minor version adds features and occasionally changes small behaviours. For this path, any reasonably recent Python 3 version is fine; you do not need to chase the newest release the day it comes out. If you are installing fresh today, picking a version from the last two or three years is a safe, boring choice, and boring is what you want from your language runtime.
On Windows and macOS, the simplest route is to download an installer from python.org and run it. On Windows, make sure you tick the option that adds Python to your PATH during installation; this is what lets you type python in a terminal from any folder and have it found. On macOS, python.org installers work well; some people instead use a package manager like Homebrew, which is also fine. On Linux, Python 3 is usually already installed, though sometimes an older or more limited version than you want, so you may still want to install a specific version through your distribution's package manager.
Checking what you have
Open a terminal. On Windows this can be Command Prompt, PowerShell, or the terminal inside an editor like VS Code. On macOS and Linux it is the Terminal application. Type the following and press enter.
python --version
You should see something like Python 3.11.4, or similar, printed back. On some systems, particularly macOS and Linux, the command might need to be python3 instead of python, because python alone is reserved for an old Python 2 or is not linked at all.
python3 --version
If neither command works, Python is not installed, or it is installed but not on your PATH, which means the terminal does not know where to find it. Reinstalling and making sure the PATH option is checked, or restarting your terminal after installation, fixes most of these cases. It is worth getting this one check working reliably before moving on, because every other command in this article assumes python (or python3) runs without error.
The problem virtual environments solve
Imagine two small projects on your laptop. One is a script that scrapes a webpage and relies on a library called requests, version 2.28. Months later you start a second project, a data analysis task, and it needs a newer library, pandas, which in turn depends on a newer version of requests, say 2.31, for a feature it needs. If you install everything directly onto your one, global Python, the second installation overwrites the first. Now your old scraping script might break, because it was tested against behaviour that version 2.28 had and 2.31 does not, or the reverse happened and it is the new project that behaves oddly. You cannot have two different versions of the same library sitting in the same global Python installation at once.
This is not a hypothetical edge case; it happens constantly in real data work, because the data science ecosystem moves quickly and different tutorials, courses and colleagues' code were often written against different library versions. The fix is to give every project its own isolated copy of Python's package storage, called a virtual environment. Inside that isolated copy, project A can have requests 2.28 and project B can have requests 2.31, and neither knows the other exists. Your actual Python interpreter is shared, but the installed packages are not.
A useful mental model is a toolbox per project. Your garage (the computer) can hold many toolboxes. Each toolbox (virtual environment) has its own wrenches and screwdrivers (packages), even if two toolboxes both happen to contain a 10mm wrench of slightly different brands. When you work on a project, you bring out that project's toolbox and put everything else away. You are never mixing tools from different boxes by accident.
Creating and using a virtual environment
Python 3 comes with a built-in module for this called venv; you do not need to install anything extra to get started. The general workflow is: create a folder for your project, create a virtual environment inside it, activate that environment, install the packages you need, and work. When you are done for the session, you deactivate it.
Creating the environment
Open a terminal, move into your project folder, and run the venv module, giving it a name for the environment. A common convention is to call it .venv, with a leading dot, so it is treated as a hidden folder on most systems.
mkdir my_project
cd my_project
python -m venv .venv
On macOS or Linux, replace python with python3 if that is what your check earlier showed you need. This command creates a folder called .venv inside my_project, containing a private copy of the Python interpreter and an empty space for packages. Nothing is active yet; you have only built the toolbox, you have not opened it.
Activating it
Activation changes your terminal session so that, for as long as it stays active, typing python or pip uses the environment's copies rather than the global ones. The command differs by operating system and shell.
# Windows, Command Prompt
.venv\Scripts\activate.bat
# Windows, PowerShell
.venv\Scripts\Activate.ps1
# macOS or Linux, bash or zsh
source .venv/bin/activate
When activation works, your terminal prompt usually changes to show the environment name in parentheses, something like (.venv) before the rest of the prompt. That small text is your signal that you are inside the toolbox. If you ever run a command and get an error about a missing package that you are sure you installed, the first thing to check is whether you forgot to activate the environment, or activated a different one than you intended. This is, by a wide margin, the most common environment mistake.
Installing packages inside it
With the environment active, you install packages using pip, Python's package installer, exactly as you would without a virtual environment, except now the packages land inside .venv instead of your system-wide Python.
pip install pandas requests
This downloads pandas and requests, along with anything they depend on, and installs them only into the active environment. If you open a new terminal window without activating the environment, or activate a different project's environment, pandas will not be there, because it never touched your global Python at all.
Leaving and removing an environment
To step out of an active environment, run deactivate. This works the same way across operating systems once the environment is active.
deactivate
To delete an environment entirely, you simply delete the .venv folder; there is no special uninstall command, because the environment is nothing more than a folder of files. This is also why it is normal practice to exclude .venv from version control systems like git: it is large, it is specific to your machine, and it can always be recreated from a short list of package names.
Recording what a project needs
If a virtual environment is a toolbox, you still want a shopping list of what should be inside it, so that you, or someone else, can rebuild the same toolbox on another machine. The conventional tool for this is a requirements file, usually named requirements.txt, listing package names and often the versions you tested against.
pip freeze > requirements.txt
pip freeze prints every package currently installed in the active environment, along with its exact version, and the redirect with the greater-than sign saves that output into a text file. A typical line inside it looks like pandas==2.1.0. Later, on a fresh environment, anyone can recreate the same setup with a single command.
pip install -r requirements.txt
You will see requirements.txt files in almost every real-world Python project, including some you will build later in this path, so it is worth getting comfortable with the two commands above now, even though we are not writing any project code yet.
Notebooks versus scripts
So far everything has happened in plain .py files run from the terminal, which is what a script is: a text file of Python code that runs from top to bottom when you execute it. There is a second common way to write Python, particularly in data work, called a notebook, and it is worth understanding what it actually is before deciding when to use it.
What a notebook actually is
A notebook, in the Jupyter sense, is a file, usually ending in .ipynb, that stores a sequence of cells. Each cell can hold code or formatted text, and cells run independently and in whatever order you choose, not necessarily top to bottom. When you run a code cell, the output, whether that is printed text, a table or a chart, appears directly underneath it and is saved as part of the file. This is the key difference from a script: a script is just instructions, while a notebook is instructions plus a running conversation with the computer, with the results kept visible alongside the code that produced them.
Underneath, a notebook is actually talking to something called a kernel, which is a running Python process, very similar to the interpreter you would get by typing python with no arguments in a terminal. The notebook interface, whether that is Jupyter Notebook, JupyterLab, or a notebook view inside an editor like VS Code, is really just a front end that sends code from a cell to the kernel and displays what comes back.
When notebooks help, and when they do not
Notebooks are well suited to exploration: looking at a new dataset, trying a transformation, plotting a quick chart, checking what a function returns, all without committing to a full program structure. The ability to run one cell, see the result, and then run the next cell based on what you just learned, is genuinely useful when you do not yet know what the final code should look like.
They are less well suited to code you want to reuse, test, or run unattended on a schedule, because the ability to run cells out of order is also their biggest trap. It is easy to end up with a notebook where cell 8 only works because you happened to run cell 3 twice earlier in the session, and the file, if you closed it and reopened it and ran everything top to bottom, would actually fail or give a different answer. Later in this path, when you reach topics like functions, modules and testing, you will be writing plain .py files, because that is where code that needs to be correct, reusable and checkable belongs. Notebooks will reappear for quick, hands-on exploration, particularly once we start working with real datasets and APIs, but they are a tool for a specific job, not a replacement for scripts.
A reasonable rule of thumb while you are learning: use a notebook when you are asking "what does this data look like" or "what does this piece of code do", and use a plain script once you are writing something you intend to run again, share, or build on top of.
Setting up Jupyter
Jupyter is not part of core Python; it is a package you install, like pandas or requests. The two common ways to install it are the classic Notebook interface and JupyterLab, a more fully featured interface built on the same underlying technology. Either is fine for this path; JupyterLab is the more actively developed of the two and gives you a file browser, multiple tabs and a console alongside your notebooks.
With your project's virtual environment activated, install one of them.
pip install jupyterlab
Once installed, launch it from the same terminal, with the environment still active.
jupyter lab
This starts a local server on your machine and opens a tab in your web browser showing the JupyterLab interface. Nothing is sent over the internet; the server runs locally and the browser is simply acting as the display for it. From the file browser panel you can create a new notebook, which gives you an empty cell ready for code. Typing a small expression and running the cell, for example 2 + 2, followed by pressing Shift and Enter together, runs that cell and shows 4 directly beneath it. This is a good first check that everything is wired together correctly.
One detail that trips people up: a notebook's kernel is tied to a specific Python environment, and it is possible to have Jupyter itself installed in one environment while trying to use a kernel from another, especially if you have several virtual environments on your machine. If a notebook cannot find a package you know you installed, check which kernel the notebook is using, visible in the interface, usually in a menu labelled something like Kernel. If you created a separate environment for a project and want it to show up as a selectable kernel inside Jupyter, install the ipykernel package inside that environment and register it.
pip install ipykernel
python -m ipykernel install --user --name=my_project
After running this, restarting JupyterLab and opening the kernel selection menu should show my_project as an option, pointing at that specific virtual environment's packages.
Where to actually write your code
For plain scripts, you need a text editor that understands Python, rather than a word processor. A widely used free option is Visual Studio Code, usually shortened to VS Code, which has an official Python extension that adds syntax highlighting, error checking, a way to run files with one click, and a built-in terminal so you do not have to switch windows to activate your environment and run commands. It also has reasonable support for opening and editing .ipynb notebook files directly, if you prefer not to run a separate browser tab for Jupyter. Other capable editors and setups exist, and if you already have a preference, there is no need to switch just for this path; the important thing is that whatever you use lets you edit a .py file and run it against the Python inside your active virtual environment.
Common mistakes
A short list of the problems that cause the most confusion when people are new to this setup, so you can recognise them quickly if they happen to you.
- Installing a package without activating the environment first, so it lands in the global Python instead of the project's toolbox, and later code that expects it to be in the environment fails to find it.
- Forgetting which terminal window or tab has the environment activated, especially when several are open at once; the prompt prefix in parentheses is there specifically so you can check at a glance.
- Running a notebook's cells out of order during exploration, then being surprised that running the whole notebook from the top produces a different result or an error.
- Using python when the system only recognises python3, or the reverse, and assuming Python is not installed when it is simply reachable under a different command name.
- Deleting a .venv folder while it is still activated in another terminal, which leaves that terminal pointing at a toolbox that no longer exists.
- Committing the .venv folder to version control, which bloats the repository with machine-specific files that a requirements.txt could recreate in seconds.
Most of these are quick to fix once you know what to look for, and all of them stem from the same root idea: your code's behaviour depends on which Python and which set of installed packages it is actually running against, and that is not always the one you assume.
Summary and what comes next
You now have Python installed and verified, you understand why virtual environments exist and how to create, activate, populate and record one with a requirements.txt file, and you know the practical difference between a script, which runs top to bottom and is suited to reusable code, and a notebook, which runs cell by cell and is suited to exploration. You have also set up Jupyter and know how to connect a notebook's kernel to a specific virtual environment.
None of this involved writing a program yet, and that is intentional. With a working Python, an isolated environment per project, and a place to run code, the next article in this path moves on to the language itself, starting with variables, types and expressions, the smallest building blocks you will use in every piece of code that follows.
Comments (0)
No comments yet. Be the first to share your thoughts.