Why the Three Titles Get Mixed Up
If you read ten job postings for 'Data Scientist', 'Machine Learning Engineer' and 'AI Engineer', you will probably find that at least three of them could have the other title swapped in without the description changing much. This is not your imagination. The titles are genuinely overlapping, companies use them inconsistently, and a five-person startup often asks one person to do all three jobs at once. That said, in organisations large enough to split the work, the three roles do diverge in a consistent way: they differ in what question they are trying to answer, what they produce at the end of the day, and how close their work sits to a running production system.
This article builds intuition first with a single running example, then gives a precise description of each role, then shows how the three roles hand work to one another on a real project, and finally covers the mistakes people commonly make when reasoning about these titles.
A Scenario to Ground the Three Roles
Imagine an online bookstore that wants to add a feature: when a customer views a book, show them three other books they might like. This single feature, in a company with a reasonably sized team, passes through distinct kinds of work. Someone has to figure out whether this is even possible with the data available and which approach would work best. Someone has to turn a promising approach into a system that runs reliably for millions of requests a day. And, increasingly, someone has to decide whether to build a custom recommendation model at all, or to wire up an existing large language model that can read a customer's recent purchases and suggest similar titles in plain language. These three kinds of work map closely onto data scientist, ML engineer and AI engineer. Keep this bookstore in mind; we will return to it after describing each role.
The Data Scientist
What the role actually does
A data scientist's core job is to turn a business question into an answer using data, and to figure out whether a data-driven approach is worth pursuing in the first place. The emphasis is on exploration, statistics, and experimentation rather than on building the system that will run in production forever. A data scientist spends a lot of time with messy, incomplete data, asking questions like: is there actually a pattern here, or are we looking at noise? Which features matter? What would a reasonable baseline look like, and does our fancy model beat it by enough to justify the extra complexity?
A large part of the job is communication, not code. A data scientist often has to explain to a product manager or an executive why a model's 72 percent accuracy is or is not good enough, what the trade-off between false positives and false negatives means for the business, and whether a proposed metric target is realistic given the data.
A worked example
Back at the bookstore: a data scientist is handed the question 'can we recommend books people will actually buy?'. They start by pulling a sample of purchase history, say 50,000 customers and their past orders. They check basic things first: how many customers have bought more than one book (if most customers buy exactly once, collaborative filtering based on co-purchase patterns will struggle), how sparse the purchase matrix is, and whether genre or author metadata is reliable enough to use.
Suppose they find that 30,000 of the 50,000 customers have bought two or more books, which is enough signal to try a co-purchase approach. They build a simple baseline: recommend the most popular book in the same genre. Then they try a slightly more sophisticated approach, item-based collaborative filtering, and measure both using an offline metric such as precision at 3 (out of the three books recommended, what fraction were actually purchased later by similar customers in a held-out test set). Say the popularity baseline scores 8 percent and the collaborative filtering approach scores 14 percent. The data scientist's job is to report this clearly: the improvement is real and roughly 75 percent relative to the baseline, but 14 percent absolute precision is still low in plain terms, so they might also suggest running a small live experiment (an A/B test) before committing engineering resources to build the full system.
Notice what the data scientist produced: an analysis, a recommendation, and a working prototype, probably in a notebook, that proves the idea has merit. They did not build the service that will serve these recommendations to live traffic.
Tools of the trade
- Python or R for analysis, commonly with pandas, numpy and scikit-learn
- SQL for pulling and shaping data from warehouses
- Statistical methods: hypothesis testing, confidence intervals, A/B test design
- Visualisation libraries such as matplotlib or seaborn, and dashboard tools like Tableau or Looker
- Notebooks (Jupyter) as the primary working environment, rather than production codebases
The Machine Learning Engineer
What the role actually does
An ML engineer's core job is to take a model or approach that has been shown to work, and turn it into software that runs reliably, at scale, over time, without someone babysitting it. The emphasis shifts from 'does this work statistically' to 'does this work as a piece of software that other systems depend on'. This includes writing the pipeline that retrains the model on fresh data every week, serving predictions fast enough that a web page does not feel slow, monitoring whether the model's performance is quietly degrading, and making sure a bug in the data pipeline does not silently corrupt recommendations for every customer.
ML engineering borrows heavily from regular software engineering: version control discipline, testing, deployment pipelines, and thinking about failure modes. The difference from a general software engineer is that an ML engineer also needs to understand what can go wrong specifically with models and data, such as a model being retrained on data that accidentally includes the answer it is supposed to predict (a common and serious bug called data leakage), or a model's accuracy dropping because the real-world data has shifted away from what it was trained on (often called data drift or concept drift).
A worked example
Continuing the bookstore example: once the data scientist's prototype has proven promising and the company decides to build it, the ML engineer takes over. Their job includes several concrete pieces. First, turning the collaborative filtering logic, which may have lived in a notebook, into a proper Python package with tests, so it can be reused and maintained. Second, building a pipeline that recomputes recommendations as new purchases come in, rather than relying on a one-time batch computed during the prototype phase. This might run nightly, recomputing similarity scores for all 50,000 customers and storing the top three recommendations per customer in a fast key-value store so the website can fetch them in a few milliseconds.
Third, the ML engineer sets up monitoring. For instance, they track the click-through rate on recommended books every day, and set up an alert if it drops more than, say, 20 percent below its usual range, since that could indicate the pipeline broke silently (a common and embarrassing failure: the retraining job fails for three days, so every customer keeps seeing the same stale recommendations, and nobody notices until someone checks). Fourth, they think about scale: if the bookstore has 2 million customers rather than 50,000, item-based similarity computed naively will not finish overnight, so the ML engineer might need to use approximate nearest-neighbour search or batch the computation across multiple machines.
The output of this work is not a report, it is a running system: a service with an API endpoint, a scheduled retraining job, logs, dashboards and alerts. The ML engineer is judged on whether that system stays up, stays fast, and keeps producing good recommendations without manual intervention.
Tools of the trade
- Software engineering fundamentals: Git, testing frameworks, code review
- Containerisation and orchestration: Docker, Kubernetes
- Pipeline and workflow tools: Airflow, or similar schedulers for retraining jobs
- Model serving: frameworks or lightweight APIs (for example, a model wrapped behind a REST endpoint using FastAPI or similar)
- Cloud infrastructure: AWS, GCP or Azure, plus monitoring tools for latency, errors and model-quality metrics over time
- ML-specific libraries for training at scale, such as scikit-learn for smaller problems or PyTorch and TensorFlow for larger, more custom models
The AI Engineer
What the role actually does
The AI engineer title is newer and emerged largely because of the rise of large, pre-trained foundation models, the kind of general-purpose language and image models accessible through an API rather than trained from scratch in-house. An AI engineer's core job is to build products and features on top of these existing models, rather than training a new model from the ground up. The central skill is less about statistics or distributed systems, and more about understanding what a given model is good at, how to get reliable behaviour out of it, how to combine it with your company's own data, and how to build guardrails around something that is, by nature, a bit unpredictable.
In practice, this often means writing careful prompts and testing them systematically, connecting a model to external data sources so its answers are grounded in facts the company actually owns (an approach commonly called retrieval-augmented generation, where relevant documents are fetched and given to the model alongside the question), chaining multiple model calls together to break a complex task into steps, and building evaluation methods for a system whose output is natural language rather than a single number, which makes 'is this correct' a much harder question than it sounds.
A worked example
Suppose the bookstore later wants to add a chat assistant: a customer can type 'I liked this mystery novel, what else would I enjoy, and do you have anything similar in audiobook format?' and get a helpful, conversational answer. This is squarely AI engineer territory. The AI engineer would likely use a general-purpose large language model through an API rather than training one from scratch, since that would be prohibitively expensive and unnecessary for this task.
Their job includes retrieving the right context before calling the model: searching the bookstore's own catalogue for books similar in genre and tone to the one the customer mentioned, and including that catalogue information, along with stock and format availability, in the prompt sent to the model, so its answer is grounded in what the store actually sells rather than invented titles. They need to design the system prompt carefully, test it against a range of realistic customer questions, including awkward or adversarial ones, and decide what to do when the model is uncertain or when it should simply say 'I don't have that information' rather than guessing. They also need an evaluation plan: perhaps a set of 100 realistic test questions with expected qualities of a good answer, checked periodically, since there is no single accuracy number the way there is for a classifier.
Notice the overlap with the ML engineer: both care about latency, monitoring, and reliability of a running system. The difference is what sits at the core of the work. The ML engineer in the earlier example was responsible for training, retraining and serving a model built from the company's own data. The AI engineer here is mostly orchestrating an existing, very large model that someone else trained, and the hard problems are about context, prompting, grounding and evaluation of open-ended text rather than about training a model from scratch.
Tools of the trade
- APIs for large language or multimodal models, such as those from OpenAI, Anthropic or open-source models served through tools like Hugging Face's infrastructure
- Frameworks for building applications on top of these models, such as LangChain or LlamaIndex, which help with chaining calls and retrieval
- Vector databases for storing and searching document embeddings used in retrieval-augmented generation, such as FAISS, Pinecone or Chroma
- Prompt testing and evaluation tooling, which is still a fast-moving area without one dominant standard
- General software engineering skills, since the final product is usually a web or mobile feature, not a notebook
Comparing the Three Roles Side by Side
- Central question: data scientist asks 'is there a pattern, and is it worth acting on'; ML engineer asks 'how do we make this run reliably at scale'; AI engineer asks 'how do we get a general-purpose model to do this task well and safely'
- Typical output: data scientist produces an analysis, a report, or a prototype in a notebook; ML engineer produces a deployed, monitored service; AI engineer produces an application feature built on top of an existing model, often also a deployed service
- Where the model comes from: data scientist often builds the first version from scratch using classical techniques; ML engineer takes that model, or a custom deep learning model, and industrialises it; AI engineer mostly uses someone else's pre-trained foundation model and adapts it through prompting, retrieval or fine-tuning
- Core skill emphasis: data scientist leans on statistics and domain reasoning; ML engineer leans on software and systems engineering; AI engineer leans on application design, prompt engineering and evaluation of open-ended outputs
- Closeness to production: data scientist is usually furthest from the live system; ML engineer and AI engineer both typically own code that runs in production
- Time horizon of a typical task: a data scientist's analysis might take days to weeks; an ML engineer's pipeline work might take weeks to get right and then run for years; an AI engineer's prompt and retrieval setup might change every few weeks as the underlying model or product requirements shift
How the Roles Work Together on One Project
It helps to see all three roles on the same timeline rather than as separate categories. Return to the bookstore's recommendation feature and watch it move through the company.
First, a product manager asks whether personalised recommendations could increase sales. A data scientist spends two weeks exploring the purchase data, builds a notebook prototype, measures precision at 3 for a couple of candidate approaches, and writes a short report recommending the item-based collaborative filtering approach along with a proposed A/B test design. This is a judgement call backed by numbers, not a final system.
Second, once the test design is approved and a small live experiment confirms the uplift in click-through rate holds up with real customers, not just in historical data, the task moves to an ML engineer. Over the following month, they rebuild the logic as a proper package, set up a nightly retraining pipeline, wrap it behind a fast API, and add monitoring dashboards. This system then runs quietly for a year, serving recommendations to every visitor, with the ML engineer occasionally called in when an alert fires or when the model needs retraining on an updated feature set.
Third, eighteen months later, the company decides it wants a conversational layer on top: a chat box where customers can ask for recommendations in natural language instead of just browsing a list. An AI engineer is brought in. They do not touch the collaborative filtering pipeline the ML engineer built; instead, they use the output of that very pipeline, the top recommended books for a given customer, as one piece of context fed into a large language model, alongside catalogue search results, so the model can phrase a natural answer grounded in real, in-stock books. The data scientist may get pulled back in here too, to help design how to measure whether the chat feature is actually helping, since 'customer satisfaction with a conversation' is a much fuzzier thing to quantify than 'precision at 3'.
This example shows the roles are not competing for the same job; they are sequential and complementary stages of turning an idea into a reliable product, each with a different definition of 'done'.
Common Misconceptions and Mistakes
- Thinking data scientists do not write real code. Many data scientists write substantial Python and SQL daily; the distinction is that their code is usually exploratory and does not need to survive in a production system the way an ML engineer's code does
- Thinking an ML engineer is just a data scientist who knows Docker. The deeper skill is understanding software reliability, failure modes, and systems thinking, which is a different discipline, not an add-on tool
- Thinking an AI engineer needs to be an expert in the mathematics behind transformers or training large models. Most of the day-to-day work is about application design, prompt iteration, retrieval systems, and evaluation, not about deriving backpropagation by hand. Deep model internals knowledge helps but is not the core job
- Assuming job titles map cleanly onto this framework everywhere. Many companies, especially smaller ones, use 'Data Scientist' to mean someone who does all three jobs, or use 'AI Engineer' simply because it sounds current. Always read the actual responsibilities in a job description rather than trusting the title alone
- Believing one role is strictly 'above' another in seniority. They require different skills, not a ranked hierarchy; a very senior ML engineer is not a 'promoted' data scientist, they followed a different specialisation
- Assuming AI engineering replaces the need for ML engineering. Plenty of problems, such as fraud detection, demand forecasting or the recommendation system in the example above, are still better solved with a custom trained model than with a general-purpose foundation model, so the ML engineering skill set remains in heavy demand alongside AI engineering, not instead of it
Which Path Fits You
If you enjoy digging into a dataset to find out whether something is actually true, designing experiments, and explaining uncertainty to people who are not technical, the data scientist path will likely suit you. Strengthen your statistics, practice framing business questions as measurable ones, and get comfortable saying 'we don't have enough evidence yet' when that is the honest answer.
If you enjoy building things that have to keep working while you are asleep, thinking about what happens when a service gets ten times the normal traffic, and writing code that other people can safely build on top of, the ML engineer path will likely suit you. Strengthen your software engineering fundamentals first; the machine learning specifics can be learned on top of a solid engineering base much more easily than the reverse.
If you enjoy working close to the product, experimenting quickly with new capabilities, and figuring out how to make an unpredictable component behave reliably inside a larger system, the AI engineer path will likely suit you. Strengthen your understanding of how to evaluate open-ended outputs, how retrieval systems work, and how to design prompts and tests systematically rather than by trial and error alone.
It is also entirely normal to want a mix of all three, and many people move between them over a career. The skills build on each other more than they compete: understanding statistics makes you a better ML engineer when debugging a model that has started behaving oddly, and understanding software engineering makes you a better AI engineer when your prompt-based feature needs to handle real traffic reliably.
Summary and What Comes Next
Data scientists answer the question of whether a data-driven approach is worth pursuing and what it should look like, usually working close to raw data and statistics and producing analysis rather than production systems. ML engineers take a promising approach and turn it into software that runs reliably at scale, owning the pipelines, deployment and monitoring that keep a model working month after month. AI engineers build products on top of existing, general-purpose foundation models, focusing on prompting, retrieval, grounding and evaluation of open-ended outputs rather than training models from scratch. The three roles are not competing labels for the same job; they are different, complementary stages in turning a rough idea into a feature that real users depend on, as the bookstore example showed from prototype, to pipeline, to chat assistant.
If you are deciding where to focus next, a practical step is to pick one small project and try to carry it through all three stages yourself at a small scale: explore a public dataset and prototype a simple model, then wrap that model in a basic API with a test suite and some logging, then, if relevant, try grounding a language model's answers in the output of your own model. Doing this once, even on a toy problem, makes the differences between these roles concrete in a way that reading about them cannot fully replace.
Comments (0)
No comments yet. Be the first to share your thoughts.