A word that gets used for everything
Open almost any product page or conference talk about AI right now and you will see the word agent. A chatbot with a system prompt is called an agent. A script that calls an API and then calls a language model is called an agent. A multi-step data pipeline with an LLM somewhere in the middle is called an agent. The word has stretched to cover so much that it has nearly stopped meaning anything specific.
This series will use a precise, narrow definition, because the precision is what makes the concept useful. Once you can tell the difference between a workflow and an agent, you can make a real engineering decision instead of a marketing one. That decision, more than any clever prompt, is usually what determines whether a project works well, costs too much, or quietly breaks in production. This article builds that definition from two examples, then gives you a way to recognize, before you write a line of code, whether your task actually needs an agent at all.
Two tasks, side by side
Task one: email me a weather summary every morning
Suppose you want a system that, every morning at seven, fetches the weather for your city, writes one friendly sentence about it using a language model, and emails it to you. Think through what this requires. You need to call a weather API. You need to pass the result to a language model with a prompt like write one cheerful sentence about this weather. You need to send an email with that sentence. The order of these three steps never changes. You know it in advance, you could write it down today, and it will be exactly the same tomorrow and next year. The only part that varies is the weather data itself, and the sentence the model writes about it.
Task two: why did signups drop last Tuesday
Now suppose you want a system that investigates a sudden drop in website signups and reports back with a likely explanation. Here you cannot write down the steps in advance, because you do not know what you will find. Maybe the first thing to check is the analytics dashboard for traffic by source. If traffic looks normal but signups dropped, the next sensible step might be to check whether a form on the signup page is broken, which means looking at error logs. If the error logs are clean, maybe the next step is to check whether a deployment happened around that time, which means looking at deployment history. Which check you do second depends entirely on what the first check showed. You could not have scripted this in order ahead of time, because the right order only becomes clear as information arrives.
These two tasks look superficially similar, both involve a language model and some external data, but they are structurally different in exactly the way that matters for this whole learning path.
What makes something an agent
Here is the definition this series will use. An agent is a system in which a language model decides, step by step, what to do next, based on what it has observed so far, choosing among a set of available actions, and continuing until it judges the task finished or some limit is reached. The defining feature is not that a language model is involved, both of the tasks above involve one. The defining feature is who decides the order and number of steps. In a workflow, the developer decides this in advance and writes it into the code. In an agent, the model decides it at run time, call by call, based on what has happened so far.
The weather email task is a workflow. There is exactly one point where the model is asked to do anything creative, write a sentence, and everything else is fixed, ordinary code. The traffic investigation task, if you want it to actually work without you scripting every possible branch in advance, needs to be an agent. The model has to look at a result, decide what it tells you, and pick the next action itself.
A short piece of pseudocode makes the contrast concrete. Here is roughly what the workflow looks like.
def send_weather_email():
weather = get_weather("London")
summary = f"Today: {weather['condition']}, high of {weather['high']}C"
sentence = llm_call(
f"Write one cheerful sentence about this weather: {summary}"
)
send_email(
to="me@example.com",
subject="Today's weather",
body=f"{sentence}\n{summary}",
)
Every run of this function does the same three things in the same order. The language model is used once, for a narrow, bounded job, writing a sentence. It never decides whether to check the weather, or whether to send the email, or in what order, that control flow belongs entirely to the code you wrote. Now compare that to a sketch of the agent for the traffic investigation.
def investigate_traffic_drop():
history = []
tools = ["query_analytics", "check_error_logs", "check_deploy_history", "answer"]
while True:
step = llm_decide(history, available_tools=tools)
if step.action == "answer":
return step.content
result = run_tool(step.action, step.arguments)
history.append((step, result))
Notice what is missing from this version. There is no fixed sequence of tool calls written into the function. The loop just keeps asking the model, given everything observed so far, what do you want to do next, until the model itself decides it has an answer and chooses the answer action. The number of iterations is not fixed either, it might take two steps or six, depending on what the earlier results turned up. This loop, and how to build it properly with real stopping conditions, is the subject of the next article in this path. For now, the point is narrower: the order and length of the process are decided by the model at run time, not by you in advance, and that is what makes it an agent rather than a workflow.
The loop, briefly
You will often see the agent pattern described as a loop of observe, think, act. The model observes the current state, which includes the original task and the results of anything it has already done. It thinks, meaning it reasons about what the observation implies and what to do next. It acts, meaning it calls a tool, asks a question, or produces a final answer. Then the result of that action becomes a new observation, and the loop repeats. That is genuinely the whole idea, and the next article in this path goes through each part of it in depth, including how the model is actually made to stop, and what goes wrong when it does not. For this article, it is enough to recognize the shape, because it is the shape that separates an agent from a workflow, regardless of how sophisticated either one looks from the outside.
Where this kind of autonomy actually helps
Letting a model decide its own next step is useful precisely when you, the developer, cannot specify the right sequence in advance. A few situations where that is genuinely true:
- The number of steps needed is not known beforehand, like the traffic investigation, which might resolve in two checks or six.
- The right tool to use depends on what an earlier tool returned, so the branching cannot be written as a simple if statement without essentially reimplementing the model's judgment in code.
- The task varies a lot from one run to the next, so a single fixed script would need an enormous number of special cases to cover the realistic range of inputs.
- Part of the task is exploration, searching, reading, or querying something until enough information has been gathered, where you do not know in advance how much is enough.
- Recovering from a failure requires trying a genuinely different approach, not just retrying the same step, which means a decision has to be made about what to try next.
Notice that none of these are about the task being hard in some abstract sense. A workflow can call a very capable model and produce excellent writing or analysis. What makes a task agent shaped is specifically that the sequence of actions cannot be pinned down ahead of time.
When you do not need one
A large share of useful, production LLM systems are workflows, not agents, and that is not a compromise, it is usually the right engineering choice. Here are the signals that point toward a workflow.
- You can write down the steps today, and they would be the same steps for every reasonable input you expect to see.
- The task is repeated many times with the same shape, like classifying a ticket, summarizing a document, or extracting fields from a form.
- Consistency matters more than adaptability, you want every run to behave the same way so you can audit, test and debug it.
- Latency and cost matter, and a fixed sequence of one or two model calls will do the job that an open-ended loop would do in more calls, with no gain in quality.
- You need to be able to explain, after the fact, exactly what the system did and why, which is far easier when the order of operations is fixed in code than when it was decided on the fly.
If most of these apply, building an agent adds moving parts, cost, and unpredictability without buying you anything, because the thing that autonomy is good for, handling an unknown sequence, is not actually present in your task. A later article in this path, Workflows Versus Agents, goes through this decision in more structural detail and looks at hybrid designs in between the two extremes. For now, this simple checklist is enough to stop you from reaching for an agent out of habit.
A worked comparison: routing 500 support tickets
To make the cost of the wrong choice concrete, consider a support team that wants every incoming ticket classified into one of five categories, billing, bug report, feature request, account access, or other, and routed to the matching queue. Say 500 tickets come in during a day.
As a workflow, this is one language model call per ticket. The prompt is fixed: here is the ticket text, choose exactly one of these five categories, respond with just the category name. The code reads the response and routes the ticket to the matching queue, with a default queue if the response does not match one of the five expected labels. That is 500 calls, each doing the same bounded job, each costing roughly the same amount and taking roughly the same time. The behaviour is consistent, because every ticket is handled by the identical prompt, and you can test the whole system against a labelled set of past tickets before it ever sees a real one.
As an agent, the same task might look reasonable on paper, give the model the ticket and a set of tools, including a search of the knowledge base, a lookup of the customer's account, and an action to assign a category, and let it decide what to check before deciding. In practice, for a task this bounded, that extra freedom tends to cost more than it returns. Some tickets might trigger three tool calls before the model commits to a category, others might trigger none, which means your cost and latency per ticket now vary in a way that is hard to predict or budget for. Two tickets with nearly identical wording might end up taking different paths through the loop and occasionally landing in different categories, which is a harder thing to debug than a wrong answer from a fixed prompt, because the inconsistency comes from the model's own run-to-run choices rather than from a classification mistake you can simply fix in the prompt. None of the extra tools actually help with the core decision, since classifying a ticket by its text does not usually require exploration, it requires reading the text carefully, which a single well-written prompt already does.
The general shape of this comparison holds beyond ticket routing. When a task is the same operation repeated many times over similar inputs, a fixed workflow tends to be cheaper, faster, and easier to test, precisely because there is nothing to explore. The autonomy of an agent earns its cost on the traffic drop kind of task, where the next step genuinely cannot be known until the previous one returns a result. It does not earn its cost on the ticket kind of task, where the next step, route after classifying, is always the same.
Common mistakes
A few patterns come up repeatedly once you start looking for them.
- Calling a single LLM call an agent. If there is one model call with no loop and no choice of what to do next, it is a model call inside a workflow, not an agent.
- Building an agent for a task whose steps you could write down today. If you can already describe the correct sequence, write that sequence in code and save the loop, the tool definitions, and the unpredictability for a task that actually needs them.
- Treating more autonomy as automatically more capable. An agent given freedom to choose its own steps is not smarter than the same model used well in a fixed sequence, it is just less constrained, which is only an advantage when the constraint was the problem.
- Reaching for an agent because it is the current trend, rather than because the task has an unknown number of steps or depends on intermediate results. The checklist in this article is a better starting point than fashion.
- Describing the model's behaviour as wanting or deciding in a human sense. The model is producing the next token, including tokens that look like a tool call, based on everything in its context. That mechanical description matters later in this path, particularly when debugging why an agent looped forever or stopped too early, because the fix is almost always about what was or was not in the context, not about the model's intentions.
Summary and what comes next
An agent, as this series uses the term, is a system where a language model decides its own next step, based on what it has observed, until it judges the task done. A workflow is a system where the developer decides the sequence of steps in advance, even if a language model does some of the work inside that sequence. The question to ask about any task is not whether a language model is clever enough to handle it, but whether you can write down the right order of actions before you run it. If you can, a workflow will usually be cheaper, faster, more consistent, and easier to debug. If you genuinely cannot, because the right next step depends on information you do not have yet, that is where an agent earns its cost.
The next article in this path opens up the loop itself in detail, observe, think, act, including how an agent actually decides when to stop, what goes into each step, and what tends to go wrong when the loop is built carelessly. Everything in this article was about recognizing when you need that loop at all, the next one is about building it properly once you do.
Comments (0)
No comments yet. Be the first to share your thoughts.