Where Part 1 left off
The first article in this path defined an agent as a system that uses a language model to decide its own next action, rather than following a sequence of steps that a programmer wrote in advance. It also made the case that many tasks do not need this: if you already know the steps and their order, a fixed script or a simple prompt will be cheaper, faster and more predictable than an agent. That article left one thing loosely defined: what actually happens inside that loop of decisions. This article opens it up.
The short version, which we will unpack for the rest of this piece, is that an agent repeats three moves. It observes its current situation: the task, the conversation so far, and the result of whatever it just did. It thinks, meaning it runs that information through the model to decide what to do next. And it acts, meaning the decision leaves the model and does something in the world, such as calling a function, searching a database, or simply replying to the user. Then the result of that action becomes the next observation, and the cycle starts again. This is usually written as observe, think, act, and sometimes called the agent loop.
A loop you already know
The thermostat loop
Before bringing in a language model, it helps to see the same three-step pattern in something much simpler: a thermostat. A thermostat observes the room temperature. It thinks, in the loosest possible sense, by comparing that reading to a target temperature. It acts by turning the heater on or off. Then it observes again, a little later, and repeats. Nobody would call a thermostat intelligent, but it already has the full shape of the loop: a sensor reading, a decision rule, an effect on the world, and repetition until some condition is reached (the room is warm enough).
What a thermostat lacks is flexibility in the thinking step. Its decision rule is a single fixed comparison: is the temperature below the target, yes or no. It cannot be asked to also open a window if it is too hot, order a blanket if the heater is broken, or explain its reasoning to someone in the room. The rule was written once, by an engineer, and it never changes based on what the thermostat is seeing. That is precisely the boundary between a fixed control loop and an agent loop.
Scaling the loop up to language
An agent built around a language model keeps the same three-step shape but replaces the fixed decision rule with a model that reads the current situation, described in natural language, and produces a decision also in natural language (or in a structured format derived from it). The model can take in a much richer observation than a thermostat's temperature reading: the original task, a transcript of what has happened so far, the output of a tool, maybe some notes retrieved from memory. And it can produce a much richer decision than on or off: write some code, search the web, ask the user a clarifying question, or declare the task finished and hand back an answer.
This is also why agents are more expensive and less predictable than fixed scripts, a point the first article made at length. Every time the loop runs, you are paying for a model call, and the model's decision is not guaranteed to be the one a human would have picked. The loop gives you flexibility in exchange for cost and uncertainty. Understanding exactly what happens at each of the three steps is what lets you judge, for a given task, whether that trade is worth it.
Observe: what the model actually sees
Observing sounds passive, like the model simply looks at something that is already there. In practice, the observation at each turn of the loop is assembled deliberately by the code that surrounds the model, and what goes into it shapes everything that follows. A typical observation at the start of a loop iteration is built from several pieces stacked together into the model's context window:
- The system prompt: standing instructions about the agent's role, the tools it can use, and the rules it should follow. This stays mostly the same across iterations.
- The user's task or question, usually fixed for the whole run.
- A record of what has happened so far in this loop: earlier decisions the model made and the results they produced.
- The most recent piece of new information: the output of the last tool call, an error message, or a reply from the user.
- Sometimes, notes or facts pulled from a separate memory store, which a later article in this path covers in detail.
The important detail is the third and fourth items. On the very first iteration, there is no history yet, so the observation is just the system prompt and the task. From the second iteration onward, the observation grows: it now includes what the model decided to do last time and what actually happened when that decision was carried out. This is the mechanism by which the loop lets the model react to the world instead of only to the original instructions. A model asked to read a file does not know what the file contains until the file has actually been read and the contents have been placed back into its context as an observation.
This also means the observation keeps growing with every iteration, since each turn adds its own decision and result to the record. A loop that runs for twenty steps is carrying twenty steps of history into the twenty-first. That growth has real consequences for cost, for how well the model can still pay attention to the original task, and for what you choose to keep versus summarise or drop. That whole topic, usually called context engineering, gets a dedicated article later in this path; for now it is enough to notice that the observation is not a static snapshot, it is a running record that the agent loop keeps adding to.
Think: turning context into a decision
The think step is where the context assembled during observation is sent to the model, and the model produces a decision about what should happen next. It is worth being precise about what counts as a decision here, because it is easy to picture the model as deliberating like a person. What actually happens is that the model generates text (or, in some setups, a mix of text and a structured object), and the surrounding code interprets that output as one of a small number of possible decisions.
Two kinds of decisions
At the end of the think step, the model's output boils down to one of two things. Either it has decided to take an action that needs to happen outside itself, such as calling a tool, in which case it needs to specify which tool and with what arguments. Or it has decided it already has enough to answer the task, in which case it produces a final answer and the loop can stop. Everything the model says during the think step, including any reasoning it writes out along the way, exists to arrive at one of these two outcomes.
Many models, when prompted to act as agents, are encouraged or trained to write out some reasoning before committing to a decision; you will sometimes see this reasoning in the raw output and sometimes it happens in a part of the response that is not shown to the user. Either way, from the point of view of the loop, that reasoning is not itself an action. Writing "I should check the file first" does not read any file. Only the structured decision that follows, such as a request to call a read_file function with a particular path, is something the surrounding code can act on. This distinction matters because it is a common source of confusion: an agent that narrates a plan in its reasoning text has not done anything yet, it has only decided what it intends to do. Turning the exact mechanics of how that intention becomes a real function call, including how tools are described to the model and how arguments are validated, is the subject of the next article in this path. For this one, it is enough to know that the think step always ends in exactly one request: run this tool with these arguments, or here is the final answer.
Act: leaving the model and touching the world
The act step is where the model's decision stops being just text and starts having an effect. This step happens entirely outside the model, in the code that is orchestrating the loop. If the decision was to call a tool, the orchestrating code looks up the real function behind that tool's name, passes it the arguments the model produced, runs it, and captures whatever it returns, including any error if something went wrong. If the decision was a final answer, the orchestrating code simply returns that answer to whoever asked the question, and the loop ends.
It is worth dwelling on a point that is easy to skip past: the model cannot act on its own. A language model, on its own, only produces text. It cannot read a real file, send a real email, or query a real database. All of that capability lives in the code around the model. The think step produces a request; the act step is the only place where that request is actually carried out, by ordinary, non-magical code that calls an ordinary function. This is why an agent is sometimes described as a loop that wraps a language model rather than a language model that happens to loop: the model supplies judgement, the surrounding program supplies hands.
One consequence of this separation is that the result of an action is ground truth, in a way that the model's own narration is not. If a tool call to read a file returns an error because the file does not exist, that is a fact about the world, captured by real code. If the model, in its reasoning text, says "I have now read the file and it contains three numbers", without that claim being backed by an actual tool call and a real result, that is not a fact, it is a guess or a hallucination dressed up as a report. A well-built loop only ever feeds real action results back into the next observation, never the model's own unchecked claims about what it has done.
A complete walk-through
It helps to see the three steps run end to end on a small, fully deterministic task, so there is no ambiguity about what each step contains. Suppose the task given to an agent is: read the file numbers.txt, sum the numbers in it, say whether the sum is divisible by 7, and if not, say what is the smallest number you would need to add to make it divisible by 7. The agent has two tools available: read_file, which returns the contents of a file, and calculate, which evaluates an arithmetic expression.
Iteration one. Observe: the context contains the system prompt describing the two available tools, and the task above. There is no history yet, because this is the first turn. Think: the model has a task that depends on the contents of a file it has not seen, so the only reasonable decision is to read that file first. It produces a decision to call read_file with the path numbers.txt. Act: the orchestrating code runs the real read_file function, which opens the file and returns something like the list 12, 45, 7, 23, 9.
Iteration two. Observe: the context now contains everything from before, plus a new entry recording that read_file was called and that it returned those five numbers. Think: the model now has the numbers it needed. It can compute the sum itself in its reasoning, 12 plus 45 plus 7 plus 23 plus 9 is 96, but to check divisibility reliably it decides to use the calculate tool rather than trust its own arithmetic, and asks for 96 modulo 7. Act: the orchestrating code runs calculate, which evaluates the expression and returns 5.
Iteration three. Observe: the context now includes the sum, 96, and the remainder, 5. Think: with a remainder of 5, the sum is not divisible by 7, and the smallest number that would need to be added to reach the next multiple of 7 is 7 minus 5, which is 2 (since 96 plus 2 is 98, and 98 divided by 7 is exactly 14). The model decides it now has everything required and produces a final answer instead of another tool call. Act: the orchestrating code sees this is a final answer, not a tool call, and returns it directly: the sum is 96, it is not divisible by 7 because the remainder is 5, and adding 2 would make it 98, which is divisible by 7. The loop stops here because there is no further action to take.
The code below reproduces the shape of this exact loop without using a real language model at all. The think step is replaced by a few if and else statements standing in for what a model would decide; everything else, including the two tools and the loop structure, is real and runs as shown. This is only meant to make the three-step cycle concrete; a version built around an actual model, with real decision-making instead of these fixed rules, is what you will build hands-on later in this path.
# A fully mocked illustration of the observe-think-act loop.
# There is no language model here; mock_think stands in for what
# an LLM would decide, using a few hardcoded rules instead.
# Run with: python3 loop_demo.py
def read_file(_path):
return [12, 45, 7, 23, 9]
def calculate(expression):
return eval(expression, {"__builtins__": {}})
def mock_think(state):
"""Stands in for an LLM deciding the next action."""
if "numbers" not in state:
return {"action": "read_file", "args": {"path": "numbers.txt"}}
if "remainder" not in state:
total = sum(state["numbers"])
return {"action": "calculate",
"args": {"expression": f"{total} % 7"}}
return {"action": "final_answer", "args": {"text": (
f"Sum is {state['total']}. Remainder when divided by 7 "
f"is {state['remainder']}.")}}
state = {}
for step in range(5): # a hard cap so the loop cannot run forever
decision = mock_think(state)
action = decision["action"]
print(f"Step {step + 1}: think -> {action}")
if action == "read_file":
state["numbers"] = read_file(decision["args"]["path"])
elif action == "calculate":
state["total"] = sum(state["numbers"])
state["remainder"] = calculate(decision["args"]["expression"])
elif action == "final_answer":
print("Final answer:", decision["args"]["text"])
break
This script needs nothing beyond a standard Python 3 installation, no external libraries. Running it prints three lines showing the decision made at each step, read_file, then calculate, then final_answer, followed by the final answer itself: sum is 96, remainder when divided by 7 is 5. Notice the for loop with range(5): even though this particular run only needs three steps, the loop has a hard upper limit built in from the start. That is not an accident. A real agent, driven by a real model instead of fixed rules, has no guarantee of stopping on its own, and a hard cap on the number of iterations is one of the simplest and most important safeguards you can add. We will return to this under stopping conditions, and again in more depth in a later article.
Why a loop, and not one long prompt
It is worth asking directly why this needs to be a loop of several separate model calls at all, rather than one long, carefully written prompt that asks the model to produce the whole answer in a single shot. The answer is that a single prompt only works when all the information needed to answer is already available to the model before it starts writing. In the numbers.txt example, the model cannot know the sum until it has actually seen the contents of the file, and it cannot reliably know whether that sum is divisible by 7 without either doing the arithmetic itself, which is a known weak spot for language models on anything beyond small, simple numbers, or delegating the arithmetic to a tool that is guaranteed to get it right.
A single prompt has no way to pause, go fetch the file, and come back with an answer informed by what it found, because by the time the model is generating text, the prompt is already fixed. The loop solves this by letting the model stop, hand off to the surrounding code, receive new information back, and continue reasoning with that information now actually present in its context. Each iteration of the loop is, in effect, one more chance for the model to be told something it did not know when the loop started.
This is also a useful test for whether a task genuinely needs a loop at all. If you, as the person designing the system, already know exactly what information will be needed and in what order it should be fetched, you can fetch all of it yourself ahead of time, in ordinary code, and hand the model a single prompt that already contains everything it needs. The loop earns its keep specifically when the next piece of information needed depends on what the previous step returned, so that the sequence of actions cannot be fully decided in advance.
Why a loop, and not a workflow
A closely related question, and one that the first article in this path raised without resolving, is how this loop differs from a workflow: a fixed pipeline of steps such as "fetch data, then summarise it, then format the summary", where each step might itself call a model, but the order and number of steps are decided by the programmer and do not change at run time. A workflow can even call a tool and feed its result into the next step, which can look superficially similar to the loop described in this article.
The difference is where the decision about what happens next is made. In a workflow, that decision was made once, in advance, by whoever wrote the code: step two always follows step one, regardless of what step one actually returned. In an agent loop, that decision is made fresh at every iteration, by the model, based on what the previous iteration actually observed. In the numbers.txt example, nothing in the code decided in advance that read_file would be followed by calculate; the model decided that, at run time, because that is what the file's contents made necessary. A workflow version of the same task could certainly exist, and for a task this predictable it would probably be the better choice: read the file, sum it with ordinary code, compute the remainder with ordinary code, format the answer. There is no real need for a model to be making decisions at every step here; this example was kept deliberately simple so the mechanics of the loop would be easy to follow, not because it is a task that calls for an agent. The question of which of the two designs to pick for a given task, and the trade-offs between them, gets its own full treatment two articles from now.
When does the loop stop
The walk-through above stopped as soon as the model decided it had a final answer, which is the clean case. In practice there are several distinct ways a loop's execution can end, and it is worth naming them even though a full treatment of error handling and stopping conditions is a later article in this path.
- The model decides it is done and produces a final answer instead of another tool call; the orchestrating code recognises this and returns it. This is what happened in the worked example.
- The loop hits a hard limit set by the system around it, such as a maximum number of iterations, a time budget, or a cost budget, and is stopped from the outside regardless of what the model wants to do next.
- A tool call fails in a way the model cannot recover from, and the surrounding code decides to stop rather than let the model keep guessing.
- A person watching the agent, in setups that include a human in the loop, steps in and ends the run manually.
The second case deserves attention here because the demonstration code above already included it: the for loop with range(5) is exactly this kind of external safeguard. Nothing in principle stops a model from deciding, forever, that it needs to call one more tool before it can answer; a badly prompted or confused agent can genuinely loop without end, each iteration adding cost and bringing the context closer to its limit without ever reaching a final answer. Relying only on the model to decide when to stop is not enough for anything you plan to run unattended. A hard cap, checked by ordinary code rather than requested of the model, is the simplest possible safety net, and it costs nothing to add.
Common mistakes in the loop
A few failure patterns show up repeatedly once you start building and watching agent loops run, and most of them trace back to blurring the clean separation between the three steps.
- No hard cap on iterations. As covered above, trusting the model alone to decide when to stop invites runs that never terminate, quietly consuming time and money. Always set an explicit maximum number of iterations, time, or cost, enforced by code outside the model.
- Feeding the entire raw history back every single turn without limit. Because each observation includes everything from before, an agent that runs for many iterations can end up with a context so large that the original task gets lost in the noise, or the context window overflows entirely. Deciding what to keep, summarise, or drop is the subject of a later article on context engineering, but it is worth knowing from the start that unbounded growth of the observation is not a neutral default, it i
- Treating the model's own narration as if it were a confirmed action. A model can write "I checked the inventory and there are 40 units left" in its reasoning without any tool having actually been called. If the surrounding code is not careful to only treat real tool results as facts, and lets this kind of unconfirmed claim flow into later decisions as if it were observed data, errors compound quickly and silently.
- Letting a failed action pass unnoticed. If a tool call returns an error, that error needs to become part of the next observation just as a successful result would. A loop that swallows errors, or that always assumes the happy path, will have the model reason confidently about data it never actually received.
- Assuming more iterations always means better answers. Every extra turn of the loop adds latency and cost, and a model that is allowed to keep going does not automatically use the extra steps productively; sometimes it repeats itself or second-guesses a correct answer it already had. More looping is not a free upgrade to quality, it is a trade that needs to be checked against whether the task actually benefited from the extra steps.
Most of these mistakes share a root cause: forgetting that the three steps have genuinely different jobs. Observing is about what information is placed in front of the model and how much of it accumulates. Thinking is about the model turning that information into exactly one decision, nothing more. Acting is about code outside the model carrying out that decision and reporting back what actually happened, truthfully, including failures. Keeping these jobs distinct, and being deliberate about the boundary between them, is most of what separates a loop that behaves predictably from one that does not.
Summary and what's next
An agent's loop repeats three steps. It observes a context built from the task, the history so far, and the result of its last action. It thinks, meaning the model turns that context into exactly one decision: call a specific tool with specific arguments, or give a final answer. It acts, meaning code outside the model carries out that decision for real and captures what actually happened, which then becomes the next observation. The loop continues until the model produces a final answer, or until an external limit on iterations, time, or cost stops it first. This is the same basic shape as a thermostat's feedback loop, scaled up by replacing a fixed comparison rule with a language model capable of far richer decisions, at the cost of far less predictability and a real need for safeguards such as hard iteration caps.
This article deliberately left two things underexplained, on purpose, because they each deserve full treatment of their own. The next article, Tools and Function Calling, goes inside the think and act steps in much more detail: how tools are described to a model so it can choose between them, how its decision is turned into a structured, safely executable call, and what can go wrong along the way. After that, Workflows Versus Agents returns to the comparison raised here and gives you a clearer basis for choosing between the two designs for a given task. With the shape of the loop now established, those two articles can go deep into its moving parts without having to first explain what the parts are for.
Comments (0)
No comments yet. Be the first to share your thoughts.