"Agent" gets used for everything from a single function-calling request to a research bot that runs for days. Underneath, the idea is narrow. An agent is a language model placed inside a loop: it can take an action, see what happened, and decide what to do next, until it proposes an answer or the runtime stops it at a turn, time, cost or policy limit.
Same model, different harness
A plain model call is one round trip. You send a prompt, the model writes text from what it already knows plus whatever you pasted in, and the interaction is over. It cannot look anything up, check its own answer, or try again unless you send another request.
Nothing about the model changes when it becomes an agent. What changes is the code around it. Instead of stopping after one response, the harness feeds the model's tool requests, and their results, back to it turn after turn.
Take a hotel guest who writes "My AC isn't working." A plain call replies with sympathy and promises to contact maintenance; no ticket exists and nobody has been told. An agent identifies the guest and the room, checks the maintenance policy, creates a repair ticket, assigns available staff, starts the SLA clock and lets the guest know.
The loop: think, act, observe
An HR assistant has exactly one tool, search_hr_docs(query). Someone asks whether fathers can carry forward unused parental leave.
- Think. The answer needs the parental leave policy, so search for it.
- Act. Call
search_hr_docs("parental leave"). - Observe. "Fathers receive 30 days of parental leave." That answers how much, not whether it carries forward.
- Act again. Call
search_hr_docs("carry forward policy"). - Observe. "Unused parental leave may carry forward up to 10 days into the next year."
- Stop. Both facts are in hand, so the model writes the answer.
If you have written retry logic, the shape is familiar, with one difference: the model proposes the next operation. The runtime still enforces a maximum number of turns, per-call deadlines, cancellation and a total cost allowance, however keen the model is to continue.
Engineering Insight
The model never touches a database, an API or a file. It emits a structured tool request, and your orchestrator validates it, runs it and returns the result. The model reaches the outside world only through the capabilities your application chooses to expose.
Five parts, and where autonomy stops
| Part | What it does |
|---|---|
| Model | Decides the next step from everything observed so far |
| Tools | The only way to read or change anything outside the model's own text |
| Memory and context | The current situation: this conversation and the tool results so far |
| Policies and guardrails | What the agent may and may not do, such as needing approval above a refund limit |
| Loop | Runs think, act, observe, and checks whether to stop |
The usual mistake is treating "agent" as all or nothing. In the hotel example the model is good at recognizing that "my AC is making a weird noise" is a maintenance request and picking the right workflow. The ticket lifecycle, the SLA calculation, permissions, approvals and audit events should stay deterministic code. Let the model decide which action fits; let your workflow engine decide how it runs.


