In this article4 sections
Start with the decision being delegated
An automation can use an AI model without needing an agent. Anthropic draws a useful architectural distinction: a workflow follows predefined paths through models and tools, while an agent lets a model decide how to proceed and which tools to use. The terminology varies, but the control boundary is a practical way to compare designs.
Consider a system that takes meeting notes, extracts action items, checks that each item has an owner, and saves a draft. The inputs contain ambiguity, yet the sequence is known. A fixed workflow can use a model for extraction while ordinary code controls validation and saving. Calling the whole system an agent would not make the process better.
Prefer a workflow when the route is known
Google’s architecture guidance suggests considering non-agentic solutions for predictable tasks, including document summaries, translation, and feedback classification. It also asks teams to define response-time expectations, costs, and the need for human involvement before selecting a pattern.
For a repeatable process, write down the steps and the possible branches. If those branches are manageable, keep them explicit. For example, an incoming request might be classified, sent through one of three predefined handlers, and placed in a review queue. An uncertain classification can go to a person rather than triggering an open-ended attempt to solve the entire request.
This is a design recommendation, not a claim that every workflow is reliable by default. Model outputs still need checks, and a fixed path can still encode the wrong policy. The benefit is that you can inspect where decisions occur and test each boundary.
Use an agent when new evidence changes the route
An agent becomes more plausible when the required steps cannot be usefully specified in advance. A research task might reveal a missing source that requires a different search. A software investigation might need to choose which files to inspect after reading an error. In both cases, intermediate results help determine the next action.
Anthropic describes this as a loop informed by feedback from the environment, with clear stopping conditions. Google recommends beginning agent development with a single agent so the core logic and tools can be refined before introducing coordination between several agents. More agents bring additional evaluation, security, reliability, and cost considerations.
Define permissions and success before the loop
AWS’s security guidance recommends limiting tool access and adding controls for sensitive actions. Translate that into concrete boundaries: which records may be read, which fields may be changed, and which actions require review? A research assistant may need browsing and draft creation without needing permission to send messages or publish results.
Set a time or step budget and an escalation path for missing information. Preserve the tool results and final outputs needed to understand failures. Then compare the agent with a simpler baseline using representative tasks, including awkward cases. Anthropic’s evaluation guidance emphasizes that agents’ multi-step behavior makes systematic evaluation particularly important.
Measure completed tasks, errors, latency, cost, and reviewer effort. Count a task as successful only when its actual outcome meets your criteria. A polished explanation of what the system intended to do is not evidence that the requested change happened. Add autonomy when the measured benefit justifies it.
Sources & further reading
This is a source-based explainer, prepared with AI assistance and checked against the references above. It is not a hands-on product test. Read our editorial approach.