What is an AI agent and how is it different from a chatbot
A chatbot is usually oriented toward producing a conversational response, while an agentic system is oriented toward achieving a goal through a sequence of steps and feedback. However, the boundary is not strict: a chatbot can invoke tools, and ordinary automation or RPA can run a multi-step process according to fixed rules without a model that independently selects actions. That is why it is more important to look at the architecture, the actions that are available, and the control points than the name of the product.
For this article, an AI agent is a software system in which a model helps choose the next steps and can call tools to achieve a given goal. This is a working definition, not a universal standard shared by all fields: under the same name, different architectures are found. Planning, state memory, interaction with the digital environment, task delegation, and the degree of autonomy can differ; not every one of these properties is required for any system that is called an agent.
International AI Safety Report 2025 links such capabilities not only to the model itself, but also to the surrounding software “scaffolding”. It provides models with tools, stores task state, organizes planning loops, and helps build a sequence of actions International AI Safety Report 2025.
Often, it is this scaffolding that turns a language model from a response system into an acting agent. For example, the model may not just suggest a draft email text, but also retrieve data from an allowed source, prepare a draft, check required fields, and pass the email to a human for approval. If the system is allowed to send, it will be able to do that as well—but at the same time, the consequences of an error will increase.
How an AI agent works
A typical cycle includes several stages:
- Goal acquisition. A user or another system sets the expected result and constraints.
- Planning. The agent breaks the goal into a sequence of sub-tasks.
- Tool selection. For each step, it can use search, a programming interface, a code interpreter, a database, or another permitted function.
- Action execution. The system calls the tool and receives the result.
- Checking the intermediate state. The agent evaluates whether the action brings it closer to the goal.
- Plan correction. If necessary, it repeats the cycle, selects another tool, or asks a human for a decision.
This scheme does not mean that the agent truly “understands” the task the way a person does. In practice, quality depends on the model, the software scaffolding, the available data, the set of tools, the way the goal is formulated, and the control mechanisms.
Also, autonomy is not a binary property. One agent only prepares recommendations. Another may read files, run code, or change records on its own. A third gets access to financial operations or critical systems. The higher the level of authority and potential harm, the stricter the constraints and oversight must be.
Use cases and the results of the 2025 report
Agents can be useful where the task can be decomposed into checkable steps and where the available actions have clear boundaries. Possible scenarios include searching and organizing information, preparing drafts, helping with program code, handling routine requests, and coordinating operations in digital systems.
International AI Safety Report 2025 describes an evaluation on 77 tasks of different types and difficulty—ranging from exploiting simple vulnerabilities in websites to training machine learning models. In the tested configuration, leading models with agentic scaffolding successfully completed almost 40% of the tasks; the result was comparable to that of people who were given 30 minutes per task. This is a result from a specific sample and experimental conditions, not an overall accuracy assessment of AI agents International AI Safety Report 2025.
Separately, the report provides a result for another, narrower group of seven difficult tasks simulating research and development in the field of AI, for example optimizing a neural network code. In two of these seven tasks, o1 made progress but did not achieve full success. These are not the same 77 tasks, and the claim is not about two fully solved tasks International AI Safety Report 2025.
These figures should be read as a historical snapshot. The study is presented in a report published in January 2025 and reflects specific models, agent scaffolding, and testing conditions at the time of the study. Therefore, it is not an assessment of the capabilities of models available in 2026, and it does not allow you to directly judge the reliability of an arbitrary production process.
A benchmark fixes specific tasks, a time limit, the environment, tools, and success criteria. In a real workflow, an agent may encounter an ambiguous goal, incomplete data, an unexpected file format, a changed interface, conflicting instructions, or malicious content. Therefore, these results help compare the tested configurations, but on their own they do not predict the reliability of a specific workflow.
Why autonomy amplifies the consequences of errors
A chatbot error can remain an incorrect text on the screen. An agent error can move into action: changing a file, calling an external service, sending a message, or executing a command. If an incorrect output becomes an input to the next step, the problem can propagate along the entire chain.
The main categories of risk include the following threats.
Errors in long chains
Even if each individual step looks plausible, small errors can accumulate. An agent can misinterpret the goal, choose an inappropriate tool, or continue execution after receiving a questionable result. The longer the task, the more points of failure there are. If multiple agents are involved, an error—or a compromised component—can pass an incorrect decision further along the chain; therefore, each agent needs separate trust boundaries and verification before transferring control OWASP AI Agent Security Cheat Sheet.
Malicious use
A user can intentionally assign the agent a dangerous task or try to bypass restrictions. The presence of tools and an autonomous loop potentially increases the scale of such actions. This does not mean that every agent is suitable for abuse or that harm is inevitable: the risk depends on the system’s capabilities, access rights, and protective measures.
Interception through external instructions
An agent working with web pages, documents, emails, or databases receives content from sources that cannot be automatically trusted. Such content may contain instructions attempting to change its behavior, make it disclose data, or trigger a dangerous tool.
International AI Safety Report 2025 notes that agents performing long-running tasks can be vulnerable to malicious instructions encountered during operation. The report also highlights risks from malicious use, reliability failures, and the weakening of human control International AI Safety Report 2025. This is a categorization of risk, not a claim that every agent is vulnerable to all kinds of influence.
Excessive authority
If an agent has more tools and permissions than are needed for the task, a single error leads to more serious consequences. Read access is different from the right to modify data; preparing a payment is different from executing it autonomously; creating a draft is different from automatically publishing.
Weakening human oversight
The less frequently a human checks intermediate decisions, the higher the chance that an incorrect action goes unnoticed. Reduced supervision is not always unacceptable, but it increases the importance of technical constraints, logging, and monitoring.
How to reduce the risks of AI agents
OWASP attributes threats to AI agents to the injection of malicious instructions, tool abuse, and excessive authority. OWASP’s recommendations are a practical reference for the community rather than a mandatory universal standard, a guarantee of eliminating risks, or an evaluation of the effectiveness of a specific deployment OWASP AI Agent Security Cheat Sheet.
1. Grant only the minimally necessary permissions
The agent should receive only the tools, data, and permissions needed for the specific task. If reading is sufficient, there is no need to provide write access. If an operation is performed within one system, access to the rest of the infrastructure is not needed.
It is helpful to separate roles and environments, limit the scope of credentials, and use time-limited permissions. This approach reduces the potential harm from an error or from intercepted control. Also consider confidentiality separately: data may enter the model’s context, be passed to a connected tool, or be stored in logs. Define what information the agent is forbidden to read, transmit, or write, and verify these boundaries during testing OWASP AI Agent Security Cheat Sheet.
2. Require explicit approval for sensitive operations
OWASP recommends requiring explicit approval for sensitive actions OWASP AI Agent Security Cheat Sheet. A person should confirm operations that can lead to financial loss, disclosure of data, irreversible changes, publication, or impact on critical processes.
Approval should describe not an abstract “next action”, but a specific operation, object, recipient, and consequences. Otherwise, the user may approve an action without understanding its scale.
3. Do not treat external content as trusted instructions
Data from emails, websites, and documents should be separated from system instructions. The agent must not automatically execute commands found in external content. It is especially important to check requests for secret disclosure, changing settings, or using new tools.
A single filter of keyword triggers is not enough: a malicious instruction can be disguised, split into parts, or embedded in a legitimate document. Protection should combine permission restriction, source validation, rules for tool invocation, and control of results.
4. Enforce limits technically
A textual prohibition inside an instruction is not sufficient. Constraints should be implemented at the architecture level: lists of allowed tools, parameter validation, an isolated execution environment, operation limits, and a ban on going beyond the established scope.
If the agent must not delete data, that function should not be available. If payments are allowed only up to a specified limit, the restriction should be checked by an external system rather than by the model itself. Minimal permissions determine what the agent can reach at all; architectural constraints determine which parameters and actions the system will accept regardless of the model’s response.
5. Log and track deviations
To investigate incidents, it is necessary to preserve the goal, plan, tool calls, results of checks, user confirmations, and final changes. At the same time, logging should account for confidentiality and should not become an additional source of secret leakage.
Monitoring helps identify unusual action frequency, access to atypical resources, recurring errors, and attempts to go beyond allowed permissions. This kind of behavior control aligns with the monitoring section in OWASP’s recommendations AI Agent Security Cheat Sheet.
6. Test before release and after changes
OWASP recommends checking an agent’s security before release and after substantial changes OWASP AI Agent Security Cheat Sheet. Re-evaluation is needed when replacing a model, connecting a new tool, expanding permissions, changing system instructions, or updating the software scaffolding.
Tests should cover not only correct task execution, but also tool failures, ambiguous commands, malicious external content, attempts to obtain additional permissions, and scenarios in which the agent is required to stop.
7. Provide for a safe stop
The agent should have stop conditions: exceeding a step limit, repeating the same error, conflicting instructions, absence of required data, or a request for a sensitive action without confirmation. Safe stopping is often better than continuing autonomously at any cost. For high-risk actions, human-in-the-loop control aligns with OWASP’s recommendations AI Agent Security Cheat Sheet.
How to assess whether an agent fits a specific task
Practical evaluation should account for not only the expected benefits, but also the consequences of a failure. Before deployment, it is useful to answer five questions:
- How unambiguous is the goal?
- Can each significant result be verified?
- Which data and tools are truly needed?
- What happens in case of erroneous or malicious actions?
- Where is a human decision required?
In a pilot, set measurable criteria in advance: the share of correctly completed tasks on a representative sample, the frequency of critical errors and unauthorized actions, the number of stops and human interventions, as well as cost and time. Separately test normal and adverse scenarios, compare with the current process, and set in advance the thresholds beyond which the agent is not allowed to take autonomous actions. This set of metrics is a practical evaluation framework for a particular deployment—not a literal requirement of NIST or OWASP; thresholds should be set with the consequences of errors in mind.
For a low-risk task, a more permissive mode can be acceptable. If an error affects money, people’s rights, safety, confidential information, or critical processes, strict constraints and independent checks are necessary.
Standards and interoperability of AI agents
On February 17, 2026 NIST announced the launch of AI Agent Standards Initiative. As of September 2026 the official NIST initiative page (updated August 14, 2026) describes ongoing work to support industry standards development, open protocols, and research on security and identification of agents, rather than a completed universal standard. The page highlights three tracks: support for the development of industry standards, contribution to open protocols, and research on security and identification of agents.
This initiative should not be presented as evidence that security problems are already solved. It should also not be conflated with an existing universal legal requirement. This material does not establish the legal status of AI agents in any country; NIST’s initiatives and OWASP recommendations are described as standardization and practical reference points, not as universal legal requirements.
Limitations of the available data
Many advanced systems and methods remain closed. Benchmarks do not cover all operational conditions, rare failures, or the risks of interaction among multiple agents; these limitations are discussed in the International AI Safety Report 2025. In addition, the rapid development of models means that individual results reflect the state of research at the time of publication, rather than a permanent limit of the technology.
Conclusion
An agentic system can go beyond dialogue: it can choose the next steps and call digital tools. But the existence of such a loop does not guarantee a successful solution to a complex task: capabilities and reliability depend on the model, the software scaffolding, the available tools, and control.
Therefore, an agent is appropriate in a specific multi-step process only after validation on representative tasks, taking the cost of errors into account. Minimal permissions, explicit confirmation of sensitive actions, technical constraints, monitoring, and repeated testing help reduce risk, but they do not eliminate it completely.