What Is an AI Agent? A Business Definition (Not a Hype One)
Everyone sells an 'agent' now. Most of them are chatbots with better marketing. Here's the working definition we use with clients.
Written, fact-checked and maintained by the gAIcko Editorial Team. Corrections: admin@gaicko.com.
What is an AI agent?
An AI agent is software that takes a goal, plans the steps to reach it, executes those steps using external tools or APIs, and reports the result. Tool use and multi-step execution are what distinguish an agent from a chatbot — not how conversational it sounds.
The short version
An AI agent is software that accepts a goal, decides on a sequence of steps, carries those steps out using external tools, and reports what happened. The defining traits are autonomy over the plan and the ability to act on systems outside itself. Everything else — how conversational it sounds, which model powers it, whether it has a friendly name — is packaging.
Why the definition matters commercially
The word "agent" now appears in almost every enterprise software launch. Buyers who cannot separate an agent from a scripted assistant end up paying agent prices for chatbot capability, then blame the technology when the business case fails. A precise definition is not pedantry; it is procurement hygiene.
In practice the confusion costs money in three ways: overpriced licences, underestimated integration work, and governance gaps where nobody realises the system can actually write to production systems.
The four capabilities that qualify
1. Goal intake
The system receives an outcome ("reconcile yesterday's failed payments"), not a command ("run script 4"). If every action must be spelled out by a human, you have an interface, not an agent.
2. Planning
It decomposes the goal into steps and can revise that plan when a step fails. A fixed if/then flowchart authored by a developer is workflow automation — valuable, cheaper, more predictable, but not agentic.
3. Tool use
It calls APIs, queries databases, sends messages, moves files. This is the capability that creates both the value and the risk. An agent that cannot act can only advise.
4. Observability and guardrails
It logs what it did, why, and at what cost, and it operates inside limits: spend caps, allow-listed tools, human approval on irreversible actions. Vendors that skip this are shipping a prototype.
What does not qualify
- Single-turn chatbots. Retrieval plus a good answer is useful; it is not agency.
- RPA scripts. Deterministic UI automation. Brittle, but predictable — and often the correct answer.
- Prompt chains without tools. Text in, text out, several times. Still text.
- "Copilots" that only draft. If a human executes every action, the human is the agent.
A worked example
Consider invoice chasing. The chatbot version answers "what is our overdue balance?" The workflow-automation version sends a templated reminder on day 30. The agent version reads the ledger, groups debtors by risk and relationship, drafts differentiated messages, checks whether a payment plan already exists in the CRM, sends the low-risk ones automatically, and routes the top ten accounts by value to a human with a recommended action and the evidence behind it. It logs every send, stops after a configured number of outbound messages per hour, and never touches accounts flagged in dispute.
That last sentence is where most vendor demos quietly end.
How to interrogate a vendor
- Which external systems does it write to, and with what credentials? Read scope only is a very different product.
- What happens when a step fails mid-plan? Retry, rollback, escalate — or silent partial completion?
- Where is the audit log, and can I export it? Governance teams will ask within a quarter.
- What are the spend and rate caps, and who can change them?
- Which decisions require human approval, and is that configurable per action type?
- How do you evaluate quality? If the answer is "customers tell us," there is no evaluation harness.
When an agent is the wrong tool
Agents earn their complexity when the path to the outcome varies. If the path is stable, a deterministic workflow is cheaper to build, cheaper to run, and far easier to certify. High-volume, low-variance, high-consequence processes — payroll, statutory filing, funds movement — usually belong in deterministic pipelines with AI confined to classification or drafting at the edges.
A reasonable rule: use an agent where the variance is genuine and the cost of a wrong step is recoverable. Use a workflow where the process is fixed or the failure mode is expensive.
Maturity ladder
- Assistive. Drafts and suggests; a human executes.
- Supervised execution. Acts on reversible tasks; humans approve the rest.
- Bounded autonomy. Acts within allow-listed tools and hard caps; humans review samples.
- Delegated ownership. Owns an outcome with SLAs and an escalation path.
Most organisations should target level two within a quarter and level three within a year. Level four is rare and, where it exists, is heavily instrumented.
Practical next step
Pick one process, write down the goal in a single sentence, list every system that must be touched, and mark each action as reversible or not. If the reversible actions cover most of the work, you have a viable first agent. If not, automate the deterministic core first and revisit.
Frequently asked questions
What is the difference between an AI agent and a chatbot?
A chatbot responds within a conversation. An agent plans multiple steps and executes them against external systems, then reports the outcome. Tool use and multi-step execution are the dividing line.
Do AI agents replace RPA?
No. RPA remains better for stable, high-volume, deterministic processes. Agents add value where the path to the outcome varies and judgement is required. Most mature stacks run both.
Are AI agents safe to let act autonomously?
Only within bounds. Safe deployments allow-list tools, cap spend and rate, require approval for irreversible actions, and keep an exportable audit log of every step.
What does an AI agent cost to run?
Costs split into model inference, integration engineering and ongoing evaluation. Inference is usually the smallest line. Budget most of the first-year cost for integration, monitoring and change management.
How do you measure whether an agent is working?
Track task completion rate, escalation rate, cost per completed task, and a quality score from sampled human review. Improvement in all four over time is the real signal.
Which processes are best for a first AI agent?
Processes with variable paths, reversible actions, clear success criteria and reasonable volume — support triage, lead qualification, invoice chasing and data reconciliation are common starting points.
Sources and further reading
- Anthropic — Building effective agents
- NIST AI Risk Management Framework
- OpenAI — Practices for governing agentic AI systems
Revision history
- — Published in full with worked examples, FAQs and sources.