Guide

AI copilot vs AI agent vs chatbot: how to choose

A chatbot answers in a chat window. A copilot answers and drafts from your own documents and systems, cites its sources, and leaves the decision to a person. An agent goes further: it calls tools and takes steps toward a goal within limits you set. Choose by who acts. If a person acts on the answer, build a copilot. If the system acts, build an agent, with approval gates.

By NEXIRT engineering · Last reviewed

Three words, three different jobs

"Chatbot," "copilot" and "agent" get used as if they were the same thing. They are not, and picking the wrong one can sink a project. The cleanest way to tell them apart is to ask who acts on the output.

A chatbot answers in a chat window. Whatever it says, the reader decides what to do with it, and the chatbot may know little about your business beyond a few pages of text.

A copilot works with your own documents and systems. It answers questions and drafts work, shows the source behind each answer, respects who is allowed to see what, and leaves the decision to a person. It is a very good assistant that never presses the button.

An agent presses the button, within limits. It calls tools, such as looking up a customer, creating a draft or filing a ticket, and takes a sequence of steps toward a goal. Anthropic's engineering team describes the distinction in its guide to building effective agents: workflows are systems where the model and the tools follow code paths written in advance, while agents are systems where the model directs its own process and tool use. Many useful business systems sit between the two: a fixed workflow with a model step or two inside it.

Compared onChatbotCopilotAgent
What it doesAnswers in a chat windowAnswers and drafts from your documents and systems, with sourcesTakes steps across tools toward a goal
Who acts on the resultThe readerA person, who decidesThe system, within limits and approval gates
Data it usesGeneral knowledge, maybe a few pagesYour documents and records, under your access rulesYour tools and records, through narrowly scoped functions
Main riskA confident wrong answerA wrong answer that looks well sourcedA wrong action
A good first useCommon public questions on a website"What does our policy say about this?"Triage, drafting and filing a request, with approval

How to choose

Four questions help settle it.

  1. Does the answer live in your own documents or systems? If yes, a general chatbot will guess. You need retrieval, which makes it a copilot at minimum.
  2. Does the job span several tools or steps? If the work is "read the request, check the account, draft a reply, log it," that is agent territory. If the work is "answer this question," it is a copilot.
  3. Is a wrong step reversible? A wrong answer is cheap to correct. A wrong email to a customer or a wrong update to a record is not. The less reversible the step, the more it needs a person's approval.
  4. Can you describe what good looks like? If not, you are not ready to build either one. A short Discovery Sprint exists to answer exactly this.

The usual path is to start with a copilot, add drafting, then add actions one at a time, each behind an approval, as the team learns to trust the system. Autonomy can wait.

Grounding and citations: what makes a copilot trustworthy

A copilot is trustworthy when you can check it. That depends on a technique called retrieval-augmented generation. AWS's description of how Amazon Bedrock knowledge bases work is a clear reference: your documents are converted to text, split into chunks, turned into embeddings and written to an index that keeps a mapping back to the original document. When someone asks a question, the system finds the chunks closest in meaning to it and hands them to the model together with the question.

Four rules turn that mechanism into something you can rely on:

  • Answer only from what was retrieved. The copilot is told to use the passages it was given and nothing else.
  • Link every answer to its source. A reader should be able to open the passage and judge it.
  • Say "not in the documents" when it is not. An honest gap beats a plausible guess.
  • Apply the same access rules as your sign-in. The copilot should never quote a document the person asking cannot open.

There is also tooling for catching the failures that slip through. Amazon Bedrock Guardrails include contextual grounding checks, which help detect responses that are not grounded in the source material or are irrelevant to the question, and sensitive-information filters that block or mask personal data in inputs and outputs. They are a second line of defense, not a substitute for good retrieval.

Evaluation: how you know it works

Evaluation is the step that separates a demo from a system.

Before building, collect a test set: real questions your staff ask, each paired with the passage that holds the right answer, plus some questions the documents genuinely cannot answer. Then measure three things. Did it find the right source? Did it answer correctly from it? And when the documents had no answer, did it say so?

Keep running that set. Re-run it whenever you change the model, the prompt or the documents, so a change that helps one question does not quietly break ten others. After launch, review a sample of real conversations weekly and flag answers that were unsupported. The numbers you track belong to your business and to the workflow; we agree on them with you during the Discovery Sprint and report on them after launch.

Guardrails: limits that do not depend on the model behaving

A good prompt asks the model to behave. A guardrail makes misbehavior hard. Build the second kind.

The OWASP Top 10 for LLM applications names the risks worth designing against. Two are especially relevant to a business system. Prompt injection is text the system reads, such as an email or a web page, that contains instructions meant to steer the model. Excessive agency is the system holding more permission than its job needs. The same list includes unbounded consumption, a reminder that costs also need limits.

The practical defenses are the same across projects:

  • Treat anything the system reads as untrusted input, never as instructions.
  • Give each tool the narrowest permission that does the job. A tool that reads orders should not be able to refund them.
  • Require a person's approval before anything irreversible or customer-facing.
  • Cap usage with daily token budgets and rate limits, so one runaway loop cannot become a surprise bill.
  • Keep an audit log of what the system did, for whom and when.

In the systems we build for clients, what an agent drafts for your customers waits for a person to approve it, and each approval, edit or denial is recorded.

Designing an agent that behaves

Tool use is where an agent's limits are enforced. In Anthropic's tool-use documentation, you define each tool with a schema, Claude decides when a tool is relevant and returns a structured request for it, and your own application executes the call and returns the result. For tools you define, your code is the gatekeeper.

That is the right place to put the rules. Before running a requested action, your code can check that the arguments are valid, that this user is allowed to do this, and that a person has approved it if it is irreversible. It can write the audit entry, and it can refuse. The model proposes; your code disposes.

A sensible rollout has stages. Start read-only, so the agent can look things up but change nothing. Move to draft-only, so it prepares work for a person to approve. Then let it act with approval, one kind of action at a time. Only later, and only for steps that are low-risk and reversible, consider letting it act alone. Anthropic's own advice, to find the simplest solution possible and add complexity only when needed, is the best summary of this staging.

When you do not need a custom copilot or agent

Use a general AI assistant directly for commodity tasks such as drafting a first pass of an email or summarizing a public article. Use the AI feature in software you already pay for when it covers the job. A custom copilot earns its cost when answers must come from your own data, with sources, access rules and an audit trail. A custom agent earns its cost when a job crosses several of your tools and the steps follow your own rules.

When not to hire NEXIRT

We are the wrong fit if you want an agent to act on your customers with no person approving it, if you need a formal certification we do not claim, if you want open-ended hourly developer capacity, or if nobody on your side can spare interview time. We also will not build an agent for a job a simple fixed workflow handles; a simpler system is easier to run and easier to trust.

Where to start

Describe the question your team asks most, or the job that eats the most hours, and talk to an engineer. We reply within one business day. For a question-heavy workflow, start with AI copilots and knowledge systems and the AI operations copilot concept study, a worked design with no client behind it. For a job that spans tools, see AI agents and assistants. Both begin with a 1–2 week Discovery Sprint that ends in a working prototype and a fixed-price quote. New to the vocabulary? Start with the guide to AI integration for small business.

Where to go next

Frequently asked questions

Is a copilot just a chatbot that has read my documents?
A chatbot that answers from your documents is the core of it. A copilot also links each answer to the passage it came from, respects who may see what, says so when the documents do not answer the question, and is measured against a test set of real questions.
When do I need an agent instead of a copilot?
When the job is a sequence of steps across tools rather than a question to answer: read an incoming request, look up the customer, draft a reply and file a ticket. Start with the agent proposing and a person approving.
How do you stop a copilot making things up?
It answers only from retrieved passages, links each answer to its source, and says so when nothing in your documents answers the question. A test set of real questions, reviewed weekly, shows where it still guesses.
Can a copilot show people documents they should not see?
It should not. Retrieval applies the same access rules as your sign-in, so the copilot never quotes a document the person asking cannot open.
Why not use ChatGPT or Claude directly?
For commodity tasks such as general writing help, do. A custom copilot earns its cost when answers must come from your own data, with sources, access rules and an audit trail.

Sources

Links checked . These are the primary documents the guide relies on.

  1. Building effective agents (Anthropic Engineering)Workflows versus agents, and keeping the solution as simple as the job allows.
  2. Tool use with Claude (Anthropic docs)Claude decides when to call a tool; your application executes client tools.
  3. Amazon Bedrock Guardrails (AWS)Sensitive-information filters and contextual grounding checks.
  4. How Amazon Bedrock knowledge bases work (AWS)How retrieval-augmented generation draws on a data store.
  5. OWASP Top 10 for LLM Applications 2025Prompt injection and excessive agency as named risks.
Next step

Have a workflow in mind?

Describe the work that eats your team's week. An engineer replies within one business day.