The difference between an agent that stays in production and one abandoned after a month is not the model. It is how narrowly its purpose was defined and what data it receives.
A narrow purpose wins
An assistant that can do anything cannot be evaluated, so it cannot be improved. An agent that classifies documents, drafts an answer, or updates a record has a right result and a wrong one, and that makes measurement possible.
The first question is not what the model can do but which task costs time today and repeats often enough to be worth automating.
Data decides quality, not the model
A good model over disorganised data gives answers that are confident and wrong. Most of the work in a project like this goes into making company information findable and trustworthy, not into writing prompts.
When the answer has to come from company documents, the agent searches them and cites the source, so the person receiving it can check. An answer without a source cannot be used in a decision.
Where a human stays mandatory
Any action that moves money, writes into an external system, or reaches a client is designed with human confirmation, at least at first. Not because the model fails often, but because a rare failure costs more than all the correct answers together.
The line between what the agent does alone and what it proposes for approval is the most important decision in the project, and it moves over time based on data rather than enthusiasm.
Without evaluation there is no progress
To know that a change improved something you need a set of real cases with the expected answer. Without it every adjustment is an impression, and the team ends up fixing things it broke with the previous change one at a time.
The same observability applies in production. What the user asked, what the agent found, and what it answered have to be reconstructible, otherwise a complaint cannot be investigated.
The real cost is maintenance
Models change, documents change, processes get rewritten. An agent left unattended degrades slowly without telling anybody, and trust in it disappears before the measurements do.
That is why the budget is planned around a year of operation rather than a delivery. The part that keeps the system useful is the one after launch.
Frequently asked questions about AI agents
Where does a company with no AI experience start?
From a single repetitive task with a verifiable result and low risk. A narrow pilot shows within weeks whether the data is good enough, and that is the information almost always missing at the start.
Does company data reach the model provider?
It depends on the chosen architecture and the contract with the provider. You can work with providers that do not train on your data, or with models run on your own infrastructure. It is a decision taken at the start, not at the end.
Does an AI agent replace people?
In practice it moves time out of repetitive work and into work that needs judgement. Projects that start from headcount reduction often fail, because they ignore precisely the difficult cases people were handling.
An agent that stays in production has a clear purpose, trustworthy data, and a person at the points that matter. If you want to evaluate a use case, see how we work with AI agents and automation.