Eight controls every production AI agent needs
A useful agent has a defined job, bounded authority, reliable evidence and a clear route back to a person.
A useful agent has a defined job, bounded authority, reliable evidence and a clear route back to a person.
An AI agent can interpret language, retrieve information and use tools. That combination makes it useful inside real work. It also means the system can act on an incomplete instruction, stale record or confident mistake.
Responsible deployment is a design discipline. The controls have to sit inside the operation, where a request becomes a decision and a decision becomes an action.
1. A bounded job
State the job in operating terms. “Help with finance” is too broad. “Extract supplier invoices, match them to purchase orders and prepare clean records for approval” defines a unit of work.
The job gives evaluation a boundary. It also makes exclusions explicit. The agent may prepare a payment record without having permission to release money.
2. Authoritative sources
Name which system or document can support each fact. The agent should know whether price comes from the ERP, a signed contract or a current catalogue. It should carry source references into consequential decisions.
Retrieval quality alone is insufficient. Freshness, permissions and conflicts between sources need defined behaviour.
3. Explicit permissions
Give the agent the smallest set of actions needed for the job. Separate reading, proposing and executing. Apply value limits, customer boundaries, time windows or record types where they reduce risk.
Authority should grow after evidence, not from enthusiasm during a demo.
4. A review threshold
Decide which cases require a person. Low confidence can trigger review, but consequence matters too. A high-confidence answer may still need approval when it changes a bank account, denies a claim or commits to a price.
The reviewer should receive the recommendation, source evidence, relevant rule and consequence of delay. A raw transcript is not a review experience.
5. Deterministic checks
Use ordinary code for rules that need exact behaviour. Totals must match. Required fields must exist. A user must hold the right permission. A payment cannot exceed an approved amount.
The model can interpret an invoice or explain an exception. Deterministic checks should enforce the constraints around that interpretation.
6. A complete audit record
Record the input, sources, model or rule version, proposed action, approval and final outcome. The purpose is operational learning as much as compliance.
When a customer challenges a decision or a quality measure changes, the team needs to reconstruct what happened without guessing.
7. Failure and recovery behaviour
APIs fail. Data arrives late. A model provider becomes unavailable. The agent needs a visible waiting state, safe retry policy and a route to manual completion.
Partial execution deserves special attention. If an action updates one system and fails in the next, the operation should show the inconsistency and assign recovery. Silent completion is the dangerous state.
8. Ongoing evaluation
The first evaluation set proves only that the system worked on the examples selected before launch. Production supplies new language, exceptions and changing policies.
Review representative cases continuously. Track human corrections, escalation reasons, unsupported answers and operating outcomes. When the same correction repeats, change the system rather than teaching people another workaround.
Our finance operations case shows these controls working together: deterministic checks move clean invoices, uncertain cases retain their evidence, and finance keeps authority over consequential exceptions. The same pattern applies across sales, service and public operations.
An agent becomes dependable when its limits are part of the design. It knows the job, works from evidence, acts inside permission and makes responsibility easier to see.
