A demo of an AI agent is easy to like. You ask a question in plain language and it does something useful.
Production is different. Production has controls, audits, and users who do unexpected things. An agent that is only a model and a prompt has no good answer to any of them.
The pattern I rely on is simple. The model proposes. A state machine decides what happens next.
Split the work
A language model is very good at a few things:
- Understanding a messy request.
- Proposing an answer from the context it is given.
- Explaining what was done in words a person can follow.
It is not a reliable keeper of process. It should not be the thing that decides which step you are in, what is allowed next, or when to stop.
That is the job of a state machine. Every request moves through named states. Each state has a clear purpose and a defined set of exits. The model is called inside a state. It never controls the flow between states.
Treat retrieval as a ladder
Before the model reasons, it needs facts. I order retrieval from the most precise method to the least:
- A graph query, when the question maps to known entities and relationships.
- A keyword search, when it does not.
- A vector search, when the wording is loose.
- The model's own fallback, only when the first three return nothing useful.
Each step down is more forgiving and less exact. Starting at the top means that when a precise answer exists, you get it.
Validate before you act
The model's proposal is a draft. Before anything is posted, paid, or sent, the system checks it.
The checks are deterministic. In finance work I keep rates, thresholds, and provisions in a rules table. Every rule has an effective date and a citation to its source. Nothing is hardcoded.
When a check fails, there are only two exits. The request goes back to reasoning with the reason for the failure, or it goes to a person. It never moves forward on a failed check.
The dull parts decide whether it survives
None of this shows up in a demo:
- Connection pools that hold under real load.
- Sessions that are isolated, so one user's context never reaches another.
- Queries that stay safe when users type things you did not plan for.
These are ordinary engineering problems. They are also where agentic systems most often fail in production.
What you get for the effort
You get a system you can explain.
For any outcome, you can say which state produced it, which rule it passed, and which source that rule came from. That is what an enterprise buyer needs before they let an agent near their books.
Autonomy is the easy part to build. Governance is what makes it usable.