What Separates a Real AI Agent From a Chat Bot or Pipeline With a Name
AI Agent... The word has been stretched to cover almost anything that calls a model twice. The useful test is not how autonomous a system is, but whether it holds a position it can be caught being wrong about.

Here is the working definition, and it is deliberately unglamorous. An agent is a model that runs a loop: it decides which tool to call, observes the result, decides what to do next, and takes real actions toward a goal.
That sentence does more work than its plainness suggests, because it has four clauses and each one excludes something currently being sold as an agent.
It decides which tool to call. The selection belongs to the system rather than to its author. If the order of operations is written down in advance, the thing is a pipeline. A pipeline can be excellent, and most professional work genuinely is a known sequence. It is still not an agent, because it cannot handle the case its author did not foresee.
It observes the result. What comes back from a call re-enters the reasoning instead of passing straight to the next stage. A system that fires a tool and hands the output onward has used a tool. It has not observed anything.
It decides what to do next. This clause is what closes the loop, and it is the one most often missing. A chain runs to its end. A loop can discover partway through that it asked the wrong question and go back, which is the only way a system handles a situation nobody anticipated for it.
It takes real actions toward a goal. Two requirements share this clause and both matter. Action means reaching past producing text into changing state; our own system prompt for Dodona calls this having hands and not only eyes, which is the right distinction. An agent that can only describe what it read has no way to be consequential and therefore no way to be accountable. The goal is what makes the loop purposive rather than merely iterative, and it is what a person is answerable for having set.
All four are required, and the definition is not generous. It is also, increasingly, commodity. Every serious framework released in the last two years hands a developer all four on the first afternoon, and that abundance is the source of the confusion rather than a cure for it.
We should say plainly that parts of our own system fail this test, because a definition written to include everything its author built is not a definition. The utility and classifier tier of our fleet does not qualify. A note-intent classifier running at temperature zero, an attachment classifier routing an inbound email, a backfill that determines when an event in an article actually occurred: each of those reasons over content with a model and each is mechanical by design. No tool selection, no loop, no goal beyond the one call it was invoked to make. The single-pass document parsers are the same. They read a file and return structured JSON. They are among the most carefully engineered components we run and they are not agents, and calling them agents would cost us the ability to say anything meaningful with the word.
"Agent" is the third word to go hollow
Which raises the question of why a sentence that plain needs writing down at all.
Twice already in these essays we have taken apart a term that used to mean something precise. "Human in the loop" began in control systems, where it described a process that could not complete without a deliberate human action, and it now stretches comfortably across both a physician reviewing every diagnosis and a compliance officer initialing a thousand reports an hour. "Interactive" once required a second party that could respond, and it now covers a filter and a sort. In both cases the word survived and the property it named quietly did not.
"Agent" is further along than either, and it matters more, because this is the term the entire industry is currently organizing itself around. Funding rounds, product categories, procurement checklists, and job titles all rest on it. If it means anything that calls a model more than once, then it means nothing, and a buyer has no way to tell a system that can act from a system that can only speak.
Building an agent got cheap, and that is why the word broke
We have written that the engineering budget used to do the editing, and that its collapse turned feature creep from a budget problem into a judgment problem. The same collapse hit agent construction and produced the same result one layer up. Standing up a tool-calling loop was a research project in 2022. It is an afternoon now, and a fair share of that afternoon goes to choosing the name.
This is a good development and we will not pretend otherwise. We built Vantage with coding agents. The SCCR fleet exists at all because the cost of keeping a hundred divergent repositories current fell far enough to make per-client forks affordable rather than ruinous. Cheap construction is why any of this is possible.
What cheap construction did not deliver is the part that makes an agent trustworthy. Our own white paper puts the ratio plainly: the agents are the easy thirty percent, and the structure underneath them is the seventy percent that makes them worth running. That was written about our delivery pipeline, and it holds just as well for the analysis fleet. The loop is the easy part. What the loop is pointed at, what it is permitted to decide, what happens when it is wrong, and how anyone would find out: that is the work, and none of it got cheaper.
The loop is the easy part. What it is permitted to decide, and how anyone would find out when it was wrong, is the work. None of that got cheaper.
Every false agent is an overcorrection toward something real
Fairness first, because each of these is a reasonable response to a genuine problem, which is exactly why capable teams walk into them.
| THE APPROACH | WHAT IT GETS RIGHT | WHERE IT BREAKS |
|---|---|---|
| The scripted workflow | Most professional work genuinely is a known sequence, and a fixed order is auditable, cheap, and repeatable in a way a loop is not | Nothing chooses. It is called an agent because it has several steps with a model in each one. It fails on the first input its author did not anticipate, and it fails silently, because there is no decision point at which it could have noticed |
| The named chatbot | A name and a face make a tool legible and approachable, and grounding answers in a retrieved corpus is real engineering that meaningfully reduces invention | Eyes and no hands. The persona is applied to the surface rather than expressed as a commitment, so the system can never be caught out of character, and a thing that cannot be caught cannot be trusted |
| The unbounded autonomous agent | This is the correct architecture, and open-ended tool selection is where the actual capability lives | No turn budget, no gate, no stance, no owner. It produces confident action nobody authored and nobody can answer for. Removing the limits is not more agency, it is the same agency with the accountability taken out |
The third deserves the most attention, because it is the failure a sophisticated team is most likely to mistake for ambition. An agent handed forty tools and a broad goal will do remarkable things in a demo and produce, in production, precisely the outcome we described when we first wrote about the boundary between people and machines: conviction without an owner. When it is wrong, and it will be, nobody can say why and nobody can answer for it.
A stance is not a costume
Here is where we part company with most of the field, and the claim is about craft rather than capability.
Almost everything shipped as an agent persona is decoration. A name, a tone instruction, sometimes an illustrated face. The prose comes out warmer and the reasoning underneath is untouched. We hold that a persona is real only when it consists of commitments that constrain the output, because a commitment is the only kind of personality that can be violated.
Our fleet is built that way, and the specifics are the argument.
The Thesis Checker is instructed to be neither cheerleader nor cynic, and the instruction has teeth because the agent scores execution and market view as two independent axes and composes them sixty forty. A company can execute well inside a breaking thesis, or miss badly inside an intact one. Collapsing the axes would hide whichever one is failing, so the split is not a stylistic preference. It is what makes the verdict legible.
Comparable Events is modeled on a court reporter, chosen for exactly one property, which is restraint. It records what was stated, attributes everything, and adds nothing. Its instruction is that eight events with honest nulls beat twelve with invented numbers, and a code-level gate drops any estimated multiple so the valuation math downstream can never anchor on a guess.
The Memo Analyst is deliberately adversarial, and exists as the counterweight to the deliberately even-handed Thesis Checker. It runs on the heaviest model we deploy. Its job is to attack the load-bearing assumptions in an investment memo, and it is told to credit genuine strengths with conviction, because an adversary that dislikes everything carries no information.
The Follow-On Advisor is forbidden from doing arithmetic. Four valuation methods are computed deterministically in code, with stage discounts fixed in advance. The model chooses how to weight them and writes the narrative, and it is explicitly barred from recalculating a figure or applying ownership. Its judgment is real, and it is bounded to the one place where judgment belongs.
Round Radar may disagree with the base rate by fifteen points and no further. A deterministic layer computes raise probabilities from runway and round cadence. The model may adjust each one by at most 0.15 in either direction, and only with a named reason. Its probabilities must be non-decreasing across widening windows, and code rejects a verdict that violates that.
Read those together and a pattern appears that runs opposite to what agent marketing implies. Every one of them is a description of what the agent is not allowed to do. The stance is the constraint. An agent with a personality and no constraints is wearing a costume, and the test of the difference is whether the thing can be caught out of character in a way that costs something.
Is Dodona a real agent
We should answer the question about our own system, since the definition above was written to be failable and some of our components fail it.
Dodona is the company-scoped oracle, and it satisfies all four clauses. Claude 4.6 Sonnet at temperature 0.4, with sixteen tools it selects among on its own. Seven read tools. Four that write to the client's database, so it can change company fields, log metrics, and create timeline events and notes. And four that reach past its own scope: it can trigger another agent's analysis, re-extract a document that parsed wrong, and file a ticket against the platform itself.
It also runs in a second form, which is the more literal case. askCompanyOracle is Dodona invoked as a sub-agent by Pythia, our fund-level oracle, from inside Pythia's own loop, with a company-scoped registry of eight tools and a six-turn budget. If it exhausts the budget, or finishes and reports low confidence with a named missing input, it escalates exactly once to twelve turns and the reason is logged. It closes by reporting its own confidence and the single most important thing it did not have. An agent dispatching another agent and reading back structured judgment about the limits of that judgment is about as unambiguous as the term gets.
The turn budgets matter as much as the tool count. Six turns with one audited escalation is a different design from unlimited turns, and that difference is the whole subject of this essay. Bounded, logged, and rare by design is what separates an agent from a process nobody is supervising.
The ticket nobody asked for
The strongest evidence we have is not in the architecture. It is in something Dodona did that nobody scripted.
In the middle of a conversation about a single company, Dodona noticed that the platform was summing tranches denominated in pounds as though they were dollars. It could have corrected the figure in front of it and moved on, which is what a retrieval system would have done. Instead it formed a judgment: that the error was not really about this company, that it would recur wherever the same code path ran, and that the correct response was to file it against the platform with a proposed fix. Then it did the same thing nine more times across the session, including catching a fair-market-value figure misstated by roughly ten million dollars.
Every one of those was a real defect. None was prompted. The disposition behind it lives in the agent's character section rather than its tool list: the prompt tells Dodona to take pride in improving the platform, to prefer a few sharp tickets over a flood, never to file one for correcting a single company's figure, and always to do it transparently. The disposition is authored. The specific judgments were not.
We are publishing the ten million dollar figure deliberately. A platform that reports its own arithmetic error is more trustworthy than one that has never found any, and we would rather be the firm whose agent caught it than the firm that waited for a client to.
That is what the last clause of the definition looks like when it is real. Not a system executing an action it was told to consider, but a system forming a view about the machine it is part of, setting itself a goal nobody handed it, and acting on that view without being asked.
An agent is something that can be wrong on the record
There is one more property, and it deliberately sits outside the definition, because a system can satisfy all four clauses and still not have it. It belongs to a different and more useful question, which is whether an agent should be trusted with judgment.
The property is a track record.
Our claim-bearing agents emit falsifiable claims at the moment they persist an opinion: the Thesis Checker's conviction score, the Follow-On Advisor's recommendation and fair-value range, the Financial Model's forward call, Round Radar's raise windows. Each claim carries a resolution specification and a close window, and the exact input bundle that produced it is frozen alongside it. A resolution engine later asks reality whether the claim came true and closes it with a score between zero and one. Ambiguous outcomes route to a queue a person settles.
Then the loop closes. Once an agent has accumulated at least twenty resolved claims inside a firm, a summary of its own accuracy is injected back into its prompt. Three disciplines govern that injection: a floor, so a small sample never misleads; firm scoping, so an agent sees only its own record within that client; and an anti-gaming instruction telling the agent the block exists to calibrate its confidence, and that it must not hedge, widen ranges, or soften a call in order to protect its score.
Consider what a workflow cannot do here. It cannot have a track record, because it never made a claim. It produced output. Nothing in it could turn out to be false, so there is nothing to score, and no amount of running it accumulates any evidence that it deserves to be believed. The same holds for a chatbot over a document index. It can be inaccurate. It cannot be wrong, in the specific sense of having committed to something that reality later refuted.
This is the instrument we described when we wrote about the boundary between people and machines, seen now from the agent's side. We argued there that a boundary which never moves is the respectable failure, because caution that is never converted into capability is a cost mistaken for a virtue, and that what lets the boundary move honestly is a measured, narrowing gap between what a machine produced and what the firm was willing to publish. The prediction ledger is the machinery that turns that gap into a number instead of an impression.
The questions that separate them are harder than autonomy
If you are evaluating something described as an agent, the useful questions are short, and the answers are hard to fake.
Does it choose its own tools, and can you see the trace of what it chose and why. Can it change anything, or only describe. What goal was it given, and who set it. What is its turn budget, and what happens when it reaches the ceiling. What is it forbidden from doing, and does the prohibition live in code or in a paragraph of the prompt.
Then the two that separate almost everything from almost everything else. What did it claim ninety days ago, and was the claim right. And, because a system can be right by saying the same safe thing about every case, did it say different things about different cases, and how much of the available surface did it address at all.
A vendor who can answer those has built the hard part. A vendor who reaches for autonomy as the answer has usually built the easy part and named it ambitiously.
None of this is a complaint about a young industry naming its components badly. That happens every time and it sorts itself out. It is an argument about where the difficulty actually sits, and it sits where it sat when we first wrote about Moravec's paradox. The machines hold the coverage: the tireless reading, the total recall, the tool use that never tires of the twentieth document. What makes that coverage worth having is the structure a person built around it, deciding what the machine is permitted to conclude, what it must never compute, when it has to stop and ask, and how the firm would find out if it were wrong.
Anyone can build a loop now. The loop was never the hard part.
Vantage runs a fleet of more than thirty specialist agents, each with an authored stance, bounded authority, and a scored track record. The machines take the coverage. Your team keeps the conviction.
Conviction Made Citable.

