Every team shipping an agent eventually asks the same question: how do we know it did the right thing, every time, without a person reading every trace? The usual answers are a hand-picked test set, an LLM judge and a dashboard of thumbs-up ratings. They help, but none of them tells you, on a Tuesday afternoon, whether the thousands of things the agent did this morning were correct.
Finance operations had this problem decades ago. Batch jobs, bank feeds, interfaces between systems written by different vendors in different decades: all automated, none fully trustworthy. The answer the profession settled on was reconciliation. You don’t trust the process. You compare its output to an independent record, you agree on what counts as a match, you queue whatever doesn’t match, and a named person signs off that the books tie out. Then someone else checks that the checking works.
I spent several years building reconciliation systems before I built agents, and I keep watching agent teams reinvent those parts one at a time, usually missing the piece that made the original work. This post maps the reconciliation toolkit onto agent harnesses. It is design analysis with illustrative code: nothing here is a measured result, and the TypeScript is a sketch, not something I’ve run.
Why this is timely
Two things landed in the same week of September 2026 and they point at the same gap.
On 28 September, Trintech, a close and reconciliation software vendor, launched three finance agents under the banner “governed autonomous finance”. Its CEO, Darren Heffernan, is quoted saying: “Finance teams do not need more AI that simply tells them what to do.” What they need, he says, is AI that does the work “in ways they trust.” The release says the Exception Management Agent “groups related exceptions under a shared root cause instead of leaving a team to work through them one line at a time”, and that “every recommendation still comes with its evidence, methodology, approvals and audit trail attached”. I haven’t used the product, so I’m making no claim about how well it works. What interests me is the vocabulary. A finance vendor selling agents leads with root causes, approvals and audit trails, because that is what its buyers already check before they trust anything automated.
Four days earlier, François Zaninotto at Marmelab published The State Of AI Harness Engineering 2026. One line stood out: “Out of the 97 harnesses big enough to test, 60% have neither a test nor an eval.”
My reading of the two together: finance buyers won’t accept an agent without controls, and much of the wider agent world is still shipping harnesses without them. Part of the reason, I think, is that most domains don’t have a ledger. When the books don’t balance, an accountant knows. When a support agent gives a subtly wrong answer, nobody finds out. Reconciliation works because finance built its record of truth first and then pointed its automation at it.
The running example
To keep this concrete I’ll use one agent throughout: a refunds agent for an online shop. It reads a customer’s support ticket, looks up the order, decides whether a refund is due and for how much, and proposes it. Refunds above a limit, or for unusual reasons, need a human.
It’s deliberately a different agent from the one in my invoice-matching eval post and my permission matrix post. Those go deep on a single case-level eval and on what an agent may do. This post is about the system of controls around the agent once it’s running.
The map
| Reconciliation concept | What it does | Agent-harness equivalent | What teams usually get wrong |
|---|---|---|---|
| Ground truth: the ledger as source of record | A record maintained independently of the process being checked, which the process’s output is compared against | Checking the agent’s claimed actions against the systems it acted on (orders, payments, tickets), not against its own trace | Grading the agent on its transcript. The trace says “refund issued”; nobody checks the payment provider |
| Matching rules, strict to loose | Ordered rules: exact key and amount first, then fuzzier rules, each tagged with the rule that fired | Deterministic graders first (schema, amounts, IDs), then heuristics, and an LLM judge last, with the grader recorded per case | Starting with an LLM judge for everything, so an exact-match failure and a tone complaint look the same |
| Tolerance thresholds | A written, owned limit for acceptable differences (a rounding penny, an FX band) | Explicit pass criteria per field: amount must be exact, wording may vary, latency under a stated budget | Tolerances nobody wrote down, living in someone’s head or in a prompt |
| Exception queues | Every unmatched item becomes a record with an owner, a reason and an age | Escalations as typed records with a reason code, evidence and an owner, not a Slack message | Escalation is a free-text message the agent writes. It can’t be counted, aged or routed |
| Maker-checker / four-eyes | The person who prepares an entry can’t be the one who approves it | The agent proposes; a different principal approves, enforced by the server | “Human in the loop” means a confirm button the same session can press, or a second agent that approves the first |
| Segregation of duties | No one role can both initiate and conceal a transaction | The agent can’t write to its own logs, edit its own eval set, or change its own thresholds | The agent’s service account has admin on the tables that record what it did |
| Audit trail | An append-only record of who did what, when, on what evidence | Immutable per-action records: actor, inputs, policy decision, result, versions of model and prompt | Traces kept for 14 days in an observability tool, with no link to the business record they changed |
| Period close and sign-off | At a set point, every item is matched or explained, and a named person signs | A regular close over agent activity: all exceptions resolved or carried forward with an owner, then signed | Nobody owns “was yesterday’s agent activity correct?” Dashboards get looked at, not signed |
| Root-cause grouping of breaks | Group breaks by shared cause so one fix clears fifty lines | Cluster failures by failing check and first bad step, then fix the cluster | Reading traces one at a time and fixing whichever was most recent |
| Controls testing | Periodically prove each control works, for example by trying to push through a transaction that should be blocked | Tests that attack the harness itself: self-approval, over-limit refunds, a missing audit record | Testing the agent’s quality and never testing that the guards fire |
The rest of the post takes four of these in more depth.
Ground truth lives outside the agent
The most important property of a ledger is that the process being checked didn’t write it. A bank statement comes from the bank. A sub-ledger reconciles to a general ledger that other processes also post to. Two records that come from different places can disagree, and that disagreement is the signal.
Agent evals often lose this. The grader reads the trace, sees a tool call to issue_refund with a successful response, and marks the case correct. But the trace is the agent’s own account of events. Tool responses can be wrong, retries can double-post, and a partial failure can return success. In reconciliation terms, you’re reconciling a sub-ledger to itself.
For the refunds agent the independent record is the payment provider and the order system. So the check compares what the agent claims with what those systems say actually happened:
type Claim = { runId: string; orderId: string; refundMinor: number; currency: string };
type ProviderRefund = { orderId: string; amountMinor: number; currency: string; refundId: string };
type Break =
| { kind: "missing_at_provider"; claim: Claim }
| { kind: "unclaimed_at_provider"; refund: ProviderRefund }
| { kind: "amount_mismatch"; claim: Claim; refund: ProviderRefund }
| { kind: "duplicate_at_provider"; orderId: string; refunds: ProviderRefund[] };
export function reconcileRefunds(claims: Claim[], provider: ProviderRefund[]): Break[] {
const breaks: Break[] = [];
const byOrder = Map.groupBy(provider, r => r.orderId);
for (const c of claims) {
const found = byOrder.get(c.orderId) ?? [];
if (found.length === 0) breaks.push({ kind: "missing_at_provider", claim: c });
else if (found.length > 1) breaks.push({ kind: "duplicate_at_provider", orderId: c.orderId, refunds: found });
else if (found[0].amountMinor !== c.refundMinor || found[0].currency !== c.currency)
breaks.push({ kind: "amount_mismatch", claim: c, refund: found[0] });
}
const claimed = new Set(claims.map(c => c.orderId));
for (const r of provider) {
if (!claimed.has(r.orderId)) breaks.push({ kind: "unclaimed_at_provider", refund: r });
}
return breaks;
}
This runs over production activity, not a test set, and that’s the point. Offline evals tell you whether the agent is good on cases you picked. A reconciliation tells you whether what it did today matches reality, including on cases nobody thought to write down. The two directions matter equally. An agent that claims refunds that never happened is broken. Refunds at the provider that no run claims are worse, because it means something is acting without a record.
Not every agent has a ledger, but more have one than teams assume. If an agent changes state somewhere, that somewhere is your ground truth.
Exceptions are records, not messages
In a reconciliation system, a break doesn’t vanish into a log line. It becomes an exception with a type, an amount, an owner and a date it was raised, and it sits in a queue until someone clears it. Ageing is tracked. An exception that’s been open for thirty days is a different conversation from one raised this morning.
Agent escalations usually come out as prose. The agent writes “I wasn’t sure about this one, flagging for review” into a ticket or a channel. You can’t count prose by reason, route it, see how old it is, or tell whether it was resolved. So I’d make escalation a tool with a typed payload and a closed set of reasons:
type EscalationReason =
| "over_agent_limit"
| "policy_ambiguous"
| "order_not_found"
| "amount_disagrees_with_order"
| "suspected_injection"
| "tool_error";
type Exception = {
id: string;
runId: string;
raisedAt: string; // ISO timestamp
reason: EscalationReason;
subject: { orderId: string; ticketId: string };
proposed?: { refundMinor: number; currency: string };
evidence: { kind: "tool_result" | "document"; ref: string }[];
owner: string | null; // assigned by routing, never by the agent
status: "open" | "resolved" | "carried_forward";
resolution?: { by: string; at: string; outcome: "approved" | "rejected" | "amended"; note: string };
versions: { model: string; prompt: string; policy: string };
};
A few choices carry the weight here. The reason is an enum, so the queue can be reported on and routed by rule. owner is set by routing, not chosen by the agent. versions pins which model, prompt and policy produced the exception, so a spike in policy_ambiguous after a prompt change points straight at the change. And resolution captures what the human decided. That’s a labelled case for your eval set, which the invoice post covers under the escalation flywheel.
The failure I see most is escalation as the agent’s only uncertainty signal. The server should also raise exceptions the agent didn’t ask for: amount over limit, order in the wrong state, a reconciliation break. In reconciliation, the process being checked doesn’t get to decide what counts as an exception.
Maker-checker the agent can’t game
Four-eyes approval is simple to state: whoever prepares an entry can’t approve it. Everyone knows it. Agent systems still break it in two common ways.
The first is the confirm button in the same session. The agent proposes a refund, the UI shows “Approve?”, and the same user, or worse the same agent in a later turn, presses it. The second is approval by a second model. A reviewer agent that reads the first agent’s output and says “looks good” adds an opinion but no independent authority, and it can share the first model’s blind spots. Even a perfect reviewer model isn’t an accountable approver.
The control has to be enforced where the money moves. An approval is a token bound to a specific proposal, and the server refuses it if the approver is the proposer:
type Principal = { id: string; kind: "agent" | "human" };
type Proposal = {
id: string;
proposedBy: Principal;
orderId: string;
refundMinor: number;
currency: string;
digest: string; // hash of the fields above
};
type Approval = { proposalId: string; digest: string; approvedBy: Principal; at: string };
export function checkApproval(p: Proposal, a: Approval, limits: Map<string, number>): string | null {
if (a.proposalId !== p.id) return "approval is for a different proposal";
if (a.digest !== p.digest) return "proposal changed after approval";
if (a.approvedBy.kind !== "human") return "approver must be a human principal";
if (a.approvedBy.id === p.proposedBy.id) return "proposer cannot approve";
const limit = limits.get(a.approvedBy.id) ?? 0;
if (p.refundMinor > limit) return "above approver's limit";
return null; // ok
}
The digest matters as much as the identity check. Without it, an agent can get approval for a £40 refund and then amend the proposal to £400. The approver’s own limit matters too: four-eyes doesn’t mean any second person, it means a second person with authority for this amount.
The proposer check also needs to reach past the obvious. If the agent acts “on behalf of” a support rep, the rep is also a maker for that proposal and shouldn’t be the checker. That’s segregation of duties, the neighbouring row in the table: roles, not just individuals, get kept apart. The same principle says the agent’s service account shouldn’t be able to write to the audit table, change its own limits or edit the eval set it’s graded on.
Group breaks by root cause
When a large reconciliation throws hundreds of breaks, nobody works them in order. You sort by the rule that failed, the counterparty, the date and the source file, and you usually find the bulk come from a handful of causes: a feed arrived late, a reference format changed, a currency table wasn’t updated. Fix the cause and most of the breaks clear. Trintech’s release describes grouping exceptions “under a shared root cause” for the same reason.
Agent teams tend to debug by reading traces. Someone opens the latest failure, finds a plausible cause, tweaks the prompt, and moves on to the next. That’s working breaks one line at a time, and it overfits to whichever trace you happened to read.
The fix is to give each failure a signature you can group by. With typed exceptions and reconciliation breaks you already have most of one. The piece that’s missing is where in the run it went wrong:
type Failure = {
runId: string;
check: string; // e.g. "amount_mismatch", "proposer_cannot_approve"
firstBadStep: { tool: string; errorClass: string } | null;
versions: { model: string; prompt: string; policy: string };
};
export function groupFailures(failures: Failure[]) {
const key = (f: Failure) =>
[f.check, f.firstBadStep?.tool ?? "-", f.firstBadStep?.errorClass ?? "-", f.versions.prompt].join("|");
return [...Map.groupBy(failures, key).entries()]
.map(([signature, items]) => ({ signature, count: items.length, sample: items.slice(0, 3).map(i => i.runId) }))
.sort((a, b) => b.count - a.count);
}
firstBadStep is the hard field. Getting it means your trace records tool results in a structured way, so a script can find the first tool call whose result disagrees with ground truth, or the first call that returned an error the agent then ignored. Once you have it, the report can start with something like “most failures share amount_mismatch | get_order | stale_cache”, and you read three sample traces from that group rather than all of them.
Failures that don’t group go in a singletons bucket someone reads by hand, and the size of that bucket over time is a useful health measure.
Close, sign-off and testing the controls
The two rows I haven’t covered in depth are the ones agent teams skip most.
Period close. Finance closes the books on a schedule. At close, every item is either matched or explained, open exceptions are carried forward with a named owner, and someone signs. For an agent the equivalent is a daily or weekly close over its activity: run the reconciliation against ground truth, clear or carry forward every exception, and record who signed off and what was outstanding. It sounds like bureaucracy until someone asks whether the agent was right last Thursday.
Controls testing. Auditors don’t just check that a control exists. They test that it works, by trying to push through something it should block. Agent evals mostly test quality: did the agent make the right call? Controls tests check the harness: does the server reject a self-approval, an amended proposal, a refund above the agent’s limit, a write to the audit table from the agent’s credentials? These are ordinary integration tests and they should run in CI. They’re also the tests most likely to be missing when a report finds harnesses with “neither a test nor an eval”, because they aren’t about the model at all.
What I’d do on Monday
- Name the source of record for every state change the agent can make. If there isn’t one, that’s the first finding.
- Write a nightly reconciliation of agent claims against that record, in both directions. Start with counts and amounts before anything clever.
- Turn escalation into a tool with an enum of reasons and a typed payload. Have the server raise its own exceptions as well.
- Put an ageing report on the exception queue and give every open item an owner.
- Move approval into the server: a token bound to a proposal digest, rejected if approver and proposer match or the amount is over the approver’s limit.
- Check the agent’s credentials against its own records. It shouldn’t be able to write its audit trail, its limits or its eval set.
- Add a failure signature (check, first bad step, versions) and look at grouped failures each week instead of the latest traces.
- Write controls tests that try to break each guard, and run them in CI.
- Hold a weekly close: reconciliation run, exceptions cleared or carried forward, one named person signing it off.
Limitations
This is design analysis and illustrative code. I haven’t built this refunds agent, the TypeScript hasn’t been run (Map.groupBy also needs a recent runtime), and I have no measured results to offer about any of it.
The mapping is an analogy, and analogies have edges. Reconciliation works on structured records where “match” can be defined. Many agents produce text, plans or code where the ground truth is judgement, and for those the ledger idea only covers the state changes, not the quality of the prose. The post also doesn’t cover statistical sampling, which auditors use when full reconciliation is too expensive, or how to set tolerances and approval limits for a real business.
On the news hook: I’ve quoted Trintech’s press release, which is vendor marketing, and I haven’t used its product. The Marmelab figure applies to the open-source harnesses that report examined, not to agent systems in general, and I haven’t checked its method beyond reading the report. None of this is audit, compliance or legal advice.