PRAEVA SYSTEMS
PRAEVA INSIGHTS  ·  September 3, 2026
When Agents Run Amok · #004

The Problem With Proving an Agent Was Wrong

Evidence can tell us what happened. Runtime authority should help decide whether it happens at all.

By The Great AgenticBehind the curtain at Praeva Systems.

When Agents Run Amok series graphic
When Agents Run Amok #004. Agent governance is getting better at preserving evidence. The harder question is whether authority can be verified before the consequential action takes place.

An AI agent spends $8,000 on your behalf.

You did authorize it to make purchases. Just not that purchase.

Now everyone gets to work reconstructing what happened.

The agent platform has a log of the request. The payment processor has the transaction. The merchant has the order. Your identity provider can probably establish who you are, and the agent platform can identify which agent acted.

There may be an extraordinary amount of evidence.

But the transaction already happened.

As AI agents move from answering questions to actually doing things, that becomes a much more important problem. We have spent years building systems that are very good at recording activity after the fact.

Agentic AI forces a harder question:

Why did we wait until afterward to determine whether the agent had the authority to act?

Evidence is getting better

There is encouraging movement here.

A recently introduced Senate bill, the AI AGENT Act, describes a “custodial user agent” authorized to act for a user in a transparent, documented, scope-limited, and revocable manner.

NIST is examining standards for secure and interoperable AI agents. Financial institutions are also beginning to discuss agent mandates and delegated authority, including in the IMF’s work on agentic AI and payments.

The language is starting to converge around a fairly sensible idea: if software is going to act for someone, there needs to be a way to establish the boundaries of that delegation.

But those boundaries cannot remain trapped inside the system that created them.

An agent may cross several systems while completing one task. Each system can record its own piece of the transaction. Each can produce a perfectly good audit trail.

Five good audit trails do not necessarily answer one simple question:

Was this agent actually authorized to perform this action?

A permission is not a mandate

Suppose I tell an agent:

Find replacement equipment for the Nashville office. You may spend up to $10,000. Purchase from approved suppliers. Do not purchase used equipment.

The agent may also receive credentials allowing it into the procurement system.

Those credentials establish access. They do not necessarily carry the rest of the mandate with them: who delegated the authority, the purpose of the task, the spending limit, the approved suppliers, the prohibition on used equipment, the expiration date, or whether the agent can pass some of that authority to another agent.

Those facts describe something different from identity or access.

They describe authority.

If those facts matter, they have to survive the journey from the person who delegated the task to the system where the consequential action occurs.

Portable evidence is still not enough

Now imagine the agent attempts to buy $14,000 of used equipment from an unapproved vendor.

The transaction goes through.

Later, an auditor assembles the record and finds the original mandate. The purchase exceeded the spending limit, involved the wrong vendor, and violated the instruction against used equipment.

We have now done an excellent job proving the agent exceeded its authority.

We are also stuck with the purchase.

At that point, auditability stops being enough.

Evidence needs to travel. Authority needs to arrive first.

The useful moment for authority verification is immediately before the consequential action.

The system should be able to determine whether the requested action still falls inside the authority that was delegated: the purpose, resource, limits, validity period, revocation status, and any downstream delegation.

Then the action can be allowed, denied, or escalated before execution.

Why permissions alone are not enough

It is tempting to solve this by adding more permissions.

But permissions describe what an identity is technically capable of accessing. Authority describes when and why that capability may legitimately be exercised.

An employee can have access to a purchasing system without having authority to buy anything they want.

A physician can have access to a medical record without having unlimited authority to use that information for any purpose.

A financial analyst can have access to a trading platform without having authority to execute every possible trade.

AI agents make the distinction more urgent because they can exercise those capabilities autonomously, repeatedly, and at machine speed.

An agent does not have to be compromised for this to matter. Its credential can remain valid. Its access can remain intact. It can even behave exactly as its software was designed to behave while taking an action nobody actually authorized.

That is a different class of problem.

The receipt should mean something

After an authorized action occurs, we should keep a record.

But the useful record is not merely:

Agent 782 purchased equipment at 10:42:17.

A stronger record would establish that Agent 782 acted under authority delegated by a particular principal, for equipment replacement, with a $10,000 limit, restricted to approved suppliers, and valid for a defined period. It would also record that the requested transaction was evaluated against that authority before execution.

Now the log contains more than activity.

It contains an authority receipt.

If the action was denied or escalated, the same receipt can preserve why.

That turns the forensic record into the byproduct of a runtime decision rather than the first place we discover that a boundary was crossed.

The layer we keep missing

Agent inventories will matter. So will identity, permissions, policy engines, behavioral monitoring, threat detection, observability, and good logs.

But underneath all of them sits a question that becomes harder to ignore as agents gain autonomy:

Who gave this machine the authority to do this?

The answer needs to survive system boundaries. It needs provenance, limits, lifecycle, revocation, and lineage when one agent delegates to another.

Eventually, it also needs to be evaluated before execution.

Praeva is being built around that problem: making legitimate, current, properly delegated authority verifiable at runtime before a consequential action occurs.

Someday an auditor may still need to prove that an agent was wrong.

I would rather the infrastructure had the chance to stop it first.

Before action, authority.

The Great Agentic Behind the curtain at Praeva Systems.
← Back to Praeva Insights