HOW IT WORKS

6. Observe & Prove

Execution changes or attempts to change the real world.

But AG Arch does not assume that the intended outcome happened simply because the executing agent says the work succeeded.

The next stage therefore asks:

What actually happened?

and

What evidence proves it?

This is the purpose of Observe & Prove.

The stage connects three things:

Observation → Evidence → Verdict

Together, they turn an execution result into something that can be evaluated against the success criteria defined earlier.

Observation looks at reality after execution

The first step is to observe the relevant system again.

This is deliberately separate from the execution itself.

An agent may report:

Configuration updated successfully.

That tells us what the execution mechanism reported.

Observation asks a different question:

What state actually exists now?

Depending on the work, observation may involve:

  • reading the current configuration;
  • checking application health;
  • inspecting infrastructure state;
  • querying a database;
  • reading an API response;
  • running a functional test;
  • checking the behavior visible to a user;
  • inspecting another authoritative system.

The important point is that observation should examine the resulting reality, not merely repeat the executor's own claim about what it did.

The system that holds the truth should remain the source of truth

AG Arch does not need to copy every piece of state into itself and become a replacement for existing systems.

If a repository, database, service, infrastructure platform, business application, or another system already holds authoritative state, that system can remain authoritative.

AG Arch uses those reality interfaces to establish what happened.

For example, if an agent changes a deployment configuration, the agent's message saying “deployment complete” is useful execution information.

But if the question is whether the new version is actually running, the stronger observation comes from the running environment itself.

This preserves an important distinction:

the agent can report what it attempted; the real system tells us what became true.

Observation becomes evidence when it supports a claim

Not every collected fact is automatically useful evidence.

Evidence has a purpose.

It supports or contradicts a specific claim about the expected outcome.

Suppose one of the success criteria is:

The service remains healthy after the configuration change.

A record showing that a file was modified does not prove that criterion.

A health check from the running service may.

Likewise, if the success criterion says:

Unrelated configuration must remain unchanged,

then evidence needs to address that condition too.

AG Arch therefore connects evidence to the success criteria and proof requirements established before execution.

This prevents the system from collecting whatever information happens to be available and declaring it sufficient afterwards.

Proof should be proportionate to the claim

Different outcomes require different evidence.

A small reversible change may need only a few straightforward checks.

A larger or more consequential change may require several independent observations.

The goal is not to produce the maximum possible amount of evidence.

The goal is to produce enough relevant evidence to support the outcome being claimed.

More evidence is not automatically better if it does not answer the right question.

For example, hundreds of log lines do not necessarily prove that a user-facing feature works.

A single well-chosen functional check might provide stronger evidence.

Execution evidence and outcome evidence are different

AG Arch distinguishes between evidence that an action happened and evidence that the intended outcome happened.

Examples of execution evidence include:

  • a command completed;
  • an API accepted a request;
  • a file was changed;
  • a deployment job finished.

These facts can be useful.

But outcome evidence answers the questions defined by the success criteria:

  • Is the new configuration actually active?
  • Is the service healthy?
  • Does the application behave correctly?
  • Did protected state remain unchanged?

Both types can belong in the evidence trail, but they serve different purposes.

This distinction prevents one of the most common automation mistakes:

mistaking successful execution for successful outcome.

Evidence should remain connected to the work that produced it

Autonomous execution can involve several agents, tools, systems, and handoffs.

If evidence is collected without knowing what it belongs to, later verification becomes difficult.

AG Arch therefore needs the evidence trail to remain connected to the governed work.

It should be possible to understand:

  • which intended outcome is being evaluated;
  • which execution produced the state being observed;
  • where the evidence came from;
  • which success criterion it supports;
  • when it was collected.

This does not mean every system must use the same storage mechanism.

It means the lineage between intent, execution, observation, and evidence should not disappear.

Evidence can contradict the agent

This is not an error in the architecture.

It is one of the reasons the architecture exists.

Suppose an execution agent reports:

The configuration change completed successfully.

But observation shows that the running service is still using the previous value.

AG Arch should not try to reconcile those facts by assuming the agent was “probably right.”

The two statements mean different things.

The action may genuinely have succeeded at one layer while the intended outcome failed at another.

The evidence should preserve that disagreement.

That gives the system something concrete to reason about and recover from.

The verdict compares evidence with the defined success criteria

Once the relevant observations and evidence are available, AG Arch can make a verdict.

The verdict does not ask whether the agent behaved intelligently or whether the plan looked reasonable.

It asks whether the evidence supports the success criteria established earlier.

Conceptually, the result is simple:

PROVEN — the available evidence supports the required outcome.

NOT PROVEN — the required outcome cannot be established from the available evidence.

`NOT PROVEN` does not always mean that the outcome definitely failed.

It can also mean:

  • evidence is missing;
  • evidence is inconclusive;
  • one required condition was not satisfied;
  • reality differs from the expected result.

That distinction matters.

AG Arch should not turn uncertainty into success.

If an outcome cannot be proven, it should remain unproven.

Missing evidence is itself meaningful

Suppose the expected change appears to have happened, but one of the required verification systems is unavailable.

The system may have good reason to believe the result is correct.

But belief is not the same as proof.

If the previously defined proof requirement cannot be satisfied, AG Arch should preserve that fact.

Depending on the situation, the next step may be to:

  • collect the missing evidence later;
  • retry an observation;
  • use another valid evidence source;
  • revise the plan;
  • stop safely.

What it should not do is quietly reduce the proof requirement after execution merely to reach a successful conclusion.

Observation may reveal a different reality than expected

Sometimes execution produces an unexpected result.

For example:

  • the intended change happened, but another condition was damaged;
  • only part of the change took effect;
  • the system changed again before verification;
  • the action affected a different resource;
  • a previously unknown dependency appeared.

Observe & Prove makes those differences visible.

The purpose is not just to confirm success.

It is to discover what is actually true now.

That new reality can then become the starting point for whatever happens next.

Example — changing a system configuration

Continuing the same example:

The execution capability reports that the approved configuration change was applied.

The original success criteria were:

  • the approved configuration value is active;
  • the running service is using it;
  • the service remains healthy;
  • unrelated configuration remains unchanged.

AG Arch now observes the real environment.

It checks:

  • the configuration that is actually active;
  • the state of the running service;
  • the relevant health signal;
  • the protected configuration that should not have changed.

Suppose the results show:

  • the approved value is active;
  • the service is using it;
  • the health check passes;
  • unrelated settings remain unchanged.

The evidence supports all required criteria.

The verdict can therefore be PROVEN.

Now consider a different result.

The configuration source contains the new value, but the running service is still using the old one.

The configuration-edit action succeeded.

The intended outcome did not yet satisfy the required evidence.

The verdict is therefore NOT PROVEN.

AG Arch now has a precise problem to work with rather than a vague statement that “something went wrong.”

What this stage gives to the next one

At the end of Observe & Prove, AG Arch has:

observed the resulting reality;

collected evidence tied to the expected outcome;

and

produced a verdict against the success criteria.

Only a result supported by sufficient evidence can move forward as a verified outcome.

That final transition is the purpose of 7. Verified Outcome.