HOW IT WORKS

2. Define Success

Once AG Arch knows what we are trying to achieve and what is true now, the next step is to define what a successful result actually looks like.

This sounds obvious, but it is one of the most important parts of governed autonomous work.

An autonomous agent can complete a sequence of actions, follow a plan correctly, and still fail to produce the outcome that matters.

That is why AG Arch does not define success as:

“the agent finished the task”

or:

“the tool returned success.”

Instead, success has to describe the state we want to be true after the work is done.

A task is not the same as an outcome

Traditional automation often focuses on actions:

  • update a file;
  • call an API;
  • restart a service;
  • send a message;
  • modify a record;
  • run a deployment.

Those actions may be necessary, but none of them proves that the intended result was achieved.

For example, a deployment command can complete successfully while the application still fails to start.

A database update can return `200 OK` while the wrong record was changed.

A configuration file can be edited correctly while the running system is still using an older version.

The action tells us what was attempted.

The outcome tells us what actually became true.

AG Arch separates those two.

Success must be defined before execution

If success is only decided after the work is finished, verification becomes subjective.

The system can easily drift into:

“The agent did something reasonable, so we will call it successful.”

That is not enough for governed execution.

Before the main action begins, AG Arch defines the success criteria that will later be used to judge the result.

These criteria should answer questions such as:

  • What state should exist when the work is complete?
  • What must no longer be true?
  • Which conditions must remain unchanged?
  • What evidence would demonstrate that the intended outcome happened?
  • Are there multiple conditions that all need to be satisfied?
  • Is partial success acceptable, or must the entire outcome be achieved?

This gives the rest of the lifecycle a stable target.

Success criteria should describe reality, not effort

A useful success criterion normally describes something observable about the result.

For example, weak criteria might be:

The agent updated the configuration.

or:

The agent ran the deployment successfully.

These describe activity.

Stronger criteria would be:

The approved configuration is active in the target environment.

and:

The service is running normally after the change.

Those statements can later be checked against reality.

The difference is important because AG Arch is not trying to prove that work was performed.

It is trying to prove that the intended outcome became real.

Success can have more than one condition

Real work often has several requirements.

Changing a system safely may mean:

  • the requested value is updated;
  • the system remains available;
  • unrelated settings remain unchanged;
  • the new state persists;
  • no prohibited action occurred.

If only one of those conditions is checked, the result may look successful while still being wrong.

AG Arch therefore allows the definition of success to include multiple conditions.

Some conditions describe what must happen.

Others describe what must not happen.

Both are important.

For example:

The service must use the new configuration.

and:

The service must not lose connectivity during the change.

The first is a desired outcome.

The second is a constraint on what still counts as success.

Success should be specific enough to verify, but not tied to one method

There is another balance to maintain.

If the success criteria are too vague, they cannot be verified.

If they are too prescriptive, they can accidentally turn into a hidden implementation plan.

For example:

Improve the system.

is too vague.

But:

Edit line 42, restart process X, then call endpoint Y

is not really a success definition. It is a procedure.

A better success definition would be:

The target service is running with the approved configuration, and its health check confirms normal operation.

This tells us what outcome is required while leaving room for the execution system to choose an appropriate method.

That is important for autonomy.

The agent should be free to choose among valid execution paths, as long as it remains inside the authority and constraints that will be defined next.

Why this matters for verification later

The success criteria created here become the reference point for the later Observe & Prove stage.

That means the system does not invent a definition of success after seeing the result.

It compares reality with a target that was already defined.

This creates a clear separation:

before execution: define what success means.

after execution: observe what actually happened.

then: compare the two.

Without that separation, the system can unconsciously lower the standard after a weak result.

AG Arch is designed to avoid that.

What if success cannot be measured directly?

Not every outcome has a single perfect measurement.

Sometimes success must be established through several pieces of evidence.

Sometimes a direct measurement is impossible, but there are reliable indicators.

For example, proving that a system change worked may involve:

  • reading the current configuration;
  • checking service health;
  • confirming expected behavior;
  • checking that no related error appeared.

The important point is not that every success criterion has to be represented by one number.

The important point is that the criteria are clear enough that later evidence can support or reject them.

Success criteria can also expose ambiguity early

Defining success often reveals that the original request was not precise enough.

For example:

“Make the system faster.”

What does that mean?

Lower average response time?

Fewer timeouts?

Better performance under a specific workload?

If nobody can say what successful completion would look like, the task is not ready for governed execution.

This is useful.

It is better to discover ambiguity before an agent starts making changes than after the system has already been modified.

AG Arch therefore treats the inability to define success as a signal that more clarification or context may be required.

Example — changing a system configuration

Continuing the same example from Intent & Reality:

The intent is to correct a configuration used by a running service.

The current state has already been established.

Now AG Arch defines what success means.

For example:

Success criteria:

  • the approved configuration value is active;
  • the running service is using that value;
  • the service remains healthy after the change;
  • unrelated configuration remains unchanged.

Notice what is not listed as success:

“The agent edited the configuration file.”

That may be one step in the process, but it is not the outcome.

The agent could edit the correct file while the service continues using an old cached configuration.

In that case, the action happened, but the success criteria were not met.

What this stage gives to the next one

At the end of Define Success, AG Arch has a clear answer to:

What result are we trying to prove?

Now the system can decide how that result may be pursued.

The next stage therefore asks:

  • What is the plan?
  • What is the agent allowed to do?
  • What constraints apply?
  • What proof will be required?

That is the purpose of 3. Plan & Govern.