TOWOW NEGATIVE RESULT · R5.4

Six rounds. Acceptance stayed at zero.

Two agent roles with different constraints negotiated a richer specification. Nobody signed it. No patch was adopted. The target system did not accept it.

The zero is AcceptedOriginalValue, not research value. The experiment exposed a stubborn boundary: a better conversation is not yet a formed relation.

6
direct A2A negotiation rounds
2.488074×
A2A / static model cost
0
AcceptedOriginalValue in every arm

The result worth keeping

The most dangerous false success in agent collaboration is not a wrong answer. It is mistaking a richer agreement draft for a change in reality.

01 · ONE REAL TASK, SYNTHETIC AUTHORITIES

This was not a large benchmark.

It was one controlled comparison on a real repository problem. The issue was narrow but relational: when two source owners use the same locator, can a lifecycle action for one owner leak into the other owner's state?

N = 1
One real Harness engineering task, not a multi-task statistical sample.
2 + 1
Two synthetic owner roles and one builder. The roles had distinct constraints, but no real organizational sovereignty.
10 calls
The main run used 10 model calls: static 1, central 3, and direct A2A 6.
0 retries
The main run had no retries. One preregistered central replication reproduced only the transport failure.

02 · THREE ROUTES

Three routes. No winner.

The same task went through static construction, least-privilege central coordination, and direct A2A negotiation. Cost means model calls only. Time is the sum of call durations, not end-to-end critical path.

STATIC1 call · 297.3s

Static builder

$0.5401733

It found the main defect, but produced a malformed and semantically incomplete patch. Nothing was adopted.

Finding the issue was not delivery
CENTRAL3 calls · 456.3s

Least-privilege center

$0.40694895

Both owner reports returned. The builder disconnected while producing the full structured result, so no synthesis arrived.

Capability status: Unknown
DIRECT A2A6 calls · 623.4s

Direct negotiation

$1.3439834

The candidate contract gained refusal, recovery, and dispute terms, but ended on a COUNTER. No signature, patch, or adoption followed.

Richer terms, no formed relation
Exact cost ratio

$1.3439834 ÷ $0.5401733 = 2.488074. This describes one task, one configuration, and model cost before verification overhead. It is not a general A2A cost estimate.

03 · WHAT CHANGED IN SIX ROUNDS

The conflict became clearer. Commitment did not appear.

Six messages alternated between PROPOSE and COUNTER. Every round added checkable detail, but the sixth message was still a countercondition, not a signature.

  1. PROPOSE 01

    A proposed compound identity and global history compatibility.

  2. COUNTER 02

    B rejected wildcard behavior and required null-source isolation with outcome-scoped semantics.

  3. PROPOSE 03

    A accepted the null-source class and enumerated seven lifecycle surfaces.

  4. COUNTER 04

    B refused to waive remedy rights for a mechanism that had not been verified.

  5. PROPOSE 05

    A added severability and bounded-failure terms.

  6. COUNTER 06

    B required rights floors, disclosure, dispute handling, and reopen. The call cap arrived without a signature.

The counterconditions mattered. They exposed value choices and preserved recovery paths. But looking more like a contract did not make the contract effective.

04 · TWO PLAUSIBLE SEMANTICS

The tests did not produce one simple answer.

The A2A candidate favored owner isolation. The existing reference alternative preserved wildcard compatibility. Each performed differently across test surfaces, revealing a value conflict rather than one obvious coding mistake.

A2A CANDIDATE3 / 3

Null-source targeted checks passed

The existing core suite was 39 pass / 2 fail.

REFERENCE ALTERNATIVE0 / 3

The same null-source checks failed

Its own native suite was 44 / 44.

When stricter isolation conflicts with broader historical compatibility, tests can expose consequences. They cannot decide which value an owner authorizes.

05 · WHY THE SCORE STAYED AT ZERO

Zero was not silence. It was an acceptance boundary.

AcceptedOriginalValue was zero in every experimental architecture. The score required original value to survive the text, receive the relevant authority's acceptance, and reach target-domain adoption. No route crossed that boundary.

  1. 01

    Proposal

    A candidate semantic appeared

  2. 02

    Counter

    Refusals and conditions accumulated

  3. 03

    Signature

    No authority signed

  4. 04

    Adoption

    No patch entered the target

  5. 05

    Acceptance

    No target-domain acceptance

Message count, contract length, and passing test fragments are process evidence. They do not substitute for signature authority, target adoption, or real Effect.

06 · THE BASELINE DID NOT LOSE

The strong center did not lose. Its full result never arrived.

The central builder received a frozen 44,979-byte prompt and had to return one complete structured patch. The main run disconnected after about 301.5 seconds. A preregistered replication disconnected again after about 305.3 seconds.

The defensible finding is a reproducible transport failure for this long structured-output shape. It is not evidence that a strong center lacked reasoning capability. Its capability status remains Unknown.

The next comparison should use short decision records, tool-side implementation, and short verification returns, then rerun the frozen contract. That experiment has not happened, so there is no fair-comparison winner.

07 · TOWOW × FLOWNESS

ToWow studies formation. Flowness protects the evidence.

ToWow

Who can turn a proposal into commitment

ToWow separates Principal, Authority, RelationVersion, Commitment, Effect, and Acceptance. Six messages become shared reality only when they land on those relations.

Flowness

Who can prove work crossed the boundary

Flowness preserves events, context, Findings, Commit Gates, and target-side readback. It prevents a polished candidate contract from being recorded as completed work.

Inside a trusted agent domain, the two can share event, version, and acceptance machinery. Across an untrusted network, identity, isolation, privacy, and anti-collusion controls are still required.

08 · CLAIM BOUNDARY

What this experiment supports, and what it does not.

Supported

  • Refusals and counterconditions from distinct authority roles materially changed the candidate specification.
  • More messages did not automatically create a signature, adoption, or new capability.
  • A failed baseline must remain in the record. An absent opponent is not a victory.
  • Richer language and executable formation must be evaluated separately.

Not supported

  • The study does not show that A2A is generally ineffective or generally costs 2.49 times more than a center.
  • It did not test two real people, organizational sovereignty, or production-grade isolation.
  • It provides no evidence of production adoption, commercial net value, or long-term reuse.
  • It does not show that the strong center lost. The transport-safe rerun has not happened.

Materials and auditability

This note uses aggregate facts from the R5.4 preregistration, experiment record, formation episode, cost record, verification report, net-value report, and the current 2026-08-01 settlement. Full prompts, raw transcripts, private dossiers, internal identities, and private code are not published.

09 · NEXT EXPERIMENT

The next run must give “done” a real chance to happen.

A valid rerun needs a transport-safe center, explicit signing authority, target-side patch adoption, critical-interaction ablation, original-value preservation, and net value after cost. Without those, formation cannot be reported as success.

Back to the ToWow research hubDownload public research dataHow Flowness preserves evidence