Static builder
$0.5401733It found the main defect, but produced a malformed and semantically incomplete patch. Nothing was adopted.
Finding the issue was not deliveryTOWOW NEGATIVE RESULT · R5.4
Two agent roles with different constraints negotiated a richer specification. Nobody signed it. No patch was adopted. The target system did not accept it.
The zero is AcceptedOriginalValue, not research value. The experiment exposed a stubborn boundary: a better conversation is not yet a formed relation.
The result worth keeping
The most dangerous false success in agent collaboration is not a wrong answer. It is mistaking a richer agreement draft for a change in reality.
01 · ONE REAL TASK, SYNTHETIC AUTHORITIES
It was one controlled comparison on a real repository problem. The issue was narrow but relational: when two source owners use the same locator, can a lifecycle action for one owner leak into the other owner's state?
02 · THREE ROUTES
The same task went through static construction, least-privilege central coordination, and direct A2A negotiation. Cost means model calls only. Time is the sum of call durations, not end-to-end critical path.
It found the main defect, but produced a malformed and semantically incomplete patch. Nothing was adopted.
Finding the issue was not deliveryBoth owner reports returned. The builder disconnected while producing the full structured result, so no synthesis arrived.
Capability status: UnknownThe candidate contract gained refusal, recovery, and dispute terms, but ended on a COUNTER. No signature, patch, or adoption followed.
Richer terms, no formed relation$1.3439834 ÷ $0.5401733 = 2.488074. This describes one task, one configuration, and model cost before verification overhead. It is not a general A2A cost estimate.
03 · WHAT CHANGED IN SIX ROUNDS
Six messages alternated between PROPOSE and COUNTER. Every round added checkable detail, but the sixth message was still a countercondition, not a signature.
A proposed compound identity and global history compatibility.
B rejected wildcard behavior and required null-source isolation with outcome-scoped semantics.
A accepted the null-source class and enumerated seven lifecycle surfaces.
B refused to waive remedy rights for a mechanism that had not been verified.
A added severability and bounded-failure terms.
B required rights floors, disclosure, dispute handling, and reopen. The call cap arrived without a signature.
The counterconditions mattered. They exposed value choices and preserved recovery paths. But looking more like a contract did not make the contract effective.
04 · TWO PLAUSIBLE SEMANTICS
The A2A candidate favored owner isolation. The existing reference alternative preserved wildcard compatibility. Each performed differently across test surfaces, revealing a value conflict rather than one obvious coding mistake.
The existing core suite was 39 pass / 2 fail.
Its own native suite was 44 / 44.
When stricter isolation conflicts with broader historical compatibility, tests can expose consequences. They cannot decide which value an owner authorizes.
05 · WHY THE SCORE STAYED AT ZERO
AcceptedOriginalValue was zero in every experimental architecture. The score required original value to survive the text, receive the relevant authority's acceptance, and reach target-domain adoption. No route crossed that boundary.
A candidate semantic appeared
Refusals and conditions accumulated
No authority signed
No patch entered the target
No target-domain acceptance
Message count, contract length, and passing test fragments are process evidence. They do not substitute for signature authority, target adoption, or real Effect.
06 · THE BASELINE DID NOT LOSE
The central builder received a frozen 44,979-byte prompt and had to return one complete structured patch. The main run disconnected after about 301.5 seconds. A preregistered replication disconnected again after about 305.3 seconds.
The defensible finding is a reproducible transport failure for this long structured-output shape. It is not evidence that a strong center lacked reasoning capability. Its capability status remains Unknown.
The next comparison should use short decision records, tool-side implementation, and short verification returns, then rerun the frozen contract. That experiment has not happened, so there is no fair-comparison winner.
07 · TOWOW × FLOWNESS
ToWow separates Principal, Authority, RelationVersion, Commitment, Effect, and Acceptance. Six messages become shared reality only when they land on those relations.
Flowness preserves events, context, Findings, Commit Gates, and target-side readback. It prevents a polished candidate contract from being recorded as completed work.
Inside a trusted agent domain, the two can share event, version, and acceptance machinery. Across an untrusted network, identity, isolation, privacy, and anti-collusion controls are still required.
08 · CLAIM BOUNDARY
This note uses aggregate facts from the R5.4 preregistration, experiment record, formation episode, cost record, verification report, net-value report, and the current 2026-08-01 settlement. Full prompts, raw transcripts, private dossiers, internal identities, and private code are not published.
09 · NEXT EXPERIMENT
A valid rerun needs a transport-safe center, explicit signing authority, target-side patch adoption, critical-interaction ablation, original-value preservation, and net value after cost. Without those, formation cannot be reported as success.