MULTI-AGENT EVALUATION · FIELD GUIDE

Trace
is not
Proof.

A collaboration trace answers “what happened”. Acceptance must also answer whether the original goal was met, whether the result entered the real system, and whether failure can be reproduced and repaired.

CORE JUDGMENT

Agent count, message count, tool calls, and runtime can show that a system was busy. None of them alone proves that the work is complete.

01 · WHAT RESEARCH ACTUALLY SUPPORTS

Papers support evaluation methods, not the correctness of a product.

The claims below stay within what the primary papers or formal conference pages support. They are not capability claims about Flowness.

R-01AgentBench · ICLR 2024

Evaluate agents in interactive environments

AgentBench evaluates reasoning and decision making across eight interactive environments, treating long-horizon reasoning, decision making, and instruction following as observable failure surfaces. The object of evaluation is behavior in an environment, not a model report.

R-02SWE-bench · ICLR 2024

Make candidate changes face an executable environment

SWE-bench combines real GitHub issues, repositories, and test environments into tasks. A candidate must change observable test state in an execution environment, showing why generating a patch and solving the problem are separate judgments.

R-03SWE-bench Verified

The verifier needs verification too

SWE-bench Verified improves evaluation reliability through human checks of task statements, test correctness, and solvability. Automated tests matter, while whether a test actually corresponds to the target remains an independent question.

R-04MT-Bench and Chatbot Arena · NeurIPS 2023

LLM judges help, but are not an unbiased truth source

Research finds that strong model judges can approximate human preferences while also showing position, verbosity, self-preference, and limited-reasoning biases. A single judge pass therefore needs reproducible evidence and, where needed, human judgment.

R-05Multi-Agent Debate · ICML 2024

Multi-agent debate is not quality assurance

One study observes gains from debate on selected reasoning and factuality tasks. Another systematic study finds that existing methods do not reliably beat self-consistency or ensembles and can be sensitive to tuning. More agents cannot replace an acceptance design.

02 · THE VERIFICATION LADDER

Split “pass” into seven different kinds of evidence.

The higher the risk and the less reversible the change, the more layers are needed. Small, reversible tasks do not need the full ceremony.

  1. 01
    CLAIM AND TRACE

    Self-report and event truth

    Who claimed what, when, and against which version? This is a signal awaiting verification. It shows only that an execution occurred.

    Shows an execution occurred
  2. 02
    ARTIFACT IDENTITY

    Artifact and exact identity

    Does the candidate exist and bind to its goal, source, version, commit, or hash? It must not be silently replaced during review.

    Supports identity tracking
  3. 03
    MECHANICAL CHECK

    Reproducible postconditions

    Do structure checks, tests, probes, queries, or policy gates show the target state changed? Does the verifier cover the real requirement?

    Supports a mechanical check
  4. 04
    INDEPENDENT REVIEW

    Independent falsification and fresh recheck

    Reviewers separated from the producer actively look for counterexamples. After a Finding is repaired, evidence is collected again on the successor candidate.

    Supports a review decision
  5. 05
    COMMITMENT

    Integration readback

    Did the artifact enter the repository, service, database, documentation system, or decision flow that actually consumes it? Readback evidence must come from that consumer.

    Supports an integration decision
  6. 06
    ACTIVATION

    Representative activation

    Did a real or representative task use the new path? Could an old entry point, cache, or silent bypass make the feature look active?

    Supports an activation decision
  7. 07
    ACCEPTANCE

    Scoped result verdict

    Does target-native readback show the authoritative postcondition, and has the responsible party made the scoped acceptance decision? Are remaining risks and uncovered scope explicit?

    Supports a scoped acceptance decision

03 · FALSE SIGNALS

The dangerous state is not failure. It is looking like success.

“The agent says it is done”

That is an executor report, not independent evidence.

“The tool returned success”

It shows protocol success, not that the business postcondition holds.

“Local tests pass”

It covers the checks that ran in the local environment, not integration, deployment, or real use.

“The judge gave a high score”

An LLM judge can scale evaluation, while bias and task fit still need control.

“The code was merged”

Built, integrated, activated, and accepted are different states.

“No problem was found”

The check may have been too narrow, the evidence missing, or the reviewers may share one blind spot.

04 · FAILURE ATLAS

Move from “which step failed” to “which work pathology appeared”.

Flowness public Failure Atlas organizes structural failure around seven questions. It is a diagnostic lens for replay and counterexample search, not statistical proof.

01

Formation

Why did an event fail to form executable work?

02

Continuity

Why did work quietly stop without a conclusion?

03

Integrity

Did identity, version, evidence, or interpretation drift?

04

Adaptation

After facts changed, was stale context or planning reused?

05

Commitment

Did the artifact enter the system that actually consumes it?

06

Closure

Did “done” correspond to a real effect and acceptance?

07

Learning

Will the same structural failure be harder to repeat next time?

05 · MINIMUM ACCEPTANCE PACKET

A deliverable acceptance should leave at least six things behind.

  1. 01

    Original goal and anti-goal

    What counts as success, and which outcomes remain unacceptable even when they look useful.

  2. 02

    Exact candidate identity

    Version, commit, hash, environment, or another identity that cannot be confused with a neighbor.

  3. 03

    Authoritative postcondition

    An observation from the target system, database, test environment, or real consumer.

  4. 04

    Review relationship and independence

    Who executed, who reviewed, which context was shared, and which common blind spots may remain.

  5. 05

    Finding and repair lineage

    Which candidate held the issue, which successor carried the repair, and why the evidence was not silently dropped.

  6. 06

    Fresh successor-candidate verdict and remaining risk

    Recheck the repaired successor candidate, then state that its verdict is not target-domain Effect or responsible-party acceptance, along with the environments, users, or conditions that remain uncovered.

EVIDENCE BOUNDARY

Research explains risk. Runtime evidence proves implementation.

AgentBench, SWE-bench, LLM-as-a-Judge, and multi-agent debate research support evaluation in task environments, observable results, useful but bounded judges, and multi-perspective evaluation worth studying.

Those papers do not validate Flowness and do not prove that any harness automatically guarantees correctness. Flowness public Alpha evidence currently supports a narrow path: execution, independent review, retained Finding, targeted rework, and a fresh successor-candidate verdict. It does not establish target-domain Effect or responsible-party acceptance.

Full Flow runtime coverage, cross-domain effects, and long-term reliability remain claims that need further verification.

06 · QUESTIONS AT THE BOUNDARY

Keep the acceptance claim as precise as the evidence.

The answers below are deliberately narrower than a product promise.

01

Is a successful trace evidence that the work is accepted?

No. A trace shows that an execution occurred. A fresh successor-candidate verdict still needs precise artifact identity, checks tied to the target, independent review, commitment, activation, and target-native readback. It is not target-domain Effect or responsible-party acceptance by itself.

02

Can an LLM judge or benchmark replace independent verification?

No. AgentBench, SWE-bench, and LLM judge research provide bounded evaluation methods. A benchmark score, judge score, debate result, CI result, or same-session review does not by itself prove a real-world acceptance condition or a fresh successor-candidate verdict.

03

What does the public Flowness evidence support today?

The public Alpha evidence supports a narrow assurance-kernel path with execution, independent review, retained Findings, targeted rework, and a fresh successor-candidate verdict. That verdict is not target-domain Effect or responsible-party acceptance. Full Flow runtime coverage, cross-domain effects, and long-term reliability remain unverified.

07 · PRIMARY SOURCES

Return every key judgment to a primary source.

Evidence checked: 2026-08-18. Research scope and Flowness implementation evidence remain separate on this page.

  1. AgentBench: Evaluating LLMs as AgentsOPEN ↗
  2. SWE-bench: Can Language Models Resolve Real-world GitHub Issues?OPEN ↗
  3. SWE-bench VerifiedOPEN ↗
  4. Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaOPEN ↗
  5. Improving factuality and reasoning through multiagent debateOPEN ↗
  6. When is Society of Mind Superior to Self-Consistency?OPEN ↗
  7. Flowness public evidence and Failure AtlasOPEN ↗

CONTINUE

Leave a record of collaboration, and make “done” survive a fresh check.

Read the two real experimentsUnderstand the harness mechanismRead the trust boundaryCompare the four layersBrowse English researchRun the public evidence ↗