01 · WHAT RESEARCH ACTUALLY SUPPORTS
Papers support evaluation methods, not the correctness of a product.
The claims below stay within what the primary papers or formal conference pages support. They are not capability claims about Flowness.
R-01AgentBench · ICLR 2024
Evaluate agents in interactive environments
AgentBench evaluates reasoning and decision making across eight interactive environments, treating long-horizon reasoning, decision making, and instruction following as observable failure surfaces. The object of evaluation is behavior in an environment, not a model report.
R-02SWE-bench · ICLR 2024
Make candidate changes face an executable environment
SWE-bench combines real GitHub issues, repositories, and test environments into tasks. A candidate must change observable test state in an execution environment, showing why generating a patch and solving the problem are separate judgments.
R-03SWE-bench Verified
The verifier needs verification too
SWE-bench Verified improves evaluation reliability through human checks of task statements, test correctness, and solvability. Automated tests matter, while whether a test actually corresponds to the target remains an independent question.
R-04MT-Bench and Chatbot Arena · NeurIPS 2023
LLM judges help, but are not an unbiased truth source
Research finds that strong model judges can approximate human preferences while also showing position, verbosity, self-preference, and limited-reasoning biases. A single judge pass therefore needs reproducible evidence and, where needed, human judgment.
R-05Multi-Agent Debate · ICML 2024
Multi-agent debate is not quality assurance
One study observes gains from debate on selected reasoning and factuality tasks. Another systematic study finds that existing methods do not reliably beat self-consistency or ensembles and can be sensitive to tuning. More agents cannot replace an acceptance design.