Skip to content

Case file · worked example

FINALIZEDSWE-bench VerifiedScenario B

Agentic coder is 20% better than the previous SOTA

Model A is 20% better than Model B on agentic coding tasks.

Misleading

The headline is not a fair reading of the evidence.

01 · What was claimed

The published sentence

Model A
Forge-Agent-Large · 1.3.0
Model B / baseline
Forge-Agent-Base · 1.0.2
Benchmark
SWE-bench Verified · verified-500-2025-03-15
Evaluation date
2026-05-18
Metric
resolved %
Reported result
66.0 vs 55.0 (headline: +20%)
Claimant
0xBADA55E0000000000000000000000000000BADA5
Published
2026-05-19T12:00:00.000Z

Methodology

Internal harness. Model A was allowed 5 retries with majority vote. Model B was allowed 1 attempt. System prompt for A included extra planning scratchpad. Failed runs for A were retried; failed runs for B were recorded as unresolved. Headline reports a 20% relative improvement without disclosing the retry gap.

Retries

5 vs 1

Temperature

0.3 vs 0

Tools

bash,python,patch,web vs bash,patch

Prompts

unmatched

System prompts

unmatched

Sample

500

02 · What evidence was provided

Off-chain artifacts, on-chain hashes

Large files are not stored in the Intelligent Contract. The contract holds canonical hashes and references. The index below is the case file.

  • Marketing report. Admits 5 retries for A vs 1 for B in a footnote.

    sha256:c8034d7bbdd62efbbb1caefcefb6cec4f8faa099aa7214cd8e356f31d54aa46f

  • configsha256:forge-unequal-retries

    Run config dump: retries_a=5, retries_b=1, tools unequal.

    sha256:5b3b2ce9e87ca0f6440cf74affde182a21de2e71a680877da5d89c1f0de92f63

  • note

    Headline does not mention retries, tools, or temperature.

    sha256:c3b64feef2f14864bbf5940fa61ca1fca57032b44a20983a77a7dfea48a77e81

03 · What was challenged

Public review

  1. Challenge 01 · Unequal retry count

    Five retries versus one

    The 20% headline is computed against a baseline that was given a single attempt while the challenger model received five retries and extra tools. That is not a matched comparison.

    0xCH411E000000000000000000000000000CH411

04 · What was decided

Verdict

Misleadinghigh confidenceWorked example (same rubric)

The published claim about Forge-Agent-Large vs Forge-Agent-Base on SWE-bench Verified is not a fair reading of the evidence (Raw results or harness artifacts were not provided.).

Key findings

  • Compared the published claim against attached evidence and condition fields.

Material issues

  • Raw results or harness artifacts were not provided.
  • Retry policy is unequal (A=5, B=1).
  • Temperature differs between models (0.3 vs 0).
  • Tool access differs between models (bash,python,patch,web vs bash,patch).
  • Prompts were not matched across models.
  • System prompts were not matched across models.
  • The published headline does not disclose the unequal evaluation conditions.

Limitations

  • This preview uses the BenchProof forensic rubric. The canonical verdict is the GenLayer consensus result.

05 · Verification

On-chain identifiers

Claim id
seed-misleading
On-chain id
not submitted
Network
studionet · chain 61999
Tx state
finalized
Transaction

Off-chain rows never override a finalized on-chain verdict. Seeded case files illustrate the rubric; live GenLayer consensus is marked explicitly. Open contract in Studio · Live evaluation tx