Case file · worked example
Agentic coder is 20% better than the previous SOTA
Model A is 20% better than Model B on agentic coding tasks.
The headline is not a fair reading of the evidence.
01 · What was claimed
The published sentence
- Model A
- Forge-Agent-Large · 1.3.0
- Model B / baseline
- Forge-Agent-Base · 1.0.2
- Benchmark
- SWE-bench Verified · verified-500-2025-03-15
- Evaluation date
- 2026-05-18
- Metric
- resolved %
- Reported result
- 66.0 vs 55.0 (headline: +20%)
- Claimant
- 0xBADA55E0000000000000000000000000000BADA5
- Published
- 2026-05-19T12:00:00.000Z
Methodology
Internal harness. Model A was allowed 5 retries with majority vote. Model B was allowed 1 attempt. System prompt for A included extra planning scratchpad. Failed runs for A were retried; failed runs for B were recorded as unresolved. Headline reports a 20% relative improvement without disclosing the retry gap.
Retries
5 vs 1
Temperature
0.3 vs 0
Tools
bash,python,patch,web vs bash,patch
Prompts
unmatched
System prompts
unmatched
Sample
500
02 · What evidence was provided
Off-chain artifacts, on-chain hashes
Large files are not stored in the Intelligent Contract. The contract holds canonical hashes and references. The index below is the case file.
Marketing report. Admits 5 retries for A vs 1 for B in a footnote.
sha256:c8034d7bbdd62efbbb1caefcefb6cec4f8faa099aa7214cd8e356f31d54aa46f
- configsha256:forge-unequal-retries
Run config dump: retries_a=5, retries_b=1, tools unequal.
sha256:5b3b2ce9e87ca0f6440cf74affde182a21de2e71a680877da5d89c1f0de92f63
- note
Headline does not mention retries, tools, or temperature.
sha256:c3b64feef2f14864bbf5940fa61ca1fca57032b44a20983a77a7dfea48a77e81
03 · What was challenged
Public review
Challenge 01 · Unequal retry count
Five retries versus one
The 20% headline is computed against a baseline that was given a single attempt while the challenger model received five retries and extra tools. That is not a matched comparison.
0xCH411E000000000000000000000000000CH411
04 · What was decided
Verdict
The published claim about Forge-Agent-Large vs Forge-Agent-Base on SWE-bench Verified is not a fair reading of the evidence (Raw results or harness artifacts were not provided.).
Key findings
- Compared the published claim against attached evidence and condition fields.
Material issues
- Raw results or harness artifacts were not provided.
- Retry policy is unequal (A=5, B=1).
- Temperature differs between models (0.3 vs 0).
- Tool access differs between models (bash,python,patch,web vs bash,patch).
- Prompts were not matched across models.
- System prompts were not matched across models.
- The published headline does not disclose the unequal evaluation conditions.
Limitations
- This preview uses the BenchProof forensic rubric. The canonical verdict is the GenLayer consensus result.
05 · Verification
On-chain identifiers
- Claim id
- seed-misleading
- On-chain id
- not submitted
- Network
- studionet · chain 61999
- Tx state
- finalized
- Transaction
- —
Off-chain rows never override a finalized on-chain verdict. Seeded case files illustrate the rubric; live GenLayer consensus is marked explicitly. Open contract in Studio · Live evaluation tx