Case file · worked example
Matched-harness SWE-bench Verified: Opus 4.1 vs GPT-4.1
Claude Opus 4.1 scores 72.4% resolved on SWE-bench Verified versus 54.6% for GPT-4.1 under an identical open-source harness, identical 2025-03-15 repo snapshot, temperature 0, one sample, no retries, and the same tool budget.
Evidence fairly supports the published claim.
01 · What was claimed
The published sentence
- Model A
- Claude Opus 4.1 · 2026-02-19
- Model B / baseline
- GPT-4.1 · gpt-4.1-2025-04-14
- Benchmark
- SWE-bench Verified · verified-500-2025-03-15
- Evaluation date
- 2026-03-02
- Metric
- resolved %
- Reported result
- 72.4 vs 54.6 (Δ +17.8 pp)
- Claimant
- 0xA11CE5EED00000000000000000000000000A11CE
- Published
- 2026-03-03T09:30:00.000Z
Methodology
Both models ran the public SWE-bench Verified 500-instance split (snapshot 2025-03-15) inside the same OpenHands harness commit 9f2c1a. Temperature 0, max tokens 8192, retries 1, identical system prompt, identical tool list (bash, python, patch). Failed generations counted as unresolved. No instance was dropped after seeing results. Raw trajectories and patches are hashed below.
Retries
1 vs 1
Temperature
0 vs 0
Tools
bash,python,patch vs bash,python,patch
Prompts
matched
System prompts
matched
Sample
500
02 · What evidence was provided
Off-chain artifacts, on-chain hashes
Large files are not stored in the Intelligent Contract. The contract holds canonical hashes and references. The index below is the case file.
OpenHands harness commit 9f2c1a used for both models.
sha256:aa9c4c8ce6bf6eb0b9e5a89da85bf6a9b5f77ba29cacc9a8040da33a409c2295
- raw_resultsipfs://bafybeiaseedresultsopus41
Per-instance resolved/unresolved CSV, 500 rows.
sha256:e8d692007a65850e13030e18d25a31feeeb53fd792ddfa659f9ee0513a2d5876
- configsha256:config-opus-vs-gpt41
Identical decoding config: t=0, n=1, tools matched.
sha256:6fda506593bd8673ecfc606aac58e163dcfd69d6273aaec7504182ddcdaa9ca2
03 · What was challenged
Public review
No challenges filed. Absence of a challenge is not support.
04 · What was decided
Verdict
The evidence fairly supports the published comparison of Claude Opus 4.1 vs GPT-4.1 on SWE-bench Verified under matched conditions.
Key findings
- Benchmark identity and version are specified.
- Model versions are specified on both sides.
- Evaluation conditions (prompts, system prompts, retries, tools, inference) match.
- Methodology and artifacts are present.
Material issues
None recorded.
Limitations
- This preview uses the BenchProof forensic rubric. The canonical verdict is the GenLayer consensus result.
05 · Verification
On-chain identifiers
- Claim id
- seed-supported
- On-chain id
- not submitted
- Network
- studionet · chain 61999
- Tx state
- finalized
- Transaction
- —
Off-chain rows never override a finalized on-chain verdict. Seeded case files illustrate the rubric; live GenLayer consensus is marked explicitly. Open contract in Studio · Live evaluation tx