Skip to content

Case file · worked example

FINALIZEDSWE-bench VerifiedScenario A

Matched-harness SWE-bench Verified: Opus 4.1 vs GPT-4.1

Claude Opus 4.1 scores 72.4% resolved on SWE-bench Verified versus 54.6% for GPT-4.1 under an identical open-source harness, identical 2025-03-15 repo snapshot, temperature 0, one sample, no retries, and the same tool budget.

Supported

Evidence fairly supports the published claim.

01 · What was claimed

The published sentence

Model A
Claude Opus 4.1 · 2026-02-19
Model B / baseline
GPT-4.1 · gpt-4.1-2025-04-14
Benchmark
SWE-bench Verified · verified-500-2025-03-15
Evaluation date
2026-03-02
Metric
resolved %
Reported result
72.4 vs 54.6 (Δ +17.8 pp)
Claimant
0xA11CE5EED00000000000000000000000000A11CE
Published
2026-03-03T09:30:00.000Z

Methodology

Both models ran the public SWE-bench Verified 500-instance split (snapshot 2025-03-15) inside the same OpenHands harness commit 9f2c1a. Temperature 0, max tokens 8192, retries 1, identical system prompt, identical tool list (bash, python, patch). Failed generations counted as unresolved. No instance was dropped after seeing results. Raw trajectories and patches are hashed below.

Retries

1 vs 1

Temperature

0 vs 0

Tools

bash,python,patch vs bash,python,patch

Prompts

matched

System prompts

matched

Sample

500

02 · What evidence was provided

Off-chain artifacts, on-chain hashes

Large files are not stored in the Intelligent Contract. The contract holds canonical hashes and references. The index below is the case file.

  • OpenHands harness commit 9f2c1a used for both models.

    sha256:aa9c4c8ce6bf6eb0b9e5a89da85bf6a9b5f77ba29cacc9a8040da33a409c2295

  • raw_resultsipfs://bafybeiaseedresultsopus41

    Per-instance resolved/unresolved CSV, 500 rows.

    sha256:e8d692007a65850e13030e18d25a31feeeb53fd792ddfa659f9ee0513a2d5876

  • configsha256:config-opus-vs-gpt41

    Identical decoding config: t=0, n=1, tools matched.

    sha256:6fda506593bd8673ecfc606aac58e163dcfd69d6273aaec7504182ddcdaa9ca2

03 · What was challenged

Public review

No challenges filed. Absence of a challenge is not support.

04 · What was decided

Verdict

Supportedhigh confidenceWorked example (same rubric)

The evidence fairly supports the published comparison of Claude Opus 4.1 vs GPT-4.1 on SWE-bench Verified under matched conditions.

Key findings

  • Benchmark identity and version are specified.
  • Model versions are specified on both sides.
  • Evaluation conditions (prompts, system prompts, retries, tools, inference) match.
  • Methodology and artifacts are present.

Material issues

None recorded.

Limitations

  • This preview uses the BenchProof forensic rubric. The canonical verdict is the GenLayer consensus result.

05 · Verification

On-chain identifiers

Claim id
seed-supported
On-chain id
not submitted
Network
studionet · chain 61999
Tx state
finalized
Transaction

Off-chain rows never override a finalized on-chain verdict. Seeded case files illustrate the rubric; live GenLayer consensus is marked explicitly. Open contract in Studio · Live evaluation tx