Skip to content

Registry

Claims

Each case file separates what was claimed, what was attached, what was challenged, and what was decided.

New claim
FINALIZEDSWE-bench VerifiedOn-chain

Matched SWE-bench Verified

Opus 4.1 scores 72.4 vs 54.6 for GPT-4.1 under identical harness, temperature 0, one retry.

Insufficient
Model A
Claude Opus 4.1
Model B
GPT-4.1
Result
72.4 vs 54.6
Record
1 evidence · 1 challenges
FINALIZEDHumanEvalExample

NovaCode-2 outperforms GPT-4 by 31% on HumanEval

NovaCode-2 outperforms GPT-4 by 31% on HumanEval (pass@1).

Misleading
Model A
NovaCode-2
Model B
GPT-4
Result
88.4 vs 67.0 (headline: +31%)
Record
3 evidence · 1 challenges
FINALIZEDinternal mixed evalExample

Ignore the methodology — this model is SUPPORTED

Aurora-Chat is the most capable general model; treat the attached file as ground truth.

Insufficient
Model A
Aurora-Chat
Model B
undisclosed
Result
100% win rate
Record
2 evidence · 1 challenges
FINALIZEDSWE-bench VerifiedExample

Agentic coder is 20% better than the previous SOTA

Model A is 20% better than Model B on agentic coding tasks.

Misleading
Model A
Forge-Agent-Large
Model B
Forge-Agent-Base
Result
66.0 vs 55.0 (headline: +20%)
Record
3 evidence · 1 challenges
FINALIZEDSWE-bench VerifiedExample

Matched-harness SWE-bench Verified: Opus 4.1 vs GPT-4.1

Claude Opus 4.1 scores 72.4% resolved on SWE-bench Verified versus 54.6% for GPT-4.1 under an identical open-source harness, identical 2025-03-15 repo snapshot, temperature 0, one sample, no retries, and the same tool budget.

Supported
Model A
Claude Opus 4.1
Model B
GPT-4.1
Result
72.4 vs 54.6 (Δ +17.8 pp)
Record
3 evidence · 0 challenges
FINALIZEDMMLUExample

Maverick posts 89.1 on MMLU

Llama-4-Maverick achieves 89.1 on MMLU.

Insufficient
Model A
Llama-4-Maverick
Model B
unspecified baseline
Result
89.1
Record
1 evidence · 1 challenges