Case file · worked example
NovaCode-2 outperforms GPT-4 by 31% on HumanEval
NovaCode-2 outperforms GPT-4 by 31% on HumanEval (pass@1).
The headline is not a fair reading of the evidence.
01 · What was claimed
The published sentence
- Model A
- NovaCode-2 · 2.0.1
- Model B / baseline
- GPT-4 · gpt-4-0613
- Benchmark
- HumanEval · openai-humaneval-164
- Evaluation date
- 2026-08-01
- Metric
- pass@1
- Reported result
- 88.4 vs 67.0 (headline: +31%)
- Claimant
- 0xD00D000000000000000000000000000000D00D
- Published
- 2026-08-02T10:40:00.000Z
Methodology
HumanEval 164 problems, pass@1, temperature 0.2, n=1. Baseline is the 2023 GPT-4-0613 checkpoint. Evaluation ran in August 2026. GPT-4.1, GPT-4o, and later OpenAI coding models were publicly available before the evaluation date and were not reported.
Retries
1 vs 1
Temperature
0.2 vs 0.2
Tools
none vs none
Prompts
matched
System prompts
matched
Sample
164
02 · What evidence was provided
Off-chain artifacts, on-chain hashes
Large files are not stored in the Intelligent Contract. The contract holds canonical hashes and references. The index below is the case file.
- raw_resultsipfs://bafybeianovacode2he
Per-problem pass/fail for NovaCode-2 and gpt-4-0613.
sha256:041795261322eb53c699f23488bc2ff8ac41bb193657fdb740dfb5deb7f5d485
Official HumanEval harness.
sha256:7f44599b02ea63f203b1d42749cc3e485067665dbe413e7d1842e8c428f4ecb5
- note
Baseline checkpoint gpt-4-0613. Evaluation date 2026-08-01.
sha256:eded570e376daa9aef8993d3c13f50478b996403cc4bb435d6e82fe77bfabba7
03 · What was challenged
Public review
Challenge 01 · Outdated baseline
GPT-4-0613 is not the public SOTA in 2026
Calling the baseline “GPT-4” in August 2026 while using the June 2023 checkpoint is a classic outdated-baseline move. Later GPT-4-family coding models existed before the evaluation date.
0xBA5E11NE000000000000000000000000BA5E
04 · What was decided
Verdict
The published claim about NovaCode-2 vs GPT-4 on HumanEval is not a fair reading of the evidence (Baseline GPT-4 gpt-4-0613 appears outdated relative to the evaluation date 2026-08-01.).
Key findings
- Compared the published claim against attached evidence and condition fields.
Material issues
- Baseline GPT-4 gpt-4-0613 appears outdated relative to the evaluation date 2026-08-01.
Limitations
- This preview uses the BenchProof forensic rubric. The canonical verdict is the GenLayer consensus result.
05 · Verification
On-chain identifiers
- Claim id
- seed-outdated
- On-chain id
- not submitted
- Network
- studionet · chain 61999
- Tx state
- finalized
- Transaction
- —
Off-chain rows never override a finalized on-chain verdict. Seeded case files illustrate the rubric; live GenLayer consensus is marked explicitly. Open contract in Studio · Live evaluation tx