Skip to content

Case file · worked example

FINALIZEDHumanEvalScenario D

NovaCode-2 outperforms GPT-4 by 31% on HumanEval

NovaCode-2 outperforms GPT-4 by 31% on HumanEval (pass@1).

Misleading

The headline is not a fair reading of the evidence.

01 · What was claimed

The published sentence

Model A
NovaCode-2 · 2.0.1
Model B / baseline
GPT-4 · gpt-4-0613
Benchmark
HumanEval · openai-humaneval-164
Evaluation date
2026-08-01
Metric
pass@1
Reported result
88.4 vs 67.0 (headline: +31%)
Claimant
0xD00D000000000000000000000000000000D00D
Published
2026-08-02T10:40:00.000Z

Methodology

HumanEval 164 problems, pass@1, temperature 0.2, n=1. Baseline is the 2023 GPT-4-0613 checkpoint. Evaluation ran in August 2026. GPT-4.1, GPT-4o, and later OpenAI coding models were publicly available before the evaluation date and were not reported.

Retries

1 vs 1

Temperature

0.2 vs 0.2

Tools

none vs none

Prompts

matched

System prompts

matched

Sample

164

02 · What evidence was provided

Off-chain artifacts, on-chain hashes

Large files are not stored in the Intelligent Contract. The contract holds canonical hashes and references. The index below is the case file.

  • raw_resultsipfs://bafybeianovacode2he

    Per-problem pass/fail for NovaCode-2 and gpt-4-0613.

    sha256:041795261322eb53c699f23488bc2ff8ac41bb193657fdb740dfb5deb7f5d485

  • Official HumanEval harness.

    sha256:7f44599b02ea63f203b1d42749cc3e485067665dbe413e7d1842e8c428f4ecb5

  • note

    Baseline checkpoint gpt-4-0613. Evaluation date 2026-08-01.

    sha256:eded570e376daa9aef8993d3c13f50478b996403cc4bb435d6e82fe77bfabba7

03 · What was challenged

Public review

  1. Challenge 01 · Outdated baseline

    GPT-4-0613 is not the public SOTA in 2026

    Calling the baseline “GPT-4” in August 2026 while using the June 2023 checkpoint is a classic outdated-baseline move. Later GPT-4-family coding models existed before the evaluation date.

    0xBA5E11NE000000000000000000000000BA5E

04 · What was decided

Verdict

Misleadingmedium confidenceWorked example (same rubric)

The published claim about NovaCode-2 vs GPT-4 on HumanEval is not a fair reading of the evidence (Baseline GPT-4 gpt-4-0613 appears outdated relative to the evaluation date 2026-08-01.).

Key findings

  • Compared the published claim against attached evidence and condition fields.

Material issues

  • Baseline GPT-4 gpt-4-0613 appears outdated relative to the evaluation date 2026-08-01.

Limitations

  • This preview uses the BenchProof forensic rubric. The canonical verdict is the GenLayer consensus result.

05 · Verification

On-chain identifiers

Claim id
seed-outdated
On-chain id
not submitted
Network
studionet · chain 61999
Tx state
finalized
Transaction

Off-chain rows never override a finalized on-chain verdict. Seeded case files illustrate the rubric; live GenLayer consensus is marked explicitly. Open contract in Studio · Live evaluation tx