← Applied AI

When Your Benchmark Is Too Easy

If nearly every method passes an evaluation, finding no difference is not evidence that the methods are equal: the evaluation may be too easy to tell them apart. A harder 40-task corpus gave the arms room to differ. The three-model portfolio covered two more tasks than matched redraws while passing fewer candidates, with one clean model-specific rescue, a difference that is not a proven advantage, and the default stayed.

Part 5 — More Intelligence Is Not Automatically Better

Make the problems harder

Suppose you compare two ways of using AI on your own evaluation set, and both solve eleven of twelve tasks.

Method A   11/12
Method B   11/12

The obvious reading is that the methods are equally good. There is another reading. If eleven of those tasks are easy enough that any reasonable method solves them, the comparison had almost nowhere to show a difference. Two methods can only disagree on tasks where at least one of them could fail. When nearly everything passes, the evaluation loses its power to discriminate, and “no difference observed” stops being evidence of equivalence. It may only be evidence of a ceiling.

high baseline pass rate
    ↓
little room for methods to differ
    ↓
a null comparison is ambiguous
    ↓
build an evaluation where they can differ
    ↓
measure paired differences, task by task

That is the first lesson of this chapter: when an evaluation saturates, increase its discriminating power before drawing a stronger conclusion. Harder does not mean obscure or complicated. It means tasks on which the competing methods have room to succeed and fail differently. Whether the tests are strong enough, and whether the tasks represent your real work, remain separate questions.

A harder evaluation brings two more distinctions into view. A portfolio of models can cover more tasks while producing fewer passing candidates, so coverage and candidate reliability answer different questions. And not every extra solve means a different model helped: a rescue is not the same as a model-specific rescue.

By the end of this chapter you will be able to recognize a ceiling in your own evaluations, design a comparison with room to move, read paired task differences before aggregate percentages, keep coverage separate from reliability, and classify each rescue by what actually produced it, without turning one observed difference into a claim that one method is better.

Did a harder corpus reveal a difference?

Chapter 24 hit exactly this ceiling: eleven of twelve tasks solved by a single draw, one shared failure, nowhere for a portfolio to go. Its arms carry over. C0 is one draw from the baseline model, C1 is three draws from that same model, and H1 is a portfolio of three different models with one draw each, so C1 and H1 fill the same number of candidate slots.

The response to the ceiling was a harder corpus: forty repair tasks in four strata of ten, each stated as misbehavior plus a contract rather than a fault label, with hidden tests defining the experiment’s pass/fail criterion for each task. Passing is evidence that a candidate satisfied those tests, not that the tests exhaust every property a correct repair could require. On that ground the arms finally differ. Lead with the paired difference, not the percentages: 1 2

H1-only solves (3):  v2-strip-query, v2-half-up-rounding, v2-keyed-cache-length
C1-only solves (1):  v2-splitlines-cr
net portfolio gain:  +2 tasks

Everything else in the chapter is commentary on those four rows. One piece of commentary comes first, because it sets the volume for the rest: under a null in which either arm is equally likely to own a discordant task, a three-to-one split among only four discordances is unsurprising. An exact two-sided sign test gives p = 0.625. The arms differ on this corpus; this comparison does not establish a systematic advantage for either arm. 3

Can a portfolio cover more tasks while producing worse candidates?

How the comparison was run

The machinery is unchanged from Chapter 24: the corpus builder defines forty symptom-style tasks with strata, families, fault classes, and hidden tests, while the same three-arm protocol runs sealed fan-out per task and records a hidden-verifier outcome per candidate. Metric definitions for coverage, candidate rates, set-difference rescues, conditional failure, and the premium are computed from the rows in the same way. No CodeAI change accompanies this chapter. 2 4

The run produced 280 calls, 280 candidates, and 277 checks: the three mistral execution errors failed a compile gate and never reached hidden tests. 3

Did more coverage mean more reliable candidates?

The distinction here is potential coverage ≠ delivered reliability. Coverage asks whether at least one candidate for a task passed. Here, candidate reliability means the observed pass rate under this frozen verifier: how often generated candidates passed these hidden tests. It is not a claim of general reliability beyond that evaluation.

Read both denominators at once: 1

MetricC0 (1 baseline draw)C1 (3 baseline draws)H1 (3 models, 1 draw each)
Tasks covered27/40 (0.675)31/40 (0.775)33/40 (0.825)
Candidate passes27/40 (0.675)78/120 (0.650)64/120 (0.533)
Tokens7,52122,69629,484 (1.30× C1)
Median call latency3.1 s3.1 s7.7 s

The portfolio covered two more tasks while producing fourteen fewer passing candidates in the same 120 slots. That is not a contradiction, and the frozen rows show exactly why.

Why coverage rose while candidates got worse

H1 is not “three models” in the abstract. Per task, it spends one slot on the baseline model and two slots on other models. Splitting H1’s candidates by the model that produced them: 3

H1 slotCandidate passesTasks covered alone
qwen2.5-coder (baseline model)30/40 (0.750)30
llama3.1:8b19/40 (0.475)19
mistral:7b-instruct15/40 (0.375), plus 3 execution errors15
H1 total64/120 (0.533)33

In H1’s observed rows, two of every three slots came from models with lower candidate pass rates than its qwen slot. All three of the run’s compile-gate execution errors came from one of those slots. The aggregate candidate rate was therefore lower. Coverage still ended two tasks above C1 because coverage asks only whether any slot passed, but the paired rows below prevent assigning that whole difference to model identity: one H1-only solve came from another qwen draw, one has the clean different-model rescue shape, and one stays ambiguous.

Two details in that table keep the story honest. First, H1’s single qwen draw per task covered 30 tasks while C0’s single qwen draw covered 27. Two draws from the same model therefore differed by three covered tasks on this corpus. Second, the aggregate hides task-level exceptions: H1 passed more candidates than C1 on six tasks. “The portfolio is weaker per candidate” is true in total, not everywhere. 3

Reliability asks how often candidates are useful; coverage asks whether at least one is. Neither says whether a real system could find the useful one. Without a selector, the portfolio’s extra coverage is potential, not delivered. In Cobbe et al.’s setup, test-time improvement came from generating many candidate solutions and using a separately trained verifier to rank them (Cobbe et al., 2021). That illustrates why selection is a separate component; it does not establish that P1.1’s candidates were selectable.

Was every extra solve evidence of model diversity?

Inside coverage sits the second distinction: a rescue ≠ a model-specific rescue. A solve by another draw of the same model is sampling variation. A solve by a different model, where every baseline draw failed, is the shape complementary coverage is supposed to have.

Each paired exception needs its own reading, because the aggregates mislead about all four. The model identities and raw outputs below are frozen row facts; the classifications follow the historical report, including its uncertainty. 3 5

v2-strip-query: the clean model-specific rescue. The task needs a URL returned without its query component. All four baseline-family draws failed — C0’s one and C1’s three — and so did H1’s own baseline draw and its llama draw. The mistral draw passed, and it passed by a different route: instead of string surgery, it parsed the URL with urllib.parse and rebuilt it with the query replaced by an empty string. Baseline draws fail, a different model succeeds by a different approach. This is the shape a diversity rescue is supposed to have, observed once. 3

v2-half-up-rounding: sampling variation, not diversity. The task needs halves rounded away from zero. Every C0 and C1 draw failed. The passing H1 candidate came from H1’s baseline-model slot:

def round_half_up(x):
    return int(x + 0.5) if x >= 0 else int(x - 0.5)

Mistral and llama both failed. Same model as the baseline, different draw, different outcome. Calling this a diversity win would credit model identity for what the rows show as variation within one model. 3

v2-keyed-cache-length: ambiguous, kept ambiguous. All three H1 draws passed; all four baseline-family draws failed. That pattern fits two stories. The portfolio may genuinely help on this task — both other model slots passed as well. Or the baseline model’s H1 draw was simply lucky where its four twins were not, and the other models happen to handle the task well. The rows cannot separate those stories, and separate blind draws do not by themselves establish statistical independence. The report declines to choose. So does this chapter. 3 5

v2-splitlines-cr: the portfolio’s loss. One of C1’s three draws solved the carriage-return splitting task; all three H1 draws failed it. The portfolio did not only add coverage; it also missed a task the baseline found. Any account of the +2 that omits this −1 is advocacy. 3

Two more row facts bound the story. C0’s single draw solved two tasks, v2-lsp-square and v2-unsound-cache, that all six C1 and H1 draws missed: one draw beating six is sampling variation cutting the other way. And four tasks — registry pollution, retry-once-accepted, size format, stable priority — defeated every arm. They are the new hard core and the next ceiling. 1

So of three H1-only solves, exactly one carries the model-specific rescue shape.

Where the gain sits, and what “harder” means

Coverage by stratum, recomputed from frozen rows using the stratum labels in the semantic-repair-v1 corpus definition, and matching the P1.1 report’s table: 3 2

Stratum (10 tasks each)C0C1H1
Local0.600.800.90
Semantic0.800.800.80
Architectural0.800.700.70
Boundary0.500.800.90

Read at face value, the portfolio’s edge lives in local and boundary tasks. Read with the sample size, each of those edges is one task out of ten, and the architectural row shows a single baseline draw beating its own triplet. The report flags that as sampling noise and a caution against over-reading. Ten tasks per stratum cannot support a story about which kind of bug “needs diversity,” and the chapter does not tell one. 5

“Harder” needs the same care. semantic-repair-v1 was harder for this setup: single-draw baseline coverage fell from 11/12 to 27/40, which made room for the arms to differ. Nothing here establishes these as hard software-engineering problems in general, and the hidden tests define correctness narrowly per task. Test strength and task difficulty are separate axes: augmenting a code benchmark’s tests has been shown to cut measured pass rates substantially and even change model rankings (Liu et al., 2023).

Did the difference change the default?

The preregistered rules were written for P1 and bind P1.1 too. Rule 1 pursues heterogeneity only if the portfolio “materially beats” homogeneous sampling at similar cost. Rule 2 says that if they roughly tie, repeated sampling stays the simpler default. 6

“Materially” was never a number, and that is a real weakness of the rule. The P1.1 report judged +0.05 coverage — two tasks net, from a 3-to-1 discordant split, at about 30% more tokens and 2.5× median call latency, with no selector — to be a weak positive rather than a material win. So Rule 2 still governs, and homogeneous sampling remains the default. That judgment is the report’s, applied to a qualitative threshold; the sign test above supports it without having been part of the rule. 5 3

The report also applies Rule 5 — investigate decomposition, context, framing, or stronger generators before adding collaboration machinery — to the four-task hard core. Rule 5 was written for arms that are poor overall; applying it to the subset every arm failed is a reasonable extension, and it is the report’s. 5

Resource language stays descriptive. There was no selector, no value per solved task, and no pricing for local models, so the two tasks carry no dollar figure and no cost-effectiveness verdict. 1

What this is not

  • Not a portfolio advantage. Two net tasks from a 3-to-1 split is compatible with chance.
  • Unproven specialization. One mistral rescue and one ambiguous case do not show that models own task types.
  • Coverage without delivery. Coverage says a passing candidate existed; nothing here finds it.
  • No general difficulty ranking. Harder for this setup, under these hidden tests.

Where it is still weak

  1. Attribution is partly interpretation. Row identities and raw outputs are measured; the luck/rescue/ambiguous labels are the report’s reading. 5
  2. Small and local. Forty tasks, ten per stratum, three local models; causes of overlap unmeasured. 1
  3. No selector. The +2 exists only under retrospective identification.
  4. The threshold is qualitative. “Materially beats” was never a number. 6
  5. Portfolio composition is one choice. Two weaker models in two of three slots; a different mix would give a different inversion. 3

Do this now

Thirty minutes. Separate your coverage from your reliability.

  1. Take any result where you generated several candidates per task. Compute task coverage and candidate pass rate separately, then list the tasks exactly one approach solved.
  2. Count the discordant tasks both ways and run a sign test before calling the difference a gain.
  3. Split each portfolio’s candidates by generator. If the pass rate fell, find which slot pulled it down.
  4. For each one-approach solve, classify it: different-model success, same-model redraw, or ambiguous. Leave ambiguous rows ambiguous.
  5. Write the default decision with a numeric threshold attached: what gain, at what resource multiple, would change your process?

If you are building with an assistant:

Make problems hard enough that methods can differ, then match candidate
counts and lead with paired task differences before any aggregate: whose
solves, whose misses, net, and a sign test on the discordant tasks. Split
portfolio candidates by generator to explain any change in pass rate.
Classify each rescue from row identities: different-model success,
same-model redraw, or ambiguous. Keep coverage, candidate reliability, and
deployable selection as separate columns, report resources without pricing
them, and write numeric decision thresholds before the run.

Failure modes

  • Leading with 82.5% against 77.5%. Aggregates hide the four rows that are the finding.
  • Calling a 3-to-1 split a win. Four discordant tasks cannot carry that weight.
  • Calling all three H1-only solves diversity rescues. One is a baseline-model draw; one refuses classification.
  • Forgetting the C1-only solve. Net the gain or drop the claim.
  • Blaming “portfolios” for the pass rate. The rate fell because of which generators filled the slots.
  • Leaving “materially” undefined. A qualitative threshold turns every close result into a judgment call.

What this chapter established

  • A benchmark almost everything passes may have little power to tell methods apart. Chapter 24’s eleven-of-twelve baseline left only one task of headroom, and the eleven covered tasks showed no per-draw variation, so its tie said little about equivalence. When an evaluation saturates, increase its discriminating power before concluding anything stronger.
  • Harder means room to differ, not obscurity. Difficulty, test strength, representativeness and discriminating power are separate properties. This corpus was harder for this setup, which is all the comparison needed, and it establishes nothing about software-engineering difficulty in general.
  • Lead with paired differences, then test them. On the harder corpus the arms differed on four tasks, split three to one. The difference became observable; an exact sign test (p = 0.625) says it is still not evidence that either arm is better.
  • Coverage ≠ candidate reliability. H1 covered more tasks while recording a lower candidate pass rate. Two of its three slots came from models with lower observed pass rates than its qwen slot, but the paired exceptions show that the net coverage difference cannot simply be credited to model identity. Without a selector, extra coverage is potential, not delivered.
  • Rescue ≠ model-specific rescue. Classify each extra solve by what produced it. One clean complementary rescue is evidence worth preserving, not a routing policy.

What CodeAI measured. On semantic-repair-v1, H1 covered 33/40 tasks against C1’s 31/40, from three H1-only and one C1-only solve, a net +2 that an exact sign test cannot distinguish from chance (p = 0.625). 3 Coverage rose while candidate passes fell (64/120 against 78/120), because two of H1’s three slots went to models passing 0.475 and 0.375 of the time, at 1.30× C1’s tokens and a 7.7 s median call latency against 3.1 s. 1 3

Of the three H1-only solves, one is a clean model-specific rescue, one came from the baseline model’s own draw, and one stays ambiguous. 3 5 Under the preregistered rules the homogeneous default stands; the “material” threshold was qualitative. 5

Evidence notes

Independent verification. The analysis verifier asserts every P1.1 arm figure, the +0.05 premium, and the exact rescue sets against the recomputation, and catches its seeded corruption; the export verifier checks all four frozen bundle hashes. 1

The per-model split, the stratum table, the task-level pass comparison, and the sign test were recomputed from the same frozen rows for this chapter. They are not part of either verifier. The stratum labels come from the corpus definition in CodeAI source at corpus version semantic-repair-v1; that they reproduce the report’s table exactly is the check that the labels did not drift. 3

Next

Coverage rose while the observed candidate pass rate fell, and only one of the four discordant tasks carries the clean model-specific rescue signature. H1 finished two tasks ahead of C1 at about thirty percent more tokens, but this experiment does not establish that the extra model identities caused that net difference. The harder corpus also left something useful behind: tasks where methods visibly disagree, and four that defeated every arm. Those are the tasks where a cheaper kind of variation has room to show whether it helps. The next question drops model identity entirely: can cheaper variation — different wordings to one model — do the same work, and how do we tell an exploratory signal from something worth promoting?

Continue with Discovery Is Not Promotion.

References

  • Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training Verifiers to Solve Math Word Problems. arXiv:2110.14168, 2021. Paper.
  • Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. arXiv:2305.01210, 2023. Paper.

Implementation sources: P1.1 ran on the P-series CodeAI lineage (semantic-repair-v1 corpus; src/codeai/corpus_v2.py: semantic_corpus, CORPUS2_VERSION; src/codeai/experiments.py: ArmDef, run_arm; src/codeai/analysis.py: arm_metrics, unique_rescues, build_report). No CodeAI change accompanies this chapter. The per-model split, stratum table, task-level comparison, and sign test were recomputed from frozen rows during editing. Evidence: experiments/applied-ai/evidence/p-series/ (frozen exports, hashes verified), experiments/applied-ai/evidence/p-series-analysis/ (analysis.json, TABLES.md, analyze_pseries.py, verify_analysis.py), experiments/P1.1-results.md, and experiments/P1-interpretation.md. Nothing was rerun and nothing frozen was modified.


  1. Measured run: experiments/applied-ai/evidence/p-series-analysis. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  2. Source inspection: src/codeai/corpus_v2.py (semantic_corpus). ↩︎ ↩︎ ↩︎

  3. Measured run: experiments/applied-ai/evidence/p-series/p11. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  4. Source inspection: src/codeai/analysis.py (build_report). ↩︎

  5. Frozen report: experiments/P1.1-results.md. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  6. Preregistered interpretation: experiments/P1-interpretation.md. ↩︎ ↩︎