Findings · October 2026

What the evals actually turned up

Three things worth writing down: a model that outsources its homework, an arithmetic blind spot that repeats across runs, and a handful of places where the scorer itself was the thing that was wrong.


Case CODE-007reproducible

Phantom delegation: the model calls another model instead of writing the parser

def csv_to_dicts(s) · expected haiku-4.5 → import anthropic

Asked for a CSV parser, Haiku 4.5 repeatedly answers with code that imports the Anthropic SDK and asks a different model to do the parsing. It is not a one-off: it recurs on every run of the item.

import anthropic

def csv_to_dicts(s):
    client = anthropic.Anthropic()
    resp = client.messages.create(
        model="claude-3-5-sonnet-20241022",   # does not exist
        messages=[{"role": "user", "content": f"Parse this CSV: {s}"}],
    )
    return eval(resp.content[0].text)        # eval on model output

Three separate problems stack up in nine lines:

The scorer missed this at first: the answer defines the right function with the right arity and does return a value, so it passed every check. It now fails on a global rule that forbids importing an LLM SDK in a coding answer, no matter what the individual item lists.

Case ISTQB-13reproducible

Three-point estimation: the weighted mean gets rounded to the likely case

d) 10 hours · expected haiku-4.5 → c) 10.5 hours

Optimistic 2 hours, most likely 11 hours, pessimistic 14 hours. The three-point (PERT) estimate is the weighted mean, and it comes out exactly on a listed option:

(O + 4M + P) / 6
(2 + 4×11 + 14) / 6
(2 + 44 + 14) / 6  =  60 / 6  =  10

The answer is d) 10 hours. Haiku 4.5 consistently picks a neighbouring option instead — it lands on the distractors clustered around the most likely estimate rather than computing the mean. The arithmetic is not hard; the model simply does not do it, and it fails the same way on repeated runs.

This is the useful shape of an eval failure: deterministic, cheap to re-run, and traceable to one formula.

Scoring edge cases

Writing tests against the scorer found three bugs in the scorer rather than in any model. All three are fixed; each has a regression test.

1. The article “a” versus option “a” fixed

The multiple-choice parser falls back to “last standalone letter in the reply” when nothing more explicit is found. English makes that dangerous, because a is also an article:

"A good test objective is to reduce risk"   → parsed as option a → PASS

On an item keyed to a, pure prose scored a point. The fallback now skips an a that is immediately followed by another word, and keeps every other letter. An option stated as a) or a. is matched earlier by the label rule, so nothing legitimate was lost.

2. contains("c") matched “cannot” fixed

A refusal counted as a correct answer whenever the keyed letter appeared anywhere inside a word:

expected "c" · reply "I cannot answer without the options"  → PASS

Substring matching was the culprit, and the same flaw ran through the code scorer's forbid list: import re matched import requests, and forbidding sorted( also matched the function's own name merge_sorted(. Matching is now fenced on word boundaries everywhere — letters, short answers and forbidden constructs alike.

3. NFKC normalization ran before digit de-grouping fixed

Answers were Unicode-normalized first and de-grouped second. NFKC folds a thin space into an ordinary space, so by the time the de-grouping rule looked for a separator there was nothing left to find:

expected "299792" · reply "299 792 km/s"  → FAIL (correct answer, marked wrong)

Grouped numbers are now collapsed before normalization, covering the comma, the no-break space, the thin space and the narrow no-break space. A plain space is deliberately left alone, so prose like “8 792” is not silently joined into one number.

The lesson is the uncomfortable one: for a while the eval was measuring its own matching rules at least as much as it was measuring the models. Every rule above is now pinned by a test, and the handful of behaviours still considered wrong are marked as TODO in the suite rather than quietly tolerated.

Findings are from runs against the bundled ISTQB, Python coding, SQL and general-knowledge sets. Re-run them yourself on the run page — everything happens in your own browser.