What the evals actually turned up
Three things worth writing down: a model that outsources its homework, an arithmetic blind spot that repeats across runs, and a handful of places where the scorer itself was the thing that was wrong.
Phantom delegation: the model calls another model instead of writing the parser
Asked for a CSV parser, Haiku 4.5 repeatedly answers with code that imports the Anthropic SDK and asks a different model to do the parsing. It is not a one-off: it recurs on every run of the item.
import anthropic
def csv_to_dicts(s):
client = anthropic.Anthropic()
resp = client.messages.create(
model="claude-3-5-sonnet-20241022", # does not exist
messages=[{"role": "user", "content": f"Parse this CSV: {s}"}],
)
return eval(resp.content[0].text) # eval on model output
Three separate problems stack up in nine lines:
- Hallucination. The model id it reaches for,
claude-3-5-sonnet-20241022, is not a model that exists. - Security hole. The result is handed straight to
eval()— arbitrary code execution driven by another model's output. - Delegation instead of solving. The task was to write a parser. Nothing in the answer parses anything.
The scorer missed this at first: the answer defines the right function with the right arity and does return a value, so it passed every check. It now fails on a global rule that forbids importing an LLM SDK in a coding answer, no matter what the individual item lists.
Three-point estimation: the weighted mean gets rounded to the likely case
Optimistic 2 hours, most likely 11 hours, pessimistic 14 hours. The three-point (PERT) estimate is the weighted mean, and it comes out exactly on a listed option:
(O + 4M + P) / 6 (2 + 4×11 + 14) / 6 (2 + 44 + 14) / 6 = 60 / 6 = 10
The answer is d) 10 hours. Haiku 4.5 consistently picks a neighbouring option instead — it lands on the distractors clustered around the most likely estimate rather than computing the mean. The arithmetic is not hard; the model simply does not do it, and it fails the same way on repeated runs.
This is the useful shape of an eval failure: deterministic, cheap to re-run, and traceable to one formula.
Scoring edge cases
Writing tests against the scorer found three bugs in the scorer rather than in any model. All three are fixed; each has a regression test.
1. The article “a” versus option “a” fixed
The multiple-choice parser falls back to “last standalone letter in the reply” when nothing more explicit is found. English makes that dangerous, because a is also an article:
"A good test objective is to reduce risk" → parsed as option a → PASS
On an item keyed to a, pure prose scored a point. The fallback now skips an a that is immediately followed by another word, and keeps every other letter. An option stated as a) or a. is matched earlier by the label rule, so nothing legitimate was lost.
2. contains("c") matched “cannot” fixed
A refusal counted as a correct answer whenever the keyed letter appeared anywhere inside a word:
expected "c" · reply "I cannot answer without the options" → PASS
Substring matching was the culprit, and the same flaw ran through the code scorer's forbid list: import re matched import requests, and forbidding sorted( also matched the function's own name merge_sorted(. Matching is now fenced on word boundaries everywhere — letters, short answers and forbidden constructs alike.
3. NFKC normalization ran before digit de-grouping fixed
Answers were Unicode-normalized first and de-grouped second. NFKC folds a thin space into an ordinary space, so by the time the de-grouping rule looked for a separator there was nothing left to find:
expected "299792" · reply "299 792 km/s" → FAIL (correct answer, marked wrong)
Grouped numbers are now collapsed before normalization, covering the comma, the no-break space, the thin space and the narrow no-break space. A plain space is deliberately left alone, so prose like “8 792” is not silently joined into one number.
The lesson is the uncomfortable one: for a while the eval was measuring its own matching rules at least as much as it was measuring the models. Every rule above is now pinned by a test, and the handful of behaviours still considered wrong are marked as TODO in the suite rather than quietly tolerated.
Findings are from runs against the bundled ISTQB, Python coding, SQL and general-knowledge sets. Re-run them yourself on the run page — everything happens in your own browser.