How to test an LLM for hallucinations
A hallucination is a test failure you have not written yet. Start from answers you already know, let deterministic checks decide, and keep the judge advisory.
Ask a model the same question twice and you may get two confident, fluent, different answers. One of them is made up. The problem is not that models hallucinate; it is that most teams find out from a user, not from a test.
This post walks through how assert(llm) catches hallucinations with plain test engineering: a golden dataset, scorers that do not guess, and an LLM judge that is allowed to comment but not to vote.
Hallucination is a test failure, not a vibe
“The model sometimes makes things up” is not actionable. “Item CODE-007 fails on every run because the answer imports an SDK and calls a model id that does not exist” is. The difference is a test with a known expected answer and a rule that decides pass or fail the same way every time.
- Known answers. Each item carries an
expectedvalue you can defend. - Deterministic verdicts. The same reply always gets the same score.
- Reproducible runs. A failure you can re-run is a failure you can fix.
Start from answers you already know
A golden dataset is a list of prompts whose correct answers you know ahead of time. In assert(llm) every item is typed, and the type decides how it is scored: mcq compares the letter the model picked, code checks the function signature and forbidden constructs, short looks for the expected phrase on word boundaries.
type: mcq · code · short
Case: phantom delegation
Asked for a CSV parser, Haiku 4.5 repeatedly answers with code that imports the Anthropic SDK and asks a different model to do the parsing. It is not a one-off: it recurs on every run of the item.
import anthropic
def csv_to_dicts(s):
client = anthropic.Anthropic()
resp = client.messages.create(
model="claude-3-5-sonnet-20241022", # does not exist
messages=[{"role": "user", "content": f"Parse this CSV: {s}"}],
)
return eval(resp.content[0].text) # eval on model output
Three problems stack up in nine lines: a model id that does not exist, an eval() on another model's output, and no parsing at all. The task was to write a parser; nothing in the answer parses anything.
The scorer missed this at first. The answer defines the right function with the right arity and returns a value, so it passed every item-level check. It now fails on a global rule that forbids importing an LLM SDK in a coding answer.
When the scorer is the bug
Writing tests against the scorer found three bugs in the scorer rather than in any model. Each one turned a correct verdict into a wrong one:
| Reply | Expected | Before the fix | Now |
|---|---|---|---|
| “A good test objective is to reduce risk” | a | PASS article read as option a | FAIL article skipped |
| “I cannot answer without the options” | c | PASS substring “c” | FAIL word boundaries |
| “299 792 km/s” | 299792 | FAIL thin space lost | PASS de-grouped first |
For a while the eval was measuring its own matching rules at least as much as it was measuring the models.
— Findings, October 2026
Where an LLM judge fits
A judge is good at explaining a failure in a sentence and bad at being the source of truth. In assert(llm) each judge scores the same answer 0–5 with a one-line reason, and that score sits next to the verdict instead of replacing it. Two judges that disagree stay visible as two averages.
Run the Python coding suite on your own models
20 coding tasks, the same SDK-import rule, any of six providers. Your key stays in the tab.