Blog · notes from an LLM test bench
What the evals turned up
Findings from real runs, deep dives into how the scoring works, and the occasional update. Every post ends with something you can re-run in your own browser.
Deep dive
How to test an LLM for hallucinations
A hallucination is a test failure you have not written yet. Start from answers you already know, let deterministic checks decide, and keep the judge advisory.
Read the post →Run these tests on your own models
ISTQB, Python, SQL, general knowledge, or your own prompts. Your key stays in the tab.