Chain-of-thought (CoT)
prompting technique where the model is asked to show its reasoning step by step before giving a final answer. Improves accuracy on math, logic, and multi-step tasks. Trade-off: more tokens, higher latency.
The vocabulary you need to read an eval report without nodding along. Each entry is one paragraph; where the idea shows up somewhere in this site, there is a link to it.
prompting technique where the model is asked to show its reasoning step by step before giving a final answer. Improves accuracy on math, logic, and multi-step tasks. Trade-off: more tokens, higher latency.
a framework or script that runs a model against a test suite and collects scores. assert(llm) is a browser-based eval harness. Others: OpenAI Evals, EleutherAI lm-evaluation-harness, promptfoo.
Run the harness →whether the model’s answer is supported by the source document it was given. A model can be correct but unfaithful (right answer, wrong reasoning) or faithful but wrong (follows the document, but the document is wrong). Core metric in RAG evaluation.
including 2-5 example input→output pairs in the prompt so the model learns the expected format and behavior. Zero-shot = no examples. One-shot = one example. More shots generally improve consistency but cost tokens.
a manually curated set of question→expected_answer pairs used as ground truth for evaluation. The quality of your eval is capped by the quality of your golden dataset. Who validates the validator? This is the hardest problem in LLM testing.
Browse the bundled sets →the model generates information that is fabricated, not grounded in input or reality. Subtypes: phantom import (imports a nonexistent library), fabricated fact (invents a statistic), wrong entity (correct pattern, wrong name). Hallucination ≠ wrong answer: a wrong answer is a mistake, a hallucination is a confident invention.
See CODE-007: phantom delegation →using one LLM to evaluate another LLM’s output. Common pattern: Model A generates, Model B scores on a 1-5 scale. Risk: self-preference bias (a model rates its own family’s output higher). Mitigation: use a different model family as judge, or calibrate with human scores.
the same prompt can produce different answers on different runs (temperature > 0). Consequence: a single eval run is a sample, not a measurement. Reliable evaluation needs multiple runs and variance tracking.
testing whether a change to a prompt (system prompt, template, examples) makes outputs better or worse on a fixed dataset. Same idea as software regression testing, applied to prompts.
Retrieval-Augmented Generation Assessment. A framework for evaluating RAG pipelines on four dimensions: faithfulness, answer relevancy, context precision, context recall. Not for evaluating standalone LLMs — specifically for systems that retrieve documents and then generate.
controls randomness in token sampling. 0 = greedy (always pick the highest-probability token), 1 = sample from the full distribution. Lower temperature = more deterministic but less creative. For eval, use temperature 0 to minimize non-determinism.
Definitions are deliberately short. For the failures these concepts describe in practice, read the findings.