AI termBrowse the neighboring terms

Failures / Research term

Benchmark contamination

Overlap or close derivation between evaluation material and data used to train, tune, prompt, or select a model, which can distort the score's interpretation.

Contamination ranges from exact question-answer leakage to paraphrases, benchmark-specific coaching, or repeated development against the test set. A high score may then reflect memorization or test adaptation as well as transferable capability. The amount and effect are difficult to establish when training data and model-development procedures are not fully disclosed.

Builder example

Public scores remain one signal, but product selection needs fresh or protected cases that match the actual workflow. Even a clean benchmark can be irrelevant if its tasks, tools, latency, or scoring do not match deployment.

Common confusion: Contamination does not make every result useless, and privacy does not guarantee cleanliness. An internal eval can also be overfit through repeated prompt and model selection.