BasicsEssential terms for understanding AI and large language models

基准测试

Standardized tasks and scoring used to compare model capabilities quantitatively.

Benchmarks cover math, code, reasoning and knowledge, e.g. MMLU, HumanEval and GPQA. Leaderboard scores guide model selection, but data contamination and saturation mean real-world evaluation still matters.

Related terms