BasicsEssential terms for understanding AI and large language models
基准测试
Standardized tasks and scoring used to compare model capabilities quantitatively.
Benchmarks cover math, code, reasoning and knowledge, e.g. MMLU, HumanEval and GPQA. Leaderboard scores guide model selection, but data contamination and saturation mean real-world evaluation still matters.