BasicsEssential terms for understanding AI and large language models
对齐
The process of steering model behavior toward human intent, values and safety requirements.
Pretrained models just continue text; alignment teaches them to follow instructions, refuse harmful requests and stay honest. RLHF and DPO are mainstream methods—key to safety and usefulness.