InferenceInference performance, quantization and cost optimization
推理时扩展
Spending more compute at inference time in exchange for higher-quality output.
Inference-time scaling lets models think longer before answering: sampling multiple candidates and picking the best, self-checking, and more. Small models can approach larger ones, at significantly higher token cost.