InferenceInference performance, quantization and cost optimization

推理时扩展

Spending more compute at inference time in exchange for higher-quality output.

Inference-time scaling lets models think longer before answering: sampling multiple candidates and picking the best, self-checking, and more. Small models can approach larger ones, at significantly higher token cost.

Related terms