InferenceInference performance, quantization and cost optimization

量化

Representing weights in lower precision (8-bit, 4-bit) to shrink size and speed up inference.

Quantization compresses FP16 weights to INT8/INT4, cutting memory by 2-4x and inference cost sharply—the top choice for local deployment and cost optimization. Popular schemes: GPTQ, AWQ, GGUF.

Related terms