InferenceInference performance, quantization and cost optimization
速率限制
Provider limits on requests or tokens per unit of time.
Rate limits (RPM/TPM) protect services and control cost; exceeding them returns HTTP 429. Evaluate both price and limits when choosing—high-concurrency workloads need higher caps or multi-provider load balancing.