FPTQuant: Function-Preserving Transforms for LLM Quantization
The introduction of FPTQuant presents a significant advancement in the quantization of large language models (LLMs) by implementing three innovative function-preserving transforms. These transforms aim to enhance the efficiency of LLMs during inference without compromising performance, addressing the challenges posed by large magnitude outliers in naive quantization methods.