A new method called XOR-Trellis improves low-bit LLM weight quantization by combining a hardware-efficient dequantizer with curvature-aware path optimization, aiming to keep inference fast without costly codebooks.…
#LLMQuantization #AIHardware #EfficientInference
https://arxiv.org/abs/2610.00432
