By the numbers:
16-bit to 8-bit: Statistically indistinguishable outputs
4-bit: The knee of the curve where quality loss remains small
3-bit: Noticeably worse
2-bit: Visibly degraded text
#gpuoptimization #batchinference #quantization #llm #modelcompression #aiperformance
