Microsoft demoed DeepSeek V4 Flash running on a laptop:
284B open-weight, quantized to 1.6 bits → ~60GB.
At 8 bits the same weights need ~284GB.
Memory, not compute, decides if it runs.
On-device inference doesn't kill the bill —
it changes its unit. #LocalInference
