The ALHR method employs hierarchical routing to lower VRAM usage from 57 to 422 MB, achieving 92.1% top-1 accuracy at 1024 tokens with NlogN inference scaling.

The ALHR method employs hierarchical routing to lower VRAM usage from 57 to 422 MB, achieving 92.1% top-1 accuracy at 1024 tokens with NlogN inference scaling.