Princeton’s new Recurrent Looped Transformer (RLT) lets a decoder‑only model pull 96 blocks per token with unlimited depth—thanks to a clever attention cache and sliding‑window tricks. The future of token inference just got a serious boost. #RLT #AttentionCache #SlidingWindowAttention
🔗
