MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head Paper • 2601.07832 • Published 12 days ago • 51
Towards Automated Kernel Generation in the Era of LLMs Paper • 2601.15727 • Published 2 days ago • 13
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Paper • 2601.14724 • Published 3 days ago • 59