Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
portfolio
publications
A Survey of Efficient Attention Methods: Hardware-Efficient, Sparse, Compact, and Linear Attention
Published in ArXiv preprint, 2025
A comprehensive survey of efficient attention methods across hardware-efficient, sparse, compact, and linear designs.
Recommended citation: Zhang et al., 2025.
Download Paper
VoidPadding: Let [VOID] Handle Padding in Masked Diffusion Language Models so that [EOS] Can Focus on Semantic Termination
Published in COLM 2026 Workshop on Efficient Reasoning (Spotlight), 2026
A dedicated padding token for robust masked diffusion language model decoding.
Recommended citation: Liu et al., 2026.
Download Paper
Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory
Published in COLM 2026 Workshop on Efficient Reasoning, 2026
CoMem uses depth-partitioned memory to keep model-side read cost independent of stored-context length.
Recommended citation: Liu et al., 2026.
Download Paper
JustQuant: You Don’t Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization
Published in ArXiv preprint, 2026
Naive low-bit activation quantization for generative models without smoothing, SVD, or rotations.
Recommended citation: Yang et al., 2026.
Download Paper
Frozen in a Frame: The Velocity Blind Spot in JEPA World Models
Published in ArXiv preprint, 2026
This work identifies the velocity blind spot in single-frame JEPA world models and introduces a lightweight latent decomposition that makes motion explicitly recoverable.
Recommended citation: Zhang et al., 2026.
Download Paper