Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory

Published in COLM 2026 Workshop on Efficient Reasoning, 2026

CoMem caches mid-layer chunk representations and retrieves a bounded relevant set for long-context language modeling.

Recommended citation: Liu et al., 2026.
Download Paper