From Attention to Gluing: A Sheaf-State Architecture for Lower-Complexity Language Models
This paper proposes a "Sheaf-State Language Model" architecture that replaces computationally excessive dense self-attention with a lower-complexity framework using local state-space dynamics and sparse, typed gluing morphisms to efficiently manage context and dependencies.