📝 Blogs
Diagonal-Plus-Low-Rank Matrices and Kimi Delta Attention
How diagonal decay and rank-1 correction combine in DPLR state transitions, and how Kimi Delta Attention constrains this structure for efficient blockwise computation.
DeltaNet and Gated DeltaNet Through the Lens of Online Learning and Gradient Descent
A step-by-step view of linear attention, DeltaNet, and Gated DeltaNet as online optimization rules over fast-weight associative memory.
Rank-1 Matrices, Householder Transformations, and the WY Representation
From rank-1 outer products to Householder transformations and the compact WY representation: why low-rank structure matters for blockwise GPU computation.