Kimi Linear: An Expressive, Efficient Attention Architecture

Jul 28, 2026 05:52 PM - 1 day ago 82

[Submitted connected 30 Oct 2025 (v1), past revised 1 Nov 2025 (this version, v2)]

Authors:Kimi Team: Yu Zhang, Zongyu Lin, Xingcheng Yao, Jiaxi Hu, Fanqing Meng, Chengyin Liu, Xin Men, Songlin Yang, Zhiyuan Li, Wentao Li, Enzhe Lu, Weizhou Liu, Yanru Chen, Weixin Xu, Longhui Yu, Yejie Wang, Yu Fan, Longguang Zhong, Enming Yuan, Dehao Zhang, Yizhi Zhang, T.Y. Liu, Haiming Wang, Shengjun Fang, Weiran He, Shaowei Liu, Yiwei Li, Jianlin Su, Jiezhong Qiu, Bo Pang, Junjie Yan, Zhejun Jiang, Weixiao Huang, Bohong Yin, Jiacheng You, Chu Wei, Zhengtao Wang, Chao Hong, Yutian Chen, Guanduo Chen, Yucheng Wang, Huabin Zheng, Feng Wang, Yibo Liu, Mengnan Dong, Zheng Zhang, Siyuan Pan, Wenhao Wu, Yuhao Wu, Longyu Guan, Jiawen Tao, Guohong Fu, Xinran Xu, Yuzhi Wang, Guokun Lai, Yuxin Wu, Xinyu Zhou, Zhilin Yang, Yulun Du

View PDF

Abstract:We present Kimi Linear, a hybrid linear attraction architecture that, for the first time, outperforms afloat attraction nether adjacent comparisons crossed various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its halfway lies Kimi Delta Attention (KDA), an expressive linear attraction module that extends Gated DeltaNet pinch a finer-grained gating mechanism, enabling much effective usage of constricted finite-state RNN memory. Our bespoke chunkwise algorithm achieves precocious hardware ratio done a specialized version of the Diagonal-Plus-Low-Rank (DPLR) modulation matrices, which substantially reduces computation compared to the wide DPLR formulation while remaining much accordant pinch the classical delta rule.
We pretrain a Kimi Linear exemplary pinch 3B activated parameters and 48B full parameters, based connected a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that pinch an identical training recipe, Kimi Linear outperforms afloat MLA pinch a sizeable separator crossed each evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results show that Kimi Linear tin beryllium a drop-in replacement for afloat attraction architectures pinch superior capacity and efficiency, including tasks pinch longer input and output lengths.
To support further research, we open-source the KDA kernel and vLLM implementations, and merchandise the pre-trained and instruction-tuned exemplary checkpoints.

Submission history

From: Yulun Du [view email]
[v1] Thu, 30 Oct 2025 16:59:43 UTC (645 KB)
[v2] Sat, 1 Nov 2025 12:05:18 UTC (691 KB)

More