ChainThink 消息,8 月 10 日,英伟达最新论文提出跨模型复用 KV Cache 的方法,尝试解决缓存通常只能在同一模型内部复用的问题。
KV Cache 是模型读取上下文后的中间计算结果,长上下文会显著占用显存和带宽。
此前 DeepSeek-V2 通过 MLA 相比 DeepSeek 67B 将 KV Cache 减少 93.3%,Kimi Linear 相比全 MLA 最多减少 75%。
论文称,在同一家族、KV 结构匹配的不同大小模型之间,缓存存在明显线性关系。
团队使用 500 段、每段 1024 Token 的文本校准后,可拟合映射,将一个模型算好的 KV Cache 转给另一个模型继续使用。
在 Qwen3-14B 切换至 32B 的测试中,32K 上下文重新计算约需 7 秒,转换缓存约需 0.28 秒,速度提升约 25 倍。
若该方法成熟,模型路由中长上下文 prefill 可更多由小模型完成,大模型在需要时直接接续生成。

Disclaimer: Contains third-party opinions, does not constitute financial advice
Amazon announces a multi-billion dollar investment in Missouri to build a data center campus, expected to create over 400 long-term positions
06-16
Elon Musk's current personal net worth has risen to the range of $1.1 trillion to $1.2 trillion.
06-15
OpenAI and Anthropic Employees Cash Out $14 Billion in Five Years
06-15
Tencent Involved in Investment in Lin Junyang's AI Lab, Former Head of Alibaba's Qwen, Valued at $2 Billion
06-15
Binance will delist the CVC/USDC, RPL/USDC, RVN/USDC, and XAI/USDC leveraged trading pairs on June 19
06-15
U.S. Security Experts Unite to Urge Repeal of Anthropic Ban, Warn That Restricting Defensive Tools Will Lose the AI Race Against China
06-15
Musk vows SpaceX revenue to reach $1 trillion by 2030
06-15







Key AI Events, Outlook on Industry Trends
ETH,BNB,SOL Hot Memes
As the 2026 crypto bear market deepens, exit scams and project blowups are becoming increasingly fre
Stock, Commodity & Bond Tokenization
BTC/ETH, Major Cryptocurrencies, and Hot Altcoins Price Trends
Follow new prediction market trends, insight real event expectations.
This column focuses on the real progress of Agents: technological evolution, application implementat
Tracking on-chain movements of the smart money and institutions
Spotlight on Frontier, trending projects, and breaking events
American Crypto Act – timely interpretations of policies worldwide