DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

(zartbot.github.io)

39 points | by mfiguiere 2 hours ago

2 comments

  • arikrahman 30 minutes ago
    I am very impressed with the KV Cache Compression work as well as the prefix cacheing making queries converge on practically free.
  • smy20011 17 minutes ago
    Should we flag this since It's AI generated?
    • sebmellen 2 minutes ago
      The writing feels human to me… and I call out AI slop as much as possible.