Patel observes Asian researchers consistently produce KV cache reduction papers unknown to American researchers.
“Some paper that reduces KV cash, and no researcher in America has even heard of this paper.”
Patel observes Asian researchers consistently reference KV cache reduction papers unknown to American researchers, citing DeepSeek and TurboQuant.
“It's the fucking best thing ever. DeepSeek and TurboQuant were the, like, most precipitous ones that popped up the most,”
Patel states avoiding redundant prefill through KV cache can cut inference costs to one-fourth of previous levels.
“So you can cut your cost to one fourth of what it was previously if you just don't do the pre fill. Right?”