Patel states same-quality AI model output became 1200x cheaper than GPT-3 through successive model releases.
“Anthropic released new models. Google released new models. Meta released new models. Right? And then now it is 1200x times it's 1200x cheaper for the same output.”
Patel argues layering KV cache offload and multi-token prediction on open-source engines drastically cuts inference costs.
“So if you take the, you know, just open source inference engines off the shelf, that gets you a certain level of cost.”