Gerstner reports Google inference generation increased 100x in a year, from 9 trillion to 980 trillion tokens monthly.
“A year ago, Google per month was doing about 9,000,000,000,000 tokens a month in terms of inference generation, right, compute generation. Today, it's 980,000,000,000,000 tokens.”
Gerstner argues the difference between $3 and $15 inference costs is irrelevant when replacing a $200 per hour consultant.
“The difference between spending $3 on a cheap model or $15 on an expensive model to replace a $200 an hour consultant. It's just irrelevance. That inference cost difference is irrelevant”