On the record about
2 people · 8 quotes · 23 Dec 2024 to 14 Aug 2025
1 of 2 lane rests on fewer than 5 quotes and is marked thin. Offsets are days from the middle first-quote date, 23 Dec 2024 — a date, and nothing else. It is not a claim about who reached a view first.
Patel describes reasoning models generating thousands of thinking tokens, sometimes switching languages mid-reasoning.
“It generates tons of things. It's like it it sometimes switches between Chinese and English. Right? Like, whatever it is. It's thinking. Right? It's churning.”
Gurley explains o one reasoning models cost 50x more per query due to 10x tokens and one-fifth concurrent user capacity.
“E. Concurrent users I can have is a fraction of that. One fourth to one fifth the number of users can currently use the server.”
“You know, I can go from 10,000,000 to 100,000,000 to billion to $10,000,000,000 on reasoning in such a quick succession.”
Patel explains reasoning models dramatically increase memory usage and reduce batch size, multiplying serving costs.
“So your your memory usage is going way up with these reasoning models, and you still have a lot of users.”
Patel explains reasoning models dramatically increase serving costs due to long output context and memory constraints.
“your memory usage is going way up with these reasoning models, and you still have a lot of users. So effectively, the cost to serve multiplies by a ton.”
Patel says DeepSeek r one zero shows reasoning behaviors emerge naturally from RL on verifiable rewards without human data.
“So it's the remarkable thing about these reasoning results, and especially the DeepSeek r one paper, is this result that they call DeepSeek r one zero, which is they took one of these pretrained models.”
Patel says reasoning capability allowing models to think before answering dramatically increased capabilities in six months.
“Right? And this has just happened in six months. And the same applies to various certain coding tasks. The same applies to yeah.”
Patel reports alternative data shows reasoning models have low actual API usage despite availability.
“So this is something that I've found very interesting is that we've been trying to build a lot of alternative data sources for token usage, who's using tokens, what models, where, etcetera, why, and it's very clear that people aren't actually using the reasoning models that much in API.”