Patel says Anthropic generates $50 million per megawatt revenue, enabling 5x return on inference spending.
“In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”
Patel reveals Anthropic turned first profit in Q2 and will hit billion-dollar operating profit in Q3 before IPO.
“So Anthropic turned their first gross profit in Q2 in June. And then in Q3, they will be turning a billion dollars of operating profit, slightly over.”
Patel reports OpenAI gross margins rose from 30% to 55% overall, 50% to 65% excluding free users.
“You look at OpenAI late last year, their margins had were roughly 30% gross margin, but if you stripped away the free users, they were at 50%.”
Patel says their AI spending equals over a third of employee compensation for their 90-person firm.
“We're spending, you know, more like, you know, more than a third of the spend, you know, employee spend, a third of it is also on top of that is AI.”
Patel argues Mythos fast mode is probably cheaper than Claude 4.6 fast mode for most tasks due to token efficiency.
“the flip side is, is, is mythos is more token efficient. Methos fast mode is probably cheaper than like, four, six fast mode for most tasks.”
Patel cites Anthropic adding $67 billion ARR per month as evidence of expanding AI revenue beyond hyperscalers.
“You're starting to see it with Anthropix revenue adding $67,000,000,000 of ARR a month, but there's so many more firms coming online with revenue streaming in,”
Patel states Anthropic now adds $23 billion revenue monthly versus few hundred million earlier.
“Anthropics, you know, adding $23,000,000,000 of revenue a month now Mhmm. Versus they were just adding a few 100,000,000 of revenue a month earlier. So clearly, we're in the take off period.”
Patel reports Claude Code doubled to 4% of GitHub commits in January alone, with overall AI coding likely at 10%.
“Just in this month just in January, it went from 4% of or 2% of commits on GitHub to 4% of GitHub commits were done by Cloud Code.”
Patel explains prefill costs one-fourth of decode per token, making 30,000 token inputs costlier than 2-4,000 token outputs.
“And that's a very common ratio. Right? And then when you think about, okay, the cost of running pre fill is roughly one fourth of running decode.”
Patel states avoiding redundant prefill through KV cache can cut inference costs to one-fourth of previous levels.
“So you can cut your cost to one fourth of what it was previously if you just don't do the pre fill. Right?”
Patel says inference providers sell public endpoints at flat or negative margins, compensating through private deployments.
“Most of the inference providers are selling at flat margins or even negative for their public endpoints. And they then make it up when people do private deployments.”
Patel calculates 20% power efficiency difference translates to only 4% TCO difference on NVIDIA deployments.
“Because if you have enough power, a 20% difference in performance per watt only ends up being a 4% difference in TCO.”
Patel says the standard inference deployment unit has shifted from single nodes to hundreds of GPUs.
“the standard unit for an inference deployment being hundreds of GPUs instead of a single node. And then there's all these different things about traffic.”
Patel notes NVIDIA Blackwell performance improved dramatically from launch to present, as did AMD hardware.
“If you tried to use NVIDIA's Blackwell six months ago, the numbers were not amazing, right? But now they're actually amazing. So how did that progress over time? Same with AMD, right?”
Patel breaks down inference TCO as 20% electricity and data center, 80% hardware on NVIDIA deployments.
“20% of your cost is your electricity, your data center real estate, roughly. And then the rest of the cost is that hardware, at least on a standard NVIDIA deployment your GPUs, your networking, etcetera.”
Patel identifies GB200's power efficiency advantage but notes deployment challenges with backplane and liquid cooling.
“GB200 has a huge power efficiency advantage, right? Everyone here understands the challenges of running and deploying GB200. There's a lot of challenges with the backplane.”
Patel reports GB200 is 10x more power efficient than H200 at certain interactivity rates.
“But it turns out at certain interactivity rates, I. E. Tokens per second per user, it's 10x more efficient per watt, right, compared to H200.”
Patel finds AMD MI355 beats NVIDIA B200 on performance TCO in certain publicly usable configurations.
“So we do different scenarios. We do document processing, which is 8,000 context in, 1,000 out. We do chat, which is 1,000 in, 1,000 out.”
Patel shows B200 has 15x raw performance advantage over H100 but only 10x performance per TCO.
“If we don't divide by TCO, then it looks like the performance of B200 is actually 15x that of H100, versus the performance TCO is only 10x,”
Patel states Supermicro's liquid cooling reduces Hopper server power from 10 kilowatts to 7-8 kilowatts versus air-cooled competitors.
“Everyone else's hopper servers h 100 air cooled. And so each server consumes, you know, 10 kilowatts almost. Right?”
Patel calculates that without Supermicro's liquid cooling solution, servers would consume 30% more power.
“So that's why Shibbol Micro dedicated deep cooling so aggressively. And if the liquid cooling solution from Super Micro didn't exist, then those servers would be consuming 30% more power.”
Patel describes diagnostic challenges in AI data centers where failures occur across multiple infrastructure layers.
“And people have had problems where their data center wasn't ready and something stopped working. They don't know, is it the chip? No.”
Patel explains OpenAI's router will send low-value queries to cheaper models but use expensive compute for high-value monetizable queries.
“if the user asks a low value query like, hey, why is the sky blue? Just route them to mini. The model can answer perfectly fine, and that is a chunk of queries.”
Patel says AI model costs dropped 1200x from GPT-3's $60 per million outputs to current pricing.
“GPT three was it cost, you know, it cost $60 for the million output. And over time, that kept reducing. Right? OpenAI released new models.”
Patel explains reasoning models dramatically increase memory usage and reduce batch size, multiplying serving costs.
“So your your memory usage is going way up with these reasoning models, and you still have a lot of users.”
Patel explains reasoning models dramatically increase serving costs due to long output context and memory constraints.
“your memory usage is going way up with these reasoning models, and you still have a lot of users. So effectively, the cost to serve multiplies by a ton.”
Patel estimates Blackwell delivers 10-15x cost improvement for inference despite NVIDIA claiming 30x at GTC.
“But now, like, Blackwell, NVIDIA's pitching 10 to 15 x improvement in cost. It's like, well, you know, they're massaging the numbers marketing.”
Patel explains OpenAI's o1 model thinks for 5-20 seconds before outputting, creating inference throughput challenges.
“When you use OpenAI's o one, it thinks for ten seconds, twenty seconds, five seconds. It varies a lot, but thinks for a while and then it sends you tokens.”
Patel reports H100 rental prices have dropped from $3-4 per hour to $2.15 or less, approaching natural cost of $1.40.
“An hour. Right? For shorter term or midterm deals. Right now, it's like, if you want a six month deal, you could get, like, $2.15 or less.”
Patel argues AI software will have lower R&D costs but much higher cost of goods sold from operating services.
“Like the R and D cost is much lower in terms of people, but the cost of goods sold in terms of actually operating the service, I think will be much higher.”