Patel claims hardware improved 30x from Hopper to Blackwell for DeepSeek on optimized deployments.
“from Hopper to Blackwell, which is all we've had over the last three years, roughly 30x improvement on DeepSeek, on the most optimized deployment, which you can see on InferenceX there's about a 30x improvement.”
Patel cites DeepSeek's open-sourced inference system requiring 140 GPUs communicating over RDMA networks for single replica.
“One example is one that DeepSeek open sourced over December of last yearJanuary, February of this year, where a single replica of inference for a single model is going to be like 140 GPUs.”
Patel reveals DeepSeek inference implementation requires 160 GPUs worth over $10 million of hardware per replica.
“That's over $10,000,000 of hardware, and then that's just one replica, then you'll have a lot of replicas and you share the caching servers between them.”
Patel reports GPT-3 to Llama 3.2 costs fell 1,200x while GPT-4 to DeepSeek v3 costs fell 600x.
“when we looked at g p d three, the cost fell 1,200 x from g p d three's initial cost to what you can get Lama 3.23 b today.”
Patel says DeepSeek's cost efficiency follows the expected trend line, just from an unexpected source.
“It's not unexpected. Right? Like, this is actually within the trend line of what happened with GPT three is happening to GPT four level quality with DeepSeq.”
Patel challenges DeepSeek's GPU count claims, citing job ads promising tens of thousands of GPUs.
“The estimates that we have is, so first of all, ads in China, they say they have tens of thousands of GPUs for researchers, right?”
Patel explains DeepSeek v3 base is trained once, then post-trained differently to create chat versus reasoning models.
“This reasoning model has a lot of overlapping training steps to DeepSeek v three, and it's confusing that you have a base model called v three that you do something to to get a chat model, and then you do some different things to get a reasoning model.”
Patel says DeepSeek modifies code at or below NVIDIA's CUDA layer, a rare technical capability.
“For example, on their to get highly efficient training, they're making modifications at or below the CUDA layer for NVIDIA chips.”
Patel notes DeepSeek's routing innovation removing auxiliary loss represents compounding small improvements over time.
“this type of change can be big, it can be small, but they add up over time.”