On the record about
5 people · 21 quotes · 4 Jan 2025 to 30 Jun 2026
3 of 5 lanes rest on fewer than 5 quotes and are marked thin. Offsets are days from the middle first-quote date, 3 Feb 2025 — a date, and nothing else. It is not a claim about who reached a view first.
Baker predicts frontier AI labs will stop releasing leading-edge models to prevent IP theft via knowledge distillation, citing DeepSeek.
“I think you will see the Frontier Labs stop releasing their leading edge models to prevent knowledge distillation and their IP effectively being stolen.”
Sacks says people would have been surprised a Chinese company released the second reasoning model after OpenAI.
“if you had said to people a few weeks ago that the second company to release a reasoning model along the lines of o one would be a Chinese company, I think people would have been surprised by that.”
Sacks says the $6 million DeepSeek training cost claim should be debunked, corroborated by industry leaders.
“On this one, I'm with Palmer Luckey and Brad Gerstner and others, and I think this has been pretty much corroborated by everyone I've talked to that that number should be debunked.”
Sacks cites Dylan Patel's estimate that DeepSeek has 50,000 Hopper chips across different models.
“Dylan Patel, who's leading semiconductor analyst, has estimated that DeepSeac has about 50,000 hoppers. And specifically, he said they have about 10,000 h one hundreds, They have 10,000 h eight hundreds and 30,000 h twenties.”
Sacks says DeepSeek's compute cluster costs over $1 billion, contradicting the $6 million narrative.
“you add up the the cost of a compute cluster with 50,000 plus hoppers and it's gonna be over $1,000,000,000.”
Sacks says DeepSeek v3 self-identified as ChatGPT-4 five out of eight times when asked.
“When you would ask it, who are you? Like what model are you? Five out of eight times, v three would tell you that it was ChatGPT four.”
Patel explains DeepSeek v3 base is trained once, then post-trained differently to create chat versus reasoning models.
“This reasoning model has a lot of overlapping training steps to DeepSeek v three, and it's confusing that you have a base model called v three that you do something to to get a chat model, and then you do some different things to get a reasoning model.”
Patel says DeepSeek modifies code at or below NVIDIA's CUDA layer, a rare technical capability.
“For example, on their to get highly efficient training, they're making modifications at or below the CUDA layer for NVIDIA chips.”
Patel notes DeepSeek's routing innovation removing auxiliary loss represents compounding small improvements over time.
“this type of change can be big, it can be small, but they add up over time.”
Gurley explains DeepSeek innovated by separating parameters to work with smaller counts faster than competitors.
“They were able to do that because they figured out a way to separate the parameters and work on things with smaller parameter counts faster, which no one else had done before.”
Gurley reports DeepSeek pricing is one twentieth of OpenAI on comparable API models.
“if you look at the models that are apples to apples on the API right now, that that DeepSeek's pricing about one twentieth of OpenAI.”
Gerstner reports DeepSeek is massively compute constrained now, with GPU differential growing versus OpenAI for o3.
“On top of that, we learned, you know, they are massively compute constrained right now. So you've seen some tweets about this, people that are, you know, getting server delay and all this stuff”
Gurley reports 1,300 variants of DeepSeek R1 appeared on Hugging Face within days of release.
“forty eight hours after r one was posted, they had 500 variants on Hugging Face. And today I pinged them this morning before we started, they're up to 1,300.”
Patel challenges DeepSeek's GPU count claims, citing job ads promising tens of thousands of GPUs.
“The estimates that we have is, so first of all, ads in China, they say they have tens of thousands of GPUs for researchers, right?”
Patel reports GPT-3 to Llama 3.2 costs fell 1,200x while GPT-4 to DeepSeek v3 costs fell 600x.
“when we looked at g p d three, the cost fell 1,200 x from g p d three's initial cost to what you can get Lama 3.23 b today.”
Patel argues the surprise was a Chinese company achieving this cost reduction, not the reduction itself.
“I think what was really surprising was that it was a Chinese company for the first time. Right? Because Google and and OpenAI and Anthropic and Meta have all traded blows.”
Patel says DeepSeek's cost efficiency follows the expected trend line, just from an unexpected source.
“It's not unexpected. Right? Like, this is actually within the trend line of what happened with GPT three is happening to GPT four level quality with DeepSeq.”
Patel reveals DeepSeek inference implementation requires 160 GPUs worth over $10 million of hardware per replica.
“That's over $10,000,000 of hardware, and then that's just one replica, then you'll have a lot of replicas and you share the caching servers between them.”
Patel cites DeepSeek's open-sourced inference system requiring 140 GPUs communicating over RDMA networks for single replica.
“One example is one that DeepSeek open sourced over December of last yearJanuary, February of this year, where a single replica of inference for a single model is going to be like 140 GPUs.”
Gerstner says Chinese AI stack including DeepSeek and other models is making disturbing inroads in Middle East and Southern Hemisphere.
“No. For sure. And in fact, I I hear disturbing things all the time about the Chinese stack making great inroads in The Middle East and in the Southern Hemisphere, etcetera, because, frankly, they're working really quickly to put their chips and their models, their open source models, Deepsea, Kimi k two, Quinn, etcetera, and exporting those to the world.”
Patel claims hardware improved 30x from Hopper to Blackwell for DeepSeek on optimized deployments.
“from Hopper to Blackwell, which is all we've had over the last three years, roughly 30x improvement on DeepSeek, on the most optimized deployment, which you can see on InferenceX there's about a 30x improvement.”