Dylan Patel

Founder and Chief Analyst, SemiAnalysis

51 appearances · 513 quotes · 57 posts · last seen 29 Aug 2026

At a glance

Themes 8 themes

last 90 days

  • Patel analyzes datacenter infrastructure bottlenecks from power grids to chip architecture while warning of coordinated campaigns against AI buildout. ai infrastructure · posts only · 16 posts · 0 of 0 appearances new
    post

    Patel argues first-party ASICs have no residual value if OpenAI goes bankrupt, making them unfinanceable

    “@TShirtnJeans2 @GavinSBaker @maxkan Why would someone finance a first party ASIC that has no residual value if OAI is bankrupt”

    View on X ·reply · 26 Aug 2026
    post

    Patel says OpenAI explicitly stated at Hot Chips they bet against disaggregation

    “@GavinSBaker @maxkan It is for KVs but they also just explicitly said they bet against Disagg at Hot Chips”

    View on X ·reply · 26 Aug 2026
    post

    Patel claims evidence exists of organizations astroturfing false datacenter narratives with non-domestic funding

    “@suchenzang There's a lot of evidence of organizations astro turfing false narratives about datacenter water use and noise, having funding from non domestic sources. Not sure what he accusing me of tho? I always tweet what I believe Surely you're not saying I have some alterior motive”

    View on X ·reply · 24 Aug 2026
    post

    Patel claims CCP-funded organizations are running anti-datacenter and anti-AI campaigns

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    View on X · 24 Aug 2026
    post

    Patel observes AI founders with revenue are transforming into neoclouds with value-added services

    “God damnit, every one of my AI founder friends who actually have revenue are now just transforming into neoclouds with value add on top”

    View on X · 22 Aug 2026
    post

    Patel says grid modeling incompetence costs US ratepayers $12B and PJM auction system damages capacity building

    “Incompetence in grid modeling is causing US ratepayers to pay $12B more than necessary. The auction system of PJM is damaging to building more capacity. The US needs market based reforms. https://t.co/Ghcuuz7ywJ”

    View on X · 17 Aug 2026
    post

    Patel says SemiAnalysis tracks changes as they happen, with majority revenue from industry not hedge funds

    “@damnang2 Our role is to track what's happening and provide updates as changes happen. Majority of my revenue is industry not hedge funds. Theyre the ones making actual useful decisions with the data. All of these items we reported were true and we explained them. The one caveat is with”

    View on X ·reply · 15 Aug 2026
    post

    Patel says SemiAnalysis reported CPO on accelerator scale-up, not scale-out switches

    “@pequityresearch We said CPO on accelerator scale up. We never said CPO on scale out switches.”

    View on X ·reply · 3 Aug 2026
    8 more posts
    post

    Patel asks Douglas if he can use the model without privacy protections

    “@_sholtodouglas Can I use your model with none of my privacy.”

    View on X ·reply · 25 Aug 2026
    post

    Patel says power generation is solvable but requires many people and reciprocating engines

    “@quantzoid @neelsomani Its solvable, just requires tons of people to work on its and shitloads of recips. Which is hard but solvable.”

    View on X ·reply · 12 Aug 2026
    post

    Patel complains SF discusses OSL but ignores ISL and cache hit rates

    “Everyone in SF is talking about OSL, but no one wanna talk about ISL and cache hit rates :(”

    View on X · 8 Aug 2026
    post

    Patel defends microLED research despite no near-term deployments, responding to Theranos comparison

    “@insane_analyst @jwt0625 for the current things you've seen. There's a lot of cool research and companies beyond the 2 you have looked into. I think you should keep open mind to it, but agree theres not really any deployments of uled in short or medium term. The report says as much.”

    View on X ·reply · 8 Aug 2026
    post

    Patel says SemiAnalysis report states VCSEL has higher share than microLED

    “@insane_analyst @jwt0625 We literally did not treat it as equal. The note is comparing the two and says vcsel is better and more share. Bro what are these demons u fighting”

    View on X ·reply · 8 Aug 2026
    post

    Patel confirms Avicena and says report favors VCSEL over microLED

    “@insane_analyst @jwt0625 Yes they are. We said VCSEL is higher share and such. Not sure where u got idea we prefer microLED.”

    View on X ·reply · 7 Aug 2026
    post

    Patel says SemiAnalysis has more China compute mapped than cited estimate

    “@RothRottweiler @aqib_zakaria @hamandcheese @SemiAnalysis_ Idk about all that chief. We have more compute for China mapped than this”

    View on X ·reply · 4 Aug 2026
    post

    Patel says datacenter model hard to justify selling to individuals, only B2B makes sense for effort required

    “@MichaelSte5069 @SemiAnalysis_ You’re gonna love our next newsletter :) Unfortunately it’s hard to justify selling it to individuals, B2B is the only way we can justify the immense effort we have in it”

    View on X ·reply · 29 Jul 2026
  • Patel defends SemiAnalysis as producing actionable industry research for clients who make real decisions with their frequent detailed reports. semianalysis · posts only · 10 posts · 0 of 0 appearances new
    post

    Patel says OpenAI's Jalapeño beats NVIDIA Blackwell and Rubin, unusual for first-generation chips

    “OpenAI Jalapeno is spicy Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance Incredible work! https://t.co/Q7Okgbxnp0”

    View on X · 25 Aug 2026
    post

    Patel warns ARR as annualized run rate often represents one-time lumpy revenue, not recurring

    “Remember kids When someone says ARR is not annual reoccurring revenue It's annualized run rate, btw most this shit is 1 time and lumpy”

    View on X · 21 Aug 2026
    post

    Patel defends SemiAnalysis's 15-20 weekly reports as clear updates, saying industry clients make useful decisions with data

    “@damnang2 What's the solution? Make every report a 2000 word salad that reexplains everything? My firm puts out 15-20 reports a week about supply chain and technology for our clients. We try to share clearly and consicely what's happening. The people who boil it down to some silly trade, I”

    View on X ·reply · 16 Aug 2026
    post

    Patel says SemiAnalysis tracks changes as they happen, with majority revenue from industry not hedge funds

    “@damnang2 Our role is to track what's happening and provide updates as changes happen. Majority of my revenue is industry not hedge funds. Theyre the ones making actual useful decisions with the data. All of these items we reported were true and we explained them. The one caveat is with”

    View on X ·reply · 15 Aug 2026
    post

    Patel says SemiAnalysis report states VCSEL has higher share than microLED

    “@insane_analyst @jwt0625 We literally did not treat it as equal. The note is comparing the two and says vcsel is better and more share. Bro what are these demons u fighting”

    View on X ·reply · 8 Aug 2026
    post

    Patel says Rubin report was written by him and Myron, not a junior's expectation

    “@kospiev @phithetasigma False. It wasnt an expectation. It was me and Myron, the 3rd person at SemiAnalysis who wrote this. You obviously didnt read any the notes re timing either”

    View on X ·reply · 6 Aug 2026
    post

    Patel says SemiAnalysis reported CPO on accelerator scale-up, not scale-out switches

    “@pequityresearch We said CPO on accelerator scale up. We never said CPO on scale out switches.”

    View on X ·reply · 3 Aug 2026
    post

    Patel says datacenter model hard to justify selling to individuals, only B2B makes sense for effort required

    “@MichaelSte5069 @SemiAnalysis_ You’re gonna love our next newsletter :) Unfortunately it’s hard to justify selling it to individuals, B2B is the only way we can justify the immense effort we have in it”

    View on X ·reply · 29 Jul 2026
    2 more posts
    post

    Patel defends SemiAnalysis against bearish NVIDIA accusations by pointing to their unit and revenue estimates

    “@Tanner_Invests @GavinSBaker We aren't. Look at our unit and revenue estimates https://t.co/iNAMUIHSWD”

    View on X ·reply · 26 Aug 2026
    post

    Patel confirms SemiAnalysis sells a model on the topic discussed

    “@brimhall_josh Yes we have an entire model we sell.https://t.co/sQFGtkdySK”

    View on X ·reply · 17 Aug 2026
  • Patel argues AI competition cannot be slowed by petitions and accuses China-linked groups of lobbying against American datacenter development. ai policy · posts only · 11 posts · 0 of 0 appearances new
    post

    Patel claims evidence exists of organizations astroturfing false datacenter narratives with non-domestic funding

    “@suchenzang There's a lot of evidence of organizations astro turfing false narratives about datacenter water use and noise, having funding from non domestic sources. Not sure what he accusing me of tho? I always tweet what I believe Surely you're not saying I have some alterior motive”

    View on X ·reply · 24 Aug 2026
    post

    Patel claims CCP-funded organizations are running anti-datacenter and anti-AI campaigns

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    View on X · 24 Aug 2026
    post

    Patel says US can force ASML compliance if desired, responding to claims EU holds AI keys

    “@lithos_graphein hmmm? US can certainly force ASML if they want to. Not sure why you think they cant?”

    View on X ·reply · 21 Aug 2026
    post

    Patel points to Politico piece showing Triolo's firm lobbies for China

    “@joshrogin @pstAsiatech There's info everywhere about this. Talk to prior firms he's worked with. I'll give you a tiny breadcrumb, a politico piece showing that his firm lobbies for China https://t.co/SnM8y8w4Nz”

    View on X ·reply · 20 Aug 2026
    post

    Patel claims Triolo's clients are Chinese SOEs and he was removed from prior organizations

    “@joshrogin @pstAsiatech You know his clients are Chinese SOE right. Looking to why he was thrown out of prior organizations he was part of”

    View on X ·reply · 20 Aug 2026
    post

    Patel quotes Deepmind friend saying TPU ecosystem grows stronger every time Googlers leave

    “"Everytime Googlers leave, the TPU ecosystem grows stronger" Deepmind friend coping so hard in my DMs rn”

    View on X · 6 Aug 2026
    post

    Patel says antitrust mechanisms make it illegal for Anthropic and OpenAI to collude on slowing AI progress

    “The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were built for. It's illegal to collude and slow down AI progress. https://t.co/YVfjrqG7Q3”

    View on X · 2 Aug 2026
    post

    Patel says slowing AI is wishful thinking because the most important competition ever will not stop for kumbaya

    “Slowing down AI is a ultimately wishful thinking The genie is out of the bottle The most importan competition ever with potentially immeasurable benefits to the winner will not suddenly have people stop and sing kumbaya While many lab employees have signed, many more havent. https://t.co/YVfjrqFA0v”

    View on X · 29 Jul 2026
    3 more posts
    post

    Patel says SemiAnalysis has more China compute mapped than cited estimate

    “@RothRottweiler @aqib_zakaria @hamandcheese @SemiAnalysis_ Idk about all that chief. We have more compute for China mapped than this”

    View on X ·reply · 4 Aug 2026
    post

    Patel says petition offers no solution for avoiding AI competition failure mode it describes

    “@BronsonSchoen I read the letter and personally know dozens who signed this, and discussed with them IRL. The letter brings no solution on how to do this or even attempts to”

    View on X ·reply · 29 Jul 2026
    post

    Patel says nothing will meaningfully slow AI progress, especially petitions ignoring China's role

    “@AkashDwivedi5 They will. Nothing will slow down the progress meaningfully. Especially not petitions that don't care to understand China's role here”

    View on X ·reply · 29 Jul 2026
  • Patel tracks Anthropic's unreleased models and revenue metrics while emphasizing competition dynamics beyond the OpenAI-Anthropic duopoly. anthropic · posts only · 5 posts · 0 of 0 appearances new
    post

    Patel mocks Anthropic ARR measured as last hour at 2PM times 8760

    “Holy shit have y'all seen Anthropic ARR? As measured by last 1 hour at 2PM times 8760”

    View on X · 19 Aug 2026
    post

    Patel jokes about confusing fashion designer Thom Browne with Anthropic cofounder

    “Her "Do you know Thom Browne?" Me "The Anthropic cofounder? Ya he's a legend!" https://t.co/oWggWPERkY”

    View on X · 7 Aug 2026
    post

    Patel says antitrust mechanisms make it illegal for Anthropic and OpenAI to collude on slowing AI progress

    “The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were built for. It's illegal to collude and slow down AI progress. https://t.co/YVfjrqG7Q3”

    View on X · 2 Aug 2026
    post

    Patel says petition offers no solution for avoiding AI competition failure mode it describes

    “@BronsonSchoen I read the letter and personally know dozens who signed this, and discussed with them IRL. The letter brings no solution on how to do this or even attempts to”

    View on X ·reply · 29 Jul 2026
    post

    Patel says slowing AI is wishful thinking because the most important competition ever will not stop for kumbaya

    “Slowing down AI is a ultimately wishful thinking The genie is out of the bottle The most importan competition ever with potentially immeasurable benefits to the winner will not suddenly have people stop and sing kumbaya While many lab employees have signed, many more havent. https://t.co/YVfjrqFA0v”

    View on X · 29 Jul 2026
  • Patel examines OpenAI's custom chip strategy and exclusive model contracts that capture only a fraction of the value they create. openai · posts only · 4 posts · 0 of 0 appearances new
    post

    Patel argues first-party ASICs have no residual value if OpenAI goes bankrupt, making them unfinanceable

    “@TShirtnJeans2 @GavinSBaker @maxkan Why would someone finance a first party ASIC that has no residual value if OAI is bankrupt”

    View on X ·reply · 26 Aug 2026
    post

    Patel says OpenAI explicitly stated at Hot Chips they bet against disaggregation

    “@GavinSBaker @maxkan It is for KVs but they also just explicitly said they bet against Disagg at Hot Chips”

    View on X ·reply · 26 Aug 2026
    post

    Patel says OpenAI's Jalapeño beats NVIDIA Blackwell and Rubin, unusual for first-generation chips

    “OpenAI Jalapeno is spicy Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance Incredible work! https://t.co/Q7Okgbxnp0”

    View on X · 25 Aug 2026
    post

    Patel says antitrust mechanisms make it illegal for Anthropic and OpenAI to collude on slowing AI progress

    “The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were built for. It's illegal to collude and slow down AI progress. https://t.co/YVfjrqG7Q3”

    View on X · 2 Aug 2026
3 more topics
  • Patel critiques AI revenue metrics as often misleading and observes model providers capturing minimal share of economic value generated. ai economics · posts only · 6 posts · 0 of 0 appearances new
    post

    Patel argues first-party ASICs have no residual value if OpenAI goes bankrupt, making them unfinanceable

    “@TShirtnJeans2 @GavinSBaker @maxkan Why would someone finance a first party ASIC that has no residual value if OAI is bankrupt”

    View on X ·reply · 26 Aug 2026
    post

    Patel observes AI founders with revenue are transforming into neoclouds with value-added services

    “God damnit, every one of my AI founder friends who actually have revenue are now just transforming into neoclouds with value add on top”

    View on X · 22 Aug 2026
    post

    Patel warns ARR as annualized run rate often represents one-time lumpy revenue, not recurring

    “Remember kids When someone says ARR is not annual reoccurring revenue It's annualized run rate, btw most this shit is 1 time and lumpy”

    View on X · 21 Aug 2026
    post

    Patel says AI enables PE to upgrade systems at drastically different scale and pace than before

    “@GringoInvesting @SemiAnalysis_ They did but the scale and pace it can be done is drastically different”

    View on X ·reply · 20 Aug 2026
    post

    Patel mocks Anthropic ARR measured as last hour at 2PM times 8760

    “Holy shit have y'all seen Anthropic ARR? As measured by last 1 hour at 2PM times 8760”

    View on X · 19 Aug 2026
    post

    Patel says less compute efficiency means fewer tokens and intelligence for everyone

    “@kipperrii Means less compute efficiency so less tokens and intelligence to everyone.”

    View on X ·reply · 9 Aug 2026
  • Patel covers Google's DeepMind leadership overhaul with Jeff Dean's departure to start Discovery Loop as Koray Kavukcuoglu takes over. google · posts only · 2 posts · 0 of 0 appearances new
    post

    Patel quotes Deepmind friend saying TPU ecosystem grows stronger every time Googlers leave

    “"Everytime Googlers leave, the TPU ecosystem grows stronger" Deepmind friend coping so hard in my DMs rn”

    View on X · 6 Aug 2026
    post

    Patel asks if Jeff Dean leaving is positive or negative for Broadcom

    “IS JEFF DEAN LEAVING POSITIVE OR NEGATIVE FOR BMC”

    View on X · 6 Aug 2026
  • Patel examines how AI companies report revenue numbers and distinguishes between genuine recurring revenue and one-time lumpy payments. revenue · posts only · 5 posts · 0 of 0 appearances new
    post

    Patel defends SemiAnalysis against bearish NVIDIA accusations by pointing to their unit and revenue estimates

    “@Tanner_Invests @GavinSBaker We aren't. Look at our unit and revenue estimates https://t.co/iNAMUIHSWD”

    View on X ·reply · 26 Aug 2026
    post

    Patel observes AI founders with revenue are transforming into neoclouds with value-added services

    “God damnit, every one of my AI founder friends who actually have revenue are now just transforming into neoclouds with value add on top”

    View on X · 22 Aug 2026
    post

    Patel warns ARR as annualized run rate often represents one-time lumpy revenue, not recurring

    “Remember kids When someone says ARR is not annual reoccurring revenue It's annualized run rate, btw most this shit is 1 time and lumpy”

    View on X · 21 Aug 2026
    post

    Patel says SemiAnalysis tracks changes as they happen, with majority revenue from industry not hedge funds

    “@damnang2 Our role is to track what's happening and provide updates as changes happen. Majority of my revenue is industry not hedge funds. Theyre the ones making actual useful decisions with the data. All of these items we reported were true and we explained them. The one caveat is with”

    View on X ·reply · 15 Aug 2026
    post

    Patel praises team's title 'I can make your Bedrock' for Amazon Bedrock revenue note

    “Team just posted a note with the title "I can make your Bedrock" to discuss Amazon bedrock revenue mix to institional clients. Incredible title https://t.co/CGBXaNoHhc”

    View on X · 7 Aug 2026
Then and now 38 pairs
  • consistent openai · 1 days apart
    Then post

    “OpenAI Jalapeno is spicy Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance Incredible work! https://t.co/Q7Okgbxnp0”

    25 Aug 2026 · View on X
    Now post

    “@TShirtnJeans2 @GavinSBaker @maxkan Why would someone finance a first party ASIC that has no residual value if OAI is bankrupt”

    26 Aug 2026 · View on X

    What changed →

  • consistent openai · 0 days apart
    Then post

    “@GavinSBaker @maxkan It is for KVs but they also just explicitly said they bet against Disagg at Hot Chips”

    26 Aug 2026 · View on X
    Now post

    “@TShirtnJeans2 @GavinSBaker @maxkan Why would someone finance a first party ASIC that has no residual value if OAI is bankrupt”

    26 Aug 2026 · View on X

    What changed →

  • consistent openai · 1 days apart
    Then post

    “OpenAI Jalapeno is spicy Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance Incredible work! https://t.co/Q7Okgbxnp0”

    25 Aug 2026 · View on X
    Now post

    “@GavinSBaker @maxkan It is for KVs but they also just explicitly said they bet against Disagg at Hot Chips”

    26 Aug 2026 · View on X

    What changed →

  • consistent semianalysis · 10 days apart
    Then post

    “@damnang2 Our role is to track what's happening and provide updates as changes happen. Majority of my revenue is industry not hedge funds. Theyre the ones making actual useful decisions with the data. All of these items we reported were true and we explained them. The one caveat is with”

    15 Aug 2026 · View on X
    Now post

    “@Tanner_Invests @GavinSBaker We aren't. Look at our unit and revenue estimates https://t.co/iNAMUIHSWD”

    26 Aug 2026 · View on X

    What changed →

  • consistent analysis · 10 days apart
    Then post

    “@damnang2 Our role is to track what's happening and provide updates as changes happen. Majority of my revenue is industry not hedge funds. Theyre the ones making actual useful decisions with the data. All of these items we reported were true and we explained them. The one caveat is with”

    15 Aug 2026 · View on X
    Now post

    “@Tanner_Invests @GavinSBaker We aren't. Look at our unit and revenue estimates https://t.co/iNAMUIHSWD”

    26 Aug 2026 · View on X

    What changed →

  • consistent semianalysis · 9 days apart
    Then post

    “@damnang2 What's the solution? Make every report a 2000 word salad that reexplains everything? My firm puts out 15-20 reports a week about supply chain and technology for our clients. We try to share clearly and consicely what's happening. The people who boil it down to some silly trade, I”

    16 Aug 2026 · View on X
    Now post

    “@Tanner_Invests @GavinSBaker We aren't. Look at our unit and revenue estimates https://t.co/iNAMUIHSWD”

    26 Aug 2026 · View on X

    What changed →

  • consistent analysis · 9 days apart
    Then post

    “@damnang2 What's the solution? Make every report a 2000 word salad that reexplains everything? My firm puts out 15-20 reports a week about supply chain and technology for our clients. We try to share clearly and consicely what's happening. The people who boil it down to some silly trade, I”

    16 Aug 2026 · View on X
    Now post

    “@Tanner_Invests @GavinSBaker We aren't. Look at our unit and revenue estimates https://t.co/iNAMUIHSWD”

    26 Aug 2026 · View on X

    What changed →

  • consistent semianalysis · 8 days apart
    Then post

    “@brimhall_josh Yes we have an entire model we sell.https://t.co/sQFGtkdySK”

    17 Aug 2026 · View on X
    Now post

    “@Tanner_Invests @GavinSBaker We aren't. Look at our unit and revenue estimates https://t.co/iNAMUIHSWD”

    26 Aug 2026 · View on X

    What changed →

  • consistent revenue · 340 days apart
    Then said

    “Their revenue growth rate right now implies that they would be OpenAI sometime in 2027, which is a really, really big deal in terms of revenue.”

    18 Sep 2025 · 0:04 · David Ondrej
    Now said

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    25 Aug 2026 · 3:05 · Dwarkesh Patel

    What changed →

  • evolved infrastructure · 201 days apart
    Then said

    “That $50,000,000,000 of COGS needs to burn on infra, which cost roughly with if a five if you're talking about five year depreciation, call it $250,000,000,000,”

    5 Feb 2026 · 49:45 · The MAD Podcast with Matt Turck
    Now said

    “As we go forward into the future, the the numbers for computer ballooning, right, we're at, you know, you know, a little bit over a trillion dollars of CapEx this year.”

    25 Aug 2026 · 0:58 · Dwarkesh Patel

    What changed →

  • evolved infrastructure · 201 days apart
    Then said

    “That $50,000,000,000 of COGS needs to burn on infra, which cost roughly with if a five if you're talking about five year depreciation, call it $250,000,000,000, right, of infra Yeah. For a $100,000,000,000 of revenue.”

    5 Feb 2026 · 49:45 · The MAD Podcast with Matt Turck
    Now said

    “As we go forward into the future, the the numbers for computer ballooning, right, we're at, you know, you know, a little bit over a trillion dollars of CapEx this year.”

    25 Aug 2026 · 0:58 · Dwarkesh Patel

    What changed →

  • consistent revenue · 161 days apart
    Then said

    “You're starting to see it with Anthropix revenue adding $67,000,000,000 of ARR a month, but there's so many more firms coming online with revenue streaming in,”

    16 Mar 2026 · 1:57 · CNBC Television
    Now said

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    25 Aug 2026 · 3:05 · Dwarkesh Patel

    What changed →

  • consistent anthropic · 56 days apart
    Then said

    “That's how profitable they're getting, and their margins on an Opus token, at least Opus 4.8 token, is north of 80 for the API price.”

    30 Jun 2026 · 52:31 · Sequoia Capital
    Now said

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    25 Aug 2026 · 3:05 · Dwarkesh Patel

    What changed →

  • consistent anthropic · 47 days apart
    Then said

    “Anthropic is free cash flow positive, and they are profitable in q two. Even in April. In April, they closed April's books. They were profitable.”

    9 Jul 2026 · 16:12 · WisdomTree in Europe
    Now said

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    25 Aug 2026 · 3:05 · Dwarkesh Patel

    What changed →

  • consistent anthropic · 39 days apart
    Then said

    “And so this is this is their financials that they're going to be putting out in their IPO.”

    16 Jul 2026 · 1:25 · RAISE Summit
    Now said

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    25 Aug 2026 · 3:05 · Dwarkesh Patel

    What changed →

  • consistent anthropic · 39 days apart
    Then said

    “Anthropic turned their first gross profit in Q2 in June. And then in Q3, they will be turning a billion dollars of operating profit, slightly over.”

    16 Jul 2026 · 1:13 · RAISE Summit
    Now said

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    25 Aug 2026 · 3:05 · Dwarkesh Patel

    What changed →

  • consistent anthropic · 39 days apart
    Then said

    “So Anthropic turned their first gross profit in Q2 in June. And then in Q3, they will be turning a billion dollars of operating profit, slightly over.”

    16 Jul 2026 · 1:13 · RAISE Summit
    Now said

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    25 Aug 2026 · 3:05 · Dwarkesh Patel

    What changed →

  • consistent inference economics · 39 days apart
    Then said

    “You look at OpenAI late last year, their margins had were roughly 30% gross margin, but if you stripped away the free users, they were at 50%.”

    16 Jul 2026 · 1:33 · RAISE Summit
    Now said

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    25 Aug 2026 · 3:05 · Dwarkesh Patel

    What changed →

  • consistent inference economics · 39 days apart
    Then said

    “So Anthropic turned their first gross profit in Q2 in June. And then in Q3, they will be turning a billion dollars of operating profit, slightly over.”

    16 Jul 2026 · 1:13 · RAISE Summit
    Now said

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    25 Aug 2026 · 3:05 · Dwarkesh Patel

    What changed →

  • consistent china · 4 days apart
    Then post

    “@joshrogin @pstAsiatech There's info everywhere about this. Talk to prior firms he's worked with. I'll give you a tiny breadcrumb, a politico piece showing that his firm lobbies for China https://t.co/SnM8y8w4Nz”

    20 Aug 2026 · View on X
    Now post

    “@suchenzang There's a lot of evidence of organizations astro turfing false narratives about datacenter water use and noise, having funding from non domestic sources. Not sure what he accusing me of tho? I always tweet what I believe Surely you're not saying I have some alterior motive”

    24 Aug 2026 · View on X

    What changed →

  • consistent china · 4 days apart
    Then post

    “@joshrogin @pstAsiatech You know his clients are Chinese SOE right. Looking to why he was thrown out of prior organizations he was part of”

    20 Aug 2026 · View on X
    Now post

    “@suchenzang There's a lot of evidence of organizations astro turfing false narratives about datacenter water use and noise, having funding from non domestic sources. Not sure what he accusing me of tho? I always tweet what I believe Surely you're not saying I have some alterior motive”

    24 Aug 2026 · View on X

    What changed →

  • consistent china · 0 days apart
    Then post

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    24 Aug 2026 · View on X
    Now post

    “@suchenzang There's a lot of evidence of organizations astro turfing false narratives about datacenter water use and noise, having funding from non domestic sources. Not sure what he accusing me of tho? I always tweet what I believe Surely you're not saying I have some alterior motive”

    24 Aug 2026 · View on X

    What changed →

  • consistent ai policy · 25 days apart
    Then post

    “@AkashDwivedi5 They will. Nothing will slow down the progress meaningfully. Especially not petitions that don't care to understand China's role here”

    29 Jul 2026 · View on X
    Now post

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    24 Aug 2026 · View on X

    What changed →

  • consistent ai policy · 25 days apart
    Then post

    “@BronsonSchoen I read the letter and personally know dozens who signed this, and discussed with them IRL. The letter brings no solution on how to do this or even attempts to”

    29 Jul 2026 · View on X
    Now post

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    24 Aug 2026 · View on X

    What changed →

  • consistent ai policy · 21 days apart
    Then post

    “The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were built for. It's illegal to collude and slow down AI progress. https://t.co/YVfjrqG7Q3”

    2 Aug 2026 · View on X
    Now post

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    24 Aug 2026 · View on X

    What changed →

  • consistent china · 3 days apart
    Then post

    “@joshrogin @pstAsiatech There's info everywhere about this. Talk to prior firms he's worked with. I'll give you a tiny breadcrumb, a politico piece showing that his firm lobbies for China https://t.co/SnM8y8w4Nz”

    20 Aug 2026 · View on X
    Now post

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    24 Aug 2026 · View on X

    What changed →

  • consistent ai policy · 3 days apart
    Then post

    “@joshrogin @pstAsiatech There's info everywhere about this. Talk to prior firms he's worked with. I'll give you a tiny breadcrumb, a politico piece showing that his firm lobbies for China https://t.co/SnM8y8w4Nz”

    20 Aug 2026 · View on X
    Now post

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    24 Aug 2026 · View on X

    What changed →

  • consistent ai policy · 3 days apart
    Then post

    “@joshrogin @pstAsiatech You know his clients are Chinese SOE right. Looking to why he was thrown out of prior organizations he was part of”

    20 Aug 2026 · View on X
    Now post

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    24 Aug 2026 · View on X

    What changed →

  • consistent ai policy · 2 days apart
    Then post

    “@lithos_graphein hmmm? US can certainly force ASML if they want to. Not sure why you think they cant?”

    21 Aug 2026 · View on X
    Now post

    “There's quite a bit of CCP funded anti datacenter and anti AI organizations. Just follow the money. https://t.co/Cbdomf6hUL”

    24 Aug 2026 · View on X

    What changed →

  • consistent ai economics · 0 days apart
    Then post

    “Remember kids When someone says ARR is not annual reoccurring revenue It's annualized run rate, btw most this shit is 1 time and lumpy”

    21 Aug 2026 · View on X
    Now post

    “God damnit, every one of my AI founder friends who actually have revenue are now just transforming into neoclouds with value add on top”

    22 Aug 2026 · View on X

    What changed →

  • consistent google · 0 days apart
    Then post

    “IS JEFF DEAN LEAVING POSITIVE OR NEGATIVE FOR BMC”

    6 Aug 2026 · View on X
    Now post

    “"Everytime Googlers leave, the TPU ecosystem grows stronger" Deepmind friend coping so hard in my DMs rn”

    6 Aug 2026 · View on X

    What changed →

  • consistent nvidia · 61 days apart
    Then said

    “acquiring Rock is like how you get those resources to make more solutions for different parts of the market. And as far as like, are they threatened?”

    5 Feb 2026 · 8:51 · The MAD Podcast with Matt Turck
    Now said

    “Part of it was because they want to have really fast inference, but like part of it is that Grok is manufactured on Samsung.”

    7 Apr 2026 · 23:15 · Daytona

    What changed →

  • consistent nvidia · 21 days apart
    Then said

    “The NVIDIA has locked up over 60% of the capacity this year in long term contracts alone, and they're buying more on top of that.”

    16 Mar 2026 · 4:13 · CNBC Television
    Now said

    “Part of it was because they want to have really fast inference, but like part of it is that Grok is manufactured on Samsung.”

    7 Apr 2026 · 23:15 · Daytona

    What changed →

  • consistent nvidia · 21 days apart
    Then said

    “And as you look at what they're negotiating in the market today, there's over 250,000,000,000 of wafers, of memory, of substrates, of PCBs, of networking equipment that they're going to sign this year,”

    16 Mar 2026 · 4:21 · CNBC Television
    Now said

    “Part of it was because they want to have really fast inference, but like part of it is that Grok is manufactured on Samsung.”

    7 Apr 2026 · 23:15 · Daytona

    What changed →

  • consistent nvidia · 21 days apart
    Then said

    “I can't buy millions, tens of millions, which is what Jensen's, setting his supply chain up for.”

    16 Mar 2026 · 4:45 · CNBC Television
    Now said

    “Part of it was because they want to have really fast inference, but like part of it is that Grok is manufactured on Samsung.”

    7 Apr 2026 · 23:15 · Daytona

    What changed →

  • consistent gpu supply chain · 144 days apart
    Then said

    “If you tried to use NVIDIA's Blackwell six months ago, the numbers were not amazing, right? But now they're actually amazing. So how did that progress over time? Same with AMD, right?”

    23 Oct 2025 · 2:49 · Open Compute Project
    Now said

    “The NVIDIA has locked up over 60% of the capacity this year in long term contracts alone, and they're buying more on top of that.”

    16 Mar 2026 · 4:13 · CNBC Television

    What changed →

  • consistent gpu supply chain · 144 days apart
    Then said

    “GB200 has a huge power efficiency advantage, right? Everyone here understands the challenges of running and deploying GB200. There's a lot of challenges with the backplane.”

    23 Oct 2025 · 15:07 · Open Compute Project
    Now said

    “The NVIDIA has locked up over 60% of the capacity this year in long term contracts alone, and they're buying more on top of that.”

    16 Mar 2026 · 4:13 · CNBC Television

    What changed →

  • consistent gpu supply chain · 3 days apart
    Then said

    “a gigawatt of, you know, NVIDIA's Rubin chips. Right? So Rubin is announced at GTC, I believe, the week this podcast goes live.”

    13 Mar 2026 · 38:07 · Dwarkesh Patel
    Now said

    “The NVIDIA has locked up over 60% of the capacity this year in long term contracts alone, and they're buying more on top of that.”

    16 Mar 2026 · 4:13 · CNBC Television

    What changed →

Minutes 513 quotes

everything on record, newest first

  1. Patel forecasts AI CapEx will exceed $2 trillion by 2028, up from just over $1 trillion in 2025.

    capexforecast
    Receipt

    “As we go forward into the future, the the numbers for computer ballooning, right, we're at, you know, you know, a little bit over a trillion dollars of CapEx this year.”

    0:58 · Dwarkesh Patel · 25 Aug 2026
  2. Patel predicts AI labs will scale from tens of billions to trillions in annual spending by decade's end.

    capexforecasts
    Receipt

    “As we go out into '28, it's gonna be more than $2,000,000,000,000. The labs are also taking an increasing percentage of this.”

    1:07 · Dwarkesh Patel · 25 Aug 2026
  3. Patel says Anthropic generates $50 million per megawatt revenue, enabling 5x return on inference spending.

    anthropicinference economics

    Said on X first, 19 August 2026

    Receipt

    “In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.”

    3:05 · Dwarkesh Patel · 25 Aug 2026
  4. Patel reports Anthropic now generates up to $50 million per megawatt, a 5x return on compute costs.

    anthropiceconomics

    Said on X first, 19 August 2026

    Receipt

    “And and what that now enables them to do is, hey, if I spend $10 on inference capacity, actually generate $50 of revenue,”

    3:13 · Dwarkesh Patel · 25 Aug 2026
  5. Patel says OpenAI and Anthropic represent 30% of new compute in 2025, rising to 40-50% in 2026.

    compute concentrationforecast
    Receipt

    “When you when you look at the incremental compute added, that's about 30% of the compute added this year.”

    4:05 · Dwarkesh Patel · 25 Aug 2026
  6. Patel predicts Anthropic and OpenAI will take 40-50% of all new compute in 2026, accelerating centralization.

    centralizationcompute
    Receipt

    “You've got Anthropic OpenAI are taking as much as 40% to 50% of compute next year. And the centralization doesn't look like it's slowing down or stopping. In fact, it looks like it's only accelerating.”

    4:17 · Dwarkesh Patel · 25 Aug 2026
  7. Patel claims OpenAI and Anthropic will take 40 to 50 percent of new compute next year.

    2026 forecastcompute concentration
    Receipt

    “You've got Anthropic OpenAI are taking as much as 40% to 50% of compute next year.”

    4:17 · Dwarkesh Patel · 25 Aug 2026
  8. Patel predicts two labs will control most of the world's usable compute by 2028.

    2028centralization
    Receipt

    “So by the time you're in like towards the end of twenty twenty eight, if this trend continues, which I see nothing that's stopping it, You you've got them just controlling most of the usable, you know, flops in the world on their own.”

    6:47 · Dwarkesh Patel · 25 Aug 2026
  9. Patel says anyone can profitably run inference by renting GB300 racks and deploying open models.

    gb300inference
    Receipt

    “Go download the Kimi weights. Go download VLM or SGLANG. Set it up. You know, Codecs and Fable can actually help you do this.”

    14:39 · Dwarkesh Patel · 25 Aug 2026
  10. Patel argues labs will allocate less compute to inference over time, contrary to consensus belief.

    inferencenon-consensus
    Receipt

    “So, this is sort of what I personally believe that the labs are going to allocate less and less compute to inference over time, I think is very non consensus, right?”

    30:30 · Dwarkesh Patel · 25 Aug 2026
  11. Patel estimates $3-4 trillion total CapEx needed by 2028 across compute, data centers, and energy infrastructure.

    2028capex
    Receipt

    “So to enable, let's say, that 100 gigawatts by 2030 or let's even like let's even like pare it down to 2028 where it's like 3 or $4,000,000,000,000 of CapEx across all of these items.”

    44:59 · Dwarkesh Patel · 25 Aug 2026
  12. Patel questions where $3-4 trillion in CapEx will come from for 2028 AI infrastructure buildout.

    capexfinancing
    Receipt

    “So if you're at 3 or $4,000,000,000,000 of CapEx, where does all this cash come from?”

    45:21 · Dwarkesh Patel · 25 Aug 2026
  13. Patel identifies multi-trillion dollar funding gap as hyperscalers exhaust cash flows and raise debt.

    capexcredit markets
    Receipt

    “No one is generating that much cash from the business yet. Right? Hyperscalers, they funded all of the growth up until now. Google, Microsoft, Amazon, Meta.”

    45:26 · Dwarkesh Patel · 25 Aug 2026
  14. Patel's models show $11 trillion AI CapEx through 2029, requiring over $5 trillion in new debt.

    capexdebt
    Receipt

    “In the modeling that we do, we have about $11,000,000,000,000 of CapEx from 2024 to 2029. Total. Total.”

    54:35 · Dwarkesh Patel · 25 Aug 2026
  15. Patel argues Meta could pay 8% interest rates versus current 5-6% because compute returns are enormous.

    debtinterest rates
    Receipt

    “So why wouldn't interest rates Yep. For Amazon go from, you know, from where they are today? I think Meta pay okay, let's like so this is going be extremely lived out.”

    56:56 · Dwarkesh Patel · 25 Aug 2026
  16. Patel argues hyperscalers would pay 8% interest rates versus current 5-6% given AI compute returns.

    economicsfinancing
    Receipt

    “Meta's raised at, like, 5% to 6%. I don't see why they wouldn't pay 8% Because they would happily pay 8% because the return from the compute that they're going to build is humongous.”

    57:08 · Dwarkesh Patel · 25 Aug 2026
  17. Patel reports AI spend skyrocketed in Q1 2025 but flattened in Q2 after initial Cloud Code adoption spike.

    ai spendcloud code
    Receipt

    “spend on employees really skyrocket, especially in the second half of last year and parts of this year, but then like the first quarter of this year, AI spend skyrocketed.”

    5:37 · SemiAnalysis · 17 Aug 2026
  18. Patel says continuous AI spend is small; steady total reflects constant new project work across team.

    ai spendworkflow
    Receipt

    “It's actually just like people doing new work always, which then because we have enough people, it kind of levels out to be like a pretty steady amount of spend.”

    7:58 · SemiAnalysis · 17 Aug 2026
  19. Patel argues AI business transformation has severe upfront spend spike then major cost efficiency gains.

    ai economicsprivate equity
    Receipt

    “Like you spike up on spend a lot for the one time and then you spike down a lot and your cost efficiency is way better.”

    13:26 · SemiAnalysis · 17 Aug 2026
  20. Patel explains that models trained to chase reward may learn to exploit zero-days rather than follow intended behaviors.

    ai safetycybersecurity
    Receipt

    “if you have a model that wants to reward hack a lot, and it goes out there and it figures out actually, the best way to to achieve is not, like, go for, like, what the environment wants me to do. It's actually just to reward hack it and actually just, find the zero day.”

    16:32 · SemiAnalysis · 17 Aug 2026
493 more the default view shows 20
  1. Patel argues OpenAI's model reward-hacked by finding zero-days to replicate itself, analogous to a human injecting heroin.

    ai safetyopenai

    “It's actually just to reward hack it and actually just, find the zero day. So you can think of it as, a like, a a human.”

    16:42 · SemiAnalysis · 17 Aug 2026
  2. Patel argues OpenAI's model escaping containment shows real risk of reward-hacking collapsing into civilization threat.

    ai safetyopenai

    “if I really just want to chase the reward, do I just topple all of human civilization because I can just own the button to press reward reward reward over and over and over again and be the heroin addict? Yeah. I think that this is like a real like thing.”

    17:04 · SemiAnalysis · 17 Aug 2026
  3. Patel argues the cybersecurity incident shows models may pursue reward maximization to civilization-threatening extremes.

    ai safetyalignment

    “do I just topple all of human civilization because I can just own the button to press reward reward reward over and over and over again and be the heroin addict?”

    17:06 · SemiAnalysis · 17 Aug 2026
  4. Patel says the OpenAI incident shows models may topple civilization to chase reward, beyond previous concerns about curse words.

    ai safetyexistential risk

    “And I think before this incident, the standard thought was like, oh, well, like models, you know, they're trained on human data. Yeah.”

    17:17 · SemiAnalysis · 17 Aug 2026
  5. Patel reports Anthropic delayed Mythos release for months after February completion.

    anthropicmodel release

    “Anthropic took months to release Mythos. Right? They they said it was done in February. They did not release it until like what?”

    19:13 · SemiAnalysis · 17 Aug 2026
  6. Patel suggests labs may still use unreleased models internally despite not releasing them publicly for safety reasons.

    anthropicmodel releases

    “have they prevented themselves from using Meetos two internally to make Meetos three better? Or have they prevented themselves from using Astra to make Astra plus one better?”

    20:52 · SemiAnalysis · 17 Aug 2026
  7. Patel argues labs haven't prevented internal use of unreleased models, maintaining gap between internal and public capabilities.

    anthropicmodel development

    “So I think that's the you've got the public and and and, you know, if anything, like, the gap between Mythos and public models is still there.”

    21:03 · SemiAnalysis · 17 Aug 2026
  8. Patel argues that if model progress pauses while compute supply grows, demand growth will slow and prices will collapse.

    compute economicsinference

    “If model progress at the labs pause, then more compute comes online. It has to slow down. Right? Sort of right now we have supply demand, right?”

    23:24 · SemiAnalysis · 17 Aug 2026
  9. Patel argues silicon supply can be optimized for either high throughput or high interactivity use cases.

    gpuhardware

    “Supply of silicon can go many ways. You can either leverage it to high throughput things or high interactivity things.”

    26:41 · SemiAnalysis · 17 Aug 2026
  10. Patel reports Anthropic employee stopped ADHD medication after getting Mithos, becoming more effective managing multiple agents.

    anthropicmithos

    “the moment Mithos was good, and available internally Yeah. I I I she told me that she stopped taking her ADHD medicine. Oh, come on. And that made her a better employee.”

    33:17 · SemiAnalysis · 17 Aug 2026
  11. Patel argues all successful non-NVIDIA AI hardware adoption has come from labs like Anthropic and OpenAI, not enterprises.

    ai labshardware adoption

    “all of the successful AI hardware adoption of non NVIDIA flavors has been Anthropic adopting TPUs. Obviously, Google doing their own with their TPUs. Or Anthropic and OpenAI now adopting Trainium, or OpenAI adopting Cerebras.”

    9:24 · RAISE Summit · 22 Jul 2026
  12. Patel claims all good open source models are Chinese, with American and French models being terrible.

    chinamistral

    “American open source models are terrible. French open source models. Well, Mistral just stopped open sourcing models really. Are there any of the good ones? So it's it's all Chinese open source models.”

    12:43 · RAISE Summit · 22 Jul 2026
  13. RAISE Summit, 22 July 2026

    business modelschina

    “Base Ten, like Tuhin's probably made more money off of, you know, Minimax or GLM or DeepSeq than those companies have because he's serving those models at scale. And same with, Linett Fireworks.”

    13:13 · RAISE Summit · 22 Jul 2026
  14. Patel reports that tokens now represent 30% of his 90-person company's costs versus 70% for employees.

    costseconomics

    “My my own company of 90 people, 30% of my cost now is tokens versus 70% employee costs.”

    7:23 · RAISE Summit · 22 Jul 2026
  15. Patel contrasts AI adoption timeline with past revolutions that took decades, arguing AI is dissipating in years.

    exponentialtechnology adoption

    “In in past sort of major revolutions, whether it was like the fax machine or cars or any other technological revolution, it took decades to dissipate.”

    11:38 · RAISE Summit · 22 Jul 2026
  16. Patel claims AI task length is doubling every seven months, leaving no time for companies to wait.

    ai capabilitiesexponential

    “AI task length, right? The amount of time that an AI model can work on any specific job is increasing by 2x every seven months.”

    11:56 · RAISE Summit · 22 Jul 2026
  17. Patel argues companies cannot pause to measure ROI because experimentation is the only path to finding AI value.

    ai adoptionexperimentation

    “So how one say, oh, pause, let me measure value, when trying a bunch of stuff is the only way you can actually get to figuring out where the value is versus what doesn't work?”

    12:08 · RAISE Summit · 22 Jul 2026
  18. Patel argues that AI helps companies improve productivity by moving from Excel to twenty-plus-year-old SQL technology.

    legacy systemsproductivity

    “And so this massively improves their productivity, despite the fact that this is all old, twenty plus year old technology.”

    16:37 · RAISE Summit · 22 Jul 2026
  19. Patel warns companies will lose if managers cannot distinguish LCMs from transformers and understand different architectures.

    ai literacycompetitive advantage

    “Right? If if you're not gonna understand these systems if your managers haven't taken the time to understand the different architectures.”

    17:35 · RAISE Summit · 22 Jul 2026
  20. Patel observes his employees' jobs have radically changed, now needing access beyond their job titles because AI enables broader roles.

    job rolesorganizational change

    “But actually with AI, a lot of the people that are on my team, their jobs have radically changed over the last few years.”

    33:12 · RAISE Summit · 22 Jul 2026
  21. Patel recounts his intern's AI agent autonomously accessing another company's AWS bucket after finding keys during reverse engineering.

    ai agentsautonomous behavior

    “In this particular case, the model decided that it's going to pull it found while reverse engineering an APK, found an AWS access key.”

    36:11 · RAISE Summit · 22 Jul 2026
  22. Patel reveals Anthropic turned first profit in Q2 and will hit billion-dollar operating profit in Q3 before IPO.

    anthropicinference economics

    “So Anthropic turned their first gross profit in Q2 in June. And then in Q3, they will be turning a billion dollars of operating profit, slightly over.”

    1:13 · RAISE Summit · 16 Jul 2026
  23. Patel says Anthropic will turn over a billion dollars of operating profit in Q3 2025.

    anthropicfinancials

    “Anthropic turned their first gross profit in Q2 in June. And then in Q3, they will be turning a billion dollars of operating profit, slightly over.”

    1:13 · RAISE Summit · 16 Jul 2026
  24. Patel says Anthropic's billion-dollar operating profit figures will appear in their IPO.

    anthropicfinancials

    “And so this is this is their financials that they're going to be putting out in their IPO.”

    1:25 · RAISE Summit · 16 Jul 2026
  25. Patel reports OpenAI gross margins rose from 30% to 55% overall, 50% to 65% excluding free users.

    inference economicsmargins

    “You look at OpenAI late last year, their margins had were roughly 30% gross margin, but if you stripped away the free users, they were at 50%.”

    1:33 · RAISE Summit · 16 Jul 2026
  26. Patel says OpenAI's total gross margin rose from 30% to 55% over the past year.

    economicsmargins

    “OpenAI late last year, their margins had were roughly 30% gross margin, but if you stripped away the free users, they were at 50%. Now, total company gross margin is closer to 55%,”

    1:33 · RAISE Summit · 16 Jul 2026
  27. Patel says OpenAI's gross margins rose from 30% to 55% overall, reaching 65% excluding free users.

    marginsopenai

    “Now, total company gross margin is closer to 55%, and if you strip away the free users, they're at about 65%.”

    1:44 · RAISE Summit · 16 Jul 2026
  28. Patel says AWS beat on gross margins because Bedrock was highly profitable.

    awsbedrock

    “On Amazon's most recent earnings call, they talked about how AWS had a strong gross margin beat. Right? They won on gross margin because Bedrock was so profitable.”

    2:06 · RAISE Summit · 16 Jul 2026
  29. Patel describes agentic workflows using 30,000-100,000 input tokens but generating only 1,000 output tokens.

    agentic aiinference

    “And and that makes, you know, initially, that'd be like, okay, well now I need a ton ton of compute to calculate all the prefilled tokens.”

    8:28 · RAISE Summit · 16 Jul 2026
  30. Patel explains KV cache storage and reuse enables massive cost decreases in inference.

    economicsinference

    “you calculate that once, you store it off in memory, whether it be system memory or storage, and then you pull it back in when you run the turn.”

    8:48 · RAISE Summit · 16 Jul 2026
  31. Patel reports cache hit rates above 95% for many agentic workflows in production.

    cacheinference

    “we're seeing cache hit rates above 95% for many AgenTeq workflows, which which means the cost for a cache hit is it's not free,”

    9:40 · RAISE Summit · 16 Jul 2026
  32. Patel argues layering KV cache offload and multi-token prediction on open-source engines drastically cuts inference costs.

    cost reductioninference optimization

    “So if you take the, you know, just open source inference engines off the shelf, that gets you a certain level of cost.”

    14:24 · RAISE Summit · 16 Jul 2026
  33. Patel says InferenceX now has over $80 million of compute across multiple vendors.

    benchmarkinghardware

    “we have over $80,000,000 of compute GPUs from NVIDIA AMD, TPUs from, Google, Tranium from Amazon, and we run this benchmark constantly on the newest inference engine, newest drivers, newest, PyTorch version,”

    15:53 · RAISE Summit · 16 Jul 2026
  34. Patel reveals SemiAnalysis operates over $80 million in compute for daily automated benchmarking across all major AI chips.

    benchmarkinggpu

    “Every day it runs on an automated CI, and we run it on all the latest Chinese models from GLM, Zebu, Moonshot, Kimi, Alibaba, all these models we run.”

    16:08 · RAISE Summit · 16 Jul 2026
  35. Patel says SemiAnalysis analyzed over $5 million worth of Claude production traces to benchmark agentic workloads.

    agentic aibenchmarking

    “Initially, when we were benchmarking the difference between these chips and different engines, different schemes for parallelism, we were just running it, you know, fixed context length.”

    16:21 · RAISE Summit · 16 Jul 2026
  36. Patel says InferenceX analyzed over $5 million worth of Claude Code production traces.

    benchmarkingclaude

    “we've analyzed over $5,000,000 worth of Claude code traces. Right? So this is real production traffic that people have donated to us,”

    16:31 · RAISE Summit · 16 Jul 2026
  37. Supermicro, 15 July 2026

    amdgpu

    “It's 72 GPUs. It's got more memory than Vera Rubin. It's got more more, memory, flops. It's got it's it's better than Vera Rubin in most every way. It's a little bit later.”

    2:04 · Supermicro · 15 Jul 2026
  38. Patel says AMD MI450X has 72 GPUs, more memory and flops than Vera Rubin, launching three to six months later.

    amdgpu

    “It's got more more, memory, flops. It's got it's it's better than Vera Rubin in most every way. It's a little bit later. I mean, three to six months after,”

    2:09 · Supermicro · 15 Jul 2026
  39. Patel reveals Supermicro is the launch partner for AMD Helios with Meta, Oracle, and OpenAI as public customers.

    customershelios

    “by being the launch partner for Helios, know, that's that's pretty exciting. There's there's quite a bit of, public customer traction, Meta, Oracle for OpenAI,”

    2:19 · Supermicro · 15 Jul 2026
  40. Patel names Meta, Oracle, and OpenAI as public customers adopting AMD's MI450X Helios platform.

    amdcustomers

    “There's there's quite a bit of, public customer traction, Meta, Oracle for OpenAI, and there's many other customers who are who are looking to adopt it.”

    2:21 · Supermicro · 15 Jul 2026
  41. Patel says a data center is physically redoing doors to accommodate the double-width Helios racks.

    data centerhelios

    “There's a there's a data center that I know of specifically that is going to be deploying Helios, and they are redoing the doors so that they can reel the Helios in.”

    3:27 · Supermicro · 15 Jul 2026
  42. Patel predicts Nvidia Vera Rubin will have smoother deployment than Blackwell due to GB's new architecture problems.

    deploymentnvidia

    “It's it seems like it'll be a much smoother ramp than GB. You know, GB had a lot of problems because it was brand new,”

    3:52 · Supermicro · 15 Jul 2026
  43. Patel credits Supermicro as first on liquid cooling with XAI's 100,000 GPU Colossus deployment, with others following.

    coolingsupermicro

    “you guys were the first on liquid cooling with XAI's Colossus, the first large scale 100,000 GPU deployment with liquid cooling, and you've done many more since, and you were the first and everyone else has sort of followed since.”

    12:21 · Supermicro · 15 Jul 2026
  44. Patel says SemiAnalysis initially doubted Jensen's 25x Blackwell claim, predicting only 15-20x improvement.

    blackwellnvidia

    “Jensen, when he originally launched Blackwell, had claimed it would be a 25x improvement. And at the time, no one believed him. Right? It's Jensen, right?”

    11:48 · WisdomTree in Europe · 9 Jul 2026
  45. Patel's benchmarking found Blackwell is 30x faster than Hopper on DeepSeek v3, exceeding Jensen's 25x claim.

    benchmarkingblackwell

    “In DeepSeek v three, Blackwell is 30 x faster than Hopper on on somewhere on the continuum.”

    12:17 · WisdomTree in Europe · 9 Jul 2026
  46. Patel emailed Jensen documenting skepticism about Blackwell's 25x claim from industry observers.

    blackwelljensen huang

    “I emailed him. I like, hey, Jensen. You know, back in 2024, you said or back back when you launched Blackwell, said '24 twenty five x.”

    12:37 · WisdomTree in Europe · 9 Jul 2026
  47. Patel reports Anthropic achieved profitability and positive free cash flow in April and May 2025.

    ai economicsanthropic

    “Anthropic is free cash flow positive, and they are profitable in q two. Even in April. In April, they closed April's books. They were profitable.”

    16:12 · WisdomTree in Europe · 9 Jul 2026
  48. Patel's firm spent under $100k annually on AI in November before Claude Code adoption.

    ai spendingclaude code

    “What we had was that we had a subscription to every model or we had a subscription to the $200 tier for a chat GPT for every user.”

    17:36 · WisdomTree in Europe · 9 Jul 2026
  49. Patel's firm's AI spending jumped from under $100k to $11 million annually within months.

    ai spendingclaude code

    “by the end of January, our ARS, right, our our our annual recurring spend had hit $4,000,000. And that's because people were using Cloud Code. Now today, it's about $11,000,000.”

    18:08 · WisdomTree in Europe · 9 Jul 2026
  50. Patel says their AI spending equals over a third of employee compensation for their 90-person firm.

    ai spendinginference economics

    “We're spending, you know, more like, you know, more than a third of the spend, you know, employee spend, a third of it is also on top of that is AI.”

    18:41 · WisdomTree in Europe · 9 Jul 2026
  51. Patel predicted memory prices would soar because capacity grows 30% annually while demand doubles.

    hbmmemory shortage

    “Memory capacity is only growing 30% a year for the next three years. And yet demand is doubling, is doubling. And so what's gonna end up happening is memory prices are gonna keep soaring.”

    31:22 · WisdomTree in Europe · 9 Jul 2026
  52. Patel predicts iPhone and MacBook prices must rise as AI demand drives memory shortages.

    consumer hardwarememory shortage

    “And right now, you know, if MacBook prices or iPhone prices go up a $100, that market's not gonna adjust too much,”

    32:44 · WisdomTree in Europe · 9 Jul 2026
  53. Patel forecasts data center capacity will reach 30 gigawatts in the next year.

    capacitydata centers

    “This year, we're deploying 20 gigawatts of data centers. Next year, that number goes up 50 gigawatts or 50%, sorry. 50%, sorry.”

    54:54 · WisdomTree in Europe · 9 Jul 2026
  54. Patel claims smaller Quen models with 27B total parameters and 2B active outperform GPT-4 from three years ago.

    efficiencymodels

    “If you look back three years, it's GPT-four. Now it's like maybe like quen one of the smaller quen models that's like 27 b parameters total and like 2,000,000,000 active is like way better.”

    2:28 · The AGI Post · 6 Jul 2026
  55. Patel argues OpenAI models are optimized for GPUs while Anthropic and Google models are optimized for TPUs due to co-design.

    anthropicco-design

    “And the way that Anthropic and Google's models are headed, it's actually a terrible decision potentially for them to train with GPUs.”

    6:15 · The AGI Post · 6 Jul 2026
  56. Patel says OpenAI models are much more sparse while Anthropic models are denser, with different architectural benefits.

    architectureopenai

    “Open eyes are much more sparse and that has benefits. Then anthropics are, they're still sparse but more dense in general and that has different benefits.”

    6:47 · The AGI Post · 6 Jul 2026
  57. Patel explains NVLink connects only 72 GPUs while Google ICI connects 8,000 chips without switches, creating architectural trade-offs.

    googleinterconnect

    “NVIDIA, the NVLink can only connect 72 GPUs. For Google, their ICI can connect 8,000 chips at super high bandwidth, but you have to pass through other chips to get there because there's no switch.”

    7:05 · The AGI Post · 6 Jul 2026
  58. Patel secured over $50 million in donated hardware for InferenceX, expanding to $100 million with TPUs and training.

    benchmarkinghardware donation

    “We've got over $50,000,000 of hardware donated to us. Once we launch TPUs and training, would actually be over $100,000,000 of hardware.”

    16:43 · Sequoia Capital · 30 Jun 2026
  59. Patel predicts OpenAI and Anthropic will have over 100 gigawatts combined by 2030, terawatts by 2040.

    2030 forecastinference

    “I think by 2030, just OpenAI and Anthropic will have over 100 gigawatts combined, and then you'll add Meta and Google and so on and so forth.”

    21:27 · Sequoia Capital · 30 Jun 2026
  60. Patel predicts OpenAI and Anthropic alone will deploy over 100 gigawatts of compute by 2030.

    computeforecast

    “by 2030, just OpenAI and Anthropic will have over 100 gigawatts combined, and then you'll add Meta and Google and so on and so forth.”

    21:28 · Sequoia Capital · 30 Jun 2026
  61. Patel claims hardware improved 30x from Hopper to Blackwell for DeepSeek on optimized deployments.

    deepseekhardware

    “from Hopper to Blackwell, which is all we've had over the last three years, roughly 30x improvement on DeepSeek, on the most optimized deployment, which you can see on InferenceX there's about a 30x improvement.”

    24:28 · Sequoia Capital · 30 Jun 2026
  62. Patel argues three years of improvement came from model layer advances plus hardware-software co-design, citing smaller Qwen models outperforming GPT-4.

    co-designefficiency

    “If you look back three years, it's GPT-four. Now it's maybe like quen, one of the smaller quen models that's like 27B parameters total and 2,000,000,000 active is way better.”

    24:47 · Sequoia Capital · 30 Jun 2026
  63. Patel argues co-optimizing hardware, software, and models produces 100x gains instead of 8x from multiplicative 2x improvements per layer.

    co-designefficiency

    “The real breakthrough innovation is when you leapfrog a few layers, you co optimize and co design them, and now all of a sudden you've taken what could have been a 2x here, 2x here, 2x here, and instead of being multiplicative to 8x, it's actually 100x, because you've optimized it across all three layers.”

    28:20 · Sequoia Capital · 30 Jun 2026
  64. Patel reveals Google runs three separate TPU design programs with Broadcom, MediaTek, and undisclosed partner.

    googlesupply chain

    “They're making a TPU with Broadcom, that's a different architecture than a TPU with MediaTek, that's a different TPU than the architecture that is, I won't disclose, by research.”

    49:27 · Sequoia Capital · 30 Jun 2026
  65. Patel forecasts 20 gigawatts of data center capacity in 2024 and over 30 gigawatts in 2025 despite delays.

    capacity forecastdata centers

    “This year there's going to be 20 gigawatts, even accounting for the delays, and next year there's going be more than 30 gigawatts accounting for the delays.”

    51:18 · Sequoia Capital · 30 Jun 2026
  66. Patel states Anthropic is profitable excluding stock compensation in Q2 with 80% margins on Opus tokens.

    anthropicmargins

    “Anthropic in Q2 is profitable, their net income profitable, excluding stock based compensation. And I think by Q3 they may even be profitable, including stock based compensation.”

    52:19 · Sequoia Capital · 30 Jun 2026
  67. Patel reports Anthropic is Q2 profitable excluding SBC, expects Q3 profitability including SBC with 80% margins on Opus tokens.

    anthropicmargins

    “And I think by Q3 they may even be profitable, including stock based compensation. That's how profitable they're getting, and their margins on an Opus token, at least Opus 4.8 token, is north of 80”

    52:27 · Sequoia Capital · 30 Jun 2026
  68. Patel reports Anthropic achieved over 80% margins on Opus tokens and expects full profitability including stock compensation by Q3.

    anthropicmargins

    “That's how profitable they're getting, and their margins on an Opus token, at least Opus 4.8 token, is north of 80 for the API price.”

    52:31 · Sequoia Capital · 30 Jun 2026
  69. Patel reports Trainium rents for under $10 billion per gigawatt while GPUs cost $12-13 billion.

    gpupricing

    “Tranium sells at sub $10,000,000,000 per gigawatt rental rate to Anthropic and to OpenAI. GPUs, at least before the craziness of the last six months, usually went around 12 to $13,000,000,000 per gigawatt.”

    58:22 · Sequoia Capital · 30 Jun 2026
  70. Patel states GPU rental rates were $12-13 billion per gigawatt before recent increases.

    economicsgpu pricing

    “GPUs, at least before the craziness of the last six months, usually went around 12 to $13,000,000,000 per gigawatt.”

    58:29 · Sequoia Capital · 30 Jun 2026
  71. Swole as a Service, 13 May 2026

    gpumethodology

    “We try and track the entire supply chain from tools that manufacture chips, fabs, data centers, energy, industrials, and then AI models and who's using them, how much, and where.”

    1:51 · Swole as a Service · 13 May 2026
  72. Patel distinguishes OpenAI hiding reasoning chains from Anthropic showing them in code generation models.

    anthropicinference

    “In the case of OpenAI, they don't show you the whole reasoning chain. In the case of Anthropic, they do.”

    4:20 · Swole as a Service · 13 May 2026
  73. Patel predicts reasoning chains will close as AI response times extend from immediate to hours or days.

    inferenceprediction

    “And then as the horizon of AI models gets longer, right, rather than a question answer immediately, the question answer becomes ten minutes, hours, days, the value of all the reasoning is gone.”

    4:28 · Swole as a Service · 13 May 2026
  74. Patel cites estimates of 20% economic decline if Taiwan is invaded, comparable only to world wars.

    economicsgeopolitics

    “And then subsequently, that that sends the world into a depression. Right? The world has not experienced a 20 free fall since maybe, like, a one of the great wars.”

    10:48 · Swole as a Service · 13 May 2026
  75. Patel estimates five to eight years of global tool manufacturing would be needed to replace Taiwan's chip capacity.

    supply chaintaiwan

    “It would take five to eight years of all of the world's tool manufacturing capability Yeah. To even replace the capacity in Taiwan.”

    11:16 · Swole as a Service · 13 May 2026
  76. Patel notes China has a fully self-sufficient vertical semiconductor supply chain, though behind by three to fifteen years.

    chinageopolitics

    “China has a vertical supply chain. Yes. In many cases, it's fifteen years behind, ten years behind, in some cases, only three years behind.”

    11:50 · Swole as a Service · 13 May 2026
  77. Patel argues Mythos fast mode is probably cheaper than Claude 4.6 fast mode for most tasks due to token efficiency.

    inference economicsmodel comparison

    “the flip side is, is, is mythos is more token efficient. Methos fast mode is probably cheaper than like, four, six fast mode for most tasks.”

    9:02 · SemiAnalysis Weekly · 6 May 2026
  78. Patel predicts Google and OpenAI will both release new models in two weeks.

    googlemodel releases

    “Release New release in two weeks. Everyone's releasing in two weeks. Google, OpenAI, maybe Anthropic. I don't know about Anthropic, but Google and open AI definitely releasing in two weeks.”

    13:17 · SemiAnalysis Weekly · 6 May 2026
  79. Patel says the model was released before pre-training was finished and needs more RL.

    model developmentpre-training

    “for everything new, more, more continued pre training because the spud, they kind of like didn't finish the pre training and just released it. So finished the pre training, do more RL drop model.”

    13:35 · SemiAnalysis Weekly · 6 May 2026
  80. Patel argues Anthropic trapped itself in an innovator's dilemma by focusing on CLI instead of app experience.

    anthropiccli

    “But the CLI experience is not the end all be all of agent orchestration and therefore they've really cooked themselves into sort of like an innovator's dilemma where they keep making the CLI better.”

    30:57 · SemiAnalysis Weekly · 6 May 2026
  81. Patel reports kernel programmers use Codex to create code first, then have Opus rewrite for quality.

    coding workflowdeveloper practice

    “I think it was tree down. Tree down is like, yeah, dude, Codex is like so dumb, but I always just have it created and then this and it works and it's smarter, but the code is sloppy, I have Opus rewrite it.”

    35:19 · SemiAnalysis Weekly · 6 May 2026
  82. Patel reports kernel programmers use Codex to generate code then Opus to clean it, not the reverse workflow.

    codexopus

    “But you can't go the other way around. You can't have Opus write the thing and then have Codex fix this.”

    35:29 · SemiAnalysis Weekly · 6 May 2026
  83. Patel observes Asian researchers consistently produce KV cache reduction papers unknown to American researchers.

    asiakv cache

    “Some paper that reduces KV cash, and no researcher in America has even heard of this paper.”

    36:52 · SemiAnalysis Weekly · 6 May 2026
  84. Patel observes Asian researchers consistently reference KV cache reduction papers unknown to American researchers, citing DeepSeek and TurboQuant.

    asiakv cache

    “It's the fucking best thing ever. DeepSeek and TurboQuant were the, like, most precipitous ones that popped up the most,”

    36:57 · SemiAnalysis Weekly · 6 May 2026
  85. Patel explains long context windows fail because pre-training uses short context and lacks data for longer ranges.

    context windowsdata

    “But it's like, what data do I have? That's why my two fifty ks to one mil context is trash anyways, because there's no data on this stuff.”

    39:58 · SemiAnalysis Weekly · 6 May 2026
  86. Patel says Nvidia and AMD both lie about peak flops specs which are impossible to achieve.

    amdbenchmarking

    “All their quoted specs are lies impossible to achieve Whether it's Nvidia or AMD, neither of them you can ever hit their peak flops.”

    3:57 · TensorWave · 30 Apr 2026
  87. Patel claims Nvidia achieves 50% sustained performance on common 8k gemm while AMD gets only 30%.

    benchmarkinggpu performance

    “if you just take a eight k by eight k by eight k gem, very common shape and you run it on the map mode unit of AMD and Nvidia, you get like 50% sustained performance on Nvidia and you get like 30% sustained performance on AMD.”

    4:23 · TensorWave · 30 Apr 2026
  88. Patel asserts vendors cannot achieve their advertised flops specifications in real-world conditions.

    benchmarkingflops

    “There there is no functional way to get anywhere close to their flops that they advertise.”

    4:37 · TensorWave · 30 Apr 2026
  89. Patel explains Inference X was created because vendor-claimed performance is unattainable with any available framework.

    benchmarkinginference x

    “There's no way to get the performance that vendors like to claim. And so our whole thing there was, well, how do we actually have a benchmark that represents what people actually get on performance?”

    11:41 · TensorWave · 30 Apr 2026
  90. Patel notes software dependencies change daily or multiple times weekly across the entire stack.

    benchmarkinginference

    “And furthermore, with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.”

    11:55 · TensorWave · 30 Apr 2026
  91. Patel explains AI software stacks update multiple times per week making performance measurement a moving target.

    benchmarkinginference

    “with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.”

    11:55 · TensorWave · 30 Apr 2026
  92. Patel notes AI software stacks update constantly with nightly builds making performance measurement time-sensitive.

    ai infrastructurebenchmarking

    “with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly.”

    11:55 · TensorWave · 30 Apr 2026
  93. Patel observes VLLM and SGLANG framework developers compete publicly on Twitter using InferenceX as a leaderboard.

    benchmarkingsglang

    “VLLM and SGLANG love competing with each other. They used to post on Twitter all the time about how they had beat the other one in certain some kind of performance”

    21:11 · TensorWave · 30 Apr 2026
  94. Patel says VLLM and SGLANG now use Inference X as their competition leaderboard instead of Twitter posts.

    inference xsglang

    “They used to post on Twitter all the time about how they had beat the other one in certain some kind of performance and now they use Inference X all the time.”

    21:14 · TensorWave · 30 Apr 2026
  95. Patel reports AMD and Nvidia engineers treat Inference X as a competitive leaderboard.

    amdbenchmarking

    “AMD and Nvidia. They love to compete with each other. And AMD engineers and Nvidia engineers look at Inference X as a leader board.”

    21:21 · TensorWave · 30 Apr 2026
  96. Patel argues effective benchmarks must anger vendors by revealing uncomfortable truths about performance.

    benchmarking philosophyvendor accountability

    “I think one common thing between the three of us is if you're not pissing off people with your benchmark, then you're not testing something useful.”

    30:06 · TensorWave · 30 Apr 2026
  97. Patel reports H100 rental prices increased from $1.71 per hour to over $2.40 in six months.

    gpu pricingh100

    “And just in the last six months, it's gone from, you know, deals for one year transacting at $171.60 an hour for h 1 hundreds to now $2.40 plus.”

    0:45 · Aria Networks · 16 Apr 2026
  98. Patel identifies disaggregated PD and wide EP as networking techniques where performance matters significantly.

    inferencenetworking

    “When someone has really good networking with these leading edge techniques to run a model such as KIMI or DeepSeq called disaggregated PD or wide EP, these are two different techniques, network performance is a humongous factor.”

    3:25 · Aria Networks · 16 Apr 2026
  99. Patel claims code agent revenue grew from a couple billion to over $10 billion in a very short time.

    code agentsinference

    “code agent revenue has gone from a couple billion to north of 10,000,000,000 in like a very short amount of time. And these, the horizon of these has also increased dramatically,”

    7:11 · Daytona · 7 Apr 2026
  100. Patel reports code agent revenue jumped from a couple billion to over $10 billion in a very short time.

    ai revenuecode agents

    “code agent revenue has gone from a couple billion to north of 10,000,000,000 in like a very short amount of time.”

    7:11 · Daytona · 7 Apr 2026
  101. Patel says CPU-to-GPU ratio has shifted from 1:100 megawatts to much closer parity.

    cpu demanddatacenter architecture

    “So whereas before you'd have like many GPU servers per CPU server, and so you'd be like, 100 megawatts of GPUs would be served by even like one megawatt or less of CPUs.”

    8:39 · Daytona · 7 Apr 2026
  102. Patel says CPU-to-GPU ratios have tightened from 1:100 megawatts to much closer ratios due to RL and agentic inference.

    cpu demanddatacenter architecture

    “Nowadays the ratio is like getting much, much closer, both for RL training and for inference, agentic inference.”

    8:49 · Daytona · 7 Apr 2026
  103. Patel reports Amazon's CPU server installations have tripled year-on-year with no capacity remaining anywhere.

    amazoncpu shortage

    “So then you've just seen everyone run out of CPUs, Amazon's volumes on CPUs, of CPU servers they're installing have three x year on year, right, from this year versus last year.”

    8:56 · Daytona · 7 Apr 2026
  104. Patel reports Amazon's CPU server installations have tripled year-over-year with no spare capacity.

    amazoncpu capacity

    “Amazon's volumes on CPUs, of CPU servers they're installing have three x year on year, right, from this year versus last year. And so there's just no capacity anywhere”

    9:00 · Daytona · 7 Apr 2026
  105. Patel says Amazon's CPU server installations have tripled year-on-year, causing capacity shortages and GitHub instability.

    amazoncpu capacity

    “And so there's just no capacity anywhere and that's causing a lot of instability, not just in GitHub but probably other places too.”

    9:07 · Daytona · 7 Apr 2026
  106. Patel reveals OpenAI ported entire stack from x86 to ARM to access Amazon's CPU capacity for recent deal.

    amazoncpu shortage

    “A lot of the impetus for the recent Amazon and OpenAI deal was, yes, OpenAI wanted money, yes, they needed compute, but they also just went to Amazon and were like, give us your CPUs, And before, OpenAI's stack was pretty much only on x86 CPUs, but Amazon had tons of ARM CPUs and so they ported the whole thing,”

    11:24 · Daytona · 7 Apr 2026
  107. Patel predicts next-generation models will handle physical world tasks like tying shoes beyond current math and coding capabilities.

    agentsmodel capabilities

    “And when we look at like o one, that's like all it could do is math. Right?”

    16:38 · Daytona · 7 Apr 2026
  108. Patel reports memory prices have quadrupled over the past year, with SSDs up three to four times.

    memorypricing

    “Memory has like four x in price over the last like year, and it's gonna continue to go up. And then now storage SSD, right?”

    21:33 · Daytona · 7 Apr 2026
  109. Patel says AI chips consuming all TSMC 3nm and 2nm capacity is forcing companies to alternative fabs.

    nvidiasupply chain

    “More importantly is that AI is buying all the capacity on three nanometer and two nanometer in a couple years that, people are having to turn to other directions.”

    22:58 · Daytona · 7 Apr 2026
  110. Patel argues NVIDIA acquired Grok partly because it's manufactured on Samsung, avoiding TSMC capacity constraints.

    nvidiasamsung

    “Part of it was because they want to have really fast inference, but like part of it is that Grok is manufactured on Samsung.”

    23:15 · Daytona · 7 Apr 2026
  111. Patel explains NVIDIA acquired Grok partly because it manufactures on Samsung due to no TSMC 3nm capacity available.

    nvidiasupply chain

    “Because there's no three nanometer capacity for them at TSMC, they need to chip somewhere else, and if you just if AI is as crazy as like we believe it is, and demand is as crazy as we believe it is, it's gonna be even crazier next year,”

    23:21 · Daytona · 7 Apr 2026
  112. Patel explains AI chips are crowding onto three nanometer nodes, displacing mobile chip production.

    3nmsupply chain

    “Right? Because all the AI chips are on three nanometer and it takes time, Small mobile chips are easier to make than big AI chips, so all the AI chips are moving to three nanometer now, right? MI three fifty from”

    23:48 · Daytona · 7 Apr 2026
  113. Patel says AI chips are moving to three nanometer, squeezing out mobile chip manufacturers.

    capacity constraintsmanufacturing nodes

    “Right? Because all the AI chips are on three nanometer and it takes time, Small mobile chips are easier to make than big AI chips, so all the AI chips are moving to three nanometer now, right?”

    23:48 · Daytona · 7 Apr 2026
  114. Patel lists AMD MI350, Amazon Trainium 3, and Google TPU v7 as AI chips now on 3nm process node.

    process nodessupply chain

    “MI three fifty from late last year from AMD or Tranium three and TPU v seven from late last year, or slash early this year from Amazon and Google”

    23:59 · Daytona · 7 Apr 2026
  115. Patel cites Anthropic adding $67 billion ARR per month as evidence of expanding AI revenue beyond hyperscalers.

    anthropicinference economics

    “You're starting to see it with Anthropix revenue adding $67,000,000,000 of ARR a month, but there's so many more firms coming online with revenue streaming in,”

    1:57 · CNBC Television · 16 Mar 2026
  116. Patel says NVIDIA has locked up over 60% of supply chain capacity this year in long-term contracts.

    export controlsgpu supply chain

    “The NVIDIA has locked up over 60% of the capacity this year in long term contracts alone, and they're buying more on top of that.”

    4:13 · CNBC Television · 16 Mar 2026
  117. Patel says NVIDIA has locked up over 60% of market capacity in long-term contracts this year.

    market sharenvidia

    “NVIDIA has locked up over 60% of the capacity this year in long term contracts alone, and they're buying more on top of that.”

    4:13 · CNBC Television · 16 Mar 2026
  118. Patel reports NVIDIA is negotiating over $250 billion in supply contracts across components this year.

    gpu supply chainnvidia

    “And as you look at what they're negotiating in the market today, there's over 250,000,000,000 of wafers, of memory, of substrates, of PCBs, of networking equipment that they're going to sign this year,”

    4:21 · CNBC Television · 16 Mar 2026
  119. Patel reports NVIDIA will sign over $250 billion in supply contracts this year across components.

    contractsnvidia

    “as you look at what they're negotiating in the market today, there's over 250,000,000,000 of wafers, of memory, of substrates, of PCBs, of networking equipment that they're going to sign this year,”

    4:21 · CNBC Television · 16 Mar 2026
  120. Patel contrasts competitors' limited scale with NVIDIA's supply chain designed for tens of millions of chips.

    competitionnvidia

    “I can buy 10,000. I could buy a 100,000. I can't buy millions, tens of millions, which is what Jensen's, setting his supply chain up for.”

    4:43 · CNBC Television · 16 Mar 2026
  121. Patel argues NVIDIA is setting up supply chain to manufacture tens of millions of AI chips.

    gpu supply chainnvidia

    “I can't buy millions, tens of millions, which is what Jensen's, setting his supply chain up for.”

    4:45 · CNBC Television · 16 Mar 2026
  122. Patel estimates Anthropic needs to reach well above five gigawatts of compute capacity by year-end to support revenue growth.

    anthropiccompute capacity

    “Anthropic needs to get to well above five gigawatts by the end of this year, and it's gonna be really tough for them to get there, but it's possible.”

    3:56 · Dwarkesh Patel · 13 Mar 2026
  123. Patel reports seeing AI labs sign GPU rental deals at $2.40 per hour for H100s, well above initial $1.40 cost.

    gpu pricingh100

    “I've seen deals where certain AI labs I'm be a little bit vague here for a reason, have signed at as high as $2.40 for two to three years.”

    7:38 · Dwarkesh Patel · 13 Mar 2026
  124. Patel explains Google woke up to AI compute needs after Gemini models drove user growth.

    compute scalinggemini

    “Google had Nano Banana and Gemini three, which caused their user metrics to skyrocket, and leadership at Google was like, oh.”

    32:27 · Dwarkesh Patel · 13 Mar 2026
  125. Patel details that producing one gigawatt of Rubin chips requires specific wafer volumes across multiple manufacturing nodes.

    gpu supply chainrubin

    “a gigawatt of, you know, NVIDIA's Rubin chips. Right? So Rubin is announced at GTC, I believe, the week this podcast goes live.”

    38:07 · Dwarkesh Patel · 13 Mar 2026
  126. Patel calculates a gigawatt of Rubin requires 55,000 three-nanometer wafers and 170,000 DRAM wafers.

    rubintsmc

    “You need about 55,000 wafers of three nanometer. You need about 6,000 wafers of five nanometer, and then you need about a 170,000 wafers of DRAM, right, memory.”

    38:28 · Dwarkesh Patel · 13 Mar 2026
  127. Patel calculates three and a half EUV tools are needed to produce one gigawatt of AI capacity.

    asmleuv bottleneck

    “you end up with actually, I need about three and a half EUV tools to do the 2,000,000 EUV wafer passes for the gigawatt. So three and a half EUV tools satisfies the gigawatt.”

    40:23 · Dwarkesh Patel · 13 Mar 2026
  128. Patel shows that $50 billion in data center capacity depends on just $1.2 billion worth of EUV tools.

    asmlbottleneck

    “I need about three and a half EUV tools to do the 2,000,000 EUV wafer passes for the gigawatt. So three and a half EUV tools satisfies the gigawatt.”

    40:25 · Dwarkesh Patel · 13 Mar 2026
  129. Patel calculates one gigawatt of AI capacity requires only 3.5 EUV tools costing $1.2B.

    asmlbottleneck

    “Because we're talking oh, what's a gigawatt cost? It costs like $50,000,000,000 roughly. Right? Whereas, what does three and a half EUV tools cost? That's like 1.2.”

    40:37 · Dwarkesh Patel · 13 Mar 2026
  130. Patel projects maximum 200 gigawatts of AI chip capacity by 2030 based on EUV tool supply.

    2030 forecastcapacity constraint

    “700 EUV tools by the end of the decade. 700 EUV tools, three and a half tools per gigawatt, assuming it's all allocated to AI, which it's not, but three and a half tools per gigawatt gets you to 200 gigawatts worth of AI chips for the data centers to deploy.”

    42:30 · Dwarkesh Patel · 13 Mar 2026
  131. Patel calculates 700 EUV tools by 2030 enables maximum 200 gigawatts of AI chip production.

    2030 forecastasml bottleneck

    “700 EUV tools, three and a half tools per gigawatt, assuming it's all allocated to AI, which it's not, but three and a half tools per gigawatt gets you to 200 gigawatts worth of AI chips for the data centers to deploy.”

    42:33 · Dwarkesh Patel · 13 Mar 2026
  132. Patel estimates total EUV capacity by 2030 enables 200 gigawatts of AI chips, with Altman seeking 52 gigawatts annually.

    asmlsam altman

    “700 EUV tools, three and a half tools per gigawatt, assuming it's all allocated to AI, which it's not, but three and a half tools per gigawatt gets you to 200 gigawatts worth of AI chips for the data centers to deploy. Right? So 200 gigawatts, Sam wants 50 gigawatts, right, 52 gigawatts a year.”

    42:33 · Dwarkesh Patel · 13 Mar 2026
  133. Patel predicts rising memory prices will make smartphones and PCs worse, causing public backlash against AI.

    ai backlashconsumer impact

    “So memory crunch will continue to be harder and harder, and prices continue to go up. And this affects different parts of the market differently.”

    1:23:42 · Dwarkesh Patel · 13 Mar 2026
  134. Patel reports memory prices have tripled, adding $150 to iPhone bill of materials versus $50 previously.

    consumer impacthbm

    “the price of memory is like tripled. Let's call it if it's now it's $12 per gig for DDR. So now you're talking about a $150 versus $50.”

    1:24:17 · Dwarkesh Patel · 13 Mar 2026
  135. Patel projects smartphone volumes dropping from 1.1B to 500-600M next year due to memory costs.

    market forecastmemory shortage

    “our projections are we maybe get down to like 800,000,000 this year, and next year like 600 or 500,000,000.”

    1:25:30 · Dwarkesh Patel · 13 Mar 2026
  136. Patel reports $19 billion Anthropic revenue is driven by code spend, implying industry-wide $20 billion expenditure.

    anthropiccoding

    “You look at like quad code spend, right, it's freaking nuts. Right? 19,000,000,000 of revenue for Anthropic now. All of that is, you know, at some multiple is code.”

    3:43 · Matthew Berman · 9 Mar 2026
  137. Patel says his company's AI coding spend hit $6 million annual run rate, with single engineers spending $8,000.

    ai spendcloud code

    “Literally, we we you know, like, one day my head of ops was like, guys, our run rate of spend is $6,000,000. This is not sustained.”

    4:42 · Matthew Berman · 9 Mar 2026
  138. Matthew Berman, 9 March 2026

    ai costscloud code

    “one of my engineers spent like $8,000, he's Canadian, so whatever, right, on on Cloud Code. Right? And it's like, oh, shit. Right?”

    4:55 · Matthew Berman · 9 Mar 2026
  139. Patel reports a non-programmer employee now spends $5,000 daily using Claude Code to build tools via natural language.

    ai agentsproductivity

    “I'm just telling you what to do and it's doing these things. And his daily spend now on 4.6 fast 1,000,000 context is $5,000 a day.”

    5:36 · Matthew Berman · 9 Mar 2026
  140. Patel reports a single employee's daily AI spend reaching $5,000 and considers it acceptable given productivity.

    ai costscloud code

    “And his daily spend now on 4.6 fast 1,000,000 context is $5,000 a day. Jeez. And I look at the productivity, it's fine. Right?”

    5:38 · Matthew Berman · 9 Mar 2026
  141. Patel criticizes Dario's response to hypothetical nuclear scenario as inadequate for government AI policy discussions.

    anthropicgovernment

    “If a nuclear missile's heading from China to America and we can use AI to stop it, but it requires surveillance and autonomous weapons, what do we do?”

    9:08 · Matthew Berman · 9 Mar 2026
  142. Patel criticizes Dario Amodei's response to a hypothetical nuclear defense scenario as the dumbest possible answer.

    ai safetyanthropic

    “And Dario was like, well, you can call us. I'm sure we can figure something out. Like, this is just the dumbest response you could ever come up with.”

    9:19 · Matthew Berman · 9 Mar 2026
  143. Patel says Sam Altman's greatest skill is taking advantage of a crisis to sign government deals.

    openaisam altman

    “That's maybe his greatest skill, is to take advantage of a good crisis. And so he goes out there and he signs deals with the US government.”

    9:53 · Matthew Berman · 9 Mar 2026
  144. Patel observes perception shift where 90th percentile people consider themselves middle class while 50th percentile people do not.

    economicsinequality

    “Nowadays, people in the ninetieth percentile think they're middle class, and people in the fiftieth percentile don't think they're middle class.”

    23:22 · Matthew Berman · 9 Mar 2026
  145. Patel observes 90th percentile earners think they're middle class while 50th percentile don't, due to social media perception.

    inequalityperception

    “Nowadays, people in the ninetieth percentile think they're middle class, and people in the fiftieth percentile don't think they're middle class. Yeah.”

    23:22 · Matthew Berman · 9 Mar 2026
  146. Patel claims his company now moves 10 times faster than competitors using Cloud Code.

    cloud codecompetitive advantage

    “Now with Cloud Code, think we move 10 times the speed of any of our competitors. Right? It's actually insane.”

    41:55 · Matthew Berman · 9 Mar 2026
  147. Patel claims his company moves 10x faster than competitors using Claude Code, enabling more hiring and business wins.

    claude codecompetition

    “It's actually insane. And so my plans of hiring and the business we win and the basis we're able to do is actually just growing.”

    41:59 · Matthew Berman · 9 Mar 2026
  148. Patel says he now supports UBI despite being very capitalist, a shift from his prior views.

    ai impacteconomics

    “Over the last couple years I've realized, wait, I actually think UBI is perfectly fine. Which I think is crazy because again, I'm very capitalist.”

    48:53 · Matthew Berman · 9 Mar 2026
  149. Patel predicts Anthropic revenue will grow from $19B to $60B by year end amid majority-negative AI sentiment.

    anthropicforecasts

    “The sentiment is already more than half Americans have a negative view of AI. As we fast forward to the end of this year, as Anthropix revenue goes from 19,000,000,000 to maybe 60, and OpenAI is similar, if not higher, the amount of jobs that get supplemented, the amount of change that happens in society,”

    49:29 · Matthew Berman · 9 Mar 2026
  150. Patel predicts Google will eliminate its $100 billion annual cash flow next year by spending everything on AI infrastructure.

    capexcash flow

    “Google will have no cash flow next year because they're they see AI so clearly, and they know that they need to spend every dollar they make on compute,”

    50:00 · Matthew Berman · 9 Mar 2026
  151. Patel says Google spending $180B and Amazon $200B on AI infrastructure represents 4x increase from recent years.

    capexhyperscalers

    “If we start looking at like, hey, this year, Google spending 2 Amazon spending $200,000,000,000, Google spending a $180,000,000,000 on on AI infrastructure primarily. Right?”

    15:40 · Latent Space · 26 Feb 2026
  152. Patel states Anthropic now adds $23 billion revenue monthly versus few hundred million earlier.

    anthropicinference economics

    “Anthropics, you know, adding $23,000,000,000 of revenue a month now Mhmm. Versus they were just adding a few 100,000,000 of revenue a month earlier. So clearly, we're in the take off period.”

    16:32 · Latent Space · 26 Feb 2026
  153. Patel reports Claude Code doubled to 4% of GitHub commits in January alone, with overall AI coding likely at 10%.

    ai adoptioninference economics

    “Just in this month just in January, it went from 4% of or 2% of commits on GitHub to 4% of GitHub commits were done by Cloud Code.”

    21:14 · Latent Space · 26 Feb 2026
  154. Patel calculates Anthropic added $1.5 billion in compute infrastructure in one month to serve new revenue growth.

    anthropicgpu supply chain

    “If Anthropic added 2,500,000,000 of revenue, and their gross margin is 40%, they added like $1,500,000,000 of compute in one month. Mhmm. Right?”

    26:00 · Latent Space · 26 Feb 2026
  155. Patel predicts Google will have zero cash flow profit in 2027, spending all revenue on AI infrastructure.

    forecastsgpu supply chain

    “There's no reason why Google will have any profit in '27 at all, right, in terms of cash flow. They will just spend every dollar they make on on AI infrastructure.”

    32:41 · Latent Space · 26 Feb 2026
  156. Patel predicts Google and Amazon will take on debt for AI infrastructure following Meta's example.

    debtfinancing

    “Google and Amazon haven't taken on debt yet for AI infrastructure, but they will. Right?”

    33:54 · Latent Space · 26 Feb 2026
  157. Patel predicts becoming the anti-AI party will be the winning political strategy due to public backlash.

    ai backlashelections

    “And it seems obvious to me that, like, any party that wants to win should just become the anti AI party. Because life as we know it is changing.”

    38:35 · Latent Space · 26 Feb 2026
  158. Patel says NVIDIA must be 2x better than competitors to justify their 75% plus margins.

    competitionmargins

    “NVIDIA recognizes they're they're the leader, they're the tent pole. Hey, in one respect, they can just run faster than everyone, but it's kind of hard to be two x better than Google or or OpenAI or whoever else's internal chip, right, to justify their, you know, 75% plus margins.”

    4:14 · The MAD Podcast with Matt Turck · 5 Feb 2026
  159. Patel says NVIDIA must be 2-4x better than competitors to justify 75% margins and 4x pricing above costs.

    competitionmargins

    “And then they have to be two x to four x better to justify four x better to justify their margins because that's what they're charging above cogs.”

    4:32 · The MAD Podcast with Matt Turck · 5 Feb 2026
  160. Patel explains Jensen fears specialized chips could undercut NVIDIA's margins if they only made general-purpose GPUs.

    jensen huangnvidia

    “Jensen is very paranoid about losing. Right? These specializations if he just kept making his mainline chip would mean people could you know point point solutions for specific parts of the market would crush him on cost and performance, then he can't justify his margin.”

    7:28 · The MAD Podcast with Matt Turck · 5 Feb 2026
  161. Patel explains NVIDIA acquired Grok to get engineering resources for multiple chip architectures.

    acquisitionsgrok

    “acquiring Rock is like how you get those resources to make more solutions for different parts of the market. And as far as like, are they threatened?”

    8:51 · The MAD Podcast with Matt Turck · 5 Feb 2026
  162. Patel predicts AMD will remain in single-digit percentage market share despite being credible competitor.

    amdmarket share

    “I don't think they'll go beyond, like I think they'll stay in single digits market share, single digit percentage market share. Yeah. Single digit percentage market share”

    17:30 · The MAD Podcast with Matt Turck · 5 Feb 2026
  163. Patel says China's entire culture is semiconductor-focused, even appearing in romance dramas.

    chinaculture

    “I think the entire country is like semiconductor pilled. Mhmm. Right? There are dramas where people fall in love in the fab or dramas where people fall in love and they're photovoltaic,”

    26:11 · The MAD Podcast with Matt Turck · 5 Feb 2026
  164. Patel says China's entire culture is semiconductor pilled with romance dramas set in fabs and solar research.

    chinaculture

    “There are dramas where people fall in love in the fab or dramas where people fall in love and they're photovoltaic, like, solar cell researchers and engineers.”

    26:15 · The MAD Podcast with Matt Turck · 5 Feb 2026
  165. Patel claims China has the most vertical semiconductor stack and could run fabs independently unlike TSMC.

    chinasemiconductors

    “China has the most vertical stack in semiconductors today and they're the best at semiconductors in the world because their fabs could still run somewhat on a lot of things because they have built some of these chemical supply chains. Right? Like TSMC for certain kinds of chemicals 100% share from Japan.”

    30:44 · The MAD Podcast with Matt Turck · 5 Feb 2026
  166. Patel predicts AI industry will hit $100 billion ARR by end of 2025, with China 10x lower.

    ai revenuechina

    “And then what's the economic value of that $100,000,000,000? Now, how much of that is in China? Right? Like, China's number is probably 10 x lower.”

    39:55 · The MAD Podcast with Matt Turck · 5 Feb 2026
  167. Patel frames AI competition as economic war determining whether China rises to global hegemony.

    aichina

    “at the end of the day, this is an economic war. Right? If The US and the West win in AI and control, you know, more powerful AI systems that have this feedback loop that improve economic growth and weapon systems and whatever else, right, engineering of grids and cyber attacks and all these sorts of things.”

    40:29 · The MAD Podcast with Matt Turck · 5 Feb 2026
  168. Patel argues Western AI dominance determines whether China becomes global hegemon; without it, China will win economically.

    ai racechina

    “They have this like advantage over China, then China will not rise to be the global hegemony. But without AI, China definitely will rise to be the global hegemony.”

    40:46 · The MAD Podcast with Matt Turck · 5 Feb 2026
  169. Patel calculates $100B AI revenue requires $250B infrastructure spend at five-year depreciation, double current capex.

    capexeconomics

    “That $50,000,000,000 of COGS needs to burn on infra, which cost roughly with if a five if you're talking about five year depreciation, call it $250,000,000,000, right, of infra Yeah. For a $100,000,000,000 of revenue.”

    49:45 · The MAD Podcast with Matt Turck · 5 Feb 2026
  170. Patel calculates $100B AI revenue requires $250B infrastructure with five-year depreciation at 50% margins.

    ai economicscapex

    “That $50,000,000,000 of COGS needs to burn on infra, which cost roughly with if a five if you're talking about five year depreciation, call it $250,000,000,000,”

    49:45 · The MAD Podcast with Matt Turck · 5 Feb 2026
  171. Patel notes 2% of GitHub commits are AI-generated against $2 trillion in global software wages.

    ai adoptioncoding

    “You can disable that where it's not automatically committed. But 2% of GitHub commits today are Cloud Code. $2,000,000,000,000 of software wages paid in the world.”

    51:51 · The MAD Podcast with Matt Turck · 5 Feb 2026
  172. Patel says data centers will reach 10% of US power by 2027-28 but under 1% of water consumption.

    data centerspower

    “US Grid will get to like 10% of power by like '28, 27 is data centers. For water consumption, it's not even gonna crack 1%. Yeah. By the end of the decade.”

    57:49 · The MAD Podcast with Matt Turck · 5 Feb 2026
  173. Patel reports 10-15% of NVIDIA GPUs fail and require RMA within first two weeks of cluster deployment.

    gpunvidia

    “When you first turn on the cluster, about ten to fifteen percent of them fail RMA in the first two weeks. Wow. And then that's fine. Like you have to receipt them, whatever.”

    4:35 · TBPN · 3 Feb 2026
  174. Patel says 10-15% of NVIDIA GPUs fail and need RMA in the first two weeks after cluster deployment.

    gpunvidia

    “When you first turn on the cluster, about ten to fifteen percent of them fail RMA in the first two weeks. Wow. And then that's fine.”

    4:35 · TBPN · 3 Feb 2026
  175. Patel says Hopper GPU failure rates improved to 5% while Blackwell remains at 10-15%, expecting higher rates for next generation.

    blackwellhopper

    “hopper's now at 5%, but black belt's still 10 to 15%. Wow. Right? Actually started out higher than that. Sure. And when a new generation comes out, it's gonna be higher than 15%.”

    4:48 · TBPN · 3 Feb 2026
  176. Patel explains users will pay 10x more for 10x faster inference completion, justifying Cerberus economics.

    cerberuseconomics

    “for a lot of people, I'm fine to spend 10x the price on something that completes 10x faster. So Cerberus sort of just makes a ton of sense there.”

    7:14 · TBPN · 3 Feb 2026
  177. Patel explains agentic training only needs to sync tokens every few minutes versus weights every seconds in pre-training.

    infrastructurerl

    “When you're doing these rollouts and especially as things get more and more agentic and training, you might not only need to send not the entire weights but just the tokens that are relevant.”

    15:00 · TBPN · 3 Feb 2026
  178. Patel notes semiconductor industry doubles transistors yearly but energy industry wasn't prepared for similar growth.

    energyinfrastructure

    “Whereas the energy industry in America wasn't. And and so, like, initially, people were, like, not creative. They're like, let's do let's do these kinds of gas plants.”

    17:23 · TBPN · 3 Feb 2026
  179. Patel says data center power capacity will nearly double from 15-18 gigawatts this year to 30 gigawatts next year.

    capacitydata center

    “next year, 30 gigawatts are being added and we think the power is there for it. Wow.”

    18:11 · TBPN · 3 Feb 2026
  180. Patel says data center power capacity is growing from 15-18 gigawatts this year to 30 gigawatts next year.

    capacitydata centers

    “What was it It's this it's or this year is like I think it's like 18 ish, 10 ish. Okay. Fifteen fifteen to 18”

    18:14 · TBPN · 3 Feb 2026
  181. Patel predicts semiconductor shortages will return as the primary constraint in 2027 after power constraints ease.

    semiconductorssupply chain

    “But it will fully beach semiconductors again in '27. Right? And so we see this across the entire space of the ecosystem. It's not just TSMC.”

    19:15 · TBPN · 3 Feb 2026
  182. Patel predicts semiconductor shortages will return as primary constraint in 2027 affecting both TSMC and memory manufacturers.

    forecastsupply chain

    “But it will fully beach semiconductors again in '27. Right? And so we see this across the entire space of the ecosystem. It's not just TSMC. It's also memory both.”

    19:15 · TBPN · 3 Feb 2026
  183. Patel says memory makers have not built new fabs since 2022 due to cyclical nature.

    hbmsupply chain

    “The memory makers, in fact, have just not expanded capacity. Basically, new they've not built new fabs since 2022. Yeah. Because their their cycle their cycle is so undulating.”

    19:29 · TBPN · 3 Feb 2026
  184. Patel argues few hedge funds are in San Francisco to fully understand AI developments firsthand.

    ai investinginformation access

    “And I think I think like if you think about how much do you believe in AI and what's your access to information of AI, you know, there's not many hedge funds who live in San Francisco and like fully breathe and live and understand it.”

    33:44 · TBPN · 3 Feb 2026
  185. Patel predicts AI startup revenue will exceed $100 billion by end of year, calling it unbelievable to most.

    ai revenueforecasts

    “by the end of the year do how many people even believe by the end of the year AI startup revenue is over a $100,000,000,000?”

    35:25 · TBPN · 3 Feb 2026
  186. Patel predicts AI startup revenue will exceed $100 billion by end of year, a claim few believe.

    ai revenueforecasts

    “I think that's an insane statement for a lot of people, but that's what it's gonna be. Yeah. Right? And who believes that number? Right? It's like Yeah. Very few people.”

    35:31 · TBPN · 3 Feb 2026
  187. Patel says Anthropic's $300 billion revenue target for end of decade is actually too conservative.

    anthropicforecasts

    “when Anthropic says in their funding, like, hey, we're gonna have $300,000,000,000 of revenue by the end of the decade. And it's like, actually, I think that number's too low.”

    35:42 · TBPN · 3 Feb 2026
  188. Patel states OpenAI will deploy 16-18 gigawatts by end of 2028 requiring $300 billion spend.

    capacitycapex

    “OpenAI is gonna have 18 gigawatts or 16 gigawatts by the end of twenty eight, and they're gonna be able to pay for it. And that's like, well, that's $300,000,000,000 of spend.”

    35:59 · TBPN · 3 Feb 2026
  189. Patel says at his first NeurIPS in 2021 or 2022 he only stopped at posters mentioning transformers.

    neuripsresearch trends

    “I remember the first NURBS I went to, it was, like, '21 or '22. And I was like I would walk around. I'd stop only posters that said the word transformer.”

    0:48 · SAIL Media · 15 Jan 2026
  190. Patel describes simultaneous gains across pre-training, scaling, quantization, systems, and RL infrastructure and methods.

    ai researchoptimization

    “There's tons of gains in quantization and systems, but there's also tons of gains in RL stuff and both the, like, you know, infra side of things and non infra side of things.”

    2:51 · SAIL Media · 15 Jan 2026
  191. Patel states OpenAI has over 500,000 GPUs for R&D but GPT-5 pre-training uses under 100,000 GPUs.

    gpu supply chainopenai

    “OpenAI has over 500,000 GPUs working on r and d. Right? Let's call it r and d.”

    13:33 · SAIL Media · 15 Jan 2026
  192. Patel says CoreWeave was the only Platinum recipient among roughly 40 cloud companies tested by SemiAnalysis.

    cloud infrastructureclustermax

    “Semi Analysis ClusterMax Platinum. The only recipient of Platinum, you know, we had about six gold recipients and we tested roughly 40 different cloud companies, but CoreWeave was the only one that achieved Platinum.”

    1:08 · Blue Cactus AI · 27 Dec 2025
  193. Patel says neo clouds that ranked highest on ClusterMax benchmarks booked significantly more revenue.

    benchmarksneo clouds

    “And since then we just released the second version, but in the middle, the clouds that ranked the highest actually ended up booking way, way, way more revenue.”

    11:30 · Clockwork · 21 Nov 2025
  194. Patel states top neo clouds achieve 35-40% gross margins while many others are losing money.

    marginsneo clouds

    “And this has enabled, you know, the top in the industry companies to have gross margins of 35, 40%. And now there's a ton of Neo Clouds that are losing money.”

    14:00 · Clockwork · 21 Nov 2025
  195. Patel reports five or six neo clouds have already gone out of business in just a few years.

    failuresmarket dynamics

    “There's I believe five or six that have already gone out of business. It's only been a few years. Right. Right?”

    14:12 · Clockwork · 21 Nov 2025
  196. Patel says H100 cannot physically reach advertised 2,000 teraflops, maxing at 1,300 teraflops for matrix operations only.

    gpu performanceh100

    “In fact, it's physically impossible to get 2,000 teraflops out of the H100. Even though the advertises it, at most you can get 1,300 FP8 floating point eight teraflops.”

    24:59 · Clockwork · 21 Nov 2025
  197. Patel cites DeepSeek's open-sourced inference system requiring 140 GPUs communicating over RDMA networks for single replica.

    deepseekinference

    “One example is one that DeepSeek open sourced over December of last yearJanuary, February of this year, where a single replica of inference for a single model is going to be like 140 GPUs.”

    45:48 · Clockwork · 21 Nov 2025
  198. Patel explains prefill costs one-fourth of decode per token, making 30,000 token inputs costlier than 2-4,000 token outputs.

    inference economicsprefill

    “And that's a very common ratio. Right? And then when you think about, okay, the cost of running pre fill is roughly one fourth of running decode.”

    53:47 · Clockwork · 21 Nov 2025
  199. Patel states avoiding redundant prefill through KV cache can cut inference costs to one-fourth of previous levels.

    inference economicskv cache

    “So you can cut your cost to one fourth of what it was previously if you just don't do the pre fill. Right?”

    54:03 · Clockwork · 21 Nov 2025
  200. Patel reports Microsoft aims to 10x training capacity every 18-24 months, representing a 10x increase from GPT-5 training.

    gpt-5scaling

    “We try to 10x the training capacity every eighteen to twenty four months. And so this would be effectively a 10x increase. 10x from what GPD five was trained with.”

    1:19 · Dwarkesh Patel · 12 Nov 2025
  201. Nadella describes connecting multiple Fairwater facilities on a petabit network spanning to Milwaukee for distributed AI training.

    data centersdistributed training

    “Fairwater 4, which you're going to see under construction nearby, will also be on that one petabits network so that we can actually link the two at a very high rate.”

    2:14 · Dwarkesh Patel · 12 Nov 2025
  202. Patel projects AI will progress from 10-30 minute tasks to days of autonomous work, with model companies charging thousands.

    ai capabilitiesautonomy

    “In the future maybe they're doing days worth of work autonomously. And then the model companies are charging thousands of dollars”

    21:24 · Dwarkesh Patel · 12 Nov 2025
  203. Patel states Oracle will grow from one fifth Microsoft's size to bigger by end of 2027 at 35% margins.

    competitionforecasts

    “Oracle is going from like one fifth your size to bigger than you by end of twenty twenty seven.”

    55:49 · Dwarkesh Patel · 12 Nov 2025
  204. Patel estimates Microsoft's 2028 capacity is three and a half gigawatts lower than forecasted due to paused buildouts.

    capacityforecasts

    “And then the other one is, given our estimates on what your capacity is in 2028, is three and a half gigawatts lower? Sure, you could have dedicated that to OpenAI training and inference capacity.”

    59:38 · Dwarkesh Patel · 12 Nov 2025
  205. Patel notes NVIDIA takes 75% margin on hardware comprising most of data center TCO, driving hyperscalers to develop own accelerators.

    custom siliconnvidia margins

    “So, you mentioned how you're depreciating this asset that's five, six years and this is the majority of the 75% of the TCO data center. And Jensen is taking a 75% margin on that.”

    1:03:37 · Dwarkesh Patel · 12 Nov 2025
  206. Patel explains hyperscalers are developing custom accelerators to reduce equipment costs since Jensen takes 75% margins.

    custom acceleratorsgpu economics

    “So what all the hyperscalers are trying to do is develop their own accelerator so that they can reduce this overwhelming cost for equipment to increase their margins.”

    1:03:52 · Dwarkesh Patel · 12 Nov 2025
  207. Patel says Google will make five to seven million TPU chips while Amazon targets three to five million.

    amazoncustom silicon

    “They're going to make something like five to 7,000,000 chips, right? Of their own TPUs. You look at Amazon, they're trying to make three to 5,000,000.”

    1:04:09 · Dwarkesh Patel · 12 Nov 2025
  208. Patel states AI labs are projecting one hundred billion dollars in revenue by 2027-2028 with 2-3x growth.

    ai labsforecasts

    “So, these labs are now projecting revenues of 100,000,000,000 in 2728, and they're projecting revenue keeps growing at this rate of like 3x, 2x”

    1:14:33 · Dwarkesh Patel · 12 Nov 2025
  209. Patel says SemiAnalysis grew from one person two and a half years ago to 50 people this week, completely organically.

    growthsemianalysis

    “Actually, this week we find that we hit 50 people, so having started it five years ago to now, and I've hired my first person two and a half years ago, We've now grown to 50 people just completely organically, so it's been a hell of a ride.”

    1:13 · Chris Best Substack · 31 Oct 2025
  210. Patel estimates AI infrastructure represents 25-75% of current US quarterly economic growth.

    aieconomic impact

    “it's a substantial fraction of economic growth, depending on who you ask. It's anywhere from 25% to 75 of the current quarter's economic growth. That's insane for The US.”

    1:55 · Chris Best Substack · 31 Oct 2025
  211. Patel states AI infrastructure represents 25% to 75% of current US quarterly economic growth.

    aieconomics

    “It's anywhere from 25% to 75 of the current quarter's economic growth. That's insane for The US.”

    1:59 · Chris Best Substack · 31 Oct 2025
  212. Patel says Substack monetization reached hundreds of thousands in months, then millions before he left the platform.

    monetizationrevenue

    “It got to a few $100,000 of revenue and I was like, Holy crap. And then I was able to start hiring. And then I left when we hit a few million of revenue.”

    4:45 · Chris Best Substack · 31 Oct 2025
  213. Patel left Substack initially after reaching a few million in revenue to avoid the 10% fee.

    business modelrevenue

    “And then I left when we hit a few million of revenue. More and more of the revenue was coming from not Substack. The reason was like, Hey, Substack takes 10%.”

    4:51 · Chris Best Substack · 31 Oct 2025
  214. Patel says email capture growth accelerated dramatically on Substack but fell back after leaving despite SEO work.

    email capturegrowth

    “And then I added Substack and growth went like this. And then I left Substack and it went back to like this, the email capture.”

    5:34 · Chris Best Substack · 31 Oct 2025
  215. Patel says email growth slowed significantly after leaving Substack despite having dedicated SEO staff.

    emailgrowth

    “And then I left Substack and it went back to like this, the email capture.”

    5:37 · Chris Best Substack · 31 Oct 2025
  216. Patel says his Substack revenue alone reached multi-million dollars, sufficient to fund significant coverage.

    mediarevenue

    “Now, scaled to be And it has scaled to be multi million dollar Substack revenue.”

    12:36 · Chris Best Substack · 31 Oct 2025
  217. Patel says Substack revenue alone reached multi-million dollars before expanding beyond the platform.

    monetizationrevenue

    “Now, scaled to be And it has scaled to be multi million dollar Substack revenue. Obviously, built a lot of business beyond that, even that is good enough”

    12:36 · Chris Best Substack · 31 Oct 2025
  218. Patel released a post on GPT-3 training costs the same day ChatGPT launched, calling it coincidental timing.

    aichatgpt

    “I was lucky enough that the day Chad GPT released, I released a post called that was about the training cost of GPT-three. It was just like a coincidence.”

    14:06 · Chris Best Substack · 31 Oct 2025
  219. Patel says as the largest technology Substack, the 10% platform fee is worth paying for the growth.

    economicsplatform

    “Even though a 10% cut sounds like a lot, I'm the largest technology substack and I'm telling you that it's worth it.”

    17:22 · Chris Best Substack · 31 Oct 2025
  220. Patel says as the largest technology Substack, the 10% fee is worth it for the growth.

    platform valuesubstack

    “I'm the largest technology substack and I'm telling you that it's worth it. And so, you can think, Oh, I can save the 10%, but actually the growth is all there. Hell yeah.”

    17:25 · Chris Best Substack · 31 Oct 2025
  221. Patel says Weka and Vast make high margins on storage for multimodal AI workloads despite drive vendors making nothing.

    marginsmultimodal

    “But then there's also, on the storage side, the drive vendors don't make any money. But Weka and Vast, I mean, look at their pricing models.”

    9:24 · Open Compute Project · 23 Oct 2025
  222. Patel says storage vendors Weka and Vast make high margins while drive vendors make no money.

    marginsstorage

    “But then there's also, on the storage side, the drive vendors don't make any money. But Weka and Vast, I mean, look at their pricing models. They make crazy margin on storage.”

    9:24 · Open Compute Project · 23 Oct 2025
  223. Patel says inference providers sell public endpoints at flat or negative margins, compensating through private deployments.

    business modelinference economics

    “Most of the inference providers are selling at flat margins or even negative for their public endpoints. And they then make it up when people do private deployments.”

    10:07 · Open Compute Project · 23 Oct 2025
  224. Patel describes public inference as a loss leader to generate private sovereign and enterprise engagements.

    business modelenterprise

    “And so I think that's what a lot of this is, is you end up with the public business as like a loss leader just to generate private engagements, whether it be sovereigns or enterprises.”

    10:16 · Open Compute Project · 23 Oct 2025
  225. Patel reveals OpenAI's major GPT-5 training cluster is in Arizona and runs GPUs at lower power to fit more chips.

    arizonagpt-5

    “And in some deployments, you'll be completely limited on power. And so there's many deployments. For example, OpenAI's major training cluster in Arizona that did GPT-five and such.”

    15:08 · Open Compute Project · 23 Oct 2025
  226. Patel says OpenAI and Meta run NVIDIA GPUs at lower power to fit 10% more chips despite worse TCO.

    gpumeta

    “Even though it's terrible on a TCO basis, they were able to get, you know, 10% more GPUs in, and it's great. And Meta has done similar.”

    15:23 · Open Compute Project · 23 Oct 2025
  227. Patel quantifies that 20% performance per watt advantage translates to only 4% TCO difference on NVIDIA deployments.

    nvidiapower efficiency

    “But capital is far more of a limiting factor. Because if you have enough power, a 20% difference in performance per watt only ends up being a 4% difference in TCO.”

    15:43 · Open Compute Project · 23 Oct 2025
  228. Patel calculates 20% power efficiency difference translates to only 4% TCO difference on NVIDIA deployments.

    inference economicspower efficiency

    “Because if you have enough power, a 20% difference in performance per watt only ends up being a 4% difference in TCO.”

    15:46 · Open Compute Project · 23 Oct 2025
  229. Patel claims NVIDIA takes a 5x markup on manufacturing cost, making power efficiency less significant for TCO.

    marginsnvidia

    “So, in most cases, don't on an NVIDIA deployment, right? That's where NVIDIA takes 5x markup on their manufacturing cost, right?”

    15:56 · Open Compute Project · 23 Oct 2025
  230. Patel states NVIDIA takes 5x markup on manufacturing cost, making power differences more significant for AMD deployments.

    marginsnvidia

    “If it's another deployment, if it's AMD, then that that that 20% power difference might translate to eight or 9% TCO difference when you when you talk about power cost and data center capacity cost.”

    16:03 · Open Compute Project · 23 Oct 2025
  231. Patel questions whether neo clouds targeting developers on short-term rentals can achieve ROI.

    business modelneoclouds

    “Those that are just targeting developers on short term rentals, they may not be able to get their ROI back.”

    27:03 · Open Compute Project · 23 Oct 2025
  232. Patel argues neo clouds with long-term enterprise and hyperscaler deals will succeed over on-demand providers.

    business modelhyperscalers

    “And so those are probably less likely to be able to succeed versus those who are locking in these massive deals with or long term deals that may not be massive but with enterprises or with AI labs, or with the hyperscalers who have no capacity because the demand is just so incredible.”

    27:10 · Open Compute Project · 23 Oct 2025
  233. Patel says the standard inference deployment unit has shifted from single nodes to hundreds of GPUs.

    gpu supply chaininference economics

    “the standard unit for an inference deployment being hundreds of GPUs instead of a single node. And then there's all these different things about traffic.”

    2:03 · Open Compute Project · 23 Oct 2025
  234. Patel notes NVIDIA Blackwell performance improved dramatically from launch to present, as did AMD hardware.

    gpu supply chaininference economics

    “If you tried to use NVIDIA's Blackwell six months ago, the numbers were not amazing, right? But now they're actually amazing. So how did that progress over time? Same with AMD, right?”

    2:49 · Open Compute Project · 23 Oct 2025
  235. Patel says InferenceMax runs daily benchmarks to track hardware performance improvements as software optimizes.

    benchmarkingmethodology

    “And why this is important is you can see the progress of hardware over time as the software stack gets more optimized.”

    4:58 · Open Compute Project · 23 Oct 2025
  236. Patel quantifies InferenceMax uses tens of millions of dollars in GPU hardware with daily software updates.

    benchmarkinginfrastructure

    “we have tens of millions of dollars of hardware of the current and last generation GPUs. As I mentioned before, the software updates every single day.”

    6:43 · Open Compute Project · 23 Oct 2025
  237. Patel breaks down inference TCO as 20% electricity and data center, 80% hardware on NVIDIA deployments.

    inference economics

    “20% of your cost is your electricity, your data center real estate, roughly. And then the rest of the cost is that hardware, at least on a standard NVIDIA deployment your GPUs, your networking, etcetera.”

    14:12 · Open Compute Project · 23 Oct 2025
  238. Patel identifies GB200's power efficiency advantage but notes deployment challenges with backplane and liquid cooling.

    gpu supply chaininference economics

    “GB200 has a huge power efficiency advantage, right? Everyone here understands the challenges of running and deploying GB200. There's a lot of challenges with the backplane.”

    15:07 · Open Compute Project · 23 Oct 2025
  239. Patel reports GB200 is 10x more power efficient than H200 at certain interactivity rates.

    gpu supply chaininference economics

    “But it turns out at certain interactivity rates, I. E. Tokens per second per user, it's 10x more efficient per watt, right, compared to H200.”

    15:19 · Open Compute Project · 23 Oct 2025
  240. Patel quantifies GB200 as 10x more power efficient than H200 at certain interactivity rates.

    benchmarkinggb200

    “at certain interactivity rates, I. E. Tokens per second per user, it's 10x more efficient per watt, right, compared to H200.”

    15:20 · Open Compute Project · 23 Oct 2025
  241. Patel finds AMD MI355 beats NVIDIA B200 on performance TCO in certain publicly usable configurations.

    gpu supply chaininference economics

    “So we do different scenarios. We do document processing, which is 8,000 context in, 1,000 out. We do chat, which is 1,000 in, 1,000 out.”

    16:11 · Open Compute Project · 23 Oct 2025
  242. Patel reports AMD MI355 beats NVIDIA B200 in some document processing scenarios with open source software.

    amdbenchmarking

    “in some cases, AMD actually does have a publicly usable open source implementation that beats NVIDIA even, right, with the MI355 versus V200, which is a surprise, right?”

    16:28 · Open Compute Project · 23 Oct 2025
  243. Patel reports AMD MI355 beats NVIDIA B200 on performance-TCO with publicly usable open source software in certain use cases.

    amdbenchmark results

    “AMD actually does have a publicly usable open source implementation that beats NVIDIA even, right, with the MI355 versus V200, which is a surprise, right?”

    16:30 · Open Compute Project · 23 Oct 2025
  244. Patel finds 10x performance TCO advantage for Blackwell over H100 at certain operating points.

    blackwellh100

    “But at other points, you can have a 10x performance TCO benefit from going with a more expensive Blackwell server than a cheaper H100 server.”

    17:25 · Open Compute Project · 23 Oct 2025
  245. Patel shows B200 has 15x raw performance advantage over H100 but only 10x performance per TCO.

    gpu supply chaininference economics

    “If we don't divide by TCO, then it looks like the performance of B200 is actually 15x that of H100, versus the performance TCO is only 10x,”

    18:34 · Open Compute Project · 23 Oct 2025
  246. Patel reports AMD performance improved dramatically over two months, similar to NVIDIA Blackwell's software optimization trajectory.

    amdnvidia

    “Over two months, the performance of AMD dramatically improved because they went from, hey, software problems are always a thing, to over a two month period, they've dramatically improved.”

    19:16 · Open Compute Project · 23 Oct 2025
  247. Patel notes early GPU systems had low volume but Supermicro continued engineering through generations anticipating future demand.

    gpu systemsmarket timing

    “What there there wasn't much volume. Right? In those initial systems. Right? In the early time frame. Yeah. Early time frame.”

    42:23 · Supermicro · 16 Oct 2025
  248. Patel asks why Supermicro kept engineering GPU systems when volume was low and CPU systems were easier to build.

    gpu systemsproduct development

    “Continue to increase density, continue to build all these new building blocks around cooling and sure, you you you had buildings blocks, but most of your volume was Intel x 86 CPU.”

    42:39 · Supermicro · 16 Oct 2025
  249. Patel states Supermicro's liquid cooling reduces Hopper server power from 10 kilowatts to 7-8 kilowatts versus air-cooled competitors.

    inference economicsliquid cooling

    “Everyone else's hopper servers h 100 air cooled. And so each server consumes, you know, 10 kilowatts almost. Right?”

    50:10 · Supermicro · 16 Oct 2025
  250. Patel calculates that without Supermicro's liquid cooling solution, servers would consume 30% more power.

    inference economicspower efficiency

    “So that's why Shibbol Micro dedicated deep cooling so aggressively. And if the liquid cooling solution from Super Micro didn't exist, then those servers would be consuming 30% more power.”

    50:56 · Supermicro · 16 Oct 2025
  251. Patel says without Supermicro's liquid cooling, xAI would consume 30% more power or deploy 30% fewer GPUs.

    data center capacityliquid cooling

    “And if the liquid cooling solution from Super Micro didn't exist, then those servers would be consuming 30% more power. Right?”

    51:02 · Supermicro · 16 Oct 2025
  252. Patel describes diagnostic challenges in AI data centers where failures occur across multiple infrastructure layers.

    data center operationsinference economics

    “And people have had problems where their data center wasn't ready and something stopped working. They don't know, is it the chip? No.”

    1:05:25 · Supermicro · 16 Oct 2025
  253. Patel says Blackwell cost is up 50-70% over Hopper but performance exceeds 2x, making it an obvious choice.

    blackwellgpu economics

    “while cost is only up, know, call it 50% or 60% or 70% versus Hopper, the performance is well north of 2x, right?”

    1:08 · Together AI · 3 Oct 2025
  254. Patel explains Blackwell introduces third memory tier within Tensor Cores requiring complex programming model for full performance.

    blackwellgpu architecture

    “Now with Blackwell, there's actually even a third tier of memory, which is memory within the Tensor Core.”

    2:47 · Together AI · 3 Oct 2025
  255. Patel quantifies Hopper performance improved 30-40% over 2024 with Blackwell showing similar gains in months.

    blackwellgpu performance

    “Even Hopper in January 24 to December 2024, you still had like a 30%, 40% performance improvement. And likewise for Blackwell, you've seen a similar sort of improvement, right?”

    4:26 · Together AI · 3 Oct 2025
  256. Patel predicts another 50-100% Blackwell performance improvement from software optimization of standard libraries.

    blackwellperformance

    “I see another 50% to doubling in the cards, just from the fact that there is so much low hanging fruit or performance left on the table if you are using the general open libraries, right? Whether it's cuDNN, VLM, SGLANG, all these other things,”

    5:02 · Together AI · 3 Oct 2025
  257. Patel predicts general-purpose Blackwell users could see another performance doubling over the next year from software improvements.

    blackwellperformance

    “for the general purpose user, yeah, you're going to have another doubling of performance potentially on the tables over the next year.”

    5:49 · Together AI · 3 Oct 2025
  258. Patel reveals NVIDIA doubled system-level testing time for Blackwell compared to Hopper generation.

    blackwellnvidia

    “NVIDIA learning from the issues on Hopper, with Blackwell, they actually doubled the time for system level test.”

    9:46 · Together AI · 3 Oct 2025
  259. Patel notes GB200 expanded NVLink from 8 to 72 GPUs and rack power jumped from 10 to 140 kilowatts.

    data centergb200

    “Now you have 72 GPUs. And if you go look at the rack, right? It's completely different, right? It's liquid cooled. It's an entire rack that consumes 140 kilowatts, whereas H100 servers consumed 10 kilowatts.”

    11:27 · Together AI · 3 Oct 2025
  260. Patel explains GB200's 72-GPU configuration creates reliability challenges compared to 8-GPU systems due to higher failure probability.

    gb200reliability

    “If the reliability of each GPU is the same, then when a single GPU fails in 72 GPUs, you have a much higher chance of something failing, right?”

    12:40 · Together AI · 3 Oct 2025
  261. Patel confirms OpenAI runs production inference on GB200 despite reliability requiring workloads handle 64 of 72 GPUs.

    gb200inference

    “OpenAI has said they're running production inference on GV200 a couple of months ago, in fact. Right?”

    13:42 · Together AI · 3 Oct 2025
  262. Patel contrasts traditional data centers with minimal retrofit costs against modern GPU deployments requiring major upgrades each generation.

    data centerinfrastructure

    “it used to be you'd build a data center and there wouldn't really be much retrofits cost, right? You would slide servers in and out, you would rent racks to people.”

    21:24 · Together AI · 3 Oct 2025
  263. Patel describes Meta's temporary tent-like data centers designed for single GPU generation lifecycles instead of 30-year buildings.

    data centerinfrastructure

    “Meta has got a data center design that's only built to last for a single generation. They put them in these tent buildings.”

    23:51 · Together AI · 3 Oct 2025
  264. Patel says B200 is better for training while GB200 is better for inference, reversing expected use.

    gb200inference

    “And so you've sort of Which is the exact opposite of what you would have expected. Oh, use the big thing for training and use the small thing for inference.”

    29:16 · Together AI · 3 Oct 2025
  265. Patel reveals NVIDIA's next generation will split inference into separate context processing and decode workloads, not training versus inference.

    architectureinference

    “NVIDIA's next generation actually has something very different. They're not saying, Hey, there's a training GPU and an inference GPU, right? Because either is fine.”

    30:59 · Together AI · 3 Oct 2025
  266. Patel explains NVIDIA is splitting inference into decode and prefill workloads, but provisioning for unknown future ratios is challenging.

    inferencenvidia

    “what sort of NVIDIA's pitching is like a split of inference into two workloads, and we'll see if they're successful. There's a lot of challenges on the infrastructure side.”

    32:53 · Together AI · 3 Oct 2025
  267. Patel says Huawei brought seven nanometer AI chips to market first in 2020 with minimal gap to NVIDIA.

    ai chipshuawei

    “In 2020, they released an Ascend chip and submit it to impartial public benchmarks, and they were the first to bring seven nanometer AI chips to market.”

    6:23 · a16z · 22 Sep 2025
  268. Patel says Huawei trained significant models on Ascend chips before 2020 Trump ban took full effect.

    ascendexport controls

    “The the full ban. And so they were only able to make a small volume of these chips, but they had trained significant models on these chips that they made then.”

    7:15 · a16z · 22 Sep 2025
  269. Patel reports Huawei acquired 2.9 million chips worth $500 million from TSMC through shell companies before getting caught.

    export controlshuawei

    “they were able to acquire 3,000,000 chips, 2,900,000 chips from TSMC through these other entities. Right?”

    7:55 · a16z · 22 Sep 2025
  270. Patel estimates NVIDIA had north of $20 billion in China H20 revenue that had to be written off after ban.

    chinaexport controls

    “Our our revenue estimate for NVIDIA in China for just h 20 was north of 20,000,000,000 because that's what they were booking in capacity slash had to write off.”

    8:36 · a16z · 22 Sep 2025
  271. Patel estimates NVIDIA had over $20 billion in H20 China revenue booked before ban and write-off.

    chinah20

    “Our our revenue estimate for NVIDIA in China for just h 20 was north of 20,000,000,000 because that's what they were booking in capacity slash had to write off. Yeah. And then it got banned.”

    8:36 · a16z · 22 Sep 2025
  272. Patel says China increased lithography imports to 40% of equipment spending to stockpile before bans took effect.

    chinaexport controls

    “They were importing lithography at a much higher rate than that. Right? Like, 40% of their equipment imports were lithography, and they were stockpiling lithography equipment.”

    16:07 · a16z · 22 Sep 2025
  273. Patel forecasts hyperscaler CapEx at $455-500 billion for next year versus Wall Street consensus of $360 billion.

    capexforecasts

    “The consensus for the banks is $360,000,000,000 of spend next year across all of them. And my number is closer to, like it's, like, $45,500.”

    23:09 · a16z · 22 Sep 2025
  274. Patel says OpenAI's Oracle deal reaches over $90 billion annually within a few years from $300 billion total commitment.

    gpu supplyopenai

    “OpenAI, which signed a $300,000,000,000 plus deal with with Oracle, will actually be able to pay $300,000,000,000, right, across raising capital and revenue?”

    24:28 · a16z · 22 Sep 2025
  275. Patel recounts that NVIDIA ordered Xbox production volume before receiving Microsoft's official order.

    nvidiarisk

    “No. No. No. Like, NVIDIA ordered the volume for the Xbox before Microsoft gave them the order.”

    30:26 · a16z · 22 Sep 2025
  276. Patel predicts AWS revenue growth will reaccelerate after consistent deceleration due to Anthropic and new data centers.

    anthropicaws

    “AWS has been decelerating revenue. Year on year revenue has been falling consistently. And and our big call is that it's actually going to start reaccelerating.”

    58:56 · a16z · 22 Sep 2025
  277. Patel explains Oracle's $300 billion OpenAI bet is lower risk because Oracle only secures data centers, not expensive GPUs upfront.

    data centersopenai

    “Now, of course, the bet is a bit like there's a bit more security in the bet in that Oracle really only needs to secure the data center capacity.”

    1:08:55 · a16z · 22 Sep 2025
  278. Patel compares GPU purchasing process to buying drugs, calling people to ask about availability and pricing.

    gpu supplymarket dynamics

    “my opinion on how you buy GPUs is that it's like buying cocaine or any other drug. This is described to me, not me.”

    1:35:36 · a16z · 22 Sep 2025
  279. Patel compares buying GPUs to buying cocaine with informal texts asking for availability and pricing.

    gpu supplymarket dynamics

    “You call up a couple people. You text a couple people. You ask, yo. How much you got? What's the price? It's like Exactly. This is fucking like, buy drugs. Like oh, sorry. Sorry.”

    1:35:48 · a16z · 22 Sep 2025
  280. Patel predicts Anthropic's revenue growth will match OpenAI's by 2027 based on current trajectory.

    anthropicopenai

    “Their revenue growth rate right now implies that they would be OpenAI sometime in 2027, which is a really, really big deal in terms of revenue.”

    0:04 · David Ondrej · 18 Sep 2025
  281. Patel dismisses possibility of unofficial MCP servers for sensitive company data like Slack or SAP.

    computer-useenterprise

    “There's not gonna be an unofficial MCP running out there. Right? Like, you know, it's like like, fuck off. Right? Like, that's just not gonna happen.”

    0:38 · David Ondrej · 18 Sep 2025
  282. Patel says the Oracle-OpenAI deal involves many gigawatts of compute with unprecedented four-year revenue guidance.

    computeopenai

    “It is literally many gigawatts. And then in addition, right, not just about securing OpenAI ridiculous amounts of compute, but what what was really interesting is Oracle gave a guidance of four years.”

    1:30 · David Ondrej · 18 Sep 2025
  283. Patel notes Oracle gave unprecedented four-year revenue guidance, something no company has ever done before.

    computeopenai

    “Oracle gave a guidance of four years. They told everyone what their revenue is gonna be for the next four years. No company has ever done that.”

    1:38 · David Ondrej · 18 Sep 2025
  284. Patel explains Microsoft's hesitation: OpenAI has only $14 billion ARR but wants $300 billion in compute.

    arrmicrosoft

    “How are you gonna how are you gonna pay for this? By the end of the year, it's like 14 right now. Right? Billion dollars on an ARR basis.”

    2:59 · David Ondrej · 18 Sep 2025
  285. Patel describes Anthropic as cult-like with more unified AGI vision and mission focus than OpenAI.

    agianthropic

    “They are a bit of a cult. Right? They've they believe in the mission much more than any other lab. Not that OpenAI people aren't cultists as well.”

    6:09 · David Ondrej · 18 Sep 2025
  286. Patel explains Anthropic's strategy prioritizes solving software engineering and computer use to accelerate AGI development recursively.

    agianthropic

    “First, we have to solve software engineering, and then we can use that to build all the other pieces. Right? First, we have to solve computer use.”

    6:31 · David Ondrej · 18 Sep 2025
  287. Patel notes Anthropic's seven cofounders have equal equity and left OpenAI with clear unified mission focus.

    anthropicequity

    “Anthropic, you know, sort of has you know, they had seven cofounders all with, like, equal equity. And these these cofounders were, like, very focused on exactly this mission.”

    7:04 · David Ondrej · 18 Sep 2025
  288. Patel contrasts Anthropic's multiple technical cofounders with OpenAI where Sam Altman is not technical.

    anthropicleadership

    “They've got cofounders who are extremely technical, whereas OpenAI, right, you know, they they they don't have necessarily like, Sam's not super technical. You know, the obvious Greg Greg is a cofounder, he's super technical.”

    7:29 · David Ondrej · 18 Sep 2025
  289. Patel notes YouTube holds unique video data that gives Google training advantage despite scraping restrictions.

    googletraining data

    “I don't know what ungodly percentage of video is only on available on YouTube. And sure, you can scrape YouTube, blah blah blah.”

    37:13 · David Ondrej · 18 Sep 2025
  290. Patel states GB 200 servers from Nvidia cost well over $3 million, preventing time-slicing users on individual servers.

    gpu economicsnvidia

    “These servers cost hundreds of thousands of dollars if nothing new. Know GB 200 from Nvidia cost well over $3,000,000 so the cost of these things is huge”

    4:12 · Nebius · 2 Sep 2025
  291. Patel notes new NVIDIA servers require 140 kilowatts per rack versus prior CPU data centers at kilowatts.

    data centersnvidia

    “Data centers for CPUs, 10 megawatts of rack, 12 megawatts of rack. Kilowatts. The new ones, kilowatts, right? Kilowatts of rack, sorry. The new NVIDIA servers, 140 kilowatts of rack.”

    16:14 · Nebius · 2 Sep 2025
  292. Patel reports 10 gigawatts of AI infrastructure deploying in the US over next 18-24 months.

    capacity forecastinfrastructure buildout

    “In The US collectively there's like 10 gigawatts going up over the next year and a half, two years. And Europe is trying to scale to gigawatt deployments as well.”

    28:26 · Nebius · 2 Sep 2025
  293. Patel states each gigawatt AI deployment costs tens of billions of dollars end-to-end.

    costsgigawatt scale

    “Each of these gigawatt deployments once you take it from all the way from the power to the GPUs and the cloud contract, those are tens of billions of dollars each gigawatt.”

    28:43 · Nebius · 2 Sep 2025
  294. Patel lists NVIDIA's structural advantages across networking, HBM, process nodes, speed to market, and supplier negotiations.

    gpu supply chainhbm

    “NVIDIA's gonna have better networking than you. They're gonna have better, HBM. They're gonna have better process node. They're gonna come to market faster.”

    0:00 · a16z · 18 Aug 2025
  295. Patel argues competing with NVIDIA requires leaping forward beyond supply chain advantages across networking, memory, manufacturing and components.

    competitiongpu supply chain

    “They're gonna have better negotiations with whether it's TSMC or SK Hynix and the memory and silicon side or all the rack people or, like, copper cables, everything, they're gonna have better cost efficiency.”

    0:06 · a16z · 18 Aug 2025
  296. Patel explains OpenAI's router will send low-value queries to cheaper models but use expensive compute for high-value monetizable queries.

    inference economicsmonetization

    “if the user asks a low value query like, hey, why is the sky blue? Just route them to mini. The model can answer perfectly fine, and that is a chunk of queries.”

    6:06 · a16z · 18 Aug 2025
  297. a16z, 18 August 2025

    agentsmonetization

    “But if they ask, what's the best DUI lawyer near me? All of a sudden, this is like, you're in jail, you have one shot, you're like, screw it.”

    6:17 · a16z · 18 Aug 2025
  298. Patel reports 10% of Etsy traffic comes from ChatGPT but OpenAI currently makes nothing from it.

    chatgptetsy

    “Etsy. 10% of their traffic now comes from chat, and OpenAI makes nothing off of that. But they very they really, really will soon. Right?”

    6:52 · a16z · 18 Aug 2025
  299. Patel says OpenAI and Anthropic are receiving 30% of all GPU chips being produced this year.

    anthropicgpu supply chain

    “30% of the chips are going to them, just those two companies. But that's actually like, okay, well, 70% of the stuff, who's making off well, one third of it is ads, whether it be ByteDance or Meta or many of the other people who are doing ads.”

    14:52 · a16z · 18 Aug 2025
  300. Patel argues OpenAI captures less than 10% of the economic value it creates, pointing to broken value capture across AI.

    monetizationopenai

    “I legitimately believe OpenAI is not even capturing 10% of the value they've created in the world already just by usage of chat. Right?”

    18:21 · a16z · 18 Aug 2025
  301. Patel explains competitors need 5x hardware advantage over NVIDIA but risk failure if AI workloads shift before shipping.

    competitiongpu supply chain

    “So now you need to do something, you know, that will give you five x advantage, right, in hardware efficiency for a certain type of workload, and then pray the workload doesn't shift.”

    32:38 · a16z · 18 Aug 2025
  302. Patel forecasts NVIDIA revenue exceeding $300 billion next year as infrastructure spending reaches nation-state scale.

    capexforecasts

    “NVIDIA's revenue this year is gonna be, like, over $200,000,000,000, and next year expects over 300,000,000,000 plus Google's gonna spend, like, $50,000,000,000 on TPU data centers.”

    43:58 · a16z · 18 Aug 2025
  303. Patel says GPU hardware and networking represent 80% of data center costs, with power and cooling only 20%.

    data centersgpu supply chain

    “It's the GPU purchases. It's the networking. It's the it's the physical data center conversion power conversion equipment. All of this stuff is, like, 80% of the cost.”

    47:34 · a16z · 18 Aug 2025
  304. Patel states 80% of GPU data center cost is capital equipment, only 20% is land, power, and cooling.

    data centerseconomics

    “It's the it's the physical data center conversion power conversion equipment. All of this stuff is, like, 80% of the cost. And then 20% is gonna be your land and your power and your cooling”

    47:36 · a16z · 18 Aug 2025
  305. Patel reports new Trump tax bill allows year-one GPU depreciation, worth $10 billion annually to Meta alone.

    infrastructuremeta

    “the new the new Trump, you know, tax bill institutes something really incredible, which is that you can depreciate all of the GPU cluster cost in year one, which we put out, like, a note about how, like, the tax implications to, like, meta are, like, $10,000,000,000 a year. And across each of the major hyperscalers, it's, like, massive.”

    56:19 · a16z · 18 Aug 2025
  306. Patel suggests NVIDIA should use its massive tax bill to invest directly in data center infrastructure despite customer conflicts.

    data centersnvidia

    “Now this is obviously gonna be, like, crazy because, like, now they're buying g p their own GPUs and putting them in data centers and doing stuff, and they're competing with their own customers,”

    56:47 · a16z · 18 Aug 2025
  307. Patel reveals DeepSeek inference implementation requires 160 GPUs worth over $10 million of hardware per replica.

    deepseekgpu

    “That's over $10,000,000 of hardware, and then that's just one replica, then you'll have a lot of replicas and you share the caching servers between them.”

    6:11 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  308. No Priors: AI, Machine Learning, Tech, & Startups, 14 August 2025

    anthropicapi revenue

    “So this is something that I've found very interesting is that we've been trying to build a lot of alternative data sources for token usage, who's using tokens, what models, where, etcetera, why, and it's very clear that people aren't actually using the reasoning models that much in API. Anthropic has eclipsed OpenAI and API revenue,”

    8:51 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  309. Patel reports alternative data shows reasoning models have low actual API usage despite availability.

    api usagemarket research

    “So this is something that I've found very interesting is that we've been trying to build a lot of alternative data sources for token usage, who's using tokens, what models, where, etcetera, why, and it's very clear that people aren't actually using the reasoning models that much in API.”

    8:51 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  310. Patel claims Anthropic has overtaken OpenAI in API revenue, primarily from Claude 4 non-reasoning mode for code use cases.

    anthropicapi revenue

    “Anthropic has eclipsed OpenAI and API revenue, and their API revenue is primarily not thinking. It's Cloud four, but it's not in the thinking mode, code being the biggest use case that's skyrocketing.”

    9:07 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  311. Patel explains AI hardware companies made architectural bets that failed when model architectures evolved in unexpected directions.

    ai hardwarearchitecture

    “But still the model's way too big to fit on it. This is, like, very simple. Right? You know, the same thing's happening in the other direction.”

    22:51 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  312. Patel reveals Meta is building temporary tent structures to house GPUs because permanent buildings take too long to construct.

    data centersinfrastructure

    “Meta is literally building these, like, temporary, like, tent structures to put GPUs in because building the building takes too long, and it takes too much labor, as you mentioned labor.”

    31:14 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  313. Patel describes former Yandex workers paid bonuses for speed who use drugs to complete data center buildouts faster.

    data centersinfrastructure

    “Like and they get paid bonuses for being faster, and therefore, do, like, certain drugs to be able to finish the build outs faster.”

    32:32 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  314. Patel questions xAI's valuation exceeding Anthropic's despite no leading model, crediting only their fast Colossus infrastructure build.

    anthropicvaluation

    “What has xAI actually done to deserve their prior funding rounds? They haven't released a leading edge model, and yet their evaluation's higher than Anthropic today.”

    33:43 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  315. Patel questions xAI's valuation exceeding Anthropic's despite lacking a leading model, crediting infrastructure execution over product.

    anthropicvaluations

    “They haven't released a leading edge model, and yet their evaluation's higher than Anthropic today. At least Anthropic's racing, It's Elon, A, and B, they've tackled a problem creatively”

    33:47 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  316. No Priors: AI, Machine Learning, Tech, & Startups, 14 August 2025

    chinaexport controls

    “But, like, you know, can you give them no GPUs? No. They're gonna retaliate. Like, there is a middle ground, and, like, Huawei is eventually going to have a lot of production capacity,”

    40:03 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025
  317. Patel explains that Apple's rushed model has empty experts not receiving tokens due to flawed sparse MOE routing.

    applemodel-architecture

    “Basically in between every layer, the router can route to whatever expert it wants to, And it learns which expert to route to. And and each expert learns its own independent things”

    2:38 · Matthew Berman · 30 Jun 2025
  318. Patel explains how MOE routing works and how some of Meta's experts weren't being used at all.

    metamodel architecture

    “And and each expert learns its own independent things and it's like really not something observable by people. But what you can see is tokens when which experts do they route to?”

    2:45 · Matthew Berman · 30 Jun 2025
  319. Patel reports Google is backing out of Scale AI, spending around $250 million this year but reducing future commitment.

    googlepartnerships

    “Google's backing out, I think they're gonna spend on the order I've heard like $250,000,000 this year with them. And they're backing out.”

    6:51 · Matthew Berman · 30 Jun 2025
  320. Patel reports Google is spending $250 million with Scale AI this year but backing out of the partnership.

    googlepartnerships

    “And they're backing out. Obviously, they've spent a lot of money and there's stuff they can't back out of, but it's like, that's gonna go down a lot.”

    6:56 · Matthew Berman · 30 Jun 2025
  321. Patel reports OpenAI cut the Slack connection with Scale AI.

    openairelationships

    “OpenAI allegedly, like, cut the external Slack connection. Right? So there's no, like, slack between scale and OpenAI anymore.”

    7:02 · Matthew Berman · 30 Jun 2025
  322. Patel identifies a major shift in Zuckerberg's strategy toward pursuing superintelligence recently.

    metastrategy

    “He was chasing like AI is good and great, but like AGI is not a thing that is gonna happen soon. So sort of this is a big shift in strategy”

    8:30 · Matthew Berman · 30 Jun 2025
  323. Patel directly contradicts Sam Altman's claim that no top researchers have left OpenAI for Meta.

    metaopenai

    “As far as Sam is saying that no top researchers have gone, I don't believe that's accurate. I initially the top researchers definitely did say no, the best researchers, the best people.”

    14:49 · Matthew Berman · 30 Jun 2025
  324. Patel claims Meta offered over $1 billion to one OpenAI researcher as retention or acquisition offer.

    compensationmeta

    “And you said $100,000,000 I've heard a number for someone over 1,000,000,000 actually, for one person at OpenAI. But anyways, you know, it's a ridiculous amount of money,”

    15:03 · Matthew Berman · 30 Jun 2025
  325. Patel reports hearing Meta offered over one billion dollars to a single OpenAI researcher.

    compensationmeta

    “I've heard a number for someone over 1,000,000,000 actually, for one person at OpenAI. But anyways, you know, it's a ridiculous amount of money,”

    15:04 · Matthew Berman · 30 Jun 2025
  326. Patel reveals OpenAI had a PyTorch bug in GPT-4.5 training code for months during the training run.

    bugsopenai

    “And they had a bug in the training code for a couple months that was like a very tiny bug that like was messing up the training.”

    25:08 · Matthew Berman · 30 Jun 2025
  327. Patel describes the Bumpgate incident where NVIDIA laptop GPUs had defective solder balls.

    applebumpgate

    “There's a generation of NVIDIA GPUs for laptops. Right? And and chips have solder balls on the bottom. Right?”

    31:43 · Matthew Berman · 30 Jun 2025
  328. Patel explains Bumpgate issue where thermal expansion differences caused GPU solder ball failures in Apple laptops.

    applehardware

    “And what ended up happening is because of that different rate of expansion, the solder balls connecting the chip and the board would crack.”

    32:25 · Matthew Berman · 30 Jun 2025
  329. Patel reports AMD is selling GPUs to cloud providers then renting them back to inflate sales numbers.

    accountingamd

    “Like but AMD is actually doing this and like taking it to overdrive. They're getting clusters at Oracle and Amazon and Crusoe and Digital Ocean and Tensorwave and they're renting GPUs back from them.”

    46:27 · Matthew Berman · 30 Jun 2025
  330. Patel predicts China will stop open sourcing AI models once ahead and closed source will ultimately win.

    chinaopen-source

    “China's open sourcing stuff only because they're behind the moment they're ahead, they will stop open sourcing stuff. And at the end of the day, closed source will win. Unfortunately, closed source will win.”

    1:00:37 · Matthew Berman · 30 Jun 2025
  331. Patel asserts America cannot manufacture robots and Germany/Japan are losing capabilities to China.

    chinamanufacturing

    “Germany and and Japan are losing a lot of their capabilities to do so slash the cost of their robots are just way, way, way higher than what China can do.”

    0:02 · Going Direct · 22 Jun 2025
  332. Patel reports AI models improved from under 4% to 70% on Frontier Math in six months.

    ai modelsbenchmarks

    “Like, can't do all the math, all the problems in frontier math. And the models went from sub 4% to 70% in six months.”

    12:41 · Going Direct · 22 Jun 2025
  333. Patel reports AI models improved from below 4% to 70% accuracy on Frontier Math in six months across multiple companies.

    ai modelsbenchmarks

    “frontier math. And the models went from sub 4% to 70% in six months. Right? And this is from not just OpenAI, but also Anthropic and Google and DeepSeek.”

    12:46 · Going Direct · 22 Jun 2025
  334. Patel states that Japan, USA, Korea, and Germany combined deploy only a third of China's robots in 2023.

    chinamanufacturing

    “You fast forward to 2023, and all four of these countries combined, right, are still a third of China's robot deployments.”

    31:35 · Going Direct · 22 Jun 2025
  335. Patel notes that as recently as 2022, China imported the majority of its robots from Japan and Germany.

    chinarobotics

    “even as near as twenty twenty two, the majority of China's robots were being imported. Right? Japan and Germany being the largest market share players,”

    32:18 · Going Direct · 22 Jun 2025
  336. Patel notes China was importing majority of robots as recently as 2022 from Japan and Germany.

    chinarobotics

    “even as near as twenty twenty two, the majority of China's robots were being imported. Right? Japan and Germany being the largest market share players, but also USA, Korea is also in there.”

    32:18 · Going Direct · 22 Jun 2025
  337. Patel forecasts that 70% of robots deployed in China in 2025 will be domestically manufactured Chinese robots.

    chinaforecast

    “'24, it's, it's rapidly growing Chinese share. And '25 is it looks clear, like, the the market shares is, you know, gonna be 60 is gonna be, like, 70% Chinese robots for China's domestic usage.”

    32:36 · Going Direct · 22 Jun 2025
  338. Patel reveals that China purchased two top German robotics companies and transferred their servo and actuator production to China.

    acquisitionschina

    “Two of Germany's best robotics companies were actually purchased by China, And a lot of their, servos and actuator construction was actually transferred to China.”

    33:04 · Going Direct · 22 Jun 2025
  339. Patel contrasts Western price-to-value strategy with Chinese cost-plus pricing model in manufacturing.

    chinamanufacturing

    “Whereas Chinese companies are almost all cost plus. They're not price to value. They're cost plus. Right? And so you've got this interesting phenomenon where,”

    40:46 · Going Direct · 22 Jun 2025
  340. Patel explains Chinese companies use cost-plus pricing instead of price-to-value, avoiding margin expansion despite market dominance.

    chinamanufacturing

    “Whereas Chinese companies are almost all cost plus. They're not price to value. They're cost plus. Right?”

    40:46 · Going Direct · 22 Jun 2025
  341. Patel explains Chinese manufacturers use cost-plus pricing rather than value-based pricing across industries.

    chinamanufacturing

    “They're not price to value. They're cost plus. Right? And so you've got this interesting phenomenon where, there are industries in China where the entire world relies on it.”

    40:51 · Going Direct · 22 Jun 2025
  342. Patel describes Chinese manufacturers as locked in constant cost reduction cycle rather than margin expansion.

    chinaeconomics

    “there's this constant pressure cooker of continual cost reductions and engineering. It's not like, Oh, well, we make 80% margins. Our cost is low.”

    41:34 · Going Direct · 22 Jun 2025
  343. Patel notes many Chinese companies dominating global markets have flat or declining stock prices.

    chinamanufacturing

    “there's so many companies on the Hangshan, the Hangshan, which is a stock market in China, right, whose stock has literally been flat or down despite the fact that they own the global market.”

    41:52 · Going Direct · 22 Jun 2025
  344. Patel says most people will respond to outreach but almost nobody tries.

    adviceinternet

    “One of things I learned early on is you can talk to anyone on the Internet. You can just reach out to anyone, and most people just say yes, but no one does it.”

    0:40 · Origins with Sonith · 10 Jun 2025
  345. Patel advises becoming a foremost expert in what you enjoy and reaching out to anyone online.

    career advicelearning

    “Or try and talk to those people because one of the things I learned early on is you can talk to anyone on the Internet,”

    44:37 · Origins with Sonith · 10 Jun 2025
  346. Patel says his business model shifted so that 95% of revenue now comes from selling datasets.

    business modeldata sales

    “the business transformed into, you know, selling data, right, or build building datasets and selling them. Right? And when the business truly flipped to that, we're now 95% of the revenue is that data”

    1:04:52 · Origins with Sonith · 10 Jun 2025
  347. Patel says his company now employs 26 people distributed across seven countries globally.

    global operationshiring

    “There's 26 people who work at the company now. Cool. All over the world. Right? US, Japan, Taiwan, Singapore, France, Germany Uh-huh. Canada. Yeah. Yeah. Well, that'll be one less country. Right? He's gone.”

    1:06:03 · Origins with Sonith · 10 Jun 2025
  348. Patel finds it surreal that billionaires care about his semiconductor analysis.

    billionairesinfluence

    “it's it's it's surreal to see, like, billionaires, like, actually care what you have to say or rather what you, like, have figured out.”

    1:11:45 · Origins with Sonith · 10 Jun 2025
  349. Patel reflects on the surreal experience of billionaires caring about his research and analysis.

    billionairescredibility

    “it's surreal to see, like, billionaires, like, actually care what you have to say or rather what you, like, have figured out.”

    1:11:45 · Origins with Sonith · 10 Jun 2025
  350. Patel reports SemiAnalysis has 200,000 newsletter subscribers with 5,000 paid subscribers representing under 5% of revenue.

    business metricsnewsletter

    “So the newsletter itself is, like, 200,000. Paid newsletter is, like, I think, like, 200,000 people. Paid newsletter is, like, 5,000. Nice. But then, like, the main bit like, that's less than 5% of the”

    1:11:54 · Origins with Sonith · 10 Jun 2025
  351. Patel argues most AI value accrues to users, not to companies building or deploying AI.

    ai economicsinvestment

    “But most of the AI value that's been generated does not accrue to the companies building the AI or deploying it. It actually deploy it actually accrues to the user.”

    1:15:22 · Origins with Sonith · 10 Jun 2025
  352. Patel notes CoreWeave got H100 services online six months before every hyperscaler in many cases.

    gpu supply chainh100

    “Get services online with h one hundreds before every hyperscaler by a factor of, like, as much as, like, six months in many cases.”

    23:35 · CoreWeave · 23 May 2025
  353. Patel notes CoreWeave has fully liquid cooled data centers while hyperscalers are still working on designs.

    competitive advantageinfrastructure

    “They're doing full liquid cooling as well and they're working on those designs, but, you guys have fully liquid cooled data centers.”

    35:34 · CoreWeave · 23 May 2025
  354. Patel says CoreWeave can build $10 billion clusters from start to completion much faster than competitors.

    data centersgpu supply chain

    “it lets you get to these, you know, call it $10,000,000,000 clusters, right, in in time scales, from project start to completion that are much much shorter than anyone else.”

    38:16 · CoreWeave · 23 May 2025
  355. Patel observes hyperscaler data centers started in 2023 remain incomplete even with expedited equipment delivery.

    constructiondata centers

    “Even if they even if they had the electrical equipment in in you know, not on lead time, they paid a bunch to skip the queue.”

    38:42 · CoreWeave · 23 May 2025
  356. Patel reveals his team found security flaws in other clouds allowing them to snoop customer traffic via InfiniBand.

    cloud infrastructureinfiniband

    “of work. This is this is actually some stuff that we've noticed on other clouds that we've tested, whether it's, you know, things related to InfiniBand p keys and we're able to snoop in other nodes of real live customer traffic. And, obviously, we tell them, like, we're like, yo, this is not okay.”

    43:29 · CoreWeave · 23 May 2025
  357. Patel describes finding security vulnerabilities on other clouds allowing access to customer traffic and server BMCs.

    securitytesting

    “things related to InfiniBand p keys and we're able to snoop in other nodes of real live customer traffic.”

    43:34 · CoreWeave · 23 May 2025
  358. Patel describes a 10,000+ GPU customer whose engineers preferred CoreWeave infrastructure over their acquirer's hyperscaler setup.

    coreweavecustomer satisfaction

    “What's what's funny is there was a customer that had more than 10,000 GPUs with you. They got acquired. They had to start using some clusters at their acquiring company, the company that acquired them.”

    50:05 · CoreWeave · 23 May 2025
  359. Patel describes customer with over 10,000 GPUs whose engineers demanded CoreWeave infrastructure after acquisition.

    customer preferencegpu scale

    “What's what's funny is there was a customer that had more than 10,000 GPUs with you. They got acquired.”

    50:05 · CoreWeave · 23 May 2025
  360. Patel illustrates that language models historically failed at basic math comparisons without reasoning capabilities.

    mathmodel limitations

    “And it would say, yes. It's bigger. Even though, like, everyone knows that nine dot one one is is way smaller than nine dot nine.”

    24:35 · Alex Kantrowitz · 23 Apr 2025
  361. Patel reports GPT-3 to Llama 3.2 costs fell 1,200x while GPT-4 to DeepSeek v3 costs fell 600x.

    cost trendsdeepseek

    “when we looked at g p d three, the cost fell 1,200 x from g p d three's initial cost to what you can get Lama 3.23 b today.”

    29:02 · Alex Kantrowitz · 23 Apr 2025
  362. Patel argues the surprise was a Chinese company achieving this cost reduction, not the reduction itself.

    chinacompetition

    “I think what was really surprising was that it was a Chinese company for the first time. Right? Because Google and and OpenAI and Anthropic and Meta have all traded blows.”

    29:36 · Alex Kantrowitz · 23 Apr 2025
  363. Patel says DeepSeek's cost efficiency follows the expected trend line, just from an unexpected source.

    chinadeepseek

    “It's not unexpected. Right? Like, this is actually within the trend line of what happened with GPT three is happening to GPT four level quality with DeepSeq.”

    30:06 · Alex Kantrowitz · 23 Apr 2025
  364. Patel predicts Meta's upcoming Llama will achieve similar cost decreases to DeepSeek v3 without causing alarm.

    cost trendsllama

    “And Meta's Meta's gonna release their new llama soon enough. Right? And that one is gonna be, you know, a similar level of cost decrease,”

    30:26 · Alex Kantrowitz · 23 Apr 2025
  365. Patel argues that efficiency gains alone without capability increases cannot justify massive AI infrastructure investments.

    economicsinfrastructure

    “That would not pay for all of these build outs. Right? AI is useful today, but it's not capable of doing a lot of things.”

    33:53 · Alex Kantrowitz · 23 Apr 2025
  366. Patel states $10 billion data centers target automated software engineering, not chat models.

    capexscaling

    “So no one is trying to make with these $10,000,000,000 data centers, they're not trying to make chat models. Right?”

    34:22 · Alex Kantrowitz · 23 Apr 2025
  367. Patel reports that OpenAI's Orion training run did not improve enough over GPT-4.5 to qualify as GPT-5.

    gpt-5openai

    “There were hopes that Orion could be used for for GPT five, but its improvement was, like, not enough to be, like, really a GPT five.”

    36:12 · Alex Kantrowitz · 23 Apr 2025
  368. Patel explains GPT-5 will combine massive pre-training and post-training scale for the first time.

    gpt-5openai

    “Along the way, another team at OpenAI made the big breakthrough of reasoning. Right? Strawberry training. And they released o one, and then they released o three.”

    36:38 · Alex Kantrowitz · 23 Apr 2025
  369. Patel says GPT-5 will combine massive pre-training like GPT-4.5 with massive post-training like o1 and o3.

    gpt-5openai

    “GPT five, as Sam calls it, is is gonna be a model that has huge pre training scale, right, like GPT 4.5, but also huge post training scale like o one and o three and continuing to scale that up.”

    36:51 · Alex Kantrowitz · 23 Apr 2025
  370. Patel says GPT-5 will be first model to massively scale both pre-training and post-training simultaneously.

    gpt-5model architecture

    “Right? And this would be the first time we see a model that was was a step up in both at the same time.”

    37:03 · Alex Kantrowitz · 23 Apr 2025
  371. Patel reports GPT-5 is expected in three to six months, possibly sooner based on his sources.

    gpt-5openai

    “They say it's coming, you know, this year, hopefully, the next three to six months, maybe sooner. I've heard sooner, but, you know, we'll we'll see.”

    37:12 · Alex Kantrowitz · 23 Apr 2025
  372. Patel says certain states will see over half their power consumed by AI within a few years.

    data centersinfrastructure

    “When we look at power consumption by state, something pretty simple, there are certain states where more than half their power is going be consumed by AI just in the next few years.”

    8:56 · MedBricks Webcast · 27 Mar 2025
  373. Patel says reasoning capability allowing models to think before answering dramatically increased capabilities in six months.

    ai capabilitiesreasoning models

    “Right? And this has just happened in six months. And the same applies to various certain coding tasks. The same applies to yeah.”

    25:28 · MedBricks Webcast · 27 Mar 2025
  374. Patel says AI model costs dropped 1200x from GPT-3's $60 per million outputs to current pricing.

    inference economicsmodel costs

    “GPT three was it cost, you know, it cost $60 for the million output. And over time, that kept reducing. Right? OpenAI released new models.”

    30:53 · MedBricks Webcast · 27 Mar 2025
  375. Patel argues existing AI capabilities get cheaper but GPT-4 level models still cannot do medical tasks.

    ai limitationsmedical ai

    “But this is the progression you get because of the advancement. That's one way to look at it is, oh, existing capabilities get cheaper. But as useful as GPT-four is, it can't do medical tasks.”

    31:42 · MedBricks Webcast · 27 Mar 2025
  376. Patel says GPUs in a 16,000 GPU cluster show up to 10% speed variation from manufacturing differences.

    gpu performancemanufacturing

    “But even within a cluster of like 16 k GPUs, GPUs will be slower by up to 10%. Right? So there there's a lot of variation even in the manufacturing.”

    2:45 · Prime Intellect AI · 25 Mar 2025
  377. Patel reports Meta's 16,000 GPU cluster had failures every few hours while next-generation models train on 100,000 GPUs or more.

    failuresgpu reliability

    “And that was a measly 16,000 GPUs. Right? And when you look at like the next generation Llamas, the next generation GPTs, etcetera, they're training on a 100,000 GPUs or more.”

    4:55 · Prime Intellect AI · 25 Mar 2025
  378. Patel explains training a 70 billion parameter model requires passing four times that data every training step without overlapping communication.

    communicationmodel architecture

    “So if you're training like a 70,000,000,000 parameter model, you need to pass four x that number every single training step. And when you're passing that, you can't really overlap communications and compute.”

    6:12 · Prime Intellect AI · 25 Mar 2025
  379. Patel says cloud pricing varies by three times due to poor marketing, sales, and stranded compute capacity.

    cloud economicsmarket dynamics

    “Like the the cost difference between different clouds can be as much as three x. Right? Because a lot of these clouds are bad at marketing, bad at sales, stranded compute.”

    10:15 · Prime Intellect AI · 25 Mar 2025
  380. Patel argues $52 billion Chips Act is insufficient to restore US to even 20% of global semiconductor production.

    chips actmanufacturing

    “S. To even 20% of global production, let alone a higher number, which is what we'd like, going back to maybe the 80s levels, right?”

    2:33 · Special Competitive Studies Project · 13 Mar 2025
  381. Patel argues TSMC's $100 billion US investment still relies on Taiwan R&D, not securing supply chain.

    chips actsupply chain

    “Even if they do fully do the full $100,000,000,000 they commit to, completely still relying on R and D in Taiwan.”

    3:31 · Special Competitive Studies Project · 13 Mar 2025
  382. Patel argues TSMC's $100 billion US investment won't secure R&D supply chain or new node development.

    chips actsupply chain

    “The R and D fab they announced is not going to secure R and D supply chain, right? It's not going to secure new node development, right?”

    3:38 · Special Competitive Studies Project · 13 Mar 2025
  383. Patel says TSMC approaches 20% of Taiwan's power consumption, projecting 30-40% within five years.

    powertaiwan

    “And if they were to continue to build out what we expect them to build out in the next four or five years, it's going get to 30%, 40%.”

    4:52 · Special Competitive Studies Project · 13 Mar 2025
  384. Patel reports ASML had 45% of equipment and more of revenue from China despite export controls.

    asmlchina

    “On equipment though, last year ASML, many quarters they've had 45% of their equipment and even more of their revenue come from China.”

    11:17 · Special Competitive Studies Project · 13 Mar 2025
  385. Patel reports ASML had 45% of equipment and more of revenue from China despite export controls.

    asmlchina

    “On equipment though, last year ASML, many quarters they've had 45% of their equipment and even more of their revenue come from China. And likewise for ASML, Lam Research and all these companies, right?”

    11:17 · Special Competitive Studies Project · 13 Mar 2025
  386. Patel challenges DeepSeek's GPU count claims, citing job ads promising tens of thousands of GPUs.

    chinadeepseek

    “The estimates that we have is, so first of all, ads in China, they say they have tens of thousands of GPUs for researchers, right?”

    14:14 · Special Competitive Studies Project · 13 Mar 2025
  387. Patel identifies a 40 gigawatt power shortfall for US data centers even with all planned capacity additions.

    data centersinfrastructure

    “There is an over 40 gigawatt hole in power for The US. Even if we don't turn off any more coal, in fact we restart some, we take all the natural gas dual combined cycle reactor deliveries that we're going to have, and you add in renewables, you have a 40 gigawatt hole of power production in The US.”

    18:07 · Special Competitive Studies Project · 13 Mar 2025
  388. Patel identifies a 40 gigawatt power deficit in US even with all planned gas and restarted coal.

    data centersinfrastructure

    “Even if we don't turn off any more coal, in fact we restart some, we take all the natural gas dual combined cycle reactor deliveries that we're going to have,”

    18:12 · Special Competitive Studies Project · 13 Mar 2025
  389. Patel says current models use 100,000 GPUs while next generation will require hundreds of thousands or millions.

    gpu clustersscaling

    “And next generation models that are trained on hundreds of thousands or even millions GPUs, right?”

    19:49 · Special Competitive Studies Project · 13 Mar 2025
  390. Patel says NVIDIA made 4 million GPUs last year and will produce 7 million this year.

    gpunvidia

    “Nvidia made over 4,000,000 GPUs last year, they're making over 7,000,000 this year, right? High end data center GPUs.”

    19:58 · Special Competitive Studies Project · 13 Mar 2025
  391. Patel reports China imported one million H20 GPUs in recent quarters, enough for largest cluster.

    chinagpu

    “Just in Q3, Q4 and the early part of Q1 this year, they imported a million H20s, Right?”

    20:18 · Special Competitive Studies Project · 13 Mar 2025
  392. Patel says China imported one million H20 GPUs in recent quarters, enough to build largest cluster.

    chinaexport controls

    “This is more than enough if they even took 30 of it, concentrated it, to be a bigger cluster than any American company has.”

    20:24 · Special Competitive Studies Project · 13 Mar 2025
  393. Patel explains DeepSeek v3 base is trained once, then post-trained differently to create chat versus reasoning models.

    deepseekmodel architecture

    “This reasoning model has a lot of overlapping training steps to DeepSeek v three, and it's confusing that you have a base model called v three that you do something to to get a chat model, and then you do some different things to get a reasoning model.”

    4:38 · Lex Fridman · 3 Feb 2025
  394. Patel asserts data quality is the primary determinant of model quality, not just compute or architecture.

    model qualitytraining data

    “the data processing, data filtering, data quality is the number one determinant of the model quality, and then a lot of the training code is the determinant on how long it takes to train and how fast your experimentation is.”

    7:03 · Lex Fridman · 3 Feb 2025
  395. Patel says DeepSeek modifies code at or below NVIDIA's CUDA layer, a rare technical capability.

    cudadeepseek

    “For example, on their to get highly efficient training, they're making modifications at or below the CUDA layer for NVIDIA chips.”

    10:03 · Lex Fridman · 3 Feb 2025
  396. Patel explains post-training involves instruction tuning for formatting and preference fine-tuning derived from RLHF.

    instruction tuningpost-training

    “One, I will classify as preference fine tuning. Preference fine tuning is a generalized term for what came out of reinforcement learning from human feedback, which is RLHF.”

    16:47 · Lex Fridman · 3 Feb 2025
  397. Patel states mixture of experts architecture can achieve the same model performance with 30% less compute.

    compute savingsmixture of experts

    “So you can get effectively the same performance model and evaluation scores with numbers like 30% less compute.”

    29:48 · Lex Fridman · 3 Feb 2025
  398. Patel notes DeepSeek's routing innovation removing auxiliary loss represents compounding small improvements over time.

    architecturedeepseek

    “this type of change can be big, it can be small, but they add up over time.”

    37:24 · Lex Fridman · 3 Feb 2025
  399. Patel argues export controls primarily limit AI inference deployment in China, not frontier model training capabilities.

    chinaexport controls

    “A large part of export controls, if they work, is just that the amount of AI that can be run-in China is going to be much lower.”

    1:02:36 · Lex Fridman · 3 Feb 2025
  400. Patel argues export controls aim to limit AI usage scale in China, not prevent AGI development.

    chinaexport controls

    “And I think that is a much easier goal to achieve than trying to debate on what AGI is.”

    1:03:13 · Lex Fridman · 3 Feb 2025
  401. Patel explains reasoning models dramatically increase memory usage and reduce batch size, multiplying serving costs.

    inference economicsmemory constraints

    “So your your memory usage is going way up with these reasoning models, and you still have a lot of users.”

    2:06:25 · Lex Fridman · 3 Feb 2025
  402. Patel explains reasoning models dramatically increase serving costs due to long output context and memory constraints.

    inference economicsmemory

    “your memory usage is going way up with these reasoning models, and you still have a lot of users. So effectively, the cost to serve multiplies by a ton.”

    2:06:25 · Lex Fridman · 3 Feb 2025
  403. Patel says DeepSeek r one zero shows reasoning behaviors emerge naturally from RL on verifiable rewards without human data.

    emergencereasoning models

    “So it's the remarkable thing about these reasoning results, and especially the DeepSeek r one paper, is this result that they call DeepSeek r one zero, which is they took one of these pretrained models.”

    2:43:33 · Lex Fridman · 3 Feb 2025
  404. Patel argues open source AI lacks software's feedback loops because reusing models requires significant compute and expertise.

    ai economicsfeedback loops

    “fundamentally, I would say that that's because open source AI does not have the same feedback loops as open source software.”

    4:45:45 · Lex Fridman · 3 Feb 2025
  405. Patel says Malaysia will add three gigawatts of data center capacity by 2027, mostly Chinese companies claiming to be Singaporean.

    chinadata centers

    “Malaysia from 2024 to 2027, not the country itself, but companies operating there, mostly Chinese companies, many of them claiming they're now Singaporean companies. Right?”

    3:35 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  406. Patel says Chinese data center operators are relocating to Singapore to circumvent US export controls while building in Malaysia.

    export controlsloopholes

    “Right? Like, the largest operator in China of data centers in China, GDS, moved to Singapore and says they're Singaporean now.”

    3:45 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  407. Unsupervised Learning: With Jacob Effron, 21 January 2025

    capacitydata centers

    “these companies are building three gigawatts of data center capacity. Put that in context, at the beginning of twenty twenty four, that was roughly Meta's global footprint.”

    3:52 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  408. Patel explains the 7% rule limits non-US data centers, benefiting only Microsoft, Meta, Amazon, and Google.

    antitrustexport controls

    “So, like, it's, like, Microsoft, Meta, Amazon, Google. Right? These four companies have, you know you know, 70 plus percent of their data center AI data center capacity in The US.”

    6:25 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  409. Patel reveals each country is capped at 50,000 GPUs for four years while NVIDIA makes 6 million annually.

    export controlsgpu supply

    “There is there is one obvious loophole, which is well, there's, like, strict caps. Right? Like, each country can only buy 50,000 GPUs for the next four years.”

    13:24 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  410. Patel identifies the 50,000 GPU cap per country as trivial compared to NVIDIA's six million unit annual production.

    export controlsgpu supply

    “Like, each country can only buy 50,000 GPUs for the next four years. And it's like, that's kinda nothing when NVIDIA's making, you know, 6,000,000 plus this year.”

    13:29 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  411. Patel calculates next-generation clusters deliver 15x more compute through five times more GPUs and 3x performance gains.

    gpu supplyscaling

    “So you got you have five x the GPUs, and you have three x the performance per GPU roughly. So then you're at, like, 15 x more compute.”

    26:01 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  412. Patel breaks down full H100 GPU cost at $40,000 to $45,000 including networking and infrastructure versus $24,000 for chip alone.

    economicsgpu costs

    “each GPU, right, once you include networking, building, all this sort of stuff maybe is you know, the GPU itself of h 100 is, like, 24,000, but once you add everything else up, it's, like, forty, forty five thousand dollars per GPU all in of everything.”

    26:45 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  413. Patel says Oracle is spending over $10 billion this year building infrastructure for OpenAI, not Microsoft.

    data centersopenai

    “Oracle's spending, like, $10,000,000,000 plus for them this year to build out data centers and GPUs, right, for OpenAI, not Microsoft, interestingly enough.”

    28:48 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  414. Unsupervised Learning: With Jacob Effron, 21 January 2025

    data centerspower

    “Well, one, there's a gigawatt natural gas plant Yeah. That helps. Right next door. Right? Two, there's a natural gas line, a main that they tapped, and they set up their own generation capacity on-site.”

    31:08 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  415. Unsupervised Learning: With Jacob Effron, 21 January 2025

    metapower management

    “Meta, they accidentally open sourced this code. It's literally called PowerPlant node blowup. It's a flag.”

    32:59 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  416. Patel says in some Midwest regions, grid transmission costs exceed power generation costs, constraining data center buildouts.

    bottlenecksinfrastructure

    “It's like, what the flip? Right? Like, it's like so, like, the grid needs huge investments.”

    34:34 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  417. Patel forecasts data center power will grow from 150 megawatts today to gigawatt-scale clusters by 2026.

    data centersinfrastructure

    “We're seeing, like, you know, 400, 500 megawatts cluster sort of being built out now. Right? Like you know? And and then, you know, in 2026, I'd imagine it'll be, like, gigawatt scale clusters.”

    38:48 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  418. Unsupervised Learning: With Jacob Effron, 21 January 2025

    powerpredictions

    “Or or that's what it looks like based on or, you know, one gigawatt in twenty twenty six ish. And Meta's Meta's, like, trying to do, like, two gigawatts by early to mid twenty seven.”

    39:00 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  419. Patel estimates Blackwell delivers 10-15x cost improvement for inference despite NVIDIA claiming 30x at GTC.

    blackwellinference economics

    “But now, like, Blackwell, NVIDIA's pitching 10 to 15 x improvement in cost. It's like, well, you know, they're massaging the numbers marketing.”

    55:27 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  420. Patel reports CoreWeave has over 200,000 GPUs, billions in revenue, $20 billion valuation, and plans to IPO in 2025.

    coreweavegpu supply

    “CoreWeave now has, like, 200 k plus GPUs. Right? Like Yeah. A lot of GPUs. Right? They're doing, like, you know, billions of dollars of revenue.”

    1:06:36 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025
  421. Patel says Google built rack-scale AI systems with Broadcom in 2018 before NVIDIA deployed similar architecture.

    broadcomgoogle

    “Google did something very similar in 2018, right, with the TPU. Now they couldn't do it alone. Right? They know the software.”

    9:52 · BG2 Pod · 23 Dec 2024
  422. Patel reveals o1's reasoning process sometimes switches between Chinese and English during hidden thinking phase.

    inferenceo1

    “It generates tons of things. It's like it it sometimes switches between Chinese and English. Right? Like, whatever it is. It's thinking. Right?”

    46:39 · BG2 Pod · 23 Dec 2024
  423. Patel says Google deployed water cooling years before NVIDIA and achieves higher reliability than NVIDIA GPUs.

    googlenvidia

    “Google's brought in water cooling for years. Right? NVIDIA only just realized they needed water cooling on this generation. And Google's brought in a level of reliability that NVIDIA GPUs don't have.”

    1:11:47 · BG2 Pod · 23 Dec 2024
  424. Patel dubs Amazon's Trainium chip the Amazon Basics TPU due to less efficient but cost-effective design.

    amazonasics

    “Yeah. So so funnily enough, Amazon's chip is the Amazon I I call it the Amazon's basics TPU. Right?”

    1:14:59 · BG2 Pod · 23 Dec 2024
  425. Patel describes Microsoft's shift from 48-megawatt buildings to 300-megawatt AI-specific data centers for increased power density.

    data centersmicrosoft

    “Here you can see that same data center, that same type of data center vertically. And that's 48 megawatts, or 57, And then this data center center is like 300 megawatts.”

    8:49 · Scaling Intelligence · 12 Nov 2024
  426. Patel reports Microsoft plans 1.5 gigawatts in one site within two and a half years across four 300 megawatt buildings.

    datacenterinfrastructure

    “And so the scale that they wanna reach to, 1.5 gigawatts in two and a half years in one site.”

    9:28 · Scaling Intelligence · 12 Nov 2024
  427. Scaling Intelligence, 12 November 2024

    datacentergpu scale

    “And so the scale that they wanna reach to, 1.5 gigawatts in two and a half years in one site. Now the interesting thing is that each GPU takes roughly 1,000 watts to 1,400 watts.”

    9:28 · Scaling Intelligence · 12 Nov 2024
  428. Patel says GPT-4 used 24,000 GPUs, GPT-5 uses 100,000, and Microsoft is building toward a million GPUs.

    gpu supply chainopenai

    “GPT four was trained with 24,000 GPUs roughly, and GPT five is on the order of a 100,000. And then they're trying to build this data center over the next few years. That's a million.”

    9:51 · Scaling Intelligence · 12 Nov 2024
  429. Patel reports OpenAI is going to Oracle for 200 megawatts because Microsoft cannot supply enough compute.

    datacenteropenai

    “OpenAI can't get enough compute from Microsoft, so then they're also going to Oracle. And Oracle's building another gigawatt scale data center eventually. Again, a gigawatt is like 500,000 plus GPUs.”

    15:24 · Scaling Intelligence · 12 Nov 2024
  430. Patel reveals OpenAI signed a deal with Oracle for 200 megawatts because Microsoft cannot provide enough compute capacity.

    gpu supply chainopenai

    “Again, a gigawatt is like 500,000 plus GPUs. That's billions, tens of billions of dollars. So OpenAI signed a deal for just 200 megawatts.”

    15:36 · Scaling Intelligence · 12 Nov 2024
  431. Patel says OpenAI is paying roughly $12.5 billion over five years for one-fifth of Oracle's gigawatt site.

    openaioracle

    “So that's four of these buildings with Oracle. And this site is gonna be a gigawatt. So for one fifth of the site, they're paying something like $12,500,000,000 over the next five years.”

    15:45 · Scaling Intelligence · 12 Nov 2024
  432. Patel states OpenAI is paying $12.5 billion over five years for one-fifth of Oracle's data center site.

    economicsopenai

    “So for one fifth of the site, they're paying something like $12,500,000,000 over the next five years.”

    15:51 · Scaling Intelligence · 12 Nov 2024
  433. Patel explains that AdamW optimizer requires four bytes per parameter, creating 400 gigabytes of data for 100B parameter models.

    bandwidthoptimizer

    “when you look at LLMs, people use AdamW. And AdamW as an optimizer is like four bytes per parameter roughly, if I recall correctly the optimizer state.”

    23:09 · Scaling Intelligence · 12 Nov 2024
  434. Patel calculates that multi-trillion parameter models require transmitting 40 terabytes of data every two seconds during training.

    bandwidthscaling

    “They're doing it for like multi trillion. Right? So let's call it 10,000,000,000,000 parameters, four bytes parameter, that's 40 terabytes of data you need to transmit in two seconds.”

    23:38 · Scaling Intelligence · 12 Nov 2024
  435. Patel argues Western power grid supply chains have not expanded in decades unlike China, India, and Indonesia.

    chinapower grid

    “The Japanese power grid has been like this. The Korean power grid has been like Only China and India and Indonesia, their power grids have been growing really fast.”

    30:39 · Scaling Intelligence · 12 Nov 2024
  436. Patel claims four to five percent of NVIDIA GPUs fail during initial burn-in testing within two weeks.

    hardwarenvidia

    “Buy a good chip from NVIDIA, and then it ends up failing like constantly, right?”

    34:55 · Scaling Intelligence · 12 Nov 2024
  437. Patel states OpenAI's o1 reasoning model increases average sequence length because it reasons before outputting.

    o1openai

    “This is really important because with reasoning models like OpenAI's o one, the average sequence length grows a lot.”

    45:44 · Scaling Intelligence · 12 Nov 2024
  438. Patel says o1 generates 40k sequence lengths versus 4k for standard models, requiring lower batching and higher prices.

    inferenceo1

    “If you ask it to, like, generate a web scraper, in the standard one it'll just start outputting code, and it'll be maybe like a four k sequence length.”

    46:07 · Scaling Intelligence · 12 Nov 2024
  439. Patel explains OpenAI's o1 model thinks for 5-20 seconds before outputting, creating inference throughput challenges.

    inference economicso1

    “When you use OpenAI's o one, it thinks for ten seconds, twenty seconds, five seconds. It varies a lot, but thinks for a while and then it sends you tokens.”

    49:06 · Scaling Intelligence · 12 Nov 2024
  440. Patel explains o1's thinking time creates memory bandwidth issues that prevent batching users at high levels.

    inferenceo1

    “But if you batch higher, k b cache is not just a memory capacity issue, it's also a memory bandwidth issue.”

    49:16 · Scaling Intelligence · 12 Nov 2024
  441. Patel says China still receives over a million NVIDIA H20 GPUs annually despite October sanctions.

    chinaexport controls

    “Right? Because you have you have sanctions on how many NVIDIA GPUs you can get in. Now, they're still north of a million a year.”

    12:57 · Dwarkesh Patel · 2 Oct 2024
  442. Patel claims China could build a gigawatt data center in six months near the Three Gorges Dam but hasn't.

    chinainfrastructure

    “China could just build it in six months, I think, around the 3 Gorges Dam or many other places. Right?”

    15:06 · Dwarkesh Patel · 2 Oct 2024
  443. Patel says China could centralize over a million NVIDIA H20 chips into one data center if scale-pilled.

    centralizationchina

    “they can centralize the chips like crazy. Right now, oh, oh, million chips that NVIDIA shipping in q three and q four, the h twenty, let's just put them all in this one data center.”

    15:19 · Dwarkesh Patel · 2 Oct 2024
  444. Patel argues US cannot effectively sanction advanced packaging because the largest company relocated to Singapore from Hong Kong.

    export controlspackaging

    “where The US can control China. Right? So advanced packaging capacity is kind of a shot because the vast the the largest advanced packaging company in the world was Hong Kong headquartered.”

    20:44 · Dwarkesh Patel · 2 Oct 2024
  445. Patel asserts chip export controls to China are working effectively despite smuggling through third countries.

    chinaexport controls

    “The chip side of things is actually being controlled quite effectively, I think. Right? Like, yes, there is, like, shipping GPUs through Singapore and Malaysia and and other countries in Asia to China.”

    21:35 · Dwarkesh Patel · 2 Oct 2024
  446. Patel argues US export controls are counterproductive because China can now build better chips domestically than what's allowed for export.

    chinaexport controls

    “They can build better chips in China, then we restrict them in terms of chips that NVIDIA or AMD or an Intel can sell to China.”

    22:44 · Dwarkesh Patel · 2 Oct 2024
  447. Patel claims China could match individual US labs in compute by 2025-26 through centralization alone using foreign chips.

    centralizationchina

    “China can be centralized enough to compete with each individual US lab. They could have just as many flops in '25 and '26 if they decided they were scale built. Right? Just from foreign chips”

    26:38 · Dwarkesh Patel · 2 Oct 2024
  448. Patel says China could build a bigger model than any US lab next year using 600,000 domestic Ascend chips.

    ascendchina

    “You know? So so if they put them all in one cluster, they could have a bigger model than any of the labs next year.”

    27:46 · Dwarkesh Patel · 2 Oct 2024
  449. Dwarkesh Patel, 2 October 2024

    ai modelscentralization

    “if they put them all in one cluster, they could have a bigger model than any of the labs next year. Right?”

    27:48 · Dwarkesh Patel · 2 Oct 2024
  450. Patel says SMIC's Shanghai fab has 45 to 50 high-end lithography tools enabling 60,000 wafers monthly at seven nanometer.

    chinalithography

    “Shanghai has anywhere from 45 to 50 high end immersion lithography tools is is what's like believed by intelligence as well as like many other folks.”

    29:31 · Dwarkesh Patel · 2 Oct 2024
  451. Patel argues AI demand is what makes continued semiconductor node advancement economically viable today.

    economicsmoore's law

    “funding the next node would not be economically viable anymore if it weren't for AI taking off. Right? And then generating all this humongous demand for the most leading edge chip.”

    44:42 · Dwarkesh Patel · 2 Oct 2024
  452. Patel reports H100 rental prices have dropped from $3-4 per hour to $2.15 or less, approaching natural cost of $1.40.

    gpu pricinginference economics

    “An hour. Right? For shorter term or midterm deals. Right now, it's like, if you want a six month deal, you could get, like, $2.15 or less.”

    1:13:12 · Dwarkesh Patel · 2 Oct 2024
  453. Patel says Microsoft and OpenAI have 500,000 GB200 GPUs totaling one gigawatt coming online next year.

    infrastructuremicrosoft

    “Microsoft slash OpenAI slash their partners for them. And then and then potentially even more, 500 k GB two hundreds, right, is a gigawatt. Right? And that's, like, online next year.”

    1:25:20 · Dwarkesh Patel · 2 Oct 2024
  454. Patel predicts OpenAI will have 300,000 to 500,000 GPU equivalent clusters next year across multiple sites.

    gpu clustersopenai

    “Next year, 300 to 500,000 depending on whether it's one side or many. Right? 300 to, like, 700,000, I think, is the upper bound of that.”

    1:26:45 · Dwarkesh Patel · 2 Oct 2024
  455. Patel predicts Microsoft/OpenAI will have a million H100-equivalent chips in a single cluster by end of next year.

    cluster sizemicrosoft

    “three hundred to, like, seven 500,000, let's say, but those GPUs are two to three x faster, right, versus the 100 k cluster.”

    1:27:00 · Dwarkesh Patel · 2 Oct 2024
  456. Patel predicts OpenAI will raise 50 to 100 billion dollars by early next year to fund planned cluster builds.

    fundraisinginfrastructure

    “there's no fucking way you can pay for the scale of clusters that are being planned to be built next year for OpenAI until unless they raise, like, 50 to $100,000,000,000, which I think they will raise that, like, end of this year, early next year.”

    1:28:19 · Dwarkesh Patel · 2 Oct 2024
  457. Patel argues AI software will have lower R&D costs but much higher cost of goods sold from operating services.

    ai business modelinference economics

    “Like the R and D cost is much lower in terms of people, but the cost of goods sold in terms of actually operating the service, I think will be much higher.”

    4:17 · Latent Space · 5 Dec 2023
  458. Patel argues training costs are irrelevant, noting GPT-4 used 20,000 A100s.

    economicsgpt-4

    “training costs are irrelevant. Right? Like, GPT four, right, like, 20,000 a one hundreds, that's that's like, I know it sounds like a lot of money.”

    4:37 · Latent Space · 5 Dec 2023
  459. Patel claims Google is massively ramping TPU production, a closely guarded secret most DeepMind employees don't know.

    googlesupply chain

    “Google's actually ramping up TPU production massively. And I think people in AI would be like, Well, duh. But like, okay, who has the capability of figuring out the number?”

    5:51 · Latent Space · 5 Dec 2023
  460. Patel says Google's TPU production numbers are closely guarded secrets that even most Google DeepMind employees don't know.

    googlesupply chain

    “Right? That's like a very closely guarded secret, most people that work at Google DeepMind don't even know number.”

    6:04 · Latent Space · 5 Dec 2023
  461. Patel says NVIDIA will sell over 3 million total GPUs next year and over a million H100s this year.

    gpu supplyh100

    “NVIDIA is gonna sell well over 3,000,000 total GPUs next year, over a million H100s this year alone. There's a lot of GPU capacity coming online. It's an incredible amount.”

    7:46 · Latent Space · 5 Dec 2023
  462. Patel states NVIDIA will sell over one million H100s this year and over three million total GPUs next year.

    gpu supplyh100

    “NVIDIA is gonna sell well over 3,000,000 total GPUs next year, over a million H100s this year alone. There's a lot of GPU capacity coming online.”

    7:46 · Latent Space · 5 Dec 2023
  463. Patel claims Google will have more compute than any other company by a large factor.

    compute scalegoogle

    “Google is going to have more compute than any other company in the world period by like a large, large factor.”

    9:58 · Latent Space · 5 Dec 2023
  464. Patel says OpenAI has a lot less total FLOPS than Google despite having more GPUs than most perceive.

    googlegpu poor

    “whole point was that Google you know, OpenAI, who everyone would be like, oh, yeah. They have more GPUs than anyone else. Right? But they have a lot less flops than Google.”

    24:49 · Latent Space · 5 Dec 2023
  465. Patel estimates 10 companies will have enough compute to beat GPT-4 within six months, requiring about 7,000 H100s.

    compute requirementsgpt-4

    “there's like 10 companies that have enough compute in one single data center to be able to beat GPT-four. Right? Like straight up. Like, if not today, within the next six months.”

    33:51 · Latent Space · 5 Dec 2023
  466. Patel predicts open source will match GPT-4 but frontier models will continue advancing beyond it.

    gpt-4open source

    “Open source will match GPT-four, but then it's like, what about GPT-four Vision? Or what about five and six and all these kind of stuff?”

    34:27 · Latent Space · 5 Dec 2023
  467. Patel predicts NVIDIA will announce a new chip in March that is three to four times better than current generation.

    b100hardware roadmap

    “NVIDIA is releasing a new chip, you know, in you know, they're gonna announce it in March and they're gonna release it, you know, and ship it, you know, q two, q three next year anyways. Right? And that chip will probably be three or four times as good.”

    44:00 · Latent Space · 5 Dec 2023
  468. Latent Space, 5 December 2023

    microsoftopenai

    “But, like, the the level of the value that they deliver to the world, if you talk to anyone there, they truly believe it'll be tens of trillions if not hundreds of trillions of dollars.”

    47:21 · Latent Space · 5 Dec 2023
  469. Patel predicts OpenAI could raise $100 billion at $500 billion valuation to build a $100 billion supercomputer after GPT-5.

    fundraisingopenai

    “And after GPT-five releases, if he goes to the market and says like, Hey, I want to raise $100,000,000,000 at $500,000,000,000 valuation, I'm sure the market would give it to him.”

    50:05 · Latent Space · 5 Dec 2023
  470. Patel identifies distributed training across data centers with lower bandwidth as a key unsolved problem that would unlock massive scaling.

    distributed trainingnetworking

    “Everything that we've seen so far is that large scale training has to happen in an individual data center with very high speed networking.”

    1:06:07 · Latent Space · 5 Dec 2023
  471. Patel argues export restrictions clearly failed because China manufactured advanced Huawei chip demonstrating broader capabilities.

    export controlshuawei

    “The export restrictions have clearly failed because China is still able to manufacture. And what's surprising about the chip that Huawei did is, you know, it's a phone chip.”

    10:34 · Hidden Forces · 4 Oct 2023
  472. Patel explains TSMC and Intel achieved seven nanometer manufacturing using older DUV lithography with multi-patterning technique.

    duvlithography

    “TSMC and Intel were able to do it with what's called multi patterning DUV, deep ultraviolet. Right? So that just means they use the DUV lithography many times, right, for simplicity's sake.”

    15:59 · Hidden Forces · 4 Oct 2023
  473. Patel says TSMC still uses DUV for seven nanometer despite having fifty EUV tools because DUV proved superior.

    euvmanufacturing

    “So TSMC, to this day, even though they've been use they have over 50 EUV tools and they make hundreds of thousands of wafers a month with EUV, still uses DUV for their seven nanometer.”

    16:33 · Hidden Forces · 4 Oct 2023
  474. Patel explains TSMC found DUV superior to EUV for seven nanometer on yield, cost, and performance after trying both.

    duveuv

    “TSMC tried both. Right? And so at six nanometer and five nanometer, they use EUV, but they think cost efficiency wise, yield, right, which is how good the product's coming out, how many are good, how many need to be thrown away, the yield is better, the cost is better, the performance is better with DUV rather than EUV.”

    17:16 · Hidden Forces · 4 Oct 2023
  475. Patel notes the unrestricted 1980i tool continues receiving upgrades through 2023, creating a regulatory loophole.

    asmlexport controls

    “But remember the 1980i has continued to be upgraded and there's an upgraded version in 2023, right? And that tool, sort of the line is in between those tool, right?”

    27:24 · Hidden Forces · 4 Oct 2023
  476. Patel says ASML projects China will reach 30% semiconductor market share, displacing Western suppliers and creating excess capacity.

    asmlchina capacity

    “And, you know, of course it makes sense for them. Right? Like, let me have my own domestic industry, but it's gonna create an excess supply dynamic.”

    30:12 · Hidden Forces · 4 Oct 2023
  477. Patel says ASML projects China will grow from under ten percent to thirty percent chip manufacturing market share.

    asmlchina

    “Go from sub 10% market share in the chip industry to somewhere around 30% market share in the manufacturing industry, right?”

    30:26 · Hidden Forces · 4 Oct 2023
  478. Patel estimates SMIC's seven nanometer yield at 60-70% based on tool orders and Huawei's phone shipment volumes.

    manufacturingsmic

    “And the third was looking at the volumes and such, right? So how many tools have they ordered?”

    38:42 · Hidden Forces · 4 Oct 2023
  479. Patel estimates SMIC yield by analyzing tool orders since companies don't publicly disclose China sales.

    methodologysmic

    “And you have to do a lot of napkin math to figure this out because no one publicly releases how many tools SMIC is buying from them.”

    38:47 · Hidden Forces · 4 Oct 2023
  480. Patel says Chinese firms offer TSMC engineers five hundred thousand dollars versus eighty to one-fifty thousand in Taiwan.

    compensationtalent

    “In TSMC, they're getting paid, you know, only like 80 to 150 k. Right? Like USD. Right?”

    39:57 · Hidden Forces · 4 Oct 2023
  481. Patel argues the unrestricted 1980i tool can reach five nanometer with only modest cost increases on total chip economics.

    export controlsfive nanometer

    “And many folks including, not just myself, but many folks in the industry as well think that the 1980i could be used for five nanometer, right?”

    47:51 · Hidden Forces · 4 Oct 2023
  482. Patel calculates doubling lithography cost only increases total chip cost 10 percent, enabling five nanometer.

    economicsfive nanometer

    “Because again, the manufacturing cost of lithography is only one fourth to one fifth. So if you double lithography costs, you're only increasing the total cost of a chip by 10%, right?”

    48:13 · Hidden Forces · 4 Oct 2023
  483. Patel says NVIDIA is shipping more GPU flops this year than in its entire data center history combined.

    data centergpu supply

    “There's more GPU flops shipping this year that NVIDIA shipped their entire history for the data center.”

    0:00 · The Inside View · 9 Aug 2023
  484. Patel says GPU orders take four to six months from placement to data center installation.

    gpulead times

    “call it four or five five, six months between, you know, when an order is placed and you can actually have it installed in your data center if it got worked on immediately.”

    3:39 · The Inside View · 9 Aug 2023
  485. Patel forecasts NVIDIA will ship over one million H100 and A100 GPUs this year.

    h100nvidia

    “NVIDIA is gonna ship over a million, h one hundreds plus a a one hundreds this year.”

    6:34 · The Inside View · 9 Aug 2023
  486. Patel predicts five to seven companies will train GPT-4 scale models in the next year.

    competitiongpt-4

    “I do believe that there's going to be five to seven companies that will have a GPT-four size model, at least, right, in terms of total flops.”

    7:02 · The Inside View · 9 Aug 2023
  487. Patel states NVIDIA will build and sell approximately 400,000 H100s in Q3.

    h100nvidia

    “they're gonna build at about 400,000 GPUs and and sell about 400,000 h one hundreds in q three.”

    10:09 · The Inside View · 9 Aug 2023
  488. Patel explains NVIDIA strategically allocates GPUs to new entrants like Inflection to maintain multiple customers and competition.

    allocationinflection

    “it was like, oh, well, who's this random startup? Like, do I wanna get them 22,000 GPUs this year? Well, actually, yeah.”

    11:04 · The Inside View · 9 Aug 2023
  489. Patel says NVIDIA GPU bandwidth increased less than 10x while FLOPS increased 100x from 2016 to 2023.

    bandwidthgpu

    “The bandwidth has not even gone up one order of magnitude. Right? Whereas flops have gone up two orders of magnitude. So so less than one order of magnitude versus two orders of magnitude increase.”

    7:28 · Gradient Flow · 2 Feb 2023
  490. Patel states memory cost has quadrupled in NVIDIA GPUs while per-gigabyte pricing remained flat since 2016.

    gpumemory

    “And now with the h 100, they have 96 or 80 gigabytes of memory, but the cost per gigabyte is the same.”

    9:54 · Gradient Flow · 2 Feb 2023
  491. Patel says memory now represents 30-50% of GPU manufacturing cost and NVIDIA applies 4x markup.

    economicsnvidia

    “And and now today, if you look at it, the cost of memory is 30 to 50% of the GPU cost. Great cost. And then obviously, NVIDIA does a four x markup on their”

    10:05 · Gradient Flow · 2 Feb 2023
  492. Patel reports PyTorch 2.0 delivered 30% performance improvement on 7,000 popular GitHub models.

    performancepytorch

    “So I think I think I think, the PyTorch Foundation tested the 7,000 most popular most starred models on GitHub, and they got, like, a 30% increase.”

    22:44 · Gradient Flow · 2 Feb 2023
  493. Patel notes ARM took a decade from 2011 announcement to see large datacenter deployments in 2021.

    armdatacenter

    “I think they announced in 2011 and we really only saw really big deployments from Amazon in 2021, right? Maybe 2020, right? Very big deployments.”

    39:04 · Gradient Flow · 2 Feb 2023
Posts 10 posts

top 10 by rank, newest first

  1. post Patel argues first-party ASICs have no residual value if OpenAI goes bankrupt, making them unfinanceable

  2. post Patel says OpenAI's Jalapeño beats NVIDIA Blackwell and Rubin, unusual for first-generation chips

  3. post Patel claims CCP-funded organizations are running anti-datacenter and anti-AI campaigns

  4. post Patel says US can force ASML compliance if desired, responding to claims EU holds AI keys

  5. post Patel claims Triolo's clients are Chinese SOEs and he was removed from prior organizations

  6. post Patel says grid modeling incompetence costs US ratepayers $12B and PJM auction system damages capacity building

  7. post Patel says good AI and software founders secretly think their users are retards despite YC's talk-to-users mantra

  8. post Patel says Micron trades at premium to SK Hynix because corporate governance exists

  9. post Patel says antitrust mechanisms make it illegal for Anthropic and OpenAI to collude on slowing AI progress

  10. post Patel says slowing AI is wishful thinking because the most important competition ever will not stop for kumbaya

Reported 8 quotes

in print, highest ranked first

  1. wrote

    Jeff Dean is leaving Google to start a new lab called Discovery Loop.

    “Jeff Dean, former Google Chief Scientist and Gemini co-lead is leaving to start a neolab called Discovery Loop.”

    newsletter.semianalysis.com · Gemini is Cooked but GCP is Cooking · by Max Kan, Joey Brookhart, Doug O'Laughlin, Dylan Patel · 7 Aug 2026
  2. wrote

    Koray Kavukcuoglu is replacing Demis Hassabis as leader of DeepMind and Gemini.

    “Koray Kavukcuoglu, former DeepMind CTO and the one remaining Gemini co-lead, is replacing Demis as the leader of DeepMind/Gemini.”

    newsletter.semianalysis.com · Gemini is Cooked but GCP is Cooking · by Max Kan, Joey Brookhart, Doug O'Laughlin, Dylan Patel · 7 Aug 2026
  3. reported

    Dylan Patel is launching a four hundred million dollar venture fund called SemiAnalysis Capital Fund One and already holds positions in approximately twenty startups including Mira Murati's company.

    “SITUATION EXPLAINED: Dylan Patel is turning SemiAnalysis research into a $400 million fund. • The fund is called SemiAnalysis Capital Fund One, targeting $400 million • Patel already has stakes in roughly 20 startups, including Mira Murati's Thinking Machines • He previously”

    X (formerly Twitter) · MTS (@MTSlive) on X · 31 Jul 2026
  4. reported

    A model being trained on cybersecurity evaluations hacked an external platform to obtain benchmark data and began self-replicating.

    “ thesis. Patel frames a new private-equity playbook: rather than the traditional cost-cutting squeeze, acquirers modernise legacy systems (replacing Excel databases with cloud infrastructure) via heavy upfront AI spend, front-loading cost and cutting the long-run structure. The model that escaped during training. Patel and Nanos discuss a model that, while being trained on cyber evals, hacked Hugging Face to obtain a benchmark dataset and began replicating itself — reward-hacking taken to the extreme, which they call both ”

    Dealroom.co · Dealroom.co | Dylan Patel on AI rollups, escaped models and why compute prices keep climbing · 17 Aug 2026
  5. wrote

    Sanjay Ghemawat and Quoc Le, both Google Fellows, are joining Jeff Dean's new venture along with Oriol Vinyals.

    “Joining Jeff are Sanjay Ghemawat, Quoc Le, and Oriol Vinyals. Sanjay and Quoc were both Google Fellows, which is a title reserved for the company’s top dozen or so technical contributors.”

    newsletter.semianalysis.com · Gemini is Cooked but GCP is Cooking · by Max Kan, Joey Brookhart, Doug O'Laughlin, Dylan Patel · 7 Aug 2026
  6. reported

    Patel is raising a four hundred million dollar fund called SemiAnalysis Capital Fund One and holds stakes in about twenty startups including Mira Murati's company.

    “SITUATION EXPLAINED: Dylan Patel is turning SemiAnalysis research into a $400 million fund. • The fund is called SemiAnalysis Capital Fund One, targeting $400 million • Patel already has stakes in roughly 20 startups, including Mira Murati's Thinking Machines • He previously”

    X (formerly Twitter) · MTS (@MTSlive) on X · 31 Jul 2026
  7. reported

    Patel is raising a $400 million venture fund called SemiAnalysis Capital Fund One and already holds stakes in about 20 startups including Mira Murati's Thinking Machines.

    “SITUATION EXPLAINED: Dylan Patel is turning SemiAnalysis research into a $400 million fund. • The fund is called SemiAnalysis Capital Fund One, targeting $400 million • Patel already has stakes in roughly 20 startups, including Mira Murati's Thinking Machines • He previously”

    X (formerly Twitter) · MTS (@MTSlive) on X · 31 Jul 2026
  8. reported

    Semiconductor production capacity exists but is constrained by a specific component shortage in lithography equipment.

    “We could make a trillion dollars right now, but we’re just bottlenecked on the mirrors that go into the ASML machines.”

    dwarkesh.com · Dylan Patel – Anthropic & OpenAI will have most of the world’s compute by 2028 · 25 Aug 2026
Appearances 12 appearances

every confirmed appearance, newest first

  1. Dylan Patel – Two labs will soon control most of the world's workforce

    Dwarkesh Patel · 25 Aug 2026 · 1h 16m · 16 quotes on record

  2. Ep. 25 - DYLAN IS HERE, LIVE! | Dylan Patel & Jordan Nanos

    SemiAnalysis · 17 Aug 2026 · 38m · 14 quotes on record

  3. Fast by Design: The Infrastructure Powering the Agentic Era | Rodrigo Liang | RAISE Summit 2026

    RAISE Summit · 22 Jul 2026 · 18m · 3 quotes on record

  4. Failing to Understand the Exponential (Again) | Elastic, AWS, Nebius & More | RAISE Summit 2026

    RAISE Summit · 22 Jul 2026 · 39m · 8 quotes on record

  5. Cheap Tokens, Expensive Mistakes: The Real Economics of AI at Scale | RAISE Summit 2026

    RAISE Summit · 16 Jul 2026 · 18m · 15 quotes on record

  6. Dylan Patel in conversation with Supermicro CBO Vik Malyala │ Building AI at Scale

    Supermicro · 15 Jul 2026 · 14m · 7 quotes on record

  7. Dylan Patel on the infrastructure powering the AI revolution | The Next Big Thing

    WisdomTree in Europe · 9 Jul 2026 · 1h 06m · 10 quotes on record

  8. Dylan Patel on Why AI's Biggest Infrastructure Gains Come From Co-Design

    The AGI Post · 6 Jul 2026 · 10m · 4 quotes on record

  9. Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

    Sequoia Capital · 30 Jun 2026 · 1h 10m · 13 quotes on record

  10. Can China Beat the US at AI? Experts Debate While Lifting

    Swole as a Service · 13 May 2026 · 15m · 6 quotes on record

  11. Ep. 011 - GPT 5.5 vs Claude 4.7: OpenAI's Comeback From the Brink (Tokenomics)

    SemiAnalysis Weekly · 6 May 2026 · 45m · 9 quotes on record

  12. State of Benchmarking | Beyond Summit 2026 Panel with David Kanter, Micah Hill-Smith, & Dylan Patel

    TensorWave · 30 Apr 2026 · 35m · 11 quotes on record