Bill Gurley on

inference

2 quotes · Sep 2024 – Dec 2024

Saidverbatim, newest first

  1. Patel explains reasoning models increase cost ten times by outputting 11,000 tokens versus 1,000 for same query

    “I outputted a thousand tokens to I outputted 11,000 tokens. I've 10x'd my spend to generate no. Not the same thing. Right? It's higher quality.”

    48:17 · Bg2 Pod · 23 Dec 2024 · permalink
  2. Patel calculates reasoning models cost fifty times more per query due to batch size and token generation combined

    “Cost increase for a single token to be generated is four to five x, but then I'm generating 10 x as many tokens.”

    50:06 · Bg2 Pod · 23 Dec 2024 · permalink

Everything Bill Gurley is on record saying · RSS