Dylan Patel on

o1

5 quotes · Nov 2024 – Dec 2024

Saidverbatim, newest first

  1. Patel reveals o1's reasoning process sometimes switches between Chinese and English during hidden thinking phase.

    “It generates tons of things. It's like it it sometimes switches between Chinese and English. Right? Like, whatever it is. It's thinking. Right?”

    46:39 · BG2 Pod · 23 Dec 2024 · permalink
  2. Patel states OpenAI's o1 reasoning model increases average sequence length because it reasons before outputting.

    “This is really important because with reasoning models like OpenAI's o one, the average sequence length grows a lot.”

    45:44 · Scaling Intelligence · 12 Nov 2024 · permalink
  3. Patel says o1 generates 40k sequence lengths versus 4k for standard models, requiring lower batching and higher prices.

    “If you ask it to, like, generate a web scraper, in the standard one it'll just start outputting code, and it'll be maybe like a four k sequence length.”

    46:07 · Scaling Intelligence · 12 Nov 2024 · permalink
  4. Patel explains OpenAI's o1 model thinks for 5-20 seconds before outputting, creating inference throughput challenges.

    “When you use OpenAI's o one, it thinks for ten seconds, twenty seconds, five seconds. It varies a lot, but thinks for a while and then it sends you tokens.”

    49:06 · Scaling Intelligence · 12 Nov 2024 · permalink
  5. Patel explains o1's thinking time creates memory bandwidth issues that prevent batching users at high levels.

    “But if you batch higher, k b cache is not just a memory capacity issue, it's also a memory bandwidth issue.”

    49:16 · Scaling Intelligence · 12 Nov 2024 · permalink

Everything Dylan Patel is on record saying · RSS