Dylan Patel · Scaling Intelligence · 12 November 2024

“when you look at LLMs, people use AdamW. And AdamW as an optimizer is like four bytes per parameter roughly, if I recall correctly the optimizer state.”

Listen at the timestamp · See it in context

Verbatim excerpt with a timestamp. The full recording is at the source; we link out and do not host it. Everything Dylan Patel is on record saying.