Baker identifies memory as the single most important factor for increasing token output per compute unit.
“It's the single most important thing you could do to increase token output per unit of compute, and then that obviously, definitionally, actually lowers costs,”
Baker explains MFU runs at 35-40%, measuring the percentage of theoretical compute actually used for training.
“MFU, model flops utilization, and that generally runs around 35 to 40%. And that's literally the percentage of compute, theoretical compute flops that you're actually applying to trading.”
Baker explains higher MFU allows 25% faster time to market with same GPU and power spending.
“You have the same amount of GPUs and the same amount of power presumably. You could choose between faster time to market.”
Baker explains that 50% MFU versus 40% MFU enables 25% faster time to market for AI models.
“You could choose between faster time to market. If you run a 50% MFU and your competitor's running 40 for an equivalent amount of trading flops, you could be in market 25% faster.”
Baker proposes a new metric decomposing MFU that he developed the morning of the interview.
“MFU is the most important metric because it gives you all of these advantages and ways to differentiate yourself amongst five people, six people who've trained these GPT core class models.”
https://x.com/GavinSBaker/status/2082166566280642676
https://x.com/GavinSBaker/status/2082817582667796992
post Baker predicts vertically focused AI native companies will accelerate due to routers, open-source models and specialized post-training
https://x.com/GavinSBaker/status/2083543298992406904
https://x.com/GavinSBaker/status/2082506941226721745
post Baker argues game theory suggests memory LTAs are durable and Nvidia may have significant cloud revenue shares
https://x.com/GavinSBaker/status/2084699467010195882
post Baker hypothesizes tomorrow's Ultrafast will be GPUs or Jalapeño with CS-4 or CS-5
https://x.com/GavinSBaker/status/2092666151520411790
https://x.com/GavinSBaker/status/2082169620497403970
post Baker speculates Ultrafast will use GPUs plus CS-4/5, notes Jalapeño not financeable at GPU rates
https://x.com/GavinSBaker/status/2092667328651800694
post Baker suggests NVIDIA accelerators have residual value that distinguishes them from first-party ASICs, responding to concerns about financing OpenAI hardware
https://x.com/GavinSBaker/status/2092747643428475048
post Baker questions whether AFD is compatible with data locality, disputes OpenAI bet against disaggregation
https://x.com/GavinSBaker/status/2092630582471823849