Dylan Patel · Latent Space · 5 December 2023

“Hugging Face's libraries are actually very inefficient, like incredibly inefficient for inference. You get like 15% MBU on on on on some configurations like eight a one hundreds and Lama 70 b.”

Listen at the timestamp · See it in context

Verbatim excerpt with a timestamp. The full recording is at the source; we link out and do not host it. Everything Dylan Patel is on record saying.