Dylan Patel · Latent Space · 5 December 2023
“Hugging Face's libraries are actually very inefficient, like incredibly inefficient for inference. You get like 15% MBU on on on on some configurations like eight a one hundreds and Lama 70 b.”
Verbatim excerpt with a timestamp. The full recording is at the source; we link out and do not host it. Everything Dylan Patel is on record saying.