Dylan Patel on

gpu utilization

2 entries, 5 Dec 2023 to 16 Jul 2026

On the recordsourced and dated, oldest first

    1. spoken

      Patel explains model bandwidth utilization is the critical metric for inference, unlike training where MFU matters more.

      “But on inference, it's not being talked about much, but model MBU, model bandwidth utilization is the important factor.”

      5 Dec 2023 · Latent Space · 18:19 · source · permalink
    1. spoken

      Patel argues prefill caching drastically raises GPU utilization by eliminating repeated context recalculation that provides no value.

      “Now that means your GPU utilization rises drastically, and the amount of time that the GPUs are generating new tokens is far, far higher than the amount of times that GPUs are generating or recalculating the context, which is not necessarily providing any value to your operations.”

      16 Jul 2026 · RAISE Summit · 9:52 · source · permalink

Dylan Patel ontheir other subjects

Everything Dylan Patel is on record saying · RSS