This is prefill:

Prefill is the part of LLM inferencing where the prompt is run through the model. The products of prefill are:

  1. K and V vectors for every layer of attention and every token in the prompt
  2. The final hidden state of the model prior to output tokens being generatedok

Because the whole input prompt is passed in, it is a dense computation and is therefore compute-limited. This contrasts with the next phase, decode.

Specifically, the input of prefill is , where is the number of input tokens and is the hidden dimension. To compute the attention part of prefill,

  1. , , and projections are computed by multiplying the input matrix by three different weight matrices (, , ). The resulting and projections are then saved in KV cache.
  2. Attention itself is computed using these computed , , and according to . This is not cacheable.

The projection part turns into GEMMs (matrix-matrix multiplication). As long as is large enough to fill a tensor core tile, each GEMM takes longer than loading the inputs from HBM, so it is compute-bound.

Prefill’s compute scales as

  • when computing projections
  • when computing attention