I wanted to see what llama.cpp’s KV cache actually looks like in memory, so I dumped the cached keys and values for a short prompt in Qwen3-0.6B, a small model with 28 layers. I used an RTX 4090, an F16 cache, and a context capacity of 4,096 tokens.
The useful picture is fairly simple: every layer has its own K and V tensors, every occupied row holds one token’s numbers, and streams add another dimension. Here is that picture, followed down to an actual pair of bytes in the dump.
28 layers, 56 tensors
Attention happens in every one of this model’s 28 layers. Each layer computes its own keys and values, so the cache keeps 28 K tensors and 28 V tensors: 56 tensors in all.
With all layers offloaded to the GPU, these tensors live in one CUDA buffer. The diagram groups the keys and values by color; it is an inventory, not their address order. In this build, the live buffer contains the layer pairs in order: K0, V0, K1, V1, .... The saved file uses a different order, which we will see below.
The constructor creates the two tensors inside its layer loop. For this model, the relevant calls are:
ggml_new_tensor_3d(ctx, type_k,
n_embd_k_gqa, kv_size, n_stream);
ggml_new_tensor_3d(ctx, type_v,
n_embd_v_gqa, kv_size, n_stream);
The constructor and per-layer structure
kv_layer keeps the pair of tensors and their stream views together.One row, one token’s keys
For one token in one layer, K contains 1,024 numbers. That is eight KV heads with 128 numbers each, packed next to each other in a row. V has the same size. Each F16 number takes two bytes, so a K row occupies 2,048 bytes, or 2 KiB. A V row takes another 2 KiB.
The dimensions in the model's startup log
n_embd_k_gqa = n_head_kv × n_embd_head_k = 8 × 128 = 1024.Across all 28 layers, one token’s K and V payload is:
2 KiB × 28 layers × 2 (K and V) = 112 KiB
For c cache cells, the total is c × 112 KiB. So even eight tokens hold 896 KiB, almost a MiB, in this small model. Here KiB means 1,024 bytes and MiB means 1,024 KiB.
llama.cpp reserves space for the configured capacity, rather than growing this buffer one token at a time. My -c 4096 run therefore reserved 448 MiB, even though the prompt only occupied seven cells.
The measured allocation
Opening the .bin dump
For the experiment, I fed in “The cat sat on the mat.”, which became seven tokens, and saved sequence 0’s occupied cache across all 28 layers. I used a small helper, kvprobe, with the following command:
./kvprobe -m Qwen3-0.6B-Q8_0.gguf \
-c 4096 -ngl 99 -fa on -ctk f16 -ctv f16 \
--ops 'text:0:The cat sat on the mat.;save:0:state-f16.bin'
python3 parse_state.py state-f16.bin --cells 3
The model’s weights are Q8_0; the cache in this experiment is F16. The helper saves the state through llama.cpp’s API:
const size_t n = llama_state_seq_get_size_ext(
ctx, s, LLAMA_STATE_SEQ_FLAGS_NONE);
std::vector<uint8_t> buf(n);
const size_t got = llama_state_seq_get_data_ext(
ctx, buf.data(), buf.size(), s,
LLAMA_STATE_SEQ_FLAGS_NONE);
std::ofstream(path, std::ios::binary).write(
(const char *) buf.data(), (std::streamsize) got);
In the saved file, the K records run from layer 0 through layer 27, followed by the V records for all 28 layers. Each record has a small header and then its seven rows. This is where the “all K, then all V” layout appears.
metadata
K0, K1, ... K27 (header + 7 rows in each record)
V0, V1, ... V27 (header + 7 rows in each record)
The first K record starts at byte 108. Its 12-byte header puts the first actual K value at byte 120. Reading four F16 numbers directly from there gives:
import struct
from pathlib import Path
data = Path("state-f16.bin").read_bytes()
print(len(data))
print(data[120:128].hex(" "))
print(tuple(round(x, 4) for x in
struct.unpack_from("<4e", data, 120)))
803596
31 37 46 35 a8 b4 6f b0
(0.4495, 0.3296, -0.291, -0.1385)
The first two bytes, 31 37, are the little-endian F16 representation of approximately 0.4495. That is an actual cached key value, not an illustrative number.
The dump, full parsed output, probe source, build helper, and parser are available alongside this post. The probe uses a CUDA build of the pinned llama.cpp revision linked below; the parser needs NumPy. Set LLAMA_DIR to that checkout when using the build helper.
The tensor’s three dimensions
A layer’s K tensor has this shape, in ggml’s fastest-dimension-first notation:
[values per token, cells per stream, number of streams]
For this run, that is [1024, 4096, 1]. The row count is capacity, including empty cells. With F16 and Flash Attention enabled, V uses the same layout.
ne gives the sizes. nb gives the byte strides: two bytes per number, 2,048 per row, and 8 MiB per stream's slab.Those strides tell us where any value lives. For example, take layer 5, cell 37, KV head 3, and element 10 within that head. Counting from zero, its offset from the start of that layer’s K tensor is:
37 × 2048 + (3 × 128 + 10) × 2 = 76,564 bytes
Unified versus non-unified
I also dug into how unified and non-unified caches affect that third dimension. Here, a stream means a cache partition, not a CUDA execution stream. A sequence is one independent conversation or run of tokens.
With a unified cache, n_stream is one. Tokens from different sequences occupy rows in the same pool, with bookkeeping and attention masks keeping track of which sequence may read which cells.
With a non-unified cache, n_stream is the maximum number of sequences. Each stream gets a two-dimensional slab of every layer’s K tensor and V tensor. These are views into the larger tensors, not extra copies.
s starts at s × nb[2].The source expresses the choice directly:
n_stream(unified ? 1 : n_seq_max)
And it creates each K slab’s view with:
ggml_view_2d(ctx, k, n_embd_k_gqa, kv_size,
k->nb[1], s * k->nb[2]);
The stream-selection code
For a total capacity of 4,096 cells and four sequences, the shapes are:
| Mode | K tensor shape, per layer |
|---|---|
| Unified | [1024, 4096, 1] |
| Non-unified | [1024, 1024, 4] |
That partitioning explains the speed-versus-memory-efficiency tradeoff. Non-unified attention can read a sequence’s own slice without scanning cells belonging to other sequences, which can make independent concurrent sequences faster. But spare capacity in one stream cannot simply be used by another, and common prefixes need separate copies. Unified keeps a shared pool, making it easier to use the space when conversation lengths differ.
For me, the cache became much easier to picture once I could connect the pieces: 28 pairs of tensors, rows of 1,024 numbers, two bytes per number, and one more dimension deciding how the rows are shared.
Source
The experiment and diagrams use llama.cpp 99b95488c and Qwen3-0.6B. The main code is in llama-kv-cache.cpp, including tensor construction and state serialization.