GGUF files explained: which build of a model to download

One model on Hugging Face, twenty files, all ending in .gguf, all several gigabytes, with names like Q4_K_M, IQ3_XXS and Q6_K. Nothing on the page explains which one belongs on a 16GB Mac, and downloading the wrong one costs an hour and a model that either crawls or refuses to load. The filenames are not arbitrary. There is a published naming convention and a published table of bits per weight, and between them they turn the choice into arithmetic.

What the format guarantees

GGUF is a binary format for storing models for inference with GGML and executors built on GGML. Models are usually developed in PyTorch or a similar framework and converted to GGUF afterwards. The format was developed by the author of llama.cpp, which is the engine most local model tooling on a Mac ends up using, directly or underneath something else.

It is the successor to three earlier formats, GGML, GGMF and GGJT, and the specification lists the properties it was designed to have:

  • Single file deployment. Everything needed to load the model is in one file, with no external files for additional information.
  • Full information. The file contains everything required to load the model, so nothing has to be supplied by the user.
  • Extensible. New information can be added without breaking compatibility with existing models.
  • mmap compatibility. Models can be loaded with mmap, which is what makes loading fast, and which requires the tensors to be aligned.

The key difference from GGJT is that hyperparameters are stored as key and value pairs, now called metadata, rather than a list of untyped values. That is why a GGUF file can describe its own architecture, context length, tokenizer and chat template, and why a runner can load a model it has never seen without being told anything about it.

The contrast worth holding on to is with safetensors, which is also a recommended format on the Hub. Safetensors stores tensors only. GGUF stores the tensors and a standardized set of metadata alongside them. That is the reason the two formats are used at different stages: one for training and fine tuning, one for handing a finished model to an inference engine.

Hugging Face supports the format directly. Models with GGUF files can be filtered at hf.co/models?library=gguf, there is a viewer on the model and files pages that shows metadata and tensor information including name, shape and precision, and the ggml-org/gguf-my-repo tool converts and quantizes weights into GGUF.

The filename is a specification

The convention is published, and once it is legible, most of the guesswork disappears. The shape is:

[<Sidecar>-]<BaseName>-<SizeLabel>-<FineTune>-<Version>-<Encoding>-<Type>-<Shard>.gguf

Each present component is separated by a hyphen. What the parts mean:

Component What it tells you
Sidecar Optional. mmproj for a multimodal projector, mtp for multi-token prediction heads. Not a standalone model
BaseName The model base type or architecture
SizeLabel Parameter class, as <expertCount>x<count><prefix>, where the prefix is K, M, B, T or Q for thousand, million, billion, trillion or quadrillion
FineTune What the fine tuning aimed at, such as Chat or Instruct
Version v<Major>.<Minor>. If absent, assume v1.0
Encoding The weights encoding scheme applied, which is where Q4_K_M and friends appear
Type LoRA for an adapter, vocab for vocabulary and metadata only. Absent means an ordinary tensor model
Shard Optional, as <ShardNum>-of-<ShardTotal>, both five digits, starting at 00001

The specification's own examples make it concrete. Mixtral-8x7B-v0.1-KQ2.gguf is 8 experts of 7 billion parameters at version 0.1 with encoding KQ2. Hermes-2-Pro-Llama-3-8B-F16.gguf is 8 billion parameters, unquantized at F16, with no version number so version 1.0 is assumed. Grok-100B-v1.0-Q4_0-00003-of-00009.gguf is the third of nine shards. mtp-Qwen3-27B-v1.0-Q4_K_M.gguf is a draft module for speculative decoding, not a model to load on its own.

The specification also flags the failure mode: with the version omitted, an encoding can be mistaken for a fine tune name. At minimum a file should carry BaseName, SizeLabel and Version to be recognisable, and the specification notes that real filenames in the wild are diverse enough that the convention is a guide to reading them rather than something to parse strictly.

Quantization, in bits per weight

Quantization compresses the weights, giving up some fidelity for size. The reason a 4 bit file is not a quarter of a 16 bit file is that these schemes store block scales and minimums alongside the weights, and those cost bits too. Hugging Face publishes the resulting figure for each type, and dividing by eight converts it into bytes per parameter.

Type Bits per weight GB per billion parameters 8B model 32B model
IQ1_S 1.56 0.20 about 1.6GB about 6.2GB
IQ2_XXS 2.06 0.26 about 2.1GB about 8.2GB
IQ2_XS 2.31 0.29 about 2.3GB about 9.2GB
Q2_K 2.625 0.33 about 2.6GB about 10.5GB
IQ3_XXS 3.06 0.38 about 3.1GB about 12.2GB
Q3_K 3.4375 0.43 about 3.4GB about 13.8GB
IQ4_XS 4.25 0.53 about 4.3GB about 17GB
Q4_K 4.5 0.56 about 4.5GB about 18GB
Q5_K 5.5 0.69 about 5.5GB about 22GB
Q6_K 6.5625 0.82 about 6.6GB about 26GB
F16 16 2.0 about 16GB about 64GB

Two things to keep in mind when using this table. It gives the weights only; the context window is allocated on top, and it is not small at the context lengths that agents and coding tools need. And on Apple Silicon the memory is unified, so whatever the model takes is taken from the same pool as the browser and the editor.

LM Studio's documentation gives the guidance in one line: choose a 4 bit option or higher if the machine is capable of running it. That is the consensus position, and the table shows why. Going from Q4_K to Q3_K saves about 1.1GB on an 8B model, which is rarely the difference between fitting and not fitting, and costs quality that is noticeable. The place where the low bit types earn their keep is a model that is otherwise entirely out of reach: a 32B model at IQ3_XXS is about 12GB of weights, where at Q4_K it is about 18GB.

What the suffixes and the IQ prefix mean

The _K in Q4_K refers to the block structure, and the published descriptions are specific about it. Q4_K and Q5_K use super-blocks of 8 blocks with 32 weights each, with a 6 bit block scale and a 6 bit block minimum. Q3_K, Q2_K and Q6_K use super-blocks of 16 blocks with 16 weights each. Q8_K uses blocks of 256 weights and is described as only being used for quantizing intermediate results, not for distributing models.

The _S, _M and _L that appear on real filenames are not in the type table, because they are not types. A real file does not quantize every tensor at the same level; whoever produced it chose a mix, keeping the tensors that matter most at a higher bit rate. Small, medium and large describe that mix. Q4_K_M is the most commonly recommended build for a reason: it is the medium mix at four bits, which is where the size and quality curves cross for most people.

The IQ prefix marks a different approach. These types use super-blocks of 256 weights and derive the weight from a super-block scale and an importance matrix, which is a calibration pass over sample data that decides which weights to protect. That extra information is why an IQ type reaches a lower bit rate at comparable quality to a plain Q type. IQ4_XS at 4.25 bits per weight against Q4_K at 4.5 is a small saving; IQ2_XS at 2.31 against Q2_K at 2.625 is the same idea where every bit counts.

Two more entries appear in the published table and are worth recognising rather than choosing by accident. TQ1_0 and TQ2_0 are ternary quantization, which only applies to models trained for it. MXFP4 is 4 bit Microscaling Block Floating Point. And the types marked legacy, Q8_0 and Q8_1, use blocks of 32 weights with a single scale, which is why the newer _K types at the same nominal bit count give better results.

Shards, sidecars, and files that are not models

A repository listing often contains files that should not be downloaded as models, and the naming convention is what distinguishes them.

A file ending -00003-of-00009.gguf is one shard of nine. All of them are needed, and the shard numbering starts at 00001 rather than 00000, which is worth knowing when a download script builds the list itself.

A file starting with mmproj- is a multimodal projector: the vision or audio encoder that gets loaded alongside a base language model. A file starting with mtp- holds multi-token prediction heads, the draft module used for speculative decoding, and the specification notes that these weights are often distributed inside the base model instead, in which case no separate file exists.

A file with LoRA in the type position is an adapter, meant to be applied to a base model. A file with vocab holds only vocabulary and metadata. Neither is a model that answers questions.

Choosing on a Mac

Put the pieces together and the procedure is short.

Start from available memory, not from the model you want. Take the memory that is genuinely free with the usual applications open, subtract a few gigabytes for the context window, and the remainder is the weights budget. Read it against the GB per billion parameters column to see which combinations of size and bit rate fit.

Given a choice between a larger model at a lower bit rate and a smaller model at four or five bits, the second is the safer default, because degradation below four bits stops being subtle. The exception is the model that only fits at all through an IQ type.

Before committing to a multi gigabyte download, use the viewer on the Hugging Face file page to check the metadata: the architecture, the context length the file declares, and the tensor precisions. It costs a moment and catches the case where two files with similar names are not the same build at all. mmap support is why loading is fast once the file is local, and it is also why the file has to be on a volume that stays connected rather than one that sleeps.

What to change first

Read the filename before reading the model card: the size label and the encoding together tell you whether it fits, and the table above turns them into gigabytes. Then pick the _K_M build at four or five bits of the largest model your memory budget allows, rather than the largest model that technically loads. If the work after that is moving listings and commands between a folder, a shell and a chat window, that is the problem Atriens is built to remove, and the way different tools handle it is set out on the Compared with other file managers page.

Frequently asked questions

What does the difference between Q4_K_S and Q4_K_M actually mean?

Both are four bit builds using the same block structure, super-blocks of 8 blocks with 32 weights. The _S and _M describe the mix of per tensor precision chosen by whoever produced the file, with the medium mix keeping more of the important tensors at a higher bit rate. _M is the usual recommendation, and the size difference between the two is typically a few hundred megabytes.

How large will the file be before downloading it?

Divide the bits per weight by eight and multiply by the parameter count. Q4_K is 4.5 bits per weight, so about 0.56GB per billion parameters, making an 8B model roughly 4.5GB and a 32B model roughly 18GB. The memory the model needs while running is that figure plus the context allocation, which is why a file that fits on disk does not always load.

Is GGUF the same as safetensors?

No. Safetensors stores tensors only. GGUF stores the tensors together with a standardized set of metadata, which is why a GGUF file can describe its own architecture, tokenizer and chat template and be loaded without additional configuration. Safetensors is the common format for training and fine tuning; GGUF is the format for handing a finished model to a GGML based inference engine.

What are the files starting with mmproj and mtp?

They are sidecar files rather than models. mmproj is a multimodal projector, the vision or audio encoder loaded alongside a base language model. mtp holds multi-token prediction heads, the draft module used for speculative decoding, which is sometimes distributed inside the base model instead. Neither will run on its own.

Do the IQ quantization types need anything special to run?

Not on the user's side. They are supported by the same runners as the plain Q types. The difference is on the production side: IQ types use an importance matrix, a calibration pass that decides which weights to protect, which is what lets them reach a lower bit rate at comparable quality. IQ4_XS is 4.25 bits per weight against 4.5 for Q4_K.

Back to all posts