Glossary

GGUF

The file format that makes a language model run on a laptop. GGUF packs quantised weights, tokenizer and metadata into one file that llama.cpp, Ollama and LM Studio can load in seconds.

General definition

GGUF is the binary model file format created by the llama.cpp project in 2023 as the successor to GGML. It stores a language model’s weights, usually already quantised to fewer bits, together with the tokenizer vocabulary, the chat template and architecture metadata, so a single file is enough to load and run the model. It was designed for memory-mapped loading, which is why a multi-gigabyte model starts almost instantly, and for CPU-first inference with optional offloading of layers to a GPU.

The file name usually tells you how the weights were compressed. Quantisation trades a little accuracy for a lot less memory and faster inference:

  • Q8_0: 8-bit, near-lossless, about half the size of the 16-bit original
  • Q6_K, Q5_K_M: small quality loss, good when memory allows
  • Q4_K_M: the common default; roughly a quarter of the 16-bit size with modest quality loss, and the level most public GGUF downloads ship first
  • Q3_K, Q2_K: aggressive; fits large models into small memory at a visible cost in quality
  • F16 / BF16: unquantised, for conversion or accuracy comparison

GGUF is the format behind most local LLM tools: llama.cpp itself, Ollama, LM Studio, Jan and GPT4All all load it, and Hugging Face hosts thousands of community conversions. The other format you will meet is safetensors, the Hugging Face standard for unquantised (or GPU-quantised) weights used by PyTorch, transformers and high-throughput servers such as vLLM. The practical split: safetensors for training, fine-tuning and serving many users on data-centre GPUs; GGUF for running a model on one machine, from a developer laptop to a modest on-premise server, where a small language model at Q4 or Q5 often gives acceptable quality with no cloud dependency.

Prefer to watch? GGUF explained in about two minutes.

In the Ethora ecosystem

Ethora does not read model files directly; the AI SDK talks to whatever serves the model. That is exactly where GGUF fits: run a GGUF model in Ollama, llama.cpp’s server or LM Studio, which each expose an OpenAI-compatible endpoint, and point an Ethora agent at it as its per-agent model. The agent then answers in chat rooms, channels and the website widget with retrieval over your knowledge base, and no message text leaves your infrastructure, which is the self-hosted LLM path for healthcare, finance and insurance teams.

The choice of quantisation level is a cost and quality dial the operator controls: a Q4_K_M small model on a single GPU or a strong CPU handles FAQ-style agents well, while a larger Q5 or Q8 model, or a safetensors deployment on vLLM, suits heavier reasoning and more concurrent users. Because each Ethora agent picks its own model, both can run side by side in the same app and be swapped without touching the chat integration. The private LLM entry covers the wider deployment question.

Get started

Build AI agents on your own stack

Ethora’s AI SDK brings agents, RAG and self-hosted LLMs into your product. Talk to our team.

Start Free
Free tier available Enterprise SLA No vendor lock-in