Glossary

Local LLM

A local LLM is a large language model that runs entirely on hardware you own or control, so prompts and responses never travel to a third-party provider.

General definition

A local LLM is any large language model deployed on infrastructure you operate, whether a developer’s laptop, an on-premises server, or a private cloud instance. The defining characteristic is that inference happens inside your own network boundary: no prompt text, no user data, and no model responses are transmitted to an external API.

  • Tools such as Ollama, LM Studio and vLLM make it straightforward to download open-weight model weights and serve them with an OpenAI-compatible API
  • Performance depends on available hardware: consumer GPUs can run 7B-13B parameter models comfortably; larger models typically need multi-GPU servers
  • Quantised versions of popular models reduce memory requirements at a modest cost to output quality

Local LLMs are popular in regulated industries where sharing data with a cloud provider is restricted by policy or law, in air-gapped environments, and on teams that need to keep fine-tuned model weights proprietary. They are closely related to the broader concept of a private LLM.

In the Ethora ecosystem

Ethora’s self-hosted LLM AI agent capability is built around this pattern. You connect your Ethora deployment to a local model endpoint, and AI-powered chat features, such as in-thread summarisation, routing bots and RAG assistants, run without any message leaving your infrastructure.

For organisations building in healthcare, finance or other compliance-heavy sectors, combining Ethora’s self-hosted messaging with a local LLM gives a complete stack where patient data, financial records and conversation context remain under your control at every layer.

Get started

Chat and AI on infrastructure you control

Run Ethora self-hosted or dedicated, with full control of data, keys and residency. Talk to our team.

Start Free
Free tier available Enterprise SLA No vendor lock-in