Software & AI · From strategy to production

AI integration · Open models

Self-Hosted LLM: Deploying Open-Source AI On Premises or in Your Cloud

Self-hosted AI relies on an open language model running on your own infrastructure, in your cloud or on premises, rather than on a provider’s API. Etixio helps you decide whether it makes sense, then takes it to production: model selection and evaluation, inference server, GPU sizing, security, monitoring and a cost comparison with an API.

Understanding the technology

A model that runs on your side.

Several developers publish the weights of their models: among the most widely used families are Llama (Meta), Mistral, Qwen (Alibaba) and Gemma (Google). These models can be downloaded and run on your servers, and the data sent to them never leaves your environment.

“Open source” is often shorthand: these are more accurately open-weight models, and their licenses vary. Some are very permissive (Apache 2.0, MIT), while others impose usage or attribution conditions. We check each model’s license before selecting it for commercial use.

A developer in front of two monitors

Relevant or not

When does self-hosting make sense?

Data that must not leave

Medical records, banking data, trade secrets, classified information or customer contracts that forbid third-party services. This is the most common reason.

An isolated environment

Industrial sites, networks without internet access, embedded devices. A compact model can run locally.

High, steady volumes

Under a large, stable load, well-utilized infrastructure can cost less than per-token billing. The math has to be done case by case.

A need for control or specialization

A frozen version with no provider-imposed retirement, fine-tuning a model on your business vocabulary, full control of the stack.

Conversely, self-hosting is rarely the right choice for a first project with uncertain volume, for tasks that require the most advanced models on the market, or if nobody can operate GPU infrastructure over time. A European API, such as Mistral AI’s, or a major provider’s API with suitable contractual commitments may be enough.

What we work on

The technical stack of a self-hosted LLM.

Model selection and evaluation

Comparing several open models of different sizes on your examples, against an API-based model used as a baseline. We pick the smallest model that reaches the expected quality level.

Inference server

vLLM to serve a model to many users with good throughput, Ollama or llama.cpp for workstations or lightweight use. These tools expose an API that follows common industry conventions, which simplifies integration.

Sizing and quantization

The GPU memory required depends on model size, context length and the number of concurrent requests. Quantization reduces memory needs, with an impact on quality that has to be measured.

Infrastructure

GPUs rented from a cloud provider (including European ones), dedicated servers or on-premises hardware, containerized deployment with Docker and Kubernetes, scaling and high availability as needed.

Document search and agents

An embedding model and vector database also hosted on your side, for a fully internal RAG setup, and tool calling for agents when the model supports it.

Security

Network isolation, authentication in front of the inference server, encryption, permission management in the application, controlled logs and component updates.

The choices that matter

Self-hosting or API, comparing costs.

An API is paid per use: cost follows volume, with no upfront investment. A self-hosted model is paid per capacity: GPUs rented or bought, whether used or not, plus the engineering time to operate, monitor and update them. At low volume or with irregular load, the API is almost always cheaper; at high, steady volume, the balance can tip the other way.

We build the comparison on your estimated volumes, including operations, and keep an architecture that can move from one to the other. A hybrid approach is common: an internal model for sensitive data, an API for everything else.

A team gathered around laptops in a bright office

In the field

A document assistant with no data leaving your environment.

For an assistant that queries confidential documents, the entire chain stays within the client’s environment: document extraction and chunking, embeddings, vector database, a generation model served by vLLM, the application and its logs. The evaluation set compares the selected model to a baseline so the quality gap accepted in exchange for confidentiality is known. Monitoring tracks GPU utilization, latency and queues.

From work to deliverables

What we deliver.

Evaluation and monitoring practices are detailed on our LLMOps and AI evaluation page. If a prototype already exists, we start from it with our AI POC to production offering. For a software vendor whose customers require controlled hosting, a self-hosted model can also power AI features built into the product.

Frequently asked questions

Self-hosted AI: your questions.

What is self-hosted AI?

It is an AI solution whose language model runs on infrastructure you control (private cloud, dedicated servers or on-premises hardware) instead of being called through a provider’s API. The data processed stays within your environment.

Is an open-source LLM as good as GPT, Claude or Gemini?

On many targeted tasks (extraction, classification, document search, summarization), the best open models deliver satisfactory results. For complex reasoning or long-running agents, the most advanced proprietary models often keep an edge. Only an evaluation on your examples can settle it.

What hardware do you need to host an LLM?

It depends on the model size, document length and number of concurrent users. A small model can run on a single GPU, or even a workstation; a large model needs several high-memory GPUs. We size the setup based on load measurements.

Is self-hosting cheaper than an API?

Not always. It mainly pays off at high, steady volumes when GPUs are well utilized. For moderate or irregular use, an API is usually cheaper, especially once operating costs are included.

Can AI be deployed on premises without an internet connection?

Yes. The model, the document base and the application can run on an isolated network. You then need a procedure for updating models and components, and suitable hardware on site.

Do open model licenses allow commercial use?

Most common models allow it, but terms vary from one license to another (usage restrictions, attribution requirements, user thresholds). We check them for each model; review by your legal team is still recommended.

Let’s find out whether self-hosting fits your project.

Tell us about your use case, your data and your hosting constraints. Together we will compare the options, from a European API to an on-premises model.

Book a 30-min call with a tech lead

What are you looking for?