Data that must not leave
Medical records, banking data, trade secrets, classified information or customer contracts that forbid third-party services. This is the most common reason.
AI integration · Open models
Self-hosted AI relies on an open language model running on your own infrastructure, in your cloud or on premises, rather than on a provider’s API. Etixio helps you decide whether it makes sense, then takes it to production: model selection and evaluation, inference server, GPU sizing, security, monitoring and a cost comparison with an API.
Understanding the technology
Several developers publish the weights of their models: among the most widely used families are Llama (Meta), Mistral, Qwen (Alibaba) and Gemma (Google). These models can be downloaded and run on your servers, and the data sent to them never leaves your environment.
“Open source” is often shorthand: these are more accurately open-weight models, and their licenses vary. Some are very permissive (Apache 2.0, MIT), while others impose usage or attribution conditions. We check each model’s license before selecting it for commercial use.

Relevant or not
Medical records, banking data, trade secrets, classified information or customer contracts that forbid third-party services. This is the most common reason.
Industrial sites, networks without internet access, embedded devices. A compact model can run locally.
Under a large, stable load, well-utilized infrastructure can cost less than per-token billing. The math has to be done case by case.
A frozen version with no provider-imposed retirement, fine-tuning a model on your business vocabulary, full control of the stack.
Conversely, self-hosting is rarely the right choice for a first project with uncertain volume, for tasks that require the most advanced models on the market, or if nobody can operate GPU infrastructure over time. A European API, such as Mistral AI’s, or a major provider’s API with suitable contractual commitments may be enough.
What we work on
Comparing several open models of different sizes on your examples, against an API-based model used as a baseline. We pick the smallest model that reaches the expected quality level.
vLLM to serve a model to many users with good throughput, Ollama or llama.cpp for workstations or lightweight use. These tools expose an API that follows common industry conventions, which simplifies integration.
The GPU memory required depends on model size, context length and the number of concurrent requests. Quantization reduces memory needs, with an impact on quality that has to be measured.
GPUs rented from a cloud provider (including European ones), dedicated servers or on-premises hardware, containerized deployment with Docker and Kubernetes, scaling and high availability as needed.
An embedding model and vector database also hosted on your side, for a fully internal RAG setup, and tool calling for agents when the model supports it.
Network isolation, authentication in front of the inference server, encryption, permission management in the application, controlled logs and component updates.
The choices that matter
An API is paid per use: cost follows volume, with no upfront investment. A self-hosted model is paid per capacity: GPUs rented or bought, whether used or not, plus the engineering time to operate, monitor and update them. At low volume or with irregular load, the API is almost always cheaper; at high, steady volume, the balance can tip the other way.
We build the comparison on your estimated volumes, including operations, and keep an architecture that can move from one to the other. A hybrid approach is common: an internal model for sensitive data, an API for everything else.

In the field
For an assistant that queries confidential documents, the entire chain stays within the client’s environment: document extraction and chunking, embeddings, vector database, a generation model served by vLLM, the application and its logs. The evaluation set compares the selected model to a baseline so the quality gap accepted in exchange for confidentiality is known. Monitoring tracks GPU utilization, latency and queues.
From work to deliverables
Evaluation and monitoring practices are detailed on our LLMOps and AI evaluation page. If a prototype already exists, we start from it with our AI POC to production offering. For a software vendor whose customers require controlled hosting, a self-hosted model can also power AI features built into the product.
Frequently asked questions
It is an AI solution whose language model runs on infrastructure you control (private cloud, dedicated servers or on-premises hardware) instead of being called through a provider’s API. The data processed stays within your environment.
On many targeted tasks (extraction, classification, document search, summarization), the best open models deliver satisfactory results. For complex reasoning or long-running agents, the most advanced proprietary models often keep an edge. Only an evaluation on your examples can settle it.
It depends on the model size, document length and number of concurrent users. A small model can run on a single GPU, or even a workstation; a large model needs several high-memory GPUs. We size the setup based on load measurements.
Not always. It mainly pays off at high, steady volumes when GPUs are well utilized. For moderate or irregular use, an API is usually cheaper, especially once operating costs are included.
Yes. The model, the document base and the application can run on an isolated network. You then need a procedure for updating models and components, and suitable hardware on site.
Most common models allow it, but terms vary from one license to another (usage restrictions, attribution requirements, user thresholds). We check them for each model; review by your legal team is still recommended.
Tell us about your use case, your data and your hosting constraints. Together we will compare the options, from a European API to an on-premises model.