Decide based on measurements
Choosing a model, editing a prompt or changing the document retrieval strategy becomes a quantified decision, compared against a baseline.
AI industrialization · LLMOps
LLMOps covers the practices that let you run an AI solution built on language models like serious software: evaluating answer quality, testing every change, adding guardrails, monitoring costs and latency, managing model versions and documenting compliance. Etixio puts these practices in place on the assistants, agents and AI features we build or take over.
Understanding the discipline
A language model does not behave like a regular function: the same question can get different answers, a prompt change can fix one case and break ten others, and providers keep updating their models. Without measurement, the quality of an AI service rests on impressions.
LLMOps applies DevOps habits to these systems: automated tests, controlled deployments, monitoring and version management. It adds what is specific to generative AI: evaluation sets, answer quality criteria, guardrails and token usage tracking.

What it changes
Choosing a model, editing a prompt or changing the document retrieval strategy becomes a quantified decision, compared against a baseline.
An update that degrades answers is caught before it reaches users, not several weeks later.
Usage per feature, per customer or per user is known, capped and optimized. Response times are tracked.
Traceability, documentation and human oversight are expected under GDPR and the EU AI Act depending on the use case. LLMOps produces the evidence as you go.
What we work on
Representative examples built with your business teams, including edge cases, out-of-scope questions and situations where the system should refuse or ask for clarification.
Accuracy against an expected answer, faithfulness to sources, format compliance, rate of justified refusals. Automated checks, LLM-as-a-judge scoring calibrated on human annotations, and human review of samples.
Prompts versioned in the codebase, evaluations re-run in the CI pipeline on every change, blocking thresholds before deployment.
Validation of structured outputs, filtering of personal data, prompt injection detection, restricted tool access and human confirmation before any sensitive action.
Traces of every call (inputs, outputs, tools called, documents retrieved), errors and latency, using tools such as Langfuse or OpenTelemetry, while protecting the confidentiality of logged data.
Token tracking per feature, caps, caching, a more compact model for simple steps, and batch processing when real time is not needed.
Over time
Providers release new models and retire old ones. We pin the model version explicitly, track deprecation notices and re-run the evaluation set before any switch. A new version can be rolled out gradually or compared in parallel on part of the traffic.
User feedback (rating an answer, correcting it, flagging it) is collected in the interface, analyzed and turned into new test cases. The evaluation set thus grows with the real situations encountered in production.
In the field
A business team reports answers that are too vague for a certain type of case file. The affected cases are added to the evaluation set, the prompt is changed in a branch, and the CI pipeline re-runs all tests and compares the results with the production version: quality, format, average cost and latency. The change is merged only if it improves the targeted cases without degrading the others. After deployment, traces and user feedback confirm whether it had the intended effect.

The choices that matter
GDPR applies as soon as personal data is sent to a model or logged: legal basis, data minimization, trace retention periods and the provider’s processing terms. The EU AI Act, whose obligations are being phased in, imposes requirements depending on the risk level of the use case, such as transparency toward users, documentation, logging and human oversight.
We do not provide legal advice, but we build the service so these requirements can be verified: usable logs, system documentation and human validation where it is required. See our article on AI precautions, risks and compliance.
From work to deliverables
The setup can come with a new build or be added to an existing service: it is often the first step when taking an AI POC to production. It works with any provider: Claude, OpenAI, Gemini, Mistral AI or a self-hosted model.
Frequently asked questions
LLMOps is the set of engineering practices used to deploy, measure and maintain applications built on language models. It covers quality evaluation, regression testing, guardrails, observability, cost tracking and model version management.
With a set of examples that reflect your real requests, expected answers or criteria defined with subject-matter experts, and suitable metrics (accuracy, faithfulness to sources, format, refusals). Public benchmarks give a general indication but do not replace testing on your data.
Partly. LLM-as-a-judge scoring makes it possible to grade many answers quickly, but it must be calibrated on human annotations and checked regularly. We combine it with deterministic automated checks and human review of samples.
It depends on how varied the requests are and how risky the use case is. Start with a small set that covers the main situations and edge cases, then grow it with cases seen in production. Coverage matters more than volume.
By measuring usage per feature, capping it, caching what can be cached, reserving the most powerful models for the steps that need them, and using evaluations to confirm that a more compact model is good enough elsewhere.
Yes. We start by instrumenting the service and building a first evaluation set to measure where things stand, then add tests, guardrails and dashboards in order of priority.
Tell us about your AI service, existing or planned, and your quality requirements. Together we will identify what to measure first.