Internal documentation and procedures
Find a rule, a quality procedure or an HR answer across hundreds of documents, with a link to the exact passage.
AI integration · RAG
Etixio builds RAG (retrieval augmented generation) assistants: applications that search your documents and databases for the relevant passages, then ask a language model to write an answer that cites its sources. We handle the entire pipeline, from document ingestion to a monitored production service, while enforcing each user’s access rights.
Understanding the technology
A language model knows nothing about your procedures, your contracts or your product documentation. RAG means giving it the relevant excerpts from your sources at the moment the question is asked. The application first retrieves the useful passages, sends them to the model with precise instructions, then displays an answer along with its references.
This approach avoids retraining a model, lets you update knowledge by reindexing documents and makes every answer verifiable. It sits at the heart of many enterprise AI agents and chatbots and AI features built into products.

Business use cases
Find a rule, a quality procedure or an HR answer across hundreds of documents, with a link to the exact passage.
Draft an answer from product sheets, manuals and previously resolved tickets, before an agent validates it.
Pull together the information in a large file (contract, customer file, patient record) and return a structured summary.
Let the users of a software product query their own content, with a separate index per customer.
Our scope of work
The quality of a RAG assistant depends first and foremost on retrieval. If the right passages are not found, no model will produce a correct answer. So we work on every step of the pipeline, and we measure each one.
Text extraction from PDFs, office documents, web pages or databases, cleanup, preservation of metadata (author, date, department, confidentiality level) and chunking into passages that respect the document’s structure.
Vector representations computed with an embedding model chosen for the language and domain, then stored in a database that fits your infrastructure (PostgreSQL with pgvector, OpenSearch, Qdrant or a managed service).
Semantic search combined with keyword search, metadata filters, then reranking so that only the most relevant passages are sent to the model.
Instructions that require the model to answer only from the excerpts provided, cite its sources and say when the information is missing rather than make it up.
Documents filtered by the user’s permissions at retrieval time, based on your identity management. A user never receives an excerpt they could not open themselves.
Scheduled or event-driven synchronization, handling of modified or deleted documents and traceability of the indexed version.
Measure rather than assume
Together with your business teams, we build a reference set of questions, including questions whose answer is not in the documents and out-of-scope questions. We measure retrieval on one side (are the right passages found?) and the answer on the other (is it faithful to the sources, complete, correctly cited?).
These evaluations are rerun every time the chunking, the embedding model, the language model or the instructions change. An improvement on one case that degrades ten others is caught before the production release. See our approach to LLMOps and AI evaluation.
Operations and costs
Logging of questions, retrieved passages and answers, tracking of errors, response times and user feedback, in line with your data retention rules.
Cost depends on indexing volume, the number of passages sent and the model chosen. We tune context size, use a lighter model when it is good enough and cache what can be cached.
Choice of model provider and hosting location based on data sensitivity, up to self-hosted open-source AI or a European model such as Mistral AI when required.

In the field
For a hospital, Etixio designed a RAG-based medical AI agent that explores the patient record and returns a chronological summary to anesthesia teams. We also built Bobby, a conversational assistant connected to our ERP data. In both cases, the work was as much about data access and the interface as about the model.
The choices that matter
If questions are about figures stored in a database, a structured query generated and checked by the application will be more reliable than searching through text. If the documents fit within the model’s context, elaborate chunking may be unnecessary. If classic search does the job, an assistant does not necessarily add value.
We raise these questions during scoping. The model is then chosen based on your examples, comparing Claude, OpenAI, Gemini or a hosted model.
From work to deliverables
An existing RAG prototype can be taken over and hardened; see our AI POC to production offering. The work is delivered as a fixed-price project or with a dedicated team if the service needs to evolve over time.
Frequently asked questions
RAG (retrieval augmented generation) combines a search engine with a language model. The application retrieves the relevant passages from your sources, then the model writes an answer based on those passages and cites its references.
Fine-tuning changes a model’s behavior through additional training; it is mainly used to adjust a style or format. RAG brings in knowledge at question time, with no retraining. To answer from documents that change over time, RAG is generally a better fit and easier to maintain.
It has to, and we address it from the design stage. Permissions are enforced at retrieval time, based on your identity management, so a user only receives excerpts from documents they can already access.
No system eliminates this risk entirely. We reduce it by requiring answers based only on the excerpts provided, displaying sources, allowing the assistant to say it does not know and measuring answer faithfulness on an evaluation set.
It depends on your infrastructure and volume. If you already run PostgreSQL, the pgvector extension is often enough. For larger volumes or advanced hybrid search, a dedicated engine may be preferable. We choose with you based on what your team can operate.
Yes. We measure its results on real questions, identify whether errors come from ingestion, retrieval or generation, then add what is missing for production — access rights, evaluations, monitoring and cost control.
Tell us about your documents, the users involved and the questions they ask. Together we will define the initial scope to explore.