Software & AI · From strategy to production

AI integration · RAG

RAG Development: An AI Assistant That Answers From Your Documents

Etixio builds RAG (retrieval augmented generation) assistants: applications that search your documents and databases for the relevant passages, then ask a language model to write an answer that cites its sources. We handle the entire pipeline, from document ingestion to a monitored production service, while enforcing each user’s access rights.

Understanding the technology

The model answers from your sources, not from memory.

A language model knows nothing about your procedures, your contracts or your product documentation. RAG means giving it the relevant excerpts from your sources at the moment the question is asked. The application first retrieves the useful passages, sends them to the model with precise instructions, then displays an answer along with its references.

This approach avoids retraining a model, lets you update knowledge by reindexing documents and makes every answer verifiable. It sits at the heart of many enterprise AI agents and chatbots and AI features built into products.

Four people reviewing documents around a table

Business use cases

Where a RAG assistant is useful.

Internal documentation and procedures

Find a rule, a quality procedure or an HR answer across hundreds of documents, with a link to the exact passage.

Customer and technical support

Draft an answer from product sheets, manuals and previously resolved tickets, before an agent validates it.

Case file analysis

Pull together the information in a large file (contract, customer file, patient record) and return a structured summary.

Search inside a SaaS product

Let the users of a software product query their own content, with a separate index per customer.

Our scope of work

The building blocks of a production RAG pipeline.

The quality of a RAG assistant depends first and foremost on retrieval. If the right passages are not found, no model will produce a correct answer. So we work on every step of the pipeline, and we measure each one.

Ingestion and chunking

Text extraction from PDFs, office documents, web pages or databases, cleanup, preservation of metadata (author, date, department, confidentiality level) and chunking into passages that respect the document’s structure.

Embeddings and vector database

Vector representations computed with an embedding model chosen for the language and domain, then stored in a database that fits your infrastructure (PostgreSQL with pgvector, OpenSearch, Qdrant or a managed service).

Hybrid search and reranking

Semantic search combined with keyword search, metadata filters, then reranking so that only the most relevant passages are sent to the model.

Generation with citations

Instructions that require the model to answer only from the excerpts provided, cite its sources and say when the information is missing rather than make it up.

Access control

Documents filtered by the user’s permissions at retrieval time, based on your identity management. A user never receives an excerpt they could not open themselves.

Index updates

Scheduled or event-driven synchronization, handling of modified or deleted documents and traceability of the indexed version.

Measure rather than assume

Evaluating retrieval and answers separately.

Together with your business teams, we build a reference set of questions, including questions whose answer is not in the documents and out-of-scope questions. We measure retrieval on one side (are the right passages found?) and the answer on the other (is it faithful to the sources, complete, correctly cited?).

These evaluations are rerun every time the chunking, the embedding model, the language model or the instructions change. An improvement on one case that degrades ten others is caught before the production release. See our approach to LLMOps and AI evaluation.

Operations and costs

Monitoring the service and controlling its cost.

Monitoring

Logging of questions, retrieved passages and answers, tracking of errors, response times and user feedback, in line with your data retention rules.

Costs

Cost depends on indexing volume, the number of passages sent and the model chosen. We tune context size, use a lighter model when it is good enough and cache what can be cached.

Two developers reviewing code together

In the field

A RAG agent in a demanding context.

For a hospital, Etixio designed a RAG-based medical AI agent that explores the patient record and returns a chronological summary to anesthesia teams. We also built Bobby, a conversational assistant connected to our ERP data. In both cases, the work was as much about data access and the interface as about the model.

The choices that matter

RAG is not always the right answer.

If questions are about figures stored in a database, a structured query generated and checked by the application will be more reliable than searching through text. If the documents fit within the model’s context, elaborate chunking may be unnecessary. If classic search does the job, an assistant does not necessarily add value.

We raise these questions during scoping. The model is then chosen based on your examples, comparing Claude, OpenAI, Gemini or a hosted model.

From work to deliverables

What we deliver.

An existing RAG prototype can be taken over and hardened; see our AI POC to production offering. The work is delivered as a fixed-price project or with a dedicated team if the service needs to evolve over time.

Frequently asked questions

RAG development: your questions.

What is RAG in artificial intelligence?

RAG (retrieval augmented generation) combines a search engine with a language model. The application retrieves the relevant passages from your sources, then the model writes an answer based on those passages and cites its references.

What is the difference between RAG and fine-tuning?

Fine-tuning changes a model’s behavior through additional training; it is mainly used to adjust a style or format. RAG brings in knowledge at question time, with no retraining. To answer from documents that change over time, RAG is generally a better fit and easier to maintain.

Does a chatbot on our internal documents respect access rights?

It has to, and we address it from the design stage. Permissions are enforced at retrieval time, based on your identity management, so a user only receives excerpts from documents they can already access.

How do you keep the assistant from making up answers?

No system eliminates this risk entirely. We reduce it by requiring answers based only on the excerpts provided, displaying sources, allowing the assistant to say it does not know and measuring answer faithfulness on an evaluation set.

Which vector database should we choose?

It depends on your infrastructure and volume. If you already run PostgreSQL, the pgvector extension is often enough. For larger volumes or advanced hybrid search, a dedicated engine may be preferable. We choose with you based on what your team can operate.

Can you take over an existing RAG prototype?

Yes. We measure its results on real questions, identify whether errors come from ingestion, retrieval or generation, then add what is missing for production — access rights, evaluations, monitoring and cost control.

Let’s build an assistant that answers from your sources.

Tell us about your documents, the users involved and the questions they ask. Together we will define the initial scope to explore.

Book a 30-min call with a tech lead

What are you looking for?