Skip to content
Ali Hassan — home

Blog

RAG architecture: how a production pipeline fits together

How a production RAG architecture fits together: ingestion, chunking, hybrid retrieval, cited answers, fallbacks and evaluation, and where each part fails.

By Ali HassanSenior engineer for SaaS and AI products

RAG architecture is the set of moving parts that lets a language model answer from your own content instead of only its training data: you ingest documents, store them as searchable vectors, retrieve the relevant pieces when a question comes in, and pass those to the model to write a grounded answer. A production pipeline adds the parts a weekend demo skips: background processing, per-tenant isolation, citations, a fixed reply when nothing matches, and tracing. Here is how those stages fit together, and where each one tends to break once real users and real data arrive.

What a RAG pipeline is

RAG stands for retrieval-augmented generation. The model on its own knows what it was trained on; it does not know your handbook, your product docs or last week's support tickets. RAG closes that gap by putting a search step in front of the model. At question time you find the handful of passages most likely to hold the answer, put them in the prompt, and ask the model to answer from them and nothing else.

That gives you two things a bare model call cannot: answers grounded in content you control, and the ability to point at where each answer came from. Everything below is in service of those two properties holding up under load.

A pipeline has two halves. One runs ahead of time, turning documents into something searchable (ingestion and storage). The other runs per question (retrieval and generation). Keeping them separate is the first design decision that matters.

Ingestion (ahead of time)        Query time (per request)
  extract -> chunk -> embed        question -> retrieve -> generate -> cite
          \                       /
           vector store (scoped per tenant)

RAG architecture, stage by stage

Ingestion: extract, chunk, embed

Ingestion takes raw files and makes them searchable. You extract the text (from PDFs, web pages, documents), split it into chunks small enough to retrieve precisely but large enough to carry meaning, and turn each chunk into an embedding, a vector the store can search by similarity.

The architectural choice that matters is that this happens off the request path. A large upload produces many chunks and many embedding calls, which is far too slow to run inside the HTTP request that accepted the file. So the upload returns straight away and a background worker does the extracting, chunking and embedding, with its status visible in the app.

I built this in AI PDF Chat, a multi-document RAG assistant: the upload responds at once and each file is extracted, chunked and embedded by a background queue, one file at a time, so a big PDF never slows the app and one file failing never holds up the rest. Evoriqa does the same with a background worker that retries, so a large upload never slows the app and an edit made while a source is being processed is never lost.

Storage: a vector store, scoped per user

The embeddings live in a vector store. This can be a dedicated vector database or a relational database with a vector extension; the trade-off is one fewer system to run against a purpose-built index.

The part that is not optional is scoping. In any multi-user or multi-tenant product, every stored vector has to carry an owner, and every search has to filter by that owner. Miss it and one customer's content turns up in another's answers, which is the worst failure a knowledge product can have.

In AI PDF Chat, searches always filter by the signed-in user, and by a single file when the question is scoped to one document, so one person's uploads never surface in another's answers. Evoriqa, my own multi-tenant AI support platform, keeps retrieval per tenant on pgvector, so a tenant's documents and their embeddings fall under the same isolation.

Retrieval: semantic plus keyword search

Retrieval is where most quality problems are won or lost. The obvious approach is pure semantic search: embed the question, find the nearest chunks. It is good at paraphrase and weak at exact strings, such as a part number, an error code, a person's name, or a term in a non-Latin script. Embeddings smooth those into something approximate, which is exactly wrong when the user typed the precise term they want.

Hybrid retrieval fixes this by running semantic and keyword search together and combining the results. Keyword search catches the exact token; semantic search catches the meaning; ranked together they cover both. This also matters for scripts that a purely semantic approach can quietly drop.

Evoriqa combines semantic and keyword search on each question, scoped to one workspace and one agent, so exact terms and product names are found as well as paraphrases, and questions in Arabic or Chinese are not silently ignored. Scoping sits here too: retrieval is filtered to the tenant, and in AI PDF Chat to the chosen document, before ranking rather than after.

One more guard belongs in retrieval: a relevance floor. In AI PDF Chat, matches too weak to be useful are dropped rather than passed to the model, so a vague question does not drag in noise that the model then treats as fact.

Generation: grounded answers with citations

Now the model runs. You give it the retrieved chunks and an instruction to answer from them only, and you stream the answer back so the first words appear quickly.

Citations are the feature that makes the whole thing trustworthy. Because you know which chunks you passed in, you can show the user where the answer came from and let them check it. In AI PDF Chat every answer cites the PDF and page, and the citation opens the document with the passage highlighted; the answer streams over SSE as it is written and is restricted to the retrieved content. Without citations a RAG answer is just a confident paragraph; with them it is a claim the reader can verify.

When nothing relevant is found

Sometimes retrieval comes back with nothing good. The dangerous response is to send whatever scraps were retrieved and let the model fill the gaps from its training, which is how a grounded system starts inventing. The correct response is a fixed reply that says it does not know.

When Evoriqa's knowledge base holds no good answer, the agent gives a fixed reply instead of guessing. That one rule is the difference between a system that is honest about its limits and one that confidently misleads, and it is cheap to add.

Tracing and evaluation

A RAG pipeline has several places an answer can go wrong: the wrong chunks retrieved, the right chunks ignored, the model drifting off them. You cannot debug that from the final text alone. The fix is to record, for each answer, what it was based on.

Evoriqa stores what each answer was based on (the chunks, scores and prompt), so a team can see why it said what it said. That trace is also what lets you evaluate quality rather than guess at it: Evoriqa can run test cases against the live agent and have them scored by a model judge, so a change to chunking or retrieval can be measured instead of eyeballed. I have written separately on how to evaluate a RAG system; the short version is that without a trace and a test set, every change is a gamble.

Where production pipelines fail

The failures are predictable, and they cluster by stage. This is the table I keep in my head when reviewing a pipeline.

ComponentWhat it doesWhat goes wrong in production
IngestionExtract, chunk and embed documentsRun inside the upload request, so large files time out; chunks too big or too small; scanned PDFs with no text indexed as empty
EmbeddingTurn text into vectorsA rate-limited provider returns empty vectors instead of an error, quietly filling the index with blanks
Vector storeHold and search embeddingsNo owner filter, so one tenant's content leaks into another's answers; stale vectors left behind after a delete
RetrievalFind the relevant chunksPure semantic search misses exact terms and non-Latin scripts; weak matches passed through as if relevant
GenerationWrite a grounded answerModel answers beyond the retrieved text; no citations, so nothing can be checked
FallbackHandle "no good answer"Guesses from training data instead of declining
Tracing and evalRecord and score answersNone exists, so quality is a matter of opinion and regressions ship unnoticed

Most of these are failures of the plumbing around the model, not the model itself. A stronger model does not save a pipeline that leaks across tenants or indexes empty vectors.

When to bring in help

If you are building a RAG feature and these stages are where you are getting stuck, ingestion that blocks, retrieval that misses the obvious term, answers nobody can verify, that is the work I do. I build retrieval over your own content with tenant isolation, citations and a measurable quality check; you can see the shape of it on RAG development. No pressure either way. If your pipeline is close, a short review often beats a rebuild.

From the work

  • PDF chat with document scope selector, landing page and upload queue of a RAG assistant that cites its pages.

    AI PDF Chat

    Upload several PDFs and ask questions across all of them or one; every answer cites the PDF and page, and the citation opens the passage highlighted.

    • Next.js
    • Express
    • LangChain
  • AI reply grounded in cited help-centre content, shadow-mode draft metrics per channel, and metered AI credits in billing.

    Evoriqa

    My own SaaS: a multi-tenant AI customer-support platform, live in production, built by the person who also pays its inference bill.

    • Next.js
    • TypeScript
    • PostgreSQL

The offer this article leads to: RAG development, typically $4k–$10k.

Tell me what you’re building and where it’s stuck.

I’ll tell you the cleanest path forward, including if it’s “don’t build that.”

Or write tocontact@alihassan.dev

Ali Hassan in a dark winter jacket, looking off to one side, standing in a stone courtyard with a minaret and cloudy sky behind him.