Skip to content
Ali Hassan — home

RAG development 2–5 weeks

RAG development for a product you already run

I add retrieval, an AI agent or an LLM feature to a product you already run, usually in 2–5 weeks. Answers come from your own content and say where they came from, the way AI PDF Chat cites the PDF and page.

Who it is for

  • SaaS products that want AI answers over their own documents, help centre or data.
  • Teams with a working product and no one who has shipped retrieval to production.

What the work covers

  • Ingestion of your content in the background: files, pages and records, chunked and embedded.
  • Retrieval scoped to each customer, combining semantic and keyword search.
  • Answers that cite their sources, and a fixed reply when nothing relevant is found.
  • A test set built from real questions, so every change is checked before release.
  • Usage metered per customer, so cost per request stays visible.

Proof

  • PDF chat with document scope selector, landing page and upload queue of a RAG assistant that cites its pages.

    AI PDF Chat

    Upload several PDFs and ask questions across all of them or one; every answer cites the PDF and page, and the citation opens the passage highlighted.

    • Next.js
    • Express
    • LangChain
  • AI reply grounded in cited help-centre content, shadow-mode draft metrics per channel, and metered AI credits in billing.

    Evoriqa

    My own SaaS: a multi-tenant AI customer-support platform, live in production, built by the person who also pays its inference bill.

    • Next.js
    • TypeScript
    • PostgreSQL

Questions

Do we need a separate vector database?
Often not. If you already run PostgreSQL, pgvector keeps documents and their embeddings under the same access rules; a dedicated vector store earns its place at larger scale. My pgvector vs Pinecone article sets out the trade-off.
How do you know the answers are right?
A test set built from real questions, run on every change to the prompt, model or chunking, plus a check that each answer is grounded in the sources it cites.
Is this the same as fine-tuning?
No. RAG changes what the model can see when it answers; fine-tuning changes how it behaves. For facts that change, or differ per customer, RAG is usually the right place to start.

Tell me what you’re building and where it’s stuck.

I’ll tell you the cleanest path forward, including if it’s “don’t build that.”

Or write tocontact@alihassan.dev

Ali Hassan in a dark winter jacket, looking off to one side, standing in a stone courtyard with a minaret and cloudy sky behind him.