Skip to content
Ali Hassan — home

Work

AI PDF Chat

Upload several PDFs and ask questions across all of them or one; every answer cites the PDF and page, and the citation opens the document with the passage highlighted.

Role
Full-stack engineer
Status
Built
  • Next.js
  • Express
  • LangChain
  • Qdrant
  • Gemini embeddings
  • Groq (Llama)
  • Zod
  • Docker Compose
PDF chat with document scope selector, landing page and upload queue of a RAG assistant that cites its pages.

The problem

Context
A signed-in user uploads several PDFs and asks questions across all of them or inside one. Every answer cites the PDF and page, and the citation opens the document with the passage highlighted.
Constraints
The embedding provider's free tier rate-limits hard and, through the library in use, returns empty results instead of an error when it does. Large PDFs produce many chunks, so processing cannot happen inside the upload request. Many users share one vector store, and the model provider had to be swappable.
What was at stake
Embedding during the upload would time out on big files. A missed rate limit would quietly fill the index with empty vectors and degrade every answer. A missing user filter would show one person's documents to another, and vectors left behind after a delete would keep surfacing deleted content.

What I built

Background ingestion

  • Uploads return at once

    The upload responds immediately and each file is extracted, chunked and embedded by a background queue, one file at a time, with its status (queued, processing, ready, failed) shown in the app. One file failing never holds up the rest.

  • Work resumes after a restart

    On startup, files that were mid-processing are put back in the queue, so a restart never leaves a document stuck.

  • A swappable model provider

    Chat and embeddings are chosen from configuration, with a hosted provider by default and a local model as an option, so the vendor can change without touching the pipeline.

Reliability under rate limits

  • Catching the silent rate limit

    Empty or short vectors are treated as a hidden rate-limit response and retried with exponential backoff; if it never recovers, the user sees a plain message about the limit. Batches are paced to stay under it in the first place.

  • Safe to retry, safe to crash

    Each re-ingest clears the file's old vectors first, so retries never create duplicates. A file deleted mid-processing is cleaned up instead of resurrected, and a scanned PDF with no text is marked failed with a useful message rather than shown as ready.

Isolation and clean deletes

  • Every search filtered by user

    Searches always filter by the signed-in user, and by the file in single-document mode, on indexed fields, with a relevance threshold so weak matches are not passed to the model.

  • Ownership checked everywhere

    Viewing, deleting and retrying files, and reading chat history, are all scoped to the signed-in user; the raw PDF route replaced an earlier open folder. Sign-in endpoints are rate-limited separately.

  • Deletes that leave nothing behind

    Deleting a PDF removes the file, its record and its vectors, and the conversations that cited only that file, while keeping conversations that drew on other documents too.

Answers you can check

  • Streamed answers with citations

    Answers stream as they are written and are restricted to the retrieved content. Citations are grouped per file and page and carry the best-matching passage, which the viewer highlights; recent turns are kept so follow-up questions make sense.

Tell me what you’re building and where it’s stuck.

I’ll tell you the cleanest path forward, including if it’s “don’t build that.”

Or write tocontact@alihassan.dev

Ali Hassan in a dark winter jacket, looking off to one side, standing in a stone courtyard with a minaret and cloudy sky behind him.