Skip to content
Ali Hassan — home

Work

AI Doctor

Describe a concern by voice, attach a photo, and get a streamed, structured assessment with a Low / Moderate / High urgency badge and an optional spoken reply.

A technology demonstration. Not a medical device, and not medical advice.

Role
Full-stack engineer
Status
Live demo
  • Next.js
  • React
  • TypeScript
  • Groq (vision, Whisper, Orpheus)
  • ElevenLabs (Scribe, TTS)
View live
Consultation with an uploaded photo and urgency badge, the push-to-talk recording state, and the app's start screen.

The problem

Context
A person describes a health concern by voice, optionally attaches a photo, and gets back a structured assessment with an urgency level, possible causes, self-care steps and when to seek care, read aloud if they want. The conversation is multi-turn: the photo and earlier answers stay in context for follow-up questions.
Constraints
It runs on serverless hosting, so nothing durable can live on the filesystem and concurrent users must never share state. Three model calls are chained (speech-to-text, a vision model, text-to-speech), yet it has to feel responsive. Two providers per voice step had to be swappable behind one switch, within each provider's input limits.
What was at stake
The first version passed audio and images through files on disk under shared names: broken on serverless, and a way for one person's photo or recording to surface in someone else's session. Speech models also invent filler phrases on silent audio, which would turn into made-up symptoms.

What I built

Architecture

  • A stateless, in-memory pipeline

    Audio arrives as in-memory form data and the image travels inside the chat message itself, so nothing is written to the server. That made it safe on serverless and impossible for concurrent users to collide; the trade-off is that a conversation does not survive a reload.

  • Streaming the answer as it is written

    The assessment streams token by token from the vision model, with proxy buffering switched off so the first words appear immediately. A provider error mid-answer is appended as a short note instead of dropping the connection.

Structure without losing streaming

  • Urgency parsed from the first line

    The model is instructed to open every reply with its urgency level, followed by fixed sections. The client turns that first line into a Low, Moderate or High badge while the rest renders as it arrives, and hides the half-typed line until it resolves.

  • Two providers per voice step

    Speech-to-text and text-to-speech each have two providers behind one toggle. Switching re-voices the latest reply, so cost and quality can be compared on exactly the same answer.

Accuracy

  • A filter for invented transcripts

    Transcripts are normalised and checked against the filler phrases speech models produce on silence; a match is treated as no speech rather than as a symptom. Transcription also runs at temperature zero. The limit: it catches a phrase that is the whole transcript, not filler inside real speech.

Access

  • A gate the browser cannot bypass

    The app's code is only sent to a visitor holding a valid signed session; everyone else receives the sign-in screen alone. Every model endpoint checks the session again on its own, so no single check is load-bearing.

  • Signed sessions and constant-time checks

    Sessions are HMAC-signed with an expiry, credentials are compared in constant time, and five failed attempts lock an address out for three hours. It fails closed if credentials are not configured.

Tell me what you’re building and where it’s stuck.

I’ll tell you the cleanest path forward, including if it’s “don’t build that.”

Or write tocontact@alihassan.dev

Ali Hassan in a dark winter jacket, looking off to one side, standing in a stone courtyard with a minaret and cloudy sky behind him.