AI Doctor
Describe a concern by voice, attach a photo, and get a streamed, structured assessment with a Low / Moderate / High urgency badge and an optional spoken reply.
A technology demonstration. Not a medical device, and not medical advice.
- Role
- Full-stack engineer
- Status
- Live demo
- Next.js
- React
- TypeScript
- Groq (vision, Whisper, Orpheus)
- ElevenLabs (Scribe, TTS)

The problem
- Context
- A person describes a health concern by voice, optionally attaches a photo, and gets back a structured assessment with an urgency level, possible causes, self-care steps and when to seek care, read aloud if they want. The conversation is multi-turn: the photo and earlier answers stay in context for follow-up questions.
- Constraints
- It runs on serverless hosting, so nothing durable can live on the filesystem and concurrent users must never share state. Three model calls are chained (speech-to-text, a vision model, text-to-speech), yet it has to feel responsive. Two providers per voice step had to be swappable behind one switch, within each provider's input limits.
- What was at stake
- The first version passed audio and images through files on disk under shared names: broken on serverless, and a way for one person's photo or recording to surface in someone else's session. Speech models also invent filler phrases on silent audio, which would turn into made-up symptoms.
What I built
Architecture
A stateless, in-memory pipeline
Audio arrives as in-memory form data and the image travels inside the chat message itself, so nothing is written to the server. That made it safe on serverless and impossible for concurrent users to collide; the trade-off is that a conversation does not survive a reload.
Streaming the answer as it is written
The assessment streams token by token from the vision model, with proxy buffering switched off so the first words appear immediately. A provider error mid-answer is appended as a short note instead of dropping the connection.
Structure without losing streaming
Urgency parsed from the first line
The model is instructed to open every reply with its urgency level, followed by fixed sections. The client turns that first line into a Low, Moderate or High badge while the rest renders as it arrives, and hides the half-typed line until it resolves.
Two providers per voice step
Speech-to-text and text-to-speech each have two providers behind one toggle. Switching re-voices the latest reply, so cost and quality can be compared on exactly the same answer.
Accuracy
A filter for invented transcripts
Transcripts are normalised and checked against the filler phrases speech models produce on silence; a match is treated as no speech rather than as a symptom. Transcription also runs at temperature zero. The limit: it catches a phrase that is the whole transcript, not filler inside real speech.
Access
A gate the browser cannot bypass
The app's code is only sent to a visitor holding a valid signed session; everyone else receives the sign-in screen alone. Every model endpoint checks the session again on its own, so no single check is load-bearing.
Signed sessions and constant-time checks
Sessions are HMAC-signed with an expiry, credentials are compared in constant time, and five failed attempts lock an address out for three hours. It fails closed if credentials are not configured.
Tell me what you’re building and where it’s stuck.
I’ll tell you the cleanest path forward, including if it’s “don’t build that.”
Or write tocontact@alihassan.dev
