
The problem
- Context
- Evoriqa is a white-label AI customer-support platform for two kinds of customer: support teams running their own AI agent, and agencies running many client workspaces under their own brand. A business uploads its knowledge (files, web pages, FAQs), brands a chat widget, embeds it with one script tag, and connects the same agent to WhatsApp, Instagram, Messenger, Slack, SMS, email and a phone line, with a person able to take over at any point.
- Constraints
- Every tenant's data has to stay isolated, and the agent may only answer from that tenant's own knowledge. Every conversation has a real model cost that the platform pays, so usage has to be metered. Each messaging channel has its own webhooks and limits.
- What was at stake
- A retrieval or isolation bug would put one business's content into another's answers, or let the agent invent facts customers act on, which ends trust in a support product. Uncontrolled model cost turns a busy or abusive tenant into a loss. And an unsupervised agent answering wrongly on a live channel damages the business's own customer relationships.
What I built
Architecture
One Postgres for data and vectors
Tenants, billing and vector search all live in PostgreSQL with pgvector, so a tenant's documents and their embeddings fall under the same isolation. The trade-off is one database to scale instead of a dedicated vector store.
Knowledge processed off the request path
Uploads are processed by a background worker with retries, so a large upload never slows the app, and an edit made while a source is being processed is never lost.
Retrieval and accuracy
Hybrid search, ranked together
Each question combines semantic and keyword search, scoped to one workspace and one agent, so exact terms and product names are found as well as paraphrases, and questions in Arabic or Chinese are not silently ignored.
Grounded, or a fixed fallback
When the knowledge base holds no good answer, the agent gives a fixed reply instead of guessing, and spends nothing on the model to do it.
Safe across embedding changes
Switching embedding providers does not scramble search results while the knowledge base is being re-indexed.
Measured quality, not assumed
Test cases can be run against the live agent and scored by a model judge, and individual replies can be scored in the background. Each answer stores what it was based on (the chunks, scores and prompt) so a team can see why it said what it said.
Cost control
One credit balance for every AI action
Chat, voice, actions and knowledge ingestion all draw from a single credit balance, so a workspace can see exactly what its AI use costs.
Caps and warnings, not surprise bills
Workspaces can set a monthly spend cap and owners are emailed as usage crosses thresholds. When credits run out, customer messages still reach a person instead of being lost.
Earning autonomy
Shadow mode before autonomy
Between human-only and fully automatic, the agent can draft every reply into the team inbox, and nothing reaches the customer until a person sends it. Per-channel numbers (drafts, approvals and how much people edit) show when a channel is ready to run on its own.
Live human takeover
A teammate can claim any conversation and reply in real time, with every open inbox kept up to date.
Security
Secrets encrypted, staff kept apart
Integration credentials and sign-in secrets are encrypted at rest, users can turn on two-factor sign-in, and platform staff use their own accounts, separate from customers.
Channels and resale
One agent on every channel
The web widget, WhatsApp, Messenger, Instagram, Slack, SMS, email and a phone line all feed one workspace, with business hours and human handoff per agent.
Agencies resell under their own brand
Agencies hold a pooled credit balance, allocate it to client workspaces, connect their own Stripe account to bill those clients at their own prices, and can send email through their own provider.
A working agent with no sign-up
A visitor pastes a website address and gets a live, shareable agent built from that site, which can be claimed into a real workspace with its knowledge intact.
How it works
Channels in
Web chat, WhatsApp, Instagram
Tenant router
Isolated workspace
Retrieval
pgvector, per tenant
Model
Grounded answer
Shadow-mode gate
A person approves
Metering
Credits, soft caps
Reply out
Tell me what you’re building and where it’s stuck.
I’ll tell you the cleanest path forward, including if it’s “don’t build that.”
Or write tocontact@alihassan.dev
