Skip to content
Ali Hassan — home

Blog

pgvector vs Pinecone: choosing a vector store

pgvector vs Pinecone for RAG: vectors kept in PostgreSQL or in a managed vector database, compared on isolation, filtering, hybrid search, operations and scale.

By Ali HassanSenior engineer for SaaS and AI products

If you already run PostgreSQL and your vectors sit next to data you query together, pgvector is usually the simpler choice: one database to operate, back up and reason about. Pinecone makes more sense when vector search is the dominant workload, needs to scale on its own, and you would rather not run the storage and indexing yourself. Most of the pgvector vs Pinecone decision comes down to a single question: do you want one system or two.

What each one is

pgvector is an open-source extension for PostgreSQL. It adds a vector column type and similarity search, with approximate indexes such as HNSW and IVFFlat so a query does not have to scan every row. Because it lives inside Postgres, your embeddings sit in the same database as the rest of your data, and you search them with SQL.

Pinecone is a managed, hosted vector database. You send it vectors and metadata through its API, and it handles the indexing, storage and scaling for you. There is no server for you to run or tune; you work through its client libraries and pay for the service.

Both do the same core job well: store embeddings and return the nearest matches for a query vector. The difference is everything around that job.

One system or two

This is the trade-off that decides most cases.

With pgvector, your relational data and your vectors are one system. A document, its owner, its permissions and its embeddings are all rows in the same database, so you can filter and join them in a single query, wrap a write to all of them in one transaction, and back them up together. There is one thing to run, one connection pool, one set of credentials, one restore to test.

A similarity search becomes an ordinary SQL query that can join and filter as it ranks:

-- Nearest chunks for one tenant, joined to the document they came from
SELECT d.title, c.content
FROM chunks c
JOIN documents d ON d.id = c.document_id
WHERE c.tenant_id = $1
  AND d.status = 'published'
ORDER BY c.embedding <=> $2
LIMIT 5;

With Pinecone you have two systems. Your relational data lives in your main database and your vectors live in the service, so a join like the one above happens in your application code or through metadata stored alongside each vector. In return, the operational work of running and scaling vector search is handled for you, and that search scales on its own rather than competing with the rest of your database for memory and CPU.

Neither is the right answer in the abstract. If your vectors are one feature inside a product that is mostly relational, keeping them in Postgres removes a moving part. If vector search is the product, a service built only for it earns its place.

Keeping each customer's data isolated

In a multi-tenant product, the rule is usually that one customer's content must never appear in another's results. Where you enforce that rule depends on the store.

With pgvector, a tenant's documents and their embeddings are rows in the same database, so they fall under the same isolation as everything else you already control there, enforced with the same tools you use for any other table. On Evoriqa, my own multi-tenant support platform, tenants, billing and vector search all live in PostgreSQL with pgvector, so a tenant's documents and their embeddings fall under the same isolation, and the trade-off is one database to scale instead of a dedicated vector store.

Pinecone isolates tenants with namespaces or metadata filters: you partition each customer's vectors and scope every query to one of them. It works, but the isolation lives in a second system that you keep in step with your main database rather than in the database itself.

The isolation question applies whatever you pick. AI PDF Chat, a RAG assistant I built, keeps per-user vector isolation in its vector store, which is Qdrant. The store differs; the discipline of scoping every search to its owner does not.

Pure vector search finds paraphrases well and exact terms badly. A product name, an error code or an acronym is often matched better by keyword search than by similarity, so most real systems want both.

This is where Postgres being one system pays off again. You already have full-text search in the same database, so you can run keyword search alongside vector search and combine the two rankings, all in SQL, all filtered by the same tenant and permission rules. On Evoriqa, each question combines semantic and keyword search, scoped to one workspace and one agent, so exact terms and product names are found as well as paraphrases.

Pinecone offers metadata filtering and its own support for sparse-dense hybrid search. The capability is there; the difference is that your keyword signal and your relational filters come from a separate place rather than from the database that already holds your text.

Scale and performance

I will not quote numbers here, because the honest answer is that it depends on your data, your index settings and your hardware. But the shape of the difference is clear.

pgvector runs inside your Postgres instance and shares its memory and CPU with every other query. Performance is a function of how you size that instance and how you tune the index: HNSW and IVFFlat each trade build time, memory and recall against query speed, and you own those choices. For many products this is more than enough, and it keeps vector search close to the data.

Pinecone scales vector search independently of your relational database. When the vector workload is very large or the query rate is high, moving it off your main database so it cannot contend with ordinary traffic is a real advantage. You are paying for someone else to solve the scaling problem, which is worth it precisely when that problem is hard.

Cost shape

The two have different cost shapes, which matters more than any single figure.

pgvector adds no separate bill. It is part of a database you already run, so the cost is the storage, memory and compute of that Postgres instance, which grows as your vectors and indexes do. There is nothing new to buy if your database has headroom.

Pinecone is a separate, usage-based service billed on top of whatever you already pay for your database. That can be the cheaper path when it saves you from over-provisioning Postgres for a demanding vector workload, and the more expensive one when your needs are modest and a database you already run would have absorbed them.

Migration effort later

Picking one now does not lock you in forever, but moving later is real work, so it is worth knowing the shape of it.

Your embeddings are portable as long as you keep the same embedding model, so the vectors themselves can be exported and re-loaded. What changes is everything around them: the query interface (SQL against pgvector versus Pinecone's API), how you express filters, and how tenant isolation is enforced. Moving from Pinecone to pgvector gains you joins and transactions with your relational data; moving the other way gives those up and shifts the ranking and filtering logic into your application. Plan for rewriting the retrieval layer, not just copying vectors.

pgvector vs Pinecone at a glance

ConcernpgvectorPinecone
Systems to runOne (your Postgres)Two (database plus service)
Data and vectorsTogether, in one databaseVectors separate from relational data
Joins and transactionsNative SQL, across bothIn application code
Operations and scalingYou run and tune itManaged; scales on its own
Tenant isolationSame rules as the rest of the databaseNamespaces or metadata filters
Keyword and hybrid searchPostgres full-text search in the same queryMetadata filters and hybrid support in the service
Cost shapePart of an existing databaseSeparate, usage-based bill
Best fitVectors are one feature of a relational productVector search is the dominant workload

Choosing between them

A short decision guide:

  • You already run PostgreSQL and vectors are one feature among many relational ones. Use pgvector. One system is less to operate, and you get joins, transactions and full-text search for free.
  • You need vector search to scale independently, at a size or query rate that would strain your database, and you would rather not run it yourself. Use Pinecone.
  • Your isolation, filtering and ranking are tightly bound to relational data. Lean towards pgvector, so that logic stays in one place.
  • Vector search is effectively the product, not a feature of something larger. A dedicated service is easier to justify.

My own default, for products that are mostly a normal application with a retrieval feature bolted on, is to start in Postgres with pgvector and only reach for a dedicated service when there is a measured reason to. Fewer moving parts is a feature. For how retrieval fits the wider system around whichever store you pick, see RAG architecture.

When to bring in help

If you are weighing this choice for a product with real users, the store is only one of several decisions that have to agree with each other: how you chunk and embed, how you filter, how you keep tenants apart, and how you check that answers are actually grounded. I build retrieval features end to end, and I am happy to talk it through if a second opinion would help. That is what my RAG development work covers.

From the work

  • AI reply grounded in cited help-centre content, shadow-mode draft metrics per channel, and metered AI credits in billing.

    Evoriqa

    My own SaaS: a multi-tenant AI customer-support platform, live in production, built by the person who also pays its inference bill.

    • Next.js
    • TypeScript
    • PostgreSQL
  • PDF chat with document scope selector, landing page and upload queue of a RAG assistant that cites its pages.

    AI PDF Chat

    Upload several PDFs and ask questions across all of them or one; every answer cites the PDF and page, and the citation opens the passage highlighted.

    • Next.js
    • Express
    • LangChain

The offer this article leads to: RAG development, typically $4k–$10k.

Tell me what you’re building and where it’s stuck.

I’ll tell you the cleanest path forward, including if it’s “don’t build that.”

Or write tocontact@alihassan.dev

Ali Hassan in a dark winter jacket, looking off to one side, standing in a stone courtyard with a minaret and cloudy sky behind him.