pgvector vs Pinecone: choosing a vector store
pgvector vs Pinecone for RAG: vectors kept in PostgreSQL or in a managed vector database, compared on isolation, filtering, hybrid search, operations and scale.
By Ali HassanSenior engineer for SaaS and AI products
If you already run PostgreSQL and your vectors sit next to data you query together, pgvector is usually the simpler choice: one database to operate, back up and reason about. Pinecone makes more sense when vector search is the dominant workload, needs to scale on its own, and you would rather not run the storage and indexing yourself. Most of the pgvector vs Pinecone decision comes down to a single question: do you want one system or two.
What each one is
pgvector is an open-source extension for PostgreSQL. It adds a vector column type and similarity search, with approximate indexes such as HNSW and IVFFlat so a query does not have to scan every row. Because it lives inside Postgres, your embeddings sit in the same database as the rest of your data, and you search them with SQL.
Pinecone is a managed, hosted vector database. You send it vectors and metadata through its API, and it handles the indexing, storage and scaling for you. There is no server for you to run or tune; you work through its client libraries and pay for the service.
Both do the same core job well: store embeddings and return the nearest matches for a query vector. The difference is everything around that job.
One system or two
This is the trade-off that decides most cases.
With pgvector, your relational data and your vectors are one system. A document, its owner, its permissions and its embeddings are all rows in the same database, so you can filter and join them in a single query, wrap a write to all of them in one transaction, and back them up together. There is one thing to run, one connection pool, one set of credentials, one restore to test.
A similarity search becomes an ordinary SQL query that can join and filter as it ranks:
-- Nearest chunks for one tenant, joined to the document they came from
SELECT d.title, c.content
FROM chunks c
JOIN documents d ON d.id = c.document_id
WHERE c.tenant_id = $1
AND d.status = 'published'
ORDER BY c.embedding <=> $2
LIMIT 5;
With Pinecone you have two systems. Your relational data lives in your main database and your vectors live in the service, so a join like the one above happens in your application code or through metadata stored alongside each vector. In return, the operational work of running and scaling vector search is handled for you, and that search scales on its own rather than competing with the rest of your database for memory and CPU.
Neither is the right answer in the abstract. If your vectors are one feature inside a product that is mostly relational, keeping them in Postgres removes a moving part. If vector search is the product, a service built only for it earns its place.
Keeping each customer's data isolated
In a multi-tenant product, the rule is usually that one customer's content must never appear in another's results. Where you enforce that rule depends on the store.
With pgvector, a tenant's documents and their embeddings are rows in the same database, so they fall under the same isolation as everything else you already control there, enforced with the same tools you use for any other table. On Evoriqa, my own multi-tenant support platform, tenants, billing and vector search all live in PostgreSQL with pgvector, so a tenant's documents and their embeddings fall under the same isolation, and the trade-off is one database to scale instead of a dedicated vector store.
Pinecone isolates tenants with namespaces or metadata filters: you partition each customer's vectors and scope every query to one of them. It works, but the isolation lives in a second system that you keep in step with your main database rather than in the database itself.
The isolation question applies whatever you pick. AI PDF Chat, a RAG assistant I built, keeps per-user vector isolation in its vector store, which is Qdrant. The store differs; the discipline of scoping every search to its owner does not.
Filtering and hybrid search
Pure vector search finds paraphrases well and exact terms badly. A product name, an error code or an acronym is often matched better by keyword search than by similarity, so most real systems want both.
This is where Postgres being one system pays off again. You already have full-text search in the same database, so you can run keyword search alongside vector search and combine the two rankings, all in SQL, all filtered by the same tenant and permission rules. On Evoriqa, each question combines semantic and keyword search, scoped to one workspace and one agent, so exact terms and product names are found as well as paraphrases.
Pinecone offers metadata filtering and its own support for sparse-dense hybrid search. The capability is there; the difference is that your keyword signal and your relational filters come from a separate place rather than from the database that already holds your text.
Scale and performance
I will not quote numbers here, because the honest answer is that it depends on your data, your index settings and your hardware. But the shape of the difference is clear.
pgvector runs inside your Postgres instance and shares its memory and CPU with every other query. Performance is a function of how you size that instance and how you tune the index: HNSW and IVFFlat each trade build time, memory and recall against query speed, and you own those choices. For many products this is more than enough, and it keeps vector search close to the data.
Pinecone scales vector search independently of your relational database. When the vector workload is very large or the query rate is high, moving it off your main database so it cannot contend with ordinary traffic is a real advantage. You are paying for someone else to solve the scaling problem, which is worth it precisely when that problem is hard.
Cost shape
The two have different cost shapes, which matters more than any single figure.
pgvector adds no separate bill. It is part of a database you already run, so the cost is the storage, memory and compute of that Postgres instance, which grows as your vectors and indexes do. There is nothing new to buy if your database has headroom.
Pinecone is a separate, usage-based service billed on top of whatever you already pay for your database. That can be the cheaper path when it saves you from over-provisioning Postgres for a demanding vector workload, and the more expensive one when your needs are modest and a database you already run would have absorbed them.
Migration effort later
Picking one now does not lock you in forever, but moving later is real work, so it is worth knowing the shape of it.
Your embeddings are portable as long as you keep the same embedding model, so the vectors themselves can be exported and re-loaded. What changes is everything around them: the query interface (SQL against pgvector versus Pinecone's API), how you express filters, and how tenant isolation is enforced. Moving from Pinecone to pgvector gains you joins and transactions with your relational data; moving the other way gives those up and shifts the ranking and filtering logic into your application. Plan for rewriting the retrieval layer, not just copying vectors.
pgvector vs Pinecone at a glance
| Concern | pgvector | Pinecone |
|---|---|---|
| Systems to run | One (your Postgres) | Two (database plus service) |
| Data and vectors | Together, in one database | Vectors separate from relational data |
| Joins and transactions | Native SQL, across both | In application code |
| Operations and scaling | You run and tune it | Managed; scales on its own |
| Tenant isolation | Same rules as the rest of the database | Namespaces or metadata filters |
| Keyword and hybrid search | Postgres full-text search in the same query | Metadata filters and hybrid support in the service |
| Cost shape | Part of an existing database | Separate, usage-based bill |
| Best fit | Vectors are one feature of a relational product | Vector search is the dominant workload |
Choosing between them
A short decision guide:
- You already run PostgreSQL and vectors are one feature among many relational ones. Use pgvector. One system is less to operate, and you get joins, transactions and full-text search for free.
- You need vector search to scale independently, at a size or query rate that would strain your database, and you would rather not run it yourself. Use Pinecone.
- Your isolation, filtering and ranking are tightly bound to relational data. Lean towards pgvector, so that logic stays in one place.
- Vector search is effectively the product, not a feature of something larger. A dedicated service is easier to justify.
My own default, for products that are mostly a normal application with a retrieval feature bolted on, is to start in Postgres with pgvector and only reach for a dedicated service when there is a measured reason to. Fewer moving parts is a feature. For how retrieval fits the wider system around whichever store you pick, see RAG architecture.
When to bring in help
If you are weighing this choice for a product with real users, the store is only one of several decisions that have to agree with each other: how you chunk and embed, how you filter, how you keep tenants apart, and how you check that answers are actually grounded. I build retrieval features end to end, and I am happy to talk it through if a second opinion would help. That is what my RAG development work covers.


