Tools/RAG & retrieval
Pinecone: a review of the managed vector database
What Pinecone really costs: read units scale with namespace size, a schema cannot be changed after creation, and where a self-hosted vector database is the better buy.
- Type
- Vector database
- Pricing
- Free tier · from $2 per month
Balázs Csorba··11 min read
- Vector search
- Embeddings
- RAG
- Namespaces
- Hybrid search

Key takeaways
- Pinecone bills reads, not vectors. A query costs one read unit per gigabyte of the namespace it scans, with a floor of 0.25 read units, so partitioning is the single decision that determines the bill.
- Read units cost four times write units on Standard, at 16 to 18 US dollars per million against 4 to 4.50. A read-heavy retrieval application is dominated by one line item.
- One million queries against a 50 GB namespace is roughly 800 US dollars a month on Standard. The same million queries against 100 namespaces of 0.5 GB is about 8 dollars. Nothing about the application changes.
- Schema migration is not supported on document indexes, and an index created before API version 2026-07 can never move to the Documents API. Adding a field means creating a new index and re-ingesting.
- There is no uptime SLA below Enterprise, which starts at a 500 US dollar monthly minimum, and the Starter plan is restricted to a single region.
PINECONE is the managed vector database most retrieval prototypes start on, and for a defensible reason: an index is one POST away, there is nothing to run, and the bill is metered rather than provisioned. That combination is genuinely hard to beat for the first year of a retrieval feature. It is also why Pinecone so often stops being the right default once a product has real traffic, because its pricing model rewards one decision teams rarely make on purpose — how the data is partitioned.
It competes with two different kinds of thing. On one side the general-purpose databases that grew a vector column: pgvector inside PostgreSQL, MongoDB Atlas, Redis. On the other side the purpose-built engines, Qdrant and Weaviate, self-hosted or in the cloud. The pitch against all of them is operational: no cluster, no reindex, no ANN tuning, capacity that appears when it is needed. The counter-pitch is lock-in, the cost per query at scale, and a data model that is deliberately not SQL.
What it is
The index is the unit of everything. Records go into one, queries hit one, and inside it the data is partitioned into namespaces, with every upsert, query, fetch and list targeting exactly one namespace. Since API version 2026-07 there are two kinds of index, and telling them apart is the most important thing to understand before writing code against Pinecone.
- A vector index is the classic one: dense and optional sparse vectors, created with
dimension,metricandvector_type, read and written through the Vectors API. - A document index is the new one: a schema declaring
dense_vector,sparse_vectorand full-textstringfields, read and written through the Documents API. - Namespaces are created implicitly on first upsert. One per tenant is the documented multitenancy pattern, and partitioning by namespace also cuts read cost, which is the point made below.
- Metadata is a flat JSON object: no nested objects, no null values, keys cannot start with
$, integers are stored as 64-bit floats, and 40 KB is the ceiling per record. Everything is indexed for filtering unless configured otherwise. - The filter language is a documented subset:
$eq,$ne,$gt,$gte,$lt,$lte,$in,$nin,$existsplus$and,$orand$not. The list operators accept at most 10,000 values. - Sparse indexes are narrow: at most 2,048 non-zero values per vector, 10 upserts per second, 100 queries per second, a
top_kceiling of 10,000, and dot-product as the only metric.
How it works
Three ranking signals can live in one index: BM25 over full-text fields, dense vectors for meaning, and sparse vectors for learned lexical importance. On a vector index a hybrid query combines dense and sparse in a single call. On a document index a search ranks by one scoring type chosen with score_by, so hybrid search there means either a text-match filter that narrows candidates before a dense search, or two searches fused with reciprocal rank fusion in the application.
The cost model follows from the same structure. A query costs one read unit per gigabyte of the namespace it scans, with a floor of 0.25 read units. top_k, include_metadata and include_values do not change that number; only namespace size does. Egress is metered separately on the bytes returned, IDs and scores included. Around the index sit three more services worth naming, because they change what the application has to do: Pinecone Inference hosts embedding and reranking models and meters them by token or by rerank request, Pinecone Assistant turns an uploaded document set into a chat endpoint with its own token and ingestion metering, and Pinecone Nexus, listed as a knowledge engine for agents, pushes the same idea further by doing retrieval once when the data changes instead of on every agent call.
Where it breaks
Four things will bite, and three of them are structural rather than fixable.
- Schema migration is not supported. Once a document index exists, fields cannot be added, removed or modified, and the documented remedy is to delete the index and create a new one.
- You cannot convert. An index's data plane is fixed when it is created, so upgrading the SDK never moves an old vector index onto the Documents API. A team that wants full-text search has to build a second index and ingest again.
- Cloud and region cannot be changed after creation, and the Starter plan is limited to
awsus-east-1. Anything with a data residency requirement is a paid-plan decision on day one. - Read cost scales with namespace size rather than with effort. A 50 GB namespace costs 50 read units per query; the same data in 100 namespaces of 0.5 GB costs 0.5 each. That is a hundredfold difference in the largest line item.
- There is no uptime SLA below Enterprise, which starts at a 500 dollar monthly minimum. Standard, the plan most production workloads land on, has no SLA and no audit logs.
- Pinecone Local, the local emulator, runs the 2025-01 API version, caps an index at 100,000 records, ignores API keys and supports no namespaces, no backups and no Pinecone Inference. It is a CI convenience, not a faithful local environment.
Take the documented rates and do the arithmetic. On Standard, read units are $16 to $18 per million depending on cloud and region. One million queries against a 50 GB namespace is 50 million read units, roughly $800 a month in reads alone. The same million queries against 100 namespaces of 0.5 GB is 500,000 read units, about $8. Nothing about the application changed; only the partitioning did. On Enterprise, where read units start at $24 per million, the same two figures become $1,200 and $12.
Getting started
The smallest useful Pinecone program is a vector index with your own embeddings, one namespace per tenant, and a metadata filter on the query. That is the shape most production code converges on, and it is the shape the cost model rewards.
import os
from pinecone import Pinecone, ServerlessSpec
pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])
if not pc.has_index("products"):
pc.create_index(
name="products",
vector_type="dense",
dimension=1536,
metric="cosine",
spec=ServerlessSpec(cloud="aws", region="eu-central-1"),
deletion_protection="disabled",
)
index = pc.Index("products")
# One namespace per tenant keeps each read inside a small slice of the index.
index.upsert(
namespace="tenant-42",
vectors=[("p-1001", embedding, {"category": "pumps", "price": 249.0})],
)
hits = index.query(
namespace="tenant-42",
vector=query_embedding,
top_k=8,
filter={"category": {"$eq": "pumps"}, "price": {"$lte": 500}},
include_metadata=True,
include_values=False, # omit the vector: egress is billed on the response
)Two details in that snippet are load-bearing. include_values=False is the default and keeps several kilobytes per hit out of the egress bill, while a fetch always returns vector values, so use query when searching and only need identifiers or metadata. And filter does not reduce read units: the namespace does. Filters narrow what is scanned inside the namespace you already chose, which is one more reason not to put everything into a single namespace.
Pricing
Four plans, and each has a different shape. Starter is free with hard ceilings, Builder is a flat $20 with hard ceilings, and Standard and Enterprise are usage-based behind a monthly minimum that behaves as a commitment rather than a fee.
| Starter | Standard | Enterprise | |
|---|---|---|---|
| Monthly minimum | $0 | $50 | $500 |
| Storage | Up to 2 GB | Unlimited at $0.33 per GB per month | Unlimited at $0.33 per GB per month |
| Write units | Up to 2M per month | Unlimited at $4 to $4.50 per million | Unlimited at $6 to $6.75 per million |
| Read units | Up to 1M per month | Unlimited at $16 to $18 per million | Unlimited at $24 to $27 per million |
| Egress | 1 GB per month | 100 GB included, then $0.10 per GB | 100 GB included, then $0.10 per GB |
| Governance | Community support, one project | SSO, RBAC, backups and restore | Audit logs, private endpoints, CMK, SCIM, 99.95% SLA |
Three notes. The $50 and $500 minimums are billed as a top-up line when usage is below them, so a quiet month still costs the minimum. Read units cost four times what write units cost on Standard, which means a read-heavy retrieval application is dominated by a single line. And the flat-fee plans do not degrade gracefully — past the allowance, in-scope reads are blocked with a 429 rather than billed, which is a very different failure mode from a surprise invoice, and arguably the better of the two.
Standard and Enterprise also unlock what a production deployment eventually needs: import from object storage at $0.25 per GB, backups at $0.10 per GB per month, restores at $0.15 per GB, and dedicated read nodes, which are not metered in read units at all and are priced by provisioned capacity instead. Dedicated read nodes are the feature to understand before concluding that Pinecone is too expensive, because they break the read-unit formula that drives the rest of this article.
Alternatives
The comparison worth making is not against every vector database. It is against the two that break the Pinecone model in a specific, structural way.
| Pinecone | Qdrant | pgvector | |
|---|---|---|---|
| Licence and hosting | Proprietary; managed serverless on AWS, Azure or GCP | Apache-2.0; self-hosted or Qdrant Cloud | PostgreSQL licence; a column in your own database |
| Hybrid search | One index holds dense, sparse and BM25, but a document index ranks by one signal per request | One query mixes dense, sparse and BM25 scores | You combine ts_vector and a vector index yourself |
| Cost of a query | 1 read unit per GB of namespace, 0.25 minimum | The machines you run, or Cloud nodes | Part of the Postgres bill you already pay |
| Changing the shape | Schema cannot change; recreate the index and re-ingest | Payload fields are added without a rebuild | ALTER TABLE, then build the ANN index |
| SQL and joins | None; no joins, no transactions | None; filtering through payload | Full SQL over vectors and rows together |
The honest summary of that table: pgvector wins on everything except scale, and the threshold sits somewhere around ten million vectors, past which index build times and recall tuning in PostgreSQL stop being a reasonable afternoon. Qdrant wins on control and on hybrid search in a single query, and loses on the fact that somebody has to run it. Pinecone wins on time to first query and on not needing an on-call rotation for a search service, and loses on cost per query once the index grows and on the fact that the shape of the data cannot be changed after the fact. Those are the trade-offs; none of them is a bug.
Verdict
Pinecone is the right default for the first ninety days of a retrieval feature, and a defensible choice for years if the workload is small, well partitioned and heavy enough to sit on the read-unit floor. It stops being the right answer when the index passes a few gigabytes, when a compliance requirement needs an SLA or audit logs, or when the shape of the corpus is still moving.
- Adopt it when nobody is going to own a search cluster. The free tier is genuinely useful and the first paid step is $20 flat.
- Partition by namespace before the index is large, not after. That single decision sets the read bill and is much harder to change later.
- Do not adopt it for a corpus whose fields are still in flux on a document index, because schema migration is not supported.
- Run the numbers with your own namespace size. The read-unit formula in the documentation is short enough to work out by hand.
- Look at Qdrant when the index will pass ten million vectors, and at pgvector when it will stay well under a million and a database administrator already exists.
Sources
- Pinecone docs: Index data overview, indexes, namespaces and metadata
- Pinecone docs: Adopt the Documents API, API version 2026-07
- Pinecone docs: Create an index, schemas, regions and metrics
- Pinecone docs: Understanding Pinecone cost, read units, write units and egress
- Pinecone pricing: plans, limits and rates
- Pinecone docs: Semantic search
- Pinecone docs: Local development with Pinecone Local
- Pinecone docs: 2026 release notes
Frequently asked questions
What is Pinecone and what is it used for?
Pinecone is a fully managed vector database. You store embeddings in an index and retrieve them by similarity, which is the retrieval step in a RAG pipeline, in semantic product search and in agent memory. It runs on managed serverless infrastructure on AWS, Azure or GCP, so there is no cluster to operate.
How much does Pinecone cost per month?
The Starter plan is free and includes up to 2 GB of storage, 2 million write units, 1 million read units and 1 GB of egress per month. Builder costs a flat 20 dollars. Standard has a 50 dollar monthly minimum and prices storage at 0.33 dollars per GB per month, write units at 4 to 4.50 per million, read units at 16 to 18 per million and egress at 0.10 per GB after 100 GB included. Enterprise starts at 500 dollars per month with read units at 24 to 27 per million.
What is the difference between a Pinecone vector index and a document index?
A vector index is the classic one, created with dimension, metric and vector_type, holding dense and sparse vectors through the Vectors API, with hybrid search in a single query. A document index is created with a schema through the Documents API and can hold dense vector, sparse vector and full-text fields in one index, but a single search ranks by one scoring type, so hybrid search means either a text-match filter followed by a dense search or two searches fused in the application.
Are Pinecone namespaces just for multi-tenancy?
They are documented for multitenancy and query speed, but the cost model makes them a cost control too, which is the least obvious thing about Pinecone to discover. A query costs one read unit per gigabyte of the namespace it scans, so a per-tenant namespace keeps a large tenant from setting the read price for everyone else.
Can I change the schema of a Pinecone index later?
No. Pinecone's documentation is explicit that schema migration is not yet supported: once a document index is created you cannot add, remove or modify fields, and the documented remedy is to delete the index and create a new one. The same applies to cloud and region, which cannot be changed after a serverless index exists.
Is there a free tier, and can I develop locally?
The Starter plan is free with hard ceilings and one project. For local development there is Pinecone Local, an in-memory Docker emulator, but it runs the 2025-01 API version rather than the current one, caps an index at 100,000 records, ignores API keys and supports no namespace management, no backups and no Pinecone Inference, so treat it as a CI convenience rather than a faithful local environment.