Tools/RAG & retrieval

Pinecone: a review of the managed vector database

What Pinecone really costs: read units scale with namespace size, a schema cannot be changed after creation, and where a self-hosted vector database is the better buy.

Type
Vector database
Pricing
Free tier · from $2 per month

··11 min read

  • Vector search
  • Embeddings
  • RAG
  • Namespaces
  • Hybrid search
Diagram: documents, dense vectors and sparse vectors upserted into a serverless index partitioned by namespace, a query naming one namespace and returning scored hits.

Key takeaways

  • Pinecone bills reads, not vectors. A query costs one read unit per gigabyte of the namespace it scans, with a floor of 0.25 read units, so partitioning is the single decision that determines the bill.
  • Read units cost four times write units on Standard, at 16 to 18 US dollars per million against 4 to 4.50. A read-heavy retrieval application is dominated by one line item.
  • One million queries against a 50 GB namespace is roughly 800 US dollars a month on Standard. The same million queries against 100 namespaces of 0.5 GB is about 8 dollars. Nothing about the application changes.
  • Schema migration is not supported on document indexes, and an index created before API version 2026-07 can never move to the Documents API. Adding a field means creating a new index and re-ingesting.
  • There is no uptime SLA below Enterprise, which starts at a 500 US dollar monthly minimum, and the Starter plan is restricted to a single region.

PINECONE is the managed vector database most retrieval prototypes start on, and for a defensible reason: an index is one POST away, there is nothing to run, and the bill is metered rather than provisioned. That combination is genuinely hard to beat for the first year of a retrieval feature. It is also why Pinecone so often stops being the right default once a product has real traffic, because its pricing model rewards one decision teams rarely make on purpose — how the data is partitioned.

It competes with two different kinds of thing. On one side the general-purpose databases that grew a vector column: pgvector inside PostgreSQL, MongoDB Atlas, Redis. On the other side the purpose-built engines, Qdrant and Weaviate, self-hosted or in the cloud. The pitch against all of them is operational: no cluster, no reindex, no ANN tuning, capacity that appears when it is needed. The counter-pitch is lock-in, the cost per query at scale, and a data model that is deliberately not SQL.

What it is

The index is the unit of everything. Records go into one, queries hit one, and inside it the data is partitioned into namespaces, with every upsert, query, fetch and list targeting exactly one namespace. Since API version 2026-07 there are two kinds of index, and telling them apart is the most important thing to understand before writing code against Pinecone.

  • A vector index is the classic one: dense and optional sparse vectors, created with dimension, metric and vector_type, read and written through the Vectors API.
  • A document index is the new one: a schema declaring dense_vector, sparse_vector and full-text string fields, read and written through the Documents API.
  • Namespaces are created implicitly on first upsert. One per tenant is the documented multitenancy pattern, and partitioning by namespace also cuts read cost, which is the point made below.
  • Metadata is a flat JSON object: no nested objects, no null values, keys cannot start with $, integers are stored as 64-bit floats, and 40 KB is the ceiling per record. Everything is indexed for filtering unless configured otherwise.
  • The filter language is a documented subset: $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $exists plus $and, $or and $not. The list operators accept at most 10,000 values.
  • Sparse indexes are narrow: at most 2,048 non-zero values per vector, 10 upserts per second, 100 queries per second, a top_k ceiling of 10,000, and dot-product as the only metric.

How it works

Three ranking signals can live in one index: BM25 over full-text fields, dense vectors for meaning, and sparse vectors for learned lexical importance. On a vector index a hybrid query combines dense and sparse in a single call. On a document index a search ranks by one scoring type chosen with score_by, so hybrid search there means either a text-match filter that narrows candidates before a dense search, or two searches fused with reciprocal rank fusion in the application.

Pinecone: documents and vectors into namespaces, queries outDocuments with full-text fields, dense vectors and sparse vectors are upserted into a serverless index that is partitioned into namespaces. A query names one namespace, costs one read unit per gigabyte of that namespace with a floor of 0.25 read units, can be narrowed with a metadata filter, and returns scored hits whose response bytes are billed as egress.Pinecone: write once, query a namespacedocs.pinecone.ioUPSERTDocumentsBM25 full-text fieldsDense vectorscosine or euclideanSparse vectorsdotproduct onlyONE SERVERLESS INDEX, PARTITIONED BY NAMESPACEIndexnamespace tenant-1 · tenant-2 · tenant-100 · the default namespaceQUERYNamespace1 read unit per GB, min 0.25Metadata filter$eq, $in, $and, $notScored hitsplus egress on the bytesOne data plane per index · fixed at creation · no schema migration
The billing model follows from the same structure: the query pays for the slice of the index it was pointed at, not for the work it did.

The cost model follows from the same structure. A query costs one read unit per gigabyte of the namespace it scans, with a floor of 0.25 read units. top_k, include_metadata and include_values do not change that number; only namespace size does. Egress is metered separately on the bytes returned, IDs and scores included. Around the index sit three more services worth naming, because they change what the application has to do: Pinecone Inference hosts embedding and reranking models and meters them by token or by rerank request, Pinecone Assistant turns an uploaded document set into a chat endpoint with its own token and ingestion metering, and Pinecone Nexus, listed as a knowledge engine for agents, pushes the same idea further by doing retrieval once when the data changes instead of on every agent call.

Where it breaks

Four things will bite, and three of them are structural rather than fixable.

  • Schema migration is not supported. Once a document index exists, fields cannot be added, removed or modified, and the documented remedy is to delete the index and create a new one.
  • You cannot convert. An index's data plane is fixed when it is created, so upgrading the SDK never moves an old vector index onto the Documents API. A team that wants full-text search has to build a second index and ingest again.
  • Cloud and region cannot be changed after creation, and the Starter plan is limited to aws us-east-1. Anything with a data residency requirement is a paid-plan decision on day one.
  • Read cost scales with namespace size rather than with effort. A 50 GB namespace costs 50 read units per query; the same data in 100 namespaces of 0.5 GB costs 0.5 each. That is a hundredfold difference in the largest line item.
  • There is no uptime SLA below Enterprise, which starts at a 500 dollar monthly minimum. Standard, the plan most production workloads land on, has no SLA and no audit logs.
  • Pinecone Local, the local emulator, runs the 2025-01 API version, caps an index at 100,000 records, ignores API keys and supports no namespaces, no backups and no Pinecone Inference. It is a CI convenience, not a faithful local environment.

Take the documented rates and do the arithmetic. On Standard, read units are $16 to $18 per million depending on cloud and region. One million queries against a 50 GB namespace is 50 million read units, roughly $800 a month in reads alone. The same million queries against 100 namespaces of 0.5 GB is 500,000 read units, about $8. Nothing about the application changed; only the partitioning did. On Enterprise, where read units start at $24 per million, the same two figures become $1,200 and $12.

Getting started

The smallest useful Pinecone program is a vector index with your own embeddings, one namespace per tenant, and a metadata filter on the query. That is the shape most production code converges on, and it is the shape the cost model rewards.

import os
from pinecone import Pinecone, ServerlessSpec

pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])

if not pc.has_index("products"):
    pc.create_index(
        name="products",
        vector_type="dense",
        dimension=1536,
        metric="cosine",
        spec=ServerlessSpec(cloud="aws", region="eu-central-1"),
        deletion_protection="disabled",
    )

index = pc.Index("products")

# One namespace per tenant keeps each read inside a small slice of the index.
index.upsert(
    namespace="tenant-42",
    vectors=[("p-1001", embedding, {"category": "pumps", "price": 249.0})],
)

hits = index.query(
    namespace="tenant-42",
    vector=query_embedding,
    top_k=8,
    filter={"category": {"$eq": "pumps"}, "price": {"$lte": 500}},
    include_metadata=True,
    include_values=False,   # omit the vector: egress is billed on the response
)

Two details in that snippet are load-bearing. include_values=False is the default and keeps several kilobytes per hit out of the egress bill, while a fetch always returns vector values, so use query when searching and only need identifiers or metadata. And filter does not reduce read units: the namespace does. Filters narrow what is scanned inside the namespace you already chose, which is one more reason not to put everything into a single namespace.

Pricing

Four plans, and each has a different shape. Starter is free with hard ceilings, Builder is a flat $20 with hard ceilings, and Standard and Enterprise are usage-based behind a monthly minimum that behaves as a commitment rather than a fee.

StarterStandardEnterprise
Monthly minimum$0$50$500
StorageUp to 2 GBUnlimited at $0.33 per GB per monthUnlimited at $0.33 per GB per month
Write unitsUp to 2M per monthUnlimited at $4 to $4.50 per millionUnlimited at $6 to $6.75 per million
Read unitsUp to 1M per monthUnlimited at $16 to $18 per millionUnlimited at $24 to $27 per million
Egress1 GB per month100 GB included, then $0.10 per GB100 GB included, then $0.10 per GB
GovernanceCommunity support, one projectSSO, RBAC, backups and restoreAudit logs, private endpoints, CMK, SCIM, 99.95% SLA

Three notes. The $50 and $500 minimums are billed as a top-up line when usage is below them, so a quiet month still costs the minimum. Read units cost four times what write units cost on Standard, which means a read-heavy retrieval application is dominated by a single line. And the flat-fee plans do not degrade gracefully — past the allowance, in-scope reads are blocked with a 429 rather than billed, which is a very different failure mode from a surprise invoice, and arguably the better of the two.

Standard and Enterprise also unlock what a production deployment eventually needs: import from object storage at $0.25 per GB, backups at $0.10 per GB per month, restores at $0.15 per GB, and dedicated read nodes, which are not metered in read units at all and are priced by provisioned capacity instead. Dedicated read nodes are the feature to understand before concluding that Pinecone is too expensive, because they break the read-unit formula that drives the rest of this article.

Alternatives

The comparison worth making is not against every vector database. It is against the two that break the Pinecone model in a specific, structural way.

PineconeQdrantpgvector
Licence and hostingProprietary; managed serverless on AWS, Azure or GCPApache-2.0; self-hosted or Qdrant CloudPostgreSQL licence; a column in your own database
Hybrid searchOne index holds dense, sparse and BM25, but a document index ranks by one signal per requestOne query mixes dense, sparse and BM25 scoresYou combine ts_vector and a vector index yourself
Cost of a query1 read unit per GB of namespace, 0.25 minimumThe machines you run, or Cloud nodesPart of the Postgres bill you already pay
Changing the shapeSchema cannot change; recreate the index and re-ingestPayload fields are added without a rebuildALTER TABLE, then build the ANN index
SQL and joinsNone; no joins, no transactionsNone; filtering through payloadFull SQL over vectors and rows together

The honest summary of that table: pgvector wins on everything except scale, and the threshold sits somewhere around ten million vectors, past which index build times and recall tuning in PostgreSQL stop being a reasonable afternoon. Qdrant wins on control and on hybrid search in a single query, and loses on the fact that somebody has to run it. Pinecone wins on time to first query and on not needing an on-call rotation for a search service, and loses on cost per query once the index grows and on the fact that the shape of the data cannot be changed after the fact. Those are the trade-offs; none of them is a bug.

Verdict

Pinecone is the right default for the first ninety days of a retrieval feature, and a defensible choice for years if the workload is small, well partitioned and heavy enough to sit on the read-unit floor. It stops being the right answer when the index passes a few gigabytes, when a compliance requirement needs an SLA or audit logs, or when the shape of the corpus is still moving.

  1. Adopt it when nobody is going to own a search cluster. The free tier is genuinely useful and the first paid step is $20 flat.
  2. Partition by namespace before the index is large, not after. That single decision sets the read bill and is much harder to change later.
  3. Do not adopt it for a corpus whose fields are still in flux on a document index, because schema migration is not supported.
  4. Run the numbers with your own namespace size. The read-unit formula in the documentation is short enough to work out by hand.
  5. Look at Qdrant when the index will pass ten million vectors, and at pgvector when it will stay well under a million and a database administrator already exists.

Sources

  1. Pinecone docs: Index data overview, indexes, namespaces and metadata
  2. Pinecone docs: Adopt the Documents API, API version 2026-07
  3. Pinecone docs: Create an index, schemas, regions and metrics
  4. Pinecone docs: Understanding Pinecone cost, read units, write units and egress
  5. Pinecone pricing: plans, limits and rates
  6. Pinecone docs: Semantic search
  7. Pinecone docs: Local development with Pinecone Local
  8. Pinecone docs: 2026 release notes

Frequently asked questions

What is Pinecone and what is it used for?

Pinecone is a fully managed vector database. You store embeddings in an index and retrieve them by similarity, which is the retrieval step in a RAG pipeline, in semantic product search and in agent memory. It runs on managed serverless infrastructure on AWS, Azure or GCP, so there is no cluster to operate.

How much does Pinecone cost per month?

The Starter plan is free and includes up to 2 GB of storage, 2 million write units, 1 million read units and 1 GB of egress per month. Builder costs a flat 20 dollars. Standard has a 50 dollar monthly minimum and prices storage at 0.33 dollars per GB per month, write units at 4 to 4.50 per million, read units at 16 to 18 per million and egress at 0.10 per GB after 100 GB included. Enterprise starts at 500 dollars per month with read units at 24 to 27 per million.

What is the difference between a Pinecone vector index and a document index?

A vector index is the classic one, created with dimension, metric and vector_type, holding dense and sparse vectors through the Vectors API, with hybrid search in a single query. A document index is created with a schema through the Documents API and can hold dense vector, sparse vector and full-text fields in one index, but a single search ranks by one scoring type, so hybrid search means either a text-match filter followed by a dense search or two searches fused in the application.

Are Pinecone namespaces just for multi-tenancy?

They are documented for multitenancy and query speed, but the cost model makes them a cost control too, which is the least obvious thing about Pinecone to discover. A query costs one read unit per gigabyte of the namespace it scans, so a per-tenant namespace keeps a large tenant from setting the read price for everyone else.

Can I change the schema of a Pinecone index later?

No. Pinecone's documentation is explicit that schema migration is not yet supported: once a document index is created you cannot add, remove or modify fields, and the documented remedy is to delete the index and create a new one. The same applies to cloud and region, which cannot be changed after a serverless index exists.

Is there a free tier, and can I develop locally?

The Starter plan is free with hard ceilings and one project. For local development there is Pinecone Local, an in-memory Docker emulator, but it runs the 2025-01 API version rather than the current one, caps an index at 100,000 records, ignores API keys and supports no namespace management, no backups and no Pinecone Inference, so treat it as a CI convenience rather than a faithful local environment.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.