Tools/RAG & retrieval
Chroma: simple vector search with one catch
Chroma is an Apache-2.0 vector database that runs embedded, single-node or as Chroma Cloud. Where it is pleasant, and where the good search features stop at the cloud boundary.
- Type
- Vector database
- Pricing
- Apache-2.0 · Cloud paid
Balázs Csorba··10 min read
- Vector search
- Embeddings
- RAG
- Hybrid search
- Metadata filter

Key takeaways
- Chroma is Apache-2.0 with about 27,000 GitHub stars and a single API across Python, TypeScript and Rust clients.
- The embedded client starts inside the application process, so a first collection needs no server and no container.
- Hybrid search with reciprocal rank fusion exists only in Chroma Cloud; the docs list single-node support as future work.
- Cloud pricing is usage-based: $2.50 per GiB written, $0.33 per GiB stored per month, $0.0075 per TiB queried and $0.09 per GiB returned.
- The honest engineering verdict: excellent default for a first retrieval layer, weak as the long-term home of a hybrid search stack.
Chroma is an open-source vector database, Apache-2.0 licensed, that runs as an embedded library inside an application, as a single `chroma run` process, or as Chroma Cloud. The position this review takes is straightforward: it is the best default for the first retrieval layer of a RAG system, because the friction between an empty repository and a working query is close to zero, and it is a poor long-term home for a serious hybrid search stack, because the interesting parts of the query language are documented as cloud-only.
In the stack it occupies one slot: storage and retrieval for embeddings. It competes with pgvector when the application already runs Postgres, with Qdrant or Milvus when retrieval needs its own service, and with Pinecone when nobody wants to operate anything. Chroma's distinctive bet is not the index but the ergonomics: an embedding function attached to a collection, a filter language that needs no index declarations, and clients for Python, TypeScript and Rust that all speak the same verbs.
What it is
The data model has three levels. A tenant holds databases, a database holds collections, and a collection is the unit of storage and querying: each item has an id, an embedding, and optionally a document and a metadata object. Access control, quota and billing are scoped at the tenant level, which matters for multi-tenant SaaS and is quietly absent from most single-node vector stores.
- Licence: Apache 2.0, with roughly 27,000 stars on the GitHub repository and a tagged 1.5.9 release from May 2026.
- Three deployment modes: an embedded library, a single node, and a distributed deployment that Chroma Cloud operates.
- Three clients — Python, TypeScript and Rust — with the same collection, add, query and get shape, plus an async HTTP client.
- Four search modes in one collection: dense vectors, sparse vectors, full-text and regex over documents, and metadata filtering.
- Embedding functions as an interface, so a collection embeds on write and `query_texts` needs no model call in application code.
- Two query APIs: the classic `query` and `get`, and a newer Search API with ranking expressions, available in Chroma Cloud.
How it works
Chroma delegates durability to subsystems it does not have to reinvent: SQLite locally, and cloud object storage in the distributed build, where hot data sits in SSD caches. The vendor's argument is economic rather than exotic — vectors are large and memory is expensive, so keeping the source of truth in object storage at a fraction of the cost of RAM is what makes the cloud price list possible.
What the vendor does not publish is a public benchmark with methodology behind it. The Cloud documentation claims that production systems exceed 90 per cent recall and that the storage design makes Chroma Cloud an order of magnitude cheaper than alternatives; both are plausible and neither is reproducible from the docs. Treat them as a starting point for your own evaluation, not as a result.
Getting started
The embedded client starts a server inside the process and loses everything when the program exits, which is exactly what you want for a test and exactly what you do not want for production. A minimal query needs a client, a collection, an upsert and a query — no server, no container, no index tuning.
import chromadb
client = chromadb.Client()
collection = client.get_or_create_collection("docs")
collection.upsert(
ids=["d1", "d2"],
documents=[
"Refunds are issued within five business days.",
"Support answers within one working day.",
],
metadatas=[{"topic": "billing"}, {"topic": "support"}],
)
hits = collection.query(
query_texts=["how fast is a refund?"],
where={"topic": "billing"},
n_results=1,
)
# Results are column-major: one list per query, one entry per result.
print(hits["documents"][0][0])Three details in that snippet cause most of the friction later. Results come back column-major, so every consumer has to zip parallel arrays rather than iterate records. `n_results` defaults to 10, which silently returns nothing useful on a small test collection. And the filter goes in `where` against metadata, while text matching against the stored document goes in `where_document` with `$contains` or a regex.
Hybrid search and the cloud line
The Search API replaces `query` and `get` with a composable expression: `Search` builds the filter and the limit, `Knn` supplies a ranking, and `Rrf` fuses several rankings. Reciprocal rank fusion scores each candidate as the negative sum of weight divided by the smoothing constant plus its rank, with k defaulting to 60, which is why it works across dense and sparse results without normalising two different score scales.
from chromadb import Search, K, Knn, Rrf
dense = Knn(
query="how fast is a refund?",
key="#embedding",
return_rank=True,
limit=200,
)
sparse = Knn(
query="how fast is a refund?",
key="sparse_embedding",
return_rank=True,
limit=200,
)
search = (
Search()
.where(K("topic") == "billing")
.rank(Rrf(ranks=[dense, sparse], weights=[0.7, 0.3], k=60))
.limit(10)
.select(K.DOCUMENT, K.SCORE)
)
rows = collection.search(search).rows()[0]
for row in rows:
print(row["score"], row["document"][:60])Pricing and what it costs per query
Chroma Cloud bills four separate meters, and the interesting one is not the vector store. Writes are $2.50 per GiB, storage is $0.33 per GiB per month, queries are $0.0075 per TiB, and network egress is $0.09 per GiB returned. A Starter plan costs nothing per month and includes $5 in credits, ten databases and ten team members; Team is $250 per month with $100 in credits, a hundred databases, thirty members and SOC II.
| Cloud planPriceUsage meterIncluded | Starter$0 per monthwrite, storage, query, network$5 credits, 10 databases, 10 members | Team$250 per monthsame four meters$100 credits, 100 databases, 30 members, SOC II | EnterpriseCustomsame four metersunlimited databases, single tenant, BYOC, SLAs |
|---|---|---|---|
| Write | per GiB written | $2.50 | billed once per ingest |
| Storage | per GiB per month | $0.33 | vectors plus documents plus metadata |
| Query | per TiB queried | $0.0075 | effectively free at most corpus sizes |
| Network | per GiB returned | $0.09 | the meter that punishes large results |
| Self-hosted | $0 | your own disk | Apache-2.0, single node, no Search API |
The vendor's own calculator is instructive. At 1536 dimensions with 8 KiB documents across 500 collections, one million documents cost about $34 to write, six million stored documents about $27 a month, and ten million queries about $19 — roughly $79 a month for a corpus of six million chunks. Note the direction of the economics: 1 GiB of text becomes about 15 GiB of vectors, so storage is billed on the number that inflates fastest, and returning whole documents rather than short chunks is what pushes the egress meter up.
Self-hosting: what you actually run
Self-hosting is not one thing. The architecture documentation separates three modes with different ceilings, and the difference between them is larger than the branding suggests.
- Embedded: `chromadb.Client()` runs in-process, keeps data in memory or on a local path, and dies with the program. Fine for tests, not for a service.
- Single node: `chroma run --path` behind an `HttpClient`, documented at fewer than 10 million records across a handful of collections. Durable, simple, single point of failure.
- Distributed: the deployment Chroma Cloud operates, with object storage persistence and SSD caches, pinned to a single region per database.
- Data residency: Chroma Cloud runs in AWS us-east-1 and, since April 2026, GCP europe-west1. The region is fixed at creation and moving means creating a new database and reindexing.
- Compliance: Chroma Cloud is SOC 2 Type II certified; customer-managed encryption keys arrived in December 2025 and private networking in January 2026.
Where it falls short
The weaknesses come first, because they decide the choice. Hybrid ranking is cloud-only. Metadata filtering is convenient rather than tunable, and a filter that slows down gives you no index to adjust. The result shape is column-major, which adds a small but permanent tax on every consumer. And the price list punishes exactly the pattern that retrieval-augmented generation encourages: returning a large context window per query.
| DeploymentHybrid searchLock-in profile | Chromaembedded, single node, cloudRRF in Chroma Cloud onlyApache-2.0 server; the Search API is not | pgvectorextension inside existing Postgresmanual: vector index plus SQL full textPostgres licence; nothing to leave | Qdrantself-hosted or Cloudnative dense plus sparseApache-2.0; filter-heavy design to learn | Pineconemanaged onlynative in the hosted APIno self-hosted path |
|---|---|---|---|---|
| Licence | Apache 2.0 | PostgreSQL licence | Apache 2.0 | proprietary |
| Operational cost | lowest | none, same instance | a service to run | none |
| Dimensional limits | none documented | 2,000 on vector, 4,000 on halfvec | none documented | none documented |
| Best fit | first retrieval layer | vectors beside the row | filter-heavy retrieval | no operations at all |
The comparison that matters is with pgvector. If the corpus already lives in a Postgres table and the vectors belong in the same transaction, a dedicated vector store is an extra service to back up, secure and monitor for a capability Postgres can already provide, with the hard ceiling of 2,000 dimensions on the standard vector type and 4,000 with half-precision. Chroma earns its place when retrieval is the product: multi-modal collections, regex over stored documents, collection forking for experiments, and an embedding function that removes a model call from application code.
Verdict
Chroma is a well-judged default with a specific blind spot. The ergonomics are real: three clients, one API, a filter language that needs no schema, and a path from `import chromadb` to a ranked result in about ten lines. The blind spot is equally real: the query features that separate a demo from a retrieval system are the ones you cannot self-host today. That is a reasonable trade for a first layer and a poor one for a system expected to last.
- Pick it when the retrieval layer is still being designed and the fastest path to a measurable baseline matters most.
- Pick it when embeddings, documents and metadata must live in one collection and the team is multi-tenant from the start.
- Pick it for prototyping that will later be replaced; the collection API is small enough that a rewrite is a day, not a quarter.
- Skip it when hybrid search has to run inside your own network boundary today — the Search API is not available there.
- Skip it when the corpus is already in Postgres and pgvector's dimension ceiling does not bite.
- Reconsider it once the corpus passes a few million chunks and single-node ceilings, or once egress per GiB returned dominates the bill.
Sources
Frequently asked questions
Is Chroma free to self-host?
Yes. The server is Apache-2.0 and runs as a local library, a single `chroma run` process or a Docker container. Chroma Cloud is the paid, managed layer on top, and the two share an API.
Does Chroma support hybrid search?
Only in Chroma Cloud. The Search API exposes reciprocal rank fusion over dense and sparse embeddings, and its documentation states plainly that single-node support is planned for a future release.
How large a corpus can one Chroma instance hold?
The architecture documentation puts single-node Chroma at fewer than 10 million records across a handful of collections. Larger workloads are meant for the distributed deployment behind Chroma Cloud.
Chroma or pgvector?
pgvector wins when the vectors already live in a Postgres row and transactional consistency matters. Chroma wins when retrieval is the product: richer search modes, multi-modal collections and a client that embeds for you.