[{"data":1,"prerenderedAt":697},["ShallowReactive",2],{"tool-pgvector-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":21,"sources":27,"cover":42,"og":43,"expertise":44,"locales":45,"lang":46,"title":49,"description":50,"coverAlt":51,"url":23,"pricing":52,"kind":53,"metaTitle":54,"takeaways":55,"faq":61,"toc":74,"blocks":99,"others":492},"pgvector","2026-09-21",10,"rag",[9,10,11,12,13],"Vector search","Postgres","HNSW","RAG","Quantisation",[4,15,16,17,18,19,20],"pgvector vs qdrant","postgres vector search","hnsw index postgres","iterative index scans","binary quantization postgres","vector database postgres",[22,24],{"name":4,"url":23},"https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector",{"name":25,"url":26},"PostgreSQL","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPostgreSQL",[28,30,33,36,39],{"title":29,"url":23},"pgvector README",{"title":31,"url":32},"pgvector changelog","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FCHANGELOG.md",{"title":34,"url":35},"PostgreSQL news: pgvector 0.8.2 released","https:\u002F\u002Fwww.postgresql.org\u002Fabout\u002Fnews\u002Fpgvector-082-released-3245\u002F",{"title":37,"url":38},"AWS: Scale pgvector with binary quantization","https:\u002F\u002Faws.amazon.com\u002Fblogs\u002Fdatabase\u002Fscale-pgvector-with-binary-quantization-on-amazon-aurora-postgresql\u002F",{"title":40,"url":41},"pgvector licence","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FLICENSE","\u002Fimages\u002Fblog\u002Fpgvector\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fpgvector\u002Fog.jpg","ai-engineer",[46,47,48],"en","de","hu","pgvector, reviewed: the vector database you do not have to run","A review of pgvector 0.8.7: iterative scans for filtered search, HNSW and IVFFlat, binary quantisation at 100M vectors, and the CVE that made index builds a patch item.","A query enters at the top and splits into an exact sequential scan, an HNSW graph walk and an IVFFlat probe; a band below shows an iterative scan continuing until the limit is full.","PostgreSQL licence","Vector database extension","pgvector: vector search inside Postgres · Balázs Csorba",[56,57,58,59,60],"pgvector is a PostgreSQL extension under the PostgreSQL licence, so vectors live in ordinary tables and tenant filters, JOINs and cascading deletes stay transactional.","Iterative index scans, added in 0.8.0, are the fix for filtered search: without them a filter matching 10% of rows at the default ef_search of 40 returns about four rows.","AWS measured a 367 GB HNSW index for 100M vectors at 768 dimensions against a 38 GB binary-quantised index that built in 1.1 hours instead of 16.1.","CVE-2026-3172, fixed in 0.8.2 in February 2026, was a buffer overflow in parallel HNSW index builds that could leak data from other relations; the fix needs no reindex.","The ceiling is memory rather than the API: once the index stops fitting shared_buffers the choices are halfvec, binary quantisation, partitioning or a second system.",[62,65,68,71],{"q":63,"a":64},"Is pgvector free to use?","Yes. It is released under the PostgreSQL licence with no fee, no open-core split and no paid tier, and it is preinstalled by an increasing number of hosted Postgres providers. Check which version they ship: 0.8.0 or later is needed for iterative index scans.",{"q":66,"a":67},"When should I use HNSW instead of IVFFlat?","HNSW gives a better speed-recall trade-off, can be created on an empty table because it has no training step, and is slower to build and hungrier for memory. IVFFlat builds faster and needs data present first; the README suggests lists = rows \u002F 1000 up to one million rows, sqrt(rows) above that, and probes starting at sqrt(lists).",{"q":69,"a":70},"Does pgvector work with a WHERE clause?","Yes, but with an approximate index the filter runs after the index scan, so a selective condition can return fewer rows than the LIMIT asks for. The documented answers are a B-tree on the filter column, iterative index scans with hnsw.iterative_scan, partial indexes per value, and list partitioning for many distinct values.",{"q":72,"a":73},"How many vectors can pgvector handle?","A single table is limited by PostgreSQL's 32 TB relation limit and by how much of the index fits in memory; AWS measured 367 GB of index for 100 million 768-dimensional vectors and recommends partitioning at billion scale. Beyond that the README points at replicas, Citus, PgDog or list partitioning rather than a built-in sharding layer.",[75,78,81,84,87,90,93,96],{"id":76,"title":77},"what-it-is","What it is",{"id":79,"title":80},"how-it-works","How it works",{"id":82,"title":83},"filtered-search","Filtered search",{"id":85,"title":86},"getting-started","Getting started",{"id":88,"title":89},"performance","Performance",{"id":91,"title":92},"where-it-shingles","Where it falls short",{"id":94,"title":95},"verdict","Verdict",{"id":97,"title":98},"sources","Sources",[100,104,107,110,113,185,186,197,206,217,218,221,255,270,271,274,276,283,297,298,301,353,356,362,388,389,392,436,442,443,446,462,468,469],{"type":101,"content":102},"paragraph",[103],"pgvector is a PostgreSQL extension that puts vectors in an ordinary table and searches them with ordinary SQL. It is not a vector database with a query language bolted on: it adds four column types, six distance operators and two index types to a database most teams already operate, under the PostgreSQL licence, which is about as permissive as licensing gets. The position taken here is that this is the right default for almost every retrieval workload up to tens of millions of vectors, and that a team standing up a dedicated vector database at that size is buying operational burden rather than capability.",{"type":101,"content":105},[106],"It competes with Qdrant, Weaviate, Chroma and the fully managed vector services, and with every hosted Postgres that now ships the extension by default. What it removes is a whole class of infrastructure: no second service to run, no second wire protocol to authenticate, no second backup schedule, no consistency gap between the rows and the embeddings that describe them. What it keeps is every Postgres constraint, one node's memory, one node's vacuum, one node's write throughput, and those become the design inputs the moment the index stops fitting in RAM.",{"type":108,"level":109,"id":76,"text":77},"heading",2,{"type":101,"content":111},[112],"The first release, 0.1.0, was published on 20 April 2021, and the current release is 0.8.7, published on 1 October 2026; the repository showed 23,300 stars and 1,300 forks when checked. It supports PostgreSQL 13 and later, ships as Docker image, PGXN, APT, Yum, Homebrew and conda-forge package, and comes preinstalled on a growing list of hosted providers, which matters because a managed provider stuck on an old version is the most common way a team ends up without a security fix it believes it has.",{"type":114,"ordered":115,"items":116},"list",false,[117,124,129,148,157,183],[118,119,123],"Licence: ",{"tag":120,"children":121},"strong",[122],"the PostgreSQL licence",", the same permissive text PostgreSQL itself uses. No open core, no commercial tier, no feature gated behind a paid plan.",[125,126,128],"Version 0.8.7 of 1 October 2026; the 0.8 line has carried ",{"tag":120,"children":127},[18]," since 0.8.0 in October 2024, which is the feature that made filtered search predictable.",[130,131,135,136,139,140,143,144,147],"Four types: ",{"tag":132,"children":133},"code",[134],"vector"," at 4 bytes per dimension, ",{"tag":132,"children":137},[138],"halfvec"," at 2, ",{"tag":132,"children":141},[142],"bit"," at one bit per dimension, and ",{"tag":132,"children":145},[146],"sparsevec"," for sparse vectors with up to 1,000 non-zero elements.",[149,150,152,153,156],"Two index types: ",{"tag":120,"children":151},[11]," for the best speed-recall trade-off and no training step, ",{"tag":120,"children":154},[155],"IVFFlat"," for faster builds and a smaller memory footprint.",[158,159,162,163,166,167,170,171,174,175,178,179,182],"Six operators usable directly in ORDER BY: L2 (",{"tag":132,"children":160},[161],"\u003C->","), inner product (",{"tag":132,"children":164},[165],"\u003C#>","), cosine (",{"tag":132,"children":168},[169],"\u003C=>","), L1 (",{"tag":132,"children":172},[173],"\u003C+>","), Hamming (",{"tag":132,"children":176},[177],"\u003C~>",") and Jaccard (",{"tag":132,"children":180},[181],"\u003C%>",").",[184],"Storage and access come from Postgres itself: ACID, WAL replication, point-in-time recovery, JOINs and row-level security on the same table as the embeddings.",{"type":108,"level":109,"id":79,"text":80},{"type":101,"content":187},[188,189,192,193,196],"Without an index, a vector query is a sequential scan with an ORDER BY over the distance function: exact results, perfect recall, and a cost that tracks the table. An approximate index changes that bargain. HNSW builds a multi-layer graph over the vectors and walks it, trading a little recall for a query cost that no longer grows with the table; IVFFlat groups vectors into lists and probes a subset of them. The detail that decides everything is that the planner only reaches for an index when the query looks like ",{"tag":132,"children":190},[191],"ORDER BY embedding \u003C=> $1 LIMIT n",". Write the same query as ",{"tag":132,"children":194},[195],"ORDER BY 1 - (embedding \u003C=> $1) DESC"," and the README states plainly that no index will be used.",{"type":198,"attrs":199,"inner":203,"caption":204},"diagram",{"viewBox":200,"role":201,"aria-labelledby":202},"0 0 720 376","img","pgv-t pgv-d","\u003Ctitle id=\"pgv-t\">Three execution paths for one vector query\u003C\u002Ftitle>\u003Cdesc id=\"pgv-d\">A query carrying an ORDER BY over a distance operator and a LIMIT enters at the top and splits three ways. Left, no index: a sequential scan with perfect recall whose cost grows with every row. Middle, HNSW: a walk over a multilayer graph with m set to 16 and ef_search set to 40 by default, where any WHERE clause is applied only after the index scan has produced its candidates. Right, IVFFlat: vectors grouped into lists and only the closest lists probed, built after the data exists and tuned through ivfflat.probes, at a lower speed-recall trade-off. A band below: iterative index scans, added in 0.8.0, keep walking while the filter removes rows, in strict_order or relaxed_order, bounded by hnsw.max_scan_tuples with a default of 20,000.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"26\" class=\"d-title\">One query, three execution paths\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"26\" text-anchor=\"end\" class=\"d-label\">the ORDER BY decides\u003C\u002Ftext>\u003Crect x=\"260\" y=\"44\" width=\"200\" height=\"50\" rx=\"10\" class=\"d-accent\"\u002F>\u003Ctext x=\"360\" y=\"66\" text-anchor=\"middle\" class=\"d-text\">Query\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"84\" text-anchor=\"middle\" class=\"d-small\">ORDER BY distance, LIMIT 10\u003C\u002Ftext>\u003Cpath d=\"M360 94 V112\" class=\"d-line\"\u002F>\u003Cpath d=\"M126 112 H594\" class=\"d-line\"\u002F>\u003Cpath d=\"M126 112 V124\" class=\"d-line\"\u002F>\u003Cpath d=\"M360 112 V124\" class=\"d-line\"\u002F>\u003Cpath d=\"M594 112 V124\" class=\"d-line\"\u002F>\u003Crect x=\"20\" y=\"124\" width=\"213\" height=\"140\" rx=\"10\" class=\"d-box\"\u002F>\u003Ctext x=\"126\" y=\"154\" text-anchor=\"middle\" class=\"d-text\">Exact scan\u003C\u002Ftext>\u003Ctext x=\"126\" y=\"182\" text-anchor=\"middle\" class=\"d-small\">sequential scan\u003C\u002Ftext>\u003Ctext x=\"126\" y=\"206\" text-anchor=\"middle\" class=\"d-small\">perfect recall\u003C\u002Ftext>\u003Ctext x=\"126\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">cost grows per row\u003C\u002Ftext>\u003Ctext x=\"126\" y=\"254\" text-anchor=\"middle\" class=\"d-small\">no index needed\u003C\u002Ftext>\u003Crect x=\"253\" y=\"124\" width=\"214\" height=\"140\" rx=\"10\" class=\"d-sky\"\u002F>\u003Ctext x=\"360\" y=\"154\" text-anchor=\"middle\" class=\"d-text\">HNSW\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"182\" text-anchor=\"middle\" class=\"d-small\">multilayer graph walk\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"206\" text-anchor=\"middle\" class=\"d-small\">m = 16, ef_search = 40\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">filter runs after the scan\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"254\" text-anchor=\"middle\" class=\"d-small\">the default index\u003C\u002Ftext>\u003Crect x=\"486\" y=\"124\" width=\"213\" height=\"140\" rx=\"10\" class=\"d-gold\"\u002F>\u003Ctext x=\"592\" y=\"154\" text-anchor=\"middle\" class=\"d-text\">IVFFlat\u003C\u002Ftext>\u003Ctext x=\"592\" y=\"182\" text-anchor=\"middle\" class=\"d-small\">lists, then probes\u003C\u002Ftext>\u003Ctext x=\"592\" y=\"206\" text-anchor=\"middle\" class=\"d-small\">build it after the load\u003C\u002Ftext>\u003Ctext x=\"592\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">tune ivfflat.probes\u003C\u002Ftext>\u003Ctext x=\"592\" y=\"254\" text-anchor=\"middle\" class=\"d-small\">faster build, less recall\u003C\u002Ftext>\u003Crect x=\"20\" y=\"284\" width=\"679\" height=\"76\" rx=\"10\" class=\"d-box d-dash\"\u002F>\u003Ctext x=\"40\" y=\"312\" class=\"d-text\">Iterative scan, added in 0.8.0\u003C\u002Ftext>\u003Ctext x=\"40\" y=\"336\" class=\"d-small\">the scan continues when the filter drops rows: strict_order keeps exact distance order\u003C\u002Ftext>\u003Ctext x=\"40\" y=\"356\" class=\"d-small\">relaxed_order trades order for recall; bounded by hnsw.max_scan_tuples, default 20,000\u003C\u002Ftext>",[205],"The filter is applied after the approximate index has done its work, which is why a selective WHERE clause and an approximate index disagree until the scan is allowed to continue.",{"type":101,"content":207},[208,209,212,213,216],"The second detail is where the ",{"tag":132,"children":210},[211],"WHERE"," clause lands. An approximate index produces candidates and the filter is applied to them afterwards, so selectivity and recall interact: with the default ",{"tag":132,"children":214},[215],"hnsw.ef_search"," of 40 and a condition matching 10% of rows, about four rows come back. Nothing is broken, the index simply never saw the rows that were dropped. That behaviour is the most common source of reports claiming the extension returns fewer results, and it has a documented fix rather than a workaround.",{"type":108,"level":109,"id":82,"text":83},{"type":101,"content":219},[220],"Filtering is a first-class problem here rather than an afterthought, and the documentation works through four moves in the order a reviewer would try them. Which one applies depends on how selective the filter is, how much recall the application actually needs, and how many distinct values the filter takes: a tenant id with 50,000 values behaves nothing like a country code with eight.",{"type":114,"ordered":115,"items":222},[223,229,239,253],[224,225,228],"Index the filter column with a plain ",{"tag":120,"children":226},[227],"B-tree first",". When the condition matches a small share of rows this returns exact nearest neighbours without touching an approximate index, and the README names it as the starting point.",[230,231,234,235,238],"If the filter stays broad, turn on iterative index scans with ",{"tag":132,"children":232},[233],"SET hnsw.iterative_scan = relaxed_order",": the graph is walked until the LIMIT is full, and ",{"tag":132,"children":236},[237],"strict_order"," is there when the distance order has to be exact.",[240,241,244,245,248,249,252],"Bound the work. ",{"tag":132,"children":242},[243],"hnsw.max_scan_tuples"," defaults to 20,000 and ",{"tag":132,"children":246},[247],"hnsw.scan_mem_multiplier"," to one multiple of ",{"tag":132,"children":250},[251],"work_mem",", so a selective filter degrades into a bounded scan rather than an open-ended one.",[254],"For many distinct values, use partial indexes per value or list-partition the table. The README also notes that tenants sharing one approximate index affect each other's recall, which is a partitioning argument rather than a tuning argument.",{"type":256,"variant":257,"title":258,"body":259},"callout","note","Recall is the metric nobody measures",[260],[261,262,265,266,269],"The documentation's own method for checking this is to run the same query inside a transaction with ",{"tag":132,"children":263},[264],"enable_indexscan = off"," and compare it against the approximate result. Teams tune ",{"tag":132,"children":267},[268],"ef_search"," for latency and skip the comparison, which is how an index that returns four rows out of ten gets described as fast.",{"type":108,"level":109,"id":85,"text":86},{"type":101,"content":272},[273],"The whole surface is SQL, which is the reason to prefer it over a system with its own API. The snippet below is the shape of a production table: a fixed-dimension column, an exact index on the filter, an approximate index on the vector, and the one setting that decides whether a filtered query comes back short.",{"type":132,"code":275},"CREATE EXTENSION IF NOT EXISTS vector;\n\n-- The dimension is part of the type, so every row has to match it.\nCREATE TABLE chunks (\n  id         bigserial PRIMARY KEY,\n  tenant_id  text       NOT NULL,\n  embedding  vector(1536)\n);\n\n-- Exact index on the filter first: for a selective tenant it answers the\n-- whole query and the approximate index is never consulted.\nCREATE INDEX ON chunks (tenant_id);\n\n-- Bulk load with COPY, then build the approximate index on top of the data.\nCREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops)\n  WITH (m = 16, ef_construction = 64);\n\n-- A filter matching 10% of rows with ef_search = 40 returns about four rows,\n-- so the scan has to be allowed to continue past its first pass.\nSET hnsw.iterative_scan = relaxed_order;\nSET hnsw.ef_search = 100;\n\nSELECT id\nFROM chunks\nWHERE tenant_id = 'acme'\nORDER BY embedding \u003C=> (SELECT embedding FROM chunks WHERE id = 42)\nLIMIT 10;",{"type":101,"content":277},[278,279,282],"Three choices in that snippet are worth defending. The dimension is part of the type, so a row written by a different embedding model fails at ",{"tag":132,"children":280},[281],"INSERT"," instead of at query time. The approximate index is built after the data, because an HNSW graph created on an empty table has nothing to link and would be rebuilt anyway. And the probe vector comes from a subselect, which is the form the planner accepts; the same query with an expression inside the ORDER BY falls back to a sequential scan without warning.",{"type":256,"variant":284,"title":285,"body":286},"warn","Patch before you build an index",[287],[288,289,292,293,296],"CVE-2026-3172, fixed in 0.8.2 on 25 February 2026, was a buffer overflow in parallel HNSW index builds that could leak data from other relations or crash the server. The fix is an extension update, ",{"tag":132,"children":290},[291],"ALTER EXTENSION pgvector UPDATE",", with no reindex; while an upgrade is impossible, ",{"tag":132,"children":294},[295],"max_parallel_maintenance_workers = 0"," for the duration of the build is the documented mitigation. Later releases fixed further buffer overflows in IVFFlat builds, in 0.8.6 and again in 0.8.7 of 1 October 2026, so the version number the managed provider actually ships is worth reading rather than assuming.",{"type":108,"level":109,"id":88,"text":89},{"type":101,"content":299},[300],"Memory is the whole performance story. A vector column costs 4 bytes per dimension plus an 8-byte header, so 1,536 dimensions are about 6 KB per row before any index, and a full-precision HNSW index over 100 million 768-dimensional vectors measured 367 GB in AWS's benchmark, roughly 3.7 GB per million vectors. The alternative types exist to attack exactly that number.",{"type":302,"head":303,"rows":312},"table",[304,306,308,310],[305],"Type",[307],"Bytes per dimension",[309],"Indexable limit",[311],"What it costs you",[313,323,333,343],[314,317,319,321],[315],{"tag":132,"children":316},[134],[318],"4",[320],"2,000 dims",[322],"the baseline; exact search over it is exact recall",[324,327,329,331],[325],{"tag":132,"children":326},[138],[328],"2",[330],"4,000 dims",[332],"half the index, near-zero recall loss in the AWS tests",[334,337,339,341],[335],{"tag":132,"children":336},[142],[338],"1\u002F8",[340],"64,000 dims",[342],"Hamming over sign bits, needs reranking to hold recall",[344,347,349,351],[345],{"tag":132,"children":346},[146],[348],"8 per non-zero",[350],"1,000 non-zero",[352],"sparse embeddings, L2, cosine, inner product and L1",{"type":101,"content":354},[355],"AWS published the clearest public numbers on this, running VectorDBBench v0.3.4 at top_k=100 on Aurora PostgreSQL 18.4 with pgvector 0.8.0. On LAION 100M at 768 dimensions, an r8g.4xlarge with 128 GB held a 367 GB full-precision HNSW index it could not keep in cache: 3.4 queries per second cold at concurrency 10, 3,336 once warm, recall 0.965, 16.1 hours to build. Binary quantisation with reranking cut the index to 38 GB and the build to 1.1 hours, reached 13.5 cold and 895 warm queries per second, and paid for it in recall: 0.931.",{"type":101,"content":357},[358,359,361],"The same benchmark carries the counter-example, which is why its numbers should be read with their methodology. On Cohere 10M, whose 768-dimensional embeddings cluster near zero, binary quantisation needed a 3,000-candidate rerank to reach 0.93 recall and collapsed to 16 queries per second with a p99 of 1,640 ms, while full-precision HNSW on a 384 GB instance delivered 6,930 queries per second at 0.952 recall. Quantisation is distribution-dependent: validate recall on your own embeddings, or take ",{"tag":132,"children":360},[138]," and halve the index instead of guessing.",{"type":114,"ordered":115,"items":363},[364,370,380,386],[365,366,369],"Raise ",{"tag":132,"children":367},[368],"maintenance_work_mem"," before building an HNSW index; Postgres prints a notice when the graph stops fitting, and the README warns against raising it until the server runs out of memory.",[371,372,375,376,379],"Load with ",{"tag":132,"children":373},[374],"COPY"," and index afterwards, and use ",{"tag":132,"children":377},[378],"CREATE INDEX CONCURRENTLY"," in production so the build does not block writes.",[381,382,385],"VACUUM on an HNSW index can take a while; the documented speed-up is ",{"tag":132,"children":383},[384],"REINDEX INDEX CONCURRENTLY"," first, vacuum after.",[387],"Horizontal scale is borrowed rather than built: replication and point-in-time recovery come from the WAL, and the README points at Citus, PgDog or list partitioning for sharding.",{"type":108,"level":109,"id":91,"text":92},{"type":101,"content":390},[391],"The weaknesses are structural rather than unfinished. Everything runs on one Postgres node, so the index, the heap and the buffer cache compete for the same memory, and an index that stops fitting becomes an I\u002FO problem before it becomes a recall problem. Approximate search and selective filters still fight even with iterative scans, because a bounded scan is bounded. Vacuum and index maintenance are the database's chores rather than someone else's. And there is no built-in sharding: horizontal scale means replicas, partitioning or an extension.",{"type":302,"head":393,"rows":402},[394,396,398,400],[395],"Alternative",[397],"Runs as",[399],"Operational surface",[401],"Where it pulls ahead",[403,414,425],[404,408,410,412],[405],{"tag":120,"children":406},[407],"Qdrant",[409],"A separate Rust server, or the vendor cloud",[411],"Another cluster to patch, back up and secure",[413],"Payload filtering and quantisation tuned for recall at scale",[415,419,421,423],[416],{"tag":120,"children":417},[418],"Weaviate",[420],"A separate server with a GraphQL API, or the vendor cloud",[422],"The same again, plus its own module configuration",[424],"Hybrid search and vectorisation configured in one place",[426,430,432,434],[427],{"tag":120,"children":428},[429],"Chroma",[431],"Embedded in the process, or a small standalone server",[433],"Almost none, but no Postgres either",[435],"The shortest path from a prototype to a running system",{"type":101,"content":437},[438,439,441],"The honest boundary: pgvector wins while the vector work is a column of data the team already stores, and starts losing when one query has to hold a large graph, a filtered scan and the rest of the application's working set in the same memory. AWS measured that boundary at 367 GB of index for 100 million vectors and worked around it with quantisation and partitioning. Teams that do not want to own that trade have four exits, ",{"tag":132,"children":440},[138],", binary quantisation with reranking, partitioning by tenant, or a dedicated store, and the first two are cheap enough that they should be tried before the fourth is discussed.",{"type":108,"level":109,"id":94,"text":95},{"type":101,"content":444},[445],"pgvector should be the default answer to where these embeddings go for any team that already runs Postgres, and a dedicated vector database should have to argue its way past it. The extension has the unusual property that its failure modes are the failure modes of a database the team already understands: memory pressure, maintenance windows, one node's write throughput. Choose something else deliberately, at a scale or a latency target that can be named.",{"type":114,"ordered":447,"items":448},true,[449,451,453,455,460],[450],"Take pgvector when the vectors describe rows the team already stores, and tenant isolation, cascading deletes or a JOIN with the source table have to be transactional.",[452],"Take it when the corpus is up to a few tens of millions of vectors and the filter is selective enough that a B-tree on the filter column carries most of the query.",[454],"Take it when the alternative is a second production system: the extension inherits backup, replication, monitoring and access control that already exist and adds nothing new to operate.",[456,457,459],"Do not take it when a single query has to keep a multi-hundred-gigabyte graph plus the application's working set in memory, unless ",{"tag":132,"children":458},[138]," or binary quantisation has already been measured on the actual embeddings.",[461],"Do not take it when the requirement is sustained multi-node write throughput or low-latency filtered recall across hundreds of millions of vectors; that is partitioning or a purpose-built store, and postponing it costs a migration later.",{"type":463,"content":464},"quote",[465,466,467],"Not all embedding models produce vectors that quantize well. Validate on your data before committing."," ","— AWS Database Blog, 18 August 2026",{"type":108,"level":109,"id":97,"text":98},{"type":114,"ordered":447,"items":470},[471,476,480,484,488],[472],{"tag":473,"href":23,"children":474},"a",[475],"pgvector README: types, indexing, filtering and scaling",[477],{"tag":473,"href":32,"children":478},[479],"pgvector changelog, 0.1.0 through 0.8.7",[481],{"tag":473,"href":35,"children":482},[483],"PostgreSQL news: pgvector 0.8.2 released (CVE-2026-3172)",[485],{"tag":473,"href":38,"children":486},[487],"AWS: Scale pgvector with binary quantization on Aurora PostgreSQL",[489],{"tag":473,"href":41,"children":490},[491],"pgvector licence: the PostgreSQL licence",[493,542,596,646],{"slug":494,"published":495,"minutes":496,"category":7,"tags":497,"keywords":502,"about":511,"sources":515,"cover":534,"og":535,"expertise":44,"locales":536,"lang":46,"title":537,"description":538,"coverAlt":539,"url":540,"pricing":541,"kind":498},"zep","2026-10-06",11,[498,499,500,501,12],"Agent memory","Knowledge graph","Temporal graph","Context engineering",[503,504,505,506,507,508,509,510],"zep ai","zep agent memory","graphiti knowledge graph","zep pricing","zep vs mem0","long-term memory for agents","temporal knowledge graph","zep cloud",[512],{"name":513,"url":514},"Zep","https:\u002F\u002Fwww.getzep.com\u002F",[516,519,522,525,528,531],{"title":517,"url":518},"Zep pricing: plans, credits and limits","https:\u002F\u002Fwww.getzep.com\u002Fpricing",{"title":520,"url":521},"Zep documentation","https:\u002F\u002Fhelp.getzep.com\u002F",{"title":523,"url":524},"Graphiti on GitHub","https:\u002F\u002Fgithub.com\u002Fgetzep\u002Fgraphiti",{"title":526,"url":527},"Graphiti product page","https:\u002F\u002Fwww.getzep.com\u002Fplatform\u002Fgraphiti\u002F",{"title":529,"url":530},"Announcing a new direction for Zep's open-source strategy","https:\u002F\u002Fwww.getzep.com\u002Fblog\u002Fannouncing-a-new-direction-for-zeps-open-source-strategy\u002F",{"title":532,"url":533},"Graphiti: temporal knowledge graphs for AI agents (arXiv)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2501.13956","\u002Fimages\u002Fblog\u002Fzep\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fzep\u002Fog.jpg",[46,47,48],"Zep review: agent memory on a temporal graph","Zep is a hosted agent-memory API on a temporal knowledge graph: credits on writes, retrieval free, Flex from $125 a month, Graphiti as the part you can self-host.","Diagram of how a fact reaches the prompt in Zep: messages and facts are extracted into a per-user context graph of entities and relationships, and retrieval walks the graph to return a context block with the supporting facts.","https:\u002F\u002Fwww.getzep.com","Apache-2.0 core · Cloud from $50 per month",{"slug":543,"published":5,"minutes":544,"category":7,"tags":545,"keywords":548,"about":556,"sources":563,"cover":588,"og":589,"expertise":44,"locales":590,"lang":46,"title":591,"description":592,"coverAlt":593,"url":594,"pricing":595,"kind":558},"lancedb",9,[9,546,547,12],"Hybrid search","Embedded database",[543,549,550,551,552,553,554,555],"lancedb review","lance vector database","embedded vector database","lancedb vs qdrant","hybrid search rrf","lancedb indexing","lance data format",[557,560],{"name":558,"url":559},"Vector database","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVector_database",{"name":561,"url":562},"Apache Arrow","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FApache_Arrow",[564,567,570,573,576,579,582,585],{"title":565,"url":566},"LanceDB quickstart","https:\u002F\u002Fdocs.lancedb.com\u002Fquickstart",{"title":568,"url":569},"LanceDB vector indexes","https:\u002F\u002Fdocs.lancedb.com\u002Findexing\u002Fvector-index",{"title":571,"url":572},"LanceDB indexing guide","https:\u002F\u002Fdocs.lancedb.com\u002Findexing\u002Findex",{"title":574,"url":575},"LanceDB hybrid search","https:\u002F\u002Fdocs.lancedb.com\u002Fsearch\u002Fhybrid-search",{"title":577,"url":578},"LanceDB Enterprise","https:\u002F\u002Fdocs.lancedb.com\u002Fenterprise",{"title":580,"url":581},"LanceDB frequently asked questions","https:\u002F\u002Fdocs.lancedb.com\u002Ffaq\u002Ffaq-oss",{"title":583,"url":584},"LanceDB pricing","https:\u002F\u002Flancedb.com\u002Fpricing",{"title":586,"url":587},"LanceDB on PyPI","https:\u002F\u002Fpypi.org\u002Fproject\u002Flancedb\u002F","\u002Fimages\u002Fblog\u002Flancedb\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flancedb\u002Fog.jpg",[46,47,48],"LanceDB: vector search that starts as a library","A review of LanceDB: an Apache-2.0 embedded vector library, its IVF and HNSW index choices, hybrid search with rank fusion, and what the Enterprise tier adds.","Cover art for the LanceDB review: one Lance table feeding a vector index and a full-text index into a fused ranking","https:\u002F\u002Flancedb.com","Apache-2.0 · Cloud paid",{"slug":597,"published":598,"minutes":6,"category":7,"tags":599,"keywords":601,"about":607,"sources":611,"cover":639,"og":640,"expertise":44,"locales":641,"lang":46,"title":642,"description":643,"coverAlt":644,"url":610,"pricing":645,"kind":498},"mem0","2026-09-17",[498,600,12,9],"Long-term memory",[597,602,603,604,605,508,606],"mem0 review","agent memory layer","mem0 self-hosted","mem0 pricing","mem0 alternatives",[608],{"name":609,"url":610},"Mem0","https:\u002F\u002Fmem0.ai",[612,615,618,621,624,627,630,633,636],{"title":613,"url":614},"Mem0 documentation","https:\u002F\u002Fdocs.mem0.ai\u002Fintroduction",{"title":616,"url":617},"Mem0 quickstart","https:\u002F\u002Fdocs.mem0.ai\u002Fquickstart",{"title":619,"url":620},"How Mem0 works","https:\u002F\u002Fdocs.mem0.ai\u002Fcore-concepts\u002Fhow-it-works",{"title":622,"url":623},"Mem0 pricing","https:\u002F\u002Fmem0.ai\u002Fpricing",{"title":625,"url":626},"Mem0 on GitHub","https:\u002F\u002Fgithub.com\u002Fmem0ai\u002Fmem0",{"title":628,"url":629},"mem0ai on PyPI","https:\u002F\u002Fpypi.org\u002Fproject\u002Fmem0ai\u002F",{"title":631,"url":632},"Mem0 research and benchmarks","https:\u002F\u002Fmem0.ai\u002Fresearch",{"title":634,"url":635},"Mem0 MCP server","https:\u002F\u002Fdocs.mem0.ai\u002Fplatform\u002Fmem0-mcp",{"title":637,"url":638},"Mem0 paper on arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2504.19413","\u002Fimages\u002Fblog\u002Fmem0\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmem0\u002Fog.jpg",[46,47,48],"Mem0: what an agent memory layer costs per turn","A review of Mem0: facts extracted from every turn, the April 2026 benchmark table and its platform-only caveat, four cloud tiers and what self-hosting leaves out.","A loop that turns conversation into stored facts and reads them back into the prompt","Free tier · from $19 per month",{"slug":647,"published":648,"minutes":6,"category":7,"tags":649,"keywords":652,"about":661,"sources":671,"cover":690,"og":691,"expertise":44,"locales":692,"lang":46,"title":693,"description":694,"coverAlt":695,"url":664,"pricing":696,"kind":558},"milvus-zilliz","2026-09-08",[9,546,650,651,12],"BM25 full text","Distributed",[653,654,655,656,657,658,659,660],"milvus","milvus vs qdrant","zilliz cloud pricing","vector database comparison","milvus hybrid search","apache milvus self-hosting","milvus 3.0","rag vector store",[662,665,668],{"name":663,"url":664},"Milvus","https:\u002F\u002Fmilvus.io",{"name":666,"url":667},"Zilliz Cloud","https:\u002F\u002Fzilliz.com",{"name":669,"url":670},"Retrieval-augmented generation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FRetrieval-augmented_generation",[672,675,678,681,684,687],{"title":673,"url":674},"Milvus architecture overview","https:\u002F\u002Fmilvus.io\u002Fdocs\u002Farchitecture_overview.md",{"title":676,"url":677},"Milvus release notes","https:\u002F\u002Fmilvus.io\u002Fdocs\u002Frelease_notes.md",{"title":679,"url":680},"Milvus releases on GitHub","https:\u002F\u002Fgithub.com\u002Fmilvus-io\u002Fmilvus\u002Freleases",{"title":682,"url":683},"Milvus README: features and licence","https:\u002F\u002Fgithub.com\u002Fmilvus-io\u002Fmilvus",{"title":685,"url":686},"Zilliz Cloud pricing","https:\u002F\u002Fzilliz.com\u002Fpricing",{"title":688,"url":689},"Zilliz Cloud list price","https:\u002F\u002Fzilliz.com\u002Fpricing\u002Fpricing-guide","\u002Fimages\u002Fblog\u002Fmilvus-zilliz\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmilvus-zilliz\u002Fog.jpg",[46,47,48],"Milvus review: the most complete vector database to operate","Milvus 3.0.2 is the most complete open-source vector database and the heaviest to run. A review of its architecture, hybrid search, costs and where it should not be used.","Cover artwork for the Milvus review showing a pipeline from ingest to index, search and reranking","Apache-2.0 · Zilliz Cloud free tier",1791383548705]