[{"data":1,"prerenderedAt":646},["ShallowReactive",2],{"tool-unstructured-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":23,"sources":27,"cover":46,"og":47,"expertise":48,"locales":49,"lang":50,"title":53,"description":54,"coverAlt":55,"url":56,"pricing":57,"kind":58,"metaTitle":59,"takeaways":60,"faq":66,"toc":79,"blocks":104,"others":447},"unstructured","2026-05-21",11,"rag",[9,10,11,12,13],"Document parsing","RAG ingestion","PDF extraction","Chunking","ETL",[15,16,17,18,19,20,21,22],"unstructured.io","unstructured library python","unstructured vs llamaparse","document parsing for rag","pdf partitioning hi_res","unstructured pricing","chunk by title","document ingestion pipeline",[24],{"name":25,"url":26},"Unstructured","https:\u002F\u002Funstructured.io\u002F",[28,31,34,37,40,43],{"title":29,"url":30},"Unstructured documentation: pipelines overview","https:\u002F\u002Fdocs.unstructured.io\u002Fwelcome",{"title":32,"url":33},"Unstructured documentation: transform quickstart","https:\u002F\u002Fdocs.unstructured.io\u002Fapi-reference\u002Ftransform\u002Fquickstart\u002Foverview",{"title":35,"url":36},"Unstructured pricing: free pages, rate per page and plans","https:\u002F\u002Funstructured.io\u002Fpricing",{"title":38,"url":39},"Unstructured benchmarks: 1,000+ enterprise pages against other parsers","https:\u002F\u002Funstructured.io\u002Fbenchmarks",{"title":41,"url":42},"Unstructured library on GitHub","https:\u002F\u002Fgithub.com\u002FUnstructured-IO\u002Funstructured",{"title":44,"url":45},"Unstructured partition endpoint container","https:\u002F\u002Fgithub.com\u002FUnstructured-IO\u002Funstructured-api","\u002Fimages\u002Fblog\u002Funstructured\u002Fcover.webp","\u002Fimages\u002Fblog\u002Funstructured\u002Fog.jpg","ai-engineer",[50,51,52],"en","de","hu","Unstructured review: parsing documents for RAG","Unstructured turns PDFs, Word files and images into typed elements for RAG: an Apache-2.0 library plus a platform at $0.015 per page after 10,000 free pages.","Diagram of the Unstructured pipeline: source files are partitioned into typed elements with layout and OCR, grouped into chunks, enriched with metadata and tables, embedded and loaded into one of more than twenty destinations.","https:\u002F\u002Funstructured.io","Apache-2.0 · paid tiers","Document ingestion","Unstructured review: parsing, chunking, price · Balázs Csorba",[61,62,63,64,65],"Unstructured is an Apache-2.0 Python library that turns PDFs, Word files and images into typed elements such as titles, tables and lists, with a hosted platform that adds connectors and the VLM strategy.","The platform charges $0.015 per page after the first 10,000 free pages, while the library runs locally with no page count.","In the vendor's own benchmark over 1,000 enterprise pages the open-source build reaches 0.426 table-cell content accuracy against 0.820 for the best platform pipeline.","Table structure and reading order fail without an error, so a sample of the corpus should be read before the embedding model is chosen.","Telemetry can be switched off with DO_NOT_TRACK and SCARF_NO_ANALYTICS; with the library the documents themselves never leave the machine.",[67,70,73,76],{"q":68,"a":69},"Is the unstructured library free to use?","Yes. The library is Apache-2.0, installed from PyPI, and runs without an account or a page limit. What is paid is the hosted platform: 10,000 free pages to start and $0.015 per page after that.",{"q":71,"a":72},"What is the difference between the library and the platform?","The library partitions, chunks and stages data inside your own process. The platform adds managed jobs, more than 40 maintained connectors, change detection so only new or modified files run, the VLM partitioning strategy and the compliance certifications. In the vendor's benchmark that shows up as 0.426 against 0.820 table-cell content accuracy.",{"q":74,"a":75},"Which partitioning strategy should be used?","auto is the default and picks per file, fast reads the embedded text and is the cheapest, hi_res runs layout detection and OCR and is the choice for scans and tables, and ocr_only handles pages with no text layer. The VLM strategy exists only in the platform.",{"q":77,"a":78},"How much does a large document set cost to process?","Beyond the first 10,000 free pages, $0.015 per page means a 500-page scan costs $7.50 on the pay-as-you-go plan. Self-hosting the library costs compute only, which is why large archives are usually parsed locally and only difficult documents go to the platform.",[80,83,86,89,92,95,98,101],{"id":81,"title":82},"what-it-is","What it is",{"id":84,"title":85},"how-it-works","How it works",{"id":87,"title":88},"getting-started","Getting started",{"id":90,"title":91},"performance","Performance and cost",{"id":93,"title":94},"pricing","Pricing",{"id":96,"title":97},"where-it-shingles","Where it shingles",{"id":99,"title":100},"verdict","Verdict",{"id":102,"title":103},"sources","Sources",[105,113,116,119,154,192,193,205,214,217,218,221,223,229,236,237,249,297,300,314,317,318,321,331,332,335,380,387,400,401,404,421,425,426],{"type":106,"content":107},"paragraph",[108,109],"Unstructured is the document-parsing layer most retrieval pipelines start with: an Apache-2.0 Python library that turns PDFs, Word files, HTML, images and spreadsheets into typed elements, titles, tables, lists and prose, plus a hosted platform that adds connectors, change detection and the strategies the library does not ship. ",{"tag":110,"children":111},"strong",[112],"The library is the right starting point for nearly everyone; the platform is worth paying for only when tables or scanned archives decide whether the pipeline works at all.",{"type":106,"content":114},[115],"It sits at the front of the stack, before chunking, embedding and retrieval, and it competes with LlamaParse, Docling, Reducto, Azure Document Intelligence and the parsing endpoints of the large cloud vendors. Its advantage is not that it wins every benchmark, its own published results say otherwise for the open-source build, but that the core runs locally under a permissive licence, so nothing in this layer forces a vector database, a chunker or a hosting choice.",{"type":117,"level":118,"id":81,"text":82},"heading",2,{"type":106,"content":120},[121,122,126,127,130,131,134,135,138,139,138,142,145,146,149,150,153],"Install with ",{"tag":123,"children":124},"code",[125],"pip install \"unstructured[all-docs]\"",", call ",{"tag":123,"children":128},[129],"partition()"," on a file and get back a list of element objects, each with a type, text and metadata such as page number and bounding box. The ",{"tag":123,"children":132},[133],"strategy"," argument picks the pipeline: ",{"tag":123,"children":136},[137],"auto",", ",{"tag":123,"children":140},[141],"fast",{"tag":123,"children":143},[144],"hi_res"," and ",{"tag":123,"children":147},[148],"ocr_only"," in the library, with ",{"tag":123,"children":151},[152],"vlm"," added by the platform to route pages through a vision model. Around that sits a chain, partition, chunk, enrich, embed and load, which the platform runs as a job and the library runs as ordinary functions.",{"type":155,"ordered":156,"items":157},"list",false,[158,160,162,164,178,180,182],[159],"Licence and ownership: Apache-2.0 for the library; the platform is a commercial product of Unstructured with a free tier, pay-as-you-go and a custom tier.",[161],"Coverage: the pricing page lists 45+ supported file types and 40+ connectors, split across more than 20 sources and more than 20 destinations.",[163],"Output: typed elements such as Title, NarrativeText, Table, ListItem and Image, carrying page numbers, coordinates and element metadata rather than one wall of text.",[165,166,138,168,138,170,145,172,174,175,177],"Strategies: ",{"tag":123,"children":167},[137],{"tag":123,"children":169},[141],{"tag":123,"children":171},[144],{"tag":123,"children":173},[148]," locally, plus ",{"tag":123,"children":176},[152]," and the enrichment steps in the platform.",[179],"Chunking: by title, by page, by character and by similarity, with contextual chunking offered as a platform feature.",[181],"Benchmarks: Unstructured publishes results over 1,000+ enterprise pages against Reducto, LlamaParse, Docling, Snowflake, Databricks and NVIDIA.",[183,184,187,188,191],"Telemetry: the library reports anonymous usage data unless ",{"tag":123,"children":185},[186],"DO_NOT_TRACK"," or ",{"tag":123,"children":189},[190],"SCARF_NO_ANALYTICS"," is set.",{"type":117,"level":118,"id":84,"text":85},{"type":106,"content":194},[195,196,198,199,201,202,204],"A run reads the file, renders pages when the strategy needs pixels, then classifies regions into elements. ",{"tag":123,"children":197},[141]," takes the embedded text and labels it; ",{"tag":123,"children":200},[144]," runs layout detection and OCR over the page image, which is why it is slower, why it needs the model stack installed, and why it is the strategy that still produces tables worth keeping; ",{"tag":123,"children":203},[148]," is the fallback when a page has no text layer at all. Chunking then groups elements under a character budget while trying not to separate a heading from the section it introduces.",{"type":206,"attrs":207,"inner":211,"caption":212},"diagram",{"viewBox":208,"role":209,"aria-labelledby":210},"0 0 760 290","img","d-uzz-t d-uzz-d","\u003Ctitle id=\"d-uzz-t\">How a document becomes retrievable text\u003C\u002Ftitle>\u003Cdesc id=\"d-uzz-d\">Source files enter at the top and are split into typed elements by the partition step, which reads layout and runs OCR. The elements are grouped into chunks by title, page or size, enriched with metadata and table structure, embedded into vectors, and loaded into one of more than twenty destinations.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"ah-p\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003Cmarker id=\"ah-pa\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head-accent\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Crect x=\"260\" y=\"16\" width=\"240\" height=\"42\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"380\" y=\"43\" text-anchor=\"middle\" class=\"d-text\">source files\u003C\u002Ftext>\u003Crect x=\"40\" y=\"118\" width=\"140\" height=\"76\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"110\" y=\"150\" text-anchor=\"middle\" class=\"d-text\">partition\u003C\u002Ftext>\u003Ctext x=\"110\" y=\"174\" text-anchor=\"middle\" class=\"d-small\">layout + OCR\u003C\u002Ftext>\u003Crect x=\"220\" y=\"118\" width=\"140\" height=\"76\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"290\" y=\"150\" text-anchor=\"middle\" class=\"d-text\">chunk\u003C\u002Ftext>\u003Ctext x=\"290\" y=\"174\" text-anchor=\"middle\" class=\"d-small\">title, page, size\u003C\u002Ftext>\u003Crect x=\"400\" y=\"118\" width=\"140\" height=\"76\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"470\" y=\"150\" text-anchor=\"middle\" class=\"d-text\">enrich\u003C\u002Ftext>\u003Ctext x=\"470\" y=\"174\" text-anchor=\"middle\" class=\"d-small\">metadata, tables\u003C\u002Ftext>\u003Crect x=\"580\" y=\"118\" width=\"140\" height=\"76\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"650\" y=\"150\" text-anchor=\"middle\" class=\"d-text\">embed\u003C\u002Ftext>\u003Ctext x=\"650\" y=\"174\" text-anchor=\"middle\" class=\"d-small\">vectors + keys\u003C\u002Ftext>\u003Cpath d=\"M380 58 L110 114\" class=\"d-line\" marker-end=\"url(#ah-p)\" \u002F>\u003Cpath d=\"M380 58 L290 114\" class=\"d-line\" marker-end=\"url(#ah-p)\" \u002F>\u003Cpath d=\"M380 58 L470 114\" class=\"d-line\" marker-end=\"url(#ah-p)\" \u002F>\u003Cpath d=\"M380 58 L650 114\" class=\"d-line\" marker-end=\"url(#ah-p)\" \u002F>\u003Cpath d=\"M184 156 H214\" class=\"d-line\" marker-end=\"url(#ah-p)\" \u002F>\u003Cpath d=\"M364 156 H394\" class=\"d-line\" marker-end=\"url(#ah-p)\" \u002F>\u003Cpath d=\"M544 156 H574\" class=\"d-line\" marker-end=\"url(#ah-p)\" \u002F>\u003Cpath d=\"M650 198 V242\" class=\"d-line-accent\" marker-end=\"url(#ah-pa)\" \u002F>\u003Ctext x=\"650\" y=\"266\" text-anchor=\"middle\" class=\"d-title\">20+ destinations\u003C\u002Ftext>",[213],"Every stage is a separate call in the library and a separate step in the platform, so a pipeline can be re-run from the middle without re-parsing the whole corpus.",{"type":106,"content":215},[216],"What makes or breaks the output is the element boundary. A table that comes out as HTML with its header row intact chunks into one useful node; the same table read as prose becomes several chunks that each carry half a number. That is the whole argument for caring about this layer, and also the whole reason to read the parser output before changing the embedding model.",{"type":117,"level":118,"id":87,"text":88},{"type":106,"content":219},[220],"The library quickstart is a single function call. The snippet below parses a PDF with the high-resolution strategy, keeps table structure, chunks by title and writes the result as JSON.",{"type":123,"code":222},"from unstructured.partition.auto import partition\nfrom unstructured.chunking.title import chunk_by_title\nfrom unstructured.staging.base import elements_to_json\n\n# strategy: auto, fast, hi_res or ocr_only; hi_res is the one that keeps table structure\nelements = partition(\n    filename=\"report.pdf\",\n    strategy=\"hi_res\",\n    infer_table_structure=True,\n)\nchunks = chunk_by_title(elements, max_characters=1_200, combine_under=300)\n\nelements_to_json(chunks, filename=\"report.elements.json\")\nprint(len(elements), \"elements,\", len(chunks), \"chunks\")",{"type":106,"content":224},[225,226,228],"For anything beyond a local experiment, the platform is the maintained path: the same partition step behind an API or a scheduled job, connectors to S3, SharePoint, Google Drive and the rest, change detection so only new or modified files are processed, and the ",{"tag":123,"children":227},[152]," strategy. The library remains where the parsing code itself lives, and the documentation is written around it.",{"type":230,"variant":231,"title":232,"body":233},"callout","tip","Read the elements before touching the embeddings",[234],[235],"Partition a hundred pages of your own corpus and read the output: element types, table structure, reading order. The failure at this layer is silent, a parser that drops half a table still returns clean JSON, and every metric downstream will blame the retriever instead.",{"type":117,"level":118,"id":90,"text":91},{"type":106,"content":238},[239,240,242,243,245,246,248],"The platform bills pages: 10,000 free to start and $0.015 per page after that, so a 400-page scanned archive is six dollars of parsing before anyone has written a query. The library bills in CPU and wall time instead: ",{"tag":123,"children":241},[141]," is close to I\u002FO-bound, ",{"tag":123,"children":244},[144]," runs a layout model and OCR per page, and the ",{"tag":123,"children":247},[152]," strategy moves the bill from compute to model tokens.",{"type":250,"head":251,"rows":260},"table",[252,254,256,258],[253],"Pipeline",[255],"Adjusted CCT",[257],"Tokens added",[259],"Table cell content",[261,270,279,288],[262,264,266,268],[263],"Unstructured platform",[265],"0.880",[267],"0.051",[269],"0.820",[271,273,275,277],[272],"Unstructured open source",[274],"0.715",[276],"0.119",[278],"0.426",[280,282,284,286],[281],"LlamaParse VLM",[283],"0.835",[285],"0.069",[287],"0.522",[289,291,293,295],[290],"Docling default",[292],"0.716",[294],"0.135",[296],"0.657",{"type":106,"content":298},[299],"These numbers come from Unstructured's own benchmark of 1,000+ enterprise pages, scanned invoices, nested tables and handwriting, and a vendor-run comparison is marketing, so the useful signal is the gap inside a single product: 0.426 against 0.820 table-cell content between the open-source build and the best platform pipeline, and 0.715 against 0.880 on adjusted text accuracy. Element alignment, whether a region was labelled as heading, table or paragraph, is where every tool in that table clusters between 0.53 and 0.61, which is the honest difficulty of this layer.",{"type":155,"ordered":156,"items":301},[302,304,312],[303],"Parse once and keep the element JSON: re-partitioning the same corpus while experimenting with chunking is the most common way this layer burns money.",[305,306,308,309,311],"Pick the strategy per document type rather than per corpus: ",{"tag":123,"children":307},[141]," for born-digital text, ",{"tag":123,"children":310},[144]," for scans and anything with tables.",[313],"Count pages rather than files, because a per-page price makes one 400-page PDF the budget item and not the number of documents.",{"type":106,"content":315},[316],"None of this is specific to this vendor, every parser trades recall against compute. What is specific is that the numbers are published at all, in a table with competitor names attached, which is more than most of the field discloses.",{"type":117,"level":118,"id":93,"text":94},{"type":106,"content":319},[320],"Two products share one name. The library is Apache-2.0 and free: install it, run it on your own machine, no account and no page count. The platform is the same parsing wrapped in managed jobs, connectors and compliance, metered by the page.",{"type":155,"ordered":156,"items":322},[323,325,327,329],[324],"Library: Apache-2.0, installed from PyPI, unlimited pages, your own hardware.",[326],"Free: 10,000 pages to start, no card required, all features included.",[328],"Pay-as-you-go: $0.015 per page after the first 10,000 pages, all features included.",[330],"Business: custom pricing for a dedicated instance, VPC or bare-metal deployment, multi-user accounts, role-based access control and the vendor's compliance certifications.",{"type":117,"level":118,"id":96,"text":97},{"type":106,"content":333},[334],"Start with the weaknesses. The open-source build is the weaker parser and the vendor's own table says so: 0.426 table-cell content against 0.820 for the best platform pipeline, 0.119 invented tokens against 0.051. The price is per page, which rewards born-digital PDFs and punishes scans, and a page is a poor unit of work when one page holds a paragraph and the next holds a 400-cell table. The features that make the platform worth renting, VLM partitioning, incremental change detection, 40+ maintained connectors and the compliance story, are exactly what the licence does not contain. And output still needs spot checks: reading order and table structure are the two fields that fail without an error.",{"type":250,"head":336,"rows":344},[337,339,340,342],[338],"Tool",[82],[341],"Where it wins",[343],"What you give up",[345,353,362,371],[346,347,349,351],[25],[348],"Library plus hosted platform",[350],"Local run under Apache-2.0 with a published quality benchmark",[352],"Per-page billing, and the strong numbers need the paid pipeline",[354,356,358,360],[355],"LlamaParse",[357],"Hosted parser from the LlamaIndex team",[359],"Fast setup and tight LlamaIndex integration",[361],"Hosted only, so every page leaves your network",[363,365,367,369],[364],"Docling",[366],"IBM's open-source parser",[368],"One dependency, MIT licence, strong table output",[370],"Fewer file types and no managed connector layer",[372,374,376,378],[373],"Reducto",[375],"Hosted parsing API with layout controls",[377],"Table accuracy and layout controls as a service",[379],"API only: no local run and no library to extend",{"type":106,"content":381},[382,383,386],"The real decision is who pays for quality. If the corpus is born-digital text and adequate output is enough, the library with ",{"tag":123,"children":384},[385],"strategy=\"fast\""," costs nothing per page and is the right size for the job. If tables, scans or regulated data decide whether the pipeline works, the platform or a specialist parser is the purchase, and the only test that matters is a hundred pages of your own documents scored on the fields you actually read.",{"type":230,"variant":388,"title":389,"body":390},"warn","Telemetry and what leaves your machine",[391],[392,393,145,396,399],"With the library the documents stay local, but the package reports anonymous usage data by default; ",{"tag":123,"children":394},[395],"DO_NOT_TRACK=1",{"tag":123,"children":397},[398],"SCARF_NO_ANALYTICS=1"," switch that off. The platform by design sends pages to a hosted service, and the security pages list encryption in transit, zero data retention and dedicated inference for teams that cannot do that.",{"type":117,"level":118,"id":99,"text":100},{"type":106,"content":402},[403],"Unstructured is the safest default at the front of a retrieval pipeline because it is boring, local and measurable: one function that returns typed elements, a benchmark you can argue with, and a paid tier you can move to without rewriting ingestion. What is actually being decided is how much parsing quality the product needs, and the vendor's own numbers say the free build is not the paid one.",{"type":155,"ordered":405,"items":406},true,[407,409,411,413,419],[408],"Use the library if a pipeline already exists and needs typed elements rather than raw text; it is Apache-2.0, runs locally and can be replaced without touching the rest of the stack.",[410],"Use the platform when pages arrive from S3, SharePoint or Confluence and need connectors, change detection and an audit trail.",[412],"Do not accept it as the parser of record for financial tables without scoring your own documents; the gap between strategies inside one product is larger than the gap between vendors.",[414,415,418],"Do not choose it where a heavy local dependency tree is unacceptable, because ",{"tag":123,"children":416},[417],"unstructured[all-docs]"," pulls in the OCR and layout stack while Docling or a hosted API is much lighter.",[420],"Budget in pages from the start; a per-page price is easy to model and easy to exceed with scanned archives.",{"type":422,"content":423},"quote",[424],"Treat document parsing as a measured step rather than plumbing: choose the strategy on a sample of your own pages, keep the element JSON, and change the parser only when the table-cell numbers move.",{"type":117,"level":118,"id":102,"text":103},{"type":155,"ordered":405,"items":427},[428,432,435,438,441,444],[429],{"tag":430,"href":30,"children":431},"a",[29],[433],{"tag":430,"href":33,"children":434},[32],[436],{"tag":430,"href":36,"children":437},[35],[439],{"tag":430,"href":39,"children":440},[38],[442],{"tag":430,"href":42,"children":443},[41],[445],{"tag":430,"href":45,"children":446},[44],[448,497,553,596],{"slug":449,"published":450,"minutes":6,"category":7,"tags":451,"keywords":457,"about":466,"sources":470,"cover":489,"og":490,"expertise":48,"locales":491,"lang":50,"title":492,"description":493,"coverAlt":494,"url":495,"pricing":496,"kind":452},"zep","2026-10-06",[452,453,454,455,456],"Agent memory","Knowledge graph","Temporal graph","Context engineering","RAG",[458,459,460,461,462,463,464,465],"zep ai","zep agent memory","graphiti knowledge graph","zep pricing","zep vs mem0","long-term memory for agents","temporal knowledge graph","zep cloud",[467],{"name":468,"url":469},"Zep","https:\u002F\u002Fwww.getzep.com\u002F",[471,474,477,480,483,486],{"title":472,"url":473},"Zep pricing: plans, credits and limits","https:\u002F\u002Fwww.getzep.com\u002Fpricing",{"title":475,"url":476},"Zep documentation","https:\u002F\u002Fhelp.getzep.com\u002F",{"title":478,"url":479},"Graphiti on GitHub","https:\u002F\u002Fgithub.com\u002Fgetzep\u002Fgraphiti",{"title":481,"url":482},"Graphiti product page","https:\u002F\u002Fwww.getzep.com\u002Fplatform\u002Fgraphiti\u002F",{"title":484,"url":485},"Announcing a new direction for Zep's open-source strategy","https:\u002F\u002Fwww.getzep.com\u002Fblog\u002Fannouncing-a-new-direction-for-zeps-open-source-strategy\u002F",{"title":487,"url":488},"Graphiti: temporal knowledge graphs for AI agents (arXiv)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2501.13956","\u002Fimages\u002Fblog\u002Fzep\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fzep\u002Fog.jpg",[50,51,52],"Zep review: agent memory on a temporal graph","Zep is a hosted agent-memory API on a temporal knowledge graph: credits on writes, retrieval free, Flex from $125 a month, Graphiti as the part you can self-host.","Diagram of how a fact reaches the prompt in Zep: messages and facts are extracted into a per-user context graph of entities and relationships, and retrieval walks the graph to return a context block with the supporting facts.","https:\u002F\u002Fwww.getzep.com","Apache-2.0 core · Cloud from $50 per month",{"slug":498,"published":499,"minutes":500,"category":7,"tags":501,"keywords":505,"about":513,"sources":520,"cover":545,"og":546,"expertise":48,"locales":547,"lang":50,"title":548,"description":549,"coverAlt":550,"url":551,"pricing":552,"kind":515},"lancedb","2026-09-21",9,[502,503,504,456],"Vector search","Hybrid search","Embedded database",[498,506,507,508,509,510,511,512],"lancedb review","lance vector database","embedded vector database","lancedb vs qdrant","hybrid search rrf","lancedb indexing","lance data format",[514,517],{"name":515,"url":516},"Vector database","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVector_database",{"name":518,"url":519},"Apache Arrow","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FApache_Arrow",[521,524,527,530,533,536,539,542],{"title":522,"url":523},"LanceDB quickstart","https:\u002F\u002Fdocs.lancedb.com\u002Fquickstart",{"title":525,"url":526},"LanceDB vector indexes","https:\u002F\u002Fdocs.lancedb.com\u002Findexing\u002Fvector-index",{"title":528,"url":529},"LanceDB indexing guide","https:\u002F\u002Fdocs.lancedb.com\u002Findexing\u002Findex",{"title":531,"url":532},"LanceDB hybrid search","https:\u002F\u002Fdocs.lancedb.com\u002Fsearch\u002Fhybrid-search",{"title":534,"url":535},"LanceDB Enterprise","https:\u002F\u002Fdocs.lancedb.com\u002Fenterprise",{"title":537,"url":538},"LanceDB frequently asked questions","https:\u002F\u002Fdocs.lancedb.com\u002Ffaq\u002Ffaq-oss",{"title":540,"url":541},"LanceDB pricing","https:\u002F\u002Flancedb.com\u002Fpricing",{"title":543,"url":544},"LanceDB on PyPI","https:\u002F\u002Fpypi.org\u002Fproject\u002Flancedb\u002F","\u002Fimages\u002Fblog\u002Flancedb\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flancedb\u002Fog.jpg",[50,51,52],"LanceDB: vector search that starts as a library","A review of LanceDB: an Apache-2.0 embedded vector library, its IVF and HNSW index choices, hybrid search with rank fusion, and what the Enterprise tier adds.","Cover art for the LanceDB review: one Lance table feeding a vector index and a full-text index into a fused ranking","https:\u002F\u002Flancedb.com","Apache-2.0 · Cloud paid",{"slug":554,"published":499,"minutes":555,"category":7,"tags":556,"keywords":560,"about":567,"sources":573,"cover":588,"og":589,"expertise":48,"locales":590,"lang":50,"title":591,"description":592,"coverAlt":593,"url":569,"pricing":594,"kind":595},"pgvector",10,[502,557,558,456,559],"Postgres","HNSW","Quantisation",[554,561,562,563,564,565,566],"pgvector vs qdrant","postgres vector search","hnsw index postgres","iterative index scans","binary quantization postgres","vector database postgres",[568,570],{"name":554,"url":569},"https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector",{"name":571,"url":572},"PostgreSQL","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPostgreSQL",[574,576,579,582,585],{"title":575,"url":569},"pgvector README",{"title":577,"url":578},"pgvector changelog","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FCHANGELOG.md",{"title":580,"url":581},"PostgreSQL news: pgvector 0.8.2 released","https:\u002F\u002Fwww.postgresql.org\u002Fabout\u002Fnews\u002Fpgvector-082-released-3245\u002F",{"title":583,"url":584},"AWS: Scale pgvector with binary quantization","https:\u002F\u002Faws.amazon.com\u002Fblogs\u002Fdatabase\u002Fscale-pgvector-with-binary-quantization-on-amazon-aurora-postgresql\u002F",{"title":586,"url":587},"pgvector licence","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FLICENSE","\u002Fimages\u002Fblog\u002Fpgvector\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fpgvector\u002Fog.jpg",[50,51,52],"pgvector, reviewed: the vector database you do not have to run","A review of pgvector 0.8.7: iterative scans for filtered search, HNSW and IVFFlat, binary quantisation at 100M vectors, and the CVE that made index builds a patch item.","A query enters at the top and splits into an exact sequential scan, an HNSW graph walk and an IVFFlat probe; a band below shows an iterative scan continuing until the limit is full.","PostgreSQL licence","Vector database extension",{"slug":597,"published":598,"minutes":555,"category":7,"tags":599,"keywords":601,"about":607,"sources":611,"cover":639,"og":640,"expertise":48,"locales":641,"lang":50,"title":642,"description":643,"coverAlt":644,"url":610,"pricing":645,"kind":452},"mem0","2026-09-17",[452,600,456,502],"Long-term memory",[597,602,603,604,605,463,606],"mem0 review","agent memory layer","mem0 self-hosted","mem0 pricing","mem0 alternatives",[608],{"name":609,"url":610},"Mem0","https:\u002F\u002Fmem0.ai",[612,615,618,621,624,627,630,633,636],{"title":613,"url":614},"Mem0 documentation","https:\u002F\u002Fdocs.mem0.ai\u002Fintroduction",{"title":616,"url":617},"Mem0 quickstart","https:\u002F\u002Fdocs.mem0.ai\u002Fquickstart",{"title":619,"url":620},"How Mem0 works","https:\u002F\u002Fdocs.mem0.ai\u002Fcore-concepts\u002Fhow-it-works",{"title":622,"url":623},"Mem0 pricing","https:\u002F\u002Fmem0.ai\u002Fpricing",{"title":625,"url":626},"Mem0 on GitHub","https:\u002F\u002Fgithub.com\u002Fmem0ai\u002Fmem0",{"title":628,"url":629},"mem0ai on PyPI","https:\u002F\u002Fpypi.org\u002Fproject\u002Fmem0ai\u002F",{"title":631,"url":632},"Mem0 research and benchmarks","https:\u002F\u002Fmem0.ai\u002Fresearch",{"title":634,"url":635},"Mem0 MCP server","https:\u002F\u002Fdocs.mem0.ai\u002Fplatform\u002Fmem0-mcp",{"title":637,"url":638},"Mem0 paper on arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2504.19413","\u002Fimages\u002Fblog\u002Fmem0\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmem0\u002Fog.jpg",[50,51,52],"Mem0: what an agent memory layer costs per turn","A review of Mem0: facts extracted from every turn, the April 2026 benchmark table and its platform-only caveat, four cloud tiers and what self-hosting leaves out.","A loop that turns conversation into stored facts and reads them back into the prompt","Free tier · from $19 per month",1791383549051]