Tools/RAG & retrieval

Microsoft GraphRAG: a knowledge-graph RAG priced up front

A review of Microsoft GraphRAG 3.2.0: MIT, maintenance mode, and an indexing bill that is decided before the first query runs.

Type
RAG pipeline
Pricing
MIT · API pay per token

··10 min read

  • Knowledge graph
  • RAG
  • Global search
  • Community summaries
  • Token cost
Cover art for the GraphRAG review: a document corpus folded into a graph of nodes and community rings

Key takeaways

  • GraphRAG 3.2.0 was published on 23 September 2026 under MIT, while the repository states it is largely in maintenance mode with no new features.
  • The bill lands at indexing time: one model call per TextUnit and one per community report, reported at $50-200 for a 500-page corpus against under $5 for vectors.
  • Global Search is the mode vector search cannot imitate, and dynamic community selection cuts its token cost by 77% without a measured quality loss.
  • LazyGraphRAG reaches vector-RAG indexing cost at 0.1% of full GraphRAG, but ships in Microsoft Discovery and Azure Local rather than in the MIT package.
  • The graph pays only when questions are corpus-wide or multi-hop; lookup-heavy workloads stay cheaper on plain vector search.

GraphRAG is Microsoft Research’s MIT-licensed pipeline that turns a corpus into a knowledge graph, clusters it with the Leiden algorithm and writes an LLM summary for every community before a single question is asked. The position taken here is that it remains the most rigorously documented graph RAG available, and that the interesting part is no longer the engineering: the repository is in maintenance mode, and the question a team actually has to answer is whether the indexing bill can be justified against vector search.

It competes with plain vector RAG, with lighter graph pipelines such as LightRAG, and with the graph retrieval built into LlamaIndex and Neo4j tooling. The distinction matters because GraphRAG is not a service: it is a Python package that spends the reader’s own tokens building an index and then exposes four query modes as library functions. Nothing is billed or hosted by Microsoft, and the README states that the code is a demonstration rather than an officially supported offering.

What it is

The design is deliberately front-loaded. Documents are split into TextUnits, an LLM extracts entities and relationships from each unit, duplicate descriptions are merged and summarised, Leiden clustering assigns the graph to a hierarchy of communities, and a further LLM pass writes a report for every community at every level. Only then does a question run, and by then the expensive work has already been paid for.

  • MIT licensed, shipped as graphrag on PyPI, current release 3.2.0 of 23 September 2026, Python 3.11 to 3.13.
  • About 36,200 stars, 3,800 forks and 495 commits on GitHub, with 49 open issues and pull requests.
  • TextUnits default to 1,200 tokens: larger chunks index faster and extract less precisely.
  • Community detection is hierarchical Leiden and costs no model calls; summarising every community does.
  • Four query modes ship in the package — global, local, DRIFT and basic — plus dynamic community selection for global search.
  • The CLI names four indexing methods: standard, fast, standard-update and fast-update, with a separate update command for changed documents.
  • Every call goes to the reader’s own endpoint — Azure OpenAI, OpenAI or a compatible service — so cost is a function of model price and corpus size.

Two consequences follow. The index becomes an asset rather than a by-product: entity descriptions and community reports are readable artefacts that can be reviewed, shared and versioned like any other output. And the graph is only as good as the extraction pass, so a corpus whose entity types matter has to be tuned before the first full run rather than after it.

How it works

Indexing is a fixed sequence: chunk, extract, merge, cluster, summarise, embed. Every stage either costs one model call per unit of text or costs nothing, and that boundary is where the bill comes from — one extraction call per TextUnit, one summarisation call per merged entity or relationship, and one report call per community at every level of the hierarchy.

One GraphRAG index, then four query modesIndexing splits the corpus into TextUnits, spends one model call per unit on extraction, clusters the graph with Leiden and spends one model call per community on reports. A question then routes to one of four modes: global map-reduce over the reports, local search around entities, DRIFT with follow-up questions, or basic vector search.Where the tokens gothe index is the assetINDEXING — PAID ONCECorpusyour documentsTextUnits1,200 tokensExtractone call eachReportsLeiden, then LLMQUERY — PAID PER QUESTIONQuestionplain wordsGlobalmap-reduceLocalentity walkDRIFTfollow-upsBasicvector search
The expensive row is the first one: every query reuses an index that has already been paid for.

The four modes are the product surface. Global Search runs map-reduce over community reports and is the mode vector search cannot imitate, because no chunk contains a corpus-wide answer. Local Search walks the graph around named entities and costs roughly what vector retrieval costs with traversal on top. DRIFT starts from the most relevant reports, asks follow-up questions and answers each of them locally. Basic Search is the package’s own vector RAG, included so a team can measure what the graph buys on its own data.

Getting started

The entry point is a directory, an init command and two files. The sequence below creates a workspace, points it at a workable model, indexes a small input folder and asks a global question; settings.yaml is where the bill is set, so a first run should stay on a small corpus and an inexpensive model until the prompts have been tuned.

python -m venv .venv && source .venv/bin/activate
python -m pip install graphrag

mkdir ragtest && cd ragtest
graphrag init -r . -m gpt-4.1 -e text-embedding-3-large
# write GRAPHRAG_API_KEY into the .env that init created
# drop a few .txt or .md files into ./input, then index:
graphrag index -r . -m standard

# global is the default; this prunes reports before the map-reduce step
graphrag query "What are the top themes across these documents?" -r . \
  --dynamic-community-selection

graphrag query "Who is the main character?" -r . -m local
graphrag update -r . -m standard-update

Two habits keep the first bill small. Index a sample rather than the corpus, because the pipeline calls the model once per TextUnit and once per community level whatever the model costs; and run prompt-tune before scaling, since the documentation states that prompts used out of the box rarely produce the best results.

What indexing costs

There is no licence fee and no hosted service, so the entire cost is model calls. Extraction runs once per TextUnit and community reporting once per community at every level, which means the bill scales with corpus size and with the number of communities the Leiden hierarchy produces. The figures below come from one independent comparison published in March 2026 at GPT-4 pricing, not from a Microsoft price list:

Approach500-page indexTimeWhat you get
GraphRAG full pipeline$50-200about 45 minentity graph, community reports, global search
Vector RAGunder $5minuteschunks by similarity, no global queries
LightRAGabout $0.50about 3 minflat graph, weaker global queries
LazyGraphRAGstated at 0.1% of full GraphRAGnot statednot shipped in the MIT package

The counter-measure Microsoft Research published is LazyGraphRAG, which replaces LLM extraction with noun-phrase extraction and defers every model call to query time. Its indexing cost is stated as identical to vector RAG and 0.1% of full GraphRAG, and the same evaluation reports comparable Global Search quality at more than 700 times lower query cost, or better-than-Global-Search quality at 4% of its cost. The implementation ships in Microsoft Discovery and Azure Local rather than in the MIT package, so a team running the open-source pipeline cannot install it: the number describes a direction, not an option on the shelf.

What queries cost

Query cost is where the modes differ and where the index either earns or fails to earn the money already spent on it. Global Search is the expensive one because it reads community reports in batches and then reduces them; the other three read a fraction of the graph:

ModeWhat it readsCost shapeWhen to use it
Globalcommunity reports at one level or a pruned selectiongrows with the number of reportscorpus-wide synthesis
Localentity neighbourhood and its text unitsvector retrieval plus traversalentity and relationship questions
DRIFTtop reports, then follow-up questions answered locallybetween local and globalscoped questions that still need coverage
Basicembedded text unitsthe package’s own vector baselinesingle-hop factual lookups

Dynamic community selection is the published fix for the first row: a cheaper model rates each report from the root and prunes irrelevant branches before map-reduce. Microsoft Research measured an average 77% reduction in token cost against static level-1 search over 50 global questions, with about 1,500 reports falling to 470 and no statistically significant difference in quality; letting the rating continue to level 3 cost 34% more on average and won 58.8% on comprehensiveness and 60.0% on empowerment. These are vendor figures from one dataset, but the direction is not in dispute — stop paying for reports that cannot answer the question.

Where it falls short

The weaknesses are operational rather than algorithmic. The repository is in maintenance mode, so prompt formats, model behaviour and dependency drift are the reader’s problem, and the README calls the code a demonstration. Updating is its own command rather than a background job: documents change, entities merge, communities shift, and the update methods still extract changed text at model-call prices. The graph also carries the ontology, because entity and relationship types come from open-ended extraction, so a noisy corpus produces a noisy graph. Most importantly, the index is paid for whether or not questions arrive — the opposite of vector search’s pay-per-query profile.

ToolWhat it isWhere it runsWhat you pay
GraphRAGfull pipeline with community summariesPython package on your own keysmodel calls while indexing and querying
Vector RAGchunk embeddings and similarity searchany vector storeembedding calls, cheap queries
LightRAGflat graph with lighter extractionPython package on your own keysreported at about 1/100 of the indexing cost
Graphititemporal graph for agent memoryyour stack with a Neo4j instanceextraction per interaction

Read that table as a statement about the question each system is built for. GraphRAG answers corpus-wide and multi-hop questions that no single chunk contains; vector RAG answers lookup questions faster and more cheaply, which is why GraphRAG ships its own Basic mode rather than pretending the graph wins everywhere. LightRAG is the reasonable default when a flat graph captures most of the value at a fraction of the cost, and Graphiti solves agent memory rather than document retrieval. The position taken here is that most teams reach for the full pipeline because its benchmark is impressive, when their query mix is dominated by lookups — and the cheapest first step is to measure that mix before indexing anything.

Verdict

GraphRAG is the right tool for a corpus whose questions are genuinely global, and the wrong default for a search box. It is the best-documented graph RAG available, it is MIT, and its index is a reusable artefact with reports people can read; it is also in maintenance mode, priced up front and slower to change than the corpus it indexes. Adopt it with a measured query mix and a small first corpus, or do not adopt it.

  1. Choose it when a meaningful share of questions need synthesis across the whole corpus: themes, trends, comparisons over everything.
  2. Choose it when answers depend on hops between entities that never appear in the same chunk.
  3. Choose it when the index itself has value — reports that people read, share and audit — because that is what the upfront spend buys.
  4. Do not choose it for a lookup-heavy search box: vector search is cheaper per query, and GraphRAG’s own Basic mode is that same search.
  5. Do not treat it as a maintained dependency; the repository states maintenance mode and no new features.
  6. Before the full run, measure the query mix and index a sample with dynamic community selection enabled.

One further consideration is where the index lives. It is a batch artefact, so it belongs where batch artefacts belong: built by a pipeline, versioned, reviewed and replaced, rather than rebuilt from inside a request handler. The mechanics of what the graph contains — entities, TextUnits, communities — are covered in an earlier piece on knowledge-graph RAG; this review is about the implementation, its modes and its bill.

GraphRAG indexing can be an expensive operation, please read all of the documentation to understand the process and costs involved, and start small.

Sources

  1. GraphRAG on GitHub
  2. GraphRAG documentation
  3. From Local to Global: A Graph RAG Approach to Query-Focused Summarization
  4. GraphRAG: Improving global search via dynamic community selection
  5. LazyGraphRAG: Setting a new standard for quality and cost
  6. Graph RAG in 2026: What Actually Works in Production

Frequently asked questions

How much does GraphRAG cost to run?

The package is MIT and nothing is hosted, so the whole bill is model calls: one extraction call per TextUnit of 1,200 tokens, one summarisation call per merged entity or relationship, and one report call per community at every hierarchy level. One independent comparison puts a 500-page corpus at $50-200 and about 45 minutes through the full pipeline, against under $5 to embed the same corpus for vector search.

When does GraphRAG beat vector search?

When the answer has to be synthesised across the corpus or hop between entities that never share a chunk: themes, trends, comparisons over everything. Direct factual lookups map onto single chunks, so they gain nothing from the graph, and GraphRAG ships its own Basic mode — vector retrieval over the same embeddings — precisely for measuring that difference.

Is GraphRAG still maintained?

The README says the project is largely in maintenance mode: no new pull requests, no new features, bug fixes and dependency updates as appropriate. It also describes the code as a demonstration rather than an officially supported Microsoft offering. The latest release, 3.2.0, arrived on 23 September 2026, and 36,241 stars sit above 495 commits and 49 open issues and pull requests.

What does LazyGraphRAG change?

It drops the LLM from indexing and uses noun-phrase extraction instead, so indexing cost is stated as identical to vector RAG and 0.1% of full GraphRAG. Microsoft Research reports comparable Global Search quality at more than 700 times lower query cost, and better-than-Global-Search quality at 4% of its cost. The implementation lives in Microsoft Discovery and Azure Local, not in the MIT repository, so a self-hosted team cannot install it.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.