> What Microsoft GraphRAG and LightRAG really do, what indexing costs, and when a knowledge graph beats vector RAG: multi-hop, global questions, product catalogues.
>
> Web page: https://balazscsorba.com/blog/graphrag-knowledge-graph-rag · Language: English · Also available in: [Deutsch](https://balazscsorba.com/de/blog/graphrag-knowledge-graph-rag.md) · [Magyar](https://balazscsorba.com/hu/blog/graphrag-knowledge-graph-rag.md)
> Author: Balázs Csorba · Published: 2026-10-02 · Keywords: GraphRAG, knowledge graph RAG, GraphRAG vs vector RAG, Microsoft GraphRAG explained, LightRAG vs GraphRAG, GraphRAG global vs local search, GraphRAG indexing cost, when to use GraphRAG, multi-hop RAG, LazyGraphRAG

[Blog](https://balazscsorba.com/blog)/RAG & retrieval

# GraphRAG and knowledge-graph RAG: when a graph beats vector search

What Microsoft GraphRAG and LightRAG really do, what indexing costs, and when a knowledge graph beats vector RAG: multi-hop, global questions, product catalogues.

[Balázs Csorba](https://balazscsorba.com/about)·October 2, 2026·13 min read

-   GraphRAG
-   Knowledge graphs
-   LightRAG
-   RAG

![Diagram: a knowledge graph hub linked to entities, communities, local search, global search and product parts.](https://balazscsorba.com/images/blog/graphrag-knowledge-graph-rag/cover.webp?v=28965c6dfa)

## Key takeaways

-   GraphRAG adds an LLM-built knowledge graph, Leiden communities and community reports on top of chunking. Its headline strength is global questions about a whole corpus, not better lookup of single facts.
-   Indexing is the price: at least one LLM call per chunk for extraction plus summaries for every entity and community. Microsoft itself warns that GraphRAG can consume a lot of LLM resources.
-   Local search (entity neighbourhoods) helps with relation and multi-hop questions, global search (map-reduce over community reports) with themes and overviews. For plain fact lookup, vector RAG with reranking is usually enough.
-   Lighter variants exist: LightRAG for incremental updates, LazyGraphRAG with vector-RAG indexing cost, HippoRAG for cheap multi-hop retrieval. Treat the claims as paper results until you replicate them on your own data.
-   In B2B catalogues the relations (replaced by, part of, compatible with) already live in your PIM or ERP. Load them as a graph instead of paying an LLM to rediscover them, and use vector search for the free text.

On this page

1.  [What Microsoft GraphRAG actually does](https://balazscsorba.com/#what-graphrag-does)
2.  [Global, local and DRIFT search](https://balazscsorba.com/#global-local-drift)
3.  [LightRAG, LazyGraphRAG, HippoRAG and friends](https://balazscsorba.com/#lightrag-and-friends)
4.  [The indexing bill](https://balazscsorba.com/#indexing-cost)
5.  [When graphs win, and when they do not](https://balazscsorba.com/#when-graphs-win)
6.  [A pragmatic B2B example: products and parts](https://balazscsorba.com/#b2b-product-parts)
7.  [A checklist before you build a graph](https://balazscsorba.com/#checklist)
8.  [Where this is going](https://balazscsorba.com/#where-this-is-going)
9.  [Sources](https://balazscsorba.com/#sources)

Every few months someone shows a beautiful knowledge-graph visualisation and claims that vector RAG is obsolete. I have built enough retrieval systems to distrust both halves of that sentence. Graphs genuinely solve problems that chunk similarity cannot, and they also cost real money and add real complexity for problems that a hybrid search with a reranker already handles.

This article explains what Microsoft's GraphRAG actually does under the hood, what LightRAG, LazyGraphRAG and HippoRAG change, where the indexing bill comes from, and which kinds of questions justify a graph. It ends with a pragmatic B2B example from catalogue data, where my advice is deliberately boring: use the graph you already have.

If you have not built a solid baseline yet, start with [chunking, hybrid search and reranking](https://balazscsorba.com/blog/rag-pipeline-chunking-hybrid-search-reranking) and the [overview of RAG in 2026](https://balazscsorba.com/blog/rag-2026-hybrid-agentic-long-context). A graph is an add-on to a working pipeline, not a replacement for one.

## What Microsoft GraphRAG actually does

The research behind it is the paper [From Local to Global: A Graph RAG Approach to Query-Focused Summarization](https://arxiv.org/abs/2404.16130) by Darren Edge and colleagues at Microsoft Research (April 2024). Its starting point is a limitation of ordinary RAG: it struggles with global questions about an entire corpus, such as "What are the main themes in the dataset?". Similarity search returns the chunks closest to the question, and a question about everything has no closest chunk.

The open-source implementation documents the indexing pipeline in stages. Documents are cut into text units (1,200 tokens by default). An LLM extracts entities with a title, type and description plus the relationships between them, and optionally claims. Repeated descriptions are consolidated. The hierarchical Leiden algorithm then clusters the graph into communities at several levels of granularity. Finally the LLM writes a report for every community, and text units, entity descriptions and reports are embedded into a vector store.

The expensive part happens once at indexing time; the two search modes read different artefacts of the same index.

Two things follow from that design. The graph is not a hand-modelled ontology but whatever the extraction prompt found, so its quality depends on prompt tuning for your domain (the documentation says so explicitly). And the community reports are a pre-computed, hierarchical summary of your corpus, which is exactly what makes global questions answerable.

In the paper's experiments, on two datasets of roughly 1 million tokens each (podcast transcripts and news articles), GraphRAG variants beat a vector RAG baseline in LLM-judged comparisons: 72 to 83 per cent win rates on comprehensiveness and 62 to 82 per cent on diversity, depending on dataset and variant. Root-level community summaries also needed 9 to 43 times fewer tokens than summarising the source text. The authors are careful about scope: the evaluation covers sensemaking questions on two corpora, and they say more work is needed to see how it generalises.

## Global, local and DRIFT search

The distinction between the query modes is the most useful thing to understand, because it tells you which questions justify the index.

Mode

Question it answers

What it reads

Cost profile

**Local search**

About specific things: "What do we know about customer X and its contracts?"

Entities matching the question, their neighbours, relationships, source text units and community reports

One retrieval and one answer, comparable to RAG with a bigger context

**Global search**

About the whole corpus: "What are the main risks across all reports?"

Community reports, processed by a map step and a reduce step

Many LLM calls; a lower community level is more thorough but slower and costlier

**DRIFT search**

Broad start, specific follow-up

Relevant community reports first, then local search on the follow-up questions

Between the two; documented as more comprehensive than plain local search

**Vector RAG**

Where is this stated?

The top-k most similar chunks

Cheapest at index and query time

Global search deserves a closer look. The documentation describes a map-reduce over community reports: reports are split into chunks, each chunk yields an intermediate answer with importance ratings, and the reduce step filters and aggregates them. The choice of community level is a direct dial between depth and cost, so a production system should expose it rather than hard-code it.

Local search is the mode most teams actually need, and it is also the one closest to what a classic retrieval pipeline does. It embeds the question, finds related entities, and pulls in connected entities, relationships, covariates, source chunks and community reports, trimmed to one context window.

## LightRAG, LazyGraphRAG, HippoRAG and friends

The original design is thorough and expensive, and several projects attack exactly that. A short orientation, with the usual caveat that all numbers below come from the authors themselves:

-   **[LightRAG](https://github.com/HKUDS/LightRAG)** (paper from October 2024, presented at EMNLP 2025, MIT licence) builds a graph plus vector index with dual-level retrieval, supports incremental updates and document deletion, and offers local, global, hybrid, naive and mix query modes. It can run on PostgreSQL, Neo4j, MongoDB, Milvus, Qdrant or OpenSearch. The paper reports improvements in retrieval accuracy and efficiency, but I could not verify a like-for-like cost comparison with GraphRAG, so I leave that claim open. Incremental updates are the feature I care about: a rebuilt-from-scratch index is a poor fit for living document sets.
-   **[LazyGraphRAG](https://www.microsoft.com/en-us/research/blog/lazygraphrag-setting-a-new-standard-for-quality-and-cost/)** (Microsoft Research, 25 November 2024) skips the up-front summarisation. Microsoft states that its indexing costs are identical to vector RAG and 0.1% of the costs of full GraphRAG, shifting work to query time. That is the right trade when you have a large corpus and few graph-worthy questions.
-   **[HippoRAG](https://arxiv.org/abs/2405.14831)** (NeurIPS 2024) combines an LLM-built knowledge graph with Personalized PageRank. The authors report gains of up to 20% on multi-hop question answering, and single-step retrieval that matches or beats iterative retrieval while being 10 to 30 times cheaper and 6 to 13 times faster. It targets multi-hop, not global questions.

My reading: "graph RAG" is a family, not a product. The useful question is which of the three jobs you need: global summaries, relation-aware multi-hop retrieval, or incrementally updated knowledge. Pick the variant for that job.

## The indexing bill

Microsoft's own getting-started page warns that GraphRAG can consume a lot of LLM resources and recommends trying the tutorial dataset and cheaper models first. That is the honest summary. I cannot give you a price per million tokens that stays true, because it depends on the model and your prompts, but I can show where the calls come from:

-   **Extraction:** at least one LLM call per text unit, more with self-reflection ("gleaning") passes. The paper found that extraction recall improves with extra passes and smaller chunks: with GPT-4, a 600-token chunk yielded almost twice as many entity references as a 2,400-token chunk.
-   **Description summarisation:** every entity and relationship that appears in many chunks gets its descriptions merged by an LLM.
-   **Community reports:** one LLM-written report per community, at every hierarchy level.
-   **Embeddings:** text units, entity descriptions and report content. This is the only part a vector RAG pipeline also pays.

Scale matters too. The paper's graphs for roughly 1 million tokens of text had 8,564 nodes and 20,691 edges (podcasts) and 15,754 nodes and 19,520 edges (news). Multiply that by your corpus and by every re-index. Re-indexing is where costs compound: if documents change weekly, an incremental design such as LightRAG or an on-demand design such as LazyGraphRAG changes the economics more than any prompt tuning.

**Estimate before you index**

Index a representative 1 to 5 per cent sample, record the tokens and cost per text unit, and extrapolate. Then add the cost of re-indexing at your real change rate. Compare the total with the value of the questions only the graph can answer. See [cost, latency and routing](https://balazscsorba.com/blog/llm-cost-latency-prompt-caching-routing) for the general method.

## When graphs win, and when they do not

The [GraphRAG-Bench study](https://arxiv.org/abs/2506.05690) (June 2025) starts from an uncomfortable observation: GraphRAG frequently underperforms vanilla RAG on many real-world tasks. It then evaluates fact retrieval, complex reasoning, summarisation and creative generation to find the conditions in which the graph pays off. That matches my experience. The decision depends on the shape of the question.

Question type

Vector RAG + reranker

Graph RAG

My call

Single-fact lookup ("What is the torque for model Z?")

Strong, cheap

Rarely better

Vector

Multi-hop ("Which supplier makes the part that replaced X?")

Misses the second hop unless an agent iterates

Strong when the relation is explicit

Graph, or agentic retrieval

Global ("What themes recur in 5,000 tickets?")

Weak: no closest chunk

The case GraphRAG was designed for

Graph, or summarise by clustering

Relational catalogue ("What fits, what replaces, what is part of?")

Finds text, not structure

Strong, and the data is already structured

Graph from structured data

Fast-changing documents

Easy to update

Costly unless incremental

Vector, or LightRAG-style updates

Small corpus that fits in context

Not needed

Not worth it

Long context

The pattern: graphs help when the answer is assembled from several pieces connected by relations, or from the corpus as a whole. They do not help when the answer sits in one passage. Before building one, write down twenty real user questions and label each as fact, multi-hop or global. If 80 per cent are facts, you need better chunking and reranking, not a graph. Measure the result with [evals tied to your product](https://balazscsorba.com/blog/llm-evals-for-product-features), not with a demo.

## A pragmatic B2B example: products and parts

Here is the situation I meet in B2B e-commerce and ERP projects. A manufacturer or wholesaler sells machines, spare parts and accessories. The sales team or a customer asks: "Our pump P-150 is discontinued. Which pump replaces it, and which seal kit do I need now?" The answer needs three hops: the successor product, its parts list, and the current replacement of a discontinued part. Related questions are "what is compatible with this flange?" and "which manual covers this variant?".

A vector search over datasheets retrieves the P-150 page and perhaps the P-200 page. It does not reliably follow "replaced by" to the right seal kit, because that fact is a relation between records, not a sentence similar to the question. This is a classic graph case. But notice where the graph should come from.

The relations are already fields in a PIM or ERP; only the manual needs text retrieval.

In a PIM or ERP such as Pimcore, Spryker or SAP, these relations already exist as structured data: successor links, bills of materials, compatibility tables. Paying an LLM to extract them again from PDFs is slower, costlier and less accurate than reading the fields. My recommended architecture is therefore:

1.  **Build the graph from structured data.** Nodes are products, parts and documents; edges are the relation types your master data already defines (replaced by, part of, compatible with, documented in). Keep the node identifiers equal to the SKU or article number.
2.  **Use vector or hybrid search for the text.** Manuals, datasheets and tickets stay in a chunked index. Each chunk carries the product IDs it mentions, so a graph traversal can fetch exactly the relevant chunks.
3.  **Resolve entities first, then traverse.** The agent or retrieval step maps "P-150" to a node (exact match beats embeddings for article numbers), walks one to three hops with a typed query, and hands the resulting records plus the linked chunks to the model.
4.  **Add LLM extraction only for the gaps,** for example compatibility notes that exist only in free text, and flag such edges as lower-confidence.
5.  **Check availability from the system of record.** Stock, price and validity belong to the ERP, called as a tool at answer time, never stored in the graph. See [agentic commerce protocols](https://balazscsorba.com/blog/agentic-commerce-protocols-ucp-acp-guide) for where this is heading.

This is graph RAG without the expensive part. It is also more auditable: every hop in the answer is a record someone can open. If you build such systems, my [B2B e-commerce work](https://balazscsorba.com/expertise/b2b-ecommerce-developer) is exactly this combination of master data and retrieval.

## A checklist before you build a graph

1.  Collect and label twenty to fifty real questions as fact, multi-hop or global.
2.  Build and measure the baseline: hybrid search with reranking and decent chunking.
3.  Check whether the relations already exist as structured data. If yes, import them; do not extract them.
4.  If you need global questions, pilot GraphRAG or LazyGraphRAG on a sample and extrapolate the cost, including re-indexing.
5.  If your documents change often, test incremental approaches such as LightRAG first.
6.  Tune the extraction prompt for your domain and inspect a sample of extracted entities by hand.
7.  Compare against the baseline on your own questions with an LLM judge plus human spot checks.
8.  Expose the retrieval mode (vector, local, global) as a router decision rather than forcing every query through the graph.

The last point is the one that saves money. A cheap classifier or a small model routes fact questions to vector search and sends only global or relational questions to the graph. The pipeline then costs what each question deserves.

## Where this is going

With long context windows and agentic retrieval, an agent can iterate over a plain index and approximate multi-hop reasoning, at the price of more calls per question. Graphs move that work to indexing time. Neither wins everywhere, which is why the 2026 pattern is a router over several retrieval strategies.

My recommendation is to be sceptical and stay empirical. Start with the baseline, add a graph only for the question types that need it, source its edges from structured data wherever you can, and keep the indexing bill visible. A graph that answers one class of question better, at a known price, is an asset. A graph built because it looks good in a diagram is a cost.

## Sources

1.  [Edge et al.: From Local to Global: A Graph RAG Approach to Query-Focused Summarization (arXiv:2404.16130)](https://arxiv.org/abs/2404.16130)
2.  [Microsoft GraphRAG documentation: overview](https://microsoft.github.io/graphrag/)
3.  [Microsoft GraphRAG documentation: default dataflow](https://microsoft.github.io/graphrag/index/default_dataflow/)
4.  [Microsoft GraphRAG documentation: global search](https://microsoft.github.io/graphrag/query/global_search/)
5.  [Microsoft GraphRAG documentation: local search](https://microsoft.github.io/graphrag/query/local_search/)
6.  [Microsoft GraphRAG documentation: DRIFT search](https://microsoft.github.io/graphrag/query/drift_search/)
7.  [Microsoft GraphRAG documentation: getting started](https://microsoft.github.io/graphrag/get_started/)
8.  [Microsoft Research: LazyGraphRAG, setting a new standard for quality and cost (25 November 2024)](https://www.microsoft.com/en-us/research/blog/lazygraphrag-setting-a-new-standard-for-quality-and-cost/)
9.  [Guo et al.: LightRAG: Simple and Fast Retrieval-Augmented Generation (arXiv:2410.05779)](https://arxiv.org/abs/2410.05779)
10.  [HKUDS/LightRAG on GitHub](https://github.com/HKUDS/LightRAG)
11.  [Gutiérrez et al.: HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models (arXiv:2405.14831)](https://arxiv.org/abs/2405.14831)
12.  [Xiang et al.: When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation (arXiv:2506.05690)](https://arxiv.org/abs/2506.05690)

## Frequently asked questions

What is GraphRAG and how is it different from normal RAG?

Normal RAG chunks documents, embeds the chunks and retrieves the most similar ones. GraphRAG, as published by Microsoft Research, first has an LLM extract entities and relationships from the chunks, clusters the resulting graph into communities with the Leiden algorithm, and writes an LLM summary per community. Queries can then use the graph neighbourhood of an entity (local search) or the community summaries (global search) instead of only similar chunks.

What is the difference between GraphRAG local search and global search?

Local search starts from entities that match the question and pulls in their connected entities, relationships, source text and community reports. It suits questions about specific things. Global search runs a map-reduce over community reports and suits questions about the whole dataset, such as the main themes. DRIFT search combines both: it starts from community reports and refines with local search.

How expensive is GraphRAG indexing?

It is much more expensive than embedding chunks, because an LLM reads every chunk to extract entities and relationships, then writes descriptions and community reports. Microsoft states that GraphRAG can consume a lot of LLM resources and recommends starting small with cheaper models. LazyGraphRAG, a later variant, reports indexing costs identical to vector RAG and 0.1% of full GraphRAG.

When is GraphRAG better than vector RAG?

When questions depend on relationships or on the corpus as a whole: multi-hop questions, corpus-wide themes and overviews, and data with explicit relations such as product and part hierarchies. For simple fact lookup it often is not. The GraphRAG-Bench study notes that GraphRAG frequently underperforms vanilla RAG on many real-world tasks, so measure on your own questions.

Is LightRAG a good alternative to Microsoft GraphRAG?

It is a lighter, MIT-licensed graph RAG framework with dual-level (local and global) retrieval, incremental updates and several storage backends such as PostgreSQL and Neo4j. It is worth a pilot when your documents change often. Its published comparisons are the authors’ own, so test it against your baseline before committing.

Do I need a graph database for GraphRAG?

Not necessarily. Microsoft GraphRAG writes tables and embeddings, and LightRAG ships with in-memory graph storage for testing and supports PostgreSQL for production. A dedicated graph database such as Neo4j becomes useful when you traverse relations a lot or already model them, for example product-part structures.

Written by Balázs Csorba

Senior fullstack & AI engineer in Styria, Austria – 10+ years of Vue, Nuxt, Node.js and PHP, now building tooling for AI agents.

[AI engineering & MCP servers →](https://balazscsorba.com/expertise/ai-engineer)[About me →](https://balazscsorba.com/about)

## More articles

-   [Reducing LLM hallucinations in production: grounding, citations and knowing when to say no](https://balazscsorba.com/blog/llm-hallucination-grounding-citations)
-   [pgvector or a vector database? How to choose vector storage in 2026](https://balazscsorba.com/blog/pgvector-vs-vector-databases)
-   [Evaluating RAG: retrieval metrics, faithfulness and how to tell which half failed](https://balazscsorba.com/blog/rag-evaluation-metrics)
-   [Semantic product search for B2B shops: part numbers, hybrid retrieval and what to measure](https://balazscsorba.com/blog/semantic-product-search-b2b)

## Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.

[Book a call](mailto:contact@balazscsorba.com) [Connect on LinkedIn](https://www.linkedin.com/in/balazs-csorba)
