Tools/AI agents

LlamaIndex review: the widest data toolkit, with its centre of gravity already moved

LlamaIndex in 2026: an MIT-licensed Python data and agent framework with 300+ integrations, event-driven Workflows, and a company that has moved its focus to LlamaParse.

Type
Agent framework
Pricing
MIT · hosted platform paid

··10 min read

  • Agent framework
  • RAG
  • Python
  • Document parsing
  • Workflows
Diagram: an event-driven workflow with typed steps, parallel workers, a fan-in step and a checkpoint after a restart.

Key takeaways

  • LlamaIndex is still the widest open-source data toolkit for LLM applications: more than 300 integration packages, one thin adapter each, on top of a small set of core abstractions.
  • Workflows, its orchestration layer, is event-driven: a step takes an event and returns an event, the graph is derived from type annotations, and it is validated before the run starts.
  • Durability is opt-in and hand-written. A run is ephemeral by default; snapshots come from `Context.to_dict()` plus a loop the team writes, and resume is at-least-once, so steps must be safe to repeat.
  • The company's own README now states that the focus is document parsing and extraction. LlamaParse is the paid product, priced in credits, and LiteParse is the local open-source parser.
  • The verdict: use the open-source core, and specifically Workflows as a small library, but keep the framework's blast radius small, because the maintenance direction has plainly moved.

LlamaIndex is an MIT-licensed Python framework for wiring LLM applications to data. It began as a set of connectors and index abstractions for retrieval, and it is still the widest such toolkit: the repository publishes more than 300 integration packages covering model providers, embedding models, vector stores, document loaders and rerankers. That breadth is why most teams meet the project first, and why it stays useful even though the company's own attention has moved elsewhere. The position taken here is simple: use the open-source core, and be deliberate about which hosted parts get adopted.

It sits between the model provider and the application. Nothing in the core talks to a model provider directly; every model, embedder and vector store arrives through an integration package implementing a core abstract class. That is a clean boundary and also the source of the main complaint: the same abstraction has to cover OpenAI, a local Ollama process and a dozen embedding models, so the useful common denominator is narrower than the surface area suggests. Against LangGraph it competes for the same orchestration role, against LangChain it competes on data handling, and against a hand-assembled vector store plus reranker it competes on convenience rather than capability.

What it is

Package layout is the first thing to check, and it is where most confusion starts. llama-index-core holds the abstractions and ships Workflows with them. llama-index is the starter package that installs core plus a selection of integrations. Workflows also publishes on its own as llama-index-workflows, and when it arrives through core it is imported as llama_index.core.workflow. Everything else is a separate install, which is the only reason the dependency tree stays workable.

  • MIT licence. The current llama-index-core release is 0.14.25, published on 21 September 2026.
  • More than 300 integration packages on PyPI, each one a thin adapter over a core abstract class.
  • Workflows, the orchestration layer, is event-driven: a step receives an event and returns another, and the framework routes by type annotation rather than by a declared edge.
  • The event graph is validated before a run starts. Workflows where an event has no producer, a produced event has no consumer, or no terminal event is reachable are rejected.
  • Runs are ephemeral by default. Persistence is opt-in through Context.to_dict() snapshots, or through a runtime plugin such as DBOS that journals step transitions into a database.
  • Python is the practical target for the framework. Only LiteParse, the company's new local parser, ships bindings beyond Python.

How Workflows works

Workflows is the part worth understanding, because it is also the part that outlives the framework's RAG reputation. A workflow is a subclass of Workflow whose methods are decorated with @step. Each step accepts one event type and returns another; returning the start event's type begins a run, returning StopEvent ends it. Branches are ordinary if statements that return different event types, and a loop is a step that returns an event type handled earlier in the graph. Concurrency is a step that returns list[Event], paired with another step that accepts list[Event] and acts as the fan-in.

Workflows: typed events move the run forwardA start event triggers a retrieval step that fans out over eight parallel workers, a reranking step, and a terminal stop event. Below, a checkpoint step serialises the context to JSON after each finished step, and a resume path replays only the unfinished work.Workflows: events move the run forwardllama-index-workflowsONE RUNquestionStartEventretrievetop_k = 8rerankcross-encoderanswerStopEventONE STEP, N WORKERSread #1WorkItemread #2WorkItemread #n8 in flightcollectlist[Done]AFTER A RESTARTcheckpointctx.to_dict() JSONresume from snapshotdone steps skipped, at-least-once
The graph is derived from the step signatures and validated before the first event is dispatched. Persistence is not part of the runtime: it is a loop the team writes around it.
  • Return list[Event] when a step has a finite batch and can produce every work item before downstream workers start.
  • Accept list[Event] when the step needs the whole batch of results before it can continue.
  • Call ctx.send_event(...) when the number of events is unknown in advance, or when an event has to be dispatched from outside a step.
  • Use ctx.store for shared per-run state, and Resource(...) for clients, models and configuration that must not be serialised into a snapshot.

The type annotations are load-bearing rather than decorative. Before execution the framework derives the event graph from the step signatures and refuses to start a workflow with an unproduced event, an unconsumed event or no reachable StopEvent. That catches a class of wiring bug that a hand-written function composition does not catch at all, and it is the strongest argument for the design. The price is that a deliberately dynamic workflow has to fall back on ctx.send_event and declare which static checks to skip, which in practice means a workflow becomes progressively less checkable as it becomes more capable.

Getting started

The smallest useful workflow is about twenty lines. Two install shapes are supported and they differ in the import path, which is the detail that catches people:

import asyncio

from llama_index.core import VectorStoreIndex
from workflows import Workflow, step
from workflows.events import Event, StartEvent, StopEvent


class Answered(Event):
    question: str
    answer: str


class TriageFlow(Workflow):
    index: VectorStoreIndex

    @step
    async def retrieve(self, ev: StartEvent) -> Answered:
        retriever = self.index.as_retriever(similarity_top_k=8)
        nodes = await retriever.aretrieve(ev.question)
        best = max(nodes, key=lambda n: n.score or 0.0)
        return Answered(question=ev.question, answer=best.node.get_content())

    @step
    async def answer(self, ev: Answered) -> StopEvent:
        return StopEvent(result=ev.answer)


async def main():
    flow = TriageFlow(index=VectorStoreIndex.from_documents(documents), timeout=60)
    result = await flow.run(question="What is the refund window?")
    print(result.result)

asyncio.run(main())

run() returns a WorkflowHandler. It is awaitable, and keeping the handler instead of awaiting it inline also gives access to stream_events() for progress reporting. The timeout argument is in seconds and is set on the constructor. Every step is async by design, so a standalone script needs a single asyncio.run entry point; inside FastAPI or a notebook it is not necessary.

Durability and retries

Workflows are ephemeral by default: once run() returns, the state is gone and the next run starts from nothing. For a fan-out over hundreds of documents that should not restart from zero, the documented mechanism is a checkpoint loop. Context.to_dict() serialises the in-flight events and the state store, Context.from_dict() rebuilds them, and run(ctx=...) continues in a different process. There is no built-in checkpointer to switch on: the run emits an internal StepStateChanged event when a step finishes, and that is the signal to snapshot.

import json

from workflows.events import StepState, StepStateChanged

handler = flow.run(question=q)

async for ev in handler.stream_events(expose_internal=True):
    if isinstance(ev, StepStateChanged) and ev.step_state == StepState.NOT_RUNNING:
        json.dump(handler.ctx.to_dict(), open("run.json", "w"))

result = await handler

# after a restart: resume from the last snapshot
ctx = Context.from_dict(flow, json.load(open("run.json")))
result = await flow.run(ctx=ctx)

Two properties of this design surface in operations. Resume is at-least-once: a step that was mid-execution when the snapshot was taken is rewound and runs again, so up to num_workers items are duplicated. And the snapshot is JSON, so a value the serialiser cannot encode makes to_dict() raise and the whole snapshot fail, not just the offending field. Heavy inputs therefore belong in a Resource, which is re-created on resume instead of being serialised. Teams that would rather not own the loop can use the DBOS runtime plugin, which journals step transitions into a database and needs no checkpoint code.

Where it falls short

The weaknesses are structural rather than bugs. Breadth is a maintenance surface: with hundreds of integrations, an upgrade can move the model wrapper you did not mean to touch, and pinning one integration often means pinning core. The high-level API hides enough that a prototype can reach production without anyone recording which chunk size, which top-k and which prompt template produced the answer; the defaults are convenient and are not documented as defaults. Durability is opt-in and hand-written, which is the opposite of what a long-running batch job wants. And the company behind the framework has publicly narrowed its focus, which is awkward to evaluate in 2026.

FrameworkOrchestration modelDurability and stateWhere it wins
LlamaIndex WorkflowsEvent-driven steps, control flow in plain PythonNo built-in checkpointer; context snapshots written by the team, or the DBOS runtime pluginRetrieval and document components live in the same library
LangGraphExplicit state graph with conditional edgesDurable execution and checkpointers are the defaultState transitions are the artefact, and stay inspectable as a graph
HaystackPipelines of typed componentsPer-component error branches and retriesA stable component catalogue for retrieval-heavy applications
DSPyDeclarative modules, optimised offline against a metricNone; it is not an orchestration runtimePrompt and module tuning, before any serving code exists

Read the durability column across and the practical split is legible. LangGraph makes checkpointing the default and buys that by making the graph explicit, which costs cognition and pays in legibility. Workflows keeps control in plain Python, which reads better and inspects worse: a workflow that fans out over 500 documents is 500 concurrent calls whose ordering no tool is helping you see. For a team whose work is to reason about state transitions, that is the wrong trade. For a team that wants to write ordinary Python and have it behave, it is the right one.

A second caveat belongs here. A dynamic workflow loses static analysis, and the documentation is explicit that unreachable steps and one-off events are exactly what those checks cannot see. Migrating a hand-drawn graph into Workflows tends to reach for skip_graph_checks early, which quietly disables the check that catches dead branches, which is the same check that would have caught the wiring mistake in the first place.

The pivot to document processing

The repository README now carries an unmissable note: the current focus of LlamaIndex is document parsing and extraction, and LlamaParse is the company's enterprise platform for it. That changes what a maintenance budget is funding. The integration packages still ship and still work; what is no longer guaranteed is that each of them is extended when a new provider appears.

  • LlamaParse is the paid product: agentic OCR, parsing, extraction and indexing, sold in credits rather than in seats.
  • LiteParse is the open-source counterweight: a Rust parser that runs locally with no LLM, no cloud dependency and no API key, with bindings for TypeScript, Python, Rust and browser WASM.
  • ParseBench and ExtractBench are the company's own public benchmarks for parsing and extraction.
  • The framework stays MIT-licensed and published on PyPI, with llama-index-core at 0.14.25 as of 21 September 2026.
  • The result is a split strategy: open tooling for local parsing, a hosted platform for hard documents, and the agent framework in between as the integration surface.

That is a defensible business decision and a mild warning for anyone choosing a framework now. The parts of the stack most likely to be maintained are the parts behind the paywall, and the MIT core is what remains useful as a library. Read it as an argument for keeping the framework's blast radius small: use Workflows for orchestration, keep loading and parsing in house where that is possible, and be explicit about which calls leave your infrastructure.

Verdict

The verdict is that LlamaIndex remains the most complete answer to a narrow question, how to get documents into an LLM application without writing the integration layer yourself, and a mediocre answer to the broader question of how to orchestrate an agent. Workflows is genuinely good at turning branching logic into typed, validated Python, and its at-least-once checkpoint model is honest about its own failure semantics. What the framework does not offer is a runtime to point at and watch: no state graph to inspect, no persistence on by default, and a maintenance direction that has plainly moved elsewhere.

  1. Use it when most of your data is documents and the connector, chunker, embedder and reranker abstractions are already written for you. That is a real saving, and it was the original purpose.
  2. Use Workflows on its own as a small library when the orchestration is branching, looping Python currently tangled inside asyncio. The typed event graph and the pre-run validation are the reason.
  3. Do not pick it for a graph that has to be reasoned about. If the state transitions are the artefact under scrutiny, an explicit graph beats Python control flow.
  4. Do not pick it for the hosted platform. LlamaParse, Extract and the index service are a separate paid product with credit pricing, and tying your ingestion bill to a framework is a decision rather than a default.
  5. Avoid it for a TypeScript or Go stack. The core is Python; only the new LiteParse parser ships broader bindings.
  6. Re-evaluate in a year if the integration breadth is load-bearing. With the company's focus on parsing and extraction, the long tail of vector store and embedder integrations is the part most exposed to slow maintenance.
The current focus of LlamaIndex is to build the best AI-powered engine for document parsing and extraction.

That sentence sits in the project README rather than in a blog post, which is unusual and worth noticing: the framework's own maintainers are the ones stating where the roadmap is going. Read as engineering, it says the orchestration layer is maintained but is not the destination, and the paid document pipeline is.

Sources

  1. LlamaIndex repository README and focus note
  2. Agent Workflows: introduction
  3. Agent Workflows: writing durable workflows
  4. Agent Workflows: observability
  5. LlamaAgents overview
  6. llama-index-core on PyPI
  7. LlamaIndex and LlamaParse pricing
  8. LiteParse documentation
  9. LangGraph on GitHub

Frequently asked questions

Is LlamaIndex still worth using in 2026?

Yes, with a narrowed scope. The MIT-licensed core is maintained, `llama-index-core` was at version 0.14.25 on 21 September 2026, and the Workflows orchestration layer is genuinely well designed for branching and looping logic. What has changed is the company's stated focus, which the README puts on document parsing and extraction rather than on the framework.

What are LlamaIndex Workflows and how do they work?

A workflow is a subclass of `Workflow` with methods decorated by `@step`. Each step accepts one event type and returns another, and the framework routes by type annotation, so branches are ordinary `if` statements and loops are steps returning an event handled earlier. Concurrency is a step returning `list[Event]` paired with a step accepting `list[Event]`. The event graph is validated before the run starts.

How does LlamaIndex handle durability and retries?

It does not, by default. Once `run()` returns, the state is gone. The documented mechanism is a checkpoint loop: serialise the context with `Context.to_dict()` when a `StepStateChanged` event marks a step as finished, then rebuild it with `Context.from_dict()` and call `run(ctx=...)`. Resume is at-least-once, so a step in flight when the snapshot was taken runs again. The DBOS runtime plugin is the alternative that removes the loop.

How much does LlamaParse cost?

The pricing page sells credits rather than seats: 1,000 credits at $1.25, 10,000 credits free, 40,000 on the $50 starter tier and 400,000 on the $500 pro tier, with pay-as-you-go caps of $500 and $5,000 a month. Concurrent parse jobs are 5 on free and starter, 20 on pro and 100 on enterprise. Parse, extract, classify and index all draw on the same balance.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.