Tools/AI agents

CrewAI review: crews, flows and the token bill

CrewAI is the MIT-licensed Python framework for multi-agent systems: crews for autonomous collaboration, flows for controlled state. What a run costs in tokens, and where the design hurts.

Type
Agent framework
Pricing
MIT · enterprise paid

··10 min read

  • Multi-agent
  • Agent orchestration
  • Python
  • Flows
  • Token cost
Diagram of a CrewAI flow driving a crew of agents that call the language model, memory and tools.

Key takeaways

  • CrewAI 1.15.23 is MIT-licensed Python for 3.10 to 3.13, with no execution cap of its own and extras for tools, LiteLLM, mem0, Qdrant and Bedrock.
  • Crews delegate through model calls, so the cost of a run is set by topology: role prompts, a hierarchical manager, three guardrail retries and memory analysis all add calls.
  • Memory keeps records as LanceDB vectors and calls a model on save and on deep recall, on top of the embedding call.
  • Built-in tracing goes to CrewAI's own backend and asks for an account plus a first-run consent prompt; Langfuse, Phoenix, Braintrust and MLflow work over OpenTelemetry instead.
  • The framework is at its best when the flow holds the logic and one or two agents do the open-ended part.

CrewAI is an MIT-licensed Python framework for multi-agent systems, and in this category it is the one that is immediately usable: an agent gets a role, a goal and a backstory, a task gets an expected output, and a crew runs the two together in a few lines of code. The verdict is that the framework is well made, unusually well documented and quietly expensive. Every structural feature it adds on top of a plain prompt is implemented as an extra model call, and a production bill is made of exactly those calls.

It occupies the same layer as LangGraph, AutoGen and Pydantic AI and competes on ergonomics rather than on control. The vendor splits the offer in two: the library on PyPI, which is free and unmetered, and CrewAI AMP, a hosted control plane with a visual editor, tracing and governance that is sold by quote. Everything below is about the library, because that is the part you install.

What CrewAI is

The library first appeared on PyPI in December 2023 and is developed in the open by CrewAI Inc. Version 1.15.23 was published on 28 September 2026, requires Python 3.10 to 3.13, and the repository carries 59,413 stars and 8,655 forks. It installs as a single wheel of about 1.2 MB, with optional extras for tools, LiteLLM, mem0, Qdrant, Bedrock, Anthropic and A2A support, so the base install stays small.

  • Licence MIT, no run quota, and no runtime dependency on LangChain or any other agent framework.
  • Two layers: flows, which own state and execution order, and crews, groups of agents that execute tasks.
  • An agent is a prompt with attributes: role, goal and backstory, plus an optional tool list, an LLM and a memory scope.
  • Processes are sequential or hierarchical. The hierarchical one needs a manager LLM and delegates work as model calls.
  • Results can be text, JSON or a Pydantic model, which is the clean way to hand output back to application code.
  • MCP servers attach to agents through crewai-tools over stdio, SSE or streamable HTTP.

How a run works

A flow is a class of methods wired by decorators. @start() marks the entry points, @listen() binds a method to another method's result, @router() sends execution down one of several branches, and and_ and or_ combine conditions. The state object is a Pydantic model, so the shape of the data moving between steps is declared rather than inferred, and flow.plot() renders the graph. When several @start() methods are satisfied at once, they run in parallel.

Where the model calls goFour boxes in a row: a flow that owns state and order, a crew that owns tasks and process, an agent built from role, goal and backstory, and the language model behind the provider API. An arrow returns the crew output into the flow state. A second row holds memory, which saves and recalls, and tools such as MCP servers, APIs and code.Where the model calls goeach box is a prompt the framework assemblesFlowstate and orderCrewtasks and processAgentrole, goal, storyLLMAPICrewOutput flows back into stateMemorysave and recallToolsMCP, APIs, codeEvery hop is a model call: agent step, delegation, guardrail retry, memory analysis.
The pipeline is cheap and the framework is not: each box is one more prompt the model has to read, and the boxes below the line are pure overhead a plain script would not have.

Inside a flow step, a crew assembles the prompt: each agent's role, goal and backstory is prepended to every task, the task description is interpolated, and the output of one task becomes context for the next. An agent runs for up to 20 iterations by default, retries twice on error and summarises messages to stay inside the context window. That is the mechanism to watch. With the defaults on, a four-task crew is four long prompts before anything is delegated, and a crew with memory adds an embedding call per recall plus a model call to analyse and consolidate what it stores.

Getting started

The CLI is the quick way in. uv tool install crewai puts a crewai binary on the path, crewai create crew <name> scaffolds a JSON-first project with agents/*.jsonc and crew.jsonc, and crewai create flow <name> scaffolds a flow project. crewai install resolves dependencies through uv, crewai run executes the entry point, and crewai memory opens a terminal browser for the store. The flag --classic restores the older Python class plus YAML layout.

from crewai import Agent, Crew, Process, Task
from crewai.flow import Flow, listen, start
from pydantic import BaseModel


class ReportState(BaseModel):
    topic: str = "agent memory"
    brief: str = ""


class ReportFlow(Flow[ReportState]):
    @start()
    def pick_topic(self):
        self.state.topic = "agent memory"

    @listen(pick_topic)
    def research(self):
        analyst = Agent(
            role="Research analyst",
            goal=f"Collect verifiable facts about {self.state.topic}",
            backstory="You read primary sources and quote them.",
            llm="openai/gpt-4o-mini",
        )
        task = Task(
            description="Write a brief on {topic}.",
            expected_output="Five bullets, each with a source URL.",
            agent=analyst,
        )
        crew = Crew(agents=[analyst], tasks=[task], process=Process.sequential)
        self.state.brief = crew.kickoff(inputs={"topic": self.state.topic}).raw

ReportFlow().kickoff()

That runs at most two model calls on a normal path: one for the agent's reasoning and tool loop, one for the final answer. Add a second agent, a manager or memory and the count moves quickly, which is why crew.usage_metrics and flow.usage_metrics are the first two things to put behind a dashboard.

Memory and knowledge

Memory has been reworked into a single Memory class. Records are vectors in LanceDB under .crewai/memory, ranked by a composite of semantic similarity, recency with a 30-day half-life and an importance score that the model assigns when saving. Recall has two depths: shallow is a plain vector search at roughly 200 ms with no model call, deep analyses the query first and runs only when the query is longer than 200 characters. When memory is enabled, a crew extracts facts from every task output and recalls context before every task.

  • Storage is LanceDB by default and local, and the backend is a protocol, so another vector store can be plugged in.
  • The default embedder is OpenAI text-embedding-3-large at 3,072 dimensions and the default analysis model is gpt-4o-mini. Both are configurable.
  • Knowledge sources are a separate mechanism: they land in ChromaDB collections with a default relevance cut-off of 0.35 and three documents per query.
  • Saves run on a background thread and recall waits for them, so nothing is lost at the end of a run, but the analysis still costs tokens.

Memory is also the part to read before signing off on a data protection review. The documentation states plainly that record content is sent to the configured LLM for scope, category and importance analysis, so anything sensitive needs a local model on both sides, for the LLM and for the embedder.

Running it in production

The library moves quickly. Stable releases landed on 9, 16 and 28 September 2026 and dev pre-releases are published daily, so pinning is not optional for anything long-lived. The repository has 570 open issues, and the same distribution now carries both the open-source runtime and the client for the hosted platform, which is why a changelog entry can hold a SQLite connection fix next to a platform feature.

Observability

Built-in tracing deserves a second look. It is off by default and configured separately from telemetry, but the destination is CrewAI's own backend: it needs a free AMP account, an authenticated CLI, and on the first run the process asks whether the execution trace may be shared. Traces contain prompts, inputs and outputs, and the local buffer keeps up to 1,000 spans before the oldest are dropped. For self-hosted work the OpenTelemetry integrations are the better default, and the documentation ships first-class paths for Langfuse, Arize Phoenix, Braintrust, Datadog, MLflow, Opik, Patronus, Portkey, Weave and Galileo.

  • Set CREWAI_DISABLE_TELEMETRY=1 unless you have a reason to send anonymous usage data to the vendor.
  • First-run tracing asks for consent in the terminal and discards the buffer if nobody answers; crewai traces enable and crewai traces disable change that later.
  • kickoff_async() only wraps the synchronous run in a thread. akickoff() and akickoff_for_each() are the native async paths and the ones to use under load.
  • @persist writes flow state to a local SQLite database by default. restore_from_state_id forks a run from a stored snapshot, while kickoff(inputs={"id": ...}) resumes the original.

Guardrails

Task guardrails come in two forms. A Python callable receives the task output and returns a verdict, while a plain string is turned into an LLM guardrail that judges the output with the agent's own model. Retries default to three, so a guardrail that never passes can triple the cost of a task before anything is escalated. Guardrails validate output; they are not a boundary around tool use, and a prompt injection in a tool result passes straight through them.

Licence and cost

The licence is the easy part. CrewAI is MIT with no run quota, and the only cost the framework itself adds is the model traffic its own features generate. The paid product is CrewAI AMP: a free Basic plan with 50 workflow executions a month, and an Enterprise tier quoted case by case.

What you runPriceWhat is included
CrewAI OSSFreeNo execution cap. You pay for model calls, tools and your own infrastructure.
CrewAI AMP BasicFree50 workflow executions a month, visual editor, GitHub sync, tracing and OpenTelemetry.
CrewAI AMP EnterpriseCustom quoteSSO, RBAC, workload identity, PII redaction, own VPC or on-prem, 45-day onboarding.

Where it shingle

The weaknesses are structural and worth stating before any comparison. Every layer of the abstraction is a prompt: role, goal and backstory are prepended to each task, so the context grows with the crew, and the framework's own defaults, twenty iterations, three guardrail retries and memory analysis on save and on recall, are billed separately. Delegation and the hierarchical manager are the worst offenders, because each hop is another model call whose only job is to decide who works next. Python is the only runtime, in a category where the surrounding application is often TypeScript. And outside the paid platform, debugging rests on log output, which is weaker than a graph runtime that can replay a failed node.

FrameworkHow work is wiredCost per taskDebugging story
CrewAIRoles and tasks; a flow owns state and orderHighest: delegation, memory analysis and guardrail retries add model callsAMP traces need an account; OpenTelemetry hooks for Langfuse, Phoenix, Braintrust, MLflow
LangGraphExplicit graph, typed state, checkpointsLowest: routing is plain PythonLangSmith traces, checkpoint replay from any node
AutoGenConversational teams with patterns and termination conditionsHigh and open-ended unless turn limits are setAgentChat logging; GraphFlow adds a directed graph

LangGraph is the better tool for a fixed production pipeline with loops and checkpoints, and it wins on cost because routing is plain Python. AutoGen is the better tool for open-ended conversation. CrewAI wins on the part that matters when the system has to be handed to someone who is not an engineer: the vocabulary is a job description, and the JSONC or YAML configuration is readable. The vendor's own documentation includes a LangGraph-to-CrewAI migration guide and comparison notebooks, which is worth knowing before a contract is signed.

Verdict

CrewAI is a good default for teams that want agent-shaped code without hand-wiring a graph, and a poor default wherever the token bill is the binding constraint. The framework is competent, its documentation is better than its competitors', and the flow layer earns its place on its own. The cost is that autonomy is expressed as model calls, and the framework is happy to make them on your behalf.

  1. Adopt it when the pipeline is mostly linear and the expensive part is one or two open-ended steps. Keep the crew small and let the flow hold the logic.
  2. Adopt it when non-engineers have to read or edit the configuration, because a role, a goal and a YAML file are easier to review than a node graph.
  3. Do not adopt it for a fixed, high-volume pipeline where every hop is a plain function call. The coordination overhead buys flexibility such a pipeline never uses.
  4. Do not adopt it if you need checkpoint replay and per-node cost attribution on day one, without either paying for the platform or wiring OpenTelemetry yourself.
  5. Measure before keeping it: put usage_metrics behind a dashboard in the first week and compare the crew against the same steps written as a plain flow.

Sources

  1. CrewAI documentation: introductionCrewAI on PyPI: crewai 1.15.23CrewAI docs: FlowsCrewAI docs: MemoryCrewAI docs: tracingCrewAI pricingGitHub: crewAIInc/crewAILangGraph overviewAutoGen AgentChat user guide

Frequently asked questions

Is CrewAI free for production use?

Yes. The framework is MIT-licensed with no run quota; the cost is the model calls your agents make plus the infrastructure you run them on. The commercial product, CrewAI AMP, has a free Basic plan with 50 workflow executions a month and an Enterprise tier priced by quote.

Does CrewAI depend on LangChain?

No. The package has no LangChain dependency and is positioned as a standalone framework. Model access is handled through provider SDKs, with LiteLLM available as the crewai[litellm] extra.

How do crews and flows differ?

A flow owns state and execution order through @start, @listen and @router methods; a crew is a set of agents and tasks that runs inside a flow step. The vendor's own recommendation is to start with a flow and delegate to a crew when a step needs autonomy.

Is CrewAI's memory safe for sensitive data?

Not out of the box. Record content is sent to the configured LLM for scope, category and importance analysis, and the default embedder is OpenAI text-embedding-3-large. The documentation recommends a local LLM and a local embedder such as Ollama when content is sensitive.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.