Tools/AI agents

Mastra review: TypeScript agents with a real evaluation loop

Mastra bundles agents, workflows, memory, MCP, guardrails, tracing and evals into one TypeScript framework. What the Apache-2.0 core covers, what the ee/ split costs and who should adopt it.

Type
Agent framework
Pricing
Apache-2.0 core · hosted paid

··11 min read

  • TypeScript
  • Agent framework
  • Workflows
  • MCP
  • Evals
Diagram of the Mastra runtime: an agent calling tools and a model, workflow steps, memory and storage below, and a strip of observability spans at the bottom.

Key takeaways

  • Mastra is the most complete TypeScript agent framework in one repository: agents, workflows, memory, MCP server authoring, guardrails, tracing and evals, with Apache-2.0 on everything outside the ee/ directories.
  • The licence has a commercial edge. Code under ee/ - currently auth, the agent builder and the editor - is source-available, and production use needs both a written agreement and a license key.
  • The framework moves fast: 99 releases of @mastra/core shipped in the 30 days to 7 October 2026, so pinning exact versions and running evals in CI is not optional.
  • The real strength is the loop between code, traces and evals rather than the agent loop itself, which is commodity; Studio time travel and experiments are what keep a team on the paid platform.
  • Against LangGraph.js and the Vercel AI SDK, Mastra owns more of the runtime and asks for more dependency surface: @mastra/core alone pulls 30 direct dependencies.

Mastra is a TypeScript framework for building LLM agents, and it is one of the more complete ones: agents, a workflow engine, memory, MCP server authoring, guardrails, tracing, evals and a hosted platform in a single repository. The position here is plain. Mastra is the best-supported way to ship an agent inside an existing Node or Next.js codebase, and the wrong choice for a Python shop or for a team that only wants the tool-calling loop.

It sits above the model layer and below the application: models go in as provider/model strings, and the framework takes over the agent loop, the workflow graph, thread state and the traces. The direct competitors are LangGraph.js, the Vercel AI SDK, the OpenAI Agents SDK and newer entries such as VoltAgent. The difference is not the loop, which is commodity, but how much of the surrounding runtime each framework is willing to own.

What Mastra actually is

Mastra is a monorepo of scoped npm packages rather than one library. The core is @mastra/core, and everything else hangs off it: @mastra/memory, @mastra/libsql, @mastra/pg, @mastra/observability, @mastra/evals, @mastra/mcp and the rest. The Mastra class in src/mastra/index.ts is the registry that wires agents, workflows, storage, logging and observability together, and it is where every other part of the framework resolves its services from.

  • Licence: Apache-2.0 for everything outside the ee/ directories, source-available under the Mastra Enterprise Edition License inside them.
  • Maturity: @mastra/core stood at 1.75.0 on 7 October 2026; the repository shows 28.6k stars and 2.9k forks, and the first release shipped in October 2024.
  • Adoption: 2.23 million npm downloads for @mastra/core in the week to 4 October 2026, against 4.55 million for @langchain/langgraph and 34.5 million for Vercel's ai package.
  • Release velocity: 99 releases of @mastra/core in the 30 days to 7 October 2026 and 301 in 90 days, out of 1,656 since October 2024.
  • Runtime: Node 22.18 or later, with a Hono-based server, deployers for Vercel, Netlify and Cloudflare, or any HTTP host of your own.
  • Model access: a model router that resolves a provider/model string, lists 213 providers and 7,790 models, supports per-model fallback chains and reads the provider key from the environment.
  • Types: Zod, Valibot or ArkType through Standard JSON Schema for tools, workflow steps and structured output.

How it works

Two primitives carry most of the load. An Agent binds a model, instructions and tools and iterates until the model emits a final answer or a stop condition is met. A workflow, built from createStep and createWorkflow with .then(), .branch(), .parallel() and .commit(), is the deterministic path for a process with a known sequence. The docs are explicit that the second is right for a defined pipeline and the first for open-ended work, which matches the split described in the agent loop explained.

Mastra runtime: agent path, workflow path, observabilityA request reaches an agent, which calls a tool and a model in a loop. Below it sit workflow steps, memory and storage. Every run is traced into spans, logs and metrics that go to Mastra storage or to OpenTelemetry-compatible backends.Mastra runtime@mastra/core 1.75.0AGENT PATHRequesttext or eventAgentmodel, toolsToolcreateTool, zodModelprovider/modelobserve and repeatWORKFLOW AND STATEWorkflow stepsthen, branch, parallelMemorythreads, observationsStorageLibSQL, PostgrestracedOBSERVABILITYSpans, logs, metrics, scorers to storage, Langfuse, Datadog, Arize
The Mastra runtime: an agent loops over tools and a model, workflow steps and memory sit below it, and every run is traced.

Around those two sit the parts that keep an agent alive in production. Memory splits into message history, working memory, semantic recall and Observational Memory, where background agents compress older turns into observations before the context window fills. Storage is pluggable across LibSQL, Postgres, ClickHouse, MongoDB, MSSQL and DuckDB, and the same engine backs workflow suspend and resume, so a workflow can wait for a human approval and continue hours later, as described in human-in-the-loop patterns.

MCP support is the differentiator

Mastra is one of the few frameworks that both consumes and serves MCP. MCPClient connects to stdio or Streamable HTTP servers and MCPServer exposes Mastra agents, tools, workflows, prompts and resources to other systems. Registered servers are served at /api/mcp/:serverId/mcp and speak the 2026-07-28 revision of the protocol, so a team whose internal tools already sit behind an MCP server can wire an agent to them without writing an integration layer.

import { Agent } from '@mastra/core/agent'
import { Mastra } from '@mastra/core/mastra'
import { MCPServer, MCPClient } from '@mastra/mcp'

// Expose Mastra primitives to other agents and systems
const mcpServer = new MCPServer({
  id: 'support-mcp',
  name: 'Support tools',
  version: '1.0.0',
  agents: { supportAgent },
  tools: { refundTool },
  workflows: { triage },
})

// Consume remote MCP tools, gating the destructive ones
const client = new MCPClient({
  servers: {
    jira: {
      url: new URL('https://jira.example.com/mcp'),
      requireToolApproval: ({ toolName }) => toolName.startsWith('delete_'),
    },
  },
})

const assistant = new Agent({
  id: 'assistant',
  name: 'Assistant',
  instructions: 'Answer from the available tools and cite the source.',
  model: 'anthropic/claude-sonnet-4-6',
  tools: await client.listTools(),
})

export const mastra = new Mastra({
  agents: { assistant },
  mcpServers: { mcpServer },
})

The security defaults are better than average. requireToolApproval gates tools by name or arguments, stdio subprocesses inherit only a curated environment whitelist unless inheritDefaultEnv is set, allowedHosts restricts outbound HTTP hosts, and the docs tell you to treat annotations from servers you do not control as untrusted hints. Tool results remain untrusted model input, which is why the guardrail processors, not the MCP layer, are the place where that content has to be sanitised, as in the injection patterns.

Getting started

npm create mastra@latest produces a project with the src/mastra layout, a dev server, Studio and a deployer. The smallest useful surface is small: one tool, one agent, one registry.

// src/mastra/tools/stock-price.ts
import { createTool } from '@mastra/core/tools'
import { z } from 'zod'

export const stockPrice = createTool({
  id: 'stock-price',
  description: 'Latest price for a ticker symbol',
  inputSchema: z.object({ symbol: z.string().describe('Ticker, for example INGY') }),
  outputSchema: z.object({ symbol: z.string(), price: z.number() }),
  execute: async ({ symbol }) => ({ symbol, price: await quote(symbol) }),
})

// src/mastra/index.ts
import { Mastra } from '@mastra/core'
import { Agent } from '@mastra/core/agent'
import { LibSQLStore } from '@mastra/libsql'
import { stockPrice } from './tools/stock-price'

export const agent = new Agent({
  id: 'research',
  name: 'Research agent',
  instructions: 'Look prices up with the tool. Never guess a number.',
  model: 'anthropic/claude-sonnet-4-6',
  tools: { stockPrice },
})

export const mastra = new Mastra({
  agents: { agent },
  storage: new LibSQLStore({ id: 'mastra', url: 'file:./mastra.db' }),
})

// run.ts - Node 22.18 and later run TypeScript directly
const answer = await mastra.getAgentById('research').generate('INGY price')
console.log(answer.text, answer.usage)

Two details trip newcomers up. Resolve agents through mastra.getAgentById() rather than importing them directly, because a direct import still runs but misses the instance storage, logger and telemetry, which leaves those runs invisible in traces. And the model string is the router format, provider/model read from an environment variable such as ANTHROPIC_API_KEY, not a provider object.

Observability and evals are the real product

This is the part that justifies the framework rather than the agent loop. Every agent run, workflow step, tool call and model call emits a span, and exporters write those spans to Mastra storage or to any OpenTelemetry-compatible backend, with named integrations for Langfuse, Arize Phoenix and Datadog. Metrics are derived from spans without extra instrumentation, and Studio adds a graph view, time travel for replaying a single step and an experiments tab. Teams that already built this with OpenTelemetry tracing by hand will recognise the shape.

  • Tracing and logging: spans plus structured logs correlated by trace and span ID, so a log line jumps to the run that produced it.
  • Metrics: token counts, latency and cost estimates extracted when a span closes; aggregation needs an analytics-capable store such as DuckDB, ClickHouse or Postgres.
  • Scorers: prebuilt and custom scorers attached to an agent or a single workflow step, with sampling that is deterministic per trace, so scores stay comparable across runs.
  • Guardrails as processors: prompt-injection detection, moderation, PII masking, system prompt scrubbing and a cost ceiling, each with a block, redact or warn strategy.

One sharp edge is worth knowing before the first CI run. Scorers attached to an agent or a step register themselves; scorers passed straight to runEvals() or to the quick checks must also be registered on the Mastra instance, otherwise every save fails with Scorer with id <id> not found and the scores never reach the store. The docs are candid that the results are identical either way, which is exactly what makes the failure confusing.

Pricing and the licence split

The framework is free and the hosted platform is where the money is. The pricing page lists three tiers, and the meter is the interesting part: observability events, CPU hours, data egress and a model gateway that adds 5.5 per cent to market token rates.

TierPriceObservabilityCompute and retention
Starter0 USD per month100k events, then 10 USD per 100k24 CPU hours, then 0.35 USD per hour; 15-day retention
Teams250 USD per month1M events, then 8 USD per 100k250 CPU hours, then 0.25 USD per hour; 6-month retention
EnterpriseCustomCustom volume and retentionRBAC, audit logs, uptime SLAs, on-prem deployment

Two line items deserve attention. An always-on deployment costs 100 USD per project on top of the tier, and the gateway charges market rate plus 5.5 per cent on input and output tokens with your own key. Memory is billed again, at 10 USD per million tokens beyond the first 100,000 or 1M depending on tier, so a chatty long-context agent shows up in three separate meters.

Where it shingles

The weaknesses come first. Release velocity is the operational one: 301 releases in 90 days means the framework's own API moves underneath you and reading release notes becomes part of the job. Second, the breadth cuts both ways. The repository ships harnesses, workspaces, channels, voice, browser control, code mode and agent factories, so a small team pays in review surface for features it will not use. Third, Observational Memory is not free, because background agents compress the history on top of every turn. Fourth, the Studio extras, time travel, experiments and the evaluate tab, are precisely the reason to stay on the paid tier, and they are gone if you export traces to Langfuse and self-host everything else.

CriterionMastraLangGraph.jsVercel AI SDK
LicenceApache-2.0 core, enterprise in ee/MITApache-2.0
Scope it ownsAgents, workflows, memory, MCP, evals, platformGraph runtime, durable executionModel calls, streaming, UI primitives
Evals and tracesBuilt in, plus a hosted platformLangSmith is separate and paidBring your own provider
FitProduct teams shipping agents in NodePython-first shops needing durable graphsApps that mostly stream model output
Weekly downloads, 4 Oct 20262.23M (@mastra/core)4.55M (@langchain/langgraph)34.5M (ai)

Set against the OpenAI Agents SDK, which is thinner, MIT licensed and stays close to the Responses API, Mastra is the more portable and the heavier option. The honest summary is that Mastra buys breadth and a genuine evaluation loop, and pays for it in dependency count, with 30 direct dependencies from @mastra/core alone, and in the pace at which that dependency changes.

Verdict

Mastra is a well-run, unusually complete framework with a clear centre of gravity: the loop between code, traces and evals is better than anything a TypeScript team would otherwise assemble by hand. The catch is that it is a moving target with a commercial overlay, and the two things a buyer cares about most, stable APIs and a licence that does not move, are exactly the two it is weakest on.

  1. Adopt it when the agent ships inside an existing Node or Next.js application and the team wants tracing, evals and a workflow engine without assembling six packages.
  2. Adopt it when human approval, long-running suspended workflows and MCP tool servers are requirements rather than nice-to-haves.
  3. Adopt it with pinned versions and an eval suite in CI, because the release cadence will otherwise decide your upgrade schedule.
  4. Skip it when the model provider is fixed and the surface is a single chat endpoint: the Vercel AI SDK or the OpenAI Agents SDK is smaller to reason about.
  5. Skip it when the stack is Python, or when the ee/ boundary falls inside something the product depends on.

Sources

  1. Mastra documentation: agents
  2. Mastra documentation: workflows
  3. Mastra documentation: memory
  4. Mastra documentation: MCP
  5. Mastra documentation: guardrails
  6. Mastra documentation: evals
  7. Mastra documentation: observability
  8. Mastra documentation: model providers
  9. Mastra pricing
  10. mastra-ai/mastra on GitHub
  11. Mastra Enterprise Edition License v2.0

Frequently asked questions

Is Mastra open source?

Mostly. The core framework and the vast majority of the monorepo are Apache-2.0, published as @mastra/core, which stood at 1.75.0 on 7 October 2026. Directories named ee/ - @mastra/core/auth/ee, @mastra/core/agent-builder/ee and @mastra/editor/ee - are source-available under the Mastra Enterprise Edition License, which allows development and testing but not production use without a license key and a written agreement.

How does Mastra compare with LangGraph.js?

LangGraph.js is a graph runtime built around durable execution and state, it is MIT licensed, and it pulls more weekly downloads: 4.55 million against 2.23 million for @mastra/core in the week to 4 October 2026. Mastra bundles more around the agent instead, with memory, guardrails, MCP server authoring, tracing, evals and a hosted platform. Choose on whether you need the graph primitives or the surrounding product surface.

Does Mastra support MCP?

In both directions. MCPClient connects to stdio and Streamable HTTP MCP servers, with requireToolApproval for gating and allowedHosts to restrict outbound hosts. MCPServer exposes Mastra agents, tools, workflows, prompts and resources at /api/mcp/:serverId/mcp and speaks the 2026-07-28 revision of the protocol.

What does Mastra cost?

The framework is free. The hosted platform lists a free Starter tier with 100k observability events, 24 CPU hours and 15-day retention, a 250 USD per month Teams tier with 1M events, 250 CPU hours, 6-month retention, SSO and SOC 2 documentation, and a custom Enterprise tier. An always-on deployment costs 100 USD per project on top, and the model gateway adds 5.5 per cent to market token rates.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.