Tools/Web engineering
Playwright: one browser API for tests, scripts and agents
Playwright drives Chromium, Firefox and WebKit from one Apache-2.0 API: a test runner, an agent-facing CLI and an MCP server. What it costs and where it falls short.
- Type
- Browser automation
- Pricing
- Apache-2.0 · cloud paid
Balázs Csorba··10 min read
- End-to-end testing
- Browser automation
- MCP
- Cross-browser

Key takeaways
- Playwright is Apache-2.0 with no paid edition; the money sits in CI minutes and, optionally, in Microsoft Playwright Workspaces, which bills per test minute by the second after a 30-day trial with 100 minutes.
- One API covers Chromium, Firefox and WebKit on Windows, Linux and macOS, with mobile emulation for Android Chrome and Mobile Safari.
- The agent surfaces are first-party: an MCP server of 70+ tools reading accessibility snapshots, a token-efficient CLI with a persistent daemon, and planner, generator and healer test agents.
- Accessibility snapshots are the design decision that matters for LLM use: elements carry refs instead of pixels, so interaction stays deterministic and no vision model is required.
- The runner, not the API, is where the maintenance cost lives: projects, sharding, traces and retries are free of charge but not of upkeep.
Playwright is Microsoft’s browser automation library, and by now it is three products in one repository: a test runner, a scripting library, and an agent-facing surface made of an MCP server and a command-line tool. It is Apache-2.0, it has no paid edition, and it remains the most ergonomic API in this category. The position of this review: it is the default for new end-to-end work, and its agent surfaces are the reason to look again at an existing suite.
It competes with Selenium, Cypress and Puppeteer, and increasingly with hand-written scraping and QA scripts, because the same page object now serves a test in CI, a nightly report and an LLM driving a browser through accessibility snapshots. Nothing in the core requires a model: the agent tooling sits on top and can be ignored, which is why teams adopt it as a testing tool first and discover the agent layer later.
What it is
The project ships browser automation for Chromium, Firefox and WebKit on Windows, Linux and macOS, locally or in CI, headless or headed, with native mobile emulation for Android Chrome and Mobile Safari. The current line is 1.63, published in September 2026, and the browser binaries are pinned to that release.
- Apache-2.0 with no paid edition; the npm package declares Node 20 or newer as its engine.
- One API in four languages: JavaScript and TypeScript, Python, Java and .NET.
- A test runner with auto-waiting, web-first assertions, fixtures, parallel projects, sharding, retries and an HTML report.
- Browser contexts give every test a fresh profile, so authentication state is a fixture rather than a shared global.
- Tooling around the run: codegen, UI mode, the trace viewer, network mocking, clock control and aria snapshots.
- Official container images published per release, which is what keeps CI browsers identical to the local ones.
- Three agent-facing surfaces: an MCP server, a CLI with installable skills, and planner, generator and healer test agents.
How it works
The driver speaks to each browser over its own protocol — patched Chromium and Firefox builds plus a patched WebKit — which is how one API can offer the same capabilities on all three engines instead of the lowest common denominator. Locators are resolved when the action runs and every action waits for the element to become actionable, so the usual fix for a flaky test is an assertion rather than a sleep.
Above the driver sits the runner, where most of the value accumulates: a config file declares the projects — browsers, viewports, credentials — tests are distributed across workers, sharding splits them across machines, and every failure carries a trace that can be replayed step by step. That trace is also what makes an assisted diagnosis practical, because it is structured data rather than a screenshot of a stack trace.
Getting started
One scaffold, one install, one run: npm init playwright@latest writes the config and a first spec, npx playwright install --with-deps fetches the pinned browser binaries with their system libraries, and npx playwright test runs headless and in parallel across the three engines.
import { test, expect, devices } from '@playwright/test';
test('order confirmation shows a number', async ({ page }) => {
await page.goto('/checkout');
await page.getByLabel('Email').fill('qa@example.com');
await page.getByRole('button', { name: 'Place order' }).click();
await expect(page.getByRole('status')).toContainText(/Order WS-\d+/);
});
test.describe('phone', () => {
test.use({ ...devices['iPhone 15'] });
test('checkout renders on a phone viewport', async ({ page }) => {
await page.goto('/checkout');
await expect(page.getByRole('heading', { name: 'Checkout' })).toBeVisible();
});
});What dates quickly is not the runner but the selectors. The project’s own advice is to prefer role and label locators, to keep assertions web-first so they re-check until the timeout instead of sampling once, and to open a trace on failure rather than guess from a stack.
Agents at the wheel
This is the part that changed the tool’s standing. Playwright is now a standard way for a language model to operate a browser, and the project ships three routes to it, all first-party.
# MCP server: structured tool calls, accessibility snapshots with element refs
{
"mcpServers": {
"playwright": { "command": "npx", "args": ["@playwright/mcp@latest"] }
}
}
# CLI: shell commands for a coding agent, headless by default
npm install -g @playwright/cli
playwright-cli open https://example.com
playwright-cli snapshot
playwright-cli click e5
# Test agents: planner, generator and healer, installed into the repo
npx playwright init-agents --loop=vscode- The MCP server returns accessibility snapshots: every element carries a ref such as
e5, so the model clicks a reference instead of guessing coordinates, and no vision model is needed. The documentation lists 70+ tools, with network mocking, storage, tracing and video behind capability flags. - The CLI is the token-efficient route: concise output, a persistent daemon so commands do not pay browser start-up, skills discovered on demand, and sessions that keep state between commands.
- Playwright Test agents — planner, generator and healer — write a Markdown test plan, turn it into specs and patch failing tests by replaying steps and re-inspecting the page;
npx playwright init-agentsinstalls the definitions for VS Code, Claude Code, Codex and OpenCode. - The documentation states the trade directly: MCP suits specialised agentic loops and exploratory automation, the CLI suits coding agents working with large codebases, and the CLI starts headless while MCP opens a headed browser.
Token economics decide this choice more often than features do. An MCP round trip puts tool schemas and a fresh snapshot into context on every step; a CLI command returns a path to a YAML snapshot and one line of status. For a long-running agent that only has to click through a flow, that difference is the gap between a session that fits in context and one that does not.
Screenshots still exist — they are one tool call away, and vision mode puts the mouse on coordinates — but they cost tokens by the pixel and reintroduce exactly the ambiguity the snapshot design removed. The accessibility tree is a summary the model can act on; a screenshot is evidence it has to interpret.
Pricing
The framework costs nothing, and there is no Pro tier, no seat count and no usage meter. Cost appears in three places: the CI minutes the suite burns, the engineering time that keeps it green, and, only when a managed grid is wanted, Microsoft’s first-party cloud.
- Framework: free under Apache-2.0, unlimited projects, unlimited tests, no licence to renew.
- Playwright Workspaces in Azure App Testing: pay-as-you-go per test minute, billed by the second, with a 30-day trial covering 100 minutes; beyond that the workspace converts to the metered plan, and published rates are region-dependent.
- Third-party grids such as BrowserStack and Sauce Labs bill per parallel session or subscription, and are needed only for real devices, a wider browser matrix or compliance.
- The largest line is never the licence: flaky-test triage and selector upkeep are the recurring cost, and no plan covers them.
Because the runner is free, the migration question is usually about time rather than budget. Moving an existing suite means rewriting selectors, fixtures and custom commands, and the payback arrives with the second month of maintenance rather than on the first green build.
Where it shingles
The weaknesses are known and mostly accepted. Browser binaries are pinned per release, so an upgrade is a scheduled task that drags CI images along with it. The breadth of the runner is a lot of surface before a team sees any benefit from sharding and traces. WebKit here is a patched build rather than Safari itself, so an Apple-specific defect can still surface late. And the CLI and MCP layers are new enough that their command sets move between minor releases.
| Playwright | Selenium | Cypress | |
|---|---|---|---|
| Driver | Own protocol per browser | W3C WebDriver | Proxy injected in the page |
| Browsers | Chromium, Firefox, WebKit | Any WebDriver browser | Chrome, Firefox, Edge, Electron |
| Runner | Built in: parallel, sharding, traces | Bring your own | Built in, sharding via plugins |
| Agent surfaces | MCP, CLI and test agents | Community MCP servers | Community plugins |
| Best fit | New suites and agent-driven work | Legacy stacks and vendor grids | Component-heavy front-end teams |
The honest comparison is about scope rather than quality. Selenium stays the answer when a contract demands WebDriver or a device cloud; Cypress is comfortable for teams already invested in its component test runner; Puppeteer is a library, not a framework, and is the right size for a script. Playwright claims that one tool covers all four jobs, and the maintenance cost of that claim is the browser matrix you now run in CI.
Verdict
Playwright is the strongest default in browser automation, and the interesting question is no longer whether to adopt it but which of its three surfaces to expose to whom.
- Adopt it for new end-to-end work. Runner, context isolation and the trace viewer pay off after the second month, which is exactly when a thinner tool starts costing more.
- Adopt it when an agent has to drive a browser. Refs and roles are cheaper and more deterministic than pixels, and the project maintains the MCP and CLI routes itself instead of leaving them to the community.
- Budget for the upgrade cadence: a release line that pulls new browser binaries every few weeks means lockfile, image and browsers move together or not at all.
- Keep Selenium where WebDriver or a device cloud is contractual, and do not rewrite a green Cypress suite only because the API elsewhere is nicer.
- Do not treat the test agents as review. The healer may skip a test it believes is broken, which is a sensible default and a bad one to leave unnoticed.
The browser is no longer only the thing under test; it is also the interface an agent operates. Tools that expose it as structured state — roles, refs, text — will age better than tools that expose it as pixels.
Sources
Frequently asked questions
Is Playwright free to use?
The framework is Apache-2.0 and has no paid edition. Cost appears as CI minutes and engineering time, and optionally as Microsoft Playwright Workspaces inside Azure App Testing, which bills per test minute, by the second, after a 30-day trial with 100 minutes.
Should an agent use the MCP server or the CLI?
The documentation puts the CLI ahead for coding agents: shell commands with concise output avoid loading tool schemas and accessibility trees into context. MCP suits exploratory loops that benefit from persistent state and structured parameters, and it opens a headed browser while the CLI defaults to headless.
How does Playwright differ from Selenium?
Playwright drives browsers through its own protocols and ships one API, a runner with auto-waiting and web-first assertions, and version-pinned browser binaries. Selenium speaks W3C WebDriver, which matters when a contract or a vendor grid requires it, and otherwise costs more glue for the same flows.
Can Playwright generate and repair tests with an LLM?
Yes, first-party: npx playwright init-agents installs planner, generator and healer definitions that write a Markdown test plan, turn it into specs, and patch failing tests by replaying steps and re-inspecting the page. The healer may also skip a test it believes is broken, which is a sensible default and a bad one to leave unnoticed.