Tools/RAG & retrieval

Firecrawl: a web crawling API reviewed for RAG pipelines

Firecrawl turns URLs into clean Markdown through one hosted API. What it costs, where crawl accounting breaks down, and when to run the AGPL core yourself.

Type
Web crawling API
Pricing
Free tier · from $16 per month

··10 min read

  • Web scraping
  • RAG ingestion
  • Crawling
  • MCP
  • AGPL
Abstract pipeline artwork for the Firecrawl review

Key takeaways

  • Firecrawl's own scrape benchmark reports 96% coverage but an extraction F1 of 0.638, so roughly a third of the content on a page is wrong or missing.
  • One credit covers one page on scrape, crawl and map; the JSON, Question and Highlight formats add 4 credits per page on top.
  • A finished crawl reporting completed equal to total tells you nothing about failures; only the crawl errors endpoint names the pages that dropped.
  • The self-hosted AGPL-3.0 stack covers scrape, crawl, map and search, while screenshots, page actions, fire-engine, Agent and Interact stay on Cloud.
  • Crawl jobs stay retrievable for 24 hours and credits do not roll over below the Scale plan, so both ingestion state and spend need a budget.

Firecrawl is a hosted API that turns a URL into clean Markdown, JSON or HTML, and a crawler that does the same for a whole site. It exists because almost every retrieval pipeline eventually needs the web inside it, and because turning an arbitrary page into text a language model can read is a browser-automation problem in data-cleaning clothes. The verdict here: it is the best default for teams that want web context in a product this week, and a poor fit for anyone who needs to control egress, keep the raw bytes, or pay per megabyte rather than per page.

In the stack it sits between the fetch layer and the embedding layer. A vector store never sees Firecrawl; it sees the Markdown that came out of it. That position is the whole idea: Firecrawl competes less with a general scraping library like Scrapy than with the browser fleet, the proxy rotation and the blocking layer a team would otherwise assemble to feed the same pipeline. It also overlaps with search APIs, because /search returns page content next to result metadata.

What it is

The interesting part is not that Firecrawl fetches pages. It is that it decides per request how hard to try: a static fetch first, a headless browser when the page needs one, proxy rotation when the target pushes back. The API is a thin JSON surface over that machinery, and the repository carries the whole engine under AGPL-3.0.

  • Endpoints: /scrape, /crawl, /map, /search, /parse, /batch/scrape, plus the Cloud-only /interact and /agent
  • Output formats: Markdown, HTML, raw links, screenshots, and JSON extracted against a schema or a natural-language prompt.
  • Licence: AGPL-3.0 for the core engine and the API, MIT for the SDKs.
  • Billing unit: one credit per page on scrape, crawl and map, two credits per ten search results, two credits per browser minute on Interact.
  • Free tier: 1,000 credits a month, no card, two concurrent browsers, ten requests a minute on /scrape
  • MCP: a keyless Streamable HTTP server at https://mcp.firecrawl.dev/v2/mcp covers Search, Scrape and Parse.
  • Self-hosting: the documentation pins release v2.11.162 and serves the API on port 3002 with Docker Compose.

How it works

Every page goes through the same path regardless of endpoint, and that is what makes crawl cheap to reason about: whatever scrape can do, crawl can do to every page it reaches. The stages below are the documented ones, plus the accounting step that decides the bill.

The Firecrawl page pipeline and where credits are chargedFour stages run left to right: fetch, render, extract and return. A box below the first two stages marks the resolution path, a static fetch first and a headless browser when the page needs one. A second box marks accounting: one credit per page returned, plus four credits for JSON format output. Three lines below state that a crawl runs the same path per page, that pages which fail leave the data array and appear only in the crawl errors endpoint, and that a target answering 403 or 404 still costs one credit.FIRECRAWL PAGE PIPELINEthe same path on every endpointfetchrenderextractreturnRESOLUTIONstatic fetch, browser if neededACCOUNTING1 credit per page, +4 for JSONcrawl runs this path per page and applies the same scrape options to every pagepages that fail leave the data array and appear only in the crawl errors endpointa target answering 403 or 404 still lands in data and still costs one credit
One scrape path shared by every endpoint, with the credit charged at the end of it.

Two consequences follow from that shape. First, /crawl is the same code path as /scrape with a queue in front of it, so a crawl job inherits every scrape behaviour, including the ones nobody asked for. Second, the credit is charged when a page is produced rather than when it is useful, which is where the cost model starts to bite.

The vendor publishes a scrape benchmark of its own, run on 13 January 2026 over 1,000 public URLs drawn from ten categories, scoring whether the tool returned the core page text. The numbers deserve a close reading, because the definition of success is generous.

MetricFirecrawl result
Dataset1,000 public URLs across ten categories
Coverage, success rate96%
Extraction accuracy, F10.638
Content recall0.639
Latency, P953,387 ms

A page counts as covered once at least 10% of the expected content came back, so 96% coverage is a claim about not returning nothing, not about returning the right thing. The F1 of 0.638 is the honest number, and it means roughly a third of the extracted content on an average page is wrong or missing. That is good enough for a corpus that will be chunked, embedded and filtered anyway; it is not good enough for a pipeline that needs the exact figure out of a table.

Getting started

The Python SDK wraps the job queue, pagination and polling, which makes the shortest useful snippet also the one that hides the failure accounting. The example below crawls a small documentation set and then asks the API directly which pages it never fetched.

import os
import time

import requests
from firecrawl import Firecrawl

key = os.environ["FIRECRAWL_API_KEY"]
app = Firecrawl(api_key=key)

job = app.start_crawl(
    "https://docs.example.com",
    limit=100,                                # the default is 10000 pages
    scrape_options={"formats": ["markdown"], "only_main_content": True},
)

# poll until a terminal status: completed, failed or cancelled
while True:
    status = app.get_crawl_status(job.id)
    if status.status in ("completed", "failed", "cancelled"):
        break
    time.sleep(5)

print(len(status.data), "pages scraped")

# completed == total does not mean every page arrived
res = requests.get(
    f"https://api.firecrawl.dev/v2/crawl/{job.id}/errors",
    headers={"Authorization": f"Bearer {key}"},
).json()

for page in res.get("errors", []):
    print("failed:", page["url"], page["error"])

print("robots-blocked:", len(res.get("robotsBlocked", [])))

Two details in that call matter more than they look. The crawl limit defaults to 10,000 pages, and the endpoint rejects the job with a 402 when the remaining credits cannot cover the limit that was asked for, so leaving the default in place on a trial account is a fast route to getting nothing. And only_main_content is what strips navigation and footers; without it the boilerplate is paid for and embedded as well.

Self-hosting

The core is open source under AGPL-3.0, which matters twice over: it is auditable, and a modified version served over a network owes its source to whoever talks to it. The default Compose stack runs the API on port 3002 with a PostgreSQL queue, Redis and a Playwright service, and the documentation pins a verified release instead of tracking the main branch.

  • Included by default: the scrape, crawl, map and search routes, with fetch and Playwright processing.
  • Needs a provider you attach: LLM-backed extraction and JSON formats want an OpenAI-compatible endpoint or Ollama.
  • Absent from the default stack: the fire-engine anti-bot service, screenshots and page actions.
  • Cloud only: Agent, Browser, Interact, the dashboard and the enterprise controls.
  • Not production-ready as shipped: the quickstart runs with authentication disabled and without durable volumes.

Pricing

Everything bills against one credit balance. The rates are identical on every plan, which turns cost forecasting into a page-count problem rather than a tier problem; only the limits move.

PlanMonthly creditsConcurrent browsersRate limit on /scrape, /map and /search
Free1,000210 / min
Hobby, $16 annually5,0005100 / min
Standard, $83 annually100,00025500 / min
Growth, $333 annually500,000505,000 / min
Scale, $599 annually1,000,00010010,000 / min

The unit economics are simple enough to reason about. A Hobby plan at $16 a month buys 5,000 pages, which is about $0.003 per page for the first 5,000 and then $5 per extra 1,000 credits. Three things break that arithmetic. The JSON, Question and Highlight formats add 4 credits per page on top of the base cost. Interact bills 2 credits per browser minute rather than per page. And a page that answers 403 or 404 is still returned to the caller and still charged 1 credit, so crawling a site full of soft 404s is a real invoice rather than a rounding error.

Where it fits and where it does not

The weaknesses come first. Crawl accounting cannot tell you whether a run was clean: the total counter sums completed, active, queued and backlogged pages and excludes failures, so completed equal to total on a finished job holds whether or not pages were dropped, and only the crawl errors endpoint names them. Crawl discovery is also non-deterministic, because pages are fetched concurrently and link order follows network timing. On the search side, Firecrawl's own benchmarks page reports it twelfth of sixteen configurations on multi-hop discovery at an F1 of 30.4%, while placing second on search-only coding tickets. It is a strong fetcher with a search product attached, not a strong search product.

ToolWhat it sellsBilling unitWhere it wins
FirecrawlOne API for scrape, crawl, map, search and parse, returning MarkdownCredits per pageClean Markdown with almost no cleanup code
ApifyAn automation platform: Actors, datasets, proxies, schedulingCompute units, one CU is 1 GB for an hour, plus proxies and storageOdd shapes that need custom code or a proxy choice
BrowserbaseManaged browser sessions plus Fetch and Search APIsBrowser hours, $0.12 an hour on Developer after the first 100Interactive flows: login, click-through, stateful sessions
Scrapy and PlaywrightLibraries you run yourself, with the whole pipeline under your controlYour infrastructure and your own timePredictable cost at volume, and no vendor in the loop

The comparison is not like for like, and the difference in kind is the useful part. Firecrawl and Browserbase both sell a fetch result. Apify sells compute plus a marketplace. Scrapy and Playwright sell nothing and cost only attention. For a fixed ingestion target with a stable HTML shape, a self-hosted crawler is still cheaper per page than any of them, because the marginal cost is a core and a queue rather than a credit.

There is the lock-in question too. Firecrawl Cloud is where the good parts live: fire-engine, screenshots, page actions, Interact and Agent are all documented as Cloud capabilities. Self-hosting gives you the core engine under AGPL-3.0 and not much else, which is a narrower proposition than the popularity of the repository suggests.

Verdict

Firecrawl is worth adopting for the ingestion problem and worth keeping away from the search problem. An extraction F1 of 0.638 is acceptable precisely because chunking and retrieval absorb some extraction error, the credit model is predictable, and running a browser fleet, a proxy pool and a blocking layer is real operational work that most product teams should not take on. The case against it is narrow but sharp: when page content has to be exact, or when egress and data residency are the actual constraint, the credits buy the wrong thing.

  1. Adopt it if a product needs web content now and the target sites are ordinary HTML or documentation.
  2. Adopt it for agent and MCP integrations. The keyless Streamable HTTP server and the CLI skills make it the shortest path from an agent to a readable page.
  3. Use the free tier to find out whether the extraction quality matches the corpus. 1,000 credits a month is enough for that.
  4. Do not adopt it as a search API. On the vendor's own multi-hop benchmark it sits near the bottom of the field, and it bills per result.
  5. Do not adopt it for exact extraction from tables, filings or prices. Budget for a verification pass, or use a source that serves structured data.
Firecrawl is best understood as a rendering service with a JSON API attached. Anything that would have been built from Playwright and a proxy subscription should now be bought per page. Anything that would have been built from a sitemap and a scheduled job is still worth building in house.

Sources

  1. Firecrawl pricing — credit rates, plan limits, rollover and pay-as-you-go rules, effective 4 September 2026
  2. Firecrawl crawl documentation — status counters, paging contract, the errors endpoint and the crawl configuration reference
  3. Firecrawl self-hosting guide — the pinned release, the Compose stack and the capability gaps in the default build
  4. Open source or Firecrawl Cloud — which capabilities belong to which operating model
  5. Firecrawl benchmarks — the January 2026 scrape run and the third-party search studies Firecrawl cites
  6. Apify pricing — compute units, proxy rates and prepaid usage for the comparison table
  7. Browserbase pricing — browser hours and Fetch and Search call rates for the comparison table

Frequently asked questions

How much does Firecrawl cost to crawl a documentation site?

One credit per page on scrape, crawl and map, so a 500-page crawl costs 500 credits. Hobby is $16 a month billed annually for 5,000 credits and Standard is $83 for 100,000. The JSON, Question and Highlight formats add 4 credits per page, and pay-as-you-go tops the balance up in $5 increments.

Is Firecrawl accurate enough for RAG?

Firecrawl's own benchmark, run on 13 January 2026 over 1,000 public URLs, reports 96% coverage and an extraction F1 of 0.638. A page counts as covered at 10% of expected content, so read the F1 rather than the coverage figure. That accuracy is fine for a chunked, embedded corpus and too low for exact figures lifted out of tables.

Can I self-host Firecrawl?

Yes. The core engine and API are AGPL-3.0 and the self-hosting guide runs them with Docker Compose on port 3002, pinning a verified release rather than the main branch. The default stack covers scrape, crawl, map and search with fetch and Playwright processing. Screenshots, page actions, fire-engine, Agent and Interact are Cloud features.

How do I find out which pages a crawl failed to fetch?

The status counters will not tell you. The total field sums completed, active, queued and backlogged pages and excludes failures, so completed equal to total holds on a finished job whether or not pages were dropped. Call the crawl errors endpoint and read its errors and robotsBlocked arrays.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.