Tools/RAG & retrieval
Firecrawl: a web crawling API reviewed for RAG pipelines
Firecrawl turns URLs into clean Markdown through one hosted API. What it costs, where crawl accounting breaks down, and when to run the AGPL core yourself.
- Type
- Web crawling API
- Pricing
- Free tier · from $16 per month
Balázs Csorba··10 min read
- Web scraping
- RAG ingestion
- Crawling
- MCP
- AGPL

Key takeaways
- Firecrawl's own scrape benchmark reports 96% coverage but an extraction F1 of 0.638, so roughly a third of the content on a page is wrong or missing.
- One credit covers one page on scrape, crawl and map; the JSON, Question and Highlight formats add 4 credits per page on top.
- A finished crawl reporting completed equal to total tells you nothing about failures; only the crawl errors endpoint names the pages that dropped.
- The self-hosted AGPL-3.0 stack covers scrape, crawl, map and search, while screenshots, page actions, fire-engine, Agent and Interact stay on Cloud.
- Crawl jobs stay retrievable for 24 hours and credits do not roll over below the Scale plan, so both ingestion state and spend need a budget.
Firecrawl is a hosted API that turns a URL into clean Markdown, JSON or HTML, and a crawler that does the same for a whole site. It exists because almost every retrieval pipeline eventually needs the web inside it, and because turning an arbitrary page into text a language model can read is a browser-automation problem in data-cleaning clothes. The verdict here: it is the best default for teams that want web context in a product this week, and a poor fit for anyone who needs to control egress, keep the raw bytes, or pay per megabyte rather than per page.
In the stack it sits between the fetch layer and the embedding layer. A vector store never sees Firecrawl; it sees the Markdown that came out of it. That position is the whole idea: Firecrawl competes less with a general scraping library like Scrapy than with the browser fleet, the proxy rotation and the blocking layer a team would otherwise assemble to feed the same pipeline. It also overlaps with search APIs, because /search returns page content next to result metadata.
What it is
The interesting part is not that Firecrawl fetches pages. It is that it decides per request how hard to try: a static fetch first, a headless browser when the page needs one, proxy rotation when the target pushes back. The API is a thin JSON surface over that machinery, and the repository carries the whole engine under AGPL-3.0.
- Endpoints:
/scrape, /crawl, /map, /search, /parse, /batch/scrape, plus the Cloud-only/interactand/agent - Output formats: Markdown, HTML, raw links, screenshots, and JSON extracted against a schema or a natural-language prompt.
- Licence: AGPL-3.0 for the core engine and the API, MIT for the SDKs.
- Billing unit: one credit per page on scrape, crawl and map, two credits per ten search results, two credits per browser minute on Interact.
- Free tier: 1,000 credits a month, no card, two concurrent browsers, ten requests a minute on
/scrape - MCP: a keyless Streamable HTTP server at
https://mcp.firecrawl.dev/v2/mcpcovers Search, Scrape and Parse. - Self-hosting: the documentation pins release
v2.11.162and serves the API on port 3002 with Docker Compose.
How it works
Every page goes through the same path regardless of endpoint, and that is what makes crawl cheap to reason about: whatever scrape can do, crawl can do to every page it reaches. The stages below are the documented ones, plus the accounting step that decides the bill.
Two consequences follow from that shape. First, /crawl is the same code path as /scrape with a queue in front of it, so a crawl job inherits every scrape behaviour, including the ones nobody asked for. Second, the credit is charged when a page is produced rather than when it is useful, which is where the cost model starts to bite.
The vendor publishes a scrape benchmark of its own, run on 13 January 2026 over 1,000 public URLs drawn from ten categories, scoring whether the tool returned the core page text. The numbers deserve a close reading, because the definition of success is generous.
| Metric | Firecrawl result |
|---|---|
| Dataset | 1,000 public URLs across ten categories |
| Coverage, success rate | 96% |
| Extraction accuracy, F1 | 0.638 |
| Content recall | 0.639 |
| Latency, P95 | 3,387 ms |
A page counts as covered once at least 10% of the expected content came back, so 96% coverage is a claim about not returning nothing, not about returning the right thing. The F1 of 0.638 is the honest number, and it means roughly a third of the extracted content on an average page is wrong or missing. That is good enough for a corpus that will be chunked, embedded and filtered anyway; it is not good enough for a pipeline that needs the exact figure out of a table.
Getting started
The Python SDK wraps the job queue, pagination and polling, which makes the shortest useful snippet also the one that hides the failure accounting. The example below crawls a small documentation set and then asks the API directly which pages it never fetched.
import os
import time
import requests
from firecrawl import Firecrawl
key = os.environ["FIRECRAWL_API_KEY"]
app = Firecrawl(api_key=key)
job = app.start_crawl(
"https://docs.example.com",
limit=100, # the default is 10000 pages
scrape_options={"formats": ["markdown"], "only_main_content": True},
)
# poll until a terminal status: completed, failed or cancelled
while True:
status = app.get_crawl_status(job.id)
if status.status in ("completed", "failed", "cancelled"):
break
time.sleep(5)
print(len(status.data), "pages scraped")
# completed == total does not mean every page arrived
res = requests.get(
f"https://api.firecrawl.dev/v2/crawl/{job.id}/errors",
headers={"Authorization": f"Bearer {key}"},
).json()
for page in res.get("errors", []):
print("failed:", page["url"], page["error"])
print("robots-blocked:", len(res.get("robotsBlocked", [])))Two details in that call matter more than they look. The crawl limit defaults to 10,000 pages, and the endpoint rejects the job with a 402 when the remaining credits cannot cover the limit that was asked for, so leaving the default in place on a trial account is a fast route to getting nothing. And only_main_content is what strips navigation and footers; without it the boilerplate is paid for and embedded as well.
Self-hosting
The core is open source under AGPL-3.0, which matters twice over: it is auditable, and a modified version served over a network owes its source to whoever talks to it. The default Compose stack runs the API on port 3002 with a PostgreSQL queue, Redis and a Playwright service, and the documentation pins a verified release instead of tracking the main branch.
- Included by default: the scrape, crawl, map and search routes, with fetch and Playwright processing.
- Needs a provider you attach: LLM-backed extraction and JSON formats want an OpenAI-compatible endpoint or Ollama.
- Absent from the default stack: the fire-engine anti-bot service, screenshots and page actions.
- Cloud only: Agent, Browser, Interact, the dashboard and the enterprise controls.
- Not production-ready as shipped: the quickstart runs with authentication disabled and without durable volumes.
Pricing
Everything bills against one credit balance. The rates are identical on every plan, which turns cost forecasting into a page-count problem rather than a tier problem; only the limits move.
| Plan | Monthly credits | Concurrent browsers | Rate limit on /scrape, /map and /search |
|---|---|---|---|
| Free | 1,000 | 2 | 10 / min |
| Hobby, $16 annually | 5,000 | 5 | 100 / min |
| Standard, $83 annually | 100,000 | 25 | 500 / min |
| Growth, $333 annually | 500,000 | 50 | 5,000 / min |
| Scale, $599 annually | 1,000,000 | 100 | 10,000 / min |
The unit economics are simple enough to reason about. A Hobby plan at $16 a month buys 5,000 pages, which is about $0.003 per page for the first 5,000 and then $5 per extra 1,000 credits. Three things break that arithmetic. The JSON, Question and Highlight formats add 4 credits per page on top of the base cost. Interact bills 2 credits per browser minute rather than per page. And a page that answers 403 or 404 is still returned to the caller and still charged 1 credit, so crawling a site full of soft 404s is a real invoice rather than a rounding error.
Where it fits and where it does not
The weaknesses come first. Crawl accounting cannot tell you whether a run was clean: the total counter sums completed, active, queued and backlogged pages and excludes failures, so completed equal to total on a finished job holds whether or not pages were dropped, and only the crawl errors endpoint names them. Crawl discovery is also non-deterministic, because pages are fetched concurrently and link order follows network timing. On the search side, Firecrawl's own benchmarks page reports it twelfth of sixteen configurations on multi-hop discovery at an F1 of 30.4%, while placing second on search-only coding tickets. It is a strong fetcher with a search product attached, not a strong search product.
| Tool | What it sells | Billing unit | Where it wins |
|---|---|---|---|
| Firecrawl | One API for scrape, crawl, map, search and parse, returning Markdown | Credits per page | Clean Markdown with almost no cleanup code |
| Apify | An automation platform: Actors, datasets, proxies, scheduling | Compute units, one CU is 1 GB for an hour, plus proxies and storage | Odd shapes that need custom code or a proxy choice |
| Browserbase | Managed browser sessions plus Fetch and Search APIs | Browser hours, $0.12 an hour on Developer after the first 100 | Interactive flows: login, click-through, stateful sessions |
| Scrapy and Playwright | Libraries you run yourself, with the whole pipeline under your control | Your infrastructure and your own time | Predictable cost at volume, and no vendor in the loop |
The comparison is not like for like, and the difference in kind is the useful part. Firecrawl and Browserbase both sell a fetch result. Apify sells compute plus a marketplace. Scrapy and Playwright sell nothing and cost only attention. For a fixed ingestion target with a stable HTML shape, a self-hosted crawler is still cheaper per page than any of them, because the marginal cost is a core and a queue rather than a credit.
There is the lock-in question too. Firecrawl Cloud is where the good parts live: fire-engine, screenshots, page actions, Interact and Agent are all documented as Cloud capabilities. Self-hosting gives you the core engine under AGPL-3.0 and not much else, which is a narrower proposition than the popularity of the repository suggests.
Verdict
Firecrawl is worth adopting for the ingestion problem and worth keeping away from the search problem. An extraction F1 of 0.638 is acceptable precisely because chunking and retrieval absorb some extraction error, the credit model is predictable, and running a browser fleet, a proxy pool and a blocking layer is real operational work that most product teams should not take on. The case against it is narrow but sharp: when page content has to be exact, or when egress and data residency are the actual constraint, the credits buy the wrong thing.
- Adopt it if a product needs web content now and the target sites are ordinary HTML or documentation.
- Adopt it for agent and MCP integrations. The keyless Streamable HTTP server and the CLI skills make it the shortest path from an agent to a readable page.
- Use the free tier to find out whether the extraction quality matches the corpus. 1,000 credits a month is enough for that.
- Do not adopt it as a search API. On the vendor's own multi-hop benchmark it sits near the bottom of the field, and it bills per result.
- Do not adopt it for exact extraction from tables, filings or prices. Budget for a verification pass, or use a source that serves structured data.
Firecrawl is best understood as a rendering service with a JSON API attached. Anything that would have been built from Playwright and a proxy subscription should now be bought per page. Anything that would have been built from a sitemap and a scheduled job is still worth building in house.
Sources
- Firecrawl pricing — credit rates, plan limits, rollover and pay-as-you-go rules, effective 4 September 2026
- Firecrawl crawl documentation — status counters, paging contract, the errors endpoint and the crawl configuration reference
- Firecrawl self-hosting guide — the pinned release, the Compose stack and the capability gaps in the default build
- Open source or Firecrawl Cloud — which capabilities belong to which operating model
- Firecrawl benchmarks — the January 2026 scrape run and the third-party search studies Firecrawl cites
- Apify pricing — compute units, proxy rates and prepaid usage for the comparison table
- Browserbase pricing — browser hours and Fetch and Search call rates for the comparison table
Frequently asked questions
How much does Firecrawl cost to crawl a documentation site?
One credit per page on scrape, crawl and map, so a 500-page crawl costs 500 credits. Hobby is $16 a month billed annually for 5,000 credits and Standard is $83 for 100,000. The JSON, Question and Highlight formats add 4 credits per page, and pay-as-you-go tops the balance up in $5 increments.
Is Firecrawl accurate enough for RAG?
Firecrawl's own benchmark, run on 13 January 2026 over 1,000 public URLs, reports 96% coverage and an extraction F1 of 0.638. A page counts as covered at 10% of expected content, so read the F1 rather than the coverage figure. That accuracy is fine for a chunked, embedded corpus and too low for exact figures lifted out of tables.
Can I self-host Firecrawl?
Yes. The core engine and API are AGPL-3.0 and the self-hosting guide runs them with Docker Compose on port 3002, pinning a verified release rather than the main branch. The default stack covers scrape, crawl, map and search with fetch and Playwright processing. Screenshots, page actions, fire-engine, Agent and Interact are Cloud features.
How do I find out which pages a crawl failed to fetch?
The status counters will not tell you. The total field sums completed, active, queued and backlogged pages and excludes failures, so completed equal to total holds on a finished job whether or not pages were dropped. Call the crawl errors endpoint and read its errors and robotsBlocked arrays.