[{"data":1,"prerenderedAt":720},["ShallowReactive",2],{"tool-replicate-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":23,"sources":30,"cover":61,"og":62,"expertise":63,"locales":64,"lang":65,"title":68,"description":69,"coverAlt":70,"url":71,"pricing":72,"kind":73,"metaTitle":74,"takeaways":75,"faq":81,"toc":94,"blocks":119,"others":504},"replicate","2026-05-22",10,"web",[9,10,11,12,13],"Model hosting","Serverless GPU","Diffusion","Async jobs","Webhooks",[15,16,17,18,19,20,21,22],"replicate api","replicate vs fal.ai","serverless gpu inference pricing","replicate webhooks","replicate cold start","replicate vs baseten","how to run an open source model in production","replicate prediction api",[24,27],{"name":25,"url":26},"Replicate","https:\u002F\u002Freplicate.com\u002Fdocs",{"name":28,"url":29},"Serverless computing","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FServerless_computing",[31,34,37,40,43,46,49,52,55,58],{"title":32,"url":33},"Replicate pricing","https:\u002F\u002Freplicate.com\u002Fpricing",{"title":35,"url":36},"Replicate docs: About predictions","https:\u002F\u002Freplicate.com\u002Fdocs\u002Ftopics\u002Fpredictions",{"title":38,"url":39},"Replicate docs: Create a prediction","https:\u002F\u002Freplicate.com\u002Fdocs\u002Ftopics\u002Fpredictions\u002Fcreate-a-prediction",{"title":41,"url":42},"Replicate docs: Prediction lifecycle","https:\u002F\u002Freplicate.com\u002Fdocs\u002Ftopics\u002Fpredictions\u002Flifecycle",{"title":44,"url":45},"Replicate docs: Rate limits","https:\u002F\u002Freplicate.com\u002Fdocs\u002Ftopics\u002Fpredictions\u002Frate-limits",{"title":47,"url":48},"Replicate docs: Data retention","https:\u002F\u002Freplicate.com\u002Fdocs\u002Ftopics\u002Fpredictions\u002Fdata-retention",{"title":50,"url":51},"Replicate docs: Verify webhooks","https:\u002F\u002Freplicate.com\u002Fdocs\u002Ftopics\u002Fwebhooks\u002Fverify-webhook",{"title":53,"url":54},"Replicate docs: Official models","https:\u002F\u002Freplicate.com\u002Fdocs\u002Ftopics\u002Fmodels\u002Fofficial-models",{"title":56,"url":57},"fal.ai pricing","https:\u002F\u002Ffal.ai\u002Fpricing",{"title":59,"url":60},"Baseten pricing","https:\u002F\u002Fwww.baseten.co\u002Fpricing\u002F","\u002Fimages\u002Fblog\u002Freplicate\u002Fcover.webp","\u002Fimages\u002Fblog\u002Freplicate\u002Fog.jpg","ai-engineer",[65,66,67],"en","de","hu","Replicate, reviewed: a model API priced per second, with the sharp edges named","Replicate puts thousands of open models behind one prediction API and bills per second of compute. A review of cold boots, version churn and the one-hour data deletion.","A request moves from starting to processing, then into one of four terminal states: succeeded, failed, canceled or aborted.","https:\u002F\u002Freplicate.com","Pay per second of compute","Model hosting API","Replicate, reviewed: per-second model hosting · Balázs Csorba",[76,77,78,79,80],"Replicate meters public models per second of hardware time and bills nothing while a shared model sits idle, which makes it cheap for bursty workloads and expensive for anything latency-sensitive.","Every run is a prediction object with a fixed set of statuses, and the difference between canceled and aborted decides whether you pay for the compute that ran.","Predictions created through the API are deleted an hour later, inputs and outputs included, so anything you need has to be copied out before then.","Only official models carry a stable input schema; community models are maintained by their authors and need a pinned version plus a test.","A deployment pins a version to hardware and an instance range, which removes cold boots and replaces them with a fixed hourly bill.",[82,85,88,91],{"q":83,"a":84},"What does Replicate actually charge for?","For models in the public catalogue, the price per second of the hardware the model runs on, starting at $0.000025 for a small CPU and reaching $0.001525 for an H100. Official models are the exception: they are billed per output image, per second of video or per token. Private models run on dedicated hardware, so you pay for setup and idle time as well as for the work.",{"q":86,"a":87},"Is Replicate cheaper than renting a GPU?","For spiky, low-duty-cycle work, yes, because you pay only while a prediction runs and shared capacity is reused between customers. For steady load above roughly a few percent utilisation, the arithmetic flips: a warm T4 instance on Replicate is $0.81 an hour, which is $583 for a 30-day month that never idles, and at that point a reserved instance on RunPod or Lambda is the cheaper contract.",{"q":89,"a":90},"Why does my Replicate prediction take minutes to start?","That is a cold boot. Replicate shuts down models that are not being used, and restarting one means loading several gigabytes of weights, which the documentation says can take several minutes. Official models are kept warm, and a deployment with min_instances set to 1 removes the wait for anything else.",{"q":92,"a":93},"Does Replicate keep my inputs and outputs?","Not for API predictions. Input parameters, output values, output files and logs are all removed an hour after the prediction is created, and the docs say you have to save your own copies. Predictions started from the web interface are kept indefinitely, which is worth knowing if the same model is used both ways.",[95,98,101,104,107,110,113,116],{"id":96,"title":97},"what-it-is","What Replicate actually is",{"id":99,"title":100},"how-it-works","How a prediction runs",{"id":102,"title":103},"getting-started","Getting started",{"id":105,"title":106},"pricing","What it costs",{"id":108,"title":109},"security-and-data","Data retention, tokens and webhooks",{"id":111,"title":112},"where-it-shingles","Where it hurts",{"id":114,"title":115},"verdict","Verdict",{"id":117,"title":118},"sources","Sources",[120,124,136,139,142,166,167,174,183,212,215,216,223,225,228,230,248,249,252,321,324,327,331,334,348,349,352,382,392,393,396,440,447,450,451,454,467,471,472],{"type":121,"content":122},"paragraph",[123],"Replicate is a hosted inference platform that puts thousands of open-weight and proprietary models behind a single prediction API, and it bills only for the seconds a machine spends running your request. The position here is plain: for a team that wants to try a new image or video model this afternoon without provisioning a GPU, it is the shortest path on the market. For a production endpoint with a latency budget and a predictable invoice, per-second billing is the wrong contract, and the platform only becomes defensible once a version is pinned and the model sits on a deployment.",{"type":121,"content":125},[126,127,131,132,135],"It sits between owning a GPU and buying from a model vendor. On one side it competes with ",{"tag":128,"href":57,"children":129},"a",[130],"fal.ai"," and ",{"tag":128,"href":60,"children":133},[134],"Baseten",", which sell inference rather than a catalogue; on the other it competes with renting a machine from a GPU broker and running your own container. What it does not replace is a token-billed LLM API. Replicate will happily serve a Llama model on an H100, but nobody prices a chat endpoint per second of GPU time.",{"type":137,"level":138,"id":96,"text":97},"heading",2,{"type":121,"content":140},[141],"Three things share one account and one API. First, a public catalogue of models published by their authors: fine-tunes, community implementations, research code. Second, a set of official models that Replicate maintains itself, which the documentation puts at over a hundred, and which are always warm, priced per output unit, and covered by a stable input schema. Third, private models that you package with Cog and run on dedicated hardware, where you pay for the whole life of the instance rather than for the work it does.",{"type":143,"ordered":144,"items":145},"list",false,[146,148,150,152,162,164],[147],"Every run is a prediction object with inputs, outputs, status, timing metrics and a cancel URL, whatever the model does.",[149],"Per-second billing on the published rates: $0.000025 for a small CPU, $0.000225 for a T4, $0.001400 for an A100 80GB, $0.001525 for an H100.",[151],"Official models bill per output instead: flux-1.1-pro at $0.04 per image, ideogram-v3-quality at $0.09, deepseek-r1 at $3.75 per million input tokens.",[153,157,158,161],{"tag":154,"children":155},"code",[156],"Asynchronous"," by default; synchronous when you send the ",{"tag":154,"children":159},[160],"Prefer: wait"," header, which holds the request open for 60 seconds.",[163],"Prepaid credit rather than a subscription. Credit is valid for one year from purchase and is not refundable.",[165],"Apache-2.0 clients for Node.js, Python, Swift and Go, plus a hosted MCP server at mcp.replicate.com that tracks the HTTP API.",{"type":137,"level":138,"id":99,"text":100},{"type":121,"content":168},[169,170,173],"Creating a prediction returns an object immediately with the status ",{"tag":154,"children":171},[172],"starting",", and that object stays the unit of work for the rest of the run. Three things then decide how long you wait: whether the model is warm, how long the model itself takes, and whether you poll, hold the HTTP connection, or wait for a webhook.",{"type":175,"attrs":176,"inner":180,"caption":181},"diagram",{"viewBox":177,"role":178,"aria-labelledby":179},"0 0 720 320","img","d-rep-t d-rep-d","\u003Ctitle id=\"d-rep-t\">The Replicate prediction lifecycle\u003C\u002Ftitle>\u003Cdesc id=\"d-rep-d\">A prediction is created and moves from starting to processing. Processing ends in one of three terminal states: succeeded, canceled after a deadline expired while the model was running, or failed. A deadline that expires before the model starts ends it as aborted, which is not charged.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"rep-arrow\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Crect x=\"15\" y=\"40\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"90\" y=\"68\" text-anchor=\"middle\" class=\"d-text\">prediction\u003C\u002Ftext>\u003Ctext x=\"90\" y=\"90\" text-anchor=\"middle\" class=\"d-small\">input JSON\u003C\u002Ftext>\u003Crect x=\"195\" y=\"40\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"270\" y=\"68\" text-anchor=\"middle\" class=\"d-text\">starting\u003C\u002Ftext>\u003Ctext x=\"270\" y=\"90\" text-anchor=\"middle\" class=\"d-small\">cold boot?\u003C\u002Ftext>\u003Crect x=\"375\" y=\"40\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"450\" y=\"68\" text-anchor=\"middle\" class=\"d-text\">processing\u003C\u002Ftext>\u003Ctext x=\"450\" y=\"90\" text-anchor=\"middle\" class=\"d-small\">predict() runs\u003C\u002Ftext>\u003Crect x=\"555\" y=\"40\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"630\" y=\"68\" text-anchor=\"middle\" class=\"d-text\">succeeded\u003C\u002Ftext>\u003Ctext x=\"630\" y=\"90\" text-anchor=\"middle\" class=\"d-small\">output plus metrics\u003C\u002Ftext>\u003Crect x=\"195\" y=\"220\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"270\" y=\"248\" text-anchor=\"middle\" class=\"d-text\">aborted\u003C\u002Ftext>\u003Ctext x=\"270\" y=\"270\" text-anchor=\"middle\" class=\"d-small\">not charged\u003C\u002Ftext>\u003Crect x=\"375\" y=\"220\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"450\" y=\"248\" text-anchor=\"middle\" class=\"d-text\">canceled\u003C\u002Ftext>\u003Ctext x=\"450\" y=\"270\" text-anchor=\"middle\" class=\"d-small\">billed up to here\u003C\u002Ftext>\u003Crect x=\"555\" y=\"220\" width=\"150\" height=\"64\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"630\" y=\"248\" text-anchor=\"middle\" class=\"d-text\">failed\u003C\u002Ftext>\u003Ctext x=\"630\" y=\"270\" text-anchor=\"middle\" class=\"d-small\">error payload\u003C\u002Ftext>\u003Cpath d=\"M165 72H193\" class=\"d-line\" marker-end=\"url(#rep-arrow)\" \u002F>\u003Cpath d=\"M345 72H373\" class=\"d-line\" marker-end=\"url(#rep-arrow)\" \u002F>\u003Cpath d=\"M525 72H553\" class=\"d-line\" marker-end=\"url(#rep-arrow)\" \u002F>\u003Cpath d=\"M270 104V216\" class=\"d-line d-dash\" marker-end=\"url(#rep-arrow)\" \u002F>\u003Ctext x=\"282\" y=\"166\" class=\"d-label\">deadline first\u003C\u002Ftext>\u003Cpath d=\"M450 104V216\" class=\"d-line\" marker-end=\"url(#rep-arrow)\" \u002F>\u003Ctext x=\"462\" y=\"166\" class=\"d-label\">cancel\u003C\u002Ftext>\u003Cpath d=\"M510 104V172H630V216\" class=\"d-line\" marker-end=\"url(#rep-arrow)\" \u002F>\u003Ctext x=\"574\" y=\"162\" text-anchor=\"middle\" class=\"d-label\">error\u003C\u002Ftext>\u003Cpath d=\"M195 252H90V108\" class=\"d-line d-dash\" marker-end=\"url(#rep-arrow)\" \u002F>\u003Ctext x=\"100\" y=\"186\" class=\"d-label\">no bill\u003C\u002Ftext>",[182],"A prediction is created, starts up, runs, and ends in one of four terminal states. The distinction that matters for cost is between canceled, which is billed for the time it ran, and aborted, which is not billed at all.",{"type":121,"content":184},[185,186,188,189,192,193,196,197,200,201,200,204,207,208,211],"The statuses are few and they carry the operational information. ",{"tag":154,"children":187},[172]," normally lasts well under a second; when it persists, it is a cold boot, and loading several gigabytes of weights can take minutes. ",{"tag":154,"children":190},[191],"processing"," is the model's own ",{"tag":154,"children":194},[195],"predict()"," method. Then one of ",{"tag":154,"children":198},[199],"succeeded",", ",{"tag":154,"children":202},[203],"failed",{"tag":154,"children":205},[206],"canceled"," or ",{"tag":154,"children":209},[210],"aborted"," ends the run, and the difference between the last two is whether the machine ever started.",{"type":121,"content":213},[214],"Predictions time out after 30 minutes, and a per-prediction deadline lets an application give up earlier. The billing rule follows the split exactly: a prediction aborted before it started is not charged, a canceled one is billed for the seconds it did run. That is a genuinely good design, and it is also the only reason a deadline is safe to set aggressively.",{"type":137,"level":138,"id":102,"text":103},{"type":121,"content":217},[218,219,222],"The Node.js client collapses the lifecycle into ",{"tag":154,"children":220},[221],"run()",". For a model that finishes in seconds, that is the entire integration and there is nothing else to write.",{"type":154,"code":224},"import Replicate from \"replicate\";\nimport { writeFile } from \"node:fs\u002Fpromises\";\n\n\u002F\u002F Reads REPLICATE_API_TOKEN from the environment.\nconst replicate = new Replicate();\n\nconst [image] = await replicate.run(\"black-forest-labs\u002Fflux-schnell\", {\n  input: {\n    prompt: \"An astronaut riding a rainbow unicorn, cinematic lighting\",\n  },\n});\n\nawait writeFile(\"output.png\", image);\nconsole.log(\"saved output.png\");",{"type":121,"content":226},[227],"Longer models need the asynchronous path, and a webhook is the only sane way to learn the outcome. Two details catch people out. The signature covers the raw request body, so verification has to run before any JSON parsing. And the client converts a byte body back into a string before you hand it to the HMAC, which would silently break every signature check.",{"type":154,"code":229},"import express from \"express\";\nimport Replicate from \"replicate\";\nimport { createHmac, timingSafeEqual } from \"node:crypto\";\n\nconst replicate = new Replicate();\nconst app = express();\n\n\u002F\u002F Verify first, parse second: the signature covers the raw body bytes.\napp.post(\"\u002Fwebhook\", express.raw({ type: \"application\u002Fjson\" }), (req, res) => {\n  const key = process.env.REPLICATE_WEBHOOK_KEY.replace(\"whsec_\", \"\");\n  const signed = `${req.headers[\"webhook-id\"]}.${req.headers[\"webhook-timestamp\"]}.${req.body}`;\n  const expected = createHmac(\"sha256\", key).update(signed).digest(\"base64\");\n  const seen = String(req.headers[\"webhook-signature\"]).split(\" \").map((p) => p.split(\",\")[1]);\n  const valid = seen.some((sig) => {\n    const a = Buffer.from(sig), b = Buffer.from(expected);\n    return a.length === b.length && timingSafeEqual(a, b);\n  });\n  if (!valid) return res.status(401).end();\n\n  const prediction = JSON.parse(String(req.body));\n  res.sendStatus(204); \u002F\u002F Replicate retries terminal webhooks on 4xx and 5xx.\n  console.log(prediction.id, prediction.status);\n});\n\napp.post(\"\u002Fimage\", async (req, res) => {\n  const prediction = await replicate.predictions.create({\n    model: \"black-forest-labs\u002Fflux-schnell\",\n    input: { prompt: req.body.prompt },\n    webhook: \"https:\u002F\u002Fexample.com\u002Fwebhook\",\n    webhook_events_filter: [\"completed\"],\n  });\n  res.json({ id: prediction.id });\n});",{"type":231,"variant":232,"title":233,"body":234},"callout","note","Holding the connection open is the cheapest integration and the worst one to build on",[235],[236,238,239,242,243,207,245,247],{"tag":154,"children":237},[160]," holds the request for 60 seconds by default and the header takes a value, so ",{"tag":154,"children":240},[241],"Prefer: wait=5"," gives a shorter budget. Past that the response comes back with the status still set to ",{"tag":154,"children":244},[172],{"tag":154,"children":246},[191]," and you are polling anyway.",{"type":137,"level":138,"id":105,"text":106},{"type":121,"content":250},[251],"The public catalogue meters compute by the second, and the rate belongs to the hardware rather than the model. The published table also converts each rate to an hourly figure, which is useful for comparison but is not what you are billed on. Setup time and idle time on shared capacity are free; only time spent processing is charged.",{"type":253,"head":254,"rows":263},"table",[255,257,259,261],[256],"Hardware",[258],"Per second",[260],"Per hour",[262],"Memory",[264,275,286,295,304,313],[265,269,271,273],[266],{"tag":154,"children":267},[268],"cpu-small",[270],"$0.000025",[272],"$0.09",[274],"2GB RAM, 1 vCPU",[276,280,282,284],[277],{"tag":154,"children":278},[279],"cpu",[281],"$0.000100",[283],"$0.36",[285],"8GB RAM, 4 vCPU",[287,289,291,293],[288],"Nvidia T4",[290],"$0.000225",[292],"$0.81",[294],"16GB VRAM",[296,298,300,302],[297],"Nvidia L40S",[299],"$0.000975",[301],"$3.51",[303],"48GB VRAM",[305,307,309,311],[306],"Nvidia A100 80GB",[308],"$0.001400",[310],"$5.04",[312],"80GB VRAM",[314,316,318,320],[315],"Nvidia H100",[317],"$0.001525",[319],"$5.49",[312],{"type":121,"content":322},[323],"Official models break the pattern, and that is where the bill becomes predictable. flux-1.1-pro is $0.04 per output image, flux-schnell is $3.00 per thousand images, ideogram-v3-quality is $0.09, and deepseek-r1 is priced per token. Any private model flips it back: it runs on dedicated hardware and you pay for setup, idle and active time alike, so an instance nobody calls all afternoon still costs money.",{"type":121,"content":325},[326],"Payment is prepaid. Credit is bought in advance, drawn down as you spend, valid for one year and non-refundable. When the balance reaches zero, running infrastructure is shut down and no new work starts. A prediction that overruns the balance is charged to the payment method at the end of the month. Auto-reload exists and the documented floor is a $5 threshold with a $15 reload, which is the cheapest guard against a queue of throttled requests.",{"type":137,"level":328,"id":329,"text":330},3,"deployments","Deployments and cold boots",{"type":121,"content":332},[333],"A deployment pins a model version to a hardware type and an instance range, and it answers both the cold-boot problem and the idle-cost problem at once. The price is explicit: a T4 kept warm at $0.81 an hour is $583 for a thirty-day month that never idles, against nothing at all for the same model called a hundred times a day on shared capacity. Deployment is the moment a Replicate bill stops being elastic and starts being a line item.",{"type":143,"ordered":144,"items":335},[336,338,344,346],[337],"Pin the version you tested. A new model version cannot then change your behaviour under load.",[339,340,343],"Set ",{"tag":154,"children":341},[342],"min_instances"," to one and the cold boot, including the multi-minute weight load, disappears.",[345],"Dedicated hardware means no queue shared with other users, at the price of paying for every idle hour.",[347],"The same split separates official models, which Replicate maintains and keeps warm, from community models, which their authors maintain and which may cold boot.",{"type":137,"level":138,"id":108,"text":109},{"type":121,"content":350},[351],"Everything sent through the API is transient by default. For predictions created through the API, the documentation states that input parameters, output values, output files and logs are all removed after an hour, and that you must save your own copies. Predictions created in the web interface are kept indefinitely. That asymmetry is the single most surprising operational fact in the docs, because the same model can sit under two very different retention regimes depending on which interface used it.",{"type":143,"ordered":144,"items":353},[354,360,362,374,380],[355,356,359],"API tokens are 40-character strings prefixed ",{"tag":154,"children":357},[358],"r8_",", passed as a bearer token, named per environment, and revocable individually from the account page.",[361],"Replicate scans public repositories for exposed tokens and disables compromised ones automatically, with an email explaining what happened. Convenient, and also a way for a false positive to stop production.",[363,364,200,367,131,370,373],"Webhooks carry ",{"tag":154,"children":365},[366],"webhook-id",{"tag":154,"children":368},[369],"webhook-timestamp",{"tag":154,"children":371},[372],"webhook-signature","; the signed content is the three joined by full stops and signed with HMAC-SHA256.",[375,376,379],"The signing secret comes from ",{"tag":154,"children":377},[378],"GET \u002Fv1\u002Fwebhooks\u002Fdefault\u002Fsecret"," and the docs recommend caching it rather than fetching it per delivery.",[381],"Limits are 600 prediction creations a minute and 3,000 requests a minute on every other endpoint, with a hard 429 body that names the reset window.",{"type":231,"variant":383,"title":384,"body":385},"warn","Make the webhook handler idempotent before anything else",[386],[387,388,391],"Terminal webhooks are retried on an exponential backoff with the final attempt about a minute after the prediction completed, and the docs state that identical webhooks can arrive more than once and that ordering is not guaranteed. A handler that assumes exactly-once will either lose jobs or double-count them. Key on ",{"tag":154,"children":389},[390],"prediction.id"," and ignore anything arriving after a terminal status.",{"type":137,"level":138,"id":111,"text":112},{"type":121,"content":394},[395],"The catalogue is the product, and the catalogue is the risk. A community model is a container published by a stranger: the input schema can change between versions, the author can disappear, and Replicate's own documentation says community models may have different levels of stability, documentation and support. Only official models carry a stable API, and pinning a version plus running it in a test is the minimum bar for anything a customer can see.",{"type":253,"head":397,"rows":403},[398,400,401,402],[399],"Dimension",[25],[130],[134],[404,413,422,431],[405,407,409,411],[406],"Billing unit",[408],"Per second of compute; official models per output",[410],"Per second, per image or per 1K tokens",[412],"Per 1M tokens, or per hour of GPU",[414,416,418,420],[415],"Catalogue",[417],"Thousands of community and official models",[419],"A small curated set of optimised models",[421],"Primarily models you deploy yourself",[423,425,427,429],[424],"Bring your own",[426],"Cog containers, public or private",[428],"Deployments on the fal GPU fleet",[430],"Dedicated deployments; VPC and self-host at the top tier",[432,434,436,438],[433],"Serverless GPU rate",[435],"No rental pool; per-second rates only",[437],"H100 from $2.49 an hour on custom deployments",[439],"Quoted per deployment",{"type":121,"content":441},[442,443,446],"The second problem is that the bill is never the bill finance expected. A model at $0.04 an image and a model at $5.49 an hour can both sit behind one careless loop, and the only per-run figure the API gives back is the ",{"tag":154,"children":444},[445],"metrics"," object on a single prediction, which nobody aggregates by default.",{"type":121,"content":448},[449],"Third, the client libraries are uneven. The Node.js client is at 1.4.0 and the Python client at 1.0.7, with a 2.0 line still in beta rather than released. A Python service that wants predictable behaviour should call the HTTP API directly instead of waiting, and the OpenAPI schema and the llms.txt endpoint make that painless.",{"type":137,"level":138,"id":114,"text":115},{"type":121,"content":452},[453],"Replicate is the best value available for one specific problem: running an open model without owning a GPU. It is a poor fit for anything else, and the reason is structural rather than a matter of maturity. Metered-by-second pricing optimises for the vendor's hardware utilisation, and your bill is the noise term in that optimisation.",{"type":143,"ordered":455,"items":456},true,[457,459,461,463,465],[458],"Teams prototyping a new image, video or speech model and needing an answer this week.",[460],"Teams willing to pin a version, hold the input schema stable and wrap the whole thing in a hard spend cap.",[462],"Anyone who wants breadth of models rather than a platform, and accepts that most of that catalogue is unmaintained.",[464],"Do not choose it for an interactive endpoint where tail latency is a product requirement, or for a steady load above a few percent utilisation, where a reserved GPU is simply cheaper.",[466],"Do not choose it when input files carry regulated personal data and a one-hour retention window does not clear your own policy.",{"type":231,"variant":232,"body":468},[469],[470],"The catalogue is a supermarket and a deployment is a lease on a machine. Deciding which one you are building on is most of the work, and the API will not make that decision for you.",{"type":137,"level":138,"id":117,"text":118},{"type":143,"ordered":455,"items":473},[474,477,480,483,486,489,492,495,498,501],[475],{"tag":128,"href":33,"children":476},[32],[478],{"tag":128,"href":36,"children":479},[35],[481],{"tag":128,"href":39,"children":482},[38],[484],{"tag":128,"href":42,"children":485},[41],[487],{"tag":128,"href":45,"children":488},[44],[490],{"tag":128,"href":48,"children":491},[47],[493],{"tag":128,"href":51,"children":494},[50],[496],{"tag":128,"href":54,"children":497},[53],[499],{"tag":128,"href":57,"children":500},[56],[502],{"tag":128,"href":60,"children":503},[59],[505,563,619,666],{"slug":506,"published":507,"minutes":6,"category":7,"tags":508,"keywords":514,"about":522,"sources":526,"cover":556,"og":557,"expertise":63,"locales":558,"lang":65,"title":559,"description":560,"coverAlt":561,"url":525,"pricing":562,"kind":509},"browserbase-stagehand","2026-09-30",[509,510,511,512,513],"Browser automation","Web agents","Headless browsers","Playwright","MCP",[515,516,517,518,519,520,521],"browserbase","stagehand","stagehand vs playwright","browserbase pricing","ai web agent framework","headless browser api","browserbase alternatives",[523],{"name":524,"url":525},"Browserbase","https:\u002F\u002Fwww.browserbase.com",[527,530,533,536,539,542,545,548,551,554],{"title":528,"url":529},"Browserbase pricing","https:\u002F\u002Fwww.browserbase.com\u002Fpricing",{"title":531,"url":532},"Browserbase plans and pricing docs","https:\u002F\u002Fdocs.browserbase.com\u002Fguides\u002Fplans-and-pricing",{"title":534,"url":535},"Stagehand documentation","https:\u002F\u002Fdocs.stagehand.dev\u002F",{"title":537,"url":538},"Stagehand quickstart","https:\u002F\u002Fdocs.stagehand.dev\u002Fv4\u002Ffirst-steps\u002Fquickstart",{"title":540,"url":541},"Stagehand v4 announcement","https:\u002F\u002Fwww.browserbase.com\u002Fblog\u002Fstagehand-v4",{"title":543,"url":544},"Stagehand repository","https:\u002F\u002Fgithub.com\u002Fbrowserbase\u002Fstagehand",{"title":546,"url":547},"Browser Use comparison of the two vendors","https:\u002F\u002Fbrowser-use.com\u002Fposts\u002Fbrowser-use-vs-browserbase",{"title":549,"url":550},"Browser Use pricing","https:\u002F\u002Fwww.browser-use.com\u002Fpricing",{"title":552,"url":553},"Skyvern pricing","https:\u002F\u002Fskyvern.com\u002Fpricing",{"title":512,"url":555},"https:\u002F\u002Fplaywright.dev\u002F","\u002Fimages\u002Fblog\u002Fbrowserbase-stagehand\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fbrowserbase-stagehand\u002Fog.jpg",[65,66,67],"Browserbase and Stagehand reviewed: rented Chrome for AI agents","Browserbase rents managed Chrome by the minute and Stagehand adds natural-language steps on top. What it costs, where the meters run and when plain Playwright wins.","Abstract pipeline artwork for the Browserbase and Stagehand review","From $20 per month · OSS library free",{"slug":564,"published":507,"minutes":565,"category":7,"tags":566,"keywords":571,"about":578,"sources":586,"cover":610,"og":611,"expertise":63,"locales":612,"lang":65,"title":613,"description":614,"coverAlt":615,"url":616,"pricing":617,"kind":618},"webmcp",11,[567,568,569,570,513],"WebMCP","Agent tools","JSON Schema","Chrome",[564,572,573,574,575,576,577],"webmcp vs mcp","document.modelContext registerTool","webmcp declarative api","browser agent tools","webmcp origin trial","webmcp chrome 149",[579,581,584],{"name":567,"url":580},"https:\u002F\u002Fgithub.com\u002Fwebmachinelearning\u002Fwebmcp",{"name":582,"url":583},"Model Context Protocol","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FModel_Context_Protocol",{"name":569,"url":585},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FJSON_Schema",[587,590,593,596,599,602,605,607],{"title":588,"url":589},"Chrome for Developers: WebMCP (get started)","https:\u002F\u002Fdeveloper.chrome.com\u002Fdocs\u002Fai\u002Fwebmcp",{"title":591,"url":592},"Chrome for Developers: Imperative API","https:\u002F\u002Fdeveloper.chrome.com\u002Fdocs\u002Fai\u002Fwebmcp\u002Fimperative-api",{"title":594,"url":595},"Chrome for Developers: Declarative API","https:\u002F\u002Fdeveloper.chrome.com\u002Fdocs\u002Fai\u002Fwebmcp\u002Fdeclarative-api",{"title":597,"url":598},"Chrome for Developers: WebMCP versus MCP","https:\u002F\u002Fdeveloper.chrome.com\u002Fdocs\u002Fai\u002Fwebmcp\u002Fcompare-mcp",{"title":600,"url":601},"Chrome for Developers: WebMCP tool security","https:\u002F\u002Fdeveloper.chrome.com\u002Fdocs\u002Fai\u002Fwebmcp\u002Fsecure-tools",{"title":603,"url":604},"Chrome for Developers: Build effective tools","https:\u002F\u002Fdeveloper.chrome.com\u002Fdocs\u002Fai\u002Fwebmcp\u002Fbuild-tools",{"title":606,"url":580},"WebMCP explainer and specification draft",{"title":608,"url":609},"WebMCP implementation status","https:\u002F\u002Fgithub.com\u002Fwebmachinelearning\u002Fwebmcp\u002Fblob\u002Fmain\u002Fimplementation-status.md","\u002Fimages\u002Fblog\u002Fwebmcp\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fwebmcp\u002Fog.jpg",[65,66,67],"WebMCP: publishing tools instead of pixels","WebMCP lets a page publish typed, callable tools to an in-browser agent. What the standard does, how much of it ships today, and where a plain MCP server is still the better call.","A page registering typed tools with the browser, a browser agent calling one of them, and the tool running inside the page and updating its own interface.","https:\u002F\u002Fgithub.com\u002Fwebmcp","Emerging standard, open","Browser protocol",{"slug":620,"published":621,"minutes":6,"category":7,"tags":622,"keywords":625,"about":633,"sources":637,"cover":659,"og":660,"expertise":63,"locales":661,"lang":65,"title":662,"description":663,"coverAlt":664,"url":635,"pricing":665,"kind":509},"playwright","2026-05-27",[623,509,513,624],"End-to-end testing","Cross-browser",[620,626,627,628,629,630,631,632],"playwright vs selenium","playwright mcp","browser automation testing","playwright test runner","playwright for ai agents","playwright vs cypress","playwright cloud pricing",[634,636],{"name":512,"url":635},"https:\u002F\u002Fplaywright.dev",{"name":582,"url":583},[638,641,644,647,650,653,656],{"title":639,"url":640},"Playwright documentation: installation","https:\u002F\u002Fplaywright.dev\u002Fdocs\u002Fintro",{"title":642,"url":643},"Playwright MCP introduction","https:\u002F\u002Fplaywright.dev\u002Fmcp\u002Fintroduction",{"title":645,"url":646},"Playwright CLI introduction","https:\u002F\u002Fplaywright.dev\u002Fagent-cli\u002Fintroduction",{"title":648,"url":649},"Playwright Test agents","https:\u002F\u002Fplaywright.dev\u002Fdocs\u002Ftest-agents",{"title":651,"url":652},"npm package metadata for @playwright\u002Ftest","https:\u002F\u002Fregistry.npmjs.org\u002F@playwright\u002Ftest\u002Flatest",{"title":654,"url":655},"Microsoft Learn: Playwright Workspaces free trial","https:\u002F\u002Flearn.microsoft.com\u002Fen-us\u002Fazure\u002Fapp-testing\u002Fplaywright-workspaces\u002Fhow-to-try-playwright-workspaces-free",{"title":657,"url":658},"Azure App Testing pricing","https:\u002F\u002Fazure.microsoft.com\u002Fen-us\u002Fpricing\u002Fdetails\u002Fapp-testing\u002F","\u002Fimages\u002Fblog\u002Fplaywright\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fplaywright\u002Fog.jpg",[65,66,67],"Playwright: one browser API for tests, scripts and agents","Playwright drives Chromium, Firefox and WebKit from one Apache-2.0 API: a test runner, an agent-facing CLI and an MCP server. What it costs and where it falls short.","A Playwright test run: a spec file, a runner fanning out to three browser engines, and an agent reading accessibility snapshots.","Apache-2.0 · cloud paid",{"slug":667,"published":621,"minutes":6,"category":7,"tags":668,"keywords":673,"about":680,"sources":687,"cover":711,"og":712,"expertise":63,"locales":713,"lang":65,"title":714,"description":715,"coverAlt":716,"url":717,"pricing":718,"kind":719},"elevenlabs",[669,670,671,672],"Text to speech","Voice API","Speech to text","Voice cloning",[667,674,675,676,677,678,679],"elevenlabs api pricing","elevenlabs vs openai tts","text to speech api comparison","elevenlabs flash latency","voice cloning api","eleven v4 model",[681,684],{"name":682,"url":683},"ElevenLabs","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FElevenLabs",{"name":685,"url":686},"Speech synthesis","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSpeech_synthesis",[688,691,694,697,700,702,705,708],{"title":689,"url":690},"ElevenLabs pricing: plans and credits","https:\u002F\u002Felevenlabs.io\u002Fpricing",{"title":692,"url":693},"ElevenAPI pricing: per-product rates","https:\u002F\u002Felevenlabs.io\u002Fpricing\u002Fapi",{"title":695,"url":696},"ElevenLabs documentation: models","https:\u002F\u002Felevenlabs.io\u002Fdocs\u002Fmodels",{"title":698,"url":699},"API reference: create speech","https:\u002F\u002Felevenlabs.io\u002Fdocs\u002Fapi-reference\u002Ftext-to-speech\u002Fconvert",{"title":701,"url":683},"Wikipedia: ElevenLabs",{"title":703,"url":704},"Amazon Polly pricing","https:\u002F\u002Faws.amazon.com\u002Fpolly\u002Fpricing\u002F",{"title":706,"url":707},"OpenAI API pricing","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fpricing",{"title":709,"url":710},"SILMA AI: text to speech price comparison, August 2026","https:\u002F\u002Fsilma.ai\u002Fblog\u002Ftext-to-speech-price-comparison-august-2026","\u002Fimages\u002Fblog\u002Felevenlabs\u002Fcover.webp","\u002Fimages\u002Fblog\u002Felevenlabs\u002Fog.jpg",[65,66,67],"ElevenLabs: speech synthesis as an API","A review of the ElevenLabs audio API: model lineup, latency figures, credit and per-character pricing, tier-gated formats, and where OpenAI and Amazon Polly win.","ElevenLabs cover artwork with speech API pipeline labels","https:\u002F\u002Felevenlabs.io","Free tier · from $5 per month","Voice and audio API",1791383549046]