Tools/Web engineering
Replicate, reviewed: a model API priced per second, with the sharp edges named
Replicate puts thousands of open models behind one prediction API and bills per second of compute. A review of cold boots, version churn and the one-hour data deletion.
- Type
- Model hosting API
- Pricing
- Pay per second of compute
Balázs Csorba··10 min read
- Model hosting
- Serverless GPU
- Diffusion
- Async jobs
- Webhooks

Key takeaways
- Replicate meters public models per second of hardware time and bills nothing while a shared model sits idle, which makes it cheap for bursty workloads and expensive for anything latency-sensitive.
- Every run is a prediction object with a fixed set of statuses, and the difference between canceled and aborted decides whether you pay for the compute that ran.
- Predictions created through the API are deleted an hour later, inputs and outputs included, so anything you need has to be copied out before then.
- Only official models carry a stable input schema; community models are maintained by their authors and need a pinned version plus a test.
- A deployment pins a version to hardware and an instance range, which removes cold boots and replaces them with a fixed hourly bill.
Replicate is a hosted inference platform that puts thousands of open-weight and proprietary models behind a single prediction API, and it bills only for the seconds a machine spends running your request. The position here is plain: for a team that wants to try a new image or video model this afternoon without provisioning a GPU, it is the shortest path on the market. For a production endpoint with a latency budget and a predictable invoice, per-second billing is the wrong contract, and the platform only becomes defensible once a version is pinned and the model sits on a deployment.
It sits between owning a GPU and buying from a model vendor. On one side it competes with fal.ai and Baseten, which sell inference rather than a catalogue; on the other it competes with renting a machine from a GPU broker and running your own container. What it does not replace is a token-billed LLM API. Replicate will happily serve a Llama model on an H100, but nobody prices a chat endpoint per second of GPU time.
What Replicate actually is
Three things share one account and one API. First, a public catalogue of models published by their authors: fine-tunes, community implementations, research code. Second, a set of official models that Replicate maintains itself, which the documentation puts at over a hundred, and which are always warm, priced per output unit, and covered by a stable input schema. Third, private models that you package with Cog and run on dedicated hardware, where you pay for the whole life of the instance rather than for the work it does.
- Every run is a prediction object with inputs, outputs, status, timing metrics and a cancel URL, whatever the model does.
- Per-second billing on the published rates: $0.000025 for a small CPU, $0.000225 for a T4, $0.001400 for an A100 80GB, $0.001525 for an H100.
- Official models bill per output instead: flux-1.1-pro at $0.04 per image, ideogram-v3-quality at $0.09, deepseek-r1 at $3.75 per million input tokens.
Asynchronousby default; synchronous when you send thePrefer: waitheader, which holds the request open for 60 seconds.- Prepaid credit rather than a subscription. Credit is valid for one year from purchase and is not refundable.
- Apache-2.0 clients for Node.js, Python, Swift and Go, plus a hosted MCP server at mcp.replicate.com that tracks the HTTP API.
How a prediction runs
Creating a prediction returns an object immediately with the status starting, and that object stays the unit of work for the rest of the run. Three things then decide how long you wait: whether the model is warm, how long the model itself takes, and whether you poll, hold the HTTP connection, or wait for a webhook.
The statuses are few and they carry the operational information. starting normally lasts well under a second; when it persists, it is a cold boot, and loading several gigabytes of weights can take minutes. processing is the model's own predict() method. Then one of succeeded, failed, canceled or aborted ends the run, and the difference between the last two is whether the machine ever started.
Predictions time out after 30 minutes, and a per-prediction deadline lets an application give up earlier. The billing rule follows the split exactly: a prediction aborted before it started is not charged, a canceled one is billed for the seconds it did run. That is a genuinely good design, and it is also the only reason a deadline is safe to set aggressively.
Getting started
The Node.js client collapses the lifecycle into run(). For a model that finishes in seconds, that is the entire integration and there is nothing else to write.
import Replicate from "replicate";
import { writeFile } from "node:fs/promises";
// Reads REPLICATE_API_TOKEN from the environment.
const replicate = new Replicate();
const [image] = await replicate.run("black-forest-labs/flux-schnell", {
input: {
prompt: "An astronaut riding a rainbow unicorn, cinematic lighting",
},
});
await writeFile("output.png", image);
console.log("saved output.png");Longer models need the asynchronous path, and a webhook is the only sane way to learn the outcome. Two details catch people out. The signature covers the raw request body, so verification has to run before any JSON parsing. And the client converts a byte body back into a string before you hand it to the HMAC, which would silently break every signature check.
import express from "express";
import Replicate from "replicate";
import { createHmac, timingSafeEqual } from "node:crypto";
const replicate = new Replicate();
const app = express();
// Verify first, parse second: the signature covers the raw body bytes.
app.post("/webhook", express.raw({ type: "application/json" }), (req, res) => {
const key = process.env.REPLICATE_WEBHOOK_KEY.replace("whsec_", "");
const signed = `${req.headers["webhook-id"]}.${req.headers["webhook-timestamp"]}.${req.body}`;
const expected = createHmac("sha256", key).update(signed).digest("base64");
const seen = String(req.headers["webhook-signature"]).split(" ").map((p) => p.split(",")[1]);
const valid = seen.some((sig) => {
const a = Buffer.from(sig), b = Buffer.from(expected);
return a.length === b.length && timingSafeEqual(a, b);
});
if (!valid) return res.status(401).end();
const prediction = JSON.parse(String(req.body));
res.sendStatus(204); // Replicate retries terminal webhooks on 4xx and 5xx.
console.log(prediction.id, prediction.status);
});
app.post("/image", async (req, res) => {
const prediction = await replicate.predictions.create({
model: "black-forest-labs/flux-schnell",
input: { prompt: req.body.prompt },
webhook: "https://example.com/webhook",
webhook_events_filter: ["completed"],
});
res.json({ id: prediction.id });
});What it costs
The public catalogue meters compute by the second, and the rate belongs to the hardware rather than the model. The published table also converts each rate to an hourly figure, which is useful for comparison but is not what you are billed on. Setup time and idle time on shared capacity are free; only time spent processing is charged.
| Hardware | Per second | Per hour | Memory |
|---|---|---|---|
cpu-small | $0.000025 | $0.09 | 2GB RAM, 1 vCPU |
cpu | $0.000100 | $0.36 | 8GB RAM, 4 vCPU |
| Nvidia T4 | $0.000225 | $0.81 | 16GB VRAM |
| Nvidia L40S | $0.000975 | $3.51 | 48GB VRAM |
| Nvidia A100 80GB | $0.001400 | $5.04 | 80GB VRAM |
| Nvidia H100 | $0.001525 | $5.49 | 80GB VRAM |
Official models break the pattern, and that is where the bill becomes predictable. flux-1.1-pro is $0.04 per output image, flux-schnell is $3.00 per thousand images, ideogram-v3-quality is $0.09, and deepseek-r1 is priced per token. Any private model flips it back: it runs on dedicated hardware and you pay for setup, idle and active time alike, so an instance nobody calls all afternoon still costs money.
Payment is prepaid. Credit is bought in advance, drawn down as you spend, valid for one year and non-refundable. When the balance reaches zero, running infrastructure is shut down and no new work starts. A prediction that overruns the balance is charged to the payment method at the end of the month. Auto-reload exists and the documented floor is a $5 threshold with a $15 reload, which is the cheapest guard against a queue of throttled requests.
Deployments and cold boots
A deployment pins a model version to a hardware type and an instance range, and it answers both the cold-boot problem and the idle-cost problem at once. The price is explicit: a T4 kept warm at $0.81 an hour is $583 for a thirty-day month that never idles, against nothing at all for the same model called a hundred times a day on shared capacity. Deployment is the moment a Replicate bill stops being elastic and starts being a line item.
- Pin the version you tested. A new model version cannot then change your behaviour under load.
- Set
min_instancesto one and the cold boot, including the multi-minute weight load, disappears. - Dedicated hardware means no queue shared with other users, at the price of paying for every idle hour.
- The same split separates official models, which Replicate maintains and keeps warm, from community models, which their authors maintain and which may cold boot.
Data retention, tokens and webhooks
Everything sent through the API is transient by default. For predictions created through the API, the documentation states that input parameters, output values, output files and logs are all removed after an hour, and that you must save your own copies. Predictions created in the web interface are kept indefinitely. That asymmetry is the single most surprising operational fact in the docs, because the same model can sit under two very different retention regimes depending on which interface used it.
- API tokens are 40-character strings prefixed
r8_, passed as a bearer token, named per environment, and revocable individually from the account page. - Replicate scans public repositories for exposed tokens and disables compromised ones automatically, with an email explaining what happened. Convenient, and also a way for a false positive to stop production.
- Webhooks carry
webhook-id,webhook-timestampandwebhook-signature; the signed content is the three joined by full stops and signed with HMAC-SHA256. - The signing secret comes from
GET /v1/webhooks/default/secretand the docs recommend caching it rather than fetching it per delivery. - Limits are 600 prediction creations a minute and 3,000 requests a minute on every other endpoint, with a hard 429 body that names the reset window.
Where it hurts
The catalogue is the product, and the catalogue is the risk. A community model is a container published by a stranger: the input schema can change between versions, the author can disappear, and Replicate's own documentation says community models may have different levels of stability, documentation and support. Only official models carry a stable API, and pinning a version plus running it in a test is the minimum bar for anything a customer can see.
| Dimension | Replicate | fal.ai | Baseten |
|---|---|---|---|
| Billing unit | Per second of compute; official models per output | Per second, per image or per 1K tokens | Per 1M tokens, or per hour of GPU |
| Catalogue | Thousands of community and official models | A small curated set of optimised models | Primarily models you deploy yourself |
| Bring your own | Cog containers, public or private | Deployments on the fal GPU fleet | Dedicated deployments; VPC and self-host at the top tier |
| Serverless GPU rate | No rental pool; per-second rates only | H100 from $2.49 an hour on custom deployments | Quoted per deployment |
The second problem is that the bill is never the bill finance expected. A model at $0.04 an image and a model at $5.49 an hour can both sit behind one careless loop, and the only per-run figure the API gives back is the metrics object on a single prediction, which nobody aggregates by default.
Third, the client libraries are uneven. The Node.js client is at 1.4.0 and the Python client at 1.0.7, with a 2.0 line still in beta rather than released. A Python service that wants predictable behaviour should call the HTTP API directly instead of waiting, and the OpenAPI schema and the llms.txt endpoint make that painless.
Verdict
Replicate is the best value available for one specific problem: running an open model without owning a GPU. It is a poor fit for anything else, and the reason is structural rather than a matter of maturity. Metered-by-second pricing optimises for the vendor's hardware utilisation, and your bill is the noise term in that optimisation.
- Teams prototyping a new image, video or speech model and needing an answer this week.
- Teams willing to pin a version, hold the input schema stable and wrap the whole thing in a hard spend cap.
- Anyone who wants breadth of models rather than a platform, and accepts that most of that catalogue is unmaintained.
- Do not choose it for an interactive endpoint where tail latency is a product requirement, or for a steady load above a few percent utilisation, where a reserved GPU is simply cheaper.
- Do not choose it when input files carry regulated personal data and a one-hour retention window does not clear your own policy.
Sources
Frequently asked questions
What does Replicate actually charge for?
For models in the public catalogue, the price per second of the hardware the model runs on, starting at $0.000025 for a small CPU and reaching $0.001525 for an H100. Official models are the exception: they are billed per output image, per second of video or per token. Private models run on dedicated hardware, so you pay for setup and idle time as well as for the work.
Is Replicate cheaper than renting a GPU?
For spiky, low-duty-cycle work, yes, because you pay only while a prediction runs and shared capacity is reused between customers. For steady load above roughly a few percent utilisation, the arithmetic flips: a warm T4 instance on Replicate is $0.81 an hour, which is $583 for a 30-day month that never idles, and at that point a reserved instance on RunPod or Lambda is the cheaper contract.
Why does my Replicate prediction take minutes to start?
That is a cold boot. Replicate shuts down models that are not being used, and restarting one means loading several gigabytes of weights, which the documentation says can take several minutes. Official models are kept warm, and a deployment with min_instances set to 1 removes the wait for anything else.
Does Replicate keep my inputs and outputs?
Not for API predictions. Input parameters, output values, output files and logs are all removed an hour after the prediction is created, and the docs say you have to save your own copies. Predictions started from the web interface are kept indefinitely, which is worth knowing if the same model is used both ways.