Tools/Web engineering

Replicate, reviewed: a model API priced per second, with the sharp edges named

Replicate puts thousands of open models behind one prediction API and bills per second of compute. A review of cold boots, version churn and the one-hour data deletion.

Type
Model hosting API
Pricing
Pay per second of compute

··10 min read

  • Model hosting
  • Serverless GPU
  • Diffusion
  • Async jobs
  • Webhooks
A request moves from starting to processing, then into one of four terminal states: succeeded, failed, canceled or aborted.

Key takeaways

  • Replicate meters public models per second of hardware time and bills nothing while a shared model sits idle, which makes it cheap for bursty workloads and expensive for anything latency-sensitive.
  • Every run is a prediction object with a fixed set of statuses, and the difference between canceled and aborted decides whether you pay for the compute that ran.
  • Predictions created through the API are deleted an hour later, inputs and outputs included, so anything you need has to be copied out before then.
  • Only official models carry a stable input schema; community models are maintained by their authors and need a pinned version plus a test.
  • A deployment pins a version to hardware and an instance range, which removes cold boots and replaces them with a fixed hourly bill.

Replicate is a hosted inference platform that puts thousands of open-weight and proprietary models behind a single prediction API, and it bills only for the seconds a machine spends running your request. The position here is plain: for a team that wants to try a new image or video model this afternoon without provisioning a GPU, it is the shortest path on the market. For a production endpoint with a latency budget and a predictable invoice, per-second billing is the wrong contract, and the platform only becomes defensible once a version is pinned and the model sits on a deployment.

It sits between owning a GPU and buying from a model vendor. On one side it competes with fal.ai and Baseten, which sell inference rather than a catalogue; on the other it competes with renting a machine from a GPU broker and running your own container. What it does not replace is a token-billed LLM API. Replicate will happily serve a Llama model on an H100, but nobody prices a chat endpoint per second of GPU time.

What Replicate actually is

Three things share one account and one API. First, a public catalogue of models published by their authors: fine-tunes, community implementations, research code. Second, a set of official models that Replicate maintains itself, which the documentation puts at over a hundred, and which are always warm, priced per output unit, and covered by a stable input schema. Third, private models that you package with Cog and run on dedicated hardware, where you pay for the whole life of the instance rather than for the work it does.

  • Every run is a prediction object with inputs, outputs, status, timing metrics and a cancel URL, whatever the model does.
  • Per-second billing on the published rates: $0.000025 for a small CPU, $0.000225 for a T4, $0.001400 for an A100 80GB, $0.001525 for an H100.
  • Official models bill per output instead: flux-1.1-pro at $0.04 per image, ideogram-v3-quality at $0.09, deepseek-r1 at $3.75 per million input tokens.
  • Asynchronous by default; synchronous when you send the Prefer: wait header, which holds the request open for 60 seconds.
  • Prepaid credit rather than a subscription. Credit is valid for one year from purchase and is not refundable.
  • Apache-2.0 clients for Node.js, Python, Swift and Go, plus a hosted MCP server at mcp.replicate.com that tracks the HTTP API.

How a prediction runs

Creating a prediction returns an object immediately with the status starting, and that object stays the unit of work for the rest of the run. Three things then decide how long you wait: whether the model is warm, how long the model itself takes, and whether you poll, hold the HTTP connection, or wait for a webhook.

The Replicate prediction lifecycleA prediction is created and moves from starting to processing. Processing ends in one of three terminal states: succeeded, canceled after a deadline expired while the model was running, or failed. A deadline that expires before the model starts ends it as aborted, which is not charged.predictioninput JSONstartingcold boot?processingpredict() runssucceededoutput plus metricsabortednot chargedcanceledbilled up to herefailederror payloaddeadline firstcancelerrorno bill
A prediction is created, starts up, runs, and ends in one of four terminal states. The distinction that matters for cost is between canceled, which is billed for the time it ran, and aborted, which is not billed at all.

The statuses are few and they carry the operational information. starting normally lasts well under a second; when it persists, it is a cold boot, and loading several gigabytes of weights can take minutes. processing is the model's own predict() method. Then one of succeeded, failed, canceled or aborted ends the run, and the difference between the last two is whether the machine ever started.

Predictions time out after 30 minutes, and a per-prediction deadline lets an application give up earlier. The billing rule follows the split exactly: a prediction aborted before it started is not charged, a canceled one is billed for the seconds it did run. That is a genuinely good design, and it is also the only reason a deadline is safe to set aggressively.

Getting started

The Node.js client collapses the lifecycle into run(). For a model that finishes in seconds, that is the entire integration and there is nothing else to write.

import Replicate from "replicate";
import { writeFile } from "node:fs/promises";

// Reads REPLICATE_API_TOKEN from the environment.
const replicate = new Replicate();

const [image] = await replicate.run("black-forest-labs/flux-schnell", {
  input: {
    prompt: "An astronaut riding a rainbow unicorn, cinematic lighting",
  },
});

await writeFile("output.png", image);
console.log("saved output.png");

Longer models need the asynchronous path, and a webhook is the only sane way to learn the outcome. Two details catch people out. The signature covers the raw request body, so verification has to run before any JSON parsing. And the client converts a byte body back into a string before you hand it to the HMAC, which would silently break every signature check.

import express from "express";
import Replicate from "replicate";
import { createHmac, timingSafeEqual } from "node:crypto";

const replicate = new Replicate();
const app = express();

// Verify first, parse second: the signature covers the raw body bytes.
app.post("/webhook", express.raw({ type: "application/json" }), (req, res) => {
  const key = process.env.REPLICATE_WEBHOOK_KEY.replace("whsec_", "");
  const signed = `${req.headers["webhook-id"]}.${req.headers["webhook-timestamp"]}.${req.body}`;
  const expected = createHmac("sha256", key).update(signed).digest("base64");
  const seen = String(req.headers["webhook-signature"]).split(" ").map((p) => p.split(",")[1]);
  const valid = seen.some((sig) => {
    const a = Buffer.from(sig), b = Buffer.from(expected);
    return a.length === b.length && timingSafeEqual(a, b);
  });
  if (!valid) return res.status(401).end();

  const prediction = JSON.parse(String(req.body));
  res.sendStatus(204); // Replicate retries terminal webhooks on 4xx and 5xx.
  console.log(prediction.id, prediction.status);
});

app.post("/image", async (req, res) => {
  const prediction = await replicate.predictions.create({
    model: "black-forest-labs/flux-schnell",
    input: { prompt: req.body.prompt },
    webhook: "https://example.com/webhook",
    webhook_events_filter: ["completed"],
  });
  res.json({ id: prediction.id });
});

What it costs

The public catalogue meters compute by the second, and the rate belongs to the hardware rather than the model. The published table also converts each rate to an hourly figure, which is useful for comparison but is not what you are billed on. Setup time and idle time on shared capacity are free; only time spent processing is charged.

HardwarePer secondPer hourMemory
cpu-small$0.000025$0.092GB RAM, 1 vCPU
cpu$0.000100$0.368GB RAM, 4 vCPU
Nvidia T4$0.000225$0.8116GB VRAM
Nvidia L40S$0.000975$3.5148GB VRAM
Nvidia A100 80GB$0.001400$5.0480GB VRAM
Nvidia H100$0.001525$5.4980GB VRAM

Official models break the pattern, and that is where the bill becomes predictable. flux-1.1-pro is $0.04 per output image, flux-schnell is $3.00 per thousand images, ideogram-v3-quality is $0.09, and deepseek-r1 is priced per token. Any private model flips it back: it runs on dedicated hardware and you pay for setup, idle and active time alike, so an instance nobody calls all afternoon still costs money.

Payment is prepaid. Credit is bought in advance, drawn down as you spend, valid for one year and non-refundable. When the balance reaches zero, running infrastructure is shut down and no new work starts. A prediction that overruns the balance is charged to the payment method at the end of the month. Auto-reload exists and the documented floor is a $5 threshold with a $15 reload, which is the cheapest guard against a queue of throttled requests.

Deployments and cold boots

A deployment pins a model version to a hardware type and an instance range, and it answers both the cold-boot problem and the idle-cost problem at once. The price is explicit: a T4 kept warm at $0.81 an hour is $583 for a thirty-day month that never idles, against nothing at all for the same model called a hundred times a day on shared capacity. Deployment is the moment a Replicate bill stops being elastic and starts being a line item.

  • Pin the version you tested. A new model version cannot then change your behaviour under load.
  • Set min_instances to one and the cold boot, including the multi-minute weight load, disappears.
  • Dedicated hardware means no queue shared with other users, at the price of paying for every idle hour.
  • The same split separates official models, which Replicate maintains and keeps warm, from community models, which their authors maintain and which may cold boot.

Data retention, tokens and webhooks

Everything sent through the API is transient by default. For predictions created through the API, the documentation states that input parameters, output values, output files and logs are all removed after an hour, and that you must save your own copies. Predictions created in the web interface are kept indefinitely. That asymmetry is the single most surprising operational fact in the docs, because the same model can sit under two very different retention regimes depending on which interface used it.

  • API tokens are 40-character strings prefixed r8_, passed as a bearer token, named per environment, and revocable individually from the account page.
  • Replicate scans public repositories for exposed tokens and disables compromised ones automatically, with an email explaining what happened. Convenient, and also a way for a false positive to stop production.
  • Webhooks carry webhook-id, webhook-timestamp and webhook-signature; the signed content is the three joined by full stops and signed with HMAC-SHA256.
  • The signing secret comes from GET /v1/webhooks/default/secret and the docs recommend caching it rather than fetching it per delivery.
  • Limits are 600 prediction creations a minute and 3,000 requests a minute on every other endpoint, with a hard 429 body that names the reset window.

Where it hurts

The catalogue is the product, and the catalogue is the risk. A community model is a container published by a stranger: the input schema can change between versions, the author can disappear, and Replicate's own documentation says community models may have different levels of stability, documentation and support. Only official models carry a stable API, and pinning a version plus running it in a test is the minimum bar for anything a customer can see.

DimensionReplicatefal.aiBaseten
Billing unitPer second of compute; official models per outputPer second, per image or per 1K tokensPer 1M tokens, or per hour of GPU
CatalogueThousands of community and official modelsA small curated set of optimised modelsPrimarily models you deploy yourself
Bring your ownCog containers, public or privateDeployments on the fal GPU fleetDedicated deployments; VPC and self-host at the top tier
Serverless GPU rateNo rental pool; per-second rates onlyH100 from $2.49 an hour on custom deploymentsQuoted per deployment

The second problem is that the bill is never the bill finance expected. A model at $0.04 an image and a model at $5.49 an hour can both sit behind one careless loop, and the only per-run figure the API gives back is the metrics object on a single prediction, which nobody aggregates by default.

Third, the client libraries are uneven. The Node.js client is at 1.4.0 and the Python client at 1.0.7, with a 2.0 line still in beta rather than released. A Python service that wants predictable behaviour should call the HTTP API directly instead of waiting, and the OpenAPI schema and the llms.txt endpoint make that painless.

Verdict

Replicate is the best value available for one specific problem: running an open model without owning a GPU. It is a poor fit for anything else, and the reason is structural rather than a matter of maturity. Metered-by-second pricing optimises for the vendor's hardware utilisation, and your bill is the noise term in that optimisation.

  1. Teams prototyping a new image, video or speech model and needing an answer this week.
  2. Teams willing to pin a version, hold the input schema stable and wrap the whole thing in a hard spend cap.
  3. Anyone who wants breadth of models rather than a platform, and accepts that most of that catalogue is unmaintained.
  4. Do not choose it for an interactive endpoint where tail latency is a product requirement, or for a steady load above a few percent utilisation, where a reserved GPU is simply cheaper.
  5. Do not choose it when input files carry regulated personal data and a one-hour retention window does not clear your own policy.

Sources

  1. Replicate pricing
  2. Replicate docs: About predictions
  3. Replicate docs: Create a prediction
  4. Replicate docs: Prediction lifecycle
  5. Replicate docs: Rate limits
  6. Replicate docs: Data retention
  7. Replicate docs: Verify webhooks
  8. Replicate docs: Official models
  9. fal.ai pricing
  10. Baseten pricing

Frequently asked questions

What does Replicate actually charge for?

For models in the public catalogue, the price per second of the hardware the model runs on, starting at $0.000025 for a small CPU and reaching $0.001525 for an H100. Official models are the exception: they are billed per output image, per second of video or per token. Private models run on dedicated hardware, so you pay for setup and idle time as well as for the work.

Is Replicate cheaper than renting a GPU?

For spiky, low-duty-cycle work, yes, because you pay only while a prediction runs and shared capacity is reused between customers. For steady load above roughly a few percent utilisation, the arithmetic flips: a warm T4 instance on Replicate is $0.81 an hour, which is $583 for a 30-day month that never idles, and at that point a reserved instance on RunPod or Lambda is the cheaper contract.

Why does my Replicate prediction take minutes to start?

That is a cold boot. Replicate shuts down models that are not being used, and restarting one means loading several gigabytes of weights, which the documentation says can take several minutes. Official models are kept warm, and a deployment with min_instances set to 1 removes the wait for anything else.

Does Replicate keep my inputs and outputs?

Not for API predictions. Input parameters, output values, output files and logs are all removed an hour after the prediction is created, and the docs say you have to save your own copies. Predictions started from the web interface are kept indefinitely, which is worth knowing if the same model is used both ways.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.