State-of-the-art text-to-image with painterly fidelity.
Replicate is the easiest way to run open-source machine learning models in the cloud. Pay only for the compute you use.
Generate images, edit video, transcribe speech, build agents — every model on Replicate ships behind the same predictable HTTP endpoint.
State-of-the-art text-to-image with painterly fidelity.
High-quality instruction-tuned 70B chat model.
Multilingual speech-to-text with word-level timestamps.
Fast image synthesis in four denoise steps.
No infra to provision. No Docker images to push. No GPUs to babysit. Pass inputs to the model slug, await the result.
Autoscale to thousands of GPUs, queue spikes gracefully, pay only for compute used per second. Replicate handles the cold-start so your inference looks instant.
Warm boot any model on demand.
GPUs spin up per request.
Pay only for what you use.
No seats, no minimums. The free tier covers exploration; production scales by the second.
For hackers, hobbyists, and weekend ideas.
For teams shipping AI into production.
For organizations with bespoke compute needs.
An honest look at the building blocks — cream canvas, hot-orange stamp, three families, full-rounded interactives. No surprises, no second accent.
→ Read the full DESIGN.md