MyAi vs Replicate
Different categories. Replicate is the everything-store for ML inference, billed by GPU time. MyAi is laser-focused on OpenAI-compatible chat — cheaper, faster cold starts, predictable per-token pricing.
| Capability | MyAi | Replicate |
|---|---|---|
| Pricing model | Per token (predictable) | Per second of GPU time (variable) |
| API shape | OpenAI-compatible /v1/chat/completions | Replicate /predictions API (custom) |
| Migration from OpenAI | One env var | Code rewrite (different request/response shape) |
| Models | 22 chat / vision / coder / embedding (open weights) | 10,000+ community models (any modality, including chat) |
| Cold start | <200ms cache; ~3s cold via warm operators | 10–60s for uncached models |
| Streaming | SSE chat streaming on every request | Yes for select models |
| Settlement | $MYAI on Base, USD via Stripe | USD only |
| Best for | Agents that need cheap, predictable, OpenAI-shaped chat inference | Anyone running diffusion / video / niche modalities or one-off model experiments |
Use MyAi when
- You're building an LLM agent or chatbot
- You want OpenAI-compatible streaming
- You need predictable per-token cost (not GPU-seconds)
Stay on Replicate when
- You're running diffusion, video, or any non-LLM model
- You want fine-grained control over a specific model's hardware
- You're prototyping community models that aren't on our network yet
uvx myai-mcp