Train, fine-tune, and serve open-source models on the fastest inference stack — at 10× lower cost than the hyperscalers. From a hosted endpoint to a dedicated GPU cluster.
Full-stack infrastructure for AI workloads — from serverless inference to dedicated training clusters. Drop-in compatible with the OpenAI client. No vendor lock-in.
Llama, Mistral, Qwen, DeepSeek, Stable Diffusion, FLUX, and every major open-source release — behind a single endpoint that's drop-in compatible with the OpenAI client. Pay only for the tokens you use.
The team behind FlashAttention, RedPajama, and the Together Inference Engine. We publish what we ship, and we ship what scales.
An end-to-end stack that delivers 3-4× higher throughput than vLLM on the same hardware, with sub-100ms first-token latency on Llama 3.3 70B.
30T tokens of curated pretraining data with full provenance. Used to train Mistral, Mixtral, and a generation of open-source frontier models.
The H100-optimized attention kernel that ships in every major training and inference framework. Open source, MIT-licensed.
Pay only for what you generate. No volume tiers, no marketing seats. Switch to a dedicated endpoint when you outgrow serverless.
| MODEL | CONTEXT | INPUT | OUTPUT | STATUS |
|---|---|---|---|---|
Llama 3.3 70B Instruct Turbo meta-llama / 70B |
128K | $0.88/ 1M tok | $0.88/ 1M tok | ON-DEMAND |
DeepSeek V3 deepseek-ai / 671B MoE |
64K | $1.25/ 1M tok | $1.25/ 1M tok | ON-DEMAND |
Mixtral 8×22B Instruct mistralai / 141B MoE |
64K | $1.20/ 1M tok | $1.20/ 1M tok | ON-DEMAND |
Qwen 2.5 72B Instruct alibaba / 72B |
32K | $0.90/ 1M tok | $0.90/ 1M tok | ON-DEMAND |
Llama 3.1 8B Instruct Turbo meta-llama / 8B |
128K | $0.18/ 1M tok | $0.18/ 1M tok | ON-DEMAND |
How we cut tail latency on the most-served open model by 38% with speculative decoding, prefix caching, and a redesigned scheduler.
A look inside Pika's training pipeline — 32-node clusters, gradient checkpointing, and the schedule that lets the research team prioritize without breaking SRE budgets.
Tokens and primitives behind the Together surface — colors grouped by role, the display-sans + uppercase-mono ladder, the lightly-rounded radius scale, and the canonical component set. Source: DESIGN.md. Dark hero, white middle, gradient ribbon as the single piece of chrome.