Together AI

www.together.ai/

Together AI is an AI-native cloud that runs 100+ open-source models — Llama, DeepSeek, Qwen, Mistral — via a fast serverless inference API with OpenAI-compatible endpoints, fine-tuning, and dedicated GPU clusters.

Overview

Together AI is an AI-native cloud founded in 2022 and headquartered in San Francisco, engineered for developers building with open-source and frontier models. It provides serverless inference, fine-tuning, and GPU clusters on a custom inference stack that the company says delivers 3.5x faster inference and 2.3x faster training than alternatives, with Llama 3.1 8B running at 200+ tokens/second and 70B models at 80+ tok/s. The catalog spans 100+ production-ready open-source models — Llama, DeepSeek, Qwen, Mistral, Gemma and more — accessible through OpenAI-compatible endpoints for drop-in migration. Serverless inference is billed per token with input and output priced transparently, batch mode cuts costs 50% for asynchronous workloads up to 30 billion tokens per model, and dedicated endpoints start around $0.85/hour with SLA-backed latency. A fine-tuning pipeline supports LoRA and full tuning that trains then deploys to serverless in a single workflow without weight export. It also offers self-service and reserved GPU clusters (H100/H200/B200) and a code interpreter, positioning it as a full runtime rather than just a model host.

Key Features

  • Serverless inference for 100+ open-source LLMs (Llama, DeepSeek, Qwen, Mistral, Gemma) via OpenAI-compatible endpoints — drop-in replacement
  • Custom inference engine delivering 3.5x faster inference than comparable stacks, with Llama 3.1 8B at 200+ tokens/sec
  • Transparent per-token pricing — you know the exact cost before sending a request, no per-second GPU math
  • Fine-tuning pipeline (LoRA and full) that trains then deploys to serverless in one workflow, no weight export
  • Batch API at 50% off for asynchronous workloads up to 30B tokens per model — ideal for RAG prep and bulk eval
  • Dedicated endpoints from $0.85/hr with SLA-backed latency, plus self-service and reserved GPU clusters (H100/H200/B200)

Pricing

Model tierInput / OutputNotes
Llama 3.1 8B$0.18 / 1M tokensBudget, low-latency, high-volume
Llama 3.3 70B~$0.88 / 1M tokensMid-tier general purpose, 128K ctx
DeepSeek-R1$3.00 / $7.00 / 1M tokensReasoning tier
Dedicated endpointfrom $0.85 / hrGuaranteed capacity, no cold start

Token prices shift often; confirm on together.ai/pricing before committing volume.

Comparison

vs. Replicate: Replicate’s strength is media generation (image/video/audio) across a huge community catalog; Together AI is built for fast, cost-efficient open-source LLM serving. If your app is chat or agents making many calls, Together is the sharper tool.

vs. Groq: Groq wins raw token speed on custom LPU hardware, but offers far fewer models. Together trades some latency for breadth — 100+ curated models and fine-tuning you can’t get on Groq.

vs. Hugging Face: HF is where models live and get trained; Together is the optimized runtime. Teams often train on HF and serve through Together’s inference stack for production throughput.

Compare alternatives

Side-by-side with the 3 closest alternatives.

ToolCategoryPricingVisit
Together AI (this) codeFrom $0.18/mo Site ↗
Replicatecode, image, videoFree $5 sign-up · From $0.000225/mo Site ↗
Groqcode, searchFree $0 · From $0.04/mo Site ↗
Hugging Facecode, searchFree $0 · From $0.5/mo Site ↗
Together AI Current

Together AI is an AI-native cloud that runs 100+ open-source models — Llama, DeepSeek, Qwen, Mistral — via a fast serverless inference API with OpenAI-compatible endpoints, fine-tuning, and dedicated GPU clusters.

code
From $0.18/mo

Cloud platform to run thousands of open-source ML models through a single API — no GPU infrastructure to manage, pay only for compute time.

codeimagevideo
Free $5 sign-up · From $0.000225/mo

Ultra-low-latency LLM and speech inference API running on custom LPU hardware, built for real-time AI apps. Read our hands-on Groq review and compare the b

codesearch
Free $0 · From $0.04/mo

The central hub for open ML — host and run a million-plus models, datasets, and Spaces, with serverless and dedicated inference.

codesearch
Free $0 · From $0.5/mo
Editor’s Review
4.6/5
Pros
  • +3.5x faster inference than comparable stacks on a custom NVIDIA GPU engine
  • +100+ curated open-source models behind OpenAI-compatible endpoints
  • +Transparent per-token pricing — exact cost known before each request
  • +Fine-tuning (LoRA and full) trains then deploys to serverless in one workflow
Cons
  • Limited to a curated catalog — niche models need fine-tuning from a supported base
  • Text/code focused; weaker than Replicate for video, audio, and 3D media
  • Free credit is small and runs out fast on large models; then pay-per-use only

Together AI is my default for serving open-source LLMs in production — the speed is measurable and the per-token pricing makes budgeting trivial. It is not the place for media generation or exotic models, but for chat and agent workloads on Llama or DeepSeek it is hard to beat.

See all reviews →