T
Together
AGGREGATEDINFERENCE
N/A
Uptime
N/A
Rating
30-Day Uptime
96.7%2026-07-212026-08-19
Inference Latency
LiquidAI: LFM2-24B-A2B173ms TTFT · 68 TPS
OpenAI: gpt-oss-20b222ms TTFT · 129 TPS
Google: Gemma 3n 4B303ms TTFT · 31 TPS
DeepSeek: DeepSeek V4 Flash 0731827ms TTFT · 41 TPS
Meta: Llama 3 8B Instruct791ms TTFT · 29 TPS
OpenAI: gpt-oss-120b419ms TTFT · 37 TPS
EssentialAI: Rnj 1 Instruct201ms TTFT · 121 TPS
Qwen: Qwen3.5-9B368ms TTFT · 87 TPS
Meta: Llama Guard 4 12B124ms TTFT · 16 TPS
Mistral: Mistral 7B Instruct v0.2315ms TTFT · 78 TPS
MiniMax: MiniMax M3800ms TTFT · 72 TPS
Qwen: Qwen2.5 7B Instruct234ms TTFT · 52 TPS
Meta: Muse Glimmer 30B438ms TTFT · 91 TPS
Google: Gemma 4 31B966ms TTFT · 13 TPS
Thinking Machines: Inkling Small1484ms TTFT · 63 TPS
NVIDIA: Nemotron 3 Ultra755ms TTFT · 89 TPS
MoonshotAI: Kimi K2.7 Code799ms TTFT · 261 TPS
Thinking Machines: Inkling472ms TTFT · 82 TPS
Meta: Llama 3.3 70B Instruct1098ms TTFT · 13 TPS
MoonshotAI: Kimi K2.6386ms TTFT · 99 TPS
Inference Models
| Model | Input $/M | Output $/M | TTFT | TPS |
|---|---|---|---|---|
| LiquidAI: LFM2-24B-A2B | $0.03 | $0.12 | 173ms | 68 |
| OpenAI: gpt-oss-20b | $0.05 | $0.20 | 222ms | 129 |
| Google: Gemma 3n 4B | $0.06 | $0.12 | 303ms | 31 |
| DeepSeek: DeepSeek V4 Flash 0731 | $0.14 | $0.28 | 827ms | 41 |
| Meta: Llama 3 8B Instruct | $0.14 | $0.14 | 791ms | 29 |
| OpenAI: gpt-oss-120b | $0.15 | $0.60 | 419ms | 37 |
| EssentialAI: Rnj 1 Instruct | $0.15 | $0.15 | 201ms | 121 |
| Qwen: Qwen3.5-9B | $0.17 | $0.25 | 368ms | 87 |
| Arcee AI: Spotlight | $0.18 | $0.18 | — | — |
| Meta: Llama Guard 4 12B | $0.20 | $0.20 | 124ms | 16 |
| Mistral: Mistral 7B Instruct | $0.20 | $0.20 | — | — |
| Mistral: Mistral 7B Instruct v0.3 | $0.20 | $0.20 | — | — |
| Meta: LlamaGuard 2 8B | $0.20 | $0.20 | — | — |
| Mistral: Mistral 7B Instruct v0.2 | $0.20 | $0.20 | 315ms | 78 |
| MiniMax: MiniMax M3 (batch) | $0.30 | $1.20 | — | — |
| MiniMax: MiniMax M3 | $0.30 | $1.20 | 800ms | 72 |
| Qwen: Qwen2.5 7B Instruct | $0.30 | $0.30 | 234ms | 52 |
| Meta: Muse Glimmer 30B | $0.35 | $1.50 | 438ms | 91 |
| Google: Gemma 4 31B | $0.39 | $0.97 | 966ms | 13 |
| Thinking Machines: Inkling Small | $0.50 | $1.20 | 1484ms | 63 |
| Arcee AI: Coder Large | $0.50 | $0.80 | — | — |
| NVIDIA: Nemotron 3 Ultra (batch) | $0.60 | $3.60 | — | — |
| NVIDIA: Nemotron 3 Ultra | $0.60 | $3.60 | 755ms | 89 |
| Arcee AI: Virtuoso Large | $0.75 | $1.20 | — | — |
| Arcee AI: Maestro Reasoning | $0.90 | $3.30 | — | — |
| MoonshotAI: Kimi K2.7 Code (batch) | $0.95 | $4.00 | — | — |
| MoonshotAI: Kimi K2.7 Code | $0.95 | $4.00 | 799ms | 261 |
| Thinking Machines: Inkling | $1.00 | $4.05 | 472ms | 82 |
| Thinking Machines: Inkling (batch) | $1.00 | $4.05 | — | — |
| Meta: Llama 3.3 70B Instruct | $1.04 | $1.04 | 1098ms | 13 |
| MoonshotAI: Kimi K2.6 | $1.20 | $4.50 | 386ms | 99 |
| Deep Cogito: Cogito v2.1 671B | $1.25 | $1.25 | 396ms | 26 |
| DeepSeek: DeepSeek V4 Pro 0813 | $1.32 | $3.96 | 1217ms | 99 |
| Z.ai: GLM 5.2 | $1.40 | $4.40 | 723ms | 67 |
| Z.ai: GLM 5.2 (batch) | $1.40 | $4.40 | — | — |
| DeepSeek: DeepSeek V4 Pro 0423 | $1.74 | $3.48 | 1245ms | 56 |
| Qwen: Qwen3.8 2.4T A95B | $2.50 | $6.25 | 4905ms | 78 |
| MoonshotAI: Kimi K3 | $3.00 | $15.00 | 2699ms | 49 |
Community Reviews
4.5★★★★★(2 reviews)
clouduser42
★★★★★2025-06-15
Reliable service, great API documentation.
mlresearcher
★★★★☆2025-06-10
Good performance but support could be faster.