LIVE
Models: —+Providers: —+Cheapest H100: $2.49/hrUpdated: 01:12 PMModels: —+Providers: —+Cheapest H100: $2.49/hrUpdated: 01:12 PM
Marketplace
Providers Models
N

Novita

AGGREGATEDINFERENCE
N/A
Uptime
N/A
Rating

30-Day Uptime

100%
2026-07-212026-08-19

Inference Latency

inclusionAI: Ling-2.6-flash (free)3047ms TTFT · 19 TPS
Ling-3.0-flash (free)2176ms TTFT · 74 TPS
inclusionAI: Ling-2.6-1T (free)2597ms TTFT · 25 TPS
Tencent: Hy3 (free)2717ms TTFT · 40 TPS
inclusionAI: Ling 3.0 Tiny (free)2151ms TTFT · 71 TPS
inclusionAI: Ring-2.6-1T (free)2682ms TTFT · 57 TPS
inclusionAI: Ling-2.6-flash578ms TTFT · 85 TPS
Meta: Llama 3.1 8B Instruct512ms TTFT · 60 TPS
Ling-3.0-flash712ms TTFT · 60 TPS
Mistral: Mistral Nemo5752ms TTFT · 5 TPS
OpenAI: gpt-oss-120b (exacto)633ms TTFT · 51 TPS
OpenAI: gpt-oss-20b538ms TTFT · 128 TPS
Sao10K: Llama 3 8B Lunaris355ms TTFT · 69 TPS
NVIDIA: Nemotron 3 Nano 30B A3B380ms TTFT · 126 TPS
OpenAI: gpt-oss-120b593ms TTFT · 83 TPS
Z.ai: GLM 4.7 Flash1693ms TTFT · 33 TPS
Qwen: Qwen3 Coder 30B A3B Instruct1078ms TTFT · 65 TPS
inclusionAI: Ring-2.6-1T2453ms TTFT · 79 TPS
inclusionAI: Ling-2.6-1T2375ms TTFT · 102 TPS
Qwen: Qwen3 235B A22B Instruct 2507680ms TTFT · 27 TPS

Inference Models

ModelInput $/MOutput $/MTTFTTPS
inclusionAI: Ling-2.6-flash (free)$0.00$0.003047ms19
Ling-3.0-flash (free)$0.00$0.002176ms74
inclusionAI: Ling-2.6-1T (free)$0.00$0.002597ms25
Tencent: Hy3 (free)$0.00$0.002717ms40
inclusionAI: Ling 3.0 Tiny (free)$0.00$0.002151ms71
inclusionAI: Ring-2.6-1T (free)$0.00$0.002682ms57
inclusionAI: Ling-2.6-flash$0.01$0.03578ms85
Meta: Llama 3.1 8B Instruct$0.02$0.05512ms60
Ling-3.0-flash$0.02$0.06712ms60
Mistral: Mistral Nemo$0.04$0.175752ms5
OpenAI: gpt-oss-120b (exacto)$0.04$0.20633ms51
OpenAI: gpt-oss-20b$0.04$0.15538ms128
Sao10K: Llama 3 8B Lunaris$0.05$0.05355ms69
NVIDIA: Nemotron 3 Nano 30B A3B$0.05$0.20380ms126
OpenAI: gpt-oss-120b$0.05$0.25593ms83
Z.ai: GLM 4.7 Flash$0.07$0.401693ms33
Baidu: ERNIE 4.5 21B A3B$0.07$0.28
Baidu: ERNIE 4.5 21B A3B Thinking$0.07$0.28
Qwen: Qwen3 Coder 30B A3B Instruct$0.07$0.271078ms65
inclusionAI: Ring-2.6-1T$0.08$0.632453ms79
inclusionAI: Ling-2.6-1T$0.08$0.632375ms102
Qwen: Qwen3 235B A22B Instruct 2507$0.09$0.58680ms27
Google: Gemma 3 27B$0.12$0.20898ms17
Z.ai: GLM 4.5 Air$0.13$0.85660ms37
Google: Gemma 4 26B A4B $0.13$0.401171ms22
Meta: Llama 3.3 70B Instruct$0.14$0.40744ms27
Google: Gemma 4 31B$0.14$0.402300ms4
Baidu: ERNIE 4.5 VL 28B A3B$0.14$0.56
DeepSeek: DeepSeek V4 Flash 0423$0.14$0.28999ms72
DeepSeek: DeepSeek V4 Flash 0731$0.14$0.281203ms83
NousResearch: Hermes 2 Pro - Llama-3 8B$0.14$0.14
Tencent: Hy3$0.14$0.582188ms44
Qwen: Qwen3 Next 80B A3B Instruct$0.15$1.50673ms36
Xiaomi: MiMo-V2.5$0.17$0.344677ms19
Meta: Llama 4 Scout$0.18$0.59582ms36
Qwen: Qwen3 VL 30B A3B Instruct$0.20$0.70895ms24
StepFun: Step 3.7 Flash$0.20$1.152241ms9
Qwen: Qwen3 Coder Next$0.20$1.50837ms68
Kwaipilot: KAT-Coder-Pro V1$0.21$0.831788ms56
DeepSeek: DeepSeek V3.1 Terminus (exacto)$0.22$0.802354ms26
DeepSeek: DeepSeek V3.2$0.27$0.401255ms22
DeepSeek: DeepSeek V3.1 Terminus$0.27$1.001818ms11
DeepSeek: DeepSeek V3.2 Exp$0.27$0.411522ms17
MiniMax: MiniMax M2.7$0.27$1.083739ms6
Meta: Llama 4 Maverick$0.27$0.85597ms25
DeepSeek: DeepSeek V3.1$0.27$1.001920ms23
DeepSeek: DeepSeek V3 0324$0.27$1.121484ms32
Baidu: ERNIE 4.5 300B A47B $0.28$1.101629ms21
MiniMax: MiniMax M2.5$0.30$1.20861ms72
Z.ai: GLM 4.6V$0.30$0.904215ms25
MiniMax: MiniMax M2$0.30$1.201560ms54
Qwen: Qwen3 VL 235B A22B Instruct$0.30$1.501836ms17
Qwen: Qwen3 235B A22B Thinking 2507$0.30$3.001905ms29
MiniMax: MiniMax M2.1$0.30$1.201793ms36
MiniMax: MiniMax M3$0.30$1.201860ms57
Qwen: Qwen3.5-27B$0.30$2.401007ms17
Qwen2.5 72B Instruct$0.38$0.406476ms34
Qwen: Qwen3 Coder 480B A35B$0.38$1.551628ms20
DeepSeek: DeepSeek V3$0.40$1.301175ms19
Qwen: Qwen3.5-122B-A10B$0.40$3.201035ms42
Baidu: ERNIE 4.5 VL 424B A47B $0.42$1.251326ms11
Z.ai: GLM 4.6 (exacto)$0.44$1.76834ms119
Xiaomi: MiMo-V2.5-Pro$0.48$0.962892ms23
Meta: Llama 3 70B Instruct$0.51$0.742071ms1
Z.ai: GLM 4.7$0.54$1.982397ms21
MiniMax: MiniMax M1$0.55$2.201207ms23
Z.ai: GLM 4.6$0.55$2.201071ms39
MoonshotAI: Kimi K2 0711$0.57$2.30762ms3
MoonshotAI: Kimi K2.5$0.57$2.851416ms33
Z.ai: GLM 4.5V$0.60$1.801694ms62
Qwen: Qwen3.5 397B A17B$0.60$3.601361ms46
MoonshotAI: Kimi K2 Thinking$0.60$2.50931ms38
MoonshotAI: Kimi K2 0905$0.60$2.50668ms19
WizardLM-2 8x22B$0.62$0.627269ms4
DeepSeek: R1 0528$0.70$2.50856ms26
DeepSeek: R1$0.70$2.501066ms21
Z.ai: GLM 5.2$0.74$2.332346ms38
DeepSeek: R1 Distill Llama 70B$0.80$0.80840ms24
MoonshotAI: Kimi K2.6$0.80$3.402526ms25
MoonshotAI: Kimi K2.7 Code$0.91$3.842315ms26
Qwen: Qwen3 VL 235B A22B Thinking$0.98$3.951238ms38
Z.ai: GLM 5$1.00$3.201339ms73
DeepSeek: DeepSeek V4 Pro 0813$1.32$3.961777ms50
Z.ai: GLM 5.1$1.38$4.402494ms35
DeepSeek: DeepSeek V4 Pro 0423$1.44$2.881500ms49
Sao10K: Llama 3.1 Euryale 70B v2.2$1.45$1.45638ms8
Sao10k: Llama 3 Euryale 70B v2.1$1.48$1.48

Community Reviews

4.5★★★★★(2 reviews)
clouduser42
★★★★★2025-06-15

Reliable service, great API documentation.

mlresearcher
★★★★2025-06-10

Good performance but support could be faster.