LIVE
Models: —+Providers: —+Cheapest H100: $2.49/hrUpdated: 01:10 PMModels: —+Providers: —+Cheapest H100: $2.49/hrUpdated: 01:10 PM
Marketplace
Providers Models
D

DeepInfra

AGGREGATEDINFERENCE
N/A
Uptime
N/A
Rating

30-Day Uptime

96.7%
2026-07-212026-08-19

Inference Latency

Mistral: Mistral Nemo328ms TTFT · 27 TPS
Meta: Llama 3.1 8B Instruct521ms TTFT · 21 TPS
OpenAI: gpt-oss-20b359ms TTFT · 80 TPS
OpenAI: gpt-oss-120b337ms TTFT · 42 TPS
OpenAI: gpt-oss-120b (exacto)500ms TTFT · 64 TPS
Sao10K: Llama 3 8B Lunaris137ms TTFT · 76 TPS
NVIDIA: Nemotron Nano 9B V2244ms TTFT · 119 TPS
NVIDIA: Nemotron 3 Nano 30B A3B267ms TTFT · 111 TPS
Google: Gemma 3 12B379ms TTFT · 23 TPS
Mistral: Mistral Small 3210ms TTFT · 39 TPS
Google: Gemma 3 4B837ms TTFT · 15 TPS
Ling-3.0-flash1104ms TTFT · 31 TPS
Z.ai: GLM 4.7 Flash363ms TTFT · 53 TPS
Microsoft: Phi 4324ms TTFT · 67 TPS
Google: Gemma 4 26B A4B 467ms TTFT · 19 TPS
Mistral: Mistral Small 3.2 24B425ms TTFT · 28 TPS
NVIDIA: Nemotron 3.5 Lightning333ms TTFT · 115 TPS
DeepSeek: DeepSeek V4 Flash 0731916ms TTFT · 30 TPS
Google: Gemma 3 27B637ms TTFT · 17 TPS
Qwen: Qwen3 32B227ms TTFT · 39 TPS

Inference Models

ModelInput $/MOutput $/MTTFTTPS
Mistral: Mistral Nemo$0.02$0.03328ms27
Meta: Llama 3.1 8B Instruct$0.02$0.04521ms21
OpenAI: gpt-oss-20b$0.03$0.14359ms80
OpenAI: gpt-oss-120b$0.04$0.17337ms42
OpenAI: gpt-oss-120b (exacto)$0.04$0.19500ms64
Sao10K: Llama 3 8B Lunaris$0.04$0.05137ms76
NVIDIA: Nemotron Nano 9B V2$0.04$0.16244ms119
NVIDIA: Nemotron 3 Nano 30B A3B$0.05$0.20267ms111
Google: Gemma 3 12B$0.05$0.15379ms23
Mistral: Mistral Small 3$0.05$0.08210ms39
Google: Gemma 3 4B$0.05$0.10837ms15
Ling-3.0-flash$0.06$0.181104ms31
Z.ai: GLM 4.7 Flash$0.06$0.40363ms53
Microsoft: Phi 4$0.07$0.14324ms67
Google: Gemma 4 26B A4B $0.07$0.34467ms19
Mistral: Mistral Small 3.2 24B$0.08$0.20425ms28
NVIDIA: Nemotron 3.5 Lightning$0.08$0.20333ms115
DeepSeek: DeepSeek V4 Flash 0731$0.08$0.18916ms30
Google: Gemma 3 27B$0.08$0.16637ms17
Qwen: Qwen3 32B$0.08$0.28227ms39
NVIDIA: Nemotron 3 Super$0.09$0.401434ms62
Qwen: Qwen3 235B A22B Instruct 2507$0.09$0.55444ms17
DeepSeek: DeepSeek V4 Flash 0423$0.09$0.18677ms41
Google: Gemma 4 31B$0.09$0.34453ms37
Qwen: Qwen3 Next 80B A3B Instruct$0.09$1.10402ms57
Meta: Llama 3.3 70B Instruct$0.10$0.32515ms15
Qwen: Qwen3.6 35B A3B$0.10$0.95699ms19
Meta: Llama 4 Scout$0.10$0.30222ms11
Qwen: Qwen3.5-9B$0.10$0.15557ms39
Qwen: Qwen3 14B$0.12$0.24411ms39
Qwen: Qwen3 30B A3B$0.12$0.50247ms69
Google: Gemma 4 31B$0.13$0.385959ms0
Qwen: Qwen3.5-35B-A3B$0.14$1.00348ms77
Tencent: Hy3$0.14$0.58402ms57
OpenAI: gpt-oss-120b$0.15$0.60409ms139
Qwen: Qwen3 VL 30B A3B Instruct$0.15$0.60301ms25
Meta: Llama Guard 4 12B$0.18$0.18527ms4
NVIDIA: Nemotron Nano 12B 2 VL$0.20$0.60458ms48
Qwen: Qwen2.5 VL 32B Instruct$0.20$0.60617ms26
Qwen: Qwen3 VL 235B A22B Instruct$0.20$0.881110ms13
AllenAI: Olmo 3.1 32B Instruct$0.20$0.60586ms37
StepFun: Step 3.7 Flash$0.20$1.15184ms113
Meta: Llama 4 Maverick$0.20$0.80452ms31
DeepSeek: DeepSeek V3.1 Terminus (exacto)$0.21$0.791372ms18
Qwen: Qwen3 235B A22B Thinking 2507$0.23$2.30470ms37
DeepSeek: DeepSeek V3 0324$0.24$0.904076ms6
MiniMax: MiniMax M2.7$0.25$1.00470ms32
DeepSeek: DeepSeek V3.1$0.25$0.951433ms5
Qwen: Qwen3.5-27B$0.26$2.60561ms35
DeepSeek: DeepSeek V3.2$0.26$0.383207ms8
Google: Gemma 4 31B$0.27$0.764900ms16
MiniMax: MiniMax M3$0.28$1.101332ms32
Qwen: Qwen3.5-122B-A10B$0.29$2.40564ms50
Meta: Muse Glimmer 30B$0.30$1.20427ms157
Qwen: Qwen3 Coder 480B A35B$0.30$1.00688ms63
DeepSeek: DeepSeek V3$0.32$0.89414ms25
Qwen: Qwen3.6 27B$0.32$3.20531ms33
Meta: Llama 3.2 11B Vision Instruct$0.35$0.35580ms43
Qwen2.5 72B Instruct$0.36$0.40286ms39
MiniMax: MiniMax M2.7$0.38$1.705726ms28
Meta: Llama 3.1 70B Instruct$0.40$0.40233ms19
MythoMax 13B$0.40$0.40268ms47
NVIDIA: Llama 3.3 Nemotron Super 49B V1.5$0.40$0.40142ms32
Xiaomi: MiMo-V2.5$0.40$2.00567ms24
Z.ai: GLM 4.7$0.40$1.75887ms29
MoonshotAI: Kimi K2.5$0.45$2.25678ms69
Thinking Machines: Inkling Small$0.45$1.20367ms184
Qwen: Qwen3.5 397B A17B$0.45$3.00256ms37
Z.ai: GLM 4.6$0.50$2.00466ms24
DeepSeek: R1 0528$0.50$2.15671ms24
NVIDIA: Nemotron 3 Ultra$0.50$2.202529ms53
Mistral: Mixtral 8x7B Instruct$0.54$0.54
Z.ai: GLM 5$0.60$2.08713ms33
MoonshotAI: Kimi K2.7 Code$0.68$3.40687ms48
Nous: Hermes 3 70B Instruct$0.70$0.70347ms23
MoonshotAI: Kimi K2.6$0.75$3.50985ms21
Z.ai: GLM 5.2$0.75$2.401349ms38
Sao10K: Llama 3.1 Euryale 70B v2.2$0.85$0.85207ms48
Thinking Machines: Inkling$0.95$4.05499ms56
Nous: Hermes 3 405B Instruct$1.00$1.00285ms22
Xiaomi: MiMo-V2.5-Pro$1.00$3.00670ms82
Z.ai: GLM 5.1$1.05$3.501001ms50
NVIDIA: Llama 3.1 Nemotron 70B Instruct$1.20$1.20
DeepSeek: DeepSeek V4 Pro 0423$1.30$2.60899ms19
Qwen: Qwen3.8 2.4T A95B$2.00$6.001508ms61
MoonshotAI: Kimi K3$2.85$14.25828ms12

Community Reviews

4.5★★★★★(2 reviews)
clouduser42
★★★★★2025-06-15

Reliable service, great API documentation.

mlresearcher
★★★★2025-06-10

Good performance but support could be faster.