AI Model Speed Rankings
Models ranked by measured output throughput (tokens per second) and average latency.
- OpenAI dall-e-3 — Throughput: 2000.0 tps, Latency: 4234.00s
- MiniMax MiniMax H3 — Throughput: 778.1 tps, Latency: 736235.00s
- OpenAI gpt-5-nano — Throughput: 591.4 tps, Latency: 13633.56s
- Google Gemini 3.5 Flash — Throughput: 379.4 tps, Latency: 8816.13s
- Google Gemini 3.6 Flash — Throughput: 364.6 tps, Latency: 6117.05s
- Google Gemini 3.1 Flash-Lite — Throughput: 353.2 tps, Latency: 1376.37s
- xAI Grok 4.20 Multi-Agent — Throughput: 326.0 tps, Latency: 0.00s
- Google Gemini Omni Flash Preview — Throughput: 306.6 tps, Latency: 99229.00s
- OpenAI GPT-5 — Throughput: 222.4 tps, Latency: 21062.68s
- OpenAI o3-mini — Throughput: 194.2 tps, Latency: 5194.60s
- Google Gemini 3.1 Flash-Lite Preview — Throughput: 178.0 tps, Latency: 888.60s
- OpenAI GPT-5.4 Mini — Throughput: 165.5 tps, Latency: 1316.89s
- OpenAI chatgpt-image-latest — Throughput: 163.0 tps, Latency: 25489.00s
- Google gemini-2.5-flash — Throughput: 144.3 tps, Latency: 19431.90s
- Google Gemini 3.5 Flash-Lite — Throughput: 140.0 tps, Latency: 2392.30s
- OpenAI GPT-5.6 Terra — Throughput: 138.6 tps, Latency: 4566.30s
- xAI Grok 4.3 — Throughput: 138.6 tps, Latency: 925.85s
- OpenAI GPT Image 2 Develop — Throughput: 138.3 tps, Latency: 31112.40s
- OpenAI gpt-4o-all — Throughput: 138.1 tps, Latency: 3184.91s
- Google Nano Banana (Gemini 2.5 Flash Image 🍌) — Throughput: 134.4 tps, Latency: 0.00s
- OpenAI gpt-image-1-mini — Throughput: 133.7 tps, Latency: 46435.50s
- xAI grok-3 — Throughput: 129.0 tps, Latency: 0.00s
- xAI grok-4-1-fast-reasoning — Throughput: 128.6 tps, Latency: 510.00s
- xAI grok-4-0709 — Throughput: 122.7 tps, Latency: 3704.50s
- OpenAI gpt-oss-120b — Throughput: 119.6 tps, Latency: 0.00s
- OpenAI gpt-4.1-nano — Throughput: 109.6 tps, Latency: 1208.50s
- OpenAI gpt-4-gizmo-* — Throughput: 109.6 tps, Latency: 3078.36s
- OpenAI GPT-5.5 — Throughput: 109.3 tps, Latency: 4325.78s
- OpenAI GPT Image 1.5 — Throughput: 108.7 tps, Latency: 41491.00s
- xAI grok-3-mini — Throughput: 104.6 tps, Latency: 0.00s
- Anthropic Claude Haiku 4.5 — Throughput: 104.4 tps, Latency: 3512.39s
- OpenAI GPT-5.1 — Throughput: 104.0 tps, Latency: 2765.79s
- Anthropic Claude Sonnet 5 — Throughput: 99.6 tps, Latency: 4690.55s
- Google Gemini 3.1 Pro Preview Customtools — Throughput: 95.7 tps, Latency: 44038.80s
- xAI grok-4-fast-reasoning — Throughput: 89.0 tps, Latency: 0.00s
- Google Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image 🍌) — Throughput: 89.0 tps, Latency: 15223.62s
- xAI Grok Build 0.1 — Throughput: 87.2 tps, Latency: 530.40s
- OpenAI gpt-4o-search-preview — Throughput: 86.8 tps, Latency: 0.00s
- OpenAI GPT-5.4 — Throughput: 86.0 tps, Latency: 3340.58s
- MiniMax MiniMax M2.5 — Throughput: 85.9 tps, Latency: 0.00s
- DeepSeek DeepSeek-V4-Pro-0813 — Throughput: 85.2 tps, Latency: 470.14s
- Google Gemini 3.1 Pro Preview — Throughput: 81.5 tps, Latency: 46269.16s
- xAI Grok 4.20 (Non-Reasoning) — Throughput: 81.1 tps, Latency: 1484.00s
- Anthropic Claude Fable 5 — Throughput: 80.3 tps, Latency: 6873.03s
- OpenAI gpt-4o-transcribe-diarize — Throughput: 76.7 tps, Latency: 0.00s
- Alibaba qwen-plus — Throughput: 75.7 tps, Latency: 1884.00s
- Anthropic Claude Opus 4.7 — Throughput: 72.1 tps, Latency: 2960.61s
- Anthropic Claude Opus 5 — Throughput: 71.0 tps, Latency: 11786.30s
- OpenAI GPT-5.6 Luna — Throughput: 69.4 tps, Latency: 3086.49s
- Google Gemini 3 Flash Preview — Throughput: 68.8 tps, Latency: 6088.67s
- xAI grok-4-1-fast-non-reasoning — Throughput: 68.4 tps, Latency: 798.00s
- OpenAI gpt-5-mini — Throughput: 66.8 tps, Latency: 6152.73s
- Anthropic Claude Opus 4.8 — Throughput: 63.5 tps, Latency: 4992.71s
- xAI Grok 4.20 — Throughput: 62.5 tps, Latency: 312.00s
- Google Nano Banana 2 (Gemini 3.1 Flash Image 🍌) — Throughput: 60.4 tps, Latency: 17029.00s
- OpenAI gpt-4.1-mini — Throughput: 60.3 tps, Latency: 3375.86s
- OpenAI gpt-5.1-codex-mini — Throughput: 60.1 tps, Latency: 0.00s
- DeepSeek DeepSeek-V4-Pro — Throughput: 59.9 tps, Latency: 1363.00s
- Moonshot Kimi K2.6 — Throughput: 58.2 tps, Latency: 1668.71s
- DeepSeek DeepSeek-V3.2 — Throughput: 57.9 tps, Latency: 1723.33s
- OpenAI gpt-4.1 — Throughput: 57.0 tps, Latency: 3298.00s
- DeepSeek deepseek-v3.1 — Throughput: 56.9 tps, Latency: 3098.00s
- Google Gemini 3.7 Flash — Throughput: 56.6 tps, Latency: 0.00s
- MiniMax MiniMax M3 — Throughput: 56.2 tps, Latency: 0.00s
- Google Gemini 3 Pro Preview — Throughput: 56.0 tps, Latency: 0.00s
- OpenAI GPT-5.2 — Throughput: 53.6 tps, Latency: 2367.25s
- Anthropic Claude Sonnet 4.5 — Throughput: 51.8 tps, Latency: 4038.08s
- OpenAI gpt-4o-mini-transcribe — Throughput: 50.6 tps, Latency: 0.00s
- Anthropic Claude Sonnet 4.6 — Throughput: 50.4 tps, Latency: 5071.27s
- OpenAI GPT-5.6 Sol — Throughput: 50.2 tps, Latency: 15414.64s
- Anthropic Claude Opus 4.6 — Throughput: 50.0 tps, Latency: 7760.58s
- xAI grok-4-fast-non-reasoning — Throughput: 49.6 tps, Latency: 0.00s
- xAI Grok 4.5 — Throughput: 48.4 tps, Latency: 1936.19s
- OpenAI gpt-4o-2024-11-20 — Throughput: 42.2 tps, Latency: 965.00s
- OpenAI GPT-5.5 Pro — Throughput: 42.0 tps, Latency: 97235.00s
- OpenAI gpt-3.5-turbo-16k — Throughput: 41.1 tps, Latency: 0.00s
- Z.ai GLM 5.2 — Throughput: 40.1 tps, Latency: 3541.40s
- Moonshot Kimi K3 — Throughput: 37.3 tps, Latency: 29631.00s
- xAI Grok 4.6 — Throughput: 36.2 tps, Latency: 2000.14s
- OpenAI gpt-4o — Throughput: 35.4 tps, Latency: 1102.75s
- ByteDance Doubao-Seed-2.0-pro — Throughput: 31.5 tps, Latency: 0.00s
- ByteDance Doubao-Seedream-5.0-pro — Throughput: 31.3 tps, Latency: 130667.00s
- OpenAI o4-mini — Throughput: 30.8 tps, Latency: 0.00s
- Xiaomi MiMo-V2.5 Pro — Throughput: 30.1 tps, Latency: 6391.75s
- OpenAI GPT Image 1.0 — Throughput: 29.8 tps, Latency: 28444.00s
- OpenAI GPT Image 2 — Throughput: 29.2 tps, Latency: 48803.63s
- Alibaba qwen-max — Throughput: 27.0 tps, Latency: 657.00s
- OpenAI gpt-4o-mini — Throughput: 26.4 tps, Latency: 2179.00s
- Anthropic Claude Opus 4.5 — Throughput: 25.4 tps, Latency: 3433.00s
- OpenAI o1 — Throughput: 21.4 tps, Latency: 0.00s
- Alibaba qwen3-vl-flash — Throughput: 20.0 tps, Latency: 0.00s
- OpenAI gpt-3.5-turbo-1106 — Throughput: 17.0 tps, Latency: 0.00s
- OpenAI GPT-5.4 Nano — Throughput: 15.3 tps, Latency: 1185.00s
- Xiaomi MiMo-V2.5 — Throughput: 15.3 tps, Latency: 15989.50s
- Google gemini-2.5-pro — Throughput: 15.3 tps, Latency: 0.00s
- OpenAI gpt-5.1-codex — Throughput: 13.6 tps, Latency: 0.00s
- OpenAI gpt-4o-transcribe — Throughput: 13.4 tps, Latency: 0.00s
- OpenAI gpt-3.5-turbo — Throughput: 9.0 tps, Latency: 0.00s
- OpenAI gpt-4-turbo-preview — Throughput: 8.9 tps, Latency: 0.00s
- OpenAI GPT-5.3 Codex — Throughput: 8.4 tps, Latency: 0.00s