AI Model Speed Rankings

Models ranked by measured output throughput (tokens per second) and average latency.

  1. OpenAI dall-e-3 — Throughput: 2000.0 tps, Latency: 4234.00s
  2. MiniMax MiniMax H3 — Throughput: 778.1 tps, Latency: 736235.00s
  3. OpenAI gpt-5-nano — Throughput: 591.4 tps, Latency: 13633.56s
  4. Google Gemini 3.5 Flash — Throughput: 379.4 tps, Latency: 8816.13s
  5. Google Gemini 3.6 Flash — Throughput: 364.6 tps, Latency: 6117.05s
  6. Google Gemini 3.1 Flash-Lite — Throughput: 353.2 tps, Latency: 1376.37s
  7. xAI Grok 4.20 Multi-Agent — Throughput: 326.0 tps, Latency: 0.00s
  8. Google Gemini Omni Flash Preview — Throughput: 306.6 tps, Latency: 99229.00s
  9. OpenAI GPT-5 — Throughput: 222.4 tps, Latency: 21062.68s
  10. OpenAI o3-mini — Throughput: 194.2 tps, Latency: 5194.60s
  11. Google Gemini 3.1 Flash-Lite Preview — Throughput: 178.0 tps, Latency: 888.60s
  12. OpenAI GPT-5.4 Mini — Throughput: 165.5 tps, Latency: 1316.89s
  13. OpenAI chatgpt-image-latest — Throughput: 163.0 tps, Latency: 25489.00s
  14. Google gemini-2.5-flash — Throughput: 144.3 tps, Latency: 19431.90s
  15. Google Gemini 3.5 Flash-Lite — Throughput: 140.0 tps, Latency: 2392.30s
  16. OpenAI GPT-5.6 Terra — Throughput: 138.6 tps, Latency: 4566.30s
  17. xAI Grok 4.3 — Throughput: 138.6 tps, Latency: 925.85s
  18. OpenAI GPT Image 2 Develop — Throughput: 138.3 tps, Latency: 31112.40s
  19. OpenAI gpt-4o-all — Throughput: 138.1 tps, Latency: 3184.91s
  20. Google Nano Banana (Gemini 2.5 Flash Image 🍌) — Throughput: 134.4 tps, Latency: 0.00s
  21. OpenAI gpt-image-1-mini — Throughput: 133.7 tps, Latency: 46435.50s
  22. xAI grok-3 — Throughput: 129.0 tps, Latency: 0.00s
  23. xAI grok-4-1-fast-reasoning — Throughput: 128.6 tps, Latency: 510.00s
  24. xAI grok-4-0709 — Throughput: 122.7 tps, Latency: 3704.50s
  25. OpenAI gpt-oss-120b — Throughput: 119.6 tps, Latency: 0.00s
  26. OpenAI gpt-4.1-nano — Throughput: 109.6 tps, Latency: 1208.50s
  27. OpenAI gpt-4-gizmo-* — Throughput: 109.6 tps, Latency: 3078.36s
  28. OpenAI GPT-5.5 — Throughput: 109.3 tps, Latency: 4325.78s
  29. OpenAI GPT Image 1.5 — Throughput: 108.7 tps, Latency: 41491.00s
  30. xAI grok-3-mini — Throughput: 104.6 tps, Latency: 0.00s
  31. Anthropic Claude Haiku 4.5 — Throughput: 104.4 tps, Latency: 3512.39s
  32. OpenAI GPT-5.1 — Throughput: 104.0 tps, Latency: 2765.79s
  33. Anthropic Claude Sonnet 5 — Throughput: 99.6 tps, Latency: 4690.55s
  34. Google Gemini 3.1 Pro Preview Customtools — Throughput: 95.7 tps, Latency: 44038.80s
  35. xAI grok-4-fast-reasoning — Throughput: 89.0 tps, Latency: 0.00s
  36. Google Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image 🍌) — Throughput: 89.0 tps, Latency: 15223.62s
  37. xAI Grok Build 0.1 — Throughput: 87.2 tps, Latency: 530.40s
  38. OpenAI gpt-4o-search-preview — Throughput: 86.8 tps, Latency: 0.00s
  39. OpenAI GPT-5.4 — Throughput: 86.0 tps, Latency: 3340.58s
  40. MiniMax MiniMax M2.5 — Throughput: 85.9 tps, Latency: 0.00s
  41. DeepSeek DeepSeek-V4-Pro-0813 — Throughput: 85.2 tps, Latency: 470.14s
  42. Google Gemini 3.1 Pro Preview — Throughput: 81.5 tps, Latency: 46269.16s
  43. xAI Grok 4.20 (Non-Reasoning) — Throughput: 81.1 tps, Latency: 1484.00s
  44. Anthropic Claude Fable 5 — Throughput: 80.3 tps, Latency: 6873.03s
  45. OpenAI gpt-4o-transcribe-diarize — Throughput: 76.7 tps, Latency: 0.00s
  46. Alibaba qwen-plus — Throughput: 75.7 tps, Latency: 1884.00s
  47. Anthropic Claude Opus 4.7 — Throughput: 72.1 tps, Latency: 2960.61s
  48. Anthropic Claude Opus 5 — Throughput: 71.0 tps, Latency: 11786.30s
  49. OpenAI GPT-5.6 Luna — Throughput: 69.4 tps, Latency: 3086.49s
  50. Google Gemini 3 Flash Preview — Throughput: 68.8 tps, Latency: 6088.67s
  51. xAI grok-4-1-fast-non-reasoning — Throughput: 68.4 tps, Latency: 798.00s
  52. OpenAI gpt-5-mini — Throughput: 66.8 tps, Latency: 6152.73s
  53. Anthropic Claude Opus 4.8 — Throughput: 63.5 tps, Latency: 4992.71s
  54. xAI Grok 4.20 — Throughput: 62.5 tps, Latency: 312.00s
  55. Google Nano Banana 2 (Gemini 3.1 Flash Image 🍌) — Throughput: 60.4 tps, Latency: 17029.00s
  56. OpenAI gpt-4.1-mini — Throughput: 60.3 tps, Latency: 3375.86s
  57. OpenAI gpt-5.1-codex-mini — Throughput: 60.1 tps, Latency: 0.00s
  58. DeepSeek DeepSeek-V4-Pro — Throughput: 59.9 tps, Latency: 1363.00s
  59. Moonshot Kimi K2.6 — Throughput: 58.2 tps, Latency: 1668.71s
  60. DeepSeek DeepSeek-V3.2 — Throughput: 57.9 tps, Latency: 1723.33s
  61. OpenAI gpt-4.1 — Throughput: 57.0 tps, Latency: 3298.00s
  62. DeepSeek deepseek-v3.1 — Throughput: 56.9 tps, Latency: 3098.00s
  63. Google Gemini 3.7 Flash — Throughput: 56.6 tps, Latency: 0.00s
  64. MiniMax MiniMax M3 — Throughput: 56.2 tps, Latency: 0.00s
  65. Google Gemini 3 Pro Preview — Throughput: 56.0 tps, Latency: 0.00s
  66. OpenAI GPT-5.2 — Throughput: 53.6 tps, Latency: 2367.25s
  67. Anthropic Claude Sonnet 4.5 — Throughput: 51.8 tps, Latency: 4038.08s
  68. OpenAI gpt-4o-mini-transcribe — Throughput: 50.6 tps, Latency: 0.00s
  69. Anthropic Claude Sonnet 4.6 — Throughput: 50.4 tps, Latency: 5071.27s
  70. OpenAI GPT-5.6 Sol — Throughput: 50.2 tps, Latency: 15414.64s
  71. Anthropic Claude Opus 4.6 — Throughput: 50.0 tps, Latency: 7760.58s
  72. xAI grok-4-fast-non-reasoning — Throughput: 49.6 tps, Latency: 0.00s
  73. xAI Grok 4.5 — Throughput: 48.4 tps, Latency: 1936.19s
  74. OpenAI gpt-4o-2024-11-20 — Throughput: 42.2 tps, Latency: 965.00s
  75. OpenAI GPT-5.5 Pro — Throughput: 42.0 tps, Latency: 97235.00s
  76. OpenAI gpt-3.5-turbo-16k — Throughput: 41.1 tps, Latency: 0.00s
  77. Z.ai GLM 5.2 — Throughput: 40.1 tps, Latency: 3541.40s
  78. Moonshot Kimi K3 — Throughput: 37.3 tps, Latency: 29631.00s
  79. xAI Grok 4.6 — Throughput: 36.2 tps, Latency: 2000.14s
  80. OpenAI gpt-4o — Throughput: 35.4 tps, Latency: 1102.75s
  81. ByteDance Doubao-Seed-2.0-pro — Throughput: 31.5 tps, Latency: 0.00s
  82. ByteDance Doubao-Seedream-5.0-pro — Throughput: 31.3 tps, Latency: 130667.00s
  83. OpenAI o4-mini — Throughput: 30.8 tps, Latency: 0.00s
  84. Xiaomi MiMo-V2.5 Pro — Throughput: 30.1 tps, Latency: 6391.75s
  85. OpenAI GPT Image 1.0 — Throughput: 29.8 tps, Latency: 28444.00s
  86. OpenAI GPT Image 2 — Throughput: 29.2 tps, Latency: 48803.63s
  87. Alibaba qwen-max — Throughput: 27.0 tps, Latency: 657.00s
  88. OpenAI gpt-4o-mini — Throughput: 26.4 tps, Latency: 2179.00s
  89. Anthropic Claude Opus 4.5 — Throughput: 25.4 tps, Latency: 3433.00s
  90. OpenAI o1 — Throughput: 21.4 tps, Latency: 0.00s
  91. Alibaba qwen3-vl-flash — Throughput: 20.0 tps, Latency: 0.00s
  92. OpenAI gpt-3.5-turbo-1106 — Throughput: 17.0 tps, Latency: 0.00s
  93. OpenAI GPT-5.4 Nano — Throughput: 15.3 tps, Latency: 1185.00s
  94. Xiaomi MiMo-V2.5 — Throughput: 15.3 tps, Latency: 15989.50s
  95. Google gemini-2.5-pro — Throughput: 15.3 tps, Latency: 0.00s
  96. OpenAI gpt-5.1-codex — Throughput: 13.6 tps, Latency: 0.00s
  97. OpenAI gpt-4o-transcribe — Throughput: 13.4 tps, Latency: 0.00s
  98. OpenAI gpt-3.5-turbo — Throughput: 9.0 tps, Latency: 0.00s
  99. OpenAI gpt-4-turbo-preview — Throughput: 8.9 tps, Latency: 0.00s
  100. OpenAI GPT-5.3 Codex — Throughput: 8.4 tps, Latency: 0.00s