AI Model Speed Rankings
Models ranked by measured output throughput (tokens per second) and average latency.
- MiniMax MiniMax H3 — Throughput: 1131.6 tps, Latency: 690313.00s
- Google gemini-2.5-pro — Throughput: 721.9 tps, Latency: 24393.00s
- OpenAI gpt-5-nano — Throughput: 440.6 tps, Latency: 11361.62s
- Anthropic Claude Sonnet 5 — Throughput: 385.1 tps, Latency: 7721.65s
- Google Gemini 3.5 Flash — Throughput: 370.8 tps, Latency: 8645.77s
- xAI Grok 4.20 Multi-Agent — Throughput: 326.0 tps, Latency: 0.00s
- Google Gemini Omni Flash Preview — Throughput: 306.6 tps, Latency: 99229.00s
- Google Gemini 3.1 Flash-Lite — Throughput: 282.7 tps, Latency: 1796.88s
- Google Gemini 3.6 Flash — Throughput: 279.1 tps, Latency: 7922.77s
- OpenAI GPT-5 — Throughput: 248.3 tps, Latency: 18403.50s
- OpenAI GPT-5.4 Mini — Throughput: 230.0 tps, Latency: 1300.79s
- Google gemini-2.5-flash — Throughput: 202.6 tps, Latency: 28581.58s
- Google Gemini 3 Flash Preview — Throughput: 202.4 tps, Latency: 6075.40s
- Google Gemini 3.1 Flash-Lite Preview — Throughput: 186.4 tps, Latency: 1320.66s
- OpenAI o3-mini — Throughput: 168.5 tps, Latency: 4605.00s
- OpenAI dall-e-3 — Throughput: 149.2 tps, Latency: 27980.00s
- OpenAI GPT-5.6 Terra — Throughput: 140.7 tps, Latency: 6363.70s
- OpenAI gpt-4.1 — Throughput: 131.8 tps, Latency: 693.47s
- Google Gemini 3 Pro Preview — Throughput: 129.5 tps, Latency: 0.00s
- Anthropic Claude Haiku 4.5 — Throughput: 128.5 tps, Latency: 3491.69s
- Google Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image 🍌) — Throughput: 128.1 tps, Latency: 11488.87s
- Google Gemini 3.1 Pro Preview Customtools — Throughput: 124.8 tps, Latency: 39148.25s
- OpenAI gpt-image-1-mini — Throughput: 122.7 tps, Latency: 50856.00s
- Google Gemini 3.1 Pro Preview — Throughput: 121.7 tps, Latency: 22790.45s
- OpenAI gpt-oss-120b — Throughput: 119.6 tps, Latency: 0.00s
- OpenAI GPT-5.1 — Throughput: 118.5 tps, Latency: 1704.44s
- OpenAI gpt-4o-all — Throughput: 117.7 tps, Latency: 2919.87s
- Anthropic Claude Opus 5 — Throughput: 114.7 tps, Latency: 15064.83s
- xAI grok-4-1-fast-reasoning — Throughput: 110.1 tps, Latency: 422.00s
- Google Nano Banana (Gemini 2.5 Flash Image 🍌) — Throughput: 107.9 tps, Latency: 8254.25s
- OpenAI GPT Image 1.5 — Throughput: 104.4 tps, Latency: 63894.50s
- DeepSeek DeepSeek-V4-Flash 0731 — Throughput: 103.8 tps, Latency: 886.41s
- xAI grok-4-0709 — Throughput: 97.4 tps, Latency: 1546.00s
- OpenAI GPT-5.5 — Throughput: 96.4 tps, Latency: 4919.12s
- xAI grok-3 — Throughput: 95.9 tps, Latency: 0.00s
- xAI Grok 4.3 — Throughput: 95.2 tps, Latency: 1419.00s
- Anthropic Claude Opus 4.7 — Throughput: 94.8 tps, Latency: 11498.95s
- xAI grok-3-mini — Throughput: 93.1 tps, Latency: 0.00s
- OpenAI gpt-4o-search-preview — Throughput: 86.8 tps, Latency: 0.00s
- MiniMax MiniMax M2.5 — Throughput: 85.9 tps, Latency: 0.00s
- Z.ai GLM 5.2 — Throughput: 85.8 tps, Latency: 5092.67s
- OpenAI GPT-5.4 — Throughput: 81.5 tps, Latency: 3000.35s
- OpenAI GPT-5.6 Luna — Throughput: 78.9 tps, Latency: 3829.35s
- xAI Grok 4.20 (Non-Reasoning) — Throughput: 78.5 tps, Latency: 708.53s
- Anthropic Claude Fable 5 — Throughput: 77.8 tps, Latency: 9170.78s
- Alibaba qwen-plus — Throughput: 74.2 tps, Latency: 1548.65s
- Google Nano Banana 2 (Gemini 3.1 Flash Image 🍌) — Throughput: 73.5 tps, Latency: 12459.19s
- xAI grok-4-fast-reasoning — Throughput: 68.0 tps, Latency: 0.00s
- xAI grok-4-1-fast-non-reasoning — Throughput: 66.9 tps, Latency: 470.00s
- Moonshot Kimi K2.6 — Throughput: 65.5 tps, Latency: 895.50s
- OpenAI GPT-5.2 — Throughput: 65.5 tps, Latency: 4025.70s
- DeepSeek DeepSeek-V4-Pro — Throughput: 64.1 tps, Latency: 1612.33s
- xAI Grok Build 0.1 — Throughput: 62.4 tps, Latency: 397.00s
- OpenAI gpt-5.1-codex-mini — Throughput: 60.1 tps, Latency: 0.00s
- Anthropic Claude Sonnet 4.6 — Throughput: 58.6 tps, Latency: 4509.34s
- xAI Grok 4.5 — Throughput: 58.5 tps, Latency: 2391.28s
- DeepSeek deepseek-v3.1 — Throughput: 56.9 tps, Latency: 3098.00s
- Anthropic Claude Opus 4.6 — Throughput: 56.5 tps, Latency: 4857.55s
- OpenAI GPT-5.6 Sol — Throughput: 56.2 tps, Latency: 17526.88s
- DeepSeek DeepSeek-V4-Pro-0813 — Throughput: 55.5 tps, Latency: 0.00s
- OpenAI gpt-4o-2024-11-20 — Throughput: 54.8 tps, Latency: 0.00s
- Anthropic Claude Opus 4.8 — Throughput: 54.2 tps, Latency: 13644.54s
- OpenAI GPT-5.5 Pro — Throughput: 49.9 tps, Latency: 206556.60s
- xAI Grok 4.20 — Throughput: 49.6 tps, Latency: 334.50s
- xAI Grok 4.6 — Throughput: 49.5 tps, Latency: 3461.22s
- OpenAI gpt-4o-2024-08-06 — Throughput: 49.3 tps, Latency: 0.00s
- Google Gemini 3.5 Flash-Lite — Throughput: 45.7 tps, Latency: 1321.85s
- Anthropic Claude Sonnet 4.5 — Throughput: 44.0 tps, Latency: 6310.48s
- OpenAI o1 — Throughput: 43.2 tps, Latency: 0.00s
- OpenAI gpt-3.5-turbo-16k — Throughput: 42.5 tps, Latency: 0.00s
- OpenAI gpt-5-mini — Throughput: 39.0 tps, Latency: 7740.35s
- OpenAI gpt-3.5-turbo — Throughput: 36.8 tps, Latency: 0.00s
- DeepSeek DeepSeek-V3.2 — Throughput: 35.8 tps, Latency: 1105.50s
- OpenAI GPT Image 1.0 — Throughput: 34.7 tps, Latency: 35973.67s
- MiniMax MiniMax M3 — Throughput: 34.4 tps, Latency: 0.00s
- Moonshot Kimi K3 — Throughput: 33.8 tps, Latency: 5622.19s
- OpenAI GPT Image 2 — Throughput: 33.4 tps, Latency: 52559.19s
- ByteDance Doubao-Seed-2.0-pro — Throughput: 31.7 tps, Latency: 0.00s
- OpenAI o3 — Throughput: 31.4 tps, Latency: 951.00s
- OpenAI GPT-5.4 Nano — Throughput: 31.4 tps, Latency: 1240.00s
- OpenAI gpt-4o — Throughput: 30.9 tps, Latency: 650.86s
- MiniMax MiniMax-M2.7 — Throughput: 30.8 tps, Latency: 0.00s
- Google Nano Banana Pro (Gemini 3 Pro Image 🍌) — Throughput: 30.3 tps, Latency: 35780.89s
- Xiaomi MiMo-V2.5 — Throughput: 30.1 tps, Latency: 2167.33s
- Alibaba qwen-max — Throughput: 27.0 tps, Latency: 657.00s
- OpenAI gpt-4.1-mini — Throughput: 25.1 tps, Latency: 7493.52s
- OpenAI gpt-4o-transcribe — Throughput: 24.9 tps, Latency: 0.00s
- OpenAI gpt-4o-mini — Throughput: 24.3 tps, Latency: 1978.00s
- OpenAI o4-mini — Throughput: 22.2 tps, Latency: 0.00s
- Yi yi-lightning — Throughput: 17.1 tps, Latency: 0.00s
- OpenAI gpt-3.5-turbo-1106 — Throughput: 17.0 tps, Latency: 0.00s
- OpenAI gpt-4-gizmo-* — Throughput: 15.9 tps, Latency: 44853.83s
- Anthropic Claude Opus 4.5 — Throughput: 15.5 tps, Latency: 2517.46s
- OpenAI gpt-5.1-codex — Throughput: 15.1 tps, Latency: 0.00s
- Google Gemini 3.7 Flash — Throughput: 14.8 tps, Latency: 87240.00s
- Alibaba qwen3-vl-flash — Throughput: 13.2 tps, Latency: 0.00s
- xAI grok-4-fast-non-reasoning — Throughput: 12.3 tps, Latency: 0.00s
- OpenAI gpt-4o-mini-transcribe — Throughput: 7.6 tps, Latency: 0.00s
- OpenAI gpt-5-codex — Throughput: 7.1 tps, Latency: 0.00s
- OpenAI GPT-5.2 Codex — Throughput: 6.5 tps, Latency: 0.00s