// Models
Put leading LLM models to the test on same prompt and see which one wins the AI battle
Claude Opus 4.6
Anthropic
Anthropic's most capable model for demanding reasoning and analysis tasks.
Claude Sonnet 4.6
Anthropic's balanced model offering strong quality at moderate cost.
Claude Sonnet 5
Anthropic's latest Sonnet flagship - strong quality at production-ready latency.
Gemini 3.1 Pro Preview
Google
Google's advanced preview model for complex reasoning and multi-step tasks.
Gemini 3.1 Flash Lite Preview
Google's fast, cost-efficient preview model optimized for high-throughput workloads.
GPT-6 Astra
OpenAI
GPT-6 Astra is a premium OpenAI model that accepts text and image input, with a 1.1M-token context window and up to 128K output tokens. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-4o
GPT-4o is a premium OpenAI model that accepts text and image input, with a 128K-token context window and up to 16K output tokens. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-5.6 Sol
GPT-5.6 Sol is a premium OpenAI model that accepts text and image input, with a 400K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-5.5 Pro
GPT-5.5 Pro is a premium OpenAI model that accepts text and image input, with a 400K-token context window. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-5.4 Pro
GPT-5.4 Pro is a premium OpenAI model that accepts text and image input, with a 400K-token context window. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-4.1
OpenAI's reliable GPT-4 series model for established production workloads.
GPT-5.6 Terra
GPT-5.6 Terra is a premium OpenAI model that accepts text and image input, with a 400K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-5.6 Luna
GPT-5.6 Luna is a premium OpenAI model that accepts text and image input, with a 400K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-5.5
GPT-5.5 is a premium OpenAI model that accepts text and image input, with a 400K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-5.4
GPT-5.4 is a premium OpenAI model that accepts text and image input, with a 400K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-5.4 mini
OpenAI's efficient flagship mini model for high-volume production workloads.
GPT-5.4 nano
GPT-5.4 nano is a OpenAI model that accepts text and image input, with a 400K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
GPT-5.2
OpenAI's proven workhorse model for production applications.
Claude Fable 5.1
Claude Fable 5.1 is a premium Anthropic model that accepts text and image input, with a 1M-token context window and up to 128K output tokens. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Claude Fable 5
Claude Fable 5 is a premium Anthropic model that accepts text and image input, with a 1M-token context window and up to 128K output tokens. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Claude Opus 4.8
Anthropic's top Opus model for demanding reasoning and long-context work.
Claude Opus 4.7
Anthropic Opus generation for complex analysis and high-stakes generation.
Claude Haiku 4.5
Anthropic's fastest model ideal for high-volume, cost-sensitive workflows.
Claude Opus 5
Claude Opus 5 is a premium Anthropic model that accepts text and image input, with a 200K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Gemini 3.5 Flash
Google's fast Gemini 3.5 Flash for high-throughput, cost-sensitive steps.
Grok 4.7
xAI
Grok 4.7 is a premium xAI model that accepts text and image input, with a 500K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Grok 4.6
Grok 4.6 is a premium xAI model that accepts text and image input, with a 500K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Grok 4.3
Grok 4.3 is a premium xAI model that accepts text and image input, with a 1M-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Grok 4.5
Grok 4.5 is a premium xAI model that accepts text and image input, with a 256K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Ministral 3 14B
Mistral
Compact Mistral model delivering fast responses at low cost.
Mistral Large 3
Mistral's large model for complex generation and reasoning tasks.
Mistral Medium 3.5
Mistral's medium-tier model balancing quality and cost for general workloads.
Mistral Small 4
Mistral's compact model for fast, low-cost inference at scale.
DeepSeek V4 Pro
DeepSeek
DeepSeek V4 Pro is a premium DeepSeek model for text input, with a 1M-token context window and up to 384K output tokens. On Test AI Models it lets you turn thinking on or off. Run it on your own prompt and compare its cost, speed and output side by side with other models.
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a DeepSeek model for text input, with a 1M-token context window and up to 384K output tokens. On Test AI Models it lets you turn thinking on or off. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Kimi K3
MoonshotAI
Kimi K3 is a premium MoonshotAI model that accepts text and image input, with a 256K-token context window. On Test AI Models it supports Low, Medium and High reasoning effort. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Kimi K2.6
MoonshotAI's updated Kimi model with strong cost-performance for general tasks.
Sonar
Perplexity
Perplexity's base search model for quick factual queries.
Sonar Pro
Perplexity's professional search-augmented model for research tasks.
Sonar Deep Research
Perplexity's deep research model for comprehensive information gathering.
Sonar Reasoning Pro
Perplexity's reasoning model with real-time web search integration.
Qwen3.8 Max
Alibaba
Qwen3.8 Max is a premium Alibaba model for text input, with a 1M-token context window. On Test AI Models it supports a capped thinking budget. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Qwen3.8 Flash
Qwen3.8 Flash is a Alibaba model for text input, with a 1M-token context window. On Test AI Models it supports a capped thinking budget. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Qwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is a premium Alibaba model for text input, with a 1M-token context window. On Test AI Models it supports a capped thinking budget. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Qwen3.8 27B
Qwen3.8 27B is a premium Alibaba model for text input, with a 1M-token context window. On Test AI Models it supports a capped thinking budget. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Qwen Plus
Qwen Plus is a Alibaba model for text input, with a 1M-token context window. On Test AI Models it supports a capped thinking budget. Run it on your own prompt and compare its cost, speed and output side by side with other models.
Qwen3 Coder Plus
Alibaba's specialized coding model for software development tasks.
Qwen3.6 Flash
Alibaba's fast Qwen 3.6 variant for high-volume, latency-sensitive steps.
Qwen3.7 Max
Alibaba's top Qwen 3.7 Max model for demanding reasoning and generation.
Qwen3.7 Plus
Alibaba's Qwen 3.7 Plus for strong quality at competitive inference cost.
Qwen3.6 Plus
Alibaba's Qwen 3.6 Plus for balanced generation quality and throughput.