Every AI model, in one place.
Pricing, benchmarks, provider latency, and how teams actually use each one.
- NewmetaMeta: Muse Spark 1.3 Contributor
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...
Language1049k ctx$0.10/M - NewmetaMeta: Muse Spark 1.3
Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...
Language1049k ctx$1.25/M - NewgoogleGoogle: Gemini 3.8 Flash
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Language1049k ctx$0.75/M - NewgoogleGoogle: Gemini 3.8 Flash (batch)
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Language1049k ctx$0.38/M - NewanthropicAnthropic: Claude Fable 5.1
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
Language1000k ctx$10.00/M - NewanthropicAnthropic: Claude Fable 5.1 (batch)
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
Language1000k ctx$5.00/M - NewIinceptionInception: Mercury 2.5 Preview
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
Language260k ctx$0.04/M - NewIGibm-graniteIBM: Granite 4.2 8B
Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...
Language131k ctx$0.10/M - NewTtencentTencent: Hy4 preview
Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...
Language1049k ctx$0.83/M - NewIinclusionaiLing 3.0 Flash Fin
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...
Language262k ctx$0.06/M - NewIinclusionaiLing 3.0 Flash Fin (free)
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...
Language262k ctx$0.00/M - Newz-aiZ.ai: GLM Flash Latest
This model always redirects to the latest model in the GLM Flash family.
Language1049k ctx$0.07/M - NewqwenQwen: Qwen3.8 Flash
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.
Language1000k ctx$0.15/M - Newz-aiZ.ai: GLM 5.3 Flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Language1049k ctx$0.07/M - Newz-aiZ.ai: GLM 5.3 Flash (batch)
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Language1049k ctx$0.15/M - NewmetaMeta: Muse Spark 1.2 Contributor
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...
Language1049k ctx$0.10/M - NewdeepseekDeepSeek: DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...
Language1049k ctx$0.44/M - NewTtencentTencent: Hy-MT2-1.8B
Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided...
Language8k ctx$0.04/M - NewTtencentTencent: Hy-MT2-30B-A3B
Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and...
Language8k ctx$0.07/M - Newz-aiZ.ai: GLM Latest
This model always redirects to the latest GLM model from Z.ai.
Language262k ctx$1.15/M - NewTtencentTencent: Hy-MT2-7B
Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.
Language8k ctx$0.07/M - Newz-aiZ.ai: GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
Language1049k ctx$1.40/M - NewqwenQwen: Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
Language1000k ctx$0.42/M - NewDSdots-studioDots Studio: Dots3-Note Preview (free)
Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is...
Language512k ctx$0.00/M - NewgoogleGoogle: Gemini 3.7 Flash
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
Language1049k ctx$0.75/M - NewgoogleGoogle: Gemini 3.7 Flash (batch)
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
Language1049k ctx$0.38/M - NewBSbytedance-seedByteDance Seed: Seed 2.1 Turbo
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...
Language262k ctx$0.50/M - NewqwenQwen: Qwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
Language1000k ctx$2.00/M - NewqwenQwen: Qwen3.8 2.4T A95B (batch)
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
Language1010k ctx$2.00/M - NewBSbytedance-seedByteDance Seed: Seed-2.0-Code
Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude...
Language262k ctx$0.50/M - NewdeepseekDeepSeek: DeepSeek V4 Pro 0813
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
Language1024k ctx$1.12/M - NewdeepseekDeepSeek: DeepSeek V4 Pro 0813 (batch)
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
Language1049k ctx$1.32/M - Newx-aiSpaceXAI: Grok 4.6
Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Language500k ctx$2.00/M - NewLliquidLiquidAI: LFM2.5-2.6B (free)
LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or...
Language66k ctx$0.00/M - NewnvidiaNVIDIA: Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
Language262k ctx$0.08/M - NewnvidiaNVIDIA: Nemotron 3.5 Lightning (free)
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
Language1000k ctx$0.00/M - NewSsakanaSakana: Sakana Namazu
Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...
Language262k ctx$0.95/M - NewUupstageUpstage: Solar Pro 4
Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...
Language524k ctx$0.03/M - NewmetaMeta: Muse Glimmer 30B
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...
Language131k ctx$0.30/M - NewmetaMeta: Muse Glimmer 30B (batch)
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...
Language131k ctx$0.35/M - NewmetaMeta: Muse Spark 1.2
Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...
Language1049k ctx$1.25/M - NewqwenQwen: Qwen3.8 Max
Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,...
Language1000k ctx$2.00/M - NewdeepseekDeepSeek V4 Flash Latest
This model always redirects to the latest model in the DeepSeek V4 Flash family.
Language1049k ctx$0.05/M - NewdeepseekDeepSeek: DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Language1049k ctx$0.07/M - NewdeepseekDeepSeek: DeepSeek V4 Flash 0731 (batch)
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Language1049k ctx$0.14/M - NewTthinkingmachinesThinking Machines: Inkling Small
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Language524k ctx$0.45/M - NewTthinkingmachinesThinking Machines: Inkling Small (batch)
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Language524k ctx$0.50/M - NewTthinkingmachinesThinking Machines: Inkling Small (free)
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Language1049k ctx$0.00/M
Prefer one page? Browse all AI models A–Z.