Every AI model, in one place.
Pricing, benchmarks, provider latency, and how teams actually use each one.
About llm models
- NewqwenQwen: Qwen3.7 Flash
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...
Language1000k ctx$0.03/M - NewanthropicClaude Opus 5
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Language1000k ctx$5.00/M - NewanthropicClaude Opus 5 (batch)
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Language1000k ctx$2.50/M - NewIinclusionaiLing-3.0-flash
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Language262k ctx$0.02/M - NewPpoolsidePoolside: Laguna S 2.1
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
Language1049k ctx$0.09/M - NewPpoolsidePoolside: Laguna S 2.1 (free)
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
Language262k ctx$0.00/M - NewgoogleGoogle: Gemini 3.6 Flash
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Language1049k ctx$0.75/M - NewgoogleGoogle: Gemini 3.6 Flash (batch)
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Language1049k ctx$0.38/M - NewgoogleGoogle: Gemini 3.5 Flash Lite
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Language1049k ctx$0.30/M - NewgoogleGoogle: Gemini 3.5 Flash Lite (batch)
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Language1049k ctx$0.15/M - NewMmeituanMeituan: LongCat 2.0
LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...
Language1049k ctx$0.30/M - NewTthinkingmachinesThinking Machines: Inkling
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Language524k ctx$1.00/M - NewTthinkingmachinesThinking Machines: Inkling (batch)
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Language524k ctx$1.00/M - NewTthinkingmachinesThinking Machines: Inkling (free)
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Language1049k ctx$0.00/M - NewmoonshotaiMoonshotAI: Kimi K3
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Language1049k ctx$3.00/M - NewmoonshotaiMoonshotAI: Kimi K3 (batch)
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Language1049k ctx$3.00/M - NewmetaMeta: Muse Spark 1.1
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...
Language1049k ctx$1.25/M - NewKkwaipilotKwaipilot: KAT-Coder-Pro V2.5
KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
Language262k ctx$0.74/M - NewopenaiOpenAI: GPT-5.6 Luna Pro
GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Language1050k ctx$0.20/M - NewopenaiOpenAI: GPT-5.6 Luna Pro (batch)
GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Language1050k ctx$0.10/M - NewopenaiOpenAI: GPT-5.6 Luna
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
Language1050k ctx$0.20/M - NewopenaiOpenAI: GPT-5.6 Luna (batch)
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
Language1050k ctx$0.10/M - NewopenaiOpenAI: GPT-5.6 Terra Pro
GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Language1050k ctx$2.00/M - NewopenaiOpenAI: GPT-5.6 Terra Pro (batch)
GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Language1050k ctx$1.00/M - NewopenaiOpenAI: GPT-5.6 Terra
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
Language1050k ctx$2.00/M - NewopenaiOpenAI: GPT-5.6 Terra (batch)
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
Language1050k ctx$1.00/M - NewopenaiOpenAI: GPT-5.6 Sol Pro
GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Language1050k ctx$2.00/M - NewopenaiOpenAI: GPT-5.6 Sol Pro (batch)
GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Language1050k ctx$1.00/M - NewopenaiOpenAI: GPT-5.6 Sol
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
Language1050k ctx$2.00/M - NewopenaiOpenAI: GPT-5.6 Sol (batch)
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
Language1050k ctx$1.00/M - Newx-aiSpaceXAI: Grok 4.5
Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
Language500k ctx$2.00/M - Newx-aixAI: Grok Latest
This model always redirects to the latest Grok model from xAI.
Language500k ctx$2.00/M - NewALaion-labsAionLabs: Aion-3.0-Mini
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...
Language131k ctx$0.70/M - NewALaion-labsAionLabs: Aion-3.0
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...
Language131k ctx$3.00/M - NewTtencentTencent: Hy3
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
Language262k ctx$0.13/M - PpoolsidePoolside: Laguna XS 2.1
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
Language262k ctx$0.06/M - PpoolsidePoolside: Laguna XS 2.1 (free)
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
Language262k ctx$0.00/M - anthropicAnthropic: Claude Sonnet 5
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
Language1000k ctx$2.00/M - anthropicAnthropic: Claude Sonnet 5 (batch)
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
Language1000k ctx$1.00/M - NAnex-agiNex AGI: Nex-N2-Mini
Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text and image input and is built for coding, tool use,...
Language262k ctx$0.03/M - SsakanaSakana: Fugu Ultra
Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...
Language1000k ctx$5.00/M - cohereCohere: North Mini Code (free)
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
Language256k ctx$0.00/M - z-aiZ.ai: GLM 5.2
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Language1049k ctx$0.97/M - z-aiZ.ai: GLM 5.2 (free)
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Language256k ctx$0.00/M - openrouterOpenRouter: Fusion
Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...
Language1000k ctxFree tier - moonshotaiMoonshotAI: Kimi K2.7 Code
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
Language262k ctx$0.66/M - anthropicAnthropic: Claude Fable Latest
This model always redirects to the latest model in the Claude Fable family.
Language1000k ctx$10.00/M - anthropicAnthropic: Claude Fable 5
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Language1000k ctx$10.00/M
Prefer one page? Browse all AI models A–Z.