LLMmoonshotai

MoonshotAI: Kimi K3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Anyone in the Project can @-mention MoonshotAI: Kimi K3 with the team's shared context - pooled credits, one chat, one memory.

All models

Starter is free forever - 1 Project, 100 credits/month, 1 MCP. No card.

Verdict

Kimi K3 offers a massive 1M-token context window at $3 input / $15 output per Mtok — roughly half the cost of GPT-4o for long-context work. The model handles text and image inputs, making it viable for multimodal document analysis. Without public benchmarks, you're trading proven performance data for cost savings on context-heavy tasks. Reach for this when budget and context length matter more than established track record.

Best for

  • Budget-conscious long-document analysis
  • Processing large codebases under $50
  • Multimodal research paper review
  • Cost-sensitive customer support transcripts
  • High-volume context window experiments

Strengths

The 1M-token context window matches top-tier models while undercutting them on price — you can process a 200-page technical manual for under $1 in input costs. Multimodal support means you can mix screenshots, diagrams, and text in a single prompt without preprocessing. The pricing structure favors read-heavy workloads where you're ingesting large documents but generating concise summaries or answers.

Trade-offs

No public benchmark data means you're flying blind on reasoning quality, code generation accuracy, or how it stacks up against Claude or GPT-4o on standard evals. The $15/Mtok output cost climbs fast if you're generating long responses — a 10k-token summary costs $0.15, which adds up at scale. Vision capabilities are unproven in head-to-head comparisons, so expect to validate performance on your own image-heavy tasks before committing production traffic.

Specifications

Provider
moonshotai
Category
llm
Context length
1,048,576 tokens
Max output
943,718 tokens
Modalities
text, image, video
License
proprietary
Released
2026-07-16

Pricing

Input
$3.00/Mtok
Output
$15.00/Mtok
Model ID
moonshotai/kimi-k3

Per-token prices show what the model costs upstream. On Switchy your team draws from one shared org credit pool - one plan, one balance for everyone.

Team cost calculator

Estimated monthly spend
$116.16
17.6M tokens / month
5 seats · 80 msgs/day

Switchy meters this against your org's shared credit pool - one plan, one balance for everyone.

Providers

Provider-level routing data is not available yet for this model.

Performance

Performance snapshots are collected daily. Check back after the next ingestion run.

Benchmarks

Public benchmark scores are not available yet for this model. Check back after the next ingestion run.

Works well with

Top MCPs

Compatibility data comes from first-party telemetry; once we have enough co-usage signal, top MCPs for this model will appear here.

How Switchy teams use it

Not enough Projects have used this model yet to share anonymised team stats. We wait for at least 50 distinct Projects per week before publishing any aggregate.

Starter prompts

Analyze Codebase Architecture

Review all files in this repository. Describe the overall architecture, identify the main modules and their responsibilities, and flag any circular dependencies or code smells.
Open in a Project →

Compare Research Papers

I'm attaching five papers on transformer architectures. Compare their methodologies, benchmark results, and novel contributions. Highlight where they agree or contradict each other.
Open in a Project →

Extract Data from Invoices

Extract the following fields from each invoice image: vendor name, invoice number, date, line items with quantities and prices, subtotal, tax, and total. Return results as JSON.
Open in a Project →

Audit Support Ticket History

Analyze this year's support tickets. Group issues by root cause, rank them by frequency, and recommend three process changes that would reduce ticket volume.
Open in a Project →

Compare with

More language models

See all language models

Data last verified 7 hours ago.Sources aggregated hourly to weekly. See docs/architecture/model-directory.md.