LLMthinkingmachines

Thinking Machines: Inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Anyone in the Project can @-mention Thinking Machines: Inkling with the team's shared context - pooled credits, one chat, one memory.

All models

Starter is free forever - 1 Project, 100 credits/month, 1 MCP. No card.

Verdict

Inkling offers a massive 524K token context window at aggressive pricing — $1 input / $4.05 output per million tokens — making it a strong candidate for long-document workflows where cost matters. Multimodal support (text, image, audio) adds flexibility for mixed-media tasks. The catch: no public benchmarks yet, so you're flying blind on reasoning quality and accuracy relative to established models. Reach for this when context length and budget are your top constraints and you can validate outputs internally.

Best for

  • Long-document analysis under budget
  • Multimodal tasks with audio input
  • Cost-sensitive batch processing
  • Prototyping with extended context
  • Mixed-media content workflows

Strengths

The 524K context window handles entire codebases, legal filings, or multi-hour transcripts in a single pass. At $1 input / $4.05 output per Mtok, it undercuts many competitors on long-context tasks where token volume drives cost. Native audio support means you can feed meeting recordings or podcasts directly without preprocessing. The pricing structure favors read-heavy workloads — ideal for summarization, extraction, and analysis where input tokens dominate.

Trade-offs

No public benchmarks means you can't compare reasoning quality, instruction-following, or accuracy against Claude, GPT-4, or Gemini. You'll need to run your own evals before trusting it for high-stakes tasks. The $4.05 output rate climbs quickly if you generate long responses — fine for extraction, less so for drafting. Multimodal capabilities are listed but undocumented, so expect trial-and-error on image and audio quality. Without a track record, you're an early adopter taking on validation overhead.

Specifications

Provider
thinkingmachines
Category
llm
Context length
524,288 tokens
Max output
471,859 tokens
Modalities
text, image, audio
License
proprietary
Released
2026-07-17

Pricing

Input
$1.00/Mtok
Output
$4.05/Mtok
Model ID
thinkingmachines/inkling

Per-token prices show what the model costs upstream. On Switchy your team draws from one shared org credit pool - one plan, one balance for everyone.

Team cost calculator

Estimated monthly spend
$33.70
17.6M tokens / month
5 seats · 80 msgs/day

Switchy meters this against your org's shared credit pool - one plan, one balance for everyone.

Providers

Provider-level routing data is not available yet for this model.

Performance

Performance snapshots are collected daily. Check back after the next ingestion run.

Benchmarks

Public benchmark scores are not available yet for this model. Check back after the next ingestion run.

Works well with

Top MCPs

Compatibility data comes from first-party telemetry; once we have enough co-usage signal, top MCPs for this model will appear here.

How Switchy teams use it

Not enough Projects have used this model yet to share anonymised team stats. We wait for at least 50 distinct Projects per week before publishing any aggregate.

Starter prompts

Summarize Long Transcript

You have a full transcript of a 3-hour board meeting. Extract the top 5 decisions made, any unresolved questions, and a one-paragraph summary of the overall discussion.
Open in a Project →

Extract Data from PDF

This is a 200-page clinical trial report. Extract the primary endpoint, sample size, statistical significance, and any adverse events mentioned. Return as a bulleted list.
Open in a Project →

Analyze Codebase Context

Here's the complete source code for a Python web app. Identify any repeated logic that could be refactored into shared utilities, and flag any security concerns in authentication flows.
Open in a Project →

Compare Audio Interviews

I've uploaded 8 customer interview recordings. Identify the top 3 pain points mentioned across all interviews and quote specific examples from each.
Open in a Project →

Multimodal Report Generation

You have a slide deck (images), speaker notes (text), and the recorded presentation (audio). Write a 500-word summary that integrates insights from all three sources.
Open in a Project →

Compare with

More language models

See all language models

Data last verified 7 hours ago.Sources aggregated hourly to weekly. See docs/architecture/model-directory.md.