LLMqwen

Qwen: Qwen3 VL 30B A3B Instruct

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...

Anyone in the Project can @-mention Qwen: Qwen3 VL 30B A3B Instruct with the team's shared context - pooled credits, one chat, one memory.

All models

Starter is free forever - 1 Project, 100 credits/month, 1 MCP. No card.

Verdict

Qwen3 VL 30B A3B Instruct is a mid-sized vision-language model offering a massive 262K token context window at aggressive pricing ($0.15/$0.60 per Mtok). The model handles both text and images, making it suitable for multimodal workflows where cost and long-context capability matter more than bleeding-edge accuracy. Without public benchmarks, you're trading proven performance data for price advantage. Reach for this when you need vision capabilities in high-volume applications where budget constraints are tight and you can tolerate some uncertainty around accuracy relative to established models like GPT-4V or Claude Sonnet.

Best for

  • Budget-conscious vision-language tasks
  • Long-context document analysis with images
  • High-volume multimodal processing
  • Screenshot analysis at scale
  • Cost-sensitive OCR and visual QA

Strengths

The 262K context window is exceptional for a model at this price point, enabling full-document processing with embedded images without chunking. Input pricing at $0.15 per million tokens undercuts most vision-capable models by 50-70%, making it viable for high-throughput applications. The 30B parameter count suggests reasonable capability for standard vision-language tasks like image captioning, visual question answering, and document extraction where you don't need frontier-model precision.

Trade-offs

Absence of public benchmarks means you're flying blind on accuracy relative to GPT-4V, Claude Sonnet 4.5, or Gemini Pro Vision. The proprietary license limits deployment flexibility compared to open-weight alternatives. Output pricing at $0.60 per Mtok is 4x the input rate, so verbose responses erode the cost advantage quickly. The A3B designation suggests this may be a quantized or efficiency-optimized variant, potentially sacrificing some accuracy for speed and cost.

Specifications

Provider
qwen
Category
llm
Context length
262,144 tokens
Max output
16,384 tokens
Modalities
text, image
License
proprietary
Released
2025-10-06

Pricing

Input
$0.15/Mtok
Output
$0.60/Mtok
Model ID
qwen/qwen3-vl-30b-a3b-instruct

Per-token prices show what the model costs upstream. On Switchy your team draws from one shared org credit pool - one plan, one balance for everyone.

Team cost calculator

Estimated monthly spend
$5.02
17.6M tokens / month
5 seats · 80 msgs/day

Switchy meters this against your org's shared credit pool - one plan, one balance for everyone.

Providers

ProviderContextInputOutputP50 latencyThroughput30d uptime
qwen131k$0.13/Mtok$0.52/Mtok

Performance

Performance snapshots are collected daily. Check back after the next ingestion run.

Benchmarks

Public benchmark scores are not available yet for this model. Check back after the next ingestion run.

Works well with

Top MCPs

Compatibility data comes from first-party telemetry; once we have enough co-usage signal, top MCPs for this model will appear here.

How Switchy teams use it

Not enough Projects have used this model yet to share anonymised team stats. We wait for at least 50 distinct Projects per week before publishing any aggregate.

Starter prompts

Extract Invoice Line Items

Extract all line items from this invoice image into a JSON array. For each item include: description, quantity, unit_price, and total. Return only valid JSON with no additional commentary.
Open in a Project →

Analyze Multi-Page Report

Review this multi-page report with embedded charts. Summarize the key findings in three bullet points, then identify any data visualizations that contradict the written conclusions.
Open in a Project →

Screenshot Troubleshooting

This is a screenshot of an error state in our application. Describe what the user is seeing, identify the likely cause based on visible UI elements, and suggest two troubleshooting steps.
Open in a Project →

Batch Image Captioning

Generate a concise, descriptive caption for this image in 15-25 words. Focus on the main subject, setting, and any notable actions or objects. Use natural language suitable for alt text.
Open in a Project →

Visual QA for E-commerce

Based on this product image, answer the following question accurately and concisely: [QUESTION]. If the answer isn't visible in the image, state that clearly rather than guessing.
Open in a Project →

Compare with

More language models

See all language models

Data last verified 7 hours ago.Sources aggregated hourly to weekly. See docs/architecture/model-directory.md.