Plain-text specification
Plain-text specification
Qwen 3 235B A22B Thinking 2507 API
Qwen 3 235B A22B Thinking 2507 is a large language model by Alibaba Qwen, available on the Venice API asqwen3-235b-a22b-thinking-2507. It runs privately, with zero data retention.Built for in-depth research and handling long, complex documents. Ideal for technical work, multimodal input, and high-precision tasks.Qwen 3 235B A22B Thinking 2507 API pricing
Qwen 3 235B A22B Thinking 2507 specifications
How to use the Qwen 3 235B A22B Thinking 2507 API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "qwen3-235b-a22b-thinking-2507" and your API key.Qwen 3 235B A22B Thinking 2507 API FAQ
How much does the Qwen 3 235B A22B Thinking 2507 API cost?
3.50 per 1M output tokens. Prices are in USD and can be paid in DIEM at parity.What is the Qwen 3 235B A22B Thinking 2507 model ID?
Useqwen3-235b-a22b-thinking-2507 as the model parameter.Is the Qwen 3 235B A22B Thinking 2507 API private?
Qwen 3 235B A22B Thinking 2507 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of Qwen 3 235B A22B Thinking 2507?
125K tokens of context, with up to 16K output tokens per response.What does Qwen 3 235B A22B Thinking 2507 support?
Qwen 3 235B A22B Thinking 2507 supports function calling, structured outputs, reasoning and web search. Reasoning effort is adjustable withreasoning_effort: low, medium and high (default low).Which endpoint does the Qwen 3 235B A22B Thinking 2507 API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Llama 3.3 70B API: 2.80 output per 1M tokens
- Kimi K2.5 API: 3.50 output per 1M tokens
- GLM 4.7 API: 2.65 output per 1M tokens
- Hermes 3 Llama 3.1 405b API: 3.00 output per 1M tokens