Plain-text specification
Plain-text specification
Qwen 3.8 Flash API
Qwen 3.8 Flash is a large language model by Alibaba Qwen, available on the Venice API asqwen-3-8-flash. Requests are anonymized, so the provider never sees your identity.Qwen 3.8 Flash is the latest multimodal model in the Qwen family, pairing strong reasoning and generation with remarkable speed. It natively supports a 1M-token context window, so it can process long documents, entire codebases, and complex conversations in a single pass. It excels at coding assistance, agentic workflows, and visual understanding, accepting text, image, and video input, and its thinking mode can be turned on or off per request.Qwen 3.8 Flash API pricing
Qwen 3.8 Flash specifications
How to use the Qwen 3.8 Flash API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "qwen-3-8-flash" and your API key.Qwen 3.8 Flash API FAQ
How much does the Qwen 3.8 Flash API cost?
0.49 per 1M output tokens, with cached input at $0.014 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the Qwen 3.8 Flash model ID?
Useqwen-3-8-flash as the model parameter.Is the Qwen 3.8 Flash API private?
Qwen 3.8 Flash is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.What is the context window of Qwen 3.8 Flash?
1M tokens of context, with up to 128K output tokens per response.What does Qwen 3.8 Flash support?
Qwen 3.8 Flash supports function calling, structured outputs, reasoning, image input, web search and prompt caching.Which endpoint does the Qwen 3.8 Flash API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- MiMo-V2.6-Flash API: 0.35 output per 1M tokens
- GLM 5.3 Flash API: 0.50 output per 1M tokens
- DeepSeek V4 Flash 0731 API: 0.35 output per 1M tokens
- GPT-6 Luna API: 0.63 output per 1M tokens