Plain-text specification
Plain-text specification
DeepSeek V4.1 Flash API
DeepSeek V4.1 Flash is a large language model by DeepSeek, available on the Venice API asdeepseek-v4-1-flash. It runs privately, with zero data retention.DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters (8B active on input, 16B on output) and a 1M-token context window. It natively processes images and text, with strong reasoning, coding, and agentic performance.DeepSeek V4.1 Flash API pricing
DeepSeek V4.1 Flash specifications
How to use the DeepSeek V4.1 Flash API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "deepseek-v4-1-flash" and your API key.DeepSeek V4.1 Flash API FAQ
How much does the DeepSeek V4.1 Flash API cost?
1.50 per 1M output tokens, with cached input at $0.0075 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the DeepSeek V4.1 Flash model ID?
Usedeepseek-v4-1-flash as the model parameter.Is the DeepSeek V4.1 Flash API private?
DeepSeek V4.1 Flash is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of DeepSeek V4.1 Flash?
1M tokens of context, with up to 128K output tokens per response.What does DeepSeek V4.1 Flash support?
DeepSeek V4.1 Flash supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, high and max (default high).Which endpoint does the DeepSeek V4.1 Flash API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- DeepSeek V4 Flash 0731 API: 0.35 output per 1M tokens
- DeepSeek V4 Flash 0423 API: 0.28 output per 1M tokens
- GPT-5.6 Luna API: 1.50 output per 1M tokens
- GPT-5.6 Luna Pro API: 1.50 output per 1M tokens
- MiniMax M2.7 API: 1.50 output per 1M tokens
- MiMo-V2.5 API: 2.00 output per 1M tokens