Plain-text specification
Plain-text specification
DeepSeek V4 Flash 0423 API
DeepSeek V4 Flash 0423 is a large language model by DeepSeek, available on the Venice API asdeepseek-v4-flash. Private and end-to-end encrypted variants are available.DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.DeepSeek V4 Flash 0423 API pricing
DeepSeek V4 Flash 0423 specifications
How to use the DeepSeek V4 Flash 0423 API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "deepseek-v4-flash" and your API key.DeepSeek V4 Flash 0423 API FAQ
How much does the DeepSeek V4 Flash 0423 API cost?
0.28 per 1M output tokens, with cached input at 0.18 input and $0.37 output. Prices are in USD and can be paid in DIEM at parity.What is the DeepSeek V4 Flash 0423 model ID?
Usedeepseek-v4-flash as the model parameter. Other variants: e2ee-deepseek-v4-flash (E2EE).Is the DeepSeek V4 Flash 0423 API private?
The Standard variant is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training. The E2EE variant is end-to-end encrypted: prompts are encrypted on your device and decrypted only inside an attested hardware enclave, so neither Venice nor the GPU provider can read them.What is the context window of DeepSeek V4 Flash 0423?
1M tokens of context, with up to 32K output tokens per response.What does DeepSeek V4 Flash 0423 support?
DeepSeek V4 Flash 0423 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default low).Which endpoint does the DeepSeek V4 Flash 0423 API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- DeepSeek V4.1 Flash API: 1.50 output per 1M tokens
- DeepSeek V4 Flash 0731 API: 0.35 output per 1M tokens
- Google Gemma 4 31B Instruct API: 0.36 output per 1M tokens
- Google Gemma 4 26B A4B Instruct API: 0.40 output per 1M tokens
- GLM 4.7 Flash Heretic API: 0.40 output per 1M tokens
- GLM 4.7 Flash API: 0.40 output per 1M tokens