Skip to main content

DeepSeek V4 Flash 0731 API

DeepSeek V4 Flash 0731 is a large language model by DeepSeek, available on the Venice API as deepseek-v4-flash-0731. It runs privately, with zero data retention.DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.

DeepSeek V4 Flash 0731 API pricing

DeepSeek V4 Flash 0731 specifications

How to use the DeepSeek V4 Flash 0731 API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "deepseek-v4-flash-0731" and your API key.

DeepSeek V4 Flash 0731 API FAQ

How much does the DeepSeek V4 Flash 0731 API cost?

0.17per1Minputtokensand0.17 per 1M input tokens and 0.35 per 1M output tokens, with cached input at 0.035per1M.TheFastvariantcosts0.035 per 1M. The Fast variant costs 0.35 input and $0.70 output. Prices are in USD and can be paid in DIEM at parity.

What is the DeepSeek V4 Flash 0731 model ID?

Use deepseek-v4-flash-0731 as the model parameter. Other variants: deepseek-v4-flash-0731-fast (Fast).

Is the DeepSeek V4 Flash 0731 API private?

DeepSeek V4 Flash 0731 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of DeepSeek V4 Flash 0731?

1M tokens of context, with up to 32K output tokens per response.

What does DeepSeek V4 Flash 0731 support?

DeepSeek V4 Flash 0731 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, high and max (default high).

Which endpoint does the DeepSeek V4 Flash 0731 API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models