Plain-text specification
Plain-text specification
OpenAI GPT OSS 120B API
OpenAI GPT OSS 120B is a large language model by OpenAI, available on the Venice API asopenai-gpt-oss-120b. Private and end-to-end encrypted variants are available.gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generationOpenAI GPT OSS 120B API pricing
OpenAI GPT OSS 120B specifications
How to use the OpenAI GPT OSS 120B API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "openai-gpt-oss-120b" and your API key.OpenAI GPT OSS 120B API FAQ
How much does the OpenAI GPT OSS 120B API cost?
0.30 per 1M output tokens. The E2EE variant costs 0.65 output. Prices are in USD and can be paid in DIEM at parity.What is the OpenAI GPT OSS 120B model ID?
Useopenai-gpt-oss-120b as the model parameter. Other variants: e2ee-gpt-oss-120b-p (E2EE).Is the OpenAI GPT OSS 120B API private?
The Standard variant is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training. The E2EE variant is end-to-end encrypted: prompts are encrypted on your device and decrypted only inside an attested hardware enclave, so neither Venice nor the GPU provider can read them.What is the context window of OpenAI GPT OSS 120B?
125K tokens of context, with up to 16K output tokens per response.What does OpenAI GPT OSS 120B support?
OpenAI GPT OSS 120B supports function calling, reasoning and web search. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default low).Which endpoint does the OpenAI GPT OSS 120B API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Google Gemma 3 27B Instruct API: 0.20 output per 1M tokens
- Mistral Small 3.2 24B Instruct API: 0.25 output per 1M tokens
- NVIDIA Nemotron 3 Nano 30B API: 0.30 output per 1M tokens
- GLM 4.7 Flash API: 0.40 output per 1M tokens