Plain-text specification
Plain-text specification
Inkling API
Inkling is a large language model by Inkling, available on the Venice API asinkling. It runs privately, with zero data retention.Inkling is a general-purpose multimodal model from Thinking Machines Lab that accepts text, image, and audio inputs and generates text. It is a 66-layer sparse MoE (975B total / 41B active) with hybrid local/global attention, 512K context, and variable thinking effort — suited for chat, coding, tool use, and agentic workflows. Video input is not supported on Venice.Inkling API pricing
Inkling specifications
How to use the Inkling API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "inkling" and your API key.Inkling API FAQ
How much does the Inkling API cost?
5.06 per 1M output tokens, with cached input at $0.21 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the Inkling model ID?
Useinkling as the model parameter.Is the Inkling API private?
Inkling is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of Inkling?
512K tokens of context, with up to 64K output tokens per response.What does Inkling support?
Inkling supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default high).Which endpoint does the Inkling API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- GLM 5.2 API: 4.40 output per 1M tokens
- DeepSeek V4 Pro 0813 API: 4.95 output per 1M tokens
- Gemini 3.6 Flash API: 4.69 output per 1M tokens
- DeepSeek V4 Pro API: 3.30 output per 1M tokens