Skip to main content
The model catalog is generated from GET /models every hour. This page explains how the numbers on it are derived, so you can reproduce them and know what they do and don’t mean.

Families and variants

Many model IDs are the same model offered in a different way. The catalog groups them into one family page: Every variant keeps its own model ID, price and privacy tier. Pass the variant’s ID as model. Dated snapshots such as deepseek-v4-flash-0731 stay separate families and appear under Other versions.

Normalized prices

Each modality bills in its own unit. The catalog shows the native unit and one reference price, so models can be sorted and compared.
Reasoning models can spend many more output tokens than non-reasoning models on the same task, so the blended price understates their real cost. Cost per task, measured on benchmark runs, is planned; see Benchmarks.

Video pricing

Video models don’t publish prices in GET /models. At build time, Venice calls POST /video/quote for every resolution × duration × audio combination each model supports, at its first listed aspect ratio (aspect ratio does not change the price). The catalog stores the full matrix and derives:
  • Per second: the median of price ÷ seconds across durations at a given resolution and audio setting. Prices are linear in duration for almost every model.
  • Draft clip: 5 seconds, 720p, silent.
  • Production clip: 10 seconds, 1080p, with audio.
When a model doesn’t offer the requested setting, the catalog uses the closest one it supports and marks the price. Edit, motion-control and upscale modes bill by the length of your source video, so they show By source instead of a price. Always call POST /video/quote before queueing for an exact amount.

Privacy tiers

A family lists every tier its variants offer. Filtering by a tier shows the variant in that tier, with its own price.

Reference prompt suite

Image and video models are rendered on a fixed prompt suite, the same prompts for every model of a modality, at default settings. That’s what makes the side-by-side view in Compare meaningful. Models without renders show Reference renders pending.

Performance

Not published yet. This section describes what the catalog will show.
Venice will measure models with synthetic probes, never customer prompts, and publish: Every figure will show its window (rolling 30 days for uptime, 7 days for latency), percentile, sample count and probe region. Uptime excludes 4xx errors caused by the request. Cache hit rate is cache_read ÷ (input + cache_read + cache_write) across all traffic. Models with too few samples will be marked Preliminary.

Benchmarks

Not published yet. This section describes what the catalog will show.
Benchmarks are chosen for the frontier: still discriminative at the top, resistant to contamination, and runnable against the model as Venice serves it. Every score will carry one provenance label, never averaged together:
  • Venice-verified: run on the Venice endpoint shown on the page.
  • Independent: run by a third party, usually on the lab’s own API.
  • Lab-reported: published by the model’s creator.
Each score also shows the benchmark and harness version, reasoning effort, test date and a confidence interval. Scores older than 90 days, or superseded by a new benchmark version, are flagged as stale. The set is reviewed every quarter. A benchmark becomes a floor check, no longer a headline, once frontier models saturate it or it fails a contamination audit.

Data sources

The catalog is built from public endpoints you can call yourself: Prices are in USD. DIEM is billed at parity.