> ## Documentation Index
> Fetch the complete documentation index at: https://veniceai-feat-models-redesign.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# How the model catalog works

> How Venice builds the model catalog: model families and variants, normalized prices for every modality, privacy tiers, the reference prompt suite, and how performance and benchmarks will be measured.

The [model catalog](/models/overview) is generated from [`GET /models`](/api-reference/endpoint/models/list) every hour. This page explains how the numbers on it are derived, so you can reproduce them and know what they do and don't mean.

## Families and variants

Many model IDs are the same model offered in a different way. The catalog groups them into one family page:

| Variant type | Example | Why it's separate |
| - | - | - |
| Privacy | `z-ai-glm-5-3` and `e2ee-glm-5-3-p` | The E2EE variant runs in an attested enclave and may have different limits |
| Speed | `claude-opus-5-5` and `claude-opus-5-5-fast` | Faster serving at a higher price |
| Mode | `kling-v3-pro-text-to-video`, `-image-to-video`, `-motion-control` | Different inputs and prices |
| Task | `nano-banana-pro` and `nano-banana-pro-edit` | Generation and editing endpoints |

Every variant keeps its own model ID, price and privacy tier. Pass the variant's ID as `model`. Dated snapshots such as `deepseek-v4-flash-0731` stay separate families and appear under **Other versions**.

## Normalized prices

Each modality bills in its own unit. The catalog shows the native unit and one reference price, so models can be sorted and compared.

| Modality | Native unit | Reference price |
| - | - | - |
| Text | Per 1M input and output tokens | **Blended**: `(3 × input + output) / 4`, the 3:1 mix Artificial Analysis uses. The cached-input discount is `1 − cached ÷ input`. |
| Image | Per image, by resolution tier | Price per image at the selected tier (1K, 2K or 4K). Edits add a per-image surcharge for extra inputs. |
| Video | Per clip, by resolution, duration and audio | Price per second and per clip at the selected settings, from a full quote matrix (below). |
| Text to speech | Per 1M characters | Per minute of speech at 150 words (about 825 characters) per minute. |
| Speech to text | Per audio second | Per audio hour and per 1,000 minutes. |
| Music and sound effects | Per track, per second or per duration bucket | Per minute of audio where the model allows it; otherwise per track. |
| Embeddings | Per 1M input tokens | Unchanged. |

<Note>
  Reasoning models can spend many more output tokens than non-reasoning models on the same task, so the blended price understates their real cost. Cost per task, measured on benchmark runs, is planned; see [Benchmarks](#benchmarks).
</Note>

### Video pricing

Video models don't publish prices in `GET /models`. At build time, Venice calls [`POST /video/quote`](/api-reference/endpoint/video/quote) for every resolution × duration × audio combination each model supports, at its first listed aspect ratio (aspect ratio does not change the price). The catalog stores the full matrix and derives:

* **Per second**: the median of `price ÷ seconds` across durations at a given resolution and audio setting. Prices are linear in duration for almost every model.
* **Draft clip**: 5 seconds, 720p, silent.
* **Production clip**: 10 seconds, 1080p, with audio.

When a model doesn't offer the requested setting, the catalog uses the closest one it supports and marks the price. Edit, motion-control and upscale modes bill by the length of your source video, so they show **By source** instead of a price. Always call `POST /video/quote` before queueing for an exact amount.

## Privacy tiers

| Tier | What it means |
| - | - |
| **E2EE** | Prompts are encrypted on your device and decrypted only inside an attested hardware enclave. See [TEE & E2EE models](/guides/features/tee-e2ee-models). |
| **TEE** | Inference runs in a hardware enclave with cryptographic attestation. |
| **Private** | Served on infrastructure Venice controls, with zero data retention. |
| **Anonymized** | Proxied to the upstream lab without your identity. The lab may retain prompts. |

A family lists every tier its variants offer. Filtering by a tier shows the variant in that tier, with its own price.

## Reference prompt suite

Image and video models are rendered on a fixed prompt suite, the same prompts for every model of a modality, at default settings. That's what makes the side-by-side view in [Compare](/models/compare) meaningful.

| Modality | Prompts |
| - | - |
| Image | In-image text, photorealism, instruction following, illustration style |
| Video | Cinematic landscape, seamless loop, urban cinematic |

Models without renders show **Reference renders pending**.

## Performance

<Info>
  Not published yet. This section describes what the catalog will show.
</Info>

Venice will measure models with synthetic probes, never customer prompts, and publish:

| Modality | Metrics |
| - | - |
| Text | Uptime, time to first token, output tokens per second, time to a 500-token response, cache hit rate, success rate |
| Image | Uptime, generation time for one 1K image (p50 and p95), success rate |
| Video | Uptime, queue time and generation time for a 5-second 720p clip (reported separately), success rate |
| Audio | Uptime, time to first audio, speed factor, success rate |
| Embeddings | Uptime, latency for a 1K-token input, success rate |

Every figure will show its window (rolling 30 days for uptime, 7 days for latency), percentile, sample count and probe region. Uptime excludes 4xx errors caused by the request. Cache hit rate is `cache_read ÷ (input + cache_read + cache_write)` across all traffic. Models with too few samples will be marked **Preliminary**.

## Benchmarks

<Info>
  Not published yet. This section describes what the catalog will show.
</Info>

Benchmarks are chosen for the frontier: still discriminative at the top, resistant to contamination, and runnable against the model as Venice serves it. Every score will carry one provenance label, never averaged together:

* **Venice-verified**: run on the Venice endpoint shown on the page.
* **Independent**: run by a third party, usually on the lab's own API.
* **Lab-reported**: published by the model's creator.

Each score also shows the benchmark and harness version, reasoning effort, test date and a confidence interval. Scores older than 90 days, or superseded by a new benchmark version, are flagged as stale.

| Modality | Planned headline set |
| - | - |
| Text | Artificial Analysis Intelligence Index, plus Humanity's Last Exam, Terminal-Bench, τ³-bench, AA-LCR, IFBench, AA-Omniscience and a Venice Openness Score |
| Image | Human-preference Elo from the Artificial Analysis and LMArena arenas, GenEval 2 and an over-refusal rate |
| Video | Artificial Analysis text-to-video (with and without audio) and image-to-video Elo, LMArena text-to-video |
| Text to speech | Artificial Analysis Speech Arena Elo, word error rate, time to first audio |
| Speech to text | Artificial Analysis word error rate, Open ASR Leaderboard, speed factor |
| Music | Artificial Analysis Music Arena (instrumental and vocals) |
| Embeddings | MTEB Multilingual, RTEB and a Venice held-out retrieval set |

The set is reviewed every quarter. A benchmark becomes a floor check, no longer a headline, once frontier models saturate it or it fails a contamination audit.

## Data sources

The catalog is built from public endpoints you can call yourself:

* [`GET /models`](/api-reference/endpoint/models/list) for specs, capabilities and prices.
* [`POST /video/quote`](/api-reference/endpoint/video/quote) for video prices.
* [`GET /models/traits`](/api-reference/endpoint/models/traits) for defaults such as `default_code`.

Prices are in USD. DIEM is billed at parity.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.