> ## Documentation Index
> Fetch the complete documentation index at: https://veniceai-feat-models-redesign.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# GPT-6 Luna API

> GPT-6 Luna API on Venice: 1.05M context, $0.13 input and $0.63 output per 1M tokens. Anonymized access without your identity. Pricing, specs and code examples.

export const HubMount = ({view, children, ...props}) => {
  const [hub, setHub] = useState(null);
  const [failed, setFailed] = useState(false);
  useEffect(() => {
    let alive = true;
    const w = window;
    if (!w.__veniceModelHub) {
      const urls = w.location.hostname === 'localhost' ? ['http://localhost:3333/data/model-hub.bundle.json', '/data/model-hub.bundle.json'] : ['/data/model-hub.bundle.json'];
      const load = i => fetch(urls[i], {
        cache: 'no-cache'
      }).then(res => {
        if (!res.ok) throw new Error(`bundle ${res.status}`);
        return res.json();
      }).catch(err => i + 1 < urls.length ? load(i + 1) : Promise.reject(err));
      const Frag = <></>.type;
      const h = (type, props, ...kids) => {
        const T = type;
        const {key, ...rest} = props || ({});
        if (!kids.length) return <T key={key} {...rest} />;
        if (kids.length === 1) return <T key={key} {...rest}>{kids[0]}</T>;
        return <T key={key} {...rest}>{kids.map((kid, i) => <Frag key={i}>{kid}</Frag>)}</T>;
      };
      w.__veniceModelHub = load(0).then(bundle => new Function(`return (${bundle.code})`)()({
        h,
        Fragment: Frag,
        useState,
        useEffect,
        useRef,
        useMemo,
        useCallback
      }));
    }
    w.__veniceModelHub.then(instance => {
      if (alive) setHub(instance);
    }).catch(() => {
      w.__veniceModelHub = null;
      if (alive) setFailed(true);
    });
    return () => {
      alive = false;
    };
  }, []);
  const View = hub ? hub[view] : null;
  if (View) return <View {...props}>{children}</View>;
  if (failed) {
    return <div className="vx-mount is-failed">
        <p className="vx-mount-note">The interactive model catalog could not load. The full data is below.</p>
        {children}
      </div>;
  }
  return <div className="vx-mount" aria-busy="true">
      <div className="vx-mount-skeleton" aria-hidden="true"><span /><span /><span /></div>
      <div className="vx-mount-source">{children}</div>
    </div>;
};

<HubMount view="ModelPage" data={{"family":{"slug":"gpt-6-luna","name":"GPT-6 Luna","modality":"text","task":"chat","provider":"openai","description":"GPT-6 Luna is OpenAI's most efficient GPT-6 model for focused, high-volume tasks. It is suited for chat, classification, and lightweight agentic workflows, with a 1.05M token context window (922K input, 128K output), text and image inputs, and capable reasoning for its price tier.","primary":"openai-gpt-6-luna","variants":["openai-gpt-6-luna"],"created":1790035200,"updated":1790035200,"privacy":["anonymized"]},"models":[{"id":"openai-gpt-6-luna","name":"GPT-6 Luna","type":"text","modality":"text","task":"chat","variant":"standard","provider":"openai","created":1790035200,"description":"GPT-6 Luna is OpenAI's most efficient GPT-6 model for focused, high-volume tasks. It is suited for chat, classification, and lightweight agentic workflows, with a 1.05M token context window (922K input, 128K output), text and image inputs, and capable reasoning for its price tier.","source":"https://developers.openai.com/api/docs/models/gpt-6-luna","privacy":"anonymized","text":{"context":1050000,"maxOutput":128000,"reasoning":{"supported":true,"effort":["none","low","medium","high","xhigh","max"],"defaultEffort":"high"},"caps":{"tools":true,"structured":true,"vision":true,"maxImages":10,"webSearch":true}},"pricing":{"input":0.125,"output":0.625,"cacheRead":0.0125,"cacheWrite":0.15625,"blended":0.25,"extended":{"threshold":272000,"input":0.25,"output":0.9375,"cacheRead":0.025,"cacheWrite":0.3125}},"headline":{"value":0.25,"unit":"per 1M tokens","basis":"blended"},"endpoints":[{"id":"chat","method":"POST","path":"/chat/completions","name":"Chat Completions","status":"stable"},{"id":"responses","method":"POST","path":"/responses","name":"Responses","status":"alpha","recommended":true,"recommendation":"OpenAI reasoning models are designed around the Responses API: reasoning, tool calls and messages come back as typed output items. Venice's /responses endpoint is in alpha."}],"family":"gpt-6-luna"}],"related":{"similar":[{"slug":"glm-5-3-flash","name":"GLM 5.3 Flash","provider":"zai","modality":"text","privacy":["private","e2ee"],"created":1787270400,"headline":{"value":0.2375,"unit":"per 1M tokens","basis":"blended"},"variants":2},{"slug":"qwen-3-8-flash","name":"Qwen 3.8 Flash","provider":"alibaba","modality":"text","privacy":["anonymized"],"created":1788998400,"headline":{"value":0.2275,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"mimo-v2-6-flash","name":"MiMo-V2.6-Flash","provider":"xiaomi","modality":"text","privacy":["private"],"created":1790553600,"headline":{"value":0.21875,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"deepseek-v4-flash-0731","name":"DeepSeek V4 Flash 0731","provider":"deepseek","modality":"text","privacy":["private"],"created":1785456000,"headline":{"value":0.21875,"unit":"per 1M tokens","basis":"blended"},"variants":2}],"versions":[{"slug":"gpt-5-6-luna","name":"GPT-5.6 Luna","provider":"openai","modality":"text","privacy":["anonymized"],"created":1783555200,"headline":{"value":0.5625,"unit":"per 1M tokens","basis":"blended"},"variants":1}]},"providers":{"openai":{"slug":"openai","name":"OpenAI","logo":"/images/icons/models/openai.svg"},"zai":{"slug":"zai","name":"Z.ai","logo":"/images/icons/models/Zhipu.svg"},"alibaba":{"slug":"alibaba","name":"Alibaba Qwen","logo":"/images/icons/models/qwen.svg"},"xiaomi":{"slug":"xiaomi","name":"Xiaomi","logo":"/images/icons/models/text.svg"},"deepseek":{"slug":"deepseek","name":"DeepSeek","logo":"/images/icons/models/deepseek.svg"}},"faq":[{"q":"How much does the GPT-6 Luna API cost?","a":"$0.13 per 1M input tokens and $0.63 per 1M output tokens, with cached input at $0.013 per 1M. Prices are in USD and can be paid in DIEM at parity."},{"q":"What is the GPT-6 Luna model ID?","a":"Use `openai-gpt-6-luna` as the `model` parameter."},{"q":"Is the GPT-6 Luna API private?","a":"GPT-6 Luna is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work."},{"q":"What is the context window of GPT-6 Luna?","a":"1.05M tokens of context, with up to 125K output tokens per response."},{"q":"What does GPT-6 Luna support?","a":"GPT-6 Luna supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with `reasoning_effort`: none, low, medium, high, xhigh and max (default high)."},{"q":"Which endpoint does the GPT-6 Luna API use?","a":"Call `POST /responses`. `/chat/completions` is also supported. OpenAI reasoning models are designed around the Responses API: reasoning, tool calls and messages come back as typed output items. Venice's /responses endpoint is in alpha."}]}} />

<div className="vx-static">
  <Accordion title="Plain-text specification">
    # GPT-6 Luna API

    GPT-6 Luna is a large language model by OpenAI, available on the Venice API as `openai-gpt-6-luna`. Requests are anonymized, so the provider never sees your identity.

    GPT-6 Luna is OpenAI's most efficient GPT-6 model for focused, high-volume tasks. It is suited for chat, classification, and lightweight agentic workflows, with a 1.05M token context window (922K input, 128K output), text and image inputs, and capable reasoning for its price tier.

    ## GPT-6 Luna API pricing

    | Model ID | Variant | Privacy | Price |
    | - | - | - | - |
    | `openai-gpt-6-luna` | Standard | Anonymized | $0.13 input / $0.63 output per 1M tokens |

    ## GPT-6 Luna specifications

    | Spec | Value |
    | - | - |
    | Provider | OpenAI |
    | Released | Sep 22, 2026 |
    | Privacy | Anonymized |
    | Context window | 1.05M tokens |
    | Max output | 125K tokens |
    | Reasoning effort | none, low, medium, high, xhigh, max |

    ## How to use the GPT-6 Luna API

    Send requests to `POST https://api.venice.ai/api/v1/responses` with `"model": "openai-gpt-6-luna"` and your API key.

    ```bash theme={null}
    curl https://api.venice.ai/api/v1/responses \
      -H "Authorization: Bearer $VENICE_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "openai-gpt-6-luna",
        "input": "Explain TEE attestation in two sentences.",
        "reasoning": { "effort": "high" }
      }'
    ```

    ## GPT-6 Luna API FAQ

    ### How much does the GPT-6 Luna API cost?

    $0.13 per 1M input tokens and $0.63 per 1M output tokens, with cached input at \$0.013 per 1M. Prices are in USD and can be paid in DIEM at parity.

    ### What is the GPT-6 Luna model ID?

    Use `openai-gpt-6-luna` as the `model` parameter.

    ### Is the GPT-6 Luna API private?

    GPT-6 Luna is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.

    ### What is the context window of GPT-6 Luna?

    1.05M tokens of context, with up to 125K output tokens per response.

    ### What does GPT-6 Luna support?

    GPT-6 Luna supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with `reasoning_effort`: none, low, medium, high, xhigh and max (default high).

    ### Which endpoint does the GPT-6 Luna API use?

    Call `POST /responses`. `/chat/completions` is also supported. OpenAI reasoning models are designed around the Responses API: reasoning, tool calls and messages come back as typed output items. Venice's /responses endpoint is in alpha.

    ## Related models

    * [GPT-5.6 Luna API](/models/gpt-5-6-luna): $0.25 input / $1.50 output per 1M tokens
    * [GLM 5.3 Flash API](/models/glm-5-3-flash): $0.15 input / $0.50 output per 1M tokens
    * [Qwen 3.8 Flash API](/models/qwen-3-8-flash): $0.14 input / $0.49 output per 1M tokens
    * [MiMo-V2.6-Flash API](/models/mimo-v2-6-flash): $0.17 input / $0.35 output per 1M tokens
    * [DeepSeek V4 Flash 0731 API](/models/deepseek-v4-flash-0731): $0.17 input / $0.35 output per 1M tokens
  </Accordion>
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.