> ## Documentation Index
> Fetch the complete documentation index at: https://veniceai-feat-models-redesign.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Qwen 3.8 Flash API

> Qwen 3.8 Flash API on Venice: 1M context, $0.14 input and $0.49 output per 1M tokens. Anonymized access without your identity. Pricing, specs and code examples.

export const HubMount = ({view, children, ...props}) => {
  const [hub, setHub] = useState(null);
  const [failed, setFailed] = useState(false);
  useEffect(() => {
    let alive = true;
    const w = window;
    if (!w.__veniceModelHub) {
      const urls = w.location.hostname === 'localhost' ? ['http://localhost:3333/data/model-hub.bundle.json', '/data/model-hub.bundle.json'] : ['/data/model-hub.bundle.json'];
      const load = i => fetch(urls[i], {
        cache: 'no-cache'
      }).then(res => {
        if (!res.ok) throw new Error(`bundle ${res.status}`);
        return res.json();
      }).catch(err => i + 1 < urls.length ? load(i + 1) : Promise.reject(err));
      const Frag = <></>.type;
      const h = (type, props, ...kids) => {
        const T = type;
        const {key, ...rest} = props || ({});
        if (!kids.length) return <T key={key} {...rest} />;
        if (kids.length === 1) return <T key={key} {...rest}>{kids[0]}</T>;
        return <T key={key} {...rest}>{kids.map((kid, i) => <Frag key={i}>{kid}</Frag>)}</T>;
      };
      w.__veniceModelHub = load(0).then(bundle => new Function(`return (${bundle.code})`)()({
        h,
        Fragment: Frag,
        useState,
        useEffect,
        useRef,
        useMemo,
        useCallback
      }));
    }
    w.__veniceModelHub.then(instance => {
      if (alive) setHub(instance);
    }).catch(() => {
      w.__veniceModelHub = null;
      if (alive) setFailed(true);
    });
    return () => {
      alive = false;
    };
  }, []);
  const View = hub ? hub[view] : null;
  if (View) return <View {...props}>{children}</View>;
  if (failed) {
    return <div className="vx-mount is-failed">
        <p className="vx-mount-note">The interactive model catalog could not load. The full data is below.</p>
        {children}
      </div>;
  }
  return <div className="vx-mount" aria-busy="true">
      <div className="vx-mount-skeleton" aria-hidden="true"><span /><span /><span /></div>
      <div className="vx-mount-source">{children}</div>
    </div>;
};

<HubMount view="ModelPage" data={{"family":{"slug":"qwen-3-8-flash","name":"Qwen 3.8 Flash","modality":"text","task":"chat","provider":"alibaba","description":"Qwen 3.8 Flash is the latest multimodal model in the Qwen family, pairing strong reasoning and generation with remarkable speed. It natively supports a 1M-token context window, so it can process long documents, entire codebases, and complex conversations in a single pass. It excels at coding assistance, agentic workflows, and visual understanding, accepting text, image, and video input, and its thinking mode can be turned on or off per request.","primary":"qwen-3-8-flash","variants":["qwen-3-8-flash"],"created":1788998400,"updated":1788998400,"privacy":["anonymized"]},"models":[{"id":"qwen-3-8-flash","name":"Qwen 3.8 Flash","type":"text","modality":"text","task":"chat","variant":"standard","provider":"alibaba","created":1788998400,"description":"Qwen 3.8 Flash is the latest multimodal model in the Qwen family, pairing strong reasoning and generation with remarkable speed. It natively supports a 1M-token context window, so it can process long documents, entire codebases, and complex conversations in a single pass. It excels at coding assistance, agentic workflows, and visual understanding, accepting text, image, and video input, and its thinking mode can be turned on or off per request.","source":"https://www.alibabacloud.com/help/en/model-studio/models","privacy":"anonymized","text":{"context":1000000,"maxOutput":131072,"reasoning":{"supported":true},"caps":{"tools":true,"structured":true,"vision":true,"maxImages":10,"videoInput":true,"webSearch":true,"code":true}},"pricing":{"input":0.14,"output":0.49,"cacheRead":0.014,"blended":0.2275},"headline":{"value":0.2275,"unit":"per 1M tokens","basis":"blended"},"endpoints":[{"id":"chat","method":"POST","path":"/chat/completions","name":"Chat Completions","status":"stable","recommended":true},{"id":"responses","method":"POST","path":"/responses","name":"Responses","status":"alpha"}],"family":"qwen-3-8-flash"}],"related":{"similar":[{"slug":"mimo-v2-6-flash","name":"MiMo-V2.6-Flash","provider":"xiaomi","modality":"text","privacy":["private"],"created":1790553600,"headline":{"value":0.21875,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"glm-5-3-flash","name":"GLM 5.3 Flash","provider":"zai","modality":"text","privacy":["private","e2ee"],"created":1787270400,"headline":{"value":0.2375,"unit":"per 1M tokens","basis":"blended"},"variants":2},{"slug":"deepseek-v4-flash-0731","name":"DeepSeek V4 Flash 0731","provider":"deepseek","modality":"text","privacy":["private"],"created":1785456000,"headline":{"value":0.21875,"unit":"per 1M tokens","basis":"blended"},"variants":2},{"slug":"gpt-6-luna","name":"GPT-6 Luna","provider":"openai","modality":"text","privacy":["anonymized"],"created":1790035200,"headline":{"value":0.25,"unit":"per 1M tokens","basis":"blended"},"variants":1}],"versions":[]},"providers":{"alibaba":{"slug":"alibaba","name":"Alibaba Qwen","logo":"/images/icons/models/qwen.svg"},"xiaomi":{"slug":"xiaomi","name":"Xiaomi","logo":"/images/icons/models/text.svg"},"zai":{"slug":"zai","name":"Z.ai","logo":"/images/icons/models/Zhipu.svg"},"deepseek":{"slug":"deepseek","name":"DeepSeek","logo":"/images/icons/models/deepseek.svg"},"openai":{"slug":"openai","name":"OpenAI","logo":"/images/icons/models/openai.svg"}},"faq":[{"q":"How much does the Qwen 3.8 Flash API cost?","a":"$0.14 per 1M input tokens and $0.49 per 1M output tokens, with cached input at $0.014 per 1M. Prices are in USD and can be paid in DIEM at parity."},{"q":"What is the Qwen 3.8 Flash model ID?","a":"Use `qwen-3-8-flash` as the `model` parameter."},{"q":"Is the Qwen 3.8 Flash API private?","a":"Qwen 3.8 Flash is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work."},{"q":"What is the context window of Qwen 3.8 Flash?","a":"1M tokens of context, with up to 128K output tokens per response."},{"q":"What does Qwen 3.8 Flash support?","a":"Qwen 3.8 Flash supports function calling, structured outputs, reasoning, image input, web search and prompt caching."},{"q":"Which endpoint does the Qwen 3.8 Flash API use?","a":"Call `POST /chat/completions`. `/responses` (Alpha) is also supported."}]}} />

<div className="vx-static">
  <Accordion title="Plain-text specification">
    # Qwen 3.8 Flash API

    Qwen 3.8 Flash is a large language model by Alibaba Qwen, available on the Venice API as `qwen-3-8-flash`. Requests are anonymized, so the provider never sees your identity.

    Qwen 3.8 Flash is the latest multimodal model in the Qwen family, pairing strong reasoning and generation with remarkable speed. It natively supports a 1M-token context window, so it can process long documents, entire codebases, and complex conversations in a single pass. It excels at coding assistance, agentic workflows, and visual understanding, accepting text, image, and video input, and its thinking mode can be turned on or off per request.

    ## Qwen 3.8 Flash API pricing

    | Model ID | Variant | Privacy | Price |
    | - | - | - | - |
    | `qwen-3-8-flash` | Standard | Anonymized | $0.14 input / $0.49 output per 1M tokens |

    ## Qwen 3.8 Flash specifications

    | Spec | Value |
    | - | - |
    | Provider | Alibaba Qwen |
    | Released | Sep 10, 2026 |
    | Privacy | Anonymized |
    | Context window | 1M tokens |
    | Max output | 128K tokens |
    | Reasoning effort | Not adjustable |

    ## How to use the Qwen 3.8 Flash API

    Send requests to `POST https://api.venice.ai/api/v1/chat/completions` with `"model": "qwen-3-8-flash"` and your API key.

    ```bash theme={null}
    curl https://api.venice.ai/api/v1/chat/completions \
      -H "Authorization: Bearer $VENICE_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "qwen-3-8-flash",
        "messages": [{ "role": "user", "content": "Explain TEE attestation in two sentences." }]
      }'
    ```

    ## Qwen 3.8 Flash API FAQ

    ### How much does the Qwen 3.8 Flash API cost?

    $0.14 per 1M input tokens and $0.49 per 1M output tokens, with cached input at \$0.014 per 1M. Prices are in USD and can be paid in DIEM at parity.

    ### What is the Qwen 3.8 Flash model ID?

    Use `qwen-3-8-flash` as the `model` parameter.

    ### Is the Qwen 3.8 Flash API private?

    Qwen 3.8 Flash is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.

    ### What is the context window of Qwen 3.8 Flash?

    1M tokens of context, with up to 128K output tokens per response.

    ### What does Qwen 3.8 Flash support?

    Qwen 3.8 Flash supports function calling, structured outputs, reasoning, image input, web search and prompt caching.

    ### Which endpoint does the Qwen 3.8 Flash API use?

    Call `POST /chat/completions`. `/responses` (Alpha) is also supported.

    ## Related models

    * [MiMo-V2.6-Flash API](/models/mimo-v2-6-flash): $0.17 input / $0.35 output per 1M tokens
    * [GLM 5.3 Flash API](/models/glm-5-3-flash): $0.15 input / $0.50 output per 1M tokens
    * [DeepSeek V4 Flash 0731 API](/models/deepseek-v4-flash-0731): $0.17 input / $0.35 output per 1M tokens
    * [GPT-6 Luna API](/models/gpt-6-luna): $0.13 input / $0.63 output per 1M tokens
  </Accordion>
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.