qwen3.8-flash
Alibaba · Priced per token
Qwen3.8-Flash is the latest multimodal large language model released by Tongyi Qwen, featuring robust understanding and generation capabilities as well as exceptional response speed. Natively supporting a million-level context window, the model can process ultra-long documents, code repositories and complex conversations in a single pass. It delivers particularly outstanding performance in scenarios such as programming assistance, agent collaboration, and image-text understanding — whether it comes to auto-fixing code, operating desktop applications, or analyzing charts and long videos, it can produce accurate, high-quality outputs. Meanwhile, it is compatible with the mainstream interface protocols of OpenAI and Anthropic, enabling seamless integration with developer tools including Claude Code and Codex to facilitate easy building of high-concurrency applications and intelligent workflows. With its excellent performance and highly competitive inference cost, Qwen3.8-Flash stands as the ideal choice for developers and enterprises seeking to balance effectiveness and efficiency in their AI applications.
Pricing
Base price
- Input: $0.1193 per 1M tokens
- Output: $0.4026 per 1M tokens
- Cached input: $0.0149 per 1M tokens
- Cache write: $0.1864 per 1M tokens
List rate at a 1x group multiplier, before any account-specific discount. The effective rate is this figure times the multiplier of the group the account is in.
Pricing by group
| Group | Multiplier | Input | Output |
|---|---|---|---|
| default | 1x | $0.1193 | $0.4026 |
API access
- openai — POST /v1/chat/completions
Available to groups: default