China's Best AI Models, One API Key.

Instant access to DeepSeek, Qwen, Minimax, Kimi and more — no Chinese phone number required. OpenAI-compatible, global low-latency, pay-as-you-go.

Get API Key Docs

Built for developers, designed for scale

  • Lightning Fast

    Optimized network architecture ensures millisecond response times

  • Secure & Reliable

    Enterprise-grade security with comprehensive permission management

  • Developer Friendly

    Compatible API routes for common AI application workflows

Change two lines, keep your code.

Point any OpenAI-compatible client at the gateway and go.

  1. Create an account

    Sign up in under a minute

  2. Create API Key

    Create a key for your app or service

  3. Add credits

    Keep enough balance before production traffic

  4. Send a request

    Verify routing with Playground or your client

One API gateway for your AI applications

TokenPAPA brings multiple AI providers behind a shared gateway, so an application can use one account, one API base URL and a gateway API key to reach the models enabled for that account. For developers exploring DeepSeek, Qwen, MiniMax, Kimi and other model families, this provides a common integration point while leaving model selection in the application's hands.

An OpenAI-compatible interface helps you reuse clients and tools that let you configure their base URL, API key and model name. You can compare models against the same prompts without building a separate account and authentication flow for every upstream provider. Your gateway key authenticates requests to this service; the gateway handles the configured upstream connection. Available models still depend on the account, access group and current catalogue.

A shared interface does not make every model interchangeable. Output quality, supported parameters and available endpoints can differ. Keep your model ID configurable, test representative requests and check the model's details before moving an existing workflow. The gateway's role is to simplify access and usage management while allowing you to choose the model that fits your application.

Choose capabilities for your workload

For conversational applications, start with a model that supports chat completions and evaluate it on your own tasks: answering customer questions, drafting content, summarising documents or assisting with code. Compare the quality of the output as well as the amount of input and output needed to get a useful answer. Streaming, tool use and structured output require support from the selected model and endpoint.

For search and retrieval workflows, use an embedding model when one is available in your catalogue. Image, audio and other multimodal workloads also need a matching model and request format. A chat model's presence in the catalogue does not imply that it accepts images, generates audio or supports every interface exposed by the gateway.

The gateway supports several API formats, including chat completions, Responses and provider-specific interfaces. Check the endpoint listed for a model rather than assuming the same request body works everywhere. Start with a small example, confirm the returned content and usage, then add the parameters your application needs. This makes compatibility issues easier to identify before sending production traffic.

Connect an OpenAI-compatible client

Create an account, generate a gateway API key and make sure the account has enough credit for the requests you intend to send. In a client that accepts a custom OpenAI base URL, use the gateway origin followed by /v1, then supply your gateway key and the exact model ID from the available catalogue. Some applications ask for the origin and add /v1 themselves; follow that client's configuration instructions to avoid duplicating the path.

The following cURL request illustrates the chat completions interface. Replace YOUR_API_KEY with a key created on this gateway and YOUR_MODEL_ID with an accessible model that supports chat completions. The placeholders are not working credentials or a model alias. A Responses or provider-specific client may need a different endpoint and request body.

curl 'https://tokenpapa.ai/v1/chat/completions' \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -d '{
    "model": "YOUR_MODEL_ID",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Keep API keys on your server or in your client's protected configuration, rather than in public browser code. After a test request, review its result and usage before increasing traffic. When changing models, recheck the endpoint, required parameters and response handling as well as the model ID. A successful request is a better integration check than simply saving the connection settings.

Read the documentation for account setup and API usage

Understand model pricing before you send traffic

Token-based models can have separate charges for input tokens, output tokens and cached input. Compare rates using the same unit: the public catalogue expresses these token prices per one million tokens. Your total depends on the usage of each request, so a low input rate alone does not determine the cost of a workflow that generates long answers.

Some models are billed per request instead of per token. Others use dynamic pricing that depends on the request or the applicable pricing tier. Those models cannot be described accurately by a single token price. Review their pricing details, including any relevant conditions, before estimating a budget. An unavailable price is not a promise of free usage.

The catalogue uses this gateway's configured display currency and public access-group rates. Your applicable group and the model's billing rules determine the rate for an actual request. Model availability and rates can change, so check the current pricing page when selecting a model and review your account's usage after testing. Use a representative workload when comparing expected costs.

Available models and current pricing

Browse the models currently published for anonymous visitors, grouped by provider. Each model name links to its details. The list uses the same catalogue and price format as the public pricing page; a model that requires a private access group is not included here.

Token rates are per 1M tokens at the lowest multiplier of any group an anonymous account can select on this gateway. Your actual rate follows the group selected for the request and the model's billing rules; the per-group breakdown is on each model's page.

Alibaba

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
qwen3.7-max$0.62$1.66$0.25
qwen3.7-plus$0.2$0.83$0.04
qwen3.8-flash$0.12$0.5$0.015
qwen3.8-max$1.8$5.4$0.23

Priced per request

ModelPer request
qwen-image-2.0$0.05
qwen-image-2.0-pro$0.1
qwen-image-3.0$0.03

Anthropic

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
claude-fable-5$9.5625$47.8125$0.9563
claude-opus-4-6$5.34$26.7$0.534
claude-opus-4-7$4.78$23.9$0.478
claude-opus-4-8$4.78$23.9$0.478
claude-sonnet-4-6$3.2$16$0.32

DeepSeek

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
deepseek-flash$0.3$1.2$0.01
deepseek-v4-flash$0.3$1.2$0.01
deepseek-v4-flash-vision-exp$0.3$1.2$0.01
deepseek-v4-pro$1.4$4$0.05

Doubao

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
doubao-seed-2-0-code-preview-260215$0.5$2.5-
doubao-seed-2-0-lite-260428$0.0938$0.5625$0.017
doubao-seed-2-0-mini-260428$0.031$0.312-
doubao-seed-2-0-pro-260215$0.5$2.5-
doubao-seed-2-1-pro-260628$0.9375$4.6875-
doubao-seed-2-1-turbo-260628$0.468$2.34-
doubao-seed-character-260628$0.12$0.3$0.03
doubao-seed-evolving$0.937$4.687-
doubao-seedance-2-0-260128$4.37$7.18-
doubao-seedance-2-0-fast-260128$3.43$5.78-
doubao-seedance-2-0-mini-260615$2.18$3.59-
doubao-seedance-2-5-260628$11.49$11.49-

Priced per request

ModelPer request
doubao-seedream-5-0-260128$0.034
doubao-seedream-5-0-pro-260628$0.046

Google

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
gemini-3-flash-preview$0.4$2.42$0.04
gemini-3-pro-image-preview$0.4$2.42-
gemini-3.1-flash-image-preview$0.4$2.42-
gemini-3.1-pro-preview$1.6154$9.6923-
gemini-3.5-flash$1.2115$1.2115$0.1212

Minimax

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
minimax-m2.5$0.2$0.8$0.02
minimax-m2.7$0.328$1.3$0.06
minimax-m3$0.328$1.3$0.06

Moonshot

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
kimi-k2.6$1.015$4.218$0.172
kimi-k2.7-code$1$4.21$0.2
kimi-k2.7-code-highspeed$2$8.44$0.4
kimi-k3$3.125$15.62$0.312

OpenAI

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
gpt-5.3-codex$1.88$15.04$0.2
gpt-5.4$2.32$13.92$0.232
gpt-5.4-mini$0.697$4.182$0.074
gpt-5.5$4.64$27.84$0.46
gpt-5.6-luna$0.93$7.44$0.093
gpt-5.6-sol$4.64$37.12$0.46
gpt-5.6-terra$2.28$18.24$0.22
gpt-image-2$5.3846$32.3-

Tencent

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
hy-mt2-pro$75$300-
hy3$0.14$0.58$0.035
hy3-preview$0.17$0.571$0.057
hy4-preview$0.8946$2.6837$0.4478

Priced per request

ModelPer request
hy-image-lite$0.002
hy-image-v3.0$0.1

Zhipu AI

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
glm-5.1$0.857$3.428$0.185
glm-5.2$1.23$4.3$0.3
glm-5.3$1.2$4.18$0.3
glm-5.3-flash$0.12$0.42$0.04

xiaomi

ModelInput / 1M tokensOutput / 1M tokensCached input / 1M tokens
mimo-v2.5$0.156$0.3125$0.003
mimo-v2.5-asr$75$75-
mimo-v2.5-pro$0.468$0.937$0.0039
mimo-v2.5-pro-ultraspeed$1.4$2.8$0.012
mimo-v2.5-tts$0$0-
mimo-v2.5-tts-voiceclone$0$0-
mimo-v2.5-tts-voicedesign$0$0-

Explore the model catalogue and pricing

Frequently asked questions

Can I use one gateway key with different providers?

A gateway key can call the models permitted by its configuration and your account's access group. Select the desired model ID in each request. That gives your application a common authentication entry point, while model permissions, endpoint support and any key restrictions still apply.

Do I need a Chinese phone number for the upstream providers?

The gateway's account and API key are the credentials you use here. You do not need to complete a separate Chinese-phone registration with each upstream provider to call models through the gateway. Follow this site's own registration requirements and check which models are available to your account.

Will my existing OpenAI client work?

Clients that allow a custom base URL, API key and model ID can use compatible gateway endpoints. Compatibility depends on the operations and parameters your application uses. Test a small request first, then verify features such as streaming or tools against the selected model's supported API.

How should I compare models?

Use representative prompts and judge the results against the needs of your application. Compare supported capabilities, response behaviour and total usage cost alongside the listed rates. A model that works well for short conversations may not be the right choice for document processing or a workflow that requires a particular output format.

Why might my rate differ from a catalogue price?

The public listing reflects the gateway's public group-rate selection and display currency. Your request may use another permitted group, and a dynamically priced model may apply a tier based on its billing rules. Use the model details and your account's usage records to understand the rate that applies to your workload.

What should I check if a request fails?

Check the base URL, endpoint, exact model ID, API key permissions and available account credit. Read the returned error before retrying. If a basic request works, add optional parameters gradually to identify unsupported combinations. Avoid repeatedly sending the same failing request without first checking its error and the model's API requirements.

Ready to simplify your AI integration?

Deploy your own gateway and start routing requests through your configured upstream services.

Get Started View Pricing