According to the latest data from OpenRouter, the total token usage of global AI large models from September 14 to 20 reached 129 trillion tokens, an increase of 1.57% compared to the previous week. Among them, the weekly token usage of large models in China reached 67.46 trillion tokens, an increase of 10.28% compared to the previous week; while that of large models in the United States was 14.21 trillion tokens, a decrease of 34.7% compared to the previous week. The weekly token usage of large models in China has exceeded that of the United States for 21 consecutive weeks.

In last week's global token usage ranking, four out of the top five were from China. DeepSeek V4.1-Flash ranked first with 15.8 trillion tokens, an increase of 219% compared to the previous week. The model was released on September 10, focusing on improving coding, agent, and multimodal understanding capabilities. It uses a Causal-Encoder-Decoder architecture, and the global KV Cache is reduced to about 1/4 of that in V4-Flash. The cost for input cache hit, miss, and output per million tokens during idle time are 0.02 yuan, 1 yuan, and 4 yuan respectively, which are 60%, 33.3%, and 11.1% lower than the previous generation.

However, a decrease in token price does not necessarily mean that the cost of individual tasks decreases simultaneously. According to Artificial Analysis testing, V4.1-Flash uses approximately 89,000 output tokens per task on average, higher than the 62,000 tokens of V4Flash0731, and it has been evaluated as relatively verbose in output.

Additionally, Zhipu GLM5.3Flash had a weekly token usage of 14.1 trillion tokens, ranking second; Tencent HuanYuan Hy4preview had a weekly token usage of 12.5 trillion tokens, ranking third; and DeepSeek-V4-Flash-0731 had a weekly token usage of 9.44 trillion tokens, ranking fifth.