Close-up of a laptop displaying DeepSeek interface with a chatbot prompt in dark mode.

Alibaba and DeepSeek Are Slashing AI Costs in China’s Model Race

The AI model race in China is heating up. And the big story right now is cost.

Alibaba just dropped their largest AI model yet. DeepSeek is undercutting everyone on price. And both are pushing the industry toward cheaper, more accessible AI.

Let me break down what’s happening.


Alibaba’s Big Move: Qwen3.8-Max

Alibaba launched Qwen3.8-Max, their biggest AI model to date. It has 2.4 trillion parameters and uses a mixture-of-experts architecture.

Here’s what that means. Instead of activating the entire model for every request, only about 95 billion parameters are active at a time. That reduces costs and response delays.

The model can process text, images, and video. It supports up to one million tokens of context. Alibaba also said the model completed a software engineering project over 16 days.

Pricing: $2 per million input tokens and $6 per million output tokens.

That’s cheaper than Moonshot AI’s Kimi K3, which costs $3 for input and $15 for output. Alibaba is directly competing on price.


DeepSeek Is Even Cheaper

DeepSeek took a different approach with V4-Flash. Instead of matching the scale of Alibaba’s model, they focused on driving prices down.

V4-Flash pricing: $0.14 per million input tokens and $0.28 per million output tokens.

That’s dramatically cheaper than almost everything else on the market. The model has 284 billion total parameters, with 13 billion active during inference.

Here’s where it gets really interesting. DeepSeek’s cache-hit pricing is $0.003 per million tokens for the Max Effort version. That’s 98% below their standard input rate. Cached input covers previously processed context that can be reused.

When you look at cost per task, the difference is even more dramatic. Artificial Analysis estimated V4-Flash’s average cost at three cents per test, compared with:

  • 86 cents for Kimi K3
  • $1.86 for OpenAI’s GPT-5.6 Sol
  • $3.15 for Anthropic’s Claude Fable 5

That’s a massive difference.


Token Prices Don’t Tell the Whole Story

Here’s something important to understand.

A lower per-token rate doesn’t always mean a lower total cost. If a model generates more output or requires multiple interactions, the total cost can add up quickly.

Example: Kimi K3 costs $3 per million input tokens and $15 per million output tokens. On Artificial Analysis’ AA-Briefcase benchmark, it averaged $10.57 per task. It generated around 120,000 output tokens and used an average of 83 turns per task.

So the advertised API price isn’t always what you actually pay for complex workloads.

This is why cost-per-task measurements add useful context to standard API pricing. Models with different architectures and usage patterns can consume substantially different amounts of compute and tokens while working through the same type of task.


Model Rankings

On performance, Alibaba’s model moved to the top position among Chinese text models on the crowdsourced comparison platform Arena.AI. But it still remained behind several Anthropic models in the overall rankings.

It also ranked second on Arena.AI’s leaderboard for models that analyze images and other visual material, behind an Anthropic Claude Fable 5 variant.

Kimi K3 recorded the second-highest overall score on the AA-Briefcase evaluation, behind Claude Fable 5. It also scored 57 on Artificial Analysis’ broader Intelligence Index.

DeepSeek V4-Flash scored 40 on the Intelligence Index. So it’s not the most capable model, but it’s incredibly cheap.


The Open-Weight Factor

This is where China’s approach differs from the US.

Alibaba, DeepSeek, and Moonshot AI have continued to support open-weight releases alongside hosted API access. That gives developers more options for how the models are deployed.

  • DeepSeek V4-Flash is available as an open-weight model under an MIT license
  • Kimi K3 is also available as an open-weight model under Moonshot AI’s own license
  • Weights are available through Hugging Face

Open weights let developers run models on their own infrastructure or through third-party providers instead of relying solely on a hosted API. Deployment costs still depend on the hardware and infrastructure used, but access to the model isn’t tied to a single provider.

This is very different from OpenAI, Anthropic, and Google, which generally keep their model weights closed.


What This Means for the Industry

Lian Jye Su, chief analyst at Omdia, put it well:

“Many business workflows do not need the industry’s very best model. They need models that are good enough, affordable, transparent and accessible, and open-weight models help meet that demand.”

That’s the key insight here.

Not everyone needs the most powerful model. Many businesses just need something that works well enough, costs less, and gives them flexibility. China’s AI companies are positioning themselves to meet that demand.

The race is shifting. It’s not just about who has the most capable model anymore. It’s about who can deliver good performance at the lowest cost with the most flexibility.


The Bottom Line

Alibaba and DeepSeek are pushing down AI costs in China. Qwen3.8-Max is Alibaba’s largest model yet, priced below competitors. DeepSeek V4-Flash is dramatically cheaper, costing pennies per task compared to dollars for Western models.

Both are available as open-weight models, giving developers deployment flexibility. And analysts say many businesses don’t need the best model, they need good enough at a reasonable price.

China’s AI companies are betting that cost and accessibility will win the race.

Similar Posts

One Comment

Leave a Reply

Your email address will not be published. Required fields are marked *