A year ago the standard story about Chinese AI was that it trailed the best US labs by a comfortable margin. That story is getting harder to tell. Bloomberg Intelligence says DeepSeek’s V4.1 Flash, released on September 10, has cut the gap between the top Chinese and US models on LiveBench to about 3%, the smallest spread the firm has tracked.
The short version
- DeepSeek V4.1 Flash scored 81.1 on LiveBench against 83.4 for Anthropic’s top model, a gap of roughly 3%
- That gap was about 9% in May and about 15% earlier in 2026
- On the agentic coding sub-benchmark DeepSeek scored 77.3 versus 66.1, ahead of the US leader
- The model has MIT-licensed weights and costs $0.30 in and $1.20 out per million tokens at peak hours
What the numbers say
LiveBench is a benchmark that refreshes its questions regularly to limit contamination from training data, which is why analysts like to use it for cross-lab comparisons. On the October 4 leaderboard snapshot cited in the coverage, DeepSeek V4.1 Flash sits 2.3 points behind Anthropic’s top entry overall. That is a small margin on a hundred-point scale.
The sub-benchmarks tell a more interesting story. In agentic coding, the kind of task where a model has to plan, write and fix code across several steps, DeepSeek scored 77.3 while Anthropic’s entry scored 66.1. A single leaderboard snapshot is not a verdict, and benchmarks shift from week to week. But it is one of the first times a Chinese open-weight model has led a US flagship on a category that matters for real developer work.
| Measure | DeepSeek V4.1 Flash | Top US model (Anthropic) |
|---|---|---|
| LiveBench overall | 81.1 | 83.4 |
| Agentic coding sub-score | 77.3 | 66.1 |
| Peak price per 1M tokens (in / out) | $0.30 / $1.20 | Not compared here |
| Off-peak price per 1M tokens (in / out) | $0.15 / $0.60 | Not compared here |
| License | MIT weights | Closed |
Cheap, open and aimed at agents
V4.1 Flash is the lighter member of DeepSeek’s V4 family. The family launched in April, and coverage at the time described the Flash model as a mixture-of-experts design with 284 billion total parameters and around 13 billion active per token, which is how it keeps serving costs down. The September update added native image input and kept the MIT license, so anyone can download the weights and run them.
The pricing is the part that makes rivals nervous. At peak hours the API costs $0.30 per million input tokens and $1.20 per million output tokens, and off-peak it halves to $0.15 and $0.60. We tracked this pricing pressure when DeepSeek cut its prices to almost nothing and was still valued at $74 billion. A model that is close to the frontier on quality and a fraction of the cost per task is exactly the kind that agent builders switch to, since agents burn through tokens.
Why analysts are pointing at hardware
Bloomberg Intelligence analyst Lea attributes the catch-up to Chinese labs improving technically and tuning their models for domestic hardware. DeepSeek’s V4 family has been supported on Huawei’s Ascend platform since launch, which means customers who cannot buy export-controlled Nvidia chips can still run frontier-class inference. That includes Chinese enterprises and, according to coverage of the model, buyers in parts of the Gulf and Southeast Asia.
That is the uncomfortable part for Washington. Export controls were designed to slow Chinese AI by limiting access to advanced chips. If a lab can close the benchmark gap while moving its workloads onto Huawei silicon, the controls look less like a wall and more like a speed bump. DeepSeek has said Pro-tier pricing could fall further once Huawei’s Ascend 950 supernodes are deployed at scale in the second half.
What to keep in perspective
- The gains sit at the frontier. Only three of the top 15 LiveBench entries are Chinese, so the field overall is not caught up
- One benchmark is not the whole story. LiveBench is useful, but it does not capture reliability, safety testing or enterprise support
- Money is still a problem. Analysts expect Chinese AI companies to struggle with profitability until around 2030
- Some US labs say the gains are borrowed. That claim is at the center of a separate fight, covered below
The distillation argument
US officials have a competing explanation for how fast Chinese models are improving. Earlier in September the NSA, CISA and FBI named six Chinese AI companies and said they had been extracting capability from American models since 2024. We broke that down in our piece on the advisory accusing Chinese labs of draining Claude, GPT and Gemini. Nothing in the LiveBench numbers proves or disproves that claim. They only show the end result: a smaller gap, whatever the cause.
What it means for you
- If you build with AI, V4.1 Flash is worth testing on agent and coding workloads, where the cost savings can be large
- If you are weighing risk, open weights let you self-host, which avoids sending data to a Chinese API but still requires your own security review
- If you follow policy, expect the benchmark gap to become a central argument in the next round of chip export debates
- If you invest, falling prices at the frontier squeeze margins for everyone selling tokens, American or Chinese
A 3% gap does not mean the race is over, and it does not mean it is tied. It does mean that the old assumption, that US labs hold a safe lead and can price accordingly, no longer holds without checking the leaderboard first.
Sources and further reading
- AI Weekly: DeepSeek V4.1 Flash narrows US-China AI gap to 3% on LiveBench
- Implicator: DeepSeek cuts US AI benchmark lead to about 3%
- Startup Fortune: DeepSeek narrows AI gap with US to just 3 percent, Bloomberg says
- DataCamp: DeepSeek V4.1 Flash features, benchmarks, pricing
About this article: GeekBlog covers U.S. technology news, AI, phones, smartwatches and gaming. Every story is written and checked under our Editorial Policy. Spotted a mistake or have a story tip? Contact our editors.

