GREG BAKER
GREG BAKER/Getty Images

Anthropic's newest distillation accusation, filed with the US Senate in June, named Alibaba's Qwen lab for what it called the largest such campaign in the company's history — 28.8 million exchanges pulled through roughly 25,000 fake accounts in six weeks, following a February report naming DeepSeek, Moonshot AI and MiniMax for a smaller version of the same behavior. Alibaba's Hong Kong-listed shares fell as much as 5% intraday on the news, to a four-month low of HK$94.55. None of the four companies has responded to comment requests.

That's the trigger. It's also close to beside the point.

The interesting question isn't whether Alibaba did what Anthropic says it did — that's a legal dispute, contested by at least one Beijing IP lawyer as resting on nothing more than scraping public API outputs, and it will get settled in Congress or in court, not in this column. The interesting question is why an entire tier of Chinese AI companies keeps reaching for the same shortcut, openly enough that an industry research note puts the share of independent Chinese model teams relying on distillation above 60% as of 2024. That number is common knowledge inside the industry, not a secret Anthropic uncovered. So what does it actually reveal about where China's AI sector is weak?

It's not originality. The easy story — China copies, America invents — doesn't survive contact with the technical record. DeepSeek's multi-head latent attention mechanism and its mixture-of-experts architecture are independently developed, reviewed in international journals, and acknowledged as genuine contributions even by people skeptical of the company's broader claims. If the weakness were conceptual, distillation wouldn't help: you can't copy your way to an idea you don't understand well enough to reproduce. Teams that distill still have to know what they're extracting and how to integrate it. The bottleneck sits somewhere else.

It's frontier-scale compute, and it's structural, not just a chip embargo. Training a model at the size where genuine new capability shows up costs an amount of GPU-hours that only a handful of Chinese firms can fund and only a fraction of those can source hardware for at anything close to Nvidia's top tier. China's exposure to Washington's chip restrictions — and its dependence on the deliberately downgraded H20 as the workaround — has already been tested once this year, when a brief supply disruption exposed how thin that substitution really is. But the deeper issue isn't the chip ban itself; it's that pretraining a frontier model from scratch requires a scale of sustained, coordinated compute investment that most companies, in any country, simply can't justify against the odds of failure. Distillation is what a company does when it wants frontier-level output without underwriting frontier-level risk. It's a rational response to a compute ceiling, not evidence of an idea deficit.

That's also why the industry is consolidating, and the consolidation is the real diagnostic. The same research note that put distillation reliance above 60% also tracked the number of independent large-model companies in China falling from a peak of 237 to 112, with a further drop toward under 50 projected. That's not a story about weak companies getting caught stealing. It's a story about a two-tier structure: a small group of firms — DeepSeek, Alibaba, a few others — with the balance sheet and research depth to run genuine frontier pretraining, and a much larger group that never had a realistic path to that tier and used distillation as the only way to stay competitive on benchmarks in the meantime. Anthropic's crackdown, and the export-API restrictions likely to follow it, don't hit the first group especially hard. They remove the only lifeline the second group had.

What doesn't transfer through distillation is the part that matters most going forward. A distilled model inherits its teacher's surface behavior — fluent reasoning traces, coding style, tool-use patterns — without inheriting the safety alignment work, the evaluation methodology, or the judgment about failure modes that the teacher model's own lab spent years building. That's a real, and underdiscussed, asymmetry: benchmark scores can converge even while the underlying research maturity gap stays wide, because benchmarks measure output quality, not the depth of the process that produced the ability to self-correct, refuse unsafe requests, or generalize outside the training distribution. If Chinese labs lose API-level access to frontier Western models — which is exactly the direction both Anthropic's own product restrictions and the proposed Hagerty-Kim sanctions point toward — the firms that never built that underlying research capacity independently will be the first to plateau, not the ones like DeepSeek that already have.

So the honest answer to "what's China AI's real vulnerability" isn't algorithms, and it isn't even chips in the narrow sense. It's that the industry's middle tier built its competitiveness on borrowed capability rather than owned research infrastructure, and every policy lever now being pulled in Washington — export controls, API access restrictions, and now distillation enforcement — is aimed precisely at that borrowed layer. The firms with real pretraini