Michael Kratsios, the White House science advisor, has publicly accused Moonshot, a prominent Chinese artificial intelligence firm, of illicitly developing its Kimi K3 large language model (LLM) by directly copying Anthropic’s proprietary Fable LLM and simultaneously utilizing advanced semiconductor chips that are explicitly banned for export to China. This serious allegation has ignited a fierce debate within the global AI community regarding intellectual property protection, the efficacy of export controls, and the escalating technological rivalry between the United States and China.
Kratsios, in a recent social media post, condemned the alleged actions, stating, "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable." His remarks come amid intensifying discussions in the AI sector about potential U.S. government measures, including outright bans on certain Chinese open-weight models, which have sent ripples of concern throughout the industry. Moonshot, for its part, has remained silent on the accusations regarding its training methodologies, and Kratsios has not yet disclosed further specifics or the sources underpinning his claims.
The Heart of the Allegation: Distillation and Watermarks
The accusation by Kratsios mirrors earlier sentiments expressed by Treasury Secretary Scott Bessent, who noted, "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that’s unacceptable." The precise nature of these "watermarks" – digital signatures or unique patterns embedded within the model’s output or structure that could indicate its origin – remains undefined, and the Treasury Department has not elaborated on the specifics. This ambiguity contributes to the complexity of proving intellectual property theft in the rapidly evolving field of generative AI.
At the core of the alleged copying mechanism is "distillation," a process in which one LLM (the student model) learns from another more powerful or sophisticated LLM (the teacher model). This typically involves systematically querying the teacher model to generate a vast dataset of prompts and responses, which is then used to train the student model, often through a technique known as supervised fine-tuning (SFT). In some cases, distillation might involve asking the teacher model to articulate its "chain-of-thought" to better understand its problem-solving approaches. While distillation can be a legitimate research tool for knowledge transfer, its use to replicate proprietary capabilities without authorization crosses into contentious territory, particularly when the intent is to bypass significant research and development investments.
Expert Skepticism Regarding Distillation’s Efficacy for Frontier Models
Despite the official accusations, several AI experts express skepticism that mere distillation could account for the advanced capabilities reportedly displayed by Moonshot’s Kimi K3, especially given the tight timeframe involved. Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, voiced significant doubts. "I don’t think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Hancock told TechCrunch. He highlighted the temporal impossibility: "There’s just not even frankly time, right? Fable’s only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks." This sentiment underscores the immense computational resources and time typically required to develop a frontier-level LLM from scratch or even via advanced distillation methods.
Nathan Lambert, an AI researcher at the Allen Institute for AI, echoed this skepticism in a recent podcast, suggesting that the impact of distillation diminishes as models approach the technological frontier. "I’ve been of the opinion that distillation is becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning]," Lambert stated. He argued that if simple distillation were sufficient, "everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won’t see this, from supervised fine-tuning alone."
Lambert further elaborated that while supervised fine-tuning (SFT) can imbue a model with certain "manners" or stylistic characteristics, making it appear similar to its teacher, achieving true Fable-like capabilities would likely necessitate more sophisticated and resource-intensive reinforcement learning (RL) techniques. These advanced methods often involve training an agent of the larger model to grade the smaller model’s responses and adjust its learning based on these evaluations. Such complex RL runs can demand tens of millions of agents and colossal infrastructure, making the use of a frontier lab’s public API for such extensive distillation prohibitively expensive and logistically slow, potentially offering little performance uplift compared to native development.
A Pattern of Allegations and Industry-Wide Practices
This isn’t the first time Moonshot has faced such accusations. Earlier this year, Anthropic publicly accused Moonshot, along with DeepSeek and MiniMax, of systematically distilling its models. Anthropic claimed to have detected millions of interactions between its models and users associated with these companies, identified through IP addresses and other metadata. These query patterns were described as "distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use." However, Anthropic has not yet responded to specific queries regarding the alleged distillation of its Fable LLM in relation to Kimi K3.
It is also important to note that the practice of distillation, or at least the generation of synthetic datasets from existing models, is not unique to Chinese firms. Elon Musk, for instance, testified earlier this year that his company, xAI (the developer of Grok), had utilized OpenAI models for training purposes, characterizing the practice as common within the industry. This highlights the blurry lines between legitimate knowledge transfer, competitive analysis, and outright intellectual property infringement in the rapidly evolving AI landscape, where the concept of a "copy" is far more nuanced than in traditional software.
The Semiconductor Export Control Dilemma
Beyond the intellectual property dispute, Kratsios’s accusation also points to a critical national security concern: Moonshot’s alleged acquisition and use of advanced Nvidia Grace Blackwell 300 (GB300) chips, and access to GB300-equipped servers in Thailand. These chips represent the cutting edge of AI computational power and are subject to stringent U.S. export controls designed to prevent China from gaining a technological advantage in strategic areas like artificial intelligence.
The U.S. government has progressively tightened restrictions on the sale of advanced AI chips to China, viewing control over this foundational technology as paramount to maintaining its lead in AI development. Despite these controls, a burgeoning black market for such chips has reportedly emerged, enabling Chinese entities to bypass restrictions. Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology (CSET), confirmed the existence of such illicit channels. The severity of this issue was underscored in May when the founder of Supermicro, a major U.S. server builder, was indicted for allegedly smuggling advanced chips into China.
These incidents highlight the immense challenge of enforcing export controls in a globally interconnected supply chain. Bresnick advocates for more robust oversight, stating, "I am a proponent of know your customer laws for data centers across the world. If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing." In 2024, President Joe Biden’s Department of Commerce proposed federal "know-your-customer" (KYC) rules for data centers, but progress on their implementation appears to have stalled under the current administration. Exporters shipping advanced chips abroad are already mandated to ensure their products are used only for approved purposes, but verifying compliance at every stage of the supply chain remains a complex task.
The Broader Geopolitical and Economic Implications
This latest controversy is deeply embedded within the broader geopolitical rivalry between the United States and China, particularly in the race for AI supremacy. Both nations view AI as a critical technology for economic growth, national security, and global influence. The U.S. strategy involves maintaining its technological lead and preventing adversaries from accessing critical components, while China is aggressively pursuing indigenous innovation and self-sufficiency in key technologies.
The allegations against Moonshot underscore several critical implications:
- Challenges to Intellectual Property in AI: The case highlights the difficulty of defining and protecting intellectual property in the age of generative AI. How can "copying" be proven when models learn from data that might include outputs from other models? The concept of "watermarks" offers a potential solution, but their reliability and legal standing are still being established.
- Efficacy of Export Controls: The alleged acquisition of banned chips by Moonshot raises serious questions about the effectiveness of current export control regimes. If advanced chips can still reach prohibited destinations through black markets or indirect channels, the strategic advantage sought by the U.S. could be undermined.
- The Future of Open-Weight Models: The debate over banning Chinese open-weight models could have significant repercussions for the global AI ecosystem. Open-weight models, by their nature, are more accessible, fostering wider innovation and research. However, if they are perceived as vectors for IP theft or national security risks, governments might impose stricter controls, potentially fragmenting the global AI landscape.
- Chinese AI Capabilities: The incident also prompts a reevaluation of China’s independent AI development capabilities. While the accusations suggest reliance on U.S. technology, experts like Braden Hancock caution against underestimating Chinese expertise. "In general, Americans are understating the technical expertise of these Chinese teams," Hancock asserted. "One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. …if American models ground to a halt, I think China’s progress would slow, but would still continue. They’re not just riding coattails here." This perspective suggests that while access to U.S. technology might accelerate China’s progress, their underlying research and engineering talent are substantial and capable of independent advancement.
Looking Ahead: Regulatory Scrutiny and Enforcement
The allegations against Moonshot are likely to intensify regulatory scrutiny on both intellectual property and export control enforcement. The U.S. government will face pressure to provide more concrete evidence for its claims of distillation and chip smuggling. Simultaneously, the industry will watch closely to see if the proposed KYC rules for data centers gain traction and if new mechanisms are developed to track and verify the end-use of advanced AI chips globally.
The ongoing saga between the U.S. and China in the AI domain is not merely a commercial dispute; it is a battle for technological leadership that will shape economic power, national security, and ethical standards for decades to come. The outcome of this specific controversy, and how it informs future policy, will have profound implications for the global development and deployment of artificial intelligence.
