The landscape of artificial intelligence has shifted significantly in the first half of 2026, as a trio of Chinese research laboratories—Moonshot AI, DeepSeek, and Zhipu AI—have claimed the top positions on the global open-weight model leaderboards. This emergence marks a pivotal moment in the democratization of high-compute AI, as these models, characterized by trillion-parameter architectures and million-token context windows, now rival or exceed the performance of the world’s most prominent proprietary systems. The release of Moonshot AI’s Kimi K3, DeepSeek’s V4 Pro, and Zhipu AI’s GLM-5.2 represents a concerted effort to dominate long-horizon coding and agentic workloads, providing developers with high-performance alternatives to closed-source ecosystems.
Technical Architectures and Model Specifications
The three models utilize sparse Mixture-of-Experts (MoE) architectures, a design choice that allows for massive total parameter counts while maintaining manageable computational costs during inference by activating only a fraction of the total parameters for any given token.
Moonshot AI’s Kimi K3 stands as the most ambitious of the three, boasting a total parameter count of 2.8 trillion. As a "Stable LatentMoE" model, it utilizes a sophisticated routing system that activates 16 out of 896 experts per token. While the exact active parameter count remains undisclosed, the model’s scale suggests a significant leap in reasoning depth. K3 is uniquely positioned as a multimodal powerhouse, incorporating native vision and video processing capabilities alongside its 1-million-token context window. It also features "always-on reasoning," a paradigm that allows the model to perform background computation to refine its outputs for complex queries.
DeepSeek V4 Pro, released by the DeepSeek-AI team, follows with 1.6 trillion total parameters. It employs a more granular expert structure, utilizing 384 routed experts plus one dedicated shared expert, with 49 billion parameters active during any single forward pass. DeepSeek has optimized V4 Pro for text-based workloads, offering a 1-million-token input context and a class-leading 384,000-token maximum output limit. This makes it particularly attractive for massive code generation tasks and long-form document synthesis.
Zhipu AI’s GLM-5.2 is the smallest of the trio in terms of raw parameter count, totaling 744 billion parameters (reported as 753 billion by some independent auditors). Despite its smaller footprint, it activates approximately 40 billion parameters per token, making its "active-to-total" ratio much higher than its competitors. GLM-5.2 was the first of this generation to demonstrate that open-weight models could effectively handle million-token contexts without significant degradation in retrieval accuracy.
Chronology of the 2026 AI Surge
The release cycle of these models illustrates the rapid-fire nature of the current AI arms race in China. The sequence began on April 24, 2026, with the launch of DeepSeek V4 Pro. At the time, DeepSeek set a new benchmark for cost-to-performance ratios, offering weights on Hugging Face immediately and challenging the pricing models of Western labs.
Six weeks later, on June 13, 2026, Zhipu AI released GLM-5.2. This model was positioned as a high-speed alternative, specifically optimized for throughput and low-latency agentic tasks. It briefly held the title of the highest-performing open-weight model on several independent leaderboards, including the Artificial Analysis Intelligence Index.
The current hierarchy was established on July 16, 2026, with the debut of Moonshot AI’s Kimi K3. By introducing a 2.8-trillion-parameter model with multimodal capabilities, Moonshot AI effectively bridged the gap between open-weight models and the leading proprietary "frontier" models like GPT-5.5 and Claude 5. The release of K3 was accompanied by a commitment to publish weights by late July, signaling a new era where "frontier" performance is no longer gated behind API-only access.
Performance Benchmarking and Capability Analysis
Comparing these models requires looking past vendor-reported figures, which often use differing evaluation harnesses. The Artificial Analysis Intelligence Index provides a normalized view of their capabilities. On this index, Kimi K3 achieved a score of 57, placing it third globally, behind only Claude Fable 5 and GPT-5.6 Sol. This puts K3 in direct competition with high-end proprietary models like Opus 4.8.
In contrast, GLM-5.2 scores a 51, and DeepSeek V4 Pro (in its Max reasoning mode) scores a 44. While K3 leads in general intelligence, the specific domain of software engineering shows a more nuanced competition. In Moonshot’s internal testing, K3 outperformed GLM-5.2 across all coding benchmarks. For example, on the SWE Marathon—a test of long-term autonomous coding—K3 scored 42.0 compared to GLM-5.2’s 13.0.
However, DeepSeek V4 Pro remains a formidable specialist. It achieved an 80.6% score on SWE-bench Verified, a result that tied it with Gemini 3.1 Pro at the time of its release. Furthermore, DeepSeek’s 83.5% score on the MRCR 1M (a long-context retrieval benchmark) confirms its reliability for large-scale codebase analysis. GLM-5.2 also remains competitive in coding, posting a 62.1% on SWE-bench Pro, which notably edged out the performance of GPT-5.5 (58.6%) in similar testing environments.
Licensing and Open-Weight Accessibility
The term "open-weight" carries different practical implications for each of these models, particularly regarding their commercial utility and immediate availability.

DeepSeek V4 Pro and GLM-5.2 are the most accessible for developers today. Both models are released under the MIT License, which is among the most permissive in the software industry. Their weights were uploaded to Hugging Face on their respective launch days, allowing for unrestricted commercial use, fine-tuning, and self-hosting. This transparency has led to a rapid proliferation of "quantized" versions and community-driven optimizations.
Moonshot AI has taken a more phased approach with Kimi K3. While the model is currently accessible via API and the Kimi application ecosystem, the weights are scheduled for release on July 27, 2026. Moonshot intends to use a "Modified MIT License." This license includes an attribution clause that only becomes mandatory for products exceeding 100 million monthly active users. For the vast majority of enterprise and individual developers, the license functions identically to a standard MIT agreement, though the delayed release of weights remains a point of contention for teams requiring immediate on-premises deployment.
The Economics of Inference: API Pricing and Efficiency
A critical axis for AI teams is the cost of serving these models. DeepSeek has positioned itself as the aggressive price leader. Its API pricing is an order of magnitude lower than its rivals: $0.435 per million input tokens and $0.87 per million output tokens. This makes DeepSeek V4 Pro the most viable option for high-volume, low-margin applications.
Zhipu AI’s GLM-5.2 occupies a middle ground, priced at $1.40 for input and $4.40 for output per million tokens. While more expensive than DeepSeek, it offers a significant speed advantage. Artificial Analysis recorded GLM-5.2 at 168 tokens per second, nearly triple the speed of DeepSeek V4 Pro and Kimi K3, which both hover around 62 tokens per second. For applications where latency is the primary constraint—such as real-time chat or interactive agents—GLM-5.2 offers the best value.
Kimi K3 is the premium offering in the group. Its list price is $3.00 for input and $15.00 for output per million tokens. However, Moonshot AI has introduced aggressive caching discounts. For coding workloads, where repetitive context is common, Kimi K3 can achieve cache hit rates above 90%, reducing the effective input cost to $0.30 per million tokens. This suggests that K3 is economically optimized for iterative tasks like software development rather than one-off queries.
Hardware Infrastructure and Self-Hosting Realities
While these models are "open-weight," the hardware required to run them locally is substantial, creating a high barrier to entry for self-hosting.
The 744-billion-parameter GLM-5.2 requires over 1 terabyte of VRAM when running in BF16 precision. In a production environment, this typically necessitates a cluster of eight NVIDIA H200 GPUs using FP8 quantization. DeepSeek V4 Pro, at 1.6 trillion parameters, scales this requirement even further, demanding multi-node configurations that are usually only available to well-funded research labs or large enterprises.
Kimi K3 represents the peak of hardware demand. Moonshot AI recommends a minimum of 64 accelerators for efficient serving of the 2.8-trillion-parameter model. To mitigate these requirements, K3 utilizes MXFP4 weights with MXFP8 activations. This advanced quantization scheme is designed to maximize the utility of next-generation hardware (such as Blackwell-based systems) but remains largely out of reach for consumer-grade or mid-range enterprise hardware.
Broader Impact and Industry Implications
The emergence of Kimi K3, DeepSeek V4 Pro, and GLM-5.2 signals a shift in the global AI power dynamic. Historically, the most capable "frontier" models were the exclusive domain of a few US-based labs. The fact that three Chinese labs now hold the top positions in open-weight performance suggests that the gap between proprietary and open systems is closing rapidly.
For the global developer community, these models provide a hedge against "vendor lock-in." The ability to switch from a proprietary API to a similarly capable open-weight model—or to host that model on private infrastructure—gives enterprises unprecedented control over their data and costs.
Furthermore, the focus on 1-million-token context windows and agentic coding benchmarks indicates that the industry is moving away from simple "chatbot" interactions toward autonomous "AI workers." These models are not just designed to answer questions; they are built to navigate entire codebases, manage complex workflows, and act as the central nervous system for sophisticated software agents.
As Moonshot AI prepares to release the K3 weights in late July, the industry will be watching to see if the "open-weight" movement can sustain this level of rapid innovation. For now, the choice for AI teams is clear: DeepSeek for cost, GLM for speed, and Kimi for peak capability. This trifecta ensures that the open-source ecosystem is no longer a step behind, but is instead setting the pace for the future of artificial intelligence.
