Key facts
- The performance gap between frontier AI models from US tech companies and the best open-weights models from Chinese companies has closed to just 4.4 months.
- Moonshot AI s Kimi K3, a leading open model, scores three points behind Anthropic s Fable 5 closed frontier model on the Artificial Analysis Intelligence Index.
- Kimi K3 costs 30 percent of Fable 5's price.
- The best closed AI models can handle tasks 1.7 times longer than the best open models.
- Tasks requiring eight to 12 hours are typically where closed frontier models excel over open models.
- Eight of the top 10 models by token volume on OpenRouter in August 2026 are open-weights models.
The performance gap between leading open-weights AI models from Chinese companies and frontier AI models developed by US tech giants has narrowed to approximately 4.4 months, according to a September 15 report by Mozilla. This shrinking gap, combined with the significantly lower cost of open models, is prompting many organizations to adopt them for routine tasks.
The "State of Open Source AI" report, shared with Ars prior to publication, highlights that Moonshot AI's Kimi K3, a prominent open model, achieved a composite AI performance score just three points behind Anthropic's Fable 5 closed model, while costing only 30 percent as much. This suggests that organizations should ideally use open models as the default for the majority of their work, reserving the more expensive closed models for specialized, expert-level tasks.
According to Mozilla CTO Raffi Krikorian, closed models still command a premium for expert professional work, high-intensity retrieval, and long context tasks. However, he noted that the decision to pay for closed models is becoming increasingly workload-specific rather than organization-specific. Many organizations still opt for closed models due to their out-of-the-box functionality, bundled compliance, support, and accountability, as they may lack the in-house expertise to effectively run open-weights models.
The research nonprofit METR has defined an AI model's "time horizon" as the length of tasks that AI can reliably complete with a 50 percent success rate. Currently, the best closed models can handle tasks 1.7 times longer than the best open models. Krikorian explained that tasks typically requiring eight to 12 hours are where closed frontier models currently hold an advantage, while tasks under eight hours can be handled by either type of model, making open models a more cost-effective choice.
Benchmarking company Vals AI, using its own neutral harness, found that the Chinese open-weights model GLM 5.2 from Z.ai scored within one point of Anthropic's Claude Opus 4.7 and 4.8, at approximately one-fifth the per-task cost. This indicates that the premium for closed models buys about a four-month head start for tasks in the eight-to-12-hour range.
The increasing popularity of open-weights models is evident on platforms like OpenRouter, where eight of the top 10 models by token volume in August 2026 were open-weights. However, open models still lag significantly in revenue, capturing only 4 percent of the market compared to 96 percent for closed models, according to a May-September 2025 analysis by Frank Nagle and Daniel Yue for the Linux Foundation. Krikorian anticipates this revenue share will shift as open models gain traction.
Krikorian pointed out that most open models currently in use globally are developed in China, with Chinese labs employing a strategy similar to that of Android by offering the technology freely while controlling the surrounding ecosystem. He expressed concern about this concentration of power, advocating for US and European labs to compete in the open model space to prevent any single country from dictating global defaults. He suggested that a coalition of institutions with a mission, similar to those that funded open-source software like Linux, could foster a more pluralistic ecosystem for open AI models.
