Xiaomi’s MiMo-V2.6-Pro leads open-model rankings on cost and performance, while Anthropic alleges Claude data helped train the model.

Xiaomi says its new MiMo-V2.6-Pro has become the strongest openly available AI model in current rankings, combining a reported Intelligence Index score of 46 with unusually low inference costs. The release puts Xiaomi at the center of a growing open-model race, but it also arrives amid Anthropic allegations that the Chinese company used Claude-generated exchanges to improve its own models.
The model’s reported performance is paired with a price of $0.435 per million input tokens and $0.87 per million output tokens. According to analysis cited by The Decoder, that works out to roughly $0.13 for a single benchmark task—an important figure for developers evaluating capable models for coding, agents, and other high-volume workloads.
The larger MiMo-V2.6-Pro is described as a mixture-of-experts model with 1.02 trillion total parameters, although only about 42 billion are active for an individual request. Xiaomi is also releasing MiMo-V2.6-Flash, a smaller model intended to provide a more efficient option, and a Pro-UltraSpeed variant that the company says can produce output at up to 20 times the speed of the standard version.
The Decoder reported that MiMo-V2.6-Pro ranked ahead of models including Kimi K3 and Qwen in Artificial Analysis’s comparison of openly available systems. A separate Unite.AI headline also identified a score of 46 as the model’s leading result among open-weight models, while The Next Web reported the ranking milestone without providing additional technical detail in the available source material.
Those results matter because Xiaomi is not positioning MiMo-V2.6 only as a research release. Low token prices can change the economics of products that make frequent model calls, including coding assistants, retrieval systems, workflow automation, and AI agents. The model’s mixture-of-experts design may also allow Xiaomi to advertise a large capability footprint without paying the full computational cost of activating all parameters on every request.
Xiaomi attributes the improvement to a substantially expanded reinforcement-learning phase. The company says it increased the amount of data processed in each training step, broadened the range of task environments, and devoted more computing resources to grading model outputs.
According to the figures reported by The Decoder, the reinforcement-learning run lasted fewer than six days and cost approximately $2.62 million for MiMo-V2.6-Pro and $850,000 for MiMo-V2.6-Flash. On the DeepSWE coding evaluation, Xiaomi reported that Pro improved from 58.4 to 72.6, while Flash rose from 48.8 to 65.7.
The company also described measures intended to keep training stable. Xiaomi froze the model’s internal distribution mechanism and added protections against reward hacking, in which a model maximizes a grading signal without genuinely completing the intended task.
Xiaomi is releasing more than model weights. Its package includes a technical report, a reinforcement-learning framework, a smaller model for additional training, and about 7,000 tasks with automatic graders. The tasks cover software development, cybersecurity, office work, and web design, with roughly 1,000 additional tasks focused on music composition.
The reported task sources are mixed. Some coding examples came from employee GitHub pull requests and user queries, while other descriptions were generated by a language model. Cybersecurity tasks draw on OSS-Fuzz, a collection of real software vulnerabilities, and office environments were rebuilt synthetically. For builders, access to the training framework and graders may be as significant as access to the model itself because it offers a starting point for reproducing or extending Xiaomi’s approach.
The headline performance and cost figures should be treated as reported claims rather than an independent verdict on overall model quality. Xiaomi supplied the training-cost, reinforcement-learning, and improvement figures, while Artificial Analysis supplied the cited Intelligence Index comparison as described by The Decoder. The available source material does not provide a complete methodology for the ranking, a broad set of independent evaluations, or a full breakdown of the model’s safety and reliability behavior.
The more serious uncertainty concerns how Xiaomi assembled training data. Anthropic’s threat-intelligence report, covering Claude abuse cases from December 2025 through August 2026, named Xiaomi among seven Chinese AI labs allegedly involved in campaigns to extract capabilities from Claude. Anthropic referred to the practice as illegal distillation. The other named companies were Alibaba, Moonshot AI, DeepSeek, Zhipu, MiniMax, and SenseTime.
In the case identified as GTG-16008, Anthropic said it observed more than 400,000 exchanges over 20 days in March and April 2026. The report alleged that Xiaomi routed conversations and coding sessions from its MiMo models through OpenClaw and OpenCode to Claude, with the goal of generating training data for later models.
These are Anthropic’s allegations, not independently established findings in the material available for this story. The Decoder reported that Anthropic provided limited documentation about the origins of earlier teacher-model data used in the internal distillation process. Xiaomi’s public release of reinforcement-learning tools does not, by itself, resolve questions about the provenance of data used elsewhere in its development pipeline.
If the reported pricing and ranking hold across independent tests, MiMo-V2.6-Pro could put pressure on model providers whose products charge substantially more for comparable coding or reasoning workloads. Startups could use the price to run larger evaluation suites, more frequent agent loops, or higher-volume customer workflows. Enterprise teams may view the open availability and training toolkit as an opportunity to customize the system rather than rely entirely on a hosted API.
The trade-off is that low inference cost does not automatically mean low deployment risk. Buyers still need to assess licensing terms, data governance, latency, uptime, security behavior, context handling, and output reliability on their own tasks. The allegations around Claude-derived data add another layer: companies considering MiMo-V2.6 for production use may want clear documentation of training-data provenance and legal assurances before embedding it into sensitive workflows.
For researchers, Xiaomi’s release also highlights a practical shift in open-model development. The competitive advantage may increasingly come from reinforcement-learning environments, graders, and task design rather than from pretraining scale alone. Open tools can accelerate progress, but their value depends on whether the tasks measure useful behavior and whether reward systems resist shortcuts.
The first signal will be independent replication of the Intelligence Index result and the reported DeepSWE gains. Developers should also compare MiMo-V2.6-Pro with Kimi K3 and Qwen under identical prompts, prices, hardware, and latency conditions rather than relying on a single leaderboard.
A second signal is Xiaomi’s documentation of training data and model-development practices. Anthropic’s allegations could lead to further evidence, a response from Xiaomi, or changes in how open-model companies disclose distillation and teacher-model use.
Finally, the market will reveal whether Xiaomi’s toolkit is usable outside its own infrastructure. Adoption by researchers and product teams, the quality of the automatic graders, and real-world performance in coding, cybersecurity, office automation, and agent workflows will show whether MiMo-V2.6 is more than a strong benchmark release.
MiMo-V2.6 is notable because it links three pressures shaping the open-model market: benchmark performance, inference economics, and accessible reinforcement-learning infrastructure. Xiaomi’s reported results suggest that a model can compete near the top of open rankings without matching the highest prices of proprietary systems.
But the Claude allegations make provenance part of the product story, not a footnote. For AI builders and enterprise buyers, the practical test is therefore twofold: whether MiMo-V2.6 delivers reliable value on real workloads, and whether Xiaomi can provide enough transparency for organizations to understand what they are deploying and how it was trained.