A report says Chinese AI developers publicly documented safety tests for only 3.6% of model releases, raising transparency concerns for buyers.

A report cited by Reuters, The Economic Times and marketscreener.com says Chinese AI developers publicly disclosed safety-test results for only 3.6% of their model releases. The finding points to a significant gap between the pace of model publication and the amount of safety evidence available to users, regulators and enterprise buyers.
The report’s headline statistic is the central fact available from the coverage. The supplied source material does not identify the report’s authors, sample size, time period, definition of a model release or the tests counted as publicly disclosed. Those details are important: a release could refer to a major foundation model, a fine-tuned variant, an application-facing model or another type of update, while “published safety tests” could cover anything from a formal evaluation report to a limited model card.
Even with those qualifications, the 3.6% figure is relevant to teams deciding whether a model can be deployed in sensitive workflows. Public documentation is one of the few ways outside users can assess a model’s known failure modes, test coverage and limits before they commit data, money and operational responsibility to it.
The three supplied sources carry the same news headline and appear to be separate versions of a wire report rather than three independent investigations. Reuters is the primary named wire source in the cluster, while The Economic Times and marketscreener.com republished or surfaced the same finding. Their agreement supports the existence of the reported statistic, but it does not independently validate the underlying methodology.
The evidence available here does not establish that 96.4% of Chinese models were never tested. It says safety tests were not published for those releases, which is a different claim. Developers may conduct internal evaluations without releasing the results, may publish information in formats not captured by the report, or may disclose testing to customers and authorities rather than to the public.
That distinction matters for procurement. An absence of public evidence is not proof that a model is unsafe. It does, however, limit the ability of independent researchers and buyers to verify the developer’s claims. The statistic should therefore be read as a transparency measure, not as a direct measurement of model risk.
For AI builders and product teams, safety testing is useful only when it is connected to a defined deployment context. A general-purpose model may perform acceptably in a customer-service assistant but fail in ways that create serious exposure when used for medical guidance, financial decisions, code generation or access to internal systems.
Published evaluations can help teams compare models on more than accuracy. They may reveal how a system handles harmful requests, sensitive information, prompt manipulation, hallucinations, bias or tool use. They can also show whether a developer tested the model against the kinds of adversarial inputs that are likely to appear in production.
The reported publication rate suggests that many buyers may be forced to rely on private documentation, vendor assurances or their own testing. That shifts cost and responsibility downstream. A startup integrating a model into its product may need to build an evaluation program from scratch, while a large enterprise may require contractual commitments, audit access and repeated testing as the model changes.
The immediate implication is not that Chinese AI models should be excluded from consideration. Rather, buyers may need stronger evidence requirements before placing them in high-impact workflows. Those requirements could include model cards, versioned evaluation results, incident reporting, red-team summaries and clear explanations of what was not tested.
Model developers also face a practical trade-off. Publishing detailed safety results can expose weaknesses or increase scrutiny, but it gives customers a basis for trust and can reduce duplicated testing across the market. For developers competing internationally, transparent reporting may become part of the product rather than an optional communications exercise.
For builders, the 3.6% finding reinforces the need for independent evaluation. Teams should test the exact model version they plan to use, under the prompts, tools, retrieval systems and permissions present in their application. A public benchmark cannot substitute for deployment-specific testing, and a lack of public testing should increase the level of verification rather than end the analysis.
The issue also matters to regulators and platform operators. If model releases are frequent and safety documentation is inconsistent, oversight based only on public launch announcements will provide an incomplete view of the market. Regulators may focus more closely on disclosure obligations, while cloud and application platforms could introduce their own documentation requirements for models offered to enterprise customers.
The first signal to watch is the underlying report itself: its authors, dataset, time window and criteria for counting a published safety test. Without that information, the 3.6% figure cannot be compared reliably with developers in other countries or with earlier periods.
The next is whether major Chinese model developers begin publishing standardized evaluations for new releases. Useful disclosures would identify the model version, test design, limitations, known failure cases and the date of the evaluation. Repeated reporting across versions would be more informative than a one-time safety statement.
Enterprise buyers should also watch procurement requirements from cloud providers and large software platforms. If those intermediaries begin requiring safety documentation before listing or integrating a model, transparency could become a commercial requirement even where regulation remains unclear.
Finally, researchers should look for evidence that public testing correlates with better operational outcomes. More reports alone will not prove that a model is safer; the quality, independence and relevance of the evaluations will matter more than the number of documents published.
The reported 3.6% rate is best understood as a warning about visibility, not a verdict on every Chinese AI model. Because the available coverage does not provide the report’s methodology, readers should avoid treating the number as a complete measure of developer competence or model safety.
Its importance lies in the decision burden it places on AI adopters. When public testing is rare, product teams must compensate with stronger internal evaluations, tighter deployment controls and documented monitoring. For the market, the competitive advantage may increasingly go to developers that can show not only capable models, but also repeatable evidence about where those models work and where they fail.