UN partners with Google to make global statistics usable by AI agents

The UN is building an AI-ready statistics platform with Google, adding natural-language search, provenance, and MCP access after poor model accuracy.

AI News

The United Nations is working with Google to rebuild how its statistical data can be found and used by people and AI systems. The new UN System Data Commons replaces the older UNData portal with natural-language search and direct machine access through the Model Context Protocol, or MCP.

The project responds to a practical problem: general-purpose AI models often fail to retrieve basic development statistics accurately or consistently. A UNICEF test of six leading models found an average accuracy rate of 21.2% across more than 133,000 responses, according to João Pedro Azevedo, UNICEF’s chief statistician.

For AI builders and enterprise data teams, the initiative is notable because it targets the layer beneath the model. Rather than asking an AI system to rely on training data or open-ended web search, the UN is creating a structured source that agents can query, inspect, and cite.

From UNData to the UN System Data Commons

The UN System Data Commons is built on Google’s open-source Data Commons platform, which Google launched in 2018 to organize public datasets from different providers within a common framework. The UN version brings statistics from multiple UN agencies into one searchable environment.

The previous UNData portal required users to browse and search through a more conventional database interface. The new system allows users to ask questions in natural language and find statistics across participating agencies. It also records the origin of each statistic, allowing a user—or an AI application—to trace an answer back to its underlying UN source.

The UN said 26 entities have committed to the platform, with data from nearly 20 available at launch. Its target is to bring 80% of the UN system’s statistical datasets onto the platform by 2027. Those figures describe commitments and goals, not completed migration, and the evidence does not specify how much data is already standardized across agencies.

Google.org provided $2 million in capacity-building funding and technical support for the platform’s core infrastructure. Google’s Data Commons team told TechCrunch that the system is hosted on a UN-governed instance and is intended eventually to be maintained, operated, and scaled independently by the UN.

Why the AI benchmark matters

The partnership follows a UNICEF working paper that exposed weaknesses in AI-assisted access to development data. The test covered OpenAI’s GPT-4o and GPT-4o-mini, Anthropic’s Claude Sonnet 4.5 and Haiku 4.5, and Google’s Gemini 2.5 Flash and Gemini 2.0 Flash.

Azevedo said roughly three in five responses failed to provide a usable number, often because the models hedged instead of answering. When the same questions were run again on the same model versions about two days later, models that supplied a number on both occasions returned the same number only about half the time.

The results should be treated cautiously. UNICEF has not yet published the methodology, code, or test data, and the working paper has not been peer-reviewed. The findings are therefore an early warning about the reliability of AI retrieval, not a definitive ranking of the models tested.

UNICEF also reported a sharp increase in traffic arriving from generative AI assistants. Referrals from ChatGPT answers to UNICEF’s website increased 67% year over year between January 1 and September 14, according to Azevedo. Such referrals represented 6.4% of all sessions, while UNICEF estimated that AI assistants accounted for about one in 10 visits overall. These are agency-reported traffic figures, and they do not show whether visitors trusted or acted on the answers they received.

MCP turns statistics into an agent workflow

The platform’s support for the Model Context Protocol is central to its intended use. MCP gives AI applications a standard way to connect to external tools and data sources. Google added MCP support to Data Commons last year, allowing AI agents to query statistics and retrieve their sources directly.

That changes the workflow from simple question answering to data assembly. In a Google demonstration, an AI system connected to the UN data through MCP identified statistics related to the impact of the U.S. President’s Emergency Plan for AIDS Relief in Africa. It combined indicators including HIV infections, AIDS mortality, and life expectancy to produce an infographic.

For product teams, the important distinction is between retrieving one verified value and generating an analysis from several values. A connected system may be able to create dashboards, charts, or written reports without requiring a user to manually locate and combine datasets. However, the quality of the result still depends on whether the agent chooses appropriate indicators, understands their definitions, and preserves differences in geography, time period, and methodology.

Google’s demonstration is a vendor example rather than an independent evaluation of production performance. Access to authoritative data can reduce one source of error, but it does not prevent an AI model from misinterpreting the data or drawing an unsupported conclusion.

Implications for builders and institutions

The UN project points toward a more controlled architecture for AI applications that answer questions about public policy, health, development, and economics. Builders can use a governed data source with provenance instead of treating a language model as the database itself. That could make it easier to audit answers, refresh figures, and identify the source behind a generated chart.

It also raises implementation requirements. A useful agent interface must expose more than a number: it should preserve the dataset, publisher, definition, unit, geographic scope, and reporting period. Without that context, a technically correct statistic can still be used incorrectly. The UN’s decision to track source lineage is therefore as important as its natural-language interface.

Enterprise buyers should not interpret the launch as proof that AI-generated analysis is ready for unsupervised publication. Google’s Prem Ramaswami said human review should remain in place because models can misinterpret nuance. That warning is especially relevant for official statistics, where small definitional differences can materially change a conclusion.

The project also gives Google a role in the infrastructure used to access public-sector data, even though the stated destination is a UN-governed and independently operated system. Its longer-term significance will depend on whether the UN can maintain consistent schemas, update schedules, access controls, and source quality across dozens of agencies.

What to watch next

The first signal will be the release of UNICEF’s methodology, code, and data. That material should clarify how the benchmark defined accuracy, which questions were asked, and whether the reported model differences hold under independent review.

The second is actual adoption of the UN System Data Commons. Progress toward the 2027 target, the number of datasets migrated, and the consistency of metadata across agencies will reveal whether the platform is becoming operational infrastructure or remaining a demonstration layer.

AI developers should also watch for MCP integrations beyond Google’s examples, including independent tools that test citation accuracy, update freshness, and agent behavior when data sources disagree. For institutions, the key question is whether generated outputs can be audited reliably enough for research, policy analysis, and public communication.

Creati.ai perspective

The UN’s move addresses a neglected part of the AI stack: reliable access to structured, authoritative information. Better models alone are unlikely to solve the problem shown by UNICEF’s test if the underlying retrieval path is ambiguous or difficult to verify.

The harder test will be governance, not the demonstration. If the UN can keep definitions, provenance, and updates consistent while opening its data to agents, the platform could become a practical reference architecture for public-sector AI. If it cannot, MCP may simply make it faster to retrieve inconsistent data and turn it into polished but misleading analysis.

Ads