
Thomson Reuters has spent about $40 million over more than two years building Thomson, an in-house language model for legal and professional information work. The company’s decision marks a move away from relying entirely on rented capacity from frontier-model providers such as OpenAI and Anthropic.
The model is built on Alibaba’s Qwen and is being introduced first inside CoCounsel Legal for document-focused work. Thomson Reuters says the investment is intended to improve economics, protect control over its proprietary content and let the company train models inside its own software tools. The strongest performance results reported so far, however, come when Thomson can access Thomson Reuters content and workflows rather than operating as a general-purpose model.
According to reporting by The Decoder, Thomson Reuters developed the model using company data and domain experts, with the model trained and evaluated in environments connected to products including Westlaw, Practical Law, Checkpoint and Reuters. The company says less than 10% of its available content has been used for training so far.
The development process began with Qwen as the open model foundation. Thomson Reuters worked with Imperial College on an intermediate version called Snowdon, which was retrained for safety, ethics and political neutrality, The Decoder reported. The company then added pre-training on its own material, expert-led post-training and reinforcement learning in internal tool environments.
Thomson is not being positioned as a replacement for every external model. At launch, it takes over the Tabular Analysis feature in CoCounsel Legal, where a smaller, lower-cost model can handle a defined task at high volume. Administrators can still switch models, and the product remains multi-model. Thomson Reuters also says customer data is not used for training.
A smaller version of Thomson is expected to be released on Hugging Face under a non-commercial license. The company plans to publish a technical report and create a developer portal, while early discussions about direct licensing with law firms are described as non-binding.
The evidence supplied by The Decoder suggests that Thomson’s competitive case is narrower than a claim of broad frontier-model superiority. On Stanford LegalBench, Thomson reportedly scored 0.823, behind Gemini 3.1 Pro and GPT-5.5. On the Harvey Legal Agent Benchmark, it was just behind Opus 4.8, while leading on instruction following and the difficult PrBench Legal evaluation.
The model reportedly performed less strongly on general reasoning and coding. Comparisons are also affected by evaluation methods: Thomson used test-time scaling, while GPT-5.5 was tested without a reasoning mode. The company’s own Deep Research evaluation produced a similar qualification. With web access alone, Thomson scored 0.53 for factual accuracy, compared with 0.65 for GPT-5.4. When Thomson Reuters content was added, Thomson reached 0.83, narrowly ahead of GPT-5.4 at 0.82.
Those results are company-reported benchmarks as presented through The Decoder, not independent validation. They also show that proprietary access may be doing at least as much work as model specialization. The same content gave GPT-5.4 a substantial improvement, meaning the advantage cannot yet be attributed solely to Thomson’s training or architecture.
Andrew Bean, Thomson Reuters’ evaluation lead, acknowledged that the model was not yet the leader when limited to web access. The results could improve as the company uses more of its content or moves to a newer Qwen variant, but those remain future possibilities rather than demonstrated outcomes.
The company’s argument rests on three linked issues: cost, data control and accumulated expertise. Thomson Reuters says fine-tuning an external frontier model can reduce general capability, while continued use leaves the buyer exposed to the provider’s inference pricing and product roadmap.
An internally controlled model can be smaller and cheaper for repetitive work such as document review or citation checking. That matters more when a model is deployed across high-volume workflows than when it is used occasionally for open-ended research. The relevant comparison is therefore not simply whether Thomson is smarter than GPT-5.4 or another frontier system, but whether it can deliver acceptable quality at a lower operating cost inside a defined product.
Data access is the second factor. Training and testing within Westlaw and related tools allows Thomson Reuters to optimize for tasks that outside providers cannot fully reproduce without access to the company’s protected material and workflow telemetry. The company says that advantage is central to its decision not to rely exclusively on third-party models.
The third factor is compounding development knowledge. Thomson Reuters executives Joel Hron and Jonathan Schwartz described the company’s investment as building a model-development capability rather than buying a single finished model. Each expert review, product evaluation and workflow improvement can potentially inform later releases. That claim is strategic reasoning from company executives, not a measured return on investment.
For enterprises, the announcement provides a more specific case for selective model ownership. Thomson Reuters has exclusive information assets, hundreds of domain experts and workflows where results can be evaluated against professional standards. Those conditions make a specialized model more defensible than an internal project built without proprietary data or reliable testing.
The news is less persuasive for companies that lack those ingredients. Training a model on generic corporate documents does not automatically create a durable advantage, particularly if the organization still depends on external infrastructure, cannot measure quality at scale or has too few examples to improve the system. In those cases, model customization, retrieval systems or a multi-model product strategy may offer more flexibility than owning a full training pipeline.
The deployment also illustrates why model choice is becoming a product-management decision rather than a one-time technical decision. CoCounsel Legal keeps multiple models available and assigns Thomson to a narrower workload. That arrangement lets Thomson Reuters use frontier systems where their broader reasoning is valuable while reserving its own model for tasks where cost, data access and repeatability matter more.
For AI agents and enterprise software teams, the important lesson is that tool access and evaluation environments can matter as much as the base model. A model trained to operate within a company’s own applications may outperform a stronger general model on a limited workflow, but only if the organization can maintain the tools, data permissions and quality controls around it.
The first signal will be actual performance and cost in CoCounsel Legal, especially for Tabular Analysis and other document-review tasks. Thomson Reuters has not provided independent evidence in the supplied reporting that the model produces a defined cost reduction or measurable customer-quality improvement in production.
The release of the smaller Hugging Face model and its technical report should provide more information about architecture, training methods, licensing and reproducibility. It will also show whether outside developers can use Thomson beyond Thomson Reuters’ own controlled environments.
Law-firm licensing discussions are another signal, although The Decoder characterized them as early and non-binding. Adoption by external legal organizations would offer a stronger test of whether Thomson is a product asset rather than primarily an internal optimization.
Finally, future evaluations should compare Thomson with current frontier models under equivalent reasoning settings and with clearly separated web-only, proprietary-content and tool-enabled conditions. That would make it easier to determine how much of the reported advantage comes from the model itself.
Thomson Reuters’ investment is not evidence that every enterprise should train its own language model. It is evidence that ownership can make sense when a company controls scarce data, has specialists who can supervise the system and operates repetitive workflows where inference economics are measurable.
The more important strategic move may be the combination of internal model development with continued use of external models. Thomson Reuters is not abandoning the frontier market; it is trying to own the part of the stack where proprietary content and high-volume legal tasks create leverage. For builders and enterprise buyers, that is a more credible model strategy than treating full in-house training as an end in itself.
Thomson Reuters is spending $40 million on its own legal AI model, seeking lower costs and control over proprietary data and high-volume workflows.