OguzhanTekin
What Thomson Reuters' $40 Million AI Investment Actually Shows
FinOps & Cloud StrategyAugust 27, 2026

What Thomson Reuters' $40 Million AI Investment Actually Shows

By Oguzhan TekinBack to Blog

For years, the AI industry followed a familiar formula: more data, more computing power, and more capital produced increasingly capable foundation models.

Thomson Reuters is pursuing a different enterprise strategy. Rather than training a general-purpose model from scratch, it began with an open-weight foundation and invested in the proprietary data, expert evaluation, infrastructure, and product workflows needed to make that model useful for professional work.

Thomson Reuters says it invested about $40 million in Thomson, its first proprietary large language model. Its technical report estimates that the final three-week training run for Thomson-1.0-Large cost under $450,000 in GPU compute.

Those figures do not measure the same thing.

The $450,000 represents the estimated computing cost of one final training run. The $40 million represents the broader development investment required to create an enterprise AI capability.

According to the report, that investment covered staff, computing, domain-expert compensation, vendor partnerships, reusable research, infrastructure engineering, and experimentation. It also absorbed unsuccessful experiments and organizational learning that made the final run possible.

A complete enterprise AI cost model must include more than development and training. It must also account for ongoing inference, monitoring, maintenance, human oversight, and the expected cost of errors.

In professional work, an unreliable result can create correction work, escalations, customer harm, compliance exposure, or legal risk. A model with a lower cost per request may still cost more overall if it requires more review and remediation.

Thomson Reuters' investment does not show that a frontier-quality professional model can be created for $450,000. It shows that once an organization has built the necessary research, data, evaluation, and deployment systems, the final training run may represent only a small part of the total investment.

In professional AI, the model is essential, but it is not the whole product. The defensible product is the system around it.

What Thomson Reuters Actually Built

Thomson-1.0-Large began with Alibaba's Qwen3.5-397B model.

That foundation was first adapted into Snowdon-1.0-Large by the Frontier AI Research Lab, jointly founded by Thomson Reuters and Imperial College London. Thomson Reuters then continued training it using its own content, professional tools, expert feedback, and evaluation methods.

Qwen supplied the base-model capabilities. Thomson Reuters supplied the proprietary content, expert feedback, retrieval tools, evaluations, and professional workflows that made those capabilities useful in its target areas.

Those assets include Westlaw, Practical Law, Checkpoint, Reuters, and approximately 175 years of accumulated material across legal, tax, accounting, and news. Thomson Reuters says it has trained Thomson on less than 10% of its available proprietary content so far.

Open-weight models reduce the cost and time required to begin adapting a capable general-purpose system. They do not eliminate the difficult work.

Information must be selected, cleaned, structured, governed, checked for usage rights, and connected to real tasks. Secure infrastructure must be designed. Evaluation methods must be created. The system must be monitored after deployment.

More content will not automatically produce a better model.

Training a model does not repair poor data or unclear work. It can make those problems more expensive.

The project also has a longer history than the launch figures suggest. Its research roots go back to Safe Sign Technologies, founded in 2022 and acquired by Thomson Reuters in 2024. The team had spent about three years studying model safety and reliability before Thomson was released.

Much of the reported investment appears to have built reusable research, engineering, and evaluation capabilities. That suggests future model iterations could have a lower incremental cost, although Thomson Reuters has not publicly disclosed what a later full-model update would cost.

The figure also understates the barrier facing a company starting from nothing. The final computing bill may be manageable. Building the data systems, technical team, expert network, and evaluation infrastructure can take years.

Expert Judgment Became Training Data

Companies often say their AI systems were built with expert input. That claim means little unless they explain how professional judgment became a training or evaluation signal.

Thomson Reuters provides more detail in its account of how it built Thomson.

The company employs around 1,500 attorney-editors whose work includes deciding whether legal analysis is correct, complete, current, and useful. Qualified lawyers spent thousands of hours reviewing model outputs and comparing competing answers against detailed criteria.

Some practitioners worked for months to create evaluation guides for difficult legal research questions. These guides did not simply ask whether an answer sounded convincing. They identified the specific legal and factual elements a professional response needed to contain.

Domain experts across the company also tested the system on difficult tasks and reported failures they were qualified to diagnose.

This matters because documents contain only part of an organization's knowledge.

Some of the most valuable knowledge exists between the first draft and the final version. It includes why an editor rejected a source, narrowed a claim, changed the emphasis, requested more evidence, or decided that a legally correct point was not useful in practice.

The finished document shows the result. Expert decisions reveal how that result was produced.

A blind preference study covering more than 3,000 expert-rated conversations gave Thomson Reuters another way to test whether professionals preferred Thomson's outputs. That is stronger evidence than a few selected demonstrations, but it remains part of a company-run evaluation program.

Thomson Reuters says independent academic evaluation is underway. Jonathan H. Choi of Washington University School of Law has also tested Thomson against ChatGPT and Claude using challenging corporate tax questions. Broader external validation will still be needed across different jurisdictions, users, tasks, and production environments.

The Benchmarks Require a Careful Reading

Thomson Reuters says Thomson performs competitively with leading frontier systems. Its published results support that claim, but they do not show that Thomson is best at everything.

The following figures come from Thomson Reuters' earlier highlighted seven-category comparison, rather than from the broader aggregate evaluation later published in its technical report.

In that earlier comparison, Thomson led three of seven benchmark groups. It performed especially well on a difficult professional legal benchmark, instruction following, and long-context tasks.

On Stanford LegalBench, however, Thomson scored 0.823. Gemini 3.1 Pro scored 0.843, while GPT-5.5 scored 0.832.

On general reasoning, Thomson scored 0.684. Gemini reached 0.748, and Claude Opus reached 0.737.

Thomson also finished last in the coding comparison, with a score of 0.399.

GPT-5.5 was evaluated in non-reasoning mode in that comparison, while Thomson Reuters disclosed limited use of additional computing at answer time for Thomson on certain tasks. The company later told LawNext that this increased Thomson's overall average only from 0.787 to 0.789.

The later technical report presents a broader set of tests and a different aggregate table. In that comparison, Thomson-1.0-Large recorded an overall average of 78.5%, compared with 79.5% for Claude Opus 4.8. Thomson still led or performed strongly in areas such as instruction following, document processing, retrieval-augmented generation, and selected professional tasks.

The fair conclusion is straightforward: Thomson is competitive with leading models and has clear strengths in professional work. It is not universally superior.

Benchmark results also depend on the selected tests, model versions, reasoning settings, tool access, prompts, and computing used to generate an answer.

The more important issue is the distinction between a model and a complete system.

A professional AI product combines the model with licensed content, retrieval tools, workflow instructions, product interfaces, human review, security controls, monitoring, and governance. Its performance cannot always be reduced to one model score.

The Content Test Measured a Complete System

Thomson Reuters also evaluated legal research performance using 53 questions written by internal subject-matter experts.

The test conditions were not equal. Thomson could search licensed Westlaw and Practical Law databases. Five external systems from OpenAI and Anthropic received broad access to the open web.

A separate Thomson Reuters article gave Thomson a factuality score of 0.83, compared with scores between 0.65 and 0.68 for leading frontier systems. Because the article does not clearly explain how those figures map onto the 53-question comparison, they should be treated as company-reported results rather than an independent head-to-head finding.

The company's factuality method is still useful. It extracted claims from research reports and checked whether the cited sources supported them.

An answer that sounds convincing is not enough in professional work. Its sources must support what it says.

But the result cannot be attributed to the model alone. Thomson had access to authoritative legal databases while the other systems searched the web.

The evaluation measured the full system, including the model, content, retrieval, prompts, workflow, and evaluation method.

A separate internal comparison addressed part of this limitation. Thomson Reuters gave competing models equal access to its content through a simpler research system. Overall scores ranged from 0.81 to 0.91. Thomson scored 0.89.

Thomson did not achieve the highest score when the models received equal access to Thomson Reuters content.

That result is consistent with the view that authoritative content contributed materially to the system-level advantage. But an internal comparison cannot isolate the exact contribution of the content, retrieval layer, model, prompts, citation handling, or workflow design.

The 0.89 result still matters. It shows that Thomson can remain competitive when the content advantage is shared, even if it does not always produce the best model-only result.

The enterprise question is therefore not simply, "Which model has the highest benchmark score?"

It is, "Which combination of models, information, tools, controls, and people produces the best business outcome?"

Why the First Deployment Makes Sense

Thomson's first deployment is in Tabular Analysis within CoCounsel Legal. The feature handles high-volume, structured document review, where performance can be assessed against clearer standards.

Thomson Reuters also says CoCounsel will remain multi-model by design, applying Thomson where it offers the clearest advantage and using other models elsewhere.

That is a practical business strategy.

An internally controlled model does not need to outperform every commercial system on every task. At high volumes, it can create value if it delivers acceptable quality with lower inference costs, lower latency, stronger data controls, or better workflow integration.

Thomson Reuters will still need to prove those benefits in customer deployments. Training cost does not predict the full operating cost, and a benchmark result does not guarantee customer value.

For professional workflows, the relevant measure is not simply cost per request. It is the quality-adjusted cost of completing useful work. That includes inference expense, processing time, human review, correction effort, escalations, and the consequences of errors.

A cheap answer that requires extensive review may be more expensive than a higher-cost answer that professionals can trust.

The multi-model approach also avoids a false build-versus-buy choice. Thomson Reuters can route suitable work to Thomson and continue using external models when they provide better quality, features, or economics.

Open Weight Does Not Mean Open Source

Thomson Reuters has released Thomson-1.0-Small for academic and other non-commercial uses. Its model card describes it as a 35-billion-parameter model with approximately 3 billion parameters active at a time, developed from Qwen3.6-35B-A3B through Snowdon-1.1-Small.

The model is distributed under the PolyForm Strict 1.0.0 license. The weights can be inspected and used within the license's restrictions, but the license does not grant unrestricted commercial rights.

Thomson-1.0-Small is therefore open-weight, not open source in the broader Open Source Initiative sense.

Publishing the weights supports external research and evaluation. It does not make the model freely available for every commercial purpose.

For companies considering an open-weight foundation, access to weights is only the beginning. They must also examine licensing, data rights, security, deployment requirements, operational costs, and long-term support.

What "Fully Owned" Means

Thomson Reuters says it fully owns and controls Thomson. In this context, ownership refers to the resulting model and the training the company performed. It does not mean Thomson Reuters owns the original Qwen foundation or created every layer from scratch.

AI sovereignty is not a binary condition. It has at least five layers:

  • Data sovereignty: Control over confidential, licensed, and proprietary information.
  • Model sovereignty: The ability to inspect, adapt, host, update, and govern the final weights.
  • Operational sovereignty: The ability to deploy the model without relying on one external model API.
  • Supply-chain sovereignty: Exposure to cloud providers, chips, GPUs, serving libraries, drivers, and data-center capacity.
  • Strategic sovereignty: The ability to replace a foundation model or supplier without rebuilding the entire product.

Thomson Reuters' technical report argues that its continual-learning methods can work across different open-weight foundations, model sizes, and professional fields. The company's How We Built Thomson article also says the team changed Thomson's root model several times as stronger open-weight foundations became available.

That does not eliminate external dependencies. Thomson Reuters still relies on foundation-model research, hardware, cloud or data-center capacity, and supporting software.

But those dependencies do not all have to become permanent.

Strategic sovereignty is stronger when a company can inspect, adapt, host, and eventually replace a component without rebuilding the product around it.

What Other Companies Should Learn

1. Start with the workflow

The best starting point is not the ambition to own a model. It is a valuable, repeatable, and measurable workflow supported by reliable data.

A specialized system for one economically important task can produce more value than a general assistant with no clear operating purpose.

2. Choose the right level of ownership

The choices extend beyond training from scratch or buying access to an external API.

A company can add retrieval to a commercial model, fine-tune an existing system, adapt an open-weight foundation, build its own evaluation layer, use multi-model routing, or combine internal and external models.

The right level of ownership depends on quality, privacy, usage volume, latency, licensing, technical talent, and cost.

3. Treat experts as infrastructure

Experts do more than test a finished product. They select data, define acceptable answers, identify subtle failures, create evaluation standards, and decide when human review is necessary.

Their time and judgment are part of the AI investment.

4. Evaluate the complete system

Model benchmarks are useful, but they do not capture the full business outcome.

Companies must test models with the content, retrieval tools, security controls, interfaces, workflows, and human-review processes that will exist in production. They must also distinguish company claims, internal evaluations, third-party reporting, and independent findings.

5. Compare total economics

The final training run is one line in a much larger budget.

A serious cost comparison must include data preparation, expert time, experimentation, infrastructure, security, product integration, monitoring, updates, human oversight, and the expected cost of unreliable outputs.

The key question is not whether an internally controlled model can match a commercial model on average.

It is whether greater ownership improves quality, control, privacy, speed, or operating economics enough to justify the continuing investment.

The Defensible Product Is the System

Thomson Reuters did not prove that every enterprise should build its own language model. It showed that capable open-weight foundations can reduce the cost and time required to begin building specialized AI systems, particularly for organizations that do not need to recreate general-purpose capabilities from scratch.

The base model still matters. It affects reasoning, language coverage, context handling, safety, tool use, and inference efficiency. But the model alone is rarely the finished product or the lasting source of enterprise advantage.

In professional AI, durable advantage may come from trusted proprietary data, expert judgment converted into training and evaluation signals, workflow integration, governance, distribution, and rigorous measurement where mistakes have real consequences.

Open-weight models lower the cost of entry. They do not create differentiated data, credible evaluations, trusted workflows, or customer confidence.

Those are the layers that determine whether an enterprise AI system becomes valuable, reliable, and defensible.

References