Trend Shift

Thomson Shifts Legal AI Toward Data Ownership

Thomson Reuters has released its proprietary LLM Thomson, seeking to combine exclusive legal content, model ownership and lower inference costs.

Thomson Reuters has released Thomson, its first internally developed LLM that the company fully owns and controls. The move suggests that competition in legal AI is shifting from renting the strongest general-purpose model toward combining proprietary content, model control and inference economics in one system. The thesis is falsifiable: if independent tests cannot reproduce Thomson's domain advantage, ownership may amount mainly to an expensive supply-chain substitution.

The Advantage Combines Model and Content

The Decoder reports that Thomson is built on Alibaba's Qwen and will cost about $40 million over two years. Thomson Reuters confirms an investment of roughly $40 million, including talent and compute, and positions the model as a way to reduce the inference cost associated with typical frontier systems. The company previously emphasized a multi-model strategy; training and retrieval now combine the model with Westlaw, Practical Law, Checkpoint and Reuters content.

Internal Results Point to Data Leverage

In a legal deep-research evaluation built around 53 questions from internal experts, Thomson Reuters says Thomson scored 0.83 for factuality when connected to Westlaw and Practical Law, compared with 0.65 and 0.68 for leading frontier models given unrestricted web access. The metric measures whether claims are supported by citations, not whether an answer is fully correct. Even with that limitation, the result poses a testable business hypothesis: proprietary data can improve both domain output and task economics.

Closed Evaluation Is the Main Countercase

The strongest alternative explanation is that the vendor designed the evaluation and that the compared systems did not receive equivalent data or product architectures. A preregistered Journal of Empirical Legal Studies study found hallucination rates of 17% to 33% in the earlier Westlaw AI-Assisted Research and Ask Practical Law AI products. It did not test the new Thomson model, so it does not directly refute the launch results, but it shows that content ownership and citations do not automatically eliminate reliability risk.

What to watch next

The decisive evidence will be independent testing under common corpora, equal retrieval access and blinded review, along with disclosure of inference cost per completed task, latency and professional-user adoption. If Thomson retains a factuality advantage while lowering costs, owned domain models could become a structural moat for information services companies. If the advantage disappears under equal retrieval conditions, most of the value may still reside in the content library rather than the model.

Sources