Trend Shift
Thomson Shifts Legal AI Toward Data Ownership
Thomson Reuters has released its proprietary LLM Thomson, seeking to combine exclusive legal content, model ownership and lower inference costs.
Thomson Reuters has released Thomson, its first internally developed LLM that the company fully owns and controls. The move suggests that competition in legal AI is shifting from renting the strongest general-purpose model toward combining proprietary content, model control and inference economics in one system. The thesis is falsifiable: if independent tests cannot reproduce Thomson's domain advantage, ownership may amount mainly to an expensive supply-chain substitution.
The Advantage Combines Model and Content
The Decoder reports that Thomson is built on Alibaba's Qwen and will cost about $40 million over two years. Thomson Reuters confirms an investment of roughly $40 million, including talent and compute, and positions the model as a way to reduce the inference cost associated with typical frontier systems. The company previously emphasized a multi-model strategy; training and retrieval now combine the model with Westlaw, Practical Law, Checkpoint and Reuters content.
Internal Results Point to Data Leverage
In a legal deep-research evaluation built around 53 questions from internal experts, Thomson Reuters says Thomson scored 0.83 for factuality when connected to Westlaw and Practical Law, compared with 0.65 and 0.68 for leading frontier models given unrestricted web access. The metric measures whether claims are supported by citations, not whether an answer is fully correct. Even with that limitation, the result poses a testable business hypothesis: proprietary data can improve both domain output and task economics.
Closed Evaluation Is the Main Countercase
The strongest alternative explanation is that the vendor designed the evaluation and that the compared systems did not receive equivalent data or product architectures. A preregistered Journal of Empirical Legal Studies study found hallucination rates of 17% to 33% in the earlier Westlaw AI-Assisted Research and Ask Practical Law AI products. It did not test the new Thomson model, so it does not directly refute the launch results, but it shows that content ownership and citations do not automatically eliminate reliability risk.
What to watch next
The decisive evidence will be independent testing under common corpora, equal retrieval access and blinded review, along with disclosure of inference cost per completed task, latency and professional-user adoption. If Thomson retains a factuality advantage while lowering costs, owned domain models could become a structural moat for information services companies. If the advantage disappears under equal retrieval conditions, most of the value may still reside in the content library rather than the model.
Sources
- The Decoder AI — Thomson Reuters bets $40M on owning its AI instead of renting from OpenAI or Anthropic
- Thomson Reuters — Thomson Reuters Leverages its World-Class Data Assets to Launch Its Own Frontier Model
- Thomson Reuters Institute — Thomson Reuters Built Its Own AI Model That Now Ranks Among the World’s Best
- Thomson Reuters Institute — How we built Thomson
- Journal of Empirical Legal Studies — Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
- Thomson Reuters — Thomson-1.0-Small
- Thomson Reuters — Thomson: a purpose-built foundation model for professionals