Thomson Reuters built its own AI model — what that means if you don't have $40M to spend
Thomson Reuters launched a proprietary LLM trained on its own legal and tax data. The lesson for smaller businesses is about owning domain data, not the $40M price tag.
What happened
Thomson Reuters launched Thomson, its first proprietary large language model, on August 24. The company spent $40 million on compute and talent to build it on an open-source foundation, then trained it on decades of Westlaw, Practical Law, Checkpoint and Reuters content — the same material that already sits behind its legal and tax products. According to the company's announcement, Thomson has so far used less than 10% of that content library, with hundreds of subject matter experts involved in evaluating its outputs.
The first deployment is narrow and specific: Thomson will power the tabular analysis feature inside CoCounsel Legal, a high-volume document review tool used for pulling structured data out of contracts and filings. CoCounsel stays a multi-model product — Thomson handles the tasks where a domain-trained model has an edge, and third-party frontier models handle everything else.
What is genuinely new
This isn't another company wrapping GPT-5.6 or Claude Opus 5 in a legal skin. Thomson Reuters built and owns the model. Its CTO, Joel Hron, framed the reasoning plainly: start with a strong open foundation, specialize deeply on the work that matters, and end up with something "far more efficient and entirely under your control."
That control claim has teeth. Thomson Reuters reports citation quality competitive with frontier models even on jurisdiction-specific questions — their public example was Canadian employment law — plus a bigger jump in following complex, multi-part professional instructions than general-purpose models show. Those are exactly the failure modes that make lawyers distrust AI-generated research: a citation that doesn't say what it claims to say, or an instruction that gets half-followed on step three of five.
Thomson Reuters is also releasing a smaller open-weight version on Hugging Face for academic and non-commercial use, and opening the model to outside legal and AI researchers for direct testing. A developer API is planned but not live yet.
What it means for a business owner
If you run a firm that licenses Westlaw, Practical Law, or Checkpoint, this shows up as a quality improvement inside tools you already pay for — not a new product you need to evaluate or a new vendor relationship to manage. That's worth knowing before a salesperson tries to upsell you on "the new AI."
The more useful signal is the strategic bet itself. A company sitting on decades of proprietary, structured domain data decided that owning a smaller specialized model beats renting a bigger general one, at least for the narrow slice of work where citation accuracy and instruction-following are the whole product. If your business has an equivalent asset — years of internal case files, claims history, service records, anything a general model has never seen — the same logic applies at a much smaller scale. You don't need $40 million to fine-tune or retrieval-ground a frontier model against your own data; you need to know that data is worth protecting and structuring in the first place.
The honest caveat
Thomson is not available to you unless you're already a Thomson Reuters customer, and even then only for one feature so far. The developer API that would let outside teams actually use the model doesn't exist yet — it's "planned." Early performance claims come from Thomson Reuters' own evaluations, not an independent benchmark, and "comparable to frontier models on a range of tasks" is a claim worth re-checking once outside researchers get real access. And this is a legal-and-tax story first: the lesson about proprietary domain data generalizes, but the model itself does not.
What to do about it
If you're evaluating whether a general-purpose model or a narrower, domain-grounded setup is the right foundation for something you're automating, don't default to "biggest model available." Look at what proprietary, structured data your business already holds, and ask whether grounding a smaller system in that data — through retrieval or fine-tuning, not necessarily a $40 million training run — would beat a bigger model working from public knowledge alone.
Want this kind of system in your business? Book a free scoping call.