Back to blog

Costs, Compute and the Quest for AI Sovereignty

Originally published on Substack · The Forward Curve.

Open weights are helping firms turn institutional judgment into models they can control. Every claim of ownership still carries a supplier underneath it.

A customer browses a record shop where AI models appear as competing vinyl album releases.
Each model launch makes the catalog deeper and the records cheaper.

An OpenAI-backed legal software company just placed a Chinese open-weight model underneath its first in-house legal model.

This month, Harvey introduced Tenet, a version of Moonshot AI’s Kimi K3 that Harvey and Fireworks post-trained for long, document-heavy legal work. Four days later, Thomson Reuters launched Thomson, a proprietary model built from an open-weight foundation after a $40 million investment in talent and compute.

The obvious explanations are cost and data security. Both appear in the company announcements. But a larger ambition sits underneath the move to build on top of open-weight models: these firms want to own the process that turns professional judgment into machine behavior.

Open weights are making the foundation model a replaceable input. The compounding asset is the loop connecting expert work, production traces, evaluations, and distribution.

Production traffic has started to catch up with the strategy. Vercel’s July AI Gateway index found that open-weight models handled 29% of June tokens, up from 11% in April, while generating less than 4% of normalized spend.

The curve accelerated again in August. On August 22, Vercel CEO Guillermo Rauch reported that open weights handled a record 62% of gateway tokens that day, compared with 28.4% on June 24. Rauch’s figures are daily readings. The monthly index remains the cleaner trend line; the August snapshot shows how quickly application traffic can move once gateways and developer tools become model-agnostic.

Guillermo Rauch reports open-weight models at 62 percent of Vercel AI Gateway tokens on August 22, compared with 28.4 percent on June 24.
Open weights hit 62% of Vercel AI Gateway tokens on August 22, up from 28.4% two months earlier. Daily reading, not monthly trend. Source: Guillermo Rauch, Vercel.

Open weights are gaining the volume. Frontier models are keeping the expensive work.

That split changes where value can accrue. Vertical software companies, data incumbents, and professional-services firms can pull more of the intelligence layer into their own economics while model labs and clouds supply the foundation.

The wrapper has started training the model

For the past few years, enterprise AI followed a familiar structure. A company called a frontier model through an API, added its documents and workflow, then charged the customer a software margin. Every query carried a supplier invoice, and every base-model improvement arrived on the provider’s schedule.

Harvey’s research shows how that boundary is moving. Tenet begins with Kimi K3, then trains inside simulated legal matters containing instructions, client files, tools, and attorney-authored rubrics. The agent searches the matter and drafts work product. A judge grades the result against the rubric, and reinforcement learning pushes the model toward work that satisfies more legal criteria with fewer tokens.

Harvey reports that Tenet completes almost twice as many held-out Legal Agent Benchmark tasks as the Kimi K3 base and 20% more of its contracts benchmark. A separate firm-knowledge model cut cost per query by 90%. These are company-authored results, with several tests run in Harvey’s own harness, so we’ll have to see how the production record compares against research.

Harvey's company-reported Tenet research benchmark results compared with base and frontier models.
Performance across legal benchmarks. Source: Harvey Tenet Research Preview. Link.

But the mechanism is clear. Harvey can train the model on the mistakes, review standards, and tool use that define its product. A general model provider sees tokens. Harvey sees which answer failed a partner’s rubric and which search path wasted 80,000 of them.

That is an objectively better dataset.


Cost is becoming an architectural choice

AI spending becomes harder to ignore once a pilot turns into a repeated workflow. McKinsey’s May 2026 Enterprise AI FinOps survey found that 62% of qualified respondents had moved beyond experimentation, while 93% had exceeded their AI budgets.

Ramp’s August AI Index shows buyers applying that discipline inside a single provider. Anthropic’s Fable 5 supplied 6% of tokens and 11.4% of model-attributed spending in Ramp’s July sample. The Financial Times later reported that the cheaper Opus 5 had overtaken it in business spending. Ramp’s model-level users skew technical, and cost is one possible driver among product fit, speed, and data-retention policy. The premium model still has to earn its spread on the task.

At scale, the useful unit is cost per accepted task. Token prices capture one part of it. A cheaper model that triggers longer traces, more retries, or extra human review can erase the saving. Harvey trained Tenet to prefer shorter trajectories at equivalent performance, attacking both price per token and tokens per completed legal task.

Consider 10 million model tasks a year at $1 per accepted result. A specialist that lowers the all-in figure to 35 cents creates $6.5 million of annual room for training and operations. Quality failures, migrations, and infrastructure costs can consume that room quickly.

Open weights therefore expand the set of choices. A company can host a model, modify it, reserve an expensive frontier API for difficult cases, and route routine work toward a smaller specialist. The architecture starts to look like a assembly line: each job goes to the venue offering the best mix of quality, speed, control, and cost.

The archive is raw material

Thomson Reuters makes the strategy easier to see because it already owns the ingredients most AI startups spend years trying to assemble. The company has Westlaw, Practical Law, Checkpoint, Reuters, roughly 1,500 attorney-editors, and distribution into the daily work of lawyers, tax professionals, and accountants.

Thomson starts from Snowdon, an open-weight model developed at Imperial College London’s FAIR Lab. The company has changed the root model several times and expects to keep doing so. Less than 10% of its content has entered continued pretraining.

The more revealing work happens around the archive. Experienced lawyers write evaluation rubrics for difficult research tasks. Other experts compare outputs and encode the difference between a legally relevant clause and the clause a client will actually care about. Internal users surface failures that generic benchmarks rarely capture.

A document repository can be licensed. The judgment that narrows a proposition, rejects an authority, or decides which issue deserves a partner’s attention is harder to copy. Thomson Reuters is trying to turn that intermediate work into training data while its products keep generating the next round of examples.

CLA and Digits are building a similar loop in accounting. CLA supplies the accounting knowledge and client workflow; Digits supplies the platform and model-training technology. CLA says every transaction its professionals and clients review, categorize, or correct will sharpen the firm’s model as it rolls out to thousands of clients over three years.

Security helps these companies protect the data. Ownership of the feedback loop helps them compound it.

Sovereignty comes in layers

Vercel AI Gateway June token-volume shares by model lab, with DeepSeek at 22.6 percent.
DeepSeek reached 22.6% of gateway tokens in June, less than two points behind Google. Source: Vercel AI Gateway Production Index, July 2026.

The phrase “our own model” compresses several ownership claims into three words. A firm can control where the model runs, the data it sees, the post-trained weights, the evaluation suite, the routing logic, and the schedule for upgrades. Each layer removes one constraint while leaving another supplier in place.

Harvey owns Tenet’s legal adaptation, training environments, and product integration. Kimi K3 supplies the foundation, and The Next Web has raised questions about its Chinese provenance and licensing chain. Thomson Reuters can replace Thomson’s open-weight root, while partners supply compute and infrastructure. CLA’s firm model learns inside a platform supplied by Digits.

Mistral Forge offers enterprises pretraining, post-training, reinforcement learning, internal deployment, and continuous evaluation. Amazon Nova Forge lets customers blend proprietary data with Amazon’s training data and connect reward functions to their own tools. The customer controls the adapted model while AWS retains the hosting relationship.

The “owned model” label is analogous to a cap table that way. Follow every layer until the remaining counterparties become visible.

Portability is part of the value. In May, OpenAI said it was winding down its existing fine-tuning platform. Existing fine-tuned models remain available until their base models are deprecated. A customer can invest in training data and validation, then inherit a migration date set by the provider.

Open weights give firms more influence over that clock. They also transfer responsibility for evaluation, security, infrastructure, upgrades, and failure recovery back toward the firm.

The market is splitting around the learning loop

Frontier labs will continue to own the expensive edge of general capability, where a few extra points of performance carry large economic value. Routine professional work creates a different contest because volume, latency, review cost, and domain reliability can outweigh the last increment of general intelligence.

Vercel’s spend data captures the split. Anthropic took 61% of June gateway spend on 32% of tokens and at least 72% of spend across coding agents, back-office agents, and application generation. The Wall Street Journal’s reported second-quarter figures point in the same direction: OpenAI revenue reached $6.7 billion in Q2’26, up 18% QoQ, while Anthropic reached $11.6 billion, more than doubling at a 142% increase QoQ.

Vertical software companies can protect margin by routing work across proprietary and external models. Data incumbents can convert archives and editorial standards into training assets. Professional-services firms can embed years of corrections into software that reaches every client. Model labs and clouds can sell the training, inference, and governance infrastructure underneath all three.

The countercase is expensive and real. A specialist can freeze today’s workflow into weights while frontier providers improve. Domain training can erode general capabilities. Internal benchmarks can reward the firm’s own assumptions, and model operations bring security and reliability work back in-house. Thomson Reuters and Harvey preserve multi-model strategies for a reason.

The likely enterprise architecture is therefore a portfolio. Some tasks will use frontier APIs, others will run on post-trained open models, and routing will change as price and performance move. The firms with the strongest position will own the evaluations that make those substitutions possible and the feedback generated after each task enters production.

AI sovereignty will arrive as partial ownership across a shifting stack. A company may rent the compute, inherit the base, own the post-training, control the evals, and capture the workflow. The strategic advantage comes from knowing which layers deserve capital and keeping enough of the system portable to change the answer.


Sources: