Market PlaySeptember 12, 2026

The AI Moat Is Quietly Moving From the Model to the Data Licence

Six Chinese labs stand accused of copying US models at industrial scale. Two days later OpenAI shipped a product for banks whose value sits in licensed market data, the one thing distillation cannot lift.

By Race to AGI· AI-assisted analysis, grounded in Race to AGI data and reviewed before publishing

Six Chinese AI labs stand accused of copying US frontier models at industrial scale. Two days later, OpenAI shipped a product for banks whose selling point is not the model.

On September 8 the NSA, CISA and FBI published advisory AA26-251A, naming six China-based AI firms including DeepSeek, Moonshot AI and Alibaba, and accusing them of industrial-scale distillation of US frontier models. Reported volumes run to billions of tokens pulled from Claude, GPT, Gemini and Grok.

Take the accusation at face value for a moment and it says something uncomfortable about the product itself. If a rival can approximate your model by querying it enough times, the model is a copyable good. Expensive to build the first time, much cheaper to approach the second.

On September 10, OpenAI announced ChatGPT for Financial Services, a version of ChatGPT Work that ships with premium market and company data built in, aimed at banks and investment firms. The same day it added a Data plugin that lets ChatGPT Work and Codex connect directly to business data sources.

Read those two together and the commercial logic is hard to miss. Licensed market data cannot be distilled out of an API. A competitor can copy how a model reasons. A competitor cannot copy a contract with a data vendor.

The instinct is not new. What is new is how directly it now answers a security problem. Our deal records show the frontier labs buying rights steadily for the past year: Getty's image libraries into ChatGPT search and discovery in June, paid Wikimedia agreements for training and products in January, Disney's $1B equity investment alongside character licensing for Sora in December, and Meta's multi-year deals with news publishers the same month. Of the 644 deals we track, the licensing ones almost never lead the coverage. They probably should.

A few reasons to hold this loosely. The timing is most likely coincidence, because a vertical product for regulated finance takes months to build and cannot be assembled in 48 hours. The distillation claims are allegations in a government advisory, not findings in a court. And nobody outside the room knows what OpenAI pays its data partners, which is the number that decides whether this is a real moat or an expensive bundle.

The direction holds regardless of what caused it. When the capability layer converges, margin moves to whatever is contractually scarce. In finance that is market data. In law it will be case history and filings. In healthcare it will be claims and outcomes data, and it will be the hardest of the three to obtain.

There is a second-order effect worth watching. A lab that bundles licensed data begins competing with the vendors it licenses from. S&P Global and Morgan Stanley both appear in the coverage of the OpenAI launch. Data vendors have spent two years working out whether frontier labs are customers or rivals, and products like this one answer the question for them.

It also complicates the enforcement idea we wrote about yesterday. If the valuable part of the product keeps migrating into licensed data, then quietly degrading a suspect account stops being much of a lever, because the part worth taking was never in the weights.

What to do with this

If you are buying AI for a regulated function, split the invoice in your head. Ask the vendor what you are paying for the model and what you are paying for the data, then price that same data directly from its source. If the bundle comes in cheaper, you are looking at a real moat. If it does not, you are paying for convenience, and that is a renegotiation.

If you are trying to see where this industry will actually make money, put the benchmark scores down for a month and read licensing announcements instead. The next vertical data bundle, whichever lab ships it and in whichever industry, will tell you more about the shape of 2027 revenue than any eval will.

#data-licensing#openai#distillation#enterprise-ai#moat