On September 1, 2026, Dentsu, Dentsu Digital and SB Intuitions announced they had built a dataset of over 116,000 human evaluations of Japanese advertising copy. The work supports a generative AI model specialized in Japanese copywriting, using LLM-as-a-judge methods and SoftBank-backed Sarashina models.
This article aggregates reporting from 3 news sources. The TL;DR is AI-generated from original reporting. Race to AGI's analysis provides editorial context on implications for AGI development.
This announcement is a good example of how non-English ecosystems are building their own deep data assets rather than relying on generic frontier models. By having 64 professional copywriters label more than 116,000 Japanese ad lines, Dentsu and SB Intuitions are effectively creating a high-quality reward model for ‘what good Japanese copy looks like.’ That gives them a unique lever to steer generative models toward culturally and linguistically nuanced output.
In the race to AGI, these kinds of domain and language specific datasets are a powerful counterweight to the idea that one monolithic English-first model will dominate everything. Labs that can combine frontier capabilities with very targeted preference data will have an edge in markets like Japan where phrasing, politeness and brand tone matter as much as raw fluency.
There is also an important governance angle: if ad holding companies control the labels that define ‘good’ persuasive language in a market, they hold a subtle but real power over how AI shapes consumer perception. That will matter as generative systems move from drafting slogans to running A/B tests and optimizing campaigns autonomously.