Japan is preparing to test domestic large language models against Anthropic’s Claude and Amazon Nova inside GENAI, the government’s generative AI platform.
The Digital Agency will run three Japanese-developed models on SAKURA Cloud and compare their responses with foreign systems already available to government employees. The pilot will examine whether a domestic model-and-cloud stack can meet public-sector requirements for usefulness, reliability and cost-effectiveness.
An experimental environment is scheduled to be established by August, with several rounds of blind A/B testing planned from September through November 2026. The agency said the findings will help it consider procurement approaches for fiscal 2027 and later years.
Japan adds domestic models to GENAI
The evaluation is part of a larger GENAI pilot that began making the platform available to about 100,000 government employees on May 29, 2026. The agency plans to expand access across participating ministries and agencies toward approximately 180,000 employees.
The wider initiative involves foundation models from five Japanese private-sector entities. The SAKURA Cloud deployment covers three: NTT DATA’s tsuzumi 2, Fujitsu’s Takane 32B and Preferred Networks’ PLaMo 2.0 Prime.
Employees at the Digital Agency and several other ministries will see randomized responses without being told which model produced them. As of July 2026, GENAI also offered Amazon Nova Lite and several Anthropic Claude models.
The deployment is the government’s first operational use of SAKURA Cloud within its Government Cloud environment. The provider met the government’s technical requirements for production use on March 27, 2026. The Digital Agency identifies it as the only domestically developed service among Japan’s Government Cloud options.
Japan’s approach reflects a wider push to reduce dependence on foreign-controlled AI and expand control over cloud infrastructure. Similar sovereignty concerns are shaping policy in the United Kingdom and the European Union.
What the government trial can prove
Blind preference testing can identify which responses employees favor, but preference alone cannot establish which model is most accurate or economical. The Digital Agency also plans to evaluate reliability and cost-effectiveness, although it has not published the sample size, task mix, error-rate methodology or cost framework.
Because the three announced domestic models will run on SAKURA Cloud, published results will be more useful if they separate model performance from infrastructure performance. Otherwise, organizations may struggle to apply the findings to another cloud or deployment architecture, especially when production AI requires infrastructure upgrades.
The Digital Agency says Japanese-developed models may better handle Japanese vocabulary, expressions, culture and values. Evidence that those characteristics improve administrative accuracy or usefulness would carry more weight than preference scores alone.
The program also has an industrial-policy objective. Japan wants operational feedback to improve domestic models and government procurement to create steadier demand for local AI providers. The testing period overlaps with fiscal 2027 budget planning, but the agency has not explained how the findings will affect individual ministries’ purchases.
GENAI operates under government security controls. Within the Digital Agency, it supports Confidentiality Level 2 information and single sign-on through Government Solution Services.
Model-level scores, tested tasks, error rates and cost calculations will show whether domestic providers can earn a lasting place alongside foreign systems. Fiscal 2027 procurement decisions will provide the clearest evidence.
Read more: Singapore’s investment in an applied AI hub offers another view of how APAC governments are combining public-sector deployment with efforts to build regional AI capacity.


