Japan Tests Domestic LLMs Against Claude and Amazon Nova in Government AI Pilot

Japanese Digital Agency GENAI banner showing the government AI platform and its model-selection interface

Japan’s Digital Agency is testing domestic large language models through GENAI, its government generative AI environment. Image: GENAI

Written By
eWEEK Staff
eWEEK Staff
Aug 5, 2026
3 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

Japan is preparing to test domestic large language models against Anthropic’s Claude and Amazon Nova inside GENAI, the government’s generative AI platform.

The Digital Agency will run three Japanese-developed models on SAKURA Cloud and compare their responses with foreign systems already available to government employees. The pilot will examine whether a domestic model-and-cloud stack can meet public-sector requirements for usefulness, reliability and cost-effectiveness.

An experimental environment is scheduled to be established by August, with several rounds of blind A/B testing planned from September through November 2026. The agency said the findings will help it consider procurement approaches for fiscal 2027 and later years.

Japan adds domestic models to GENAI

The evaluation is part of a larger GENAI pilot that began making the platform available to about 100,000 government employees on May 29, 2026. The agency plans to expand access across participating ministries and agencies toward approximately 180,000 employees.

The wider initiative involves foundation models from five Japanese private-sector entities. The SAKURA Cloud deployment covers three: NTT DATA’s tsuzumi 2, Fujitsu’s Takane 32B and Preferred Networks’ PLaMo 2.0 Prime.

Employees at the Digital Agency and several other ministries will see randomized responses without being told which model produced them. As of July 2026, GENAI also offered Amazon Nova Lite and several Anthropic Claude models.

The deployment is the government’s first operational use of SAKURA Cloud within its Government Cloud environment. The provider met the government’s technical requirements for production use on March 27, 2026. The Digital Agency identifies it as the only domestically developed service among Japan’s Government Cloud options.

Japan’s approach reflects a wider push to reduce dependence on foreign-controlled AI and expand control over cloud infrastructure. Similar sovereignty concerns are shaping policy in the United Kingdom and the European Union.

Advertisement

What the government trial can prove

Blind preference testing can identify which responses employees favor, but preference alone cannot establish which model is most accurate or economical. The Digital Agency also plans to evaluate reliability and cost-effectiveness, although it has not published the sample size, task mix, error-rate methodology or cost framework.

Because the three announced domestic models will run on SAKURA Cloud, published results will be more useful if they separate model performance from infrastructure performance. Otherwise, organizations may struggle to apply the findings to another cloud or deployment architecture, especially when production AI requires infrastructure upgrades.

The Digital Agency says Japanese-developed models may better handle Japanese vocabulary, expressions, culture and values. Evidence that those characteristics improve administrative accuracy or usefulness would carry more weight than preference scores alone.

The program also has an industrial-policy objective. Japan wants operational feedback to improve domestic models and government procurement to create steadier demand for local AI providers. The testing period overlaps with fiscal 2027 budget planning, but the agency has not explained how the findings will affect individual ministries’ purchases.

GENAI operates under government security controls. Within the Digital Agency, it supports Confidentiality Level 2 information and single sign-on through Government Solution Services.

Model-level scores, tested tasks, error rates and cost calculations will show whether domestic providers can earn a lasting place alongside foreign systems. Fiscal 2027 procurement decisions will provide the clearest evidence.

Read more: Singapore’s investment in an applied AI hub offers another view of how APAC governments are combining public-sector deployment with efforts to build regional AI capacity.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.