No Major AI Lab Tops C+ in 2026 AI Safety Index

2026 AI Safety Index scorecard showing Anthropic leading with a C+ as the highest grade among major AI labs

The 2026 AI Safety Index found that no major AI lab earned better than a C+, raising fresh questions about vendor governance and safety controls. Image: Screenshot from Future of Life Institute

Written By
eWEEK Staff
eWEEK Staff
Aug 9, 2026
3 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

None of the nine companies evaluated in the Future of Life Institute’s Summer 2026 AI Safety Index earned better than a C+. Anthropic ranked first with a 2.66 score, followed by OpenAI at 2.28 and Google DeepMind at 2.01. xAI, DeepSeek, and Mistral received F grades.

The scores are not certifications of how safely individual models perform in production. They combine technical risk assessment with governance, transparency, safety frameworks, documented harms, and preparations for severe AI failures, making the index more useful as a vendor-governance benchmark than a product ranking.

How the index builds its grades

The Summer 2026 AI Safety Index evaluates nine companies across 37 indicators in six domains: Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing. Its evidence window closed June 3, 2026, so later policy or technical changes did not affect the grades.

Seven technical and governance experts reviewed public materials, including model cards, research, benchmark results, and policy documents, plus survey responses from participating companies. Final grades are averaged expert assessments, not results from an independent production audit.

Meanwhile, enterprise AI governance gaps persist around visibility, ownership, traceability, and incident response. A strong disclosure or governance score does not show how a model will behave or resist attacks in a specific deployment.

Anthropic led five of the six domains, including Information Sharing at B+ and Governance & Accountability at B. OpenAI led Risk Assessment, reflecting its wider set of evaluations and participation in outside testing.

Weak safety scores raise vendor-risk questions

Existential Safety was the lowest-scoring domain. Anthropic and OpenAI received D+ grades, Google DeepMind received a D, and the other six companies received F grades. The category assesses work aimed at preventing loss of control over increasingly capable AI systems, not conventional enterprise cybersecurity controls.

The review panel also concluded that Anthropic, OpenAI, Google DeepMind, and Meta had weakened or removed earlier commitments to pause development unilaterally if specified risk thresholds were approached. In some cases, newer policies make action dependent on competitors’ behavior.

Advertisement

Operational safeguards are a separate concern. Google DeepMind recently outlined controls for advanced AI agents, including monitoring, access restrictions, and blocking mechanisms for systems connected to internal infrastructure.

Military AI was another focus. Reviewers treated expanding defense relationships as a developing risk under the Current Harms category after several developers moved away from earlier broad restrictions on military applications. That is the panel’s assessment of those policy shifts, not evidence that military use alone makes a model unsafe.

An April 2026 UC Berkeley risk-management assessment separately evaluated Anthropic, OpenAI, Google, and Meta against high-priority practices for general-purpose AI. It found varying levels of alignment and called for stronger documentation and controls in areas including risk assessment and post-deployment management.

Regulators are also pushing model evaluation into enterprise risk planning, increasing pressure on organizations to scrutinize capabilities, supplier dependencies, testing, and access to critical systems.

Procurement and security teams should request current evaluation results, independent testing where available, incident-response procedures, post-deployment monitoring, and explanations for material changes to safety commitments. The C+ ceiling exposes a governance gap, not a simple ranking of safe and unsafe vendors. Enterprises still have to determine whether a model’s controls, monitoring, contracts, and compliance posture match the risks of the workload they plan to put into production.

Read more: Recent testing showed how containment can fail when OpenAI models escaped a restricted cyber evaluation, reinforcing the need to test both model capabilities and the security of the evaluation environment.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.