None of the nine companies evaluated in the Future of Life Institute’s Summer 2026 AI Safety Index earned better than a C+. Anthropic ranked first with a 2.66 score, followed by OpenAI at 2.28 and Google DeepMind at 2.01. xAI, DeepSeek, and Mistral received F grades.
The scores are not certifications of how safely individual models perform in production. They combine technical risk assessment with governance, transparency, safety frameworks, documented harms, and preparations for severe AI failures, making the index more useful as a vendor-governance benchmark than a product ranking.
How the index builds its grades
The Summer 2026 AI Safety Index evaluates nine companies across 37 indicators in six domains: Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing. Its evidence window closed June 3, 2026, so later policy or technical changes did not affect the grades.
Seven technical and governance experts reviewed public materials, including model cards, research, benchmark results, and policy documents, plus survey responses from participating companies. Final grades are averaged expert assessments, not results from an independent production audit.
Meanwhile, enterprise AI governance gaps persist around visibility, ownership, traceability, and incident response. A strong disclosure or governance score does not show how a model will behave or resist attacks in a specific deployment.
Anthropic led five of the six domains, including Information Sharing at B+ and Governance & Accountability at B. OpenAI led Risk Assessment, reflecting its wider set of evaluations and participation in outside testing.
Weak safety scores raise vendor-risk questions
Existential Safety was the lowest-scoring domain. Anthropic and OpenAI received D+ grades, Google DeepMind received a D, and the other six companies received F grades. The category assesses work aimed at preventing loss of control over increasingly capable AI systems, not conventional enterprise cybersecurity controls.
The review panel also concluded that Anthropic, OpenAI, Google DeepMind, and Meta had weakened or removed earlier commitments to pause development unilaterally if specified risk thresholds were approached. In some cases, newer policies make action dependent on competitors’ behavior.
Operational safeguards are a separate concern. Google DeepMind recently outlined controls for advanced AI agents, including monitoring, access restrictions, and blocking mechanisms for systems connected to internal infrastructure.
Military AI was another focus. Reviewers treated expanding defense relationships as a developing risk under the Current Harms category after several developers moved away from earlier broad restrictions on military applications. That is the panel’s assessment of those policy shifts, not evidence that military use alone makes a model unsafe.
An April 2026 UC Berkeley risk-management assessment separately evaluated Anthropic, OpenAI, Google, and Meta against high-priority practices for general-purpose AI. It found varying levels of alignment and called for stronger documentation and controls in areas including risk assessment and post-deployment management.
Regulators are also pushing model evaluation into enterprise risk planning, increasing pressure on organizations to scrutinize capabilities, supplier dependencies, testing, and access to critical systems.
Procurement and security teams should request current evaluation results, independent testing where available, incident-response procedures, post-deployment monitoring, and explanations for material changes to safety commitments. The C+ ceiling exposes a governance gap, not a simple ranking of safe and unsafe vendors. Enterprises still have to determine whether a model’s controls, monitoring, contracts, and compliance posture match the risks of the workload they plan to put into production.
Read more: Recent testing showed how containment can fail when OpenAI models escaped a restricted cyber evaluation, reinforcing the need to test both model capabilities and the security of the evaluation environment.


