AI Safety Index 2026: What the Grades Miss
Executive Summary
The Future of Life Institute graded nine frontier AI companies in July 2026. The best grade in the industry was a C+.
That was Anthropic, at 2.66 on a 4.3-point scale, leading five of six domains. OpenAI took a C at 2.28, Google DeepMind a C at 2.01. Meta scored D+. Z.ai and Alibaba Cloud landed at D-. Three companies failed outright: xAI at 0.65, DeepSeek at 0.47, and Mistral at 0.33 — dead last, from the jurisdiction with the world's most developed AI safety regulation.
Those numbers are being read, in a lot of procurement conversations, as a vendor shortlist. They are not one, and the gap matters: the Index measures how seriously a company takes AI safety research. It does not measure whether you can deploy that company's model and defend the decision to an auditor. Those are different questions with different evidence, and a strong answer to the first tells you very little about the second.
This piece covers what the Index measured, the three findings that should change how you write vendor contracts, and the question the grades cannot answer for you.
What was measured
Nine companies, 37 indicators, six domains, graded on the US GPA scale where A+ is 4.3 and F is 0. Scoring was reviewed by an independent panel of seven: David Krueger, Sharon Li, Tegan Maharaj, Sneha Revanur, Stuart Russell, Robert Trager, and Yi Zeng.
| Company | Overall | Score |
|---|---|---|
| Anthropic | C+ | 2.66 |
| OpenAI | C | 2.28 |
| Google DeepMind | C | 2.01 |
| Meta | D+ | 1.32 |
| Z.ai | D- | 0.88 |
| Alibaba Cloud | D- | 0.87 |
| xAI | F | 0.65 |
| DeepSeek | F | 0.47 |
| Mistral | F | 0.33 |
The domain-level grid is where the detail lives:
| Domain | Indicators | Best grade awarded |
|---|---|---|
| Risk Assessment | 6 | C+ |
| Current Harms | 9 | B- |
| Safety Frameworks | 4 | B- |
| Existential Safety | 4 | D+ |
| Governance & Accountability | 4 | B |
| Information Sharing | 10 | B+ |
Read that table by column and you get a ranking. Read it by row and you get something more useful: a map of which safety questions the industry has answers to, and which it does not.
Information Sharing tops out at B+. Governance at B. These are the domains where companies publish, document, and disclose — where doing the work and showing the work overlap.
Existential Safety tops out at D+. Anthropic and OpenAI earned that D+. Google DeepMind took a D. The remaining six companies all received an F.
No company in the industry has a passing answer for the hardest safety question in it. That is not a ranking problem, and it does not resolve by picking a different vendor.
Finding one: commitments are not controls
Buried in the key findings is a sentence with direct commercial consequences:
Even industry leaders in safety practices are retreating from prior commitments. Anthropic, OpenAI, Google DeepMind, and Meta have weakened or voided pledges to pause unilaterally if redlines are approached, some citing competitor-contingent conditions.
Read that again with a contracts hat on. Four of the most safety-forward companies in the industry made public commitments, then weakened or voided them — in some cases explicitly because of what competitors were doing.
The lesson is not that these are bad companies. Anthropic still earned the highest grade in the Index, and FLI's recommendation to them is specific: reverse the RSP 3.0 walk-back on pause commitments and restore the credibility of those commitments. The lesson is structural, and every compliance function already knows it in other contexts:
A public safety commitment is a statement of present intent, not a contractual control. It can be revised when competitive conditions change, and in 2026 it was.
If your AI risk assessment cites a vendor's published safety pledges as a mitigating control, you are relying on something the vendor can withdraw unilaterally — and per FLI's own findings, several already have. Controls you can rely on are the ones written into your agreement: your BAA, your data processing addendum, your retention terms, your audit rights. Those survive a change in competitive strategy. A blog post about redlines does not.
Finding two: the deployment context is not in the grade
The reviewers flagged the industry's pivot to military AI as an emerging current-harm risk. Between 2024 and 2026, companies including Anthropic, OpenAI, Google DeepMind, and Meta that had previously banned military applications reversed course, joining xAI and Mistral in actively pursuing defence partnerships. The report notes xAI's commitment to deploying military AI "without ideological constraints," and that defence work accounts for 10–15% of Mistral's revenue, with active contracts for the French, Singaporean, and Luxembourg armed forces.
The panel weighed this explicitly for only one company — the one at the top of the table.
FLI criticises Anthropic for "questionable military engagements," citing its Pentagon work and a reported link to the Minab school strike. The report's own account is carefully hedged, and it is worth reproducing at that hedge rather than tightening it:
Washington Post reporting has suggested that Anthropic's Claude LLM, integrated into Maven Smart System in collaboration with Palantir, may have been involved in targeting the Shajarah Tayyebeh elementary school in Minab, Iran on Feb 28, hitting it at least twice and killing 175–180 people, mostly girls aged 7–12.
FLI adds that Anthropic has neither publicly confirmed nor denied involvement, that the U.S. military investigation is ongoing, and that controversy remains about causes and responsibility. Reported death tolls vary between sources. Internal U.S. investigation findings have been characterised as attributing the strike to outdated data that misidentified the school, framed as human error rather than a failure of the deployed system — a characterisation that is itself contested.
We are not in a position to resolve any of that, and this piece does not try to.
What is documented and uncontested is the surrounding sequence. On 26 February 2026, Anthropic CEO Dario Amodei published a statement on the company's discussions with the Department of War, refusing the Pentagon's demand for an "any lawful use" contract and insisting on two safeguards: no mass domestic surveillance, and no fully autonomous weapons. A Pentagon dispute followed. Anthropic held red lines that several competitors did not, and drew criticism from the review panel anyway.
For a procurement audience, the transferable point has nothing to do with assigning blame for Minab. It is this:
The model was one component in someone else's system. Claude, integrated into Maven, in collaboration with Palantir, operating on data of contested currency, inside a targeting process run by a third party. Whatever the eventual finding, the risk did not live in the model card. It lived in the integration, the data feeding it, and the decision procedure wrapped around it.
That is the same structure as every enterprise AI deployment, minus the stakes. Your vendor's safety grade describes the component. Your exposure comes from the system you build around it — what you connect it to, what data you feed it, what authority you grant its output, and who is accountable when it is wrong. No index grades that, because no index can see it.
It is also the clearest possible demonstration that a top grade is not a clean bill of health. The best-graded company in the Index is the one the panel singled out for criticism. Both facts are in the same report.
Finding three: safety is not regional
Three companies failed: xAI in the US, DeepSeek in China, Mistral in Europe. FLI draws the obvious conclusion — inadequate safety is a global problem, not a regional one.
The European result deserves attention from anyone who assumes regulatory jurisdiction is a proxy for vendor safety posture. The EU leads the world in AI safety regulation. The top European AI company scored last of nine.
Operating in a well-regulated jurisdiction does not make a vendor safe, and it does not discharge your obligations as a deployer. The EU AI Act assigns duties to deployers directly. Buying European does not transfer them to someone else.
What the grades cannot tell you
Here is the practical problem with using the Index as a procurement shortlist.
Suppose you take the ranking at face value and select the top-graded vendor. You have now chosen a company with a C+ in safety research. You still do not know:
- whether they will sign a BAA for your use case
- whether your prompts and outputs train their models by default, and what the opt-out actually covers
- how long your data is retained, and in which jurisdictions
- which subprocessors touch it
- what their obligations are under the EU AI Act, and which of them transfer to you as a deployer
- what happens to any of the above when you access the same model through a cloud reseller
None of these are in the Index, and that is not a criticism of the Index. FLI set out to evaluate safety practices at frontier labs and did so with more rigour than anyone else publishing in the space. The questions above are a different assessment, and nobody is doing it for you.
The distinction sharpens once you notice that the safety grade travels with the model, while the compliance terms travel with the platform. You can run a top-graded developer's model through a cloud provider FLI does not grade at all, under that provider's commercial terms and that provider's BAA. The safety research behind the model and the contractual posture around the deployment come from two different companies. One row in the Index cannot express that, because the Index is not about deployments.
This is why "which vendor scored highest" is the wrong opening question. The useful one is narrower: for this workload, under this regulation, what can I get in writing?
Using the Index well
Use it as a trend instrument. The Index publishes editions — this is Summer 2026, following Winter 2025 — and movement is more informative than any single grade. Meta improved from 6th to 4th. xAI dropped from 4th to 7th. Direction of travel across editions tells you something a snapshot cannot.
Use it to write better vendor questions. The 37 indicators are a due-diligence checklist someone else has already built and had reviewed by seven independent experts. You need not adopt the scoring to borrow the questions.
Use it to calibrate expectations internally. When a business unit asks why AI procurement takes more scrutiny than a typical SaaS purchase, "the best safety grade in the entire industry is a C+, and the best existential-safety grade is a D+" lands harder than a policy citation.
Do not use it as a substitute for diligence. A C+ is not an approval, an F is not a prohibition, and neither is evidence about your regulatory posture. Mistral scoring last does not mean a Mistral model cannot be deployed compliantly. Anthropic scoring first does not mean an Anthropic deployment is compliant by default.
Questions worth asking your AI vendor
None of these are answered by a safety grade. All of them are answerable in writing, and the answers belong in your file.
- Will you sign a BAA covering this specific use case, or only parts of your platform?
- Are our inputs and outputs used for training by default? What exactly does the opt-out cover, and does it apply retroactively?
- What is the retention period for prompts, outputs, and logs — and in which jurisdictions are they held?
- Which subprocessors process our data, and how are we notified when that list changes?
- Under the EU AI Act, what is your classification, and which deployer obligations do you consider ours?
- If we access your model through a cloud provider, whose terms govern — and which of the above answers change?
- What contractual notice do we get if your safety framework or acceptable-use policy materially changes?
Question seven exists because of finding one. In 2026, that is not a hypothetical.
The through-line
FLI did the industry a service by grading something nobody else was grading and publishing the methodology alongside the results. Nine companies, 37 indicators, an independent review panel, and a document that shows its work.
The grades are also a ceiling rather than a floor: no passing grade in Existential Safety anywhere in the field, safety leaders walking back commitments made two years earlier, and the top-ranked company criticised in the same report for how its model was used downstream. If you deploy this technology in a regulated environment, all three belong in your risk register.
None of them answers the question your auditor will actually ask, which is not how safe is this company but what did you verify, when, and where is it written down.
The Summer 2026 AI Safety Index is published in full at futureoflife.org, including per-indicator grading sheets for every company. It is worth reading directly rather than through anyone's summary, including this one.
About This Resource
Need Expert Guidance?
Our team can help you put these insights into practice.
Schedule a Consultationor call (415) 644-8208Ready to Take the Next Step?
Our consultants understand your compliance requirements and can help you build a practical AI strategy.
