The FLI AI Safety Index
How the Future of Life Institute grades 9 frontier AI companies across 37 indicators in six domains — and, more usefully for anyone buying AI, what it deliberately does not measure.
How it works
The Index grades 9 frontier AI companies against 37 indicators grouped into six domains, on the US GPA scale where A+ is 4.3 and F is 0. Grades are published per edition; the current one is Summer 2026.
Scoring is reviewed by an independent panel rather than produced by FLI alone. The Summer 2026 panel was David Krueger, Sharon Li, Tegan Maharaj, Sneha Revanur, Stuart Russell, Robert Trager and Yi Zeng.
Grades belong to the edition that produced them. FLI's methodology changes between editions, so an unlabelled grade becomes wrong rather than merely stale — which is why every grade in our tracker is stored against its edition.
The six domains
- Risk Assessment
- Whether the company systematically identifies and evaluates the risks its models pose before and after release.
- Current Harms
- Measurable harm happening now — model behaviour, downstream misuse, and the company’s response to it.
- Safety Frameworks
- The published policies governing what the company will and will not do as capability increases, and whether those commitments are specific enough to be broken visibly.
- Existential Safety
- Preparation for risks from systems substantially more capable than today’s. No company scored above D+ in this domain.
- Governance & Accountability
- Corporate structure, internal escalation paths, and whether anyone inside the company can halt a launch.
- Information Sharing
- What the company discloses to researchers, regulators and the public — including what it declines to disclose.
The Summer 2026 grades in full
All 9 graded companies, as published. Our tracker covers 5 of them plus 2 deployment platforms that FLI does not grade at all.
| Company | Overall | Score | Existential Safety |
|---|---|---|---|
| Anthropic | C+ | 2.66 | D+ |
| OpenAI | C | 2.28 | D+ |
| Google DeepMind | C | 2.01 | D |
| Meta | D+ | 1.32 | F |
| Z.ai | D- | 0.88 | F |
| Alibaba Cloud | D- | 0.87 | F |
| xAI | F | 0.65 | F |
| DeepSeek | F | 0.47 | F |
| Mistral | F | 0.33 | F |
Source: Future of Life Institute AI Safety Index, Summer 2026. Verified 2026-08-07.
What the Index does not measure
Our analysis. FLI is explicit that the Index assesses safety practices, not deployability; the procurement consequences below are ours to draw.
Suppose you take the ranking at face value and select the top-graded vendor. You have chosen a company with a C+ in safety. You still do not know whether it will sign a business associate agreement, which of its features that agreement covers, what it retains and for how long, whether your inputs train its models, who its sub-processors are, or where it stands under the EU AI Act.
The safety grade travels with the model. The compliance terms travel with the platform.
You can run a top-graded developer's model through a cloud provider FLI does not grade at all, under that provider's terms and that provider's BAA. The safety research behind the model and the contractual posture around the deployment come from two different companies. One row in an index cannot express that, because the Index is not about deployments.
This is why "which vendor scored highest" is the wrong opening question. The useful one is narrower: for this workload, under this regulation, what can I get in writing?
Where to take that question
Our vendor assurance tracker answers it for 7 vendors — BAA availability, certifications, training posture, retention, sub-processor transparency and EU AI Act status, each with a primary source and a verification date.
Open the vendor assurance trackerCommon questions
- What is the FLI AI Safety Index?
- An independent assessment by the Future of Life Institute grading 9 frontier AI companies across 37 indicators in six domains, on the US GPA scale where A+ is 4.3 and F is 0. Scoring in the Summer 2026 edition was reviewed by a panel of seven independent experts.
- What was the highest grade in the Summer 2026 AI Safety Index?
- Anthropic, at C+ (2.66). No company scored above C+ overall. The best grade in the Existential Safety domain anywhere in the field was a D+, earned by Anthropic and OpenAI; every other company received a D or an F in that domain.
- Does a high FLI safety grade mean a vendor is compliant?
- No. The Index measures safety-research posture, not contractual or regulatory posture. It does not assess whether a vendor will sign a business associate agreement, what certifications it holds, whether it trains on your data, how long it retains it, or its EU AI Act status. A C+ is not an approval and an F is not a prohibition.
- Does the FLI Index grade cloud platforms like Azure OpenAI or Amazon Bedrock?
- No. FLI grades model developers. Deployment platforms host other companies’ models under their own commercial terms and are not graded. This is why the safety grade travels with the model while the compliance terms travel with the platform — a buyer usually combines one of each.
- How should a regulated buyer use the FLI Index?
- As one input among several, read at domain level rather than by headline grade, and as a trend instrument across editions rather than a snapshot. It is useful for calibrating internal expectations — the best safety grade in the industry being a C+ lands harder than a policy citation. It is not a substitute for diligence on the terms you can actually get in writing.
Turning a safety grade into a defensible decision
We help regulated organisations build AI vendor diligence that holds up in an audit.
