Categories:
Strategy
AI Safety Governance Vendor Selection Strategy

Every Major AI Lab Failed Safety Class in 2026 — Here's How to Pick a Vendor Anyway

Feature image for Every Major AI Lab Failed Safety Class in 2026 — Here's How to Pick a Vendor Anyway

Every Major AI Lab Just Got a Report Card — and the Valedictorian Scored a C+

Anthropic earned the highest grade in the 2026 AI Safety Index. The grade was a C+.

Nobody passed. And that should change how you pick AI tools.

The Future of Life Institute (FLI) published its 2026 AI Safety Index on July 16, 2026, grading every major frontier lab on risk management, transparency, governance, and whether they actually honor the safety commitments they make during fundraising rounds. Anthropic topped the class at C+. OpenAI and Google DeepMind each pulled a C. Meta landed a D+. And xAI, DeepSeek, and Mistral failed outright.

When the best student in the class is clocking a C+, the class itself is in trouble. These are the same models being wired into customer support, cybersecurity tooling, healthcare assistants, and autonomous agents — and the institutions building them are, by their own published standards, doing mediocre work on safety.

Here’s what the grades actually tell you, and how to use them when you’re deciding which lab’s API to build on.

The Full Gradebook: Who Passed and Who Didn’t

The index breaks down each lab across four dimensions — risk management practices, transparency of safety reporting, governance structures, and follow-through on commitments. The composite grades:

  • Anthropic: C+ — Highest score, driven by relatively strong governance and the most consistent track record of publishing safety research. Still mediocre against its own bar.
  • OpenAI: C — Solid transparency in spots, but flagged for walking back commitments and inconsistent governance as the company’s structure has evolved.
  • Google DeepMind: C — Strong technical safety research undercut by uneven public reporting and Frontier Safety Framework updates that critics found vague.
  • Meta: D+ — Open-weight releases without sufficient downstream accountability, limited red-teaming disclosure.
  • xAI, DeepSeek, Mistral: Fail — Insufficient transparency, missing or vague safety frameworks, limited independent evaluation access.

The headline finding that landed hardest: several labs quietly walked back safety promises they made during earlier fundraising cycles. Commitments made to investors and the public during capital raises did not survive contact with the deployment pressure of shipping competitive models.

Why a C+ Should Worry You If You’re Building on These APIs

If you’re shipping a product on top of one of these APIs, you’ve effectively outsourced a chunk of your safety, security, and compliance posture to a vendor that independent evaluators just graded as mediocre. That has real downstream consequences:

  • Prompt injection and data exfiltration risk scales with how rigorously the lab red-teams its models before release. A lab that cut safety corners during fundraising pressure is a lab that may have cut red-teaming corners too.
  • Regulatory exposure lands on you, not the lab. If an EU AI Act audit or a sector-specific compliance review asks how you evaluated your model provider’s safety practices, “they had a nice blog post” is not a defensible answer — and right now, most labs can’t hand you much more than that.
  • Commitment drift means the safety guarantees you evaluated during vendor selection may not be the guarantees in place a year later when your product is in production.

The FLI report essentially confirms what security practitioners have suspected: independent scorecards matter more than lab-published safety blogs. A polished corporate safety newsletter is marketing. A structured independent evaluation is data.

The “Will They Cut Corners Under Pressure?” Test

The single most useful question the index surfaces isn’t which lab scored highest — it’s which labs actually did what they said they would do when it cost them something. That’s the governance and commitment follow-through dimension, and it’s the one that predicts future behavior.

Labs that honored their commitments even when shipping late or losing competitive ground scored well. Labs that quietly revised their safety frameworks to match whatever they’d already shipped scored poorly. The first group is a safer long-term bet for a production dependency.

A Practical Framework: How to Evaluate AI Safety When Choosing a Lab

Use the FLI index as a starting point, not an endpoint. When you’re vetting a model provider for a real product, run this four-point check:

1. Demand third-party evaluations, not first-party claims

Ask the vendor for independent red-teaming results, external safety audits, and participation in frameworks like the index. If the only safety evidence they can produce is their own blog, that’s a signal.

2. Audit their commitment track record

Look at what the lab publicly committed to 18–24 months ago — safety frameworks, deployment pauses, evaluation access for researchers — and check whether they followed through. FLI built an entire scoring dimension around this because it’s the best predictor of future behavior.

3. Map the governance structure

Who at the lab has the authority to delay or halt a deployment on safety grounds? Is that authority real, or is it a PR artifact? Labs with meaningful internal safety veto power scored higher. Labs where safety teams report into commercial leadership scored lower.

4. Pressure-test your own fallback

If your chosen lab walked back a safety commitment tomorrow, or shipped a model that failed an external audit, how fast could you switch providers or pin to an older model version? Vendor lock-in is a safety risk. Architect for portability from day one.

What This Means for Your AI Stack Right Now

Three practical moves:

  • Diversify your model providers if you haven’t already. The gap between a C+ lab and a failing lab is wide enough that you shouldn’t have your entire product riding on the lowest-graded vendor.
  • Weight governance and follow-through heavily when selecting a primary provider. A lab with strong technical research but weak commitment tracking is a future incident waiting to happen.
  • Document your vendor safety evaluation as part of your procurement process. Regulators are going to start asking for this, and the companies that already have a structured answer will move faster.

The FLI index doesn’t tell you which model is smartest or cheapest. It tells you which labs are run by people who take safety seriously enough to do it even when it’s inconvenient. Right now, the best of them is a C+. Build accordingly.

Related Articles