Executive Summary

In March 2026, AI model evaluation nonprofit METR suffered a significant credential theft incident when attackers exploited a fail-open authentication vulnerability in a researcher's AWS instance containing API keys for public AI models. The attackers maintained persistence for three weeks, consuming approximately $600,000 in AI model credits. A separate May incident involved sustained reconnaissance and probing of METR's infrastructure, including attempts to access internal evaluation data through an inadvertently exposed SQL endpoint. Both incidents highlight critical security gaps in AI research organizations handling sensitive frontier model evaluations for major vendors including OpenAI, Anthropic, Google, and Meta. The attacks demonstrate how conventional cloud security failures can lead to substantial financial impact and potential intellectual property exposure in the rapidly evolving AI evaluation ecosystem.

Why This Matters Now

As AI evaluation organizations become critical infrastructure for frontier AI safety, their security posture directly impacts the entire AI supply chain, making them high-value targets for threat actors seeking access to advanced models and evaluation data.

Attack Path Analysis

MITRE ATT&CK® Techniques

Potential Compliance Exposure

Sector Implications

Sources

Frequently Asked Questions

Attackers exploited a fail-open authentication vulnerability in a researcher's AWS instance that was inadvertently exposed to the public internet, then prompted an AI agent to reveal the API key.

Cloud Native Security Fabric Mitigations and ControlsCNSF

Based on the attack progression modeled above, these are the defensive controls that would constrain each stage.

Aviatrix Zero Trust CNSF would have significantly constrained this AI research infrastructure attack by enforcing microsegmentation around exposed EC2 instances and controlling east-west traffic flows. The segmented architecture would likely have reduced the attacker's blast radius and limited their ability to access model provider APIs across the compromised environment.

Initial Compromise

Control: Cloud Native Security Fabric (CNSF)

Mitigation: Zero Trust fabric policies would likely have reduced the attack surface of exposed EC2 instances by implementing default-deny postures and constraining which services could be directly accessible from the public internet.

Privilege Escalation

Control: Zero Trust Segmentation

Mitigation: Zero Trust segmentation policies would likely have constrained the agent orchestration tool's access to sensitive credential stores, reducing the scope of API keys accessible even after authentication bypass.

Lateral Movement

Control: East-West Traffic Security

Mitigation: East-west traffic controls would likely have restricted the attacker's ability to pivot between cloud services using compromised credentials, constraining their reachability to additional model provider endpoints and evaluation systems.

Command & Control

Control: Multicloud Visibility & Control

Mitigation: Continuous visibility and control mechanisms would likely have detected the unauthorized SSH key addition and anomalous outbound communication patterns, constraining the attacker's ability to maintain persistent command channels.

Exfiltration

Control: Egress Security & Policy Enforcement

Mitigation: Egress controls would likely have constrained the volume and pattern of API calls to external model providers, reducing the scope of credit consumption and limiting data exfiltration through controlled outbound traffic policies.

Impact (Mitigations)

While financial damages would likely still occur, the constrained blast radius from Zero Trust segmentation would have reduced the scope of affected systems requiring remediation and limited exposure of sensitive evaluation data.

Impact at a Glance

Affected Business Functions

  • AI Model Security Evaluation
  • Research Operations
  • Third-party AI Assessment Services
  • Intellectual Property Protection
Operational Disruption

Estimated downtime: 7 days

Financial Impact

Estimated loss: $600,000

Data Exposure

Potential exposure of unpublished AI model evaluation results, some sensitive model access data including private-model evaluation results and hidden chain-of-thought data. No evidence of actual data exfiltration in the May incident, but vulnerability existed that could have exposed category 2 and some category 3 classified data.

Recommended Actions

  • Implement Zero Trust Segmentation with identity-based policies and least privilege access to prevent lateral movement from compromised instances to sensitive AI model infrastructure
  • Deploy Multicloud Visibility & Control with centralized policy enforcement to detect anomalous API usage patterns and repeated malformed requests against evaluation endpoints
  • Enable Egress Security & Policy Enforcement with FQDN filtering to prevent unauthorized data exfiltration and monitor outbound traffic to external model providers
  • Establish Threat Detection & Anomaly Response capabilities to baseline normal API consumption patterns and alert on suspicious automation or credential abuse
  • Implement Cloud Native Security Fabric (CNSF) with distributed policy enforcement to provide real-time inspection and control of AI agent interactions and autonomous systems

Secure the Paths Between Cloud Workloads

A cloud-native security fabric that enforces Zero Trust across workload communication—reducing attack paths, compliance risk, and operational complexity.

Cta pattren Image