Anthropic published a comprehensive 186-page risk report on August 17 that fundamentally revises the company’s previous safety assessments. The document reveals three significant developments: a downward revision of the company’s own safety verdict, the shelving of a frontier model that will never ship, and an admission that bioweapon detection classifiers were disabled for eleven months.
The report marks a rare instance of an AI lab publicly acknowledging internal safety failures at this scale. According to sources familiar with the document, Anthropic’s safety team had originally assessed certain model capabilities as within acceptable risk thresholds—but the new analysis reaches different conclusions based on updated evaluation methodologies.
Perhaps most notably, the report discloses that Anthropic decided not to release a frontier-level model internally designated as “Model 2.” The decision to shelve the model came after extended safety evaluations revealed capabilities that exceeded the company’s risk tolerance thresholds. This marks one of the few documented cases of a major AI lab voluntarily withholding a trained frontier model from release.
The eleven-month gap in bioweapon classifier functionality represents a significant transparency milestone. From approximately late 2025 through mid-2026, Anthropic’s systems for detecting requests related to biological weapons creation were not operational. The report details when the gap was discovered and the steps taken to restore and improve these safety measures.
Industry observers note that the timing of the report release—coming on the heels of Anthropic’s first operating profit and amid intensifying regulatory scrutiny—adds strategic weight to the disclosure. The company has faced increasing pressure from governments worldwide to demonstrate robust AI safety practices.
The Future of Life Institute’s Summer 2026 AI Safety Index rated Anthropic at C+, the highest grade among evaluated labs but still below the B threshold that would indicate strong safety performance. The new risk report suggests the company’s self-assessment may be even more critical than external evaluations have been.
Anthropic stated that the report represents a commitment to transparency and continuous improvement of safety practices. The company indicated it will publish updated safety frameworks based on lessons learned from the review process.