arXiv: Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
AI Analysis
A new research paper, published on arXiv, challenges the reliability of internal harmfulness scores used to evaluate AI safety. The study demonstrates that these scores, which are often used to rank how dangerous a model's outputs are, can be systematically manipulated. Specifically, the authors found that successful jailbreaks—prompts designed to bypass safety filters—frequently receive low harmfulness scores, meaning the internal measurement system fails to flag them as dangerous. This indicates that relying solely on these scores for safety assurance is fundamentally flawed.
This finding directly impacts any organization deploying or developing large language models, particularly those in regulated sectors like finance, healthcare, and legal services, where compliance with AI safety standards is critical. It also affects cloud providers and AI vendors who use these internal metrics to certify their models for enterprise use. The paper suggests that current evaluation frameworks, which may be used to demonstrate compliance with emerging EU AI Act requirements, could be providing a false sense of security.
Compliance teams should immediately review their AI risk assessment procedures to determine if they depend on internal harmfulness scores. They should treat these scores as a supplementary signal, not a primary safety guarantee. The next step is to implement additional, independent testing methods, such as red-teaming with diverse jailbreak attempts and human review of edge cases. This will help ensure that safety evaluations are robust and that the organization is not unknowingly deploying models with exploitable vulnerabilities.
Get notified about AI_SAFETY changes
Subscribe to our free weekly digest covering 24 compliance frameworks.