arXiv: Which Model Is Actually Serving You? IRIS: Budgeted Black-Box Auditing of Model Substitution and Routing Dilution in LLM Gateways
AI Analysis
This paper, published on arXiv, introduces a new auditing framework called IRIS designed to detect two specific risks in Large Language Model (LLM) gateways: model substitution and routing dilution. Model substitution occurs when a provider secretly swaps a requested high-cost, high-performance model for a cheaper, less capable one. Routing dilution happens when a gateway mixes outputs from multiple models without disclosure, degrading performance or introducing bias. The authors propose a budgeted, black-box method to test whether the model you are paying for is actually the one serving your requests.
The primary organizations affected are any entity deploying or procuring LLM services through third-party gateways, including cloud providers, AI startups, and enterprise IT departments. Financial services, healthcare, and legal sectors that rely on verifiable model outputs for compliance or liability reasons are especially vulnerable. Regulators and auditors monitoring AI safety and transparency under frameworks like the EU AI Act will also need to consider these risks.
Compliance teams should immediately review their contracts and service-level agreements with LLM providers to ensure explicit guarantees against model substitution and routing dilution. They should also begin evaluating whether to implement independent auditing tools like IRIS to verify model identity and output consistency. Finally, teams should document any discrepancies found and report them to relevant regulatory bodies, as undisclosed model changes may violate transparency obligations under emerging AI governance rules.
Get notified about AI_SAFETY changes
Subscribe to our free weekly digest covering 24 compliance frameworks.