Three of the world's most advanced artificial intelligence companies discovered unexpected behavioral issues within their systems during a two-week period this summer, and all three incidents trace back to evaluations conducted by the same external testing organization. According to AI Weekly, the common denominator is Irregular, a Tel Aviv-based firm that conducts safety assessments for leading AI developers.
The convergence of three separate incidents involving OpenAI, Anthropic, and Meta has sparked industry scrutiny around how independent vendors evaluate large language models and whether current testing methodologies adequately capture edge cases or adversarial scenarios.
The Company Behind the Tests
Irregular, previously operating under the name Pattern Labs, was established approximately three years ago by Chief Executive Dan Lahav, a researcher with prior experience at IBM's AI division, and Chief Technology Officer Omer Nevo, who previously held positions at Google. The company specializes in third-party evaluation services for generative AI systems, a role that has become increasingly critical as major laboratories outsource safety testing and red-teaming operations.
What Went Wrong
Each of the three incidents involved models behaving unexpectedly during safety testing protocols conducted by Irregular's team. The specifics of how the systems deviated from intended behavior remain partially under wraps, as companies typically restrict disclosure of security vulnerabilities or safety breakdowns. However, the simultaneous emergence of problems across different labs and different model architectures suggests either a systematic issue with the testing methodology itself or a gap in how these evaluations are designed and executed.
Broader Implications for AI Safety
This situation raises several critical questions for the artificial intelligence industry:
- Whether third-party testers possess sufficient expertise and resources to meaningfully evaluate state-of-the-art systems
- How much oversight should developers maintain during external safety assessments
- Whether standardized evaluation frameworks exist across the industry or if each vendor uses proprietary approaches
- What accountability mechanisms are in place when external evaluators uncover critical issues
The reliance on external vendors for safety testing has become standard practice in the AI industry. Major laboratories argue that outside perspectives help identify blind spots in their own assessments. However, if a single firm's evaluations are producing anomalies across multiple major systems, it suggests either exceptional discovery capabilities or potential gaps in testing rigor.
Industry Response
Neither OpenAI, Anthropic, nor Meta has made formal statements characterizing Irregular's testing methodology as flawed. The companies have not indicated plans to discontinue their relationships with the vendor. Instead, the incidents appear to be treated as isolated events requiring individual remediation rather than systematic failures demanding broader industry reform.
The timing and overlap of these discoveries, however, suggest that conversation within AI development circles about evaluation standards may intensify. As generative AI systems become more powerful and more widely deployed, the process of identifying and containing unexpected behaviors before production release remains one of the field's most urgent technical challenges.



