Ridge Security Benchmark Challenges AI Model Dominance in Autonomous Red Teaming

  • Ridge Security published the first public benchmark comparing eight AI models in autonomous penetration testing, evaluating 96 model-target test runs.
  • Grok 4.5 led with 77% cumulative coverage, while GPT-OSS-120B achieved the highest efficiency at 16.9 findings per million tokens.
  • The benchmark revealed that model performance alone doesn't determine effectiveness in autonomous security workflows.
  • RidgeGen™ harness demonstrated the importance of architecture in managing execution, adaptation, and validation of AI models in security tasks.
  • Frontier models showed alignment issues, refusing certain actions mid-workflow, highlighting the need for architectural safety controls.

Ridge Security's benchmark underscores a shift in AI-powered cybersecurity from model selection to system architecture. As autonomous red teaming gains traction, the ability to verify and reproduce findings will become a critical differentiator. The findings challenge the assumption that the most capable general-purpose LLM will automatically excel in specialized security tasks, highlighting the need for tailored architectures like RidgeGen™. This trend reflects broader industry movements toward modular, adaptable AI systems in enterprise security.

Model-Agnostic Flexibility
Whether Ridge Security's model-agnostic architecture will allow organizations to adapt to evolving AI models without platform overhauls.
Cost-Efficiency Tradeoffs
How smaller and open-source models will compete with frontier models in balancing coverage, efficiency, and deployment flexibility.
Alignment Challenges
The pace at which AI models will improve in maintaining alignment with authorized security testing boundaries.