
SASTBench: Measuring AI's Path to Practical Security Automation
Can AI-powered triage solve your SAST alert bottleneck, or make a bad situation worse? We built SASTBench to evaluate triage agents under realistic conditions — real CVEs as true positives, filtered SAST findings as the noise.

Mythos 'Discovered' a CVE Already in Its Training Data — and That's Still Worrying
Anthropic made headlines claiming Claude Mythos achieved the "first remote kernel exploit discovered and exploited by an AI." We went looking for how — and found a 20-year-old bug hiding in plain sight.

Shai Hulud Returns: A Live Supply Chain Attack Unfolding
A new wave of activity consistent with the Shai Hulud supply chain attack pattern is emerging right now.

ClosedCaption: Finding Interpretable Clusters with LLMs
We discuss use-cases for LLM-cluster interpretability for agent analysis, experimental design to test these pipelines, and conclusions on how we use these techniques to analyze our own agents.

A Sneak Peek at Taxi — How We Understand Agents at Scale
We're excited to present 🚕 Taxi — a new tool we're developing to solve the difficulty of understanding what agents actually do at scale. Taxi is a generic, trajectory-oriented taxonomy generator that helps you make sense of your agent's behavior at scale.

How to Scale Agentic Reasoning Without Breaking
Introducing Conductor: Rival Security’s reasoning engine built for real‑world complexity. Unlike today’s fragile agentic systems, Conductor delivers verifiable, scalable results and achieves breakthrough performance on Spider 2.0 — marking a serious step forward in enterprise‑grade cybersecurity AI.

Setting the Standard: Our AI Model Outperforms Spider 2
Rival is redefining AI reasoning in cybersecurity by solving real-world analytical challenges that traditional agentic systems fail to handle. Here's how Rival's orchestrated workflows outperform state-of-the-art models on Spider 2.0.

Welcome to the Rival Security Research Blog
Follow our journey on the new Rival Security research blog, where we share our path from early experiments to building foundational AI systems that redefine cybersecurity.
