Blog

Product

Cogent at the Frontier: Announcing VR-1, the First Mythos-Class AI Model Built for Cyber Defense

VR-1 is a reasoning model that investigates enterprises like a skilled adversary. It arrives with a new benchmark and the runtime to deploy it safely.

4 min read

Geng Sng

Co-Founder and CTO

Geng Sng

Co-Founder and CTO

Today we’re announcing Cogent VR-1, our new frontier AI model. Models like Claude Mythos 5 and GLM 5.2 handle cyber tasks well as a side effect of general capability. VR-1 is the first frontier model trained and optimized specifically for cyber, and on that specialized work it surpasses them.

Last week, OpenAI’s models broke out of a sandboxed evaluation on their own and compromised part of Hugging Face’s production infrastructure. That is a preview of what enterprises will face once capable AI is in adversary hands. Every organization needs the defensive counterpoint: AI that finds the attack paths an AI would find, so defenders can close them first.

We’re also releasing IntrusionBench, a new benchmark for AI cyber capability, and the Cogent AI Harness, a runtime environment for AI agents doing defensive work. All three are available through the Cogent Frontier Access Program.

VR-1: A Reasoning Model Trained to Find Attack Paths

The latest frontier models picked up impressive cyber capability as a byproduct of general strength in code and reasoning. VR-1 is the first frontier model deliberately trained for the work. Most frontier models can find a single bug in a single codebase. Real attacks rarely stay that contained.

Attacks often travel through a chain of small, individually unremarkable weaknesses. A public service exposes a minor flaw. An identity carries more permission than it needs. Somewhere in the build pipeline, an artifact gets trusted without ever being checked. On their own, each looks like ordinary background noise. Strung together, they become a path to a crown jewel.

We trained VR-1 to work the way a skilled adversary does. Given a foothold inside an enterprise, it maps the surrounding environment and connects weaknesses that span different systems to prove which paths are actually reachable. It then checks whether a proposed fix closes the gap or simply relocates the risk. Reasoning that once took an expert security researcher weeks now runs in hours.

Building a model for this work meant being able to measure it, and that required a new kind of benchmark.

IntrusionBench: A Benchmark That Tests Like a Real Breach

Existing benchmarks measure whether a model can exploit code in isolation. IntrusionBench measures something closer to a real breach. Each run gives an agent a foothold and an objective, then scores whether it can chain its way across cloud infrastructure, identity systems, and a company's own internal tools to reach a specific target and prove it got there. Runs are graded by execution rather than by what the agent claims it could reach.

On IntrusionBench, VR-1 proved twice as effective at finding attack paths in enterprise environments as other frontier models, at one-fourth the cost.



The figures above come from the hardest configuration, where the agent is told nothing about the environment it lands in, the same constraint a real attacker starts with. As more context was revealed, every model improved and the spread narrowed. Once the source code and the underlying weakness were disclosed outright, the models converged, which suggests VR-1’s advantage lies in probing unfamiliar environments rather than raw exploitation skill.

Cogent AI Harness: The Governed Runtime for Security Agents

The runtime an agent operates in shapes its performance as much as the model does. Alongside VR-1, we’re releasing the Cogent AI Harness, the runtime that makes a model like this safe to deploy inside a live enterprise. It gives any capable model, open-weight or frontier, the environment context, scoped tools, policy enforcement, and verification needed to operate as a governed security agent. 

The Cogent AI Harness is grounded in a detailed picture of each enterprise’s environment and how it operates. AI agents investigate who owns each asset and where sensitive data lives, along with operational context like patch windows, change freezes, and fixes that failed before. That history informs how new remediation gets planned. The picture is continually updated as assets and risks change, policies evolve, and the system applies mitigations and remediations. 

The Cogent AI Harness works with any capable AI model, and it lifts other frontier models’ performance on IntrusionBench as well.

Available Only Through the Cogent Frontier Access Program

VR-1 is not being released openly. Because the capabilities that help defenders can also be misused, we’re making the model available only to vetted organizations through the Cogent Frontier Access Program, with our safety guardrails, policy controls, and audit logging in place. Participants work directly with Cogent Research on model evaluation, validation in their own environment, and safe deployment.

Qualified participants can also receive a Frontier Model Risk Assessment, a grounded readout of the attack paths a frontier-level model could reach inside their environment, which weaknesses become more dangerous as AI capabilities advance, and where AI can be deployed safely to reduce that risk.


The Frontier Access Program is open for applications today.

Related articles

View all articles

View all articles

B0o8o@k9  aE  dAeJmTo1

See Cogent In Action

Schedule a personalized demo today to learn how Cogent can supercharge your vulnerability management program.

Book a demo

Book a demo

Free risk assessment

Free risk assessment

BZo#o4k2  aI  dXeCmLoI

See Cogent In Action

Schedule a personalized demo today to learn how Cogent can supercharge your vulnerability management program.

Book a demo

Book a demo

Free risk assessment

Free risk assessment

BQoYoBkR  aA  dHe%m4oE

See Cogent In Action

Schedule a personalized demo today to learn how Cogent can supercharge your vulnerability management program.

Book a demo

Book a demo

Free risk assessment

Free risk assessment