Blog
Engineering
Beyond the Loop: Agentic Capabilities at Scale
As security agents take on more consequential work, each new capability must inherit the same controls and reliability as the last. Cogent’s platform uses layered abstractions to make that possible, from governed access and asynchronous communication to execution that scales with demand.
6 min read

Cybersecurity is home to some of the highest-stakes challenges in applied AI today, and those challenges keep changing as new threats and new agentic patterns emerge. Keeping up means adding new capabilities to our platform every day, and to move at blazing pace without compromising on the strict requirements of cybersecurity. The gold standard for AI platforms is exactly that combination, capability that grows as fast as the frontier and holds up to the stakes of security work, and it's the standard we're building Cogent's platform to meet.
Layers of abstraction
Building production-grade agent systems in production is not easy. Much of that difficulty lies in defining the right abstractions that are maintainable and extensible as the system grows.
At the core of every agent we run is a simple agentic loop, a pattern widely adopted across the industry today. On top of that simple loop, our harness builds multiple layers of reusable abstractions, each composed from the layer beneath it.
The foundation is security-agnostic. It starts with the standard agentic primitives: prompts, tools, MCP servers, and hooks. Runtime services build on those primitives, giving our agents sandboxes for code execution, governed access to customer systems, context, and the agent inbox. Platform capabilities such as report generation and interactive visualizations compose those services into general-purpose building blocks.
Our security capabilities, like attack path analysis, are built on top of that foundation, and that's where our domain expertise lives. At the top, each agent composes exactly the capabilities its job requires.

The payoff of abstraction compounds at each layer. One well-chosen abstraction at a low layer becomes the foundation for many capabilities above it, and each of those becomes a new building block in turn. The result is a single agent platform that is highly dynamic and can naturally support any type of agentic use case.
Governed access that every other capability inherits
Governance is a capability implemented in the lowest layers, so every capability built above them inherits it. At its core is tenant isolation: each customer's data is kept separate, so data from different customers never mixes. The tenant-isolated infrastructure guarantees that no agent can cross into another customer's data: every agent carries an identity (which agent, its exact version, which tenant, which user) that binds it to a single tenant for its entire lifecycle. Within that tenant, permissions narrow access further, and they're declared where each tool is defined:
Every call is checked against the permissions of the user the agent is acting for. An agent can never do more than the person it works for. Every action the agent takes passes through our Agent Gateway, which decides whether that specific action is allowed.
Capabilities across the stack

Some capabilities span the backend and frontend, and defining the right interfaces allows us to reuse the same patterns across many experiences. When an agent traces an attack path or the lineage of a container image, the tool it calls can return an interactive graph along with its result. The harness checks the graph against what that tool is allowed to display, sends it to the product, and keeps it out of what the model reads. The user sees the graph, and the model reasons over the underlying data without spending its limited attention on display details.
Asynchronous communication, layered on existing primitives
Some lower-level capabilities are improved and superseded by higher-level ones. Agents have traditionally taken input only at rigid turn boundaries; to go beyond that, we layered a new abstraction on primitives the harness already had. Hooks give the loop safe points between steps where additional context can be added before the next model call. Shared storage outside the agent's process gives that input somewhere to wait. Together, they form a per-session inbox that anything in the system can write to and the agent reads at every safe point.
That one abstraction unlocks a new set of capabilities one layer up, and none of them needed infrastructure of its own:
Human-to-agent steering. A user sees their AI assistant looking at the wrong data and types "use real-time data instead." The agent changes course at its next step without starting over.

Async tools. A tool that scans a large dataset or generates a complex report returns a receipt right away, and the agent keeps working. The result arrives later as an inbox message.

Agents working together. Agents can write to each other's inboxes, handing over context or intent they've uncovered. The receiving agent folds it in at its next safe point, and neither has to wait for the other to go idle.

Agent progress and output are emitted in a similar way. Everything an agent does is published to an event stream backed by Redis. The live stream is short-lived and built for fast delivery, and combined with durable progress snapshots, it provides a permanent execution record. Any person, service, or agent can follow along, and a follower that disconnects can reconnect later and ask for everything after the last event it saw.

Infrastructure that keeps up with threats
The right abstractions are only half the problem. The other half is making sure they hold up at scale. Our product regularly generates large bursts of agent work: scheduled runs across every customer, bulk triage when new findings arrive, and batches of investigations launched at once. Incoming work lands in a queue, and our worker pool scales with it, so a burst means more workers rather than slower agents. That only works because the lowest layers, like the agent inbox and event stream, scale out with them. Everything built above inherits that scale, so no capability becomes the bottleneck as the number of running agents grows.
Holding up at scale also means holding up under failure; in a distributed system, infrastructure failures are a given. In AI systems, errors outside our control are even more common, with many external dependencies on model providers and a multitude of data dependencies that span MCP servers, databases, and search engines. For our agents, failure handling spans multiple layers:
Durable execution. We use Temporal to ensure that long-running and background work survives restarts and deploys.
Idempotent side effects. Consequential actions are designed so that a retry can't apply them twice.
The agent's own reasoning. Agents themselves are uniquely good at self-recovering from intermediate states and finding fallback paths if some dependencies are unavailable. With the right guardrails to prevent inconsistent state, combined with proper visibility into what’s already happened, we don’t necessarily need to build the static error handling other systems might require.
Where capability, security, and scale meet
Our agents span a wide range of use cases: some help security engineers automate their day-to-day work, some work in the background to contextualize and prioritize security findings against business context, while others drive auto-remediation to autonomously eliminate security risks. This type of work demands frontier capabilities alongside strict permission control, agent identity, and customer isolation, all as hard requirements from day one.
That's the bar we're building Cogent's platform to meet, and what will power the next generation of cyber defense.





