Blog

Engineering

Own the Loop, Integrate the Ecosystem: Why We Built a Custom Harness

Finding a vulnerability is only the start. To carry remediation through to a verified result, Cogent built a custom agent harness that keeps context, access, and verification under one product contract, even when a specialized external runtime does the work.

9 min read

Jia Zeng

Member of Technical Staff

Anirudh Ravula

Head of AI

Jia Zeng

Member of Technical Staff

Anirudh Ravula

Head of AI

Cybersecurity agents work in adversarial environments, with incomplete context and tasks that can outlast a single session. Their work matters when it ends in verified risk reduction. Cogent is building an AI workforce for vulnerability management. Our agents investigate findings in the context of each customer’s environment, determine what matters, coordinate remediation across people and systems, and verify that the risk is actually gone.

Our agents depend on the same foundation: the right context, governed access to tools, durable state, a clear lifecycle, and evidence that the work is complete. The system that brings those pieces together is the agent harness—the point where a model becomes a product. At its center is the agent loop: the cycle that turns context into action, evaluates what happened, and continues until the work reaches a verifiable result.

We built Cogent’s custom harness so we could evolve that loop with the product. As we added more kinds of agents and execution environments, we needed a consistent way to define agents, add capabilities, and connect new runtimes without reshaping the product around each one. That need led us to build our Agent Development Platform (ADP) SDK.

A capability-first approach

Early in Cogent’s journey, we were deliberate about building the capabilities that moved agents beyond conversation. We wanted agents to work more like well-equipped interns than simple chatbots: able to take on a real security task and carry it through to a verifiable result. We built our agent environment around that distinction, giving models the context and freedom to act while keeping their work contained and verifiable. Keeping capability development close to the product let each iteration create immediate value and establish the patterns that would later shape the SDK.

A clear operating model emerged for the loop: agents worked inside sandboxes, used reusable skills for specialized tasks, and verified their work before publication. In a sanitized IntrusionBench example, VR-1 combined cloud permissions with operational documentation to reach the specified invoice; access to other sensitive data did not count as completion.

Bringing security capabilities to every agent

Consistency and reuse became the next priority. Too much behavior was coupled to a particular agent loop or runtime. Adding a capability could require changes across unrelated parts of the system, and the same behavior was often wired differently from one agent to another. New execution models and external harnesses could extend what our agents were capable of, but each one required another bespoke integration.

The result was slower iteration and more maintenance as time went by. Capabilities that should have been portable were expensive to reuse, and it was harder than it should have been to adopt improvements from the broader agent ecosystem.

The issue was architectural: capabilities lived at the wrong interface layer. We needed a common interface that made them portable across agents and runtimes while preserving the containment, reuse, and verification model that had made them effective.

Why own the harness boundary?

We evaluated existing agent frameworks and coding harnesses. They can supply valuable execution capabilities, but adopting a loop does not by itself connect an agent to our customers’ live context, identity and access controls, remediation lifecycle, or verification criteria. Those are product responsibilities we needed to preserve as models and harnesses changed. Owning the boundary lets us evolve those controls and our native loop together.

At the same time, the broader agent ecosystem was moving quickly. As frontier models continued to improve and evolve, so too did the harnesses. New models and specialized harnesses offered stronger execution for particular kinds of work. We wanted a boundary that let us use different harnesses without reshaping the product around each system’s assumptions.

We therefore chose to own Cogent’s default loop and mandatory lifecycle contract while integrating external execution engines behind it. The goal was never to rebuild the ecosystem. It was to own the boundary closest to our product.

In Cogent’s native runtime, Context Engine supplies customer context; Agent Gateway provides governed tool access; Agent Sandbox Control Plane provides isolated execution; and the verifier checks outcomes. The Agent Control Plane brings agent versioning, identity, and deployment into the lifecycle around that execution.

Identity and policy sit on the path of every tool call. The harness requests an identity for the agent from the Agent Identity Service and passes it to Agent Gateway with each call. Agent Gateway authenticates that identity and checks the requested action with the Policy Engine before the tool runs, so an agent acts only as itself and only within policy.

Own the loop, integrate the ecosystem.

ADP SDK: a stable contract around a moving target

ADP SDK is the engineering interface to that operating model. It gives existing and future capabilities a consistent way to be defined, combined, and executed.

That interface keeps the product contract stable while allowing the execution layer to change. Cogent owns the context, capabilities, behavior, lifecycle events, state, and outcomes that determine how an agent enters and exits the product. The SDK includes a native loop, but it can also delegate execution to a specialized harness behind the same runner and event boundary.

The SDK expresses that contract through a deliberately small vocabulary shared across harnesses. A tool gives an agent a capability it may choose to use. A hook shapes behavior at a known point in the lifecycle. A runner determines how and where the work executes. External harnesses fit behind the same contract, preserving one product architecture.

@tool(description="Look up current weather.")
async def weather_lookup(args: WeatherArgs) -> dict[str, object]: ...

Harness.register_tool(weather_lookup)
await Harness.initialize()

agent = Agent(
    agent_id="weather",
    prompt=PromptSpec.from_text("Help with the weather."),
    llm_config=LlmConfig(model_name="..."),
    tools=["weather_lookup"],
)

runner, _ = await Harness.create_runner(agent=agent)
stream = await runner.stream(request)
@tool(description="Look up current weather.")
async def weather_lookup(args: WeatherArgs) -> dict[str, object]: ...

Harness.register_tool(weather_lookup)
await Harness.initialize()

agent = Agent(
    agent_id="weather",
    prompt=PromptSpec.from_text("Help with the weather."),
    llm_config=LlmConfig(model_name="..."),
    tools=["weather_lookup"],
)

runner, _ = await Harness.create_runner(agent=agent)
stream = await runner.stream(request)
@tool(description="Look up current weather.")
async def weather_lookup(args: WeatherArgs) -> dict[str, object]: ...

Harness.register_tool(weather_lookup)
await Harness.initialize()

agent = Agent(
    agent_id="weather",
    prompt=PromptSpec.from_text("Help with the weather."),
    llm_config=LlmConfig(model_name="..."),
    tools=["weather_lookup"],
)

runner, _ = await Harness.create_runner(agent=agent)
stream = await runner.stream(request)

Production agents extend this declaration with model settings, credentialed integrations, approval hooks, and typed outputs.

Low ceremony supports the larger objective: a small change surface. Engineers should be able to understand and change an agent from its definition.

The contract works in both directions. Specialized harnesses retain their own execution logic behind the runner and event interfaces, while Cogent’s native harness uses SDK primitives for agent definitions, tools, hooks, runners, and lifecycle events. Broader product systems connect through those interfaces.

One product architecture supports both paths. Cogent’s native loop can continue to evolve, and specialized harnesses can be adopted wherever they are the better fit.

Four kinds of work, one shared foundation

A shared foundation matters because Cogent’s agents operate in very different settings. ADP SDK must preserve the same product lifecycle across them while allowing the runtime to adapt to the work.

Customer-facing agents demand responsiveness and control. Our flagship product agents stream progress directly to security teams, so their behavior must be predictable and each turn must reach a clear terminal state. They also resolve current tenant context and apply approval and guardrail policy.

Automations and background agents demand durability. Customers usually see the result rather than the execution, and these jobs may run for minutes or hours. They must withstand transient failures across models, tools, networks, and workers. Recovery should preserve useful state, avoid duplicated side effects, and still produce a terminal outcome the calling system can reliably retrieve.

Cloud agents demand isolation and execution flexibility. These internal task and coding agents operate inside isolated cloud environments and often run long enough to require recovery. Different tasks also benefit from different harnesses. Engineers should be able to choose native, internal, or specialized external execution without rebuilding the surrounding product integration.

Internal operations agents demand governance and recoverability. They support incident orchestration, engineering-context collection, on-call assistance, and other operational work. They need governed access to operational systems and an inspectable record of their actions. Live workflows also require explicit human checkpoints and recovery from partial failures.

All four share the same product context, lifecycle, and outcome model. Their execution requirements lead to multiple runtimes behind that common contract. For long-running work, that contract also needs to preserve what the agent has learned and where execution stopped. Recovering investigation context and resuming execution are related requirements: a restarted process is not useful if it has forgotten the evidence needed for its next decision.

The runtime should fit the work

Hosted execution favors control. The agent loop runs inside a managed service, where application-level policy can govern each action. This fits customer-facing interactions that require secure, predictable behavior and responsive progress, while placing tighter bounds on autonomy.

Ephemeral sandboxes favor autonomy within isolation. Infrastructure becomes the primary guardrail, giving an agent room to inspect its environment, execute code, and adapt its approach. When a task benefits from a specialized external harness, the same SDK contract lets that harness retain its own execution logic inside the sandbox.

Durable execution preserves progress through failure. A durable proxy can supervise and resume a long-running external harness session. Workflows that require tighter control over every transition can run the agent loop itself as a durable workflow. In both cases, transient failure interrupts the work without erasing its progress, and every run moves toward a stable terminal result.

Across these execution models, ADP SDK preserves the same agent definition and product lifecycle while the runtime adapts to the work.

The custom harness is the Cogent-native runtime

Across the execution models above, Cogent’s custom harness is the runtime we design alongside the product. It gives us direct control over the complete system around the model.

Our preliminary VR-1 evaluation reported more than 2× the black-box pass@3 of the strongest evaluated frontier baseline on IntrusionBench, within two hours or 250 turns per trajectory, whichever came first. Pass@3 counts success within three attempts. The small, controlled benchmark measures VR-1 with its harness; it does not isolate the harness’s contribution or establish production remediation performance. That is why we evaluate model × context × tools × memory × policy × verifier as a system.

Together, these systems create the foundation for managing agents across their full lifecycle. Through the harness and surrounding platform, we can version agents, define their capabilities and behavior, govern their identity and access, and deploy them to the runtimes suited to their work.

That foundation lets us focus on what matters most: what our agents can do for customers. It brings the pace of frontier AI into production, enabling Cogent’s AI workforce to take on more of vulnerability management and carry the work further—from understanding a customer’s environment through coordinated action to verified resolution—while preserving the control and evidence enterprise teams need.

The winning cybersecurity AI won’t be the system that produces the most plausible answer. It will be the system that understands the customer’s environment, takes the right action under control, survives real operational complexity, and proves the risk is gone. Cogent built that system.

Related articles

View all articles

View all articles

B9oJoFk5  a7  d&eJm5oB

See Cogent In Action

Schedule a personalized demo today to learn how Cogent can supercharge your vulnerability management program.

Book a demo

Book a demo

Free risk assessment

Free risk assessment

B1oOo&kV  aX  dOe5mAoV

See Cogent In Action

Schedule a personalized demo today to learn how Cogent can supercharge your vulnerability management program.

Book a demo

Book a demo

Free risk assessment

Free risk assessment

B8oLoOkM  aU  d9e0m3oM

See Cogent In Action

Schedule a personalized demo today to learn how Cogent can supercharge your vulnerability management program.

Book a demo

Book a demo

Free risk assessment

Free risk assessment