NetForge RL: Multi-Agent Simulation with Durative Actions for Cyber Defense

Discover NetForge RL, a multi-agent simulation environment with durative actions to train cyber defenders on OT/IT networks.

miércoles, 15 de julio de 2026 • 5 min read • Q2BSTUDIO Team

NetForge RL: Multi-Agent Defense with Durative Actions

In the fast-paced world of cybersecurity, anticipation is key. Businesses can no longer just react to attacks; They need controlled environments where they can train their defensive systems before a real threat hits. This is where NetForge RL comes in, a multi-agent simulation platform with durative actions designed specifically for cyber defense. This article explores how this tool is transforming the way organizations approach securing their networks, and how the integration of AI and AI agents is redefining the digital battlefield.

The premise is simple but powerful: simulate a realistic scenario where an attacker (red agent) tries to compromise hosts with partial observability, while several defenders (blue agents) collaborate to protect the network. Each agent acts under uncertainty, with synthetic telemetry and event logs that become dense representations, not perfect states. This approach faithfully reflects the day-to-day life of a security team, where information is noisy and incomplete. Simulation allows you to iterate quickly, testing defensive strategies without putting the actual infrastructure at risk.

For companies, this ability to experiment is an invaluable asset. Instead of relying on one-off penetration testing or waiting for an incident to occur to learn, they can integrate simulation as part of their continuous improvement cycle. This is where custom software development becomes relevant: every organization has unique needs, and a generic platform doesn't always fit. Customized solutions allow simulation parameters, threat types, and success metrics to be tailored to the concrete reality of the business.

NetForge RL is not just an academic environment. Its architecture, built on generative procedures from enterprise and OT networks, includes up to five scenarios of varying difficulty and a separate evaluation set. In addition, its vectorized transition core in JAX reaches more than 250,000 steps per second in batches of 4096, making it a fast substitute for training reinforcement learning loops. This is especially useful for teams working with cloud infrastructures such as AWS and Azure cloud services, where scalability and computational efficiency are critical.

Modern cybersecurity demands a multidisciplinary approach. It's not enough to have a firewall or antivirus; A strategy that integrates artificial intelligence, data analytics, and automation is needed. Multi-agent simulations allow AI agents to be trained to coordinate durative actions—operations that take time and require planning, such as isolating a network segment or applying emergency patches. This type of behavior is difficult to model with traditional methods, but multi-agent reinforcement learning (MARL) offers a promising path.

From a business perspective, the adoption of these technologies is not trivial. It involves investing in AI for companies that not only understand algorithms, but also integration with legacy systems and regulatory compliance. Companies that already work with business intelligence services, such as Power BI, find simulations an additional source of data to visualize threat trends or evaluate the performance of their defenses. The combination of AI agents with interactive dashboards allows managers to make informed decisions based on simulations rather than hunches.

The concept of durative actions is central to NetForge RL. Unlike board games where every move is instantaneous, in cyber defense actions have a realistic duration: a port scan can take seconds, a response to an incident can take minutes. Modeling this is crucial for agents to learn how to prioritize and manage limited resources. Cybersecurity teams can use these simulators to test response protocols, train their analysts, and evaluate security tools before deploying them to production.

At Q2BSTUDIO, we understand that technology is not an end in itself, but a means to protect critical assets. That's why we offer cybersecurity services ranging from audits to implementation of simulation environments such as NetForge RL. We also develop custom applications that integrate with cloud platforms, allowing our customers to take full advantage of the scalability of AWS and Azure without sacrificing customization. Business intelligence and tools such as Power BI become allies to monitor the status of the simulation and draw actionable conclusions.

One of the biggest challenges in cyber defense simulation is reproducibility. NetForge RL addresses this using deterministic seeds and an evaluation system with 95% confidence intervals. This allows researchers and security teams to compare strategies objectively. In addition, it includes six diagnostic probes that measure specific defensive skills, such as the ability to detect lateral movement or contain a ransomware attack. For companies, these probes are equivalent to unit tests in software development: they help identify weaknesses in the defensive posture before they are exploited.

The transition from simulation environments to real environments is another key aspect. Agents trained in NetForge RL can be moved to production systems with minimal adjustments, as long as the fidelity of the model has been taken care of. This is where collaboration with experts in custom software development makes all the difference. A team that knows both the capabilities of AI agents and the specifics of the corporate network can design integration bridges that minimize friction.

In short, NetForge RL represents a step forward in the convergence between artificial intelligence, simulation and cybersecurity. For companies looking to stay one step ahead of attackers, investing in these types of tools is not a luxury, but a necessity. Whether by adopting pre-configured environments or by developing bespoke solutions, the key is to understand that cyber defense is a dynamic process that requires continuous learning. Q2BSTUDIO is prepared to accompany organizations on this path, offering both the technology and the knowledge necessary to transform simulation into real protection.

The cybersecurity of the future will not only depend on faster firewalls or stronger encryptions, but on systems capable of learning, adapting, and collaborating. NetForge RL is an example of where we're headed: an ecosystem where AI agents, trained in realistic scenarios, become the front line. And for companies that want to lead this transition, investing in artificial intelligence, cloud services, and custom applications is the surest path to proactive and effective defense.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.