Adversarial reinforcement learning system for simulating security checkpoint environments

Inventors

LEWIS, Brian JacobDEICH, Jason AdamMELSOM, Stephen JohnDODENHOFF, Kara JeanNIGGEL, William Tyler

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Assignees

Noblis Inc

Member
Noblis
Noblis

Noblis is a nonprofit research and technical organization supporting federal missions in defense, health, environment, and security. Emphasizing applied sciences, engineering, digital transformation, artificial intelligence, cloud, and cybersecurity, Noblis provides objective solutions for government agencies confronting complex operational and scientific challenges.

Publication Number

US-12242615-B2

Patent

Publication Date

2025-03-04

Expiration Date


Abstract

An adversarial reinforcement learning system is used to simulate a spatial environment. The system includes a simulation engine configured to simulate a spatial environment and various objects therein. The system further includes a first model configured to control objects in the simulation and a second model configured to control objects in the simulation. The first model generates a threat-mitigation input to control one or more objects in the simulation, and the second model generates a threat input to control one or more objects in the simulation. The system then executes a first portion of the simulation based at least in part of the threat mitigation input and the threat input.

Core Innovation

The invention provides an adversarial reinforcement learning system and simulation method for security checkpoint environments with a spatial environment or layout. A simulation engine models a spatial checkpoint that includes threat-mitigation objects and threat objects, where the threat-mitigation objects represent detectors/personnel and the threat objects represent weapons/materials. The simulation supports modeling threat outcomes based on location, proximity, and combined threats.

In the simulation, a defense model generates a threat-mitigation input comprising instructions for controlling one or more simulated objects, and the defense model is configured to minimize one or more harm outcomes of the simulation. In parallel, an attack model distinct from the defense model generates a threat input comprising instructions for controlling one or more simulated objects, where the attack model is configured to maximize one or more harm outcomes. A first portion of the simulation is executed based at least in part on both the threat-mitigation input and the threat input.

The method runs many simulation iterations, optionally in turn-based or simultaneous input configurations, and can use outcomes from prior iterations to adapt strategies toward equilibrium or stability. The simulation may stop based on satisfaction of stability criteria and can output optimized attack/defense strategies and metrics. The document further describes using synthetic data and an automated threat-detection algorithm, receiving environment/object data and asset limitations to constrain deployments and funding, and applying results to configure real-world security checkpoint devices and deployments, including x-ray machines, metal detectors, body-scanner devices, canine security agents, and security officers.

Claims Coverage

The independent claim coverage includes three independent claim types: a method for simulating a spatial environment, an adversarial reinforcement learning system for simulating a spatial environment, and a non-transitory computer-readable storage medium storing instructions for the system. Across these independent claims, the coverage centers on two distinct models that generate a threat-mitigation input and a threat input to respectively minimize and maximize harm outcomes, followed by executing a portion of the simulation based on both inputs.

Adversarial threat mitigation and threat maximization inputs

Generating, by a first model, a threat mitigation input, wherein the threat mitigation input comprises instructions for controlling one or more simulated objects in the simulation, wherein the first model is configured to minimize one or more harm outcomes of the simulation; and generating, by a second model, a threat input, wherein the threat input comprises instructions for controlling one or more simulated objects in the simulation, wherein the second model is distinct from the first model and is configured to maximize one or more harm outcomes of the simulation.

Executing a portion of the adversarial simulation based on both inputs

Executing a first portion of the simulation based at least in part on the threat mitigation input and the threat input.

Across the independent claims, the core inventive structure is the adversarial setup of two distinct models that generate a threat-mitigation input to minimize harm outcomes and a threat input to maximize harm outcomes, followed by executing a portion of the simulation based on both inputs. The independent system and medium claims correspond to the same generate-and-execute structure, implemented as an adversarial reinforcement learning system or as stored instructions, respectively.

Stated Advantages

Documented Applications

Applying results to configure real-world security checkpoint devices and deployments, including x-ray machines, metal detectors, body-scanner devices, canine security agents, and security officers.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.