Method for detecting an application progress and handling an application failure in a distributed system

Inventors

Dasari, Dakshina NarahariHamann, ArnePEREIRA, Nuno

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Assignees

Robert Bosch GmbH

Member
Carnegie Mellon University
Carnegie Mellon University

Carnegie Mellon University is a global research institution based in Pittsburgh, Pennsylvania, recognized for interdisciplinary education, research, and innovation in science, engineering, arts, technology, and social sciences. The university leads advancements in artificial intelligence, robotics, digital health, and performing arts. Located in a technology-driven and culturally rich city, CMU powers real-world impact through research centers, industry engagement, workforce training, and initiatives that shape regional and global communities.

Publication Number

US-12626540-B2

Patent

Publication Date

2026-05-12

Expiration Date


Abstract

A method for detecting an application progress and handling an application failure in a distributed system. The method includes: monitoring an interaction between modules of at least one application, the at least one application being deployed across different physical nodes, the interaction being carried out by exchanging messages between the modules using a message broker, the monitoring being carried out at least partially using the message broker; detecting the application progress based on the monitoring; initiating a failure handling based on the detecting.

Core Innovation

The distributed system includes a message broker through which messages are transmitted between a plurality of modules of at least one application deployed across different physical nodes. The message broker is programmed with a respective application manifest for each of one or more of the modules, and each manifest specifies expected processing behavior of the respective module within a processing pipeline. When the message broker routes a predefined input message output by another module in the processing pipeline to the respective module, the respective module is expected to generate a corresponding predefined output message for routing by the message broker to the other module or to a further module in the pipeline.

Application progress detection is performed by monitoring, by the message broker, for consistency of interactions between the modules with the application manifests. The message broker detects that a respective module has stalled when an expected processing behavior defined in the respective manifest is not met. The expected processing behavior comprises any one or more of an expected timing between receiving the input message and publishing the output message, a maximum backlog of unprocessed input messages, and a minimum number of messages to be processed successfully.

In response to detecting that the respective module has stalled, the message broker initiates a recovery action comprising restarting the respective module on a same physical node on which the respective module was running prior to the detection, restarting the respective module on a different node than on which the respective module was running prior to the detection, or starting a backup module to perform a processing of the stalled module. The architecture further includes a central orchestrator and a local module manager on the physical nodes to manage recovery-related module actions based on results of the stall detection.

Claims Coverage

The independent claims are 1, 19, and 20, covering a method, a non-transitory computer-readable medium, and a data processing apparatus. Across these independent claims, the inventive features center on manifest-driven monitoring for module stalls based on expected processing behavior, followed by broker-initiated recovery actions that include restarting on the same node, restarting on a different node, or starting a backup module. An additional independent-claim-related inventive aspect is the orchestration architecture (central orchestrator and local module manager) described for recovery actions.

Manifest-defined expected processing behavior consistency monitoring

The message broker monitors for consistency of interactions between the modules with the application manifests, where each manifest specifies expected processing behavior of the respective module within a processing pipeline.

Stall detection based on manifest-defined expected processing behavior not met

The message broker detects that a respective module has stalled when an expected processing behavior defined in the respective manifest is not met, the expected processing behavior comprising any one or more of an expected timing between receiving the input message and publishing the output message, a maximum backlog of unprocessed input messages, and a minimum number of messages to be processed successfully.

Broker-initiated recovery action with restart on same node, restart on different node, or backup module

In response to detecting that the respective module has stalled, the message broker initiates a recovery action comprising restarting the respective module on a same physical node, restarting the respective module on a different node, or starting a backup module to perform processing of the stalled module.

Central orchestration with local module manager recovery actions

Detection results of a stall are sent to a central orchestrator, and the orchestrator controls recovery-related module actions by issuing commands to a local module manager that deploys, stops, or starts modules on different physical nodes and reports module and node resource status back to the orchestrator.

The independent claim set covers using per-module application manifests to monitor inter-module interactions for consistency, detect module stalls when manifest-defined expected processing behavior is not met (timing, backlog, and/or success-message minimums), and initiate recovery actions via the message broker that restart the module on the same node, restart it on a different node, or start a backup module.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.