System and methods for gesture inference using computer vision

Inventors

Ang, DexterCipoletta, DavidValk, Henry

Assignees

Pison Technology Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12340627-B2

Patent

Publication Date

2025-06-24

Expiration Date


Abstract

Disclosed are methods, systems and non-transitory computer readable memory for gesture inference. For instance, a first method may include computer vision to train and/or infer gesture inferences. For instance, a second method may include using transformations to data and/or ML models to address inter/intra-session variability of sensor data. For instance, a third method may include using ML model selection to select a ML model to address inter/intra-session variability of sensor data.

Core Innovation

The invention describes a gesture-inference system that combines video-based hand/arm key-point processing with a wearable arm device that provides biopotential data and motion data. At least one camera captures video images of an environment that include image timestamps. A first machine learning model processes a plurality of sets of key-point values determined from the image(s) to output a first gesture inference of the user's hand/arm among a plurality of defined gestures and assigns a first gesture inference timestamp based on the image timestamps.

A wearable device includes a biopotential sensor configured to obtain biopotential data indicating electrical signals generated by nerves and muscles in the arm, and a motion sensor configured to obtain motion data relating to motion of the arm portion, with the biopotential data and/or the motion data having sensor data timestamps. The system selects a subset of the biopotential data and the motion data whose sensor data timestamps overlap the first gesture inference timestamp. Using a second machine learning model, the system processes the selected subset to generate a second gesture inference using a combination at least including the biopotential data and the motion data.

The system modifies the second machine learning model based on at least a comparison between the first gesture inference and the second gesture inference. The disclosure further indicates preprocessing and feature extraction for multimodal inference, selecting or rejecting low-quality/outlier sessions, and transformations addressing inter-session and intra-session variability. It also indicates selecting improved machine learning models per user/session based on calibration or similarity to feature clusters or latent-space embeddings, and determining machine-interpretable events based on the gesture inference and executing corresponding actions.

Claims Coverage

Two independent claims are present: one system claim and one computer-implemented method claim. Both center on time-synchronized fusion of a vision-based gesture inference with wearable biopotential and motion data, followed by modifying the second machine learning model based on comparison between the two gesture inferences.

Time-synchronized dual-model gesture inference with wearable arm biopotential and motion data

A gesture-inference system that obtains video image(s) with image timestamps, determines key-point value sets from the image(s), uses a first machine learning model to output a first gesture inference and assigns a first gesture inference timestamp, selects a subset of biopotential data and motion data whose sensor data timestamps overlap the first gesture inference timestamp, and uses a second machine learning model to output a second gesture inference using at least the biopotential data and the motion data.

Model modification based on comparison between vision and biopotential-plus-motion gesture inferences

A system configured to modify the second machine learning model based on at least a comparison between the first gesture inference and the second gesture inference.

Computer-implemented time-synchronized dual-model gesture inference

A computer-implemented method that obtains video image(s) with image timestamps, determines key-point value sets indicating locations of hand portions for the image(s), uses a first machine learning model to obtain a first gesture inference from the key-point value sets, assigns a first gesture inference timestamp based on the image timestamps, obtains biopotential data from a wearable device and motion data from the wearable device with sensor data timestamps, selects a subset of the biopotential and motion data whose sensor data timestamps overlap the first gesture inference timestamp, and uses a second machine learning model to generate a second gesture inference from at least the biopotential data and the motion data.

Modifying the second machine learning model based on comparison between two gesture inferences

The computer-implemented method is configured to modify the second machine learning model based on at least a comparison between the first gesture inference and the second gesture inference.

Across both independent claims, the core inventive coverage is the same: generate a first vision-based gesture inference from key-point value sets with image timestamps, time-align wearable biopotential and motion data using overlapping sensor timestamps, fuse the aligned wearable data in a second machine learning model to produce a second gesture inference, and modify the second machine learning model based on a comparison between the two gesture inferences.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.