Intelligent surgery video management and retrieval system
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
This disclosure describes an intelligent management and retrieval system for surgery videos. The system allows user to upload their surgery videos and provide description of the surgery. A trained machine learning model detects whether there are any privacy leaking segments or frames inside the user uploaded video, and the system removes such privacy information from the video. Medical devices, tissue characteristics and events are detected in the surgery video using trained object recognition, event recognition models. Such detection result, combined with user provided description of the surgery, are utilized to construct rich description of the surgery video.
Core Innovation
The invention relates to management of medical operation videos, where a medical operation video showing a medical operation performed on a patient is received, and after the medical operation video is created, a user-provided text description of the medical operation is received. A plurality of features is detected from the medical operation video, and a rich description of the medical operation video is generated based on the after-video-creation text description and the detected features, where the content information is distinct from and complements the user-provided text description.
The management process includes performing a retrieval process by receiving a search query for video retrieval and matching the search query against the rich descriptions of the plurality of medical operation videos. The matching is performed based on both the descriptions and the detected features of the plurality of medical operation videos, so that video retrieval uses content information derived from video content features in addition to the user-provided text description.
The invention further provides enriched retrieval scenarios in which the search query includes a search video snippet that expresses a search need. In such cases, a return video or return video snippet is retrieved from among multiple medical operation videos, and the search video snippet is processed by detecting a plurality of features and determining a textual description based on the detected plurality of features, using one or more sentence embedding techniques to support matching of the search query.
Claims Coverage
The document includes three independent claims, a method, a system, and a non-transitory machine-readable medium, each organized around generating a rich description from user-provided text and detected video features and performing retrieval by matching a search query against the rich descriptions using both descriptions and detected features.
Rich description from user text and detected video features
Detecting a plurality of features from the medical operation video and generating a rich description of the medical operation video based on the after-video-creation text description of the medical operation and the detected features, where the rich description provides content information about what is in video contents of the medical operation video and complements the user-provided text description.
Retrieval by matching search query against rich descriptions using descriptions and detected features
Receiving a search query for video retrieval and matching the search query against the rich descriptions of the plurality of medical operation videos, based on both the descriptions and the detected features of the plurality of medical operation videos.
Search video snippet-based retrieval
Receiving a search query made from a search video snippet that expresses a search need to match and retrieve a return video or return video snippet from among multiple medical operation videos.
Textual description of search video snippet from detected features
Determining a textual description of the search video snippet by using the detected plurality of features from the snippet and using that textual description to match the search query and retrieve a return video or return video snippet from among multiple medical operation videos.
Semantic-space embedding for snippet textual description
Determining the textual description of the search video snippet by embedding the detected plurality of features from the snippet into a semantic space using one or more sentence embedding techniques.
Across the independent claims, the core inventive coverage is generating a rich description that complements user-provided text with detected video content features, and performing retrieval by matching a search query against those rich descriptions using both stored descriptions and stored detected features. Dependent claim coverage further includes retrieval using a search video snippet, deriving a textual description from detected snippet features, and embedding into a semantic space using sentence embedding techniques.
Stated Advantages
Provides content information about what is in video contents of the medical operation video.
Content information in the rich description is distinct from and complements the user-provided text description.
Enables video retrieval by matching a search query against the rich descriptions based on both descriptions and detected features.
Documented Applications
Medical operation video retrieval using text-based or search-video-snippet-based queries to retrieve a return video or return video snippet from among multiple medical operation videos.
Privacy-related video management by detecting portions showing image content outside the patient's body and removing or modifying those detected portions.
Interested in licensing this patent?