Document retrieval using intra-image relationships

Inventors

Flagg, Cristopher • Frieder, Ophir

Assignees

Georgetown University

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12169518-B2

Patent

Publication Date

2024-12-17

Expiration Date


Abstract

Technologies are described for retrieving documents using image representations in the documents and is based on intra-image features. The identification of elements within an image representation can allow for deeper understanding of the image representation and for better relating image representations based on their intra-image features. The intra-image features present in image representations can be used in searches. Search results can further be reranked to improve search results. For example, reranking can allow search results to conform to intra-image dominant image features.

Core Innovation

The invention provides a document retrieval approach that uses an image representation comprising two or more segmentations of an image. For a first segmentation and a second segmentation, one or more latent space representations are generated, and first and second one or more feature vectors are generated based on the latent space representations. The method determines that the first segmentation is a dominant segmentation and the second segmentation is a non-dominant segmentation in the image representation, and then assigns a greater weight to the feature vectors for the dominant segmentation and a lower weight to the feature vectors for the non-dominant segmentation.

A weighted set of feature vectors is compared to one or more other feature vectors generated from one or more latent space representations based on one or more other image representations. Based on the comparing, at least one image representation is retrieved from the one or more other image representations that has a greater similarity to the dominant segmentation than the non-dominant segmentation. The latent space representations for each segmentation are generated by determining an anchor image representation, selecting a positive image representation, selecting a negative image representation, calculating a first vector representation and a second vector representation, and generating the one or more latent space representations based on the anchor, the positive, and the negative image representations.

The invention further supports query-time search result re-ranking by obtaining a query image representation comprising two or more segmentations of a query image and generating latent space representations and feature vectors for each segmentation. The method determines a dominant segmentation and a non-dominant segmentation for the query image representation, assigns a greater weight to the feature vectors for the dominant segmentation to generate a weighted set of feature vectors, obtains a set of external search results comprising a set of image representations, and generates latent space representations and feature vectors for the external search result image representations.

An order of the set of external search results is modified based on a similarity comparison of the weighted set of feature vectors and the feature vectors generated for each external search result image representation. The approach is motivated by inter-image label locality problems, and the document retrieval is implemented using latent/feature-vector similarity rather than pixel-level matching, with segmentations and annotations represented as multi-label elements.

Claims Coverage

The independent claims are clm-00001, clm-00005, and clm-00020, each centered on generating latent space representations and feature vectors from two or more segmentations, identifying a dominant versus non-dominant segmentation, applying feature weighting, and using similarity for retrieval or re-ranking. Across the independent claims, the inventive features focus on weighted dominant/non-dominant feature comparison for retrieval and reranking of external search results using similarity between weighted dominant/non-dominant query features and external image features, with latent-space generation specified via anchor/positive/negative vector representations in clm-00001.

Dominant and non-dominant segmentation feature weighting for similarity-based retrieval

Obtain an image representation comprising two or more segmentations; generate one or more latent space representations and first and second one or more feature vectors for a first segmentation and a second segmentation; determine that the first segmentation is a dominant segmentation and the second segmentation is a non-dominant segmentation; assign a greater weight to the first feature vectors and a lower weight to the second feature vectors to generate a weighted set of feature vectors; compare the weighted set of feature vectors to one or more other feature vectors generated from latent space representations based on one or more other image representations; retrieve at least one image representation having greater similarity to the dominant segmentation than the non-dominant segmentation.

Anchor/positive/negative latent space generation from segmentations

Generate the one or more latent space representations for each of the first segmentation and the second segmentation by determining an anchor image representation from the respective segmentation; selecting a positive image representation; selecting a negative image representation; calculating a first vector representation between the anchor image representation and the positive image representation; calculating a second vector representation between the anchor image representation and the negative image representation; and generating the one or more latent space representations based on the anchor, the positive, and the negative image representations.

Weighted dominant and non-dominant query features for modifying external search results order

Obtain a query image representation comprising two or more segmentations of a query image; generate one or more latent space representations for each segmentation; generate first and second one or more feature vectors for the first and second segmentations based on the latent space representations; determine that the first segmentation is a dominant segmentation and the second segmentation is a non-dominant segmentation; assign a greater weight to the feature vectors for the dominant segmentation and a lower weight to the feature vectors for the non-dominant segmentation to generate a weighted set of feature vectors; obtain a set of external search results comprising a set of image representations; generate one or more latent space representations for each image representation in the external search results; generate feature vectors based on the latent space representations generated for each external image representation; and modify an order of the external search results based on a similarity comparison of the weighted set of feature vectors and the feature vectors.

Computing devices for reranking based on similarity of weighted segment features

One or more computing devices configured to perform search results order modification comprising obtaining a query image representation with two or more segmentations; generating one or more latent space representations for each segmentation; generating first and second one or more feature vectors for a first and second segmentation; determining the first segmentation as a dominant segmentation and the second segmentation as a non-dominant segmentation; assigning a greater weight to the first feature vectors and a lower weight to the second feature vectors to generate a weighted set of feature vectors; obtaining a set of external search results comprising image representations; generating latent space representations for each external image representation; generating feature vectors; and modifying an order of the external search results based on similarity comparison between the weighted set of feature vectors and the feature vectors.

Across the independent claims, the core claim coverage is directed to generating latent space representations and feature vectors for two or more segmentations, determining a dominant segmentation and a non-dominant segmentation, weighting dominant versus non-dominant feature vectors, and using similarity comparisons to retrieve images (clm-00001) or to modify an order of external search results for re-ranking (clm-00005 and clm-00020). clm-00001 additionally specifies anchor/positive/negative selections and vector calculations to generate the latent space representations.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.