Multi-stage narrative analysis
Inventors
Cho, Simon • Clark, Chae A. • Gordon-Sarney, Reed • Gentile, James
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
An extractive summarization model provides extraction and classification of research assertions, or claims, made by a documented body of research such as a scientific paper or article. Modern electronic publication and dissemination allows tremendous capability for researching and scrutinizing previous documented efforts for further research and study. Accordingly, a substantial volume of material is easily obtained in response to research efforts. The extractive summarization model provides a summarization of this scientific literature by identifying and classifying asserted claims made by a particular research document. Researchers may quickly identify relevant documents based on the extracted claims asserted by the document, facilitating substantive review.
Core Innovation
The invention relates to extractive summarization of scientific documents and uses a multi-stage classifier to detect claim assertions and to classify them into claim types. A first model computes a first group of sentences from a narrative representation of a scientific effort, based on a comparison using a first set of annotated features derived from an annotated corpus of statements derived from stored documents. The first stage provides a binary designation associated with which sentences are treated as claim assertions.
A second model is trained from the annotated corpus of statements, and the second model is trained on a subset of the annotated corpus determined based on the first set of annotated features. Using the first group of sentences, the second stage computes a classification of each sentence based on comparing the second model using a second set of annotated features derived from the annotated corpus. In some embodiments, the second stage produces multiclass designation into a plurality of classification groups determined by a multiclass designation from the subset based on the binary designation.
The resulting per-sentence classifications are defined by a probability having a higher accuracy than a classification or a probability computed from a single model trained on both the first set and the second set of annotated features, or from a single model trained from the annotations applied to an entire corpus of statements. The document also describes extracting claim segments and matching them to candidate sentences in the annotated corpus to compute a probability that each candidate sentence corresponds to its segment, and training data generation that includes negative examples indicative of prose but not annotated as scientific assertions.
Claims Coverage
The document describes independent multi-stage classification claims for determining claim assertions of a scientific effort using a first stage to compute a binary designation and a second stage to compute a multiclass designation. Across the independent claims, the core inventive framework includes training a first model on a first set of annotated features, selecting a subset of the annotated corpus for training a second model based on the first set, and computing sentence-level or sentence-group probabilities characterized as higher-accuracy than a single-model approach.
Multi-stage sentence classification with subset-trained second model for higher-accuracy probabilities
computing, from a narrative representation of a scientific effort, a first group of sentences based on a comparison using a first model of a first set of annotated features derived from the annotated corpus of statements; training a second model from the annotated corpus; and computing, from the first group of sentences, a classification of each sentence based on comparing the second model of a second set of annotated features derived from the annotated corpus, the second model trained on a subset of the annotated corpus determined based on the first set of annotated features, where the classification is defined by a probability having a higher accuracy than a classification based only on a single model derived from both the first set and the second set of annotated features.
Binary designation followed by multiclass classification with higher-accuracy probabilities
receiving an annotated corpus of statements, the annotations indicative of a scientific assertion proposed by a respective statement in the corpus of statements; training a model from an annotated corpus of statements derived from stored documents; computing, from a training narrative of a scientific effort, a group of sentences based on a binary designation determined by the model; and computing, from the group of sentences, a plurality of classification groups based on a multiclass designation determined by a model trained from a subset of the annotated corpus, the subset based on the binary designation, where the multiclass designation for each sentence is computed as a probability having a higher accuracy than a probability computed from a single model trained from the annotations applied to the entire corpus of statements.
Server device executing first and second stage models in series for multiclass designation rendering with higher-accuracy probabilities
a first stage model responsive to a training narrative of a scientific effort for computing a group of sentences based on a binary designation, the first model trained from an annotated corpus of statements derived from stored documents; a second stage model trained from a subset of the annotated corpus of statements, the second stage model responsive to the group of sentences for computing a plurality of classification groups based on a multiclass designation determined from the subset of the group of sentences, the subset based on the binary designation; the server device configured for executing the first stage model and the second stage model in series; and a rendering device for receiving the sentences based on the multiclass designation, the multiclass designation for each sentence computed as a probability having a higher accuracy than a probability computed from a single model trained from a single model defining the multiclass designation.
Computer program on a non-transitory medium implementing multi-stage classification with subset-trained second model for higher-accuracy probabilities
training a first model from an annotated corpus of statements derived from stored documents; computing, from a narrative representation of a scientific effort, a first group of sentences, based on a comparison of the first model of a first set of annotated features derived from the annotated corpus of statements; training a second model from the annotated corpus of statements; and computing, from the first group of sentences, a classification of each of the sentences in the first group based on comparing the second model of a second set of annotated features derived from the annotated corpus of statements, the second model trained on a subset of the annotated corpus, the subset determined based on the first set of annotated features, where the classification of each sentence is defined by a probability having a higher accuracy than a classification based only on a single model derived from both the first set and the second set of annotated features.
Binary prose-or-claims classification followed by multiclass claim-type classification with subset-trained multiclass stage
computing, from a training narrative of a scientific effort, a group of sentences based on a binary designation determined by a model derived from the annotated corpus; training the model for binary designations based on annotations designating statements in a corpus as one of either prose or claims to generate a binary classifier model; training the model for multiclass designations based on the annotations designating the claims further designating a type of claim to generate a multiclass classifier model; and computing, from the group of sentences, a plurality of classification groups based on a multiclass designation determined by a model trained from a subset of the annotated corpus, the subset based on the binary designation, where the multiclass designation for each sentence is computed as a probability having a higher accuracy than a probability computed from a single model trained from the annotations applied to the entire corpus of statements.
Across the independent claims, the inventive concept is a multi-stage classification pipeline that first computes a binary designation of sentences, then computes multiclass claim-type classifications using a second model trained on a subset determined by the binary stage, with output probabilities characterized as having higher accuracy than single-model approaches.
Stated Advantages
the per-sentence classifications are defined by a probability having a higher accuracy than a classification based only on a single model
the multiclass designation for each sentence is computed as a probability having a higher accuracy than a probability computed from a single model trained from the annotations applied to the entire corpus of statements
Documented Applications
extractive summarization of scientific documents, using claim assertion detection and claim-type classification as part of a multi-stage classifier pipeline
Interested in licensing this patent?