Systems and methods for training a statistical model to predict tissue characteristics for a pathology image
Inventors
Beck, Andrew H. • Khosla, Aditya
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
In some aspects, the described systems and methods provide for a method for training a statistical model to predict tissue characteristics for a pathology image. The method includes accessing annotated pathology images. Each of the images includes an annotation describing a tissue characteristic category for a portion of the image. A set of training patches and a corresponding set of annotations are defined using an annotated pathology image. Each of the training patches in the set includes values obtained from a respective subset of pixels in the annotated pathology image and is associated with a corresponding patch annotation determined based on an annotation associated with the respective subset of pixels. The statistical model is trained based on the set of training patches and the corresponding set of patch annotations. The trained statistical model is stored on at least one storage device.
Core Innovation
The invention provides a method for predicting an entity of interest for a pathology image by using a two-stage statistical modeling approach. It accesses a pathology image, retrieves a first trained statistical model trained on a plurality of annotated pathology images, and processes the pathology image with the first trained statistical model to generate an annotated pathology image including one or more annotations describing tissue characteristics for one or more portions of the image. The method then extracts values for one or more features from the annotated pathology image.
The extracted feature values are used as input to a second trained statistical model that is trained on extracted values from the plurality of annotated pathology images. The second trained statistical model is then used to predict an entity of interest selected from a group consisting of survival time, drug response, patient level phenotype/molecular characteristics, mutational burden, tumor molecular characteristics, transcriptomic features, protein expression features, and patient clinical outcomes. The predicted entity of interest is stored on the at least one storage device.
In some embodiments, the first trained statistical model is a convolutional neural network with multiple layers and no padding applied to the output of any layer, and the convolutional layers are configured so that (N−K)/S is an integer based on input dimension size N, convolution filter size K, and stride size S. The annotated pathology output can be processed to associate predicted annotations with portions of the pathology image and to store the resulting associations to generate an annotated pathology image. Feature values can include selected histological features such as area measurements, numbers of mitotic figures, nuclear grade metrics, and distance metrics between defined cell types and structures.
Claims Coverage
Independent claim clm-00001 defines a complete pipeline with two trained statistical models: a first model generates annotations on a pathology image, and a second model predicts an entity of interest from extracted feature values. The dependent claims refine the first model architecture and constraints, the feature/annotation representation, the allowable second-model types, and how predicted annotations are associated with image portions.
Two-stage statistical-model prediction from annotated pathology
Accessing a pathology image; retrieving a first trained statistical model trained on a plurality of annotated pathology images where each image includes at least one annotation describing tissue characteristics for one or more portions of the image; processing the pathology image with the first trained statistical model to generate an annotated pathology image including one or more annotations; extracting values for one or more features from the annotated pathology image; retrieving a second trained statistical model trained on the extracted values from the plurality of annotated pathology images; processing the extracted values with the second trained statistical model to predict an entity of interest selected from a group consisting of survival time, drug response, patient level phenotype/molecular characteristics, mutational burden, tumor molecular characteristics, transcriptomic features, protein expression features, and patient clinical outcomes; and storing the predicted entity of interest on the at least one storage device.
Convolutional neural network with no padding on layer outputs
The method wherein the first trained statistical model is a convolutional neural network with multiple layers and no padding is applied to the output of any layer.
Integer alignment constraint (N−K)/S for convolution layers
The method wherein the first trained statistical model is a convolutional neural network with layers configured so that (N−K)/S is an integer, where N is the input dimension size, K is the convolution filter size, and S is the stride size.
Histological feature extraction including area, grade, counts, and inter-structure distances
The method using one or more selected histological features including area measurements, counts including mitotic figures, nuclear grade metrics, and distance measurements between defined cell types and structures.
Second trained statistical model selected from generalized linear model, random forest, support vector machine, and/or gradient boosted tree
The method wherein the second trained statistical model is a generalized linear model, random forest, support vector machine, and/or gradient boosted tree.
Associating predicted annotations to image portions and storing the associations
The method wherein the first trained statistical model processes the pathology image, associates portions of the image with predicted annotations, and stores the associations on at least one storage device to generate an annotated pathology image.
Overall, the claims cover a two-stage prediction framework in which a first trained model generates tissue-characteristic annotations on pathology images, feature values are extracted from those annotations, and a second trained model predicts a selected entity of interest such as survival time or drug response from the extracted feature values. Refinements include CNN constraints (no padding; (N−K)/S integer alignment), specific histological features for extraction, specific allowable second-model types, and storage of associations between predicted annotations and image portions.
Stated Advantages
Documented Applications
No documented applications found
Interested in licensing this patent?