Method for identifying microorganisms by mass spectrometry and score normalization
Inventors
Strubel, Grégory • Arsac, Maud • Desseree, Denis • Cotte-Pattat, Pierre-Jean • Mahe, Pierre
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
An identification by mass spectrometry of a microorganism from among reference microorganisms represented by reference data sets includes: determining a set of data of the microorganism according to a spectrum; for each reference microorganism, calculating a distance between the determined and reference sets; and calculating a probability ƒ(m) according to relationf(m)=pN(m|μ,σ)pN(m|μ,σ)+(1-p)N(m|μ_,σ_) where: m is the distance calculated for the reference microorganism; N(m|μ,σ) is the value, for m, of a random variable modeling the distance between a reference microorganism to be identified and the reference microorganism, when the microorganism is the reference microorganism; N(m|μ,σ) is the value, for m, of a random variable modeling the distance between a microorganism to be identified and the reference microorganism, when the microorganism is not the reference microorganism; and p is a scalar in the range from 0 to 1.
Core Innovation
The invention relates to identifying by mass spectrometry a microorganism from among a predetermined set of reference microorganisms. Each reference microorganism is represented by a set of reference data obtained from mass spectrometry measurements, and a data processing unit processes a set of data representative of the microorganism to be identified from a mass spectrometry measurement acquired from a mass spectrometer.
For each reference microorganism, the method calculates a distance between the determined set of data and the set of reference data using classification tools, and the distance is representative of the microorganism to be identified. The method then calculates a probability that the microorganism to be identified is the reference microorganism using a probability relation f(m), with the distance modeled by two Gaussian random variables N(m|μ,σ) and N(m|μ′,σ′).
After calculating probabilities for the microorganism to be identified for each reference microorganism, the method compares the probabilities and provides an identification decision based on the comparing. The disclosed approach supports calibration/training separation, vectorization of spectra into a vectorial space, and classification frameworks such as OVA SVM or tolerant-distance/vector-boundary classification to resolve ambiguity among similarly matching reference species.
Learned probability parameters and/or boundaries are stored in a knowledge base, and the method can use a threshold-based “none of the reference microorganisms” rejection, together with normalization of similarity measures to obtain reliable, comparable similarity scores.
Claims Coverage
The independent claim covers a complete identification pipeline in which distances between mass-spectrometry-derived data and reference data are transformed into probabilities using a two-Gaussian random-variable model, followed by probability comparison to provide an identification decision. The claim is built from three inventive features covering the distance computation, the Gaussian probability transformation, and the decision by comparing probabilities across all reference microorganisms.
Mass-spectrometry distance against reference microorganism data
Calculating, for each reference microorganism, a distance between the determined set of data representative of the microorganism to be identified and the set of reference data of the reference microorganism, said distance being representative of the microorganism to be identified.
Two-Gaussian random-variable probability mapping of distance
Calculating, for each reference microorganism, a probability for the microorganism to be identified to be the reference microorganism, according to relation f(m)=p·N/(p·N+(1−p)·N'), where N and N' model the distance m with Gaussian random variables N(m|μ,σ) and N(m|μ',σ').
Probability comparison to provide an identification decision
Comparing, for the microorganism to be identified, the calculated probabilities for the microorganism to be identified to be each of the reference microorganisms, and providing an identification decision for the microorganism based on the comparing of the calculated probabilities.
Overall, the claim coverage centers on converting each reference-specific distance into a reference-vs-non-reference probability using two Gaussian random-variable models and then selecting an identification decision by comparing the probabilities across all reference microorganisms.
Stated Advantages
Provides reliable, comparable similarity scores through normalization of similarity scores.
Resolves ambiguity among similarly matching reference species.
Supports threshold-based “none of the reference microorganisms” rejection.
Documented Applications
No documented applications found
Interested in licensing this patent?