Systems and methods for de novo peptide sequencing using deep learning and spectrum pairs

Inventors

Qiao, RuiTRAN, Ngoc HieuXIN, LeiChen, XinSHAN, BaozhenGhodsi, AliLi, Ming

Assignees

BIOINFORMATICS SOLUTIONS Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11644470-B2

Patent

Publication Date

2023-05-09

Expiration Date


Abstract

The present systems and methods are directed to de novo identification of peptide sequences from tandem mass spectrometry data. The systems and methods uses unconverted mass spectrometry data from which features are extracted. Using unconverted mass spectrometry data reduces the loss of information and provides more accurate sequencing of peptides. The systems and methods combine deep learning and neural networks to sequencing of peptides.

Core Innovation

The invention concerns a computer implemented system and corresponding method for de novo sequencing of peptides from a sample mass spectrometry data using neural networks. The system generates a probability measure for one or more candidates to a next amino acid in an amino acid sequence by using a plurality of layered nodes configured as an artificial neural network, and the neural network is trained on known mass spectrometry data representing theoretical fragment ion peaks of known sequences differing in length and differing by one or more amino acids.

The system receives sample mass spectrometry data comprising a plurality of coordinate data pairs representing m/z and intensity values of observed fragment ion peaks and receives second coordinate data representing theoretical fragment ion peaks. The layered nodes include at least one convolutional and/or fully connected layer for comparing at least one of the observed fragment ion peaks against at least one of the theoretical fragment ion peaks, and identifies the difference in m/z between the observed fragment ion peaks and the theoretical fragment ion peaks by constraining a difference tensor (D).

After the peptide-sequencing neural network produces output, the system obtains an input prefix representing a determined amino acid sequence of the peptide. The system identifies a next amino acid based on a candidate next amino acid having a greatest probability measure based on the output of the artificial neural network and the sample mass spectrometry data, and updates the determined amino acid sequence with the next amino acid. In the described architecture, an LSTM is used as a language model to iteratively predict the next amino acid symbol from an input prefix and update the peptide sequence, with learned representations including a difference tensor constraint and an m/z positional embedding.

Claims Coverage

The partial content provides two independent claims: one system claim and one method claim. Across these independent claims, the main inventive features involve neural-network-based comparison of observed and theoretical fragment ion peaks from sample mass spectrometry coordinate pairs, constrained difference-tensor matching of m/z differences, and iterative next-amino-acid probability prediction from an input prefix to update a determined amino acid sequence.

Neural-network de novo peptide sequencing from sample mass spectrometry coordinate pairs

A computer implemented system configured to de novo sequence peptides from sample mass spectrometry data using neural networks, where a plurality of layered nodes form an artificial neural network that generates a probability measure for one or more candidates for a next amino acid in an amino acid sequence, and where the artificial neural network is trained on known mass spectrometry data representing theoretical fragment ion peaks of known sequences.

Observed-versus-theoretical fragment ion comparison via convolutional and/or fully connected layers

The plurality of layered nodes is configured to receive sample mass spectrometry data comprising coordinate data pairs representing m/z and intensity values of observed fragment ion peaks and receive second coordinate data representing theoretical fragment ion peaks, and to include at least one convolutional and/or fully connected layer comparing at least one observed fragment ion peak against at least one theoretical fragment ion peak.

Constrained difference tensor m/z matching for theoretical fragment ion peak identification

The plurality of layered nodes identifies the difference in m/z between observed fragment ion peaks and theoretical fragment ion peaks and identifies a matching theoretical fragment ion peak to the at least one observed fragment ion peak by constraining a difference tensor (D) representing the difference in m/z between the observed and theoretical fragment ion peaks.

Iterative next-amino-acid probability prediction from an input prefix

The processor obtains an input prefix representing a determined amino acid sequence, identifies a next amino acid based on a candidate next amino acid having a greatest probability measure based on the output of the artificial neural network and the sample mass spectrometry data, and updates the determined amino acid sequence with the next amino acid.

Layered-nodes method for de novo peptide sequencing with probability measures for candidate next amino acids

A method for de novo sequencing peptides from mass spectrometry data using a plurality of layered nodes configured to form an artificial neural network for generating a probability measure for one or more candidates to a next amino acid in an amino acid sequence, including receiving sample mass spectrometry data with coordinate data pairs for observed fragment ion peaks and receiving second coordinate data representing theoretical fragment ion peaks.

Matching theoretical fragment ion peaks using constrained difference tensor m/z differences

The method compares at least one observed fragment ion peak against at least one theoretical fragment ion peak by convolutional and/or fully connected layers, identifies the difference in m/z between observed and theoretical fragment ion peaks, identifies a matching theoretical fragment ion peak by constraining a difference tensor (D) representing the difference in m/z, and outputs an amino acid sequence corresponding to the at least one observed fragment ion peak.

Candidate next-amino-acid selection and sequence updating

The method obtains an input prefix representing a determined amino acid sequence, outputs a probability measure for each candidate of a next amino acid, identifies a next amino acid based on a candidate next amino acid having a greatest probability measure based on the output of the artificial neural network and the sample mass spectrometry data, and updates the determined amino acid sequence with the next amino acid.

Together, the independent claims cover de novo peptide sequencing by using layered neural networks to compare observed and theoretical fragment ion peaks represented as m/z and intensity coordinate pairs, constrain a difference tensor (D) to identify matching theoretical fragment ion peaks, and iteratively select the next amino acid by choosing the candidate with the greatest probability measure based on an input prefix and the neural network outputs.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.