Systems and methods for de novo peptide sequencing from data-independent acquisition using deep learning
Inventors
SHAN, Baozhen • TRAN, Ngoc Hieu • Li, Ming • XIN, Lei • Qiao, Rui • Chen, Xin • Liu, Chuyi
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
The present systems and methods introduce deep learning to de novo peptide sequencing from tandem mass spectrometry data, and in particular mass spectrometry data obtained by data-independent acquisition. The systems and methods achieve improvements in sequencing accuracy over existing systems and methods and enables complete assembly of novel protein sequences without assisting databases. To sequence peptides from mass spectrometry data obtained by data-independent acquisition, precursor profiles representing intensities of one or more precursor ion signals associated with a precursor retention time and fragment ion spectra representing signals from fragment ions and fragment retention times are fed into a neural network.
Core Innovation
The invention provides a computer implemented system for de novo sequencing of a peptide from mass spectrometry data acquired by data-independent acquisition using neural networks. A first input represents at least one precursor profile with precursor retention time, and a second input represents a plurality of fragment ion spectra for each precursor profile with fragment retention time. The neural network generates a probability measure for one or more candidates to a next amino acid in an amino acid sequence, and a processor updates a determined amino acid sequence to generate an output signal representing a final determined sequence.
The system uses precursor profiles and fragment ion spectra to restructure DIA extra dimensionality into learned three-dimensional fragment-ion shapes and precursor-fragment correlations. The layered artificial neural network uses at least one convolutional layer configured to filter mass spectrometry spectrum data to detect fragment ion peaks, and it is trained on mass spectrometry data containing retention time and fragment ion peaks of sequences differing in length and differing by one or more amino acids.
For sequencing, the processor receives an input prefix representing a determined amino acid sequence, provides the mass spectrometry spectrum data based on the first and second inputs to the plurality of layered nodes, and identifies the next amino acid by selecting a candidate next amino acid with the greatest probability measure. The layered nodes receive matrix data representing the mass spectrometry spectrum data and output a probability measure vector, and the document additionally describes training using focal loss and selecting a set of fragment ion spectra nearest the precursor retention time.
Claims Coverage
The partial content includes three independent claims: a system for de novo peptide sequencing using neural networks, a method for the same task, and non-transitory computer-readable media storing instructions to perform the method. Across these independent claims, the core inventive features center on receiving precursor profiles and fragment ion spectra with retention times, using a layered artificial neural network with at least one convolutional layer to detect fragment ion peaks and generate probability measures for candidate next amino acids, and iteratively selecting the next amino acid with the greatest probability measure while updating an amino-acid sequence and outputting a final determined sequence.
Neural-network de novo peptide sequencing from DIA precursor profiles and fragment ion spectra
A system configured to receive a first input representing at least one precursor profile with precursor retention time and intensities, and a second input representing a plurality of fragment ion spectra for each precursor profile with signals from fragment ions and a fragment retention time, and to generate an output signal representing a final determined sequence by identifying a next amino acid from a probability measure produced by an artificial neural network.
Convolutional filtering of fragment-ion peaks to generate next-amino-acid probability measures
The plurality of layered nodes includes at least one convolutional layer configured to filter mass spectrometry spectrum data to detect fragment ion peaks, and the neural network is trained on mass spectrometry data containing retention time and a plurality of fragment ion peaks of sequences differing in length and differing by one or more amino acids, producing a probability measure for one or more candidates to a next amino acid.
Matrix-based neural network inputs and next-amino-acid selection using greatest probability
The plurality of layered nodes receives matrix data representing the mass spectrometry spectrum data and outputs a probability measure vector, wherein the second input comprises matrix data including batch size, number of amino acids, ion types, number of fragment ion spectra associated with a precursor profile, and window size for filtering fragment ion peaks; the processor identifies a next amino acid based on a candidate next amino acid having a greatest probability measure from the neural network output.
Iterative prefix-driven sequencing and final determined sequence output
The processor receives an input prefix representing a determined amino acid sequence of the peptide, provides the mass spectrometry spectrum data to the plurality of layered nodes, updates the determined amino acid sequence with the next amino acid, and generates an output signal representing a final determined sequence based on the neural network probability measures.
Method of de novo sequencing using convolutional layered nodes to detect fragment-ion peaks
A method that receives a precursor profile input with precursor retention time and intensities and a fragment ion spectra input with fragment retention time, filters the mass spectrometry spectrum data to detect fragment ion peaks by at least one convolutional layer of a plurality of layered nodes configured as an artificial neural network to generate a probability measure for candidate next amino acids, obtains an input prefix, identifies the next amino acid by greatest probability, updates the determined sequence, and generates an output signal representing a final determined sequence.
Non-transitory computer readable media implementing the layered convolutional sequencing method
Non-transitory computer-readable media storing instructions that cause a processor to perform receiving precursor profiles and fragment ion spectra inputs, filtering spectrum data using at least one convolutional layer of layered nodes to detect fragment ion peaks and generate probability measures for candidates to a next amino acid, obtaining an input prefix, providing spectrum data to the layered nodes, identifying the next amino acid using a greatest probability measure, updating the determined sequence, and generating an output signal representing a final determined sequence.
Across the independent claims, the coverage is directed to de novo peptide sequencing from DIA mass spectrometry using neural networks that take precursor profiles and fragment ion spectra with retention times, use layered nodes with at least one convolutional layer to filter and detect fragment ion peaks, and output probability measures for candidate next amino acids. Sequencing is carried out by iteratively selecting the next amino acid with the greatest probability measure based on a prefix-driven determined sequence and producing a final determined sequence output.
Stated Advantages
Improved sequencing accuracy.
Identification of novel peptides.
Validated results using confidence-score filtering and augmented database re-search, including 1% FDR filtering.
Documented Applications
Evaluation of de novo peptide sequencing on datasets including ovarian cyst, urinary tract infection, plasma, HLA peptides, and Jurkat (Jurkat-Oxford).
Interested in licensing this patent?