Methods and systems for assembly of protein sequences
Inventors
TRAN, Ngoc Hieu • RAHMAN, Mohammad Ziaur • He, Lin • XIN, Lei • SHAN, Baozhen • Li, Ming
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Methods and systems for determining amino acid sequence of a polypeptide or protein from mass spectrometry data is provided, using a weighted de Bruijn graph. Extracted and purified protein is cleaved into a mixture of peptide and then analyzed using mass spectrometry. A list of peptide sequences is derived from mass spectrometry fragment data by de novo sequencing, and amino acid confidence scores are determined from peak fragment ion intensity. A weighted de Bruijn graph is constructed for the list of peptide sequences having node weights defined by k−1 mer confidence scores. At least one contig is assembled from the de Bruijn graph by identifying node weights having the highest k−1 mer confidence scores.
Core Innovation
The invention determines an amino acid sequence of a polypeptide from mass spectrometry data by generating mass spectrometry fragment ion data of peptides cleaved from the polypeptide. The fragment ion data are converted into a first list of peptide sequences as sequence reads of the mass spectrometry fragment ions, and for each amino acid position in the peptide sequences, the system determines an amino acid confidence score as a probability that an amino acid at that position is correct based on a corresponding fragment ion intensity.
A first weighted de Bruijn graph is constructed from the peptide sequence reads by identifying a k-amino acid length substring and an adjacent k-amino acid length substring with overlapping sequence of length k-1. Each k-1 mer is assigned as a node, and relationships between k-1 mers are represented as paths connecting the nodes, where multiple possible relationships are represented by junctions with multiple paths. Each path is assigned a k-1 mer confidence score according to Equation I using peptide intensities based on logarithm of precursor intensity and using amino-acid confidence and a weight ratio.
The amino-acid weight ratio assigns higher weight to amino acids at both ends of the k-1 mer than to amino acids at middle positions of the k-1 mer. At each junction, contigs are assembled by connecting nodes using paths having the highest k-1 mer confidence score to assemble at least one contig from the first weighted de Bruijn graph.
In the described applications, an integrated system (ALPS) is applied to monoclonal antibody light/heavy chains, using hybrid peptide spectrum match sets with database search (PEAKS DB) and homology/variant detection (SPIDER) with false discovery rate filtering. The approach reports full-length assembly for light and heavy chains and quantified accuracy/coverage improvements compared to de novo-only, including 100% accuracy for light chains and approximately 99.09% accuracy for one heavy chain.
Claims Coverage
The independent claims cover a computer implemented system and a corresponding method that determine an amino acid sequence from mass spectrometry data using a weighted de Bruijn graph with per-amino-acid confidence scoring and junction path selection. Across the independent claims, the inventive features include fragment-ion-to-sequence-read conversion, per-amino-acid confidence scores derived from fragment ion intensity, weighted de Bruijn graph construction using k-mers/k-1-mers with k-1-mer confidence scoring based on Equation I, and contig assembly by selecting highest-confidence k-1-mer paths at junctions.
Weighted de Bruijn graph amino-acid sequence determination from mass spectrometry fragment ions
A processor converts mass spectrometry fragment ion data of peptides into a first list of peptide sequences as sequence reads of the mass spectrometry fragment ions, determines an amino acid confidence score for each amino acid using corresponding fragment ion intensity as a probability of correctness, constructs a first weighted de Bruijn graph by mapping k-mers and adjacent overlapping k-1-mers into nodes and paths, assigns each path a k-1-mer confidence score according to Equation I using amino-acid confidence and a weight ratio, and assembles at least one contig by connecting nodes using paths having the highest k-1-mer confidence score at each junction.
Weighted de Bruijn graph contig assembly for amino-acid sequence from mass spectrometry method
A method purifies a polypeptide and cleaves it into peptides, analyzes the peptides through mass spectrometry to obtain mass spectrometry fragment ion data, converts the fragment ion data into a first list of peptide sequences as sequence reads, determines an amino acid confidence score for each amino acid based on corresponding fragment ion intensity as a probability of correctness, constructs a first weighted de Bruijn graph by identifying k-mers and adjacent k-1-mers and representing nodes and paths with junctions for multiple relationships, assigns each path a k-1-mer confidence score according to Equation I using amino-acid confidence and end-weighting of amino acids, and assembles at least one contig by connecting nodes using paths having the highest k-1-mer confidence score at each junction.
Both independent claims require deriving sequence reads and per-amino-acid confidence from mass spectrometry fragment ion intensity, building a weighted de Bruijn graph whose k-1-mer path confidences follow Equation I with end-weighted amino acids, and assembling contigs by selecting the highest-confidence paths at junctions.
Stated Advantages
Reports full-length assembly for monoclonal antibody light and heavy chains.
Quantified accuracy/coverage improvements compared to de novo-only.
100% accuracy for light chains.
Approximately 99.09% accuracy for one heavy chain.
Documented Applications
Monoclonal antibody light/heavy chain amino-acid sequence determination using an integrated system (ALPS) with hybrid peptide spectrum match sets (PEAKS DB and SPIDER) and false discovery rate filtering.
Interested in licensing this patent?