Protein structure prediction system
Inventors
Blattner, Frederick R. • DARNELL, Steven J. • Larson, Matthew R. • MITCHELL, Amanda E. • Schroeder, John L.
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
The present invention is an accelerated conformational sampling method for predicting target peptide and protein structures comprising a process of determining energy minimized synthetic templates using a simple system for modeling individual molecular bonds within the subject peptide or protein. Use of these synthetic templates greatly reduces the computational resources necessary for optimally determining structural features of the target peptide or protein. The present invention also provides methods for rapid and efficient analysis of the effect of mutations on target peptides and proteins.
Core Innovation
The invention provides a conformational sampling method for predicting the structure of an amino acid sequence by creating a sequence profile matrix and determining an alignment of residues against the sequence profile matrix. Internal residue contacts are identified using the alignment, and features from the sequence profile matrix and the internal residue contacts are collected into a collected feature matrix. The collected feature matrix is aligned using one or more threading models with a structural feature database of original templates, and a plurality of optimally aligned original templates is selected.
For each optimally aligned original template, normal modes of motion are calculated and the original template is perturbed along each pair of calculated normal modes to collectively create a plurality of synthetic templates. The energy difference between each original template and its corresponding synthetic template is scored, and a subset of synthetic templates is selected based on satisfaction of a predetermined cut-off criterion. The selected synthetic templates are used to replace or supplement the original templates to generate a plurality of modeling templates.
Within the modeling templates, distance and contact restraints are calculated and Markov Chain Monte Carlo simulations are performed to obtain simulation results. The simulation results are clustered into a plurality of clusters, representative models are selected from each cluster, and the representative models are refined by energy minimization. The lowest energy refined representative model is selected as the predicted structure of the amino acid sequence.
Claims Coverage
The partial content includes three independent claims: a conformational sampling method, a computer system, and a non-transitory computer readable storage medium. Across these independent claims, the same core inventive features are covered, including generating synthetic templates via normal-mode perturbations and selecting them using an energy-difference score with a predetermined cut-off criterion, followed by distance/contact restraints and Markov Chain Monte Carlo with clustering and energy-minimized model selection.
Sequence profile, internal contacts, and threading against original templates
Creating a sequence profile matrix of the amino acid sequence; determining an alignment of each respective residue of the amino acid sequence against the sequence profile matrix; identifying internal residue contacts for one or more residues using the alignment; collecting features into a collected feature matrix; aligning the collected feature matrix using one or more threading models with a structural feature database of original templates; selecting a plurality of optimally aligned original templates.
Synthetic templates via normal-mode perturbations and energy-difference scoring
Calculating normal modes of motion for each original template; perturbing each respective original template for each pair of calculated normal modes to collectively create a plurality of synthetic templates; scoring the energy difference between each original template and the corresponding synthetic template; selecting a subset of synthetic templates from the plurality based on satisfaction of a predetermined cut-off criterion.
Modeling templates using selected synthetic templates, then distance/contact restraints
Replacing or supplementing the original templates with the corresponding selected subset of synthetic templates to generate a plurality of modeling templates; calculating distance and contact restraints within the modeling templates.
Markov Chain Monte Carlo with clustering and lowest-energy refined representative selection
Performing Markov Chain Monte Carlo simulations with the modeling templates to obtain simulation results; clustering the simulation results into a plurality of clusters representing models; selecting representative models of each cluster; refining the representative models by energy minimization; selecting the lowest energy refined representative model as the predicted structure.
All three independent claims (method, computer system, and non-transitory computer readable storage medium) share the same inventive pipeline: residue feature construction via sequence profile and internal residue contacts, threading to select original templates, generation of synthetic templates by normal-mode perturbations with energy-difference scoring and cut-off based selection, creation of modeling templates by replacing or supplementing templates, and Markov Chain Monte Carlo with clustering and energy-minimized representative model selection.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?