PREDICTING CHOLESTEROL BINDING SITES IN MEMBRANE PROTEINS WITH MACHINE LEARNING MODELS

Loading...
Thumbnail Image

Date

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Cholesterol is the most well-known and abundant of all the sterols in mam- malians. Its unique chemical structure is involved in modulating the material properties of plasma membranes (e.g. membrane thickness, bending rigidity), while influencing the functions of membrane proteins, even acting as an allosteric agonist or antagonist. These specific interactions, although similar to those of proteins with small drug-like molecules, have important differences due to cholesterol concentration and ordering within the membrane. Traditionally, the study of cholesterol-protein interactions has been done through the lenses of binding motifs, sequences of amino acids identified to be present in cholesterol-binding proteins. Nevertheless, there is inconclusive evidence that this is a necessary condition for interaction, and even more, motifs have been observed in proteomes with no known cholesterol interactions. In this work, we tackle these shortcomings and approach the problem of pre- dicting protein-cholesterol binding sites by combining novel deep-learning methods for protein representation, experimental structures for membrane proteins with choles- terol bound, and molecular dynamics simulations of predicted and observed cholesterol binding sites. We tested two different algorithms. We first tried GearNet, a Graph Neural Network model which uses protein coordinates to build a protein graph, and at- tempted to train the model to perform a binary node-level classification to discriminate cholesterol interacting residues from non-interacting residues. However, our model was unable to capture the relationship between amino acids that make up a binding pocket, and therefore could not predict cholesterol binding residues. We then tested DiffDock, a generative diffusion model that was designed for predicting small molecule binding sites. We applied it directly without modification to predict cholesterol binding sites, given as input the protein coordinates. In this case, the model’s predictions were quite good, reproducing many observed cholesterol binding sites. These predictions were further validated with MD simulations of a selected group of proteins. For each case, we simulated both the experimentally bound pose and the DiffDock predicted pose. Our simulation results showed that in most cases, the predicted poses remain stably bound to the protein, and the resulting binding pocket matched well to that of the experimental control simulation. These results establish a proof of concept for the combination of diffusion model prediction of cholesterol binding sites and molecular simulations to test these predic- tions. Further work will more extensively test this workflow, particularly regarding elimination of false positive predictions of cholesterol binding sites.

Description

Keywords

Citation

Endorsement

Review

Supplemented By

Referenced By