Understanding signal sequences with machine learning

  • Authors:
  • Jean-Luc Falcone;Renée Kreuter;Dominique Belin;Bastien Chopard

  • Affiliations:
  • Département d'informatique, Université de Genève, Genève 4, Switzerland;Département de Pathologie et d'Immunologie, Université de Genève, Switzerland;Département de Pathologie et d'Immunologie, Université de Genève, Switzerland;Département d'informatique, Université de Genève, Genève 4, Switzerland

  • Venue:
  • EvoBIO'07 Proceedings of the 5th European conference on Evolutionary computation, machine learning and data mining in bioinformatics
  • Year:
  • 2007

Quantified Score

Hi-index 0.00

Visualization

Abstract

Protein translocation, the transport of newly synthesized proteins out of the cell, is a fundamental mechanism of life. We are interested in understanding how cells recognize the proteins that are to be exported and how the necessary information is encoded in the so called "Signal Sequences". In this paper, we address these problems by building a physico-chemical model of signal sequence recognition, using experimental data. This model was built using decision trees. In a first phase the classifier were built from a set of features derived from the current knowledge about signal sequences. It was then expanded by feature generation with genetic algorithms. The resulting predictors are efficient, achieving an accuracy of more than 99% with our wild-type proteins set. Furthermore the generated features can give us a biological insight about the export mechanism. Our tool is freely available through a web interface.