Architecting dependable systems with proactive fault management

  • Authors:
  • Felix Salfner;Miroslaw Malek

  • Affiliations:
  • Humboldt-Universität zu Berlin, Institut für Informatik, Berlin, Germany;Humboldt-Universität zu Berlin, Institut für Informatik, Berlin, Germany

  • Venue:
  • Architecting dependable systems VII
  • Year:
  • 2010

Quantified Score

Hi-index 0.00

Visualization

Abstract

Management of an ever-growing complexity of computing systems is an everlasting challenge for computer system engineers. We argue that we need to resort to predictive technologies in order to harness the system's complexity and transform a vision of proactive system and failure management into reality. We describe proactive fault management, provide an overview and taxonomy for online failure prediction methods and present a classification of failure prediction-triggered methods. We present a model to assess the effects of proactive fault management on system reliability and show that overall dependability can significantly be enhanced. After having shown the methods and potential of proactive fault management we describe a blueprint how proactive fault management can be incorporated into a dependable system's architecture.