Yemanja—A Layered Fault Localization System for Multi-Domain Computing Utilities

  • Authors:
  • K. Appleby;G. Goldszmidt;M. Steinder

  • Affiliations:
  • IBM T.J. Watson Research Center, 30 Saw Mill River Road, Hawthorne, New York 10532/ applebyk@us.ibm.com;IBM T.J. Watson Research Center, 30 Saw Mill River Road, Hawthorne, New York 10532;Computer and Information Sciences, University of Delaware, Newark, Delaware 19716

  • Venue:
  • Journal of Network and Systems Management
  • Year:
  • 2002

Quantified Score

Hi-index 0.00

Visualization

Abstract

Yemanja is a model-based event correlation engine for multi-layer fault diagnosis. It targets complex propagating fault scenarios, and can smoothly correlate low-level network events with high-level application performance alerts related to quality-of-service violations. Entity-models that represent devices or abstract components encapsulate their behavior. Distantly associated entity-models are not explicitly aware of each other, and communicate through internal event chains. Yemanja's state-based engine supports generic scenario definitions, prioritization of alternate solutions, integrated problem and device testing, and simultaneous analysis of overlapping problems. The system of correlation rules was developed based on the analysis of device and layer functions, and the dependencies among physical and abstract system components. The primary objectives of this research include the development of reusable, configuration independent, correlation scenarios, adaptability and extensibility of the engine to match the constantly changing topology of a multi-domain server farm, and development of a concise specification language that is relatively simple yet powerful.