CPPC-G: fault-tolerant applications on the grid

  • Authors:
  • Daniel Díaz;Xoán C. Pardo;María J. Martín;Patricia González;Gabriel Rodríguez

  • Affiliations:
  • Computer Architecture Group, University of A Coruña, Spain;Computer Architecture Group, University of A Coruña, Spain;Computer Architecture Group, University of A Coruña, Spain;Computer Architecture Group, University of A Coruña, Spain;Computer Architecture Group, University of A Coruña, Spain

  • Venue:
  • PPAM'07 Proceedings of the 7th international conference on Parallel processing and applied mathematics
  • Year:
  • 2007

Quantified Score

Hi-index 0.00

Visualization

Abstract

The Grid community has made an important effort in developing middleware to provide different functionalities, such as resource discovery, resource management, job submission, execution monitoring. As part of this effort this paper addresses the design and implementation of an architecture (CPPC-G) based on services to manage the execution of fault tolerant applications on Grids. The CPPC (Controller/Precompiler for Portable Checkpointing) framework is used to insert checkpoint instrumentation into the application code. Designed services will be in charge of submission and monitoring of the execution of the application, management of checkpoint files and detection and automatic restart of failed executions.