Performance and scalability evaluation of the Ceph parallel file system

  • Authors:
  • Feiyi Wang;Mark Nelson;Sarp Oral;Scott Atchley;Sage Weil;Bradley W. Settlemyer;Blake Caldwell;Jason Hill

  • Affiliations:
  • Oak Ridge National Laboratory, Oak Ridge, Tennessee;Inktank Inc., Los Angeles, CA;Oak Ridge National Laboratory, Oak Ridge, Tennessee;Oak Ridge National Laboratory, Oak Ridge, Tennessee;Inktank Inc., Los Angeles, CA;Oak Ridge National Laboratory, Oak Ridge, Tennessee;Oak Ridge National Laboratory, Oak Ridge, Tennessee;Oak Ridge National Laboratory, Oak Ridge, Tennessee

  • Venue:
  • PDSW '13 Proceedings of the 8th Parallel Data Storage Workshop
  • Year:
  • 2013

Quantified Score

Hi-index 0.00

Visualization

Abstract

Ceph is an emerging open-source parallel distributed file and storage system. By design, Ceph leverages unreliable commodity storage and network hardware, and provides reliability and fault-tolerance via controlled object placement and data replication. This paper presents our file and block I/O performance and scalability evaluation of Ceph for scientific high-performance computing (HPC) environments. Our work makes two unique contributions. First, our evaluation is performed under a realistic setup for a large-scale capability HPC environment using a commercial high-end storage system. Second, our path of investigation, tuning efforts, and findings made direct contributions to Ceph's development and improved code quality, scalability, and performance. These changes should benefit both Ceph and the HPC community at large.