Graywulf: a platform for federated scientific databases and services

Authors:
László Dobos;István Csabai;Alexander S. Szalay;Tamás Budavári;Nolan Li
Affiliations:
Eötvös Loránd University, Budapest, Hungary;Eötvös Loránd University, Budapest, Hungary;The Johns Hopkins University, Baltimore, MD;The Johns Hopkins University, Baltimore, MD;The Johns Hopkins University, Baltimore, MD
Venue:
Proceedings of the 25th International Conference on Scientific and Statistical Database Management
Year:
2013

Citing 13
Cited 0

The multidimensional database system RasDaMan

SIGMOD '98 Proceedings of the 1998 ACM SIGMOD international conference on Management of data
Astronomical archives of the future: a virtual observatory

Future Generation Computer Systems
Building a prototype for network measurement virtual observatory

Proceedings of the 3rd annual ACM workshop on Mining network data
CasJobs and MyDB: A Batch Query Workbench

Computing in Science and Engineering
The Catalog Archive Server Database Management System

Computing in Science and Engineering
The sqlLoader Data-Loading Pipeline

Computing in Science and Engineering
Wireless sensor networks for soil science

International Journal of Sensor Networks
SciQL, a query language for science applications

Proceedings of the EDBT/ICDT 2011 Workshop on Array Databases
Array requirements for scientific applications and an implementation for microsoft SQL server

Proceedings of the EDBT/ICDT 2011 Workshop on Array Databases
The architecture of SciDB

SSDBM'11 Proceedings of the 23rd international conference on Scientific and statistical database management
Implementing a general spatial indexing library for relational databases of large numerical simulations

SSDBM'11 Proceedings of the 23rd international conference on Scientific and statistical database management
The Future of Scientific Data Bases

ICDE '12 Proceedings of the 2012 IEEE 28th International Conference on Data Engineering
SkyQuery: an implementation of a parallel probabilistic join engine for cross-identification of multiple astronomical databases

SSDBM'12 Proceedings of the 24th international conference on Scientific and Statistical Database Management

Quantified Score

Hi-index	0.00

Visualization

Abstract

Many fields of science rely on relational database management systems to analyze, publish and share data. Since RDBMS are originally designed for, and their development directions are primarily driven by, business use cases they often lack features very important for scientific applications. Horizontal scalability is probably the most important missing feature which makes it challenging to adapt traditional relational database systems to the ever growing data sizes. Due to the limited support of array data types and metadata management, successful application of RDBMS in science usually requires the development of custom extensions. While some of these extensions are specific to the field of science, the majority of them could easily be generalized and reused in other disciplines. With the Graywulf project we intend to target several goals. We are building a generic platform that offers reusable components for efficient storage, transformation, statistical analysis and presentation of scientific data stored in Microsoft SQL Server. Graywulf also addresses the distributed computational issues arising from current RDBMS technologies. The current version supports load balancing of simple queries and parallel execution of partitioned queries over a set of mirrored databases. Uniform user access to the data is provided through a web based query interface and a data surface for software clients. Queries are formulated in a slightly modified syntax of SQL that offers a transparent view of the distributed data. The software library consists of several components that can be reused to develop complex scientific data warehouses: a system registry, administration tools to manage entire database server clusters, a sophisticated workflow execution framework, and a SQL parser library.