Using Hadoop File System and MapReduce in a small/medium Grid site

H Riahi,Giacinto Donvito,Livio Fano,Massimiliano Fasi,Giovanni Marzulli,Daniele Spiga,Andrea Valentini

Using Hadoop File System and MapReduce in a small/medium Grid site

2012

Data storage and data access represent the key of CPU-intensive and data-intensive high performance Grid computing. Hadoop is an open-source data processing framework that includes fault-tolerant and scalable distributed data processing model and execution environment, named MapReduce, and distributed File System, named Hadoop distributed File System (HDFS). HDFS was deployed and tested within the Open Science Grid (OSG) middleware stack. Efforts have been taken to integrate HDFS with gLite middleware. We have tested the File System thoroughly in order to understand its scalability and fault-tolerance while dealing with small/medium site environment constraints. To benefit entirely from this File System, we made it working in conjunction with Hadoop Job scheduler to optimize the executions of the local physics analysis workflows. The performance of the analysis jobs which used such architecture seems to be promising, making it useful to follow up in the future.

Keywords:

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations