Abstract
Large-scale ad hoc analytics of genomic data is popular using the R-programming language supported by over 700 software packages provided by Bioconductor. More recently, analytical jobs are benefitting from on-demand computing and storage, their scalability and their low maintenance cost, all of which are offered by the cloud. While Biologists and Bioinformaticists can take an analytical job and execute it on their personal workstations, it remains challenging to seamlessly execute the job on the cloud infrastructure without extensive knowledge of the cloud dashboard. How analytical jobs can not only with minimum effort be executed on the cloud, but also how both the resources and data required by the job can be managed is explored in this paper. An open-source light-weight framework for executing R-scripts using Bioconductor packages, referred to as ‘RBioCloud’, is designed and developed. RBioCloud offers a set of simple command-line tools for managing the cloud resources, the data and the execution of the job. Three biological test cases validate the feasibility of RBioCloud. The framework is available from http://www.rbiocloud.com.
Original language | English |
---|---|
Pages (from-to) | 871-878 |
Number of pages | 9 |
Journal | IEEE/ACM Transactions on Computational Biology And Bioinformatics |
Volume | 12 |
Issue number | 4 |
DOIs | |
Publication status | Published - 2 Oct 2014 |
Keywords
- Cloud computing
- R programming
- Bioconductor
- Amazon Web Services
- Data analytics