Bwasw-Cloud: Efficient sequence alignment algorithm for two big data with MapReduce.
Bwasw-Cloud: Efficient sequence alignment algorithm for two big data with MapReduce.The recent next-generation sequencing machines generate sequences at an unprecedented rate, and a sequence is not short any more called read. The reference sequences which are aligned reads against are also increasingly large. Efficiently mapping large number of long sequences with big reference sequences poses a new challenge to sequence alignment. Sequence alignment algorithms become to match on two big data.
To address the above problem, we propose a new parallel sequence alignment algorithm called Bwasw-Cloud, optimized for aligning long reads against a large sequence data (e.g. the human genome). It is modeled after the widely used BWA-SW algorithm and uses the open-source Hadoop implementation of MapReduce. The results show that Bwasw-Cloud can effectively and quickly match two big data in common cluster.
Similar IEEE Project Titles
- Towards a Collective Layer in the Big Data Stack.
- Probe into setting up big data processing specialty in Chinese universities
- Language based web crawling on big data.
- A Big Data Architecture for Large Scale Security Monitoring.
- Experiments with computing similarity coefficient over big data.