The amount of RDF data being published on the Web is increasing at a massive rate. MapReduce-based distributed frameworks have become the general trend in processing SPARQL queries against the RDF data. Currently, query processing systems that use MapReduce have not been able to keep up with increases in semantic annotated data, resulting in non-interactive SPARQL query processing. The principal reason is that intermediate query results from join operations in a MapReduce framework are so massive that network bandwidth and hard disk drive I/O speeds may not keep pace with the processing speed. In this paper, we present an efficient SPARQL processing system that uses MapReduce and HBase. The system runs a job optimized query plan using our proposed abstract RDF data to decrease the amount of intermediate data, thus resulting in faster query processing performance. We also present an efficient algorithm of using Map-side joins while also using the abstract RDF data to filter out unneeded RDF data. Experimental results show that the proposed approach demonstrates better performance when processing queries with a large set of inputs than those found in previous works.
|Title of host publication||Proceedings - 2015 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology, WI-IAT 2015|
|Publisher||Institute of Electrical and Electronics Engineers Inc.|
|Number of pages||8|
|Publication status||Published - 2016 Feb 2|
|Event||2015 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology Workshops, WI-IAT Workshops 2015 - Singapore, Singapore|
Duration: 2015 Dec 6 → 2015 Dec 9
|Name||Proceedings - 2015 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology, WI-IAT 2015|
|Other||2015 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology Workshops, WI-IAT Workshops 2015|
|Period||15/12/6 → 15/12/9|
Bibliographical notePublisher Copyright:
© 2015 IEEE.
All Science Journal Classification (ASJC) codes
- Computer Networks and Communications