Data Placement and Query Processing Based on RPE Parallelisms

Yaxin Yu; Guoren Wang; Ge Yu; Gang Wu; Junan Hu; Nan Tang

Data Placement and Query Processing Based on RPE Parallelisms

Yaxin Yu^*, Guoren Wang, Ge Yu, Gang Wu, Junan Hu, Nan Tang

^*Corresponding author for this work

Northeastern University China

Research output: Contribution to journal › Conference article › peer-review

4 Citations (Scopus)

Abstract

The basic idea behind parallel database systems is to perform operations in parallel to reduce the response time and improve the system throughput. Data placement is a key factor on the performance of parallel database systems. This paper proposes two data partition strategies to decluster XML documents with very large size, Path Schema based Path Instance Balancing (PSPIB) strategy, in which all path instances with the same path schema in a data tree are declustered evenly over all sites, and Node Schema based Node Round-Robin (NSNRR) strategy, in which all node objects with the same node schema in a data tree are declustered over all sites in a round-robin way. Accordingly, two query processing algorithms are proposed based on the two partition methods, Parallel Path Merge (PPM) algorithm and Parallel Pipelining Path Join (PPPJ) algorithm. The performance analysis and evaluation on the two data placement strategies and corresponding query processing algorithms are given in this paper.

Original language	English
Pages (from-to)	151-156
Number of pages	6
Journal	Proceedings - IEEE Computer Society's International Computer Software and Applications Conference
Publication status	Published - 2003
Externally published	Yes
Event	Proceedings: 27th Annual International Computer Software and Applications Conference, COMPSAC 2003 - Dallas, TX, United States Duration: 3 Nov 2003 → 6 Nov 2003

Cite this

Yu, Y., Wang, G., Yu, G., Wu, G., Hu, J., & Tang, N. (2003). Data Placement and Query Processing Based on RPE Parallelisms. Proceedings - IEEE Computer Society's International Computer Software and Applications Conference, 151-156.

@article{6cc8bfa8e15942ea9267cf0947a04903,

title = "Data Placement and Query Processing Based on RPE Parallelisms",

abstract = "The basic idea behind parallel database systems is to perform operations in parallel to reduce the response time and improve the system throughput. Data placement is a key factor on the performance of parallel database systems. This paper proposes two data partition strategies to decluster XML documents with very large size, Path Schema based Path Instance Balancing (PSPIB) strategy, in which all path instances with the same path schema in a data tree are declustered evenly over all sites, and Node Schema based Node Round-Robin (NSNRR) strategy, in which all node objects with the same node schema in a data tree are declustered over all sites in a round-robin way. Accordingly, two query processing algorithms are proposed based on the two partition methods, Parallel Path Merge (PPM) algorithm and Parallel Pipelining Path Join (PPPJ) algorithm. The performance analysis and evaluation on the two data placement strategies and corresponding query processing algorithms are given in this paper.",

author = "Yaxin Yu and Guoren Wang and Ge Yu and Gang Wu and Junan Hu and Nan Tang",

year = "2003",

language = "English",

pages = "151--156",

journal = "Proceedings - IEEE Computer Society's International Computer Software and Applications Conference",

issn = "0730-3157",

publisher = "Institute of Electrical and Electronics Engineers Inc.",

note = "Proceedings: 27th Annual International Computer Software and Applications Conference, COMPSAC 2003 ; Conference date: 03-11-2003 Through 06-11-2003",

}

TY - JOUR

T1 - Data Placement and Query Processing Based on RPE Parallelisms

AU - Yu, Yaxin

AU - Wang, Guoren

AU - Yu, Ge

AU - Wu, Gang

AU - Hu, Junan

AU - Tang, Nan

PY - 2003

Y1 - 2003

N2 - The basic idea behind parallel database systems is to perform operations in parallel to reduce the response time and improve the system throughput. Data placement is a key factor on the performance of parallel database systems. This paper proposes two data partition strategies to decluster XML documents with very large size, Path Schema based Path Instance Balancing (PSPIB) strategy, in which all path instances with the same path schema in a data tree are declustered evenly over all sites, and Node Schema based Node Round-Robin (NSNRR) strategy, in which all node objects with the same node schema in a data tree are declustered over all sites in a round-robin way. Accordingly, two query processing algorithms are proposed based on the two partition methods, Parallel Path Merge (PPM) algorithm and Parallel Pipelining Path Join (PPPJ) algorithm. The performance analysis and evaluation on the two data placement strategies and corresponding query processing algorithms are given in this paper.

AB - The basic idea behind parallel database systems is to perform operations in parallel to reduce the response time and improve the system throughput. Data placement is a key factor on the performance of parallel database systems. This paper proposes two data partition strategies to decluster XML documents with very large size, Path Schema based Path Instance Balancing (PSPIB) strategy, in which all path instances with the same path schema in a data tree are declustered evenly over all sites, and Node Schema based Node Round-Robin (NSNRR) strategy, in which all node objects with the same node schema in a data tree are declustered over all sites in a round-robin way. Accordingly, two query processing algorithms are proposed based on the two partition methods, Parallel Path Merge (PPM) algorithm and Parallel Pipelining Path Join (PPPJ) algorithm. The performance analysis and evaluation on the two data placement strategies and corresponding query processing algorithms are given in this paper.

UR - http://www.scopus.com/inward/record.url?scp=0345529057&partnerID=8YFLogxK

M3 - Conference article

AN - SCOPUS:0345529057

SN - 0730-3157

SP - 151

EP - 156

JO - Proceedings - IEEE Computer Society's International Computer Software and Applications Conference

JF - Proceedings - IEEE Computer Society's International Computer Software and Applications Conference

T2 - Proceedings: 27th Annual International Computer Software and Applications Conference, COMPSAC 2003

Y2 - 3 November 2003 through 6 November 2003

ER -

Data Placement and Query Processing Based on RPE Parallelisms

Abstract

Other files and links

Fingerprint

Cite this