RelJoin: Relative-cost-based selection of distributed join methods for query plan optimization

Feng Liang; Francis C.M. Lau; Heming Cui; Yupeng Li; Bing Lin; Chengming Li; Xiping Hu

doi:10.1016/j.ins.2023.120022

RelJoin: Relative-cost-based selection of distributed join methods for query plan optimization

Feng Liang, Francis C.M. Lau, Heming Cui, Yupeng Li, Bing Lin, Chengming Li, Xiping Hu^*

^*此作品的通讯作者

医学技术学院

科研成果: 期刊稿件 › 文章 › 同行评审

1 引用（Scopus）

摘要

Selecting appropriate distributed join methods for logical join operations in a query plan is crucial for the performance of data-intensive scalable computing (DISC). Different network communication patterns in the data exchange phase generate varying network communication workloads and significantly affect the distributed join performance. However, most cost-based query optimizers focus on the local computing cost and do not precisely model the network communication cost. We propose a cost model for various distributed join methods to optimize join queries in DISC platforms. Our method precisely measures the network and local computing workloads in different execution phases, using information on the size and cardinality statistics of datasets and cluster join parallelism. Our cost model reveals the importance of the relative size of the joining datasets. We implement an efficient distributed join selection strategy, known as RelJoin in SparkSQL, which is an industry-prevalent distributed data processing framework. RelJoin uses runtime adaptive statistics for accurate cost estimation and selects optimal distributed join methods for logical joins to optimize the physical query plan. The evaluation results on the TPC-DS benchmark show that RelJoin performs best in 62 of the 97 queries and can reduce the average query time by 21% compared with other strategies.¹

源语言	英语
文章编号	120022
期刊	Information Sciences
卷	658
DOI	https://doi.org/10.1016/j.ins.2023.120022
出版状态	已出版 - 2月 2024

访问文件

10.1016/j.ins.2023.120022

其它文件与链接

链接到 Scopus 的出版物

引用此

@article{0344bcf7b5214f0eac54c912d70f9dc6,

title = "RelJoin: Relative-cost-based selection of distributed join methods for query plan optimization",

abstract = "Selecting appropriate distributed join methods for logical join operations in a query plan is crucial for the performance of data-intensive scalable computing (DISC). Different network communication patterns in the data exchange phase generate varying network communication workloads and significantly affect the distributed join performance. However, most cost-based query optimizers focus on the local computing cost and do not precisely model the network communication cost. We propose a cost model for various distributed join methods to optimize join queries in DISC platforms. Our method precisely measures the network and local computing workloads in different execution phases, using information on the size and cardinality statistics of datasets and cluster join parallelism. Our cost model reveals the importance of the relative size of the joining datasets. We implement an efficient distributed join selection strategy, known as RelJoin in SparkSQL, which is an industry-prevalent distributed data processing framework. RelJoin uses runtime adaptive statistics for accurate cost estimation and selects optimal distributed join methods for logical joins to optimize the physical query plan. The evaluation results on the TPC-DS benchmark show that RelJoin performs best in 62 of the 97 queries and can reduce the average query time by 21% compared with other strategies.1",

keywords = "Adaptive statistics, Cost-based, Distributed join, Query plan optimization",

author = "Feng Liang and Lau, {Francis C.M.} and Heming Cui and Yupeng Li and Bing Lin and Chengming Li and Xiping Hu",

note = "Publisher Copyright: {\textcopyright} 2023 Elsevier Inc.",

year = "2024",

month = feb,

doi = "10.1016/j.ins.2023.120022",

language = "English",

volume = "658",

journal = "Information Sciences",

issn = "0020-0255",

publisher = "Elsevier Inc.",

}

TY - JOUR

T1 - RelJoin

T2 - Relative-cost-based selection of distributed join methods for query plan optimization

AU - Liang, Feng

AU - Lau, Francis C.M.

AU - Cui, Heming

AU - Li, Yupeng

AU - Lin, Bing

AU - Li, Chengming

AU - Hu, Xiping

PY - 2024/2

Y1 - 2024/2

N2 - Selecting appropriate distributed join methods for logical join operations in a query plan is crucial for the performance of data-intensive scalable computing (DISC). Different network communication patterns in the data exchange phase generate varying network communication workloads and significantly affect the distributed join performance. However, most cost-based query optimizers focus on the local computing cost and do not precisely model the network communication cost. We propose a cost model for various distributed join methods to optimize join queries in DISC platforms. Our method precisely measures the network and local computing workloads in different execution phases, using information on the size and cardinality statistics of datasets and cluster join parallelism. Our cost model reveals the importance of the relative size of the joining datasets. We implement an efficient distributed join selection strategy, known as RelJoin in SparkSQL, which is an industry-prevalent distributed data processing framework. RelJoin uses runtime adaptive statistics for accurate cost estimation and selects optimal distributed join methods for logical joins to optimize the physical query plan. The evaluation results on the TPC-DS benchmark show that RelJoin performs best in 62 of the 97 queries and can reduce the average query time by 21% compared with other strategies.1

AB - Selecting appropriate distributed join methods for logical join operations in a query plan is crucial for the performance of data-intensive scalable computing (DISC). Different network communication patterns in the data exchange phase generate varying network communication workloads and significantly affect the distributed join performance. However, most cost-based query optimizers focus on the local computing cost and do not precisely model the network communication cost. We propose a cost model for various distributed join methods to optimize join queries in DISC platforms. Our method precisely measures the network and local computing workloads in different execution phases, using information on the size and cardinality statistics of datasets and cluster join parallelism. Our cost model reveals the importance of the relative size of the joining datasets. We implement an efficient distributed join selection strategy, known as RelJoin in SparkSQL, which is an industry-prevalent distributed data processing framework. RelJoin uses runtime adaptive statistics for accurate cost estimation and selects optimal distributed join methods for logical joins to optimize the physical query plan. The evaluation results on the TPC-DS benchmark show that RelJoin performs best in 62 of the 97 queries and can reduce the average query time by 21% compared with other strategies.1

KW - Adaptive statistics

KW - Cost-based

KW - Distributed join

KW - Query plan optimization

UR - http://www.scopus.com/inward/record.url?scp=85180369465&partnerID=8YFLogxK

U2 - 10.1016/j.ins.2023.120022

DO - 10.1016/j.ins.2023.120022

M3 - Article

AN - SCOPUS:85180369465

SN - 0020-0255

VL - 658

JO - Information Sciences

JF - Information Sciences

M1 - 120022

ER -

RelJoin: Relative-cost-based selection of distributed join methods for query plan optimization

摘要

访问文件

其它文件与链接

指纹

引用此