跳到主要导航 跳到搜索 跳到主要内容

Revisiting thread configuration of SpMV kernels on GPU: A machine learning based approach

  • Jianhua Gao
  • , Weixing Ji*
  • , Jie Liu
  • , Yizhuo Wang
  • , Feng Shi
  • *此作品的通讯作者
  • Beijing Institute of Technology

科研成果: 期刊稿件文章同行评审

摘要

Sparse matrix-vector multiplication (SpMV) optimization on GPUs has been challenging due to irregular memory accesses and unbalanced workloads. The majority of existing solutions assign a fixed number of threads to one or more rows of sparse matrices according to empirical formulas. However, this method does not give the optimal thread configuration and results in a significant performance loss. This paper proposes a new machine learning-based thread assignment strategy for SpMV on GPU, predicting the near-optimal thread configuration for matrices. Further, we partition irregular sparse matrices into blocks according to the distribution of non-zero elements and predict the optimal thread configuration for each block. A new SpMV kernel is designed to accelerate the execution of different blocks. Experimental results show that our machine learning-based approach can select the near-optimal thread configuration for most matrices. The efficiency of SpMV for irregular matrices is also improved by matrix partitioning and blockwise prediction. Finally, we dive into the trained model to find out the connection between the features of a sparse matrix and its optimal thread configuration.

源语言英语
期刊论文编号104799
期刊Journal of Parallel and Distributed Computing
185
DOI
出版状态已出版 - 3月 2024

学术指纹

探究 'Revisiting thread configuration of SpMV kernels on GPU: A machine learning based approach' 的科研主题。它们共同构成独一无二的学术指纹。

引用此