Honghe Zhang

dblp:166/7116 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
autotuning
0.812024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024
Compilers and program optimization › domain-specific compilation
tensor algebra compilation
0.812024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024
High-performance computing › sparse linear algebra
sparse matrix multiplication
0.812024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024
High-performance computing › sparse linear algebra › sparse matrix multiplication
SpMM
0.812024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024
Bioinformatics and computational biology › cancer genomics
cancer driver gene identification
0.712023
WMDS.net: a network control framework for identifying key players in transcriptome programs · Bioinform. 2023
Bioinformatics and computational biology
cancer genomics
0.712023
WMDS.net: a network control framework for identifying key players in transcriptome programs · Bioinform. 2023
Bioinformatics and computational biology › biological network
network biology
0.712023
WMDS.net: a network control framework for identifying key players in transcriptome programs · Bioinform. 2023
Machine learning › Efficient and distributed learning › inference efficiency
sparse neural network inference
0.212024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024

Methods — techniques the papers use, named apart from their topics

search space reduction · 2.3rank-based cost model · 2.3proxy estimation · 2.3structural controllability · 0.7minimum dominating set · 0.7
YearPublicationVenuePosition
2026 QTLNetwork-MP: integrative mapping of additive, epistatic, and G × E effects for complex traits in multiparent advanced generation intercross populations
abstract
Quantitative trait locus (QTL) mapping is a powerful approach to uncover the genetic architecture of complex traits. Traditional biparental populations, though widely used, offer limited genetic diversity and resolution, making them less effective for robust genetic architecture dissection. In contrast, multiparent advanced generation intercross (MAGIC) populations possess more allelic diversity, enabling the detection of intricate gene interactions including dominance and epistasis in QTL mapping. However, QTL mapping in MAGIC populations encounters some great challenges, particularly in inferring parental origin of QTL and estimating their effects unbiasedly. We tackled these challenges by developing a novel QTL mapping method for pure-line MAGIC populations, accompanied by the software QTLNetwork-MP. This method extends the mixed linear model framework by integrating a Markov chain-based algorithm and orthogonal QTL effect decomposition. This allows for the estimation of QTL effects, including additive and additive-additive epistasis as well as environmental interactions for MAGIC populations derived from four-way or eight-way crosses. Simulation studies show that QTLNetwork-MP achieves high statistical power and well-controlled false discovery rates across varying heritability, population size, and parent number. We applied QTLNetwork-MP to map QTLs for seed size in an eight-parent recombinant inbred line MAGIC cowpea population, identifying three significant additive QTLs and accurately estimating their genetic effects, validating the value of the method in practical application.
Tianneng Zhu, Zhijun Tong, Mingzhe Suo, Mingzi Xu, Honghe Zhang, Bingguang Xiao
Briefings Bioinform.7
2024 LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs
abstract
As deep neural networks (DNNs) become increasingly large and complicated, pruning techniques are proposed for lower memory footprint and more efficient inference. The most critical kernel to execute pruned sparse DNNs on GPUs is Sparse-dense Matrix Multiplication (SpMM). To maximize the performance of SpMM, despite the high-performance implementation generated from advanced tensor compilers, they often take a long time to iteratively search tuning configurations. Such a long time slows down the cycle of exploring better DNN architectures or pruning algorithms. In this article, we propose LO-SpMM to efficiently generate high-performance SpMM implementations for sparse DNN inference. Based on the analysis of nonzero elements’ layout, the characterization of the GPU architecture, and a rank-based cost model, LO-SpMM can effectively reduce the search space and eliminate possibly low-performance candidates. Besides, rather than generating complete SpMM implementations for evaluation, LO-SpMM constructs simplified proxies to quickly estimate performance, thereby substantially reducing compilation and execution costs. Experimental results show that LO-SpMM can reduce the search time by 281× at most, while the performance of generated SpMM implementations is comparable to or better than the state-of-the-art sparse tensor compiling solutions.
Junqing Lin, Jingwei Sun 0001, Honghe Zhang, Xianzhi Yu, Guangzhong Sun
ACM Trans. Archit. Code Optim.4
2023 EC-SpMM: Efficient Compilation of SpMM Kernel on GPUs
abstract
As deep neural networks (DNNs) become increasingly large and complicated, pruning techniques are proposed for lower memory footprint and more efficient inference. The most critical kernel to execute pruned sparse DNNs on GPUs is Sparse-dense Matrix Multiplication (SpMM). To maximize the performance of SpMM, despite the high-performance code generated from recent tensor compilers, they often take a long time for iteratively searching candidate configurations. Such a long time slows down the cycle of exploring better DNN architectures or pruning algorithms. In this paper, we propose EC-SpMM to efficiently generate high-performance SpMM kernels for sparse DNN inference. Based on the analysis of nonzero elements’ layout, the characterization of GPU architecture, and a rank-based cost model, EC-SpMM can effectively reduce the search space and eliminate possibly low-performance candidates. Experimental results show that EC-SpMM can reduce the compilation time by a factor of 35 ×, while the performance of generated SpMM kernels is comparable or even better, compared with the state-of-the-art sparse tensor compiling solution.
Junqing Lin, Honghe Zhang, Jingwei Sun 0001, Xianzhi Yu, Guangzhong Sun
ICPP2
2023 WMDS.net: a network control framework for identifying key players in transcriptome programs
abstract
MOTIVATION: Mammalian cells can be transcriptionally reprogramed to other cellular phenotypes. Controllability of such complex transitions in transcriptional networks underlying cellular phenotypes is an inherent biological characteristic. This network controllability can be interpreted by operating a few key regulators to guide the transcriptional program from one state to another. Finding the key regulators in the transcriptional program can provide key insights into the network state transition underlying cellular phenotypes. RESULTS: To address this challenge, here, we proposed to identify the key regulators in the transcriptional co-expression network as a minimum dominating set (MDS) of driver nodes that can fully control the network state transition. Based on the theory of structural controllability, we developed a weighted MDS network model (WMDS.net) to find the driver nodes of differential gene co-expression networks. The weight of WMDS.net integrates the degree of nodes in the network and the significance of gene co-expression difference between two physiological states into the measurement of node controllability of the transcriptional network. To confirm its validity, we applied WMDS.net to the discovery of cancer driver genes in RNA-seq datasets from The Cancer Genome Atlas. WMDS.net is powerful among various cancer datasets and outperformed the other top-tier tools with a better balance between precision and recall. AVAILABILITY AND IMPLEMENTATION: https://github.com/chaofen123/WMDS.net. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Md Amanullah, Weigang Liu, Xiaoqing Pan, Honghe Zhang, Pengyuan Liu 0003
Bioinform.6