Shuhong Huang

dblp:38/6377 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
automatic differentiation
1.012026
ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct Indexing · PPoPP 2026
Compilers and program optimization › automatic differentiation
reverse-mode automatic differentiation
1.012026
ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct Indexing · PPoPP 2026
Compilers and program optimization › program transformation
compiler transformations
0.712023
EINNET: Optimizing Tensor Programs with Derivation-Based Transformations · OSDI 2023
Compilers and program optimization › deep learning compiler
tensor program optimization
0.712023
EINNET: Optimizing Tensor Programs with Derivation-Based Transformations · OSDI 2023

Methods — techniques the papers use, named apart from their topics

tape restructuring · 2.0direct indexing · 2.0
YearPublicationVenuePosition
2026 ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct Indexing
abstract
Automatic Differentiation (AD) is a technique that computes the derivatives of numerical programs by systematically applying the chain rule, playing a critical role in domains such as machine learning, simulation, and control systems. However, parallelizing differentiated programs remains a significant challenge due to the conflict between tapes (a data structure for intermediate variable storage) and summations: the differentiation process inherently introduces inter-thread summation patterns, which require prohibitively expensive atomic operations; and traditional tape designs tightly couple data retrieval with the program’s control flow, preventing code restructuring needed to eliminate these costly dependencies.
Shuhong Huang, Shizhi Tang, Yuan Wen, Huanqi Cao, Ruibai Tang, Yidong Chen 0003, Jiping Yu, Jidong Zhai
PPoPP1
2025 IntelliGen: Instruction-Level Auto-tuning for Tensor Program with Monotonic Memory Optimization
abstract
Tensor compilers play a critical role in optimizing deep neural networks (DNNs), with memory performance emerging as a key bottleneck in code generation for DNN models. Existing tensor compilers are constrained by inefficient auto-tuning algorithms. They either must deploy coarse-grained descriptions, thus miss potential optimization, or struggle with vast search spaces, rendering auto-tuning inapplicable. Tensor compilers require a more holistic optimization of memory performance to overcome these constraints. To address this issue, we focus our optimization objective on memory performance, which allows us to design monotonic optimization methods, significantly enhancing the efficiency of auto-tuning and thus enabling auto-tuning on a fine-granularity description. Based on these observations, we propose IntelliGen, a tensor compiler with instruction-level auto-tuning and monotonic memory optimization. We design an instruction-level graph description, and a monotonic optimization method for optimization on . Benefiting from auto-tuning techniques with fine-grained description, IntelliGen demonstrates significant speedup of up to 3.13×, 3.55×, and 16.9× (averaging 1.46×, 1.85×, and 2.30×, respectively) on NVIDIA GPUs, AMD GPUs, and Cambricon MLUs over the most efficient existing frameworks.
Zixuan Ma, Haojie Wang 0004, Jingze Xing, Shuhong Huang, Liyan Zheng 0001, Chen Zhang 0001, Huanqi Cao, Kezhao Huang, Mingshu Zhai, Shizhi Tang, Penghan Wang, Jidong Zhai
CGO4
2023 EINNET: Optimizing Tensor Programs with Derivation-Based Transformations
Liyan Zheng 0001, Haojie Wang 0004, Jidong Zhai, Muyan Hu, Zixuan Ma, Tuowei Wang, Shuhong Huang, Xupeng Miao, Shizhi Tang, Kezhao Huang
OSDI7
2012 Derivation of Reliability and Variance Estimates for Multi-State Systems With Binary-Capacitated Components
abstract
This paper analytically derives the system reliability estimate, and the associated variance estimate for multi-state systems respectively using reliability, and variance estimates of binary-capacitated components. The derivation utilizes the universal generating function method to formulate a state table and a product expectation table when replacing two components with an equivalent virtual component. Closed-form expressions of the system reliability estimate and the associated variance estimate are formulated through an iterative derivation process. The derivation can be applied to multi-state systems with series-parallel configurations. Three example systems in the literature are used to illustrate the effectiveness and accuracy of the proposed analytical estimation approach. The confidence interval for the system reliability estimate is developed based on the derived results. The developed interval is then compared with another interval from the literature that approximated the variance estimate using a pseudo binomial distribution. Comparisons through Monte Carlo simulations on the example systems indicate that the coverage probabilities have been significantly improved by the interval constructed based on the proposed derivation.
Lin Li 0004, Shuhong Huang
IEEE Trans. Reliab.3
2009 Feature selection using tabu search with long-term memories and probabilistic neural networks
Lin Li 0004, Shuhong Huang
Pattern Recognit. Lett.4