Huaping Chen 0001

dblp:66/2271 · DBLP profile ↗
← Back
28ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-2058-7086ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 since 2021Systems, architecture and hardware · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Out-of-Memory Graph Processing Acceleration via Algorithmic-Hardware Codesign on FPGAs
abstract
Emerging FPGA devices are widely used in graph processing and achieve high performance by exploiting fine-grained parallelism and high-bandwidth memory (HBM) subsystems with dozens of channels. However, the graph size increases rapidly and often exceeds the capacity of the FPGA device or one of the memory channels, incurring out-of-memory (OoM) faults and performance degradation. Existing approaches focus on improving the performance of on-device full graph processing, ignoring the support for hyper-scale graphs and cross-iteration acceleration. In this work, we propose SoGraph, a graph processing architecture for large graph acceleration on HBM-equipped FPGAs, cooperating with the host. To minimize PCIe traffic during the processing, we adopt three hardware-oriented optimizations: 1) activeness-driven subgraph scheduler, which minimizes on-device graph resident set; 2) lightweight state-aware processing extension in edge-centric paradigm; 3) bit-level update analyzer for to-be-transfer subgraph reduction. The results show that the prototype on Alveo U280 FPGA achieves an average of 22.3x performance speedup over the modified state-of-the-art FPGA design and 1.3x device energy efficiency over the GPU solution.
Qianyu Cheng, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Huaping Chen 0001, Xuehai Zhou
IEEE Trans. Computers8
2025 Efficient and Fast High-Performance Library Generation for Deep Learning Accelerators
abstract
The widespread adoption of deep learning accelerators (DLAs) underscores their pivotal role in improving the performance and energy efficiency of neural networks. To fully leverage the capabilities of these accelerators, exploration-based library generation approaches have been widely used to substantially reduce software development overhead. However, these approaches have been challenged by issues related to sub-optimal optimization results and excessive optimization overheads. In this paper, we proposeHeronto generate high-performance libraries of DLAs in an efficient and fast way. The key is automatically enforcing massive constraints through the entire program generation process and guiding the exploration with an accurate pre-trained cost model.Heronrepresents the search space as a constrained satisfaction problem (CSP) and explores the space via evolving the CSPs. Thus, the sophisticated constraints of the search space are strictly preserved during the entire exploration process. The exploration algorithm has the flexibility to engage in space exploration using either online-trained models or pre-trained models. Experimental results demonstrate thatHeronaveragely achieves 2.71$\times$speedup over three state-of-the-art automatic generation approaches. Also, compared to vendor-provided hand-tuned libraries,Heronachieves a 2.00$\times$speedup on average. When employing a pre-trained model,Heronachieves 11.6$\times$compilation time speedup, incurring a minor impact on execution time.
Jun Bi, Yuanbo Wen 0001, Xiaqing Li, Yongwei Zhao 0001, Enshuai Zhou, Xing Hu 0001, Zidong Du, Ling Li 0001, Huaping Chen 0001, Tianshi Chen 0002, Qi Guo 0001
IEEE Trans. Computers10
2025 Harmonia: A Unified Architecture for Efficient Deep Symbolic Regression
abstract
Symbolic regression (SR), the process of formulating a mathematical expression based on observed data points, is a fundamental task in artificial intelligence but is often hindered by its intense computational demands. Deep-learning-based SR methods (DSR) aim to alleviate these demands by breaking down the SR process into two stages: 1) neural network (NN) inference and 2) Broyden-Fletcher–Goldfarb-Shanno (BFGS) optimization. Although NN accelerators can expedite the NN stage, the performance of the BFGS optimization is compromised due to its poor performance for the variety of transcendental functions. Moreover, the distinct computational characteristics of NN inference and BFGS cause not only low hardware utilization but also significant area waste. To address these issues, we propose Harmonia, a unified architecture with the neural transcendental function unit (NTFU) and the Unified Array for efficient DSR. The NTFU utilizes the radial basis function network (RBFN) as a universal approximator for various transcendental functions, which significantly reduces the heavy transcendental function computation cost. We further propose an efficient training algorithm called random nonlinear optimization (RNO) to obtain a lightweight RBFN without accuracy loss. Moreover, Harmonia supports configurable dataflow which integrates the two computing stages into the Unified Array. Experimental results show that Harmonia achieves hardware utilization of 83.83%, on average. Compared to the GPU baseline, Harmonia achieves$4.8\times $speedup and$47.6\times $energy saving, alongside considerable low area cost.
Tianyun Ma, Yuanbo Wen 0001, Xinkai Song, Pengwei Jin, Husheng Han, Ziyuan Nan, Zhongkai Yu, Shaohui Peng, Yongwei Zhao 0001, Huaping Chen 0001, Zidong Du, Xing Hu 0001, Qi Guo 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.11
2024 Cambricon-C: Efficient 4-Bit Matrix Unit via Primitivization
abstract
Deep learning trends to use low precision numeral formats to cope with the ever-growing model sizes. For example, the large language model LLaMA2 has been widely deployed in 4-bit precision. With larger models and fewer unique values caused by low precision, an increasing proportion of arithmetic in matrix multiplication is repeating. Although discussed in prior works, such value redundancy has not been fully exploited, and the cost to leverage the value redundancy often offsets any advantages. In this paper, we propose to primitivize the matrix multiplication, that is decomposing it down to the 1-ary successor function (a.k.a. counting) to merge repeating arithmetic. We revisited various techniques to propose Cambricon-C SA, a 4-bit primitive matrix multiplication unit that doubles the energy efficiency over conventional systolic arrays. Experimental results show that Cambricon-C SA can achieve$\mathbf{1}.\mathbf{95}\times$energy efficiency improvement compared with MAC-based systolic array.
Yongwei Zhao 0001, Yifan Hao 0001, Yuanbo Wen 0001, Yuntao Dai, Xiaqing Li, Yang Liu 0466, Rui Zhang 0040, Mo Zou, Xinkai Song, Xing Hu 0001, Zidong Du, Huaping Chen 0001, Qi Guo 0001, Tianshi Chen 0002
MICRO13
2023 Heron: Automatically Constrained High-Performance Library Generation for Deep Learning Accelerators
abstract
Deep Learning Accelerators (DLAs) are effective to improve both performance and energy efficiency of compute-intensive deep learning algorithms. A flexible and portable mean to exploit DLAs is using high-performance software libraries with well-established APIs, which are typically either manually implemented or automatically generated by exploration-based compilation approaches. Though exploration-based approaches significantly reduce programming efforts, they fail to find optimal or near-optimal programs from a large but low-quality search space because the massive inherent constraints of DLAs cannot be accurately characterized.
Jun Bi, Qi Guo 0001, Xiaqing Li, Yongwei Zhao 0001, Yuanbo Wen 0001, Enshuai Zhou, Xing Hu 0001, Zidong Du, Ling Li 0001, Huaping Chen 0001, Tianshi Chen 0002
ASPLOS (3)11
2022 Scheduling a single batch processing machine with non-identical two-dimensional job sizes
Shengchao Zhou, Mingzhou Jin, Huaping Chen 0001
Expert Syst. Appl.5
2022 ViA: A Novel Vision-Transformer Accelerator Based on FPGA
abstract
Since Google proposed Transformer in 2017, it has made significant natural language processing (NLP) development. However, the increasing cost is a large amount of calculation and parameters. Previous researchers designed and proposed some accelerator structures for transformer models in field-programmable gate array (FPGA) to deal with NLP tasks efficiently. Now, the development of Transformer has also affected computer vision (CV) and has rapidly surpassed convolution neural networks (CNNs) in various image tasks. And there are apparent differences between the image data used in CV and the sequence data in NLP. The details in the models contained with transformer units in these two fields are also different. The difference in terms of data brings about the problem of the locality. The difference in the model structure brings about the problem of path dependence, which is not noticed in the existing related accelerator design. Therefore, in this work, we propose the ViA, a novel vision transformer (ViT) accelerator architecture based on FPGA, to execute the transformer application efficiently and avoid the cost of these challenges. By analyzing the data structure in the ViT, we design an appropriate partition strategy to reduce the impact of data locality in the image and improve the efficiency of computation and memory access. Meanwhile, by observing the computing flow of the ViT, we use the half-layer mapping and throughput analysis to reduce the impact of path dependence caused by the shortcut mechanism and fully utilize hardware resources to execute the Transformer efficiently. Based on optimization strategies, we design two reuse processing engines with the internal stream, different from the previous overlap or stream design patterns. In the stage of the experiment, we implement the ViA architecture in Xilinx Alveo U50 FPGA and finally achieved ~5.2 times improvement of energy efficiency compared with NVIDIA Tesla V100, and 4–10 times improvement of performance compared with related accelerators based on FPGA, that obtained nearly 309.6 GOP/s computing performance in the peek.
Lei Gong 0003, Chao Wang 0003, Yang Yang 0080, Yingxue Gao, Xuehai Zhou, Huaping Chen 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2021 A scheduling decision support model for minimizing the number of drones with dynamic package arrivals and personalized deadlines
Huaping Chen 0001, Xueping Li 0002, Zeyu Liu 0002
Expert Syst. Appl.2
2019 Design Exploration of Multi-FPGAs for Accelerating Deep Learning
abstract
Due to the low power consumption and reconfigurability of FPGA, the use of FPGA for accelerating calculations is becoming more and more hot, including deep learning. However, due to limited hardware resource, single FPGA-based accelerator cannot configure optimal parameters for each layer, and its performance is also limited by the data memory bandwidth. To accelerate the calculation of neural network, this work designs a calculation module, and based on this module, further optimizes the data transmission path between multi-FPGA, thus achieving nearly linear performance growth between performance and the number of FPGAs.
Lei Gong 0003, Chao Wang 0003, Xuehai Zhou, Huaping Chen 0001
CLUSTER5
2019 Pseudo transformation mechanism between resource allocation and bin-packing in batching environments
Xinle Liang, Shengchao Zhou, Huaping Chen 0001, Rui Xu 0004
Future Gener. Comput. Syst.3
2019 Tabu search algorithms for minimizing total completion time on a single machine with an actual time-dependent learning effect
Chunhui Zheng, Huaping Chen 0001, Rui Xu 0004
Nat. Comput.2
2018 MALOC: A Fully Pipelined FPGA Accelerator for Convolutional Neural Networks With All Layers Mapped on Chip
abstract
Recently, field-programmable gate arrays (FPGAs) have been widely used in the implementations of hardware accelerator for convolutional neural networks (CNNs). However, most of these existing accelerators are designed in the same idea as their ASIC counterparts, in which all operations from different layers are mapped to the same hardware units and working in a multiplexed way. This manner does not take full advantage of reconfigurability and customizability of FPGAs, resulting in a certain degree of computational efficiency degradation. In this paper, we propose a new architecture for FPGA-based CNN accelerator that maps all the layers to their own on-chip units and working concurrently as a pipeline. A comprehensive mapping and optimizing methodology based on establishing roofline model oriented optimization model is proposed, which can achieve maximum resource utilization as well as optimal computational efficiency. Besides, to ease the programming burden, we propose a design framework which can provide a one-stop function for developers to generate the accelerator with our optimizing methodology. We evaluate our proposal by implementing different modern CNN models on Xilinx Zynq-7020 and Virtex-7 690t FPGA platforms. Experimental results show that our implementations can achieve a peak performance of 910.2 GOPS on Virtex-7 690t, and 36.36 GOP/s/W energy efficiency on Zynq-7020, which are superior to the previous approaches.
Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Huaping Chen 0001, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 Thermal Augmented Expression Recognition
abstract
Visible facial images provide geometric and appearance patterns of facial expressions and are sensitive to illumination changes. Thermal facial images record facial temperature distribution and are robust to light conditions. Therefore, expression recognition is enhanced by visible and thermal image fusion. In most cases, only visible images are available due to the widespread popularity of visible cameras and the high cost of thermal cameras. Thus, we propose a novel visible expression recognition method by using thermal infrared (IR) data as privileged information, which is only available during training. Specifically, we first learn a deep model for visible images and thermal images. Then we use the learned feature representations to train support vector machine (SVM) classifiers for expression classification. We jointly refine the deep models as well as the SVM classifiers for both thermal images and visible images by imposing the constraint that the outputs of the SVM classifiers from two views are similar. Thermal IR images during training are then exploited to construct better facial representations and expression classifiers from visible images. We extend the proposed thermal augmented expression recognition method for partially unpaired data, acknowledging that visible images and thermal images maybe not be recorded synchronously. Experimental resulton the MAHNOB laughter database demonstrate that the proposed thermal augmented expression recognition method can effectively exploit thermal IR images' supplementary role for visible facial expression recognition during training to obtain better facial representations and a better visible expression classifier. The proposed thermal augmented expression recognition method achieves state-of-the-art expression recognition performance for both paired and unpaired facial images.
Shangfei Wang, Bowen Pan, Huaping Chen 0001
IEEE Trans. Cybern.3
2016 Facial Expression Recognition with Deep two-view Support Vector Machine
abstract
This paper proposes a novel deep two-view approach to learn features from both visible and thermal images and leverage the commonality among visible and thermal images for facial expression recognition from visible images. The thermal images are used as privileged information, which is required only during training to help visible images learn better features and classifier. Specifically, we first learn a deep model for visible images and thermal images respectively, and use the learned feature representations to train SVM classifiers for expression classification. We then jointly refine the deep models as well as the SVM classifiers for both thermal images and visible images by imposing the constraint that the outputs of the SVM classifiers from two views are similar. Therefore, the resulting representations and classifiers capture the inherent connections among visible facial image, infrared facial image and target expression labels, and hence improve the recognition performance for facial expression recognition from visible images during testing. Experimental results on the benchmark expression database demonstrate the effectiveness of our proposed method.
Chongliang Wu, Shangfei Wang, Bowen Pan, Huaping Chen 0001
ACM Multimedia4
2016 Priority-based constructive algorithms for scheduling agile earth observation satellites with total priority maximization
Rui Xu 0004, Huaping Chen 0001, Xinle Liang
Expert Syst. Appl.2
2015 A Boltzmann-Based Estimation of Distribution Algorithm for a General Resource Scheduling Model
abstract
Most researchers employed common functional models when managing scheduling problems with controllable processing times. However, in many complicated manufacturing systems with a high diversity of jobs, these functional resource models fail to reflect their specific characteristics. To fulfill these requirements, we apply a more general model, the discrete model. Traditional functional models can be viewed as special cases of such model. In this paper, the discrete model is implemented on a problem of minimizing the weighted resource allocation subject to a common deadline on a single machine. By reducing the problem to a partition problem, we demonstrate that it is NP-complete, which addresses the difficult issue of the guarantee of both the solution quality and time cost. In order to tackle the problem, we develop an estimation of distribution algorithm based on an approximation of the Boltzmann distribution. The approximation strategy represents a tradeoff between complexity and solution accuracy. The results of the experiments conducted on benchmarks show that, compared with other alternative approaches, the proposed algorithm has competitive behavior, obtaining 74 best solutions out of 90 instances.
Xinle Liang, Huaping Chen 0001, José Antonio Lozano 0001
IEEE Trans. Evol. Comput.2
2013 An Estimation of Distribution Algorithm for the 3D Bin Packing Problem with Various Bin Sizes
Yaxiong Cai, Huaping Chen 0001, Rui Xu 0004, Hao Shao, Xueping Li 0002
IDEAL2
2013 An Effective Ant Colony Approach for Scheduling Parallel Batch-Processing Machines
Rui Xu 0004, Huaping Chen 0001, Hao Shao
IDEAL2
2013 A social network-empowered research analytics framework for project selection
Thushari P. Silva, Zhiling Guo, Jian Ma 0008, Hongbing Jiang, Huaping Chen 0001
Decis. Support Syst.5
2011 Repurchase intention in B2C e-commerce - A relationship quality perspective
Yulin Fang, Kwok Kee Wei, Elaine Ramsey, Patrick McCole, Huaping Chen 0001
Inf. Manag.6
2009 How do mediated and non-mediated power affect electronic supply chain management system adoption? The mediating effects of trust and institutional pressures
Weiling Ke, Hefu Liu, Kwok Kee Wei, Jibao Gu, Huaping Chen 0001
Decis. Support Syst.5
2009 Understanding the role of gender in bloggers' switching behavior
Kem Z. K. Zhang, Matthew K. O. Lee, Christy M. K. Cheung, Huaping Chen 0001
Decis. Support Syst.4
2008 A chaotic Ant Colony Optimization method for scheduling a single batch-processing machine with non-identical job sizes
abstract
The problem of minimizing makespan on a single batch-processing machine with non-identical job sizes is strongly NP-hard. This paper proposes an ant colony optimization (ACO) algorithm with chaotic control to solve the problem. The metropolis criterion is adopted to select the paths of ants to escape immature convergence. In order to improve the solutions of ACO, a chaotic optimizer is designed and integrated into ACO to reinforce the capacity of global optimization. Batch first fit is introduced to decode the paths into feasible solutions of the problem. In the experiment, the instances of 24 levels are simulated and the results show that the proposed CACO outperforms genetic algorithm and simulated annealing on all the instances.
Bayi Cheng, Huaping Chen 0001, Hao Shao, Rui Xu 0004, George Q. Huang
IEEE Congress on Evolutionary Computation2
2008 Fuzzy scheduling for single batch-processing machine with non-identical job sizes
abstract
In this paper, we introduce the fuzzy model of the makespan on a single batch-processing machine with non-identical job sizes and propose an improved DNA evolutionary algorithm (IDEA) solution approach. The model is based on fuzzy batch processing time and fuzzy intervals between batches. DEA is improved by integrating the crossover operator to overcome the immature convergence caused by the determinate selection of vertical operator in DEA. To decode the permutations of jobs searched by IDEA, the heuristic first fit decreasing (FFD) is applied to produce batches. In the experiment, the results of the fuzzy makespan demonstrate the proposed algorithm outperforms GA and SA on all instances.
Bayi Cheng, Huaping Chen 0001, Shuan-shi Wang
FUZZ-IEEE2
2006 Comparison Model and Algorithm for Distributed Firewall Policy
Weiping Wang 0006, Zhepeng Li, Huaping Chen 0001
ICIC (2)4
2006 An Empirical Study of What Drives Users to Share Knowledge in Virtual Communities
Shun Ye, Huaping Chen 0001, Xiaoling Jin
KSEM2
2004 VAST: A Service Based Resource Integration System for Grid Society
Jiulong Shan, Huaping Chen 0001, Guangzhong Sun
ISPA2
2000 A Fast Algorithm for Mining Association Rules
Liusheng Huang, Huaping Chen 0001, Wang Xun, Guoliang Chen 0001
J. Comput. Sci. Technol.2