EDBT 2026 Demo / reviewers in the wild / expert
Haining Fang
dblp:88/10207
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0007-8579-9681ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D2 Prune: Sparsifying Large Language Models via Dual Taylor Expansion and Attention Distribution AwarenessabstractLarge language models (LLMs) face significant deployment challenges due to their massive computational demands. While pruning offers a promising compression solution, existing methods suffer from two critical limitations: (1) They neglect activation distribution shifts between calibration data and test data, resulting in inaccurate error estimations; (2) Overlooking the long-tail distribution characteristics of activations in the attention module. To address these limitations, this paper proposes a novel pruning method, D²Prune. First, we propose a dual Taylor expansion-based method that jointly models weight and activation perturbations for precise error estimation, leading to precise pruning mask selection and weight updating and facilitating error minimization during pruning. Second, we propose an attention-aware dynamic update strategy that preserves the long-tail attention pattern by jointly minimizing the KL divergence of attention distributions and the reconstruction error. Extensive experiments show that D²Prune consistently outperforms SOTA methods across various LLMs (e.g., OPT-125M, LLaMA2/3, Qwen3). Moreover, the dynamic attention update mechanism also generalizes well to ViT-based vision models like DeiT, achieving superior accuracy on ImageNet-1K. Lang Xiong, Ning Liu 0007, Ao Ren, Yuheng Bai, Haining Fang, Binyan Zhang, Yujuan Tan, Duo Liu 0002 |
AAAI | 5 |
| 2026 | LiPRA: Lightweight pruning rate allocation for LLMs via global sensitivity measurement
Haining Fang, Ning Liu 0007, Lang Xiong, Zhenyu Wang 0002, Xianzhang Chen, Ao Ren, Yujuan Tan |
Neurocomputing | 1 |
| 2025 | CoSF: A Co-Optimization Framework for Operator Splitting and Fusion
Wei Li 0322, Ao Ren, Qingqiu Lan, Haining Fang, Zhenyu Wang 0002, Yujuan Tan, Kan Zhong, Duo Liu 0002 |
Euro-Par (1) | 4 |
| 2025 | DSAV: A Deep Sparse Acceleration Framework for Voxel-Based 3-D Object DetectionabstractVoxel-based 3-D object detection has been widely applied in robotics, virtual reality, and autonomous driving. However, inefficiency in the voxelization and backbone-network computation, which are the main components of the voxel-based models, prevents efficient 3-D object detection. First, due to the high sparsity and irregularity of the point cloud, the voxelization process usually requires generalized platforms, such as CPUs, and causes low voxelization speed. Second, the voxel-based models contain considerable transposed convolutional layers, and existing accelerators introduce considerable additional hardware to support both the convolution and transposed convolution operations. Nonetheless, this strategy incurs significant hardware costs. Besides, transposed convolutions result in various patterns of sparse feature maps, and pruning as a representative model compression technique, results in sparse weight matrices. The two types of sparsity impose challenges in accelerating the voxel-based models, including activation-weight matching efficiency, low partial-sum accumulation efficiency, and workload imbalance issues. In this work, we propose DSAV, a 3-D object detection accelerator to address these obstacles. Specifically, we first propose a hash-based voxelizer for efficient voxelization, by storing and indexing voxels hierarchically. Then, we collaboratively design the transposed convolution acceleration method, structured pruning method, and accelerator architecture for the voxel-based models. As a result, the accelerator can fully leverage the sparsity lies in both feature maps and weight matrices. Experimental results show that the proposed accelerator can outperform the prior studies by$19{\times } \sim 19.8{\times }$faster in voxelization and$4.29{\times } \sim 38.01\times $faster in backbone inference. Finally, the accelerator achieves$4.61{\times } \sim 31.63{\times }$speedups than its counterparts in 3-D object detection tasks. Haining Fang, Yujuan Tan, Ao Ren, ZhiYong Qin, Duo Liu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | RACI: A Resource-Aware Cooperative Inference Framework on Heterogeneous Edge DevicesabstractCooperative inference for deep neural networks (DNNs) across edge devices has received increasing attention, due to the benefits of low latency, low power consumption, and privacy preservation. Cooperative inference partitions a DNN model into multiple segments, which will then be allocated to distributed devices for parallel inference. Nonetheless, prior works fail to comprehensively study the impact of layer configurations, dynamic network bandwidths, and heterogeneous device capabilities on the inference speed, resulting in suboptimal inference performance. In this work, we conduct a comprehensive analysis of these key factors and figure out the limitations of conventional transfer-based and redundant computation-based methods. Based on the analysis, we first propose a latency prediction agent that accounts for the layer configurations, network bandwidths, and device computing capabilities, aiming to quickly evaluate the inference latency. Furthermore, we propose RACI, a resource-aware cooperative DNNs inference framework on heterogeneous edge devices. It co-trains a model agent for model partition and a workload agent for workload allocation to generate co-optimized model partition and workload allocation strategies, leading to high cooperation inference acceleration. Experimental results demonstrate that RACI outperforms the state-of-the-art approaches by 1.1× -5.2× in terms of inference speedup for three representative DNN models. Zhenyu Wang 0002, Ao Ren, Duo Liu 0002, Haining Fang, Jiaxing Shi, Yujuan Tan, Xianzhang Chen |
ICCAD | 4 |
| 2024 | FreePrune: An Automatic Pruning Framework Across Various Granularities Based on Training-Free EvaluationabstractNetwork pruning is an effective technique that reduces the computational costs of networks while maintaining accuracy. However, pruning requires expert knowledge and hyperparameter tuning, such as determining the pruning rate for each layer. Automatic pruning methods address this challenge by proposing an effective training-free metric to quickly evaluate the pruned network without fine-tuning. However, most existing automatic pruning methods only investigate a certain pruning granularity, and it remains unclear whether metrics benefit automatic pruning at different granularities. Neural architecture search also studies training-free metrics to accelerate network generation. Nevertheless, whether they apply to pruning needs further investigation. In this study, we first systematically analyze various advanced training-free metrics for various granularities in pruning, and then we investigate the correlation between the training-free metric score and the after-fine-tuned model accuracy. Based on the analysis, we proposed FreePrune score, a more general metric compatible with all pruning granularities. Aiming at generating high-quality pruned networks and unleashing the power of FreePrune score, we further propose FreePrune, an automatic framework that can rapidly generate and evaluate the candidate networks, leading to a final pruned network with both high accuracy and pruning rate. Experiments show that our method achieves high correlation on various pruning granularities and comprehensively improves the accuracy. Ning Liu 0007, Haining Fang, Qiu Lin, Yujuan Tan, Xianzhang Chen, Duo Liu 0002, Kan Zhong, Ao Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |