Haonan Zhang 0002

dblp:238/5984-2 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-4239-6141ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Explicit-implicit dense to sparse knowledge distillation for efficient sparse 3D detection
Haonan Zhang 0002, Kang Ke, Yingke Gao, Longjun Liu
Pattern Recognit.2
2026 Efficient vision-based occupancy prediction with knowledge distillation
Kang Ke, Haonan Zhang 0002, Penghui Fan, Longjun Liu
Pattern Recognit.2
2026 RFIDet: Visual Prior-Guided Rail Fastener Integrity Detection for UAV-Based Aerial Railroad Inspection
abstract
Rail fasteners are crucial railroad infrastructure and their health status is directly connected to the safety of traveling trains. UAV-based rail fastener visual inspection has shown strong advantages over traditional inspection techniques. However, existing general-purpose object detection architectures inevitably suffer from the limitation that they can only detect objects that are clearly visible. Consequently, they tend to neglect objects affected by occlusion or shadows and lead to missed detection problem. The industrial application of such models can bring huge safety risks to long-term operations of safety-sensitive railroad systems. Concerning the issues, this paper proposes a visual prior-guided rail fastener integrity detection architecture (RFIDet) to realize coarse-to-fine detection of all rail fasteners, whether normally visible or visually obscured. RFIDet employs a two-stage pipeline: the visual prior guidance (VPG) stage generates standard rail fastener layout representation (SRFLR) for coarse priors, while the precise location search (PLS) stage enables NMS-free refinement using adaptive anchors designed from actual physical distance priors. SRFLR takes full advantage of inherent spatial priors of all rail fasteners to perceive a unified, interconnected, and coarse location distribution. Then all rail fastener candidates activated by those coarse locations are further trained to search and regress refined offsets to the final bounding boxes. Structural loss functions for both stages are customized to facilitate the detection of individual fasteners while constraining the overall spatial distribution of all fasteners. Experiments have verified the effectiveness and better robustness of the proposed RFIDet with the mAP50value increased by at least 5.9% compared to a series of general-purpose SOTA YOLO detectors. RFIDet outperforms the comparing algorithms especially when coming across unexpected occlusions or shadows.
Limin Jia 0002, Honggui Han, Yong Qin 0002, Haonan Zhang 0002, Zhipeng Wang 0002
IEEE Trans. Intell. Transp. Syst.6
2026 Object-Guided Semi-Supervised Bird's-Eye View 3D Object Detection With 3D Box Refinement
abstract
Recently, Bird’s-Eye View (BEV) representation has received increasing attention in multi-view 3D object detection. However, training high-performance camera-based BEV 3D detectors typically requires cumbersome annotated training samples. In this challenging scenario, we are the pioneers in addressing the problem of semi-supervised BEV 3D object detection. To this end, we propose object-guided semi-supervised BEV 3D object detection (OSS3D), which trains the 3D detectors by using a small amount of labeled data and a large amount of unlabeled data to reduce the cost of data labeling. Firstly, our approach employs an object-guided Gaussian-like mask generation paired with multi-perspective geometric correspondence to emphasize foreground regions and mitigate background noise, ensuring accurate object alignment across different spaces. Additionally, this Gaussian-like mask is integral to our object-guided deep feature map supervision strategy. Secondly, in order to avoid the impact of unreliable pseudo-labels on student network training and improve the importance of object depth information in real space, we propose a self-training depth-based refinement algorithm to generate high-quality and stable pseudo-labels. Finally, our method can be easily inserted into a variety of BEV 3D detection networks. Several experiments demonstrate that the proposed method outperforms current semi-supervised approaches by 2.59% mAP and 2.94% NDS in the nuScenes dataset, achieving new state-of-the-art (SOTA) compared to various camera-based 3D detectors. Code will be available athttps://github.com/yangzhaojason/OSS3D
Yinan Shi, Jiangtong Zhu, Haonan Zhang 0002, Longjun Liu
IEEE Trans. Intell. Transp. Syst.5
2025 D2S: Towards Efficient Sparse 3D Object Detection via Dense to Sparse Knowledge Distillation
abstract
LiDAR-based 3D object detection is widely used in high-level autonomous driving schemes. However, the cumbersome modules in most 3D detectors lead to substantial computational overhead. Despite knowledge distillation (KD) is an effective approach for compressing models, previous methods cannot be extended to the dense-to-sparse paradigm. To this end, we propose a simple yet effective Dense to Sparse Knowledge Distillation (D2S) framework for accelerating 3D detectors. Firstly, to compensate for the difference in predicted location between dense and sparse detectors, we introduce a lightweight feature diffusion (FeaD) module for spreading important features. Secondly, to achieve high performance, we propose a dual-stream distillation scheme to transfer knowledge. In this scheme, we align both of the feature and category prediction between distillation pairs at important positions. Extensive experiments on KITTI and Waymo Open Dataset demonstrate the effectiveness of our method. For example, on KITTI dataset, the sparse detector we obtained surpasses VoxelNeXt with around 2.0× fewer parameters and 1.6× fewer FLOPs.
Longjun Liu, Yingke Gao, Haonan Zhang 0002, Haoteng Li
ICASSP4
2025 Self-supervised motion forecasting with local information interaction in autonomous driving
Longjun Liu, Haoteng Li, Haonan Zhang 0002
Appl. Intell.4
2025 CLEAN: Category Knowledge-Driven Compression Framework for Efficient 3D Object Detection
abstract
Deep neural networks (DNNs) are potent in LiDAR-based 3D object detection (LiDAR-3DOD), yet their deployment remains daunting due to their cumbersome parameters and computations. Knowledge distillation (KD) is promising for compressing DNNs in LiDAR-3DOD. However, most existing KD methods transfer inadequate knowledge between homogeneous detectors, and do not thoroughly explore optimal student architectures, resulting in insufficient gains for compact student detectors. To this end, we propose a category knowledge-driven compression framework to achieve efficient LiDAR-based 3D detectors. Firstly, we distill knowledge from two-stage teacher detectors to one-stage student detectors, overcoming the limitations of homogeneous pairs. To conduct KD in these heterogeneous pairs, we explore the gap between heterogeneous detectors, and introduce category knowledge-driven KD (CaKD), which includes both student-oriented distillation and two-stage-oriented label assignment distillation. Secondly, to search for the optimal architecture of compact student detectors, we introduce a masked category knowledge-driven structured pruning scheme. This scheme evaluates filter importance by analyzing the changes in category predictions related to foreground regions before and after filter removal, and prunes the less important filters accordingly. Finally, we propose a modified IoU-aware redundancy elimination module to remove redundant false positive samples, thereby further improving the accuracy of detectors. Experiments on various point cloud datasets demonstrate that our method delivers impressive results. For example, on KITTI, several compressed one-stage detectors outperform two-stage detectors in both efficiency and accuracy. Besides, on WOD-mini, our framework reduces the memory footprint of CenterPoint by 5.2× and improves the L2 mAPH by 0.55$\%$%.
Haonan Zhang 0002, Longjun Liu, Fei Hui, Hengmin Zhang, Zhiyuan Zha
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 DenseKD: Dense Knowledge Distillation by Exploiting Region and Sample Importance
abstract
Knowledge distillation (KD) can compress deep neural networks (DNNs) by transferring the knowledge of the redundant teacher model to the resource-friendly student model, where cross-layer KD (CKD) conducts KD between each stage of students and the multiple stages of teachers. However, previous CKD schemes select the coarse-grained stagewise features of teachers to teach students, leading to improper channel alignment. Also, most of these methods conduct uniform distillation for all the knowledge, limiting students to focus more on important knowledge. To address these problems, we propose a dense KD (DenseKD) in this article, dubbed as DenseKD. First, to achieve more accurate feature alignment in CKD, we construct the learnable dense architecture to make each channel of student flexibly capture more diverse channelwise features from teacher. Moreover, we introduce region importance to investigate the region's guiding potential, it distinguishes the influence of different regions by the variation of representations of teacher models. In addition, to make students pay more attention to useful samples in KD, we calculate sample importance by the loss of teacher models. Consistent improvements over state-of-the-art approaches are observed in experiments on multiple vision tasks. For example, in the classification task, DenseKD achieves 72.30% accuracy of ResNet-20 on CIFAR-100, which is higher than the results of previous CKD methods. In addition, in the object detection task, DenseKD gains 2.84% mean average precision (mAP) improvements of Faster R-CNN with ResNet-18 against vanilla KD.
Haonan Zhang 0002, Longjun Liu, Yi Zhang 0140, Fei Hui, Bihan Wen
IEEE Trans. Neural Networks Learn. Syst.1
2025 LeOp-GS: Learned Optimizer With Dynamic Gradient Update for Sparse-View 3DGS
abstract
3D Gaussian Splatting (3DGS) achieves remarkable speed and performance in novel view synthesis but suffers from overfitting and degraded reconstruction when handling sparse-view inputs. This paper innovatively addresses this challenge from a learning-to-optimize perspective by leveraging a learned optimizer (i.e., a multi-layer perceptron, MLP) to update the relevant parameters of 3DGS during the optimization process. Evidently, using a single MLP to handle all optimization variables, whose numbers may even vary during the optimization process, is impossible. Therefore, we present a point-wise position-aware optimizer that updates the parameters for each 3DGS point individually. Specifically, it takes the point coordinates and corresponding parameter values as input to predict the updates, thereby allowing efficient and adaptive optimization. In the case of sparse view modeling, the learned optimizer imposes position-aware constraints on the parameter updates during optimization. This effectively encourages the relevant parameters to converge stably to better solutions. To update the optimizer's parameters, we propose a dynamic gradient update strategy based on spatial perturbation and weighted fusion, enabling the optimizer to capture broader contextual information. Experiments demonstrate that our method effectively addresses the problem of modeling 3DGS from sparse training views, achieving state-of-the-art results across multiple datasets.
Longjun Liu, Haoteng Li, Haonan Zhang 0002
IEEE Trans. Vis. Comput. Graph.5
2024 IS-DARTS: Stabilizing DARTS through Precise Measurement on Candidate Importance
abstract
Among existing Neural Architecture Search methods, DARTS is known for its efficiency and simplicity. This approach applies continuous relaxation of network representation to construct a weight-sharing supernet and enables the identification of excellent subnets in just a few GPU days. However, performance collapse in DARTS results in deteriorating architectures filled with parameter-free operations and remains a great challenge to the robustness. To resolve this problem, we reveal that the fundamental reason is the biased estimation of the candidate importance in the search space through theoretical and experimental analysis, and more precisely select operations via information-based measurements. Furthermore, we demonstrate that the excessive concern over the supernet and inefficient utilization of data in bi-level optimization also account for suboptimal results. We adopt a more realistic objective focusing on the performance of subnets and simplify it with the help of the informationbased measurements. Finally, we explain theoretically why progressively shrinking the width of the supernet is necessary and reduce the approximation error of optimal weights in DARTS. Our proposed method, named IS-DARTS, comprehensively improves DARTS and resolves the aforementioned problems. Extensive experiments on NAS-Bench-201 and DARTS-based search space demonstrate the effectiveness of IS-DARTS.
Hongyi He, Longjun Liu, Haonan Zhang 0002, Nanning Zheng 0001
AAAI3
2024 CaKDP: Category-Aware Knowledge Distillation and Pruning Framework for Lightweight 3D Object Detection
abstract
Knowledge distillation (KD) possesses immense potential to accelerate the deep neural networks (DNNs) for LiDAR-based 3D detection. However, in most of prevailing approaches, the suboptimal teacher models and insufficient student architecture investigations limit the performance gains. To address these issues, we propose a simple yet effective Category-aware Knowledge Distillation and Pruning (CaKDP) framework for compressing 3D detectors. Firstly, CaKDP transfers the knowledge of two-stage detector to one-stage student one, mitigating the impact of inadequate teacher models. To bridge the gap between the heterogeneous detectors, we investigate their differences, and then introduce the student-motivated category-aware KD to align the category prediction between distillation pairs. Secondly, we propose a category-aware pruning scheme to obtain the customizable architecture of compact student model. The method calculates the category prediction gap before and after removing each filter to evaluate the importance of filters, and retains the important filters. Finally, to further improve the student performance, a modified IOU-aware refinement module with negligible computations is leveraged to remove the redundant false positive predictions. Experiments demonstrate that CaKDP achieves the compact detector with high performance. For example, on WOD, CaKDP accelerates CenterPoint by half while boosting L2 mAPH by 1.61%. The code is available at https://github.com/zhnxjtu/CaKDP.
Haonan Zhang 0002, Longjun Liu, Bihan Wen
CVPR1
2023 Cross-Layer Patch Alignment and Intra-and-Inter Patch Relations for Knowledge Distillation
abstract
To date, most existing distillation methods exploit coarse-grained instance-level information as valuable knowledge to transfer, such as instance logits, instance features and instance relations. However, the fine-grained knowledge of internal regions and relationships between semantic entities within a single instance are overlooked and not fully explored. To address above limitations, we propose a novel fine-grained patch-level distillation method, dubbed as Patch Aware Knowledge Distillation. PAKD rethink knowledge distillation from a new perspective regarding the significance of cross-layer patch alignment and patch relations within and across instances. Specifically, we first devise a novel cross-layer architecture to fuse patches across stages, which is capable of utilizing multi-level information of the teacher to guide one-level learning of the student. Then, we propose cross-layer patch alignment, allowing the student to be aware of patches discriminatively and find the best way to learn from the teacher. Besides, patch relations within and across instances are leveraged to supervise the structural knowledge distillation in the manifold space. We apply our method to image classification and object detection tasks. Consistent improvements over state-of-the-art approaches on different datasets and diverse teacher-student combinations manifest the great potential of our proposed PAKD.
Yi Zhang 0140, Yingke Gao, Haonan Zhang 0002, Longjun Liu
ICIP3
2023 RepCo: Replenish sample views with better consistency for contrastive learning
Longjun Liu, Yi Zhang 0140, Puhang Jia, Haonan Zhang 0002, Nanning Zheng 0001
Neural Networks5
2023 Hierarchical Model Compression via Shape-Edge Representation of Feature Maps - an Enlightenment From the Primate Visual System
abstract
The cumbersome computation of deep neural networks (DNNs) limits their practical deployment on resource-constrained mobile multimedia devices. To deploy DNNs on devices with limited computing resources, model compression techniques are leveraged to accelerate the networks, where network pruning can improve the inference efficiency of DNNs by removing redundant weights and structures. As one of the important components of DNNs, the feature maps (FMs) can be leveraged to evaluate the importance of network structures for DNN pruning. However, previous methods neglect to fully explore the characteristics of FMs in network pruning. In this paper, we investigate the high capacity and resource efficient analogy-ventral dual-pathway primates visual system (PVS) to propose a hierarchical pruning framework (dubbed as HPSE). In an efficient PVS, the analog pathway analyzes low-frequency information to facilitate the high-frequency information inference in ventral stream. In HPSE, we extract the low-frequency shape information and high-frequency edge information from FMs to present a novel pruning pipeline that resembles the analysis mechanism of PVS. In particular, we first imitate the analogy pathway to group different FMs in each layer by calculating the shape-feature overlap. Secondly, we leverage the edge information modulated by the grouping results of the first step to prune the network. The effectiveness of HPSE is verified by pruning various DNNs on different benchmarks. For example, for ResNet-56 on CIFAR-10, HPSE reduces 52.9% of FLOPs with a slight accuracy improvement; for ResNet-50 on ImageNet, we achieve 54.3%-FLOPs drop with only 0.49% Top-1 accuracy loss.
Haonan Zhang 0002, Longjun Liu, Bingyao Kang, Nanning Zheng 0001
IEEE Trans. Multim.1
2022 A Novel Differentiable Mixed-Precision Quantization Search Framework for Alleviating the Matthew Effect and Improving Robustness
Hengyi Zhou, Hongyi He, Wanchen Liu, Yuhai Li, Haonan Zhang 0002, Longjun Liu
ACML5
2022 CMB: A Novel Structural Re-parameterization Block without Extra Training Parameters
abstract
Structural re-parameterization is a raising field, which aims at improving the performance of convolutional neural networks (CNNs) through training an over-parameterization model and transferring it into a compact inference model. However, the performance improvements of prior structural re-parameterization works often come at the cost of heavy extra training resources, which increases carbon emissions and limits the potential applications on large-scale industrial tasks. To this end, first, we conduct experiments with a series of blocks composed of multiple identical branches to investigate the mechanism behind the structural re-parameterization, and then provide an interpretation. Moreover, motivated by the studies of effective receptive fields in the biological visual systems and neural networks, we propose a novel compact block named circular mask block (CMB). Given a neural network, we replace the regular convolutional layer with CMB to construct a training architecture, which can be trained to gain an accuracy boost with No extra training parameters and limited extra training FLOPs. After training, the training architecture can be transformed into the original architecture for inference. Extensive experiments are performed on CIFAR-10 and ImageNet to evaluate the effectiveness of our method. For example, we improve 0.85% top-1 accuracy of ResNet-50 on ImageNet without extra training parameters and only 11.32M extra training FLOPs, which saves 434x training FLOPs compared with prior works.
Hengyi Zhou, Longjun Liu, Haonan Zhang 0002, Hongyi He, Nanning Zheng 0001
IJCNN3
2022 Rethinking the Mechanism of the Pattern Pruning and the Circle Importance Hypothesis
abstract
Network pruning is an effective and widely-used model compression technique. Pattern pruning is a new sparsity dimension pruning approach whose compression ability has been proven in some prior works. However, a detailed study on "pattern" and pattern pruning is still lacking. In this paper, we analyze the mechanism behind pattern pruning. Our analysis reveals that the effectiveness of pattern pruning should be attributed to finding the less important weights even before training. Then, motivated by the fact that the retinal ganglion cells in the biological visual system have approximately concentric receptive fields, we further investigate and propose the Circle Importance Hypothesis to guide the design of efficient patterns. We also design two series of special efficient patterns - circle patterns and semicircle patterns. Moreover, inspired by the neural architecture search technique, we propose a novel one-shot gradient-based pattern pruning algorithm. Besides, we also expand depthwise convolutions with our circle patterns, which improves the accuracy of networks with little extra memory cost. Extensive experiments are performed to validate our hypotheses and the effectiveness of the proposed methods. For example, we reduce the 44.0% FLOPS of ResNet-56 while improving its accuracy to 94.38% on CIFAR-10. And we reduce the 41.0% FLOPS of ResNet-18 with only a 1.11% accuracy drop on ImageNet.
Hengyi Zhou, Longjun Liu, Haonan Zhang 0002, Nanning Zheng 0001
ACM Multimedia3
2022 DFSNet: Dividing-fuse deep neural networks with searching strategy for distributed DNN architecture
Wenxuan Hou 0001, Longjun Liu, Haonan Zhang 0002, Hongbin Sun 0001, Nanning Zheng 0001
Neurocomputing3
2022 CMD: controllable matrix decomposition with global optimization for deep neural network compression
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Hongbin Sun 0001, Nanning Zheng 0001
Mach. Learn.1
2022 FCHP: Exploring the Discriminative Feature and Feature Correlation of Feature Maps for Hierarchical DNN Pruning and Compression
abstract
Pruning can remove the redundant parameters and structures of Deep Neural Networks (DNNs) to reduce inference time and memory overhead. As one of the important components of DNN, feature maps (FMs) have been widely used in network pruning. However, previous approaches do not fully investigate the discriminative features in FMs, and also do not explicitly utilize all the features associated with each layer in the pruning procedure. In this paper, we explore the discriminative feature of FMs and explicitly investigate the two-adjacent-layer features of each layer to propose a three-phase hierarchical pruning framework, dubbed as FCHP. Firstly, we decompose each FM into several components to extract the discriminative feature. After that, since pruning each layer is related to the FMs of adjacent layers, we explicitly calculate the feature correlation of discriminative features of two adjacent layers, and then use the feature correlation to cluster FMs into several hierarchies to guide subsequent pruning. Finally, we compute the content of discriminative features, and remove channels corresponding to FMs with fewer discriminative features in each hierarchy, respectively. In the experiment, we prune DNNs with the multiple types of architecture on different benchmarks, and the results have achieved the state-of-the-arts in terms of compressed parameters and FLOPs drop. For example, as for ResNet-56 on CIFAR-10, FCHP respectively obtains 50% of parameters and FLOPs reduction with negligible accuracy loss. Besides, as for ResNet-50 on ImageNet, FCHP reduces 40.5% of parameters and 44.1% of FLOPs with 0.43% of Top-1 accuracy drop.
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Liang Si, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 AKECP: Adaptive Knowledge Extraction from Feature Maps for Fast and Efficient Channel Pruning
abstract
Pruning can remove redundant parameters and structures of Deep Neural Networks (DNNs) to reduce inference time and memory overhead. As an important component of neural networks, the feature map (FM) has stated to be adopted for network pruning. However, the majority of FM-based pruning methods do not fully investigate effective knowledge in the FM for pruning. In addition, it is challenging to design a robust pruning criterion with a small number of images and achieve parallel pruning due to the variability of FMs. In this paper, we propose Adaptive Knowledge Extraction for Channel Pruning (AKECP), which can compress the network fast and efficiently. In AKECP, we first investigate the characteristics of FMs and extract effective knowledge with an adaptive scheme. Secondly, we formulate the effective knowledge of FMs to measure the importance of corresponding network channels. Thirdly, thanks to the effective knowledge extraction, AKECP can efficiently and simultaneously prune all the layers with extremely few or even one image. Experimental results show that our method can compress various networks on different datasets without introducing additional constraints, and it has advanced the state-of-the-arts. Notably, for ResNet-110 on CIFAR-10, AKECP achieves 59.9% of parameters and 59.8% of FLOPs reduction with negligible accuracy loss. For ResNet-50 on ImageNet, AKECP saves 40.5% of memory footprint and reduces 44.1% of FLOPs with only 0.32% of Top-1 accuracy drop.
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Wenxuan Hou 0001, Hongbin Sun 0001, Nanning Zheng 0001
ACM Multimedia1