EDBT 2026 Demo / reviewers in the wild / expert
Longjun Liu
dblp:85/1790
· DBLP profile ↗
45ranked-venue papers
8as first author
30since 2021 · last 2026
0000-0002-7467-4994ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 17 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explicit-implicit dense to sparse knowledge distillation for efficient sparse 3D detection
Haonan Zhang 0002, Kang Ke, Yingke Gao, Longjun Liu |
Pattern Recognit. | 5 |
| 2026 | Efficient vision-based occupancy prediction with knowledge distillation
Kang Ke, Haonan Zhang 0002, Penghui Fan, Longjun Liu |
Pattern Recognit. | 5 |
| 2026 | Object-Guided Semi-Supervised Bird's-Eye View 3D Object Detection With 3D Box RefinementabstractRecently, Bird’s-Eye View (BEV) representation has received increasing attention in multi-view 3D object detection. However, training high-performance camera-based BEV 3D detectors typically requires cumbersome annotated training samples. In this challenging scenario, we are the pioneers in addressing the problem of semi-supervised BEV 3D object detection. To this end, we propose object-guided semi-supervised BEV 3D object detection (OSS3D), which trains the 3D detectors by using a small amount of labeled data and a large amount of unlabeled data to reduce the cost of data labeling. Firstly, our approach employs an object-guided Gaussian-like mask generation paired with multi-perspective geometric correspondence to emphasize foreground regions and mitigate background noise, ensuring accurate object alignment across different spaces. Additionally, this Gaussian-like mask is integral to our object-guided deep feature map supervision strategy. Secondly, in order to avoid the impact of unreliable pseudo-labels on student network training and improve the importance of object depth information in real space, we propose a self-training depth-based refinement algorithm to generate high-quality and stable pseudo-labels. Finally, our method can be easily inserted into a variety of BEV 3D detection networks. Several experiments demonstrate that the proposed method outperforms current semi-supervised approaches by 2.59% mAP and 2.94% NDS in the nuScenes dataset, achieving new state-of-the-art (SOTA) compared to various camera-based 3D detectors. Code will be available athttps://github.com/yangzhaojason/OSS3D Yinan Shi, Jiangtong Zhu, Haonan Zhang 0002, Longjun Liu |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | D2S: Towards Efficient Sparse 3D Object Detection via Dense to Sparse Knowledge DistillationabstractLiDAR-based 3D object detection is widely used in high-level autonomous driving schemes. However, the cumbersome modules in most 3D detectors lead to substantial computational overhead. Despite knowledge distillation (KD) is an effective approach for compressing models, previous methods cannot be extended to the dense-to-sparse paradigm. To this end, we propose a simple yet effective Dense to Sparse Knowledge Distillation (D2S) framework for accelerating 3D detectors. Firstly, to compensate for the difference in predicted location between dense and sparse detectors, we introduce a lightweight feature diffusion (FeaD) module for spreading important features. Secondly, to achieve high performance, we propose a dual-stream distillation scheme to transfer knowledge. In this scheme, we align both of the feature and category prediction between distillation pairs at important positions. Extensive experiments on KITTI and Waymo Open Dataset demonstrate the effectiveness of our method. For example, on KITTI dataset, the sparse detector we obtained surpasses VoxelNeXt with around 2.0× fewer parameters and 1.6× fewer FLOPs. Longjun Liu, Yingke Gao, Haonan Zhang 0002, Haoteng Li |
ICASSP | 2 |
| 2025 | Dualdiff: Dual-Branch Diffusion Model for Autonomous Driving with Semantic FusionabstractAccurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and background control, fall short in capturing the complexity of the scene and integrating multi-modal information. In this paper, we propose DualDiff, a dual-branch conditional diffusion model designed to enhance multi-view driving scene generation. We introduce Occupancy Ray Sampling (ORS), a semantic-rich 3D representation, alongside numerical driving scene representation, for comprehensive foreground and background control. To improve cross-modal information integration, we propose a Semantic Fusion Attention (SFA) mechanism that aligns and fuses features across modalities. Furthermore, we design a foreground-aware masked (FGM) loss to enhance the generation of tiny objects. DualDiff achieves state-of-the-art performance in FID score, as well as consistently better results in downstream BEV segmentation and 3D object detection tasks. Haoteng Li, Zezhong Qian, Gongpeng Zhao, Jun Yu 0001, Huazheng Zhou, Longjun Liu |
ICRA | 8 |
| 2025 | Towards Accurate Semi-Supervised BEV 3D Object Detection with Depth-Aware Refinement and Denoising-Aided AlignmentabstractRecently, camera-based Bird's-Eye View (BEV) representation has gained significant traction in 3D object detection. However, training high-performance BEV 3D detectors typically requires a large number of annotated samples, which can be costly. Traditional semi-supervised methods for BEV 3D object detection face challenges including loss of rich depth information, inconsistent object representations across spaces, and unreliable pseudo label generation, leading to decreased accuracy and performance. Addressing this challenge, we pioneer the introduction of a semi-supervised BEV 3D object detection framework. Our approach leverages a small set of labeled data alongside a larger set of unlabeled data, significantly reducing annotation costs while maintaining robust detection performance. Firstly, we propose a depth-based self-refinement module to generate high-quality and stable pseudo labels, which can effectively regulate training with noisy labels. Secondly, we designed a denoising labels regression module that integrates denoising for both labeled and unlabeled data. Thirdly, in order to alleviate object inconsistency, we propose a consistent object-guided alignment method to ensure the consistency of objects in multi-spaces. Finally, our method can be easily plugged into various BEV 3D detection networks. Extensive experiments show that the proposed method achieves a new state-of-the-art compared to various camera-based 3D detectors tested on multiple public autonomous driving datasets. Yinan Shi, Jiangtong Zhu, Longjun Liu |
ICRA | 5 |
| 2025 | An FPGA-based Real-Time Optical Flow Accelerator for Recurrent All-Pairs Field TransformsabstractOptical flow can capture the positional changes of pixels between two consecutive frames, thereby enabling the extraction of motion information for objects. Real-time optical flow estimation is widely applied in tasks such as motion estimation, object detection, and tracking. Deep neural network-based optical flow algorithms have made a significant breakthrough in accuracy compared to traditional methods; however, their dense computational requirements hinder real-time deployment on resource-constrained embedded platforms. In this paper, we present ERAFT: a novel lightweight deep neural network architecture based on the Recurrent All-Pairs Field Transforms (RAFT) algorithm, which is more suitable for hardware deployment. Furthermore, we propose a specialized optical flow accelerator based on a prediction mechanism, enabling real-time and efficient optical flow estimation computations. The hardware accelerator was evaluated on the Xilinx VCK190 evaluation board. The results indicate that this accelerator achieves high accuracy on the Middlebury dataset, with speeds of up to 86 frames/s for 640 × 480 pixel images. Xiaoliang Jia, Xiqin Zheng, Yingke Gao, Longjun Liu |
ISCAS | 5 |
| 2025 | Self-supervised motion forecasting with local information interaction in autonomous driving
Longjun Liu, Haoteng Li, Haonan Zhang 0002 |
Appl. Intell. | 2 |
| 2025 | CLEAN: Category Knowledge-Driven Compression Framework for Efficient 3D Object DetectionabstractDeep neural networks (DNNs) are potent in LiDAR-based 3D object detection (LiDAR-3DOD), yet their deployment remains daunting due to their cumbersome parameters and computations. Knowledge distillation (KD) is promising for compressing DNNs in LiDAR-3DOD. However, most existing KD methods transfer inadequate knowledge between homogeneous detectors, and do not thoroughly explore optimal student architectures, resulting in insufficient gains for compact student detectors. To this end, we propose a category knowledge-driven compression framework to achieve efficient LiDAR-based 3D detectors. Firstly, we distill knowledge from two-stage teacher detectors to one-stage student detectors, overcoming the limitations of homogeneous pairs. To conduct KD in these heterogeneous pairs, we explore the gap between heterogeneous detectors, and introduce category knowledge-driven KD (CaKD), which includes both student-oriented distillation and two-stage-oriented label assignment distillation. Secondly, to search for the optimal architecture of compact student detectors, we introduce a masked category knowledge-driven structured pruning scheme. This scheme evaluates filter importance by analyzing the changes in category predictions related to foreground regions before and after filter removal, and prunes the less important filters accordingly. Finally, we propose a modified IoU-aware redundancy elimination module to remove redundant false positive samples, thereby further improving the accuracy of detectors. Experiments on various point cloud datasets demonstrate that our method delivers impressive results. For example, on KITTI, several compressed one-stage detectors outperform two-stage detectors in both efficiency and accuracy. Besides, on WOD-mini, our framework reduces the memory footprint of CenterPoint by 5.2× and improves the L2 mAPH by 0.55$\%$%. Haonan Zhang 0002, Longjun Liu, Fei Hui, Hengmin Zhang, Zhiyuan Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | DenseKD: Dense Knowledge Distillation by Exploiting Region and Sample ImportanceabstractKnowledge distillation (KD) can compress deep neural networks (DNNs) by transferring the knowledge of the redundant teacher model to the resource-friendly student model, where cross-layer KD (CKD) conducts KD between each stage of students and the multiple stages of teachers. However, previous CKD schemes select the coarse-grained stagewise features of teachers to teach students, leading to improper channel alignment. Also, most of these methods conduct uniform distillation for all the knowledge, limiting students to focus more on important knowledge. To address these problems, we propose a dense KD (DenseKD) in this article, dubbed as DenseKD. First, to achieve more accurate feature alignment in CKD, we construct the learnable dense architecture to make each channel of student flexibly capture more diverse channelwise features from teacher. Moreover, we introduce region importance to investigate the region's guiding potential, it distinguishes the influence of different regions by the variation of representations of teacher models. In addition, to make students pay more attention to useful samples in KD, we calculate sample importance by the loss of teacher models. Consistent improvements over state-of-the-art approaches are observed in experiments on multiple vision tasks. For example, in the classification task, DenseKD achieves 72.30% accuracy of ResNet-20 on CIFAR-100, which is higher than the results of previous CKD methods. In addition, in the object detection task, DenseKD gains 2.84% mean average precision (mAP) improvements of Faster R-CNN with ResNet-18 against vanilla KD. Haonan Zhang 0002, Longjun Liu, Yi Zhang 0140, Fei Hui, Bihan Wen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | LeOp-GS: Learned Optimizer With Dynamic Gradient Update for Sparse-View 3DGSabstract3D Gaussian Splatting (3DGS) achieves remarkable speed and performance in novel view synthesis but suffers from overfitting and degraded reconstruction when handling sparse-view inputs. This paper innovatively addresses this challenge from a learning-to-optimize perspective by leveraging a learned optimizer (i.e., a multi-layer perceptron, MLP) to update the relevant parameters of 3DGS during the optimization process. Evidently, using a single MLP to handle all optimization variables, whose numbers may even vary during the optimization process, is impossible. Therefore, we present a point-wise position-aware optimizer that updates the parameters for each 3DGS point individually. Specifically, it takes the point coordinates and corresponding parameter values as input to predict the updates, thereby allowing efficient and adaptive optimization. In the case of sparse view modeling, the learned optimizer imposes position-aware constraints on the parameter updates during optimization. This effectively encourages the relevant parameters to converge stably to better solutions. To update the optimizer's parameters, we propose a dynamic gradient update strategy based on spatial perturbation and weighted fusion, enabling the optimizer to capture broader contextual information. Experiments demonstrate that our method effectively addresses the problem of modeling 3DGS from sparse training views, achieving state-of-the-art results across multiple datasets. Longjun Liu, Haoteng Li, Haonan Zhang 0002 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | IS-DARTS: Stabilizing DARTS through Precise Measurement on Candidate ImportanceabstractAmong existing Neural Architecture Search methods, DARTS is known for its efficiency and simplicity. This approach applies continuous relaxation of network representation to construct a weight-sharing supernet and enables the identification of excellent subnets in just a few GPU days. However, performance collapse in DARTS results in deteriorating architectures filled with parameter-free operations and remains a great challenge to the robustness. To resolve this problem, we reveal that the fundamental reason is the biased estimation of the candidate importance in the search space through theoretical and experimental analysis, and more precisely select operations via information-based measurements. Furthermore, we demonstrate that the excessive concern over the supernet and inefficient utilization of data in bi-level optimization also account for suboptimal results. We adopt a more realistic objective focusing on the performance of subnets and simplify it with the help of the informationbased measurements. Finally, we explain theoretically why progressively shrinking the width of the supernet is necessary and reduce the approximation error of optimal weights in DARTS. Our proposed method, named IS-DARTS, comprehensively improves DARTS and resolves the aforementioned problems. Extensive experiments on NAS-Bench-201 and DARTS-based search space demonstrate the effectiveness of IS-DARTS. Hongyi He, Longjun Liu, Haonan Zhang 0002, Nanning Zheng 0001 |
AAAI | 2 |
| 2024 | CaKDP: Category-Aware Knowledge Distillation and Pruning Framework for Lightweight 3D Object DetectionabstractKnowledge distillation (KD) possesses immense potential to accelerate the deep neural networks (DNNs) for LiDAR-based 3D detection. However, in most of prevailing approaches, the suboptimal teacher models and insufficient student architecture investigations limit the performance gains. To address these issues, we propose a simple yet effective Category-aware Knowledge Distillation and Pruning (CaKDP) framework for compressing 3D detectors. Firstly, CaKDP transfers the knowledge of two-stage detector to one-stage student one, mitigating the impact of inadequate teacher models. To bridge the gap between the heterogeneous detectors, we investigate their differences, and then introduce the student-motivated category-aware KD to align the category prediction between distillation pairs. Secondly, we propose a category-aware pruning scheme to obtain the customizable architecture of compact student model. The method calculates the category prediction gap before and after removing each filter to evaluate the importance of filters, and retains the important filters. Finally, to further improve the student performance, a modified IOU-aware refinement module with negligible computations is leveraged to remove the redundant false positive predictions. Experiments demonstrate that CaKDP achieves the compact detector with high performance. For example, on WOD, CaKDP accelerates CenterPoint by half while boosting L2 mAPH by 1.61%. The code is available at https://github.com/zhnxjtu/CaKDP. Haonan Zhang 0002, Longjun Liu, Bihan Wen |
CVPR | 2 |
| 2024 | AMA: An Analytical Approach to Maximizing the Efficiency of Deep Learning on Versal AI EngineabstractThe traditional cache-based multi-core architecture represented by CUDA has been plagued by the “memory wall” problem, and a large number of applications represented by large language model inference are unable to meet the high computational intensity requirements. The Versal AI Engine architecture provides a variety of rich inter-core connections in the processor array, increasing many data reuse opportunities and potentially alleviating the “memory wall” issue. However, traditional parallel programming models cannot be directly applied to this architecture, and how to map computations to achieve high computational utilization becomes a new challenge. To address this, we propose AMA, a hierarchical performance analysis model built on the Versal AI Engine architecture, designed to maximize the efficiency of typical deep learning applications. Experiments show that AMA modeling is accurate and efficient. On the VCK190 platform, we achieved a matrix multiplication throughput of 5867.29 GFLOPS in fp32 and 88.55 TOPS in int8, and a convolution throughput of 99.6770 TOPS. In terms of energy efficiency, AMA achieved 142.68 GFLOPS/W in fp32 precision and 1.416 TOPS/W in int8 matrix multiplication. Compared to the current state-of-the-art methods, we achieved a $\mathbf{1 4. 9 9 \%}$ increase in throughput and a 22.92% increase in energy efficiency, providing new analytical performance model and practical guidance for efficient deep learning deployment on AI Engine. Xiaodong Deng, Longjun Liu, Nanning Zheng 0001 |
FPL | 5 |
| 2024 | Compensation Architecture to Alleviate Noise Effects in RRAM-based Computing-in-memory Chips with Residual ResourceabstractResistive random access memory (RRAM) is a promising technology for energy-efficient in-memory computing. However, due to technology limits, RRAM device faces a series of reliability issues. Deep neural network (DNN) computing based on RRAM suffers from accuracy degradation. On the one hand, offline DNN training solutions are difficult to fully consider and simulate all nonidealities. Worse still, new error or nonideality may come up with the usage of RRAM, which further deteriorates the effectiveness of offline training. On the other hand, online training poses great challenges on programming overhead and device lifetime. The iterative write-verify technique to program multi-bit RRAM cells prolongs write latency more than 10× longer than read latency. To overcome these issues, we propose a compensation architecture and a software and hardware co-training design to mitigate the realistic network accuracy loss in RRAM-based computing-in-memory chips. Firstly, we add trainable compensation channels in crossbars utilizing the residual resource after original weight mapping. Secondly, an offline training procedure with computing output from hardware is triggered to settle down appropriate weight value in compensation channels. Experimental results demonstrate that the proposed design can guarantee ≤ 0.8% loss of accuracy in DNN on MNIST and CIFAR10 dataset even when nonidealities reduce the original accuracy down to ≤73%. Longjun Liu, Yuyi Liu, Bin Gao 0006, Hongbin Sun 0001 |
ISCAS | 2 |
| 2023 | Cross-Layer Patch Alignment and Intra-and-Inter Patch Relations for Knowledge DistillationabstractTo date, most existing distillation methods exploit coarse-grained instance-level information as valuable knowledge to transfer, such as instance logits, instance features and instance relations. However, the fine-grained knowledge of internal regions and relationships between semantic entities within a single instance are overlooked and not fully explored. To address above limitations, we propose a novel fine-grained patch-level distillation method, dubbed as Patch Aware Knowledge Distillation. PAKD rethink knowledge distillation from a new perspective regarding the significance of cross-layer patch alignment and patch relations within and across instances. Specifically, we first devise a novel cross-layer architecture to fuse patches across stages, which is capable of utilizing multi-level information of the teacher to guide one-level learning of the student. Then, we propose cross-layer patch alignment, allowing the student to be aware of patches discriminatively and find the best way to learn from the teacher. Besides, patch relations within and across instances are leveraged to supervise the structural knowledge distillation in the manifold space. We apply our method to image classification and object detection tasks. Consistent improvements over state-of-the-art approaches on different datasets and diverse teacher-student combinations manifest the great potential of our proposed PAKD. Yi Zhang 0140, Yingke Gao, Haonan Zhang 0002, Longjun Liu |
ICIP | 5 |
| 2023 | Design Hybrid Computing Architecture for Accelerating Point Cloud RegistrationabstractHigh-precision simultaneous localization and mapping (SLAM) is one of the core technologies of unmanned driving. LiDAR-based SLAM algorithms are often complex and computationally intensive, and usually are deployed on high performance CPU or GPU computing architecture with high power consumption and low energy efficiency ratio, which is not conducive to vehicle-level applications. In this paper, we design and implement a low power CPU and FPGA hybrid computing architecture for accelerating the key algorithm of LiDAR-based localization scheme. More specifically, we propose a software and hardware co-design strategy: (1) we first propose chain representation as a new type of map representation, which uses the depth discontinuity region as the segmentation location to segment the point cloud data. Our method not only reduces noise issues for down-sampling operation in point cloud representation, but also has the same computational and storage overhead as point cloud representation. (2) We further exploit the inherent parallelism in the algorithms to design a pipeline hardware architecture, which can effectively improve the speed of the algorithm in the embedded platform. Deployed on the Xilinx ZCU102 platform, our system achieves 24.4x and 3.2x speedups compared to the ARM Cortex A53 processor and the Intel i7-10700 processor, respectively, at 4.204W power consumption without severely degrading the final output quality. Xiao Wang 0002, Xiaodong Deng, Yingxiang Li, Shi-tao Chen, Longjun Liu, Nanning Zheng 0001 |
IV | 5 |
| 2023 | RepCo: Replenish sample views with better consistency for contrastive learning
Longjun Liu, Yi Zhang 0140, Puhang Jia, Haonan Zhang 0002, Nanning Zheng 0001 |
Neural Networks | 2 |
| 2023 | Hierarchical Model Compression via Shape-Edge Representation of Feature Maps - an Enlightenment From the Primate Visual SystemabstractThe cumbersome computation of deep neural networks (DNNs) limits their practical deployment on resource-constrained mobile multimedia devices. To deploy DNNs on devices with limited computing resources, model compression techniques are leveraged to accelerate the networks, where network pruning can improve the inference efficiency of DNNs by removing redundant weights and structures. As one of the important components of DNNs, the feature maps (FMs) can be leveraged to evaluate the importance of network structures for DNN pruning. However, previous methods neglect to fully explore the characteristics of FMs in network pruning. In this paper, we investigate the high capacity and resource efficient analogy-ventral dual-pathway primates visual system (PVS) to propose a hierarchical pruning framework (dubbed as HPSE). In an efficient PVS, the analog pathway analyzes low-frequency information to facilitate the high-frequency information inference in ventral stream. In HPSE, we extract the low-frequency shape information and high-frequency edge information from FMs to present a novel pruning pipeline that resembles the analysis mechanism of PVS. In particular, we first imitate the analogy pathway to group different FMs in each layer by calculating the shape-feature overlap. Secondly, we leverage the edge information modulated by the grouping results of the first step to prune the network. The effectiveness of HPSE is verified by pruning various DNNs on different benchmarks. For example, for ResNet-56 on CIFAR-10, HPSE reduces 52.9% of FLOPs with a slight accuracy improvement; for ResNet-50 on ImageNet, we achieve 54.3%-FLOPs drop with only 0.49% Top-1 accuracy loss. Haonan Zhang 0002, Longjun Liu, Bingyao Kang, Nanning Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | A Novel Differentiable Mixed-Precision Quantization Search Framework for Alleviating the Matthew Effect and Improving Robustness
Hengyi Zhou, Hongyi He, Wanchen Liu, Yuhai Li, Haonan Zhang 0002, Longjun Liu |
ACML | 6 |
| 2022 | CMB: A Novel Structural Re-parameterization Block without Extra Training ParametersabstractStructural re-parameterization is a raising field, which aims at improving the performance of convolutional neural networks (CNNs) through training an over-parameterization model and transferring it into a compact inference model. However, the performance improvements of prior structural re-parameterization works often come at the cost of heavy extra training resources, which increases carbon emissions and limits the potential applications on large-scale industrial tasks. To this end, first, we conduct experiments with a series of blocks composed of multiple identical branches to investigate the mechanism behind the structural re-parameterization, and then provide an interpretation. Moreover, motivated by the studies of effective receptive fields in the biological visual systems and neural networks, we propose a novel compact block named circular mask block (CMB). Given a neural network, we replace the regular convolutional layer with CMB to construct a training architecture, which can be trained to gain an accuracy boost with No extra training parameters and limited extra training FLOPs. After training, the training architecture can be transformed into the original architecture for inference. Extensive experiments are performed on CIFAR-10 and ImageNet to evaluate the effectiveness of our method. For example, we improve 0.85% top-1 accuracy of ResNet-50 on ImageNet without extra training parameters and only 11.32M extra training FLOPs, which saves 434x training FLOPs compared with prior works. Hengyi Zhou, Longjun Liu, Haonan Zhang 0002, Hongyi He, Nanning Zheng 0001 |
IJCNN | 2 |
| 2022 | Rethinking the Mechanism of the Pattern Pruning and the Circle Importance HypothesisabstractNetwork pruning is an effective and widely-used model compression technique. Pattern pruning is a new sparsity dimension pruning approach whose compression ability has been proven in some prior works. However, a detailed study on "pattern" and pattern pruning is still lacking. In this paper, we analyze the mechanism behind pattern pruning. Our analysis reveals that the effectiveness of pattern pruning should be attributed to finding the less important weights even before training. Then, motivated by the fact that the retinal ganglion cells in the biological visual system have approximately concentric receptive fields, we further investigate and propose the Circle Importance Hypothesis to guide the design of efficient patterns. We also design two series of special efficient patterns - circle patterns and semicircle patterns. Moreover, inspired by the neural architecture search technique, we propose a novel one-shot gradient-based pattern pruning algorithm. Besides, we also expand depthwise convolutions with our circle patterns, which improves the accuracy of networks with little extra memory cost. Extensive experiments are performed to validate our hypotheses and the effectiveness of the proposed methods. For example, we reduce the 44.0% FLOPS of ResNet-56 while improving its accuracy to 94.38% on CIFAR-10. And we reduce the 41.0% FLOPS of ResNet-18 with only a 1.11% accuracy drop on ImageNet. Hengyi Zhou, Longjun Liu, Haonan Zhang 0002, Nanning Zheng 0001 |
ACM Multimedia | 2 |
| 2022 | DFSNet: Dividing-fuse deep neural networks with searching strategy for distributed DNN architecture
Wenxuan Hou 0001, Longjun Liu, Haonan Zhang 0002, Hongbin Sun 0001, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2022 | CMD: controllable matrix decomposition with global optimization for deep neural network compression
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Hongbin Sun 0001, Nanning Zheng 0001 |
Mach. Learn. | 2 |
| 2022 | FCHP: Exploring the Discriminative Feature and Feature Correlation of Feature Maps for Hierarchical DNN Pruning and CompressionabstractPruning can remove the redundant parameters and structures of Deep Neural Networks (DNNs) to reduce inference time and memory overhead. As one of the important components of DNN, feature maps (FMs) have been widely used in network pruning. However, previous approaches do not fully investigate the discriminative features in FMs, and also do not explicitly utilize all the features associated with each layer in the pruning procedure. In this paper, we explore the discriminative feature of FMs and explicitly investigate the two-adjacent-layer features of each layer to propose a three-phase hierarchical pruning framework, dubbed as FCHP. Firstly, we decompose each FM into several components to extract the discriminative feature. After that, since pruning each layer is related to the FMs of adjacent layers, we explicitly calculate the feature correlation of discriminative features of two adjacent layers, and then use the feature correlation to cluster FMs into several hierarchies to guide subsequent pruning. Finally, we compute the content of discriminative features, and remove channels corresponding to FMs with fewer discriminative features in each hierarchy, respectively. In the experiment, we prune DNNs with the multiple types of architecture on different benchmarks, and the results have achieved the state-of-the-arts in terms of compressed parameters and FLOPs drop. For example, as for ResNet-56 on CIFAR-10, FCHP respectively obtains 50% of parameters and FLOPs reduction with negligible accuracy loss. Besides, as for ResNet-50 on ImageNet, FCHP reduces 40.5% of parameters and 44.1% of FLOPs with 0.43% of Top-1 accuracy drop. Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Liang Si, Hongbin Sun 0001, Nanning Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Exploring Effective DNN Models for Forensic Age Estimation based on Panoramic Radiograph ImagesabstractDental age estimation is widely used in forensic identification, but the accuracy of traditional methods cannot satisfy the demand for accuracy, especially for age estimation of adults. We introduce a deep learning-based methodology to estimate the age based on collected X-ray images of the teeth. We present a new dental dataset, which contains labeled orthopan-tomograms (OPGs) of 27,957 people, including 16,383 OPGs for females as well as 11,574 OPGs for males. All ages range from 0 to 93-year-old with a median of 27. The accuracy of the age labels is guaranteed by the ID card information. Aiming at the characteristics of the dental data itself, we explore various neural network elements that are effective for age estimation, including proper network depth, convolution kernel size, multi-branch structure, and the feature reusing of early layers. Based on the characteristic exploration, we further search models for dental age estimation by using the popular Neural Architecture Search (NAS) method. Experiment results show that our model achieves a mean absolute error (MAE) of 1.64 years, surpass all existing CNN models. Compared with Inception-v4 with an MAE of 1.70 and 20.46B FLOPs (inputs size 384×384), the FLOPs of our model can be reduced by 2.7 times (7.49B FLOPs). To our best knowledge, this is the first study for age estimation by exploring and searching the DNN model. Our results have surpassed legal medical expert-level performance (with an MAE of more than 2) for age estimation. Our methodology and results in this paper are very meaningful to forensic medicine for aging estimation with panoramic radiograph images. Wenxuan Hou 0001, Longjun Liu, Jinxia Gao, Anguo Zhu, Keyang Pan, Hongbin Sun 0001, Nanning Zheng 0001 |
IJCNN | 2 |
| 2021 | AKECP: Adaptive Knowledge Extraction from Feature Maps for Fast and Efficient Channel PruningabstractPruning can remove redundant parameters and structures of Deep Neural Networks (DNNs) to reduce inference time and memory overhead. As an important component of neural networks, the feature map (FM) has stated to be adopted for network pruning. However, the majority of FM-based pruning methods do not fully investigate effective knowledge in the FM for pruning. In addition, it is challenging to design a robust pruning criterion with a small number of images and achieve parallel pruning due to the variability of FMs. In this paper, we propose Adaptive Knowledge Extraction for Channel Pruning (AKECP), which can compress the network fast and efficiently. In AKECP, we first investigate the characteristics of FMs and extract effective knowledge with an adaptive scheme. Secondly, we formulate the effective knowledge of FMs to measure the importance of corresponding network channels. Thirdly, thanks to the effective knowledge extraction, AKECP can efficiently and simultaneously prune all the layers with extremely few or even one image. Experimental results show that our method can compress various networks on different datasets without introducing additional constraints, and it has advanced the state-of-the-arts. Notably, for ResNet-110 on CIFAR-10, AKECP achieves 59.9% of parameters and 59.8% of FLOPs reduction with negligible accuracy loss. For ResNet-50 on ImageNet, AKECP saves 40.5% of memory footprint and reduces 44.1% of FLOPs with only 0.32% of Top-1 accuracy drop. Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Wenxuan Hou 0001, Hongbin Sun 0001, Nanning Zheng 0001 |
ACM Multimedia | 2 |
| 2021 | HSC: Leveraging horizontal shortcut connections for improving accuracy and computational efficiency of lightweight CNN
Anguo Zhu, Longjun Liu, Wenxuan Hou 0001, Hongbin Sun 0001, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2021 | Dynamic Dataflow Scheduling and Computation Mapping Techniques for Efficient Depthwise Separable Convolution AccelerationabstractDepthwise separable convolution (DSC) has become one of the essential structures for lightweight convolutional neural networks. Nevertheless, its hardware architecture has not received much attention. Several previous hardware designs incur either high off-chip memory traffic or large on-chip memory usage, and hence have deficiency in terms of hardware efficiency as well as performance. This paper proposes two efficient dynamic design techniques, i.e. adaptive row-based dataflow scheduling and adaptive computation mapping, to achieve a much better trade-off between hardware efficiency and performance for DSC-based lightweight CNN accelerator. The effectiveness and efficiency of the proposed dynamic design techniques have been extensively evaluated using six DSC-based lightweight CNNs. Compared with the reference architectures, the simulation results show the proposed architectural techniques can at least reduce on-chip buffer size by 50.4% and improve the performance of convolution calculation by 1.18× while maintaining the minimum off-chip memory traffic. MobileNetV2 is implemented on Zynq UltraScale+ ZCU102 SoC FPGA, and the results show the proposed accelerator can achieve 381.7 frames per second (fps), which is 1.43× of the reference design, and it can save about 36.3% on-chip buffer size compared with the reference design, while maintaining the same off-chip memory traffic. Baoting Li, Xuchong Zhang, Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | Exploring Highly Dependable and Efficient Datacenter Power System Using Hybrid and Hierarchical Energy BuffersabstractThe massive and irregular load surges challenge datacenter power infrastructures. As a result, power mismatching between supply and demand has emerged as a crucial availability issue in modern datacenters which are either under-provisioned or powered by intermittent power sources. Recent proposals have employed energy storage devices such as the uninterruptible power supply (UPS) to address this issue. However, current approaches lack the capacity of efficiently handling the irregular and unpredictable power mismatches. In this paper, we propose Hybrid and Hierarchical Energy Buffering (HHEB), a novel heterogeneous and adaptive scheme that could enable various energy storage devices (ESDs) to be efficiently integrated into existing datacenters for dynamically dealing with power mismatches. Our techniques exploit the diverse characteristics of different ESDs and intelligent load assignment algorithms to improve the dependability and efficiency of datacenter power systems. We evaluate the HHEB design with a prototype. Compared with a homogenous battery energy buffering system, HHEB could improve energy efficiency by 39.7 percent, extend UPS lifetime by 4.7X, promote energy availability by 3.2X, reduce system downtime by 41 percent, and effectively improve the energy availability of various energy buffers in different hierarchies. It allows datacenters to adapt to various power supply anomalies, thereby improving operational efficiency, dependability and availability. Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001 |
IEEE Trans. Sustain. Comput. | 1 |
| 2020 | Designing Efficient Shortcut Architecture for Improving the Accuracy of Fully Quantized Neural Networks AcceleratorabstractNetwork quantization is an effective solution to compress Deep Neural Networks (DNN) that can be accelerated with custom circuit. However, existing quantization methods suffer from significant loss in accuracy. In this paper, we propose an efficient shortcut architecture to enhance the representational capability of DNN between different convolution layers. We further implement the shortcut hardware architecture to effectively improve the accuracy of fully quantized neural networks accelerator. The experimental results show that our shortcut architecture can obviously improve network accuracy while increasing very few hardware resources ( 0.11 × and 0.17 × for LUT and FF respectively) compared with the whole accelerator. Baoting Li, Longjun Liu, Yanming Jin, Hongbin Sun 0001, Nanning Zheng 0001 |
ASP-DAC | 2 |
| 2020 | Toward Customized Hybrid Fuel-Cell and Battery-powered Mobile Device for Individual UsersabstractRapidly evolving technologies and applications of mobile devices inevitably increase the power demands on the battery. However, the development of batteries can hardly keep pace with the fast-growing demands, leading to short battery life, which becomes the top complaints from customers. In this article, we investigate a novel energy supply technology, fuel cell (FC), and leverage its advantages of providing long-term energy storage to build a hybrid FC-battery power system. Therefore, mobile device operation time is dramatically extended, and users are no longer bothered by battery recharging. We examine real-world smartphone usage data and find that a naive hybrid power system cannot meet many users’ highly diversified power demands. We thus propose an OS-level power management policy that reduces the device power consumption for each power peak to solve this mismatch. This technique trades the quality-of-service (QoS) for a larger FC ratio in the system and thus much longer device operation time. We further observe that the user’s personality largely determines his/her satisfaction with the QoS degradation and the operation time extension. Thus, applying a hybrid system with fixed configuration (i.e., peak throttling level coupled with corresponding FC/battery ratio) fails to satisfy every user. We then explore customized hybrid system configuration based on each individual user’s personality to deliver the optimal satisfaction for him/her. The experimental results show that our personality-aware hybrid FC-battery solution can achieve 4× longer operation time and 25% higher satisfaction score compared to the common setting for state-of-the-art mobile devices. Kaige Yan, Jingweijia Tan, Longjun Liu, Xingyao Zhang 0002, Stanko R. Brankovic, Jinghong Chen, Xin Fu 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2019 | Exploring Hardware Friendly Bottleneck Architecture in CNN for Embedded Computing SystemsabstractIn this paper, we explore how to design lightweight CNN architecture for embedded computing systems. We propose L-Mobilenet model for ZYNQ based hardware platform. L-Mobilenet can adapt well to hardware computing and accelerating, and its network structure is inspired by the state-of-the-art work of Inception-Resnet and Mobilenet-V2, which can effectively reduce parameters and delay while maintaining the accuracy of inference. We deploy our L-Mobilenet model to GPU and ZYNQ embedded platform for fully evaluating the performance of our design. By measuring with cifar10 and cifar100 datasets, L-Mobilenet model is able to gain 3× speed up and 3.7× fewer parameters than MobileNet-V2 while maintaining a similar accuracy. It also can obtain 2× speed up and 1.5× fewer parameters than Shufflenet-V2 while maintaining the same accuracy. Experiments show that our network model can obtain better performance because of the special considerations for hardware accelerating and software-hardware co-design strategies in our L-Mobilenet bottleneck architecture. Xing Lei, Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001 |
ICIP | 2 |
| 2019 | REcache: Efficient Sustainable Energy Management Circuits and Policies for Computing SystemsabstractThe rapidly growing computing systems, such as AI server cluster, IoT devices etc. are facing increasing energy expenditure pressure and the warning of carbon footprint. Designing eco-friendly computing systems which integrated renewable energy sources have attracted considerable attentions recently. Existing schemes either incur green energy efficiency degradation or sacrifice workload performance. This paper proposes REcache (Renewable Energy cache), a sustainable energy management scheme to efficiently utilize green energy for computing systems. Compared to previous proposals, we present a dedicated circuit and energy-aware management policies to coordinate energy harvesting, power management and workload scheduling. We evaluate our scheme through both prototyping and simulation. The experimental results show that the REcache could effectively improve the energy availability 10%, workload performance 5% for different workloads on average. Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001, Tao Li 0006 |
ISCAS | 1 |
| 2019 | Efficient Compression-Based Line Buffer Design for Image/Video Processing CircuitsabstractLine buffer is a typical and major on-chip memory design architecture for image/video processing circuits. As it usually occupies very large on-chip circuit area, it is of great importance to reduce its hardware cost through efficient architecture design. Data compression is a promising technique to improve the hardware efficiency of line buffer architecture. Nevertheless, the previously proposed data compression technique for line buffer architecture only exploits fixed length code (FLC), which actually has the deficiency on compression performance. Instead, this paper explores to efficiently use variable length code in line buffer architecture. By restricting variable length coding within small compression granularity (CG), the proposed compression algorithm not only significantly improves compression performance but also meets the specific requirements in line buffer architecture design. The simple compression algorithm further enables the efficient and fully pipelined VLSI architecture and circuits. Experimental results demonstrate that the proposed compression algorithm achieves 6.67-dB peak signal-to-noise ratio improvement at the compression ratio of 50% and the CG of 16 pixels, compared with FLC design. The VLSI circuits of the proposed compression can achieve the throughput of 4K × 2K at 60 fps with reasonable hardware cost. The use of the proposed compression technique in line buffer architecture can significantly reduce on-chip memory cost while maintaining satisfactory visual quality. Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Exploring Customizable Heterogeneous Power Distribution and Management for DatacenterabstractLarge-scale datacenters are facing increasing pressure of capping their carbon emission and power cost. Many leading-edge studies have started to explore server clusters running on multiple power sources. Existing approaches do not sufficiently consider the fine-grained power delivery to satisfy diverse requirements in datacenter, especially in the multi-tenant/colocation datacenter, which may yield low energy utilization. To address the emerging trend and new requirements, this article proposes a novel Datacenter inner Power Switch Network (DiPSN) to improve datacenter power efficiency and user satisfaction. DiPSN is a reconfigurable and easy-to-scale-out power architecture, which enables datacenter to distribute various power sources in a fine-grained manner. Moreover, a tailored machine learning based power source management framework is proposed for DiPSN to dynamically optimize user customized performance metrics and maximize datacenter revenue. Compared with conventional single-switch power distribution system, our DiPSN can be configured to improve solar energy utilization by 39.6 percent, reduce utility power cost by 11.1 percent and improve workload performance by 33.8 percent. Meanwhile, our design can extend battery lifetime by 9.3 percent. This work could provide valuable guidelines for designing heterogeneous power distribution architecture and management methodology in datacenters for improving user-customizable efficiency, sustainability and economy. Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Tao Li 0006, Nanning Zheng 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Managing Battery Aging for High Energy Availability in Green DatacentersabstractEnergy storage devices (ESD), such as UPS batteries, have been repurposed in datacenter as a promising tuning knob for peak power shaving and power cost reducing. However, batteries progressively aging due to irregular usage patterns, which result in less effective capacity and even pose serious threat to server availability. Nevertheless, prior proposals largely ignore the aging issues of battery which may lead to low energy availability for datacenter servers. To fill this critical void, we thoroughly investigate battery aging on a heavily instrumented prototype system over an observation period of ten months. We propose Battery Anti-Aging Treatment Plus (BAAT-P), a novel power delivery architecture included aging management algorithms from the perspective of computing system to hide, reduce, mitigate and plan the battery aging effects for high energy availability in datacenter. Our techniques exploit diverse battery aging mechanisms and dynamic aging management algorithms to provide system-level availability guarantee for datacenter. We evaluate the BAAT-P design with a real prototype. Compared with a battery powered datacenter without aging management policies, the results show that BAAT-P can extend battery lifetime by 72 percent, reduce battery cost by 33 percent and effectively improve energy availability for datacenter servers while maintaining workload performance for the performance critical workloads. Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | HOPE: Enabling Efficient Service Orchestration in Software-Defined Data CentersabstractThe functional scope of today's software-defined data centers (SDDC) has expanded to such an extent that servers face a growing amount of critical background operational tasks like load monitoring, logging, migration, and duplication, etc. These ancillary operations, which we refer to as management operations, often nibble the stringent data center power envelope and exert a tremendous amount of pressure on front-end user tasks. However, existing power capping, peak shaving, and time shifting mechanisms mainly focus on managing data center power demand at the "macro level" -- they do not distinguish ancillary background services from user tasks, and therefore often incur significant performance degradation and energy overhead. Yang Hu 0001, Chao Li 0009, Longjun Liu, Tao Li 0006 |
ICS | 3 |
| 2016 | Towards an Adaptive Multi-Power-Source DatacenterabstractBig data and cloud computing are accelerating the capacity growth of datacenters all over the world. Their energy costs and environmental issues have pushed datacenter operators to explore and integrate alternative energy sources, such as various renewable energy supplies and energy storage devices. Designing datacenters powered by multi-power supplies in the smart grid environment is becoming a promising trend in the next few decades. However, gracefully provisioning various power sources and efficiently manage them in datacenter is a significant challenge. Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Nanning Zheng 0001, Tao Li 0006 |
ICS | 1 |
| 2016 | RE-UPS: an adaptive distributed energy storage system for dynamically managing solar energy in green datacenters
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Jingmin Xin, Nanning Zheng 0001, Tao Li 0006 |
J. Supercomput. | 1 |
| 2015 | BAAT: Towards Dynamically Managing Battery Aging in Green DatacentersabstractEnergy storage devices (batteries) have shown great promise in eliminating supply/demand power mismatch and reducing energy/power cost in green datacenters. These important components progressively age due to irregular usage patterns, which result in less effective capacity and even pose serious threat to server availability. Nevertheless, prior proposals largely ignore the aging issue of batteries or simply use ad-hoc discharge capping to extend their lifetime. To fill this critical void, we thoroughly investigate battery aging on a heavily instrumented prototype over an observation period of six months. We propose battery anti-aging treatment (BAAT), a novel framework for hiding, reducing, and planning the battery aging effects. We show that BAAT can extend battery lifetime by 69%. It enables datacenters to maximally utilize energy storage resources to enhance availability and boost performance. Moreover, it reduces 26% battery cost and allows datacenters to economically scale in the big data era. Longjun Liu, Chao Li 0009, Hongbin Sun 0001, Yang Hu 0001, Juncheng Gu, Tao Li 0006 |
DSN | 1 |
| 2015 | Towards sustainable in-situ server systems in the big data eraabstractRecent years have seen an explosion of data volumes from a myriad of distributed sources such as ubiquitous cameras and various sensors. The challenges of analyzing these geographically dispersed datasets are increasing due to the significant data movement overhead, time-consuming data aggregation, and escalating energy needs. Rather than constantly move a tremendous amount of raw data to remote warehouse-scale computing systems for processing, it would be beneficial to leverage in-situ server systems (InS) to pre-process data, i.e., bringing computation to where the data is located. Chao Li 0009, Yang Hu 0001, Longjun Liu, Juncheng Gu, Mingcong Song, Xiaoyao Liang, Jingling Yuan, Tao Li 0006 |
ISCA | 3 |
| 2015 | HEB: deploying and managing hybrid energy buffers for improving datacenter efficiency and economyabstractToday, an increasing number of applications and services are being hosted by large-scale data centers. The massive and irregular load surges challenge data center power infrastructures. As a result, power mismatching between supply and demand has emerged as a crucial issue in modern data centers which are either under-provisioned or powered by intermittent power sources. Recent proposals have employed energy storage devices such as the uninterruptible power supply (UPS) systems to address this issue. However, current approaches lack the capacity of efficiently handling the irregular and unpredictable power mismatches. Longjun Liu, Chao Li 0009, Hongbin Sun 0001, Yang Hu 0001, Juncheng Gu, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001 |
ISCA | 1 |
| 2014 | Towards Automated Provisioning and Emergency Handling in Renewable Energy Powered Datacenters
Chao Li 0009, Rui Wang 0014, Yang Hu 0001, Ruijin Zhou, Ming Liu 0006, Longjun Liu, Jingling Yuan, Tao Li 0006, Depei Qian 0001 |
J. Comput. Sci. Technol. | 6 |
| 2013 | Enabling datacenter servers to scale out economically and sustainablyabstractAs cloud applications proliferate and data-processing demands increase, server resources must grow to unleash the performance of emerging workloads that scale well with large number of compute nodes. Nevertheless, power has become a crucial bottleneck that restricts horizontal scaling (scale out) of server systems, especially in datacenters that employ power over-subscription. When a datacenter hits the maximum capacity of its power provisioning equipment, the owner has to either build another facility or upgrade existing utility power infrastructure -- both approaches add huge capital expenditure, require significant construction lead time, and can further increase the owner's carbon footprint. Chao Li 0009, Yang Hu 0001, Ruijin Zhou, Ming Liu 0006, Longjun Liu, Jingling Yuan, Tao Li 0006 |
MICRO | 5 |