Hongbin Sun 0001

dblp:98/6690-1 · DBLP profile ↗
← Back
85ranked-venue papers
4as first author
44since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 30 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 11 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Expand Your SCOPE: Semantic Cognition over Potential-Based Exploration for Embodied Visual Navigation
abstract
Embodied visual navigation remains a challenging task, as agents must explore unknown environments with limited knowledge. Existing zero-shot studies have shown that incorporating memory mechanisms to support goal-directed behavior can improve long-horizon planning performance. However, they overlook visual frontier boundaries, which fundamentally dictate future trajectories and observations, and fall short of inferring the relationship between partial visual observations and navigation goals. In this paper, we propose Semantic Cognition Over Potential-based Exploration (SCOPE), a zero-shot framework that explicitly leverages frontier information to drive potential-based exploration, enabling more informed and goal-relevant decisions. SCOPE estimates exploration potential with a Vision-Language Model and organizes it into a spatio-temporal potential graph, capturing boundary dynamics to support long-horizon planning. In addition, SCOPE incorporates a self-reconsideration mechanism that revisits and refines prior decisions, enhancing reliability and reducing overconfident errors. Experimental results on two diverse embodied navigation tasks show that SCOPE outperforms state-of-the-art baselines by 4.6% in accuracy. Further analysis demonstrates that its core components lead to improved calibration, stronger generalization, and higher decision quality.
Ningnan Wang, Weihuang Chen, Haoxuan Ji, Zhongyu Guo, Xuchong Zhang, Hongbin Sun 0001
AAAI7
2026 A Low-Error Approximate Logarithmic Multiplier with Symmetric LUT for Efficient DNN Training
Baoting Li, Tai Yu, Xuchong Zhang, Hongbin Sun 0001
ISCAS6
2026 MFF-DCNet: A Network With Multifeature Focus and Depth-Wise Cross-Stage Transformer for UAV Infrared Small-Object Detection
abstract
The detection of infrared small objects from unmanned aerial vehicles (UAVs) is critical for a wide range of Internet of Things (IoT) applications, including reconnaissance, surveillance, and security monitoring. However, existing methods for small object detection are primarily designed for visible light images and exhibit poor performance when applied to infrared images due to their distinct characteristics such as lower resolution, lack of color and texture information, and higher noise levels. Most existing infrared small object detection algorithms are based on segmentation networks, which often struggle with false alarms when processing UAV-captured imagery with complex backgrounds. Moreover, these segmentation networks are computationally intensive, making them unsuitable for deployment on IoT edge devices. To address these challenges, we propose MFF-DCNet, an efficient network specifically designed for infrared small object detection in UAVs. The proposed network comprises a novel Depth-wise Cross-stage Transformer enhanced backbone and a Multi-Feature Focus neck structure, collectively strengthening multiscale feature extraction and representation. Evaluations on the HIT-UAV and DroneVehicle dataset demonstrate that the proposed network achieves state-of-the-art performance with anAP50−95of 57.4%, representing a 5.8% improvement over specific UAV imagery detectors while simultaneously achieving a 10% increase in FPS. Furthermore, our method achieves real-time performance of 39.6 FPS on the NVIDIA Jetson Orin NX, demonstrating its practical deployment capability in resource constrained IoT environments.
Xuchong Zhang, Hongbin Sun 0001
IEEE Internet Things J.4
2025 VLR-Driver: Large Vision-Language-Reasoning Models for Embodied Autonomous Driving
Fanjie Kong, Weihuang Chen, Zhongyu Guo, Hongbin Sun 0001
ICCV9
2025 Latent Feature and Attention Dual Erasure Attack against Multi-View Diffusion Models for 3D Assets Protection
abstract
Multi-View Diffusion Models (MVDMs) enable remarkable improvements in the field of 3D geometric reconstruction, but the issue regarding intellectual property has received increasing attention due to unauthorized imitation. Recently, some works have utilized adversarial attacks to protect copyright. However, all these works focus on single-image generation tasks which only need to consider the inner feature of images. Previous methods are inefficient in attacking MVDMs because they lack the consideration of disrupting the geometric and visual consistency among the generated multi-view images. This paper is the first to address the intellectual property infringement issue arising from MVDMs. Accordingly, we propose a novel latent feature and attention dual erasure attack to disrupt the distribution of latent feature and the consistency across the generated images from multi-view and multi-domain simultaneously. The experiments conducted on SOTA MVDMs indicate that our approach achieves superior performances in terms of attack effectiveness, transferability, and robustness against defense methods. Therefore, this paper provides an efficient solution to protect 3D assets from MVDMs-based 3D geometry reconstruction. The code is publicly available at: https://github.com/super-jw/LFADEA
Xuchong Zhang, Changfeng Sun, Qicheng Bai, Hongbin Sun 0001
ICME5
2025 Object-fabrication targeted attack for object detection
Xuchong Zhang, Changfeng Sun, Haoliang Han, Hongbin Sun 0001
Neurocomputing4
2025 TBag: Three Recipes for Building up a Lightweight Hybrid Network for Real-Time SISR
abstract
The prevalent convolution neural network (CNN) and Transformer have revolutionized the area of single-image super-resolution (SISR). Though these models have significantly improved performance, they often struggle with real-time applications or on resource-constrained platforms due to their complexity. In this paper, we propose TBag, a lightweight hybrid network that combines the strengths of CNN and Transformer to address these challenges. Our method simplifies the Transformer block with three key optimizations: 1) No projection layer is applied to the value in the original self-attention operation; 2) The number of tokens is rescaled before the self-attention operation and then rescaled back for easing of computation; 3) The expansion factor of the original feed-forward network (FFN) is adjusted. These optimizations enable the development of an efficient hybrid network tailored for real-time SISR. Notably, the hybrid design of CNN and Transformer further enhances both local detail recovery and global feature modeling. Extensive experiments show that TBag achieves a competitive trade-off between effectiveness and efficiency compared to previous lightweight SISR methods (e.g.,+0.42 dBPSNR with an86.7%reduction in latency). Moreover, TBag's real-time capabilities make it highly suitable for practical applications, with the TBag-Tiny version achieving up to59 FPSon hardware devices. Future work will explore the potential of this hybrid approach in other image restoration tasks, such as denoising and deblurring.
Ruoyi Xue, Hongbin Sun 0001
IEEE Trans. Multim.4
2024 An Efficient Sparse-Aware Summation Optimization Strategy for DNN Accelerator
abstract
Due to the various applications and high sparsity of deep neural network (DNN), a lot of sparse-aware DNN accelerators have been proposed to exploit the sparsity in DNN. Furthermore, it is essential to optimize for the accumulations and inter-channel aggregations in DNN to reduce memory overhead and improve performance of DNN accelerator. However, the uncertain number and location of non-zero element in DNN pose critical challenges for optimizing such accelerators and this inspires us to explore an efficient spare-aware summation optimization strategy for DNN accelerator. In this paper, we leverage the strategy that trading higher cost memory storage/access for lower cost computation to propose a random index based sparse-aware adder tree (RAT), which achieves a better trade-off among performance, hardware resource overhead and adaptability. Synthesis and simulation results demonstrate that, compared with reference design, the proposed design achieves 1.71× and 1.52× the normalized area efficiency and energy efficiency improvement on ResNet18, respectively.
Danqing Zhang, Baoting Li, Xuchong Zhang, Hongbin Sun 0001
ISCAS5
2024 Compensation Architecture to Alleviate Noise Effects in RRAM-based Computing-in-memory Chips with Residual Resource
abstract
Resistive random access memory (RRAM) is a promising technology for energy-efficient in-memory computing. However, due to technology limits, RRAM device faces a series of reliability issues. Deep neural network (DNN) computing based on RRAM suffers from accuracy degradation. On the one hand, offline DNN training solutions are difficult to fully consider and simulate all nonidealities. Worse still, new error or nonideality may come up with the usage of RRAM, which further deteriorates the effectiveness of offline training. On the other hand, online training poses great challenges on programming overhead and device lifetime. The iterative write-verify technique to program multi-bit RRAM cells prolongs write latency more than 10× longer than read latency. To overcome these issues, we propose a compensation architecture and a software and hardware co-training design to mitigate the realistic network accuracy loss in RRAM-based computing-in-memory chips. Firstly, we add trainable compensation channels in crossbars utilizing the residual resource after original weight mapping. Secondly, an offline training procedure with computing output from hardware is triggered to settle down appropriate weight value in compensation channels. Experimental results demonstrate that the proposed design can guarantee ≤ 0.8% loss of accuracy in DNN on MNIST and CIFAR10 dataset even when nonidealities reduce the original accuracy down to ≤73%.
Longjun Liu, Yuyi Liu, Bin Gao 0006, Hongbin Sun 0001
ISCAS5
2024 MDC-Net: Multi-domain constrained kernel estimation network for blind image super resolution
Zhenyu Ding, Yuhai Li, Hongbin Sun 0001
Comput. Vis. Image Underst.5
2024 Targeted context attack for object detection
Changfeng Sun, Xuchong Zhang, Haoliang Han, Hongbin Sun 0001
Neurocomputing4
2024 DQ-STP: An Efficient Sparse On-Device Training Processor Based on Low-Rank Decomposition and Quantization for DNN
abstract
Due to the bottleneck problems such as scenario-varying application, significant data communication overhead and privacy protection between off-line training and on-line inference, intelligent edge devices capable of adaptively fine-tuning the deep neural network (DNN) models for specific tasks have become the most urgent need. However, the computational cost is intolerable for ordinary on-device training (ODT), which inspires us to explore an efficient ODT processor, named DQ-STP. In this paper, we leverage a series of optimization techniques using software-hardware co-design. On the one hand, the proposed design incorporates SVD-based low-rank decomposition,$2^{n}$quantization and ACBN algorithm on the software side. This unifies the sparse computing mode of convolutional layers and enhancing weight sparsity. On the other hand, the proposed design effectively leverages data sparsity on the hardware side through four techniques: 1) The flag compressed sparse row is proposed to compress input feature maps and gradient maps. 2) A unified processing element (PE) array comprising shifters and adders is proposed to expedite forward and error propagation steps. 3) The PE arrays for error propagation and weight gradients generation are separated to enhance throughput. 4) A sparse alignment strategy is proposed to further enhance PE utilization. Through these software and hardware co-optimization, the proposed DQ-STP achieves an area efficiency and peak energy efficiency of 41.2 GOPS/mm2 and 90.63 TOPS/W. In comparison to state-of-the-art reference designs, the proposed DQ-STP demonstrates a$2.19\times $improvement in normalized area efficiency and a$1.85\times $enhancement in energy efficiency.
Baoting Li, Danqing Zhang, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2024 Toward Robust LiDAR-Camera Fusion in BEV Space via Mutual Deformable Attention and Temporal Aggregation
abstract
LiDAR and camera are two critical sensors that can provide complementary information for accurate 3D object detection. Most works are devoted to improving the detection performance of fusion models on the clean and well-collected datasets. However, the collected point clouds and images in real scenarios may be corrupted to various degrees due to potential sensor malfunctions, which greatly affects the robustness of the fusion model and poses a threat to safe deployment. In this paper, we first analyze the shortcomings of most fusion detectors, which rely mainly on the LiDAR branch, and the potential of the bird’s eye-view (BEV) paradigm in dealing with partial sensor failures. Based on that, we present a robust LiDAR-camera fusion pipeline in unified BEV space with two novel designs under four typical LiDAR-camera malfunction cases. Specifically, a mutual deformable attention is proposed to dynamically model the spatial feature relationship and reduce the interference caused by the corrupted modality, and a temporal aggregation module is devised to fully utilize the rich information in the temporal domain. Together with the decoupled feature extraction for each modality and holistic BEV space fusion, the proposed detector, termed RobBEV, can work stably regardless of single-modality data corruption. Extensive experiments on the large-scale nuScenes dataset under robust settings demonstrate the effectiveness of our approach.
Jian Wang 0113, Fan Li 0003, Yi An, Xuchong Zhang, Hongbin Sun 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Physical Strip Attack for Object Detection in Optical Remote Sensing
abstract
A growing trend in the field of adversarial attacks is evolving from the digital domain to the more challenging physical domain. The previous works mainly employ printable adversarial patches with special textures in real-world physical attacks. However, due to lighting conditions and atmospheric scattering, the texture-based patches are prone to distortion in the long-range situation than in the close-range case, resulting in poor physical attack performance in remote sensing scenarios. Therefore, this article proposes a new physical attack method using single-color strip-based patches to hide the objects from being detected correctly in optical aerial detection. Specifically, we design a differentiable representation and an optimization method to optimize the position, thickness, and color of the adversarial strips. Compared with the traditional complex texture-based patch, the proposed strip-based patch is more robust when mapping from the digital domain to the physical domain. Extensive experiments are conducted on multiple datasets and real-world scenarios to evaluate the attack performance of various attack methods. The results show that the proposed strip-based adversarial patch has better attack performance against white-box, black-box, and even defense detectors. Furthermore, we can improve the physical attack success rate (ASR) in remote sensing scenarios by about 70% compared with previous texture-based methods.
Changfeng Sun, Xuchong Zhang, Qicheng Bai, Hongbin Sun 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Adversarial Obstacle Generation Against LiDAR-Based 3D Object Detection
abstract
LiDAR sensors are widely used in many safety-critical applications such as autonomous driving and drone control, and the collected data called point clouds are subsequently processed by 3D object detectors for visual perception. Recent works have shown that attackers can inject virtual points into LiDAR sensors by strategically transmitting laser pulses to them; additionally, deep visual models have been found to be vulnerable to carefully crafted adversarial examples. Therefore, a LiDAR-based perception may be maliciously attacked with serious safety consequences. In this article, we present a highly-deceptive adversarial obstacle generation algorithm against deep 3D detection models, to mimic fake obstacles within the effective detection range of LiDAR using a limited number of points. To achieve this goal, we first perform a physical LiDAR simulation to construct sparse obstacle point clouds. Then, we devise a strong attack strategy to adversarially perturb prototype points along each direction of the ray. Our method achieves a high attack success rate while complying with physical laws at the hardware level. We perform comprehensive experiments on different types of 3D detectors and determine that the voxel-based detectors are more vulnerable to adversarial attacks than the point-based methods. For example, our approach achieves an 89% mean attack success rate against PV-RCNN by using only 20 points to spoof a fake car.
Jian Wang 0113, Fan Li 0003, Xuchong Zhang, Hongbin Sun 0001
IEEE Trans. Multim.4
2023 DBQ-SSD: Dynamic Ball Query for Efficient 3D Object Detection
Lin Song 0002, Weixin Mao, Xiaoping Li 0005, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
ICLR7
2023 CoCoSoDa: Effective Contrastive Learning for Code Search
abstract
Code search aims to retrieve semantically relevant code snippets for a given natural language query. Recently, many approaches employing contrastive learning have shown promising results on code representation learning and greatly improved the performance of code search. However, there is still a lot of room for improvement in using contrastive learning for code search. In this paper, we propose CoCoSoDa to effectively utilize contrastive learning for code search via two key factors in contrastive learning: data augmentation and negative samples. Specifically, soft data augmentation is to dynamically masking or replacing some tokens with their types for input sequences to generate positive samples. Momentum mechanism is used to generate large and consistent representations of negative samples in a mini-batch through maintaining a queue and a momentum encoder. In addition, multimodal contrastive learning is used to pull together representations of code-query pairs and push apart the unpaired code snippets and queries. We conduct extensive experiments to evaluate the effectiveness of our approach on a large-scale dataset with six programming languages. Experimental results show that: (1) CoCoSoDa outperforms 18 baselines and especially exceeds CodeBERT, GraphCodeBERT, and UniXcoder by 13.3%, 10.5%, and 5.9% on average MRR scores, respectively. (2) The ablation studies show the effectiveness of each component of our approach. (3) We adapt our techniques to several different pre-trained models such as RoBERTa, CodeBERT, and GraphCodeBERT and observe a significant boost in their performance in code search. (4) Our model performs robustly under different hyper-parameters. Furthermore, we perform qualitative and quantitative analyses to explore reasons behind the good performance of our model.
Ensheng Shi, Yanlin Wang 0001, Lun Du, Hongyu Zhang 0002, Shi Han, Dongmei Zhang 0001, Hongbin Sun 0001
ICSE8
2023 Boosting Lidar 3D Object Detection with Point Cloud Semantic Segmentation
abstract
The integration of semantic information can effectively enhance the performance of 3D object detection based on lidar point cloud. Most of previous researches utilize camera-lidar fusion to improve detection accuracy for distant or small objects. However, this approach is typically unsuitable for real-time applications due to the large amount of input data. Recently, a multi-task framework using only Iidar has emerged as an alternative that employs the same feature extraction backbone with different heads to simultaneously output detection and semantic segmentation results for lidar point clouds. Nonetheless, some previous works have failed to achieve an optimal balance between accuracy and speed. To address this issue, we propose a multi-task framework which leverages the Cartesian pillar and a multi-scale semantic segmentation head to overcome the shortcomings of existing works and improve the detection accuracy. We evaluate the proposed method using typical pillar-based and voxel-based detection models on the nuScenes dataset. The experimental results demonstrate that the proposed design achieves better performance especially on small objects, compared to single-task models. Moreover, the proposed network increases mAP and NDS by 3.1 % and 2.5 % respectively on the nuScenes test set, compared to the representative multi-task network.
Xuchong Zhang, Chong Min, Yijie Jia, Jingmin Zhang, Hongbin Sun 0001
IROS6
2023 Towards Efficient Fine-Tuning of Pre-trained Code Models: An Experimental Study and Beyond
abstract
Recently, fine-tuning pre-trained code models such as CodeBERT on downstream tasks has achieved great success in many software testing and analysis tasks. While effective and prevalent, fine-tuning the pre-trained parameters incurs a large computational cost. In this paper, we conduct an extensive experimental study to explore what happens to layer-wise pre-trained representations and their encoded code knowledge during fine-tuning. We then propose efficient alternatives to fine-tune the large pre-trained code model based on the above findings. Our experimental study shows that (1) lexical, syntactic and structural properties of source code are encoded in the lower, intermediate, and higher layers, respectively, while the semantic property spans across the entire model. (2) The process of fine-tuning preserves most of the code properties. Specifically, the basic code properties captured by lower and intermediate layers are still preserved during fine-tuning. Furthermore, we find that only the representations of the top two layers change most during fine-tuning for various downstream tasks. (3) Based on the above findings, we propose Telly to efficiently fine-tune pre-trained code models via layer freezing. The extensive experimental results on five various downstream tasks demonstrate that training parameters and the corresponding time cost are greatly reduced, while performances are similar or better.
Ensheng Shi, Yanlin Wang 0001, Hongyu Zhang 0002, Lun Du, Shi Han, Dongmei Zhang 0001, Hongbin Sun 0001
ISSTA7
2023 Meta-Adapter: An Online Few-shot Learner for Vision-Language Model
abstract
The contrastive vision-language pre-training, known as CLIP, demonstrates remarkable potential in perceiving open-world visual concepts, enabling effective zero-shot image recognition. Nevertheless, few-shot learning methods based on CLIP typically require offline fine-tuning of the parameters on few-shot samples, resulting in longer inference time and the risk of overfitting in certain domains. To tackle these challenges, we propose the Meta-Adapter, a lightweight residual-style adapter, to refine the CLIP features guided by the few-shot samples in an online manner. With a few training samples, our method can enable effective few-shot learning capabilities and generalize to unseen data or tasks without additional fine-tuning, achieving competitive performance and high efficiency. Without bells and whistles, our approach outperforms the state-of-the-art online few-shot learning method by an average of 3.6\% on eight image classification datasets with higher inference speed. Furthermore, our model is simple and flexible, serving as a plug-and-play module directly applicable to downstream tasks. Without further fine-tuning, Meta-Adapter obtains notable performance improvements in open-vocabulary object detection and segmentation tasks.
Lin Song 0002, Ruoyi Xue, Hongbin Sun 0001, Yixiao Ge, Ying Shan
NeurIPS5
2023 CPNet: Continuity Preservation Network for infrared video colorization
Hongbin Sun 0001
Comput. Vis. Image Underst.5
2023 Efficient multi-stage network with pixel-wise degradation prediction for real-time motion deblurring
Zeyu Hao, Xuchong Zhang, Yuhai Li, Hongbin Sun 0001
Comput. Vis. Image Underst.5
2023 CoCoAST: Representing Source Code via Hierarchical Splitting and Reconstruction of Abstract Syntax Trees
Ensheng Shi, Yanlin Wang 0001, Lun Du, Hongyu Zhang 0002, Shi Han, Dongmei Zhang 0001, Hongbin Sun 0001
Empir. Softw. Eng.7
2023 Multimodal Pedestrian Trajectory Prediction Using Probabilistic Proposal Network
abstract
Forecasting multiple pedestrian trajectories is a challenging task for real-world applications, as the motion patterns of pedestrian are essentially stochastic and uncertain. Previous works have demonstrated that predicting diverse goals in advance can effectively improve the performance of pedestrian trajectory prediction. However, these methods are either unable to perform probabilistic and high-efficiency trajectory prediction, or mainly rely on the predefined template trajectories which are not high-performance and insufficient to represent the possible pedestrian behaviors. In this paper, we propose a new Probabilistic Proposal Network (PPNet) to concentrate on the generation of goals and the utilization of goal guidance. PPNet firstly generates multiple weighted goals based on the diverse latent intentions automatically obtained by unsupervised learning, and then designs the goal-conditioned Transformer networks to predict probabilistic proposals as the final trajectories. Extensive experimental results on ETH/UCY datasets and Stanford Drone Dataset indicate that PPNet achieves both state-of-the-art performance and high efficiency on pedestrian trajectory prediction.
Weihuang Chen, Lingyang Xue, Jinghai Duan, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 A Low-Cost Reduced-Latency DRAM Architecture With Dynamic Reconfiguration of Row Decoder
abstract
DRAM latency has remained almost constant over decades and has become a performance bottleneck of computing systems. In this study, we propose a low-cost DRAM architecture enabling dynamic reconfiguring of row decoder to provide reduced latency with high flexibility and reliability. We apply minimum changes to row decoders and allow dynamic reconfiguration to switch array blocks between two modes: 1) normal mode, where the DRAM array behaves in the same manner as the conventional DRAM does and 2) low latency mode, where two DRAM cells in the neighbor array blocks are coupled to operate as a logical cell and reduce latency reliably according to the differential principle. On the basis of an industrial open bitline (BL) cell array, we only change the word-line decoding scheme but keep the cell array and sense amplifiers (SAs) untouched to avoid modifications to the DRAM process for cost and reliability considerations. Our circuit simulation shows that the low-latency mode can reduce row-to-column delay and row access strobe time by 25.7% and 23.2%, respectively. We evaluate the reduced-latency LPDDR4 DRAM on various workloads. Compared with the JEDEC standard DRAM, our proposal provides a maximum system performance improvement of 8.5%. We believe that our proposal is a reliable and cost-friendly solution to DRAM latency reduction.
Fujun Bai, Xuerong Jia, Cong Lai, Qiwei Ren, Hongbin Sun 0001
IEEE Trans. Very Large Scale Integr. Syst.9
2023 ACBN: Approximate Calculated Batch Normalization for Efficient DNN On-Device Training Processor
abstract
Batch normalization (BN) has been established as a very effective component in deep learning, largely helping accelerate the convergence of deep neural network (DNN) training. Nevertheless, its hardware architecture has not received much attention in the field of DNN on-device training processors. Several previous designs incur either high off-chip memory traffic or high circuit complexity, and hence have deficiencies in terms of hardware efficiency and performance. This article proposes approximately calculated BN (ACBN) to achieve a much better tradeoff between hardware efficiency and performance for DNN on-device training processors. The accuracy and convergence rate of the proposed ACBN have been extensively evaluated using four typical DNN models. Compared with the state-of-the-art reference design, the hardware simulation results show the proposed ACBN can at least reduce floating point operations by 22.2% and save external memory access by 33.3% on average. Moreover, the proposed ACBN introduces 63.6% data sparsity for the backward propagation of BN layers of VGG16 on average. To the best of our knowledge, we are the first to introduce data sparsity for the backward propagation of BN layers. The ACBN module is implemented on Zynq UltraScale+ ZCU102 system-on-chip (SoC) field-programmable gate array (FPGA), and the results show that the implementation of ACBN hardware module saves 33.9% look-up table (LUT), 49.4% flip-flop (FF), 75% digital signal processor (DSP), and reduces the power by 12.4% compared with the reference design while achieving better performance.
Baoting Li, Fujie Luo, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2022 RACE: Retrieval-augmented Commit Message Generation
abstract
Commit messages are important for software development and maintenance.Many neural network-based approaches have been proposed and shown promising results on automatic commit message generation.However, the generated commit messages could be repetitive or redundant.In this paper, we propose RACE, a new retrieval-augmented neural commit message generation method, which treats the retrieved similar commit as an exemplar and leverages it to generate an accurate commit message.As the retrieved commit message may not always accurately describe the content/intent of the current code diff, we also propose an exemplar guider, which learns the semantic similarity between the retrieved and current code diff and then guides the generation of commit message based on the similarity.We conduct extensive experiments on a large public dataset with five programming languages.Experimental results show that RACE can outperform all baselines.Furthermore, RACE can boost the performance of existing Seq2Seq models in commit message generation.
Ensheng Shi, Yanlin Wang 0001, Wei Tao 0003, Lun Du, Hongyu Zhang 0002, Shi Han, Dongmei Zhang 0001, Hongbin Sun 0001
EMNLP8
2022 On the Evaluation of Neural Code Summarization
abstract
Source code summaries are important for program comprehension and maintenance. However, there are plenty of programs with missing, outdated, or mismatched summaries. Recently, deep learning techniques have been exploited to automatically generate summaries for given code snippets. To achieve a profound understanding of how far we are from solving this problem and provide suggestions to future research, in this paper, we conduct a systematic and in-depth analysis of 5 state-of-the-art neural code summarization models on 6 widely used BLEU variants, 4 pre-processing operations and their combinations, and 3 widely used datasets. The evaluation results show that some important factors have a great influence on the model evaluation, especially on the performance of models and the ranking among the models. However, these factors might be easily overlooked. Specifically, (1) the BLEU metric widely used in existing work of evaluating code summarization models has many variants. Ignoring the differences among these variants could greatly affect the validity of the claimed results. Besides, we discover and resolve an important and previously unknown bug in BLEU calculation in a commonly-used software package. Furthermore, we conduct human evaluations and find that the metric BLEU-DC is most correlated to human perception; (2) code preprocessing choices can have a large (from -18% to +25%) impact on the summarization performance and should not be neglected. We also explore the aggregation of pre-processing combinations and boost the performance of models; (3) some important characteristics of datasets (corpus sizes, data splitting methods, and duplication ratios) have a significant impact on model evaluation. Based on the experimental results, we give actionable suggestions for evaluating code summarization and choosing the best method in different scenarios. We also build a shared code summarization toolbox to facilitate future research.
Ensheng Shi, Yanlin Wang 0001, Lun Du, Junjie Chen 0003, Shi Han, Hongyu Zhang 0002, Dongmei Zhang 0001, Hongbin Sun 0001
ICSE8
2022 DFSNet: Dividing-fuse deep neural networks with searching strategy for distributed DNN architecture
Wenxuan Hou 0001, Longjun Liu, Haonan Zhang 0002, Hongbin Sun 0001, Nanning Zheng 0001
Neurocomputing4
2022 Towards high-quality thermal infrared image colorization via attention-based hierarchical network
Xuchong Zhang, Hongbin Sun 0001
Neurocomputing4
2022 End-to-end learning of self-rectification and self-supervised disparity prediction for stereo vision
Xuchong Zhang, Han Zhai, Hongbin Sun 0001, Nanning Zheng 0001
Neurocomputing5
2022 CMD: controllable matrix decomposition with global optimization for deep neural network compression
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Hongbin Sun 0001, Nanning Zheng 0001
Mach. Learn.4
2022 Adaptive Disparity Candidates Prediction Network for Efficient Real-Time Stereo Matching
abstract
Efficient real-time disparity estimation is critical for the application of stereo vision systems in various areas. Recently, stereo network based on coarse-to-fine method has largely relieved the memory constraints and speed limitations of large-scale network models. Nevertheless, all of the previous coarse-to-fine designs employ constant offsets and three or more stages to progressively refine the coarse disparity map, still resulting in unsatisfactory computation accuracy and inference time when deployed on mobile devices. This paper claims that the coarse matching errors can be corrected efficiently with fewer stages as long as more accurate disparity candidates can be provided. Therefore, we propose a dynamic offset prediction module to meet different correction requirements of diverse objects and design an efficient two-stage framework. In addition, a disparity-independent convolution is proposed to regularize the compact cost volume efficiently and further improve the overall performance. The disparity quality and efficiency of various stereo networks are evaluated on multiple datasets and platforms. Evaluation results demonstrate that, the disparity error rate of the proposed network achieves 2.66% and 2.71% on KITTI 2012 and 2015 test sets respectively, where the computation speed is$2\times $faster than the state-of-the-art lightweight models on high-end and source-constrained GPUs.
He Dai, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 FCHP: Exploring the Discriminative Feature and Feature Correlation of Feature Maps for Hierarchical DNN Pruning and Compression
abstract
Pruning can remove the redundant parameters and structures of Deep Neural Networks (DNNs) to reduce inference time and memory overhead. As one of the important components of DNN, feature maps (FMs) have been widely used in network pruning. However, previous approaches do not fully investigate the discriminative features in FMs, and also do not explicitly utilize all the features associated with each layer in the pruning procedure. In this paper, we explore the discriminative feature of FMs and explicitly investigate the two-adjacent-layer features of each layer to propose a three-phase hierarchical pruning framework, dubbed as FCHP. Firstly, we decompose each FM into several components to extract the discriminative feature. After that, since pruning each layer is related to the FMs of adjacent layers, we explicitly calculate the feature correlation of discriminative features of two adjacent layers, and then use the feature correlation to cluster FMs into several hierarchies to guide subsequent pruning. Finally, we compute the content of discriminative features, and remove channels corresponding to FMs with fewer discriminative features in each hierarchy, respectively. In the experiment, we prune DNNs with the multiple types of architecture on different benchmarks, and the results have achieved the state-of-the-arts in terms of compressed parameters and FLOPs drop. For example, as for ResNet-56 on CIFAR-10, FCHP respectively obtains 50% of parameters and FLOPs reduction with negligible accuracy loss. Besides, as for ResNet-50 on ImageNet, FCHP reduces 40.5% of parameters and 44.1% of FLOPs with 0.43% of Top-1 accuracy drop.
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Liang Si, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 S2TNet: Spatio-Temporal Transformer Networks for Trajectory Prediction in Autonomous Driving
abstract
To safely and rationally participate in dense and heterogeneous traffic, autonomous vehicles require to sufficiently analyze the motion patterns of surrounding traffic-agents and accurately predict their future trajectories. This is challenging because the trajectories of traffic-agents are not only influenced by the traffic-agents themselves but also by spatial interaction with each other. Previous methods usually rely on the sequential step-by-step processing of Long Short-Term Memory networks (LSTMs) and merely extract the interactions between spatial neighbors for single type traffic-agents. We propose the Spatio-Temporal Transformer Networks (S2TNet), which models the spatio-temporal interactions by spatio-temporal Transformer and deals with the temporel sequences by temporal Transformer. We input additional category, shape and heading information into our networks to handle the heterogeneity of traffic-agents. The proposed methods outperforms state-of-the-art methods on ApolloScape Trajectory dataset by more than 7% on both the weighted sum of Average and Final Displacement Error.
Weihuang Chen, Hongbin Sun 0001
ACML3
2021 End-to-End Object Detection With Fully Convolutional Network
abstract
Mainstream object detectors based on the fully convolutional network has achieved impressive performance. While most of them still need a hand-designed non-maximum suppression (NMS) post-processing, which impedes fully end-to-end training. In this paper, we give the analysis of discarding NMS, where the results reveal that a proper label assignment plays a crucial role. To this end, for fully convolutional detectors, we introduce a Prediction-aware One-To-One (POTO) label assignment for classification to enable end-to-end detection, which obtains comparable performance with NMS. Besides, a simple 3D Max Filtering (3DMF) is proposed to utilize the multi-scale features and improve the discriminability of convolutions in the local region. With these techniques, our end-to-end framework achieves competitive performance against many state-of-the-art detectors with NMS on COCO and CrowdHuman datasets. The code is available at https://github.com/Megvii-BaseDetection/DeFCN.
Lin Song 0002, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
CVPR4
2021 CAST: Enhancing Code Summarization with Hierarchical Splitting and Reconstruction of Abstract Syntax Trees
abstract
Code summarization aims to generate concise natural language descriptions of source code, which can help improve program comprehension and maintenance.Recent studies show that syntactic and structural information extracted from abstract syntax trees (ASTs) is conducive to summary generation.However, existing approaches fail to fully capture the rich information in ASTs because of the large size/depth of ASTs.In this paper, we propose a novel model CAST that hierarchically splits and reconstructs ASTs.First, we hierarchically split a large AST into a set of subtrees and utilize a recursive neural network to encode the subtrees.Then, we aggregate the embeddings of subtrees by reconstructing the split ASTs to get the representation of the complete AST.Finally, AST representation, together with source code embedding obtained by a vanilla code token encoder, is used for code summarization.Extensive experiments, including the ablation study and the human evaluation, on benchmarks have demonstrated the power of CAST.To facilitate reproducibility, our code and data are available at https://github.com/ DeepSoftwareAnalytics/CAST.
Ensheng Shi, Yanlin Wang 0001, Lun Du, Hongyu Zhang 0002, Shi Han, Dongmei Zhang 0001, Hongbin Sun 0001
EMNLP (1)7
2021 Exploring Effective DNN Models for Forensic Age Estimation based on Panoramic Radiograph Images
abstract
Dental age estimation is widely used in forensic identification, but the accuracy of traditional methods cannot satisfy the demand for accuracy, especially for age estimation of adults. We introduce a deep learning-based methodology to estimate the age based on collected X-ray images of the teeth. We present a new dental dataset, which contains labeled orthopan-tomograms (OPGs) of 27,957 people, including 16,383 OPGs for females as well as 11,574 OPGs for males. All ages range from 0 to 93-year-old with a median of 27. The accuracy of the age labels is guaranteed by the ID card information. Aiming at the characteristics of the dental data itself, we explore various neural network elements that are effective for age estimation, including proper network depth, convolution kernel size, multi-branch structure, and the feature reusing of early layers. Based on the characteristic exploration, we further search models for dental age estimation by using the popular Neural Architecture Search (NAS) method. Experiment results show that our model achieves a mean absolute error (MAE) of 1.64 years, surpass all existing CNN models. Compared with Inception-v4 with an MAE of 1.70 and 20.46B FLOPs (inputs size 384×384), the FLOPs of our model can be reduced by 2.7 times (7.49B FLOPs). To our best knowledge, this is the first study for age estimation by exploring and searching the DNN model. Our results have surpassed legal medical expert-level performance (with an MAE of more than 2) for age estimation. Our methodology and results in this paper are very meaningful to forensic medicine for aging estimation with panoramic radiograph images.
Wenxuan Hou 0001, Longjun Liu, Jinxia Gao, Anguo Zhu, Keyang Pan, Hongbin Sun 0001, Nanning Zheng 0001
IJCNN6
2021 AKECP: Adaptive Knowledge Extraction from Feature Maps for Fast and Efficient Channel Pruning
abstract
Pruning can remove redundant parameters and structures of Deep Neural Networks (DNNs) to reduce inference time and memory overhead. As an important component of neural networks, the feature map (FM) has stated to be adopted for network pruning. However, the majority of FM-based pruning methods do not fully investigate effective knowledge in the FM for pruning. In addition, it is challenging to design a robust pruning criterion with a small number of images and achieve parallel pruning due to the variability of FMs. In this paper, we propose Adaptive Knowledge Extraction for Channel Pruning (AKECP), which can compress the network fast and efficiently. In AKECP, we first investigate the characteristics of FMs and extract effective knowledge with an adaptive scheme. Secondly, we formulate the effective knowledge of FMs to measure the importance of corresponding network channels. Thirdly, thanks to the effective knowledge extraction, AKECP can efficiently and simultaneously prune all the layers with extremely few or even one image. Experimental results show that our method can compress various networks on different datasets without introducing additional constraints, and it has advanced the state-of-the-arts. Notably, for ResNet-110 on CIFAR-10, AKECP achieves 59.9% of parameters and 59.8% of FLOPs reduction with negligible accuracy loss. For ResNet-50 on ImageNet, AKECP saves 40.5% of memory footprint and reduces 44.1% of FLOPs with only 0.32% of Top-1 accuracy drop.
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Wenxuan Hou 0001, Hongbin Sun 0001, Nanning Zheng 0001
ACM Multimedia5
2021 Dynamic Grained Encoder for Vision Transformers
abstract
Transformers, the de-facto standard for language modeling, have been recently applied for vision tasks. This paper introduces sparse queries for vision transformers to exploit the intrinsic spatial redundancy of natural images and save computational costs. Specifically, we propose a Dynamic Grained Encoder for vision transformers, which can adaptively assign a suitable number of queries to each spatial region. Thus it achieves a fine-grained representation in discriminative regions while keeping high efficiency. Besides, the dynamic grained encoder is compatible with most vision transformer frameworks. Without bells and whistles, our encoder allows the state-of-the-art vision transformers to reduce computational complexity by 40%-60% while maintaining comparable performance on image classification. Extensive experiments on object detection and segmentation further demonstrate the generalizability of our approach. Code is available at https://github.com/StevenGrove/vtpack.
Lin Song 0002, Songyang Zhang 0001, Xuming He 0001, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
NeurIPS6
2021 HSC: Leveraging horizontal shortcut connections for improving accuracy and computational efficiency of lightweight CNN
Anguo Zhu, Longjun Liu, Wenxuan Hou 0001, Hongbin Sun 0001, Nanning Zheng 0001
Neurocomputing4
2021 Efficient Repair Analysis Algorithm Exploration for Memory With Redundancy and In-Memory ECC
abstract
In-memory error correction code (ECC) is a promising technique to improve the yield and reliability of high density memory design. However, the use of in-memory ECC poses a new problem to memory repair analysis algorithm, which has not been explored before. This article first makes a quantitative evaluation and demonstrates that the straightforward algorithms for memory with redundancy and in-memory ECC have serious deficiency on either repair rate or repair analysis speed. Accordingly, an optimal repair analysis algorithm that leverages preprocessing/filter algorithms, hybrid search tree, and depth-first search strategy is proposed to achieve low computational complexity and optimal repair rate in the meantime. In addition, a heuristic repair analysis algorithm that uses a greedy strategy is proposed to efficiently find repair solutions. Experimental results demonstrate that the proposed optimal repair analysis algorithm can achieve optimal repair rate and increase the repair analysis speed by up to 105×105× compared with the straightforward exhaustive search algorithm. The proposed heuristic repair analysis algorithm is approximately 28 percent faster than the proposed optimal algorithm, at the expense of 5.8 percent repair rate loss.
Minjie Lv, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Computers2
2021 Dynamic Dataflow Scheduling and Computation Mapping Techniques for Efficient Depthwise Separable Convolution Acceleration
abstract
Depthwise separable convolution (DSC) has become one of the essential structures for lightweight convolutional neural networks. Nevertheless, its hardware architecture has not received much attention. Several previous hardware designs incur either high off-chip memory traffic or large on-chip memory usage, and hence have deficiency in terms of hardware efficiency as well as performance. This paper proposes two efficient dynamic design techniques, i.e. adaptive row-based dataflow scheduling and adaptive computation mapping, to achieve a much better trade-off between hardware efficiency and performance for DSC-based lightweight CNN accelerator. The effectiveness and efficiency of the proposed dynamic design techniques have been extensively evaluated using six DSC-based lightweight CNNs. Compared with the reference architectures, the simulation results show the proposed architectural techniques can at least reduce on-chip buffer size by 50.4% and improve the performance of convolution calculation by 1.18× while maintaining the minimum off-chip memory traffic. MobileNetV2 is implemented on Zynq UltraScale+ ZCU102 SoC FPGA, and the results show the proposed accelerator can achieve 381.7 frames per second (fps), which is 1.43× of the reference design, and it can save about 36.3% on-chip buffer size compared with the reference design, while maintaining the same off-chip memory traffic.
Baoting Li, Xuchong Zhang, Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2021 Exploring Highly Dependable and Efficient Datacenter Power System Using Hybrid and Hierarchical Energy Buffers
abstract
The massive and irregular load surges challenge datacenter power infrastructures. As a result, power mismatching between supply and demand has emerged as a crucial availability issue in modern datacenters which are either under-provisioned or powered by intermittent power sources. Recent proposals have employed energy storage devices such as the uninterruptible power supply (UPS) to address this issue. However, current approaches lack the capacity of efficiently handling the irregular and unpredictable power mismatches. In this paper, we propose Hybrid and Hierarchical Energy Buffering (HHEB), a novel heterogeneous and adaptive scheme that could enable various energy storage devices (ESDs) to be efficiently integrated into existing datacenters for dynamically dealing with power mismatches. Our techniques exploit the diverse characteristics of different ESDs and intelligent load assignment algorithms to improve the dependability and efficiency of datacenter power systems. We evaluate the HHEB design with a prototype. Compared with a homogenous battery energy buffering system, HHEB could improve energy efficiency by 39.7 percent, extend UPS lifetime by 4.7X, promote energy availability by 3.2X, reduce system downtime by 41 percent, and effectively improve the energy availability of various energy buffers in different hierarchies. It allows datacenters to adapt to various power supply anomalies, thereby improving operational efficiency, dependability and availability.
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Sustain. Comput.2
2020 Designing Efficient Shortcut Architecture for Improving the Accuracy of Fully Quantized Neural Networks Accelerator
abstract
Network quantization is an effective solution to compress Deep Neural Networks (DNN) that can be accelerated with custom circuit. However, existing quantization methods suffer from significant loss in accuracy. In this paper, we propose an efficient shortcut architecture to enhance the representational capability of DNN between different convolution layers. We further implement the shortcut hardware architecture to effectively improve the accuracy of fully quantized neural networks accelerator. The experimental results show that our shortcut architecture can obviously improve network accuracy while increasing very few hardware resources ( 0.11 × and 0.17 × for LUT and FF respectively) compared with the whole accelerator.
Baoting Li, Longjun Liu, Yanming Jin, Hongbin Sun 0001, Nanning Zheng 0001
ASP-DAC5
2020 Rethinking Learnable Tree Filter for Generic Feature Transform
abstract
The Learnable Tree Filter presents a remarkable approach to model structure-preserving relations for semantic segmentation. Nevertheless, the intrinsic geometric constraint forces it to focus on the regions with close spatial distance, hindering the effective long-range interactions. To relax the geometric constraint, we give the analysis by reformulating it as a Markov Random Field and introduce a learnable unary term. Besides, we propose a learnable spanning tree algorithm to replace the original non-differentiable one, which further improves the flexibility and robustness. With the above improvements, our method can better capture long range dependencies and preserve structural details with linear complexity, which is extended to several vision tasks for more generic feature transform. Extensive experiments on object detection/instance segmentation demonstrate the consistent improvements over the original version. For semantic segmentation, we achieve leading performance (82.1% mIoU) on the Cityscapes benchmark without bells-and whistles. Code is available at https://github.com/StevenGrove/LearnableTreeFilterV2.
Lin Song 0002, Zhengkai Jiang 0001, Xiangyu Zhang 0005, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
NeurIPS6
2020 Fine-Grained Dynamic Head for Object Detection
abstract
The Feature Pyramid Network (FPN) presents a remarkable approach to alleviate the scale variance in object representation by performing instance-level assignments. Nevertheless, this strategy ignores the distinct characteristics of different sub-regions in an instance. To this end, we propose a fine-grained dynamic head to conditionally select a pixel-level combination of FPN features from different scales for each instance, which further releases the ability of multi-scale feature representation. Moreover, we design a spatial gate with the new activation function to reduce computational complexity dramatically through spatially sparse convolutions. Extensive experiments demonstrate the effectiveness and efficiency of the proposed method on several state-of-the-art detection benchmarks. Code is available at https://github.com/StevenGrove/DynamicHead.
Lin Song 0002, Zhengkai Jiang 0001, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
NeurIPS5
2020 Spatiotemporal neural networks for action recognition based on joint loss
Chao Jing, Ping Wei 0001, Hongbin Sun 0001, Nanning Zheng 0001
Neural Comput. Appl.3
2020 Algorithm and VLSI Architecture Co-Design on Efficient Semi-Global Stereo Matching
abstract
Semi-global matching (SGM) is favored for high accuracy real-time stereo matching design as it achieves a good trade-off between disparity image quality and computational complexity. Nevertheless, most of previous SGM designs so far are restricted to the real-time processing of small image resolution and disparity range, or achieve high throughput by simplifying the original algorithm at the penalty of significant disparity image quality degradation. We analyze that the major challenge to efficient SGM design is its memory architecture, including both on-chip memory cost and off-chip memory bandwidth. We address the memory architecture challenge by algorithm and architecture co-design. Based on two observed features of SGM algorithm, i.e. incompleteness and inaccuracy, this paper proposes several efficient techniques to reduce on-chip memory cost and compress off-chip memory bandwidth respectively. Moreover, we also design high throughput and pipelined architecture to implement the proposed techniques. The disparity image quality and hardware efficiency of the proposed SGM design are evaluated on both KITTI2015 and Middlebury V3 stereo datasets. Evaluation results demonstrate that, the throughput of the proposed circuit designs can easily achieve 1080P@30fps at the disparity range of 128, and can reduce the on-chip memory cost and off-chip memory bandwidth by up to 4× and 2× respectively while achieving better or the same disparity image quality, compared with the best reference design techniques.
Xuchong Zhang, He Dai, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 TACNet: Transition-Aware Context Network for Spatio-Temporal Action Detection
abstract
Current state-of-the-art approaches for spatio-temporal action detection have achieved impressive results but remain unsatisfactory for temporal extent detection. The main reason comes from that, there are some ambiguous states similar to the real actions which may be treated as target actions even by a well trained network. In this paper, we define these ambiguous samples as “transitional states”, and propose a Transition-Aware Context Network (TACNet) to distinguish transitional states. The proposed TACNet includes two main components, i.e., temporal context detector and transition-aware classifier. The temporal context detector can extract long-term context information with constant time complexity by constructing a recurrent network. The transition-aware classifier can further distinguish transitional states by classifying action and transitional states simultaneously. Therefore, the proposed TACNet can substantially improve the performance of spatio-temporal action detection. We extensively evaluate the proposed TACNet on UCF101-24 and J-HMDB datasets. The experimental results demonstrate that TACNet obtains competitive performance on JHMDB and significantly outperforms the state-of-the-art methods on the untrimmed UCF101 24 in terms of both frame-mAP and video-mAP.
Lin Song 0002, Shiwei Zhang 0001, Gang Yu 0002, Hongbin Sun 0001
CVPR4
2019 Exploring Hardware Friendly Bottleneck Architecture in CNN for Embedded Computing Systems
abstract
In this paper, we explore how to design lightweight CNN architecture for embedded computing systems. We propose L-Mobilenet model for ZYNQ based hardware platform. L-Mobilenet can adapt well to hardware computing and accelerating, and its network structure is inspired by the state-of-the-art work of Inception-Resnet and Mobilenet-V2, which can effectively reduce parameters and delay while maintaining the accuracy of inference. We deploy our L-Mobilenet model to GPU and ZYNQ embedded platform for fully evaluating the performance of our design. By measuring with cifar10 and cifar100 datasets, L-Mobilenet model is able to gain 3× speed up and 3.7× fewer parameters than MobileNet-V2 while maintaining a similar accuracy. It also can obtain 2× speed up and 1.5× fewer parameters than Shufflenet-V2 while maintaining the same accuracy. Experiments show that our network model can obtain better performance because of the special considerations for hardware accelerating and software-hardware co-design strategies in our L-Mobilenet bottleneck architecture.
Xing Lei, Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001
ICIP4
2019 REcache: Efficient Sustainable Energy Management Circuits and Policies for Computing Systems
abstract
The rapidly growing computing systems, such as AI server cluster, IoT devices etc. are facing increasing energy expenditure pressure and the warning of carbon footprint. Designing eco-friendly computing systems which integrated renewable energy sources have attracted considerable attentions recently. Existing schemes either incur green energy efficiency degradation or sacrifice workload performance. This paper proposes REcache (Renewable Energy cache), a sustainable energy management scheme to efficiently utilize green energy for computing systems. Compared to previous proposals, we present a dedicated circuit and energy-aware management policies to coordinate energy harvesting, power management and workload scheduling. We evaluate our scheme through both prototyping and simulation. The experimental results show that the REcache could effectively improve the energy availability 10%, workload performance 5% for different workloads on average.
Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001, Tao Li 0006
ISCAS2
2019 A Hardware-Efficient Post-Processing Algorithm for Motion Compensated Frame Rate Up-Conversion
abstract
Post-processing is an important module in motion compensated frame rate up-conversion design, and has a direct impact on the image quality of interpolated frame. However, how to balance between image quality and computational efficiency is still very challenging for post-processing, especially in hardware design. This paper proposes a hardware-efficient post-processing algorithm which leverages the temporal and spatial constraints to locally refine interpolated pixels. Moreover, we employ the quantization and approximation techniques to further reduce the computational intensity of the proposed post-processing algorithm. The quality of interpolated frame has been greatly improved both in objective and subjective aspects. The proposed post-processing algorithm is extensively evaluated by a set of video test sequences. Evaluation results demonstrate that, compared with the reference designs, the proposed algorithm can improve PSNR by at least 1.69 dB with comparable computational complexity.
Yunqi Mi, Hongbin Sun 0001, Nanning Zheng 0001
ISCAS4
2019 Scene-Guided Region Proposal Re-ranking Method for On-road Vehicle Candidate Generation
abstract
Vehicle candidate generation is important for vehicle detection. Existing vehicle detection studies usually employ general-purpose region proposal methods to generate vehicle candidates, which do not consider the specificity of on-road vehicles in traffic scenes. In this paper, we propose a model to re-rank the candidates that are generated by general-purpose region proposal methods. Our model considers the specificity of on-road vehicle candidate generation in traffic scenes by encoding global-local semantic context and location-size geometric compatibility. In the experiments, we test our model on three art-of-the-state region proposal methods using two public datasets. The results show the significant performance improvement is gained after applying our model.
Zhixiong Nan, Jiawei He 0002, Ping Wei 0001, Linhai Xu, Hongbin Sun 0001, Nanning Zheng 0001
IV6
2019 Learnable Tree Filter for Structure-preserving Feature Transform
abstract
Learning discriminative global features plays a vital role in semantic segmentation. And most of the existing methods adopt stacks of local convolutions or non-local blocks to capture long-range context. However, due to the absence of spatial structure preservation, these operators ignore the object details when enlarging receptive fields. In this paper, we propose the learnable tree filter to form a generic tree filtering module that leverages the structural property of minimal spanning tree to model long-range dependencies while preserving the details. Furthermore, we propose a highly efficient linear-time algorithm to reduce resource consumption. Thus, the designed modules can be plugged into existing deep neural networks conveniently. To this end, tree filtering modules are embedded to formulate a unified framework for semantic segmentation. We conduct extensive ablation studies to elaborate on the effectiveness and efficiency of the proposed method. Specifically, it attains better performance with much less overhead compared with the classic PSP block and Non-local operation under the same backbone. Our approach is proved to achieve consistent improvements on several benchmarks without bells-and-whistles. Code and models are available at https://github.com/StevenGrove/TreeFilter-Torch.
Lin Song 0002, Gang Yu 0002, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
NeurIPS5
2019 NIPM-sWMF: Toward Efficient FPGA Design for High-Definition Large-Disparity Stereo Matching
abstract
Large disparity stereo matching is critical to the application of a stereo vision system especially for outdoor scenes. Nevertheless, how to efficiently design high accuracy large-disparity stereo matching on a field-programmable gate array (FPGA) is still a grand challenge. The computational complexity of previously proposed stereo matching is inevitably proportional to disparity range; hence their hardware designs become very inefficient when the disparity range is large. Motivated by the original PatchMatch and weighted median filtering (WMF) algorithms, this paper proposes a non-iterative PatchMatch and separable WMF (NIPM-sWMF) algorithm to significantly reduce the computational complexity of stereo matching and make it independent of disparity range. Moreover, we also propose a fully pipelined architecture design on FPGA that employs several hardware techniques to efficiently implement the proposed NIPM-sWMF. The disparity quality of the proposed NIPM-sWMF algorithm is evaluated on both KITTI2015 and Middlebury V3 stereo data sets, and the proposed architecture design is implemented and synthesized on Xilinx FPGA. Evaluation results demonstrate that the proposed NIPM-sWMF design on FPGA reaches the real-time performance of 1920 × 1080@60 Hz at the disparity range of 128, and can achieve almost the same disparity estimation accuracy, 4.5× processing throughput, while reducing the hardware cost of LUT, Register, DSP, and BRAM by 40%, 47%, 100%, and 68%, respectively, compared with the reference stereo matching design. Therefore, the proposed NIPM-sWMF design is an efficient way to address the challenge of large-disparity stereo matching.
Xuchong Zhang, Hongbin Sun 0001, Shiqiang Chen, Lin Song 0002, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2019 Learning Composite Latent Structures for 3D Human Action Representation and Recognition
abstract
3D human action representation and recognition are important issues in many multimedia applications. While latent state approaches have been widely used for action modeling, previous works assume the latent states of actions are single attribute. This assumption is inaccurate for representing structures of complex actions. In this paper, we propose that latent states have composite attributes and introduce a novel composite latent structure (CLS) model to represent and recognize 3D human actions with skeleton sequences. A human action is modeled with a hierarchical graph, which represents the action sequence as sequential atomic actions. An atomic action is represented as a composite latent state, which is composed of a latent semantic attribute and a latent geometric attribute. A discriminative EM-like algorithm is proposed to learn the model parameters and the composite latent structures of human actions. Given a 3D skeleton sequence, a composite attribute iterative programming algorithm is proposed to recognize the action and infer the action's latent temporal structure. We evaluate the proposed method on three challenging 3D action datasets-MSR 3D Action Dataset, Multiview 3D Event Dataset, and UTKinect-Action 3D Dataset. Extensive experimental results on these datasets demonstrate the effectiveness and advantage of the proposed method.
Ping Wei 0001, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Multim.2
2019 Efficient Compression-Based Line Buffer Design for Image/Video Processing Circuits
abstract
Line buffer is a typical and major on-chip memory design architecture for image/video processing circuits. As it usually occupies very large on-chip circuit area, it is of great importance to reduce its hardware cost through efficient architecture design. Data compression is a promising technique to improve the hardware efficiency of line buffer architecture. Nevertheless, the previously proposed data compression technique for line buffer architecture only exploits fixed length code (FLC), which actually has the deficiency on compression performance. Instead, this paper explores to efficiently use variable length code in line buffer architecture. By restricting variable length coding within small compression granularity (CG), the proposed compression algorithm not only significantly improves compression performance but also meets the specific requirements in line buffer architecture design. The simple compression algorithm further enables the efficient and fully pipelined VLSI architecture and circuits. Experimental results demonstrate that the proposed compression algorithm achieves 6.67-dB peak signal-to-noise ratio improvement at the compression ratio of 50% and the CG of 16 pixels, compared with FLC design. The VLSI circuits of the proposed compression can achieve the throughput of 4K × 2K at 60 fps with reasonable hardware cost. The use of the proposed compression technique in line buffer architecture can significantly reduce on-chip memory cost while maintaining satisfactory visual quality.
Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2018 Exploring the Potential of Using Semantic Context and Common Sense in On-Road Vehicle Detection
abstract
Vehicle detection is an important research topic for autonomous driving community. Since the great success of deep learning on object detection, almost all vehicle detection methods go along with this line. However, deep learning methods heavily rely on the training data, and the whole mechanism is like a “black box” Therefore, in this paper, we explore a vehicle detection method using traffic semantic context and human common sense instead of relying on the training data. To verify our idea, we compare our method with two classic machine learning methods as well as three state- of-the-art deep learning methods on a dataset collected in real traffics. The results show that our method outperforms others on this dataset. The deep learning methods may exceed ours after enlarging the training data or testing on more complicated datasets. However, the main contribution of this paper is providing inspiration for learning methods, and we believe their performance can be greatly improved after considering the idea of this paper.
Zhixiong Nan, Menghan Pan, Xiao Wang 0002, Ping Wei 0001, Linhai Xu, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium6
2018 Leveraging Spatio-Temporal Evidence and Independent Vision Channel to Improve Multi-Sensor Fusion for Vehicle Environmental Perception
abstract
For intelligent vehicles, multi-sensor fusion is of great importance to perceive traffic environment with high accuracy and robustness. In this paper, we propose two effective methods, i.e. spatio-temporal evidence generating and independent vision channel, to improve multi-sensor track-level fusion for vehicle environmental perception. The spatio-temporal evidence includes instantaneous evidence, tracking evidence and tracks matching evidence to improve existence fusion. Independent vision channel leverages the specific advantage of vision processing on object recognition to improve classification fusion. The proposed methods are evaluated by using the multi-sensor dataset collected from real traffic environment. Experimental results demonstrate that the proposed methods can significantly improve the multi-sensor track-level fusion in terms of both detection accuracy and classification accuracy.
Juwang Shi, Wenxiu Wang, Xiao Wang 0002, Hongbin Sun 0001, Xuguang Lan, Jingmin Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium4
2018 Efficient Rectangle Fitting of Sparse Laser Data for Robust On-Road Obiect Detection
abstract
On-road object detection is one of the most important tasks for the autonomous driving of intelligent vehicle. Nevertheless, the previous methods based on 2D LIDAR sensor only focus on the detection of vehicles, and show severe limitations on the detection of other objects. Accordingly, this paper proposes an on-road object detection method, which employs rectangle fitting and concavity determination to improve the robustness of ob- ject detection. The proposed approaches are extensively evaluated by using the sparse laser data collected by 2D LIDAR from real traffic environment. Experimental results demonstrate that the proposed rectangle fitting outperforms the previous approaches in terms of both detection accuracy and computational efficiency.
Zhaohong Xiang, Xiao Wang 0002, Hongbin Sun 0001, Jinming Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium5
2018 VLSI Architecture Exploration of Guided Image Filtering for 1080P@60Hz Video Processing
abstract
Guided image filtering (GIF) is a promising edge-preserving filtering technique that has been applied in a variety of applications. Nevertheless, an efficient very-large-scale integration (VLSI) architecture design of GIF is still very challenging for the real-time processing of full-high definition videos. Previously proposed architectures are somewhat inefficient in terms of either on-chip memory usage or off-chip memory bandwidth. This paper aims to improve the balance between on-chip memory usage and off-chip memory bandwidth through architecture exploration. Three critical architectural tradeoffs in the VLSI design of GIF are explored, and two efficient VLSI architectures, namely sequential line-based and parallel line-based architectures, are proposed. Experimental results demonstrate that the proposed VLSI design only consumes 34.1-K logic gates, 25.4-KB on-chip memories, and 373-MB/s off-chip memory bandwidth while achieving a real-time video processing of 1080P@60Hz at the maximum clock frequency of 297-MHz. Moreover, the proposed VLSI circuits are fully pipelined and synchronized to the pixel clock of output video, so can be seamlessly integrated into diverse real-time video processing systems.
Xuchong Zhang, Hongbin Sun 0001, Shiqiang Chen, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Worst Case Driven Display Frame Compression for Energy-Efficient Ultra-HD Display Processing
abstract
Display frame compression is an effective technique to address the challenge of external memory access in ultrahigh definition video display system. Nevertheless, previously proposed display frame compression designs are inadequate in terms of either energy efficiency or throughput. This paper aims to exploit the algorithm and very large scale integration (VLSI) architecture of a worst case driven display frame compression. By using a prediction-and-compression framework and a semi-fixed length coding scheme, the proposed design can achieve the much better balance between compression efficiency and throughput, and substantially reduce the bandwidth requirement and energy consumption of external memory system in the meanwhile. Extensive experiments demonstrate that the proposed display frame compression achieves 5.7-dB peak signal-to-noise ratio improvement, 3.1% compression ratio reduction, 3 × throughput, and 66.4% hardware cost saving, compared with the best previous work. In addition, the proposed VLSI design can support the throughput of 4 K × 2 K@60 Hz and reduce at least 17.6% energy consumption of external memory system by exploiting dynamic voltage and frequency scaling, compared with conventional display frame compression works.
Qiubo Chen, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Multim.2
2018 Exploring Customizable Heterogeneous Power Distribution and Management for Datacenter
abstract
Large-scale datacenters are facing increasing pressure of capping their carbon emission and power cost. Many leading-edge studies have started to explore server clusters running on multiple power sources. Existing approaches do not sufficiently consider the fine-grained power delivery to satisfy diverse requirements in datacenter, especially in the multi-tenant/colocation datacenter, which may yield low energy utilization. To address the emerging trend and new requirements, this article proposes a novel Datacenter inner Power Switch Network (DiPSN) to improve datacenter power efficiency and user satisfaction. DiPSN is a reconfigurable and easy-to-scale-out power architecture, which enables datacenter to distribute various power sources in a fine-grained manner. Moreover, a tailored machine learning based power source management framework is proposed for DiPSN to dynamically optimize user customized performance metrics and maximize datacenter revenue. Compared with conventional single-switch power distribution system, our DiPSN can be configured to improve solar energy utilization by 39.6 percent, reduce utility power cost by 11.1 percent and improve workload performance by 33.8 percent. Meanwhile, our design can extend battery lifetime by 9.3 percent. This work could provide valuable guidelines for designing heterogeneous power distribution architecture and management methodology in datacenters for improving user-customizable efficiency, sustainability and economy.
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Tao Li 0006, Nanning Zheng 0001
IEEE Trans. Parallel Distributed Syst.2
2017 sWMF: Separable weighted median filter for efficient large-disparity stereo matching
abstract
Although large disparity stereo matching is critical to the practical application of stereo vision system especially for outdoor scenes, its efficient hardware design is still a grand challenge. Motivated by the discovery that well-designed weighted median filter (WMF) can achieve satisfactory accuracy with simple box-filter aggregation, this paper proposes a separable weighted median filter (sWMF) that only has the computational complexity of O(r) and is independent of disparity range. Moreover, the proposed sWMF can be efficiently implemented as a fully pipelined architecture. Evaluation results demonstrate that, at the penalty of only 0.06% disparity error rate, the proposed sWMF design can save 12.9% Slice LUTs, 76.7% DSPs and 64.0% Block RAMs at the disparity range of 128, compared with previous WMF implementation on FPGA.
Shiqiang Chen, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
ISCAS3
2017 Improving 3D DRAM Fault Tolerance Through Weak Cell Aware Error Correction
abstract
Although the emerging 3D DRAM products can significantly improve the computing system performance, the relatively high cost is one of the most critical issues that prevent their wide real-life adoption. Intuitively, a strong memory fault tolerance can be leveraged to reduce the fabrication cost of DRAM dies, and the total cost will reduce if the fabrication cost saving can off-set the cost overhead of memory fault tolerance. Nevertheless, such a simple concept can be a practically viable option only for 3D DRAM because: (1) The stacked logic die can solely implement memory fault tolerance inside 3D DRAM chips, obviating any changes on the host CPUs and CPU-DRAM interfaces. (2) With the total ownership of both the logic die and DRAM dies inside 3D DRAM chips, DRAM manufacturers can fully exploit the potential to truly minimize the 3D DRAM bit cost. Following this intuition, we developed a 3D DRAM fault tolerance design strategy. It can achieve a very strong tolerance to weak DRAM cells at very small redundancy and latency overhead. The key is to cohesively leverage the detectability of weak cells and runtime configurability of error correction code (ECC) decoding. In addition, this design strategy can gracefully embrace the inaccuracy of weak cell detection (e.g., weak cell miss-detection and false-detection). We carried out thorough mathematical analysis, and the results show that, under the redundancy overhead of 1:8 (same as today's ECC DIMM), this design strategy can tolerate the weak cell rate of as high as 10-4 and 6x10-5 if 100 and 90 percent of all the weak cells are known in prior. Using Micron's hybrid memory cube (HMC) 3D DRAM chips as the test vehicle, we evaluated the implementation cost and the results show that it only consumes less than 0.4 mm2 (45 nm node) on the logic die. Using CPU and DRAM simulators, we further carried out simulations over a variety of computing benchmarks and the results show that this design solution only incurs less than 2 percent performance degradation on average.
Hao Wang 0042, Kai Zhao 0005, Minjie Lv, Hongbin Sun 0001, Tong Zhang 0002
IEEE Trans. Computers5
2017 Managing Battery Aging for High Energy Availability in Green Datacenters
abstract
Energy storage devices (ESD), such as UPS batteries, have been repurposed in datacenter as a promising tuning knob for peak power shaving and power cost reducing. However, batteries progressively aging due to irregular usage patterns, which result in less effective capacity and even pose serious threat to server availability. Nevertheless, prior proposals largely ignore the aging issues of battery which may lead to low energy availability for datacenter servers. To fill this critical void, we thoroughly investigate battery aging on a heavily instrumented prototype system over an observation period of ten months. We propose Battery Anti-Aging Treatment Plus (BAAT-P), a novel power delivery architecture included aging management algorithms from the perspective of computing system to hide, reduce, mitigate and plan the battery aging effects for high energy availability in datacenter. Our techniques exploit diverse battery aging mechanisms and dynamic aging management algorithms to provide system-level availability guarantee for datacenter. We evaluate the BAAT-P design with a real prototype. Compared with a battery powered datacenter without aging management policies, the results show that BAAT-P can extend battery lifetime by 72 percent, reduce battery cost by 33 percent and effectively improve energy availability for datacenter servers while maintaining workload performance for the performance critical workloads.
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Parallel Distributed Syst.2
2016 Towards an Adaptive Multi-Power-Source Datacenter
abstract
Big data and cloud computing are accelerating the capacity growth of datacenters all over the world. Their energy costs and environmental issues have pushed datacenter operators to explore and integrate alternative energy sources, such as various renewable energy supplies and energy storage devices. Designing datacenters powered by multi-power supplies in the smart grid environment is becoming a promising trend in the next few decades. However, gracefully provisioning various power sources and efficiently manage them in datacenter is a significant challenge.
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Nanning Zheng 0001, Tao Li 0006
ICS2
2016 Algorithm and VLSI Architecture of Edge-Directed Image Upscaling for 4k Display System
abstract
High-quality and cost-efficient image upscaling design is very important for many real-time video processing applications, especially when the display panel resolution reaches ultrahigh definition. Compared with New Edge-Directed Interpolation (NEDI) based implicit edge directional upscaling, explicit methods require less computational resource and more easily reach real-time performance, especially when the required image definition and upscaling ratio are very high. Nevertheless, the investigation of applications of explicit methods in video processing systems remains largely missing arguably because it is commonly believed that explicit edge-directed interpolation tends to introduce unexpected artifacts because of inaccurate detection and hence its image quality is relatively poor. This paper proposes an explicit edge-directed adaptive interpolation method that leverages more sophisticated edge detection and orientation estimation algorithms to avoid misinterpolation, thereby providing similar or even better image quality than those with implicit methods. Targeting the real-time 4K video display system, the proposed edge-directed image upscaling algorithm is further implemented with an efficient very-large-scale integration (VLSI) architecture. The experimental results demonstrate that the proposed interpolation algorithm outperforms previous explicit and implicit edge-directed methods in both objective and subjective tests. The presented VLSI implementation further demonstrates that the maximum output video sequence of the proposed interpolation method can reach 4k × 2k@60 Hz with a reasonable hardware cost.
Qiubo Chen, Hongbin Sun 0001, Xuchong Zhang, Huibin Tao, Jie Yang 0001, Jizhong Zhao, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2016 On-Road Vehicle Detection and Tracking Using MMW Radar and Monovision Fusion
abstract
With the potential to increase road safety and provide economic benefits, intelligent vehicles have elicited a significant amount of interest from both academics and industry. A robust and reliable vehicle detection and tracking system is one of the key modules for intelligent vehicles to perceive the surrounding environment. The millimeter-wave radar and the monocular camera are two vehicular sensors commonly used for vehicle detection and tracking. Despite their advantages, the drawbacks of these two sensors make them insufficient when used separately. Thus, the fusion of these two sensors is considered as an efficient way to address the challenge. This paper presents a collaborative fusion approach to achieve the optimal balance between vehicle detection accuracy and computational efficiency. The proposed vehicle detection and tracking design is extensively evaluated with a real-world data set collected by the developed intelligent vehicle. Experimental results show that the proposed system can detect on-road vehicles with 92.36% detection rate and 0% false alarm rate, and it only takes ten frames (0.16 s) for the detection and tracking of each vehicle. This system is installed on Kuafu-II intelligent vehicle for the fourth and fifth autonomous vehicle competitions, which is called “Intelligent Vehicle Future Challenge” in China.
Xiao Wang 0002, Linhai Xu, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.3
2016 Integrated Longitudinal and Lateral Control for Kuafu-II Autonomous Vehicle
abstract
Over the past decades, there has been significant research effort dedicated to the development of autonomous vehicles and advanced driver assistance systems. The driving control system, which is responsible for trajectory tracking and driving safety, is one of the most important technologies for autonomous vehicles. This paper describes the design of driving control system, including both longitudinal and lateral controllers, for the Kuafu-II autonomous vehicle. Compared with most of the previous researches that inevitably require a large amount of parameters, the presented control system design in this paper integrates several typical and efficient controllers to significantly reduce the system sensitivity to these parameters, and it is able to achieve the system robustness under diversified circumstances. The effectiveness of the presented control system design has been extensively evaluated under simulation and on road tests.
Linhai Xu, Yingzhou Wang, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.3
2016 RE-UPS: an adaptive distributed energy storage system for dynamically managing solar energy in green datacenters
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Jingmin Xin, Nanning Zheng 0001, Tao Li 0006
J. Supercomput.2
2016 Exploiting Intracell Bit-Error Characteristics to Improve Min-Sum LDPC Decoding for MLC NAND Flash-Based Storage in Mobile Device
abstract
A multilevel per cell (MLC) technique significantly improves the storage density, but also poses serious data integrity challenge for NAND flash memory. This consequently makes the low-density parity-check (LDPC) code and the soft-decision memory sensing become indispensable in the next-generation flash-based solid-state storage devices. However, the use of LDPC codes inevitably increases memory read latency and, hence, degrades speed performance. Motivated by the observation of intracell unbalanced bit error probability and data dependence in the MLC NAND flash memory, this paper proposes two techniques, i.e., intracell data placement interleaving and intracell data dependence aware LDPC decoding, to efficiently improve the LDPC decoding throughput and energy efficiency for the MLC NAND flash-based storage in a mobile device. Experimental results show that, by exploiting the intracell bit-error characteristics, the proposed techniques together can improve the LDPC decoding throughput by up to 84.6% and reduce the energy consumption by up to 33.2% while only incurring less than 0.2% silicon area overhead.
Hongbin Sun 0001, Minjie Lv, Guiqiang Dong, Nanning Zheng 0001, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Logic-DRAM co-design to efficiently repair stacked DRAM with unused spares
abstract
Three dimensional (3D) integration is promising to provide dramatic performance and energy efficiency improvement to 3D logic-DRAM integrated computing system, but also poses significant challenge to the yield and reliability. By leveraging logic-DRAM co-design, this paper exploits the cost efficient approach to repair 3D integration induced defective cells in stacked DRAM with unused spares. In particular, we propose to make the DRAM array open its redundancy to off-chip access by small architecture modification, and further design the defective address comparison and redundant address remapping with very efficient architecture on logic die to achieve the equivalent memory repair. Simulation results have demonstrated that the proposed repair technique for DRAM after die stacking is able to significantly alleviate the yield loss, with very low area and power consumption overhead and negligible timing penalty.
Minjie Lv, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
ASP-DAC2
2015 BAAT: Towards Dynamically Managing Battery Aging in Green Datacenters
abstract
Energy storage devices (batteries) have shown great promise in eliminating supply/demand power mismatch and reducing energy/power cost in green datacenters. These important components progressively age due to irregular usage patterns, which result in less effective capacity and even pose serious threat to server availability. Nevertheless, prior proposals largely ignore the aging issue of batteries or simply use ad-hoc discharge capping to extend their lifetime. To fill this critical void, we thoroughly investigate battery aging on a heavily instrumented prototype over an observation period of six months. We propose battery anti-aging treatment (BAAT), a novel framework for hiding, reducing, and planning the battery aging effects. We show that BAAT can extend battery lifetime by 69%. It enables datacenters to maximally utilize energy storage resources to enhance availability and boost performance. Moreover, it reduces 26% battery cost and allows datacenters to economically scale in the big data era.
Longjun Liu, Chao Li 0009, Hongbin Sun 0001, Yang Hu 0001, Juncheng Gu, Tao Li 0006
DSN3
2015 HEB: deploying and managing hybrid energy buffers for improving datacenter efficiency and economy
abstract
Today, an increasing number of applications and services are being hosted by large-scale data centers. The massive and irregular load surges challenge data center power infrastructures. As a result, power mismatching between supply and demand has emerged as a crucial issue in modern data centers which are either under-provisioned or powered by intermittent power sources. Recent proposals have employed energy storage devices such as the uninterruptible power supply (UPS) systems to address this issue. However, current approaches lack the capacity of efficiently handling the irregular and unpredictable power mismatches.
Longjun Liu, Chao Li 0009, Hongbin Sun 0001, Yang Hu 0001, Juncheng Gu, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001
ISCA3
2015 Exploiting bit-depth scaling for quality-scalable energy efficient display processing
abstract
The energy efficiency of video display processing is critical for its integration into mobile SoC. This paper proposes a bit-depth scaling scheme to enable dynamic voltage and frequency scaling for external memory system and hence provides a quality-scalable energy efficient display processing solution. Experimental results have demonstrated that the proposed scheme is able to steadily reduce the energy consumption of overall external memory system by up to 23.8%, with graceful image quality degradation and very low cost overhead.
Qiubo Chen, Hengyu Zhao, Hongbin Sun 0001, Nanning Zheng 0001
ISCAS3
2014 Improving min-sum LDPC decoding throughput by exploiting intra-cell bit error characteristic in MLC NAND flash memory
abstract
Multi-level per cell (MLC) technique significantly improves storage density, but also poses new challenge to data integrity in NAND flash memory. Therefore, low-density parity-check (LDPC) code and soft-decision memory sensing have become indispensable in future NAND flash-based solid state drive design. However, these more powerful technologies inevitably increase the memory read latency and hence degrade the decoding throughput. Motivated by intra-cell unbalanced bit error probability and data dependency in MLC NAND flash memory, this paper proposes two techniques, i.e. intra-cell data placement interleaving and intra-cell data dependency aware min-sum decoding, to effectively improve the throughput of LDPC decoding. Experimental results show that, the proposed techniques used in an integrated way can improve the LDPC decoding throughput by up to 85% when the MLC NAND flash chip is heavily cycled, compared with conventional design practice.
Hongbin Sun 0001, Minjie Lv, Guiqiang Dong, Nanning Zheng 0001, Tong Zhang 0002
MSST2
2013 LDPC-in-SSD: making advanced error correction codes work effectively in solid state drives
Kai Zhao 0005, Hongbin Sun 0001, Tong Zhang 0002, Xiaodong Zhang 0001, Nanning Zheng 0001
FAST3
2013 Scheduling Algorithms for Handling Updates in Shingled Magnetic Recording
abstract
Shingled recording has recently emerged as one promising candidate to sustain the historical growth of magnetic recording storage areal density. However, since the convenient update-in-place feature is no longer available in shingled recording, many sectors must be read and written back in order to update one sector. This leads to a significant update-induced latency overhead and makes conventional hard disk drive scheduling algorithms perform poorly. This paper concerns with the development of appropriate scheduling algorithms for shingled recording based hard disk drives. We first present a simple partial-update scheduling algorithm that can naturally embrace the update latency issue and achieves significant gains over conventional scheduling algorithms. We enhance this algorithm by incorporating a shortest update first policy, which can further reduce the update response time on an average by 70%. Finally, motivated by abundant workload spatial and temporal locality, we develop a spatio-temporal band coalescing scheme that can achieve an additional reduction of update response time of up to 96.8%.
Kalyana Sundaram Venkataraman, Tong Zhang 0002, Hongbin Sun 0001, Nanning Zheng 0001
NAS4
2013 Exploring the Use of Emerging Nonvolatile Memory Technologies in Future FPGAs
abstract
As new nonvolatile memory technologies become increasingly mature, there has been a growing interest on investigating their use in future field-programmable gate arrays (FPGAs). Similar to existing FPGAs with embedded Flash memory, future FPGAs can embed these new nonvolatile memories to persistently store configuration data. By comparing with prior work, we first propose the more appropriate design style for new nonvolatile configuration data storage memory. Moreover, this brief studies a dynamic random-access memory (DRAM)-based FPGA design strategy enabled by high-density embedded nonvolatile memory. Existing FPGAs do not use on-chip DRAM cells for configuration data storage mainly because DRAM self-refresh involves destructive DRAM read. This problem can be solved, if we use embedded nonvolatile memory as primary FPGA configuration data storage and externally refresh on-chip DRAM cells. Analysis and simulations have been carried out to demonstrate the potential advantages of such a design strategy.
Yangyang Pan, Yiran Li 0001, Hongbin Sun 0001, Wei Xu 0021, Nanning Zheng 0001, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2012 Using Magnetic RAM to Build Low-Power and Soft Error-Resilient L1 Cache
abstract
Due to its great scalability, fast read access, low leakage power, and nonvolatility, magnetic random access memory (MRAM) appears to be a promising memory technology for on-chip cache memory in microprocessors. However, the write-to-MRAM process is relatively slow and results in high dynamic power consumption. Such inherent disadvantages of MRAM make researchers easily conclude that MRAM can only be used for low-level caches (e.g., L2 or L3 cache), where cache memories are less frequently accessed and slow write to MRAM can be more easily compensated using simple architectural techniques. By developing a hybrid cache architecture, this paper attempts to show that, with appropriate architecture design, MRAM can also be used in L1 cache to improve both the energy efficiency and soft error immunity. The basic idea is to supplement the MRAM L1 cache with several small SRAM buffers, which can substantially mitigate the performance degradation and dynamic energy overhead induced by MRAM write operations. Moreover, the proposed hybrid cache architecture is also an efficient solution to protect cache memory from radiation-induced soft errors, as MRAM is inherently invulnerable to emissive particles. Simulation results show that, with only less than 2% performance degradation, the proposed design approach can reduce the power consumption by up to 76.1% on average compared with the traditional SRAM L1 cache. In addition, the architectural vulnerability factor of L1 data cache is reduced from 28.3% to as low as 0.5%.
Hongbin Sun 0001, Chuanyin Liu, Wei Xu 0021, Jizhong Zhao, Nanning Zheng 0001, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.1
2011 Design techniques to improve the device write margin for MRAM-based cache memory
abstract
As one promising non-volatile memory technology, magnetoresistive RAM (MRAM) based on magnetic tunneling junctions (MTJs) has recently attracted much attention. However, latest device research has discovered that, in order to maintain sufficient MTJ write margin to prevent device breakdown, MTJs will be subject to unconventionally high random write error rates (e.g., 10-3 and above) as memory cell size is being scaled down. This new discovery seriously threatens the scalability of MRAM, and the material/device research community is actively searching for solutions to largely reduce MTJ write error rates and meanwhile maintain sufficient device write margin. In this paper, we attempt to address this challenge from the architecture level when using MRAM to implement cache memory. In particular, we show that two simple cache architecture design techniques can be used to effectively tolerate high MTJ write error rates at small performance and implementation cost, which makes it much easier to maintain sufficient MTJ write margin and hence push the MRAM scalability envelope. Using the full system simulator PTLsim and a variety of benchmarks, we show that the proposed design techniques can readily accommodate MTJ write error rate up to 0.75% at the penalty of less than 4% processor performance degradation, less than 10% silicon area overhead, and 6% energy consumption overhead.
Hongbin Sun 0001, Chuanyin Liu, Nanning Zheng 0001, Tai Min, Tong Zhang 0002
ACM Great Lakes Symposium on VLSI1
2011 Design of Last-Level On-Chip Cache Using Spin-Torque Transfer RAM (STT RAM)
abstract
Because of its high storage density with superior scalability, low integration cost and reasonably high access speed, spin-torque transfer random access memory (STT RAM) appears to have a promising potential to replace SRAM as last-level on-chip cache (e.g., L2 or L3 cache) for microprocessors. Due to unique operational characteristics of its storage device magnetic tunneling junction (MTJ), STT RAM is inherently subject to a write latency versus read latency tradeoff that is determined by the memory cell size. This paper first quantitatively studies how different memory cell sizing may impact the overall computing system performance, and shows that different computing workloads may have conflicting expectations on memory cell sizing. Leveraging MTJ device switching characteristics, we further propose an STT RAM architecture design method that can make STT RAM cache with relatively small memory cell size perform well over a wide spectrum of computing benchmarks. This has been well demonstrated using CACTI-based memory modeling and computing system performance simulations using SimpleScalar. Moreover, we show that this design method can also reduce STT RAM cache energy consumption by up to 30% over a variety of benchmarks.
Wei Xu 0021, Hongbin Sun 0001, Xiaobin Wang, Yiran Chen 0001, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.2
2009 Leveraging Access Locality for the Efficient Use of Multibit Error-Correcting Codes in L2 Cache
abstract
It is almost evident that SRAM-based cache memories will be subject to a significant degree of parametric random defects if one wants to leverage the technology scaling to its full extent. Although strong multibit error-correcting codes (ECC) appear to be a natural choice to handle a large number of random defects, investigation of their applications in cache remains largely missing arguably because it is commonly believed that multibit ECC may incur prohibitive performance degradation and silicon/energy cost. By developing a cost-effective L2 cache architecture using multibit ECC, this paper attempts to show that, with appropriate cache architecture design, this common belief may not necessarily hold true for L2 cache. The basic idea is to supplement a conventional L2 cache core with several special-purpose small caches/buffers, which can greatly reduce the silicon cost and minimize the probability of explicitly executing multibit ECC decoding on the cache read critical path, and meanwhile, maintain soft error tolerance. Experiments show that, at the random defect density of 0.5 percent, this design approach can maintain almost the same instruction per cycle (IPC) performance over a wide spectrum of benchmarks compared with ideal defect-free L2 cache, while only incurring less than 3 percent of silicon area overhead and 36 percent power consumption overhead.
Hongbin Sun 0001, Nanning Zheng 0001, Tong Zhang 0002
IEEE Trans. Computers1