Xuchong Zhang

dblp:128/0520 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0003-2772-2700ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
abstract
Multimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs. The limitation arises from MLLMs' context limit and training costs, necessitating sparse frame sampling before feeding videos into MLLMs. However, building a trainable sampling method remains challenging due to the unsupervised and non-differentiable nature of sparse frame sampling in Video-MLLMs. To address these problems, we propose Temporal Sampling Policy Optimization (**TSPO**), advancing MLLMs' long-form video-language understanding via reinforcement learning. Specifically, we first propose a trainable event-aware temporal agent, which captures event-query correlation for performing probabilistic keyframe selection. Then, we propose the TSPO reinforcement learning paradigm, which models keyframe selection and language generation as a joint decision-making process, enabling end-to-end group relative optimization for the temporal sampling policy. Furthermore, we propose a dual-style long video training data construction pipeline, balancing comprehensive temporal understanding and key segment localization. Finally, we incorporate rule-based answering accuracy and temporal locating reward mechanisms to optimize the temporal sampling policy. Comprehensive experiments show that our TSPO achieves state-of-the-art performance across multiple long video understanding benchmarks, and shows transferable ability across different cutting-edge Video-MLLMs.
Canhui Tang, Zifan Han, Sanping Zhou, Xuchong Zhang, Jinglin Xu
AAAI5
2026 Expand Your SCOPE: Semantic Cognition over Potential-Based Exploration for Embodied Visual Navigation
abstract
Embodied visual navigation remains a challenging task, as agents must explore unknown environments with limited knowledge. Existing zero-shot studies have shown that incorporating memory mechanisms to support goal-directed behavior can improve long-horizon planning performance. However, they overlook visual frontier boundaries, which fundamentally dictate future trajectories and observations, and fall short of inferring the relationship between partial visual observations and navigation goals. In this paper, we propose Semantic Cognition Over Potential-based Exploration (SCOPE), a zero-shot framework that explicitly leverages frontier information to drive potential-based exploration, enabling more informed and goal-relevant decisions. SCOPE estimates exploration potential with a Vision-Language Model and organizes it into a spatio-temporal potential graph, capturing boundary dynamics to support long-horizon planning. In addition, SCOPE incorporates a self-reconsideration mechanism that revisits and refines prior decisions, enhancing reliability and reducing overconfident errors. Experimental results on two diverse embodied navigation tasks show that SCOPE outperforms state-of-the-art baselines by 4.6% in accuracy. Further analysis demonstrates that its core components lead to improved calibration, stronger generalization, and higher decision quality.
Ningnan Wang, Weihuang Chen, Haoxuan Ji, Zhongyu Guo, Xuchong Zhang, Hongbin Sun 0001
AAAI6
2026 A Low-Error Approximate Logarithmic Multiplier with Symmetric LUT for Efficient DNN Training
Baoting Li, Tai Yu, Xuchong Zhang, Hongbin Sun 0001
ISCAS5
2026 QB-MOTR: A simple query bootstrapping end-to-end multi-object tracking method with transformer
Zifan Han, Xuchong Zhang
Comput. Vis. Image Underst.2
2026 MFF-DCNet: A Network With Multifeature Focus and Depth-Wise Cross-Stage Transformer for UAV Infrared Small-Object Detection
abstract
The detection of infrared small objects from unmanned aerial vehicles (UAVs) is critical for a wide range of Internet of Things (IoT) applications, including reconnaissance, surveillance, and security monitoring. However, existing methods for small object detection are primarily designed for visible light images and exhibit poor performance when applied to infrared images due to their distinct characteristics such as lower resolution, lack of color and texture information, and higher noise levels. Most existing infrared small object detection algorithms are based on segmentation networks, which often struggle with false alarms when processing UAV-captured imagery with complex backgrounds. Moreover, these segmentation networks are computationally intensive, making them unsuitable for deployment on IoT edge devices. To address these challenges, we propose MFF-DCNet, an efficient network specifically designed for infrared small object detection in UAVs. The proposed network comprises a novel Depth-wise Cross-stage Transformer enhanced backbone and a Multi-Feature Focus neck structure, collectively strengthening multiscale feature extraction and representation. Evaluations on the HIT-UAV and DroneVehicle dataset demonstrate that the proposed network achieves state-of-the-art performance with anAP50−95of 57.4%, representing a 5.8% improvement over specific UAV imagery detectors while simultaneously achieving a 10% increase in FPS. Furthermore, our method achieves real-time performance of 39.6 FPS on the NVIDIA Jetson Orin NX, demonstrating its practical deployment capability in resource constrained IoT environments.
Xuchong Zhang, Hongbin Sun 0001
IEEE Internet Things J.3
2025 Latent Feature and Attention Dual Erasure Attack against Multi-View Diffusion Models for 3D Assets Protection
abstract
Multi-View Diffusion Models (MVDMs) enable remarkable improvements in the field of 3D geometric reconstruction, but the issue regarding intellectual property has received increasing attention due to unauthorized imitation. Recently, some works have utilized adversarial attacks to protect copyright. However, all these works focus on single-image generation tasks which only need to consider the inner feature of images. Previous methods are inefficient in attacking MVDMs because they lack the consideration of disrupting the geometric and visual consistency among the generated multi-view images. This paper is the first to address the intellectual property infringement issue arising from MVDMs. Accordingly, we propose a novel latent feature and attention dual erasure attack to disrupt the distribution of latent feature and the consistency across the generated images from multi-view and multi-domain simultaneously. The experiments conducted on SOTA MVDMs indicate that our approach achieves superior performances in terms of attack effectiveness, transferability, and robustness against defense methods. Therefore, this paper provides an efficient solution to protect 3D assets from MVDMs-based 3D geometry reconstruction. The code is publicly available at: https://github.com/super-jw/LFADEA
Xuchong Zhang, Changfeng Sun, Qicheng Bai, Hongbin Sun 0001
ICME2
2025 BAP-DETR: Efficient drone object detection network based on bipartite attentive processing and dual fusion encoder
Xuchong Zhang
Comput. Vis. Image Underst.3
2025 Object-fabrication targeted attack for object detection
Xuchong Zhang, Changfeng Sun, Haoliang Han, Hongbin Sun 0001
Neurocomputing1
2024 An Efficient Sparse-Aware Summation Optimization Strategy for DNN Accelerator
abstract
Due to the various applications and high sparsity of deep neural network (DNN), a lot of sparse-aware DNN accelerators have been proposed to exploit the sparsity in DNN. Furthermore, it is essential to optimize for the accumulations and inter-channel aggregations in DNN to reduce memory overhead and improve performance of DNN accelerator. However, the uncertain number and location of non-zero element in DNN pose critical challenges for optimizing such accelerators and this inspires us to explore an efficient spare-aware summation optimization strategy for DNN accelerator. In this paper, we leverage the strategy that trading higher cost memory storage/access for lower cost computation to propose a random index based sparse-aware adder tree (RAT), which achieves a better trade-off among performance, hardware resource overhead and adaptability. Synthesis and simulation results demonstrate that, compared with reference design, the proposed design achieves 1.71× and 1.52× the normalized area efficiency and energy efficiency improvement on ResNet18, respectively.
Danqing Zhang, Baoting Li, Xuchong Zhang, Hongbin Sun 0001
ISCAS4
2024 Targeted context attack for object detection
Changfeng Sun, Xuchong Zhang, Haoliang Han, Hongbin Sun 0001
Neurocomputing2
2024 DQ-STP: An Efficient Sparse On-Device Training Processor Based on Low-Rank Decomposition and Quantization for DNN
abstract
Due to the bottleneck problems such as scenario-varying application, significant data communication overhead and privacy protection between off-line training and on-line inference, intelligent edge devices capable of adaptively fine-tuning the deep neural network (DNN) models for specific tasks have become the most urgent need. However, the computational cost is intolerable for ordinary on-device training (ODT), which inspires us to explore an efficient ODT processor, named DQ-STP. In this paper, we leverage a series of optimization techniques using software-hardware co-design. On the one hand, the proposed design incorporates SVD-based low-rank decomposition,$2^{n}$quantization and ACBN algorithm on the software side. This unifies the sparse computing mode of convolutional layers and enhancing weight sparsity. On the other hand, the proposed design effectively leverages data sparsity on the hardware side through four techniques: 1) The flag compressed sparse row is proposed to compress input feature maps and gradient maps. 2) A unified processing element (PE) array comprising shifters and adders is proposed to expedite forward and error propagation steps. 3) The PE arrays for error propagation and weight gradients generation are separated to enhance throughput. 4) A sparse alignment strategy is proposed to further enhance PE utilization. Through these software and hardware co-optimization, the proposed DQ-STP achieves an area efficiency and peak energy efficiency of 41.2 GOPS/mm2 and 90.63 TOPS/W. In comparison to state-of-the-art reference designs, the proposed DQ-STP demonstrates a$2.19\times $improvement in normalized area efficiency and a$1.85\times $enhancement in energy efficiency.
Baoting Li, Danqing Zhang, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 Toward Robust LiDAR-Camera Fusion in BEV Space via Mutual Deformable Attention and Temporal Aggregation
abstract
LiDAR and camera are two critical sensors that can provide complementary information for accurate 3D object detection. Most works are devoted to improving the detection performance of fusion models on the clean and well-collected datasets. However, the collected point clouds and images in real scenarios may be corrupted to various degrees due to potential sensor malfunctions, which greatly affects the robustness of the fusion model and poses a threat to safe deployment. In this paper, we first analyze the shortcomings of most fusion detectors, which rely mainly on the LiDAR branch, and the potential of the bird’s eye-view (BEV) paradigm in dealing with partial sensor failures. Based on that, we present a robust LiDAR-camera fusion pipeline in unified BEV space with two novel designs under four typical LiDAR-camera malfunction cases. Specifically, a mutual deformable attention is proposed to dynamically model the spatial feature relationship and reduce the interference caused by the corrupted modality, and a temporal aggregation module is devised to fully utilize the rich information in the temporal domain. Together with the decoupled feature extraction for each modality and holistic BEV space fusion, the proposed detector, termed RobBEV, can work stably regardless of single-modality data corruption. Extensive experiments on the large-scale nuScenes dataset under robust settings demonstrate the effectiveness of our approach.
Jian Wang 0113, Fan Li 0003, Yi An, Xuchong Zhang, Hongbin Sun 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Physical Strip Attack for Object Detection in Optical Remote Sensing
abstract
A growing trend in the field of adversarial attacks is evolving from the digital domain to the more challenging physical domain. The previous works mainly employ printable adversarial patches with special textures in real-world physical attacks. However, due to lighting conditions and atmospheric scattering, the texture-based patches are prone to distortion in the long-range situation than in the close-range case, resulting in poor physical attack performance in remote sensing scenarios. Therefore, this article proposes a new physical attack method using single-color strip-based patches to hide the objects from being detected correctly in optical aerial detection. Specifically, we design a differentiable representation and an optimization method to optimize the position, thickness, and color of the adversarial strips. Compared with the traditional complex texture-based patch, the proposed strip-based patch is more robust when mapping from the digital domain to the physical domain. Extensive experiments are conducted on multiple datasets and real-world scenarios to evaluate the attack performance of various attack methods. The results show that the proposed strip-based adversarial patch has better attack performance against white-box, black-box, and even defense detectors. Furthermore, we can improve the physical attack success rate (ASR) in remote sensing scenarios by about 70% compared with previous texture-based methods.
Changfeng Sun, Xuchong Zhang, Qicheng Bai, Hongbin Sun 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Adversarial Obstacle Generation Against LiDAR-Based 3D Object Detection
abstract
LiDAR sensors are widely used in many safety-critical applications such as autonomous driving and drone control, and the collected data called point clouds are subsequently processed by 3D object detectors for visual perception. Recent works have shown that attackers can inject virtual points into LiDAR sensors by strategically transmitting laser pulses to them; additionally, deep visual models have been found to be vulnerable to carefully crafted adversarial examples. Therefore, a LiDAR-based perception may be maliciously attacked with serious safety consequences. In this article, we present a highly-deceptive adversarial obstacle generation algorithm against deep 3D detection models, to mimic fake obstacles within the effective detection range of LiDAR using a limited number of points. To achieve this goal, we first perform a physical LiDAR simulation to construct sparse obstacle point clouds. Then, we devise a strong attack strategy to adversarially perturb prototype points along each direction of the ray. Our method achieves a high attack success rate while complying with physical laws at the hardware level. We perform comprehensive experiments on different types of 3D detectors and determine that the voxel-based detectors are more vulnerable to adversarial attacks than the point-based methods. For example, our approach achieves an 89% mean attack success rate against PV-RCNN by using only 20 points to spoof a fake car.
Jian Wang 0113, Fan Li 0003, Xuchong Zhang, Hongbin Sun 0001
IEEE Trans. Multim.3
2023 Boosting Lidar 3D Object Detection with Point Cloud Semantic Segmentation
abstract
The integration of semantic information can effectively enhance the performance of 3D object detection based on lidar point cloud. Most of previous researches utilize camera-lidar fusion to improve detection accuracy for distant or small objects. However, this approach is typically unsuitable for real-time applications due to the large amount of input data. Recently, a multi-task framework using only Iidar has emerged as an alternative that employs the same feature extraction backbone with different heads to simultaneously output detection and semantic segmentation results for lidar point clouds. Nonetheless, some previous works have failed to achieve an optimal balance between accuracy and speed. To address this issue, we propose a multi-task framework which leverages the Cartesian pillar and a multi-scale semantic segmentation head to overcome the shortcomings of existing works and improve the detection accuracy. We evaluate the proposed method using typical pillar-based and voxel-based detection models on the nuScenes dataset. The experimental results demonstrate that the proposed design achieves better performance especially on small objects, compared to single-task models. Moreover, the proposed network increases mAP and NDS by 3.1 % and 2.5 % respectively on the nuScenes test set, compared to the representative multi-task network.
Xuchong Zhang, Chong Min, Yijie Jia, Jingmin Zhang, Hongbin Sun 0001
IROS1
2023 Efficient multi-stage network with pixel-wise degradation prediction for real-time motion deblurring
Zeyu Hao, Xuchong Zhang, Yuhai Li, Hongbin Sun 0001
Comput. Vis. Image Underst.3
2023 ACBN: Approximate Calculated Batch Normalization for Efficient DNN On-Device Training Processor
abstract
Batch normalization (BN) has been established as a very effective component in deep learning, largely helping accelerate the convergence of deep neural network (DNN) training. Nevertheless, its hardware architecture has not received much attention in the field of DNN on-device training processors. Several previous designs incur either high off-chip memory traffic or high circuit complexity, and hence have deficiencies in terms of hardware efficiency and performance. This article proposes approximately calculated BN (ACBN) to achieve a much better tradeoff between hardware efficiency and performance for DNN on-device training processors. The accuracy and convergence rate of the proposed ACBN have been extensively evaluated using four typical DNN models. Compared with the state-of-the-art reference design, the hardware simulation results show the proposed ACBN can at least reduce floating point operations by 22.2% and save external memory access by 33.3% on average. Moreover, the proposed ACBN introduces 63.6% data sparsity for the backward propagation of BN layers of VGG16 on average. To the best of our knowledge, we are the first to introduce data sparsity for the backward propagation of BN layers. The ACBN module is implemented on Zynq UltraScale+ ZCU102 system-on-chip (SoC) field-programmable gate array (FPGA), and the results show that the implementation of ACBN hardware module saves 33.9% look-up table (LUT), 49.4% flip-flop (FF), 75% digital signal processor (DSP), and reduces the power by 12.4% compared with the reference design while achieving better performance.
Baoting Li, Fujie Luo, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2022 Towards high-quality thermal infrared image colorization via attention-based hierarchical network
Xuchong Zhang, Hongbin Sun 0001
Neurocomputing3
2022 End-to-end learning of self-rectification and self-supervised disparity prediction for stereo vision
Xuchong Zhang, Han Zhai, Hongbin Sun 0001, Nanning Zheng 0001
Neurocomputing1
2022 Adaptive Disparity Candidates Prediction Network for Efficient Real-Time Stereo Matching
abstract
Efficient real-time disparity estimation is critical for the application of stereo vision systems in various areas. Recently, stereo network based on coarse-to-fine method has largely relieved the memory constraints and speed limitations of large-scale network models. Nevertheless, all of the previous coarse-to-fine designs employ constant offsets and three or more stages to progressively refine the coarse disparity map, still resulting in unsatisfactory computation accuracy and inference time when deployed on mobile devices. This paper claims that the coarse matching errors can be corrected efficiently with fewer stages as long as more accurate disparity candidates can be provided. Therefore, we propose a dynamic offset prediction module to meet different correction requirements of diverse objects and design an efficient two-stage framework. In addition, a disparity-independent convolution is proposed to regularize the compact cost volume efficiently and further improve the overall performance. The disparity quality and efficiency of various stereo networks are evaluated on multiple datasets and platforms. Evaluation results demonstrate that, the disparity error rate of the proposed network achieves 2.66% and 2.71% on KITTI 2012 and 2015 test sets respectively, where the computation speed is$2\times $faster than the state-of-the-art lightweight models on high-end and source-constrained GPUs.
He Dai, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Dynamic Dataflow Scheduling and Computation Mapping Techniques for Efficient Depthwise Separable Convolution Acceleration
abstract
Depthwise separable convolution (DSC) has become one of the essential structures for lightweight convolutional neural networks. Nevertheless, its hardware architecture has not received much attention. Several previous hardware designs incur either high off-chip memory traffic or large on-chip memory usage, and hence have deficiency in terms of hardware efficiency as well as performance. This paper proposes two efficient dynamic design techniques, i.e. adaptive row-based dataflow scheduling and adaptive computation mapping, to achieve a much better trade-off between hardware efficiency and performance for DSC-based lightweight CNN accelerator. The effectiveness and efficiency of the proposed dynamic design techniques have been extensively evaluated using six DSC-based lightweight CNNs. Compared with the reference architectures, the simulation results show the proposed architectural techniques can at least reduce on-chip buffer size by 50.4% and improve the performance of convolution calculation by 1.18× while maintaining the minimum off-chip memory traffic. MobileNetV2 is implemented on Zynq UltraScale+ ZCU102 SoC FPGA, and the results show the proposed accelerator can achieve 381.7 frames per second (fps), which is 1.43× of the reference design, and it can save about 36.3% on-chip buffer size compared with the reference design, while maintaining the same off-chip memory traffic.
Baoting Li, Xuchong Zhang, Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2020 Algorithm and VLSI Architecture Co-Design on Efficient Semi-Global Stereo Matching
abstract
Semi-global matching (SGM) is favored for high accuracy real-time stereo matching design as it achieves a good trade-off between disparity image quality and computational complexity. Nevertheless, most of previous SGM designs so far are restricted to the real-time processing of small image resolution and disparity range, or achieve high throughput by simplifying the original algorithm at the penalty of significant disparity image quality degradation. We analyze that the major challenge to efficient SGM design is its memory architecture, including both on-chip memory cost and off-chip memory bandwidth. We address the memory architecture challenge by algorithm and architecture co-design. Based on two observed features of SGM algorithm, i.e. incompleteness and inaccuracy, this paper proposes several efficient techniques to reduce on-chip memory cost and compress off-chip memory bandwidth respectively. Moreover, we also design high throughput and pipelined architecture to implement the proposed techniques. The disparity image quality and hardware efficiency of the proposed SGM design are evaluated on both KITTI2015 and Middlebury V3 stereo datasets. Evaluation results demonstrate that, the throughput of the proposed circuit designs can easily achieve 1080P@30fps at the disparity range of 128, and can reduce the on-chip memory cost and off-chip memory bandwidth by up to 4× and 2× respectively while achieving better or the same disparity image quality, compared with the best reference design techniques.
Xuchong Zhang, He Dai, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 NIPM-sWMF: Toward Efficient FPGA Design for High-Definition Large-Disparity Stereo Matching
abstract
Large disparity stereo matching is critical to the application of a stereo vision system especially for outdoor scenes. Nevertheless, how to efficiently design high accuracy large-disparity stereo matching on a field-programmable gate array (FPGA) is still a grand challenge. The computational complexity of previously proposed stereo matching is inevitably proportional to disparity range; hence their hardware designs become very inefficient when the disparity range is large. Motivated by the original PatchMatch and weighted median filtering (WMF) algorithms, this paper proposes a non-iterative PatchMatch and separable WMF (NIPM-sWMF) algorithm to significantly reduce the computational complexity of stereo matching and make it independent of disparity range. Moreover, we also propose a fully pipelined architecture design on FPGA that employs several hardware techniques to efficiently implement the proposed NIPM-sWMF. The disparity quality of the proposed NIPM-sWMF algorithm is evaluated on both KITTI2015 and Middlebury V3 stereo data sets, and the proposed architecture design is implemented and synthesized on Xilinx FPGA. Evaluation results demonstrate that the proposed NIPM-sWMF design on FPGA reaches the real-time performance of 1920 × 1080@60 Hz at the disparity range of 128, and can achieve almost the same disparity estimation accuracy, 4.5× processing throughput, while reducing the hardware cost of LUT, Register, DSP, and BRAM by 40%, 47%, 100%, and 68%, respectively, compared with the reference stereo matching design. Therefore, the proposed NIPM-sWMF design is an efficient way to address the challenge of large-disparity stereo matching.
Xuchong Zhang, Hongbin Sun 0001, Shiqiang Chen, Lin Song 0002, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 VLSI Architecture Exploration of Guided Image Filtering for 1080P@60Hz Video Processing
abstract
Guided image filtering (GIF) is a promising edge-preserving filtering technique that has been applied in a variety of applications. Nevertheless, an efficient very-large-scale integration (VLSI) architecture design of GIF is still very challenging for the real-time processing of full-high definition videos. Previously proposed architectures are somewhat inefficient in terms of either on-chip memory usage or off-chip memory bandwidth. This paper aims to improve the balance between on-chip memory usage and off-chip memory bandwidth through architecture exploration. Three critical architectural tradeoffs in the VLSI design of GIF are explored, and two efficient VLSI architectures, namely sequential line-based and parallel line-based architectures, are proposed. Experimental results demonstrate that the proposed VLSI design only consumes 34.1-K logic gates, 25.4-KB on-chip memories, and 373-MB/s off-chip memory bandwidth while achieving a real-time video processing of 1080P@60Hz at the maximum clock frequency of 297-MHz. Moreover, the proposed VLSI circuits are fully pipelined and synchronized to the pixel clock of output video, so can be seamlessly integrated into diverse real-time video processing systems.
Xuchong Zhang, Hongbin Sun 0001, Shiqiang Chen, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 sWMF: Separable weighted median filter for efficient large-disparity stereo matching
abstract
Although large disparity stereo matching is critical to the practical application of stereo vision system especially for outdoor scenes, its efficient hardware design is still a grand challenge. Motivated by the discovery that well-designed weighted median filter (WMF) can achieve satisfactory accuracy with simple box-filter aggregation, this paper proposes a separable weighted median filter (sWMF) that only has the computational complexity of O(r) and is independent of disparity range. Moreover, the proposed sWMF can be efficiently implemented as a fully pipelined architecture. Evaluation results demonstrate that, at the penalty of only 0.06% disparity error rate, the proposed sWMF design can save 12.9% Slice LUTs, 76.7% DSPs and 64.0% Block RAMs at the disparity range of 128, compared with previous WMF implementation on FPGA.
Shiqiang Chen, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
ISCAS2
2016 Algorithm and VLSI Architecture of Edge-Directed Image Upscaling for 4k Display System
abstract
High-quality and cost-efficient image upscaling design is very important for many real-time video processing applications, especially when the display panel resolution reaches ultrahigh definition. Compared with New Edge-Directed Interpolation (NEDI) based implicit edge directional upscaling, explicit methods require less computational resource and more easily reach real-time performance, especially when the required image definition and upscaling ratio are very high. Nevertheless, the investigation of applications of explicit methods in video processing systems remains largely missing arguably because it is commonly believed that explicit edge-directed interpolation tends to introduce unexpected artifacts because of inaccurate detection and hence its image quality is relatively poor. This paper proposes an explicit edge-directed adaptive interpolation method that leverages more sophisticated edge detection and orientation estimation algorithms to avoid misinterpolation, thereby providing similar or even better image quality than those with implicit methods. Targeting the real-time 4K video display system, the proposed edge-directed image upscaling algorithm is further implemented with an efficient very-large-scale integration (VLSI) architecture. The experimental results demonstrate that the proposed interpolation algorithm outperforms previous explicit and implicit edge-directed methods in both objective and subjective tests. The presented VLSI implementation further demonstrates that the maximum output video sequence of the proposed interpolation method can reach 4k × 2k@60 Hz with a reasonable hardware cost.
Qiubo Chen, Hongbin Sun 0001, Xuchong Zhang, Huibin Tao, Jie Yang 0001, Jizhong Zhao, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3