Sheng Zhong 0001

dblp:53/4506-1 · DBLP profile ↗
← Back
37ranked-venue papers
1as first author
24since 2021 · last 2026
0000-0003-2865-8202ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FineEdu: a fine-grained class students behavior understanding dataset with jointly action and attention annotations
Zhijun Zhang 0009, Ziyue Feng, Zhiying Yan, Xu Zou 0002, Sheng Zhong 0001
Neural Comput. Appl.5
2026 Supervisory feedback for high-resolution low-textured large-scale multi-view stereo
Yongjian Liao, Shixiang Huang, Chunxi Li, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
Pattern Recognit.8
2025 DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes
abstract
Vision-centric autonomous driving systems require diverse data for robust training and evaluation, which can be augmented by manipulating object positions and appearances within existing scene captures. While recent advancements in diffusion models have shown promise in video editing, their application to object manipulation in driving scenarios remains challenging due to imprecise positional control and difficulties in preserving high-fidelity object appearances. To address these challenges in position and appearance control, we introduce DriveEditor, the first diffusion-based framework for object editing in driving videos. DriveEditor offers a unified framework for comprehensive object editing operations, including repositioning, replacement, deletion, and insertion. These diverse manipulations are all achieved through a shared set of varying inputs, processed by identical position control and appearance maintenance modules. The position control module projects the given 3D bounding box while preserving depth information and hierarchically injects it into the diffusion process, enabling precise control over object position and orientation. The appearance maintenance module preserves consistent attributes with a single reference image by employing a three-tiered approach: low-level detail preservation, high-level semantic maintenance, and the integration of 3D priors from a novel view synthesis model. Extensive qualitative and quantitative evaluations on the nuScenes dataset demonstrate DriveEditor's exceptional fidelity and controllability in generating diverse driving scene edits, as well as its remarkable ability to facilitate downstream tasks.
Yiyuan Liang 0001, Zhiying Yan, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
AAAI6
2025 High-dimension Prototype is a Better Incremental Object Detection Learner
abstract
Incremental object detection (IOD), surpassing simple classification, requires the simultaneous overcoming of catastrophic forgetting in both recognition and localization tasks, primarily due to the significantly higher feature space complexity. Integrating Knowledge Distillation (KD) would mitigate the occurrence of catastrophic forgetting. However, the challenge of knowledge shift caused by invisible previous task data hampers existing KD-based methods, leading to limited improvements in IOD performance. This paper aims to alleviate knowledge shift by enhancing the accuracy and granularity in describing complex high-dimensional feature spaces. To this end, we put forth a novel higher-dimension-prototype learning approach for KD-based IOD, enabling a more flexible, accurate, and fine-grained representation of feature distributions without the need to retain any previous task data. Existing prototype learning methods calculate feature centroids or statistical Gaussian distributions as prototypes, disregarding actual irregular distribution information or leading to inter-class feature overlap, which is not directly applicable to the more difficult task of IOD with complex feature space. To address the above issue, we propose a Gaussian Mixture Distribution-based Prototype (GMDP), which explicitly models the distribution relationships of different classes by directly measuring the likelihood of embedding from new and old models into class distribution prototypes in a higher dimension manner. Specifically, GMDP dynamically adapts the component weights and corresponding means/variances of class distribution prototypes to represent both intra-class and inter-class variability more accurately. Progressing into a new task, GMDP constrains the distance between the distribution of new and previous task classes, minimizing overlap with existing classes and thus striking a balance between stability and adaptability. GMDP can be readily integrated into existing IOD methods to enhance performance further. Extensive experiments on the PASCAL VOC and MS-COCO show that our method consistently exceeds four baselines by a large margin and significantly outperforms other SOTA results under various settings.
Tianming Zhao 0003, Tao Zhang 0147, Guodong Wang 0001, Luxin Yan, Sheng Zhong 0001, Jiahuan Zhou, Xu Zou 0002
ICLR7
2025 Divide-And-Conquer: Dual-Hierarchical Optimization for Semantic 4D Gaussian Spatting
abstract
Semantic 4D Gaussians can be used for reconstructing and understanding dynamic scenes, with temporal variations than static scenes. Directly applying static methods to understand dynamic scenes will fail to capture the temporal features. Few works focus on dynamic scene understanding based on Gaussian Splatting, since once the same update strategy is employed for both dynamic and static parts, regardless of the distinction and interaction between Gaussians, significant artifacts and noise appear. We propose Dual-Hierarchical Optimization (DHO), which consists of Hierarchical Gaussian Flow and Hierarchical Gaussian Guidance in a divide-and-conquer manner. The former implements effective division of static and dynamic rendering and features. The latter helps to mitigate the issue of dynamic foreground rendering distortion in textured complex scenes. Extensive experiments show that our method consistently outperforms the baselines on both synthetic and real-world datasets, and supports various downstream tasks. Project Page: https://sweety-yan.github.io/DHO/
Zhiying Yan, Yiyuan Liang 0001, Shilv Cai, Tao Zhang 0147, Sheng Zhong 0001, Luxin Yan, Xu Zou 0002
ICME5
2025 Consistent Learning of Sparse Background Features for Infrared Small-Target Labeling
abstract
Recent years have witnessed many remarkable achievements in infrared small target detection based on deep learning. To achieve high performance in real-world applications, deep learning methods require a substantial number of accurate labels. However, infrared small target annotating is labor intensive as they are very small. Determining and annotating edge pixels demands considerable time and effort, which slows down data expansion and further research in infrared small target detection. To mitigate the issue, we propose a pseudo-label generation method named Consistent Learning of Sparse Background Feature (CLSBF). This approach models the generation of infrared small target pseudo-labels as a domain transformation from local target maps to background ones. It can relax the supervision requirements of deep learning methods, transitioning from absolute pixel-level supervision to point supervision. This method employs an unsupervised approach primarily, supplemented by semi-simulated supervision, to achieve mutual conversion of local images from different domains and obtain the final target pseudo-labels through the differences between target images and transformed background ones. Experiments show that equipped with our method, models trained with the input of a coarsely accurate center label can achieve performance up to 99.94% compared to models trained with official accurate labels. Furthermore, when the official labels are not accurate enough, models trained with pseudo-labels generated by CLSBF consistently show a performance improvement of 1.11% to 4.09% when they are evaluated based on re-labeled bounding boxes. Extensive experiments demonstrate the reliability of CLSBF in generating pseudo-labels and its potential to alleviate the labor-intensive process of manual labeling significantly. We have released an infrared small target labeling tool with CLSBF as an assistant at https://github.com/SeaHifly/CLSBF_software.git.
Sheng Zhong 0001, Luxin Yan, Xu Zou 0002
IEEE Trans. Geosci. Remote. Sens.3
2025 Robust Point Cloud Registration via Patch Matching
abstract
We study the problem of exacting accurate correspondence pairs for point cloud registration. The existing correspondence methods focus on constructing point descriptors and then extracting correspondence point pairs. However, this process encounters two main issues: 1) point features are unstable and susceptible to noise, leading to a low inlier ratio (IR) for correspondence pairs and 2) the positional deviation of correspondence point pair results in accuracy errors when computing rigid transformations. To address these issues, we propose a robust point cloud registration framework based on patch matching, achieving high positional accuracy and high inlier-rate prediction of correspondence pairs. Specifically, we design a dual-branch point cloud registration network, with one branch dedicated to patch matching and the other branch to predicting the patch anchor, i.e., the coordinates used for patch matching. For patch matching, we integrate the topology of patches into the attention mechanism and adopt a multilevel patch-matching strategy to enhance the matching success rate. For coordinate prediction, we introduce graph convolutional network (GCN) and cross-attention mechanisms to explore local similar points through information interaction and feature correlation of patch pairs. Thanks to the stability of patch descriptors, our method demonstrates higher robustness compared to existing correspondence methods. Extensive experiments conducted on indoor, outdoor, synthetic, and deformable benchmarks validate the superiority of our method. Additionally, our method achieves certain effectiveness in cross-source point clouds.
Tianming Zhao 0003, Tian Tian 0006, Xu Zou 0002, Luxin Yan, Sheng Zhong 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Hunt Camouflaged Objects via Revealing Mutation Regions
abstract
Due to the high similarity between hidden objects and the surrounding background, camouflaged object detection (COD) remains a challenge. While many recently proposed methods have shown remarkable performance, most of them begin object perception by indiscriminately considering every pixel of the image. However, these early-stage region-insensitive perception methods still struggle to resist background interference, potentially missing subtle pixel changes by not prioritizing potential camouflaged areas initially. Fortunately, we reveal that the availability of an accurate mutation map can significantly enhance camouflaged discrimination ability. To this end, we propose MRNet (Mutation Region Network). MRNet initially generates a mutation map that identifies potential mutation regions exhibiting subtle pixel changes. The generation method involves amplifying and differing pixel changes based on the position and corresponding values of pixels. Subsequently, the selective expansion search operation utilizes the mutation map to extract the mapped graph, effectively reducing interference from background pixels that are distant from the mutation regions. Finally, decoding the mapped graph generates precise masks. Furthermore, we have created the largest test dataset with known categories to advance community research. Extensive experiments conducted on three widely used datasets and our proposed dataset show that MRNet surpasses other methods with superior performance. Source code is publicly available athttps://github.com/XinyueZhangHust/MRNet
Xinyue Zhang 0009, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
IEEE Trans. Inf. Forensics Secur.4
2024 Make Lossy Compression Meaningful for Low-Light Images
abstract
Low-light images frequently occur due to unavoidable environmental influences or technical limitations, such as insufficient lighting or limited exposure time. To achieve better visibility for visual perception, low-light image enhancement is usually adopted. Besides, lossy image compression is vital for meeting the requirements of storage and transmission in computer vision applications. To touch the above two practical demands, current solutions can be categorized into two sequential manners: ``Compress before Enhance (CbE)'' or ``Enhance before Compress (EbC)''. However, both of them are not suitable since: (1) Error accumulation in the individual models plagues sequential solutions. Especially, once low-light images are compressed by existing general lossy image compression approaches, useful information (e.g., texture details) would be lost resulting in a dramatic performance decrease in low-light image enhancement. (2) Due to the intermediate process, the sequential solution introduces an additional burden resulting in low efficiency. We propose a novel joint solution to simultaneously achieve a high compression rate and good enhancement performance for low-light images with much lower computational cost and fewer model parameters. We design an end-to-end trainable architecture, which includes the main enhancement branch and the signal-to-noise ratio (SNR) aware branch. Experimental results show that our proposed joint solution achieves a significant improvement over different combinations of existing state-of-the-art sequential ``Compress before Enhance'' or ``Enhance before Compress'' solutions for low-light images, which would make lossy low-light image compression more meaningful. The project is publicly available at: https://github.com/CaiShilv/Joint-IC-LL.
Shilv Cai, Sheng Zhong 0001, Luxin Yan, Jiahuan Zhou, Xu Zou 0002
AAAI3
2024 SNIDA: Unlocking Few-Shot Object Detection with Non-Linear Semantic Decoupling Augmentation
abstract
Once only a few-shot annotated samples are available, the performance of learning-based object detection would be heavily dropped. Many few-shot object detection (FSOD) methods have been proposed to tackle this issue by adopting image-level augmentations in linear manners. Nevertheless, those handcrafted enhancements often suf-fer from limited diversity and lack of semantic awareness, resulting in unsatisfactory performance. To this end, we propose a Semantic-guided Nonlinear Instance-level Data Augmentation method (SNIDA) for FSOD by decoupling the foreground and background to increase their diversities respectively. We design a semantic awareness enhancement strategy to separate objects from backgrounds. Concretely, masks of instances are extracted by an unsupervised semantic segmentation module. Then the diversity of samples would be improved by fusing instances into different backgrounds. Considering the shortcomings of augmenting images in a limited transformation space of existing traditional data augmentation methods, we introduce an object reconstruction enhancement module. The aim of this module is to generate sufficient diversity and nonlinear training data at the instance level through a semantic-guided masked autoencoder. In this way, the potential of data can be fully exploited in various object detection scenarios. Extensive experiments on PASCAL VOC and MS-COCO demonstrate that the proposed method outperforms base-lines by a large margin and achieves new state-of-the-art results under different shot settings.
Xu Zou 0002, Luxin Yan, Sheng Zhong 0001, Jiahuan Zhou
CVPR4
2024 Powerful Lossy Compression for Noisy Images
abstract
Image compression and denoising represent fundamental challenges in image processing with many real-world applications. To address practical demands, current solutions can be categorized into two main strategies: 1) sequential method; and 2) joint method. However, sequential methods have the disadvantage of error accumulation as there is information loss between multiple individual models. Recently, the academic community began to make some attempts to tackle this problem through end-to-end joint methods. Most of them ignore that different regions of noisy images have different characteristics. To solve these problems, in this paper, our proposed signal-to-noise ratio (SNR) aware joint solution exploits local and non-local features for image compression and denoising simultaneously. We design an end-to-end trainable network, which includes the main encoder branch, the guidance branch, and the signal-to-noise ratio (SNR) aware branch. We conducted extensive experiments on both synthetic and real-world datasets, demonstrating that our joint solution outperforms existing state-of-the-art methods.
Shilv Cai, Xiaoguo Liang, Shuning Cao, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
ICME5
2024 Perceptual-Distortion Balanced Image Super-Resolution is a Multi-Objective Optimization Problem
abstract
Training Single-Image Super-Resolution (SISR) models using pixel-based regression losses can achieve high distortion metrics scores (e.g., PSNR and SSIM), but often results in blurry images due to insufficient recovery of high-frequency details. Conversely, using GAN or perceptual losses can produce sharp images with high perceptual metric scores (e.g., LPIPS), but may introduce artifacts and incorrect textures. Balancing these two types of losses can help achieve a trade-off between distortion and perception, but the challenge lies in tuning the loss function weights. To address this issue, we propose a novel method that incorporates Multi-Objective Optimization (MOO) into the training process of SISR models to balance perceptual quality and distortion. We conceptualize the relationship between loss weights and image quality assessment (IQA) metrics as black-box objective functions to be optimized within our Multi-Objective Bayesian Optimization Super-Resolution (MOBOSR) framework. This approach automates the hyperparameter tuning process, reduces overall computational cost, and enables the use of numerous loss functions simultaneously. Extensive experiments demonstrate that MOBOSR outperforms state-of-the-art methods in terms of both perceptual quality and distortion, significantly advancing the perception-distortion Pareto frontier. Our work points towards a new direction for future research on balancing perceptual quality and fidelity in nearly all image restoration tasks. The source code and pretrained models are available at: https://github.com/ZhuKeven/MOBOSR.
Qiwen Zhu, Shilv Cai, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
ACM Multimedia7
2024 I2C: Invertible Continuous Codec for High-Fidelity Variable-Rate Image Compression
abstract
Lossy image compression is a fundamental technology in media transmission and storage. Variable-rate approaches have recently gained much attention to avoid the usage of a set of different models for compressing images at different rates. During the media sharing, multiple re-encodings with different rates would be inevitably executed. However, existing Variational Autoencoder (VAE)-based approaches would be readily corrupted in such circumstances, resulting in the occurrence of strong artifacts and the destruction of image fidelity. Based on the theoretical findings of preserving image fidelity via invertible transformation, we aim to tackle the issue of high-fidelity fine variable-rate image compression and thus propose the Invertible Continuous Codec (I2C). We implement the I2C in a mathematical invertible manner with the core Invertible Activation Transformation (IAT) module. I2C is constructed upon a single-rate Invertible Neural Network (INN) based model and the quality level (QLevel) would be fed into the IAT to generate scaling and bias tensors. Extensive experiments demonstrate that the proposed I2C method outperforms state-of-the-art variable-rate image compression methods by a large margin, especially after multiple continuous re-encodings with different rates, while having the ability to obtain a very fine variable-rate control without any performance compromise.
Shilv Cai, Zhijun Zhang 0009, Xiangyun Zhao, Jiahuan Zhou, Yuxin Peng 0001, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
IEEE Trans. Pattern Anal. Mach. Intell.8
2024 Learning Oriented Object Detection via Naive Geometric Computing
abstract
Detecting oriented objects along with estimating their rotation information is one crucial step for image analysis, especially for remote sensing images. Despite that many methods proposed recently have achieved remarkable performance, most of them directly learn to predict object directions under the supervision of only one (e.g., the rotation angle) or a few (e.g., several coordinates) groundtruth (GT) values individually. Oriented object detection would be more accurate and robust if extra constraints, with respect to proposal and rotation information regression, are adopted for joint supervision during training. To this end, we propose a mechanism that simultaneously learns the regression of horizontal proposals, oriented proposals, and rotation angles of objects in a consistent manner, via naive geometric computing, as one additional steady constraint. An oriented center prior guided label assignment strategy is proposed for further enhancing the quality of proposals, yielding better performance. Extensive experiments on six datasets demonstrate the model equipped with our idea significantly outperforms the baseline by a large margin and several new state-of-the-art results are achieved without any extra computational burden during inference. Our proposed idea is simple and intuitive that can be readily implemented. Source codes are publicly available at: https://github.com/wangWilson/CGCDet.git.
Zhijun Zhang 0009, Guodong Wang 0001, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
IEEE Trans. Neural Networks Learn. Syst.7
2023 Multiscale Multilevel Residual Feature Fusion for Real-Time Infrared Small Target Detection
abstract
Detecting infrared dim and small targets is one crucial step for many tasks such as early warning. It remains a continuing challenge since characteristics of infrared small targets, usually represented by only a few pixels, are generally not salient. Despite that many traditional methods have significantly advanced the community, their robustness or efficiency is still lacking. Most recently, CNN-based object detection has achieved remarkable performance and some researchers focus on it. However, these methods are not computationally efficient when implemented on some CPU-only machines and few datasets are available publicly. To promote the detection of infrared small targets in complex backgrounds, we propose a new lightweight CNN-based architecture. The network contains three modules: the feature extraction module is designed for representing multi-scale and multi-level features, the grid resample operation module is proposed to fuse features from all scales, and a decoupled head to distinguish infrared small targets from backgrounds. Moreover, we collect a brand-new infrared small target detection dedicated dataset which consists of 68311 practical captured images with complex backgrounds for alleviating the data dilemma. To validate the proposed model, 54758 images are used for training and 13553 images are used for testing respectively. Extensive experimental results demonstrate that the proposed method outperforms all traditional methods by a large margin and runs much faster than other CNN methods with high precision. The proposed model can be implemented on the Intel i7-10850H CPU (2.3GHz) platform and Jetson Nano for real-time infrared small target detection at 44 FPS and 27 FPS, respectively. It can be even deployed on an Atom x5-Z8500 (1.44GHz) machine at about 25 FPS with 128×128 local images. The source codes and the dataset have been made publicly available at https://github.com/SeaHifly/Infrared-Small-Target.
Sheng Zhong 0001, Tianxu Zhang, Xu Zou 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 Category-Aware Transformer Network for Better Human-Object Interaction Detection
abstract
Human-Object Interactions (HOI) detection, which aims to localize a human and a relevant object while recognizing their interaction, is crucial for understanding a still image. Recently, tranformer-based models have significantly advanced the progress of HOI detection. However, the capability of these models has not been fully explored since the Object Query of the model is always simply initialized as just zeros, which would affect the performance. In this paper, we try to study the issue of promoting transformer-based HOI detectors by initializing the Object Query with category-aware semantic information. To this end, we innovatively propose the Category-Aware Transformer Network (CATN). Specifically, the Object Query would be initialized via category priors represented by an external object detection model to yield a better performance. Moreover, such category priors can be further used for enhancing the representation ability of features via the attention mechanism. We have firstly verified our idea via the Oracle experiment by initializing the Object Query with the groundtruth category information. And then extensive experiments have been conducted to show that a HOI detection model equipped with our idea outperforms the baseline by a large margin to achieve a new state-of-the-art result.
Leizhen Dong, Kunlun Xu, Zhijun Zhang 0009, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
CVPR6
2022 High-Fidelity Variable-Rate Image Compression via Invertible Activation Transformation
abstract
Learning-based methods have effectively promoted the community of image compression. Meanwhile, variational autoencoder(VAE) based variable-rate approaches have recently gained much attention to avoid the usage of a set of different networks for various compression rates. Despite the remarkable performance that has been achieved, these approaches would be readily corrupted once multiple compression/decompression operations are executed, resulting in the fact that image quality would be tremendously dropped and strong artifacts would appear. Thus, we try to tackle the issue of high-fidelity fine variable-rate image compression and propose the Invertible Activation Transformation(IAT) module. We implement the IAT in a mathematical invertible manner on a single rate Invertible Neural Network(INN) based model and the quality level(QLevel) would be fed into the IAT to generate scaling and bias tensors. IAT and QLevel together give the image compression model the ability of fine variable-rate control while better maintaining the image fidelity. Extensive experiments demonstrate that the single rate image compression model equipped with our IAT module has the ability to achieve variable-rate control without any compromise. And our IAT-embedded model obtains comparable rate-distortion performance with recent learning-based image compression methods. Furthermore, our method outperforms the state-of-the-art variable-rate image compression method by a large margin, especially after multiple re-encodings.
Shilv Cai, Zhijun Zhang 0009, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
ACM Multimedia5
2022 Effective actor-centric human-object interaction detection
Kunlun Xu, Zhijun Zhang 0009, Leizhen Dong, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
Image Vis. Comput.7
2022 Learning dynamic background for weakly supervised moving object detection
Zhijun Zhang 0009, Yi Chang 0002, Sheng Zhong 0001, Luxin Yan, Xu Zou 0002
Image Vis. Comput.3
2021 A New Orientation Estimation Method Based on Rotation Invariant Gradient for Feature Points
abstract
For many remote sensing image applications, orientation estimation is a crucial step for feature points extraction and matching, but it has attracted little attention. Due to the intensity differences between remote sensing image pairs, it is difficult to estimate the orientations of corresponding points accurately, resulting in performance degradation of feature matching. Thus, encountering the intensity differences, a robust spatial structure description for feature regions, and an effective calculation manner from description to orientation play a key role in accurate orientation estimation. To this end, in this letter, we first define a plausible orientation for feature points by the total gradient, offering an effective way to convert the gradient trend of the feature region to orientation. Therefore, we further propose a novel orientation estimation method, in which the rotation invariant gradient is introduced to improve the accuracy of gradient calculation and robustness of spatial structure description. Experimental results on multisensor remote sensing images demonstrate that our method increases the orientation estimation accuracy remarkably and outperforms other orientation estimation methods by a large margin, and effectively improves the performance of feature matching.
Sheng Zhong 0001, Luxin Yan
IEEE Geosci. Remote. Sens. Lett.2
2021 Category-Aware Aircraft Landmark Detection
abstract
Aircraft landmark detection (ALD) aims at detecting the keypoints of aircraft, which can serve as an important role for subsequent applications such as fine-grained aircraft recognition. In ALD, the physical size discrepancy between different kinds of aircraft may lead to inconsistent landmark structure, which significantly harms landmark detection results. In this letter, we take advantage of the category prior to alleviate the size discrepancy in ALD. The proposed category-aware landmark detection network (CALDN) possesses two streams: a classification stream for size categorization and a localization stream for landmark detection. Instance-level size category information captured by classification stream serves as the guidance in the localization stream for robust landmark detection. Moreover, a category attention module (CAM) is proposed for better-utilizing category information to guide ALD. Benefitting from the adaptive attention mechanism, CAM can automatically highlight category-specific features for ulteriorly reducing the influence of size discrepancy. Furthermore, to advance ALD research, we contribute the first perspective-variant aircraft landmark dataset. Solid experiments demonstrate the superiority of our method.
Yi Li 0033, Yi Chang 0002, Yuntong Ye, Xu Zou 0002, Sheng Zhong 0001, Luxin Yan
IEEE Signal Process. Lett.5
2021 Towards Unconstrained Facial Landmark Detection Robust to Diverse Cropping Manners
abstract
Facial landmark detection is one crucial step for face-based image/video analysis. Despite the fact that recently many facial landmark detection models have achieved remarkable performance, most state-of-the-art heatmap regression-based methods heavily rely on initialization of the face detector. However, there inevitably exists semantic gaps among different annotators or face detectors. An improper facial bounding box will tremendously drop off the performance of the facial landmark detection model. Facial landmark detection would be more practical if robust to face images cropped by diverse manners (see Figure 1, the col.1 shows face images cropped by a proper bounding box, col.2 and col.3 show face images cropped by an oversize and a small bounding boxes respectively). To this end, we present a “Unconstrained Facial Landmark Detection(UFLD)” mechanism, that aims at enhancing the robustness of facial landmark detection, to deal with the inconsistent cropping manner issue. UFLD consists of two aspects: a Transformation-Invariant Landmark Detector(TILD) and an Availability-Guided Solver(AGS). TILD gives the ability to detect consistent landmarks for face images cropped by diverse manners. And AGS can alleviate the by-effect of “landmarks outside the image” caused by improper cropping results or TILD, and further promote the performance. The proposed mechanism achieved above 6.5% improvement in standard normalized landmarks mean error reduction on face images cropped by diverse manners compared to baselines.
Xu Zou 0002, Luxin Yan, Sheng Zhong 0001, Ying Wu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 Hyperspectral Image Restoration: Where Does the Low-Rank Property Exist
abstract
Hyperspectral image (HSI) restoration is to recover the clean image from degraded version, such as the noisy, blurred, or damaged. Recent low-rank tensor-based recovery methods have been widely explored in HSIs restoration. Most of previous methods, however, neglect an inconspicuous but important phenomenon that the physical meaning and dimension along the spatial, spectral, and nonlocal mode are markedly different. In this work, we discover the low-rank property discrepancy along spatial, spectral, and nonlocal self-similarity mode in the HSIs, and argue that the intrinsic low-rank correlations along each mode contribute different to the final restoration results. Consequently, we figure out that the combination of the spectral and nonlocal-induced low-rank is most beneficial for HSIs modeling, and propose an optimal low-rank tensor (OLRT) model for HSIs restoration. Furthermore, we not only explore the low-rank property in the image component, but also in the sparse error component (stripe noise in HSIs). Thus, we extend OLRT to the OLRT-robust principal component analysis (RPCA) with low-rank tensor priors for both the HSIs and sparse error. Besides, previous methods are usually designed for one specific HSI task, which is less robust to various tasks. We prove that the proposed optimal low-rank prior is very flexible for various HSI restoration problems including denoising, deblurring, inpainting, and destriping. The proposed methods have been extensively evaluated on several benchmarks and tasks, and greatly outperform state-of-the-art (STOA). We show the simple yet effective OLRT strategy is also beneficial to STOA.
Yi Chang 0002, Luxin Yan, Bingling Chen, Sheng Zhong 0001, Yonghong Tian 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Visual Navigation and Landing Control of an Unmanned Aerial Vehicle on a Moving Autonomous Surface Vehicle via Adaptive Learning
abstract
This article presents a visual navigation and landing control paradigm for an unmanned aerial vehicle (UAV) to land on a moving autonomous surface vehicle (ASV). Therein, an adaptive learning navigation rule with a multilayer nested guidance is designed to pinpoint the position of the ASV and to guide and control the UAV to fulfill horizontal tracking and vertical descending in a narrow landing region of the ASV by means of merely relative position feedback. To ensure the feasibility of the proposed control law, asymptotical stability conditions are derived based on Lyapunov stability theory. Landing experimental results are reported for a UAV-ASV system consisting of an M-100 UAV and a self-developed three-meters-long HUSTER-30 ASV on a lake to substantiate the efficacy of the proposed landing control method.
Hai-Tao Zhang, Binbin Hu, Zhecheng Xu, Zhi Cai, Tao Geng, Sheng Zhong 0001
IEEE Trans. Neural Networks Learn. Syst.8
2020 Weighted Low-Rank Tensor Recovery for Hyperspectral Image Restoration
abstract
Hyperspectral imaging, providing abundant spatial and spectral information simultaneously, has attracted a lot of interest in recent years. Unfortunately, due to the hardware limitations, the hyperspectral image (HSI) is vulnerable to various degradations, such as noises (random noise), blurs (Gaussian and uniform blur), and downsampled (both spectral and spatial downsample), each corresponding to the HSI denoising, deblurring, and super-resolution tasks, respectively. Previous HSI restoration methods are designed for one specific task only. Besides, most of them start from the 1-D vector or 2-D matrix models and cannot fully exploit the structurally spectral-spatial correlation in 3-D HSI. To overcome these limitations, in this article, we propose a unified low-rank tensor recovery model for comprehensive HSI restoration tasks, in which nonlocal similarity within spectral-spatial cubic and spectral correlation are simultaneously captured by third-order tensors. Furthermore, to improve the capability and flexibility, we formulate it as a weighted low-rank tensor recovery (WLRTR) model by treating the singular values differently. We demonstrate the reweighed strategy, which has been extensively studied in the matrix, also greatly benefits the tensor modeling. We also consider the stripe noise in HSI as the sparse error by extending WLRTR to robust principal component analysis (WLRTR-RPCA). Extensive experiments demonstrate the proposed WLRTR models consistently outperform state-of-the-art methods in typical HSI low-level vision tasks, including denoising, destriping, deblurring, and super-resolution.
Yi Chang 0002, Luxin Yan, Xi-Le Zhao, Houzhang Fang, Zhijun Zhang 0009, Sheng Zhong 0001
IEEE Trans. Cybern.6
2020 Toward Universal Stripe Removal via Wavelet-Based Deep Convolutional Neural Network
abstract
Stripe noise from different remote sensing imaging systems varies considerably in terms of response, length, angle, and periodicity. Due to the complex distributions of different stripes, the destriping results of previous methods may be oversmoothed or contain residual stripe. To overcome this key problem, we provide a comprehensive analysis of existing destriping methods and propose a deep convolutional neural network (CNN) for handling various kinds of stripes. Moreover, previous methods individually model the stripe or the image priors, which may lose the relationship between them. In this article, a two-stream CNN is designed to simultaneously model the stripe and image, which better facilitates distinguishing them from each other. Moreover, we incorporate the wavelet into our CNN model for better directional feature representation. Therefore, the CNN learns the discriminative representation from the external data set, while the wavelet models the internal directionality of the stripe, in which both the internal and external priors are beneficial to the destriping task. In addition, the wavelet extracts the multiscale information with a larger receptive field for global contextual information modeling; thus, we can better distinguish the stripe from the similar image line pattern structures. The proposed method has been extensively evaluated on a number of data sets and outperforms the state-of-the-art methods by substantially a large margin in terms of quantitative and qualitative assessments, speed, and robustness.
Yi Chang 0002, Meiya Chen, Luxin Yan, Xi-Le Zhao, Yi Li 0033, Sheng Zhong 0001
IEEE Trans. Geosci. Remote. Sens.6
2019 Learning Robust Facial Landmark Detection via Hierarchical Structured Ensemble
abstract
Heatmap regression-based models have significantly advanced the progress of facial landmark detection. However, the lack of structural constraints always generates inaccurate heatmaps resulting in poor landmark detection performance. While hierarchical structure modeling methods have been proposed to tackle this issue, they all heavily rely on manually designed tree structures. The designed hierarchical structure is likely to be completely corrupted due to the missing or inaccurate prediction of landmarks. To the best of our knowledge, in the context of deep learning, no work before has investigated how to automatically model proper structures for facial landmarks, by discovering their inherent relations. In this paper, we propose a novel Hierarchical Structured Landmark Ensemble (HSLE) model for learning robust facial landmark detection, by using it as the structural constraints. Different from existing approaches of manually designing structures, our proposed HSLE model is constructed automatically via discovering the most robust patterns so HSLE has the ability to robustly depict both local and holistic landmark structures simultaneously. Our proposed HSLE can be readily plugged into any existing facial landmark detection baselines for further performance improvement. Extensive experimental results demonstrate our approach significantly outperforms the baseline by a large margin to achieve a state-of-the-art performance.
Xu Zou 0002, Sheng Zhong 0001, Luxin Yan, Xiangyun Zhao, Jiahuan Zhou, Ying Wu 0001
ICCV2
2019 Directional-Aware Automatic Defect Detection in High-Speed Railway Catenary System
abstract
It is crucial to detect the defect objects that could be a hidden danger in high-speed catenary system. Instead of detecting objects based on horizontal rectangle, we propose an end-to-end trainable directional-aware defect detection network (D3-Net) which can automatic select the defective components. D3-Net is composed of two streams, including a directional-aware locator that regresses an inclined rectangle and a channel-wise classifier to diagnose the defects of object. Specifically, we make full use of the directionality characteristic of the man-made objects in the localization stream, which can predict the inclined rectangle with an angle parameter. Besides, we introduce a channel attention module (CAM) in the classification stream to obtain discriminative features for better distinguishing the normal and defective objects. Experimental results on two datasets, Dropper and Insulator, demonstrate that our proposed model outperforms the traditional horizontal detection methods.
Yi Chang 0002, Ziqin Li, Sheng Zhong 0001, Luxin Yan
ICIP4
2019 Infrared Aerothermal Nonuniform Correction via Deep Multiscale Residual Network
abstract
In the infrared focal plane arrays imaging systems, the temperature-dependent nonuniformity effects severely degrade the image quality. In this letter, we propose a very deep convolutional neural network for unified infrared aerothermal nonuniform correction. Our network is built with the multiscale and residual training. The multiscale subnetworks utilize the multiscale property in the images, and the long-short-term residual learning contributes to the information propagation. Compared with the previous methods, the proposed method is more robust to various nonuniform artifacts and more efficient at processing time. Experimental results validate the superiority of our method for infrared nonuniform correction.
Yi Chang 0002, Luxin Yan, Li Liu 0050, Houzhang Fang, Sheng Zhong 0001
IEEE Geosci. Remote. Sens. Lett.5
2019 HSI-DeNet: Hyperspectral Image Restoration via Convolutional Neural Network
abstract
The spectral and the spatial information in hyperspectral images (HSIs) are the two sides of the same coin. How to jointly model them is the key issue for HSIs' noise removal, including random noise, structural stripe noise, and dead pixels/lines. In this paper, we introduce the deep convolutional neural network (CNN) to achieve this goal. The learned filters can well extract the spatial information within their local receptive filed. Meanwhile, the spectral correlation can be depicted by the multiple channels of the learned 2-D filters, namely, the number of filters in each layer. The consequent advantages of our CNN-based HSI denoising method (HSI-DeNet) over previous methods are threefold. First, the proposed HSI-DeNet can be regarded as a tensor-based method by directly learning the filters in each layer without damaging the spectral-spatial structures. Second, the HSI-DeNet can simultaneously accommodate various kinds of noise in HSIs. Moreover, our method is flexible for both single image and multiple images by slightly modifying the channels of the filters in the first and last layers. Last but not least, our method is extremely fast in the testing phase, which makes it more practical for real application. The proposed HSI-DeNet is extensively evaluated on several HSIs, and outperforms the state-of-the-art HSI-DeNets in terms of both speed and performance.
Yi Chang 0002, Luxin Yan, Houzhang Fang, Sheng Zhong 0001, Wenshan Liao
IEEE Trans. Geosci. Remote. Sens.4
2018 Motion Correlation Discovery for Visual Tracking
abstract
Motion information plays an important role in identifying moving objects, which has not been well utilized in state-of-the-art tracking algorithms. In this letter, we propose a unified framework integrating two tracking problems, i.e., pixel-level foreground probabilistic inference and motion parameter estimation. Our model employs motion fields to propagate probability forward, and discovers motion patterns in the spatial domain to distinguish targets from the background. It takes advantage of continuity and inertia of both target and camera motion, and provides reliable evidence to resolve confusion caused by appearance similarity between targets and the background. Target localization is effectively achieved from the pixel-level foreground probabilistic map. Experimental results demonstrate that the proposed method significantly improves our baseline method and achieves performance comparable to state-of-the-art tracking methods with more complex features.
Sheng Zhong 0001, Ying Wu 0001
IEEE Signal Process. Lett.2
2017 Hyper-Laplacian Regularized Unidirectional Low-Rank Tensor Recovery for Multispectral Image Denoising
abstract
Recent low-rank based matrix/tensor recovery methods have been widely explored in multispectral images (MSI) denoising. These methods, however, ignore the difference of the intrinsic structure correlation along spatial sparsity, spectral correlation and non-local self-similarity mode. In this paper, we go further by giving a detailed analysis about the rank properties both in matrix and tensor cases, and figure out the non-local self-similarity is the key ingredient, while the low-rank assumption of others may not hold. This motivates us to design a simple yet effective unidirectional low-rank tensor recovery model that is capable of truthfully capturing the intrinsic structure correlation with reduced computational burden. However, the low-rank models suffer from the ringing artifacts, due to the aggregation of overlapped patches/cubics. While previous methods resort to spatial information, we offer a new perspective by utilizing the exclusively spectral information in MSIs to address the issue. The analysis-based hyper-Laplacian prior is introduced to model the global spectral structures, so as to indirectly alleviate the ringing artifacts in spatial domain. The advantages of the proposed method over the existing ones are multi-fold: more reasonably structure correlation representability, less processing time, and less artifacts in the overlapped regions. The proposed method is extensively evaluated on several benchmarks, and significantly outperforms state-of-the-art MSI denoising methods.
Yi Chang 0002, Luxin Yan, Sheng Zhong 0001
CVPR3
2017 Transformed Low-Rank Model for Line Pattern Noise Removal
abstract
This paper addresses the problem of line pattern noise removal from a single image, such as rain streak, hyperspectral stripe and so on. Most of the previous methods model the line pattern noise in original image domain, which fail to explicitly exploit the directional characteristic, thus resulting in a redundant subspace with poor representation ability for those line pattern noise. To achieve a compact subspace for the line pattern structure, in this work, we incorporate a transformation into the image decomposition model so that maps the input image to a domain where the line pattern appearance has an extremely distinct low-rank structure, which naturally allows us to enforce a low-rank prior to extract the line pattern streak/stripe from the noisy image. Moreover, the random noise is usually mixed up with the line pattern noise, which makes the challenging problem much more difficult. While previous methods resort to the spectral or temporal correlation of the multi-images, we give a detailed analysis between the noisy and clean image in both local gradient and nonlocal domain, and propose a compositional directional total variational and low-rank prior for the image layer, thus to simultaneously accommodate both types of noise. The proposed method has been evaluated on two different tasks, including remote sensing image mixed random-stripe noise removal and rain streak removal, all of which obtain very impressive performances.
Yi Chang 0002, Luxin Yan, Sheng Zhong 0001
ICCV3
2017 Hyperspectral image denoising via spectral and spatial low-rank approximation
abstract
Hyperspectral images (HSI) unavoidably suffer from degradations such as random noise, due to photon effects, calibration error, and so on. Most of existing HSI denoising methods focus on utilizing the spectral correlation or the spatial nonlocal self-similarity individually. In this paper, we propose an unified low-rank recovery framework for HSI denoising, in which taking both the underlying characteristics of high correlation across spectra and non-local self-similarity over the space cubic of HSI into consideration simultaneously. Our work rely on a basic observation that both the multiple spectral bands and similar spatial structures are lying on low-rank subspaces and can facilitate to remove the noise jointly. Experimental results on both simulated and real HSI demonstrate that the proposed method can significantly outperform the state-of-the-art methods on several datasets in terms of both visual and quantitative assessment.
Yi Chang 0002, Luxin Yan, Sheng Zhong 0001
IGARSS3
2016 Remote Sensing Image Stripe Noise Removal: From Image Decomposition Perspective
abstract
Stripe noise removal (destriping) is a fundamental problem in remote sensing image processing that holds significant practical importance for subsequent applications. These variational destriping methods have obtained impressive results and attracted widely studied research interests. However, most of them are dedicated to estimate the clear image from the striped one, paying much attention to the image itself, while ignoring the structural characteristic of stripe, which would easily cause damages to the image structure and leave residual stripes in image recovery. In this paper, we treat the image and stripe components equally and convert the image destriping task as an image decomposition problem naturally. We first give a detailed analysis about the structural characteristic of stripes and the prior knowledge about the remote sensing images. Then, incorporating them, we propose a low-rank-based single-image decomposition model (LRSID) to separate the original image from the stripe component perfectly. This low-rank constraint for the stripe perfectly matches the fact that only parts of data vectors are corrupted but the others are not. Moreover, we further utilize the spectral information of the remote sensing images, and we extend our 2-D image decomposition method to the 3-D case. Extensive experiments on both simulated and real data have been carried out to validate the effectiveness and efficiency of the proposed algorithms.
Yi Chang 0002, Luxin Yan, Sheng Zhong 0001
IEEE Trans. Geosci. Remote. Sens.4
2014 An Embedded System-on-Chip Architecture for Real-time Visual Detection and Matching
abstract
Detecting and matching image features is a fundamental task in video analytics and computer vision systems. It establishes the correspondences between two images taken at different time instants or from different viewpoints. However, its large computational complexity has been a challenge to most embedded systems. This paper proposes a new FPGA-based embedded system architecture for feature detection and matching. It consists of scale-invariant feature transform (SIFT) feature detection, as well as binary robust independent elementary features (BRIEF) feature description and matching. It is able to establish accurate correspondences between consecutive frames for 720-p (1280x720) video. It optimizes the FPGA architecture for the SIFT feature detection to reduce the utilization of FPGA resources. Moreover, it implements the BRIEF feature description and matching on FPGA. Due to these contributions, the proposed system achieves feature detection and matching at 60 frame/s for 720-p video. Its processing speed can meet and even exceed the demand of most real-life real-time video analytics applications. Extensive experiments have demonstrated its efficiency and effectiveness.
Sheng Zhong 0001, Luxin Yan, Zhiguo Cao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2013 A real-time embedded architecture for SIFT
Sheng Zhong 0001, Luxin Yan, Lie Kang, Zhiguo Cao 0001
J. Syst. Archit.1