Luxin Yan

dblp:81/9161 · DBLP profile ↗
← Back
80ranked-venue papers
2as first author
58since 2021 · last 2026
0000-0002-5445-2702ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 31 since 2021Artificial intelligence and machine learning · 39 · 33 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 2 first-author · 12 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Spatio-Temporal Context Learning with Temporal Difference Convolution for Moving Infrared Small Target Detection
abstract
Moving infrared small target detection (IRSTD) plays a critical role in practical applications, such as surveillance of unmanned aerial vehicles (UAVs) and UAV-based search system. Moving IRSTD still remains highly challenging due to weak target features and complex background interference. Accurate spatio-temporal feature modeling is crucial for moving target detection, typically achieved through either temporal differences or spatio-temporal (3D) convolutions. Temporal difference can explicitly leverage motion cues but exhibits limited capability in extracting spatial features, whereas 3D convolution effectively represents spatio-temporal features yet lacks explicit awareness of motion dynamics along the temporal dimension. In this paper, we propose a novel moving IRSTD network (TDCNet), which effectively extracts and enhances spatio-temporal features for accurate target detection. Specifically, we introduce a novel temporal difference convolution (TDC) re-parameterization module that comprises three parallel TDC blocks designed to capture contextual dependencies across different temporal ranges. Each TDC block fuses temporal difference and 3D convolution into a unified spatio-temporal convolution representation. This re-parameterized module can effectively capture multi-scale motion contextual features while suppressing pseudo-motion clutter in complex backgrounds, significantly improving detection performance. Moreover, we propose a TDC-guided spatio-temporal attention mechanism that performs cross-attention between the spatio-temporal features extracted from the TDC-based backbone and a parallel 3D backbone. This mechanism models their global semantic dependencies to refine the current frame’s features, thereby guiding the model to focus more accurately on critical target regions. To facilitate comprehensive evaluation, we construct a new challenging benchmark, IRSTD-UAV, consisting of 15,106 real infrared images with diverse low signal-to-clutter ratio scenarios and complex backgrounds. Extensive experiments on IRSTD-UAV and public infrared datasets demonstrate that our TDCNet achieves state-of-the-art detection performance in moving target detection.
Houzhang Fang, Shukai Guo, Qiuhuan Chen, Yi Chang 0002, Luxin Yan
AAAI5
2026 Blur-Robust Detection via Feature Restoration: An End-to-End Framework for Prior-Guided Infrared UAV Target Detection
abstract
Infrared unmanned aerial vehicle (UAV) target images often suffer from motion blur degradation caused by rapid sensor movement, significantly reducing contrast between target and background. Generally, detection performance heavily depends on the discriminative feature representation between target and background. Existing methods typically treat deblurring as a preprocessing step focused on visual quality, while neglecting the enhancement of task-relevant features crucial for detection. Improving feature representation for detection under blur conditions remains challenging. In this paper, we propose a novel Joint Feature-Domain Deblurring and Detection end-to-end framework, dubbed JFD³. We design a dual-branch architecture with shared weights, where the clear branch guides the blurred branch to enhance discriminative feature representation. Specifically, we first introduce a lightweight feature restoration network, where features from the clear branch serve as feature-level supervision to guide the blurred branch, thereby enhancing its distinctive capability for detection. We then propose a frequency structure guidance module that refines the structure prior from the restoration network and integrates it into shallow detection layers to enrich target structural information. Finally, a feature consistency self-supervised loss is imposed between the dual-branch detection backbones, driving the blurred branch to approximate the feature representations of the clear one. We also construct a benchmark, named IRBlurUAV, containing 30,000 simulated and 4,118 real infrared UAV target images with diverse motion blur. Extensive experiments on IRBlurUAV demonstrate that JFD³ achieves superior detection performance while maintaining real-time efficiency.
Xiaolin Wang 0006, Houzhang Fang, Qingshan Li, Lu Wang 0014, Yi Chang 0002, Luxin Yan
AAAI6
2026 Supervisory feedback for high-resolution low-textured large-scale multi-view stereo
Yongjian Liao, Shixiang Huang, Chunxi Li, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
Pattern Recognit.7
2026 A 240 × 180 Event-Based Vision Sensor ROIC With Global Threshold Voltage Calibration Technology
abstract
This paper presents a$240\times 180$event-based vision sensor (EVS) readout integrated circuit (ROIC) that incorporates global threshold voltage calibration technology. Conventional EVS suffer from error events and intra-array mismatch. This paper delves into four types of error events and presents effective solutions for them. This paper proposes a global threshold voltage calibration (GTVC) technique to minimize the impact of mismatch in the pixel array. The pixel size measures$10\times 10~\mu $m${}^{\mathbf {2}}$, using 40 nm CMOS technology. The analog pixel circuits work with supply voltages of 2.5 V and 1.1 V, while the power consumption of this chip is 12.16 mW. By using the global threshold voltage calibration technology, which incorporates eight reference pixels for determining the input voltage of event comparators, the error in the detection result is reduced to 2.84 mV. The maximum event rate reaches 360 Meps with a 40 MHz system clock, boasting a dynamic range of 97.3 dB. Further, the event power efficiency stands at 29.6 Ge/W.
Yanwen Su, Hao Li 0098, Kaiyue Li, Ang Hu, Zhichen Yang, Luxin Yan, Dongsheng Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 Infrared UAV Target Tracking With Dynamic Feature Refinement and Global Contextual Attention Knowledge Distillation
abstract
Unmanned aerial vehicle (UAV) target tracking based on thermal infrared imaging has been one of the most important sensing technologies in anti-UAV applications. However, the infrared UAV targets often exhibit weak features and complex backgrounds, posing significant challenges to accurate tracking. To address these problems, we introduce SiamDFF, a novel dynamic feature fusion Siamese network that integrates feature enhancement and global contextual attention knowledge distillation for infrared UAV target (IRUT) tracking. The SiamDFF incorporates a selective target enhancement network (STEN), a dynamic spatial feature aggregation module (DSFAM), and a dynamic channel feature aggregation module (DCFAM). The STEN employs intensity-aware multi-head cross-attention to adaptively enhance important regions for both template and search branches. The DSFAM enhances multi-scale UAV target features by integrating local details with global features, utilizing spatial attention guidance within the search frame. The DCFAM effectively integrates the mixed template generated from STEN in the template branch and original template, avoiding excessive background interference with the template and thereby enhancing the emphasis on UAV target region features within the search frame. Furthermore, to enhance the feature extraction capabilities of the network for IRUT without adding extra computational burden, we propose a novel tracking-specific target-aware contextual attention knowledge distiller (TCAKD). It transfers the target prior from the teacher network to the student model, significantly improving the student network's focus on informative regions at each hierarchical level of the backbone network. Extensive experiments on real infrared UAV datasets demonstrate that the proposed approach outperforms state-of-the-art target trackers under complex backgrounds while achieving a real-time tracking speed.
Houzhang Fang, Chenxing Wu, Xiaolin Wang 0006, Yi Chang 0002, Luxin Yan
IEEE Trans. Multim.8
2026 Hierarchical Semantic-Visual Fusion of Visible and Near-Infrared Images for Long-Range Haze Removal
abstract
While image dehazing has advanced substantially in the past decade, most efforts have focused on short-range scenarios, leaving long-range haze removal under-explored. As distance increases, intensified scattering leads to severe haze and signal loss, making it impractical to recover distant details solely from visible images. Near- infrared, with superior fog penetration, offers critical complementary cues through multimodal fusion. However, existing methods focus on content integration while often neglecting haze embedded in visible images, leading to results with residual haze. In this work, we argue that the infrared and visible modalities not only provide complementary low-level visual features, but also share high-level semantic consistency. Motivated by this, we propose a Hierarchical Semantic-Visual Fusion (HSVF) framework, comprising a semantic stream to reconstruct haze-free scenes and a visual stream to incorporate structural details from the near- infrared modality. The semantic stream first acquires haze-robust semantic prediction by aligning modality-invariant intrinsic representations. Then the shared semantics act as strong priors to restore clear and high-contrast distant scenes under severe haze degradation. In parallel, the visual stream focuses on recovering lost structural details from near- infrared by fusing complementary cues from both visible and near- infrared images. Through the cooperation of dual streams, HSVF produces results that exhibit both high-contrast scenes and rich texture details. Moreover, we introduce a novel pixel-aligned visible- infrared haze dataset with semantic labels to facilitate benchmarking. Extensive experiments demonstrate the superiority of our method over state-of-the-art approaches in real-world long-range haze removal.
Yi Li 0033, Xiaoxiong Wang, Yi Chang 0002, Luxin Yan
IEEE Trans. Multim.6
2025 DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes
abstract
Vision-centric autonomous driving systems require diverse data for robust training and evaluation, which can be augmented by manipulating object positions and appearances within existing scene captures. While recent advancements in diffusion models have shown promise in video editing, their application to object manipulation in driving scenarios remains challenging due to imprecise positional control and difficulties in preserving high-fidelity object appearances. To address these challenges in position and appearance control, we introduce DriveEditor, the first diffusion-based framework for object editing in driving videos. DriveEditor offers a unified framework for comprehensive object editing operations, including repositioning, replacement, deletion, and insertion. These diverse manipulations are all achieved through a shared set of varying inputs, processed by identical position control and appearance maintenance modules. The position control module projects the given 3D bounding box while preserving depth information and hierarchically injects it into the diffusion process, enabling precise control over object position and orientation. The appearance maintenance module preserves consistent attributes with a single reference image by employing a three-tiered approach: low-level detail preservation, high-level semantic maintenance, and the integration of 3D priors from a novel view synthesis model. Extensive qualitative and quantitative evaluations on the nuScenes dataset demonstrate DriveEditor's exceptional fidelity and controllability in generating diverse driving scene edits, as well as its remarkable ability to facilitate downstream tasks.
Yiyuan Liang 0001, Zhiying Yan, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
AAAI5
2025 Detection-Friendly Nonuniformity Correction: A Union Framework for Infrared UAV Target Detection
abstract
Infrared unmanned aerial vehicle (UAV) images captured using thermal detectors are often affected by temperature-dependent low-frequency nonuniformity, which significantly reduces the contrast of the images. Detecting UAV targets under nonuniform conditions is crucial in UAV surveillance applications. Existing methods typically treat infrared nonuniformity correction (NUC) as a preprocessing step for detection, which leads to suboptimal performance. Balancing the two tasks while enhancing detection-beneficial information remains challenging. In this paper, we present a detection-friendly union framework, termed UniCD, that simultaneously addresses both infrared NUC and UAV target detection tasks in an end-to-end manner. We first model NUC as a small number of parameter estimation problem jointly driven by priors and data to generate detection-conducive images. Then, we incorporate a new auxiliary loss with target mask supervision into the backbone of the infrared UAV target detection network to strengthen target features while suppressing the background. To better balance correction and detection, we introduce a detection-guided self-supervised loss to reduce feature discrepancies between the two tasks, thereby enhancing detection robustness to varying nonuniformity levels. Additionally, we construct a new benchmark composed of 50,000 infrared images in various nonuniformity types, multi-scale UAV targets and rich backgrounds with target annotations, called IRBFD. Extensive experiments on IRBFD demonstrate that our UniCD is a robust union framework for NUC and UAV target detection while achieving real-time processing capabilities. Dataset can be available at https://github.com/IVPLaboratory/UniCD.
Houzhang Fang, Xiaolin Wang 0006, Zengyang Li, Lu Wang 0014, Qingshan Li, Yi Chang 0002, Luxin Yan
CVPR7
2025 TimeTracker: Event-based Continuous Point Tracking for Video Frame Interpolation with Non-linear Motion
abstract
Video frame interpolation (VFI) that leverages the bio-inspired event cameras as guidance has recently shown better performance and memory efficiency than the frame-based methods, thanks to the event cameras’ advantages, such as high temporal resolution. A hurdle for event-based VFI is how to effectively deal with non-linear motion, caused by the dynamic changes in motion direction and speed within the scene. Existing methods either use events to estimate sparse optical flow or fuse events with image features to estimate dense optical flow. Unfortunately, motion errors often degrade the VFI quality as the continuous motion cues from events do not align with the dense spatial information of images in the temporal dimension. In this paper, we find that object motion is continuous in space, tracking local regions over continuous time enables more accurate identification of spatiotemporal feature correlations. In light of this, we propose a novel continuous point tracking-based VFI framework, named TimeTracker. Specifically, we first design a Scene-Aware Region Segmentation (SARS) module to divide the scene into similar patches. Then, a Continuous Trajectory guided Motion Estimation (CTME) module is proposed to track the continuous motion trajectory of each patch through events. Finally, intermediate frames at any given time are generated through global motion optimization and frame refinement. Moreover, we collect a real-world dataset that features fast non-linear motion. Extensive experiments show that our method outperforms prior arts in both motion estimation and frame interpolation quality.
Haoyue Liu 0001, Yi Chang 0002, Hanyu Zhou, Haozhi Zhao, Luxin Yan
CVPR7
2025 Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical Flow
abstract
High-dynamic scene optical flow is a challenging task, which suffers spatial blur and temporal discontinuous motion due to large displacement in frame imaging, thus deteriorating the spatiotemporal feature of optical flow. Typically, existing methods mainly introduce event camera to directly fuse the spatiotemporal features between the two modalities. However, this direct fusion is ineffective, since there exists a large gap due to the heterogeneous data representation between frame and event modalities. To address this issue, we explore a common-latent space as an intermediate bridge to mitigate the modality gap. In this work, we propose a novel common spatiotemporal fusion between frame and event modalities for high-dynamic scene optical flow, including visual boundary localization and motion correlation fusion. Specifically, in visual boundary localization, we figure out that frame and event share the similar spatiotemporal gradients, whose similarity distribution is consistent with the extracted boundary distribution. This motivates us to design the common spatiotemporal gradient to constrain the reference boundary localization. In motion correlation fusion, we discover that the frame-based motion possesses spatially dense but temporally discontinuous correlation, while the event-based motion has spatially sparse but temporally continuous correlation. This inspires us to use the reference boundary to guide the complementary motion knowledge fusion between the two modalities. Moreover, common spatiotemporal fusion can not only relieve the cross-modal feature discrepancy, but also make the fusion process interpretable for dense and continuous optical flow. Extensive experiments have been performed to verify the superiority of the proposed method.
Hanyu Zhou, Haoyue Liu 0001, Yuxing Duan, Yi Chang 0002, Luxin Yan
CVPR6
2025 STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
abstract
High-dynamic scene reconstruction aims to represent static background with rigid spatial features and dynamic objects with deformed continuous spatiotemporal features. Typically, existing methods adopt unified representation model (e.g., Gaussian) to directly match the spatiotemporal features of dynamic scene from frame camera. However, this unified paradigm fails in the potential discontinuous temporal features of objects due to frame imaging and the heterogeneous spatial features between background and objects. To address this issue, we disentangle the spatiotemporal features into various latent representations to alleviate the spatiotemporal mismatching between background and objects. In this work, we introduce event camera to compensate for frame camera, and propose a spatiotemporal-disentangled Gaussian splatting framework for high-dynamic scene reconstruction. As for dynamic scene, we figure out that background and objects have appearance discrepancy in frame-based spatial features and motion discrepancy in event-based temporal features, which motivates us to distinguish the spatiotemporal features between background and objects via clustering. As for dynamic object, we discover that Gaussian representations and event data share the consistent spatiotemporal characteristic, which could serve as a prior to guide the spatiotemporal disentanglement of object Gaussians. Within Gaussian splatting framework, the cumulative scene-object disentanglement can improve the spatiotemporal discrimination between background and objects to render the time-continuous dynamic scene. Extensive experiments have been performed to verify the superiority of the proposed method.
Hanyu Zhou, Haoyue Liu 0001, Yuxing Duan, Luxin Yan, Gim Hee Lee
ICCV5
2025 High-dimension Prototype is a Better Incremental Object Detection Learner
abstract
Incremental object detection (IOD), surpassing simple classification, requires the simultaneous overcoming of catastrophic forgetting in both recognition and localization tasks, primarily due to the significantly higher feature space complexity. Integrating Knowledge Distillation (KD) would mitigate the occurrence of catastrophic forgetting. However, the challenge of knowledge shift caused by invisible previous task data hampers existing KD-based methods, leading to limited improvements in IOD performance. This paper aims to alleviate knowledge shift by enhancing the accuracy and granularity in describing complex high-dimensional feature spaces. To this end, we put forth a novel higher-dimension-prototype learning approach for KD-based IOD, enabling a more flexible, accurate, and fine-grained representation of feature distributions without the need to retain any previous task data. Existing prototype learning methods calculate feature centroids or statistical Gaussian distributions as prototypes, disregarding actual irregular distribution information or leading to inter-class feature overlap, which is not directly applicable to the more difficult task of IOD with complex feature space. To address the above issue, we propose a Gaussian Mixture Distribution-based Prototype (GMDP), which explicitly models the distribution relationships of different classes by directly measuring the likelihood of embedding from new and old models into class distribution prototypes in a higher dimension manner. Specifically, GMDP dynamically adapts the component weights and corresponding means/variances of class distribution prototypes to represent both intra-class and inter-class variability more accurately. Progressing into a new task, GMDP constrains the distance between the distribution of new and previous task classes, minimizing overlap with existing classes and thus striking a balance between stability and adaptability. GMDP can be readily integrated into existing IOD methods to enhance performance further. Extensive experiments on the PASCAL VOC and MS-COCO show that our method consistently exceeds four baselines by a large margin and significantly outperforms other SOTA results under various settings.
Tianming Zhao 0003, Tao Zhang 0147, Guodong Wang 0001, Luxin Yan, Sheng Zhong 0001, Jiahuan Zhou, Xu Zou 0002
ICLR6
2025 Divide-And-Conquer: Dual-Hierarchical Optimization for Semantic 4D Gaussian Spatting
abstract
Semantic 4D Gaussians can be used for reconstructing and understanding dynamic scenes, with temporal variations than static scenes. Directly applying static methods to understand dynamic scenes will fail to capture the temporal features. Few works focus on dynamic scene understanding based on Gaussian Splatting, since once the same update strategy is employed for both dynamic and static parts, regardless of the distinction and interaction between Gaussians, significant artifacts and noise appear. We propose Dual-Hierarchical Optimization (DHO), which consists of Hierarchical Gaussian Flow and Hierarchical Gaussian Guidance in a divide-and-conquer manner. The former implements effective division of static and dynamic rendering and features. The latter helps to mitigate the issue of dynamic foreground rendering distortion in textured complex scenes. Extensive experiments show that our method consistently outperforms the baselines on both synthetic and real-world datasets, and supports various downstream tasks. Project Page: https://sweety-yan.github.io/DHO/
Zhiying Yan, Yiyuan Liang 0001, Shilv Cai, Tao Zhang 0147, Sheng Zhong 0001, Luxin Yan, Xu Zou 0002
ICME6
2025 Injecting Frame-Event Complementary Fusion into Diffusion for Optical Flow in Challenging Scenes
abstract
Optical flow estimation has achieved promising results in conventional scenes but faces challenges in high-speed and low-light scenes, which suffer from motion blur and insufficient illumination. These conditions lead to weakened texture and amplified noise and deteriorate the appearance saturation and boundary completeness of frame cameras, which are necessary for motion feature matching. In degraded scenes, the frame camera provides dense appearance saturation but sparse boundary completeness due to its long imaging time and low dynamic range. In contrast, the event camera offers sparse appearance saturation, while its short imaging time and high dynamic range gives rise to dense boundary completeness. Traditionally, existing methods utilize feature fusion or domain adaptation to introduce event to improve boundary completeness. However, the appearance features are still deteriorated, which severely affects the mostly adopted discriminative models that learn the mapping from visual features to motion fields and generative models that generate motion fields based on given visual features. So we introduce diffusion models that learn the mapping from noising flow to clear flow, which is not affected by the deteriorated visual features. Therefore, we propose a novel optical flow estimation framework Diff-ABFlow based on diffusion models with frame-event appearance-boundary fusion. Inspired by the appearance-boundary complementarity of frame and event, we propose an Attention-Guided Appearance-Boundary Fusion module to fuse frame and event. Based on diffusion models, we propose a Multi-Condition Iterative Denoising Decoder. Our proposed method can effectively utilize the respective advantages of frame and event, and shows great robustness to degraded input. In addition, we propose a dual-modal optical flow dataset for generalization experiments. Extensive experiments have verified the superiority of our proposed method. The code is released at <https://github.com/Haonan-Wang-aurora/Diff-ABFlow>.
Hanyu Zhou, Haoyue Liu 0001, Luxin Yan
NeurIPS4
2025 UniAnimate: taming unified video diffusion models for consistent human image animation
Xiang Wang 0012, Shiwei Zhang 0001, Changxin Gao, Xiaoqiang Zhou, Yingya Zhang, Luxin Yan, Nong Sang
Sci. China Inf. Sci.7
2025 NER-Net+: Seeing Motion at Nighttime With an Event Camera
abstract
We focus on a very challenging task: imaging at nighttime dynamic scenes. Conventional RGB cameras struggle with the trade-off between long exposure for low-light imaging and short exposure for capturing dynamic scenes. Event cameras react to dynamic changes, with their high temporal resolution (microsecond) and dynamic range (120 dB), and thus offer a promising alternative. However, existing methods are mostly based on simulated datasets due to the lack of paired event-clean image data for nighttime conditions, where the domain gap leads to performance limitations in real-world scenarios. Moreover, most existing event reconstruction methods are tailored for daytime data, overlooking issues unique to low-light events at night, such as strong noise, temporal trailing, and spatial non-uniformity, resulting in unsatisfactory reconstruction results. To address these challenges, we construct the first real paired low-light event dataset (RLED) through a co-axial imaging system, comprising 80,400 spatially and temporally aligned image GTs and low-light events, which provides a unified training and evaluation dataset for existing methods. We further conduct a comprehensive analysis of the causes and characteristics of strong noise, temporal trailing, and spatial non-uniformity in nighttime events, and propose a nighttime event reconstruction network (NER-Net+). It includes a learnable event timestamps calibration module (LETC) to correct the temporal trailing events and a non-stationary spatio-temporal information enhancement module (NSIE) to suppress sensor noise and spatial non-uniformity. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art methods in visual quality and generalization on real-world nighttime datasets.
Haoyue Liu 0001, Shihan Peng, Yi Chang 0002, Hanyu Zhou, Yuxing Duan, Lin Zhu 0012, Yonghong Tian 0001, Luxin Yan
IEEE Trans. Pattern Anal. Mach. Intell.9
2025 Adverse Weather Optical Flow: Cumulative Homogeneous-Heterogeneous Adaptation
abstract
Optical flow has made great progress in clean scenes, while suffers degradation under adverse weather due to the violation of the brightness constancy and gradient continuity assumptions of optical flow. Typically, existing methods mainly adopt domain adaptation to transfer motion knowledge from clean to degraded domain through one-stage adaptation. However, this direct adaptation is ineffective, since there exists a large gap due to adverse weather and scene style between clean and real degraded domains. Moreover, even within the degraded domain itself, static weather (e.g., fog) and dynamic weather (e.g., rain) have different impacts on optical flow. To address above issues, we explore synthetic degraded domain as an intermediate bridge between clean and real degraded domains, and propose a cumulative homogeneous-heterogeneous adaptation framework for real adverse weather optical flow. Specifically, for clean-degraded transfer, our key insight is that static weather possesses the depth-association homogeneous feature which does not change the intrinsic motion of the scene, while dynamic weather additionally introduces the heterogeneous feature which results in a significant boundary discrepancy in warp errors between clean and degraded domains. For synthetic-real transfer, we figure out that cost volume correlation shares a similar statistical histogram between synthetic and real degraded domains, benefiting to holistically aligning the homogeneous correlation distribution for synthetic-real knowledge distillation. Under this unified framework, the proposed method can progressively and explicitly transfer knowledge from clean scenes to real adverse weather. In addition, we further collect a real adverse weather dataset with manually annotated optical flow labels and perform extensive experiments to verify the superiority of the proposed method.
Hanyu Zhou, Yi Chang 0002, Zhiwei Shi 0001, Wending Yan, Gang Chen 0023, Yonghong Tian 0001, Luxin Yan
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Consistent Learning of Sparse Background Features for Infrared Small-Target Labeling
abstract
Recent years have witnessed many remarkable achievements in infrared small target detection based on deep learning. To achieve high performance in real-world applications, deep learning methods require a substantial number of accurate labels. However, infrared small target annotating is labor intensive as they are very small. Determining and annotating edge pixels demands considerable time and effort, which slows down data expansion and further research in infrared small target detection. To mitigate the issue, we propose a pseudo-label generation method named Consistent Learning of Sparse Background Feature (CLSBF). This approach models the generation of infrared small target pseudo-labels as a domain transformation from local target maps to background ones. It can relax the supervision requirements of deep learning methods, transitioning from absolute pixel-level supervision to point supervision. This method employs an unsupervised approach primarily, supplemented by semi-simulated supervision, to achieve mutual conversion of local images from different domains and obtain the final target pseudo-labels through the differences between target images and transformed background ones. Experiments show that equipped with our method, models trained with the input of a coarsely accurate center label can achieve performance up to 99.94% compared to models trained with official accurate labels. Furthermore, when the official labels are not accurate enough, models trained with pseudo-labels generated by CLSBF consistently show a performance improvement of 1.11% to 4.09% when they are evaluated based on re-labeled bounding boxes. Extensive experiments demonstrate the reliability of CLSBF in generating pseudo-labels and its potential to alleviate the labor-intensive process of manual labeling significantly. We have released an infrared small target labeling tool with CLSBF as an assistant at https://github.com/SeaHifly/CLSBF_software.git.
Sheng Zhong 0001, Luxin Yan, Xu Zou 0002
IEEE Trans. Geosci. Remote. Sens.4
2025 Robust Point Cloud Registration via Patch Matching
abstract
We study the problem of exacting accurate correspondence pairs for point cloud registration. The existing correspondence methods focus on constructing point descriptors and then extracting correspondence point pairs. However, this process encounters two main issues: 1) point features are unstable and susceptible to noise, leading to a low inlier ratio (IR) for correspondence pairs and 2) the positional deviation of correspondence point pair results in accuracy errors when computing rigid transformations. To address these issues, we propose a robust point cloud registration framework based on patch matching, achieving high positional accuracy and high inlier-rate prediction of correspondence pairs. Specifically, we design a dual-branch point cloud registration network, with one branch dedicated to patch matching and the other branch to predicting the patch anchor, i.e., the coordinates used for patch matching. For patch matching, we integrate the topology of patches into the attention mechanism and adopt a multilevel patch-matching strategy to enhance the matching success rate. For coordinate prediction, we introduce graph convolutional network (GCN) and cross-attention mechanisms to explore local similar points through information interaction and feature correlation of patch pairs. Thanks to the stability of patch descriptors, our method demonstrates higher robustness compared to existing correspondence methods. Extensive experiments conducted on indoor, outdoor, synthetic, and deformable benchmarks validate the superiority of our method. Additionally, our method achieves certain effectiveness in cross-source point clouds.
Tianming Zhao 0003, Tian Tian 0006, Xu Zou 0002, Luxin Yan, Sheng Zhong 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Hunt Camouflaged Objects via Revealing Mutation Regions
abstract
Due to the high similarity between hidden objects and the surrounding background, camouflaged object detection (COD) remains a challenge. While many recently proposed methods have shown remarkable performance, most of them begin object perception by indiscriminately considering every pixel of the image. However, these early-stage region-insensitive perception methods still struggle to resist background interference, potentially missing subtle pixel changes by not prioritizing potential camouflaged areas initially. Fortunately, we reveal that the availability of an accurate mutation map can significantly enhance camouflaged discrimination ability. To this end, we propose MRNet (Mutation Region Network). MRNet initially generates a mutation map that identifies potential mutation regions exhibiting subtle pixel changes. The generation method involves amplifying and differing pixel changes based on the position and corresponding values of pixels. Subsequently, the selective expansion search operation utilizes the mutation map to extract the mapped graph, effectively reducing interference from background pixels that are distant from the mutation regions. Finally, decoding the mapped graph generates precise masks. Furthermore, we have created the largest test dataset with known categories to advance community research. Extensive experiments conducted on three widely used datasets and our proposed dataset show that MRNet surpasses other methods with superior performance. Source code is publicly available athttps://github.com/XinyueZhangHust/MRNet
Xinyue Zhang 0009, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
IEEE Trans. Inf. Forensics Secur.3
2024 Make Lossy Compression Meaningful for Low-Light Images
abstract
Low-light images frequently occur due to unavoidable environmental influences or technical limitations, such as insufficient lighting or limited exposure time. To achieve better visibility for visual perception, low-light image enhancement is usually adopted. Besides, lossy image compression is vital for meeting the requirements of storage and transmission in computer vision applications. To touch the above two practical demands, current solutions can be categorized into two sequential manners: ``Compress before Enhance (CbE)'' or ``Enhance before Compress (EbC)''. However, both of them are not suitable since: (1) Error accumulation in the individual models plagues sequential solutions. Especially, once low-light images are compressed by existing general lossy image compression approaches, useful information (e.g., texture details) would be lost resulting in a dramatic performance decrease in low-light image enhancement. (2) Due to the intermediate process, the sequential solution introduces an additional burden resulting in low efficiency. We propose a novel joint solution to simultaneously achieve a high compression rate and good enhancement performance for low-light images with much lower computational cost and fewer model parameters. We design an end-to-end trainable architecture, which includes the main enhancement branch and the signal-to-noise ratio (SNR) aware branch. Experimental results show that our proposed joint solution achieves a significant improvement over different combinations of existing state-of-the-art sequential ``Compress before Enhance'' or ``Enhance before Compress'' solutions for low-light images, which would make lossy low-light image compression more meaningful. The project is publicly available at: https://github.com/CaiShilv/Joint-IC-LL.
Shilv Cai, Sheng Zhong 0001, Luxin Yan, Jiahuan Zhou, Xu Zou 0002
AAAI4
2024 Seeing Motion at Nighttime with an Event Camera
abstract
We focus on a very challenging task: imaging at night-time dynamic scenes. Most previous methods rely on the low-light enhancement of a conventional RGB camera. However, they would inevitably face a dilemma between the long exposure time of nighttime and the motion blur of dynamic scenes. Event cameras react to dynamic changes with higher temporal resolution (microsecond) and higher dynamic range (120dB), offering an alternative solution. In this work, we present a novel nighttime dynamic imaging method with an event camera. Specifically, we discover that the event at nighttime exhibits temporal trailing characteristics and spatial non-stationary distribution. Consequently, we propose a nighttime event reconstruction network (NER-Net) which mainly includes a learnable event timestamps calibration module (LETC) to align the temporal trailing events and a non-uniform illumination aware module (NIAM) to stabilize the spatiotemporal distribution of events. Moreover, we construct a paired real low-light event dataset (RLED) through a co-axial imaging system, including 64,200 spatially and temporally aligned image GTs and low-light events. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art methods in terms of visual quality and generalization ability on real-world nighttime datasets. The project are available at: https://github.com/Liu-haoyue/NER-Net.
Haoyue Liu 0001, Shihan Peng, Lin Zhu 0012, Yi Chang 0002, Hanyu Zhou, Luxin Yan
CVPR6
2024 SNIDA: Unlocking Few-Shot Object Detection with Non-Linear Semantic Decoupling Augmentation
abstract
Once only a few-shot annotated samples are available, the performance of learning-based object detection would be heavily dropped. Many few-shot object detection (FSOD) methods have been proposed to tackle this issue by adopting image-level augmentations in linear manners. Nevertheless, those handcrafted enhancements often suf-fer from limited diversity and lack of semantic awareness, resulting in unsatisfactory performance. To this end, we propose a Semantic-guided Nonlinear Instance-level Data Augmentation method (SNIDA) for FSOD by decoupling the foreground and background to increase their diversities respectively. We design a semantic awareness enhancement strategy to separate objects from backgrounds. Concretely, masks of instances are extracted by an unsupervised semantic segmentation module. Then the diversity of samples would be improved by fusing instances into different backgrounds. Considering the shortcomings of augmenting images in a limited transformation space of existing traditional data augmentation methods, we introduce an object reconstruction enhancement module. The aim of this module is to generate sufficient diversity and nonlinear training data at the instance level through a semantic-guided masked autoencoder. In this way, the potential of data can be fully exploited in various object detection scenarios. Extensive experiments on PASCAL VOC and MS-COCO demonstrate that the proposed method outperforms base-lines by a large margin and achieves new state-of-the-art results under different shot settings.
Xu Zou 0002, Luxin Yan, Sheng Zhong 0001, Jiahuan Zhou
CVPR3
2024 Long-Range Turbulence Mitigation: A Large-Scale Dataset and A Coarse-to-Fine Framework
Shengqi Xu, Run Sun, Yi Chang 0002, Shuning Cao, Xueyao Xiao, Luxin Yan
ECCV (42)6
2024 Exploring the Common Appearance-Boundary Adaptation for Nighttime Optical Flow
abstract
We investigate a challenging task of nighttime optical flow, which suffers from weakened texture and amplified noise. These degradations weaken discriminative visual features, thus causing invalid motion feature matching. Typically, existing methods employ domain adaptation to transfer knowledge from auxiliary domain to nighttime domain in either input visual space or output motion space. However, this direct adaptation is ineffective, since there exists a large domain gap due to the intrinsic heterogeneous nature of the feature representations between auxiliary and nighttime domains. To overcome this issue, we explore a common-latent space as the intermediate bridge to reinforce the feature alignment between auxiliary and nighttime domains. In this work, we exploit two auxiliary daytime and event domains, and propose a novel common appearance-boundary adaptation framework for nighttime optical flow. In appearance adaptation, we employ the intrinsic image decomposition to embed the auxiliary daytime image and the nighttime image into a reflectance-aligned common space. We discover that motion distributions of the two reflectance maps are very similar, benefiting us to consistently transfer motion appearance knowledge from daytime to nighttime domain. In boundary adaptation, we theoretically derive the motion correlation formula between nighttime image and accumulated events within a spatiotemporal gradient-aligned common space. We figure out that the correlation of the two spatiotemporal gradient maps shares significant discrepancy, benefitting us to contrastively transfer boundary knowledge from event to nighttime domain. Moreover, appearance adaptation and boundary adaptation are complementary to each other, since they could jointly transfer global motion and local boundary knowledge to the nighttime domain. Extensive experiments have been performed to verify the superiority of the proposed method.
Hanyu Zhou, Yi Chang 0002, Haoyue Liu 0001, Wending Yan, Yuxing Duan, Zhiwei Shi 0001, Luxin Yan
ICLR7
2024 Powerful Lossy Compression for Noisy Images
abstract
Image compression and denoising represent fundamental challenges in image processing with many real-world applications. To address practical demands, current solutions can be categorized into two main strategies: 1) sequential method; and 2) joint method. However, sequential methods have the disadvantage of error accumulation as there is information loss between multiple individual models. Recently, the academic community began to make some attempts to tackle this problem through end-to-end joint methods. Most of them ignore that different regions of noisy images have different characteristics. To solve these problems, in this paper, our proposed signal-to-noise ratio (SNR) aware joint solution exploits local and non-local features for image compression and denoising simultaneously. We design an end-to-end trainable network, which includes the main encoder branch, the guidance branch, and the signal-to-noise ratio (SNR) aware branch. We conducted extensive experiments on both synthetic and real-world datasets, demonstrating that our joint solution outperforms existing state-of-the-art methods.
Shilv Cai, Xiaoguo Liang, Shuning Cao, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
ICME4
2024 JSTR: Joint Spatio-Temporal Reasoning for Event-based Moving Object Detection
abstract
Event-based moving object detection is a challenging task, where static background and moving object are mixed together. Typically, existing methods mainly align the background events to the same spatial coordinate system via motion compensation to distinguish the moving object. However, they neglect the potential spatial tailing effect of moving object events caused by excessive motion, which may affect the structure integrity of the extracted moving object. We discover that the moving object has a complete columnar structure in the point cloud composed of motion-compensated events along the timestamp. Motivated by this, we propose a novel joint spatio-temporal reasoning method for event-based moving object detection. Specifically, we first compensate the motion of background events using inertial measurement unit. In spatial reasoning stage, we project the compensated events into the same image coordinate, discretize the timestamp of events to obtain a time image that can reflect the motion confidence, and further segment the moving object through adaptive threshold on the time image. In temporal reasoning stage, we construct the events into a point cloud along timestamp, and use RANSAC algorithm to extract the columnar shape in the cloud for peeling off the background. Finally, we fuse the results from the two reasoning stages to extract the final moving object region. This joint spatio-temporal reasoning framework can effectively detect the moving object from motion confidence and geometric structure. Moreover, we conduct extensive experiments on various datasets to verify that the proposed method can improve the moving object detection accuracy by 13%.
Hanyu Zhou, Zhiwei Shi 0001, Hao Dong 0013, Shihan Peng, Yi Chang 0002, Luxin Yan
ICRA6
2024 Perceptual-Distortion Balanced Image Super-Resolution is a Multi-Objective Optimization Problem
abstract
Training Single-Image Super-Resolution (SISR) models using pixel-based regression losses can achieve high distortion metrics scores (e.g., PSNR and SSIM), but often results in blurry images due to insufficient recovery of high-frequency details. Conversely, using GAN or perceptual losses can produce sharp images with high perceptual metric scores (e.g., LPIPS), but may introduce artifacts and incorrect textures. Balancing these two types of losses can help achieve a trade-off between distortion and perception, but the challenge lies in tuning the loss function weights. To address this issue, we propose a novel method that incorporates Multi-Objective Optimization (MOO) into the training process of SISR models to balance perceptual quality and distortion. We conceptualize the relationship between loss weights and image quality assessment (IQA) metrics as black-box objective functions to be optimized within our Multi-Objective Bayesian Optimization Super-Resolution (MOBOSR) framework. This approach automates the hyperparameter tuning process, reduces overall computational cost, and enables the use of numerous loss functions simultaneously. Extensive experiments demonstrate that MOBOSR outperforms state-of-the-art methods in terms of both perceptual quality and distortion, significantly advancing the perception-distortion Pareto frontier. Our work points towards a new direction for future research on balancing perceptual quality and fidelity in nearly all image restoration tasks. The source code and pretrained models are available at: https://github.com/ZhuKeven/MOBOSR.
Qiwen Zhu, Shilv Cai, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
ACM Multimedia6
2024 I2C: Invertible Continuous Codec for High-Fidelity Variable-Rate Image Compression
abstract
Lossy image compression is a fundamental technology in media transmission and storage. Variable-rate approaches have recently gained much attention to avoid the usage of a set of different models for compressing images at different rates. During the media sharing, multiple re-encodings with different rates would be inevitably executed. However, existing Variational Autoencoder (VAE)-based approaches would be readily corrupted in such circumstances, resulting in the occurrence of strong artifacts and the destruction of image fidelity. Based on the theoretical findings of preserving image fidelity via invertible transformation, we aim to tackle the issue of high-fidelity fine variable-rate image compression and thus propose the Invertible Continuous Codec (I2C). We implement the I2C in a mathematical invertible manner with the core Invertible Activation Transformation (IAT) module. I2C is constructed upon a single-rate Invertible Neural Network (INN) based model and the quality level (QLevel) would be fed into the IAT to generate scaling and bias tensors. Extensive experiments demonstrate that the proposed I2C method outperforms state-of-the-art variable-rate image compression methods by a large margin, especially after multiple continuous re-encodings with different rates, while having the ability to obtain a very fine variable-rate control without any performance compromise.
Shilv Cai, Zhijun Zhang 0009, Xiangyun Zhao, Jiahuan Zhou, Yuxin Peng 0001, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Unsupervised Deraining: Where Asymmetric Contrastive Learning Meets Self-Similarity
abstract
Most existing learning-based deraining methods are supervisedly trained on synthetic rainy-clean pairs. The domain gap between the synthetic and real rain makes them less generalized to complex real rainy scenes. Moreover, the existing methods mainly utilize the property of the image or rain layers independently, while few of them have considered their mutually exclusive relationship. To solve above dilemma, we explore the intrinsic intra-similarity within each layer and inter-exclusiveness between two layers and propose an unsupervised non-local contrastive learning (NLCL) deraining method. The non-local self-similarity image patches as the positives are tightly pulled together and rain patches as the negatives are remarkably pushed away, and vice versa. On one hand, the intrinsic self-similarity knowledge within positive/negative samples of each layer benefits us to discover more compact representation; on the other hand, the mutually exclusive property between the two layers enriches the discriminative decomposition. Thus, the internal self-similarity within each layer (similarity) and the external exclusive relationship of the two layers (dissimilarity) serving as a generic image prior jointly facilitate us to unsupervisedly differentiate the rain from clean image. We further discover that the intrinsic dimension of the non-local image patches is generally higher than that of the rain patches. This insight motivates us to design an asymmetric contrastive loss that precisely models the compactness discrepancy of the two layers, thereby improving the discriminative decomposition. In addition, recognizing the limited quality of existing real rain datasets, which are often small-scale or obtained from the internet, we collect a large-scale real dataset under various rainy weathers that contains high-resolution rainy images. Extensive experiments conducted on different real rainy datasets demonstrate that the proposed method obtains state-of-the-art performance in real deraining.
Yi Chang 0002, Yun Guo, Yuntong Ye, Changfeng Yu, Lin Zhu 0012, Xi-Le Zhao, Luxin Yan, Yonghong Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 SCINet: Spatial and Contrast Interactive Super-Resolution Assisted Infrared UAV Target Detection
abstract
Unmanned aerial vehicle (UAV) detection based on thermal infrared imaging has been one of the most important sensing technologies in the anti-UAV system. However, the technical limitations and long-range detection of thermal sensors often lead to acquiring low-resolution (LR) infrared images, thereby bringing great challenges for the subsequent target detection task. In this article, we propose a novel spatial and contrast interactive super-resolution network (SCINet) for assisting infrared UAV target detection. The network consists of two main subnetworks: a spatial enhancement branch (SEB) and a contrast enhancement branch (CEB). The SEB embeds the lightweight convolution module and attention mechanism to highlight the spatial structure detail features of infrared UAV targets. The proposed CEB incorporates the center-oriented contrast-aware module and multibranch collapsible module, which can provide local contrast priors to reconstruct the super-resolved UAV target image. The spatial features of the intermediate layers from the SEB are integrated into the CEB as a rich gradient prior. Besides, the output features of the CEB are aggregated into those of the SEB for further supplementing the contrast of spatial features in return. Dual-branch feature interaction (DBFI) of the SEB and CEB can further enhance the spatial details and target saliency of the targets. In addition, we also introduce an infrared UAV detection network via a new dual-dimensional feature calibration module (DFCM) for boosting the detection performance. Extensive experiments demonstrate that the SCINet outperforms the state-of-the-art (SOTA) SR methods on real infrared UAV sequences and improves the detection performance of infrared small UAV targets. The code is available athttps://github.com/IVPLaboratory/SCINet.
Houzhang Fang, Lan Ding, Xiaolin Wang 0006, Yi Chang 0002, Luxin Yan, Li Liu 0050, Jinrui Fang
IEEE Trans. Geosci. Remote. Sens.5
2024 Online Infrared UAV Target Tracking With Enhanced Context-Awareness and Pixel-Wise Attention Modulation
abstract
Unmanned aerial vehicles (UAVs) have been popular in many commercial and industrial applications, but they also pose great threats to urban safety and aerial security. Intelligent UAV surveillance based on thermal infrared (TIR) imaging has attracted increasing attention for its long-range monitoring ability in both day and night scenarios. However, weak UAV target features, dynamically changing UAV states, and complex background interferences present serious challenges to the accurate tracking of UAVs. To tackle these problems, we propose a novel online multiscale infrared UAV target (IRUT) tracking network (SiamCAP) incorporating enhanced context feature awareness and pixel-wise attention modulation. We first introduce a novel contrast-enhanced multiscale online re-parameterization block (CMORB) to effectively extract contrast difference intensity information between the target and the background, and transform it into a single branch for both training and inference without introducing computational overhead. Then, we construct a feature fusion modulation module (FFMM) to guide cross-layer feature aggregation. It uses low-level attention to highlight the UAV target feature in the deep layer with the novel full spatial resolution channel attention (FSRCA), which calculates pixel-wise importance without dimensionality compression. Finally, we propose a cross-attention-based updatable feature interaction module (CUFIM) to model the correlation between online updating multitemplate and search frame, which improves the model’s robustness to changes in the state of UAVs and complex backgrounds. Extensive experiments on real infrared UAV datasets demonstrate that the proposed approach outperforms the state-of-the-art (SOTA) target trackers under complex backgrounds while achieving a real-time tracking speed.
Houzhang Fang, Chenxing Wu, Xiaolin Wang 0006, Yi Chang 0002, Luxin Yan
IEEE Trans. Geosci. Remote. Sens.6
2024 Direction and Residual Awareness Curriculum Learning Network for Rain Streaks Removal
abstract
Single-image rain streaks' removal has attracted great attention in recent years. However, due to the highly visual similarity between the rain streaks and the line pattern image edges, the over-smoothing of image edges or residual rain streaks' phenomenon may unexpectedly occur in the deraining results. To overcome this problem, we propose a direction and residual awareness network within the curriculum learning paradigm for the rain streaks' removal. Specifically, we present a statistical analysis of the rain streaks on large-scale real rainy images and figure out that rain streaks in local patches possess principal directionality. This motivates us to design a direction-aware network for rain streaks' modeling, in which the principal directionality property endows us with the discriminative representation ability of better differing rain streaks from image edges. On the other hand, for image modeling, we are motivated by the iterative regularization in classical image processing and unfold it into a novel residual-aware block (RAB) to explicitly model the relationship between the image and the residual. The RAB adaptively learns balance parameters to selectively emphasize informative image features and better suppress the rain streaks. Finally, we formulate the rain streaks' removal problem into the curriculum learning paradigm which progressively learns the directionality of the rain streaks, rain streaks' appearance, and the image layer in a coarse-to-fine, easy-to-hard guidance manner. Solid experiments on extensive simulated and real benchmarks demonstrate the visual and quantitative improvement of the proposed method over the state-of-the-art methods.
Yi Chang 0002, Meiya Chen, Changfeng Yu, Yi Li 0033, Luxin Yan
IEEE Trans. Neural Networks Learn. Syst.6
2024 Learning Oriented Object Detection via Naive Geometric Computing
abstract
Detecting oriented objects along with estimating their rotation information is one crucial step for image analysis, especially for remote sensing images. Despite that many methods proposed recently have achieved remarkable performance, most of them directly learn to predict object directions under the supervision of only one (e.g., the rotation angle) or a few (e.g., several coordinates) groundtruth (GT) values individually. Oriented object detection would be more accurate and robust if extra constraints, with respect to proposal and rotation information regression, are adopted for joint supervision during training. To this end, we propose a mechanism that simultaneously learns the regression of horizontal proposals, oriented proposals, and rotation angles of objects in a consistent manner, via naive geometric computing, as one additional steady constraint. An oriented center prior guided label assignment strategy is proposed for further enhancing the quality of proposals, yielding better performance. Extensive experiments on six datasets demonstrate the model equipped with our idea significantly outperforms the baseline by a large margin and several new state-of-the-art results are achieved without any extra computational burden during inference. Our proposed idea is simple and intuitive that can be readily implemented. Source codes are publicly available at: https://github.com/wangWilson/CGCDet.git.
Zhijun Zhang 0009, Guodong Wang 0001, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
IEEE Trans. Neural Networks Learn. Syst.6
2023 Unsupervised Hierarchical Domain Adaptation for Adverse Weather Optical Flow
abstract
Optical flow estimation has made great progress, but usually suffers from degradation under adverse weather. Although semi/full-supervised methods have made good attempts, the domain shift between the synthetic and real adverse weather images would deteriorate their performance. To alleviate this issue, our start point is to unsupervisedly transfer the knowledge from source clean domain to target degraded domain. Our key insight is that adverse weather does not change the intrinsic optical flow of the scene, but causes a significant difference for the warp error between clean and degraded images. In this work, we propose the first unsupervised framework for adverse weather optical flow via hierarchical motion-boundary adaptation. Specifically, we first employ image translation to construct the transformation relationship between clean and degraded domains. In motion adaptation, we utilize the flow consistency knowledge to align the cross-domain optical flows into a motion-invariance common space, where the optical flow from clean weather is used as the guidance-knowledge to obtain a preliminary optical flow for adverse weather. Furthermore, we leverage the warp error inconsistency which measures the motion misalignment of the boundary between the clean and degraded domains, and propose a joint intra- and inter-scene boundary contrastive adaptation to refine the motion boundary. The hierarchical motion and boundary adaptation jointly promotes optical flow in a unified framework. Extensive quantitative and qualitative experiments have been performed to verify the superiority of the proposed method.
Hanyu Zhou, Yi Chang 0002, Gang Chen 0023, Luxin Yan
AAAI4
2023 Unsupervised Cumulative Domain Adaptation for Foggy Scene Optical Flow
abstract
Optical flow has achieved great success under clean scenes, but suffers from restricted performance under foggy scenes. To bridge the clean-to-foggy domain gap, the existing methods typically adopt the domain adaptation to transfer the motion knowledge from clean to synthetic foggy domain. However, these methods unexpectedly neglect the synthetic-to-real domain gap, and thus are erroneous when applied to real-world scenes. To handle the practical optical flow under real foggy scenes, in this work, we propose a novel unsupervised cumulative domain adaptation optical flow (UCDA-Flow) framework: depth-association motion adaptation and correlation-alignment motion adaptation. Specifically, we discover that depth is a key ingredient to in-fluence the optical flow: the deeper depth, the inferior optical flow, which motivates us to design a depth-association motion adaptation module to bridge the clean-to-foggy domain gap. Moreover, we figure out that the cost volume correlation shares similar distribution of the synthetic and real foggy images, which enlightens us to devise a correlation-alignment motion adaptation module to distill motion knowledge of the synthetic foggy domain to the real foggy domain. Note that synthetic fog is designed as the intermediate domain. Under this unified framework, the proposed cumulative adaptation progressively transfers knowledge from clean scenes to real foggy scenes. Extensive experiments have been performed to verify the superiority of the proposed method.
Hanyu Zhou, Yi Chang 0002, Wending Yan, Luxin Yan
CVPR4
2023 From Sky to the Ground: A Large-scale Benchmark and Simple Baseline Towards Real Rain Removal
abstract
Learning-based image deraining methods have made great progress. However, the lack of large-scale high-quality paired training samples is the main bottleneck to hamper the real image deraining (RID). To address this dilemma and advance RID, we construct a Large-scale High-quality Paired real rain benchmark (LHP-Rain), including 3000 video sequences with 1 million high-resolution (1920*1080) frame pairs. The advantages of the proposed dataset over the existing ones are three-fold: rain with higher-diversity and larger-scale, image with higher-resolution and higher-quality ground-truth. Specifically, the real rains in LHP-Rain not only contain the classical rain streak/veiling/occlusion in the sky, but also the splashing on the ground overlooked by deraining community. Moreover, we propose a novel robust low-rank tensor recovery model to generate the GT with better separating the static background from the dynamic rain. In addition, we design a simple transformer-based single image deraining baseline, which simultaneously utilize the self-attention and cross-layer attention within the image and rain layer with discriminative feature representation. Extensive experiments verify the superiority of the proposed dataset and deraining method over state-of-the-art.
Yun Guo, Xueyao Xiao, Yi Chang 0002, Shumin Deng, Luxin Yan
ICCV5
2023 Both Diverse and Realism Matter: Physical Attribute and Style Alignment for Rainy Image Generation
abstract
Although considerable progress has been made in image deraining under synthetic data, real rain removal is still a tough problem due to the huge domain gap between synthetic and real data. Besides, difficulties in collecting and labeling diverse real rain images hinder the progress of this field. Consequently, we attempt to promote real rain removal from rain image generation (RIG) perspective. Existing RIG methods mainly focus on diversity but miss realistic, or the realistic but neglect diversity of the generation. To solve this dilemma, we propose a physical alignment and controllable generation network (PCGNet) for diverse and realistic rain generation. Our key idea is to simultaneously utilize the controllability of attributes from synthetic and the realism of appearance from real data. Specifically, we devise a unified framework to disentangle background, rain attributes, and appearance style from synthetic and real data. Then we collaboratively align the factors with a novel semi-supervised weight moving strategy for attribute, an explicit distribution modeling method for real rain style. Furthermore, we pack these aligned factors into the generation model, achieving physical controllable mapping from the attributes to real rain with image-level and attribute-level consistency loss. Extensive experiments show that PCGNet can effectively generate appealing rainy results, which significantly improve the performance under synthetic and real scenes for all existing deraining methods.
Changfeng Yu, Shiming Chen 0002, Yi Chang 0002, Yibing Song, Luxin Yan
ICCV5
2023 DANet: Multi-scale UAV Target Detection with Dynamic Feature Perception and Scale-aware Knowledge Distillation
abstract
Multi-scale infrared unmanned aerial vehicle (UAV) targets (IRUTs) detection under dynamic scenarios remains a challenging task due to weak target features, varying shapes and poses, and complex background interference. Current detection methods find it difficult to address the above issues accurately and efficiently. In this paper, we design a dynamic attentive network (DANet) incorporating a scale-adaptive feature enhancement mechanism (SaFEM) and an attention-guided cross-weighting feature aggregator (ACFA). The SaFEM adaptively adjusts the network's receptive fields at hierarchical network levels leveraging separable deformable convolution (SDC), which enhances the network's multi-scale IRUT awareness. The ACFA, modulated by two crossing attention mechanisms, strengthens structural and semantic properties on neighboring levels for the accurate representation of multi-scale IRUT features from different levels. A plug-and-play anti-distractor contrastive regularization (ADCR) is also imposed on our DANet, which enforces similarity on features of targets and distractors from a new uncompressed feature projector (UFP) to increase the network's anti-distractor ability in complex backgrounds. To further increase the multi-scale UAV detection performance of DANet while maintaining its efficiency superiority, we propose a novel scale-specific knowledge distiller (SSKD) based on a divide-and-conquer strategy. For the "divide'' stage, we intendedly construct three task-oriented teachers to learn tailored knowledge for small-, medium-, and large-scale IRUTs. For the "conquer'' stage, we propose a novel element-wise attentive distillation module (EADM), where we employ a pixel-wise attention mechanism to highlight teacher and student IRUT features, and incorporate IRUT-associated prior knowledge for the collaborative transfer of refined multi-scale IRUT features to our DANet. Extensive experiments on real infrared UAV datasets demonstrate that our DANet is able to detect multi-scale UAVs with a satisfactory balance between accuracy and efficiency.
Houzhang Fang, Zikai Liao, Lu Wang 0014, Qingshan Li, Yi Chang 0002, Luxin Yan, Xuhua Wang
ACM Multimedia6
2023 Content-Aware Subspace Low-Rank Tensor Recovery for Hyperspectral Image Restoration
abstract
The low-rank tensor model has made great progress for hyperspectral image (HSI) restoration. Recently, the low-rank tensor methods have further been boosted with subspace learning by transforming original HSI into a low-dimensional subspace with reduced computational burden and discriminative feature representation. However, existing subspace-based methods consistently employ a fixed subspace dimension for all patches, which may violate the intrinsic dimension discrepancy of different image content, leading to information loss or redundancy. In this work, our key observation is that the intrinsic subspace of different image patches along different dimensions is different, which should be adaptively modelled for compact feature extraction. Therefore, we propose a content-aware subspace low-rank tensor recovery (CSLRTR) methods by leveraging both deep network and low-rank tensor model. Specifically, we first analyze the intrinsic discrepancy of different HSI patches among both spatial and spectral dimensions, and design a simple network to adaptively learn the optimal subspace dimension. The adaptive subspace learning and low-rank tensor recovery are iteratively performed and mutually promote each other. On one hand, the learned subspace would contribute to more compact low-rank representation for better restoration; on the other hand, the low rank tensor recovery with less degradations would definitely ease the difficulty of the subspace estimation. Note that, the adaptive content-aware subspace strategy has been simultaneously employed on both spectral and nonlocal dimensions, where the spectral-spatial relationship has been further strengthened with better restoration. We have performed extensive experiments on different datasets and restoration tasks, and extended the content-aware subspace strategy to previous methods.
Xueyao Xiao, Wei Zhang 0161, Yi Chang 0002, Shuning Cao, Wei He 0003, Houzhang Fang, Luxin Yan
IEEE Trans. Geosci. Remote. Sens.7
2023 Differentiated Attention Guided Network Over Hierarchical and Aggregated Features for Intelligent UAV Surveillance
abstract
Intelligent unmanned aerial vehicle (UAV) surveillance based on infrared imaging has wide applications in the anti-UAV system for protecting urban security and aerial safety. However, weak target features and complex background distraction pose great challenges for the accurate detection of UAVs. To address this issue, we propose a novel differentiated attention guided network to adaptively strengthen the discriminative features between UAV targets and complex background. First, a novel spatial-aware channel attention (SCA) is introduced into deep layers via preserving critical spatial features and leveraging channel interdependencies to focus on the large-scale targets. The channel-modulated deformable spatial attention is introduced into shallow layers via refining channel context and dynamically perceiving the spatial features for focusing on the small-scale targets. A combination of the above two attention mechanisms is employed in intermediate layers of the network for concentrating on the medium-scale targets. Then, we embed a feature aggregator at the detection branches to guide the information exchange of high-level feature maps and low-level feature maps with a bottom-up context modulation, and integrate an SCA at the end to further boost the distinctive feature representation for task-awareness. The above design can adaptively enhance multiscale UAV target features and suppress complex background interferences, leading to better detection performance, especially for small targets. Extensive experiments on real infrared UAV datasets reveal that the proposed method outperforms the baseline object detectors by a large margin, validating its feasibility in real-world infrared UAV detection. The source code can be found athttps://github.com/KALEIDOSCOPEIP/DAGNet.
Houzhang Fang, Zikai Liao, Xuhua Wang, Yi Chang 0002, Luxin Yan
IEEE Trans. Ind. Informatics5
2022 Close the Loop: A Unified Bottom-Up and Top-Down Paradigm for Joint Image Deraining and Segmentation
abstract
In this work, we focus on a very practical problem: image segmentation under rain conditions. Image deraining is a classic low-level restoration task, while image segmentation is a typical high-level understanding task. Most of the existing methods intuitively employ the bottom-up paradigm by taking deraining as a preprocessing step for subsequent segmentation. However, our statistical analysis indicates that not only deraining would benefit segmentation (bottom-up), but also segmentation would further improve deraining performance (top-down) in turn. This motivates us to solve the rainy image segmentation task within a novel top-down and bottom-up unified paradigm, in which two sub-tasks are alternatively performed and collaborated with each other. Specifically, the bottom-up procedure yields both clearer images and rain-robust features from both image and feature domains, so as to ease the segmentation ambiguity caused by rain streaks. The top-down procedure adopts semantics to adaptively guide the restoration for different contents via a novel multi-path semantic attentive module (SAM). Thus the deraining and segmentation could boost the performance of each other cooperatively and progressively. Extensive experiments and ablations demonstrate that the proposed method outperforms the state-of-the-art on rainy image segmentation.
Yi Li 0033, Yi Chang 0002, Changfeng Yu, Luxin Yan
AAAI4
2022 Category-Aware Transformer Network for Better Human-Object Interaction Detection
abstract
Human-Object Interactions (HOI) detection, which aims to localize a human and a relevant object while recognizing their interaction, is crucial for understanding a still image. Recently, tranformer-based models have significantly advanced the progress of HOI detection. However, the capability of these models has not been fully explored since the Object Query of the model is always simply initialized as just zeros, which would affect the performance. In this paper, we try to study the issue of promoting transformer-based HOI detectors by initializing the Object Query with category-aware semantic information. To this end, we innovatively propose the Category-Aware Transformer Network (CATN). Specifically, the Object Query would be initialized via category priors represented by an external object detection model to yield a better performance. Moreover, such category priors can be further used for enhancing the representation ability of features via the attention mechanism. We have firstly verified our idea via the Oracle experiment by initializing the Object Query with the groundtruth category information. And then extensive experiments have been conducted to show that a HOI detection model equipped with our idea outperforms the baseline by a large margin to achieve a new state-of-the-art result.
Leizhen Dong, Kunlun Xu, Zhijun Zhang 0009, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
CVPR5
2022 Physically Disentangled Intra- and Inter-domain Adaptation for Varicolored Haze Removal
abstract
Learning-based image dehazing methods have achieved marvelous progress during the past few years. On one hand, most approaches heavily rely on synthetic data and may face difficulties to generalize well in real scenes, due to the huge domain gap between synthetic and real images. On the other hand, very few works have considered the varicolored haze, caused by chromatic casts in real scenes. In this work, our goal is to handle the new task: real-world varicolored haze removal. To this end, we propose a physically disentangled joint intra- and inter-domain adaptation paradigm, in which intra-domain adaptation focuses on color correction and inter-domain procedure transfers knowledge between synthetic and real domains. We first learn to physically disentangle haze images into three components complying with the scattering model: background, transmission map, and atmospheric light. Since haze color is determined by atmospheric light, we perform intra-domain adaptation by specifically translating atmospheric light from varicolored space to unified color-balanced space, and then reconstructing color-balanced haze image through the scattering model. Consequently, we perform inter-domain adaptation between the synthetic and real images by mutually exchanging the background and other two components. Then we can reconstruct both identity and domain-translated haze images with self-consistency and adversarial loss. Extensive experiments demonstrate the superiority of the proposed method over the state-of-the-art for real varicolored image dehazing.
Yi Li 0033, Yi Chang 0002, Changfeng Yu, Luxin Yan
CVPR5
2022 Unsupervised Deraining: Where Contrastive Learning Meets Self-similarity
abstract
Image deraining is a typical low-level image restoration task, which aims at decomposing the rainy image into two distinguishable layers: clean image layer and rain layer. Most of the existing learning-based deraining methods are supervisedly trained on synthetic rainy-clean pairs. The domain gap between the synthetic and real rains makes them less generalized to different real rainy scenes. Moreover, the existing methods mainly utilize the property of the two layers independently, while few of them have considered the mutually exclusive relationship between the two layers. In this work, we propose a novel non-local contrastive learning (NLCL) method for unsupervised image deraining. Consequently, we not only utilize the intrinsic self-similarity property within samples, but also the mutually exclusive property between the two layers, so as to better differ the rain layer from the clean image. Specifically, the non-local self-similarity image layer patches as the positives are pulled together and similar rain layer patches as the negatives are pushed away. Thus the similar positive/negative samples that are close in the original space benefit us to enrich more discriminative representation. Apart from the self-similarity sampling strategy, we analyze how to choose an appropriate feature encoder in NLCL. Extensive experiments on different real rainy datasets demonstrate that the proposed method obtains state-of-the-art performance in real deraining.
Yuntong Ye, Changfeng Yu, Yi Chang 0002, Lin Zhu 0012, Xi-Le Zhao, Luxin Yan, Yonghong Tian 0001
CVPR6
2022 High-Fidelity Variable-Rate Image Compression via Invertible Activation Transformation
abstract
Learning-based methods have effectively promoted the community of image compression. Meanwhile, variational autoencoder(VAE) based variable-rate approaches have recently gained much attention to avoid the usage of a set of different networks for various compression rates. Despite the remarkable performance that has been achieved, these approaches would be readily corrupted once multiple compression/decompression operations are executed, resulting in the fact that image quality would be tremendously dropped and strong artifacts would appear. Thus, we try to tackle the issue of high-fidelity fine variable-rate image compression and propose the Invertible Activation Transformation(IAT) module. We implement the IAT in a mathematical invertible manner on a single rate Invertible Neural Network(INN) based model and the quality level(QLevel) would be fed into the IAT to generate scaling and bias tensors. IAT and QLevel together give the image compression model the ability of fine variable-rate control while better maintaining the image fidelity. Extensive experiments demonstrate that the single rate image compression model equipped with our IAT module has the ability to achieve variable-rate control without any compromise. And our IAT-embedded model obtains comparable rate-distortion performance with recent learning-based image compression methods. Furthermore, our method outperforms the state-of-the-art variable-rate image compression method by a large margin, especially after multiple re-encodings.
Shilv Cai, Zhijun Zhang 0009, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
ACM Multimedia4
2022 Effective actor-centric human-object interaction detection
Kunlun Xu, Zhijun Zhang 0009, Leizhen Dong, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002
Image Vis. Comput.6
2022 Learning dynamic background for weakly supervised moving object detection
Zhijun Zhang 0009, Yi Chang 0002, Sheng Zhong 0001, Luxin Yan, Xu Zou 0002
Image Vis. Comput.4
2022 Infrared Small UAV Target Detection Based on Residual Image Prediction via Global and Local Dilated Residual Networks
abstract
Thermal infrared imaging possesses the ability to monitor unmanned aerial vehicles (UAVs) in both day and night conditions. However, long-range detection of the infrared UAVs often suffers from small/dim targets, heavy clutter, and noise in the complex background. The conventional local prior-based and the nonlocal prior-based methods commonly have a high false alarm rate and low detection accuracy. In this letter, we propose a model that converts small UAV detection into a problem of predicting the residual image (i.e., background, clutter, and noise). Such novel reformulation allows us to directly learn a mapping from the input infrared image to the residual image. The constructed image-to-image network integrates the global and the local dilated residual convolution blocks into the U-Net, which can capture local and contextual structure information well and fuse the features at different scales both for image reconstruction. Additionally, subpixel convolution is utilized to upscale the image and avoid image distortion during upsampling. Finally, the small UAV target image is obtained by subtracting the residual image from the input infrared image. The comparative experiments demonstrate that the proposed method outperforms state-of-the-art ones in detecting real-world infrared images with heavy clutter and dim targets.
Houzhang Fang, Mingjiang Xia, Yi Chang 0002, Luxin Yan
IEEE Geosci. Remote. Sens. Lett.5
2022 Robust Blind Deblurring Under Stripe Noise for Remote Sensing Images
abstract
The blind image deblurring methods have achieved great progress for Gaussian random noise. Few works have paid attention to the image deblurring under the structural noise, which is a very common degradation in multi-detectors imaging systems. This paper considers the practical yet challenging problem of blind deblurring in presence of the line-pattern stripe noise for remote sensing images. To overcome this issue, we explicitly formulate the structural noise into a novel and robust blind image deblurring framework. We observe that the structural line-pattern stripe noise would deteriorate both the kernel estimation and non-blind deblurring, and propose a three-stage restoration framework to progressively estimate the blur kernel and clean image. Specifically, we first estimate an intermediate blur kernel by getting rid of the negative influence of the stripe noise in the unidirectional gradient domain. Next, a learning-based kernel refinement network is introduced to rectify the missing details of the inaccurate kernel. Finally, a low-rank decomposition-based non-blind deblurring model is proposed to simultaneously estimate the clean image and stripe noise. Experimental results on real and synthetic datasets demonstrate that the proposed RBDS method outperforms the state-of-the-art blind deblurring methods.
Shuning Cao, Houzhang Fang, Wei Zhang 0161, Yi Chang 0002, Luxin Yan
IEEE Trans. Geosci. Remote. Sens.6
2022 On-Orbit Real-Time Variational Image Destriping: FPGA Architecture and Implementation
abstract
On-orbit real-time image processing is of increasing demands due to requirements for quick-response missions. Image destriping is usually an important pre-processing step to improve image quality in practice. The unidirectional variational models have shown impressive destriping performance. However, they are not easy for real-time implementation for high-computation complexity and there are few attempts for hardware implementation. This article is the first hardware implementation of variational image destriping algorithm, which achieves high throughput for large-swath remote-sensing images. In this article, a fully pipelined hardware architecture is proposed. First, the involved iteration loop is unrolled and a coarse-grained parallelism is obtained. Second, for each deployed iteration computation blocks (ICBs), a dedicated timing arrangement is designed to alleviate the bottleneck caused by the data dependency within each ICB, obtaining a fine-grained parallelism. Moreover, to further optimize the critical path, an approximate simplification scheme is proposed, saving the resource usage and reducing computing delay. The proposed architecture is implemented and verified on a XILINX 6vcx240t field programmable gate array (FPGA); it achieves a maximum frame rate up to 41.9 frames/s with delay of only tens of row cycles for 8-bit${{2048}} \times {{2048}}$images. It performs all the processing on the pixels in raster scan order on- the-fly as they are being transmitted from camera payload, which significantly facilitates on- orbit real-time processing of large-swath remote-sensing images with high data rate.
Yi Chang 0002, Luxin Yan
IEEE Trans. Geosci. Remote. Sens.3
2021 Closing the Loop: Joint Rain Generation and Removal via Disentangled Image Translation
abstract
Existing deep learning-based image deraining methods have achieved promising performance for synthetic rainy images, typically rely on the pairs of sharp images and simulated rainy counterparts. However, these methods suffer from significant performance drop when facing the real rain, because of the huge gap between the simplified synthetic rain and the complex real rain. In this work, we argue that the rain generation and removal are the two sides of the same coin and should be tightly coupled. To close the loop, we propose to jointly learn real rain generation and removal procedure within a unified disentangled image translation framework. Specifically, we propose a bidirectional disentangled translation network, in which each unidirectional network contains two loops of joint rain generation and removal for both the real and synthetic rain image, respectively. Meanwhile, we enforce the disentanglement strategy by decomposing the rainy image into a clean background and rain layer (rain removal), in order to better preserve the identity background via both the cycle-consistency loss and adversarial loss, and ease the rain layer translating between the real and synthetic rainy image. A counterpart composition with the entanglement strategy is symmetrically applied for rain generation. Extensive experiments on synthetic and real-world rain datasets show the superiority of proposed method compared to state-of-the-arts.
Yuntong Ye, Yi Chang 0002, Hanyu Zhou, Luxin Yan
CVPR4
2021 Skeleton-Aware Network for Aircraft Landmark Detection
Yuntong Ye, Yi Chang 0002, Yi Li 0033, Luxin Yan
ICIG (1)4
2021 Unsupervised Image Deraining: Optimization Model Driven Deep CNN
abstract
The deep convolutional neural network has achieved significant progress for single image rain streak removal. However, most of the data-driven learning methods are full-supervised or semi-supervised, unexpectedly suffering from significant performance drop when dealing with the real rain. These data-driven learning methods are representative yet generalize poor for real rain. The opposite holds true for the model-driven unsupervised optimization methods. To overcome these problems, we propose a unified unsupervised learning framework which inherits the generalization and representation merits for real rain removal. Specifically, we first discover a simple yet important domain knowledge that directional rain streak is anisotropic while the natural clean image is isotropic, and formulate the structural discrepancy into the energy function of the optimization model. Consequently, we design an optimization model driven deep CNN in which the unsupervised loss function of the optimization model is enforced on the proposed network for better generalization. In addition, the architecture of the network mimics the main role of the optimization models with better feature representation. On one hand, we take advantage of the deep network to improve the representation. On the other hand, we utilize the unsupervised loss of the optimization model for better generalization. Overall, the unsupervised learning framework achieves good generalization and representation: unsupervised training (loss) with only a few real rainy images (input) and physical meaning network (architecture). Extensive experiments on synthetic and real-world rain datasets show the superiority of the proposed method.
Changfeng Yu, Yi Chang 0002, Yi Li 0033, Xi-Le Zhao, Luxin Yan
ACM Multimedia5
2021 A New Orientation Estimation Method Based on Rotation Invariant Gradient for Feature Points
abstract
For many remote sensing image applications, orientation estimation is a crucial step for feature points extraction and matching, but it has attracted little attention. Due to the intensity differences between remote sensing image pairs, it is difficult to estimate the orientations of corresponding points accurately, resulting in performance degradation of feature matching. Thus, encountering the intensity differences, a robust spatial structure description for feature regions, and an effective calculation manner from description to orientation play a key role in accurate orientation estimation. To this end, in this letter, we first define a plausible orientation for feature points by the total gradient, offering an effective way to convert the gradient trend of the feature region to orientation. Therefore, we further propose a novel orientation estimation method, in which the rotation invariant gradient is introduced to improve the accuracy of gradient calculation and robustness of spatial structure description. Experimental results on multisensor remote sensing images demonstrate that our method increases the orientation estimation accuracy remarkably and outperforms other orientation estimation methods by a large margin, and effectively improves the performance of feature matching.
Sheng Zhong 0001, Luxin Yan
IEEE Geosci. Remote. Sens. Lett.5
2021 Category-Aware Aircraft Landmark Detection
abstract
Aircraft landmark detection (ALD) aims at detecting the keypoints of aircraft, which can serve as an important role for subsequent applications such as fine-grained aircraft recognition. In ALD, the physical size discrepancy between different kinds of aircraft may lead to inconsistent landmark structure, which significantly harms landmark detection results. In this letter, we take advantage of the category prior to alleviate the size discrepancy in ALD. The proposed category-aware landmark detection network (CALDN) possesses two streams: a classification stream for size categorization and a localization stream for landmark detection. Instance-level size category information captured by classification stream serves as the guidance in the localization stream for robust landmark detection. Moreover, a category attention module (CAM) is proposed for better-utilizing category information to guide ALD. Benefitting from the adaptive attention mechanism, CAM can automatically highlight category-specific features for ulteriorly reducing the influence of size discrepancy. Furthermore, to advance ALD research, we contribute the first perspective-variant aircraft landmark dataset. Solid experiments demonstrate the superiority of our method.
Yi Li 0033, Yi Chang 0002, Yuntong Ye, Xu Zou 0002, Sheng Zhong 0001, Luxin Yan
IEEE Signal Process. Lett.6
2021 Towards Unconstrained Facial Landmark Detection Robust to Diverse Cropping Manners
abstract
Facial landmark detection is one crucial step for face-based image/video analysis. Despite the fact that recently many facial landmark detection models have achieved remarkable performance, most state-of-the-art heatmap regression-based methods heavily rely on initialization of the face detector. However, there inevitably exists semantic gaps among different annotators or face detectors. An improper facial bounding box will tremendously drop off the performance of the facial landmark detection model. Facial landmark detection would be more practical if robust to face images cropped by diverse manners (see Figure 1, the col.1 shows face images cropped by a proper bounding box, col.2 and col.3 show face images cropped by an oversize and a small bounding boxes respectively). To this end, we present a “Unconstrained Facial Landmark Detection(UFLD)” mechanism, that aims at enhancing the robustness of facial landmark detection, to deal with the inconsistent cropping manner issue. UFLD consists of two aspects: a Transformation-Invariant Landmark Detector(TILD) and an Availability-Guided Solver(AGS). TILD gives the ability to detect consistent landmarks for face images cropped by diverse manners. And AGS can alleviate the by-effect of “landmarks outside the image” caused by improper cropping results or TILD, and further promote the performance. The proposed mechanism achieved above 6.5% improvement in standard normalized landmarks mean error reduction on face images cropped by diverse manners compared to baselines.
Xu Zou 0002, Luxin Yan, Sheng Zhong 0001, Ying Wu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 Hyperspectral Image Restoration: Where Does the Low-Rank Property Exist
abstract
Hyperspectral image (HSI) restoration is to recover the clean image from degraded version, such as the noisy, blurred, or damaged. Recent low-rank tensor-based recovery methods have been widely explored in HSIs restoration. Most of previous methods, however, neglect an inconspicuous but important phenomenon that the physical meaning and dimension along the spatial, spectral, and nonlocal mode are markedly different. In this work, we discover the low-rank property discrepancy along spatial, spectral, and nonlocal self-similarity mode in the HSIs, and argue that the intrinsic low-rank correlations along each mode contribute different to the final restoration results. Consequently, we figure out that the combination of the spectral and nonlocal-induced low-rank is most beneficial for HSIs modeling, and propose an optimal low-rank tensor (OLRT) model for HSIs restoration. Furthermore, we not only explore the low-rank property in the image component, but also in the sparse error component (stripe noise in HSIs). Thus, we extend OLRT to the OLRT-robust principal component analysis (RPCA) with low-rank tensor priors for both the HSIs and sparse error. Besides, previous methods are usually designed for one specific HSI task, which is less robust to various tasks. We prove that the proposed optimal low-rank prior is very flexible for various HSI restoration problems including denoising, deblurring, inpainting, and destriping. The proposed methods have been extensively evaluated on several benchmarks and tasks, and greatly outperform state-of-the-art (STOA). We show the simple yet effective OLRT strategy is also beneficial to STOA.
Yi Chang 0002, Luxin Yan, Bingling Chen, Sheng Zhong 0001, Yonghong Tian 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Learning Blind Denoising Network for Noisy Image Deblurring
abstract
Noisy image deblurring is to recover the blurry image in the presence of the random noise. One key to this problem is to know the noise level in each iteration. The existing methods manually adjust the regularization parameter for varying noise levels, which is quite inaccuracy and tedious for practical application. In this work, we discover that the noise level and the denoiser is tightly coupled. Consequently, we propose efficient blind denoising convolutional neural network (BDCNN) consisting of two stages: a down-sampling regression network for estimating noise level and a fully convolutional network for denoising, such that our model could adaptively handle the unknown noise level during iteration. Further, the BDCNN functions as a discriminative prior and is plugged into the iterative deblurring framework for noisy image deblurring. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods in terms of practicability and performance.
Meiya Chen, Yi Chang 0002, Shuning Cao, Luxin Yan
ICASSP4
2020 Intelligent multi-agent based C-RAN architecture for 5G radio resource management
Zbigniew Dziong, Luxin Yan, Adnane Cabani
Comput. Networks3
2020 Weighted Low-Rank Tensor Recovery for Hyperspectral Image Restoration
abstract
Hyperspectral imaging, providing abundant spatial and spectral information simultaneously, has attracted a lot of interest in recent years. Unfortunately, due to the hardware limitations, the hyperspectral image (HSI) is vulnerable to various degradations, such as noises (random noise), blurs (Gaussian and uniform blur), and downsampled (both spectral and spatial downsample), each corresponding to the HSI denoising, deblurring, and super-resolution tasks, respectively. Previous HSI restoration methods are designed for one specific task only. Besides, most of them start from the 1-D vector or 2-D matrix models and cannot fully exploit the structurally spectral-spatial correlation in 3-D HSI. To overcome these limitations, in this article, we propose a unified low-rank tensor recovery model for comprehensive HSI restoration tasks, in which nonlocal similarity within spectral-spatial cubic and spectral correlation are simultaneously captured by third-order tensors. Furthermore, to improve the capability and flexibility, we formulate it as a weighted low-rank tensor recovery (WLRTR) model by treating the singular values differently. We demonstrate the reweighed strategy, which has been extensively studied in the matrix, also greatly benefits the tensor modeling. We also consider the stripe noise in HSI as the sparse error by extending WLRTR to robust principal component analysis (WLRTR-RPCA). Extensive experiments demonstrate the proposed WLRTR models consistently outperform state-of-the-art methods in typical HSI low-level vision tasks, including denoising, destriping, deblurring, and super-resolution.
Yi Chang 0002, Luxin Yan, Xi-Le Zhao, Houzhang Fang, Zhijun Zhang 0009, Sheng Zhong 0001
IEEE Trans. Cybern.2
2020 Toward Universal Stripe Removal via Wavelet-Based Deep Convolutional Neural Network
abstract
Stripe noise from different remote sensing imaging systems varies considerably in terms of response, length, angle, and periodicity. Due to the complex distributions of different stripes, the destriping results of previous methods may be oversmoothed or contain residual stripe. To overcome this key problem, we provide a comprehensive analysis of existing destriping methods and propose a deep convolutional neural network (CNN) for handling various kinds of stripes. Moreover, previous methods individually model the stripe or the image priors, which may lose the relationship between them. In this article, a two-stream CNN is designed to simultaneously model the stripe and image, which better facilitates distinguishing them from each other. Moreover, we incorporate the wavelet into our CNN model for better directional feature representation. Therefore, the CNN learns the discriminative representation from the external data set, while the wavelet models the internal directionality of the stripe, in which both the internal and external priors are beneficial to the destriping task. In addition, the wavelet extracts the multiscale information with a larger receptive field for global contextual information modeling; thus, we can better distinguish the stripe from the similar image line pattern structures. The proposed method has been extensively evaluated on a number of data sets and outperforms the state-of-the-art methods by substantially a large margin in terms of quantitative and qualitative assessments, speed, and robustness.
Yi Chang 0002, Meiya Chen, Luxin Yan, Xi-Le Zhao, Yi Li 0033, Sheng Zhong 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 Learning Robust Facial Landmark Detection via Hierarchical Structured Ensemble
abstract
Heatmap regression-based models have significantly advanced the progress of facial landmark detection. However, the lack of structural constraints always generates inaccurate heatmaps resulting in poor landmark detection performance. While hierarchical structure modeling methods have been proposed to tackle this issue, they all heavily rely on manually designed tree structures. The designed hierarchical structure is likely to be completely corrupted due to the missing or inaccurate prediction of landmarks. To the best of our knowledge, in the context of deep learning, no work before has investigated how to automatically model proper structures for facial landmarks, by discovering their inherent relations. In this paper, we propose a novel Hierarchical Structured Landmark Ensemble (HSLE) model for learning robust facial landmark detection, by using it as the structural constraints. Different from existing approaches of manually designing structures, our proposed HSLE model is constructed automatically via discovering the most robust patterns so HSLE has the ability to robustly depict both local and holistic landmark structures simultaneously. Our proposed HSLE can be readily plugged into any existing facial landmark detection baselines for further performance improvement. Extensive experimental results demonstrate our approach significantly outperforms the baseline by a large margin to achieve a state-of-the-art performance.
Xu Zou 0002, Sheng Zhong 0001, Luxin Yan, Xiangyun Zhao, Jiahuan Zhou, Ying Wu 0001
ICCV3
2019 Directional-Aware Automatic Defect Detection in High-Speed Railway Catenary System
abstract
It is crucial to detect the defect objects that could be a hidden danger in high-speed catenary system. Instead of detecting objects based on horizontal rectangle, we propose an end-to-end trainable directional-aware defect detection network (D3-Net) which can automatic select the defective components. D3-Net is composed of two streams, including a directional-aware locator that regresses an inclined rectangle and a channel-wise classifier to diagnose the defects of object. Specifically, we make full use of the directionality characteristic of the man-made objects in the localization stream, which can predict the inclined rectangle with an angle parameter. Besides, we introduce a channel attention module (CAM) in the classification stream to obtain discriminative features for better distinguishing the normal and defective objects. Experimental results on two datasets, Dropper and Insulator, demonstrate that our proposed model outperforms the traditional horizontal detection methods.
Yi Chang 0002, Ziqin Li, Sheng Zhong 0001, Luxin Yan
ICIP5
2019 Infrared Aerothermal Nonuniform Correction via Deep Multiscale Residual Network
abstract
In the infrared focal plane arrays imaging systems, the temperature-dependent nonuniformity effects severely degrade the image quality. In this letter, we propose a very deep convolutional neural network for unified infrared aerothermal nonuniform correction. Our network is built with the multiscale and residual training. The multiscale subnetworks utilize the multiscale property in the images, and the long-short-term residual learning contributes to the information propagation. Compared with the previous methods, the proposed method is more robust to various nonuniform artifacts and more efficient at processing time. Experimental results validate the superiority of our method for infrared nonuniform correction.
Yi Chang 0002, Luxin Yan, Li Liu 0050, Houzhang Fang, Sheng Zhong 0001
IEEE Geosci. Remote. Sens. Lett.2
2019 Motion-blur kernel size estimation via learning a convolutional neural network
Lerenhan Li, Nong Sang, Luxin Yan, Changxin Gao
Pattern Recognit. Lett.3
2019 High-Throughput Architecture for Both Lossless and Near-lossless Compression Modes of LOCO-I Algorithm
abstract
Real-time image lossless and near-lossless compression based on LOCO-I algorithm is in great demand in many critical missions, such as satellite remote sensing, space exploration, and nuclear medical imaging, for its excellent complexity/compression rate tradeoff. However, the real-time implementation of LOCO-I encounters two bottlenecks in error prediction: the context conflict in context update and the pixel reconstruction loop in near-lossless mode, which always limit the overall throughput in practice. This paper adopts a patch-wise compression method and proposes a high-performance globally pipelined hardware architecture with spatial parallelism in local, in which patch-wise dual parallel error prediction (PEP) modules are designed. The PEP gains an extra pixel cycle and thus, alleviates both the two bottlenecks. Besides, an equivalent simplification scheme is designed to accelerate the error quantization computation, the most time and resource consuming procedure in the pixel reconstruction loop, saving the resource usage, and reducing computing delay. The proposed encoder is able to compress each image patch independently in either lossless or near-lossless mode by parameter setting. It facilitates the Region of Interests compression so as to improve the overall compression ratio and achieves excellent information fidelity. Moreover, the possible error propagation can be prevented and constrained within one single patch. The proposed architecture is implemented on a XILINX Virtex6-75t FPGA and achieves a maximum throughput up to 51.684 MPixel/s. It is the fastest architecture reported in the literature which implements both lossless and near-lossless compression modes.
Luxin Yan, Hongshi Sang, Tianxu Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2019 HSI-DeNet: Hyperspectral Image Restoration via Convolutional Neural Network
abstract
The spectral and the spatial information in hyperspectral images (HSIs) are the two sides of the same coin. How to jointly model them is the key issue for HSIs' noise removal, including random noise, structural stripe noise, and dead pixels/lines. In this paper, we introduce the deep convolutional neural network (CNN) to achieve this goal. The learned filters can well extract the spatial information within their local receptive filed. Meanwhile, the spectral correlation can be depicted by the multiple channels of the learned 2-D filters, namely, the number of filters in each layer. The consequent advantages of our CNN-based HSI denoising method (HSI-DeNet) over previous methods are threefold. First, the proposed HSI-DeNet can be regarded as a tensor-based method by directly learning the filters in each layer without damaging the spectral-spatial structures. Second, the HSI-DeNet can simultaneously accommodate various kinds of noise in HSIs. Moreover, our method is flexible for both single image and multiple images by slightly modifying the channels of the filters in the first and last layers. Last but not least, our method is extremely fast in the testing phase, which makes it more practical for real application. The proposed HSI-DeNet is extensively evaluated on several HSIs, and outperforms the state-of-the-art HSI-DeNets in terms of both speed and performance.
Yi Chang 0002, Luxin Yan, Houzhang Fang, Sheng Zhong 0001, Wenshan Liao
IEEE Trans. Geosci. Remote. Sens.2
2017 Hyper-Laplacian Regularized Unidirectional Low-Rank Tensor Recovery for Multispectral Image Denoising
abstract
Recent low-rank based matrix/tensor recovery methods have been widely explored in multispectral images (MSI) denoising. These methods, however, ignore the difference of the intrinsic structure correlation along spatial sparsity, spectral correlation and non-local self-similarity mode. In this paper, we go further by giving a detailed analysis about the rank properties both in matrix and tensor cases, and figure out the non-local self-similarity is the key ingredient, while the low-rank assumption of others may not hold. This motivates us to design a simple yet effective unidirectional low-rank tensor recovery model that is capable of truthfully capturing the intrinsic structure correlation with reduced computational burden. However, the low-rank models suffer from the ringing artifacts, due to the aggregation of overlapped patches/cubics. While previous methods resort to spatial information, we offer a new perspective by utilizing the exclusively spectral information in MSIs to address the issue. The analysis-based hyper-Laplacian prior is introduced to model the global spectral structures, so as to indirectly alleviate the ringing artifacts in spatial domain. The advantages of the proposed method over the existing ones are multi-fold: more reasonably structure correlation representability, less processing time, and less artifacts in the overlapped regions. The proposed method is extensively evaluated on several benchmarks, and significantly outperforms state-of-the-art MSI denoising methods.
Yi Chang 0002, Luxin Yan, Sheng Zhong 0001
CVPR2
2017 Transformed Low-Rank Model for Line Pattern Noise Removal
abstract
This paper addresses the problem of line pattern noise removal from a single image, such as rain streak, hyperspectral stripe and so on. Most of the previous methods model the line pattern noise in original image domain, which fail to explicitly exploit the directional characteristic, thus resulting in a redundant subspace with poor representation ability for those line pattern noise. To achieve a compact subspace for the line pattern structure, in this work, we incorporate a transformation into the image decomposition model so that maps the input image to a domain where the line pattern appearance has an extremely distinct low-rank structure, which naturally allows us to enforce a low-rank prior to extract the line pattern streak/stripe from the noisy image. Moreover, the random noise is usually mixed up with the line pattern noise, which makes the challenging problem much more difficult. While previous methods resort to the spectral or temporal correlation of the multi-images, we give a detailed analysis between the noisy and clean image in both local gradient and nonlocal domain, and propose a compositional directional total variational and low-rank prior for the image layer, thus to simultaneously accommodate both types of noise. The proposed method has been evaluated on two different tasks, including remote sensing image mixed random-stripe noise removal and rain streak removal, all of which obtain very impressive performances.
Yi Chang 0002, Luxin Yan, Sheng Zhong 0001
ICCV2
2017 Hyperspectral image denoising via spectral and spatial low-rank approximation
abstract
Hyperspectral images (HSI) unavoidably suffer from degradations such as random noise, due to photon effects, calibration error, and so on. Most of existing HSI denoising methods focus on utilizing the spectral correlation or the spatial nonlocal self-similarity individually. In this paper, we propose an unified low-rank recovery framework for HSI denoising, in which taking both the underlying characteristics of high correlation across spectra and non-local self-similarity over the space cubic of HSI into consideration simultaneously. Our work rely on a basic observation that both the multiple spectral bands and similar spatial structures are lying on low-rank subspaces and can facilitate to remove the noise jointly. Experimental results on both simulated and real HSI demonstrate that the proposed method can significantly outperform the state-of-the-art methods on several datasets in terms of both visual and quantitative assessment.
Yi Chang 0002, Luxin Yan, Sheng Zhong 0001
IGARSS2
2017 Joint local and non-local priors for ground-based astronomical image denoising
abstract
Ground-based astronomical images usually suffer from the degradation of random noise, which has a negative impact on the subsequent performance. It is expected that local and nonlocal variations are the two sides of the same coin in image modeling. However, most of existing denoising methods focus on modeling the sparsity of local patches or the nonlocal self-similarity information individually, which may result in poor image structure representation ability. To remedy this issue, we propose a unified approach via combining the framelet and the low-rank prior scheme for astronomical image denoising. On one hand, the synthesis-based low-rank prior is employed to reveal the intrinsic low-dimensional subspace (namely the main edges) of the similar patches. On the other hand, the analysis-based framelet prior is introduced to capture the local subtle texture structures. Experimental results validate the effectiveness of this combination, and the proposed method outperforms the state-of-the-art denoising methods.
Luxin Yan, Wenshan Liao, Yi Chang 0002, Chunan Luo
IGARSS1
2016 Remote Sensing Image Stripe Noise Removal: From Image Decomposition Perspective
abstract
Stripe noise removal (destriping) is a fundamental problem in remote sensing image processing that holds significant practical importance for subsequent applications. These variational destriping methods have obtained impressive results and attracted widely studied research interests. However, most of them are dedicated to estimate the clear image from the striped one, paying much attention to the image itself, while ignoring the structural characteristic of stripe, which would easily cause damages to the image structure and leave residual stripes in image recovery. In this paper, we treat the image and stripe components equally and convert the image destriping task as an image decomposition problem naturally. We first give a detailed analysis about the structural characteristic of stripes and the prior knowledge about the remote sensing images. Then, incorporating them, we propose a low-rank-based single-image decomposition model (LRSID) to separate the original image from the stripe component perfectly. This low-rank constraint for the stripe perfectly matches the fact that only parts of data vectors are corrupted but the others are not. Moreover, we further utilize the spectral information of the remote sensing images, and we extend our 2-D image decomposition method to the 3-D case. Extensive experiments on both simulated and real data have been carried out to validate the effectiveness and efficiency of the proposed algorithms.
Yi Chang 0002, Luxin Yan, Sheng Zhong 0001
IEEE Trans. Geosci. Remote. Sens.2
2015 Anisotropic Spectral-Spatial Total Variation Model for Multispectral Remote Sensing Image Destriping
abstract
Multispectral remote sensing images often suffer from the common problem of stripe noise, which greatly degrades the imaging quality and limits the precision of the subsequent processing. The conventional destriping approaches usually remove stripe noise band by band, and show their limitations on different types of stripe noise. In this paper, we tentatively categorize the stripes in remote sensing images in a more comprehensive manner. We propose to treat the multispectral images as a spectral-spatial volume and pose an anisotropic spectral-spatial total variation regularization to enhance the smoothness of solution along both the spectral and spatial dimension. As a result, a more comprehensive stripes and random noise are perfectly removed, while the edges and detail information are well preserved. In addition, the split Bregman iteration method is employed to solve the resulting minimization problem, which highly reduces the computational load. We extensively validate our method under various stripe categories and show comparison with other approaches with respect to result quality, running time, and quantitative assessments.
Yi Chang 0002, Luxin Yan, Houzhang Fang, Chunan Luo
IEEE Trans. Image Process.2
2014 Global and local exploitation for saliency using bag-of-words
abstract
The guidance of attention helps human vision system to detect objects rapidly. In this study, the authors present a new saliency detection algorithm by using bag‐of‐words (BOW) representation. The authors regard salient regions as coming from globally rare features and regions locally differ from their surroundings. Our approach consists of three stages: first, calculate global rarity of visual words. A vocabulary, a group of visual words, is generated from the given image and a rarity factor for each visual word is introduced according to its occurrence. Second, calculate local contrast. Representations of local patch are achieved from the histograms of words. Then, local contrast is computed by the difference between the two BOW histograms of a patch and its surroundings. Finally, saliency is measured by the combination of global rarity and local patch contrast. We compare our model with the previous methods on natural images, and experimental results demonstrate good performance of our model and fair consistency with human eye fixations.
Zhenzhu Zheng, Yun Zhang 0002, Luxin Yan
IET Comput. Vis.3
2014 Simultaneous Destriping and Denoising for Remote Sensing Images With Unidirectional Total Variation and Sparse Representation
abstract
Remote sensing images destriping and denoising are both classical problems, which have attracted major research efforts separately. This letter shows that the two problems can be successfully solved together within a unified variational framework. To do this, we proposed a joint destriping and denoising method by integrating the unidirectional total variation and sparse representation regularizations. Experimental results on simulated and real data in terms of qualitative and quantitative assessments show significant improvements over conventional methods.
Yi Chang 0002, Luxin Yan, Houzhang Fang, Hai Liu 0004
IEEE Geosci. Remote. Sens. Lett.2
2014 An Embedded System-on-Chip Architecture for Real-time Visual Detection and Matching
abstract
Detecting and matching image features is a fundamental task in video analytics and computer vision systems. It establishes the correspondences between two images taken at different time instants or from different viewpoints. However, its large computational complexity has been a challenge to most embedded systems. This paper proposes a new FPGA-based embedded system architecture for feature detection and matching. It consists of scale-invariant feature transform (SIFT) feature detection, as well as binary robust independent elementary features (BRIEF) feature description and matching. It is able to establish accurate correspondences between consecutive frames for 720-p (1280x720) video. It optimizes the FPGA architecture for the SIFT feature detection to reduce the utilization of FPGA resources. Moreover, it implements the BRIEF feature description and matching on FPGA. Due to these contributions, the proposed system achieves feature detection and matching at 60 frame/s for 720-p video. Its processing speed can meet and even exceed the demand of most real-life real-time video analytics applications. Extensive experiments have demonstrated its efficiency and effectiveness.
Sheng Zhong 0001, Luxin Yan, Zhiguo Cao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2013 Joint blind deblurring and destriping for remote sensing images
abstract
Deblurring and destriping are both classical problems for remote sensing images, which are known to be difficult. Treating deblurring and destriping separately, such a straightforward approach, however, suffers greatly from the defective output. This paper shows that the two problems can be successfully solved together and benefit greatly from each other within a unified variational framework. To do this, we propose a joint deblurring and destriping method by combining the framelet regularization and unidirectional total variation. Extensive experiments on simulation and real remote sensing images are carried out and the results of our joint model show significant improvement over conventional methods of treating the two tasks separately.
Yi Chang 0002, Houzhang Fang, Luxin Yan, Hai Liu 0004
ICIP3
2013 A real-time embedded architecture for SIFT
Sheng Zhong 0001, Luxin Yan, Lie Kang, Zhiguo Cao 0001
J. Syst. Archit.3
2012 Atmospheric-Turbulence-Degraded Astronomical Image Restoration by Minimizing Second-Order Central Moment
abstract
Atmospheric turbulence affects imaging systems by virtue of wave propagation through a medium with a nonuniform index of refraction. It can lead to blurring in images acquired from a long distance away. In this letter, it is observed that blurring increases the second-order central moment (SOCM) of images, and we introduce a new parametric blur identification method by minimizing SOCM. The method applies to finite-support images, in which the scene consists of a finite-extent object against a uniformly black, gray, or white background. The SOCM method has been validated by direct comparisons with other methods on simulated and real degraded images.
Luxin Yan, Mingzhi Jin, Houzhang Fang, Hai Liu 0004, Tianxu Zhang
IEEE Geosci. Remote. Sens. Lett.1