Xin Fan 0001

dblp:87/3021-1 · DBLP profile ↗
← Back
217ranked-venue papers
17as first author
131since 2021 · last 2026
0000-0002-8991-4188ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 155 · 13 first-author · 91 since 2021Artificial intelligence and machine learning · 75 · 5 first-author · 48 since 2021Computer networks · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels
abstract
Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their inability to provide precise annotation data for sonar images. Therefore, designing effective object detection methods for sonar images with extremely limited labels is particularly important. To address this, we propose a teacher-student framework called RSOD, which aims to fully learn the characteristics of sonar images and develop a pseudo-label strategy suitable for these images to mitigate the impact of limited labels. First, RSOD calculates a reliability score by assessing the consistency of the teacher's predictions across different views. To leverage this score, we introduce an object mixed pseudo-label method to tackle the shortage of labeled data in sonar images. Finally, we optimize the performance of the student by implementing a reliability-guided adaptive constraint. By taking full advantage of unlabeled data, the student can perform well even in situations with extremely limited labels. Notably, on the UATD dataset, our method, using only 5% of labeled data, achieves results that can compete against those of our baseline algorithm trained on 100% labeled data. We also collected a new dataset to provide more valuable data for research in the field of sonar.
Chengzhou Li, Guanchen Meng, Qi Jia 0001, Jinyuan Liu 0001, Zhu Liu 0004, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001
AAAI10
2026 Underwater Data Collection Scheme based on LLMs
Kunhong Ji, Chi Lin 0001, Jiankang Ren, Xin Fan 0001, Zhongxuan Luo
INFOCOM5
2026 Versatile Luminosity Tuning: Relighting Illumination via Dual-Prompt Exposure Correction
Jinyuan Liu 0001, Gehui Li, Zhiying Jiang, Long Ma 0002, Miao Zhang 0004, Xin Fan 0001, Risheng Liu
Int. J. Comput. Vis.6
2026 Cross-domain time-frequency Mamba: A more effective model for long-term time series forecasting
Yuhang Duan, Lin Lin 0008, Jinyuan Liu 0001, Xin Fan 0001
Knowl. Based Syst.5
2026 Energy-Aware Adaptive Topology Control for UOWSNs
abstract
Underwater Optical Wireless Sensor Network (UOWSN) is a promising technology as it can achieve high-speed communication in underwater environment. However, affected by the uncertainty of complex underwater environment, the network topology of UOWSN is highly dynamic, making it difficult to quantify flexibility or further optimize the topological structure. Additionally, node mobility and energy constraints pose significant challenges to reliable communication. In this paper, we propose a mobility-aware and energy-efficient flexibility-based network topology evaluation model (ME-FEM) for UOWSNs. Then, a reinforcement learning model, termed ME-FEM-DRL, for optimizing the network topology based on ME-FEM is developed, which enables UOWSN to maintain an optimal topology when working in harsh underwater environments. Theoretical analysis proves the NP-hardness of the optimization problem and demonstrates that our algorithm achieves an approximation ratio of$O(\log N)$with optimal parameter boundaries. Simulation results demonstrate that the proposed method can significantly improve the network flexibility. Compared with the five baseline algorithms in simulations, ME-FEM-DRL reduces normalized topology optimization time cost by 64% and extends network lifetime by 95% on average. Test-bed experiments verify the applicability and effectiveness in practical applications for detecting emergent events.
Yang Chi, Chi Lin 0001, Haipeng Dai 0001, Yu Tian 0014, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Mob. Comput.5
2026 Adaptive Interference Alignment for Underwater Optical Wireless Sensor Networks
abstract
Underwater Optical Wireless Sensor Networks (UOWSNs) have emerged as a promising solution for high-speed underwater communication. However, these networks face a critical challenge of mutual interference among optical nodes, which occurs when the directional optical beams intersect or coverage areas overlap due to node mobility in dynamic underwater environments. Existing interference management approaches demonstrate limited effectiveness due to their reliance on simplified channel models and inability to handle rapid topology changes, resulting in significant network performance degradation. This paper presents a novel framework that systematically addresses interference management in UOWSNs through two key innovations. First, we propose a Sparse Bayesian Learning-based Interference Detection (SBL-ID) algorithm that enables real-time identification and characterization of interference patterns under complex underwater channel conditions. Second, we develop an Adaptive Interference Alignment and Delay Compensation (AIADC) algorithm that projects interference signals into a reduced-dimensional subspace, thereby enhancing the signal-to-interference ratio and facilitating accurate detection of desired signals amid interference. Our framework transforms the NP-hard interference management problem into tractable optimizations, achieving near-optimal solutions with polynomial time complexity. Extensive simulations demonstrate that our approach reduces BER by 95% and improves network throughput by 67% compared to state-of-the-art techniques. Testbed experiments conducted in both pool and lake further validate our framework's effectiveness, maintaining consistent performance improvements under diverse underwater conditions.
Yang Chi, Chi Lin 0001, Fengqi Li, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Mob. Comput.5
2025 Personalizing Federated Learning Guided by Site-Aggregated Representation for Multi-Site One-Shot Medical Image Segmentation
abstract
Personalized federated learning for medical image segmentation enables collaborative model training across multiple clinical sites without sharing patient data, by synchronizing a subset of global model parameters while retaining others for local adaptation. However, previous methods adopt paired labeled images to train the model, which is hard to apply in real scenarios due to time-consuming medical experts on manual annotation. Additionally, these methods focus on local parameter learning ignoring inter-site consistencies during local training. To address these challenges, we propose a personalizing Federated framework guided by Site-aggregated Representation (FedSR) for multi-site one-shot medical image segmentation, which exploits the site-invariant latent information to boost segmentation performance. Specifically, we propose to learn an omniscient encoder by federated learning, which can not only model the data distribution between multi-site datasets but also adapt the multi-task in an efficient way. With the learned robust representation, we further propose to learn site-aggregated representation between multi-site data by mutual information maximization, and then adopt such site-aggregated latent representation to guide the personalized dual-task head decoder. Extensive experiments conducted on two MIS tasks demonstrate that the proposed FedSR outperforms state-of-the-art one-shot MIS methods on segmentation.
Jia Wang 0036, Yunan Mei, Xin Fan 0001
BIBM5
2025 3D CT Reconstruction from X-Ray Projections Under Simulated Clinical Conditions
abstract
Reconstructing 3D CT (Computed tomography) from 2D X-rays is a critical task in medical imaging. However, existing methods often struggle in real-world scenarios, where even slight variations in posture can lead to degradation in reconstruction quality. Consequently, improving reconstruction quality in practical clinical settings continues to be a significant challenge. In this study, we first investigate the impact of small postural rotations on the quality of CT reconstruction. Based on this analysis, we propose a novel method for reconstructing 3D CT images from biplanar X-rays, which incorporates implicit 2D-to-3D registration. By embedding the concept of X-ray-to-CT registration directly into the reconstruction process, we improve robustness against pose variations. Additionally, we introduce differentiable perspective projection constraints during training, enabling the model to effectively learn the correspondence between misaligned input X-rays and the ground-truth CT volumes. The experimental findings indicate that even minor posture rotations can lead to noticeable degradation in reconstruction performance. However, the proposed method outperforms the existing approach by offering a slight but consistent improvement. This work simulates a real clinical setting to analyze how posture rotations affect the quality of learning-based CT reconstruction and proposes a potential solution to improve the robustness.
Shuqiong Wu, Xin Fan 0001
BIBM3
2025 DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resulting in fused images that offer only marginal gains in task performance and fail to provide constructive feedback for optimizing the fusion process. To overcome these limitations, we propose a Discriminative Cross-Dimension Evolutionary Learning Framework, termed DCEvo, which simultaneously enhances visual quality and perception accuracy. Leveraging the robust search capabilities of Evolutionary Learning, our approach formulates the optimization of dual tasks as a multi-objective problem by employing an Evolutionary Algorithm (EA) to dynamically balance loss function parameters. Inspired by visual neuroscience, we integrate a Discriminative Enhancer (DE) within both the encoder and decoder, enabling the effective learning of complementary features from different modalities. Additionally, our Cross-Dimensional Embedding (CDE) block facilitates mutual enhancement between high-dimensional task features and low-dimensional fusion features, ensuring a cohesive and efficient feature integration process. Experimental results on three benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, achieving an average improvement of 9.32% in visual quality while also enhancing subsequent high-level tasks. The code is available at https://github.com/Beate-Suy-Zhang/DCEvo.
Jinyuan Liu 0001, Qingyun Mei, Xingyuan Li 0005, Yang Zou 0004, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001
CVPR9
2025 Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond
abstract
Multi-modality image fusion, particularly infrared and visible, plays a crucial role in integrating diverse modalities to enhance scene understanding. Although early research prioritized visual quality, preserving fine details and adapting to downstream tasks remains challenging. Recent approaches attempt task-specific design but rarely achieve "The Best of Both Worlds" due to inconsistent optimization goals. To address these issues, we propose a novel method that leverages the semantic knowledge from the Segment Anything Model (SAM) to Grow the quality of fusion results and Enable downstream task adaptability, namely SAGE. Specifically, we design a Semantic Persistent Attention (SPA) Module that efficiently maintains source information via the persistent repository while extracting high-level semantic priors from SAM. More importantly, to eliminate the impractical dependence on SAM during inference, we introduce a bi-level optimization-driven distillation mechanism with triplet losses, which allow the student network to effectively extract knowledge. Extensive experiments show that our method achieves a balance between high-quality visual results and downstream task adaptability while maintaining practical deployment efficiency. The code is available at https://github.com/RollingPlain/SAGE_IVIF.
Guanyao Wu, Hongming Fu, Yichuan Peng, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
CVPR6
2025 TextMEF: Text-guided Prompt Learning for Multi-exposure Image Fusion
abstract
Multi-exposure image fusion~(MEF) aims to integrate a set of low dynamic range images, producing a single image with a higher dynamic range than either one. Despite significant advancements, current MEF approaches still struggle to handle extremely over- or under-exposed conditions, resulting in unsatisfactory visual effects such as hallucinated details and distorted color tones. With this regard, we propose TextMEF, a prompt-driven fusion method enhanced by prompt learning, for multi-exposure image fusion. Specifically, we learn a set of prompts based on text-image similarity among negative and positive samples (over-exposed, under-exposed images, and well-exposed ones). These learned prompts are seamlessly integrated into the loss function, providing high-level guidance for constraining non-uniform exposure regions. Furthermore, we develop a attention Mamba module effectively translates over-/under- exposed regional features into exposure invariant space and ensure them to build efficient long-range dependency to high dynamic range image. Extensive experimental results on three publicly available benchmarks demonstrate that our TextMEF significantly outperforms state-of-the-art approaches in both visual inspection and objective analysis.
Jinyuan Liu 0001, Qianjun Huang, Guanyao Wu, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001
IJCAI8
2025 LiDAR-Track: Multi-Person Positioning and Tracking Using LiDAR
Kunhong Ji, Chi Lin 0001, Jie Xiong 0001, Liming Chen 0001, Xin Fan 0001, Guowei Wu 0001
INFOCOM5
2025 EchoCardMAE: Video Masked Auto-Encoders Customized for Echocardiography
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Miao Zhang 0004, Yi Wang 0037, Xin Fan 0001, Hongkai Wang 0002, Qingxiong Yue, Xiangjian He, Yen-Wei Chen 0001
MICCAI (13)7
2025 TNT-GS: Truncated and Tailored Gaussian Splatting
abstract
Gaussian Splatting (GS) is widely used for efficient 3D scene representation and rendering by modeling scenes as continuous Gaussian distributions. However, GS struggles with high-frequency details and sharp transitions due to its low-pass filtering effect, often requiring multiple Gaussian stacking, which increases computational and memory costs. To overcome these limitations, we propose Truncated and Tailored Gaussian Splatting (TNT-GS), a novel approach that enhances shape complexity and preserves sharp boundaries. Our method truncates Gaussians to generate sharp edges and flexible shapes without excessive stacking, improving efficiency. We also introduce learnable parameters to dynamically tailor the receptive field of the primitives, optimizing the balance between high-frequency details and smooth regions. Furthermore, we employ specialized densification strategies to further improve efficiency during tile computation. Experimental results show that TNT-GS outperforms state-of-the-art methods in storage efficiency and rendering speed, offering a robust solution for real-time rendering. The code of TNT-GS is available at https://github.com/GoogolplexGoodenough/TNT-GS.
Xiaofeng Liu 0001, Guanchen Meng, Chongyang Feng, Risheng Liu, Zhongxuan Luo, Xin Fan 0001
ACM Multimedia6
2025 Physics-Guided Sonar Image Fine-grained Recognition under Scarce Annotations
abstract
Sonar image recognition is a key technology in underwater exploration systems. Compared with natural images, sonar images have fewer texture details and are easily affected by heavy noise, making it more challenging for specialists to distinguish the subtle differences among classes. In view of this, studying fine-grained classification methods for sonar images with scarce annotations is of significant importance. To address this issue, we propose a Physics-Guided Teacher-Student (PGTS) framework to explore the unique physical information of sonar images while simultaneously mitigating the effects of limited annotations. First, PGTS reconstructs sonar signals through physical simulation and a specially designed physics-guided feature generation module, which allows it to bypass the time-consuming physical simulation during inference. Then, we design a multi-modal teacher model combines the reconstructed sonar signals and sonar images to extract discriminative features to generate robust pseudo labels for fine-grained target categories. Finally, the knowledge is transferred to a single-modal student model through consistency loss. Under the joint constraints of the teacher model and the reconstructed sonar physical signals, the student model continuously improves its performance in annotation-scarce scenarios. Notably, when merely 1% of the data is labeled, our method outperforms other state-of-the-art approaches by 12.46% in terms of accuracy.
Chengzhou Li, Qi Jia 0001, Jinyuan Liu 0001, Zhiying Jiang, Longhan Feng, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001
ACM Multimedia9
2025 Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark
abstract
We engage in the relatively underexplored task named thermal infrared image enhancement. Existing infrared image enhancement methods primarily focus on tackling individual degradations, such as noise, contrast, and blurring, making it difficult to handle coupled degradations. Meanwhile, all-in-one enhancement methods, commonly applied to RGB sensors, often demonstrate limited effectiveness due to the significant differences in imaging models. In sight of this, we first revisit the imaging mechanism and introduce a Recurrent Prompt Fusion Network (RPFN). Specifically, the RPFN initially establishes prompt pairs based on the thermal imaging process. For each type of degradation, we fuse the corresponding prompt pairs to modulate the model's features, providing adaptive guidance that enables the model to better address specific degradations under single or multiple conditions.In addition, a selective recurrent training mechanism is introduced to gradually refine the model's handling of composite cases to align the enhancement process, which not only allows the model to remove camera noise and retain key structural details, but also enhancing the overall contrast of the thermal image. Furthermore, we introduce the most comprehensive high-quality infrared benchmark covering a wide range of scenarios. Extensive experiments substantiate that our approach not only delivers promising visual results under specific degradation but also significantly improves performance on complex degradation scenes, achieving a notable 8.76% improvement.
Jinyuan Liu 0001, Zhu Liu 0004, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Risheng Liu
NeurIPS6
2025 Bilevel Fast Scene Adaptation for Low-Light Image Enhancement
Long Ma 0002, Dian Jin 0003, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
Int. J. Comput. Vis.5
2025 HUPE: Heuristic Underwater Perceptual Enhancement with Semantic Collaborative Learning
Zengxi Zhang, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
Int. J. Comput. Vis.5
2025 Infrared and Visible Image Fusion: From Data Compatibility to Task Adaption
abstract
Infrared-visible image fusion (IVIF) is a fundamental and critical task in the field of computer vision. Its aim is to integrate the unique characteristics of both infrared and visible spectra into a holistic representation. Since 2018, growing amount and diversity IVIF approaches step into a deep-learning era, encompassing introduced a broad spectrum of networks or loss functions for improving visual enhancement. As research deepens and practical demands grow, several intricate issues like data compatibility, perception accuracy, and efficiency cannot be ignored. Regrettably, there is a lack of recent surveys that comprehensively introduce and organize this expanding domain of knowledge. Given the current rapid development, this paper aims to fill the existing gap by providing a comprehensive survey that covers a wide array of aspects. Initially, we introduce a multi-dimensional framework to elucidate the prevalent learning-based IVIF methodologies, spanning topics from basic visual enhancement strategies to data compatibility, task adaptability, and further extensions. Subsequently, we delve into a profound analysis of these new approaches, offering a detailed lookup table to clarify their core ideas. Last but not the least, We also summarize performance comparisons quantitatively and qualitatively, covering registration, fusion and follow-up high-level tasks. Beyond delving into the technical nuances of these learning-based fusion approaches, we also explore potential future directions and open issues that warrant further exploration by the community.
Jinyuan Liu 0001, Guanyao Wu, Zhu Liu 0004, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Risheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 Learning With Self-Calibrator for Fast and Robust Low-Light Image Enhancement
abstract
Convolutional Neural Networks (CNNs) have shown significant success in the low-light image enhancement task. However, most of existing works encounter challenges in balancing quality and efficiency simultaneously. This limitation hinders practical applicability in real-world scenarios and downstream vision tasks. To overcome these obstacles, we propose a Self-Calibrated Illumination (SCI) learning scheme, introducing a new perspective to boost the model's capability. Based on a weight-sharing illumination estimation process, we construct an embedded self-calibrator to accelerate stage-level convergence, yielding gains that utilize only a single basic block for inference, which drastically diminishes computation cost. Additionally, by introducing the additivity condition on the basic block, we acquire a reinforced version dubbed SCI++, which disentangles the relationship between the self-calibrator and illumination estimator, providing a more interpretable and effective learning paradigm with faster convergence and better stability. We assess the proposed enhancers on standard benchmarks and in-the-wild datasets, confirming that they can restore clean images from diverse scenes with higher quality and efficiency. The verification on different levels of low-light vision tasks shows our applicability against other methods.
Long Ma 0002, Tengyu Ma 0004, Chengpei Xu, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 DRNet: Learning a dynamic recursion network for chaotic rain streak removal
Zhiying Jiang, Risheng Liu, Shuzhou Yang, Zengxi Zhang, Xin Fan 0001
Pattern Recognit.5
2025 Striving for Faster and Better: A One-Layer Architecture With Auto Re-Parameterization for Low-Light Image Enhancement
abstract
Deep learning-based low-light image enhancers have made significant progress in recent years, with a trend towards achieving satisfactory visual quality while gradually reducing the number of parameters and improving computational efficiency. In this work, we aim to delving into the limits of image enhancers both from visual quality and computational efficiency, while striving for both better performance and faster processing. To be concrete, by rethinking the task demands, we build an explicit connection, i.e., visual quality and computational efficiency are corresponding to model learning and structure design, respectively. Around this connection, we enlarge parameter space by introducing the re-parameterization for ample model learning of a pre-defined minimalist network (e.g., just one layer), to avoid falling into a local solution. To strengthen the structural representation, we define a hierarchical search scheme for discovering a task-oriented re-parameterized structure, which also provides powerful support for efficiency. Ultimately, this achieves efficient low-light image enhancement using only a single convolutional layer, while maintaining excellent visual quality. Experimental results show our sensible superiority both in quality and efficiency against recently-proposed methods. Especially, our running time on various platforms (e.g., CPU, GPU, NPU, DSP) consistently moves beyond the existing fastest scheme. The source code will be released athttps://github.com/vis-opt-group/AR-LLIE.
Long Ma 0002, Guangchao Han, Xin Fan 0001, Risheng Liu
IEEE Trans. Circuits Syst. Video Technol.4
2025 Edge Computing Underwater Optical Wireless Sensor Networks
abstract
Underwater Optical Wireless Sensor Networks (UOWSNs) play important roles in resource exploration and maritime rescue. However, they face significant challenges in real-time data transmission due to the limited propagation range of optical signals (typically 10-100 m), frequent link disconnections caused by node mobility, and the extended distances to onshore servers. Traditional cloud computing solutions, designed for stable terrestrial networks with stationary edge servers and continuous connectivity, experience high latency (3-15 s) in UOWSNs, rendering them unsuitable for real-time applications in underwater environments. To address this issue, we propose a cloud-edge-end architecture tailored for UOWSNs, which can not only combat unique underwater environmental interference on link connection and topological changes but also guarantee robust and real-time communication. We develop a dynamic link-stability-based task offloading path selection (DLS-TOPS) algorithm for maximizing network resource profits. Afterward, we propose an online primal-dual task offloading (OPD-TO) algorithm for minimizing task completion time. Simulation results indicate that the proposed method significantly improves the real-time performance and resource profits of the network, reducing the total task completion time by more than 50% compared to baseline algorithms. We implemented a UOWSN with a cloud-edge-end architecture using commercial off-the-shelf and verified the applicability and effectiveness of the proposed scheme in emergency detection through testbed experiments.
Yang Chi, Chi Lin 0001, Jing Deng 0001, Kaiwen Ning, Xin Fan 0001, Guowei Wu 0001
IEEE Trans. Mob. Comput.5
2025 Wireless Charging for Uncertain Location Nodes
abstract
Benefiting from Wireless Power Transfer (WPT) technology, Wireless Rechargeable Sensor Networks (WRSNs) effectively address the lifetime bottleneck of sensor nodes, enabling them to work perpetually. Most state-of-the-art studies assume that all WRSNs’ information is known or precise in advance. However, sensor nodes may be deployed randomly in a large-scale area, and some critical information (such as node location) may be unavailable or difficult to obtain precisely. In this work, we eliminate the effect of uncertain or imprecise node location and formalize theMaximizingChargingEnergy utility for uncertain location nodesproblem (i.e., MCE problem). With magnetic resonance coupling and beamforming technologies, we propose a novel node localization method to determine precise node location information. In addition, we present a reinforcement learning framework and a charging path scheduling method to maximize charging energy. To validate the effectiveness of our proposed scheme in real-world scenarios, we conduct test-bed experiments. The results demonstrate that our approach significantly improves charging efficiency by an average of 20.9% in a large-scale network, even when the locations of sensors are entirely unknown.
Chi Lin 0001, Shibo Hao, Yi Wang 0037, Lei Wang 0005, Xin Fan 0001, Guowei Wu 0001
IEEE Trans. Mob. Comput.7
2025 Compromising Rechargeable Sensor Networks in Marine Environment
abstract
Marine Wireless Rechargeable Sensor Networks (MWRSNs), enhanced by recent Wireless Power Transfer (WPT) technology, present a significant advancement in extending network life. Traditional methods improve network performance through algorithm optimization, but neglect charging security, exposing networks to potential attacks. This paper addresses this problem from an adversarial view and develops a novel attack for MWRSN through Denying of Charge (DoC) to maximize network destructiveness. We start by establishing a generalized on-demand charging model, essential for developing DoC tactics. Subsequently, we unveil the Collaborative DoC (CoDoC) algorithm, capable of manipulating and falsifying charging requests. Central to CoDoC is the Request Prediction Method (RPM), which forecasts the initiation of charging requests and facilitates rapid request surges to enhance the attack's efficacy. CoDoC is able to disguise the presence of the attack, which is able to escape from being detected by the base station. Theoretical analyses are provided to explore the features of the proposed scheme. To demonstrate the outperformed features of the proposed schemes, extensive simulations and test-bed experiments are conducted. Our analysis and extensive simulations demonstrate that CoDoC increases sensor node failures by 20% to 142% compared to traditional methods, highlighting its effectiveness in marine environments.
Chi Lin 0001, Haipeng Dai 0001, Mohammad S. Obaidat, Kuei-Fang Hsiao, Xin Fan 0001
IEEE Trans. Mob. Comput.6
2025 Robust One-Stop Multi-Modality Image Registration-Fusion-Segmentation Framework Against Misalignments and Adversarial Attacks
abstract
In complex open scenes, multi-modality image fusion and segmentation encounter two challenges: i) Imaging misalignments, manifested as pixel shifts and structural distortions, are perceptible. ii) Human-crafted adversarial attacks, reflected in pixel distribution variations, are imperceptible. They not only degrade the visual quality of fused images, e.g., noticeable edge ghosts but more critically undermine semantic perception. However, none of the existing works considered the coupled effect of these degradations. This paper proposes a One-Stop framework incorporating sequential task flows of “Registration-Fusion-Segmentation”, termed OS-RFS. Registration aims to mitigate the chained impact of misalignment on fusion and segmentation. We follow a coarse-to-fine registration paradigm and develop a Global-Local Incremental Registration (GLoIR) model, where the global shift registration (GSR) is performed initially for long-range pixel shifts, followed by incremental local deformation registration (LDR) for subtle local deformations. To improve segmentation robustness, we innovatively introduce auxiliary positive attacks and build a Cancellation Defense Strategy (CDS) in the fusion model. The CDS constrains the fusion model to fit fused images to the distribution of positive attacks, endowing fused images with a robust defense ability against adversarial attacks. This significantly mitigates the impact of adversarial attacks on semantic segmentation. Extensive experimental results reveal that our OS-RFS performs remarkable robustness on multi-modality image fusion and semantic segmentation against imaging misalignments and adversarial attacks.
Di Wang 0018, Xianghao Jiao, Jinyuan Liu 0001, Xin Fan 0001
IEEE Trans. Multim.4
2025 A Dual-Stream-Modulated Learning Framework for Illuminating and Super-Resolving Ultra-Dark Images
abstract
Enhancement of image resolution for scenes captured under extremely dim conditions represents a practical yet challenging problem that has received little attention. In such low-light scenarios, the limited lighting and minimal signal clarity tend to intensify issues such as diminished detail visibility and altered color accuracy, which are often more severe during the image enhancement process than in scenarios with adequate lighting. Consequently, standard methods for enhancing low-light images or improving their resolution, whether implemented independently or through a combined approach, generally face challenges in effectively restoring luminance, preserving color integrity, and detailing intricate features. To conquer these issues, this article introduces an innovative dual-stream (DS) modulated learning framework designed to tackle the real-world coupled degradation issues in super-resolution (SR) under low-light conditions. Leveraging natural image color characteristics, we introduce a self-regularized luminance constraint to specifically target uneven illumination. We develop illumination-semantic dual modulator (ISDM), a refinement middleware embedded in the decoding stage to bridge illumination and semantic features concurrently, aimed at safeguarding the integrity of lighting and color details at the feature level. Our approach replaces simple upsampling methods with the resolution-sensitive merging upsampler (RSMU) module, which integrates diverse sampling techniques to effectively reduce artifacts and halo effects. Comprehensive experiments on three benchmarks showcase the applicability and generalizability of our approach to diverse and challenging ultra-poorly lit settings, outperforming state-of-the-art methods with a notable improvement. The code and benchmark are publicly available at https://github.com/moriyaya/UltraIS.
Jiaxin Gao 0001, Ziyu Yue, Sihan Xie, Xin Fan 0001, Risheng Liu
IEEE Trans. Neural Networks Learn. Syst.5
2024 Trash to Treasure: Low-Light Object Detection via Decomposition-and-Aggregation
abstract
Object detection in low-light scenarios has attracted much attention in the past few years. A mainstream and representative scheme introduces enhancers as the pre-processing for regular detectors. However, because of the disparity in task objectives between the enhancer and detector, this paradigm cannot shine at its best ability. In this work, we try to arouse the potential of enhancer + detector. Different from existing works, we extend the illumination-based enhancers (our newly designed or existing) as a scene decomposition module, whose removed illumination is exploited as the auxiliary in the detector for extracting detection-friendly features. A semantic aggregation module is further established for integrating multi-scale scene-related semantic information in the context space. Actually, our built scheme successfully transforms the "trash" (i.e., the ignored illumination in the detector) into the "treasure" for the detector. Plenty of experiments are conducted to reveal our superiority against other state-of-the-art methods. The code will be public if it is accepted.
Xiaohan Cui, Long Ma 0002, Tengyu Ma 0004, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
AAAI5
2024 Towards Robust Image Stitching: An Adaptive Resistance Learning against Compatible Attacks
abstract
Image stitching seamlessly integrates images captured from varying perspectives into a single wide field-of-view image. Such integration not only broadens the captured scene but also augments holistic perception in computer vision applications. Given a pair of captured images, subtle perturbations and distortions which go unnoticed by the human visual system tend to attack the correspondence matching, impairing the performance of image stitching algorithms. In light of this challenge, this paper presents the first attempt to improve the robustness of image stitching against adversarial attacks. Specifically, we introduce a stitching-oriented attack (SoA), tailored to amplify the alignment loss within overlapping regions, thereby targeting the feature matching procedure. To establish an attack resistant model, we delve into the robustness of stitching architecture and develop an adaptive adversarial training (AAT) to balance attack resistance with stitching precision. In this way, we relieve the gap between the routine adversarial training and benign models, ensuring resilience without quality compromise. Comprehensive evaluation across real-world and synthetic datasets validate the deterioration of SoA on stitching performance. Furthermore, AAT emerges as a more robust solution against adversarial perturbations, delivering superior stitching results. Code is available at: https://github.com/Jzy2017/TRIS.
Zhiying Jiang, Xingyuan Li 0005, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
AAAI4
2024 Hybrid-Supervised Dual-Search: Leveraging Automatic Learning for Loss-Free Multi-Exposure Image Fusion
abstract
Multi-exposure image fusion (MEF) has emerged as a prominent solution to address the limitations of digital imaging in representing varied exposure levels. Despite its advancements, the field grapples with challenges, notably the reliance on manual designs for network structures and loss functions, and the constraints of utilizing simulated reference images as ground truths. Consequently, current methodologies often suffer from color distortions and exposure artifacts, further complicating the quest for authentic image representation. In addressing these challenges, this paper presents a Hybrid-Supervised Dual-Search approach for MEF, dubbed HSDS-MEF, which introduces a bi-level optimization search scheme for automatic design of both network structures and loss functions. More specifically, we harness a unique dual research mechanism rooted in a novel weighted structure refinement architecture search. Besides, a hybrid supervised contrast constraint seamlessly guides and integrates with searching process, facilitating a more adaptive and comprehensive search for optimal loss functions. We realize the state-of-the-art performance in comparison to various competitive schemes, yielding a 10.61% and 4.38% improvement in Visual Information Fidelity (VIF) for general and no-reference scenarios, respectively, while providing results with high contrast, rich details and colors. The code is available at https://github.com/RollingPlain/HSDS_MEF.
Guanyao Wu, Hongming Fu, Jinyuan Liu 0001, Long Ma 0002, Xin Fan 0001, Risheng Liu
AAAI5
2024 Construction of Simulated 3D CT from Multiple-View X-ray Images
Liang Zhao 0005, Sijia Hou, Xin Fan 0001, Zhikui Chen, Shuqiong Wu
BIBM4
2024 Bi-level Learning of Task-Specific Decoders for Joint Registration and One-Shot Medical Image Segmentation
abstract
One-shot medical image segmentation (MIS) aims to cope with the expensive, time-consuming, and inherent human bias annotations. One prevalent method to address one-shot MIS is joint registration and segmentation (JRS) with a shared encoder, which mainly explores the voxel-wise correspondence between the labeled data and unlabeled data for better segmentation. However, this method omits underlying connections between task-specific decoders for segmentation and registration, leading to unstable training. In this paper, we propose a novel Bi-level Learning of Task-Specific Decoders for one-shot MIS, employing a pretrained fixed shared encoder that is proved to be more quickly adapted to brand-new datasets than existing JRS without fixed shared encoder paradigm. To be more specific, we introduce a bi-level optimization training strategy considering registration as a major objective and segmentation as a learnable constraint by leveraging inter-task coupling dependencies. Furthermore, we design an appearance conformity constraint strategy that learns the backward transformations generating the fake labeled data used to perform data augmentation instead of the labeled image, to avoid performance degradation caused by inconsistent styles between unlabeled data and labeled data in previous methods. Extensive experiments on the brain MRI task across ABIDE, ADNI, and PPMI datasets demonstrate that the proposed Bi-JROS outperforms state-of-the-art one-shot MIS methods for both segmentation and registration tasks. The code will be available at https://github.com/Coradlut/Bi-JROS.
Xin Fan 0001, Jiaxin Gao 0001, Jia Wang 0036, Zhongxuan Luo, Risheng Liu
CVPR1
2024 Contourlet Residual for Prompt Learning Enhanced Infrared Image Super-Resolution
Xingyuan Li 0005, Jinyuan Liu 0001, Zhixin Chen, Yang Zou 0004, Long Ma 0002, Xin Fan 0001, Risheng Liu
ECCV (3)6
2024 Depth-Guided Dominant Plane Perception for Unsupervised Homography Estimation
abstract
Homography describes the mapping relations of the same plane across views. In scenarios with multiple planes, single homography estimation aims to obtain the optimal solution generated by the largest consistent plane to obey the coplanar constraints. However, existing methods typically consider all planes equally, neglecting the negative impact of regions that differ significantly from the largest approximate planar areas (dominant plane). In this work, we propose a depth-guided dominant plane perception network to achieve unsupervised homography estimation with additional attention on the dominant plane. Specifically, we leverage the depth-wise prior to adaptively detecting the approximate dominant plane, invoking essential scene structures for unsupervised homography estimation. Then, we enhance the corresponding features of the dominant plane and explore their correlations through a specially designed perceptual module. Finally, we employ dominant plane perception on multi-scale features progressively to estimate the homography in a coarse-to-fine manner. Extensive experiments on a large parallax dataset demonstrate that our method improves the alignment performance by 10.29%, yielding more accurate alignment than previous competitive methods.
Xiaomei Feng, Qi Jia 0001, Yu Liu 0012, Xin Fan 0001, Longin Jan Latecki
ICASSP4
2024 CSCNet: Class-Specified Cascaded Network for Compositional Zero-Shot Learning
abstract
Attribute and object (A-O) disentanglement is a fundamental and critical problem for Compositional Zero-shot Learning (CZSL), whose aim is to recognize novel A-O compositions based on foregone knowledge. Existing methods based on disentangled representation learning lose sight of the contextual dependency between the A-O primitive pairs. Inspired by this, we propose a novel A-O disentangled framework for CZSL, namely Class-specified Cascaded Network (CSC-Net). The key insight is to firstly classify one primitive and then specifies the predicted class as a priori for guiding another primitive recognition in a cascaded fashion. To this end, CSCNet constructs Attribute-to-Object and Object-to- Attribute cascaded branches, in addition to a composition branch modeling the two primitives as a whole. Notably, we devise a parametric classifier (ParamCls) to improve the matching between visual and semantic embeddings. By improving the A-O disentanglement, our framework achieves superior results than previous competitive methods.
Yanyi Zhang, Qi Jia 0001, Xin Fan 0001, Yu Liu 0012
ICASSP3
2024 Impossible Trinity in Underwater Optical Wireless Communication
abstract
Underwater Optical Wireless Communication (UOWC) is considered a promising approach, offering the potential for flexible and high-speed communication under the surface of the water. However, the interdependent relationship among three key performance elements, namely communication distance, bit error rate, and communication rate, has been largely overlooked. This oversight impedes the complete utilization of the system performance. In this work, we innovatively introduce a “UOWC Impossible Trinity” model and theorems to establish relationships among the three key performance elements, which clarify the inherent constraints within UOWC system optimization. Moreover, we formulate the Underwater Optical Communication Trade-offs (UOCT) Problem to maximize communication performance. Furthermore, we provide feasible non-dominated solution sets, considering the constraints of real environments and user demands of specific scenarios. Our model has been validated by extensive simulations, demonstrating that our approach not only clarifies fundamental limitations of UOWC systems, but also provides practical guidelines for designing and optimizing the systems. Our approach has been experimentally validated with an impressive accuracy of over 95%, surpassing conventional models, which not only enhances the understanding of UOWC system optimization but also validates the existence of inherent trade-offs. Furthermore, our approach demonstrates a significant increase in communication distance, outperforming traditional methods by more than 20%.
Chi Lin 0001, Yi Wang 0037, Yu Sun 0077, Lei Wang 0005, Xin Fan 0001, Guowei Wu 0001
ICNP6
2024 Where Elegance Meets Precision: Towards a Compact, Automatic, and Flexible Framework for Multi-modality Image Fusion and Applications
Jinyuan Liu 0001, Guanyao Wu, Zhu Liu 0004, Long Ma 0002, Risheng Liu, Xin Fan 0001
IJCAI6
2024 Advancing Generalized Transfer Attack with Initialization Derived Bilevel Optimization and Dynamic Sequence Truncation
Jiaxin Gao 0001, Xuan Liu 0011, Xianghao Jiao, Xin Fan 0001, Risheng Liu
IJCAI5
2024 UWBeacon: Lighting up Centimeter-Level Underwater Positioning
abstract
Underwater positioning plays a key role in many underwater operations. This paper presents the design, implementation, and evaluation of UWBeacon, a centimeter-level visible light-based underwater positioning system. UWBeacon consists of LED beacons as the light signal transmitter and a camera-based receiver as the target. To address unique challenges in underwater environment such as limited visibility and strong ambient interference, we exploit a novel design that utilizes polarized lights of different colors with different polarization angles for background subtraction. UWBeacon is implemented with commercial-off-the-shelf LEDs and cameras. Comprehensive experiments conducted in various real underwater environments show that UWBeacon can achieve a mean positioning error below 6 cm and an orientation error below 1.5° at a distance of 10 meters.
Chi Lin 0001, Jie Xiong 0001, Lei Wang 0005, Guowei Wu 0001, Xin Fan 0001, Zhongxuan Luo
MobiCom7
2024 CoCoNet: Coupled Contrastive Learning Network with Multi-level Feature Ensemble for Multi-modality Image Fusion
Jinyuan Liu 0001, Runjia Lin, Guanyao Wu, Risheng Liu, Zhongxuan Luo, Xin Fan 0001
Int. J. Comput. Vis.6
2024 Breaking the water dilemma: Transmission-guided bilevel adaptive learning for underwater imagery
Sihan Xie, Peiming Li, Jiaxin Gao 0001, Ziyu Yue, Xin Fan 0001, Risheng Liu
Neurocomputing5
2024 Learning With Constraint Learning: New Perspective, Solution Strategy and Various Applications
abstract
The complexity of learning problems, such as Generative Adversarial Network (GAN) and its variants, multi-task and meta-learning, hyper-parameter learning, and a variety of real-world vision applications, demands a deeper understanding of their underlying coupling mechanisms. Existing approaches often address these problems in isolation, lacking a unified perspective that can reveal commonalities and enable effective solutions. Therefore, in this work, we proposed a new framework, named Learning with Constraint Learning (LwCL), that can holistically examine challenges and provide a unified methodology to tackle all the above-mentioned complex learning and vision problems. Specifically, LwCL is designed as a general hierarchical optimization model that captures the essence of these diverse learning and vision problems. Furthermore, we develop a gradient-response based fast solution strategy to overcome optimization challenges of the LwCL framework. Our proposed framework efficiently addresses a wide range of applications in learning and vision, encompassing three categories and nine different problem types. Extensive experiments on synthetic tasks and real-world applications verify the effectiveness of our approach. The LwCL framework offers a comprehensive solution for tackling complex machine learning and computer vision problems, bridging the gap between theory and practice.
Risheng Liu, Jiaxin Gao 0001, Xuan Liu 0011, Xin Fan 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 A Task-Guided, Implicitly-Searched and Meta-Initialized Deep Model for Image Fusion
abstract
Image fusion plays a key role in a variety of multi-sensor-based vision systems, especially for enhancing visual quality and/or extracting aggregated features for perception. However, most existing methods just consider image fusion as an individual task, thus ignoring its underlying relationship with these downstream vision problems. Furthermore, designing proper fusion architectures often requires huge engineering labor. It also lacks mechanisms to improve the flexibility and generalization ability of current fusion approaches. To mitigate these issues, we establish a Task-guided, Implicit-searched and Meta-initialized (TIM) deep model to address the image fusion problem in a challenging real-world scenario. Specifically, we first propose a constrained strategy to incorporate information from downstream tasks to guide the unsupervised learning process of image fusion. Within this framework, we then design an implicit search scheme to automatically discover compact architectures for our fusion model with high efficiency. In addition, a pretext meta initialization technique is introduced to leverage divergence fusion data to support fast adaptation for different kinds of image fusion tasks. Qualitative and quantitative experimental results on different categories of image fusion problems and related downstream tasks (e.g., visual enhancement and semantic understanding) substantiate the flexibility and effectiveness of our TIM.
Risheng Liu, Zhu Liu 0004, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Collaborative brightening and amplification of low-light imagery via bi-level adversarial learning
Jiaxin Gao 0001, Ziyu Yue, Xin Fan 0001, Risheng Liu
Pattern Recognit.4
2024 WBNet: Weakly-supervised salient object detection via scribble and pseudo-background priors
abstract
Weakly supervised salient object detection (WSOD) methods endeavor to boost sparse labels to get more salient cues in various ways. Among them, an effective approach is using pseudo labels from multiple unsupervised self-learning methods, but inaccurate and inconsistent pseudo labels could ultimately lead to detection performance degradation. To tackle this problem, we develop a new multi-source WSOD framework, WBNet, that can effectively utilize pseudo-background (non-salient region) labels combined with scribble labels to obtain more accurate salient features. We first design a comprehensive salient pseudo-mask generator from multiple self-learning features. Then, we pioneer the exploration of generating salient pseudo-labels via point-prompted and box-prompted Segment-Anything Models (SAM). Then, WBNet leverages a pixel-level Feature Aggregation Module (FAM), a mask-level Transformer-decoder (TFD), and an auxiliary Boundary Prediction Module (EPM) with a hybrid loss function to handle complex saliency detection tasks. Comprehensively evaluated with state-of-the-art methods on five widely used datasets, the proposed method significantly improves saliency detection performance. The code and results are publicly available at https://github.com/yiwangtz/WBNet.
Yi Wang 0037, Ruili Wang 0001, Xiangjian He, Chi Lin 0001, Tianzhu Wang, Qi Jia 0001, Xin Fan 0001
Pattern Recognit.7
2024 Searching a Compact Architecture for Robust Multi-Exposure Image Fusion
abstract
In recent years, learning-based methods have achieved significant advancements in multi-exposure image fusion. However, two major stumbling blocks hinder the development, including pixel misalignment and inefficient inference. Reliance on aligned image pairs in existing methods causes susceptibility to artifacts due to device motion. Additionally, existing techniques often rely on handcrafted architectures with huge network engineering, resulting in redundant parameters, adversely impacting inference efficiency and flexibility. To mitigate these limitations, this study introduces an architecture search-based paradigm incorporating self-alignment and detail repletion modules for robust multi-exposure image fusion. Specifically, targeting the extreme discrepancy of exposure, we propose the self-alignment module, leveraging scene relighting to constrain the illumination degree for following alignment and feature extraction. Detail repletion is proposed to enhance the texture details of scenes. Additionally, incorporating a hardware-sensitive constraint, we present the fusion-oriented architecture search to explore compact and efficient networks for fusion. The proposed method outperforms various competitive schemes, achieving a noteworthy 3.19% improvement in PSNR for general scenarios and an impressive 23.5% enhancement in misaligned scenarios. Moreover, it significantly reduces inference time by 69.1%. The code will be available at https://github.com/LiuZhu-CV/CRMEF.
Zhu Liu 0004, Jinyuan Liu 0001, Guanyao Wu, Xin Fan 0001, Risheng Liu
IEEE Trans. Circuits Syst. Video Technol.5
2024 Edge-Aware Correlation Learning for Unsupervised Progressive Homography Estimation
abstract
Homography estimation aligns image pairs in cross-views, which is a crucial and fundamental computer vision problem. Existing methods only consider correspondences of texture features for homography estimation, leading to unpleasant artifacts and misalignments introduced by mismatches, especially for low-texture image pairs. In contrast to others, we introduce intuitive structural information as an additional clue that is more sensitive to human vision and low-texture scenarios. In this paper, we propose an edge-aware unsupervised progressive network that couples texture and edge correlation to comprehensively explore potential matching features for homography estimation. To explore robust edge and texture features, we employ a multiscale network to capture feature pyramids with different receptive fields. Then, we design an edge-aware correlation module tailored for homography regression, which plugs in multiscale features to capture accurate correlation maps. Specifically, the edge-aware correlation module leverages the feature-selecting strategy for edge features to capture discriminative matching edges and further guides the texture correlation unit to focus on correctly matched textures. Finally, we leverage multiscale edge-aware correlation maps to predict homography progressively from coarse to fine. Experimental results demonstrate that our proposed method improves PSNR by 11.09% on the real large parallax dataset and reduces matching error by 32.04% on the synthetic COCO dataset, yielding more accurate alignment results than previous state-of-the-art methods.
Xiaomei Feng, Qi Jia 0001, Zikun Zhao, Yu Liu 0012, Xinwei Xue, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Learning Cruxes to Push for Object Detection in Low-Quality Images
abstract
Highly degraded images greatly challenge existing algorithms to detect objects of interest in adverse scenarios, such as rain, fog, and underwater. Recently, researchers develop sophisticated deep architectures in order to enhance image quality. Unfortunately, the visually appealing output of the enhancement module does not necessarily generate high accuracy for deep detectors. Another feasible solution for low-quality image detection is to transform it into a domain adaptation problem. Typically, these approaches invoke complicated training strategies such as adversarial learning and graph matching. False detection is likely to occur in local regions of a low-quality image. In this paper, we propose a simple yet effective strategy with two learners for low-quality image detection. We devise the crux learner to generate cruxes that have great impacts on detection performance. The catch-up leaner with a simple residual transfer mechanism maps the feature distributions of crux regions to those favouring a deep detector. These two learners can be plugged into any CNN-based feature extraction networks, e.g., ResNetXT101 and ResNet50, and yield high detection accuracy on various degraded scenarios. Extensive experiments on several public datasets demonstrate that our method achieves more promising results than state-of-the-art detection approaches. The codes:https://github.com/xiaoDetection/learning-cruxes-to-push.
Chenping Fu, Jiewen Xiao, Wanqi Yuan, Risheng Liu, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Learning to Search a Lightweight Generalized Network for Medical Image Fusion
abstract
Image fusion is indispensable in a comprehensive medical imaging pipeline. By embracing deep learning technology, medical image fusion has achieved tremendous progress over the past few years. However, existing approaches make efforts on the specific type of medical image fusion task and may face difficulties in generalizing well. Moreover, most of them strain every nerve to design various architectures with an increase of the width of depth, placing an obstacle in running efficiency. To address the above problems, we propose an Auto-searching Light-weighted Multi-source Fusion network, namely ALMFnet, aiming at incorporating both software and hardware knowledge in a network architecture searching manner for medical image fusion. Specifically, the ALMFnet, consisting of two different feature-extracting modules and one fusion module, is developed to extract and refine multi-source features in a generalized model. Besides, motivated by the collaborative principle, we introduce hardware constraints for sufficient searching the each particular component, further reducing the complexity of the obtained model. Furthermore, to preserve important details in pathological image areas, we introduce a segmentation mask into the developed method. Experimental results demonstrate that our generalized model outperforms previous methods not only in terms of quantitative scores but also in model complexity. Source code will be available at https://github.com/RollingPlain/ALMFnet.
Pan Mu, Guanyao Wu, Jinyuan Liu 0001, Yuduo Zhang, Xin Fan 0001, Risheng Liu
IEEE Trans. Circuits Syst. Video Technol.5
2024 Improving Misaligned Multi-Modality Image Fusion With One-Stage Progressive Dense Registration
abstract
Misalignments between multi-modality images pose challenges in image fusion, manifesting as structural distortions and edge ghosts. Existing efforts commonly resort to registering first and fusing later, typically employing two separate stages for registration, i.e., coarse registration and fine registration. Both stages directly estimate the respective target deformation fields. This paper contends that the separate two-stage registration lacks compactness, and the direct estimation of their target deformation fields falls short in accuracy. To tackle these challenges, we introduce IMF, a framework for improving misaligned multi-modality image fusion. Central to IMF is a One-stage Progressive Dense Registration (OPDR) scheme, which accomplishes the coarse-to-fine registration through only a one-stage optimization. Specifically, two pivotal components are involved in OPDR, a dense Deformation Field Fusion (DFF) module and a Progressive Feature Fine (PFF) module. The DFF aggregates the predicted multi-scale deformation sub-fields at the current scale, while the PFF progressively refines the remaining misaligned features. Together, they effectively and accurately estimate the final deformation fields. In addition, we develop a Transformer-Conv-based Fusion (TCF) subnetwork that considers local and long-range feature dependencies, allowing us to capture more informative features from the registered infrared and visible images for the generation of high-quality fused images. Extensive experimental analysis demonstrates the superiority of the proposed method in the fusion of misaligned cross-modality images. The code will be available athttps://github.com/wdhudiekou/IMF.
Di Wang 0018, Jinyuan Liu 0001, Long Ma 0002, Risheng Liu, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Multispectral Image Stitching via Global-Aware Quadrature Pyramid Regression
abstract
Image stitching is a critical task in panorama perception that involves combining images captured from different viewing positions to reconstruct a wider field-of-view (FOV) image. Existing visible image stitching methods suffer from performance drops under severe conditions since environmental factors can easily impair visible images. In contrast, infrared images possess greater penetrating ability and are less affected by environmental factors. Therefore, we propose an infrared and visible image-based multispectral image stitching method to achieve all-weather, broad FOV scene perception. Specifically, based on two pairs of infrared and visible images, we employ the salient structural information from the infrared images and the textual details from the visible images to infer the correspondences within different modality-specific features. For this purpose, a multiscale progressive mechanism coupled with quadrature correlation is exploited to improve regression in different modalities. Exploiting the complementary properties, accurate and credible homography can be obtained by integrating the deformation parameters of the two modalities to compensate for the missing modality-specific information. A global-aware guided reconstruction module is established to generate an informative and broad scene, wherein the attentive features of different viewpoints are introduced to fuse the source images with a more seamless and comprehensive appearance. We construct a high-quality infrared and visible stitching dataset for evaluation, including real-world and synthetic sets. The qualitative and quantitative results demonstrate that the proposed method outperforms the intuitive cascaded fusion-stitching procedure, achieving more robust and credible panorama generation. Code and dataset are available at https://github.com/Jzy2017/MSGA.
Zhiying Jiang, Zengxi Zhang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
IEEE Trans. Image Process.4
2024 AirWrite: An Aerial Handwriting Trajectory Tracking and Recognition System With mmWave
abstract
In the field of human-computer interaction (HCI), handwriting trajectory tracking and recognition have attracted significant attention due to their wide range of applications. However, many existing approaches rely on handheld devices and are highly susceptible to factors such as environmental conditions, location, and writing style. To overcome these limitations, we propose AirWrite, a novel contactless aerial system for handwriting trajectory tracking and recognition using mmWave technology. We introduce a signal clipping method based on the Doppler effect caused by user actions to accurately remove non-handwriting signals in the time domain. Additionally, we analyze power variations within the signal frequency interval to determine the handwriting frequency and employ a band-pass filter to eliminate dynamic environmental noise effectively. Through extensive experiments, we demonstrate that AirWrite can precisely track handwriting trajectories in noisy environments regardless of distance, angle, handwriting speed, character size, or in the presence of obstacles. Furthermore, we present an effective handwritten character recognition method for AirWrite that recognizes alphabets, numbers, and words. AirWrite can achieve an average accuracy of over 96% with only a 34 KB small dataset within 0.15 s for recognition.
Chi Lin 0001, Zhouhe Sun, Asfandeyar Ahmad, Xinxin Fan, Yi Wang 0037, Lei Wang 0005, Xin Fan 0001, Guowei Wu 0001
IEEE Trans. Mob. Comput.7
2024 Hierarchical Similarity Learning for Aliasing Suppression Image Super-Resolution
abstract
As a highly ill-posed issue, single-image super-resolution (SISR) has been widely investigated in recent years. The main task of SISR is to recover the information loss caused by the degradation procedure. According to the Nyquist sampling theory, the degradation leads to the aliasing effect and makes it hard to restore the correct textures from low-resolution (LR) images. In practice, there are correlations and self-similarities among the adjacent patches in the natural images. This article considers the self-similarity and proposes a hierarchical image super-resolution network (HSRNet) to suppress the influence of aliasing. We consider the SISR issue in the optimization perspective and propose an iterative solution pattern based on the half-quadratic splitting (HQS) method. To explore the texture with local image prior, we design a hierarchical exploration block (HEB) and progressive increase the receptive field. Furthermore, multilevel spatial attention (MSA) is devised to obtain the relations of adjacent feature and enhance the high-frequency information, which acts as a crucial role for visual experience. The experimental result shows that HSRNet achieves better quantitative and visual performance than other works and remits the aliasing more effectively.
Yuqing Liu 0001, Qi Jia 0001, Jian Zhang 0018, Xin Fan 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 AutoDLAR: A Semi-supervised Cross-modal Contact-free Human Activity Recognition System
abstract
WiFi-based human activity recognition (HAR) plays an essential role in various applications such as security surveillance, health monitoring, and smart home. Existing HAR methods, though yielding promising performance in indoor scenarios, highly depend on a massive labeled dataset for training which is extremely difficult to acquire in practical applications. In this paper, we present an automatic data labeling and HAR system, termed AutoDLAR. Taking a semi-supervised cross-modal learning framework with a hybrid loss function as the core, AutoDLAR transfers rich visual information to automatically label WiFi signals for WiFi-based HAR. Specifically, we devise a lightweight and multi-view WiFi sensing model with a parallel feature embedding method to accurately identify activities and accelerate recognition speed. Then, we exploit the video data to fine-tune a well-established visual HAR model, generating effective pseudo-labels for guiding the WiFi model’s training. We also build a synchronized Video-WiFi dataset with seven types of human activities under different scenarios to enable training and validating the semi-supervised HAR system. Extensive experiments on our collected activity dataset and the emotion recognition benchmark demonstrate that AutoDLAR attains an average accuracy of over 95.89% without manual labeling and only spends the inference time of 3.35 ms, outperforming the state-of-the-art (SOTA) methods.
Xinxin Lu, Lei Wang 0005, Chi Lin 0001, Xin Fan 0001, Zhenquan Qin
ACM Trans. Sens. Networks4
2024 A rotation robust shape transformer for cartoon character recognition
Qi Jia 0001, Yi Wang 0037, Xin Fan 0001, Haibin Ling, Longin Jan Latecki
Vis. Comput.4
2023 Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection
abstract
Salient object detection (SOD) aims to mimic the human visual system (HVS) and cognition mechanisms to identify and segment salient objects. However, due to the complexity of these mechanisms, current methods are not perfect. Accuracy and robustness need to be further improved, particularly in complex scenes with multiple objects and background clutter. To address this issue, we propose a novel approach called Multiple Enhancement Network (MENet) that adopts the boundary sensibility, content integrity, iterative refinement, and frequency decomposition mechanisms of HVS. A multi-level hybrid loss is firstly designed to guide the network to learn pixel-level, region-level, and object-level features. A flexible multiscale feature enhancement module (ME-Module) is then designed to gradually aggregate and refine global or detailed features by changing the size order of the input feature sequence. An iterative training strategy is used to enhance boundary features and adaptive features in the dual-branch decoder of MENet. Comprehensive evaluations on six challenging benchmark datasets show that MENet achieves state-of-the-art results. Both the codes and results are publicly available at https://github.com/yiwangtz/MENet.
Yi Wang 0037, Ruili Wang 0001, Xin Fan 0001, Tianzhu Wang, Xiangjian He
CVPR3
2023 Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation
abstract
Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, e.g., fusion or segmentation, making it hard to reach ‘Best of Both Worlds’. To overcome this issue, in this paper, we propose a Multi-interactive Feature learning architecture for image fusion and Segmentation, namely SegMiF, and exploit dual-task correlation to promote the performance of both tasks. The SegMiF is of a cascade structure, containing a fusion sub-network and a commonly used segmentation sub-network. By slickly bridging intermediate features between two components, the knowledge learned from the segmentation task can effectively assist the fusion task. Also, the benefited fusion network supports the segmentation one to perform more pretentiously. Besides, a hierarchical interactive attention block is established to ensure fine-grained mapping of all the vital information between two tasks, so that the modality/semantic features can be fully mutual-interactive. In addition, a dynamic weight factor is introduced to automatically adjust the corresponding weights of each task, which can balance the interactive feature correspondence and break through the limitation of laborious tuning. Furthermore, we construct a smart multi-wave binocular imaging system and collect a full-time multi-modality benchmark with 15 annotated pixel-level categories for image fusion and segmentation. Extensive experiments on several public datasets and our benchmark demonstrate that the proposed method outputs visually appealing fused images and perform averagely 7.66% higher segmentation mIoU in the real-world scene than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/JinyuanLiu-CV/SegMiF.
Jinyuan Liu 0001, Zhu Liu 0004, Guanyao Wu, Long Ma 0002, Risheng Liu, Zhongxuan Luo, Xin Fan 0001
ICCV8
2023 Bi-level Dynamic Learning for Jointly Multi-modality Image Fusion and Beyond
abstract
Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and neglecting others, seldom investigating their underlying connections for joint promotion. To overcome these limitations, we establish the hierarchical dual tasks-driven deep model to bridge these tasks. Concretely, we firstly construct an image fusion module to fuse complementary characteristics and cascade dual task-related modules, including a discriminator for visual effects and a semantic network for feature measurement. We provide a bi-level perspective to formulate image fusion and follow-up downstream tasks. To incorporate distinct task-related responses for image fusion, we consider image fusion as a primary goal and dual modules as learnable constraints. Furthermore, we develop an efficient first-order approximation to compute corresponding gradients and present dynamic weighted aggregation to balance the gradients for fusion learning. Extensive experiments demonstrate the superiority of our method, which not only produces visually pleasant fused results but also realizes significant promotion for detection and segmentation than the state-of-the-art approaches.
Zhu Liu 0004, Jinyuan Liu 0001, Guanyao Wu, Long Ma 0002, Xin Fan 0001, Risheng Liu
IJCAI5
2023 Learning Pixel-wise Alignment for Unsupervised Image Stitching
abstract
Image stitching aims to align a pair of images in the same view. Generating precise alignment with natural structures is challenging for image stitching, as there is no wider field-of-view image as a reference, especially in non-coplanar practical scenarios. In this paper, we propose an unsupervised image stitching framework, breaking through the coplanar constraints in homography estimation, yielding accurate pixel-wise alignment under limited overlapping regions. First, we generate a global transformation by an iterative dense feature matching combined with an error control strategy to alleviate the difference introduced by large parallax. Second, we propose a pixel-wise warping network embedded within a large-scale feature extractor and a correlative feature enhancement module to explicitly learn correspondences between the inputs, and generate accurate pixel-level offsets upon novel constraints on both overlapping and non-overlapping regions. Notably, we leverage the pixel-level offsets in the overlapping area to guide the adjustment in the non-overlapping area upon content and structure consistency constraints, rendering a natural transition between two regions and distortions suppression over the entire stitched image. The proposed method achieves state-of-the-art performance that surpasses both traditional and deep learning approaches by a large margin. It also achieves the shortest execution time and has the best generalization ability on the traditional dataset.
Qi Jia 0001, Xiaomei Feng, Yu Liu 0012, Xin Fan 0001, Longin Jan Latecki
ACM Multimedia4
2023 Multi-Spectral Image Stitching via Spatial Graph Reasoning
abstract
Multi-spectral image stitching leverages the complementarity between infrared and visible images to generate a robust and reliable wide field-of-view~(FOV) scene. The primary challenge of this task is to explore the relations between multi-spectral images for aligning and integrating multi-view scenes. Capitalizing on the strengths of Graph Convolutional Networks (GCNs) in modeling feature relationships, we propose a spatial graph reasoning based multi-spectral image stitching method that effectively distills the deformation and integration of multi-spectral images across different viewpoints. To accomplish this, we embed multi-scale complementary features from the same view position into a set of nodes. The correspondence across different views is learned through powerful dense feature embeddings, where both inter- and intra-correlations are developed to exploit cross-view matching and enhance inner feature disparity. By introducing long-range coherence along spatial and channel dimensions, the complementarity of pixel relations and channel interdependencies aids in the reconstruction of aligned multi-view features, generating informative and reliable wide FOV scenes. Moreover, we release a challenging dataset named ChaMS, comprising both real-world and synthetic sets with significant parallax, providing a new option for comprehensive evaluation. Extensive experiments demonstrate that our method surpasses the state-of-the-arts.
Zhiying Jiang, Zengxi Zhang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
ACM Multimedia4
2023 PEARL: Preprocessing Enhanced Adversarial Robust Learning of Image Deraining for Semantic Segmentation
abstract
In light of the significant progress made in the development and application of semantic segmentation tasks, there has been increasing attention towards improving the robustness of segmentation models against natural degradation factors (e.g., rain streaks) or artificially attack factors (e.g., adversarial attack). Whereas, most existing methods are designed to address a single degradation factor and are tailored to specific application scenarios. In this work, we present the first attempt to improve the robustness of semantic segmentation tasks by simultaneously handling different types of degradation factors. Specifically, we introduce the Preprocessing Enhanced Adversarial Robust Learning (PEARL) framework based on the analysis of our proposed Naive Adversarial Training (NAT) framework. Our approach effectively handles both rain streaks and adversarial perturbation by transferring the robustness of the segmentation model to the image derain model. Furthermore, as opposed to the commonly used Negative Adversarial Attack (NAA), we design the Auxiliary Mirror Attack (AMA) to introduce positive information prior to the training of the PEARL framework, which improves defense capability and segmentation performance. Our extensive experiments and ablation studies based on different derain methods and segmentation models have demonstrated the significant performance improvement of PEARL with AMA in defense against various adversarial attacks and rain streaks while maintaining high generalization performance across different datasets. The source codes are available at https://github.com/JiaoXianghao/PEARL.
Xianghao Jiao, Jiaxin Gao 0001, Xinyuan Chu, Xin Fan 0001, Risheng Liu
ACM Multimedia5
2023 Fearless Luminance Adaptation: A Macro-Micro-Hierarchical Transformer for Exposure Correction
abstract
Photographs taken with less-than-ideal exposure settings often display poor visual quality. Since the correction procedures vary significantly, it is difficult for a single neural network to handle all exposure problems. Moreover, the inherent limitations of convolutions, hinder the models ability to restore faithful color or details on extremely over-/under-exposed regions. To overcome these limitations, we propose a Macro-Micro-Hierarchical transformer, which consists of a macro attention to capture long-range dependencies, a micro attention to extract local features, and a hierarchical structure for coarse-to-fine correction. In specific, the complementary macro-micro attention designs enhance locality while allowing global interactions. The hierarchical structure enables the network to correct exposure errors of different scales layer by layer. Furthermore, we propose a contrast constraint and couple it seamlessly in the loss function, where the corrected image is pulled towards the positive sample and pushed away from the dynamically generated negative samples. Thus the remaining color distortion and loss of detail can be removed. We also extend our method as an image enhancer for low-light face recognition and low-light semantic segmentation. Experiments demonstrate that our approach obtains more attractive results than state-of-the-art methods quantitatively and qualitatively.
Gehui Li, Jinyuan Liu 0001, Long Ma 0002, Zhiying Jiang, Xin Fan 0001, Risheng Liu
ACM Multimedia5
2023 Bilevel Generative Learning for Low-Light Vision
abstract
Recently, there has been a growing interest in constructing deep learning schemes for Low-Light Vision (LLV). Existing techniques primarily focus on designing task-specific and data-dependent vision models on the standard RGB domain, which inherently contain latent data associations. In this study, we propose a generic low-light vision solution by introducing a generative block to convert data from the RAW to the RGB domain. This novel approach connects diverse vision problems by explicitly depicting data generation, which is the first in the field. To precisely characterize the latent correspondence between the generative procedure and the vision task, we establish a bilevel model with the parameters of the generative block defined as the upper level and the parameters of the vision task defined as the lower level. We further develop two types of learning strategies targeting different goals, namely low cost and high accuracy, to acquire a new bilevel generative learning paradigm. The generative blocks embrace a strong generalization ability in other low-light vision tasks through the bilevel optimization on enhancement tasks. Extensive experimental evaluations on three representative low-light vision tasks, namely enhancement, detection, and segmentation, fully demonstrate the superiority of our proposed approach. The code will be available at https://github.com/Yingchi1998/BGL.
Yingchi Liu, Zhu Liu 0004, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
ACM Multimedia5
2023 PAIF: Perception-Aware Infrared-Visible Image Fusion for Attack-Tolerant Semantic Segmentation
abstract
Infrared and visible image fusion is a powerful technique that combines complementary information from different modalities for downstream semantic perception tasks. Existing learning-based methods show remarkable performance, but are suffering from the inherent vulnerability of adversarial attacks, causing a significant decrease in accuracy. In this work, a perception-aware fusion framework is proposed to promote segmentation robustness in adversarial scenes. We first conduct systematic analyses about the components of image fusion, investigating the correlation with segmentation robustness under adversarial perturbations. Based on these analyses, we propose a harmonized architecture search with a decomposition-based structure to balance standard accuracy and robustness. We also propose an adaptive learning strategy to improve the parameter robustness of image fusion, which can learn effective feature extraction under diverse adversarial perturbations. Thus, the goals of image fusion (i.e., extracting complementary features from source modalities and defending attack) can be realized from the perspectives of architectural and learning strategies. Extensive experimental results demonstrate that our scheme substantially enhances the robustness, with gains of 15.3% mIOU of segmentation in the adversarial scene, compared with advanced competitors. The source codes are available at https://github.com/LiuZhu-CV/PAIF.
Zhu Liu 0004, Jinyuan Liu 0001, Benzhuang Zhang, Long Ma 0002, Xin Fan 0001, Risheng Liu
ACM Multimedia5
2023 WaterFlow: Heuristic Normalizing Flow for Underwater Image Enhancement and Beyond
abstract
Underwater images suffer from light refraction and absorption, which impairs visibility and interferes the subsequent applications. Existing underwater image enhancement methods mainly focus on image quality improvement, ignoring the effect on practice. To balance the visual quality and application, we propose a heuristic normalizing flow for detection-driven underwater image enhancement, dubbed WaterFlow. Specifically, we first develop an invertible mapping to achieve the translation between the degraded image and its clear counterpart. Considering the differentiability and interpretability, we incorporate the heuristic prior into the data-driven mapping procedure, where the ambient light and medium transmission coefficient benefit credible generation. Furthermore, we introduce a detection perception module to transmit the implicit semantic guidance into the enhancement procedure, where the enhanced images hold more detection-favorable features and are able to promote the detection performance. Extensive experiments prove the superiority of our WaterFlow, against state-of-the-art methods quantitatively and qualitatively.
Zengxi Zhang, Zhiying Jiang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
ACM Multimedia4
2023 Exploring a Distillation with Embedded Prompts for Object Detection in Adverse Environments
Hao Fu 0004, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
PRCV (10)4
2023 Rethinking general underwater object detection: Datasets, challenges, and solutions
Chenping Fu, Risheng Liu, Xin Fan 0001, Puyang Chen, Hao Fu 0004, Wanqi Yuan, Ming Zhu 0001, Zhongxuan Luo
Neurocomputing3
2023 Learning With Nested Scene Modeling and Cooperative Architecture Search for Low-Light Vision
abstract
Images captured from low-light scenes often suffer from severe degradations, including low visibility, color casts, intensive noises, etc. These factors not only degrade image qualities, but also affect the performance of downstream Low-Light Vision (LLV) applications. A variety of deep networks have been proposed to enhance the visual quality of low-light images. However, they mostly rely on significant architecture engineering and often suffer from the high computational burden. More importantly, it still lacks an efficient paradigm to uniformly handle various tasks in the LLV scenarios. To partially address the above issues, we establish Retinex-inspired Unrolling with Architecture Search (RUAS), a general learning framework, that can address low-light enhancement task, and has the flexibility to handle other challenging downstream vision tasks. Specifically, we first establish a nested optimization formulation, together with an unrolling strategy, to explore underlying principles of a series of LLV tasks. Furthermore, we design a differentiable strategy to cooperatively search specific scene and task architectures for RUAS. Last but not least, we demonstrate how to apply RUAS for both low- and high-level LLV applications (e.g., enhancement, detection and segmentation). Extensive experiments verify the flexibility, effectiveness, and efficiency of RUAS.
Risheng Liu, Long Ma 0002, Tengyu Ma 0004, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Investigating intrinsic degradation factors by multi-branch aggregation for real-world underwater image enhancement
Xinwei Xue, Long Ma 0002, Qi Jia 0001, Risheng Liu, Xin Fan 0001
Pattern Recognit.6
2023 Breaking Free From Fusion Rule: A Fully Semantic-Driven Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion plays a vital role in the field of computer vision. Previous approaches make efforts to design various fusion rules in the loss functions. However, these experimental designed fusion rules make the methods more and more complex. Besides, most of them only focus on boosting the visual effects, thus showing unsatisfactory performance for the follow-up high-level vision tasks. To address these challenges, in this letter, we develop a semantic-level fusion network to sufficiently utilize the semantic guidance, emancipating the experimental designed fusion rules. In addition, to achieve a better semantic understanding of the feature fusion process, a fusion block based on the transformer is presented in a multi-scale manner. Moreover, we devise a regularization loss function, together with a training strategy, to fully use semantic guidance from the high-level vision tasks. Compared with state-of-the-art methods, our method does not depend on the hand-crafted fusion loss function. Still, it achieves superior performance on visual quality along with the follow-up high-level vision tasks.
Zhu Liu 0004, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
IEEE Signal Process. Lett.4
2023 Learning Heavily-Degraded Prior for Underwater Object Detection
abstract
Underwater object detection suffers from low detection performance because the distance and wavelength dependent imaging process yield evident image quality degradations such as haze-like effects, low visibility, and color distortions. Therefore, we commit to resolving the issue of underwater object detection with compounded environmental degradations. Typical approaches attempt to develop sophisticated deep architecture to generate high-quality images or features. However, these methods are only work for limited ranges because imaging factors are either unstable, too sensitive, or compounded. Unlike these approaches catering for high-quality images or features, this paper seeks transferable prior knowledge from detector-friendly images. The prior guides detectors removing degradations that interfere with detection. It is based on statistical observations that, the heavily degraded regions of detector-friendly (DFUI) and underwater images have evident feature distribution gaps while the lightly degraded regions of them overlap each other. Therefore, we propose a residual feature transference module (RFTM) to learn a mapping between deep representations of the heavily degraded patches of DFUI- and underwater-images, and make the mapping as a heavily degraded prior (HDP) for underwater detection. Since the statistical properties are independent to image content, HDP can be learned without the supervision of semantic labels and plugged into popular CNN-based feature extraction networks to improve their performance on underwater object detection. Without bells and whistles, evaluations on URPC2020 and UODD show that our methods outperform CNN-based detectors by a large margin. Our method with higher speeds and less parameters still performs better than transformer-based detectors. Our code and DFUI dataset can be found inhttps://github.com/xiaoDetection/Learning-Heavily-Degraed-Prior.
Chenping Fu, Xin Fan 0001, Jiewen Xiao, Wanqi Yuan, Risheng Liu, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.2
2023 Automated Learning for Deformable Medical Image Registration by Jointly Optimizing Network Architectures and Objective Functions
abstract
Deformable image registration plays a critical role in various tasks of medical image analysis. A successful registration algorithm, either derived from conventional energy optimization or deep networks, requires tremendous efforts from computer experts to well design registration energy or to carefully tune network architectures with respect to medical data available for a given registration task/scenario. This paper proposes an automated learning registration algorithm (AutoReg) that cooperatively optimizes both architectures and their corresponding training objectives, enabling non-computer experts to conveniently find off-the-shelf registration algorithms for various registration scenarios. Specifically, we establish a triple-level framework to embrace the searching for both network architectures and objectives with a cooperating optimization. Extensive experiments on multiple volumetric datasets and various registration scenarios demonstrate that AutoReg can automatically learn an optimal deep registration network for given volumes and achieve state-of-the-art performance. The automatically learned network also improves computational efficiency over the mainstream UNet architecture from 0.558 to 0.270 seconds for a volume pair on the same configuration.
Xin Fan 0001, Risheng Liu, Zhongxuan Luo, Hao Huang 0016
IEEE Trans. Image Process.1
2023 Characteristic Mapping for Ellipse Detection Acceleration
abstract
It is challenging to characterize the intrinsic geometry of high-degree algebraic curves with lower-degree algebraic curves. The reduction in the curve's degree implies lower computation costs, which is crucial for various practical computer vision systems. In this paper, we develop a characteristic mapping (CM) to recursively degenerate 3n points on a planar curve of n th order to 3(n-1) points on a curve of (n-1) th order. The proposed characteristic mapping enables curve grouping on a line, a curve of the lowest order, that preserves the intrinsic geometric properties of a higher-order curve (ellipse). We prove a necessary condition and derive an efficient arc grouping module that finds valid elliptical arc segments by determining whether the mapped three points are colinear, invoking minimal computation. We embed the module into two latest arc-based ellipse detection methods, which reduces their running time by 25% and 50% on average over five widely used data sets. This yields faster detection than the state-of-the-art algorithms while keeping their precision comparable or even higher. Two CM embedded methods also significantly surpass a deep learning method on all evaluation metrics.
Qi Jia 0001, Xin Fan 0001, Yang Yang 0120, Xuxu Liu, Zhongxuan Luo, Xinchen Zhou, Longin Jan Latecki
IEEE Trans. Image Process.2
2023 Optimization-Inspired Learning With Architecture Augmentations and Control Mechanisms for Low-Level Vision
abstract
In recent years, there has been a growing interest in combining learnable modules with numerical optimization to solve low-level vision tasks. However, most existing approaches focus on designing specialized schemes to generate image/feature propagation. There is a lack of unified consideration to construct propagative modules, provide theoretical analysis tools, and design effective learning mechanisms. To mitigate the above issues, this paper proposes a unified optimization-inspired learning framework to aggregate Generative, Discriminative, and Corrective (GDC for short) principles with strong generalization for diverse optimization models. Specifically, by introducing a general energy minimization model and formulating its descent direction from different viewpoints (i.e., in a generative manner, based on the discriminative metric and with optimality-based correction), we construct three propagative modules to effectively solve the optimization models with flexible combinations. We design two control mechanisms that provide the non-trivial theoretical guarantees for both fully- and partially-defined optimization formulations. Under the support of theoretical guarantees, we can introduce diverse architecture augmentation strategies such as normalization and search to ensure stable propagation with convergence and seamlessly integrate the suitable modules into the propagation respectively. Extensive experiments across varied low-level vision tasks validate the efficacy and adaptability of GDC.
Risheng Liu, Zhu Liu 0004, Pan Mu, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.4
2023 Low-Light Image Enhancement via Self-Reinforced Retinex Projection Model
abstract
Low-light image enhancement aims to improve the quality of images captured under low-lightening conditions, which is a fundamental problem in computer vision and multimedia areas. Although many efforts have been invested over the years, existing illumination-based models tend to generate unnatural-looking results (e.g., over-exposure). It is because that the widely-adopted illumination adjustment (e.g., Gamma Correction) breaks down the favorable smoothness property of the original illumination derived from the well-designed illumination estimation model. To settle this issue, a great-efficiency and high-quality Self-Reinforced Retinex Projection (SRRP) model is developed in this paper, which contains optimization modules of both illumination and reflectance layers. Specifically, we construct a new fidelity term with the self-reinforced function for the illumination optimization to eliminate the dependence of the illumination adjustment to obtain a desired illumination with the excellent smoothing property. By introducing a flexible feasible constraint, we obtain a reflectance optimization module with projection. Owing to its flexibility, we can extend our model to an enhanced version by integrating a data-driven denoising mechanism as the projection, which is able to effectively handle the generated noises/artifacts in the enhanced procedure. In the experimental part, on one side, we make ample comparative assessments on multiple benchmarks with considerable state-of-the-art methods. These evaluations fully verify the outstanding performance of our method, in terms of the qualitative and quantitative analyses and execution efficiency. On the other side, we also conduct extensive analytical experiments to indicate the effectiveness and advantages of our proposed model.
Long Ma 0002, Risheng Liu, Yiyang Wang 0001, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Multim.4
2023 Learning adaptive hyper-guidance via proxy-based bilevel optimization for image enhancement
Jiaxin Gao 0001, Xiaokun Liu, Risheng Liu, Xin Fan 0001
Vis. Comput.4
2023 A unified image fusion framework with flexible bilevel paradigm integration
Jinyuan Liu 0001, Zhiying Jiang, Guanyao Wu, Risheng Liu, Xin Fan 0001
Vis. Comput.5
2023 SSoB: searching a scene-oriented architecture for underwater object detection
Wanqi Yuan, Chenping Fu, Risheng Liu, Xin Fan 0001
Vis. Comput.4
2022 Segment, Magnify and Reiterate: Detecting Camouflaged Objects the Hard Way
abstract
It is challenging to accurately detect camouflaged objects from their highly similar surroundings. Existing methods mainly leverage a single-stage detection fashion, while neglecting small objects with low-resolution fine edges requires more operations than the larger ones. To tackle camouflaged object detection (COD), we are inspired by humans attention coupled with the coarse-to-fine detection strategy, and thereby propose an iterative refinement framework, coined SegMaR, which integrates Segment, Magnify and Reiterate in a multi-stage detection fashion. Specifically, we design a new discriminative mask which makes the model attend on the fixation and edge regions. In addition, we leverage an attention-based sampler to magnify the object region progressively with no need of enlarging the image size. Extensive experiments show our SegMaR achieves remarkable and consistent improvements over other state-of-the-art methods. Especially, we surpass two competitive methods 7.4% and 20.0% respectively in average over standard evaluation metrics on small camouflaged objects. Additional studies provide more promising insights into Seg-MaR, including its effectiveness on the discriminative mask and its generalization to other network architectures. Code is available at https://github.com/dlut-dimt/SegMaR.
Qi Jia 0001, Shuilian Yao, Yu Liu 0012, Xin Fan 0001, Risheng Liu, Zhongxuan Luo
CVPR4
2022 Toward Fast, Flexible, and Robust Low-Light Image Enhancement
abstract
Existing low-light image enhancement techniques are mostly not only difficult to deal with both visual quality and computational efficiency but also commonly invalid in unknown complex scenarios. In this paper, we develop a new Self-Calibrated Illumination (SCI) learning framework for fast, flexible, and robust brightening images in real-world low-light scenarios. To be specific, we establish a cascaded illumination learning process with weight sharing to handle this task. Considering the computational burden of the cascaded pattern, we construct the self-calibrated module which realizes the convergence between results of each stage, producing the gains that only use the single basic block for inference (yet has not been exploited in previous works), which drastically diminishes computation cost. We then define the unsupervised training loss to elevate the model capability that can adapt general scenes. Further, we make comprehensive explorations to excavate SCI's inherent properties (lacking in existing works) including operation-insensitive adaptability (acquiring stable performance under the settings of different simple operations) and model-irrelevant generality (can be applied to illumination-based existing works to improve performance). Finally, plenty of experiments and ablation studies fully indicate our superiority in both quality and efficiency. Applications on low-light face detection and nighttime semantic segmentation fully reveal the latent practical values for SCI. The source code is available at https://github.com/vis-opt-group/SCI.
Long Ma 0002, Tengyu Ma 0004, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
CVPR4
2022 Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object Detection
abstract
This study addresses the issue of fusing infrared and visible images that appear differently for object detection. Aiming at generating an image of high visual quality, previous approaches discover commons underlying the two modalities and fuse upon the common space either by iterative optimization or deep networks. These approaches neglect that modality differences implying the complementary information are extremely important for both fusion and subsequent detection task. This paper proposes a bilevel optimization formulation for the joint problem of fusion and detection, and then unrolls to a target-aware Dual Adversarial Learning (TarDAL) network for fusion and a commonly used detection network. The fusion network with one generator and dual discriminators seeks commons while learning from differences, which preserves structural information of targets from the infrared and textural details from the visible. Furthermore, we build a synchronized imaging system with calibrated infrared and optical sensors, and collect currently the most comprehensive benchmark covering a wide range of scenarios. Extensive experiments on several public datasets and our benchmark demonstrate that our method outputs not only visually appealing fusion but also higher detection mAP than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/dlut-dimt/TarDAL.
Jinyuan Liu 0001, Xin Fan 0001, Zhanbo Huang, Guanyao Wu, Risheng Liu, Zhongxuan Luo
CVPR2
2022 ReCoNet: Recurrent Correction Network for Fast and Efficient Multi-modality Image Fusion
Zhanbo Huang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu, Zhongxuan Luo
ECCV (18)3
2022 Learning to Fuse Heterogeneous Features for Low-Light Image Enhancement
abstract
To see clearly in low-light scenarios, a series of learning-based techniques have been developed to improve visual quality. However, due to the absence of semantic-level features, the existing methods are perhaps less effective on semantic-oriented visual analysis tasks (e.g., saliency detection). To break down the limitation, we propose a new classification-driven enhancement method with heterogeneous feature fusion. Specifically, we construct a new low-light image enhancement network by integrating features acquired from the pre-trained classification network. Then, to better exploit the semantic-level information, we establish a Heterogeneous Feature Fusion (HF2) operation with channel-and-spatial attention to strength the effects of cross-domain features. HF2 acts on not only the fusion between classification and encoded features but also the fusion between encoded and decoded features. Extensive experiments are conducted to indicate our superiority against other state-of-the-art methods. The application on saliency detection further reveals our effectiveness in settling the semantic-oriented visual tasks.
Zhenyu Tang 0004, Long Ma 0002, Xiaoke Shang, Xin Fan 0001
ICASSP4
2022 Semantic-aware Texture-Structure Feature Collaboration for Underwater Image Enhancement
abstract
Underwater image enhancement has become an attractive topic as a significant technology in marine engi-neering and aquatic robotics. However, the limited number of datasets and imperfect hand-crafted ground truth weaken its robustness to unseen scenarios, and hamper the application to high-level vision tasks. To address the above limitations, we develop an efficient and compact enhancement network in collaboration with a high-level semantic-aware pretrained model, aiming to exploit its hierarchical feature representation as an auxiliary for the low-level underwater image enhance-ment. Specifically, we tend to characterize the shallow layer features as textures while the deep layer features as structures in the semantic-aware model, and propose a multi-path Contextual Feature Refinement Module (CFRM) to refine features in multiple scales and model the correlation between different features. In addition, a feature dominative network is devised to perform channel-wise modulation on the aggregated texture and structure features for the adaptation to different feature patterns of the enhancement network. Extensive experiments on benchmarks demonstrate that the proposed algorithm achieves more appealing results and outperforms state-of-the-art meth-ods by large margins. We also apply the proposed algorithm to the underwater salient object detection task to reveal the favorable semantic-aware ability for high-level vision tasks.
Di Wang 0018, Long Ma 0002, Risheng Liu, Xin Fan 0001
ICRA4
2022 Unsupervised Misaligned Infrared and Visible Image Fusion via Cross-Modality Image Generation and Registration
abstract
Recent learning-based image fusion methods have marked numerous progress in pre-registered multi-modality data, but suffered serious ghosts dealing with misaligned multi-modality data, due to the spatial deformation and the difficulty narrowing cross-modality discrepancy. To overcome the obstacles, in this paper, we present a robust cross-modality generation-registration paradigm for unsupervised misaligned infrared and visible image fusion (IVIF). Specifically, we propose a Cross-modality Perceptual Style Transfer Network (CPSTN) to generate a pseudo infrared image taking a visible image as input. Benefiting from the favorable geometry preservation ability of the CPSTN, the generated pseudo infrared image embraces a sharp structure, which is more conducive to transforming cross-modality image alignment into mono-modality registration coupled with the structure-sensitive of the infrared image. In this case, we introduce a Multi-level Refinement Registration Network (MRRN) to predict the displacement vector field between distorted and pseudo infrared images and reconstruct registered infrared image under the mono-modality setting. Moreover, to better fuse the registered infrared images and visible images, we present a feature Interaction Fusion Module (IFM) to adaptively select more meaningful features for fusion in the Dual-path Interaction Fusion Network (DIFN). Extensive experimental results suggest that the proposed method performs superior capability on misaligned cross-modality image fusion.
Di Wang 0018, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
IJCAI3
2022 Hierarchical Bilevel Learning with Architecture and Loss Search for Hadamard-based Image Restoration
abstract
In the past few decades, Hadamard-based image restoration problems (e.g., low-light image enhancement) attract wide concerns in multiple areas related to artificial intelligence. However, existing works mostly focus on heuristically defining architecture and loss by the engineering experiences that came from extensive practices. This way brings about expensive verification costs for seeking out the optimal solution. To this end, we develop a novel hierarchical bilevel learning scheme to discover the architecture and loss simultaneously for different Hadamard-based image restoration tasks. More concretely, we first establish a new Hadamard-inspired neural unit to aggregate domain knowledge into the network design. Then we model a triple-level optimization that consists of the architecture, loss and parameters optimizations to deliver a macro perspective for network learning. Then we introduce a new hierarchical bilevel learning scheme for solving the built triple-level model to progressively generate the desired architecture and loss. We also define an architecture search space consisting of a series of simple operations and an image quality-oriented loss search space. Extensive experiments on three Hadamard-based image restoration tasks (including low-light image enhancement, single image haze removal and underwater image enhancement) fully verify our superiority against state-of-the-art methods.
Guijing Zhu, Long Ma 0002, Xin Fan 0001, Risheng Liu
IJCAI3
2022 PIA: Parallel Architecture with Illumination Allocator for Joint Enhancement and Detection in Low-Light
abstract
Visual perception in low-light conditions (e.g., nighttime) plays an important role in various multimedia-related applications (e.g., autonomous driving). The enhancement (provides a visual-friendly appearance) and detection (detects the instances of objects) in low-light are two fundamental and crucial visual perception tasks. In this paper, we make efforts on how to simultaneously realize low-light enhancement and detection from two aspects. First, we define a parallel architecture to satisfy the task demand for both two tasks. In which, a decomposition-type warm-start acting on the entrance of parallel architecture is developed to narrow down the adverse effects brought by low-light scenes to some extent. Second, a novel illumination allocator is designed by encoding the key illumination component (the inherent difference between normal-light and low-light) to extract hierarchical features for assisting in enhancement and detection. Further, we make a substantive discussion for our proposed method. That is, we solve enhancement in a coarse-to-fine manner and handle detection in a decomposed-to-integrated fashion. Finally, multidimensional analytical and evaluated experiments are performed to indicate our effectiveness and superiority. The code is available at \urlhttps://github.com/tengyu1998/PIA
Tengyu Ma 0004, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
ACM Multimedia3
2022 Towards All Weather and Unobstructed Multi-Spectral Image Stitching: Algorithm and Benchmark
abstract
Image stitching is a fundamental task that requires multiple images from different viewpoints to generate a wide field-of-viewing~(FOV) scene. Previous methods are developed on RGB images. However, the severe weather and harsh conditions, such as rain, fog, low light, strong light, etc., on visible images may introduce evident interference, leading to the distortion and misalignment of the stitched results. To remedy the deficient imaging of optical sensors, we investigate the complementarity across infrared and visible images to improve the perception of scenes in terms of visual information and viewing ranges. Instead of the cascaded fusion-stitching process, where the inaccuracy accumulation caused by image fusion hinders the stitch performance, especially content loss and ghosting effect, we develop a learnable feature adaptive network to investigate a stitch-oriented feature representation and perform the information complementary at the feature-level. By introducing a pyramidal structure along with the global fast correlation regression, the quadrature attention based correspondence is more responsible for feature alignment, and the estimation of sparse offsets can be realized in a coarse-to-fine manner. Furthermore, we propose the first infrared and visible image based multi-spectral image stitching dataset, covering a more comprehensive range of scenarios and diverse viewing baselines. Extensive experiments on real-world data demonstrate that our method reconstructs the wide FOV images with more credible structure and complementary information against state-of-the-arts.
Zhiying Jiang, Zengxi Zhang, Xin Fan 0001, Risheng Liu
ACM Multimedia3
2022 Best of Both Worlds: See and Understand Clearly in the Dark
abstract
Recently, with the development of intelligent technology, the perception of low-light scenes has been gaining widespread attention. However, existing techniques usually focus on only one task (e.g., enhancement) and lose sight of the others (e.g., detection), making it difficult to perform all of them well at the same time. To overcome this limitation, we propose a new method that can handle visual quality enhancement and semantic-related tasks (e.g., detection, segmentation) simultaneously in a unified framework. Specifically, we build a cascaded architecture to meet the task requirements. To better enhance the entanglement in both tasks and achieve mutual guidance, we develop a new contrastive-alternative learning strategy for learning the model parameters, to largely improve the representational capacity of the cascaded architecture. Notably, the contrastive learning mechanism establishes the communication between two objective tasks in essence, which actually extends the capability of contrastive learning to some extent. Finally, extensive experiments are performed to fully validate the advantages of our method over other state-of-the-art works in enhancement, detection, and segmentation. A series of analytical evaluations are also conducted to reveal our effectiveness. The code is available at https://github.com/k914/contrastive-alternative-learning.
Xinwei Xue, Long Ma 0002, Yi Wang 0037, Xin Fan 0001, Risheng Liu
ACM Multimedia5
2022 DeepEnReg: Joint Enhancement and Affine Registration for Low-contrast Medical Images
Yun Peng 0005, Huanyu Luo, Huanjie Li, Zhongxuan Luo, Xin Fan 0001
PRCV (2)9
2022 Efficient Representation and Optimization of TPMS-Based Porous Structures for 3D Heat Dissipation
Shengfa Wang, Yu Jiang 0019, Jiangbei Hu, Xin Fan 0001, Zhongxuan Luo, Ligang Liu 0001
Comput. Aided Des.4
2022 Learning Deformable Image Registration From Optimization: Perspective, Modules, Bilevel Training and Beyond
abstract
Conventional deformable registration methods aim at solving an optimization model carefully designed on image pairs and their computational costs are exceptionally high. In contrast, recent deep learning-based approaches can provide fast deformation estimation. These heuristic network architectures are fully data-driven and thus lack explicit geometric constraints which are indispensable to generate plausible deformations, e.g., topology-preserving. Moreover, these learning-based approaches typically pose hyper-parameter learning as a black-box problem and require considerable computational and human effort to perform many training runs. To tackle the aforementioned problems, we propose a new learning-based framework to optimize a diffeomorphic model via multi-scale propagation. Specifically, we introduce a generic optimization model to formulate diffeomorphic registration and develop a series of learnable architectures to obtain propagative updating in the coarse-to-fine feature space. Further, we propose a new bilevel self-tuned training strategy, allowing efficient search of task-specific hyper-parameters. This training strategy increases the flexibility to various types of data while reduces computational and human burdens. We conduct two groups of image registration experiments on 3D volume datasets including image-to-atlas registration on brain MRI data and image-to-image registration on liver CT data. Extensive results demonstrate the state-of-the-art performance of the proposed method with diffeomorphic guarantee and extreme efficiency. We also apply our framework to challenging multi-modal image registration, and investigate how our registration to support the down-streaming tasks for medical image analysis including multi-modal fusion and image segmentation.
Risheng Liu, Xin Fan 0001, Chenying Zhao, Hao Huang 0016, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Learn to Search a Lightweight Architecture for Target-Aware Infrared and Visible Image Fusion
abstract
Deep learning technology has recently achieved remarkable progress in infrared and visible image fusion. Nevertheless, existing methods have encountered blurred targets or unfaithful textural details on their fused results. They suffer from high computational expenses, thus unable to directly serve the subsequent high-level vision tasks. In this letter, to alleviate this issue, we proposed leveraging a lightweight architecture based on Neural Architecture Search (NAS) to realize the infrared and visible image fusion in an end-to-end manner, significantly reducing the computational expenses and runtime. Concretely, we construct a search-based architecture to explore the feature representation across different modalities automatically. Then a saliency-based loss function is designed to retain both the distinct target and texture details. Motivated by the cooperative principle, we also formulate a flexible hardware-sensitive regularization constraint in our loss function for discovering efficient operations. As a result, we can generate a target-distinct fused result with high efficiency. Extensive qualitative and quantitative experiments reveal that our method has superior performance against the state-of-the-art methods, especially highlighting the target, retaining realistic details, and achieving fast running speed. Specifically, our method increases by 150% in time, reduces the FLOPS by 21.3% and reduces the model parameters by 25%.
Jinyuan Liu 0001, Guanyao Wu, Risheng Liu, Xin Fan 0001
IEEE Signal Process. Lett.5
2022 Target Oriented Perceptual Adversarial Fusion Network for Underwater Image Enhancement
abstract
Due to the refraction and absorption of light by water, underwater images usually suffer from severe degradation, such as color cast, hazy blur, and low visibility, which would degrade the effectiveness of marine applications equipped on autonomous underwater vehicles. To eliminate the degradation of underwater images, we propose a target oriented perceptual adversarial fusion network, dubbed TOPAL. Concretely, we consider the degradation factors of underwater images in terms of turbidity and chromatism. And according to the degradation issues, we first develop a multi-scale dense boosted module to strengthen the visual contrast and a deep aesthetic render module to perform the color correction, respectively. After that, we employ the dual channel-wise attention module and guide the adaptive fusion of latent features, in which both diverse details and credible appearance are integrated. To bridge the gap between synthetic and real-world images, a global-local adversarial mechanism is introduced in the reconstruction. Besides, perceptual information is also embedded into the process to assist the understanding of scenery content. To evaluate the performance of TOPAL, we conduct extensive experiments on several benchmarks and make comparisons among state-of-the-art methods. Quantitative and qualitative results demonstrate that our TOPAL improves the quality of underwater images greatly and achieves superior performance than others.
Zhiying Jiang, Zhuoxiao Li, Shuzhou Yang, Xin Fan 0001, Risheng Liu
IEEE Trans. Circuits Syst. Video Technol.4
2022 Learning a Deep Multi-Scale Feature Ensemble and an Edge-Attention Guidance for Image Fusion
abstract
Image fusion integrates a series of images acquired from different sensors,e.g., infrared and visible, outputting an image with richer information than either one. Traditional and recent deep-based methods have difficulties in preserving prominent structures and recovering vital textural details for practical applications. In this article, we propose a deep network for infrared and visible image fusion cascading a feature learning module with a fusion learning mechanism. Firstly, we apply a coarse-to-fine deep architecture to learn multi-scale features for multi-modal images, which enables discovering prominent common structures for later fusion operations. The proposed feature learning module requires no well-aligned image pairs for training. Compared with the existing learning-based methods, the proposed feature learning module can ensemble numerous examples from respective modals for training, increasing the ability of feature representation. Secondly, we design an edge-guided attention mechanism upon the multi-scale features to guide the fusion focusing on common structures, thus recovering details while attenuating noise. Moreover, we provide a new aligned infrared and visible image fusion dataset, RealStreet, collected in various practical scenarios for comprehensive evaluation. Extensive experiments on two benchmarks, TNO and RealStreet, demonstrate the superiority of the proposed method over the state-of-the-art in terms of both visual inspection and objective analysis on six evaluation metrics. We also conduct the experiments on the FLIR and NIR datasets, containing foggy weather and poor light conditions, to verify the generalization and robustness of the proposed method.
Jinyuan Liu 0001, Xin Fan 0001, Ji Jiang, Risheng Liu, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.2
2022 Cross-SRN: Structure-Preserving Super-Resolution Network With Cross Convolution
abstract
It is challenging to restore low-resolution (LR) images to super-resolution (SR) images with correct and clear details. Existing deep learning works almost neglect the inherent structural information of images, which acts as an important role for visual perception of SR results. In this paper, we design a hierarchical feature exploitation network to probe and preserve structural information in a multi-scale feature fusion manner. First, we propose a cross convolution upon traditional edge detectors to localize and represent edge features. Then, cross convolution blocks (CCBs) are designed with feature normalization and channel attention to consider the inherent correlations of features. Finally, we leverage multi-scale feature fusion group (MFFG) to embed the cross convolution blocks and develop the relations of structural features in different scales hierarchically, invoking a lightweight structure-preserving network named as Cross-SRN. Experimental results demonstrate the Cross-SRN achieves competitive or superior restoration performances against the state-of-the-art methods with accurate and clear structural details. Moreover, we set a criterion to select images with rich structural textures. The proposed Cross-SRN outperforms the state-of-the-art methods on the selected benchmark, which demonstrates that our network has a significant advantage in preserving edges.
Yuqing Liu 0001, Qi Jia 0001, Xin Fan 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Attention-Guided Global-Local Adversarial Learning for Detail-Preserving Multi-Exposure Image Fusion
abstract
Deep learning networks have recently demonstrated yielded impressive progress for multi-exposure image fusion. However, how to restore realistic texture details while correcting color distortion is still a challenging problem to be solved. To alleviate the aforementioned issues, in this paper, we propose an attention-guided global-local adversarial learning network for fusing extreme exposure images in a coarse-to-fine manner. Firstly, the coarse fusion result is generated under the guidance of attention weight maps, which acquires the essential region of interest from both sides. Secondly, we formulate an edge loss function, along with a spatial feature transform layer, for refining the fusion process. So that it can take full use of the edge information to deal with blurry edges. Moreover, by incorporating global-local learning, our method can balance pixel intensity distribution and correct the color distortion on spatially varying source images from both image/patch perspectives. Such a global-local discriminator ensures all the local patches of the fused images align with realistic normal-exposure ones. Extensive experimental results on two publicly available datasets show that our method drastically outperforms state-of-the-art methods in visual inspection and objective analysis. Furthermore, sufficient ablation experiments prove that our method has significant advantages in generating high-quality fused results with appealing details, clear targets, and faithful color. Source code will be available athttps://github.com/JinyuanLiu-CV/AGAL.
Jinyuan Liu 0001, Jingjie Shang, Risheng Liu, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 A New Dataset, Poisson GAN and AquaNet for Underwater Object Grabbing
abstract
To boost the object grabbing capability of underwater robots for open-sea farming, we propose a new dataset (UDD) consisting of three categories (seacucumber, seaurchin, and scallop) with 2,227 images. To the best of our knowledge, it is the first 4K HD dataset collected in a real open-sea farm. We also propose a novel Poisson-blending Generative Adversarial Network (Poisson GAN) and an efficient object detection network (AquaNet) to address two common issues within related datasets: the class-imbalance problem and the problem of mass small object, respectively. Specifically, Poisson GAN combines Poisson blending into its generator and employs a new loss called Dual Restriction loss (DR loss), which supervises both implicit space features and image-level features during training to generate more realistic images. By utilizing Poisson GAN, objects of minority class like seacucumber or scallop could be added into an image naturally and annotated automatically, which could increase the loss of minority classes during training detectors to eliminate the class-imbalance problem; AquaNet is a high-efficiency detector to address the problem of detecting mass small objects from cloudy underwater pictures. Within it, we design two efficient components: a depth-wise-convolution-based Multi-scale Contextual Features Fusion (MFF) block and a Multi-scale Blursampling (MBP) module to reduce the parameters of the network to 1.3 million. Both two components could provide multi-scale features of small objects under a short backbone configuration without any loss of accuracy. In addition, we construct a large-scale augmented dataset (AUDD) and a pre-training dataset via Poisson GAN from UDD. Extensive experiments show the effectiveness of the proposed Poisson GAN, AquaNet, UDD, AUDD, and pre-training dataset.
Chongwei Liu, Zhihui Wang 0001, Shijie Wang 0003, Yulong Tao, Caifei Yang, Xing Liu 0005, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.9
2022 Twin Adversarial Contrastive Learning for Underwater Image Enhancement and Beyond
abstract
Underwater images suffer from severe distortion, which degrades the accuracy of object detection performed in an underwater environment. Existing underwater image enhancement algorithms focus on the restoration of contrast and scene reflection. In practice, the enhanced images may not benefit the effectiveness of detection and even lead to a severe performance drop. In this paper, we propose an object-guided twin adversarial contrastive learning based underwater enhancement method to achieve both visual-friendly and task-orientated enhancement. Concretely, we first develop a bilateral constrained closed-loop adversarial enhancement module, which eases the requirement of paired data with the unsupervised manner and preserves more informative features by coupling with the twin inverse mapping. In addition, to confer the restored images with a more realistic appearance, we also adopt the contrastive cues in the training phase. To narrow the gap between visually-oriented and detection-favorable target images, a task-aware feedback module is embedded in the enhancement process, where the coherent gradient information of the detector is incorporated to guide the enhancement towards the detection-pleasing direction. To validate the performance, we allocate a series of prolific detectors into our framework. Extensive experiments demonstrate that the enhanced results of our method show remarkable amelioration in visual quality, the accuracy of different detectors conducted on our enhanced images has been promoted notably. Moreover, we also conduct a study on semantic segmentation to illustrate how object guidance improves high-level tasks. Code and models are available at https://github.com/Jzy2017/TACL.
Risheng Liu, Zhiying Jiang, Shuzhou Yang, Xin Fan 0001
IEEE Trans. Image Process.4
2022 Triple-Level Model Inferred Collaborative Network Architecture for Video Deraining
abstract
Video deraining is an important issue for outdoor vision systems and has been investigated extensively. However, designing optimal architectures by the aggregating model formation and data distribution is a challenging task for video deraining. In this paper, we develop a model-guided triple-level optimization framework to deduce network architecture with cooperating optimization and auto-searching mechanism, named Triple-level Model Inferred Cooperating Searching (TMICS), for dealing with various video rain circumstances. In particular, to mitigate the problem that existing methods cannot cover various rain streaks distribution, we first design a hyper-parameter optimization model about task variable and hyper-parameter. Based on the proposed optimization model, we design a collaborative structure for video deraining. This structure includes Dominant Network Architecture (DNA) and Companionate Network Architecture (CNA) that is cooperated by introducing an Attention-based Averaging Scheme (AAS). To better explore inter-frame information from videos, we introduce a macroscopic structure searching scheme that searches from Optical Flow Module (OFM) and Temporal Grouping Module (TGM) to help restore latent frame. In addition, we apply the differentiable neural architecture searching from a compact candidate set of task-specific operations to discover desirable rain streaks removal architectures automatically. Extensive experiments on various datasets demonstrate that our model shows significant improvements in fidelity and temporal consistency over the state-of-the-art works. Source code is available at https://github.com/vis-opt-group/TMICS.
Pan Mu, Zhu Liu 0004, Risheng Liu, Xin Fan 0001
IEEE Trans. Image Process.5
2022 Underexposed Image Correction via Hybrid Priors Navigated Deep Propagation
abstract
Enhancing visual quality for underexposed images is an extensively concerning task that plays an important role in various areas of multimedia and computer vision. Most existing methods often fail to generate high-quality results with appropriate luminance and abundant details. To address these issues, we develop a novel framework, integrating both knowledge from physical principles and implicit distributions from data to address underexposed image correction. More concretely, we propose a new perspective to formulate this task as an energy-inspired model with advanced hybrid priors. A propagation procedure navigated by the hybrid priors is well designed for simultaneously propagating the reflectance and illumination toward desired results. We conduct extensive experiments to verify the necessity of integrating both underlying principles (i.e., with knowledge) and distributions (i.e., from data) as navigated deep propagation. Plenty of experimental results of underexposed image correction demonstrate that our proposed method performs favorably against the state-of-the-art methods on both subjective and objective assessments. In addition, we execute the task of face detection to further verify the naturalness and practical value of underexposed image correction. What is more, we apply our method to solve single-image haze removal whose experimental results further demonstrate our superiorities.
Risheng Liu, Long Ma 0002, Yuxi Zhang 0001, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.4
2022 Learning Deep Context-Sensitive Decomposition for Low-Light Image Enhancement
abstract
Enhancing the quality of low-light (LOL) images plays a very important role in many image processing and multimedia applications. In recent years, a variety of deep learning techniques have been developed to address this challenging task. A typical framework is to simultaneously estimate the illumination and reflectance, but they disregard the scene-level contextual information encapsulated in feature spaces, causing many unfavorable outcomes, e.g., details loss, color unsaturation, and artifacts. To address these issues, we develop a new context-sensitive decomposition network (CSDNet) architecture to exploit the scene-level contextual dependencies on spatial scales. More concretely, we build a two-stream estimation mechanism including reflectance and illumination estimation network. We design a novel context-sensitive decomposition connection to bridge the two-stream mechanism by incorporating the physical principle. The spatially varying illumination guidance is further constructed for achieving the edge-aware smoothness property of the illumination component. According to different training patterns, we construct CSDNet (paired supervision) and context-sensitive decomposition generative adversarial network (CSDGAN) (unpaired supervision) to fully evaluate our designed architecture. We test our method on seven testing benchmarks [including massachusetts institute of technology (MIT)-Adobe FiveK, LOL, ExDark, and naturalness preserved enhancement (NPE)] to conduct plenty of analytical and evaluated experiments. Thanks to our designed context-sensitive decomposition connection, we successfully realized excellent enhanced results (with sufficient details, vivid colors, and few noises), which fully indicates our superiority against existing state-of-the-art approaches. Finally, considering the practical needs for high efficiency, we develop a lightweight CSDNet (named LiteCSDNet) by reducing the number of channels. Furthermore, by sharing an encoder for these two components, we obtain a more lightweight version (SLiteCSDNet for short). SLiteCSDNet just contains 0.0301M parameters but achieves the almost same performance as CSDNet. Code is available at https://github.com/KarelZhang/CSDNet-CSDGAN.
Long Ma 0002, Risheng Liu, Jiaao Zhang, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.4
2022 Global structure-guided learning framework for underwater image enhancement
Runjia Lin, Jinyuan Liu 0001, Risheng Liu, Xin Fan 0001
Vis. Comput.4
2021 Physics-inspired Learning for Structure-Aware Texture-Sensitive Underwater Image Enhancement
abstract
Recently, improving the visual quality of underwater images using deep learning-based methods has drawn considerable attention. Unfortunately, diverse environmental factors (e.g., blue/green color distortion) severely limit their performance in real-world environments. Therefore, strengthening the superiority of the underwater image enhancement method is critical. In this paper, we devote ourselves to develop a new architecture with strong superiority and adaptability. Inspired by the underwater imaging principle, we establish a novel physics-inspired learning model that is easy to realize. A Structure-Aware Texture-Sensitive Network (SATS-Net) is further developed to portray the model. The structure-aware module is responsible for structural information, and the texture-sensitive module is responsible for textural information. Thus, SATS-Net successfully incorporates robust characterization absorbed from the physical principle to achieve strong robustness and adaptability. We conduct extensive experiments to demonstrate that SATS-Net outperforms existing advanced techniques in various real-world underwater environments.
Xinwei Xue, Long Ma 0002, Risheng Liu, Xin Fan 0001
ACML5
2021 Leveraging Line-Point Consistence To Preserve Structures for Wide Parallax Image Stitching
abstract
Generating high-quality stitched images with natural structures is a challenging task in computer vision. In this paper, we succeed in preserving both local and global geometric structures for wide parallax images, while reducing artifacts and distortions. A projective invariant, Characteristic Number, is used to match co-planar local sub-regions for input images. The homography between these well-matched sub-regions produces consistent line and point pairs, suppressing artifacts in overlapping areas. We explore and introduce global collinear structures into an objective function to specify and balance the desired characters for image warping, which can preserve both local and global structures while alleviating distortions. We also develop comprehensive measures for stitching quality to quantify the collinearity of points and the discrepancy of matched line pairs by considering the sensitivity to linear structures for human vision. Extensive experiments demonstrate the superior performance of the proposed method over the state-of-the-art by presenting sharp textures and preserving prominent natural structures in stitched images. Especially, our method not only exhibits lower errors but also the least divergence across all test images. Code is available at https://github.com/dut-media-lab/Image-Stitching.
Qi Jia 0001, Zhengjun Li, Xin Fan 0001, Shiyu Teng, Xinchen Ye, Longin Jan Latecki
CVPR3
2021 Retinex-Inspired Unrolling With Cooperative Prior Architecture Search for Low-Light Image Enhancement
abstract
Low-light image enhancement plays very important roles in low-level vision areas. Recent works have built a great deal of deep learning models to address this task. However, these approaches mostly rely on significant architecture engineering and suffer from high computational burden. In this paper, we propose a new method, named Retinex-inspired Unrolling with Architecture Search (RUAS), to construct lightweight yet effective enhancement network for low-light images in real-world scenario. Specifically, building upon Retinex rule, RUAS first establishes models to characterize the intrinsic underexposed structure of low-light images and unroll their optimization processes to construct our holistic propagation structure. Then by designing a cooperative reference-free learning strategy to discover low-light prior architectures from a compact search space, RUAS is able to obtain a top-performing image enhancement network, which is with fast speed and requires few computational resources. Extensive experiments verify the superiority of our RUAS framework against recently proposed state-of-the-art methods. The project page is available at http://dutmedia.org/RUAS/.
Risheng Liu, Long Ma 0002, Jiaao Zhang, Xin Fan 0001, Zhongxuan Luo
CVPR4
2021 NASA: A Noise-Adaptive and Structure-Aware Learning Framework for Image Deblurring
abstract
Image deblurring is a classical low-level visual processing task, which aims to recover a potentially noise-free sharp image from the blurred image. Existing prior-based and learning-based methods usually need to manually set some vital auxiliary components (e.g., noise level). It brings about extremely weak adaptability and flexibility. To settle this issue, we develop a Noise-Adaptive Structure-Aware learning framework (NASA) to achieve fully intelligent manufacturing. Concretely, by introducing a new task-assisted module, we define a novel robust image deblurring model derived from a MAP-based energy function. Consequently, we establish the NASA which consists of three basic modules including the task-assisted, fidelity-term, and regularization-term modules, to solve our designed model. The task-assisted module generates the noise-adaptive and structure-aware maps, which are fed to the other two modules. By end-to-end training our NASA, we successfully avoid the cumbersome manually parameters-adjustment process. Quantitative and qualitative experiments demonstrate our superiority compared to the state-of-the-art methods, both in visual effect and numerical scores. A series of ablation study also verify the effectiveness and necessity of our designed mechanism.
Xiaokun Liu, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICASSP5
2021 Temporal Rain Decomposition with Spatial Structure Guidance for Video Deraining
abstract
Recently, removing rain streaks from videos has drawn wide concerns in vision and multimedia communities. But existing works ignore the depicts of image inherent structure and rain location to cause details loss, and their adopted manners of exploiting temporal information are still insufficient. In this work, we propose a multi-frame deraining network with temporal rain decomposition and spatial structure guidance to more effectively accomplish video deraining. A learnable decomposition method is defined to learn the distribution of rain, where the location map acts on a single-frame deraining block. We construct a multi-frame fusion module with a detailed guidance map to integrate temporal and spatial information. Many evaluated experiments demonstrate that our algorithm performs favorably on video deraining tasks compared with other methods. The elaborate ablation study in terms of network architecture fully indicates the effectiveness of our network.
Xinwei Xue, Ying Ding 0006, Long Ma 0002, Yi Wang 0037, Risheng Liu, Xin Fan 0001
ICASSP6
2021 GTA-Net: Gradual Temporal Aggregation Network for Fast Video Deraining
abstract
Recently, the development of intelligent technology arouses the requirements of high-quality videos. Rain streak is a frequent and inevitable factor to degrade the video. Many researchers have put their energies into eliminating the adverse effects of rainy video. Unfortunately, how to fully utilize the temporal information from rainy video is still in suspense. In this work, to effectively exploit temporal information, we develop a simple but effective network, Gradual Temporal Aggregation Network (GTA-Net for short). To be specific, according to the temporal distance between rainy frames and the reference frame, we divide the rainy frames into different groups. A multi-stream coarse temporal aggregation module is first performed to aggregate different temporal information with equal status and importance. Then we design a single-stream fine temporal aggregation module to further fuse the integrated frames that maintain the different distances with the target frame. In this way of coarse-to-fine, we not only achieve superior performance, but also gain the surprising execution speed owing to abandon the time-consuming alignment operation. Plenty of experimental results demonstrate that our GTA-Net performs favorably compared to other state-of-the-art approaches. The meticulous ablation study further indicates the effectiveness of our designed GTA-Net.
Xinwei Xue, Long Ma 0002, Risheng Liu, Xin Fan 0001
ICASSP5
2021 Video Deraining Via Temporal Aggregation-and-Guidance
abstract
Learning-based video deraining methods generally integrate temporal correlation within the network. But their non-transparency (i.e., difficult to comprehend how to exploit temporal correlation) seriously limits the development for video deraining. To conquer it, this paper proposes a novel Temporal Aggregation-and-Guidance Network (TAG-Net). Concretely, we define a new temporal ensemble model with set representation by modeling correspondence between rain regions of the current frame and rain-free regions of adjacent frames. Further, we build a TAG-Net that contains: 1) temporal aggregation network derived from the ensemble model, which is with newly-designed self-directed attention acting on video sequences, it automatically learns temporal correlation from multiple adjacent frames to optimize the current frame. 2) temporal guidance network, which aims at eliminating rain streaks in intersected rain regions between the current and adjacent frames to enhance the previously-recovered frame. Extensive evaluations verify that TAG-Net yields the best performance against other advanced methods.
Long Ma 0002, Risheng Liu, Xin Fan 0001
ICME5
2021 Multiple Task-Oriented Encoders for Unified Image Fusion
abstract
Image fusion methods have achieved incredible progress, but they are vulnerable to handling a certain type of fusion task rather than considering deeper relations between cross-realm task correlations. To achieve this, we integrate different image fusion tasks into a unified network. Our method is accomplished through multiple task-oriented encoders and a generic decoder, in addition to a self-adapting loss function. The taskoriented encoders are trained to learn task-specific features, while the generic decoder reconstructs the fused features to generate a comprehensive image. Subsequently, by introducing the self-adapting loss in our method, it can automatically adjust itself to source data characteristics on different tasks. Besides, we formulate a training strategy based on bilevel optimization to update the multi-encoder and generic decoder in an alternative manner. Extensive experimental results demonstrate the superior performance of our method over the stateof-the-art methods.
Zhuoxiao Li, Jinyuan Liu 0001, Risheng Liu, Xin Fan 0001, Zhongxuan Luo, Wen Gao 0001
ICME4
2021 Spatial-Temporal Integration Network with Self-Guidance for Robust Video Deraining
abstract
Recently, video deraining has become a research focus. Network-based approaches are continuously showing extrusive performance. However, they lack precise control over the motion consistency in temporal information and characterize spatial distribution, so that their results are unsatisfying, especially in some real-world scenarios. To settle them, we develop a spatial-temporal integration network with self-guidance. It contains flow-induced alignment, self-guidance generation, and spatial-temporal integration modules. The alignment module not only preliminarily removes rain to provide more effective temporal correlation but also accurately keeps motion consistency between frames. The self-guidance map characterizes the pixel-level spatial distribution for the target to avoid injuring the background. Finally, we concatenate adjacent aligned frames, self-guidance map, and original current rain frame into the integration module to progressively fuse them in a coarse-to-fine way. Extensive evaluations demonstrate our superiority against other state-of-the-art methods qualitatively and quantitatively.
Xiaokun Liu, Risheng Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICME4
2021 Halder: Hierarchical Attention-Guided Learning with Detail-Refinement for Multi-Exposure Image Fusion
abstract
Deep learning techniques have yielded impressive progress in the field of computational imaging. Existing approaches ignore designing specific constrain on illumination or edges, making them limited in handling asymmetric halos and more likely to generate a fusion result with color discrepancy or blurred edges. To alleviate these issues, we propose a hierarchical attention-guided learning with detail-refinement, termed as HALDeR, to tackle the multi-exposure fusion (MEF) task in a coarse-to-fine manner. Firstly, a hierarchical attention network is designed to produce a fusion result by calculating well-exposed areas under different illumination. Secondly, we develop a collaborative-refine module for preventing the missing details and correcting distorted color simultaneously. Moreover, adversarial learning is employed at end of our network, which can effectively alleviate other remaining artifacts (e.g., ringing effect and noises). Extensive quantitative and qualitative results on two publicly available datasets demonstrate that our HALDeR performs favorably against the state-of-the-art methods in generating vivid color and faithful detail. Source code will be available at https://github.com/JinyuanLiu-CV/HALDeR.
Jinyuan Liu 0001, Jingjie Shang, Risheng Liu, Xin Fan 0001
ICME4
2021 Searching Frame-Recurrent Attentive Deformable Network for Real-Time Video Deraining
abstract
Video deraining has become an issue of great interest since rain streaks inevitably affect video quality. Most of the existing works focus on heuristically designing the network architecture to integrate available information derived from the temporal dimension. However, their inferences take a long time, so that the practicability is somewhat ignored. To solve this problem, we develop a real-time video deraining network in a frame-recurrent manner. It includes a fast attentive deformable alignment module and an automatically-discovered spatial-temporal reconstruction module. In which, the alignment is composed of a single newly-built deformable convolution under the channel attention mechanism to keep the accurate motion consistency and reduce time-consuming by a wide margin. The reconstruction part for the first time introduces the architecture search technique for video deraining to automatically discover a high-effective architecture by designing an effective and compact search space. Experimental results demonstrate remarkable superiority both in computational efficiency and actual performance compared to other state-of-the-art approaches.
Xinwei Xue, Long Ma 0002, Yi Wang 0037, Risheng Liu, Xin Fan 0001
ICME6
2021 Hardware-Aware Low-Light Image Enhancement via One-Shot Neural Architecture Search with Shrinkage Sampling
abstract
Low-light image enhancement has traditionally been tackled by training a heuristically designed neural network architecture. Despite the success of these approaches, the heuristic design pattern inherently not only hinders further optimization of network architectures, but also limits the factors that the designer can take into consideration. As a result, these methods are difficult to achieve a balance between enhancing performance and hardware related performance. In this paper, we equip a basic enhancing algorithm with a neural architecture search technique. This technique helps to automatically search an optimal hardware-aware architecture while also increases neglectable computation burden. In this work, we propose a shrinkage sampling strategy to drastically decrease the computation cost of neural architecture search while improving the quality of search. Extensive experiments on various benchmarks demonstrate that our algorithm achieves state-of-the-art performance with higher speed.
Yuansheng Yao, Risheng Liu, Jiaao Zhang, Xin Fan 0001, Zhongxuan Luo
ICME5
2021 Star-Net: Spatial-Temporal Attention Residual Network for Video Deraining
abstract
Learning-based video deraining has recently drawn increasing attention. They tend to directly package aligned frames to input a fully end-to-end network. However, the network is generally object-driven and cannot recognize how to utilize temporal information so that the results are unsatisfied. In this work, we design a novel Spatial-Temporal Attention Network (STAR-Net) to explicitly utilize the temporal information. Concretely, we define the self-spatial attention to characterizing the rain region of the target frame, and the temporal-spatial attention to learn the profitable information for remedying the rain region of the target frame from the adjacent frame. We also introduce a simple residual network to further strengthen the relationship between the target and the adjacent frame. These addressed frames are fused by a three-layers convolutional module to further improve the capability. Extensive evaluations indicate our superiority against state-of-the-art methods.
Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME5
2021 Collaborative Reflectance-And-Illumination Learning For High-Efficient Low-Light Image Enhancement
abstract
In this paper, we settle the low-light image enhancement problem by developing a collaborative learning framework, which not only improves lightness and suppresses noises simultaneously but also with fast speed and requires few computational resources. The approach is inspired by the fact that reflectance and illumination are highly correlated to satisfy the well-known Retinex decomposition principle. With this in mind, we establish a Reflectance-and-Illumination Collaborative (RIC) block to depict the compact physical relationship between reflectance and illumination. By cascading multiple RIC blocks, we obtain an end-to-end RICNet to interactively optimize these two components in a collaborative manner. Benefiting from the RIC block that integrates powerful task cues, RICNet just needs few parameters to simultaneously improve brightness and remove noises. Extensive experiments demonstrate our superiority against existing state-of-the-art methods. We also make meticulous analysis for the RIC block. The results reveal the rationality and effectiveness of our built mechanism.
Guijing Zhu, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME4
2021 Underwater Species Detection using Channel Sharpening Attention
abstract
With the continuous exploration of marine resources, underwater artificial intelligent robots play an increasingly important role in the fish industry. However, the detection of underwater objects is a very challenging problem due to the irregular movement of underwater objects, the occlusion of sand and rocks, the diversity of water illumination, and the poor visibility and low color contrast in the underwater environment. In this article, we first propose a real-world underwater object detection dataset (UODD), which covers more than 3K images of the most common aquatic products. Then we propose Channel Sharpening Attention Module (CSAM) as a plug-and-play module to further fuse high-level image information, providing the network with the privilege of selecting feature maps. Fusion of original images through CSAM can improve the accuracy of detecting small and medium objects, thereby improving the overall detection accuracy. We also use Water-Net as a preprocessing method to remove the haze and color cast in complex underwater scenes, which shows a satisfactory detection result on small-sized objects. In addition, we use the class weighted loss as the training loss, which can accurately describe the relationship between classification and precision of bounding boxes of targets, and the loss function converges faster during the training process. Experimental results show that the proposed method reaches a maximum AP of 50.1%, outperforming other traditional and state-of-the-art detectors. In addition, our model only needs an average inference time of 25.4 ms per image, which is quite fast and might suit the real-time scenario.
Lihao Jiang, Yi Wang 0037, Qi Jia 0001, Shengwei Xu, Yu Liu 0012, Xin Fan 0001, Risheng Liu, Xinwei Xue, Ruili Wang 0001
ACM Multimedia6
2021 Bridging the Gap between Low-Light Scenes: Bilevel Learning for Fast Adaptation
abstract
Brightening low-light images of diverse scenes is a challenging but widely concerned task in the multimedia community. Convolutional Neural Networks (CNNs) based approaches mostly acquire the enhanced model by learning the data distribution from the specific scenes. However, these works present poor adaptability (even fail) when meeting real-world scenarios that never encountered before. To conquer it, we develop a novel bilevel learning scheme for fast adaptation to bridge the gap between low-light scenes. Concretely, we construct a Retinex-induced encoder-decoder with an adaptive denoising mechanism, aiming at covering more practical cases. Different from existing works that directly learn model parameters by using the massive data, we provide a new hyperparameter optimization perspective to formulate a bilevel learning scheme towards general low-light scenarios. This scheme depicts the latent correspondence (i.e., scene-irrelevant encoder) and the respective characteristic (i.e., scene-specific decoder) among different data distributions. Due to the expensive inner optimization, estimating the hyper-parameter gradient exactly can be prohibitive, we develop an approximate hyper-parameter gradient method by introducing the one-step forward approximation and finite difference approximation to ensure the high-efficient inference. Extensive experiments are conducted to reveal our superiority against other state-of-the-art methods. A series of analytical experiments are also executed to verify our effectiveness.
Dian Jin 0003, Long Ma 0002, Risheng Liu, Xin Fan 0001
ACM Multimedia4
2021 Searching a Hierarchically Aggregated Fusion Architecture for Fast Multi-Modality Image Fusion
abstract
Multi-modality image fusion refers to generating a complementary image that integrates typical characteristics from source images. In recent years, we have witnessed the remarkable progress of deep learning models for multi-modality fusion. Existing CNN-based approaches strain every nerve to design various architectures for realizing these tasks in an end-to-end manner. However, these handcrafted designs are unable to cope with the high demanding fusion tasks, resulting in blurred targets and lost textural details. To alleviate these issues, in this paper, we propose a novel approach, aiming at searching effective architectures according to various modality principles and fusion mechanisms.
Risheng Liu, Zhu Liu 0004, Jinyuan Liu 0001, Xin Fan 0001
ACM Multimedia4
2021 Latency-Constrained Spatial-Temporal Aggregated Architecture Search for Video Deraining
Zhu Liu 0004, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo, Yuduo Zhang
PRCV (3)4
2021 Semantic-Driven Context Aggregation Network for Underwater Image Enhancement
Dongxiang Shi, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
PRCV (3)4
2021 Deep Multi-Illumination Fusion for Low-Light Image Enhancement
Long Ma 0002, Risheng Liu, Xin Fan 0001
PRCV (3)5
2021 Novelty Detection and Online Learning for Chunk Data Streams
abstract
Datastream analysis aims at extracting discriminative information for classification from continuously incoming samples. It is extremely challenging to detect novel data while incrementally updating the model efficiently and stably, especially for high-dimensional and/or large-scale data streams. This paper proposes an efficient framework for novelty detection and incremental learning for unlabeled chunk data streams. First, an accurate factorization-free kernel discriminative analysis (FKDA-X) is put forward through solving a linear system in the kernel space. FKDA-X produces a Reproducing Kernel Hilbert Space (RKHS), in which unlabeled chunk data can be detected and classified by multiple known-classes in a single decision model with a deterministic classification boundary. Moreover, based on FKDA-X, two optimal methods FKDA-CX and FKDA-C are proposed. FKDA-CX uses the micro-cluster centers of original data as the input to achieve excellent performance in novelty detection. FKDA-C and incremental FKDA-C (IFKDA-C) using the class centers of original data as their input have extremely fast speed in online learning. Theoretical analysis and experimental validation on under-sampled and large-scale real-world datasets demonstrate that the proposed algorithms make it possible to learn unlabeled chunk data streams with significantly lower computational costs and comparable accuracies than the state-of-the-art approaches.
Yi Wang 0037, Xiangjian He, Xin Fan 0001, Chi Lin 0001, Fengqi Li, Tianzhu Wang, Zhongxuan Luo, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Learning to Discover a Unified Architecture for Low-Level Vision
abstract
Neural Architecture Search (NAS) has pioneered various constructive principles to push forward the development of deep learning and achieved dramatic performances for diverse tasks recently. Existing NAS methods mainly focus on a single specific task to discover the architecture automatically. But actually, these methods lack ample exploitation and exploration for the latent ability of architecture search mechanism, e.g., from diverse cross-task distributions to discover a unified architecture automatically. In this work, we propose a Cross-task Differentiable ARchiTecture Search (Cross-DARTS for short) framework to discover a unified architecture for different low-level vision tasks automatically, to further widen the capacity of NAS. Specifically, we establish a new model to bridge different low-level vision tasks under the architecture search perspective. By performing a new data construction that integrates multi-task distributions, Cross-DARTS is obtained based on the differentiable search scheme. A multi-scale fusion cell with powerful contextual representation capacity is designed as the basic component of search space towards the low-level vision. Consistent achievements of promising results on three vision tasks, including noise, rain, joint rain and haze removal fully show our superiority.
Zhu Liu 0004, Long Ma 0002, Risheng Liu, Xin Fan 0001
IEEE Signal Process. Lett.4
2021 SMoA: Searching a Modality-Oriented Architecture for Infrared and Visible Image Fusion
abstract
Nowadays, driven by the high demand for autonomous driving and surveillance, infrared and visible image fusion (IVIF) has attracted significant attention from both the industry and research community. Existing learning-based IVIF methods tried to design various architectures to extract features. Still, these hand-crafted designed architectures cannot adequately represent the typical features of different modalities, resulting in undesirable artifacts on the fused results. To alleviate this issue, we propose a Neural Architecture Search (NAS)-based deep learning network to realize the IVIF task, which can automatically discover the modality-oriented feature representation. Our network is accomplished through two modality-oriented encoders and a unified decoder, in addition to a self-visual saliency weight module (SvSW). The two modality-oriented encoders target to learn different intrinsic feature representations automatically from infrared-/visible- modality images. Subsequently, these intermediate features are merged via the SvSW module. Finally, the fused image is recovered by a unified decoder. Extensive experiments demonstrate that our method outperforms the state-of-the-art approaches by a large margin, especially in generating distinct targets and abundant details.
Jinyuan Liu 0001, Zhanbo Huang, Risheng Liu, Xin Fan 0001
IEEE Signal Process. Lett.5
2021 Learning Hadamard-Product-Propagation for Image Dehazing and Beyond
abstract
Image dehazing has evolved into an attractive research field in the computer vision community in the past few decades. Previous traditional approaches attempt to design energy-based objective functions. However, they cannot accurately express the intrinsic characteristics of the images, posing weak adaptation ability for real-world complex scenarios. More recently, deep learning techniques for image dehazing have matured and become more reliable, showing outstanding performance. Nevertheless, these methods heavily depend on training data, restricting their application ranges. More importantly, both traditional and deep learning approaches all ignore a common issue, noises/artifacts always appear in the recovery process. To this end, a new Hadamard-Product (HP) model is proposed, which consists of a series of data-driven priors. Based on this model, we derive a Learnable Hadamard-Product-Propagation (LHPP) by cascading a series of principle-inspired guidance and recovery modules. In which, the principle-inspired guidance related to transmission is endowed the smoothness property, the other recovery module satisfies the distribution of natural images. The Hadamard-product-based propagations is generated in our developed learnable framework for the task of image dehazing. In this way, we can eliminate noises/artifacts in the recovery procedure to obtain the ideal outputs. Subsequently, since the generality of our HP model, we successfully extend our LHPP to settle low-light image enhancement and underwater image enhancement problems. A series of analytical experiments are performed to verify our effectiveness. Plenty of performance evaluations on three complex tasks fully reveal our superiority against multiple state-of-the-art methods.
Risheng Liu, Jinyuan Liu 0001, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.5
2021 A Bilevel Integrated Model With Data-Driven Layer Ensemble for Multi-Modality Image Fusion
abstract
Image fusion plays a critical role in a variety of vision and learning applications. Current fusion approaches are designed to characterize source images, focusing on a certain type of fusion task while limited in a wide scenario. Moreover, other fusion strategies (i.e., weighted averaging, choose-max) cannot undertake the challenging fusion tasks, which furthermore leads to undesirable artifacts facilely emerged in their fused results. In this paper, we propose a generic image fusion method with a bilevel optimization paradigm, targeting on multi-modality image fusion tasks. Corresponding alternation optimization is conducted on certain components decoupled from source images. Via adaptive integration weight maps, we are able to get the flexible fusion strategy across multi-modality images. We successfully applied it to three types of image fusion tasks, including infrared and visible, computed tomography and magnetic resonance imaging, and magnetic resonance imaging and single-photon emission computed tomography image fusion. Results highlight the performance and versatility of our approach from both quantitative and qualitative aspects.
Risheng Liu, Jinyuan Liu 0001, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.4
2021 Unsupervised Monocular Depth Estimation via Recursive Stereo Distillation
abstract
Existing unsupervised monocular depth estimation methods resort to stereo image pairs instead of ground-truth depth maps as supervision to predict scene depth. Constrained by the type of monocular input in testing phase, they fail to fully exploit the stereo information through the network during training, leading to the unsatisfactory performance of depth estimation. Therefore, we propose a novel architecture which consists of a monocular network (Mono-Net) that infers depth maps from monocular inputs, and a stereo network (Stereo-Net) that further excavates the stereo information by taking stereo pairs as input. During training, the sophisticated Stereo-Net guides the learning of Mono-Net and devotes to enhance the performance of Mono-Net without changing its network structure and increasing its computational burden. Thus, monocular depth estimation with superior performance and fast runtime can be achieved in testing phase by only using the lightweight Mono-Net. For the proposed framework, our core idea lies in: 1) how to design the Stereo-Net so that it can accurately estimate depth maps by fully exploiting the stereo information; 2) how to use the sophisticated Stereo-Net to improve the performance of Mono-Net. To this end, we propose a recursive estimation and refinement strategy for Stereo-Net to boost its performance of depth estimation. Meanwhile, a multi-space knowledge distillation scheme is designed to help Mono-Net amalgamate the knowledge and master the expertise from Stereo-Net in a multi-scale fashion. Experiments demonstrate that our method achieves the superior performance of monocular depth estimation in comparison with other state-of-the-art methods.
Xinchen Ye, Xin Fan 0001, Mingliang Zhang 0002, Rui Xu 0002
IEEE Trans. Image Process.2
2021 Dual Neural Networks Coupling Data Regression With Explicit Priors for Monocular 3D Face Reconstruction
abstract
We address the challenging issue of reconstructing a 3D face from one single image under various expressions and illuminations, which is widely applied in multimedia tasks. Methods built upon classical parametric morphable models (3DMMs) gain success on reconstructing the global geometry of a 3D face, but fail to precisely characterize local facial details. Recently, deep neural networks (DNN) have been applied to the reconstruction that directly predicts depth maps, showing compelling performance on detail recovery. Unfortunately, their reconstruction is prone to structural distortions owing to the lack of explicit prior constraints. In this paper, we propose dual neural networks that optimize one energy coupling data fitting with local explicit geometric prior. Specifically, we build one residual network upon traditional convolution layers in order to directly predict 3D structures by fitting an input image. Meanwhile, we devise a novel architecture stacking shallow networks to refine 3D clouds with geometric priors given by Markov random fields (MRFs). Quantitative evaluations demonstrate the superior performance of the dual networks over either end-to-end DNNs or parametric models. Comparisons with the state-of-the-art also show competitive reconstruction quality on various conditions.
Xin Fan 0001, Shichao Cheng, Kang Huyan, Minjun Hou, Risheng Liu, Zhongxuan Luo
IEEE Trans. Multim.1
2021 Location-Aware and Regularization-Adaptive Correlation Filters for Robust Visual Tracking
abstract
Correlation filter (CF) has recently been widely used for visual tracking. The estimation of the search window and the filter-learning strategies is the key component of the CF trackers. Nevertheless, prevalent CF models separately address these issues in heuristic manners. The commonly used CF models directly set the estimated location in the previous frame as the search center for the current one. Moreover, these models usually rely on simple and fixed regularization for filter learning, and thus, their performance is compromised by the search window size and optimization heuristics. To break these limits, this article proposes a location-aware and regularization-adaptive CF (LRCF) for robust visual tracking. LRCF establishes a novel bilevel optimization model to address simultaneously the location-estimation and filter-training problems. We prove that our bilevel formulation can successfully obtain a globally converged CF and the corresponding object location in a collaborative manner. Moreover, based on the LRCF framework, we design two trackers named LRCF-S and LRCF-SA and a series of comparisons to prove the flexibility and effectiveness of the LRCF framework. Extensive experiments on different challenging benchmark data sets demonstrate that our LRCF trackers perform favorably against the state-of-the-art methods in practice.
Risheng Liu, Qianru Chen, Yuansheng Yao, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.4
2020 Image Restoration Via Data-Dependent Proximal Averaged Optimization
abstract
Maximum A Posterior (MAP) acts as one of the most popular modeling scheme in image restoration and is usually reduced to a separable optimization model. Unfortunately, it is challenging to establish exact regularization term and the model with complex priors is hard to optimize. In additionally, it is still hard to incorporate different domain knowledge and data-dependent information into MAP model without changing the property of the objective. To partially address the above issues, we develop a Data-dependent Proximal Averaged (DPA) paradigm through optimizing objective and data-dependent feasibility constraint for the challenging Image Restoration (IR) tasks. Both visual and quantitative comparison results demonstrate that our method outperforms the state of the art.
Pan Mu, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICASSP5
2020 Sequential Deep Unrolling With Flow Priors For Robust Video Deraining
abstract
Video deraining has attracted wide attention since the urgent demand of high-quality video in recent years. The indistinct details and nonideal deraining effects are the most common defects in existing techniques, whose cause lies in the insufficient usage of single-frame image and temporal information. To effectively settle video deraining, we establish a new deraining model with flow priors to simultaneously introduce spatial and temporal information for accurately depicting the enhancement model of the current frame. A sequential deep unrolling framework is substantially presented by solving this model based on optimization techniques. The ablation study indicates our effectiveness as far as the design of architecture. Plenty of subjective and objective evaluations fully demonstrate our superiority in detail recovery and deraining effects against other state-of-the-are video deraining approaches.
Xinwei Xue, Ying Ding 0006, Pan Mu, Long Ma 0002, Risheng Liu, Xin Fan 0001
ICASSP6
2020 Principle-Inspired Multi-Scale Aggregation Network for Extremely Low-Light Image Enhancement
abstract
The under-exposure and low-light environments are common to degrade the image-quality with invisible information. To ameliorate this case, a copious of low-light image enhancement methods are developed. However, these existing works are hard to handle extremely low-light conditions with noises, even well-known network-based methods. To address this issue, we develop a Principle-inspired Multi-scale Aggregation Network (PMA-Net) to simultaneously achieve the exposure enhancement and noises removal. Specifically, we establish a pioneering principle-inspired connection to present the physical principle in the inside of the network, to strengthen the structural depict. Subsequently, we propose a multi-scale aggregation strategy to eliminate the noises in the enhanced results. Sufficient ablation studies manifest the effectiveness of our PMA-Net. Extensive qualitative and quantitative comparisons with other state-of-the-art methods are conducted to fully indicates our outstanding performance.
Jiaao Zhang, Risheng Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICASSP5
2020 An Efficient Ellipse Detector Based On Region Detection And Arc Pruning
abstract
Detecting ellipses accurately and efficiently for real-world images is crucial for various visual-based applications. Most existing methods employ detection strategies throughout images, while most time is spent on the non-ellipse region. Meanwhile, small ellipses are often miss-detected due to the low resolution and fixed parameters of detectors. In this paper, we proposed an effective ellipse detector benefiting from the region detection method, which provides a basic estimation on the region and size of ellipses. Then, a two-level arc pruning strategy is proposed to detect ellipses efficiently while limiting false-positive and false-negative results. Furthermore, for the pre-estimated region without detected ellipses, interpolation method is employed to enlarge the target region, which makes small and blur ellipses to be detected. Experimental results demonstrate that the proposed method achieves competitive accuracy compared with the state-of-the-art methods.
Ruike Zhang, Jingchao Liang, Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo
ICIP5
2020 Flexible Bilevel Image Layer Modeling For Robust Deraining
abstract
Visual quality degradation by rain streaks in images/videos is a significant factor that makes many computer vision systems fail to function properly. However, existing rain removal methods tend to remove a specific type of rain streaks while cannot deal with diverse real rainy images. In this paper, we formulate a novel rain model collectively with two contrasting rain streaks and a weighting map. To self-adaptively handle the rain removal problem in the presence of various types of rain streaks, we further propose a bilevel optimization learning framework. Then, we synthesize a new dataset to evaluate the ability of our method to deal with diverse rain streaks. Extensive experiments show that our method can make better performance on both synthesized and real rainy images.
Pan Mu, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME4
2020 Cascaded Detail-Aware Network for Unsupervised Monocular Depth Estimation
abstract
Existing unsupervised learning methods usually reformulate the depth estimation into the image reconstruction problem by training on stereo image pairs to circumvent the need of dense labeled ground truth depth information. Most of them are designed based on a simple encoder-decoder backbone architecture, which has limited expression for context information and suffers from the loss of depth details. In this paper, we propose a cascaded detail-aware network which contains a contextual network (CN) followed by consecutive spatial networks (SNs) to make an unsupervised coarse-to-fine prediction. CN aims to provide good initialized depth estimation results by introducing a multi-scale attention fusion module to enhance the ability of feature representation. Then, SN is progressively applied on the coarse depth map to produce refined depth outputs by exploiting abundant spatial details from input color image. Moreover, we design a robust loss function that further considers the penalty of photometric errors and the occlusion, and strengthens the recovery of spatial details for better depth estimation. Experimental results show that the proposed method achieves promising performance.
Xinchen Ye, Mingliang Zhang 0002, Xin Fan 0001, Rui Xu 0002, Juncheng Pu, Ruoke Yan
ICME3
2020 Bi-level Probabilistic Feature Learning for Deformable Image Registration
abstract
We address the challenging issue of deformable registration that robustly and efficiently builds dense correspondences between images. Traditional approaches upon iterative energy optimization typically invoke expensive computational load. Recent learning-based methods are able to efficiently predict deformation maps by incorporating learnable deep networks. Unfortunately, these deep networks are designated to learn deterministic features for classification tasks, which are not necessarily optimal for registration. In this paper, we propose a novel bi-level optimization model that enables jointly learning deformation maps and features for image registration. The bi-level model takes the energy for deformation computation as the upper-level optimization while formulates the maximum \emph{a posterior} (MAP) for features as the lower-level optimization. Further, we design learnable deep networks to simultaneously optimize the cooperative bi-level model, yielding robust and efficient registration. These deep networks derived from our bi-level optimization constitute an unsupervised end-to-end framework for learning both features and deformations. Extensive experiments of image-to-atlas and image-to-image deformable registration on 3D brain MR datasets demonstrate that we achieve state-of-the-art performance in terms of accuracy, efficiency, and robustness.
Risheng Liu, Yuxi Zhang 0001, Xin Fan 0001, Zhongxuan Luo
IJCAI4
2020 Coupling Deep Textural and Shape Features for Sketch Recognition
abstract
Recognizing freehand sketches with high arbitrariness is such a great challenge that the automatic recognition rate has reached a ceiling in recent years. In this paper, we explicitly explore the shape properties of sketches, which has almost been neglected before in the context of deep learning, and propose a sequential dual learning strategy that combines both shape and texture features. We devise a two-stage recurrent neural network to balance these two types of features. Our architecture also considers stroke orders of sketches to reduce the intra-class variations of input features. Extensive experiments on the TU-Berlin benchmark set show that our method achieves over 90% recognition rate for the first time on this task, outperforming both humans and state-of-the-art algorithms by over 19 and 7.5 percentage points, respectively. Especially, our approach can distinguish the sketches with similar textures but different shapes more effectively than recent deep networks. Based on the proposed method, we develop an on-line sketch retrieval and imitation application to teach children or adults to draw. The application is available as Sketch.Draw.
Qi Jia 0001, Xin Fan 0001, Meiyu Yu, Yuqing Liu 0001, Dingrong Wang, Longin Jan Latecki
ACM Multimedia2
2020 Learning Multi-scale Retinex with Residual Network for Low-Light Image Enhancement
Long Ma 0002, Jingjie Shang, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
PRCV (1)5
2020 Blind image deblurring via hybrid deep priors modeling
Shichao Cheng, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
Neurocomputing4
2020 Unsupervised detail-preserving network for high quality monocular depth estimation
Mingliang Zhang 0002, Xinchen Ye, Xin Fan 0001
Neurocomputing3
2020 Unsupervised depth estimation from monocular videos with hybrid geometric-refined loss and contextual attention
Mingliang Zhang 0002, Xinchen Ye, Xin Fan 0001
Neurocomputing3
2020 On the Convergence of Learning-Based Iterative Methods for Nonconvex Inverse Problems
abstract
Numerous tasks at the core of statistics, learning and vision areas are specific cases of ill-posed inverse problems. Recently, learning-based (e.g., deep) iterative methods have been empirically shown to be useful for these problems. Nevertheless, integrating learnable structures into iterations is still a laborious process, which can only be guided by intuitions or empirical insights. Moreover, there is a lack of rigorous analysis about the convergence behaviors of these reimplemented iterations, and thus the significance of such methods is a little bit vague. This paper moves beyond these limits and proposes Flexible Iterative Modularization Algorithm (FIMA), a generic and provable paradigm for nonconvex inverse problems. Our theoretical analysis reveals that FIMA allows us to generate globally convergent trajectories for learning-based iterative methods. Meanwhile, the devised scheduling policies on flexible modules should also be beneficial for classical numerical methods in the nonconvex scenario. Extensive experiments on real applications verify the superiority of FIMA.
Risheng Liu, Shichao Cheng, Xin Fan 0001, Zhouchen Lin, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 A sparsity-promoting image decomposition model for depth recovery
Xinchen Ye, Mingliang Zhang 0002, Jing-Yu Yang 0002, Xin Fan 0001, Fangfang Guo
Pattern Recognit.4
2020 Joint Over and Under Exposures Correction by Aggregated Retinex Propagation for Image Enhancement
abstract
Since the interference of ambient light and the limitation of physical devices, it is quite a common phenomenon that images taken in real-world scenarios turn out to be incorrectly exposed. Most existing techniques emphasize underexposed image correction. On one hand, these works ignore the correction of over-exposure regions in the original input. On the other hand, it is likely to generate over-exposure images. To mitigate these issues, we have developed a novel aggregated Retinex propagations to simultaneously correct over and under-exposure correction of a single image. Concretely, we first manifest the necessity of concurrently correcting under and over-exposure appearances. We establish a Retinex image propagation framework with shared weights to correct different levels of exposure. Then by introducing the fusion computational module, we achieve the accurate exposure correction for a single image. Plenty of quantitative and qualitative comparisons are conducted to fully indicate our superiority against other state-of-the-art algorithms. The elaborated algorithmic analyses show our effectiveness. Experiments on face detection further verify our practicability.
Long Ma 0002, Dian Jin 0003, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
IEEE Signal Process. Lett.4
2020 Real-World Underwater Enhancement: Challenges, Benchmarks, and Solutions Under Natural Light
abstract
Underwater image enhancement is such an important low-level vision task with many applications that numerous algorithms have been proposed in recent years. These algorithms developed upon various assumptions demonstrate successes from various aspects using different data sets and different metrics. In this work, we setup an undersea image capturing system, and construct a large-scale Real-world Underwater Image Enhancement (RUIE) data set divided into three subsets. The three subsets target at three challenging aspects for enhancement, i.e., image visibility quality, color casts, and higher-level detection/classification, respectively. We conduct extensive and systematic experiments on RUIE to evaluate the effectiveness and limitations of various algorithms to enhance visibility and correct color casts on images with hierarchical categories of degradation. Moreover, underwater image enhancement in practice usually serves as a preprocessing step for mid-level and high-level vision tasks. We thus exploit the object detection performance on enhanced images as a brand new task-specific evaluation criterion. The findings from these evaluations not only confirm what is commonly believed, but also suggest promising solutions and new directions for visibility enhancement, color correction, and object detection on real-world underwater images. The benchmark is available at: https://github.com/dlut-dimt/Realworld-Underwater-Image-Enhancement-RUIE-Benchmark.
Risheng Liu, Xin Fan 0001, Ming Zhu 0001, Minjun Hou, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.2
2020 Deep Joint Depth Estimation and Color Correction From Monocular Underwater Images Based on Unsupervised Adaptation Networks
abstract
Degraded visibility and geometrical distortion typically make the underwater vision more intractable than open air vision, which impedes the development of underwater-related machine vision and robotic perception. Therefore, this paper addresses the problem of joint underwater depth estimation and color correction from monocular underwater images, which aims at enjoying the mutual benefits between these two related tasks from a multi-task perspective. Our core ideas lie in our new deep learning architecture. Due to the lack of effective underwater training data, and the weak generalization to the real-world underwater images trained on synthetic data, we consider the problem from a novel perspective of style-level and feature-level adaptation, and propose an unsupervised adaptation network to deal with the joint learning problem. Specifically, a style adaptation network (SAN) is first proposed to learn a style-level transformation to adapt in-air images to the style of underwater domain. Then, we formulate a task network (TN) to jointly estimate the scene depth and correct the color from a single underwater image by learning domain-invariant representations. The whole framework can be trained end-to-end in an adversarial learning manner. Extensive experiments are conducted under air-to-water domain adaptation settings. We show that the proposed method performs favorably against state-of-the-art methods in both depth estimation and color correction tasks.
Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Rui Xu 0002, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.7
2020 Detection and Defense of Cache Pollution Attacks Using Clustering in Named Data Networks
abstract
Named Data Network (NDN), as a promising information-centric networking architecture, is expected to support next-generation of large-scale content distribution with open in-network cachings. However, such open in-network caches are vulnerable against Cache Pollution Attacks (CPAs) with the goal of filling cache storage with non-popular contents. The detection and defense against such attacks are especially difficult because of CPA's similarities with normal fluctuations of content requests. In this work, we use a clustering technique to detect and defend against CPAs. By clustering the content interests, our scheme is able to distinguish whether they have followed the Zipf-like distribution or not for accurate detections. Once any attack is detected, an attack table will be updated to record the abnormal requests. While such requests are still forwarded, the corresponding content chunks are not cached. Extensive simulations in ndnSIM demonstrate that our scheme can resist CPA effectively with higher cache hit, higher detecting ratio, lower hop count, and lower algorithm complexity compared to other state-of-the-art schemes.
Lin Yao 0001, Zhenzhen Fan, Jing Deng 0001, Xin Fan 0001, Guowei Wu 0001
IEEE Trans. Dependable Secur. Comput.4
2020 Investigating Task-Driven Latent Feasibility for Nonconvex Image Modeling
abstract
Properly modeling latent image distributions plays an important role in a variety of image-related vision problems. Most exiting approaches aim to formulate this problem as optimization models (e.g., Maximum A Posterior, MAP) with handcrafted priors. In recent years, different CNN modules are also considered as deep priors to regularize the image modeling process. However, these explicit regularization techniques require deep understandings on the problem and elaborately mathematical skills. In this work, we provide a new perspective, named Task-driven Latent Feasibility (TLF), to incorporate specific task information to narrow down the solution space for the optimization-based image modeling problem. Thanks to the flexibility of TLF, both designed and trained constraints can be embedded into the optimization process. By introducing control mechanisms based on the monotonicity and boundedness conditions, we can also strictly prove the convergence of our proposed inference process. We demonstrate that different types of image modeling problems, such as image deblurring and rain streaks removals, can all be appropriately addressed within our TLF framework. Extensive experiments also verify the theoretical results and show the advantages of our method against existing state-of-the-art approaches.
Risheng Liu, Pan Mu, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.4
2020 A Deep Framework Assembling Principled Modules for CS-MRI: Unrolling Perspective, Convergence Behaviors, and Practical Modeling
abstract
Compressed Sensing Magnetic Resonance Imaging (CS-MRI) significantly accelerates MR acquisition at a sampling rate much lower than the Nyquist criterion. A major challenge for CS-MRI lies in solving the severely ill-posed inverse problem to reconstruct aliasing-free MR images from the sparse k -space data. Conventional methods typically optimize an energy function, producing restoration of high quality, but their iterative numerical solvers unavoidably bring extremely large time consumption. Recent deep techniques provide fast restoration by either learning direct prediction to final reconstruction or plugging learned modules into the energy optimizer. Nevertheless, these data-driven predictors cannot guarantee the reconstruction following principled constraints underlying the domain knowledge so that the reliability of their reconstruction process is questionable. In this paper, we propose a deep framework assembling principled modules for CS-MRI that fuses learning strategy with the iterative solver of a conventional reconstruction energy. This framework embeds an optimal condition checking mechanism, fostering efficient and reliable reconstruction. We also apply the framework to three practical tasks, i.e., complex-valued data reconstruction, parallel imaging and reconstruction with Rician noise. Extensive experiments on both benchmark and manufacturer-testing images demonstrate that the proposed method reliably converges to the optimal solution more efficiently and accurately than the state-of-the-art in various scenarios.
Risheng Liu, Yuxi Zhang 0001, Shichao Cheng, Zhongxuan Luo, Xin Fan 0001
IEEE Trans. Medical Imaging5
2020 Knowledge-Driven Deep Unrolling for Robust Image Layer Separation
abstract
Single-image layer separation targets to decompose the observed image into two independent components in terms of different application demands. It is known that many vision and multimedia applications can be (re)formulated as a separation problem. Due to the fundamentally ill-posed natural of these separations, existing methods are inclined to investigate model priors on the separated components elaborately. Nevertheless, it is knotty to optimize the cost function with complicated model regularizations. Effectiveness is greatly conceded by the settled iteration mechanism, and the adaption cannot be guaranteed due to the poor data fitting. What is more, for a universal framework, the most taxing point is that one type of visual cue cannot be shared with different tasks. To partly overcome the weaknesses mentioned earlier, we delve into a generic optimization unrolling technique to incorporate deep architectures into iterations for adaptive image layer separation. First, we propose a general energy model with implicit priors, which is based on maximum a posterior, and employ the extensively accepted alternating direction method of multiplier to determine our elementary iteration mechanism. By unrolling with one general residual architecture prior and one task-specific prior, we attain a straightforward, flexible, and data-dependent image separation framework successfully. We apply our method to four different tasks, including single-image-rain streak removal, high-dynamic-range tone mapping, low-light image enhancement, and single-image reflection removal. Extensive experiments demonstrate that the proposed method is applicable to multiple tasks and outperforms the state of the arts by a large margin qualitatively and quantitatively.
Risheng Liu, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.3
2019 A Theoretically Guaranteed Deep Optimization Framework for Robust Compressive Sensing MRI
abstract
Magnetic Resonance Imaging (MRI) is one of the most dynamic and safe imaging techniques available for clinical applications. However, the rather slow speed of MRI acquisitions limits the patient throughput and potential indications. Compressive Sensing (CS) has proven to be an efficient technique for accelerating MRI acquisition. The most widely used CS-MRI model, founded on the premise of reconstructing an image from an incompletely filled k-space, leads to an ill-posed inverse problem. In the past years, lots of efforts have been made to efficiently optimize the CS-MRI model. Inspired by deep learning techniques, some preliminary works have tried to incorporate deep architectures into CS-MRI process. Unfortunately, the convergence issues (due to the experience-based networks) and the robustness (i.e., lack real-world noise modeling) of these deeply trained optimization methods are still missing. In this work, we develop a new paradigm to integrate designed numerical solvers and the data-driven architectures for CS-MRI. By introducing an optimal condition checking mechanism, we can successfully prove the convergence of our established deep CS-MRI optimization scheme. Furthermore, we explicitly formulate the Rician noise distributions within our framework and obtain an extended CS-MRI network to handle the real-world nosies in the MRI process. Extensive experimental results verify that the proposed paradigm outperforms the existing state-of-theart techniques both in reconstruction accuracy and efficiency as well as robustness to noises in real scene.
Risheng Liu, Yuxi Zhang 0001, Shichao Cheng, Xin Fan 0001, Zhongxuan Luo
AAAI4
2019 Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous Driving
abstract
In this paper, we propose a monocular 3D object detection framework in the domain of autonomous driving. Unlike previous image-based methods which focus on RGB feature extracted from 2D images, our method solves this problem in the reconstructed 3D space in order to exploit 3D contexts explicitly. To this end, we first leverage a stand-alone module to transform the input data from 2D image plane to 3D point clouds space for a better input representation, then we perform the 3D detection using PointNet backbone net to obtain objects' 3D locations, dimensions and orientations. To enhance the discriminative capability of point clouds, we propose a multi-modal feature fusion module to embed the complementary RGB cue into the generated point clouds representation. We argue that it is more effective to infer the 3D bounding boxes from the generated 3D scene space (i.e., X,Y, Z space) compared to the image plane (i.e., R,G,B image plane). Evaluation on the challenging KITTI dataset shows that our approach boosts the performance of state-of-the-art monocular approach by a large margin.
Xinzhu Ma, Zhihui Wang 0001, Wanli Ouyang, Xin Fan 0001
ICCV6
2019 Compounded Layer-Prior Unrolling: A Unified Transmission-Based Image Enhancement Framework
abstract
Improving the quality of images degraded by various transmission media has important practical significance. Such enhancement tasks involve resolving both transmission degradation and residual contamination including imaging noise, color distortion, and occlusions. Existing methods typically develop the priors on natural scenes to resolve ill-posed problems separately. However, the solutions derived from hand-crafted priors may fail on specific regions where a priori assumptions break, and recent data-driven methods highly depend on training data owing to the absence of effective priors. Based on a unified formulation for transmission-based image enhancement tasks, we develop a compounded unrolling framework to generate hybrid image layer propagations. Specifically, as multiple deeply-trained priors are integrated into the iterative propagation scheme, the deep model can recognize specific task properties and data distributions for different applications. Both quantitative and qualitative experiments demonstrate the superior performance of the proposed framework on various transmission-based tasks (haze removal, underwater image enhancement and rain removal).
Risheng Liu, Minjun Hou, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo
ICME4
2019 Enhanced Residual Dense Intrinsic Network for Intrinsic Image Decomposition
abstract
Intrinsic image decomposition is a challenging task, which aims at recovering intrinsic components from the observation. Hand-crafted priors have been widely used in traditional methods, yet with unsatisfactory performance of quality and runtime. Recently, network-based approaches have been greatly developed, but the physical imaging principle is ignored causing the multiplication of estimated components is hard to reconstruct the observation. To overcome these limitations, we develop an enhanced residual dense intrinsic network (ERDIN) for intrinsic decomposition. Specifically, we construct the basic module (i.e., enhanced residual dense block (ERDB)) to fully exploit the hierarchical features. The physical imaging principle is designed as the reconstruction loss to ensure the consistency between the observation and the multiplication of estimated components, which is of equal importance with the data loss. Extensive experimental results illustrate our excellent performance compared with other state-of-the-art methods.
Risheng Liu, Long Ma 0002, Miao Zhang 0004, Xin Fan 0001, Zhongxuan Luo
ICME5
2019 Unsupervised Monocular Depth Estimation Based on Dual Attention Mechanism and Depth-Aware Loss
abstract
Most existing monocular depth estimation approaches are su- pervised, but enough quantities of ground truth depth data are required during training. To cope with this, recent techniques deal with the depth estimation task in an unsupervised man- ner, i.e., replacing the use of depth data with easily obtained stereo images for training. Based on this, we propose a nov- el unsupervised learning architecture, which integrates dual attention mechanism into the framework and designs a depth- aware loss for better depth estimation. Specifically, to en- hance the ability of feature representations, we introduce a d- ual attention module to capture global feature dependencies in spatial and channel dimensions for scene understanding and depth estimation. Meanwhile, we propose a depth-aware loss that fully addresses the occlusion problem in brightness con- stancy assumption, the intrinsic characteristics of depth map, and the left-right consistency problem, respectively. Besides, an adversarial loss is employed to discriminate synthetic or realistic depth maps by training a discriminator so as to pro- duce better results. Extensive experiments on KITTI dataset show that our approach achieves state-of-the-art performance compared with other monocular depth estimation methods.
Xinchen Ye, Mingliang Zhang 0002, Rui Xu 0002, Xin Fan 0001, Zhu Liu 0004, Jiaao Zhang
ICME5
2019 Fast density-peaks clustering for registration-free pediatric white matter tract analysis
Xin Fan 0001, Yuzhuo Duan, Shichao Cheng, Yuxi Zhang 0001
Artif. Intell. Medicine1
2019 Learning Bilevel Layer Priors for Single Image Rain Streaks Removal
abstract
Rain streaks removal is an important issue of the outdoor vision system and recently has been investigated extensively. In the past decades, maximum a posterior and network-based architecture have been attracting considerable attention for this problem. However, it is challenging to establish effective regularization priors and the cost function with complex prior is hard to optimize. On the other hand, it is still hard to incorporate data-dependent information into conventional numerical iterations. To partially address the above limits and inspired by the leader-follower gaming perspective, we introduce an unrolling strategy to incorporate data-dependent network architectures into the established iterations, i.e., a learning bilevel layer priors method to jointly investigate the learnable feasibility and optimality of rain streaks removal problem. Both visual and quantitative comparison results demonstrate that our method outperforms the state of the art.
Pan Mu, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
IEEE Signal Process. Lett.4
2019 Deep Proximal Unrolling: Algorithmic Framework, Convergence Analysis and Applications
abstract
Deep learning models have gained great success in many real-world applications. However, most existing networks are typically designed in heuristic manners, thus these approaches lack of rigorous mathematical derivations and clear interpretations. Several recent studies try to build deep models by unrolling a particular optimization model that involves task information. Unfortunately, due to the dynamic nature of network parameters, their resultant deep propagations do not possess the nice convergence property as the original optimization scheme does. In this work, we develop a generic paradigm to unroll nonconvex optimization for deep model design. Different from most existing frameworks, which just replace the iterations by network architectures, we prove in theory that the propagation generated by our proximally unrolled deep model can globally converge to the critical-point of the original optimization model. Moreover, even if the task information is only partially available (e.g., no prior regularization), we can still train a convergent deep propagations. We also extend these theoretical investigations on the more general multi-block models and thus a lot of real-world applications can be successfully handled by the proposed framework. Finally, we conduct experiments on various low-level vision tasks (i.e., non-blind deconvolution, dehazing, and low-light image enhancement) and demonstrate the superiority of our proposed framework, compared with existing state-of-the-art approaches.
Risheng Liu, Shichao Cheng, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.4
2019 Learning Aggregated Transmission Propagation Networks for Haze Removal and Beyond
abstract
Single-image dehazing is an important low-level vision task with many applications. Early studies have investigated different kinds of visual priors to address this problem. However, they may fail when their assumptions are not valid on specific images. Recent deep networks also achieve a relatively good performance in this task. But unfortunately, due to the disappreciation of rich physical rules in hazes, a large amount of data are required for their training. More importantly, they may still fail when there exist completely different haze distributions in testing images. By considering the collaborations of these two perspectives, this paper designs a novel residual architecture to aggregate both prior (i.e., domain knowledge) and data (i.e., haze distribution) information to propagate transmissions for scene radiance estimation. We further present a variational energy-based perspective to investigate the intrinsic propagation behavior of our aggregated deep model. In this way, we actually bridge the gap between prior-driven models and data-driven networks and leverage advantages but avoid limitations of previous dehazing approaches. A lightweight learning framework is proposed to train our propagation network. Finally, by introducing a task-aware image separation formulation with a flexible optimization scheme, we extend the proposed model for more challenging vision tasks, such as underwater image enhancement and single-image rain removal. Experiments on both synthetic and real-world images demonstrate the effectiveness and efficiency of the proposed framework.
Risheng Liu, Xin Fan 0001, Minjun Hou, Zhiying Jiang, Zhongxuan Luo, Lei Zhang 0006
IEEE Trans. Neural Networks Learn. Syst.2
2019 Fast example searching for input-adaptive data-driven dehazing with Gaussian process regression
Xin Fan 0001, Xianxuan Tang, Minjun Hou, Zhongxuan Luo
Vis. Comput.1
2018 Self-Reinforced Cascaded Regression for Face Alignment
Xin Fan 0001, Risheng Liu, Kang Huyan, Yuyao Feng, Zhongxuan Luo
AAAI1
2018 Proximal Alternating Direction Network: A Globally Converged Deep Unrolling Framework
abstract
Deep learning models have gained great success in many real-world applications. However, most existing networks are typically designed in heuristic manners, thus lack of rigorous mathematical principles and derivations. Several recent studies build deep structures by unrolling a particular optimization model that involves task information. Unfortunately, due to the dynamic nature of network parameters, their resultant deep propagation networks do not possess the nice convergence property as the original optimization scheme does. This paper provides a novel proximal unrolling framework to establish deep models by integrating experimentally verified network architectures and rich cues of the tasks. More importantly,we prove in theory that 1) the propagation generated by our unrolled deep model globally converges to a critical-point of a given variational energy, and 2) the proposed framework is still able to learn priors from training data to generate a convergent propagation even when task information is only partially available. Indeed, these theoretical results are the best we can ask for, unless stronger assumptions are enforced. Extensive experiments on various real-world applications verify the theoretical convergence and demonstrate the effectiveness of designed deep models.
Risheng Liu, Xin Fan 0001, Shichao Cheng, Zhongxuan Luo
AAAI2
2018 Deep Layer Prior Optimization for Single Image Rain Streaks Removal
abstract
Visible distortions caused by rain streaks have significant negative effects on the performance of many vision and learning algorithms. Most of the existing deraining approaches propose to build complex prior models to formulate the appearance of rain streaks. Unfortunately, these human-designed priors tend to over-smooth the background and leave too many rain streaks since the distribution of rain streaks is complex and disordered. In this work, we exploit a deep layer prior under the maximum a posterior framework to recover the intrinsic rain structure. The optimization of the resulted variational energy can be understood as simultaneously performing rain and image propagations based on data-dependent residual networks and task cues (e.g., total variation regularization), respectively. Experimental results on both synthetic and real test images demonstrate the effectiveness of our approach against both designed priors and fully data-dependent convolutional neural networks.
Risheng Liu, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICASSP4
2018 Robust Haze Removal Via Joint Deep Transmission and Scene Propagation
abstract
Haze is one of the most important factors which reduce the outdoor image quality. Existing approaches often aim to design their models based on principles of hazes. However, even with exactly modeled haze distribution, it is still a challenging task due to factors in real scenario, such as noises, halos and artifacts. To address limitations of existing approaches for real-world hazy removal problem, this paper proposes a novel framework to incorporate deep residual architectures into a propagation scheme to jointly estimate transmission and clean scene. We evaluate the proposed framework on both widely used benchmarks and real-world low-quality hazy images. Extensive experimental results demonstrate that our method performs favorably against approaches designed only based on haze cues and achieves the state-of-the-art results, compared with both conventional shallow models and deep dehzaing networks.
Risheng Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICASSP4
2018 Gradient Regression for Brain Landmark Localization on Magnetic Resonance Imaging
abstract
Landmark localization in human brain from Magnetic Resonance Imaging (MRI) is primarily important for numerous medical analysis applications. Recently developed regression based (including deep networks) methods typically learn a mapping from input features to individual landmark positions or transform parameters. These methods neglect the geometric correlations among landmarks, thus resulting in inaccurate localization, especially for the parcellation functional regions whose boundaries are composed of a bunch of landmarks. In this paper, we build a shape energy for landmarks on 3D M-RI features and learn the gradient regression for the energy. Our method accelerates the iterative gradient calculation and accurately detect brain landmarks. We validate the algorithm on two localization tasks for two key points, anterior commissure (AC) and posterior commissure (PC), and for three functional regions on the OASIS TI-weighted MR data set. Experimental results demonstrate its efficiency and effectiveness by comparing with the state-of-the-art.
Yuzhuo Duan, Xin Fan 0001, Huiying Kang
ICIP2
2018 Joint Residual Learning for Underwater Image Enhancement
abstract
Improving the quality of underwater image has a significant impact on many signal processing and computer vision applications, while haze-effect and color shift are main handicaps need to be surmounted. Due to the complexity of the underwater environmental factors, most existing image enhancement techniques cannot be directly applied to address this task. In this work, we develop a novel framework to jointly performing residual learning on transmission and image domains for underwater scene entrenchment. Indeed, our deep model consists of a data-driven residual architecture for transmission estimation and a knowledge-driven scene residual formulation for underwater illumination balance. Therefore, we can aggregate the prior knowledge and data information to investigate the underlying underwater image distribution. Moreover, by introducing adaptive exposure map, image colors will also be corrected accordingly. Experimentally, both quantitative and qualitative analysis can indicate outstanding effectiveness of the proposed algorithm, against state-of-the-art approaches.
Minjun Hou, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICIP3
2018 Single Image Layer Separation via Deep Admm Unrolling
abstract
Single image layer separation aims to divide the observed image into two independent components according to special task requirements and has been widely used in many vision and multimedia applications. Because this task is fundamentally ill-posed, most existing approaches tend to design complex priors on the separated layers. However, the cost function with complex prior regularization is hard to optimize. The performance is also compromised by fixed iteration schemes and less data fitting ability. More importantly, it is also challenging to design a unified framework to separate image layers for different applications. To partially mitigate the above limitations, we develop a flexible optimization unrolling technique to incorporate deep architectures into iterations for adaptive image layer separation. Specifically, we first design a general energy model with implicit priors and adopt the widely used alternating direction method of multiplier (ADMM) to establish our basic iteration scheme. By unrolling with residual convolution architectures, we successfully obtain a simple, flexible, and data-dependent image separation method. Extensive experiments on the tasks of rain streak removal and reflection removal validate the effectiveness of our approach.
Risheng Liu, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo
ICME3
2018 Toward Designing Convergent Deep Operator Splitting Methods for Task-specific Nonconvex Optimization
abstract
Operator splitting methods have been successfully used in computational sciences, statistics, learning and vision areas to reduce complex problems into a series of simpler subproblems. However, prevalent splitting schemes are mostly established only based on the mathematical properties of some general optimization models. So it is a laborious process and often requires many iterations of ideation and validation to obtain practical and task-specific optimal solutions, especially for nonconvex problems in real-world scenarios. To break through the above limits, we introduce a new algorithmic framework, called Learnable Bregman Splitting (LBS), to perform deep-architecture-based operator splitting for nonconvex optimization based on specific task model. Thanks to the data-dependent (i.e., learnable) nature, our LBS can not only speed up the convergence, but also avoid unwanted trivial solutions for real-world tasks. Though with inexact deep iterations, we can still establish the global convergence and estimate the asymptotic convergence rate of LBS only by enforcing some fairly loose assumptions. Extensive experiments on different applications (e.g., image completion and deblurring) verify our theoretical results and show the superiority of LBS against existing methods.
Risheng Liu, Shichao Cheng, Xin Fan 0001, Zhongxuan Luo
IJCAI4
2018 Fast Factorization-free Kernel Learning for Unlabeled Chunk Data Streams
abstract
Data stream analysis aims at extracting discriminative information for classification from continuously incoming samples. It is extremely challenging to detect novel data while updating the model in an efficient and stable fashion, especially for the chunk data. This paper proposes a fast factorization-free kernel learning method to unify novelty detection and incremental learning for unlabeled chunk data streams in one framework. The proposed method constructs a joint reproducing kernel Hilbert space from known class centers by solving a linear system in kernel space. Naturally, unlabeled data can be detected and classified among multi-classes by a single decision model. And projecting samples into the discriminative feature space turns out to be the product of two small-sized kernel matrices without needing such time-consuming factorization like QR-decomposition or singular value decomposition. Moreover, the insertion of a novel class can be treated as the addition of a new orthogonal basis to the existing feature space, resulting in fast and stable updating schemes. Both theoretical analysis and experimental validation on real-world datasets demonstrate that the proposed methods learn chunk data streams with significantly lower computational costs and comparable or superior accuracy than the state of the art.
Yi Wang 0037, Nan Xue 0004, Xin Fan 0001, Jiebo Luo 0001, Risheng Liu, Zhongxuan Luo
IJCAI3
2018 Learning Collaborative Generation Correction Modules for Blind Image Deblurring and Beyond
abstract
Blind image deblurring plays a very important role in many vision and multimedia applications. Most existing works tend to introduce complex priors to estimate the sharp image structures for blur kernel estimation. However, it has been verified that directly optimizing these models is challenging and easy to fall into degenerate solutions. Although several experience-based heuristic inference strategies, including trained networks and designed iterations, have been developed, it is still hard to obtain theoretically guaranteed accurate solutions. In this work, a collaborative learning framework is established to address the above issues. Specifically, we first design two modules, named Generator and Corrector, to extract the intrinsic image structures from the data-driven and knowledge-based perspectives, respectively. By introducing a collaborative methodology to cascade these modules, we can strictly prove the convergence of our image propagations to a deblurring-related optimal solution. As a nontrivial byproduct, we also apply the proposed method to address other related tasks, such as image interpolation and edge-preserved smoothing. Plenty of experiments demonstrate that our method can outperform the state-of-the-art approaches on both synthetic and real datasets.
Risheng Liu, Shichao Cheng, Xin Fan 0001, Zhongxuan Luo
ACM Multimedia4
2018 A Bridging Framework for Model Optimization and Deep Propagation
abstract
Optimizing task-related mathematical model is one of the most fundamental methodologies in statistic and learning areas. However, generally designed schematic iterations may hard to investigate complex data distributions in real-world applications. Recently, training deep propagations (i.e., networks) has gained promising performance in some particular tasks. Unfortunately, existing networks are often built in heuristic manners, thus lack of principled interpretations and solid theoretical supports. In this work, we provide a new paradigm, named Propagation and Optimization based Deep Model (PODM), to bridge the gaps between these different mechanisms (i.e., model optimization and deep propagation). On the one hand, we utilize PODM as a deeply trained solver for model optimization. Different from these existing network based iterations, which often lack theoretical investigations, we provide strict convergence analysis for PODM in the challenging nonconvex and nonsmooth scenarios. On the other hand, by relaxing the model constraints and performing end-to-end training, we also develop a PODM based strategy to integrate domain knowledge (formulated as models) and real data distributions (learned by networks), resulting in a generic ensemble framework for challenging real-world applications. Extensive experiments verify our theoretical results and demonstrate the superiority of PODM against these state-of-the-art approaches.
Risheng Liu, Shichao Cheng, Xiaokun Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
NeurIPS5
2018 Designing a stable feedback control system for blind image deconvolution
Shichao Cheng, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
Neural Networks3
2018 Line matching based on line-points invariant and local homography
Qi Jia 0001, Xin Fan 0001, Xinkai Gao, Meiyu Yu, Zhongxuan Luo
Pattern Recognit.2
2018 Explicit Shape Regression With Characteristic Number for Facial Landmark Localization
abstract
Robustly localizing facial landmarks plays a very important role in many multimedia and vision applications. Most recently proposed regression-based methods prevailing in the community lack explicit shape constraints for faces and require a large number of facial images to cover great appearance variations. To address these limitations, this paper introduces a novel projective invariant called characteristic number (CN) to explicitly characterize the intrinsic geometries of facial points shared by human faces. It can be verified that the shape priors from CN are inherently invariant to pose changes. By further developing a shape-to-gradient regression framework, we provide a robust and efficient landmark detector for facial images in the wild. The computation of our model can be successfully addressed by learning the descent directions using point-CN pairs without the need for large collections for appearance training. As a nontrivial byproduct, this paper also builds a face dataset, where each face has 15 well-defined viewpoints (poses) to quantitatively analyze the effects of different poses on localization methods. Extensive experiments on challenging benchmarks and our newly built dataset demonstrate the effectiveness of our proposed detector against other state-of-the-art approaches.
Xin Fan 0001, Risheng Liu, Zhongxuan Luo, Yuyao Feng
IEEE Trans. Multim.1
2017 Fast Online Incremental Learning on Mixture Streaming Data
abstract
The explosion of streaming data poses challenges to feature learning methods including linear discriminant analysis (LDA). Many existing LDA algorithms are not efficient enough to incrementally update with samples that sequentially arrive in various manners. First, we propose a new fast batch LDA (FLDA/QR) learning algorithm that uses the cluster centers to solve a lower triangular system that is optimized by the Cholesky-factorization. To take advantage of the intrinsically incremental mechanism of the matrix, we further develop an exact incremental algorithm (IFLDA/QR). The Gram-Schmidt process with reorthogonalization in IFLDA/QR significantly saves the space and time expenses compared with the rank-one QR-updating of most existing methods. IFLDA/QR is able to handle streaming data containing 1) new labeled samples in the existing classes, 2) samples of an entirely new (novel) class, and more significantly, 3) a chunk of examples mixed with those in 1) and 2). Both theoretical analysis and numerical experiments have demonstrated much lower space and time costs (2~10 times faster) than the state of the art, with comparable classification accuracy.
Yi Wang 0037, Xin Fan 0001, Zhongxuan Luo, Tianzhu Wang, Maomao Min, Jiebo Luo 0001
AAAI2
2017 Incremental zero-shot learning based on attributes for image classification
abstract
Instead of assuming a closed-world environment comprising a fixed number of objects, modern pattern recognition systems need to recognize outliers, identify anomalies, or discover entirely new objects, which is known as zero-shot object recognition. However, many existing zero-shot learning methods are not efficient enough to incrementally update themselves with new samples mixed with known or novel class labels. In this paper, we propose an incremental zero-shot learning framework (IIAP/QR) based on indirect-attribute-prediction (IAP) model. Firstly, a fast incremental classifier based on null space based linear discriminant analysis with QR-updating (NLDA/QR) is put forward, which can solve small-sample-size (SSS) problem and unequal-sample-size (USS) problem that usually occur in incremental learning using the centroid of each class as input. Then with the probabilistic inference of Class-Attribute layer and Attribute-Zero shot classification layer, IIAP/QR model can efficiently update itself for the insertion of both new samples to the existing class and totally novel classes with comparable recognition accuracy for zero-shot object recognition.
Nan Xue 0004, Yi Wang 0037, Xin Fan 0001, Maomao Min
ICIP3
2017 Leveraging geometric correlation for input-adaptive facial landmark regression
abstract
Facial analysis plays very important role in many vision applications, such as authentication and entertainments. The very early works in the 1990s mostly focus on estimating geometric deformations of facial landmarks to address this task. While in the past several years, more and more efforts have been made to directly learn an appearance regression for facial analysis. Though training regressions on controlled facial images can successfully capture the appearance variations, the performance of these appearance-based models are tightly related to the quantity and quality of the training data. In this paper, we develop a novel framework, named geometric correlated landmark regression (GCLR), to inherit the advantages but overcome limitations of these two categories of methods. Specifically, we first establish a landmark-to-landmark regression to estimate the geometry of facial images. By further incorporating a sparse coding term into the regression framework, we can successfully leverage the geometric correlations between the test image and the shape dictionary, thus significantly enhance the geometry regression performance. Experimental results on various challenging facial data sets verify the effectiveness and efficiency of GCLR.
Yuyao Feng, Risheng Liu, Xin Fan 0001, Kang Huyan, Zhongxuan Luo
ICME3
2017 Blind image deblurring via adaptive dynamical system learning
abstract
Blind image deblurring is one of the main phases in most media analysis tasks. Many existing works aim to simultaneously estimate the latent image and the blur kernel under a MAP framework. However, it has been demonstrated that such joint estimation strategies may lead to the undesired trivial solution. In this paper, we propose a learnable nonlinear dynamical system to formulate the image propagation so that the blur kernel estimation can be efficiently controlled by both cues and training data. Our analysis also indicates that the proposed dynamical system is feasible on image modeling socialities. Experimental results on different benchmark image sets evaluate the effectiveness of our proposed approach.
Risheng Liu, Shichao Cheng, Xin Fan 0001, Zhongxuan Luo
ICME3
2017 Deep hybrid residual learning with statistic priors for single image super-resolution
abstract
This paper considers single image super-resolution (SISR), which is an important low-level vision task and has various applications in multimedia society. Recently, deep neural networks have archived good performance on this field. But most of existing deep models are based on the fully data-dependent network architecture, thus missing majority of domain-knowledge of the super-resolution task. To address this limitation, we develop a new hybrid residual learning approach to leverage priors of SISR within the maximum a posteriori framework for network architecture design. We demonstrate that it can incorporate both image priors and data fidelity into the network, leading to a novel cascaded residual learning system for SISR process. Extensive experimental results on real-world images show that the proposed algorithm performs favorably against state-of-the-art methods.
Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME3
2017 Compact CNN Based Video Representation for Efficient Video Copy Detection
Xin Fan 0001, Zhongxuan Luo
MMM (1)4
2017 Adaptive low-rank subspace learning with online optimization for robust visual tracking
Risheng Liu, Di Wang 0018, Yuzhuo Han, Xin Fan 0001, Zhongxuan Luo
Neural Networks4
2017 Two-Layer Gaussian Process Regression With Example Selection for Image Dehazing
abstract
Researchers have devoted great efforts to image dehazing with prior assumptions in the past decade. Recently developed example-based approaches typically lack elegant models for the hazy process and meanwhile demand synthetic hazy images by manual selection. The priors from observations, and those trained from synthetic images cannot always reflect true structural information of natural images in practice. In this paper, we present a learning model for haze removal by using two-layer Gaussian process regression (GPR). By using training examples, the two-layer GPR establishes a direct relationship from the input image to the depth-dependent transmission, and learns local image priors to further improve the estimation. We also provide a systematic scheme to automatically collect suitable training pairs, which works for both simulated examples and images of natural scenes. Both qualitative and quantitative comparisons on real-world and synthetic hazy images demonstrate the effectiveness of the proposed approach, especially for white or bright objects and heavy haze regions in which traditional methods may fail.
Xin Fan 0001, Yi Wang 0037, Xianxuan Tang, Renjie Gao, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.1
2017 A Fast Ellipse Detector Using Projective Invariant Pruning
abstract
Detecting elliptical objects from an image is a central task in robot navigation and industrial diagnosis, where the detection time is always a critical issue. Existing methods are hardly applicable to these real-time scenarios of limited hardware resource due to the huge number of fragment candidates (edges or arcs) for fitting ellipse equations. In this paper, we present a fast algorithm detecting ellipses with high accuracy. The algorithm leverages a newly developed projective invariant to significantly prune the undesired candidates and to pick out elliptical ones. The invariant is able to reflect the intrinsic geometry of a planar curve, giving the value of -1 on any three collinear points and +1 for any six points on an ellipse. Thus, we apply the pruning and picking by simply comparing these binary values. Moreover, the calculation of the invariant only involves the determinant of a 3×3 matrix. Extensive experiments on three challenging data sets with 648 images demonstrate that our detector runs 20%-50% faster than the state-of-the-art algorithms with the comparable or higher precision.
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Lianbo Song, Tie Qiu 0001
IEEE Trans. Image Process.2
2016 Novel Coplanar Line-Points Invariants for Robust Line Matching Across Views
Qi Jia 0001, Xinkai Gao, Xin Fan 0001, Zhongxuan Luo, Ziyao Chen
ECCV (8)3
2016 Discriminative Feature Learning with an Optimal Pattern Model for Image Classification
Xin Fan 0001, Zhongxuan Luo
MMM (1)4
2016 Discriminant Manifold Learning via Sparse Coding for Image Analysis
Binghui Wang, Xin Fan 0001, Chuang Lin 0001
MMM (2)3
2016 An efficient mesh-based face beautifier on mobile devices
Xin Fan 0001, Yuyao Feng, Yi Wang 0037, Shengfa Wang, Zhongxuan Luo
Neurocomputing1
2016 Cross-view action matching using a novel projective invariant on non-coplanar space-time points
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Kang Huyan, Zezhou Li
Multim. Tools Appl.2
2016 Hierarchical projective invariant contexts for shape recognition
Qi Jia 0001, Xin Fan 0001, Yu Liu 0012, Zhongxuan Luo, He Guo 0001
Pattern Recognit.2
2016 3D facial landmark localization using texture regression via conformal mapping
Xin Fan 0001, Qi Jia 0001, Kang Huyan, Xianfeng Gu, Zhongxuan Luo
Pattern Recognit. Lett.1
2016 Image morphing with conformal welding
Xin Fan 0001, Yuyao Feng, Xianfeng Gu, Zhongxuan Luo
Vis. Comput.1
2016 Haze editing with natural transmission
Xin Fan 0001, Yi Wang 0037, Renjie Gao, Zhongxuan Luo
Vis. Comput.1
2016 Generalized rational Bézier curves for the rigid body motion design
Zhongxuan Luo, Xin Fan 0001, Yaqi Gao, Panpan Shui
Vis. Comput.3
2015 Sparse concept discriminant matrix factorization for image representation
abstract
Over the past few decades, matrix factorization has attracted considerable attention for image representation. It is desired for a matrix factorization technique to find the basis that is able to capture highly discriminant information as well as to preserve the intrinsic manifold structure. Besides, the basis has to generate a sparse representation for a given image. In this paper, we propose a matrix factorization method called Sparse concept Discriminant Matrix Factorization (SDMF) by combining a novel fisher-like criterion with the sparse coding. The criterion is discriminant enough across different feature spaces, and meanwhile maintains locally neighboring structures. The proposed method is general for both cases with and without class labels, hence yielding supervised and un-supervised SDMFs. Experimental results show that SDMF provides better representation with higher performance on two tasks (image recognition and clustering) compared with the existing matrix factorization methods.
Chuang Lin 0001, Risheng Liu, Xin Fan 0001, Jifeng Jiang, Zhongxuan Luo
ICIP4
2015 Characteristic number regression for facial feature extraction
abstract
Facial feature extraction plays an important role in many multimedia and vision applications. Recent regression methods for extraction lack the explicit shape constraints for faces, and require a large number of facial images covering great appearance variations. This paper introduces a novel projective invariant, named characteristic number (CN), to explicitly characterize the intrinsic geometries of facial points shared by human faces, which is inherently invariant to pose changes. By further developing a shape-to-gradient regression framework, we provide a robust and efficient feature extractor for facial images in the wild. The computation of our model can be successfully addressed by learning the descent directions using point-CN pairs without the need of large collections for appearance training. Extensive experiments on challenging benchmark data sets demonstrate the effectiveness of our proposed detector against other state-of-the-art approaches.
Xin Fan 0001, Risheng Liu, Yuyao Feng, Zhongxuan Luo, Zezhou Li
ICME2
2015 Online gesture-based interaction with visual oriental characters based on manifold learning
Yi Wang 0037, Xin Fan 0001, Xiangjian He, Qi Jia 0001, Renjie Gao
Signal Process.3
2015 Fiducial Facial Point Extraction Using a Novel Projective Invariant
abstract
Automatic extraction of fiducial facial points is one of the key steps to face tracking, recognition, and animation.Great facial variations, especially pose or viewpoint changes,typically degrade the performance of classical methods. Recent learning or regression-based approaches highly rely on the availability of a training set that covers facial variations as wide as possible. In this paper, we introduce and extend a novel projective invariant, named the characteristic number (CN), which unifies the collinearity, cross ratio, and geometrical characteristics given by more (6) points. We derive strong shape priors from CN statistics on a moderate size (515) of frontal upright faces in order to characterize the intrinsic geometries shared by human faces. We combine these shape priors with simple appearance based constraints, e.g., texture, edge, and corner, into a quadratic optimization. Thereafter, the solution to facial point extraction can be found by the standard gradient descent. The inclusion of these shape priors renders the robustness to pose changes owing to their invariance to projective transformations. Extensive experiments on the Labeled Faces in the Wild, Labeled Face Parts in the Wild and Helen database, and cross-set faces with various changes demonstrate the effectiveness of the CN-based shape priors compared with the state of the art.
Xin Fan 0001, Zhongxuan Luo, Daiyun Luo
IEEE Trans. Image Process.1
2014 Fiducial facial point extraction with cross ratio
abstract
Automatic extraction of fiducial facial points is one of the key steps to face tracking, recognition and animation as well as video communication. In this paper, we present a method to localize 8 fiducial points in a face image with cross ratio (CR), a fundamental projective invariant. We derive strong shape priors, which characterize the intrinsic geometries shared by human faces, from CR statistics on a moderate size (515) of frontal upright faces. We combine these shape priors with Gabor textural features and edge/corner into a convex optimization. The Gabor features of local patches and geometric constraints from CR are insensitive to global perspective transformations. Thereafter, the proposed approach renders the robustness to great pose or viewpoint changes. Extensive experiments on facial images from several data sets with great variations on expressions, illuminations and poses demonstrate the effectiveness of the proposed approach.
Xin Fan 0001
ICIP2
2014 Real-time estimation of hand gestures based on manifold learning from monocular videos
Yi Wang 0037, Zhongxuan Luo, Xin Fan 0001, Yunzhen Wu
Multim. Tools Appl.4
2014 A new geometric descriptor for symbols with affine deformations
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Yu Liu 0012, He Guo 0001
Pattern Recognit. Lett.2
2014 Hierarchical Bayes based Adaptive Sparsity in Gaussian Mixture Model
Binghui Wang, Chuang Lin 0001, Xin Fan 0001, Ning Jiang 0001, Dario Farina
Pattern Recognit. Lett.3
2013 A shape matching framework using metric partition constraint
abstract
The crucial problem for shape matching is to balance between discrimination power and computation complexity. Popular solutions mainly rely on either global or local information of shape contours, and neglect their intrinsic correlation. But the methods that combine both information may bring high computation complexity. In this paper, we present a shape matching framework, in which a novel shape descriptor named metric partition constraint (MPC) is proposed, and many metric methods can be included. The metric information is used to bridge the local points and the global shape. Meanwhile, we devise a partition smoothing process to improve the robustness to local deformation. Finally, Comprehensive comparisons with the classical shape context and other latest methods on standard datasets show the excellent performance in terms of precision while retaining computational efficiency.
Yu Liu 0012, Qi Jia 0001, He Guo 0001, Xin Fan 0001
ICIP4
2013 A shape descriptor based on new projective invariants
abstract
Great attention has been devoted to the development of shape descriptors that is the key to object recognition. Previous works have great success on either relatively simple shapes or limited transformations, e.g., translation, rotation and scaling. We propose a new projective invariant, named characteristic number (CN) that includes more points for complex shapes with rich inner structures. Moreover, we build a novel shape descriptor with CN values calculated on triangles that cover the convex hull of a shape. The matching based on the descriptor also runs fast since only one initial point for the triangular coverage needs to align based on its CN value prior to the matching. The performance of the proposed descriptor is validated by the experiments compared with the classical shape context (SC) and recently developed cross ratio spectrum (CRS) on 32 logos of television networks with a wide range of transformations (512 images in total).
Zhongxuan Luo, Daiyun Luo, Xin Fan 0001, Xinchen Zhou, Qi Jia 0001
ICIP3
2012 Haze filtering with aerial perspective
abstract
In this paper, we present haze filtering that is capable of editing the amount of haze in an image given a haze observation. Aerial perspective is taken into account to generate depth dependent haze. We re-formulate the transmission estimation, the key to haze removal or filtering so that users are able to change the amount of haze in an image by tuning maximum visibility. The guided filter is employed in order to efficiently refine the estimated transmission. Additionally, we develop color correction and sky compensation based on physical priors for quality improvements. Experimental results show that the proposed method is able to generate images with various degree of haze in a natural and efficient fashion. The results are also free of color distortion that typically occurs when shooting in fog weather.
Renjie Gao, Xin Fan 0001, Jielin Zhang, Zhongxuan Luo
ICIP2
2012 Face recognition using average invariant factor
abstract
The recent developed intrinsic discriminate analysis (IDA) demonstrates superior recognition rate compared with classical methods such as PCA and LDA. In this paper, we not only re-prove the core theorem of IDA from a new perspective, but also define the Average Invariant Factor (AIF) that generalizes IDA. Two new algorithms for face recognition are built upon the AIF by using SVD and QR decomposition. Moreover, this new formulation facilitates the kernel extensions for the recognition algorithms, which relax the linear assumption for IDA. The presented kernel based AIF algorithms also significantly lower down the computational expenses of the original IDA method. A series of experiments on YALE and ORL sets demonstrate higher performance in terms of recognition rate and efficiency compared with classical statistical analysis methods (e.g., PCA, KPCA and 2DPCA) and the IDA algorithm.
Zhongxuan Luo, Xin Fan 0001, Jielin Zhang
ICIP3
2012 Adaptive Kalman Filtering for Histogram-Based Appearance Learning in Infrared Imagery
abstract
Targets of interest in video acquired from imaging infrared sensors often exhibit profound appearance variations due to a variety of factors, including complex target maneuvers, ego-motion of the sensor platform, background clutter, etc., making it difficult to maintain a reliable detection process and track lock over extended time periods. Two key issues in overcoming this problem are how to represent the target and how to learn its appearance online. In this paper, we adopt a recent appearance model that estimates the pixel intensity histograms as well as the distribution of local standard deviations in both the foreground and background regions for robust target representation. Appearance learning is then cast as an adaptive Kalman filtering problem where the process and measurement noise variances are both unknown. We formulate this problem using both covariance matching and, for the first time in a visual tracking application, the recent autocovariance least-squares (ALS) method. Although convergence of the ALS algorithm is guaranteed only for the case of globally wide sense stationary process and measurement noises, we demonstrate for the first time that the technique can often be applied with great effectiveness under the much weaker assumption of piecewise stationarity. The performance advantages of the ALS method relative to the classical covariance matching are illustrated by means of simulated stationary and nonstationary systems. Against real data, our results show that the ALS-based algorithm outperforms the covariance matching as well as the traditional histogram similarity-based methods, achieving sub-pixel tracking accuracy against the well-known AMCOM closure sequences and the recent SENSIAC automatic target recognition dataset.
Vijay Venkataraman, Guoliang Fan 0001, Joseph P. Havlicek, Xin Fan 0001, Yan Zhai, Mark B. Yeary
IEEE Trans. Image Process.4
2006 Mining Text and Visual Links to Browse TV Programs in a Web-Like Way
abstract
As the amount of receded TV content is increasing rapidly, people need active and interactive browsing methods. In this paper, we use both text information from closed captions and visual information from video frames to generate links to enable one to explore not only the original video content but also augmented information from the Web. This solution especially shows its superiority when the video content cannot be well represented only by closed captions. A prototype system was implemented and some experiments were carried out to prove the effectiveness and efficiency
Xin Fan 0001, Hisashi Miyamori, Katsumi Tanaka, Mingjing Li
ICME1
2006 Inquiring of the Sights from the Web via Camera Mobiles
abstract
In this paper, we presented an image search service for mobile users. It can be used to acquire related information by taking and sending pictures to the server, for example, getting book reviews by a photo of the cover. The key problem here is to find images that contain the same prominent object as that in the query image. In the literature, local feature based image matching has been proven to outperform those based on global features. When using local features, however, one query image may contain thousands of high dimensional feature vectors. Each feature vector needs to match against millions of features in the database. Therefore, it is critical to design an efficient search scheme. Our proposed matching approach was based on identifying semi-local visual parts from multiple query images. Experiments on two real-world datasets showed that this approach was superior to conventional solutions
Yinghua Zhou, Xin Fan 0001, Xing Xie 0001, Yuchang Gong, Wei-Ying Ma
ICME2
2006 Detecting The Sufficient Display Resolution For Image Browsing
abstract
In image browsing, the resolution greatly affects user’s experience. If an image is down-scaled too much, a considerable amount of information within it will be lost. In this paper, we studied the problem of "What is a sufficient display resolution or scale for an image or an image region?" This problem arises in many real-life applications including image browsing on mobile devices, image adaptation and progressive image delivery. Kullback-Leibler (K-L) distance is employed to measure the information loss and the sufficient display scale is selected based on the information loss curve during image down-sampling. Since the images are presented to viewers finally, some visual characteristics are also taken into account to ensure the precision of the measurement. A user study was carried out to evaluate the performance of our approach. Experimental results show that the approach is in good accord with human perception.
Xin Fan 0001, Xing Xie 0001, Wei-Ying Ma
MDM1
2006 Photo-to-Search: Using Camera Phones to Inquire of the Surrounding World
abstract
With the pervasive use of camera phones, the embedded camera has been considered as a promising HCI manner for mobiles. With necessary technologies, it is possible to become a powerful tool to acquire the information in daily life. We have designed and implemented a system named Photo-to-Search to carry out queries from camera phones simply by taking some photos of interested objects. The captured pictures are compared with a large amount of Web images to select the ones which contain the same prominent object. Consequently, the related information is extracted from the Web pages where the matched images locate. In our demo, data of large buildings, storefronts and products are collected and these kinds of queries are specifically demonstrated to show the efficiency and the effectiveness of our system.
Menglei Jia, Xin Fan 0001, Xing Xie 0001, Mingjing Li, Wei-Ying Ma
MDM2
2006 An attention based spatial adaptation scheme for H.264 videos on mobiles
abstract
When browsing videos in mobile devices, people often feel that resolution greatly affects their perceptual experience in the limited screen size. In this paper, an attention based spatial video adaptation scheme is proposed to overcome display constraints by producing the region of interest. According to the size of the target display, we automatically detect and crop the informative region in each frame to generate a smooth sequence. To avoid costly fully encoding operations, we employ a set of transcoding techniques based on the H.264 standard. Experimental results show that this approach not only improves the perceptual quality but also saves the bandwidth and computation, especially for the videos which are not well edited
Yi Wang 0037, Xin Fan 0001, Houqiang Li, Zhengkai Liu, Mingjing Li
MMM2
2006 An Attention Based Spatial Adaptation Scheme for H.264 Videos on Mobiles
abstract
With the growing popularity of personal digital assistant devices and smart phones, consumers have become increasingly enthusiastic to watching videos from these mobile devices. However, when browsing videos in mobiles, users often feel that the display resolution greatly affects their perceptual experience with the limited screen size. In this paper, an attention based spatial video adaptation scheme is proposed to overcome the display constraints by producing and displaying the region of interest. According to the size of the target display, we automatically detect and crop the informative region in each frame to generate a smooth sequence. To avoid costly full encoding operations, we develop a set of transcoding techniques based on the H.264 standard. Experimental results show that this approach not only improves the perceptual quality but also saves the bandwidth and computation, especially for the videos which have not been well edited.
Yi Wang 0037, Houqiang Li, Xin Fan 0001, Chang Wen Chen
Int. J. Pattern Recognit. Artif. Intell.3
2003 Visual attention based image browsing on mobile devices
abstract
Images have become more and more common in mobile communications. People now can easily take and exchange pictures on the move using their mobile devices and digital cameras. However, a crucial challenge is to provide a better user experience for browsing large images on limited and heterogeneous screen sizes of mobile devices. In this paper, we propose a novel image viewing technique based on an adaptive attention shifting model. A presentation technique named rapid serial visual presentation (RSVP), borrowed from the UI community, is used to simulate the attention shifting process. We show a prototype image viewer developed for pocket PC and conduct some evaluations to demonstrate the effectiveness of our approach.
Xin Fan 0001, Xing Xie 0001, Wei-Ying Ma, HongJiang Zhang, He-Qin Zhou
ICME1
2003 Looking into video frames on small displays
abstract
With the growing popularity of personal digital assistants and smart phones, people have become enthusiastic to watch videos through these mobile devices. However, a crucial challenge is to provide a better user experience for browsing videos on the limited and heterogeneous screen sizes. In this paper, we present a novel approach which allows users to overcome the display constraints by zooming into video frames while browsing. An automatic approach for detecting the focus regions is introduced to minimize the amount of user interaction. In order to improve the quality of output stream, virtual camera control is employed in the system. Preliminary evaluation shows that this approach is an effective way for video browsing on small displays.
Xin Fan 0001, Xing Xie 0001, He-Qin Zhou, Wei-Ying Ma
ACM Multimedia1
2003 A visual attention model for adapting images on small displays
Li-Qun Chen, Xing Xie 0001, Xin Fan 0001, Wei-Ying Ma, HongJiang Zhang, He-Qin Zhou
Multim. Syst.3