EDBT 2026 Demo / reviewers in the wild / expert
Xinghao Ding
dblp:80/7693
· DBLP profile ↗
223ranked-venue papers
8as first author
139since 2021 · last 2026
0000-0003-2288-5287ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 125 · 6 first-author · 71 since 2021Artificial intelligence and machine learning · 95 · 1 first-author · 62 since 2021Applied, interdisciplinary, general and emerging computing · 37 · 1 first-author · 28 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image EnhancementabstractPan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the two modalities. Traditional CNN-based methods typically rely on channel-wise concatenation with fixed convolutional operators, which limits their adaptability to diverse spatial and spectral variations. While cross-attention mechanisms enable global interactions, they are computationally inefficient and may dilute fine-grained correspondences, making it difficult to capture complex semantic relationships. Recent advances in the Multimodal Diffusion Transformer (MMDiT) architecture have demonstrated impressive success in image generation and editing tasks. Unlike cross-attention, MMDiT employs in-context conditioning to facilitate more direct and efficient cross-modal information exchange. In this paper, we propose MMMamba, a cross-modal in-context fusion framework for pan-sharpening, with the flexibility to support image super-resolution in a zero-shot manner. Built upon the Mamba architecture, our design ensures linear computational complexity while maintaining strong cross-modal interaction capacity. Furthermore, we introduce a novel multimodal interleaved (MI) scanning mechanism that facilitates effective information exchange between the PAN and MS modalities. Extensive experiments demonstrate the superior performance of our method compared to existing state-of-the-art (SOTA) techniques across multiple tasks and benchmarks. Yingying Wang 0005, Xuanhua He, Jialing Huang, Suiyun Zhang, Xinghao Ding, Haoxuan Che |
AAAI | 7 |
| 2026 | Self-supervised Multiplex Consensus Mamba for General Image FusionabstractImage fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, general image fusion needs to address a wide range of tasks while improving performance without increasing complexity. To achieve this, we propose SMC-Mamba, a Self-supervised Multiplex Consensus Mamba framework for general image fusion. Specifically, the Modality-Agnostic Feature Enhancement (MAFE) module preserves fine details through adaptive gating and enhances global representations via spatial-channel and frequency rotational scanning. The Multiplex Consensus Cross-modal Mamba (MCCM) module enables dynamic collaboration among experts, reaching a consensus to efficiently integrate complementary information from multiple modalities. The cross-modal scanning within MCCM further strengthens feature interactions across modalities, facilitating seamless integration of critical information from both sources. Additionally, we introduce a Bi-level Self-supervised Contrastive Learning Loss (BSCL), which preserves high-frequency information without increasing computational overhead while simultaneously boosting performance in downstream tasks. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) image fusion algorithms in tasks such as infrared-visible, medical, multi-focus, and multi-exposure fusion, as well as downstream visual tasks. Yingying Wang 0005, Rongjin Zhuang, Hui Zheng 0003, Xuanhua He, Ke Cao 0001, Xiaotong Tu, Xinghao Ding |
AAAI | 7 |
| 2026 | Decompose and Attribute: Boosting Generalizable Open-Set Object Detection via Objectness ScoreabstractOpen-set object detection (OSOD) aims to recognize known object categories while localizing previously unseen instances. However, real-world scenarios often involve co-occurring domain shifts and novel object categories. Existing OSOD methods typically overlook domain shifts, relying on source-trained representations that entangle domain-specific style with semantic content, thereby hindering generalization to both unseen domains and novel categories. To address this challenge, we propose a unified framework, termed DecOmpose and ATtribute (DOAT), which disentangles domain-specific style from semantic structure, thereby facilitating generalizable object detection. DOAT employs wavelet-based feature decomposition to separate style information from high-frequency structural details, thus enabling an explicit separation of domain and category shifts. To account for domain shift, the low-frequency components are perturbed within a style subspace to simulate diverse domain appearances. For unknown object discovery, the high-frequency components are utilized to estimate objectness scores via an attribution mechanism that fuses wavelet energy with semantic distance to known-category prototypes. Extensive experiments on standard open-set benchmarks have demonstrated the superior generalization performance of DOAT. Lichen Wei, Luyao Tang, Chaoqi Chen, Zheyuan Cai, Yue Huang 0001, Xinghao Ding |
AAAI | 7 |
| 2026 | Exploiting point-language models with dual-prompts for 3D anomaly detection
Haote Xu, Xiaolu Chen, Haodi Xu, Yue Huang 0001, Xinghao Ding, Xiaotong Tu |
Expert Syst. Appl. | 6 |
| 2026 | Bridging the synthetic-to-real gap in quantitative MRI mapping via frequency-guided domain adaptation
Linyu Fan, Qizhi Yang, Zejun Wu, Xinghao Ding, Yue Huang 0001, Jianfeng Bao, Shuhui Cai, Congbo Cai |
Pattern Recognit. | 7 |
| 2026 | Temporal-invariant video contrastive learning: A novel perspective from brain knowledge
Wei Lin 0021, Changxing Jing, Xinghao Ding, Huanqiang Zeng |
Pattern Recognit. | 3 |
| 2026 | SimpleRM: A Lightweight Reconstruction-Residual Refinement Module for Time Series Forecasting
Xiaowei Lin, Zheyuan Cai, Yue Huang 0001, Xinghao Ding |
IEEE Signal Process. Lett. | 5 |
| 2026 | Accelerating Adaptive Diffusion and Uncertainty Modeling for Underwater Image EnhancementabstractUnderwater image enhancement (UIE) aims to mitigate wavelength-dependent absorption and multi-path scattering effects, enabling the recovery of natural colors and rich details. Despite notable progress, consistently achieving high-quality enhancement in both fidelity and perceptual clarity remains a fundamental challenge. To address this, we propose the Laplacian domain Dual-Focus Enhancer (DFE), an innovative framework consisting of two stages: adaptive diffusion-accelerated low frequency enhancement (ADALE) and progressive uncertainty driven high-frequency enhancement (PUHE). Specifically, DFE applies a Laplacian transform to decouple the frequency-specific degradations in underwater images, supporting fidelity- and clarity-oriented enhancement along separate pathways. To facilitate high-fidelity restoration, ADALE incorporates an HSV guided optimization mechanism (HSV-OM) to establish a robust color and brightness calibration baseline for the low-frequency diffusion model, adaptively managing basic degradations with minimal sampling steps. Furthermore, to enhance contour and detail perception, PUHE models the uncertainty of reference textures and integrates it with feature modulation to progressively reconstruct multi-scale high-frequency structures. The multi reference underwater texture enhancement (MUTE) dataset fur ther improves image clarity. Extensive experiments demonstrate that our DFE outperforms state-of-the-art (SOTA) methods in both quantitative metrics and visual quality. Xiuna Zeng, Jiaao Peng, Zhenqi Fu, Linyu Fan, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Multim. | 7 |
| 2025 | AGLLDiff: Guiding Diffusion Models Towards Unsupervised Training-free Real-world Low-light Image EnhancementabstractExisting low-light image enhancement (LIE) methods have achieved noteworthy success in solving synthetic distortions, yet they often fall short in practical applications. The limitations arise from two inherent challenges in real-world LIE: 1) the collection of distorted/clean image pairs is often impractical and sometimes even unavailable, and 2) accurately modeling complex degradations presents a non-trivial problem. To overcome them, we propose the Attribute Guidance Diffusion framework (AGLLDiff), a training-free method for effective real-world LIE. Instead of specifically defining the degradation process, AGLLDiff shifts the paradigm and models the desired attributes, such as image exposure, structure and color of normal-light images. These attributes are readily available and impose no assumptions about the degradation process, which guides the diffusion sampling process to a reliable high-quality solution space. Extensive experiments demonstrate that our approach outperforms the current leading unsupervised LIE methods across benchmarks in terms of distortion-based and perceptual-based metrics, and it performs well even in sophisticated wild degradation. Yunlong Lin, Tian Ye 0001, Sixiang Chen, Zhenqi Fu, Yingying Wang 0005, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003, Xinghao Ding |
AAAI | 10 |
| 2025 | DPLUT: Unsupervised Low-light Image Enhancement with Lookup Tables and Diffusion PriorsabstractLow-light image enhancement (LIE) aims at precisely and efficiently recovering an image degraded in poor illumination environments. Recent advanced LIE techniques are using deep neural networks, which require lots of low-normal light image pairs, network parameters, and computational resources. As a result, their practicality is limited. In this work, we devise a novel unsupervised LIE framework based on diffusion priors and lookup tables (DPLUT) to achieve efficient low-light image recovery. The proposed approach comprises two critical components: a light adjustment lookup table (LLUT) and a noise suppression lookup table (NLUT). LLUT is optimized with a set of unsupervised losses. It aims at predicting pixel-wise curve parameters for the dynamic range adjustment of a specific image. NLUT is designed to remove the amplified noise after the light brightens. As diffusion models are sensitive to noise, diffusion priors are introduced to achieve high-performance noise suppression. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in terms of visual quality and efficiency. Yunlong Lin, Zhenqi Fu, Kairun Wen, Tian Ye 0001, Sixiang Chen, Ge Meng, Yingying Wang 0005, Chui Kong, Yue Huang 0001, Xiaotong Tu, Xinghao Ding |
AAAI | 11 |
| 2025 | Accelerated Diffusion via High-Low Frequency Decomposition for Pan-SharpeningabstractPan-sharpening aims to preserve the spectral information of the multi-spectral (MS) image while leveraging the high-frequency details from the guided high-resolution panchromatic (PAN) image to enhance its spatial resolution. The key challenge is how to preserve the spectral information from the MS image and the spatial details from the PAN image as much as possible. Diffusion models have achieved favorable results in image restoration and synthesis tasks but suffer from excessive computational resource and time consumption. In this paper, we design a novel and computationally efficient diffusion-based pan-sharpening network that achieves accelerated diffusion while reducing task complexity by decoupling the high and low-frequency components of the fused image. Specifically, leveraging the information-preserving characteristic of the wavelet transformation, we introduce a Wavelet-based Low-frequency Diffusion Model (WLDM). WLDM generates the low-frequency coefficient of high-resolution MS (HRMS) image from the low-resolution MS (LRMS) image. This approach significantly reduces computational resources and complexity compared to the direct restoration of the HRMS image. Furthermore, we have devised a High-frequency Information Restoration Module (HIRM) to restore the high-frequency information in the HRMS image through the interaction of high-frequency coefficients from the PAN image in three directions. Extensive experiments on three different datasets demonstrate that our method outperforms existing approaches in both quantitative metrics, qualitative metrics, and inference efficiency. Ge Meng, Jingjia Huang, Jingyan Tu, Yingying Wang 0005, Yunlong Lin, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
AAAI | 8 |
| 2025 | Sp3ctralMamba: Physics-Driven Joint State Space Model for Hyperspectral Image ReconstructionabstractHyperspectral image (HSI) reconstruction aims to restore the original 3D HSIs from the 2D hyperspectral snapshot compressive images (SCIs). The key to high-fidelity HSI reconstruction lies in designing refined spatial and spectral attention mechanisms, which are crucial for generating fine-grained representations of HSI based on the limited spatial and spectral information available in SCI. Recently, Mamba has demonstrated remarkable performance and efficiency in modeling spatial correlations. Its implicit attention mechanism generates three orders of magnitude more attention matrices than transformers, significantly raising the performance ceiling for HSI reconstruction. In this paper, we propose a novel joint SSM network named Sp3ctralMamba for HSI reconstruction. Sp3ctralMamba integrates frequency domain knowledge and physical priors to enhance reconstruction quality. Specifically, we first perform hierarchical decomposition of the 3D HSI embedding to mitigate the negative impact of distant bands on reconstruction. Next, we design a joint SSM block S3Mamba (S3MAB) to perform parallel scans of the embeddings from different bands. In addition to the conventional vanilla scan, S3MAB introduces a local scanning scheme to address the reconstruction challenges posed by the spatial sparsity of spectral information. Furthermore, a spiral scanning scheme in the frequency domain is incorporated to enhance the order correlation between different frequency signals. Finally, we introduce energy priors and structural priors to constrain the generation of spectral and spatial representations during the training process. Extensive experiments on both simulated and real datasets demonstrate that Sp3ctralMamba significantly elevates HSI reconstruction performance to a new level, surpassing SOTA methods in both quantitative and qualitative metrics. Ge Meng, Jingyan Tu, Jingjia Huang, Yunlong Lin, Yingying Wang 0005, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
AAAI | 8 |
| 2025 | Bayesian Gaussian Process ODEs via Double Normalizing FlowsabstractGaussian processes have been used to model the vector field of continuous dynamical systems, which are characterized by a probabilistic ordinary differential equation (GP-ODE). Bayesian inference for these models has been extensively studied and applied in tasks such as time series prediction. However, the use of standard GPs with basic kernels like squared exponential kernels has been common in GP-ODE research, limiting the model’s ability to represent complex scenarios. To address this limitation, we introduce normalizing flows to reparameterize the ODE vector field, resulting in a data-driven prior distribution, thereby increasing flexibility and expressive power. We develop a variational inference algorithm that utilizes analytically tractable probability density functions of normalizing flows. Additionally, we also apply normalizing flows to the posterior inference of GP-ODEs to resolve the issue of strong mean-field assumptions. By applying normalizing flows in these ways, our model improves accuracy and uncertainty estimates for Bayesian GP-ODEs. We validate the effectiveness of our approach on simulated dynamical systems and real-world human motion data, including time series prediction and missing data recovery tasks. Jian Xu 0021, Shian Du, Junmei Yang, Xinghao Ding, Delu Zeng, John W. Paisley |
AISTATS | 4 |
| 2025 | Track Any Anomalous Object: A Granular Video Anomaly Detection PipelineabstractVideo anomaly detection (VAD) is crucial in scenarios such as surveillance and autonomous driving, where timely detection of unexpected activities is essential. Albeit existing methods have primarily focused on detecting anomalous objects in videos—either by identifying anomalous frames or objects—they often neglect finer-grained analysis, such as anomalous pixels, which limits their ability to capture a broader range of anomalies. To address this challenge, we propose an innovative VAD framework called Track Any Anomalous Object (TAO), which introduces a Granular Video Anomaly Detection Framework that, for the first time, integrates the detection of multiple fine-grained anomalous objects into a unified framework. Unlike methods that assign anomaly scores to every pixel at each moment, our approach transforms the problem into pixel-level tracking of anomalous objects. By linking anomaly scores to subsequent tasks such as image segmentation and video tracking, our method eliminates the need for threshold selection and achieves more precise anomaly localization, even in long and challenging video sequences. Experiments on extensive datasets demonstrate that TAO achieves state-of-the-art performance, setting a new progress for VAD by providing a practical, granular, and holistic solution. For more information, visit the project page at: https://tao-25.github.io/ Yuzhi Huang, Chenxin Li, Zixu Lin, Yunlong Lin, Hengyu Liu 0007, Wuyang Li, Xinyu Liu 0001, Jiechao Gao, Yue Huang 0001, Xinghao Ding, Yixuan Yuan |
CVPR | 11 |
| 2025 | JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image RestorationabstractVision-centric perception systems struggle with unpredictable and coupled weather degradations in the wild. Current solutions are often limited, as they either depend on specific degradation priors or suffer from significant domain gaps. To enable robust and autonomous operation in real-world conditions, we propose JarvisIR, a VLM-powered agent that leverages the VLM as a controller to manage multiple expert restoration models. To further enhance system robustness, reduce hallucinations, and improve generalizability in real-world adverse weather, JarvisIR employs a novel two-stage framework consisting of supervised fine-tuning and human feedback alignment. Specifically, to address the lack of paired data in real-world scenarios, the human feedback alignment enables the VLM to be fine-tuned effectively on large-scale real-world data in an unsupervised manner. To support the training and evaluation of JarvisIR, we introduce CleanBench, a comprehensive dataset consisting of high-quality and large-scale instruction-responses pairs, including 150K synthetic entries and 80K real entries. Extensive experiments demonstrate that JarvisIR exhibits superior decision-making and restoration capabilities. Compared with existing methods, it achieves a 50% improvement in the average of all perception metrics on CleanBench-Real. Yunlong Lin, Zixu Lin, Haoyu Chen 0003, Panwang Pan, Chenxin Li, Sixiang Chen, Kairun Wen, Yeying Jin, Wenbo Li 0002, Xinghao Ding |
CVPR | 10 |
| 2025 | CLIP-Guided Frequency-Aware Representation Learning for Generalizable Remote-Sensing Image Tampering Detection
Qingyao Wu, Xinghao Ding, Yue Huang 0001, Xiaotong Tu |
ICANN (2) | 5 |
| 2025 | Dynamic Category Queries Transformer for Generalized Few-shot Semantic SegmentationabstractFew-shot segmentation (FSS) tackles data scarcity using multiple priors, but its simplicity limits handling base and novel classes with limited data access. Generalized few-shot semantic segmentation (GFSS) enhances model performance for base classes with abundant data, while novel classes have limited data access, improving generalization with scarce data. Building on the design of query-based segmentation models, which decouple the mask and classification tasks for individual optimization, we here present the Dynamic Category Queries Transformer (DCQ-Former) which forms a novel approach to the GFSS. The proposed DCQ-Former first uses category suggested dynamic queries to perform mask segmentation and category classification tasks on a large amount of base class data. Considering the case when the novel classes only have access to a limited amount of training data, the queries for the novel classes are instead dynamically composed from the base classes in order to prevent the category suggested module from providing limited suggestion queries given the representativeness of the fewshot samples. Extensive experiments on COCO-20iand Pascal-5idatasets show that DCQ-Former achieves superior accuracy and generalization than current state-of-the-art methods. Our code are available at https://github.com/fallpavilion/DCQ-Former. Kunze Huang, Jieyuan Yang, Andreas Jakobsson, Luyao Tang, Xiaotong Tu, Xinghao Ding, Yue Huang 0001 |
ICASSP | 6 |
| 2025 | Efficient Dataset Distillation through Low-Rank Space SamplingabstractHuge amount of data is the key of the success of deep learning, however, redundant information impairs the generalization ability of the model and increases the burden of calculation. Dataset Distillation (DD) compresses the original dataset into a smaller but representative subset for high-quality data and efficient training strategies. Existing works for DD generate synthetic images by treating each image as an independent entity, thereby overlooking the common features among data. This paper proposes a dataset distillation method based on Matching Training Trajectories with Low-rank Space Sampling(MTT-LSS), which uses low-rank approximations to capture multiple low-dimensional manifold subspaces of the original data. The synthetic data is represented by basis vectors and shared dimension mappers from these subspaces, reducing the cost of generating individual data points while effectively minimizing information redundancy. The proposed method is tested on CIFAR-10, CIFAR-100, and SVHN datasets, and outperforms the baseline methods by an average of 9.9%. Hangyang Kong, Xuxiang He, Xiaotong Tu, Xinghao Ding |
ICASSP | 5 |
| 2025 | Efficient Infrared Image Super-Resolution Reconstruction via Guided Filter Coefficients Estimation with Parallax Attention MechanismabstractDue to the spectral range mismatch between the images, building an efficient infrared (IR) image super-resolution algorithm suitable for embedded devices remains a significant challenge. Given that visible images possess more abundant high-frequency information compared to infrared images, we utilize the visible light to guide infrared image super-resolution reconstruction. Specifically, we transfer the reconstruction task to a guided filter learning process, whose coefficients are estimated by joint learning of visible and infrared image to complete the reconstruction through homologous constraints. In order to efficiently predict guided filter coefficients, we design a lightweight network which incorporates reparameterized differential convolution blocks and a feature fusion strategy. Striving to enhance the fusion strategy performance, we utilize parallax attention mechanism to solve the non-pixel registration problem between infrared and visible images. Extensive experiments on two challenging IR image datasets show that our method performs SOTA in terms of PSNR, SSIM and LPIPS as compared to current state-of-the-art approaches while showing its effectiveness and practicality in the edge platform of RK3588. Qingyao Wu, Bosheng Chen, Xiaotong Tu, Xinghao Ding, Yue Huang 0001 |
ICASSP | 5 |
| 2025 | Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
Luyao Tang, Kunze Huang, Chaoqi Chen, Chenxin Li, Xiaotong Tu, Xinghao Ding, Yue Huang 0001 |
ICCV | 7 |
| 2025 | ASGS: Single-Domain Generalizable Open-Set Object Detection via Adaptive Subgraph Searching
Luyao Tang, Chaoqi Chen, Yue Huang 0001, Xinghao Ding |
ICCV | 6 |
| 2025 | Self-supervised Sound Source Localization for UAVs Using GCC-PHAT in Low SNR Environments
Shengbin Ma, Saqlain Abbas, Xinghao Ding, Xiaotong Tu |
ICIC (12) | 4 |
| 2025 | A Single-Channel Drone Noise Reduction Algorithm Based on Speech Harmonic Features
Shengbin Ma, Saqlain Abbas, Xinghao Ding, Xiaotong Tu |
ICIC (9) | 4 |
| 2025 | Demeaned Sparse: Efficient Anomaly Detection by Residual EstimateabstractFrequency-domain image anomaly detection methods can substantially enhance anomaly detection performance, however, they still lack an interpretable theoretical framework to guarantee the effectiveness of the detection process. We propose a novel test to detect anomalies in structural image via a Demeaned Fourier transform (DFT) under factor model framework, and we proof its effectiveness. We also briefly give the asymptotic theories of our test, the asymptotic theory explains why the test can detect anomalies at both the image and pixel levels within the theoretical lower bound. Based on our test, we derive a module called Demeaned Fourier Sparse (DFS) that effectively enhances detection performance in unsupervised anomaly detection tasks, which can construct masks in the Fourier domain and utilize a distribution-free sampling method similar to the bootstrap method. The experimental results indicate that this module can accurately and efficiently generate effective masks for reconstruction-based anomaly detection tasks, thereby enhancing the performance of anomaly detection methods and validating the effectiveness of the theoretical framework. Yifan Fang, Yifei Fang, Ruizhe Chen, Haote Xu, Xinghao Ding, Yue Huang 0001 |
ICML | 5 |
| 2025 | Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up TablesabstractRecently, deep learning-based pan-sharpening algorithms have achieved notable advancements over traditional methods. However, deep learning-based methods incur substantial computational overhead during inference, especially with large images. This excessive computational demand limits the applicability of these methods in real-world scenarios, particularly in the absence of dedicated computing devices such as GPUs and TPUs. To address these challenges, we propose Pan-LUT, a novel learnable look-up table (LUT) framework for pan-sharpening that strikes a balance between performance and computational efficiency for large remote sensing images. Our method makes it possible to process 15K$\times$15K remote sensing images on a 24GB GPU. To finely control the spectral transformation, we devise the PAN-guided look-up table (PGLUT) for channel-wise spectral mapping. To effectively capture fine-grained spatial details, we introduce the spatial details look-up table (SDLUT). Furthermore, to adaptively aggregate channel information for generating high-resolution multispectral images, we design an adaptive output look-up table (AOLUT). Our model contains fewer than 700K parameters and processes a 9K$\times$9K image in under 1 ms using one RTX 2080 Ti GPU, demonstrating significantly faster performance compared to other methods. Experiments reveal that Pan-LUT efficiently processes large remote sensing images in a lightweight manner, bridging the gap to real-world applications. Furthermore, our model surpasses SOTA methods in full-resolution scenes under real-world conditions, highlighting its effectiveness and efficiency. We also extend our method to general image fusion tasks. Zhongnan Cai, Yingying Wang 0005, Hui Zheng 0003, Panwang Pan, Zixu Lin, Ge Meng, Chenxin Li, Chunming He, Jiaxin Xie, Yunlong Lin, Junbin Lu, Yue Huang 0001, Xinghao Ding |
NeurIPS | 13 |
| 2025 | JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching AgentabstractPhoto retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based solutions provide automation but often suffer from limited adjustability and poor generalization, failing to meet diverse and personalized editing needs. To bridge this gap, we introduce JarvisArt, a multi-modal large language model (MLLM)-driven agent that understands user intent, mimics the reasoning process of professional artists, and intelligently coordinates over 200 retouching tools within Lightroom. JarvisArt undergoes a two-stage training process: an initial Chain-of-Thought supervised fine-tuning to establish basic reasoning and tool-use skills, followed by Group Relative Policy Optimization for Retouching (GRPO-R) to further enhance its decision-making and tool proficiency. We also propose the Agent-to-Lightroom Protocol to facilitate seamless integration with Lightroom. To evaluate performance, we develop MMArt-Bench, a novel benchmark constructed from real-world user edits. JarvisArt demonstrates user-friendly interaction, superior generalization, and fine-grained control over both global and local adjustments, paving a new avenue for intelligent photo retouching. Notably, it outperforms GPT-4o with a 60\% improvement in average pixel-level metrics on MMArt-Bench for content fidelity, while maintaining comparable instruction-following capabilities. Yunlong Lin, Zixu Lin, Kunjie Lin, Jinbin Bai, Panwang Pan, Chenxin Li, Haoyu Chen 0003, Zhongdao Wang, Xinghao Ding, Wenbo Li 0002, Shuicheng Yan |
NeurIPS | 9 |
| 2025 | FRN: Fractal-Based Recursive Spectral Reconstruction NetworkabstractGenerating hyperspectral images (HSIs) from RGB images through spectral reconstruction can significantly reduce the cost of HSI acquisition. In this paper, we propose a Fractal-Based Recursive Spectral Reconstruction Network (FRN), which differs from existing paradigms that attempt to directly integrate the full-spectrum information from the R, G, and B channels in a one-shot manner. Instead, it treats spectral reconstruction as a progressive process, predicting from broad to narrow bands or employing a coarse-to-fine approach for predicting the next wavelength. Inspired by fractals in mathematics, FRN establishes a novel spectral reconstruction paradigm by recursively invoking an atomic reconstruction module. In each invocation, only the spectral information from neighboring bands is used to provide clues for the generation of the image at the next wavelength, which follows the low-rank property of spectral data. Moreover, we design a band-aware state space model that employs a pixel-differentiated scanning strategy at different stages of the generation process, further suppressing interference from low-correlation regions caused by reflectance differences. Through extensive experimentation across different datasets, FRN achieves superior reconstruction performance compared to state-of-the-art methods. Code is available at https://github.com/mongko007/frn. Ge Meng, Zhongnan Cai, Ruizhe Chen, Jingyan Tu, Yingying Wang 0005, Yue Huang 0001, Xinghao Ding |
NeurIPS | 7 |
| 2025 | DynamicVerse: A Physically-Aware Multimodal Framework for 4D World ModelingabstractUnderstanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act within real environments with human‑like capabilities. However, existing datasets are often derived from limited simulators or utilize traditional Structure-from-Motion for up-to-scale annotation and offer limited descriptive captioning, which restricts the capacity of foundation models to accurately interpret real-world dynamics from monocular videos, commonly sourced from the internet. To bridge these gaps, we introduce **DynamicVerse**, a physical‑scale, multimodal 4D world modeling framework for dynamic real-world video. We employ large vision, geometric, and multimodal models to interpret metric-scale static geometry, real-world dynamic motion, instance-level masks, and holistic descriptive captions. By integrating window-based Bundle Adjustment with global optimization, our method converts long real-world video sequences into a comprehensive 4D multimodal format. DynamicVerse delivers a large-scale dataset consists of 100K+ videos with 800K+ annotated masks and 10M+ frames from internet videos. Experimental evaluations on three benchmark tasks, namely video depth estimation, camera pose estimation, and camera intrinsics estimation, demonstrate that our 4D modeling achieves superior performance in capturing physical-scale measurements with greater global accuracy than existing methods. Kairun Wen, Yuzhi Huang, Runyu Chen, Hui Zheng 0003, Yunlong Lin, Panwang Pan, Chenxin Li, Wenyan Cong, Junbin Lu, Chenguo Lin, Dilin Wang, Zhicheng Yan 0001, Hongyu Xu, Justin Theiss, Yue Huang 0001, Xinghao Ding, Zhiwen Fan |
NeurIPS | 17 |
| 2025 | SOMA: A semantic-guided Order-aware Mamba Architecture for multivariate time series forecasting
Jinkai Zhang, Yingying Wang 0005, Shengbin Ma, Xinghao Ding, Xiaotong Tu |
Adv. Eng. Informatics | 4 |
| 2025 | Spatial-frequency dual-domain Kolmogorov-Arnold networks for multimodal medical image fusion
Lewu Lin, Jiaxin Xie, Yingying Wang 0005, Jialing Huang, Rongjin Zhuang, Xiaotong Tu, Xinghao Ding, Na Shen |
Neurocomputing | 7 |
| 2025 | Generalizable Prompts Guided by Image-Redundant Separation for Vehicle ReidentificationabstractVehicle reidentification (reID) is a critical computer vision task with applications in video surveillance and autonomous vehicles. While significant progress has been made in recent years, domain generalization (DG) in reID remains a challenging and valuable research direction. Learning discriminative features that capture the intrinsic characteristics of vehicles, rather than domain-specific details, is paramount in addressing the domain shift problem, which encompasses disparities in data distribution, feature distribution, and label distribution. Recently, contrastive language image pretraining (CLIP) has attracted widespread attention because of its capacity to generalize knowledge across different domains or contexts. When fine-tuned for DG tasks, it can leverage this broad knowledge to perform well in domains or on tasks it has not specifically seen during training. The foremost work in this context is CLIP-reID, showcasing outstanding experimental performance on vehicle datasets through the integration of learnable prompts. However, the process of acquiring learnable prompts inevitably incorporates noisy text descriptions, such as background and camera style information, resulting in its limitations in DG tasks. To address this distinctive issue, we propose a CLIP-based Image-Redundant Separation (CIRS) framework to remove redundant domain-specific information and then implement visual-text alignment of CLIP. Specifically, we employ a classic variational autoencoder for image reconstruction, which can encourage the images generated by the vector quantized-variational autoencoder (VQ-VAE) network to contain features unrelated to vehicle IDs. Under the precise guidance of the image-redundant separation framework, a set of generalizable and learnable prompts for each vehicle can be effectively generated for reID. Extensive experimental results indicate that our method has achieved remarkable performance on several public datasets. Zhenyu Kuang, Lidong Cheng, Yue Huang 0001, Xinghao Ding |
IEEE Internet Things J. | 5 |
| 2025 | PORSCHE: Progressive Optimization and Robust Spatial Convolution for Hybrid Enhancement in Visible-Infrared Vehicle Re-IdentificationabstractVisible-infrared vehicle re-identification has become crucial for stable 24-h surveillance of Visual Internet of Things (VIoT). It aims to leverage the shared information between different modalities to retrieve specific objects. Previous works primarily focus on projecting images from two modalities into a common embedding space to measure their similarity scores. However, the inherent distribution discrepancies between different modalities often lead models to rely on spurious features that are unrelated to vehicle identity, making effective feature fusion challenging. To address this unique problem, we propose the Progressive Optimization and Robust Spatial Convolution for Hybrid Enhancement (PORSCHE) model, which reduces the negative effects of spurious correlations and biases toward training pairs. Specifically, we introduce the Patch-wise Matching (PAM) module, which performs initial coarse-grained alignment between different modalities. Building upon this foundation, we develop the Point-wise Matching (POM) module to achieve fine-grained discriminative alignment through precise point-level feature matching, thereby enhancing identity-specific representation. To optimize these complementary PAM and POM components effectively, we implement a progressive training strategy that hierarchically refines feature representations from local patches to global structures, ensuring stable learning of modality-invariant characteristics. This coarse-to-fine architecture enables gradual fusion and alignment across modalities at both patch and point levels, effectively capturing the essential discriminative features required for robust cross-modality retrieval. Extensive experimental results on MSVR310, WMVeID863, and RGBN300 benchmarks demonstrate the effectiveness of our proposed method. The code will be available at https://github.com/HowardLiu28/PORSCHE. Yinhao Liu, Zhenyu Kuang, Yige Ma, Xinghao Ding, Yue Huang 0001, Congbo Cai, Xiaosong Li 0004 |
IEEE Internet Things J. | 5 |
| 2025 | Open world out-of-distribution generalization via dream open and sustain close
Kunze Huang, Luyao Tang, Jieyuan Yang, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
Knowl. Based Syst. | 7 |
| 2025 | CVC: Further aligning LLMs via cross-view correction for time series forecasting
Haote Xu, Haodi Xu, Yinhao Liu, Xinghao Ding |
Knowl. Based Syst. | 5 |
| 2025 | PRADA: Prompt-guided Representation Alignment and Dynamic Adaption for time series forecasting
Yinhao Liu, Zhenyu Kuang, Xinghao Ding |
Knowl. Based Syst. | 6 |
| 2025 | Random Response-Based SAR Purification DefenseabstractDeep learning-driven synthetic aperture radar automatic target recognition (SAR ATR) has gained increasing attention recently. However, current methods remain highly vulnerable to adversarial attacks, limiting their practical application. Most adversarial defense methods rely on adversarial training or attack detection, which tend to overfit and result in poor robustness against different attack types. Moreover, these methods show limited resilience to query-based attacks, where attackers iteratively probe the model to identify vulnerabilities and generate precise adversarial samples. To overcome these challenges, we propose a random response-based SAR purification defense (RRPD) framework composed of two key components: a diffusion purification module (DPM) and a multiexpert randomized response module (MRRM). The DPM removes adversarial noise through diffusion denoising while preserving critical information for accurate recognition. The MRRM introduces multiple expert models to increase response randomness, thus reducing the effectiveness of query-based attacks. Experimental results show that the proposed framework significantly enhances robustness against adversarial attacks on public SAR datasets, improving system security and reliability. The code is available athttps://github.com/SmartDSP2024/RRPD. Zixu Lin, Jiewei Zheng, Jingchao Guo, Huangxing Lin, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Lightweight SAR Ship Detection via Pearson Correlation and Nonlocal DistillationabstractAiming to the challenge of efficient synthetic aperture radar (SAR) ship detection, knowledge distillation recently gained increasing attention as an effective model lightweight approach. SAR ship detection faces challenges including small target detection and complex background clutter. Most existing knowledge distillation methods impose overly strict constraints on the student model, leading to insufficient extraction of detailed features for small target detection. In addition, current mainstream convolutional neural networks (CNNs) primarily focus on extracting local features, which are often inadequate for effectively distinguishing targets from background in complex environments. To address these issues, this proposed work proposes the Pearson correlation distillation and nonlocal distillation (PND) algorithm for SAR ship detection. The Pearson correlation coefficient (PCC) is utilized to model features, relaxing the constraints on the magnitude of the student model’s features. The nonlocal module captures long-range dependencies, enhancing adaptability to complex backgrounds. Experimental results on the SSDD and AIR-SARShip-1.0 datasets demonstrate that our method effectively improves the detection performance of the student model, while also facilitating its transfer to other detectors. Weimin Cai, Jingchao Guo, Hangyang Kong, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Vision-Language Model Priors-Driven State Space Model for Infrared-Visible Image FusionabstractInfrared and visible image fusion (IVIF) aims to effectively integrate complementary information from both infrared and visible modalities, enabling a more comprehensive understanding of the scene and improving downstream semantic tasks. Recent advancements in Mamba have shown remarkable performance in image fusion, owing to its linear complexity and global receptive fields. However, leveraging Vision-Language Model (VLM) priors to drive Mamba for modality-specific feature extraction and using them as constraints to enhance fusion results has not been fully explored. To address this gap, we introduce VLMPD-Mamba, a Vision-Language Model Priors-Driven Mamba framework for IVIF. Initially, we employ the VLM to adaptively generate modality-specific textual descriptions, which enhance image quality and highlight critical target information. Next, we present Text-Controlled Mamba (TCM), which integrates textual priors from the VLM to facilitate effective modality-specific feature extraction. Furthermore, we design the Cross-modality Fusion Mamba (CFM) to fuse features from different modalities, utilizing VLM priors as constraints to enhance fusion outcomes while preserving salient targets with rich details. In addition, to promote effective cross modality feature interactions, we introduce a novel bi-modal interaction scanning strategy within the CFM. Extensive experiments on various datasets for IVIF, as well as downstream visual tasks, demonstrate the superiority of our approach over state-of-the-art (SOTA) image fusion algorithms. Rongjin Zhuang, Yingying Wang 0005, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
IEEE Signal Process. Lett. | 5 |
| 2025 | Fusion2Void: Unsupervised Multi-Focus Image Fusion Based on Image InpaintingabstractMulti-focus image fusion aims to integrate clear segments from different partially focused images, creating an ‘all-in-focus’ composite. Due to the lack of ground-truth for multi-focus image fusion, supervised deep learning methods are deemed inappropriate for this task. In this paper, we present an unsupervised approach for multi-focus image fusion, named Fusion2Void. Fusion2Void ingeniously tackles the challenge of missing ground-truth by framing image inpainting as an auxiliary task. Specifically, Fusion2Void utilizes a fusion network to merge focused regions from multiple source images. Following the fusion process, image patches in the source images are randomly dropped to construct an additional image inpainting task. Subsequently, an image inpainting network uses the fused image as a guide to restore the missing content in the source images. The missing content in the source images includes both focused and defocused regions. Restoring focused image patches is significantly more challenging than restoring their defocused counterparts due to their inclusion of more high-frequency details. If the focused image patches are effectively restored, the repair of the defocused image patches becomes notably easier. Therefore, the image inpainting network implicitly compels the fused image to incorporate all focused content from the source images, as these can be utilized to restore the missing focused regions in the source images perfectly. Based on image inpainting, the fusion network generates ‘all-in-focus’ images in an unsupervised manner. Experiments on several synthetic and real-world datasets highlight Fusion2Void’s state-of-the-art performance relative to other methods. Huangxing Lin, Yunlong Lin, Jingyuan Xia, Linyu Fan, Yingying Wang 0005, Xinghao Ding |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Domain Adaptive Oriented Object Detection From Optical to SAR ImagesabstractOriented object detection in synthetic aperture radar (SAR) images presents significant challenges due to the scarcity of labeled data. In contrast, acquiring labeled optical remote sensing images is considerably easier. This article proposes a domain adaptive oriented object detection (DAOOD) model, termed the pixel-instance information transfer-based model (PITM). PITM aims to transfer knowledge from optical to SAR domains, thereby reducing the dependency of oriented SAR object detection on labels. Given the pronounced domain disparity between optical and SAR images, the efficient migration of both visual content and rotating instances is incorporated to bridge the gap in their information distribution simultaneously. Specifically, regarding pixel-level information transfer, speckle noise from SAR images is mixed into the optical domain to form an intermediate domain, thus compensating for the visual difference between the two domains. For instance-level information transfer, considering the angle diversity of rotating objects, multiscale and multidirectional spatial information extraction is combined with decoupled instance-invariant features, enhancing the cross-domain discernment capacity of rotating instances. Experimental results on four DAOOD benchmarks (i.e., two optical datasets to two SAR datasets) demonstrate that the proposed PITM significantly improves oriented object detection performance, even in the absence of labeled SAR images. Specifically, it individually outperforms two source-only models by 88.16% and 54.62% in average precision (AP). Hailiang Huang 0002, Jingchao Guo, Huangxing Lin, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Learning Diffusion High-Quality Priors for Pan-Sharpening: A Two-Stage Approach With Time-Aware Adapter Fine-TuningabstractPan-sharpening aims to enhance the spatial resolution of the low-resolution multispectral (LRMS) image by incorporating high-frequency details from the panchromatic (PAN) image, while maintaining the spectral qualities of the LRMS image. Recent advancements in diffusion models have shown remarkable capabilities in image restoration and generation. However, simply applying diffusion models in pan-sharpening yields suboptimal outcomes in terms of fine-grained details and spectral fidelity. To this end, we introduce TA-DiffHQP, a two-stage approach that integrates the diffusion high-quality priors model (DiffHQP) and the time-aware adapter (TA-Adapter). Initially, we perform self-reconstruction pretraining DiffHQP with a fixed sampling strategy on approximately 24K high-resolution remote sensing datasets to explicitly model the high-quality texture details and spectral fidelity, after which we freeze most of DiffHQP’s parameters. In stage two, we integrate time-aware fusion adapters with the DiffHQP, enabling rapid adaptation to the pan-sharpening task. The TA-Adapters prioritize low-frequency main scenes during the early phases of the denoising process and refine high-frequency details in the later phases, achieving cross-modal information fusion from coarse to fine. Extensive experiments conducted on three satellite datasets demonstrate that our approach attains state-of-the-art (SOTA) performance over existing methods, revealing superior fusion outcomes in pan-sharpening. Yingying Wang 0005, Yunlong Lin, Xuanhua He, Hui Zheng 0003, Linyu Fan, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Toward Generalizable Pansharpening: Conditional Flow-Based Learning Guided by Implicit High-Frequency PriorsabstractThe goal of pansharpening is to restore the missing high-frequency details in the low-resolution multispectral (LRMS) image to generate its high-resolution multispectral (HRMS) counterpart by exploiting the high-resolution panchromatic (PAN) image as guidance. Previous research has predominantly focused on improving pansharpening performance for single satellites, often neglecting the challenge of generalization. Moreover, pansharpening is inherently an ill-posed problem. Precise and generalizable prior guidance is crucial for effectively addressing this issue. To this end, we propose conditional flow-based learning guided by implicit high-frequency priors (CFLIHPs) toward generalizable pansharpening. Specifically, we utilize implicit neural representation (INR) to precisely align implicit high-frequency texture priors from LRMS and PAN images within Fourier and gradient domains. The flow-based restoration module then leverages these priors as the guiding condition to restore domain-irrelevant high-frequency details, thereby facilitating effective cross-satellite generalization. Furthermore, to tackle the complex degradation process in real-world scenarios, we introduce noise perturbation to the high-frequency learning part, enhancing generalizability across diverse spatial resolutions and improving the robustness of our framework. Extensive experiments conducted on multiple satellite datasets demonstrate that our proposed framework outperforms state-of-the-art (SOTA) methods, achieving superior performance and excellent generalization results in both cross-satellite scenarios and full-resolution scenes. Yingying Wang 0005, Hui Zheng 0003, Yunlong Lin, Linyu Fan, Xuanhua He, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Toward Better Generalization Using Synthetic Data: A Domain Adaptation Framework for T2 Mapping via Multiple Overlapping-Echo AcquisitionabstractThe generation of synthetic data using physics-based modeling provides a solution to limited or lacking real-world training samples in deep learning methods for rapid quantitative magnetic resonance imaging (qMRI). However, synthetic data distribution differs from real-world data, especially under complex imaging conditions, resulting in gaps between domains and limited generalization performance in real scenarios. Recently, a single-shot qMRI method, multiple overlapping-echo detachment imaging (MOLED), was proposed, quantifying tissue transverse relaxation time ( $\text {T}_{{2}}$ ) in the order of milliseconds with the help of a trained network. Previous works leveraged a Bloch-based simulator to generate synthetic data for network training, which leaves the domain gap between synthetic and real-world scenarios and results in limited generalization. In this study, we proposed a $\text {T}_{{2}}$ mapping method via MOLED from the perspective of domain adaptation, which obtained accurate mapping performance without real-label training and reduced the cost of sequence research at the same time. Experiments demonstrate that our method outshined in the restoration of MR anatomical structures. Qizhi Yang, Linyu Fan, Shaocong Yu, Liyan Sun, Congbo Cai, Xinghao Ding |
IEEE Trans. Medical Imaging | 7 |
| 2025 | SRCD: Semantic Reasoning With Compound Domains for Single-Domain Generalized Object DetectionabstractThis article provides a novel framework for single-domain generalized object detection (i.e., Single-DGOD), where we are interested in learning and maintaining the semantic structures of self-augmented compound cross-domain samples to enhance the model's generalization ability. Different from domain generalized object detection (DGOD) trained on multiple source domains, Single-DGOD is far more challenging to generalize well to multiple target domains with only one single source domain. Existing methods mostly adopt a similar treatment from DGOD to learn domain-invariant features by decoupling or compressing the semantic space. However, there may exist two potential limitations: 1) pseudo attribute-label correlation due to extremely scarce single-domain data and 2) the semantic structural information is usually ignored, i.e., we found the affinities of instance-level semantic relations in samples are crucial to model generalization. In this article, we introduce semantic reasoning with compound domains (SRCD) for Single-DGOD. Specifically, our SRCD contains two main components, namely, the texture-based self-augmentation (TBSA) module and the local-global semantic reasoning (LGSR) module. TBSA aims to eliminate the effects of irrelevant attributes associated with labels, such as light, shadow, and color, at the image level by a light-yet-efficient self-augmentation. Moreover, LGSR is used to further model the semantic relationships on instance features to uncover and maintain the intrinsic semantic structures. Extensive experiments on multiple benchmarks demonstrate the effectiveness of the proposed SRCD. Code is available at github.com/zjrao/SRCD. Zhijie Rao, Jingcai Guo, Luyao Tang, Yue Huang 0001, Xinghao Ding, Song Guo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Unsupervised Pan-Sharpening via Mutually Guided Detail RestorationabstractPan-sharpening is a task that aims to super-resolve the low-resolution multispectral (LRMS) image with the guidance of a corresponding high-resolution panchromatic (PAN) image. The key challenge in pan-sharpening is to accurately modeling the relationship between the MS and PAN images. While supervised deep learning methods are commonly employed to address this task, the unavailability of ground-truth severely limits their effectiveness. In this paper, we propose a mutually guided detail restoration method for unsupervised pan-sharpening. Specifically, we treat pan-sharpening as a blind image deblurring task, in which the blur kernel can be estimated by a CNN. Constrained by the blur kernel, the pan-sharpened image retains spectral information consistent with the LRMS image. Once the pan-sharpened image is obtained, the PAN image is blurred using a pre-defined blur operator. The pan-sharpened image, in turn, is used to guide the detail restoration of the blurred PAN image. By leveraging the mutual guidance between MS and PAN images, the pan-sharpening network can implicitly learn the spatial relationship between the two modalities. Extensive experiments show that the proposed method significantly outperforms existing unsupervised pan-sharpening methods. Huangxing Lin, Xinghao Ding, Tianpeng Liu, Yongxiang Liu |
AAAI | 3 |
| 2024 | Progressive High-Frequency Reconstruction for Pan-Sharpening with Implicit Neural RepresentationabstractPan-sharpening aims to leverage the high-frequency signal of the panchromatic (PAN) image to enhance the resolution of its corresponding multi-spectral (MS) image. However, deep neural networks (DNNs) tend to prioritize learning the low-frequency components during the training process, which limits the restoration of high-frequency edge details in MS images. To overcome this limitation, we treat pan-sharpening as a coarse-to-fine high-frequency restoration problem and propose a novel method for achieving high-quality restoration of edge information in MS images. Specifically, to effectively obtain fine-grained multi-scale contextual features, we design a Band-limited Multi-scale High-frequency Generator (BMHG) that generates high-frequency signals from the PAN image within different bandwidths. During training, higher-frequency signals are progressively injected into the MS image, and corresponding residual blocks are introduced into the network simultaneously. This design enables gradients to flow from later to earlier blocks smoothly, encouraging intermediate blocks to concentrate on missing details. Furthermore, to address the issue of pixel position misalignment arising from multi-scale features fusion, we propose a Spatial-spectral Implicit Image Function (SIIF) that employs implicit neural representation to effectively represent and fuse spatial and spectral features in the continuous domain. Extensive experiments on different datasets demonstrate that our method outperforms existing approaches in terms of quantitative and visual measurements for high-frequency detail recovery. Ge Meng, Jingjia Huang, Yingying Wang 0005, Zhenqi Fu, Xinghao Ding, Yue Huang 0001 |
AAAI | 5 |
| 2024 | Mixstyle-Entropy: Whole Process Domain Generalization with Causal Intervention and Perturbation
Luyao Tang, Chaoqi Chen, Xinghao Ding, Yue Huang 0001 |
BMVC | 4 |
| 2024 | Implicit Foreground-Guided Network for Anomaly Detection and LocalizationabstractAnomaly detection plays an essential role in large-scale industrial manufacturing. However, reconstruction-based anomaly detection methods, as one of the mainstream methods, are prone to incorrectly detecting background noise as anomalous regions. Therefore, inspired by multi-task learning, we propose an Implicit Foreground-guided Network (IFgNet), which consists of a Multi-Task Attention Shared (MTAS) sub-network and a discriminative sub-network. Specifically, the MTAS sub-network implements the foreground detection and reconstruction tasks within the shared network, while the discriminative sub-network performs the final anomaly detection. In the MTAS sub-network, multiple task-specific attention blocks are applied to learn task-specific features while allowing features to be shared between different tasks. Consequently, the features that contain both semantic and edge structure information are learned through the foreground detection task, which also facilitates the reconstruction task. Furthermore, the outputs of foreground detection can be utilized to refine the anomaly detection results. In this way, IFgNet effectively mitigates the influence of background noise and achieves competitive performance on the VisA and BTAD datasets with existing methods. Xiaolu Chen, Haote Xu, Chenghao Deng, Xiaotong Tu, Xinghao Ding, Yue Huang 0001 |
ICASSP | 5 |
| 2024 | Dataset Distillation with Channel Efficient ProcessabstractThe success of deep learning is primarily attributed to the vast amount of data used for training, which comes with massive computation costs and storage. Dataset distillation(DD) aims to reduce the dependency on such massive data by learning a small synthetic dataset that preserves most information from the original dataset. Recent work has proposed a new condensation framework that generates multiple synthetic data with a limited storage budget. However, they only focus on the synthetic data’s spatial regularity and ignore the compressible space on the channel. In this paper, we propose a novel channel-efficient process that augments the number of condensed data and trains the synthetic data in a channel information-intensive mode. We design the process as a plugand-play strategy that is portable to any existing DD baseline, and our experiment results demonstrate that it can yield significant improvement on downstream classification tasks compared with previous DD methods. Guoqing Zheng, Xinghao Ding |
ICASSP | 3 |
| 2024 | High-Resolution Wideband DOA Estimation Based on Multi-Frequency Cyclic Rank-MinimizationabstractWideband DOA estimation has been applied in various signal source location scenarios, e.g., in wireless communication systems to improve the capacity of communication. Existing wideband DOA methods often require prior knowledge such as the number of sources as well as pre-estimations. Moreover, they may suffer from model-mismatch problem. In this paper, we employ the manifold separation technique and Jacobi-Anger expansion to allow multi-frequency joint processing of wideband DOA, which alleviates the challenge of model-mismatch and leads to a much higher DOA resolution. The proposed method is further formulated to be a multi-convex rank-minimization problem to facilitate the analysis of the problem and to improve the convergence performance. The superior performance of the proposed multi-frequency joint processing method has been demonstrated by several numerical studies. Hedeng Yu, Zhenlong Xiao, Xinghao Ding, Xianbin Wang 0001 |
ICC | 3 |
| 2024 | S3AHI: Source-Free Domain Adaptive Small Object Detection with Slicing Aided Hyper InferenceabstractThe time-consuming and laborious annotation of small objects has resulted in a relative scarcity of datasets specifically designed for small objects. Additionally, variations in data acquisition devices and application scenarios often cause a domain shift between source-trained data and target data. Unsupervised Domain Adaptation (UDA) is extensively applied to alleviate the domain shift between two domains based on the assumption that source data is accessible during the adaptation process. However, source data may be unavailable in some scenarios due to data privacy or data transmission issues. In this paper, we propose a Source-free domain adaptive framework for Small object detection with Slicing Aided Hyper Inference (S3AHI) in the test-time training stage. Without access to source data, Source-Free Domain Adaptation (SFDA) only employs unlabeled target data to adapt a source-trained model to the target domain during the test phase. SAHI provides a generic and effective solution to detect small objects and can be seamlessly integrated into nearly any pipeline. Motivated by contrastive learning, we build instance-level correlation graphs with the semantic features of proposals and learn high-quality pseudolabels. The S3AHI distills target domain knowledge to the sourcetrained model under the mean-teacher framework. Extensive experiments reveal that our approach outperforms existing SFDA and UDA methods significantly. Haizhou Ding, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
IJCNN | 5 |
| 2024 | An Adaptive Spatio-Temporal Graph Structure Learning Model for Lithium-Ion Battery Pack State of Health EstimationabstractCurrent research on lithium-ion battery state of health (SOH) estimation predominantly focuses on a single battery, not on an entire battery pack, which makes these methods inadequate for describing the SOH of energy systems that work in real-life situations. Furthermore, the current SOH estimation methods using graph neural networks (GNNs) inherently adopt suboptimal graph construction approaches, which makes them fail to accurately extract the most pertinent spatial dependencies, resulting in a reduction in prediction accuracy. In this work, we propose a pioneering GNN-based model to predict battery pack SOH. Specifically, an optimal graph structure, learned by the introduced optimal graph extractor, is used to capture the feature dependencies that are most appropriate for the downstream SOH prediction task. Additionally, a Graph Attention network (GAT) and a Gated Recurrent Unit (GRU) are employed to learn the spatial and temporal features, respectively. Lastly, we introduce an effective and simple spatio-temporal feature fusion module to ensure that the extracted features are fully utilized, and the fused features are then used for prediction. Experimental results on the NASA and CALCE datasets validate the superiority of the proposed approach over existing state-of-the-art methods for battery SOH estimation. Canxing Lai, Xiaotong Tu, Andreas Jakobsson, Xinghao Ding, Yue Huang 0001 |
IJCNN | 4 |
| 2024 | Class Incremental Aerial Scene Recognition Under Long-Tailed DistributionabstractDeep learning excels in aerial scene recognition (ASR) but struggles with learning from sequential data due to catastrophic forgetting. Class incremental learning (CIL) can address this but often assumes a balanced data distribution. Real-world aerial scenes exhibit long-tailed properties, the undesirable bias toward the head classes as well as overfitting for the tail classes aggravate the challenge of class incremental ASR. So here, we introduce a two-stage framework for class incremental ASR under long-tailed distribution. First, the encoder is trained by conventional CIL method. Next, keep the encoder fixed and then distribution transfer via inter-class similarity is used to generate a sufficient number of features for tail classes, thus a more balanced classifier can be trained. Our framework can be easily integrated into existing CIL methods. Experiments on public datasets demonstrated the superior performance of our proposed framework. Haizhou Ding, Xiaotong Tu, Lexing Huang, Yue Huang 0001, Xinghao Ding |
IJCNN | 6 |
| 2024 | SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly DetectionabstractRecently, large pre-trained vision-language models, such as CLIP, have demonstrated significant potential in zero-/few-shot anomaly detection tasks. However, existing methods not only rely on expert knowledge to manually craft extensive text prompts but also suffer from a misalignment of high-level language features with fine-level vision features in anomaly segmentation tasks. In this paper, we propose a method, named SimCLIP, which focuses on refining the aforementioned misalignment problem through bidirectional adaptation of both Multi-Hierarchy Vision Adapter (MHVA) and Implicit Prompt Tuning (IPT). In this way, our approach requires only a simple binary prompt to efficiently accomplish anomaly classification and segmentation tasks in zero-shot scenarios. Furthermore, we introduce its few-shot extension, SimCLIP+, integrating the relational information among vision embeddings and skillfully merging the cross-modal synergy information between vision and language to address downstream anomaly detection tasks. Extensive experiments on two challenging datasets prove the more remarkable generalization capacity of our method compared to the current SOTA approaches. Our code is available at https://github.com/CH-ORGI/SimCLIP. Chenghao Deng, Haote Xu, Xiaolu Chen, Haodi Xu, Xiaotong Tu, Xinghao Ding, Yue Huang 0001 |
ACM Multimedia | 6 |
| 2024 | P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical ImagesabstractGenerating diverse plausible outputs from a single input is crucial for addressing visual ambiguities, exemplified in medical imaging where experts may provide varying semantic segmentation annotations for the same image.Existing methods handles ambiguous segmentation relying on probabilistic modeling and extensive multi-output annotated data while often struggles with limited ambiguously labeled datasets common in real-world applications.To surmount the challenge, we propose P²SAM, a novel framework that leverages the Segment Anything Model (SAM)'s prior knowledge for ambiguous object segmentation. By transforming SAM's sensitivity to prompts into an advantage, we introduce a prior probabilistic space for prompts.Experimental results show that P²SAM significantly enhances medical segmentation precision and diversity using minimal ambiguously annotated samples. Benchmarking against state-of-the-art methods demonstrates superior performance with just 5.5% of the training data (+12% Dmax). This approach marks a significant advancement towards deploying probabilistic models in data-limited real-world scenarios. Yuzhi Huang, Chenxin Li, Zixu Lin, Hengyu Liu 0007, Haote Xu, Yifan Liu 0010, Yue Huang 0001, Xinghao Ding, Xiaotong Tu, Yixuan Yuan |
ACM Multimedia | 8 |
| 2024 | Efficient Perceiving Local Details via Adaptive Spatial-Frequency Information Integration for Multi-focus Image FusionabstractMulti-focus image fusion (MFIF) aims to combine multiple images with different focused regions into a single all-in-focus image. Existing unsupervised deep learning-based methods only fuse structural information of images in the spatial domain, neglecting potential solutions from the frequency domain exploration. In this paper, we make the first attempt to integrate spatial-frequency information to achieve high-quality MFIF. We propose a novel unsupervised spatial-frequency interaction MFIF network named SFIMFN, which consists of three key components: Adaptive Frequency Domain Information Interaction Module (AFIM), Ret-Attention-Based Spatial Information Extraction Module (RASEM), and Invertible Dual-domain Feature Fusion Module (IDFM). Specifically, in AFIM, we interactively explore global contextual information by combining the amplitude and phase information of multiple images separately. In RASEM, we design a customized transformer to encourage the network to capture important local high-frequency information by redesigning the self-attention mechanism with a bidirectional, two-dimensional form of explicit decay. Finally, we employ IDFM to fuse spatial-frequency information without information loss to generate the desired all-in-focus image. Extensive experiments on different datasets demonstrate that our method significantly outperforms state-of-the-art unsupervised methods in terms of qualitative and quantitative metrics as well as the generalization ability. Jingjia Huang, Jingyan Tu, Ge Meng, Yingying Wang 0005, Xiaotong Tu, Xinghao Ding, Yue Huang 0001 |
ACM Multimedia | 7 |
| 2024 | Source-free cross-domain fault diagnosis of rotating machinery using the Siamese framework
Chenyu Ma, Xiaotong Tu, Guanxing Zhou, Yue Huang 0001, Xinghao Ding |
Knowl. Based Syst. | 5 |
| 2024 | AIVR-Net: Attribute-based invariant visual representation learning for vehicle re-identification
Zhenyu Kuang, Lidong Cheng, Yinhao Liu, Xinghao Ding, Yue Huang 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Learning to sound imaging by a model-based interpretable network
Xiaotong Tu, Saqlain Abbas, Hao Liang 0011, Yue Huang 0001, Xinghao Ding |
Signal Process. | 6 |
| 2024 | Efficient Training Acceleration via Sample-Wise Dynamic Probabilistic PruningabstractData pruning is observed to substantially reduce the computation and memory costs of model training. Previous studies have primarily focused on constructing a series of coresets with representative samples by leveraging predefined rules for evaluating sample importance. Learning dynamics and selection bias, however, are rarely being considered. In this letter, a novel Sample-wise Dynamic Probabilistic Pruning (SwDPP) method is proposed for efficient training. Specifically, instead of hard-pruning the samples that are considered easy or well-learned, we formulate the pruning process as a probabilistic sampling problem. This is achieved by a carefully-designed soft-selection mechanism, which constantly expresses learning dynamics and relaxes selection bias. Moreover, to alleviate the accuracy drop under high pruning rates, we introduce a probabilistic Mixup strategy for information diversity maintenance. Extensive experiments conducted on CIFAR-10, CIFAR-100 and Tiny-ImageNet show that, the proposed SwDPP outperforms current state-of-the-art methods across various pruning settings. Notably, on CIFAR-10 and CIFAR-100, SwDPP achieves lossless training acceleration using only 70% of the data per epoch. Feicheng Huang, Yue Huang 0001, Xinghao Ding |
IEEE Signal Process. Lett. | 4 |
| 2024 | Diffusion-Based Continuous Feature Representation for Infrared Small-Dim Target DetectionabstractInfrared small-dim target detection plays a pivotal role in missions involving rescue, surveillance, and early warning systems. Despite remarkable strides made by existing methods, certain limitations still hinder the detection accuracy, including deficiency in high-resolution representation, inadequacy in addressing dim targets, and difficulty in tackling low-contrast targets against complex backgrounds. To overcome these limitations, we propose a diffusion-based continuous feature representation network (DCFR-Net), comprising two crucial branches: diffusion-based continuous high-resolution feature representation (DCHFR) and infrared small-dim target detection (ISDTD). Specifically, to precisely capture extremely small target contours, DCHFR integrates implicit neural representation (INR) into a conditional denoising diffusion model, super-resolving infrared targets in a self-supervised strategy. ISDTD leverages the shared encoder from DCHFR to construct high-resolution feature representation, which is fed into multi-scale implicit feature alignment (MIFA) and spatial-frequency feature interaction (SFFI). To alleviate the impact of dim and vulnerable targets, MIFA delicately aggregates different-layer features in a resolution-free manner. Furthermore, to enhance the contrast between infrared targets and intricate backgrounds, SFFI achieves profound spatial-frequency feature interaction and global-local receptive field mixture. Extensive experiments conducted on three challenging datasets of NUAA-SIRST, IRSTD-1k and NUDT-SIRST reveal that our DCFR-Net outperforms the state-of-the-art (SOTA) methods, demonstrating the superiority and robustness of our approach in infrared small-dim target detection. Code will be available at https://github.com/flyannie/DCFR-Net. Linyu Fan, Yingying Wang 0005, Guoliang Hu, Hui Zheng 0003, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2024 | Cross-Modality Interaction Network for Pan-SharpeningabstractPan-sharpening seeks to generate a high-resolution multispectral (HRMS) image by merging the high-resolution panchromatic (PAN) image and its low-resolution multispectral (LRMS) counterpart. The main challenge lies in enhancing modality-aware features and efficiently integrating complementary information between PAN and MS pairs. To achieve desired fusion results, it is crucial to fully utilize both intramodality characteristics and intermodality relationships. Current research often overlooks the exploration of cross-modality relationships and neglects the enhancement of modality-aware features in pan-sharpening. In this work, we introduce an innovative pan-sharpening framework, named cross-modality interaction network (CMINet), which comprises three core designs: a modality-aware feature enhancement (MAFE) module to enhance the feature representation of both modalities, a cross-modality attention (CMA) module that effectively extracts the intramodality features and fully leverages the intermodality complementary information, and a modality alignment (MA) module to address modality-aware misalignment issue during fusion. Extensive experiments are conducted to verify the effectiveness of our proposed network and showcase its superior performance in comparison to other state-of-the-art approaches. Yingying Wang 0005, Xuanhua He, Yunlong Lin, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | AFSC: Adaptive Fourier Space Compression for Anomaly DetectionabstractThe primary challenge faced by reconstruction-based anomaly detection (AD) methods is that neural networks exhibit strong generalization, resulting in a high probability and accuracy of anomaly reconstruction. Several existing methods attempt to alleviate this problem by randomly masking partial image regions and reconstructing the image from partial inpaintings. However, local masking in spatial space is not guaranteed to remove anomalous regions during the testing phase and poses the risk of normal regions being inaccurately reconstructed. Hence, we explore an approach to compress the global information of the image while ensuring the loss of partial anomaly information renders it difficult to reconstruct. Inspired by the fact that each Fourier coefficient contains global information of the image, we propose an adaptive Fourier space compression (AFSC) method. Specifically, the Fourier coefficients of the input image are sparsely sampled by binary masks obtained from the AFSC module (AFSCm). In AFSCm, the masks are jointly optimized with the reconstruction network subject to sparsity constraint. The learned masks are forced to selectively retain part of the global information that is favourable to recovering normal images. In addition, we introduce an efficient Fourier convolution module that enables the network to accurately reconstruct normal regions under conditions of losing partial information. Experimental results on three benchmarks of industrial scenarios demonstrate our method (without external prior) achieves competitive results compared with recent methods. Haote Xu, Xiaolu Chen, Changxing Jing, Liyan Sun, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Ind. Informatics | 7 |
| 2023 | Self-Supervised Image Denoising Using Implicit Deep Denoiser PriorabstractWe devise a new regularization for denoising with self-supervised learning. The regularization uses a deep image prior learned by the network, rather than a traditional predefined prior. Specifically, we treat the output of the network as a ``prior'' that we again denoise after ``re-noising.'' The network is updated to minimize the discrepancy between the twice-denoised image and its prior. We demonstrate that this regularization enables the network to learn to denoise even if it has not seen any clean images. The effectiveness of our method is based on the fact that CNNs naturally tend to capture low-level image statistics. Since our method utilizes the image prior implicitly captured by the deep denoising CNN to guide denoising, we refer to this training strategy as an Implicit Deep Denoiser Prior (IDDP). IDDP can be seen as a mixture of learning-based methods and traditional model-based denoising methods, in which regularization is adaptively formulated using the output of the network. We apply IDDP to various denoising tasks using only observed corrupted data and show that it achieves better denoising results than other self-supervised denoising methods. Huangxing Lin, Yihong Zhuang, Xinghao Ding, Delu Zeng, Yue Huang 0001, Xiaotong Tu, John W. Paisley |
AAAI | 3 |
| 2023 | Learning a Simple Low-Light Image Enhancer from Paired Low-Light InstancesabstractLow-light Image Enhancement (LIE) aims at improving contrast and restoring details for images captured in lowlight conditions. Most of the previous LIE algorithms adjust illumination using a single input image with several handcrafted priors. Those solutions, however, often fail in revealing image details due to the limited information in a single image and the poor adaptability of handcrafted priors. To this end, we propose PairLIE, an unsupervised approach that learns adaptive priors from low-light image pairs. First, the network is expected to generate the same clean images as the two inputs share the same image content. To achieve this, we impose the network with the Retinex theory and make the two reflectance components consistent. Second, to assist the Retinex decomposition, we propose to remove inappropriate features in the raw image with a simple self-supervised mechanism. Extensive experiments on public datasets show that the proposed PairLIE achieves comparable performance against the state-of-the-art approaches with a simpler network and fewer handcrafted priors. Code is available at: https://github.com/zhenqifu/PairLIE. Zhenqi Fu, Xiaotong Tu, Yue Huang 0001, Xinghao Ding, Kai-Kuang Ma |
CVPR | 5 |
| 2023 | Topology Design for Robust IoT Data Gathering via Bayesian NetworksabstractInternet of Things (IoT) systems have become the critical platform to enable a wide variety of smart applications. During IoT data gathering over wireless network, data may be missing due to the constraints of sensors as well as the reliability of communications. From a graph signal processing perspective, recovery of missing data may be strongly affected by the IoT system topology, which can be characterized by a directed adjacency matrix. To guarantee a robust data gathering, we propose a novel method in this paper to design the optimal topology for IoT networks via Bayesian networks, where the designed directed adjacency matrix is with orthogonal graph frequency components. Moreover, the gathering of IoT data becomes sparser in the graph frequency domain using the designed adjacency matrix and may hence improve the recovery performance of missing data. Experimental results show that our proposed methods outperform several existing algorithms. Haiyan Wei, Zhenlong Xiao, Xinghao Ding, Xianbin Wang 0001 |
GLOBECOM | 3 |
| 2023 | Hint-Dynamic Knowledge DistillationabstractKnowledge Distillation (KD) transfers the knowledge from a high-capacity teacher model to promote a smaller student model. Existing efforts guide the distillation by matching their prediction logits, feature embedding, etc., while leaving how to efficiently utilize them in junction less explored. In this paper, we propose Hint-dynamic Knowledge Distillation, dubbed HKD, which excavates the knowledge from the teacher’s hints in a dynamic scheme. The guidance effect from the knowledge hints usually varies in different instances and learning stages, which motivates us to customize a specific hint-learning manner for each instance adaptively. Specifically, a meta-weight network is introduced to generate the instance-wise weight coefficients about knowledge hints in the perception of the dynamical learning progress of the student model. We further present a weight ensembling strategy to eliminate the potential bias of coefficient estimation by exploiting the historical statics. Experiments on standard benchmarks of CIFAR-100 and Tiny-ImageNet manifest that the proposed HKD well boost the effect of knowledge distillation tasks. Chenxin Li, Xiaotong Tu, Xinghao Ding, Yue Huang 0001 |
ICASSP | 4 |
| 2023 | Underwater Image Enhancement and Super-Resolution Using Implicit Neural NetworksabstractUnderwater images are often notably degraded by light scattering and absorption. To improve image quality and object details, we present a novel unsupervised underwater image enhancement and super-resolution method using implicit neural networks. Concretely, taking low-resolution coordinates as the inputs, we first leverage Fourier feature mapping to encode the coordinates. Then, three implicit neural networks are applied to estimate each component (i.e., the global background light, the transmission map, and the scene radiance) of the underwater formation model. Those components are further used to reconstruct the raw underwater image in a self-supervised fashion. In the inference stage, high-resolution coordinates are employed to predict a high-quality and high- resolution underwater image. Extensive experiments show that our method achieves a favorable performance in terms of both super-resolution and quality enhancement as compared with current approaches. Xueye Chu, Zhenqi Fu, Shaocong Yu, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
ICIP | 6 |
| 2023 | Radar HRRP Unseen Class Recognition Based on the Joint Dictionary LearningabstractExisting task settings and methods for radar high resolution range profile (HRRP) recognition are limited in addressing open challenges. To avoid labor-intensive data collection and model retraining, we formulate a new task called HRRP unseen class recognition, where the testing classes are unknown during training. To perform this task, we utilize metric learning to explore the potential information of unseen categories. Due to the target-aspect sensitivity problem of HRRP, feature extraction is a key step for recognition. Therefore, we take into account the time, frequency and aspect-angle characteristics of targets. Then a joint dictionary learning method is proposed to align different modalities in the latent common space to capture the intrinsic invariant representations of the unseen class targets. A variety of experiments demonstrate the effectiveness of our method. Chuchu He, Zhenyu Kuang, Yijin Zhong, Xinghao Ding, Yue Huang 0001 |
ICIP | 4 |
| 2023 | Joint Under-Sampling Pattern Optimization and Content-Based Reconstruction Network for Fast MRI ReconstructionabstractMagnetic Resonance Imaging (MRI) is frequently used by physicians for diagnosing human tissue, which requires focusing on specific content regions. However, existing compressed-sensing MRI (CS-MRI) reconstruction methods do not optimize under-sampling adaptively based on content or utilize k-space sensing resources effectively. To address these issues, we propose CBRecNet, a model that combines under-sampling pattern optimization with content-based reconstruction network. We extract multi-scale features of the content and calibrate the features of the reconstruction network using a pixel attention mechanism, which can guide the optimizable pattern to get a more suitable sampling pattern. To validate the proposed CBRecNet, we conduct extensive experiments on the CHAOS dataset with a four-stage learning strategy and the result demonstrates a very favorable outcome in CS-MRI. Our model provides a new approach to calibrated feature schemes based on meaningful medical content and demonstrates promising applications in CS-MRI. Shaocong Yu, Xueye Chu, Zhenglin Zhou, Jiefeng Guo, Xinghao Ding |
ICIP | 5 |
| 2023 | Domain Generalization via Implicit Domain Augmentation
Zhijie Rao, Chaoqi Chen, Yue Huang 0001, Xinghao Ding |
ICONIP (8) | 5 |
| 2023 | Domain Generalized Object Detection with Triple Graph Reasoning Network
Zhijie Rao, Luyao Tang, Yue Huang 0001, Xinghao Ding |
ICONIP (3) | 4 |
| 2023 | Domain-irrelevant Feature Learning for Generalizable Pan-sharpeningabstractPan-sharpening aims to spatially enhance the low-resolution multispectral image (LRMS) by transferring high-frequency details from a panchromatic image (PAN) while preserving the spectral characteristics of LRMS. Previous arts mainly focus on how to learn a high-resolution multispectral image (HRMS) on the i.i.d. assumption. However, the distribution of training and testing data often encounters significant shifts in different satellites. To this end, this paper proposes a generalizable pan-sharpening network via domain-irrelevant feature learning. On the one hand, a structural preservation module (STP) is designed to fuse high-frequency information of PAN and LRMS. Our STP is performed on the gradient domain because it consists of structure and texture details that can generalize well on different satellites. On the other hand, to avoid spectral distortion while promoting the generalization ability, a spectral preservation module (SPP) is developed. The key design of SPP is to learn a phase fusion network of PAN and LRMS. The amplitude of LRMS, which contains 'satellite style' information is directly injected in different fusion stages. Extensive experiments have demonstrated the effectiveness of our method against state-of-the-art methods in both single-satellite and cross-satellite scenarios. Code is available at: https://github.com/LYL1015/DIRFL. Yunlong Lin, Zhenqi Fu, Ge Meng, Yingying Wang 0005, Linyu Fan, Hedeng Yu, Xinghao Ding |
ACM Multimedia | 8 |
| 2023 | Learning High-frequency Feature Enhancement and Alignment for Pan-sharpeningabstractPan-sharpening aims to utilize the high-resolution panchromatic (PAN) image as a guidance to super-resolve the spatial resolution of the low-resolution multispectral (MS) image. The key challenge in pan-sharpening is how to effectively and precisely inject high-frequency edges and textures from the PAN image into the low-resolution MS image. To address this issue, we propose a High-frequency Feature Enhancement and Alignment Network (HFEAN) for effectively encouraging the high-frequency learning. To implement it, three core designs are customized: a Fourier convolution based efficient feature enhancement module (FEM), an implicit neural alignment module (INA), and a preliminary alignment module (Pre-align). To be specific, FEM employs the fast Fourier convolution with attention mechanism to achieve the mixed global-local receptive field on each scale of the high-frequency domain, thus yielding the informative latent codes. INA leverages implicit neural function to precisely align the latent codes from different scales in the continuous domain. In this way, the high frequency signals at different scales are represented as functions of continuous coordinates, enabling a precise feature alignment in a resolution-free manner. Pre-align is developed to further address the inherent misalignment between PAN and MS pairs. Extensive experiments over multiple satellite datasets validate the effectiveness of the proposed network and demonstrate its favorable performance against the existing state-of-the-art methods both visually and quantitatively. Code is available at: https://github.com/Gracewangyy/HFEAN. Yingying Wang 0005, Yunlong Lin, Ge Meng, Zhenqi Fu, Linyu Fan, Hedeng Yu, Xinghao Ding, Yue Huang 0001 |
ACM Multimedia | 8 |
| 2023 | Boosting Generalization Performance in Person Re-identification
Lidong Cheng, Zhenyu Kuang, Xinghao Ding, Yue Huang 0001 |
PRCV (10) | 4 |
| 2023 | High-Resolution Feature Representation Driven Infrared Small-Dim Object Detection
Yingying Wang 0005, Linyu Fan, Xinghao Ding, Yue Huang 0001 |
PRCV (12) | 4 |
| 2023 | DP-INNet: Dual-Path Implicit Neural Network for Spatial and Spectral Features Fusion in Pan-Sharpening
Jingjia Huang, Ge Meng, Yingying Wang 0005, Yunlong Lin, Yue Huang 0001, Xinghao Ding |
PRCV (8) | 6 |
| 2023 | Adversarial Robustness via Multi-experts Framework for SAR Recognition with Class Imbalanced
Chuyang Lin, Senlin Cai, Hailiang Huang 0002, Xinghao Ding, Yue Huang 0001 |
PRCV (4) | 4 |
| 2023 | A Two-Stage Federated Learning Framework for Class Imbalance in Aerial Scene Classification
Zhengpeng Lv, Yihong Zhuang, Yue Huang 0001, Xinghao Ding |
PRCV (4) | 5 |
| 2023 | Recognizer Embedding Diffusion Generation for Few-Shot SAR Recognization
Chuyang Lin, Yijin Zhong, Yue Huang 0001, Xinghao Ding |
PRCV (4) | 5 |
| 2023 | Infrared and Visible Image Fusion via Test-Time Training
Guoqing Zheng, Zhenqi Fu, Xiaopeng Lin, Xueye Chu, Yue Huang 0001, Xinghao Ding |
PRCV (10) | 6 |
| 2023 | Graph-Based Dependency-Aware Non-Intrusive Load Monitoring
Guoqing Zheng, Yuming Hu, Zhenlong Xiao, Xinghao Ding |
PRCV (10) | 4 |
| 2023 | AnoCSR-A Convolutional Sparse Reconstructive Noise-Robust Framework for Industrial Anomaly Detection
Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
PRCV (5) | 4 |
| 2023 | Distributed representation learning with skip-gram model for trained random forests
Chao Ma 0006, Le Zhang 0001, Zhiguang Cao, Yue Huang 0001, Xinghao Ding |
Neurocomputing | 6 |
| 2023 | Interclass Similarity Transfer for Imbalanced Aerial Scene ClassificationabstractImbalanced class distributions widely exist in real-world aerial images, which brings a significant challenge to aerial scene classification due to the undesirable bias toward the majority classes as well as overfitting for the minority classes. Although the similarity between different scene classes may be inconsistent, they can be measured by the mean of feature statistics. This motivates us to transfer the statistics of the majority class to the minority class having similar feature statistics. Specifically, based on the observation that the feature statistics of each class may follow the Gaussian distribution, the similarity across different classes would thus be described by the mean of feature statistics. The distributions of minority classes would afterward be calibrated by statistical transfer via interclass similarity (STAIRS), and a sufficient number of features could hence be generated for the minority class to improve its performance in classifier learning. We demonstrate the effectiveness of the proposed method for imbalanced aerial scene classification on the imbalanced aerial image dataset (AID) and NWPU-RESISC45 datasets. The proposed method outperforms alternatives by a large margin in both overall performance and minority classification performance of imbalanced aerial scenes. Changxing Jing, Lexing Huang, Senlin Cai, Yihong Zhuang, Zhenlong Xiao, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2023 | A Contrastive-Based Adversarial Training Algorithm for HRRP Target RecognitionabstractIn recent years, deep learning methods have significantly improved the recognition performance of high-resolution range profiles (HRRP). However, the vulnerability of the deep network to attacks poses a serious threat to the security of radar target recognition systems. In this paper, an adversarial training algorithm based on contrastive learning is proposed that introduces the N-pair loss function and balances the feature space to smooth the decision boundary and improve the robustness. The experimental analysis demonstrates that the proposed method achieves better defense performance than the traditional adversarial training algorithms. The work presented in this paper provides an important step towards improving the security and reliability of deep learning-based radar recognition systems. Liangchao Shi, Chuyang Lin, Senlin Cai, Wei Lin 0021, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2023 | Adaptive Ship Detection From Optical to SAR ImagesabstractRecent advances in Synthetic Aperture Radar (SAR) ship detection have witnessed remarkable success by using large-scale annotated datasets. However, the annotation of SAR images requires strong domain-specific expertise, significantly hindering the prompt adoption of modern object detectors in this regime. Compared to SAR data, the optical data in geoscience are considerably easier to label. Motivated by this, we investigate a new and challenging problem – adaptive ship detection – with the goal of enhancing ship detection performance on SAR images by leveraging knowledge transferred from optical images. Considering the large distributional discrepancy between the source (optical) and target (SAR) domains, we presentOmniAdapt, a novel framework that progressively narrows the distance between the two types of images at the pixel, feature, and classifier levels. Specifically, OmniAdapt consists of three main modules, Target-like Generation Module (TLGM), Multi-feature Alignment Module (MFAM), and Common Specific Decomposition Module (CSDM). TLGM minimizes the visual disparity by infusing the target domain style into the source domain. MFAM aligns local- and global-level feature representations in an adversarial manner. Finally, CSDM decomposes the classifier into two independent components,i.e., the domain-common component and the domain-specific component, and promotes the recognition ability of the former via regularization learning. Experimental results demonstrate the effectiveness of the proposed method. Zhijie Rao, Chuyang Lin, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Contrastive Learning for Radar HRRP Recognition With Missing AspectsabstractHigh resolution range profile (HRRP) has attracted increasing attention in radar automatic target recognition (RATR). However, the target-aspect missing problem in non-cooperative targets recognition, which is one of the most challenging tasks in RATR, has received very few contributions recently. The proposed work is motivated by a very simple observation, i.e., as compared with HRRP signals of interesting targets, sufficient unlabeled HRRP signals are much easier to acquire. However, these signals are often neglected since there is no label information. This work focuses on the target-aspect missing problem by using a dual self-supervised contrastive learning framework. The proposed model takes advantage of massive unlabeled HRRP signals to enhance the generalization ability. Specifically, we employ self-supervised contrastive learning and online clustering module for extracting target-aspect invariant representations. Additionally, a data augmentation strategy is used as a simple yet effective module. The experimental results demonstrate that the proposed method with fewer labels enhances the recognition performance as compared to supervised and self-supervised methods in missing aspect problem. Yijin Zhong, Wei Lin 0021, Lexing Huang, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2023 | Exploring personalization via federated representation Learning on non-IID data
Changxing Jing, Yan Huang 0032, Yihong Zhuang, Liyan Sun, Zhenlong Xiao, Yue Huang 0001, Xinghao Ding |
Neural Networks | 7 |
| 2023 | Relation Matters: Foreground-Aware Graph-Based Relational Reasoning for Domain Adaptive Object DetectionabstractDomain Adaptive Object Detection (DAOD) focuses on improving the generalization ability of object detectors via knowledge transfer. Recent advances in DAOD strive to change the emphasis of the adaptation process from global to local in virtue of fine-grained feature alignment methods. However, both the global and local alignment approaches fail to capture the topological relations among different foreground objects as the explicit dependencies and interactions between and within domains are neglected. In this case, only seeking one-vs-one alignment does not necessarily ensure the precise knowledge transfer. Moreover, conventional alignment-based approaches may be vulnerable to catastrophic overfitting regarding those less transferable regions (e.g., backgrounds) due to the accumulation of inaccurate localization results in the target domain. To remedy these issues, we first formulate DAOD as an open-set domain adaptation problem, in which the foregrounds and backgrounds are seen as the "known classes" and "unknown class" respectively. Accordingly, we propose a new and general framework for DAOD, named Foreground-aware Graph-based Relational Reasoning (FGRR), which incorporates graph structures into the detection pipeline to explicitly model the intra- and inter-domain foreground object relations on both pixel and semantic spaces, thereby endowing the DAOD model with the capability of relational reasoning beyond the popular alignment-based paradigm. FGRR first identifies the foreground pixels and regions by searching reliable correspondence and cross-domain similarity regularization respectively. The inter-domain visual and semantic correlations are hierarchically modeled via bipartite graph structures, and the intra-domain relations are encoded via graph attention mechanisms. Through message-passing, each node aggregates semantic and contextual information from the same and opposite domain to substantially enhance its expressive power. Empirical results demonstrate that the proposed FGRR exceeds the state-of-the-art performance on four DAOD benchmarks. Chaoqi Chen, Jiongcheng Li, Xiaoguang Han 0001, Yue Huang 0001, Xinghao Ding, Yizhou Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Adaptive nonlinear group delay mode estimationabstractThe decomposition of non-stationary signals remains a challenge in a wide variety of fields. Especially, the impulse or cross-mode signals are difficult to be reconstructed by recent methods due to their transient characteristic. Moreover, most methods rely heavily on the user-defined settings of the regularized parameter for the convex optimization algorithm. In this work, an adaptive nonlinear group delay mode estimation (ANGDME) algorithm is proposed by exploiting the sparsity of signals formulated as the nonlinear group delay model. The ANGDME introduces a complex Bayesian compressive sensing (CBCS) framework to process the sparse reconstruction. Then, a hierarchical Laplace scale mixture (LSM) prior is utilized to model dependencies among coefficients and provides superior probabilistic predictions. Furthermore, the estimator of amplitudes and group delays (GDs) of signals are obtained from the posterior distribution by Bayesian inference instead of point estimation. Finally, the proposed method updates the dictionary matrix in a traditional data-driven manner, resulting in a high-resolution time-frequency representation. Both simulated and experimental results illustrate the adaptability and effectiveness of the proposed method. Yijin Mao, Xiaotong Tu, Saqlain Abbas, Hao Liang 0011, Yue Huang 0001, Xinghao Ding |
Signal Process. | 6 |
| 2023 | Enhanced features in image manipulation detection
Chuchu He, Yunshu Chen, Yue Huang 0001, Xiaotong Tu, Xinghao Ding |
Signal Process. Image Commun. | 6 |
| 2023 | Unsupervised Video-Based Action Recognition With Imagining Motion and Perceiving AppearanceabstractVideo-based action recognition is a challenging task, which demands carefully considering the temporal property of videos in addition to the appearance attributes. Particularly, the temporal domain of raw videos usually contains significantly more redundant or irrelevant information than still images. For that, this paper proposes an unsupervised video-based action recognition approach with imagining motion and perceiving appearance, called IMPA, by comprehensively learning the spatio-temporal characteristics inherited in videos, with a particular emphasis on the moving object for action recognition. Specifically, a self-supervised Motion Extracting Block (MEB) is designed to extract the principal motion features by focusing on the large movement of the moving object, based on the observation that humans can infer complete motion trajectories from partial moving objects. To further take the indispensable appearance attribute in videos into account, an unsupervised Appearance Learning Block (ALB) is developed to perceive the static appearance, thus in combination with the MEB to recognize actions. Extensive validation experiments and ablation studies on multiple datasets demonstrate that our proposed IMPA approach obtains superior performance and surpasses other classical and state-of-the-art unsupervised action recognition methods. Wei Lin 0021, Yihong Zhuang, Xinghao Ding, Xiaotong Tu, Yue Huang 0001, Huanqiang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Unpaired Speckle Extraction for SAR DespecklingabstractSpeckle suppression is a critical step in synthetic aperture radar (SAR) imaging. Since speckle-free SAR images are inaccessible, supervised denoising methods are not suitable for this task. To exploit the strong capabilities of convolutional neural networks (CNNs), we propose Unpaired Speckle Extraction (SAR-USE), an unsupervised method for SAR despeckling. Our method utilizes unpaired SAR and clean optical images to extract “real” speckle for learning despeckling. First, a CNN that has never seen clean SAR images is employed to extract speckle from the SAR image. Then, the extracted speckle is multiplied with a random optical image to synthesize paired data for learning speckle removal. Through a Siamese network, speckle extraction and learning despeckling are performed alternately and promote each other. To make the extracted speckle more visually and statistically realistic, it is constrained by a noise correction module to be unit mean while maintaining spatial correlation. After convergence, the CNN is a good denoiser that can effectively extract speckle from SAR images. Experiments on synthetic datasets show that the denoising ability of the proposed method is as good as its supervised counterpart. More importantly, SAR-USE is very efficient for removing the spatially correlated speckle in real data that supervised learning methods cannot. Huangxing Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Self-Supervised Video-Based Action Recognition With DisturbancesabstractSelf-supervised video-based action recognition is a challenging task, which needs to extract the principal information characterizing the action from content-diversified videos over large unlabeled datasets. However, most existing methods choose to exploit the natural spatio-temporal properties of video to obtain effective action representations from a visual perspective, while ignoring the exploration of the semantic that is closer to human cognition. For that, a self-supervised Video-based Action Recognition method with Disturbances called VARD, which extracts the principal information of the action in terms of the visual and semantic, is proposed. Specifically, according to cognitive neuroscience research, the recognition ability of humans is activated by visual and semantic attributes. An intuitive impression is that minor changes of the actor or scene in video do not affect one person's recognition of the action. On the other hand, different humans always make consistent opinions when they recognize the same action video. In other words, for an action video, the necessary information that remains constant despite the disturbances in the visual video or the semantic encoding process is sufficient to represent the action. Therefore, to learn such information, we construct a positive clip/embedding for each action video. Compared to the original video clip/embedding, the positive clip/embedding is disturbed visually/semantically by Video Disturbance and Embedding Disturbance. Our objective is to pull the positive closer to the original clip/embedding in the latent space. In this way, the network is driven to focus on the principal information of the action while the impact of sophisticated details and inconsequential variations is weakened. It is worthwhile to mention that the proposed VARD does not require optical flow, negative samples, and pretext tasks. Extensive experiments conducted on the UCF101 and HMDB51 datasets demonstrate that the proposed VARD effectively improves the strong baseline and outperforms multiple classical and advanced self-supervised action recognition methods. Wei Lin 0021, Xinghao Ding, Yue Huang 0001, Huanqiang Zeng |
IEEE Trans. Image Process. | 2 |
| 2023 | A Multiscale Approach to Deep Blind Image Quality AssessmentabstractFaithful measurement of perceptual quality is of significant importance to various multimedia applications. By fully utilizing reference images, full-reference image quality assessment (FR-IQA) methods usually achieve better prediction performance. On the other hand, no-reference image quality assessment (NR-IQA), also known as blind image quality assessment (BIQA), which does not consider the reference image, makes it a challenging but important task. Previous NR-IQA methods have focused on spatial measures at the expense of information in the available frequency bands. In this paper, we present a multiscale deep blind image quality assessment method (BIQA, M.D.) with spatial optimal-scale filtering analysis. Motivated by the multi-channel behavior of the human visual system and contrast sensitivity function, we decompose an image into a number of spatial frequency bands through multiscale filtering and extract features to map an image to its subjective quality score by applying convolutional neural network. Experimental results show that BIQA, M.D. compares well with existing NR-IQA methods and generalizes well across datasets. Manni Liu, Jiabin Huang 0003, Delu Zeng, Xinghao Ding, John W. Paisley |
IEEE Trans. Image Process. | 4 |
| 2023 | Joint Image and Feature Levels Disentanglement for Generalizable Vehicle Re-identificationabstractDomain generalization (DG), which doesn’t require any data from target domains during training, is more challenging but practical than unsupervised domain adaptation (UDA). Since different vehicles of the same type have a similar appearance, neural networks always rely on a small amount of useful information to distinguish them, meaning that is more significant to remove ID-unrelated information for vehicle re-Identification (re-ID). Therefore, it is the key to eliminating the interference of a large amount of redundant information for the generalizable vehicle re-ID method. To address this unique challenge, we propose a novel disentanglement learning method that encourages variational autoencoder (VAE) network to reduce ID-unrelated features of vehicles by minimizing image reconstruction errors and providing sufficient representation to vehicle labels. To capture the intrinsic characteristics associated with the DG task, our core idea is to build the identity information streaming framework to separate ID-related and ID-unrelated information at the image and feature levels. In contrast with the general decoupling methods, our method leverages the decoupling of joint image and feature levels to extract more generalizable features. Furthermore, we present a brand-new vehicle dataset of truck types named “Optimus Prime (Opri)”, which includes multiple images of each truck captured by cameras at different high-speed toll gates. Experimental results on public datasets demonstrate that our method can achieve promising results and outperform several state-of-the-art approaches. Our codes and models are available at JIFD. Zhenyu Kuang, Chuchu He, Yue Huang 0001, Xinghao Ding, Huafeng Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Learning Rate DropoutabstractOptimization algorithms are of great importance to efficiently and effectively train a deep neural network. However, the existing optimization algorithms show unsatisfactory convergence behavior, either slowly converging or not seeking to avoid bad local optima. Learning rate dropout (LRD) is a new gradient descent technique to motivate faster convergence and better generalization. LRD aids the optimizer to actively explore in the parameter space by randomly dropping some learning rates (to 0); at each iteration, only parameters whose learning rate is not 0 are updated. Since LRD reduces the number of parameters to be updated for each iteration, the convergence becomes easier. For parameters that are not updated, their gradients are accumulated (e.g., momentum) by the optimizer for the next update. Accumulating multiple gradients at fixed parameter positions gives the optimizer more energy to escape from the saddle point and bad local optima. Experiments show that LRD is surprisingly effective in accelerating training while preventing overfitting. Huangxing Lin, Weihong Zeng, Yihong Zhuang, Xinghao Ding, Yue Huang 0001, John W. Paisley |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Enhanced Deep Blind Hyperspectral Image FusionabstractThe goal of hyperspectral image fusion (HIF) is to reconstruct high spatial resolution hyperspectral images (HR-HSI) via fusing low spatial resolution hyperspectral images (LR-HSI) and high spatial resolution multispectral images (HR-MSI) without loss of spatial and spectral information. Most existing HIF methods are designed based on the assumption that the observation models are known, which is unrealistic in many scenarios. To address this blind HIF problem, we propose a deep learning-based method that optimizes the observation model and fusion processes iteratively and alternatively during the reconstruction to enforce bidirectional data consistency, which leads to better spatial and spectral accuracy. However, general deep neural network inherently suffers from information loss, preventing us to achieve this bidirectional data consistency. To settle this problem, we enhance the blind HIF algorithm by making part of the deep neural network invertible via applying a slightly modified spectral normalization to the weights of the network. Furthermore, in order to reduce spatial distortion and feature redundancy, we introduce a Content-Aware ReAssembly of FEatures module and an SE-ResBlock model to our network. The former module helps to boost the fusion performance, while the latter make our model more compact. Experiments demonstrate that our model performs favorably against compared methods in terms of both nonblind HIF fusion and semiblind HIF fusion. Xueyang Fu, Weihong Zeng, Liyan Sun, Ronghui Zhan, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | A Real-Time Global Inference Network for One-Stage Referring Expression ComprehensionabstractReferring expression comprehension (REC) is an emerging research topic in computer vision, which refers to the detection of a target region in an image given a test description. Most existing REC methods follow a multistage pipeline, which is computationally expensive and greatly limits the applications of REC. In this article, we propose a one-stage model toward real-time REC, termed real-time global inference network (RealGIN). RealGIN addresses the issues of expression diversity and complexity of REC with two innovative designs: adaptive feature selection (AFS) and Global Attentive ReAsoNing (GARAN). Expression diversity concerns varying expression content, which includes information such as colors, attributes, locations, and fine-grained categories. To address this issue, AFS adaptively fuses features of different semantic levels to tackle the changes in expression content. In contrast, expression complexity concerns the complex relational conditions in expressions that are used to identify the referent. To this end, GARAN uses the textual feature as a pivot to collect expression-aware visual information from all regions and then diffuses this information back to each region, which provides sufficient context for modeling the relational conditions in expressions. On five benchmark datasets, i.e., RefCOCO, RefCOCO+, RefCOCOg, ReferIT, and Flickr30k, the proposed RealGIN outperforms most existing methods and achieves very competitive performances against the most advanced one, i.e., MAttNet. More importantly, under the same hardware, RealGIN can boost the processing speed by 10-20 times over the existing methods. Yiyi Zhou, Rongrong Ji, Gen Luo, Xiaoshuai Sun, Jinsong Su, Xinghao Ding, Chia-Wen Lin, Qi Tian 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Unsupervised Underwater Image Restoration: From a Homology PerspectiveabstractUnderwater images suffer from degradation due to light scattering and absorption. It remains challenging to restore such degraded images using deep neural networks since real-world paired data is scarcely available while synthetic paired data cannot approximate real-world data perfectly. In this paper, we propose an UnSupervised Underwater Image Restoration method (USUIR) by leveraging the homology property between a raw underwater image and a re-degraded image. Specifically, USUIR first estimates three latent components of the raw underwater image, i.e., the global background light, the transmission map, and the scene radiance (the clean image). Then, a re-degraded image is generated by randomly mixing up the estimated scene radiance and the raw underwater image. We demonstrate that imposing a homology constraint between the raw underwater image and the re-degraded image is equivalent to minimizing the restoration error and hence can be used for the unsupervised restoration. Extensive experiments show that USUIR achieves promising performance in both inference time and restoration quality. Zhenqi Fu, Huangxing Lin, Shu Chai, Liyan Sun, Yue Huang 0001, Xinghao Ding |
AAAI | 7 |
| 2022 | EffiSeaNet: Pioneering Lightweight Network for Underwater Salient Object Detection
Qingyao Wu, Zhenqi Fu, Chenyu Ma, Xiaotong Tu, Xinghao Ding |
ACCV (4) | 6 |
| 2022 | Uncertainty Inspired Underwater Image Enhancement
Zhenqi Fu, Yue Huang 0001, Xinghao Ding, Kai-Kuang Ma |
ECCV (18) | 4 |
| 2022 | Knowledge Condensation Distillation
Chenxin Li, Mingbao Lin, Zhiyuan Ding, Nie Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Liujuan Cao |
ECCV (11) | 7 |
| 2022 | Unsupervised and Untrained Underwater Image Restoration Based on Physical Image Formation ModelabstractUnderwater images suffer from degradation caused by light scattering and absorption. Training a deep neural network to restore underwater images is challenging due to the labor-intensive data collection and the lack of paired data. To this end, we propose an unsupervised and untrained underwater image restoration method based on the layer disentanglement and the underwater image formation model. Specifically, our network disentangles an underwater image into four components, i.e., the scene radiance, the direct transmission map, the backscatter transmission map, and the global background light, which are further combined to reconstruct the underwater image in a self-supervised manner. Our method can avoid using paired training data and large-scale datasets, benefiting from the unsupervised and untrained characteristics. Extensive experiments demonstrated that our method obtains promising performance compared with six methods on three real-world underwater image databases. Shu Chai, Zhenqi Fu, Yue Huang 0001, Xiaotong Tu, Xinghao Ding |
ICASSP | 5 |
| 2022 | A Robust Object Segmentation Network for UnderWater ScenesabstractUnderwater object segmentation is one of the key technologies in the fields of marine biology research and autonomous underwater vehicles. The challenges of underwater object segmentation originate from two aspects, 1) the complex underwater environment and 2) the camouflage characteristics of marine animals. In this paper, we propose WaterSNet, an underwater object segmentation network to address these challenges. Specified, we propose a random style adaption (RSA) module as well as a siamese structure to reduce the impact of water degradation diversity. We also extract multi-scale features via the receptive field block (RFB) module, and then fuses multi-level features to better utilize global context information via the attention fusion block (AFB) module. Experimental results on marine animal dataset MAS3K demonstrate that the proposed method outperforms other state-of-the-art methods significantly. The code will be available at: https://github.com/ruizhechen/WaterSNet/ Ruizhe Chen, Zhenqi Fu, Yue Huang 0001, En Cheng, Xinghao Ding |
ICASSP | 5 |
| 2022 | Underwater Image Enhancement Via Learning Water Type Desensitized RepresentationsabstractWe present a novel underwater image enhancement method termed SCNet to improve the image quality meanwhile cope with the degradation diversity caused by the water. SCNet is based on normalization schemes across both spatial and channel dimensions with the key idea of learning water type desensitized features. Specifically, we apply whitening to de-correlate activations across spatial dimensions for each instance in a mini-batch. We also eliminate channel-wise correlation by standardizing and re-injecting the first two moments of the activations across channels. The normalization schemes of spatial and channel dimensions are performed at each scale of the U-Net to obtain multi-scale representations. With such water type irrelevant encodings, the decoder can easily reconstruct the clean signal and be unaffected by the distortion types. Experimental results on two real-world underwater image datasets show that our approach can successfully enhance images with diverse water types, and achieves competitive performance in visual quality improvement. Zhenqi Fu, Xiaopeng Lin, Yue Huang 0001, Xinghao Ding |
ICASSP | 5 |
| 2022 | A Two-Stage Contrastive Learning Framework For Imbalanced Aerial Scene RecognitionabstractIn real-world scenarios, aerial image datasets are generally class imbalanced, where the majority classes have rich samples, while the minority classes only have a few samples. Such class imbalanced datasets bring great challenges to aerial scene recognition. In this paper, we explore a novel two-stage contrastive learning framework, which aims to take care of representation learning and classifier learning, thereby boosting aerial scene recognition. Specifically, in the representation learning stage, we design a data augmentation policy to improve the potential of contrastive learning according to the characteristics of aerial images. And we employ supervised contrastive learning to learn the association between aerial images of the same scene. In the classification learning stage, we fix the encoder to maintain good representation and use the re-balancing strategy to train a less biased classifier. A variety of experimental results on the imbalanced aerial image datasets show the advantages of the proposed two-stage contrastive learning framework for the imbalanced aerial scene recognition. Lexing Huang, Senlin Cai, Yihong Zhuang, Changxing Jing, Yue Huang 0001, Xiaotong Tu, Xinghao Ding |
ICASSP | 7 |
| 2022 | Adaptive Variational Nonlinear Chirp Mode DecompositionabstractVariational nonlinear chirp mode decomposition (VNCMD) is a recently introduced method for nonlinear chirp signal decomposition that has aroused notable attention in various fields. One limiting aspect of the method is that its performance relies heavily on the setting of the bandwidth parameter. To overcome this problem, we here propose a Bayesian implementation of the VNCMD, which can adaptively estimate the instantaneous amplitudes and frequencies of the nonlinear chirp signals, and then learn the active dictionary in a data-driven manner, thereby enabling a high-resolution time-frequency representation. Numerical example of both simulated and measured data illustrate the resulting improvement performance of the proposed method. Hao Liang 0011, Xinghao Ding, Andreas Jakobsson, Xiaotong Tu, Yue Huang 0001 |
ICASSP | 2 |
| 2022 | A Self-Supervised Method for Infrared and Visible Image FusionabstractInfrared and visible image fusion (IVIF) plays important roles in many applications. Since there is no ground-truth, the fusion performance measurement is a difficult but important problem for the task. Previous unsupervised deep learning based fusion methods depend on a hand-crafted loss function to define the distance between the fused image and two types of source images, which still cannot well preserve the vital information in the fused images. To address these issues, we propose an image fusion performance measurement between the fused image and the decomposition of the fused image. A novel self-supervised network for infrared and visible image fusion is designed to preserve the vital information of source images by narrowing the distance between the source images and the decomposed ones. Extensive experimental results demonstrate that our proposed measurement has the ability in improving the performance of backbone network in both subjective and objective evaluations. Xiaopeng Lin, Guanxing Zhou, Weihong Zeng, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
ICIP | 6 |
| 2022 | A Simple Siamese Framework for Vibration Signal RepresentationsabstractSiamese networks are widely used in various contrastive learning methods for recognition tasks, with few labeled data and abundant unlabeled data. In the field of fault diagnosis, it is universal to face the problem that large collections of common fault data and few catastrophic fault samples result in the imbalanced distribution of fault data collection. In this paper, a simple Siamese framework is proposed to learn meaningful signal representations using the differently augmented views of the signals only in the time domain. The industrial fault diagnosis including class balanced and imbalanced motor fault diagnosis is performed to verify the validity of the signal representations. The results demonstrate that the proposed method can significantly balance the representations of both the major and minor classes, which proves the capability of the Siamese framework for class imbalanced classification. Guanxing Zhou, Yihong Zhuang, Xinghao Ding, Yue Huang 0001, Saqlain Abbas, Xiaotong Tu |
ICIP | 3 |
| 2022 | Unsupervised Anomaly Segmentation for Brain Lesions Using Dual Semantic-Manifold Reconstruction
Zhiyuan Ding, Haote Xu, Chenxin Li, Xinghao Ding, Yue Huang 0001 |
ICONIP (3) | 5 |
| 2022 | A Hybrid Framework Based on Classifier Calibration for Imbalanced Aerial Scene Recognition
Yihong Zhuang, Changxing Jing, Senlin Cai, Lexing Huang, Yue Huang 0001, Xiaotong Tu, Xinghao Ding |
ICONIP (3) | 7 |
| 2022 | Recognition-Aware HRRP Generation With Generative Adversarial NetworkabstractExisting works on radar high-resolution range profile (HRRP) recognition commonly focus on utilizing data and various deep learning models in achieving high classification accuracy. However, in practical applications, it is often difficult to obtain HRRP signals, especially for noncooperative targets. Such lack of data dramatically decreases the recognition performance, so this letter applies data augmentation to address small-sample problems. A recognition-aware HRRP generation framework based on a generative adversarial network is proposed for data augmentation, which generates discriminative samples by decomposing and reorganizing signal’s characteristics. The proposed model increases the generated signals’ discriminative power, thus meeting the application requirements. Experiments show that the generated HRRP signals can not only accurately expand the data set but also improve the recognition system’s performance. Besides, the developed model outperforms traditional data augmentation methods and other generative methods. To the best of our knowledge, this is the first work on HRRP signal generation in radar automatic target recognition systems. Yue Huang 0001, Yi Wen 0003, Liangchao Shi, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Self-Supervised SAR Despeckling Powered by Implicit Deep Denoiser PriorabstractSpeckle removal is an important preprocessing step for synthetic aperture radar (SAR) imaging. Since speckle-free SAR images do not exist, supervised methods are not applicable. In this letter, we propose implicit deep denoiser prior (SAR-IDDP), a self-supervised method for SAR despeckling. SAR-IDDP uses a deep image prior (DIP) implicitly captured by the convolutional neural network (CNN) to formulate regularization instead of traditional hand-crafted priors. Specifically, we treat the output of the CNN as a “prior” that we denoise again after “renoising.” The CNN is updated to maximize the similarity between the again denoised image and its prior. The renoising procedure is designed based on the assumption of unit mean noise, while the spatial correlation of speckle is also involved. The despeckling ability of our method stems from CNN’s natural tendency to capture low-level image statistics. Experiments show that SAR-IDDP achieves significant improvements over existing model-based and self-supervised despeckling methods on both synthetic and real SAR images. Huangxing Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | One-Shot HRRP Generation for Radar Target RecognitionabstractInsufficient data of a noncooperative target seriously affect the performance of radar automatic target recognition (RATR) using the high-resolution range profile (HRRP), especially when the noncooperative target has only one sample. To this end, we propose an unsupervised data generation method to generate noncooperative HRRP signals. We utilize the pretrained generative adversarial networks (GANs) model to learn the HRRP general probability distribution. To emphasize the representative and discriminative power of generated HRRP signals, a joint optimization method is proposed to preserve category information. Moreover, a feature diversification method is proposed to make the generated samples have sufficient aspect characteristics to further fit the probability distribution of the noncooperative target. Thus, the generated HRRP signals can effectively improve the recognition performance of noncooperative target. Extensive experiments on HRRP data sets demonstrate the superior performance of our method over other state-of-the-art methods. Liangchao Shi, Zhehan Liang, Yi Wen 0003, Yihong Zhuang, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Deep Multiscale Feedback Network for Hyperspectral Image FusionabstractHyperspectral imaging is useful in many remote sensing tasks. However, it is often challenging to obtain high-resolution images in both the spatial and spectral domains due to hardware limitations. Hyperspectral image fusion (HIF) solves this problem by fusing a low spatial resolution hyperspectral image (LR-HSI) and a high spatial resolution multispectral images (HR-MSI) to obtain a high spatial resolution hyperspectral image (HR-HSI). Many methods have been proposed for HIF, but few approaches have explored the multiscale mutual dependencies between LR-HSI, HR-MSI, and HR-HSI. This kind of mutual dependencies come from the fact that LR-HSI, HR-MSI, and HR-HSI capture the same scene with different spatial or spectral resolutions. To this end, we propose a deep multiscale feedback network (DMFBN) that iteratively learns image fusion and degeneration for HIF. We further equip the network with an error feedback mechanism coupled with multiscale feature learning. Both strategies help better learn the mutual dependencies. Extensive quantitative and qualitative evaluations on two public datasets show that the proposed method performs favorably against the state-of-the-art (SOTA) methods. Weihong Zeng, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Cardiac segmentation on late gadolinium enhancement MRI: A benchmark study from multi-sequence cardiac MR segmentation challenge
Xiahai Zhuang, Jiahang Xu, Xinzhe Luo, Chen Chen 0042, Cheng Ouyang, Daniel Rueckert, Víctor M. Campello, Karim Lekadir, Sulaiman Vesal, Nishant Ravikumar, Yashu Liu 0003, Gongning Luo, Jingkun Chen, Hongwei Li 0004, Buntheng Ly, Maxime Sermesant, Holger Roth, Wentao Zhu 0001, Jiexiang Wang, Xinghao Ding, Sen Yang 0006, Lei Li 0020 |
Medical Image Anal. | 20 |
| 2022 | Hierarchical deep network with uncertainty-aware semi-supervised learning for vessel segmentation
Chenxin Li, Wenao Ma, Liyan Sun, Xinghao Ding, Yue Huang 0001, Guisheng Wang, Yizhou Yu |
Neural Comput. Appl. | 4 |
| 2022 | A teacher-student framework for liver and tumor segmentation under mixed supervision from abdominal CT scans
Liyan Sun, Jianxiong Wu, Xinghao Ding, Yue Huang 0001, Zhong Chen 0005, Guisheng Wang, Yizhou Yu |
Neural Comput. Appl. | 3 |
| 2022 | High-resolution source localization exploiting the sparsity of the beamforming map
Xinghao Ding, Hao Liang 0011, Andreas Jakobsson, Xiaotong Tu, Yue Huang 0001 |
Signal Process. | 1 |
| 2022 | Twice Mixing: A rank learning based quality assessment approach for underwater image enhancement
Zhenqi Fu, Xueyang Fu, Yue Huang 0001, Xinghao Ding |
Signal Process. Image Commun. | 4 |
| 2022 | Dual Domain Multi-Task Model for Vehicle Re-IdentificationabstractVehicle re-identification (re-id) is an essential task in the field of intelligent transportation systems (ITS). The main goal of re-id is to find the same vehicle in different scenarios, which can is still a challenging task in both ITS and computer vision (CV). The existing vehicle re-identification methods simply combine the coarse-grained and the fine-grained attributes together with multi-task training. However, such combination may still have limited performance in vehicles with trivial appearance differences, or with rare models and colors. To solve this problem, we propose a simple yet effective framework, called dual domain multi-task model (DDM), that divides the vehicle images into two domains based on the frequency. And then two parallel branches are proposed to recover the two domains. Furthermore, a multi-task method is proposed, which combines the classification loss in color and model together with triplet loss for fine-grained distance measurement. Besides, a progressive strategy is used in the training process. Two public datasets, PKU VehicleID and VeRi are used to validate the proposed DDM. The experimental results demonstrate that the proposed approach outperforms the existing methods on both datasets. Yue Huang 0001, Borong Liang, Weiping Xie, Yinghao Liao, Zhenyu Kuang, Yihong Zhuang, Xinghao Ding |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | Harmonizing Pathological and Normal Pixels for Pseudo-Healthy SynthesisabstractSynthesizing a subject-specific pathology-free image from a pathological image is valuable for algorithm development and clinical practice. In recent years, several approaches based on the Generative Adversarial Network (GAN) have achieved promising results in pseudo-healthy synthesis. However, the discriminator (i.e., a classifier) in the GAN cannot accurately identify lesions and further hampers from generating admirable pseudo-healthy images. To address this problem, we present a new type of discriminator, the segmentor, to accurately locate the lesions and improve the visual quality of pseudo-healthy images. Then, we apply the generated images into medical image enhancement and utilize the enhanced results to cope with the low contrast problem existing in medical image segmentation. Furthermore, a reliable metric is proposed by utilizing two attributes of label noise to measure the health of synthetic images. Comprehensive experiments on the T2 modality of BraTS demonstrate that the proposed method substantially outperforms the state-of-the-art methods. The method achieves better performance than the existing methods with only 30% of the training data. The effectiveness of the proposed method is also demonstrated on the LiTS and the T1 modality of BraTS. The code and the pre-trained model of this study are publicly available at https://github.com/Au3C2/Generator-Versus-Segmentor. Yihong Zhuang, Liyan Sun, Yue Huang 0001, Xinghao Ding, Guisheng Wang, Lin Yang 0002, Yizhou Yu |
IEEE Trans. Medical Imaging | 6 |
| 2022 | A Model-Driven Deep Unfolding Method for JPEG Artifacts RemovalabstractDeep learning-based methods have achieved notable progress in removing blocking artifacts caused by lossy JPEG compression on images. However, most deep learning-based methods handle this task by designing black-box network architectures to directly learn the relationships between the compressed images and their clean versions. These network architectures are always lack of sufficient interpretability, which limits their further improvements in deblocking performance. To address this issue, in this article, we propose a model-driven deep unfolding method for JPEG artifacts removal, with interpretable network structures. First, we build a maximum posterior (MAP) model for deblocking using convolutional dictionary learning and design an iterative optimization algorithm using proximal operators. Second, we unfold this iterative algorithm into a learnable deep network structure, where each module corresponds to a specific operation of the iterative algorithm. In this way, our network inherits the benefits of both the powerful model ability of data-driven deep learning method and the interpretability of traditional model-driven method. By training the proposed network in an end-to-end manner, all learnable modules can be automatically explored to well characterize the representations of both JPEG artifacts and image content. Experiments on synthetic and real-world datasets show that our method is able to generate competitive or even better deblocking results, compared with state-of-the-art methods both quantitatively and qualitatively. Xueyang Fu, Menglu Wang 0003, Xiangyong Cao, Xinghao Ding, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Rain Streak Removal via Dual Graph Convolutional NetworkabstractDeep convolutional neural networks (CNNs) have become dominant in the single image de-raining area. However, most deep CNNs-based de-raining methods are designed by stacking vanilla convolutional layers, which can only be used to model local relations. Therefore, long-range contextual information is rarely considered for this specific task. To address the above problem, we propose a simple yet effective dual graph convolutional network (GCN) for single image rain removal. Specifically, we design two graphs to perform global relational modeling and reasoning. The first GCN is used to explore global spatial relations among pixels in feature maps, while the second GCN models the global relations across the channels. Compared to standard convolutional operations, the proposed two graphs enable the network to extract representations from new dimensions. To achieve the image rain removal, we further embed these two graphs and multi-scale dilated convolution into a symmetrically skip-connected network architecture. Therefore, our dual graph convolutional network is able to well handle complex and spatially long rain streaks by exploring multiple representations, e.g., multi-scale local feature, global spatial coherence and cross-channel correlation. Meanwhile, our model is easy to implement, end-to-end trainable and computationally efficient. Extensive experiments on synthetic and real data demonstrate that our method achieves significant improvements over the recent state-of-the-art methods. Xueyang Fu, Qi Qi 0005, Zhengjun Zha, Yurui Zhu, Xinghao Ding |
AAAI | 5 |
| 2021 | Unsupervised Large-Scale Social Network Alignment via Cross Network EmbeddingabstractNowadays, it is common for a person to possess different identities on multiple social platforms. Social network alignment aims to match the identities that from different networks. Recently, unsupervised network alignment methods have received significant attention since no identity anchor is required. However, to capture the relevance between identities, the existing unsupervised methods generally rely heavily on user profiles, which is unobtainable and unreliable in real-world scenarios. In this paper, we propose an unsupervised alignment framework named Large-Scale Network Alignment (LSNA) to integrate the network information and reduce the requirement on user profile. The embedding module of LSNA, named Cross Network Embedding Model (CNEM), aims to integrate the topology information and the network correlation to simultaneously guide the embedding process. Moreover, in order to adapt LSNA to large-scale networks, we propose a network disassembling strategy to divide the costly large-scale network alignment problem into multiple executable sub-problems. The proposed method is evaluated over multiple real-world social network datasets, and the results demonstrate that the proposed method outperforms the state-of-the-art methods. Zhehan Liang, Yu Rong 0001, Chenxin Li, Yue Huang 0001, Tingyang Xu, Xinghao Ding, Junzhou Huang |
CIKM | 7 |
| 2021 | I3Net: Implicit Instance-Invariant Network for Adapting One-Stage Object DetectorsabstractRecent works on two-stage cross-domain detection have widely explored the local feature patterns to achieve more accurate adaptation results. These methods heavily rely on the region proposal mechanisms and ROI-based instance-level features to design fine-grained feature alignment modules with respect to the foreground objects. However, for one-stage detectors, it is hard or even impossible to obtain explicit instance-level features in the detection pipelines. Motivated by this, we propose an Implicit Instance-Invariant Network (I3Net), which is tailored for adapting one-stage detectors and implicitly learns instance-invariant features via exploiting the natural characteristics of deep features in different layers. Specifically, we facilitate the adaptation from three aspects: (1) Dynamic and Class-Balanced Reweighting (DCBR) strategy, which considers the coexistence of intra-domain and intra-class variations to assign larger weights to those sample-scarce categories and easy-to-adapt samples; (2) Category-aware Object Pattern Matching (COPM) module, which boosts the cross-domain foreground objects matching guided by the categorical information and suppresses the uninformative background features; (3) Regularized Joint Category Alignment (RJCA) module, which jointly enforces the category alignment at different domain-specific layers with a consistency regularization. Experiments reveal that I3Net exceeds the state-of-the-art performance on benchmark datasets. Chaoqi Chen, Zebiao Zheng, Yue Huang 0001, Xinghao Ding, Yizhou Yu |
CVPR | 4 |
| 2021 | Dual Bipartite Graph Learning: A General Approach for Domain Adaptive Object DetectionabstractDomain Adaptive Object Detection (DAOD) relieves the reliance on large-scale annotated data by transferring the knowledge learned from a labeled source domain to a new unlabeled target domain. Recent DAOD approaches resort to local feature alignment in virtue of domain adversarial training in conjunction with the ad-hoc detection pipelines to achieve feature adaptation. However, these methods are limited to adapt the specific types of object detectors and do not explore the cross-domain topological relations. In this paper, we first formulate DAOD as an open-set domain adaptation problem in which foregrounds (pixel or region) can be seen as the “known class”, while backgrounds (pixel or region) are referred to as the “unknown class”. To this end, we present a new and general perspective for DAOD named Dual Bipartite Graph Learning (DBGL), which captures the cross-domain interactions on both pixel-level and semantic-level via increasing the distinction between foregrounds and backgrounds and modeling the cross-domain dependencies among different semantic categories. Experiments reveal that the proposed DBGL in conjunction with one-stage and two-stage detectors exceeds the state-of-the-art performance on standard DAOD benchmarks. Chaoqi Chen, Jiongcheng Li, Zebiao Zheng, Yue Huang 0001, Xinghao Ding, Yizhou Yu |
ICCV | 5 |
| 2021 | TRAR: Routing the Attention Spans in Transformer for Visual Question AnsweringabstractDue to the superior ability of global dependency modeling, Transformer and its variants have become the primary choice of many vision-and-language tasks. However, in tasks like Visual Question Answering (VQA) and Referring Expression Comprehension (REC), the multimodal prediction often requires visual information from macro- to micro-views. Therefore, how to dynamically schedule the global and local dependency modeling in Transformer has become an emerging issue. In this paper, we propose an example-dependent routing scheme called TRAnsformer Routing (TRAR) to address this issue1. Specifically, in TRAR, each visual Transformer layer is equipped with a routing module with different attention spans. The model can dynamically select the corresponding attentions based on the output of the previous inference step, so as to formulate the optimal routing path for each example. Notably, with careful designs, TRAR can reduce the additional computation and memory overhead to almost negligible. To validate TRAR, we conduct extensive experiments on five benchmark datasets of VQA and REC, and achieve superior performance gains than the standard Transformers and a bunch of state-of-the-art methods. Yiyi Zhou, Tianhe Ren, Xiaoshuai Sun, Jianzhuang Liu, Xinghao Ding, Mingliang Xu 0001, Rongrong Ji |
ICCV | 6 |
| 2021 | Consistent Posterior Distributions Under Vessel-Mixing: A Regularization For Cross-Domain Retinal Artery/Vein ClassificationabstractRetinal artery/vein (A/V) classification is a critical technique for diagnosing diabetes and cardiovascular diseases. Although deep learning based methods achieve impressive results in A/V classification, the performance usually degrades when directly apply the models that trained on one dataset to another set, due to the domain shift, e.g., caused by the variations in imaging protocols. In this paper, we propose a novel method to improve cross-domain generalization for pixel-wise retinal A/V classification. That is, vessel-mixing based consistency regularization, which regularizes the models to give consistent posterior distributions for vessel-mixing samples. The proposed method achieves the state-of-the-art performance on extensive experiments for cross-domain A/V classification, which is even close to the performance of fully supervised learning on target domain in some cases. Chenxin Li, Zhehan Liang, Wenao Ma, Yue Huang 0001, Xinghao Ding |
ICIP | 6 |
| 2021 | Noise2Grad: Extract Image Noise to DenoiseabstractIn many image denoising tasks, the difficulty of collecting noisy/clean image pairs limits the application of supervised CNNs. We consider such a case in which paired data and noise statistics are not accessible, but unpaired noisy and clean images are easy to collect. To form the necessary supervision, our strategy is to extract the noise from the noisy image to synthesize new data. To ease the interference of the image background, we use a noise removal module to aid noise extraction. The noise removal module first roughly removes noise from the noisy image, which is equivalent to excluding much background information. A noise approximation module can therefore easily extract a new noise map from the removed noise to match the gradient of the noisy input. This noise map is added to a random clean image to synthesize a new data pair, which is then fed back to the noise removal module to correct the noise removal process. These two modules cooperate to extract noise finely. After convergence, the noise removal module can remove noise without damaging other background details, so we use it as our final denoising network. Experiments show that the denoising performance of the proposed method is competitive with other supervised CNNs. Huangxing Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Yizhou Yu |
IJCAI | 4 |
| 2021 | Fast Magnetic Resonance Imaging on Regions of Interest: From Sensing to Reconstruction
Liyan Sun, Xinghao Ding, Yue Huang 0001, Yizhou Yu |
MICCAI (6) | 3 |
| 2021 | Generator Versus Segmentor: Pseudo-healthy Synthesis
Chenxin Li, Liyan Sun, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Yizhou Yu |
MICCAI (6) | 7 |
| 2021 | Successive Graph Convolutional Network for Image De-raining
Xueyang Fu, Qi Qi 0005, Zhengjun Zha, Xinghao Ding, Feng Wu 0001, John W. Paisley |
Int. J. Comput. Vis. | 4 |
| 2021 | Hard class rectification for domain adaptation
Changxing Jing, Huangxing Lin, Chaoqi Chen, Yue Huang 0001, Xinghao Ding, Yang Zou 0003 |
Knowl. Based Syst. | 6 |
| 2021 | Curriculum Feature Alignment Domain Adaptation for Epithelium-Stroma Classification in Histopathological ImagesabstractIn recent years, deep learning methods have received more attention in epithelial-stroma (ES) classification tasks. Traditional deep learning methods assume that the training and test data have the same distribution, an assumption that is seldom satisfied in complex imaging procedures. Unsupervised domain adaptation (UDA) transfers knowledge from a labelled source domain to a completely unlabeled target domain, and is more suitable for ES classification tasks to avoid tedious annotation. However, existing UDA methods for this task ignore the semantic alignment across domains. In this paper, we propose a Curriculum Feature Alignment Network (CFAN) to gradually align discriminative features across domains through selecting effective samples from the target domain and minimizing intra-class differences. Specifically, we developed the Curriculum Transfer Strategy (CTS) and Adaptive Centroid Alignment (ACA) steps to train our model iteratively. We validated the method using three independent public ES datasets, and experimental results demonstrate that our method achieves better performance in ES classification compared with commonly used deep learning methods and existing deep domain adaptation methods. Qi Qi 0005, Chaoqi Chen, Weiping Xie, Yue Huang 0001, Xinghao Ding, Yizhou Yu |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Crowd Density Estimation Using Fusion of Multi-Layer FeaturesabstractCrowd counting is very important in many tasks such as video surveillance, traffic monitoring, public security, and urban planning, so it is a very important part of the intelligent transportation system. However, achieving an accurate crowd counting and generating a precise density map are still challenging tasks due to the occlusion, perspective distortion, complex backgrounds, and varying scales. In addition, most of the existing methods focus only on the accuracy of crowd counting without considering the correctness of a density distribution; namely, there are many false negatives and false positives in a generated density map. To address this issue, we propose a novel encoder-decoder Convolution Neural Network (CNN) that fuses the feature maps in both encoding and decoding sub-networks to generate a more reasonable density map and estimate the number of people more accurately. Furthermore, we introduce a new evaluation method named the Patch Absolute Error (PAE) which is more appropriate to measure the accuracy of a density map. The extensive experiments on several existing public crowd counting datasets demonstrate that our approach achieves better performance than the current state-of-the-art methods. Lastly, considering the cross-scene crowd counting in practice, we evaluate our model on some cross-scene datasets. The results show our method has a good performance in cross-scene datasets. Xinghao Ding, Fujin He, Zhirui Lin, Yu Wang 0160, Huimin Guo, Yue Huang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Deep Multiscale Detail Networks for Multiband Spectral Image SharpeningabstractWe introduce a new deep detail network architecture with grouped multiscale dilated convolutions to sharpen images contain multiband spectral information. Specifically, our end-to-end network directly fuses low-resolution multispectral and panchromatic inputs to produce high-resolution multispectral results, which is the same goal of the pansharpening in remote sensing. The proposed network architecture is designed by utilizing our domain knowledge and considering the two aims of the pansharpening: spectral and spatial preservations. For spectral preservation, the up-sampled multispectral images are directly added to the output for lossless spectral information propagation. For spatial preservation, we train the proposed network in the high-frequency domain instead of the commonly used image domain. Different from conventional network structures, we remove pooling and batch normalization layers to preserve spatial information and improve generalization to new satellites, respectively. To effectively and efficiently obtain multiscale contextual features at a fine-grained level, we propose a grouped multiscale dilated network structure to enlarge the receptive fields for each network layer. This structure allows the network to capture multiscale representations without increasing the parameter burden and network complexity. These representations are finally utilized to reconstruct the residual images which contain spatial details of PAN. Our trained network is able to generalize different satellite images without the need for parameter tuning. Moreover, our model is a general framework, which can be directly used for other kinds of multiband spectral image sharpening, e.g., hyperspectral image sharpening. Experiments show that our model performs favorably against compared methods in terms of both qualitative and quantitative qualities. Xueyang Fu, Yue Huang 0001, Xinghao Ding, John W. Paisley |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Harmonizing Transferability and Discriminability for Adapting Object DetectorsabstractRecent advances in adaptive object detection have achieved compelling results in virtue of adversarial feature adaptation to mitigate the distributional shifts along the detection pipeline. Whilst adversarial adaptation significantly enhances the transferability of feature representations, the feature discriminability of object detectors remains less investigated. Moreover, transferability and discriminability may come at a contradiction in adversarial adaptation given the complex combinations of objects and the differentiated scene layouts between domains. In this paper, we propose a Hierarchical Transferability Calibration Network (HTCN) that hierarchically (local-region/image/instance) calibrates the transferability of feature representations for harmonizing transferability and discriminability. The proposed model consists of three components: (1) Importance Weighted Adversarial Training with input Interpolation (IWAT-I), which strengthens the global discriminability by re-weighting the interpolated image-level features; (2) Context-aware Instance-Level Alignment (CILA) module, which enhances the local discriminability by capturing the underlying complementary effect between the instance-level feature and the global context information for the instance-level feature alignment; (3) local feature masks that calibrate the local transferability to provide semantic guidance for the following discriminative pattern alignment. Experimental results show that HTCN significantly outperforms the state-of-the-art methods on benchmark datasets. Chaoqi Chen, Zebiao Zheng, Xinghao Ding, Yue Huang 0001, Qi Dou 0001 |
CVPR | 3 |
| 2020 | K-armed Bandit based Multi-Modal Network Architecture Search for Visual Question AnsweringabstractIn this paper, we propose a cross-modal network architecture search (NAS) algorithm for VQA, termed as k-Armed Bandit based NAS (KAB-NAS). KAB-NAS regards the design of each layer as a k-armed bandit problem and updates the preference of each candidate via numerous samplings in a single-shot search framework. To establish an effective search space, we further propose a new architecture termed Automatic Graph Attention Network (AGAN), and extend the popular self-attention layer with three graph structures, denoted as dense-graph, co-graph and separate-graph.These graph layers are used to form the direction of information propagation in the graph network, and their optimal combinations are searched by KAB-NAS. To evaluate KAB-NAS and AGAN, we conduct extensive experiments on two VQA benchmark datasets, i.e., VQA2.0 and GQA, and also test AGAN with the popular BERT-style pre-training. The experimental results show that with the help of KAB-NAS, AGAN can achieve the state-of-the-art performance on both benchmark datasets with much fewer parameters and computations. Yiyi Zhou, Rongrong Ji, Xiaoshuai Sun, Gen Luo, Xiaopeng Hong, Jinsong Su, Xinghao Ding, Ling Shao 0001 |
ACM Multimedia | 7 |
| 2020 | A dual-domain deep lattice network for rapid MRI reconstruction
Liyan Sun, Yawen Wu, Binglin Shu, Xinghao Ding, Congbo Cai, Yue Huang 0001, John W. Paisley |
Neurocomputing | 4 |
| 2020 | Multiple-source domain adaptation with generative adversarial nets
Chaoqi Chen, Weiping Xie, Yi Wen 0003, Yue Huang 0001, Xinghao Ding |
Knowl. Based Syst. | 5 |
| 2020 | Underwater image enhancement using an edge-preserving filtering Retinex algorithm
Peixian Zhuang, Xinghao Ding |
Multim. Tools Appl. | 2 |
| 2020 | Correction to: Underwater image enhancement using an edge-preserving filtering Retinex algorithm
Peixian Zhuang, Xinghao Ding |
Multim. Tools Appl. | 2 |
| 2020 | Rain O'er Me: Synthesizing Real Rain to Derain With Data DistillationabstractWe present a weakly-supervised technique for learning to remove rain from images without using synthetic rain software. The method is based on a two-stage data distillation approach, which requires only some unpaired rainy and clean images to generate supervision. First, a rainy image is paired with a coarsely derained version using on a simple filtering technique (“rain-to-clean”). Then a clean image is randomly matched with the rainy soft-labeled pair. Through a shared deep neural network, the rain that is removed from the first image is then added to the clean image to generate a second pair (“clean-to-rain”). The neural network simultaneously learns to map both images such that high resolution structure in the clean images can inform the deraining of the rainy images. Demonstrations show that this approach can address those visual characteristics of rain not easily synthesized by software in the usual way. Huangxing Lin, Xueyang Fu, Xinghao Ding, Yue Huang 0001, John W. Paisley |
IEEE Trans. Image Process. | 4 |
| 2020 | An Adversarial Learning Approach to Medical Image Synthesis for Lesion DetectionabstractThe identification of lesion within medical image data is necessary for diagnosis, treatment and prognosis. Segmentation and classification approaches are mainly based on supervised learning with well-paired image-level or voxel-level labels. However, labeling the lesion in medical images is laborious requiring highly specialized knowledge. We propose a medical image synthesis model named abnormal-to-normal translation generative adversarial network (ANT-GAN) to generate a normal-looking medical image based on its abnormal-looking counterpart without the need for paired training data. Unlike typical GANs, whose aim is to generate realistic samples with variations, our more restrictive model aims at producing a normal-looking image corresponding to one containing lesions, and thus requires a special design. Being able to provide a "normal" counterpart to a medical image can provide useful side information for medical imaging tasks like lesion segmentation or classification validated by our experiments. In the other aspect, the ANT-GAN model is also capable of producing highly realistic lesion-containing image corresponding to the healthy one, which shows the potential in data augmentation verified in our experiments. Liyan Sun, Jiexiang Wang, Yue Huang 0001, Xinghao Ding, Hayit Greenspan, John W. Paisley |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Cost-Effective Vehicle Type Recognition in Surveillance Images With Deep Active Learning and Web DataabstractRecently, vehicle type recognition in surveillance images with deep learning has received significant attention in various applications of intelligent transportation systems. However, annotating large-scale images from many surveillance images is tedious and time-consuming, which impedes its application in the real world. This paper aims to resolve this problem by reducing manual labeling in surveillance images, and then maximizing the effect of the few tagged data. Thus, a deep active learning method with a new query strategy is proposed in this paper for vehicle type recognition in surveillance images. First, the proposed method constructs a memory space using a large-scale fully labeled auxiliary dataset collected from the Internet. Subsequently, two metrics, the similarity measurement in memory space and the entropy, are used to simultaneously emphasize the diversity and uncertainty in the query strategy. Moreover, an additional label-consistent term apart from the hyper-parameters is used to adaptively adjust the combination of the two principles in active learning. The proposed method was evaluated on the Comprehensive Cars dataset. The experimental results demonstrated that the proposed method could effectively reduce the annotation cost by up to 40% in surveillance vehicle type recognition compared with the random selection method. Yue Huang 0001, Minghui Jiang 0004, Xinghao Ding |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | A 3D Spatially Weighted Network for Segmentation of Brain Tissue From MRIabstractThe segmentation of brain tissue in MRI is valuable for extracting brain structure to aid diagnosis, treatment and tracking the progression of different neurologic diseases. Medical image data are volumetric and some neural network models for medical image segmentation have addressed this using a 3D convolutional architecture. However, this volumetric spatial information has not been fully exploited to enhance the representative ability of deep networks, and these networks have not fully addressed the practical issues facing the analysis of multimodal MRI data. In this paper, we propose a spatially-weighted 3D network (SW-3D-UNet) for brain tissue segmentation of single-modality MRI, and extend it using multimodality MRI data. We validate our model on the MRBrainS13 and MALC12 datasets. This unpublished model ranked first on the leaderboard of the MRBrainS13 Challenge. Liyan Sun, Wenao Ma, Xinghao Ding, Yue Huang 0001, Dong Liang 0001, John W. Paisley |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Lightweight Pyramid Networks for Image DerainingabstractExisting deep convolutional neural networks (CNNs) have found major success in image deraining, but at the expense of an enormous number of parameters. This limits their potential applications, e.g., in mobile devices. In this paper, we propose a lightweight pyramid networt (LPNet) for single-image deraining. Instead of designing a complex network structure, we use domain-specific knowledge to simplify the learning process. In particular, we find that by introducing the mature Gaussian-Laplacian image pyramid decomposition technology to the neural network, the learning problem at each pyramid level is greatly simplified and can be handled by a relatively shallow network with few parameters. We adopt recursive and residual network structures to build the proposed LPNet, which has less than 8K parameters while still achieving the state-of-the-art performance on rain removal. We also discuss the potential value of LPNet for other low- and high-level vision tasks. Xueyang Fu, Borong Liang, Yue Huang 0001, Xinghao Ding, John W. Paisley |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Progressive Feature Alignment for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) transfers knowledge from a label-rich source domain to a fully-unlabeled target domain. To tackle this task, recent approaches resort to discriminative domain transfer in virtue of pseudo-labels to enforce the class-level distribution alignment across the source and target domains. These methods, however, are vulnerable to the error accumulation and thus incapable of preserving cross-domain category consistency, as the pseudo-labeling accuracy is not guaranteed explicitly. In this paper, we propose the Progressive Feature Alignment Network (PFAN) to align the discriminative features across domains progressively and effectively, via exploiting the intra-class variation in the target domain. To be specific, we first develop an Easy-to-Hard Transfer Strategy (EHTS) and an Adaptive Prototype Alignment (APA) step to train our model iteratively and alternatively. Moreover, upon observing that a good domain adaptation usually requires a non-saturated source classifier, we consider a simple yet efficient way to retard the convergence speed of the source classification loss by further involving a temperature variate into the soft-max function. The extensive experimental results reveal that the proposed PFAN exceeds the state-of-the-art performance on three UDA datasets. Chaoqi Chen, Weiping Xie, Wenbing Huang 0001, Yu Rong 0001, Xinghao Ding, Yue Huang 0001, Tingyang Xu, Junzhou Huang |
CVPR | 5 |
| 2019 | A Variational Pan-Sharpening With Local Gradient ConstraintsabstractPan-sharpening aims at fusing spectral and spatial information, which are respectively contained in the multispectral (MS) image and panchromatic (PAN) image, to produce a high resolution multi-spectral (HRMS) image. In this paper, a new variational model based on a local gradient constraint for pan-sharpening is proposed. Different with previous methods that only use global constraints to preserve spatial information, we first consider gradient difference of PAN and HRMS images in different local patches and bands. Then a more accurate spatial preservation based on local gradient constraints is incorporated into the objective to fully utilize spatial information contained in the PAN image. The objective is formulated as a convex optimization problem which minimizes two leastsquares terms and thus very simple and easy to implement. A fast algorithm is also designed to improve efficiency. Experiments show that our method outperforms previous variational algorithms and achieves better generalization than recent deep learning methods. Xueyang Fu, Zihuang Lin, Yue Huang 0001, Xinghao Ding |
CVPR | 4 |
| 2019 | Look More Than Once: An Accurate Detector for Text of Arbitrary ShapesabstractPrevious scene text detection methods have progressed substantially over the past years. However, limited by the receptive field of CNNs and the simple representations like rectangle bounding box or quadrangle adopted to describe text, previous methods may fall short when dealing with more challenging text instances, such as extremely long text and arbitrarily shaped text. To address these two problems, we present a novel text detector namely LOMO, which localizes the text progressively for multiple times (or in other word, LOok More than Once). LOMO consists of a direct regressor (DR), an iterative refinement module (IRM) and a shape expression module (SEM). At first, text proposals in the form of quadrangle are generated by DR branch. Next, IRM progressively perceives the entire long text by iterative refinement based on the extracted feature blocks of preliminary proposals. Finally, a SEM is introduced to reconstruct more precise representation of irregular text by considering the geometry properties of text instance, including text region, text center line and border offsets. The state-of-the-art results on several public benchmarks including ICDAR2017-RCTW, SCUT-CTW1500, Total-Text, ICDAR2015 and ICDAR17-MLT confirm the striking robustness and effectiveness of LOMO. Chengquan Zhang, Borong Liang, Zuming Huang, Mengyi En, Junyu Han, Errui Ding, Xinghao Ding |
CVPR | 7 |
| 2019 | Lung Nodule Detection with a 3D ConvNet via IoU Self-normalization and Maxout UnitabstractThe automatic pulmonary nodule detection in thoracic computed tomography (CT) scans plays a crucial role in the early diagnosis of lung cancer. In this paper, we propose a novel framework with a 3D convolutional network (ConvNet) for pulmonary nodule detection. To improve the efficiency and flexibility, we adopt one-stage process without the false positive reduction stage. Specially, the great challenge of the nodule detection is the recall rate of small nodules. We propose two methods to solve this issue. Firstly, we set the classification label by the intersection over union (IoU) self-normalization, which enables to eliminate the loss of regression information caused by misleading classification confidence. Secondly, pulmonary nodules differ in size, shape and density, leading to large intra-class variations. We introduce maxout unit to solve this problem. Overall, we achieve an average FROC score of 0.912 on LUNA16 dataset, outperforming all other one-stage models as far as we know. Fei Li 0021, Yawen Wu, Congbo Cai, Yue Huang 0001, Xinghao Ding |
ICASSP | 6 |
| 2019 | Two-stream Multi-focus Image Fusion Based on the Latent Decision MapabstractThe multi-focus image fusion with deep learning methods is mostly regarded as a two or three-category problem. Current systems utilize sliding windows to classify each pixel into focused or defocused, which is time consuming and requires post-processing such as denoising. In this paper, we propose a novel network architecture for multi-focus image fusion based on the latent decision map. For a regression task instead of a classification problem, we focus on learning the latent spatial decision map. This decision map indicates the degree of each focused pixel. To further improve the fusion result, we utilize the ResNet blocks to extract image features, and then combine low-level features with high-level semantic information. Our apporach makes the learning process easier and has better robustness and efficiency as well. Experimental results demonstrate that our framework has ability of achieving the state-of-the-art in terms of both qualitative and quantitative measures. Weihong Zeng, Fei Li 0021, Yue Huang 0001, Xinghao Ding |
ICASSP | 5 |
| 2019 | JPEG Artifacts Reduction via Deep Convolutional Sparse CodingabstractTo effectively reduce JPEG compression artifacts, we propose a deep convolutional sparse coding (DCSC) network architecture. We design our DCSC in the framework of classic learned iterative shrinkage-threshold algorithm. To focus on recognizing and separating artifacts only, we sparsely code the feature maps instead of the raw image. The final de-blocked image is directly reconstructed from the coded features. We use dilated convolution to extract multi-scale image features, which allows our single model to simultaneously handle multiple JPEG compression levels. Since our method integrates model-based convolutional sparse coding with a learning-based deep neural network, the entire network structure is compact and more explainable. The resulting lightweight model generates comparable or better de-blocking results when compared with state-of-the-art methods. Xueyang Fu, Zhengjun Zha, Feng Wu 0001, Xinghao Ding, John W. Paisley |
ICCV | 4 |
| 2019 | Deep Blind Hyperspectral Image FusionabstractHyperspectral image fusion (HIF) reconstructs high spatial resolution hyperspectral images from low spatial resolution hyperspectral images and high spatial resolution multispectral images. Previous works usually assume that the linear mapping between the point spread functions of the hyperspectral camera and the spectral response functions of the conventional camera is known. This is unrealistic in many scenarios. We propose a method for blind HIF problem based on deep learning, where the estimation of the observation model and fusion process are optimized iteratively and alternatingly during the super-resolution reconstruction. In addition, the proposed framework enforces simultaneous spatial and spectral accuracy. Using three public datasets, the experimental results demonstrate that the proposed algorithm outperforms existing blind and non-blind methods. Weihong Zeng, Yue Huang 0001, Xinghao Ding, John W. Paisley |
ICCV | 4 |
| 2019 | Compressed Sensing MRI with Joint Image-Level and Patch-Level PriorsabstractWe develop a novel compressed sensing magnetic resonance imaging (CSMRI) algorithm with joint image and patch priors, where the total variation (TV) is adopted for image-level sparse prior and the expected patch log likelihood (EPLL) is used for patch-level sparse prior. The proposed joint priors capture global and local sparse nature of the MR image to promote image structures and suppress artifacts or noise, which can alleviate the limitations of previous CSMRI methods. And we derive an appropriate cost function which can be addressed by an efficient optimization scheme that iteratively alternates among l1norm approximation, latent patch reconstruction and ideal image reconstruction. Final experiments are provided to show the satisfactory performance of the proposed method in MRI reconstruction, which outperforms other competitive CSMRI approaches in both subjective results and objective assessments. Peixian Zhuang, Xinghao Ding |
ICIP | 2 |
| 2019 | Pay Attention to Deep Feature Fusion in Crowd Density Estimation
Huimin Guo, Fujin He, Xinghao Ding, Yue Huang 0001 |
ICONIP (4) | 4 |
| 2019 | G-HAPNet: A Novel Structure for Single Image Super-Resolution
Mingyong Zhuang, Congbo Cai, Yue Huang 0001, Xinghao Ding |
ICONIP (5) | 5 |
| 2019 | T-SAMnet: A Segmentation Driven Network for Image Manipulation Detection
Yunshu Chen, Yue Huang 0001, Xinghao Ding, En Cheng |
ICONIP (5) | 4 |
| 2019 | A Robustness and Low Bit-Rate Image Compression Network for Underwater Acoustic Communication
Mingyong Zhuang, Xinghao Ding, Yue Huang 0001, Yinghao Liao |
ICONIP (2) | 3 |
| 2019 | Multi-task Neural Networks with Spatial Activation for Retinal Vessel Segmentation and Artery/Vein Classification
Wenao Ma, Kai Ma 0002, Jiexiang Wang, Xinghao Ding, Yefeng Zheng 0001 |
MICCAI (1) | 5 |
| 2019 | Divide-and-conquer framework for image restoration and enhancement
Peixian Zhuang, Xinghao Ding |
Eng. Appl. Artif. Intell. | 2 |
| 2019 | Prominent edge detection with deep metric expression and multi-scale features
Shulian Cai, Jiabin Huang 0003, Yue Huang 0001, Xinghao Ding, Delu Zeng |
Multim. Tools Appl. | 5 |
| 2019 | Pan-GGF: A probabilistic method for pan-sharpening with gradient domain guided image filtering
Peixian Zhuang, Qingshan Liu 0001, Xinghao Ding |
Signal Process. | 3 |
| 2019 | MRI reconstruction with an edge-preserving filtering prior
Peixian Zhuang, Xinghao Ding |
Signal Process. | 3 |
| 2019 | A Deep Information Sharing Network for Multi-Contrast Compressed Sensing MRI ReconstructionabstractCompressed sensing (CS) theory can accelerate multi-contrast magnetic resonance imaging (MRI) by sampling fewer measurements within each contrast. However, conventional optimization-based reconstruction models suffer several limitations, including a strict assumption of shared sparse support, time-consuming optimization, and "shallow" models with difficulties in encoding the patterns contained in massive MRI data. In this paper, we propose the first deep learning model for multi-contrast CS-MRI reconstruction. We achieve information sharing through feature sharing units, which significantly reduces the number of model parameters. The feature sharing unit combines with a data fidelity unit to comprise an inference block, which are then cascaded with dense connections, allowing for efficient information transmission across different depths of the network. Experiments on various multi-contrast MRI datasets show that the proposed model outperforms both state-of-the-art single-contrast and multi-contrast MRI methods in accuracy and efficiency. We demonstrate that improved reconstruction quality can bring benefits to subsequent medical image analysis. Furthermore, the robustness of the proposed model to misregistration shows its potential in real MRI applications. Liyan Sun, Zhiwen Fan, Xueyang Fu, Yue Huang 0001, Xinghao Ding, John W. Paisley |
IEEE Trans. Image Process. | 5 |
| 2019 | Label-Efficient Breast Cancer Histopathological Image ClassificationabstractThe automatic classification of breast cancer histopathological images has great significance in computer-aided diagnosis. Recently, deep learning via neural networks has enabled pattern detection and prediction using large, labeled datasets; whereas, collecting and annotating sufficient histological data using professional pathologists is time consuming, tedious, and extremely expensive. In the proposed paper, a deep active learning framework is designed and implemented for classification of breast cancer histopathological images, with the goal of maximizing the learning accuracy from very limited labeling. This method involves manual annotation of the most valuable unlabeled samples, which are then integrated into the training set. The model is then iteratively updated with an increasing training set. Here, two selection strategies are discussed for the proposed deep active learning framework: An entropy-based strategy and a confidence-boosting strategy. The proposed method has been validated using a publicly available breast cancer histopathological image dataset, wherein each image patch is binarily classified as benign or malignant. The experimental results demonstrate that, compared with a random selection, our proposed framework can reduce annotation costs up to 66.67%, with higher accuracy and less expensive annotation than standard query strategy. Qi Qi 0005, Jitian Wang, Han Zheng 0004, Yue Huang 0001, Xinghao Ding, Gustavo K. Rohde |
IEEE J. Biomed. Health Informatics | 6 |
| 2018 | Compressed Sensing MRI Using a Recursive Dilated NetworkabstractCompressed sensing magnetic resonance imaging (CS-MRI) is an active research topic in the field of inverse problems. Conventional CS-MRI algorithms usually exploit the sparse nature of MRI in an iterative manner. These optimization-based CS-MRI methods are often time-consuming at test time, and are based on fixed transform bases or shallow dictionaries, which limits modeling capacity. Recently, deep models have been introduced to the CS-MRI problem. One main challenge for CS-MRI methods based on deep learning is the trade off between model performance and network size. We propose a recursive dilated network (RDN) for CS-MRI that achieves good performance while reducing the number of network parameters. We adopt dilated convolutions in each recursive block to aggregate multi-scale information within the MRI. We also adopt a modified shortcut strategy to help features flow into deeper layers. Experimental results show that the proposed RDN model achieves state-of-the-art performance in CS-MRI while using far fewer parameters than previously required. Liyan Sun, Zhiwen Fan, Yue Huang 0001, Xinghao Ding, John W. Paisley |
AAAI | 4 |
| 2018 | A Segmentation-Aware Deep Fusion Network for Compressed Sensing MRI
Zhiwen Fan, Liyan Sun, Xinghao Ding, Yue Huang 0001, Congbo Cai, John W. Paisley |
ECCV (6) | 3 |
| 2018 | Breast Cancer Histopathological Image Classification via Deep Active Learning and Confidence Boosting
Baolin Du, Qi Qi 0005, Han Zheng 0004, Yue Huang 0001, Xinghao Ding |
ICANN (2) | 5 |
| 2018 | A Deeply-Recursive Convolutional Network For Crowd CountingabstractThe estimation of crowd count in images has a wide range of applications such as video surveillance, traffic monitoring, public safety and urban planning. Recently, the convolutional neural network (CNN) based approaches have been shown to be more effective in crowd counting than traditional methods that use handcrafted features. However, the existing CNN-based methods still suffer from large number of parameters and large storage space, which require high storage and computing resources and thus limit the real-world application. Consequently, we propose a deeply-recursive network (DR-ResNet) based on ResNet blocks for crowd counting. The recursive structure makes the network deeper while keeping the number of parameters unchanged, which enhances network capability to capture statistical regularities in the context of the crowd. Besides, we generate a new dataset from the video-monitoring data of Beijing bus station. Experimental results have demonstrated that proposed method outperforms most state-of-the-art methods with far less number of parameters. Xinghao Ding, Zhirui Lin, Fujin He, Yu Wang 0160, Yue Huang 0001 |
ICASSP | 1 |
| 2018 | Bindctnet: A Simple Binary Dct Network for Image ClassificationabstractConvolution neural networks play an important role in the image classification tasks. However, it is time consuming to train the network and the cost of memory resources is usually high. In this paper, a simple and effective network named BinDCTNet is presented by using the binary discrete cosine transform(BinDCT) to extract the feature-maps and a hyper-parameter to reduce dimension of the extracted feature. The proposed network has extremely low computing complexity and there is almost no parameters needed to be stored. Experiments are carried out on the hand written digit dataset MNIST and the vehicle logo VLOGO dataset. The results show that the proposed network achieves the state-of-the-art accuracy with fast speed and low memory cost, which makes it applicable on mobile and embedded devices. Xiangrui Xing, Wenao Ma, Yue Huang 0001, Delu Zeng, Xinghao Ding |
ICASSP | 6 |
| 2018 | Man-Made Object Recognition from Underwater Optical Images Using Deep Learning and Transfer LearningabstractWith the development of underwater optical sensors, manmade object recognition from underwater optical images has attracted wide attention. Deep learning methods have demonstrated impressive performance in object recognition tasks from natural images. However, it is difficult to collect large-scale labeled underwater optical images for training such a model. Based on the assumption that it is possible to acquire sufficient labeled in-air images, the proposed work leverages a combination of deep learning and transfer learning to develop a novel recognition system for man-made object from underwater optical images. The extracted features from the proposed network have high representative power, and demonstrate robustness in both in-air and underwater imaging modalities. Therefore, our proposed framework has the ability to recognize underwater man-made objects using only labeled in -air images. The results of experiments on simulated data demonstrate that the proposed method outperforms traditional deep learning methods in the task of underwater man-made object recognition. Xiangrui Xing, Han Zheng 0004, Xueyang Fu, Yue Huang 0001, Xinghao Ding |
ICASSP | 6 |
| 2018 | Age Estimation from MR Images via 3D Convolutional Neural Network and Densely Connect
Qi Qi 0005, Baolin Du, Mingyong Zhuang, Yue Huang 0001, Xinghao Ding |
ICONIP (7) | 5 |
| 2018 | Weakly-Supervised Man-Made Object Recognition in Underwater Optimal Image Through Deep Domain Adaptation
Chaoqi Chen, Weiping Xie, Yue Huang 0001, Xinghao Ding |
ICONIP (5) | 5 |
| 2018 | High Efficient Reconstruction of Single-Shot Magnetic Resonance T_2 Mapping Through Overlapping Echo Detachment and DenseNet
Yawen Wu, Xinghao Ding, Yue Huang 0001, Congbo Cai |
ICONIP (6) | 3 |
| 2018 | A Deep Ensemble Network for Compressed Sensing MRI
Huafeng Wu, Yawen Wu, Liyan Sun, Congbo Cai, Yue Huang 0001, Xinghao Ding |
ICONIP (1) | 6 |
| 2018 | MEnet: A Metric Expression Network for Salient Object SegmentationabstractRecent CNN-based saliency models have achieved excellent performance on public datasets, but most are sensitive to distortions from noise or compression. In this paper, we propose an end-to-end generic salient object segmentation model called Metric Expression Network (MEnet) to overcome this drawback. We construct a topological metric space where the implicit metric is determined by a deep network. In this latent space, we can group pixels within an observed image semantically into two regions, based on whether they are in a salient region or a non-salient region in the image. We carry out all feature extractions at the pixel level, which makes the output boundaries of the salient object finely-grained. Experimental results show that the proposed metric can generate robust salient maps that allow for object segmentation. By testing the method on several public benchmarks, we show that the performance of MEnet achieves excellent results. We also demonstrate that the proposed method outperforms previous CNN-based methods on distorted images. Shulian Cai, Jiabin Huang 0003, Delu Zeng, Xinghao Ding, John W. Paisley |
IJCAI | 4 |
| 2018 | Residual-Guide Network for Single Image DerainingabstractSingle image rain streaks removal is extremely important since rainy condition adversely affects many computer vision systems. Deep learning based methods have great success in image deraining tasks. In this paper, we propose a novel residual-guide feature fusion network, called ResGuideNet, for single image deraining that progressively predicts high-quality reconstruction while using fewer parameters than previous methods. Specifically, we propose a cascaded network and adopt residuals from shallower blocks to guide deeper blocks. We can obtain a coarse-to-fine estimation of negative residual as the blocks go deeper with this strategy. The outputs of different blocks are merged into the final reconstruction. We adopt recursive convolution to build each block and apply supervision to intermediate de-rained results. ResGuideNet is detachable to meet different rainy conditions. For images with light rain streaks and limited computational resource at test time, we can obtain a decent performance even with several building blocks. Experiments validate that ResGuideNet can benefit other low- and high-level vision tasks. Zhiwen Fan, Huafeng Wu, Xueyang Fu, Yue Huang 0001, Xinghao Ding |
ACM Multimedia | 5 |
| 2018 | Vehicle Type Recognition in Surveillance Images From Labeled Web-Nature Data Using Deep Transfer LearningabstractVehicle type recognition from surveillance images represents a challenging task in the domain of intelligent monitoring systems. Recently, deep learning methods have been applied to solve this problem. The existing deep learning methods, such as convolutional neural networks (CNN), assume that the training and test data are generated from the same or similar imaging systems. They also require a lot of manual annotations for each task. In this paper, we aim to create an improved deep learning method for vehicle type recognition from surveillance images and propose a system based on CNN and transfer learning. Labeled image data of different types of vehicles are easy to acquire from both vehicle manufacturers and Internet sources. Therefore, our proposed surveillance-based vehicle type recognition system is implemented using only labels from Web data. This allows us to overcome the task of manually labeling the data from surveillance images during the training phase. We need to overcome the gap in the types of vehicles between two different imaging systems. For this, a regularization technique in transfer learning is introduced to the objective function of the traditional convolutional neural network. The proposed method was verified through experiments with the public data set comprehensive cars. The experimental results demonstrate that our proposed recognition method outperforms existing deep learning methods when the training and test data are taken from different imaging systems. Jitian Wang, Han Zheng 0004, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2017 | Removing Rain from Single Images via a Deep Detail NetworkabstractWe propose a new deep network architecture for removing rain streaks from individual images based on the deep convolutional neural network (CNN). Inspired by the deep residual network (ResNet) that simplifies the learning process by changing the mapping form, we propose a deep detail network to directly reduce the mapping range from input to output, which makes the learning process easier. To further improve the de-rained result, we use a priori image domain knowledge by focusing on high frequency detail during training, which removes background interference and focuses the model on the structure of rain in images. This demonstrates that a deep architecture not only has benefits for high-level vision tasks but also can be used to solve low-level imaging problems. Though we train the network on synthetic data, we find that the learned network generalizes well to real-world test images. Experiments show that the proposed method significantly outperforms state-of-the-art methods on both synthetic and real-world images in terms of both qualitative and quantitative measures. We discuss applications of this structure to denoising and JPEG artifact reduction at the end of the paper. Xueyang Fu, Jiabin Huang 0003, Delu Zeng, Yue Huang 0001, Xinghao Ding, John W. Paisley |
CVPR | 5 |
| 2017 | Epithelium-stroma classification in histopathological images via convolutional neural networks and self-taught learningabstractEpithelium-stroma classification is always considered as an important preprocessing step for morphological quantitative analysis in image-based histological researches of oncologic diseases. However, large-scale accurate ground-truth labeling is expensive in histopathological image analysis, thus the classification performances will still be limited with the insufficient labeled training samples. Considering that acquisition of public unlabeled histopathological images is much cheaper, an epithelium-stroma classification framework is developed, based on the deep convolutional neural network framework and the strategies of self-taught learning. The method has the ability of taking advantage of large-scale unlabeled public histopathological data as auxiliary data, and then transferring the knowledge to enhance the performances in epithelium-stroma classification with limited labeled training data. The experiments demonstrate that the proposed method outperforms traditional CNNs when the labeled training data size is decreasing dramatically. Yue Huang 0001, Han Zheng 0004, Gustavo K. Rohde, Delu Zeng, Xinghao Ding |
ICASSP | 7 |
| 2017 | PanNet: A Deep Network Architecture for Pan-SharpeningabstractWe propose a deep network architecture for the pan-sharpening problem called PanNet. We incorporate domain-specific knowledge to design our PanNet architecture by focusing on the two aims of the pan-sharpening problem: spectral and spatial preservation. For spectral preservation, we add up-sampled multispectral images to the network output, which directly propagates the spectral information to the reconstructed image. To preserve spatial structure, we train our network parameters in the high-pass filtering domain rather than the image domain. We show that the trained network generalizes well to images from different satellites without needing retraining. Experiments show significant improvement over state-of-the-art methods visually and in terms of standard quality metrics. Xueyang Fu, Yuwen Hu, Yue Huang 0001, Xinghao Ding, John W. Paisley |
ICCV | 5 |
| 2017 | Camera model identification with residual neural networkabstractWith the development of multimedia, camera model identification from given images has attract large attentions in cyber-forensic area recently. The task has achieved a great improvement due to some deep learning methods, where the features are extracted with the stacked architectures. However, it should be considered that both low-level and highlevel features have contributions to the recognition. In this paper, we investigate the task with another deep learning model, residual neural network (ResNet). Proposed framework has been evaluated on the experiments of brand-attribution, model-attribution and device-attribution. Besides, we also include cell phone model identification in the brand-attribution experiment for the first time. The classification results have demonstrated that the proposed work has the ability of enhancing the identification performances compared with existing methods in each specific task. The proposed work can be considered as an effective approach on image forensics. Yunshu Chen, Yue Huang 0001, Xinghao Ding |
ICIP | 3 |
| 2017 | Compressed sensing MRI using total variation regularization with K-space decompositionabstractCompressed sensing theory facilitates the fast magnetic resonance imaging by reducing the required number of measurements for reconstruction. Conventional compressed sensing magnetic resonance imaging(CSMRI) method utilize the partial k-space measurements as a whole without considering their intrinsic property. Some recent researches have shown the advantage of dealing the high and low frequency image content separately. Based on this, we propose a novel CSMRI algorithm based on total variation regularization with k-space decomposition. First we decompose k-space into high frequency band and low frequency band, then we reconstruct the corresponding high and low MR images which will be used for integration later. All the steps can be unified into a objective function. We will show that the proposed objective function can be split into several subproblems to solve iteratively using ADMM technique. The experimental results show that the proposed method outperforms the conventional CSMRI method. Besides, the proposed method can be extended to other image processing applications as well. Liyan Sun, Yue Huang 0001, Congbo Cai, Xinghao Ding |
ICIP | 4 |
| 2017 | A Simple Convolutional Transfer Neural Networks in Vision Tasks
Wenlei Wu, Zhaohang Lin, Xinghao Ding, Yue Huang 0001 |
ICONIP (4) | 3 |
| 2017 | Multiple-Instance feature extraction at the bag and instance levels using the maximum trace-difference criterion
Jing Chai, Bo Chen 0001, Xinghao Ding |
Inf. Sci. | 5 |
| 2017 | Image enhancement using divide-and-conquer strategy
Peixian Zhuang, Xueyang Fu, Yue Huang 0001, Xinghao Ding |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | Non-blind deconvolution with ℓ 1 -norm of high-frequency fidelity
Peixian Zhuang, Yue Huang 0001, Delu Zeng, Xinghao Ding |
Multim. Tools Appl. | 4 |
| 2017 | Clearing the Skies: A Deep Network Architecture for Single-Image Rain RemovalabstractWe introduce a deep network architecture called DerainNet for removing rain streaks from an image. Based on the deep convolutional neural network (CNN), we directly learn the mapping relationship between rainy and clean image detail layers from data. Because we do not possess the ground truth corresponding to real-world rainy images, we synthesize images with rain for training. In contrast to other common strategies that increase depth or breadth of the network, we use image processing domain knowledge to modify the objective function and improve deraining with a modestly sized CNN. Specifically, we train our DerainNet on the detail (high-pass) layer rather than in the image domain. Though DerainNet is trained on synthetic data, we find that the learned network translates very effectively to real-world images for testing. Moreover, we augment the CNN framework with image enhancement to improve the visual results. Compared with the state-of-the-art single image de-raining methods, our method has improved rain removal and much faster computation time after network training. Xueyang Fu, Jiabin Huang 0003, Xinghao Ding, Yinghao Liao, John W. Paisley |
IEEE Trans. Image Process. | 3 |
| 2017 | Epithelium-Stroma Classification via Convolutional Neural Networks and Unsupervised Domain Adaptation in Histopathological ImagesabstractEpithelium-stroma classification is a necessary preprocessing step in histopathological image analysis. Current deep learning based recognition methods for histology data require collection of large volumes of labeled data in order to train a new neural network when there are changes to the image acquisition procedure. However, it is extremely expensive for pathologists to manually label sufficient volumes of data for each pathology study in a professional manner, which results in limitations in real-world applications. A very simple but effective deep learning method, that introduces the concept of unsupervised domain adaptation to a simple convolutional neural network (CNN), has been proposed in this paper. Inspired by transfer learning, our paper assumes that the training data and testing data follow different distributions, and there is an adaptation operation to more accurately estimate the kernels in CNN in feature extraction, in order to enhance performance by transferring knowledge from labeled data in source domain to unlabeled data in target domain. The model has been evaluated using three independent public epithelium-stroma datasets by cross-dataset validations. The experimental results demonstrate that for epithelium-stroma classification, the proposed framework outperforms the state-of-the-art deep neural network model, and it also achieves better performance than other existing deep domain adaptation methods. The proposed model can be considered to be a better option for real-world applications in histopathological image analysis, since there is no longer a requirement for large-scale labeled data in each specified domain. Yue Huang 0001, Han Zheng 0004, Xinghao Ding, Gustavo K. Rohde |
IEEE J. Biomed. Health Informatics | 4 |
| 2016 | A Weighted Variational Model for Simultaneous Reflectance and Illumination EstimationabstractWe propose a weighted variational model to estimate both the reflectance and the illumination from an observed image. We show that, though it is widely adopted for ease of modeling, the log-transformed image for this task is not ideal. Based on the previous investigation of the logarithmic transformation, a new weighted variational model is proposed for better prior representation, which is imposed in the regularization terms. Different from conventional variational models, the proposed model can preserve the estimated reflectance with more details. Moreover, the proposed model can suppress noise to some extent. An alternating minimization scheme is adopted to solve the proposed model. Experimental results demonstrate the effectiveness of the proposed model with its algorithm. Compared with other variational methods, the proposed method yields comparable or better results on both subjective and objective assessments. Xueyang Fu, Delu Zeng, Yue Huang 0001, Xiao-Ping Zhang 0002, Xinghao Ding |
CVPR | 5 |
| 2016 | A fusion-based method for single backlit image enhancementabstractIn this work, a new simple but effective fusion-based strategy for enhancing single backlit image is proposed. The fundamental idea of proposed strategy is to blend different features into a single one to improve the specific quality of image. Most of existing methods are based on the modification of histogram to enhance the contrast of low light images. However, the backlit images are different from low light images, which have wide dynamic ranges of light regions, thus the existing methods cannot achieve good enhanced results of backlit images. To improve performance of enhanced results, the proposed method considers numerous features of images and processes the dark and bright regions, respectively. Furthermore, proposed method introduces weight maps to increase the visibility. Experimental results show that proposed method is superior to existing methods, which achieves better results both in visual effects and processing time. Xueyang Fu, Xiao-Ping Zhang 0002, Xinghao Ding |
ICIP | 4 |
| 2016 | Mixed noise removal based on a novel non-parametric Bayesian sparse outlier model
Peixian Zhuang, Yue Huang 0001, Delu Zeng, Xinghao Ding |
Neurocomputing | 4 |
| 2016 | Designing bag-level multiple-instance feature-weighting algorithms based on the large margin principle
Jing Chai, Hongtao Chen, Xinghao Ding |
Inf. Sci. | 4 |
| 2016 | Single image rain and snow removal via guided L0 smoothing filter
Xinghao Ding, Liqin Chen, Xianhui Zheng, Yue Huang 0001, Delu Zeng |
Multim. Tools Appl. | 1 |
| 2016 | A harmonic means pooling strategy for structural similarity index measurement in image quality assessment
Yue Huang 0001, Xin Chen 0006, Xinghao Ding |
Multim. Tools Appl. | 3 |
| 2016 | A fusion-based enhancing method for weakly illuminated images
Xueyang Fu, Delu Zeng, Yue Huang 0001, Yinghao Liao, Xinghao Ding, John W. Paisley |
Signal Process. | 5 |
| 2016 | A novel framework method for non-blind deconvolution using subspace images priors
Peixian Zhuang, Xueyang Fu, Yue Huang 0001, Delu Zeng, Xinghao Ding |
Signal Process. Image Commun. | 5 |
| 2016 | Saliency Detection With Spaces of Background-Based DistributionabstractIn this letter, an effective image saliency detection method is proposed by constructing some novel spaces to model the background and redefine the distance of the salient patches away from the background. Concretely, given the backgroundness prior, eigendecomposition is utilized to create four spaces of background-based distribution (SBD) to model the background, in which a more appropriate metric (Mahalanobis distance) is quoted to delicately measure the saliency of every image patch away from the background. After that, a coarse saliency map is obtained by integrating the four adjusted Mahalanobis distance maps, each of which is formed by the distances between all the patches and background in the corresponding SBD. To be more discriminative, the coarse saliency map is further enhanced into the posterior probability map within Bayesian perspective. Finally, the final saliency map is generated by properly refining the posterior probability map with geodesic distance. Experimental results on two usual datasets show that the proposed method is effective compared with the state-of-the-art algorithms. Lin Li 0032, Xinghao Ding, Yue Huang 0001, Delu Zeng |
IEEE Signal Process. Lett. | 3 |
| 2015 | A novel pooling strategy for Full Reference Image Quality Assessment based on harmonic meansabstractThe most perceptual Full Reference Image Quality Assessment metrics (FR-IQA) shared a common two-step model; local quality measurement, and pooling. In this letter, a novel pooling strategy based on harmonic mean is proposed to predict the final quality score in FR-IQA. In contrast to arithmetic mean, the harmonic mean tends to emphasize the contributions from the local severely distorted regions or pixels in the definition of assessment function using reciprocal transformation. It is derived from the observations that humans visual attention is mostly affected with the region having severely distorted points or regions. In addition, the relationship of subjective visual quality with the quality score against different levels of distortion in the images is described as a non-linear procedure by introducing another reciprocal transformation in harmonic mean. The proposed pooling strategy is applied to some popular FR-IQA metrics, including SSIM, GSSIM, and FSIM. The experimental results have demonstrated that the metrics with proposed pooling strategy have better performances compared to the standard versions, especially on the images with small but seriously distorted regions. The proposed pooling strategy is computationally very efficient since only one averaging operation and two reciprocal transformations are required. Xinghao Ding, Xin Chen 0006, Yue Huang 0001 |
ICASSP | 1 |
| 2015 | Fast magnetic susceptibility reconstruction using L0 norm of gradientabstractThere is a growing interest in quantifying tissue susceptibility in MRI. However, the zeros in the dipole kernel makes the calculation of the magnetic susceptibility from the measured field to be an ill-posed problem. Recently, Bayesian regularization approaches have been utilized to enable accurate quantitative susceptibility mapping(QSM), such as L2 norm gradient minimization and TV. In this work, we propose an efficient QSM method by using a sparsity promoting regularization which called L0 norm of gradient to reconstruct susceptibility map. The use of L0 norm allows us to yield high quality image and prevent penalizing salient edges. Since the L0 minimization is an NP-hard problem, a special alternating optimization strategy by introducing an auxiliary variable is adopted to solve the problem and it only takes 1-2 mins to reconstruct the whole 3D susceptibility data. Both numerical phantom simulations and human brain tests are performed to demonstrate the superior performance of the proposed method compared with previous methods. Jianzhong Lin, Congbo Cai, Delu Zeng, Xinghao Ding |
ICASSP | 5 |
| 2015 | Pan-Sharpening with a Hyper-Laplacian PenaltyabstractPan-sharpening is the task of fusing spectral information in low resolution multispectral images with spatial information in a corresponding high resolution panchromatic image. In such approaches, there is a trade-off between spectral and spatial quality, as well as computational efficiency. We present a method for pan-sharpening in which a sparsity-promoting objective function preserves both spatial and spectral content, and is efficient to optimize. Our objective incorporates the l1/2-norm in a way that can leverage recent computationally efficient methods, and l1for which the alternating direction method of multipliers can be used. Additionally, our objective penalizes image gradients to enforce high resolution fidelity, and exploits the Fourier domain forfurther computational efficiency. Visual quality metrics demonstrate that our proposed objective function can achieve higher spatial and spectral resolution than several previous well-known methods with competitive computational efficiency. Yiyong Jiang, Xinghao Ding, Delu Zeng, Yue Huang 0001, John W. Paisley |
ICCV | 2 |
| 2015 | Patch-based nonlocal dynamic MRI reconstruction with low-rank priorabstractCompressed sensing utilizes the sparsity of Magnetic resonance (MR) images to obtain accurate reconstructions from undersampled k-space data. In this paper, a novel nonlocal dynamic MRI reconstruction method with low-rank regularization is developed to exploit the spatiotemporal structural sparsity of a MRI sequence. The nonlocal prior and low rank prior are combined organically by grouping similar patches in both spatial and temporal domain. The low-rank regularization can be approximated by nuclear norm minimization solved by a singular value thresholding (SVT) method with adaptive thresholds estimation. The objective function is divided into several sub-problems that are easier to solve by alternative direction multiplier method (ADMM). Extensive experiments show that the new method outperforms commonly used classical dynamic MRI reconstruction algorithms. Liyan Sun, Jinchu Chen, Xiao-Ping Zhang 0002, Xinghao Ding |
MMSP | 4 |
| 2015 | Single-trial ERPs denoising via collaborative filtering on ERPs images
Yue Huang 0001, Xin Chen 0006, Delu Zeng, Xinghao Ding |
Neurocomputing | 6 |
| 2015 | Remote Sensing Image Enhancement Using Regularized-Histogram Equalization and DCTabstractIn this letter, an effective enhancement method for remote sensing images is introduced to improve the global contrast and the local details. The proposed method constitutes an empirical approach by using the regularized-histogram equalization (HE) and the discrete cosine transform (DCT) to improve the image quality. First, a new global contrast enhancement method by regularizing the input histogram is introduced. More specifically, this technique uses the sigmoid function and the histogram to generate a distribution function for the input image. The distribution function is then used to produce a new image with improved global contrast by adopting the standard lookup table-based HE technique. Second, the DCT coefficients of the previous contrast improved image are automatically adjusted to further enhance the local details of the image. Compared with conventional methods, the proposed method can generate enhanced remote sensing images with higher contrast and richer details without introducing saturation artifacts. Xueyang Fu, Jiye Wang, Delu Zeng, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2015 | A Probabilistic Method for Image Enhancement With Simultaneous Illumination and Reflectance EstimationabstractIn this paper, a new probabilistic method for image enhancement is presented based on a simultaneous estimation of illumination and reflectance in the linear domain. We show that the linear domain model can better represent prior information for better estimation of reflectance and illumination than the logarithmic domain. A maximum a posteriori (MAP) formulation is employed with priors of both illumination and reflectance. To estimate illumination and reflectance effectively, an alternating direction method of multipliers is adopted to solve the MAP problem. The experimental results show the satisfactory performance of the proposed method to obtain reflectance and illumination with visually pleasing enhanced results and a promising convergence rate. Compared with other testing methods, the proposed method yields comparable or better results on both subjective and objective assessments. Xueyang Fu, Yinghao Liao, Delu Zeng, Yue Huang 0001, Xiao-Ping Zhang 0002, Xinghao Ding |
IEEE Trans. Image Process. | 6 |
| 2015 | Vehicle Logo Recognition System Based on Convolutional Neural Networks With a Pretraining StrategyabstractSince a vehicle logo is the clearest indicator of a vehicle manufacturer, most vehicle manufacturer recognition (VMR) methods are based on vehicle logo recognition. Logo recognition can be still a challenge due to difficulties in precisely segmenting the vehicle logo in an image and the requirement for robustness against various imaging situations simultaneously. In this paper, a convolutional neural network (CNN) system has been proposed for VMR that removes the requirement for precise logo detection and segmentation. In addition, an efficient pretraining strategy has been introduced to reduce the high computational cost of kernel training in CNN-based systems to enable improved real-world applications. A data set containing 11 500 logo images belonging to 10 manufacturers, with 10 000 for training and 1500 for testing, is generated and employed to assess the suitability of the proposed system. An average accuracy of 99.07% is obtained, demonstrating the high classification potential and robustness against various poor imaging situations. Yue Huang 0001, Ruiwen Wu, Wei Wang 0155, Xinghao Ding |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2014 | Pan-sharpening with a Bayesian nonparametric dictionary learning modelabstractPan-sharpening, a method for constructing high resolution images from low resolution observations, has recently been explored from the perspective of compressed sensing and sparse representation theory. We present a new pan-sharpening algorithm that uses a Bayesian nonparametric dictionary learning model to give an underlying sparse representation for image reconstruction. In contrast to existing dictionary learning methods, the proposed method infers parameters such as dictionary size, patch sparsity and noise variances. In addition, our regularization includes image constraints such as a total variation penalization term and a new gradient penalization on the reconstructed PAN image. Our method does not require high resolution multiband images for dictionary learning, which are unavailable in practice, but rather the dictionary is learned directly on the reconstructed image as part of the inversion process. We present experiments on several images to validate our method and compare with several other well-known approaches. Xinghao Ding, Yiyong Jiang, Yue Huang 0001, John W. Paisley |
AISTATS | 1 |
| 2014 | A novel retinex based approach for image enhancement with illumination adjustmentabstractRetinex based algorithms have been widely used among in image enhancement. Since many retinex based algorithms remove illumination and regard the reflectance as enhancement, over-enhancement and unnaturalness are inevitable. In this paper, a novel retinex based image enhancement using illumination adjustment is proposed. Different from existing variational retinex models, a new model without the logarithmic transformation is established and can well preserve the edge. A fast alternating direction optimization method is used to solve this problem. After the decomposition of illumination and reflectance, a simple and effective post-processing method for illumination adjustment is adopted for the enhancement to make the result more natural. The proposed method can deal with many kinds of image, such as high dynamic range (HDR) images and non-uniform illumination images. Experimental results illustrate that the naturalness can be preserved while details are enhanced by the presented new approach. Xueyang Fu, Minghui LiWang, Yue Huang 0001, Xiao-Ping Zhang 0002, Xinghao Ding |
ICASSP | 6 |
| 2014 | A retinex-based enhancing approach for single underwater imageabstractSince the light is absorbed and scattered while traveling in water, color distortion, under-exposure and fuzz are three major problems of underwater imaging. In this paper, a novel retinex-based enhancing approach is proposed to enhance single underwater image. The proposed approach has mainly three steps to solve the problems mentioned above. First, a simple but effective color correction strategy is adopted to address the color distortion. Second, a variational framework for retinex is proposed to decompose the reflectance and the illumination, which represent the detail and brightness respectively, from single underwater image. An effective alternating direction optimization strategy is adopted to solve the proposed model. Third, the reflectance and the illumination are enhanced by different strategies to address the under-exposure and fuzz problem. The final enhanced image is obtained by combining use the enhanced reflectance and illumination. The enhanced result is improved by color correction, lightens dark regions, naturalness preservation, and well enhanced edges and details. Moreover, the proposed approach is a general method that can enhance other kinds of degraded image, such as sandstorm image. Xueyang Fu, Peixian Zhuang, Yue Huang 0001, Yinghao Liao, Xiao-Ping Zhang 0002, Xinghao Ding |
ICIP | 6 |
| 2014 | A compressed sensing-based pan-sharpening using joint data fidelity and blind blurring kernel estimationabstractPan-sharpening is an approach that fuse low resolution multi-spectral (LRMS) images with a high spatial detail of panchromatic (PAN) image to obtain the high resolution multispectral (HRMS) images. In this paper, we present a compressed sensing-based pan-sharpening method that include joint data fidelity and blind blurring kernel estimation. The joint data fidelity contain following three fidelity terms: (1) the LRMS images could be the decimated form of the HRMS images by convolving a blurring kernel, (2) the gradient of HRMS images in the spectrum direction could be proximity to those of the LRMS images, (3) the high frequency part of linear combination of HRMS image bands is approximate to the corresponding parts of the PAN image. Different from other methods which simply apply average blurring kernel for pan-sharpening, a blind deconvolution algorithm is introduced to estimate the blurring kernel from different satellites respectively. We also include a novel anisotropic total variation (TV) prior term to better reconstruct the image edges. The alternating direction method of multipliers (ADMM) is used to solve the proposed model efficiently. Finally, a Pléiades satellite image is employed to demonstrate that the proposed method achieve effective and efficient results simultaneously compared with other existing methods. Yiyong Jiang, Liqin Chen, Wei Wang 0155, Xinghao Ding, Yue Huang 0001 |
ICIP | 4 |
| 2014 | A fusion-based enhancing approach for single sandstorm imageabstractIn this paper, a novel image enhancing approach focuses on single sandstorm image is proposed. The degraded image has some problems, such as color distortion, low-visibility, fuzz and non-uniform luminance, due to the light is absorbed and scattered by particles in sandstorm. The proposed approach based on fusion principles aims to overcome the aforementioned limitations. First, the degraded image is color corrected by adopting a statistical strategy. Then two inputs, which represent different brightness, are derived only from the color corrected image by applying Gamma correction. Three weighted maps (sharpness, chromaticity and prominence), which contain important features to increase the quality of the degraded image, are computed from the derived inputs. Finally, the enhanced image is obtained by fusing the inputs with the weight maps. The proposed method is the first to adopt a fusion-based method for enhancing single sandstorm image. Experimental results show that enhanced results can be improved by color correction, well enhanced details and local contrast while promoted global brightness, increasing the visibility, naturalness preservation. Moreover, the proposed algorithm is mostly calculated by per-pixel operation, which is appropriate for real-time applications. Xueyang Fu, Yue Huang 0001, Delu Zeng, Xiao-Ping Zhang 0002, Xinghao Ding |
MMSP | 5 |
| 2014 | Robust mixed noise removal with non-parametric Bayesian sparse outlier modelabstractThis paper proposes a novel non-parametric Bayesian framework for solving mixed noise removal problem. In order to removing unstable effects of outlier noise such as salt-and-pepper in the training data, we decompose the observed data model into three components terms of ideal data, Gaussian noise and sparse outlier. And the proposed model employs spike-slab sparse prior to find the sparser coefficients of desired data term and outlier noise. Note that the proposed non-parametric Bayesian model can infer the noise statistics from the training data and have been robust to the mixed noise without tuning of model parameters. Experimental results demonstrate our proposed algorithm performs well with mixed noise and achieves better performance over other state-of-the-art methods. Peixian Zhuang, Wei Wang 0155, Delu Zeng, Xinghao Ding |
MMSP | 4 |
| 2014 | Multiple-instance discriminant analysis
Jing Chai, Xinghao Ding, Hongtao Chen |
Pattern Recognit. | 2 |
| 2014 | Bayesian Nonparametric Dictionary Learning for Compressed Sensing MRIabstractWe develop a Bayesian nonparametric model for reconstructing magnetic resonance images (MRIs) from highly undersampled k -space data. We perform dictionary learning as part of the image reconstruction process. To this end, we use the beta process as a nonparametric dictionary learning prior for representing an image patch as a sparse combination of dictionary elements. The size of the dictionary and patch-specific sparsity pattern are inferred from the data, in addition to other dictionary learning variables. Dictionary learning is performed directly on the compressed image, and so is tailored to the MRI being considered. In addition, we investigate a total variation penalty term in combination with the dictionary learning model, and show how the denoising property of dictionary learning removes dependence on regularization parameters in the noisy setting. We derive a stochastic optimization algorithm based on Markov chain Monte Carlo for the Bayesian model, and use the alternating direction method of multipliers for efficiently performing total variation minimization. We present empirical results on several MRI, which show that the proposed regularization framework can improve reconstruction accuracy over other methods. Yue Huang 0001, John W. Paisley, Xinghao Ding, Xueyang Fu, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 4 |
| 2013 | Compressed sensing MRI with Bayesian dictionary learningabstractWe present an inversion algorithm for magnetic resonance images (MRI) that are highly undersampled in k-space. The proposed method incorporates spatial finite differences (total variation) and patch-wise sparsity through in situ dictionary learning. We use the beta-Bernoulli process as a Bayesian prior for dictionary learning, which adaptively infers the dictionary size, the sparsity of each patch and the noise parameters. In addition, we employ an efficient numerical algorithm based on the alternating direction method of multipliers (ADMM). We present empirical results on two MR images. Xinghao Ding, John W. Paisley, Yue Huang 0001, Xianbo Chen, Xiao-Ping Zhang 0002 |
ICIP | 1 |
| 2013 | Pan-sharpening based on nonparametric Bayesian adaptive dictionary learningabstractPan-sharpening based on compressed sensing (CS) theory has been widely studied in recent years. In this paper, we present a novel CS-based pan-sharpening method based on nonparametric Bayesian adaptive dictionary learning. In contrast to existing optimization methods, the proposed method adaptively infers parameters such as dictionary size, patch sparsity and noise variances. In addition, high resolution multiband images, which are unavailable in practice, are not required to learn the dictionary anymore. An IKONOS satellite image is employed to validate the method. Both visual results and quality metrics demonstrate that proposed method is able to achieve higher spatial and spectral resolution simultaneously, compared with other well-known methods. Yue Huang 0001, John W. Paisley, Xinghao Ding, Xiao-Ping Zhang 0002 |
ICIP | 4 |
| 2013 | Single-Trial Event-Related Potentials Classification via a Discriminative Dictionary Learning Scheme
Yue Huang 0001, Xin Chen 0006, Delu Zeng, Xinghao Ding, Qingfeng Cai |
ICONIP (1) | 5 |
| 2013 | Single-Image-Based Rain and Snow Removal Using Multi-guided Filter
Xianhui Zheng, Yinghao Liao, Xueyang Fu, Xinghao Ding |
ICONIP (3) | 5 |
| 2011 | Bayesian Robust Principal Component AnalysisabstractA hierarchical Bayesian model is considered for decomposing a matrix into low-rank and sparse components, assuming the observed matrix is a superposition of the two. The matrix is assumed noisy, with unknown and possibly non-stationary noise statistics. The Bayesian framework infers an approximate representation for the noise statistics while simultaneously inferring the low-rank and sparse-outlier contributions; the model is robust to a broad range of noise levels, without having to change model hyperparameter settings. In addition, the Bayesian framework allows exploitation of additional structure in the matrix. For example, in video applications each row (or column) corresponds to a video frame, and we introduce a Markov dependency between consecutive rows in the matrix (corresponding to consecutive frames in the video). The properties of this Markov process are also inferred based on the observed matrix, while simultaneously denoising and recovering the low-rank and sparse components. We compare the Bayesian model to a state-of-the-art optimization-based implementation of robust PCA; considering several examples, we demonstrate competitive performance of the proposed model. Xinghao Ding, Lihan He, Lawrence Carin |
IEEE Trans. Image Process. | 1 |