Xiaotong Tu

dblp:207/0123 · DBLP profile ↗
← Back
46ranked-venue papers
1as first author
45since 2021 · last 2026
0000-0002-7190-2429ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 1 first-author · 32 since 2021Artificial intelligence and machine learning · 17 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Self-supervised Multiplex Consensus Mamba for General Image Fusion
abstract
Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, general image fusion needs to address a wide range of tasks while improving performance without increasing complexity. To achieve this, we propose SMC-Mamba, a Self-supervised Multiplex Consensus Mamba framework for general image fusion. Specifically, the Modality-Agnostic Feature Enhancement (MAFE) module preserves fine details through adaptive gating and enhances global representations via spatial-channel and frequency rotational scanning. The Multiplex Consensus Cross-modal Mamba (MCCM) module enables dynamic collaboration among experts, reaching a consensus to efficiently integrate complementary information from multiple modalities. The cross-modal scanning within MCCM further strengthens feature interactions across modalities, facilitating seamless integration of critical information from both sources. Additionally, we introduce a Bi-level Self-supervised Contrastive Learning Loss (BSCL), which preserves high-frequency information without increasing computational overhead while simultaneously boosting performance in downstream tasks. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) image fusion algorithms in tasks such as infrared-visible, medical, multi-focus, and multi-exposure fusion, as well as downstream visual tasks.
Yingying Wang 0005, Rongjin Zhuang, Hui Zheng 0003, Xuanhua He, Ke Cao 0001, Xiaotong Tu, Xinghao Ding
AAAI6
2026 Exploiting point-language models with dual-prompts for 3D anomaly detection
Haote Xu, Xiaolu Chen, Haodi Xu, Yue Huang 0001, Xinghao Ding, Xiaotong Tu
Expert Syst. Appl.7
2026 Accelerating Adaptive Diffusion and Uncertainty Modeling for Underwater Image Enhancement
abstract
Underwater image enhancement (UIE) aims to mitigate wavelength-dependent absorption and multi-path scattering effects, enabling the recovery of natural colors and rich details. Despite notable progress, consistently achieving high-quality enhancement in both fidelity and perceptual clarity remains a fundamental challenge. To address this, we propose the Laplacian domain Dual-Focus Enhancer (DFE), an innovative framework consisting of two stages: adaptive diffusion-accelerated low frequency enhancement (ADALE) and progressive uncertainty driven high-frequency enhancement (PUHE). Specifically, DFE applies a Laplacian transform to decouple the frequency-specific degradations in underwater images, supporting fidelity- and clarity-oriented enhancement along separate pathways. To facilitate high-fidelity restoration, ADALE incorporates an HSV guided optimization mechanism (HSV-OM) to establish a robust color and brightness calibration baseline for the low-frequency diffusion model, adaptively managing basic degradations with minimal sampling steps. Furthermore, to enhance contour and detail perception, PUHE models the uncertainty of reference textures and integrates it with feature modulation to progressively reconstruct multi-scale high-frequency structures. The multi reference underwater texture enhancement (MUTE) dataset fur ther improves image clarity. Extensive experiments demonstrate that our DFE outperforms state-of-the-art (SOTA) methods in both quantitative metrics and visual quality.
Xiuna Zeng, Jiaao Peng, Zhenqi Fu, Linyu Fan, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
IEEE Trans. Multim.5
2025 DPLUT: Unsupervised Low-light Image Enhancement with Lookup Tables and Diffusion Priors
abstract
Low-light image enhancement (LIE) aims at precisely and efficiently recovering an image degraded in poor illumination environments. Recent advanced LIE techniques are using deep neural networks, which require lots of low-normal light image pairs, network parameters, and computational resources. As a result, their practicality is limited. In this work, we devise a novel unsupervised LIE framework based on diffusion priors and lookup tables (DPLUT) to achieve efficient low-light image recovery. The proposed approach comprises two critical components: a light adjustment lookup table (LLUT) and a noise suppression lookup table (NLUT). LLUT is optimized with a set of unsupervised losses. It aims at predicting pixel-wise curve parameters for the dynamic range adjustment of a specific image. NLUT is designed to remove the amplified noise after the light brightens. As diffusion models are sensitive to noise, diffusion priors are introduced to achieve high-performance noise suppression. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in terms of visual quality and efficiency.
Yunlong Lin, Zhenqi Fu, Kairun Wen, Tian Ye 0001, Sixiang Chen, Ge Meng, Yingying Wang 0005, Chui Kong, Yue Huang 0001, Xiaotong Tu, Xinghao Ding
AAAI10
2025 Accelerated Diffusion via High-Low Frequency Decomposition for Pan-Sharpening
abstract
Pan-sharpening aims to preserve the spectral information of the multi-spectral (MS) image while leveraging the high-frequency details from the guided high-resolution panchromatic (PAN) image to enhance its spatial resolution. The key challenge is how to preserve the spectral information from the MS image and the spatial details from the PAN image as much as possible. Diffusion models have achieved favorable results in image restoration and synthesis tasks but suffer from excessive computational resource and time consumption. In this paper, we design a novel and computationally efficient diffusion-based pan-sharpening network that achieves accelerated diffusion while reducing task complexity by decoupling the high and low-frequency components of the fused image. Specifically, leveraging the information-preserving characteristic of the wavelet transformation, we introduce a Wavelet-based Low-frequency Diffusion Model (WLDM). WLDM generates the low-frequency coefficient of high-resolution MS (HRMS) image from the low-resolution MS (LRMS) image. This approach significantly reduces computational resources and complexity compared to the direct restoration of the HRMS image. Furthermore, we have devised a High-frequency Information Restoration Module (HIRM) to restore the high-frequency information in the HRMS image through the interaction of high-frequency coefficients from the PAN image in three directions. Extensive experiments on three different datasets demonstrate that our method outperforms existing approaches in both quantitative metrics, qualitative metrics, and inference efficiency.
Ge Meng, Jingjia Huang, Jingyan Tu, Yingying Wang 0005, Yunlong Lin, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
AAAI6
2025 Sp3ctralMamba: Physics-Driven Joint State Space Model for Hyperspectral Image Reconstruction
abstract
Hyperspectral image (HSI) reconstruction aims to restore the original 3D HSIs from the 2D hyperspectral snapshot compressive images (SCIs). The key to high-fidelity HSI reconstruction lies in designing refined spatial and spectral attention mechanisms, which are crucial for generating fine-grained representations of HSI based on the limited spatial and spectral information available in SCI. Recently, Mamba has demonstrated remarkable performance and efficiency in modeling spatial correlations. Its implicit attention mechanism generates three orders of magnitude more attention matrices than transformers, significantly raising the performance ceiling for HSI reconstruction. In this paper, we propose a novel joint SSM network named Sp3ctralMamba for HSI reconstruction. Sp3ctralMamba integrates frequency domain knowledge and physical priors to enhance reconstruction quality. Specifically, we first perform hierarchical decomposition of the 3D HSI embedding to mitigate the negative impact of distant bands on reconstruction. Next, we design a joint SSM block S3Mamba (S3MAB) to perform parallel scans of the embeddings from different bands. In addition to the conventional vanilla scan, S3MAB introduces a local scanning scheme to address the reconstruction challenges posed by the spatial sparsity of spectral information. Furthermore, a spiral scanning scheme in the frequency domain is incorporated to enhance the order correlation between different frequency signals. Finally, we introduce energy priors and structural priors to constrain the generation of spectral and spatial representations during the training process. Extensive experiments on both simulated and real datasets demonstrate that Sp3ctralMamba significantly elevates HSI reconstruction performance to a new level, surpassing SOTA methods in both quantitative and qualitative metrics.
Ge Meng, Jingyan Tu, Jingjia Huang, Yunlong Lin, Yingying Wang 0005, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
AAAI6
2025 CLIP-Guided Frequency-Aware Representation Learning for Generalizable Remote-Sensing Image Tampering Detection
Qingyao Wu, Xinghao Ding, Yue Huang 0001, Xiaotong Tu
ICANN (2)7
2025 Dynamic Category Queries Transformer for Generalized Few-shot Semantic Segmentation
abstract
Few-shot segmentation (FSS) tackles data scarcity using multiple priors, but its simplicity limits handling base and novel classes with limited data access. Generalized few-shot semantic segmentation (GFSS) enhances model performance for base classes with abundant data, while novel classes have limited data access, improving generalization with scarce data. Building on the design of query-based segmentation models, which decouple the mask and classification tasks for individual optimization, we here present the Dynamic Category Queries Transformer (DCQ-Former) which forms a novel approach to the GFSS. The proposed DCQ-Former first uses category suggested dynamic queries to perform mask segmentation and category classification tasks on a large amount of base class data. Considering the case when the novel classes only have access to a limited amount of training data, the queries for the novel classes are instead dynamically composed from the base classes in order to prevent the category suggested module from providing limited suggestion queries given the representativeness of the fewshot samples. Extensive experiments on COCO-20iand Pascal-5idatasets show that DCQ-Former achieves superior accuracy and generalization than current state-of-the-art methods. Our code are available at https://github.com/fallpavilion/DCQ-Former.
Kunze Huang, Jieyuan Yang, Andreas Jakobsson, Luyao Tang, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ICASSP5
2025 Efficient Dataset Distillation through Low-Rank Space Sampling
abstract
Huge amount of data is the key of the success of deep learning, however, redundant information impairs the generalization ability of the model and increases the burden of calculation. Dataset Distillation (DD) compresses the original dataset into a smaller but representative subset for high-quality data and efficient training strategies. Existing works for DD generate synthetic images by treating each image as an independent entity, thereby overlooking the common features among data. This paper proposes a dataset distillation method based on Matching Training Trajectories with Low-rank Space Sampling(MTT-LSS), which uses low-rank approximations to capture multiple low-dimensional manifold subspaces of the original data. The synthetic data is represented by basis vectors and shared dimension mappers from these subspaces, reducing the cost of generating individual data points while effectively minimizing information redundancy. The proposed method is tested on CIFAR-10, CIFAR-100, and SVHN datasets, and outperforms the baseline methods by an average of 9.9%.
Hangyang Kong, Xuxiang He, Xiaotong Tu, Xinghao Ding
ICASSP4
2025 PANDA: Patch-Aware Graph Network with Dual Alignment for Time Series Forecasting
abstract
Multivariate time series (MTS) forecasting aims to predict future patterns by extracting features from multivariate history. Predominant methods face challenges in learning spatial dependencies while capturing long-term trends and local details, leading to suboptimal performance in MTS forecasting. To address this problem, we propose a Patch-Aware graph Network with Dual Alignment (PANDA) of features in both time and frequency domain. Specifically, we utilize a patch-aware graph neural network to extract spatio-temporal dependencies. Additionally, dual alignment is employed for multi-perspective optimization. By performing multi-view joint optimization, our model reduces feature redundancy and improves the extraction of structured spatio-temporal patterns. Extensive experiments demonstrate that PANDA achieves superior forecasting accuracy in both long- and short-time series forecasting. Code is available at this repository: https://github.com/lichen0620/PANDA.
Saqlain Abbas, Chenyu Ma, Yinhao Liu, Xiaotong Tu
ICASSP6
2025 Efficient Infrared Image Super-Resolution Reconstruction via Guided Filter Coefficients Estimation with Parallax Attention Mechanism
abstract
Due to the spectral range mismatch between the images, building an efficient infrared (IR) image super-resolution algorithm suitable for embedded devices remains a significant challenge. Given that visible images possess more abundant high-frequency information compared to infrared images, we utilize the visible light to guide infrared image super-resolution reconstruction. Specifically, we transfer the reconstruction task to a guided filter learning process, whose coefficients are estimated by joint learning of visible and infrared image to complete the reconstruction through homologous constraints. In order to efficiently predict guided filter coefficients, we design a lightweight network which incorporates reparameterized differential convolution blocks and a feature fusion strategy. Striving to enhance the fusion strategy performance, we utilize parallax attention mechanism to solve the non-pixel registration problem between infrared and visible images. Extensive experiments on two challenging IR image datasets show that our method performs SOTA in terms of PSNR, SSIM and LPIPS as compared to current state-of-the-art approaches while showing its effectiveness and practicality in the edge platform of RK3588.
Qingyao Wu, Bosheng Chen, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ICASSP4
2025 Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
Luyao Tang, Kunze Huang, Chaoqi Chen, Chenxin Li, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ICCV6
2025 Self-supervised Sound Source Localization for UAVs Using GCC-PHAT in Low SNR Environments
Shengbin Ma, Saqlain Abbas, Xinghao Ding, Xiaotong Tu
ICIC (12)5
2025 A Single-Channel Drone Noise Reduction Algorithm Based on Speech Harmonic Features
Shengbin Ma, Saqlain Abbas, Xinghao Ding, Xiaotong Tu
ICIC (9)5
2025 Feature Reconstruction via Reverse Distillation for Multi-class Anomaly Detection
Haodi Xu, Zheyuan Cai, Xiaotong Tu
ICIC (16)4
2025 SOMA: A semantic-guided Order-aware Mamba Architecture for multivariate time series forecasting
Jinkai Zhang, Yingying Wang 0005, Shengbin Ma, Xinghao Ding, Xiaotong Tu
Adv. Eng. Informatics5
2025 Spatial-frequency dual-domain Kolmogorov-Arnold networks for multimodal medical image fusion
Lewu Lin, Jiaxin Xie, Yingying Wang 0005, Jialing Huang, Rongjin Zhuang, Xiaotong Tu, Xinghao Ding, Na Shen
Neurocomputing6
2025 Open world out-of-distribution generalization via dream open and sustain close
Kunze Huang, Luyao Tang, Jieyuan Yang, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
Knowl. Based Syst.5
2025 Vision-Language Model Priors-Driven State Space Model for Infrared-Visible Image Fusion
abstract
Infrared and visible image fusion (IVIF) aims to effectively integrate complementary information from both infrared and visible modalities, enabling a more comprehensive understanding of the scene and improving downstream semantic tasks. Recent advancements in Mamba have shown remarkable performance in image fusion, owing to its linear complexity and global receptive fields. However, leveraging Vision-Language Model (VLM) priors to drive Mamba for modality-specific feature extraction and using them as constraints to enhance fusion results has not been fully explored. To address this gap, we introduce VLMPD-Mamba, a Vision-Language Model Priors-Driven Mamba framework for IVIF. Initially, we employ the VLM to adaptively generate modality-specific textual descriptions, which enhance image quality and highlight critical target information. Next, we present Text-Controlled Mamba (TCM), which integrates textual priors from the VLM to facilitate effective modality-specific feature extraction. Furthermore, we design the Cross-modality Fusion Mamba (CFM) to fuse features from different modalities, utilizing VLM priors as constraints to enhance fusion outcomes while preserving salient targets with rich details. In addition, to promote effective cross modality feature interactions, we introduce a novel bi-modal interaction scanning strategy within the CFM. Extensive experiments on various datasets for IVIF, as well as downstream visual tasks, demonstrate the superiority of our approach over state-of-the-art (SOTA) image fusion algorithms.
Rongjin Zhuang, Yingying Wang 0005, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
IEEE Signal Process. Lett.3
2024 Implicit Foreground-Guided Network for Anomaly Detection and Localization
abstract
Anomaly detection plays an essential role in large-scale industrial manufacturing. However, reconstruction-based anomaly detection methods, as one of the mainstream methods, are prone to incorrectly detecting background noise as anomalous regions. Therefore, inspired by multi-task learning, we propose an Implicit Foreground-guided Network (IFgNet), which consists of a Multi-Task Attention Shared (MTAS) sub-network and a discriminative sub-network. Specifically, the MTAS sub-network implements the foreground detection and reconstruction tasks within the shared network, while the discriminative sub-network performs the final anomaly detection. In the MTAS sub-network, multiple task-specific attention blocks are applied to learn task-specific features while allowing features to be shared between different tasks. Consequently, the features that contain both semantic and edge structure information are learned through the foreground detection task, which also facilitates the reconstruction task. Furthermore, the outputs of foreground detection can be utilized to refine the anomaly detection results. In this way, IFgNet effectively mitigates the influence of background noise and achieves competitive performance on the VisA and BTAD datasets with existing methods.
Xiaolu Chen, Haote Xu, Chenghao Deng, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ICASSP4
2024 S3AHI: Source-Free Domain Adaptive Small Object Detection with Slicing Aided Hyper Inference
abstract
The time-consuming and laborious annotation of small objects has resulted in a relative scarcity of datasets specifically designed for small objects. Additionally, variations in data acquisition devices and application scenarios often cause a domain shift between source-trained data and target data. Unsupervised Domain Adaptation (UDA) is extensively applied to alleviate the domain shift between two domains based on the assumption that source data is accessible during the adaptation process. However, source data may be unavailable in some scenarios due to data privacy or data transmission issues. In this paper, we propose a Source-free domain adaptive framework for Small object detection with Slicing Aided Hyper Inference (S3AHI) in the test-time training stage. Without access to source data, Source-Free Domain Adaptation (SFDA) only employs unlabeled target data to adapt a source-trained model to the target domain during the test phase. SAHI provides a generic and effective solution to detect small objects and can be seamlessly integrated into nearly any pipeline. Motivated by contrastive learning, we build instance-level correlation graphs with the semantic features of proposals and learn high-quality pseudolabels. The S3AHI distills target domain knowledge to the sourcetrained model under the mean-teacher framework. Extensive experiments reveal that our approach outperforms existing SFDA and UDA methods significantly.
Haizhou Ding, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
IJCNN3
2024 An Adaptive Spatio-Temporal Graph Structure Learning Model for Lithium-Ion Battery Pack State of Health Estimation
abstract
Current research on lithium-ion battery state of health (SOH) estimation predominantly focuses on a single battery, not on an entire battery pack, which makes these methods inadequate for describing the SOH of energy systems that work in real-life situations. Furthermore, the current SOH estimation methods using graph neural networks (GNNs) inherently adopt suboptimal graph construction approaches, which makes them fail to accurately extract the most pertinent spatial dependencies, resulting in a reduction in prediction accuracy. In this work, we propose a pioneering GNN-based model to predict battery pack SOH. Specifically, an optimal graph structure, learned by the introduced optimal graph extractor, is used to capture the feature dependencies that are most appropriate for the downstream SOH prediction task. Additionally, a Graph Attention network (GAT) and a Gated Recurrent Unit (GRU) are employed to learn the spatial and temporal features, respectively. Lastly, we introduce an effective and simple spatio-temporal feature fusion module to ensure that the extracted features are fully utilized, and the fused features are then used for prediction. Experimental results on the NASA and CALCE datasets validate the superiority of the proposed approach over existing state-of-the-art methods for battery SOH estimation.
Canxing Lai, Xiaotong Tu, Andreas Jakobsson, Xinghao Ding, Yue Huang 0001
IJCNN2
2024 Class Incremental Aerial Scene Recognition Under Long-Tailed Distribution
abstract
Deep learning excels in aerial scene recognition (ASR) but struggles with learning from sequential data due to catastrophic forgetting. Class incremental learning (CIL) can address this but often assumes a balanced data distribution. Real-world aerial scenes exhibit long-tailed properties, the undesirable bias toward the head classes as well as overfitting for the tail classes aggravate the challenge of class incremental ASR. So here, we introduce a two-stage framework for class incremental ASR under long-tailed distribution. First, the encoder is trained by conventional CIL method. Next, keep the encoder fixed and then distribution transfer via inter-class similarity is used to generate a sufficient number of features for tail classes, thus a more balanced classifier can be trained. Our framework can be easily integrated into existing CIL methods. Experiments on public datasets demonstrated the superior performance of our proposed framework.
Haizhou Ding, Xiaotong Tu, Lexing Huang, Yue Huang 0001, Xinghao Ding
IJCNN3
2024 SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly Detection
abstract
Recently, large pre-trained vision-language models, such as CLIP, have demonstrated significant potential in zero-/few-shot anomaly detection tasks. However, existing methods not only rely on expert knowledge to manually craft extensive text prompts but also suffer from a misalignment of high-level language features with fine-level vision features in anomaly segmentation tasks. In this paper, we propose a method, named SimCLIP, which focuses on refining the aforementioned misalignment problem through bidirectional adaptation of both Multi-Hierarchy Vision Adapter (MHVA) and Implicit Prompt Tuning (IPT). In this way, our approach requires only a simple binary prompt to efficiently accomplish anomaly classification and segmentation tasks in zero-shot scenarios. Furthermore, we introduce its few-shot extension, SimCLIP+, integrating the relational information among vision embeddings and skillfully merging the cross-modal synergy information between vision and language to address downstream anomaly detection tasks. Extensive experiments on two challenging datasets prove the more remarkable generalization capacity of our method compared to the current SOTA approaches. Our code is available at https://github.com/CH-ORGI/SimCLIP.
Chenghao Deng, Haote Xu, Xiaolu Chen, Haodi Xu, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ACM Multimedia5
2024 P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images
abstract
Generating diverse plausible outputs from a single input is crucial for addressing visual ambiguities, exemplified in medical imaging where experts may provide varying semantic segmentation annotations for the same image.Existing methods handles ambiguous segmentation relying on probabilistic modeling and extensive multi-output annotated data while often struggles with limited ambiguously labeled datasets common in real-world applications.To surmount the challenge, we propose P²SAM, a novel framework that leverages the Segment Anything Model (SAM)'s prior knowledge for ambiguous object segmentation. By transforming SAM's sensitivity to prompts into an advantage, we introduce a prior probabilistic space for prompts.Experimental results show that P²SAM significantly enhances medical segmentation precision and diversity using minimal ambiguously annotated samples. Benchmarking against state-of-the-art methods demonstrates superior performance with just 5.5% of the training data (+12% Dmax). This approach marks a significant advancement towards deploying probabilistic models in data-limited real-world scenarios.
Yuzhi Huang, Chenxin Li, Zixu Lin, Hengyu Liu 0007, Haote Xu, Yifan Liu 0010, Yue Huang 0001, Xinghao Ding, Xiaotong Tu, Yixuan Yuan
ACM Multimedia9
2024 Efficient Perceiving Local Details via Adaptive Spatial-Frequency Information Integration for Multi-focus Image Fusion
abstract
Multi-focus image fusion (MFIF) aims to combine multiple images with different focused regions into a single all-in-focus image. Existing unsupervised deep learning-based methods only fuse structural information of images in the spatial domain, neglecting potential solutions from the frequency domain exploration. In this paper, we make the first attempt to integrate spatial-frequency information to achieve high-quality MFIF. We propose a novel unsupervised spatial-frequency interaction MFIF network named SFIMFN, which consists of three key components: Adaptive Frequency Domain Information Interaction Module (AFIM), Ret-Attention-Based Spatial Information Extraction Module (RASEM), and Invertible Dual-domain Feature Fusion Module (IDFM). Specifically, in AFIM, we interactively explore global contextual information by combining the amplitude and phase information of multiple images separately. In RASEM, we design a customized transformer to encourage the network to capture important local high-frequency information by redesigning the self-attention mechanism with a bidirectional, two-dimensional form of explicit decay. Finally, we employ IDFM to fuse spatial-frequency information without information loss to generate the desired all-in-focus image. Extensive experiments on different datasets demonstrate that our method significantly outperforms state-of-the-art unsupervised methods in terms of qualitative and quantitative metrics as well as the generalization ability.
Jingjia Huang, Jingyan Tu, Ge Meng, Yingying Wang 0005, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ACM Multimedia6
2024 Source-free cross-domain fault diagnosis of rotating machinery using the Siamese framework
Chenyu Ma, Xiaotong Tu, Guanxing Zhou, Yue Huang 0001, Xinghao Ding
Knowl. Based Syst.2
2024 Learning to sound imaging by a model-based interpretable network
Xiaotong Tu, Saqlain Abbas, Hao Liang 0011, Yue Huang 0001, Xinghao Ding
Signal Process.2
2023 Self-Supervised Image Denoising Using Implicit Deep Denoiser Prior
abstract
We devise a new regularization for denoising with self-supervised learning. The regularization uses a deep image prior learned by the network, rather than a traditional predefined prior. Specifically, we treat the output of the network as a ``prior'' that we again denoise after ``re-noising.'' The network is updated to minimize the discrepancy between the twice-denoised image and its prior. We demonstrate that this regularization enables the network to learn to denoise even if it has not seen any clean images. The effectiveness of our method is based on the fact that CNNs naturally tend to capture low-level image statistics. Since our method utilizes the image prior implicitly captured by the deep denoising CNN to guide denoising, we refer to this training strategy as an Implicit Deep Denoiser Prior (IDDP). IDDP can be seen as a mixture of learning-based methods and traditional model-based denoising methods, in which regularization is adaptively formulated using the output of the network. We apply IDDP to various denoising tasks using only observed corrupted data and show that it achieves better denoising results than other self-supervised denoising methods.
Huangxing Lin, Yihong Zhuang, Xinghao Ding, Delu Zeng, Yue Huang 0001, Xiaotong Tu, John W. Paisley
AAAI6
2023 Learning a Simple Low-Light Image Enhancer from Paired Low-Light Instances
abstract
Low-light Image Enhancement (LIE) aims at improving contrast and restoring details for images captured in lowlight conditions. Most of the previous LIE algorithms adjust illumination using a single input image with several handcrafted priors. Those solutions, however, often fail in revealing image details due to the limited information in a single image and the poor adaptability of handcrafted priors. To this end, we propose PairLIE, an unsupervised approach that learns adaptive priors from low-light image pairs. First, the network is expected to generate the same clean images as the two inputs share the same image content. To achieve this, we impose the network with the Retinex theory and make the two reflectance components consistent. Second, to assist the Retinex decomposition, we propose to remove inappropriate features in the raw image with a simple self-supervised mechanism. Extensive experiments on public datasets show that the proposed PairLIE achieves comparable performance against the state-of-the-art approaches with a simpler network and fewer handcrafted priors. Code is available at: https://github.com/zhenqifu/PairLIE.
Zhenqi Fu, Xiaotong Tu, Yue Huang 0001, Xinghao Ding, Kai-Kuang Ma
CVPR3
2023 Hint-Dynamic Knowledge Distillation
abstract
Knowledge Distillation (KD) transfers the knowledge from a high-capacity teacher model to promote a smaller student model. Existing efforts guide the distillation by matching their prediction logits, feature embedding, etc., while leaving how to efficiently utilize them in junction less explored. In this paper, we propose Hint-dynamic Knowledge Distillation, dubbed HKD, which excavates the knowledge from the teacher’s hints in a dynamic scheme. The guidance effect from the knowledge hints usually varies in different instances and learning stages, which motivates us to customize a specific hint-learning manner for each instance adaptively. Specifically, a meta-weight network is introduced to generate the instance-wise weight coefficients about knowledge hints in the perception of the dynamical learning progress of the student model. We further present a weight ensembling strategy to eliminate the potential bias of coefficient estimation by exploiting the historical statics. Experiments on standard benchmarks of CIFAR-100 and Tiny-ImageNet manifest that the proposed HKD well boost the effect of knowledge distillation tasks.
Chenxin Li, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ICASSP3
2023 Underwater Image Enhancement and Super-Resolution Using Implicit Neural Networks
abstract
Underwater images are often notably degraded by light scattering and absorption. To improve image quality and object details, we present a novel unsupervised underwater image enhancement and super-resolution method using implicit neural networks. Concretely, taking low-resolution coordinates as the inputs, we first leverage Fourier feature mapping to encode the coordinates. Then, three implicit neural networks are applied to estimate each component (i.e., the global background light, the transmission map, and the scene radiance) of the underwater formation model. Those components are further used to reconstruct the raw underwater image in a self-supervised fashion. In the inference stage, high-resolution coordinates are employed to predict a high-quality and high- resolution underwater image. Extensive experiments show that our method achieves a favorable performance in terms of both super-resolution and quality enhancement as compared with current approaches.
Xueye Chu, Zhenqi Fu, Shaocong Yu, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
ICIP4
2023 AnoCSR-A Convolutional Sparse Reconstructive Noise-Robust Framework for Industrial Anomaly Detection
Xiaotong Tu, Yue Huang 0001, Xinghao Ding
PRCV (5)2
2023 Adaptive nonlinear group delay mode estimation
abstract
The decomposition of non-stationary signals remains a challenge in a wide variety of fields. Especially, the impulse or cross-mode signals are difficult to be reconstructed by recent methods due to their transient characteristic. Moreover, most methods rely heavily on the user-defined settings of the regularized parameter for the convex optimization algorithm. In this work, an adaptive nonlinear group delay mode estimation (ANGDME) algorithm is proposed by exploiting the sparsity of signals formulated as the nonlinear group delay model. The ANGDME introduces a complex Bayesian compressive sensing (CBCS) framework to process the sparse reconstruction. Then, a hierarchical Laplace scale mixture (LSM) prior is utilized to model dependencies among coefficients and provides superior probabilistic predictions. Furthermore, the estimator of amplitudes and group delays (GDs) of signals are obtained from the posterior distribution by Bayesian inference instead of point estimation. Finally, the proposed method updates the dictionary matrix in a traditional data-driven manner, resulting in a high-resolution time-frequency representation. Both simulated and experimental results illustrate the adaptability and effectiveness of the proposed method.
Yijin Mao, Xiaotong Tu, Saqlain Abbas, Hao Liang 0011, Yue Huang 0001, Xinghao Ding
Signal Process.2
2023 Enhanced features in image manipulation detection
Chuchu He, Yunshu Chen, Yue Huang 0001, Xiaotong Tu, Xinghao Ding
Signal Process. Image Commun.5
2023 Unsupervised Video-Based Action Recognition With Imagining Motion and Perceiving Appearance
abstract
Video-based action recognition is a challenging task, which demands carefully considering the temporal property of videos in addition to the appearance attributes. Particularly, the temporal domain of raw videos usually contains significantly more redundant or irrelevant information than still images. For that, this paper proposes an unsupervised video-based action recognition approach with imagining motion and perceiving appearance, called IMPA, by comprehensively learning the spatio-temporal characteristics inherited in videos, with a particular emphasis on the moving object for action recognition. Specifically, a self-supervised Motion Extracting Block (MEB) is designed to extract the principal motion features by focusing on the large movement of the moving object, based on the observation that humans can infer complete motion trajectories from partial moving objects. To further take the indispensable appearance attribute in videos into account, an unsupervised Appearance Learning Block (ALB) is developed to perceive the static appearance, thus in combination with the MEB to recognize actions. Extensive validation experiments and ablation studies on multiple datasets demonstrate that our proposed IMPA approach obtains superior performance and surpasses other classical and state-of-the-art unsupervised action recognition methods.
Wei Lin 0021, Yihong Zhuang, Xinghao Ding, Xiaotong Tu, Yue Huang 0001, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.5
2022 EffiSeaNet: Pioneering Lightweight Network for Underwater Salient Object Detection
Qingyao Wu, Zhenqi Fu, Chenyu Ma, Xiaotong Tu, Xinghao Ding
ACCV (4)5
2022 Unsupervised and Untrained Underwater Image Restoration Based on Physical Image Formation Model
abstract
Underwater images suffer from degradation caused by light scattering and absorption. Training a deep neural network to restore underwater images is challenging due to the labor-intensive data collection and the lack of paired data. To this end, we propose an unsupervised and untrained underwater image restoration method based on the layer disentanglement and the underwater image formation model. Specifically, our network disentangles an underwater image into four components, i.e., the scene radiance, the direct transmission map, the backscatter transmission map, and the global background light, which are further combined to reconstruct the underwater image in a self-supervised manner. Our method can avoid using paired training data and large-scale datasets, benefiting from the unsupervised and untrained characteristics. Extensive experiments demonstrated that our method obtains promising performance compared with six methods on three real-world underwater image databases.
Shu Chai, Zhenqi Fu, Yue Huang 0001, Xiaotong Tu, Xinghao Ding
ICASSP4
2022 A Two-Stage Contrastive Learning Framework For Imbalanced Aerial Scene Recognition
abstract
In real-world scenarios, aerial image datasets are generally class imbalanced, where the majority classes have rich samples, while the minority classes only have a few samples. Such class imbalanced datasets bring great challenges to aerial scene recognition. In this paper, we explore a novel two-stage contrastive learning framework, which aims to take care of representation learning and classifier learning, thereby boosting aerial scene recognition. Specifically, in the representation learning stage, we design a data augmentation policy to improve the potential of contrastive learning according to the characteristics of aerial images. And we employ supervised contrastive learning to learn the association between aerial images of the same scene. In the classification learning stage, we fix the encoder to maintain good representation and use the re-balancing strategy to train a less biased classifier. A variety of experimental results on the imbalanced aerial image datasets show the advantages of the proposed two-stage contrastive learning framework for the imbalanced aerial scene recognition.
Lexing Huang, Senlin Cai, Yihong Zhuang, Changxing Jing, Yue Huang 0001, Xiaotong Tu, Xinghao Ding
ICASSP6
2022 Adaptive Variational Nonlinear Chirp Mode Decomposition
abstract
Variational nonlinear chirp mode decomposition (VNCMD) is a recently introduced method for nonlinear chirp signal decomposition that has aroused notable attention in various fields. One limiting aspect of the method is that its performance relies heavily on the setting of the bandwidth parameter. To overcome this problem, we here propose a Bayesian implementation of the VNCMD, which can adaptively estimate the instantaneous amplitudes and frequencies of the nonlinear chirp signals, and then learn the active dictionary in a data-driven manner, thereby enabling a high-resolution time-frequency representation. Numerical example of both simulated and measured data illustrate the resulting improvement performance of the proposed method.
Hao Liang 0011, Xinghao Ding, Andreas Jakobsson, Xiaotong Tu, Yue Huang 0001
ICASSP4
2022 A Self-Supervised Method for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion (IVIF) plays important roles in many applications. Since there is no ground-truth, the fusion performance measurement is a difficult but important problem for the task. Previous unsupervised deep learning based fusion methods depend on a hand-crafted loss function to define the distance between the fused image and two types of source images, which still cannot well preserve the vital information in the fused images. To address these issues, we propose an image fusion performance measurement between the fused image and the decomposition of the fused image. A novel self-supervised network for infrared and visible image fusion is designed to preserve the vital information of source images by narrowing the distance between the source images and the decomposed ones. Extensive experimental results demonstrate that our proposed measurement has the ability in improving the performance of backbone network in both subjective and objective evaluations.
Xiaopeng Lin, Guanxing Zhou, Weihong Zeng, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
ICIP4
2022 A Simple Siamese Framework for Vibration Signal Representations
abstract
Siamese networks are widely used in various contrastive learning methods for recognition tasks, with few labeled data and abundant unlabeled data. In the field of fault diagnosis, it is universal to face the problem that large collections of common fault data and few catastrophic fault samples result in the imbalanced distribution of fault data collection. In this paper, a simple Siamese framework is proposed to learn meaningful signal representations using the differently augmented views of the signals only in the time domain. The industrial fault diagnosis including class balanced and imbalanced motor fault diagnosis is performed to verify the validity of the signal representations. The results demonstrate that the proposed method can significantly balance the representations of both the major and minor classes, which proves the capability of the Siamese framework for class imbalanced classification.
Guanxing Zhou, Yihong Zhuang, Xinghao Ding, Yue Huang 0001, Saqlain Abbas, Xiaotong Tu
ICIP6
2022 A Hybrid Framework Based on Classifier Calibration for Imbalanced Aerial Scene Recognition
Yihong Zhuang, Changxing Jing, Senlin Cai, Lexing Huang, Yue Huang 0001, Xiaotong Tu, Xinghao Ding
ICONIP (3)6
2022 High-resolution source localization exploiting the sparsity of the beamforming map
Xinghao Ding, Hao Liang 0011, Andreas Jakobsson, Xiaotong Tu, Yue Huang 0001
Signal Process.4
2021 Estimating nonlinear chirp modes exploiting sparsity
Xiaotong Tu, Johan Sward, Andreas Jakobsson, Fucai Li
Signal Process.1
2020 Gaussian-modulated linear group delay model: Application to second-order time-reassigned synchrosqueezing transform
Zhoujie He, Xiaotong Tu, Wenjie Bao, Yue Hu 0007, Fucai Li
Signal Process.2