Rui Huang 0006

dblp:56/2875-6 · DBLP profile ↗
← Back
31ranked-venue papers
21as first author
21since 2021 · last 2025
0000-0002-3343-066XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Gaussian Difference: Find Any Change Instance in 3D Scenes
abstract
Instance-level change detection in 3D scenes presents significant challenges, particularly under uncontrolled conditions without labeled image pairs, varying camera poses, or restricted lighting. This paper addresses this challenge by developing a novel approach to detect changes in real-world scenarios. Leveraging 4D Gaussians to embed multiple images into 3D Gaussian distributions, our method enables the rendering of two coherent image sequences. By segmenting each image and assigning a unique identifier to each instance, we can efficiently identify changed instances through ID comparison. Additionally, we utilize change maps and classification encodings to categorize the 4D Gaussians as changed or unchanged, allowing for the rendering of a comprehensive change map from any view direction. Through extensive experiments on various instance-level change detection datasets, our method demonstrates significant improvements in detection accuracy over state-of-the-art methods like C-NERF and CYWS-3D, particularly in scenarios with large lighting variations.
Binbin Jiang, Rui Huang 0006, Qingyi Zhao, Yuxiang Zhang 0003
ICASSP2
2025 SFFCE-CD: Spatial And Frequency Feature Cross Enhancement For Change Detection
abstract
Most existing change detection (CD) methods focus on spatial domain modeling, while ignore the rich information of frequency domain. In this paper, we enhance the feature representative ability with spatial and frequency feature cross enhancement. Specifically, we propose a Change Feature Extract Module (CFEM) to obtain high-quality change features from the bi-temporal images. These features are then merged together and refined through three parallel branches: the Local branch uses multiscale max-pooling operations to generate multiscale feature; the Wavelet Transform Decomposer (WTD) branch decomposes feature into low-frequency and high-frequency signals with Haar wavelet transform; the Global branch adopts Mamba to capture the long-range dependencies. To bridge the semantic gap between frequency and spatial features, we design Dual-Representation Aggregation Module (DRAM) to promote the combination of features from different representation domains in a flow-based manner and dual cross attention. Extensive experiments demonstrate our method outperforms 11 SOTA CD methods on three remote sensing CD datasets.
Jiali Hu, Binbin Jiang, Qingyi Zhao, Longxi Feng, Rui Huang 0006
ICASSP6
2025 GTPC-SSCD: Gate-guided Two-level Perturbation Consistency-based Semi-Supervised Change Detection
abstract
Semi-supervised change detection (SSCD) utilizes partially labeled data and abundant unlabeled data to detect differences between multi-temporal remote sensing images. The mainstream SSCD methods based on consistency regularization have limitations. They perform perturbations mainly at a single level, restricting the utilization of unlabeled data and failing to fully tap its potential. In this paper, we introduce a novel Gate-guided Two-level Perturbation Consistency regularization-based SSCD method (GTPC-SSCD). It simultaneously maintains strong-to-weak consistency at the image level and perturbation consistency at the feature level, enhancing the utilization efficiency of unlabeled data. Moreover, we develop a hardness analysis-based gating mechanism to assess the training complexity of different samples and determine the necessity of performing feature perturbations for each sample. Through this differential treatment, the network can explore the potential of unlabeled data more efficiently. Extensive experiments conducted on six benchmark CD datasets demonstrate the superiority of our GTPC-SSCD over seven state-of-the-art methods.
Qi'ao Xu, Zongyu Guo, Rui Huang 0006, Yuxiang Zhang 0003
ICME4
2025 Mining and Integrating Spatiotemporal Continuity Information for Remote Sensing Change Detection
abstract
Existing remote sensing (RS) change detection (CD) methods typically extract multi-scale features from bi-temporal images, followed by feature fusion and decoding. However, they often overlook spatiotemporal continuity and external correlations, leading to insufficient ability to mine spatiotemporal context. In this paper, we propose a Continuity Information Mining and Integrating Network (CIMINet), which captures and integrates spatiotemporal continuity to reduce pseudo-changes caused by scale differences and limited temporal cues. Spatially, bi-temporal images are processed by Adjacent-Scale Feature Fusion Module (ASFFM) and Spatial Change Extraction Module (SCEM) to extract continuous spatial features and spatial differences. Temporally, pseudo-video frames generated via interpolation help construct temporal continuity, which is further exploited by Temporal Change Extraction Module (TCEM). Finally, we jointly fuse spatial and temporal differences by Spatiotemporal Fusion Module (STFM), highlighting the change regions. Various experimental results indicate that CIMINet outperforms 12 SOTA CD methods on four RS datasets. The code can be found at https://github.com/Aosion518/CIMINet.
Rui Huang 0006, Haojie Tao, Yunan Jia
IEEE Geosci. Remote. Sens. Lett.1
2025 Frequency-Enhanced Mamba for Remote Sensing Change Detection
abstract
Remote sensing (RS) change detection (CD) is a critical task in monitoring surface dynamics. Recently, Mamba-based methods have shown promising performance and are quickly adopted in change detection. However, when addressing the task of CD in complex scenarios, existing methods have limitations in capturing features of minor and texture changes due to the lack of frequency information. To address these challenges, we propose a frequency-enhanced Mamba for RSCD (FEMCD). First, we design a difference-guided state-space model (DGSSM) to extract change-related features. DGSSM takes the features of bitemporal images as input and uses absolute-difference features to guide the network to focus on change regions. Second, we develop a DCT-aided Mamba decoder (DCTMD) for feature decoding and refinement. DCTMD uses the omnidirectional selective scan module (OSSM) to refine the change-related features and DCT to capture minor change details. Finally, we use a simple classifier to generate the final change map. We have conducted extensive experiments on five RSCD datasets, comparing FEMCD with 11 SOTA change detectors. The experimental results show that our proposed FEMCD method outperforms other compared methods. The code can be found at:https://github.com/JYN712/FEMCD.
Yunan Jia, Jiali Hu, Rui Huang 0006
IEEE Geosci. Remote. Sens. Lett.5
2025 Consistent Bokeh for Multi-View Images With 3D Gaussian Splatting
abstract
Bokeh refocuses the desired regions and generates out-of-focus blur in the remaining regions, which has been well studied for single-image. However, when processing multi-view images of a given scene, the existing Bokeh methods might focus on different objects due to inconsistencies in salient object detection across different views. Additionally, the salient objects with larger depth might be blurred because they are not on the in-focus plane. In this letter, we propose a framework to generate consistent Bokeh across multi-view images. We utilize 3D Gaussian Splatting to render a group of images with a small viewpoint span. To guarantee that the salient objects are identical, we propose aConsistent Saliency Map Generation(CSMG) method with mask tracking. We also propose aDepth Value Reassignment(DVR) method to endow the salient objects with new depth values. With the consistent salient objects and reassigned depth values, we adopt Dr.Bokeh to generate consistent Bokeh effects across multi-view images. Various experiments conducted on the DoF-NeRF dataset demonstrate that our proposed framework outperforms four state-of-the-art bokeh methods on four no-reference image quality evaluation metrics. Code can be found athttps://github.com/jiutaojushi/CBMI.
Rui Huang 0006, Haojie Tao, Liangying Tang, Jingcheng Zeng
IEEE Signal Process. Lett.1
2025 C-NeRF: Representing Scene Changes as Directional Consistency Difference-Based NeRF
abstract
In this work, we aim to detect the changes caused by object variations in a scene represented by the neural radiance fields (NeRFs). Given an arbitrary view and two sets of scene images captured at different timestamps, we can predict scene changes in that view, which has significant potential applications in scene monitoring and measuring. We conducted preliminary studies and found that such an exciting task cannot be easily achieved by utilizing existing NeRFs and 2D change detection (CD) methods with many false or missing detections. The main reason is that the 2D CD is based on the pixel appearance difference between spatial-aligned image pairs and neglects the stereo information in the NeRF. To address the limitations, we propose the C-NeRF to represent scene changes as directional consistency difference-based NeRF, which mainly contains three modules. We first build two aligned NeRFs from pre-change and post-change scenes. Then, we identify the change points based on the direction-consistent constraint; that is, real change points have similar change representations across view directions, but fake change points do not. Finally, we design the change map rendering process based on the built NeRFs and can generate the change map of an arbitrarily specified view direction. To validate the effectiveness, we build a new dataset containing ten scenes covering diverse scenarios with different changing objects. Our approach surpasses state-of-the-art CD methods and NeRF-based methods by a significant margin.
Rui Huang 0006, Haojie Tao, Binbin Jiang, Qingyi Zhao, Liang Wang 0001, Qing Guo 0005
IEEE Trans. Image Process.1
2024 Fine-grained classification for aero-engine borescope images based on the fusion of local and global features
abstract
Borescope image classification is an essential preprocessing step for damage detection in aero-engines, which plays a crucial role in setting the detection threshold for later stages. However, traditional image classification methods such as ResNet or fine-grained image classification methods like FBSD perform poorly in classifying borescope images. To address this issue, we propose a novel and effective fine-grained classification method for aero-engine borescope images based on the fusion of local and global features. Our approach leverages a region proposal network to identify key local regions and combines the local features with global features for better classification. This enables us to distinguish between regions with high similarity and small inter-class differences. Additionally, we have created an aero-engine borescope image dataset with 5,158 images from 14 aero-engine components. Our method achieves a high classification accuracy of 98.31% on this dataset, outperforming other image classification and fine-grained image classification methods. Our implementation code is available at https://github.com/LonelyProceduralApe/LG_fusion
Rui Huang 0006, Jingcheng Zeng, Xuyi Cheng, Jieda Wei
CSCWD1
2024 PAPnet: A Plug-and-play Virus Network for Backdoor Attack
abstract
Most existing backdoor attacks focus on designing various trigger injection methods and fine-tuning victim networks, which are difficult to deploy in real-world applications. In this paper, we propose a plug-and-play virus network, dubbed PAPnet, for backdoor attack. PAPnet is a lightweight network with the same dimensional output as the victim network. In the training stage, we only need the output of the victim network and train PAPnet to learn from poisoned data and clean data. This makes PAPnet easier to learn than fine-tunebased backdoor attack methods. Besides, PAPnet can be easily attached to the different classification network models without modifying the architecture of the victim network and fine-tuning processing. We have conducted various experiments on four datasets with four classical classification networks. Experimental results demonstrate the superiority of our proposed method.
Rui Huang 0006, Zongyu Guo, Qingyi Zhao, Wei Fan 0001
CSCWD1
2024 SPY-Watermark: Robust Invisible Watermarking for Backdoor Attack
abstract
Backdoor attack aims to deceive a victim model when facing backdoor instances while maintaining its performance on benign data. Current methods use manual patterns or special perturbations as triggers, while they often overlook the robustness against data corruption, making backdoor attacks easy to defend in practice. To address this issue, we propose a novel backdoor attack method named Spy-Watermark, which remains effective when facing data collapse and backdoor defense. Therein, we introduce a learnable watermark embedded in the latent domain of images, serving as the trigger. Then, we search for a watermark that can withstand collapse during image decoding, cooperating with several anti-collapse operations to further enhance the resilience of our trigger against data corruption. Extensive experiments are conducted on CIFAR10, GTSRB, and ImageNet datasets, demonstrating that Spy-Watermark overtakes ten state-of-the-art methods in terms of robustness and stealthiness.
Ruofei Wang, Renjie Wan, Zongyu Guo, Qing Guo 0005, Rui Huang 0006
ICASSP5
2024 Instance-Level Detection and Region Partition of HPT Blade by Slot Localization
Rui Huang 0006, Xuyi Cheng, Jingcheng Zeng
ICIC (6)1
2024 PBIM: Paired Backdoor Injection Method for Change Detection
Rui Huang 0006, Mengjia Hao, Zongyu Guo
ICIC (3)1
2024 AutoClick: Auto Seed Selection for Interactive Segmentation
Rui Huang 0006, Jingcheng Zeng
ICIC (6)1
2024 AdaAug+: A Reinforcement Learning-Based Adaptive Data Augmentation for Change Detection
abstract
Data augmentation (DA) increases the diversity of training data to improve the model generalization ability. Most of the DA methods focus on image classification or object detection. Directly using the existing DA methods on change detection tasks not only falls short of fully exploring the specificity of the change image pairs but also leads to longer training times. In this article, we first propose a mask-guided mixing (MGM) DA for change detection, which mixes the change regions of the current training sample based on prediction results and labels to generate high-quality samples with more positive samples. We then propose a new reinforcement learning (RL)-based Adaptive DA method, AdaAug+, to adaptively select the optimal DA policy for the training samples. An actor selects the best augmentation operation from the operation set according to the image pair. The augmented image pairs make it easier for the change detector to learn the optimal parameters and improve the final detection performance. To reduce the training time, we identify and remove the redundant training samples during the training process by our redundancy searching policy. We have conducted various experiments on four remote sensing change detection datasets with different change detectors. The experimental results demonstrate that AdaAug+ achieves promising performance compared to the state-of-the-art DA methods and requires less training time.
Rui Huang 0006, Jieda Wei, Qing Guo 0005
IEEE Trans. Geosci. Remote. Sens.1
2023 Background-Mixed Augmentation for Weakly Supervised Change Detection
abstract
Change detection (CD) is to decouple object changes (i.e., object missing or appearing) from background changes (i.e., environment variations) like light and season variations in two images captured in the same scene over a long time span, presenting critical applications in disaster management, urban development, etc. In particular, the endless patterns of background changes require detectors to have a high generalization against unseen environment variations, making this task significantly challenging. Recent deep learning-based methods develop novel network architectures or optimization strategies with paired-training examples, which do not handle the generalization issue explicitly and require huge manual pixel-level annotation efforts. In this work, for the first attempt in the CD community, we study the generalization issue of CD from the perspective of data augmentation and develop a novel weakly supervised training algorithm that only needs image-level labels. Different from general augmentation techniques for classification, we propose the background-mixed augmentation that is specifically designed for change detection by augmenting examples under the guidance of a set of background changing images and letting deep CD models see diverse environment variations. Moreover, we propose the augmented & real data consistency loss that encourages the generalization increase significantly. Our method as a general framework can enhance a wide range of existing deep learning-based detectors. We conduct extensive experiments in two public datasets and enhance four state-of-the-art methods, demonstrating the advantages of our method. We release the code at https://github.com/tsingqguo/bgmix.
Rui Huang 0006, Ruofei Wang, Qing Guo 0005, Jieda Wei, Yuxiang Zhang 0003, Wei Fan 0001, Yang Liu 0003
AAAI1
2023 Multi-scale Convolutional Feature Approximation for Defocus Blur Detection
abstract
Deep learning technology has promoted the performance of defocus blur detection. However, blur detectors suffer from background clutter, scale ambiguity and blurred boundaries of the defocus blur regions. To conquer these issues, previous methods propose to use multi-scale image patches or images for blur detection, which costs much computation time. In this paper, we propose a deep neural network that takes a single-scale image as input to generate robust defocus blur detection. Specifically, we first extract multi-scale convolutional features by a feature extraction network. And then we resize the convolutional features of each layer by a fixed ratio to approximate convolutional features that extracted from a resized image with the same ratio. By approximation, it not only generates features extracted from a scaled image but also reduces the computation of feature extraction from multi-scale images. We concatenate the features extracted from the original image with the approximated features at the corresponding layers by convolutional layers to increase the blur distinguish ability. We gradually fuse the convolutional features from top-to-bottom by Conv-LSTMs to refine the blur predictions. We compare our method with nine state-of-the-art defocus blur detectors on two defocus blur detection benchmark datasets. Experiment results demonstrate the effectiveness of our proposed defocus blur detector.
Rui Huang 0006, Huan Lu, Wei Fan 0001
CSCWD1
2023 HQFS: High-Quality Feature Selection for Accurate Change Detection
Qi'ao Xu, Qingyi Zhao, Rui Huang 0006, Yuxiang Zhang 0003
ICIG (1)4
2023 SpanMTL: a span-based multi-table labeling for aspect-oriented fine-grained opinion extraction
Yuexuan Zhu, Wei Fan 0001, Yuxiang Zhang 0003, Rui Huang 0006, Zhaojun Gu, Andrew W. H. Ip, Kai-Leung Yung
Soft Comput.5
2022 Selecting change image for efficient change detection
abstract
Abstract Change detection (CD) is a fundamental problem that aims at detecting changed objects from two observations. Previous CNN‐based CD methods detect changes through multi‐scale deep convolutional features extracted from two images. However, we find that change always occurs in the ‘Query’ image for fixed cameras. This condition means that changes can be detected in advance from a single image with a coarse change. In this paper, we propose an efficient CD method to detect precise changes from the change image. First, a change image selector is designed to identify the image containing changes. Second, a coarse change prior map generator is proposed to generate coarse change prior to indicate the position of changes. Then, we introduce a simple multi‐scale CD module to refine the coarse change detection. As only one image is used in the multi‐scale CD module, our method is more efficient in training and testing than other compared methods. Numerous experiments have been conducted to analyse the effectiveness of the proposed method. Experimental results show that the proposed method achieves superior detection performance and higher speed than other compared CD methods.
Rui Huang 0006, Ruofei Wang, Yuxiang Zhang 0003, Wei Fan 0001, Kai-Leung Yung
IET Signal Process.1
2021 Change detection with cross enhancement of high- and low-level change-related features
abstract
Abstract Change detection (CD) is a fundamental yet challenging problem, which aims at detecting changed object in two observations. Recent CD methods are designed based on the off‐the‐shelf semantic segmentation network architectures, which is not optimal for extracting and using change‐related features. In this paper, a novel CD network architecture is proposed, including change‐related feature extraction, cross feature enhancement, and multi‐level supervision. Absolute difference of the features of different convolutional layers is first computed from a Unet‐like network for two observations. The features are partitioned into high‐ and low‐level features according to their functionalities. Then the high‐ and low‐level features are recurrently refined by cross feature enhancement to increase the representational ability of the features. The network learns change‐related features with multi‐level supervisions. The final CD result can be obtained by fusing multiple predictions. Experimental results on three CD benchmark datasets indicate the superiority of the authors' method when compared with six state‐of‐the‐art deep learning‐based CD methods.
Rui Huang 0006, Ruofei Wang
IET Image Process.1
2021 Change detection with various combinations of fluid pyramid integration networks
Rui Huang 0006, Yaobin Zou, Wei Fan 0001
Neurocomputing1
2020 Exemplar-based image saliency and co-saliency detection
Rui Huang 0006, Wei Feng 0005, Yaobin Zou
Neurocomputing1
2020 Change detection with absolute difference of multiscale deep features
Rui Huang 0006, Qiang Zhao 0005, Yaobin Zou
Neurocomputing1
2020 A network representation method based on edge information extraction
Wei Fan 0001, Hui Min Wang, Rui Huang 0006, Andrew W. H. Ip, Kai-Leung Yung
Soft Comput.4
2020 Triple-Complementary Network for RGB-D Salient Object Detection
abstract
Most of the existing RGB-D saliency detectors have tried different strategies to fuse RGB and depth information to generate better saliency detection results. However, the ability of only using RGB image to detect the salient object is ignored in RGB-D saliency detectors. In this paper, we propose a triple-complementary network for fully exploring the RGB and depth information. The first two sub-networks are used for extracting saliency from RGB image and RGB-D image pair, respectively. The third sub-network refines the saliency map by comprehensively considering RGB image, depth map and the refined saliency map of the first two sub-networks. We propose depth weighted refinement to suppress the high contrast objects with large depth values. The proposed method not only fully explores the information of RGB but also makes the RGB and depth tightly coupled. The experiments on five RGB-D saliency detection datasets show the superiority of the proposed method over the state-of-the-art RGB-D saliency detectors.
Rui Huang 0006, Yaobin Zou
IEEE Signal Process. Lett.1
2019 Predicting Scientific Impact via Heterogeneous Academic Network Embedding
Chunjing Xiao, Jianing Han, Wei Fan 0001, Senzhang Wang, Rui Huang 0006, Yuxiang Zhang 0003
PRICAI (2)5
2019 RGB-D Salient Object Detection by a CNN With Multiple Layers Fusion
abstract
Recent saliency detectors use depth to improve the precision of the results. However, most of the existing RGB-D saliency detectors only treat depth as an additional feature, which cannot explore the distinguishing ability of the depth map. In this letter, we propose a deep convolutional neural network for RGB-D saliency detection. The network takes RGB-D as inputs and produces a saliency prediction in an end-to-end manner. To solve the scale problem, we fuse the features of higher layers to the features of lower layers gradually. Our method not only fully explores the information of the depth, but also makes the RGB and depth tightly coupled. Compared with the state-of-the-art RGB saliency detectors and depth-aware saliency detectors on two benchmark datasets, our method outperforms the competitors on mean absolute error and F-measure by large margins.
Rui Huang 0006
IEEE Signal Process. Lett.1
2018 Multiscale blur detection by learning discriminative deep features
Rui Huang 0006, Wei Feng 0005, Mingyuan Fan 0001
Neurocomputing1
2017 Learning Dynamic Siamese Network for Visual Object Tracking
abstract
How to effectively learn temporal variation of target appearance, to exclude the interference of cluttered background, while maintaining real-time response, is an essential problem of visual object tracking. Recently, Siamese networks have shown great potentials of matching based trackers in achieving balanced accuracy and beyond realtime speed. However, they still have a big gap to classification & updating based trackers in tolerating the temporal changes of objects and imaging conditions. In this paper, we propose dynamic Siamese network, via a fast transformation learning model that enables effective online learning of target appearance variation and background suppression from previous frames. We then present elementwise multi-layer fusion to adaptively integrate the network outputs using multi-level deep features. Unlike state-of-theart trackers, our approach allows the usage of any feasible generally- or particularly-trained features, such as SiamFC and VGG. More importantly, the proposed dynamic Siamese network can be jointly trained as a whole directly on the labeled video sequences, thus can take full advantage of the rich spatial temporal information of moving objects. As a result, our approach achieves state-of-the-art performance on OTB-2013 and VOT-2015 benchmarks, while exhibits superiorly balanced accuracy and real-time response over state-of-the-art competitors.
Qing Guo 0005, Wei Feng 0005, Ce Zhou, Rui Huang 0006, Song Wang 0002
ICCV4
2017 Color Feature Reinforcement for Cosaliency Detection Without Single Saliency Residuals
abstract
Cosaliency detects the common salient objects within a group of images. Hence, those objects that are salient only in individual image or small portion of the image group should conceptually be treated as background. However, most state-of-the-art methods cannot do this well because they measure cosaliency as an explicit combination of single-image saliency and interimage similarity, thus inevitably leaving single saliency residuals into the cosaliency maps. In this letter, we show such problem can be solved by color feature reinforcement, based on a simple observation that cosalient objects usually have similar color distributions in an abundant color feature space. Since we model the cosaliency of an image w.r.t. another one as a reinforced product of the foreground dictionary and sparsely coded saliency coefficients of the two images, respectively, within a same rich feature space, we can effectively eliminate the single saliency residual effect in cosaliency detection. Extensive experiments validate the superior performance of the proposed approach on benchmark datasets.
Rui Huang 0006, Wei Feng 0005
IEEE Signal Process. Lett.1
2015 Saliency and co-saliency detection by low-rank multiscale fusion
abstract
To facilitate efficiency, most recent successful saliency detection methods are built on superpixel level. However, saliency detection with single-scale superpixel segmentation may fail in capturing the intrinsic salient objects in complex natural scenes with small-scale high-contrast backgrounds. To tackle this problem and realize more reliable saliency detection, we present a simple strategy using multiscale superpixels to jointly detect salient object via low-rank analysis. Specifically, we construct a multiscale superpixel pyramid and derive the corresponding saliency map using multiple saliency features and priors for each single scale at first. Then, we show that by joint low-rank analysis of multiscale saliency maps, we can obtain a more reliable adaptively fused saliency map that takes all scales saliency results into account. We further propose a GMM-based co-saliency prior to enable the above approach to detecting co-salient objects from multiple images. Extensive experiments on benchmark datasets validate the effectiveness and superiority of the proposed approach over state-of-the-art methods.
Rui Huang 0006, Wei Feng 0005
ICME1