Shuwei Huo

dblp:202/6581 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0002-7290-7838ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Query-calibrated attention and visibility-aware anisotropic regression for small object detection
Ruxiang Duan, Qiliang Du, Lianfang Tian, Shikun Feng, Shuwei Huo
Neurocomputing6
2025 Skim-and-scan transformer: A new transformer-inspired architecture for video-query based video moment retrieval
Shuwei Huo, Yuan Zhou 0006, Keran Chen, Wei Xiang 0001
Expert Syst. Appl.1
2025 Adaptive motion enhancement for passive non-line-of-sight action recognition
Zhongqi Sun, Yuan Zhou 0006, Shuwei Huo, Sun-Yuan Kung
Neurocomputing3
2025 Focal Spotter: Small-Object Detection via Inner Scale and Interscale Feature Focalization
abstract
Small object detection is a challenging but vital task for applications such as search and rescue, surveillance, and remote sensing, where the goal is to detect small targets such as people, vehicles, or boats in complex environments. The minimal pixel representation of such targets, combined with cluttered backgrounds and environmental variations like lighting changes, makes accurate detection particularly difficult. To address these challenges, we propose Focal Spotter, a novel transformer-based detector that incorporates two key modules:the Energy-Based Focal Module (EFM) and the Focalized Inter-Scale Feature Complementary Module (FICM). The EFM leverages an Energy-Based Model(EBM) to achieve inner-scale feature focalization. The EBM dynamically allocates weights through an energy function, focusing on low-energy target signals (such as small object features), significantly enhancing the model’s sensitivity to sparse or weak signals. This capability allows EFM to outperform conventional attention methods in both robustness and precision.The FICM facilitates inter-scale feature focalization by aligning and integrating high-level semantic and low-level detailed features spatially and channel-wise. By preemptively extracting and merging compensatory information across scales, it boosts feature fusion efficiency, minimizes redundancy, ensures semantic consistency, and achieves precise spatial localization, significantly enhancing small object detection in complex scenes. Extensive experiments on multiple benchmark datasets, including SeaDronesSee and VisDrone, demonstrate that Focal Spotter outperforms state-of-the-art methods in small object detection. Ablation studies highlight the critical contributions of EFM and FICM, with EFM showing superior performance over other attention mechanisms in capturing sparse small targets. The results underscore the robustness and effectiveness of our approach across diverse and challenging scenarios, from maritime to urban environments.
Ruxiang Duan, Qiliang Du, Lianfang Tian, Shuwei Huo
IEEE Trans. Geosci. Remote. Sens.5
2024 Weakly Supervised Video Re-Localization Through Multi-Agent-Reinforced Switchable Network
abstract
The objective of video re-localization (VRL) is to localize a successive sequence of frames, namely, the target moment, from untrimmed reference videos that semantically correspond to a given query video. During training, the weakly supervised setting of VRL provides only coarse-grained video-level rather than frame-level annotations. For the weakly supervised VRL (WS-VRL) task, obtaining effective video feature representations that can be used to evaluate the relevance between videos and localizing the accurate temporal boundaries of the target moment remain challenging. In this paper, a novel multi-agent-reinforced switchable network (MARS) is proposed to address these challenges. MARS can adaptively guide video feature encoding and moment localization using multiple learned agents. Specifically, an agent-controlled switchable encoder is used to obtain effective video feature representations, and an agent-reinforced boundary localizer is used to determine accurate localized moments through progressive refinement. Furthermore, a relevance-oriented reward generator was designed to estimate the relevance of the localized moment to the query video and assign a reward to multiple agents. The effectiveness of the proposed MARS model was verified through extensive experiments on the ActivityNet-VRL dataset.
Yuan Zhou 0006, Axin Guo, Shuwei Huo, Yu Liu 0004, Sun-Yuan Kung
IEEE Trans. Circuits Syst. Video Technol.3
2024 Geometric Variation Adaptive Network for Remote Sensing Image Change Detection
abstract
Change detection identifies surface changes on the earth by comparing two images from the same area at different times. To generate smooth change maps, a common method is fusing information from neighboring areas around each pixel. While the conventional fusion methods primarily rely on fixed regular-shaped neighboring areas, which may be inadequate in capturing the diverse and irregular geometric structures of changed ground objects. To address this limitation, we propose a novel Geometric Variation Adaptive Change Detector (GVA-CD), which adaptively adjusts the shape and size of neighboring areas based on the geometrical structure of ground objects. More specifically, we design a new geometric variation adaptive module (GVAM) as a component of GVA-CD. GVAM captures the structure of the ground objects to constructs geometrically flexible neighboring areas for each pixel, enabling the model to adapt to different ground object structures and generate discriminative difference features. We further propose a new difference measurement module to compute the difference between the features of pre-and post-change images by leveraging the adaptive neighboring areas. In addition, the GVA-CD introduces a multi-stage cross-scale fusion mechanism in both feature extraction and change map generation, to enhance the scale adaption ability of the feature extraction and change map generation. Extensive experiments on three large datasets demonstrate that our GVA-CD can outperform existing methods in change detection.
Shuwei Huo, Yuan Zhou 0006, Lei Zhang 0202, Yanjie Feng, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Geosci. Remote. Sens.1
2024 Dynamic View Aggregation for Multi-View 3D Shape Recognition
abstract
In the field of 3D shape recognition, the view-based approach has achieved state-of-the-art performance. A major challenge that needs to be addressed by the view-based approach is how to effectively aggregate multi-view features to obtain a better 3D shape representation. Existing methods which rely on networks with static parameters for feature aggregation adversely coerce the network to learn a general feature aggregation strategy for all inputs, ignoring the diversity of input 3D shapes in real-world scenarios. In this work, we propose a novelDynamic View Aggregation NetworkcalledDVA-Netto address this challenge. DVA-Net can dynamically adjust the network parameter depending on the input 3D shapes to flexibly fuse multi-view information. The shape-specific parameter adaptation is achieved by our designedDynamic Relation-aware Aggregationmodule, dubbedDRAmodule. It is responsible for learning relations among views and adaptively integrating multi-view features. Comprehensive experiments on benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance for 3D shape classification and retrieval.
Yuan Zhou 0006, Zhongqi Sun, Shuwei Huo, Sun-Yuan Kung
IEEE Trans. Multim.3
2024 Automatic Metric Search for Few-Shot Learning
abstract
Few-shot learning (FSL) aims to learn a model that can identify unseen classes using only a few training samples from each class. Most of the existing FSL methods adopt a manually predefined metric function to measure the relationship between a sample and a class, which usually require tremendous efforts and domain knowledge. In contrast, we propose a novel model called automatic metric search (Auto-MS), in which an Auto-MS space is designed for automatically searching task-specific metric functions. This allows us to further develop a new searching strategy to facilitate automated FSL. More specifically, by incorporating the episode-training mechanism into the bilevel search strategy, the proposed search strategy can effectively optimize the network weights and structural parameters of the few-shot model. Extensive experiments on the miniImageNet and tieredImageNet datasets demonstrate that the proposed Auto-MS achieves superior performance in FSL problems.
Yuan Zhou 0006, Jieke Hao, Shuwei Huo, Boyu Wang 0004, Leijiao Ge, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.3
2023 Hierarchical full-attention neural architecture search based on search space compression
abstract
Neural architecture search (NAS) has significantly advanced the automatic design of convolutional neural architectures. However, it is challenging to directly extend existing NAS methods to attention networks because of the uniform structure of the search space and the lack of long-range feature extraction. To address these issues, we construct a hierarchical search space that allows various attention operations to be adopted for different layers of a network. To reduce the complexity of the search, a low-cost search space compression method is proposed to automatically remove the unpromising candidate operations for each layer. Furthermore, we propose a novel search strategy combining a self-supervised search with a supervised one to simultaneously capture long-range and short-range dependencies. To verify the effectiveness of the proposed methods, we conduct extensive experiments on various learning tasks, including image classification , fine-grained image recognition, and zero-shot image retrieval . The empirical results show strong evidence that our method is capable of discovering high-performance full-attention architectures while guaranteeing the required search efficiency.
Yuan Zhou 0006, Shuwei Huo, Boyu Wang 0004
Knowl. Based Syst.3
2023 Weakly-supervised content-based video moment retrieval using low-rank video representation
abstract
Content-based video moment retrieval (CVMR) aims to localize a successive sequence of frames in an untrimmed reference video, called target moment, that is semantically corresponding to a given query video. Current state-of-the-art CVMR methods are mainly developed using frame-level annotation, which is often quite expensive to collect. In this paper, we aim to develop a weakly-supervised CVMR method, which uses coarse-grained video-level annotations during training. Under weak supervision, video localizers require more discriminative frame-level video features. To achieve this goal, we proposed a novel prior, termed low-rank prior, based on an observation that the frame-level feature of a video should have low-rank properties. We demonstrated that the low-rank features are more discriminative and are beneficial to accurately localize the action boundaries. To produce a low-rank feature, we designed a low-rank feature reconstruction (LFR) operator. A new differentiable matrix decomposition approach is proposed to generate the low-rank reconstruction of the input matrix, meanwhile ensuring that the matrix decomposition process is differentiable. Based on the LFR, we developed a new weakly-supervised CVMR model which produces low-rank video representation and performs semantic consistency measures to discover the semantically matched segment in the reference video to the query video. Extensive experiments demonstrate that our method outperforms state-of-the-art weakly-supervised methods consistently and even achieves competing performance to fully-supervised baselines.
Shuwei Huo, Yuan Zhou 0006, Wei Xiang 0001, Sun-Yuan Kung
Knowl. Based Syst.1
2023 Hyperspectral Band Selection With Iterative Graph Autoencoder
abstract
Hyperspectral band selection (BS) is an important task for hyperspectral image (HSI) processing, which aims to select a discriminative and low-redundant band subset. As a significant cue for BS, structure information describes the cross-band correlation which brings the redundancy of HSI. Existing methods model structure information via manual rule-based graph construction. However, such a graph construction method fails to model complex and diverse structural relationships of HSI data. To address this problem, we propose a data-driven method, named iterative graph auto-encoder for band selection (IGAEBS). It adaptively captures structure information by a data-specific automatic construction process, rather than by a fixed empirical design. Specifically, we propose a new unsupervised pretext task to train graph convolution neural network to extract HSI features. These features are used to construct a graph to represent the structural relationships among bands. To enhance the reliability of the graph, we further design an iterative graph improvement mechanism to progressively refine the structure representation. Using the derived graph, we partition the bands into several clusters and select a representative band from each cluster. During the selection process, both intra- and inter-cluster information are considered to improve the discriminativeness of band subset. Extensive experiments are conducted on three public data sets to validate the superiority of the proposed method compared to other state-of-the-art methods.
Yuan Zhou 0006, Qingren Yao, Shuwei Huo, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Semantic Relevance Learning for Video-Query Based Video Moment Retrieval
abstract
The task of video-query based video moment retrieval (VQ-VMR) aims to localize the segment in the reference video, which matches semantically with a short query video. This is a challenging task due to the rapid expansion and massive growth of online video services. With accurate retrieval of the target moment, we propose a new metric to effectively assess the semantic relevance between the query video and segments in the reference video. We also develop a new VQ-VMR framework to discover the intrinsic semantic relevance between a pair of input videos. It comprises two key components: a Fine-grained Feature Interaction (FFI) module and a Semantic Relevance Measurement (SRM) module. Together they can effectively deal with both the spatial and temporal dimensions of videos. First, the FFI module computes the semantic similarity between videos at a local frame level, mainly considering the spatial information in the videos. Subsequently, the SRM module learns the similarity between videos from a global perspective, taking into account the temporal information. We have conducted extensive experiments on two key datasets which demonstrate noticeable improvements of the proposed approach over the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Ruolin Wang, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Multim.1
2022 YFormer: A New Transformer Architecture for Video-Query Based Video Moment Retrieval
Shuwei Huo, Yuan Zhou 0006
PRCV (3)1
2022 Robust object tracking via deformation samples generator
Xuesong Gao, Yuan Zhou 0006, Shuwei Huo, Zizi Li, Keqiu Li
J. Vis. Commun. Image Represent.3
2022 Cross-Scale Residual Network: A General Framework for Image Super-Resolution, Denoising, and Deblocking
abstract
In general, image restoration involves mapping from low-quality images to their high-quality counterparts. Such optimal mapping is usually nonlinear and learnable by machine learning. Recently, deep convolutional neural networks have proven promising for such learning processing. It is desirable for an image processing network to support well with three vital tasks, namely: 1) super-resolution; 2) denoising; and 3) deblocking. It is commonly recognized that these tasks have strong correlations, which enable us to design a general framework to support all tasks. In particular, the selection of feature scales is known to significantly impact the performance on these tasks. To this end, we propose the cross-scale residual network to exploit scale-related features among the three tasks. The proposed network can extract spatial features across different scales and establish cross-temporal feature reusage, so as to handle different tasks in a general framework. Our experiments show that the proposed approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations for multiple image restoration tasks.
Yuan Zhou 0006, Xiaoting Du, Mingfei Wang, Shuwei Huo, Yeda Zhang, Sun-Yuan Kung
IEEE Trans. Cybern.4
2022 Joint Frequency-Spatial Domain Network for Remote Sensing Optical Image Change Detection
abstract
Change detection for remote sensing images involves detecting regional surface changes of interest between two images taken of the same geographical area but at different times. In image processing, the spatial domain uses grayscale values to describe an image. The frequency is directly related to the spatial change rate, so the frequency domain can be intuitively associated with patterns of intensity variations in the image. These two domains provide different perspectives for image interpretation. Most existing deep-learning-based methods formulate change detection as a pixel-wise binary classification problem and utilize various strategies to extract information in the spatial domain. However, they rarely pay attention to the rich information in the frequency domain. To address this problem, we propose an end-to-end joint frequency-spatial domain network (JFSDNet) to implement remote sensing optical image change detection. Specifically, we introduce frequency information into the change detection to supplement the loss of image details caused by down-sampling. In addition, we employ a frequency selection module to adaptively discriminate and choose frequency clues by reducing the complexity of the frequency features. The JFSDNet is applied to two publicly available datasets: the CDD dataset and the LEVIR-CD dataset. Compared with other methods, both visual interpretation and quantitative assessment confirmed that our proposed method achieved a favorable performance.
Yuan Zhou 0006, Yanjie Feng, Shuwei Huo, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Shape autotuning activation function
Yuan Zhou 0006, Shuwei Huo, Sun-Yuan Kung
Expert Syst. Appl.3
2020 Adaptive Irregular Graph Construction-Based Salient Object Detection
abstract
Saliency detection represents a vital pre-processing stage of computer vision. Most existing propagation-based salient object detection methods construct a k-regular graph for saliency propagation. Applying a regular graph to a vast smooth region is potentially prone to unnecessary or prolonged propagation errors, leading to the excessive highlighting of the background regions. To mitigate such problems, we substitute the conventional k-regular graph with an adaptive irregular graph for saliency value propagation, thereby avoiding unnecessary iterations over a vast smooth region. We first perform a clustering analysis based on the smoothness, color, and other features of regions. The new graph boosts an adaptive link density by considering the clustering result. In addition, we propose a seeding strategy for the propagation. Based on our experimental studies of six major benchmark datasets, our method performed favorably against the other state-of-the-art methods, both quantitatively and qualitatively.
Yuan Zhou 0006, Shuwei Huo, Chunping Hou, Sun-Yuan Kung
IEEE Trans. Circuits Syst. Video Technol.3
2019 Semi-Supervised Salient Object Detection Using a Linear Feedback Control System Model
abstract
To overcome the challenging problems in saliency detection, we propose a novel semi-supervised classifier which makes good use of a linear feedback control system (LFCS) model by establishing a relationship between control states and salient object detection. First, we develop a boundary homogeneity model to estimate the initial saliency and background likelihoods, which are regarded as the labeled samples in our semi-supervised learning procedure. Then in order to allocate an optimized saliency value to each superpixel, we present an iterative semi-supervised learning framework which integrates multiple saliency cues and image features using an LFCS model. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. This paper also covers comprehensive simulation study based on public datasets, which demonstrates the superiority of the proposed approach.
Yuan Zhou 0006, Shuwei Huo, Wei Xiang 0001, Chunping Hou, Sun-Yuan Kung
IEEE Trans. Cybern.2
2019 Salient Object Detection via Fuzzy Theory and Object-Level Enhancement
abstract
This paper proposes a bottom-up saliency detection method via effective integration of regional saliency measure and object-level information using fuzzy theory. First, we generate an initial saliency map by fusing multiple prior maps. Second, to emphasize the object-level concept of saliency, we further generate many object proposals of the input image. A fuzzy set theory is then applied to measure the objectness score of the object proposals and integrate them into an objectness map. Third, an optimization framework is proposed to effectively fuse various prior saliency cues and object-level information to produce a clean and uniform saliency map as well as to maintain the salient object completeness. Experimental studies in several benchmark datasets confirmed the superiority of the proposed method over state-of-the-art saliency detection methods.
Yuan Zhou 0006, Ailing Mao, Shuwei Huo, Jianjun Lei 0001, Sun-Yuan Kung
IEEE Trans. Multim.3
2019 Semisupervised Learning Based on a Novel Iterative Optimization Model for Saliency Detection
abstract
In this paper, we propose a novel iterative optimization model for bottom-up saliency detection. By exploring bottom-up saliency principles and semisupervised learning approaches, we design a high-performance saliency analysis method for wide ranging scenes. The proposed algorithm consists of two stages: 1) we develop a boundary homogeneity model to characterize the general position and the contour of the salient objects and 2) we propose a novel iterative optimization model, termed gradual saliency optimization, for further performance improvement. Our main contribution falls on the second stage, where we propose an iterative framework with self-repairing mechanisms for refining saliency maps. In this framework, we further develop a more comprehensive optimization function applying a novel semisupervised learning scheme to enhance the traditional saliency measure. More elaborately, the iterative method can gradually improve the output in each iteration and finally converge to high-quality saliency maps. Based on our experiments on four different public data sets, it can be demonstrated that our approach significantly outperforms the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.1
2018 Super-resolution Imaging Based on Global Interpolation and Structural Similarities
abstract
In this paper, we propose a double dictionary learning method for image super-resolution (SR) reconstruction. Different from existing dictionary learning based super-resolution, we combine both self-similarity and external images to construct a double dictionary learning method. A new optimization model is established using self-similarities and external-similarities as regularization terms. Furthermore, we propose a global interpolation method to reconstruct an accurate initial estimation at the edges. Experimental results show that the proposed algorithm can produce high-quality reconstruction results both perceptually and quantitatively in terms of peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), as compared to existing algorithms.
Yuan Zhou 0006, Shuwei Huo, Sun-Yuan Kung
ICPR2
2018 Iterative Feedback Control-Based Salient Object Segmentation
abstract
In this paper, we establish a mathematical model that relates the control states and the saliency values in salient object detection. We show that a linear feedback control system (LFCS) is amenable to saliency detection tasks owing to its functional properties. This inspired us to employ an LFCS to detect salient objects in static images. Based on the novel iteration method, the system gradually converges to an optimized stable state, which is associated with an accurate saliency map. In addition, to initialize the system, we propose a so-called boundary homogeneity based on a priori knowledge of the boundary in order to estimate the background likelihood and indirectly obtain a foreground (saliency) map. The experimental results indicate that such a feedback control model can offer significant improvement in salient object detection performance.
Shuwei Huo, Yuan Zhou 0006, Jianjun Lei 0001, Nam Ling, Chunping Hou
IEEE Trans. Multim.1
2017 Salient object detection via a linear feedback control system
abstract
Linear feedback control systems (LFCS) have been widely applied in signal analysis, filtering, and error correction. Many functional properties of LFCS are amenable to numerous object recognition and detection tasks. In fact, there exists an intimate relationship between control states and salient values. This prompts us to adopt the linear feedback control system to detect salient object in static images. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. In addition, to initialize the system, we propose the so called boundary homogeneity based on a priori knowledge on the boundary to estimate the background likelihood and indirectly depict a foreground (saliency) map. By our experimental results, we demonstrates that such feedback control model can bring about noticable improvement in salient object detection.
Shuwei Huo, Yuan Zhou 0006, Sun-Yuan Kung
ICIP1
2017 Label propagation based saliency detection via graph design
abstract
Saliency detection has been widely used as the pre-processing of the computer-vision tasks. Existing propagation based saliency detection methods simply select a k-regular graph for saliency propagation, which usually leads to the mistaken highlighting of the long-range smooth background regions. In this paper, we design a novel graph for label propagation based saliency detection by considering the local consistency and the global symmetry of the image scene and updating the graph model based on smoothness assumption and cluster assumption. Then, we label the reliable seeds and propagate the saliency value through the designed graph. On two widely used large open benchmark data sets, the proposed method significantly outperforms thirteen state-of-the-arts under either quantitative or qualitative evaluation.
Yuan Zhou 0006, Shuwei Huo, Chunping Hou
ICIP3
2017 Semi-supervised saliency classifier based on a linear feedback control system model
abstract
Linear feedback control systems (LFCS) are amenable to numerous object recognition and detection tasks on account of its functional properties in signal filtering and error correction. In fact, there exists an intimate relationship between control states and salient values. Therefore, we propose a novel semi-supervised classifier which makes use of linear feedback control theory to improve saliency detection performance. First, we develop a boundary homogeneity model to estimate the initial saliency and background likelihoods, which may lead to the labeled samples in our semi-supervised learning procedure. Then in order to allocate an optimized saliency value to each superpixel, we present an iterative semi-supervised learning framework which integrates multiple saliency cues and image features using a LCSF model. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. Based on our experiments on public datasets, it can be demonstrated that our approach significantly outperforms the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Sun-Yuan Kung
IJCNN1