Yuan Zhou 0006

dblp:40/7018-6 · DBLP profile ↗
← Back
71ranked-venue papers
30as first author
42since 2021 · last 2027
0000-0002-6072-337XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 9 first-author · 7 since 2021Artificial intelligence and machine learning · 23 · 12 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 7 first-author · 15 since 2021Computer networks · 7 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2027 Auto-metric network for open-set recognition
Shuoshi Li, Yuan Zhou 0006, Ke Zhang 0005, Sun-Yuan Kung
Expert Syst. Appl.2
2026 Large-model-based smart agent for time series anomaly detection in power systems
Bingrui Wang, Yuan Zhou 0006, Leijiao Ge, Sun-Yuan Kung
Expert Syst. Appl.2
2026 EViMNet: An Efficient Visual Model for Radar-Based Human Activity Recognition
abstract
In recent years, Radar-based human action recognition has recently gained attention due to its high precision and robustness, making it well-suited for Internet of Things (IoT) applications. However, radar spectrograms often contain noise and redundant information, while conventional classification models exhibit limitations in capturing complex nonlinear patterns, which hinders recognition performance. To address these challenges, we propose EViMNet, an efficient vision-based model for robust Radar-based human behavior recognition. EViMNet integrates two core components. The first is a Stem Filtering Block (SFB) that applies Gaussian filtering to enhance target features and suppress source noise. The second is an improved Multi-Layer Perceptron (MLP) architecture, where traditional linear layers are replaced by Kolmogorov–Arnold Networks (KANs). These KANs utilize B-spline basis functions to enable flexible nonlinear modeling and increase sensitivity to subtle signal variations. Experimental results show that EViMNet achieves 95.20% accuracy on the Glasgow dataset. For generalization evaluation, it further attains 99.18% cross-validation accuracy on the Doppler Gesture dataset and achieves the best cross-dataset performance when trained on Glasgow and tested on the corresponding actions in CI4R-MIX77. These results verify the strong generalization and practical applicability of the proposed model in IoT scenarios.
Liyuan Lin, Guanhua Qiao, Jingpeng Yan, Leguang Wang, Weibin Zhou, Yuan Zhou 0006
IEEE Internet Things J.9
2025 Skim-and-scan transformer: A new transformer-inspired architecture for video-query based video moment retrieval
Shuwei Huo, Yuan Zhou 0006, Keran Chen, Wei Xiang 0001
Expert Syst. Appl.2
2025 Adaptive motion enhancement for passive non-line-of-sight action recognition
Zhongqi Sun, Yuan Zhou 0006, Shuwei Huo, Sun-Yuan Kung
Neurocomputing2
2025 Illumination guided domain adaptation object detection in thermal imagery
Yuan Zhou 0006, Yu Liu 0004, Sun-Yuan Kung
Neurocomputing1
2025 Adaptive Control Scheme for USV Trajectory Tracking Under Complex Environmental Disturbances via Deep Reinforcement Learning
abstract
Unmanned surface vehicles (USVs) have demonstrated impressive practical value and potential in Marine Internet of Things (MIoT) system. Although trajectory-tracking control is among the most common practical technology of USVs, various limitations remain unaddressed. Existing studies have employed simple mathematical models to simulate marine environment without utilizing actual data, resulting in a lack of environmental authenticity. Moreover, a complex marine environment increases the need for robustness and adaptability of the control policy. To overcome these limitations, this study proposes a deep reinforcement learning (DRL)-based policy for USV trajectory-tracking control, which can effectively adapt to complex environmental disturbances. First, we use actual marine data, including ocean currents and winds, to construct a time-varying multi-element marine environment model. Next, an effective Markov decision processes (MDPs) formulation integrating LOS guidance law is elaborately proposed, in which the composite reward function and state transition function are used to avoid ineffective exploration and achieve better convergence ability. Furthermore, a USV trajectory-tracking controller based on hybrid priority twin-delayed deep deterministic policy gradient (TD3) agent is designed; specifically, a hybrid priority experience replay mechanism is developed and integrated within the TD3. It evaluates the significance of an experience by weighing the temporal-difference (TD) error and reward value, thus enabling the USV agent to explore optimal control policies and further accelerate the network convergence. Experimental results show that our method achieves better trajectory-tracking performance than mainstream DRL-based and model-based control approaches, and adapts to different reference trajectories and different intensities of environmental disturbances with high tracking accuracy.
Yuan Zhou 0006, Chongwei Gong, Keran Chen
IEEE Internet Things J.1
2025 Multiprior Knowledge-Guided Deep Learning Model for Kuroshio Loop Current Intrusion Prediction in South China Sea
abstract
The Kuroshio intrusion into the South China Sea via the Luzon Strait significantly influences regional ocean dynamics. However, predicting this intrusion, especially the Kuroshio Loop Current (KLC), remains challenging due to its complex mesoscale and submesoscale processes. Traditional physical models struggle to capture the nonlinear and multiscale features of the Kuroshio intrusion, while deep learning approaches face challenges in incorporating the essential physical processes that characterize the KLC. To address these challenges, we developed the Kuroshio Intrusion Forecast Network (KIFnet), a multi-prior knowledge guided deep learning model. KIFnet integrates physical oceanographic principles with data-driven predictions, enhancing its ability to capture complex ocean dynamics. KIFnet incorporates an SST-guided SSH prediction module and a vorticity-guided loss function to explicitly model thermal and dynamic features of the KLC, advancing the challenging task of forecasting KLC intrusion events. Experimental results demonstrate the model achieves an accuracy of 88% for KLC intrusion events and provides reliable predictions up to 10 days ahead. Prior limitations in KLC forecasting have constrained SCS climate modeling and marine ecosystem management. KIFnet provides accurate KLC predictions, supporting proactive climate adaptation and sustainable ecosystem strategies.
Yuan Zhou 0006, Mingzhe Yang, Keran Chen, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Contrastive learning based open-set recognition with unknown score
Yuan Zhou 0006, Songyu Fang, Shuoshi Li, Boyu Wang 0004, Sun-Yuan Kung
Knowl. Based Syst.1
2024 Weakly Supervised Video Re-Localization Through Multi-Agent-Reinforced Switchable Network
abstract
The objective of video re-localization (VRL) is to localize a successive sequence of frames, namely, the target moment, from untrimmed reference videos that semantically correspond to a given query video. During training, the weakly supervised setting of VRL provides only coarse-grained video-level rather than frame-level annotations. For the weakly supervised VRL (WS-VRL) task, obtaining effective video feature representations that can be used to evaluate the relevance between videos and localizing the accurate temporal boundaries of the target moment remain challenging. In this paper, a novel multi-agent-reinforced switchable network (MARS) is proposed to address these challenges. MARS can adaptively guide video feature encoding and moment localization using multiple learned agents. Specifically, an agent-controlled switchable encoder is used to obtain effective video feature representations, and an agent-reinforced boundary localizer is used to determine accurate localized moments through progressive refinement. Furthermore, a relevance-oriented reward generator was designed to estimate the relevance of the localized moment to the query video and assign a reward to multiple agents. The effectiveness of the proposed MARS model was verified through extensive experiments on the ActivityNet-VRL dataset.
Yuan Zhou 0006, Axin Guo, Shuwei Huo, Yu Liu 0004, Sun-Yuan Kung
IEEE Trans. Circuits Syst. Video Technol.1
2024 Geometric Variation Adaptive Network for Remote Sensing Image Change Detection
abstract
Change detection identifies surface changes on the earth by comparing two images from the same area at different times. To generate smooth change maps, a common method is fusing information from neighboring areas around each pixel. While the conventional fusion methods primarily rely on fixed regular-shaped neighboring areas, which may be inadequate in capturing the diverse and irregular geometric structures of changed ground objects. To address this limitation, we propose a novel Geometric Variation Adaptive Change Detector (GVA-CD), which adaptively adjusts the shape and size of neighboring areas based on the geometrical structure of ground objects. More specifically, we design a new geometric variation adaptive module (GVAM) as a component of GVA-CD. GVAM captures the structure of the ground objects to constructs geometrically flexible neighboring areas for each pixel, enabling the model to adapt to different ground object structures and generate discriminative difference features. We further propose a new difference measurement module to compute the difference between the features of pre-and post-change images by leveraging the adaptive neighboring areas. In addition, the GVA-CD introduces a multi-stage cross-scale fusion mechanism in both feature extraction and change map generation, to enhance the scale adaption ability of the feature extraction and change map generation. Extensive experiments on three large datasets demonstrate that our GVA-CD can outperform existing methods in change detection.
Shuwei Huo, Yuan Zhou 0006, Lei Zhang 0202, Yanjie Feng, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Geosci. Remote. Sens.2
2024 PSRNet: A Progressive Self-Refine Network for Lightweight Optical Remote Sensing Image Dehazing
abstract
In this article, we proposed a lightweight method for remote sensing image (RSI) dehazing, termed progressive self-refine network (PSR-Net). Image dehazing is an effective means to enhance the quality of images obscured by haze. Given that RSIs typically encompass extensive area with high resolution, RSI dehazing requires a lightweight model to cope with the complex RSIs. Existing methods focused mainly on improving dehazing performance without considering the computational overhead or model size. To remedy the deficiency, the proposed PSR-Net designs a lightweight framework incorporating a predehaze module (PDM) and a progressive self-refine module (PSRM). It first generates the initial dehazing results with low computation requirement, then iteratively refines the initial dehazing results by using the proposed restoration attention block (RAB). Meanwhile, considering the broad scope and varied object sizes inherent in RSIs, we designed a spatial adaptive feature extraction block which dynamically adjust its receptive field according to the input image. Extensive experiments are conducted to demonstrate that our proposed PSR-Net performs favorably against state-of-the-art methods on three popular RS image dehazing benchmark datasets.
Shuoshi Li, Yuan Zhou 0006, Sun-Yuan Kung
IEEE Trans. Geosci. Remote. Sens.2
2024 DeepSeaNet: A Bio-Detection Network Enabling Species Identification in the Deep Sea Imagery
abstract
The detection and preservation of marine biodiversity has garnered global attention. The incorporation of deep learning methodologies can elevate the efficiency of species detection. In this study, we developed a DeepSeaNet for effective localization and accurate classification of organisms based on deep-sea images, as well as for hinting at unknown organisms (new species). The DeepSeaNet fully accommodates the unique characteristics of deep-sea organisms and imaging environment, leading to remarkable advancements in fine-grained analysis and accuracy. The DeepSeaNet comprises two network components: a deep-sea Classes Detection Network (CDN) and an unsupervised Species Clustering Network (SCN). CDN is used for biological class detection and is specifically tailored for deep-sea environments. It incorporates modules for feature fusion, multi-scale analysis, and self-attention. SCN is specifically designed to detect and identify new species by utilizing the location information extracted from the CDN output results. It is composed of a feature extraction module and a clustering module. By collecting deep-sea image data from the “KeXue” Science Research Vessel, we constructed a dataset totaling 29,436 images of deep-sea organisms covering more than 500 species of deep-sea seamount organisms. This dataset serves as the foundational dataset for our experiment. As a result, our model achieves an 82.18% mean average precision for class detection and a 43.4% accuracy for species detection. Furthermore, the model has the capability to identify new species through the computation of inter-species distances.
Aiyue Liu, Yuhai Liu, Kuidong Xu, Yuan Zhou 0006, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Reconstructing 3-D Thermohaline Structures for Mesoscale Eddies Using Satellite Observations and Deep Learning
abstract
Mesoscale eddies are circular water currents found widely in the ocean and significantly impact the ocean’s circulation, water distribution, and biology. However, our comprehension of eddies’ three-dimensional (3D) structures remains constrained due to the scarcity of in-situ data. Therefore, we introduce a novel deep learning model, 3D-EddyNet, designed for reconstructing the 3D thermohaline structure of mesoscale eddies. Utilizing multi-source satellite data and Argo profiles collected from eddies in the North Pacific Ocean between 2000 and 2015, we optimized the 3D-EddyNet model by adjusting image sizes, introducing a Convolutional Block Attention Module, and incorporating eddy physical parameters. Results demonstrate remarkable accuracy, with an average root mean square error (RMSE) of 0.32 °C (0.03 psu) for temperature (salinity) within anticyclonic eddies and 0.41 °C (0.04 psu) within cyclonic eddies in the upper 1000 m. We applied 3D-EddyNet to reconstruct 3D eddy structures in the Kuroshio Extension (KE) and the Oyashio Current (OC) regions, demonstrating its capability to accurately represent the 3D thermohaline eddy structures both vertically and horizontally. The consistency in the averaged 3D eddy structures between our 3D-EddyNet and the ARMOR3D dataset in the KE and OC regions underscores the robust generalizability of our model, indicating the model’s ability to infer 3D eddy structures when Argo profiles are unavailable. The distinctive advantage offered by 3D-EddyNet enhances our ability to understand mesoscale eddy dynamics, overcoming challenges posed by the limited availability of in-situ data.
Yuan Zhou 0006, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 End-to-End Hyperspectral Image Change Detection Based on Band Selection
abstract
Change detection (CD) aims to identify differences in the same scene at different times. With the increasing amount of hyperspectral images (HSIs), more and more change detection techniques use HSIs as the raw data. HSIs often contain redundant bands, where only a few are crucial for CD while others may be detrimental. However, most existing HSI-CD methods extract features directly from full-dimensional HSIs, leading to a degradation of feature discrimination. To tackle this issue, in this paper, we propose an end-to-end hyperspectral image change detection network based on band selection (ECDBS), unlocking the potential synergy between band selection and CD. The network compromises a deep learning based band selection module and cascaded band-specific spatial attention (BSA) blocks. The band selection module selectively retains bands favourable to CD according to the importance of the bands measured based on band correlation. The BSA block tailors the feature extraction strategy for each band based on its feature distribution, allowing extracting sufficient features from each band. Experimental evaluations were conducted on three widely used HSI-CD datasets, demonstrating the effectiveness and superiority of our proposed method over other state-of-the-art techniques.
Qingren Yao, Yuan Zhou 0006, Chang Tang, Wei Xiang 0001, Gang Zheng 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Purposive Data Augmentation Strategy and Lightweight Classification Model for Small Sample Industrial Defect Dataset
abstract
Industrial defect detection plays a critical role in controlling product quality. Obtaining industrial defects with diverse and balanced classes in natural environments is often challenging. Most methods tend to uniformly augment all classes in small-sample datasets, which wastes computing resources and the classification performance is not always good. To achieve the purposive data augmentation, we propose a minority class imbalance rate (MiCIR) and an MiCIR-based data augmentation strategy that can determine the class and the number of samples to be augmented. In addition, to address the misclassification problem of classes with relatively large sample sizes, we introduce a lightweight classification model, ShcNet. We construct convolution-batchnorm-hard-swish (CBH) and convolution-batchnorm-hard-swish-convolutional block attention mechanism (CBHC) modules in ShcNet to improve classification performance. Experimental results demonstrate that our data augmentation strategy can significantly improve the classification results with generalizability across different datasets. The ShcNet outperforms the baseline models on classification accuracy while maintaining fewer parameters and model complexity.
Liyuan Lin, Shuxian Zhao, Aolin Wen, Jingpeng Yan, Ying Wang 0087, Yuan Zhou 0006
IEEE Trans. Ind. Informatics8
2024 Digital Twin Empowered Industrial IoT Based on Credibility-Weighted Swarm Learning
abstract
Driven by digital twin (DT) technology, the industrial Internet of Things (IIoT) is expanding to open up new frontiers in industrial applications. However, traditional DT modeling approaches require synchronizing massive amounts of data, resulting in high communications overhead and privacy vulnerability. To address this problem, this article proposes a novel DT architecture for IIoT, where the DT can showcase the real-time operating status of the industrial environment. Swarm learning (SL) is an emerging decentralized federated learning (FL) technique that eliminates the need of a centralized server. We present a novel credibility-weighted SL scheme to construct the DT models, which improves data security while ensuring the fairness of participants as opposed to conventional FL. In addition, we develop a DT-assisted deep reinforcement learning algorithm for simultaneously optimizing the system reliability and energy consumption of IIoT. Simulation comparisons demonstrate that the proposed scheme outperforms some state-of-the-art benchmarks in terms of both reliability and energy consumption.
Wei Xiang 0001, Jie Li 0019, Yuan Zhou 0006, Peng Cheng 0002, Jiong Jin, Kan Yu 0002
IEEE Trans. Ind. Informatics3
2024 Securing Multi-Source Domain Adaptation With Global and Domain-Wise Privacy Demands
abstract
Making available a large size of training data for deep learning models and preserving data privacy are two ever-growing concerns in the machine learning community.Multi-source domain adaptation(MDA) leverages the data information from different domains and aggregates them to improve the performance in the target task, while the privacy leakage risk of publishing models under malicious attacker for membership or attribute inference is even more complicated than the one faced by single-source domain adaptation. In this paper, we tackle the problem of effectively protecting data privacy while training and aggregating multi-source information, where each source domain enjoys an independent privacy budget. Specifically, we develop adifferentially private MDA(DPMDA) algorithm to provide domain-wise privacy protection with adaptive weighting scheme based on task similarity and task-specific privacy budget. We evaluate our algorithm on three benchmark tasks and show that DPMDA can effectively leverage different private budgets from source domains and consistently outperforms the existing private baselines with a reasonable gap with non-private state-of-the-art.
Shuwen Chai, Yutang Xiao, Jian Zhu 0001, Yuan Zhou 0006
IEEE Trans. Knowl. Data Eng.5
2024 Dynamic View Aggregation for Multi-View 3D Shape Recognition
abstract
In the field of 3D shape recognition, the view-based approach has achieved state-of-the-art performance. A major challenge that needs to be addressed by the view-based approach is how to effectively aggregate multi-view features to obtain a better 3D shape representation. Existing methods which rely on networks with static parameters for feature aggregation adversely coerce the network to learn a general feature aggregation strategy for all inputs, ignoring the diversity of input 3D shapes in real-world scenarios. In this work, we propose a novelDynamic View Aggregation NetworkcalledDVA-Netto address this challenge. DVA-Net can dynamically adjust the network parameter depending on the input 3D shapes to flexibly fuse multi-view information. The shape-specific parameter adaptation is achieved by our designedDynamic Relation-aware Aggregationmodule, dubbedDRAmodule. It is responsible for learning relations among views and adaptively integrating multi-view features. Comprehensive experiments on benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance for 3D shape classification and retrieval.
Yuan Zhou 0006, Zhongqi Sun, Shuwei Huo, Sun-Yuan Kung
IEEE Trans. Multim.1
2024 Automatic Metric Search for Few-Shot Learning
abstract
Few-shot learning (FSL) aims to learn a model that can identify unseen classes using only a few training samples from each class. Most of the existing FSL methods adopt a manually predefined metric function to measure the relationship between a sample and a class, which usually require tremendous efforts and domain knowledge. In contrast, we propose a novel model called automatic metric search (Auto-MS), in which an Auto-MS space is designed for automatically searching task-specific metric functions. This allows us to further develop a new searching strategy to facilitate automated FSL. More specifically, by incorporating the episode-training mechanism into the bilevel search strategy, the proposed search strategy can effectively optimize the network weights and structural parameters of the few-shot model. Extensive experiments on the miniImageNet and tieredImageNet datasets demonstrate that the proposed Auto-MS achieves superior performance in FSL problems.
Yuan Zhou 0006, Jieke Hao, Shuwei Huo, Boyu Wang 0004, Leijiao Ge, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.1
2023 Hierarchical full-attention neural architecture search based on search space compression
abstract
Neural architecture search (NAS) has significantly advanced the automatic design of convolutional neural architectures. However, it is challenging to directly extend existing NAS methods to attention networks because of the uniform structure of the search space and the lack of long-range feature extraction. To address these issues, we construct a hierarchical search space that allows various attention operations to be adopted for different layers of a network. To reduce the complexity of the search, a low-cost search space compression method is proposed to automatically remove the unpromising candidate operations for each layer. Furthermore, we propose a novel search strategy combining a self-supervised search with a supervised one to simultaneously capture long-range and short-range dependencies. To verify the effectiveness of the proposed methods, we conduct extensive experiments on various learning tasks, including image classification , fine-grained image recognition, and zero-shot image retrieval . The empirical results show strong evidence that our method is capable of discovering high-performance full-attention architectures while guaranteeing the required search efficiency.
Yuan Zhou 0006, Shuwei Huo, Boyu Wang 0004
Knowl. Based Syst.1
2023 Dual-branch cross-dimensional self-attention-based imputation model for multivariate time series
abstract
In real-world scenarios, partial information losses of multivariate time series degrade the time series analysis. Hence, the time series imputation technique has been adopted to compensate for the missing values. Existing methods focus on investigating temporal correlations, cross-variable correlations, and bidirectional dynamics of time series, and most of these methods rely on recurrent neural networks (RNNs) to capture temporal dependency. However, the RNN-based models suffer from the common problems of slow speed and high complexity when dealing with long-term dependency. While some self-attention-based models without any recurrent structures can tackle long-term dependency with parallel computing, they do not fully learn and utilize correlations across the temporal and cross-variable dimensions. To address the limitations of existing methods, we propose a novel so-called dual-branch cross-dimensional self-attention-based imputation (DCSAI) model for multivariate time series, which is capable of performing global and auxiliary cross-dimensional analyses when imputing the missing values. In particular, this model contains masked multi-head self-attention-based encoders aligned with auxiliary generators to obtain global and auxiliary correlations in two dimensions, and these correlations are then combined into one final representation through three weighted combinations. Extensive experiments are presented to show that our model performs better than other state-of-the-art benchmarkers on three real-world public datasets under various missing rates. Furthermore, ablation study results demonstrate the efficacy of each component of the model.
Le Fang 0001, Wei Xiang 0001, Yuan Zhou 0006, Juan Fang 0004, Lianhua Chi, ZongYuan Ge
Knowl. Based Syst.3
2023 Weakly-supervised content-based video moment retrieval using low-rank video representation
abstract
Content-based video moment retrieval (CVMR) aims to localize a successive sequence of frames in an untrimmed reference video, called target moment, that is semantically corresponding to a given query video. Current state-of-the-art CVMR methods are mainly developed using frame-level annotation, which is often quite expensive to collect. In this paper, we aim to develop a weakly-supervised CVMR method, which uses coarse-grained video-level annotations during training. Under weak supervision, video localizers require more discriminative frame-level video features. To achieve this goal, we proposed a novel prior, termed low-rank prior, based on an observation that the frame-level feature of a video should have low-rank properties. We demonstrated that the low-rank features are more discriminative and are beneficial to accurately localize the action boundaries. To produce a low-rank feature, we designed a low-rank feature reconstruction (LFR) operator. A new differentiable matrix decomposition approach is proposed to generate the low-rank reconstruction of the input matrix, meanwhile ensuring that the matrix decomposition process is differentiable. Based on the LFR, we developed a new weakly-supervised CVMR model which produces low-rank video representation and performs semantic consistency measures to discover the semantically matched segment in the reference video to the query video. Extensive experiments demonstrate that our method outperforms state-of-the-art weakly-supervised methods consistently and even achieves competing performance to fully-supervised baselines.
Shuwei Huo, Yuan Zhou 0006, Wei Xiang 0001, Sun-Yuan Kung
Knowl. Based Syst.2
2023 M2SCN: Multi-Model Self-Correcting Network for Satellite Remote Sensing Single-Image Dehazing
abstract
Remote sensing (RS) image dehazing is an effective means to enhance the quality of hazy RS images. However, existing dehazing methods are ineffective in dealing with nonhomogeneous RS haze scenes. To tackle this deficiency, we design a multi-model joint estimation (M2JE) module and a self-correcting (SC) module to construct a unified end-to-end network for RS image dehazing, termed the multi-model SC network (M2SCN). Specifically, the M2JE module regards the dehazing process as a multi-model ensemble problem, so as to improve the generalization ability of the model. The SC module can gradually correct the error in the intermedia features extracted by the network, thus enabling the network to deal with nonhomogeneous hazy images. Extensive experiments are conducted to demonstrate that our proposed M2SCN performs favorably against state-of-the-art methods on popular RS image dehazing benchmark datasets.
Shuoshi Li, Yuan Zhou 0006, Wei Xiang 0001
IEEE Geosci. Remote. Sens. Lett.2
2023 Exploiting Frequency-Domain Information of GNSS Reflectometry for Sea Surface Wind Speed Retrieval
abstract
Global navigation satellite system reflectometry (GNSS-R) Delay-Doppler map measures the sea surface roughness, which has recently been applied to retrieve sea surface wind speed. However, current studies on GNSS-R wind speed retrieval only use the spatial domain of the delay-Doppler map without considering the variations patterns in the map, which is regarded as frequency domain information of the map. In this study, we propose a joint frequency-spatial-domain wind speed retrieval network (FSNet) based on reflectivity data provided by the Cyclone Global Navigation Satellite System (CyGNSS) mission. We construct a matchup dataset between the CyGNSS satellite data and ECMWF model data from January 1, 2018, to December 31, 2019. The wind speed range is 0–25 m/s. Using the proposed FSNet, frequency and spatial features are simultaneously extracted. The frequency domain feature supplements the spatial-domain information of the mid and high-level features in the neural network. Rather than directly concatenating the frequency-domain features with the spatial-domain features, we designed a feature fusion module to fuse frequency and spatial features for wind speed retrieval adaptively. Experiments show that our FSNet wind speed retrieval has a root mean square error (RMSE) of 1.63 m/s for a wind range of 0-25 m/s. This accuracy is 25.4% better than the operational algorithm provided by the CyGNSS Level 2 wind speed product. For a higher wind range of 16-25m/s, FSNet performed even better, improving the RMSE by 31%.
Keran Chen, Yuan Zhou 0006, Shuoshi Li, Ping Wang 0015, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Progressive Learning for Unsupervised Change Detection on Aerial Images
abstract
This article focuses on unsupervised methods for optical aerial image change detection. Existing unsupervised change detection techniques are mainly categorized as patch-based methods and transfer-learning-based methods. However, the first type ignores the spatial information in the images, and the second type may introduce new errors due to knowledge extracted from additional datasets. To effectively tackle these problems, we propose an unsupervised progressive learning framework (UPLF). We first use original estimated change maps as the labeled samples and choose the reliable regions from samples to train the network. We then propose a progressive learning method to expand the reliable labeled region. Briefly, we apply a label selection filter to filter out incorrect change information from the regions to help rectify incorrect labeling in the regions. This leads to a more reliable labeled region and thus, in turn, more accurate detection results. Compared with the patch-based and transfer-learning-based unsupervised techniques, our method takes the entire map as the training sample to avoid the problem associated with using small patches; moreover, our iterative and progressive methods further enhance the change detection performance without involving external knowledge. Indeed, based on our experimental results on the real datasets, the proposed method demonstrates highly competitive performance compared with the state-of-the-art.
Yuan Zhou 0006, Xiangrui Li, Keran Chen, Sun-Yuan Kung
IEEE Trans. Geosci. Remote. Sens.1
2023 Hyperspectral Band Selection With Iterative Graph Autoencoder
abstract
Hyperspectral band selection (BS) is an important task for hyperspectral image (HSI) processing, which aims to select a discriminative and low-redundant band subset. As a significant cue for BS, structure information describes the cross-band correlation which brings the redundancy of HSI. Existing methods model structure information via manual rule-based graph construction. However, such a graph construction method fails to model complex and diverse structural relationships of HSI data. To address this problem, we propose a data-driven method, named iterative graph auto-encoder for band selection (IGAEBS). It adaptively captures structure information by a data-specific automatic construction process, rather than by a fixed empirical design. Specifically, we propose a new unsupervised pretext task to train graph convolution neural network to extract HSI features. These features are used to construct a graph to represent the structural relationships among bands. To enhance the reliability of the graph, we further design an iterative graph improvement mechanism to progressively refine the structure representation. Using the derived graph, we partition the bands into several clusters and select a representative band from each cluster. During the selection process, both intra- and inter-cluster information are considered to improve the discriminativeness of band subset. Extensive experiments are conducted on three public data sets to validate the superiority of the proposed method compared to other state-of-the-art methods.
Yuan Zhou 0006, Qingren Yao, Shuwei Huo, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 PFONet: A Progressive Feedback Optimization Network for Lightweight Single Image Dehazing
abstract
Image dehazing is an effective means to enhance the quality of images captured in foggy or hazy weather conditions. However, existing image dehazing methods are either ineffective in dealing with complex haze scenes, or incurring too much computation. To overcome these deficiencies, we propose a progressive feedback optimization network (PFONet) which is lightweight yet effective for image dehazing. The PFONet consists of a multi-stream dehazing module and a progressive feedback module. The progressive feedback module feeds the output dehazed image back to the intermedia features extracted by the network, thus enabling the network to gradually reconstruct a complex degraded image. Considering both the effectiveness and efficiency of the network, we also design a lightweight hybrid residual dense block serving as the basic feature extraction module of the proposed PFONet. Extensive experimental results are presented to demonstrate that the proposed model outperforms its state-of-the-art single-image dehazing competitors for both synthetic and real-world images.
Shuoshi Li, Yuan Zhou 0006, Wenqi Ren, Wei Xiang 0001
IEEE Trans. Image Process.2
2023 On the Benefits of Two Dimensional Metric Learning
abstract
In this paper, we study two dimensional metric learning (2DML) for matrix data from both theoretical and algorithmic perspectives. We first investigate the generalization bounds of 2DML based on the notion of Rademacher complexity, which theoretically justifies the benefits of learning from matrices directly. Furthermore, we present a novel boosting-based algorithm that scales well with the feature dimension. Finally, we introduce an efficient rank-one correction algorithm, which is tailored to our boosting learning procedure to produce a low-rank solution to 2DML. As our algorithm works directly on the data in matrix representation, it scales well with the feature dimension, keeps the structure and dependence in the data, and has a more compact structure and much fewer parameters to optimize. Extensive evaluations on several benchmark data sets also empirically verify the effectiveness and efficiency of our algorithm.
Di Wu 0044, Fan Zhou 0006, Boyu Wang 0004, Qicheng Lao, Chiman Wong, Changjian Shui, Yuan Zhou 0006, Feng Wan 0003
IEEE Trans. Knowl. Data Eng.7
2023 Semantic Relevance Learning for Video-Query Based Video Moment Retrieval
abstract
The task of video-query based video moment retrieval (VQ-VMR) aims to localize the segment in the reference video, which matches semantically with a short query video. This is a challenging task due to the rapid expansion and massive growth of online video services. With accurate retrieval of the target moment, we propose a new metric to effectively assess the semantic relevance between the query video and segments in the reference video. We also develop a new VQ-VMR framework to discover the intrinsic semantic relevance between a pair of input videos. It comprises two key components: a Fine-grained Feature Interaction (FFI) module and a Semantic Relevance Measurement (SRM) module. Together they can effectively deal with both the spatial and temporal dimensions of videos. First, the FFI module computes the semantic similarity between videos at a local frame level, mainly considering the spatial information in the videos. Subsequently, the SRM module learns the similarity between videos from a global perspective, taking into account the temporal information. We have conducted extensive experiments on two key datasets which demonstrate noticeable improvements of the proposed approach over the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Ruolin Wang, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Multim.2
2022 YFormer: A New Transformer Architecture for Video-Query Based Video Moment Retrieval
Shuwei Huo, Yuan Zhou 0006
PRCV (3)2
2022 Robust object tracking via deformation samples generator
Xuesong Gao, Yuan Zhou 0006, Shuwei Huo, Zizi Li, Keqiu Li
J. Vis. Commun. Image Represent.2
2022 Cross-Scale Residual Network: A General Framework for Image Super-Resolution, Denoising, and Deblocking
abstract
In general, image restoration involves mapping from low-quality images to their high-quality counterparts. Such optimal mapping is usually nonlinear and learnable by machine learning. Recently, deep convolutional neural networks have proven promising for such learning processing. It is desirable for an image processing network to support well with three vital tasks, namely: 1) super-resolution; 2) denoising; and 3) deblocking. It is commonly recognized that these tasks have strong correlations, which enable us to design a general framework to support all tasks. In particular, the selection of feature scales is known to significantly impact the performance on these tasks. To this end, we propose the cross-scale residual network to exploit scale-related features among the three tasks. The proposed network can extract spatial features across different scales and establish cross-temporal feature reusage, so as to handle different tasks in a general framework. Our experiments show that the proposed approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations for multiple image restoration tasks.
Yuan Zhou 0006, Xiaoting Du, Mingfei Wang, Shuwei Huo, Yeda Zhang, Sun-Yuan Kung
IEEE Trans. Cybern.1
2022 Dual-Branch Neural Network for Sea Fog Detection in Geostationary Ocean Color Imager
abstract
Sea fog significantly threatens the safety of maritime activities. This paper develops a sea fog dataset (SFDD) and a dual branch sea fog detection network (DB-SFNet). We investigate all the observed sea fog events in the Yellow Sea and the Bohai Sea (118.1°E-128.1°E, 29.5°N-43.8°N) from 2010 to 2020, and collect the sea fog images for each event from the Geostationary Ocean Color Imager (GOCI) to comprise the dataset SFDD. The location of the sea fog in each image in SFDD is accurately marked. The proposed dataset is characterized by a long-time span, large number of samples, and accurate labeling, that can substantially improve the robustness of various sea fog detection models. Furthermore, this paper proposes a dual branch sea fog detection network to achieve accurate and holistic sea fog detection. The poporsed DB-SFNet is composed of a knowledge extraction module and a dual branch optional encoding decoding module. The two modules jointly extracts discriminative features from both visual and statistical domain. Experiments show promising sea fog detection results with an F1-score of 0.77 and a critical success index of 0.63. Compared with existing advanced deep learning networks, DB-SFNet is superior in detection performance and stability, particularly in the mixed cloud and fog areas.
Yuan Zhou 0006, Keran Chen, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Joint Frequency-Spatial Domain Network for Remote Sensing Optical Image Change Detection
abstract
Change detection for remote sensing images involves detecting regional surface changes of interest between two images taken of the same geographical area but at different times. In image processing, the spatial domain uses grayscale values to describe an image. The frequency is directly related to the spatial change rate, so the frequency domain can be intuitively associated with patterns of intensity variations in the image. These two domains provide different perspectives for image interpretation. Most existing deep-learning-based methods formulate change detection as a pixel-wise binary classification problem and utilize various strategies to extract information in the spatial domain. However, they rarely pay attention to the rich information in the frequency domain. To address this problem, we propose an end-to-end joint frequency-spatial domain network (JFSDNet) to implement remote sensing optical image change detection. Specifically, we introduce frequency information into the change detection to supplement the loss of image details caused by down-sampling. In addition, we employ a frequency selection module to adaptively discriminate and choose frequency clues by reducing the complexity of the frequency features. The JFSDNet is applied to two publicly available datasets: the CDD dataset and the LEVIR-CD dataset. Compared with other methods, both visual interpretation and quantitative assessment confirmed that our proposed method achieved a favorable performance.
Yuan Zhou 0006, Yanjie Feng, Shuwei Huo, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Multilayer Fusion Recurrent Neural Network for Sea Surface Height Anomaly Field Prediction
abstract
Sea surface height anomaly (SSHA) is vitally important for climate and marine ecosystems. This article develops a multilayer fusion recurrent neural network (MLFrnn) to achieve an accurate and holistic prediction of the SSHA field, given only as a series of past SSHA observations. The proposed approach learns long-term dependencies within the SSHA time series and spatial correlations among neighboring and remote regions. A new multilayer fusion cell as the building block of the MLFrnn model was designed, which fully fused spatial and temporal features. The daily average satellite altimeter SSHA data in the South China Sea from January 1, 2001, to May 13, 2019, were used to train and test the model. We performed a 21-day ahead SSHA prediction and our MLFrnn model has very high accuracy, with a root mean square error (RMSE) of 0.027 m. Compared with existing deep learning networks, the proposed model was superior both in prediction performance and stability, especially on the wide-scale and long-term predictions.
Yuan Zhou 0006, Chang Lu 0001, Keran Chen, Xiaofeng Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Exploiting Operation Importance for Differentiable Neural Architecture Search
abstract
Recently, differentiable neural architecture search (NAS) methods have made significant progress in reducing the computational costs of NASs. Existing methods search for the best architecture by choosing candidate operations with higher architecture weights. However, architecture weights cannot accurately reflect the importance of each operation, that is, the operation with the highest weight might not be related to the best performance. To circumvent this deficiency, we propose a novel indicator that can fully represent the operation importance and, thus, serve as an effective metric to guide the model search. Based on this indicator, we further develop a NAS scheme for "exploiting operation importance for effective NAS" (EoiNAS). More precisely, we propose a high-order Markov chain-based strategy to slim the search space to further improve search efficiency and accuracy. To evaluate the effectiveness of the proposed EoiNAS, we applied our method to two tasks: image classification and semantic segmentation. Extensive experiments on both tasks provided strong evidence that our method is capable of discovering high-performance architectures while guaranteeing the requisite efficiency during searching.
Yuan Zhou 0006, Xukai Xie, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.1
2021 Shape autotuning activation function
Yuan Zhou 0006, Shuwei Huo, Sun-Yuan Kung
Expert Syst. Appl.1
2021 Image super-resolution based on dense convolutional auto-encoder blocks
Yuan Zhou 0006, Yeda Zhang, Xukai Xie, Sun-Yuan Kung
Neurocomputing1
2021 Anypath Routing Protocol Design via Q-Learning for Underwater Sensor Networks
abstract
As a promising technology in the Internet of Underwater Things, underwater sensor networks (UWSNs) have drawn a widespread attention from both academia and industry. However, designing a routing protocol for UWSNs is a great challenge due to high energy consumption and large latency in the underwater environment. This article proposes a Q-learning -based localization-free anypath routing (QLFR) protocol to prolong the lifetime as well as reduce the end-to-end delay for UWSNs. Aiming at optimal routing policies, the Q-value is calculated by jointly considering the residual energy and depth information of sensor nodes throughout the routing process. More specifically, we define two reward functions (i.e., depth-related and energy-related rewards) for Q-learning with the objective of reducing latency and extending network lifetime. In addition, a new holding time mechanism for packet forwarding is designed according to the priority of forwarding candidate nodes. Furthermore, mathematical analyses are presented to analyze the performance and computational complexity of the proposed routing protocol. Extensive simulation results demonstrate the superiority performance of the proposed routing protocol in terms of the end-to-end delay and the network lifetime.
Yuan Zhou 0006, Wei Xiang 0001
IEEE Internet Things J.1
2021 Adversarial Learning for Multiscale Crowd Counting Under Complex Scenes
abstract
In this article, a multiscale generative adversarial network (MS-GAN) is proposed for generating high-quality crowd density maps of arbitrary crowd density scenes. The task of crowd counting has many challenges, such as severe occlusions in extremely dense crowd scenes, perspective distortion, and high visual similarity between the pedestrians and background elements. To address these problems, the proposed MS-GAN combines a multiscale convolutional neural network (generator) and an adversarial network (discriminator) to generate a high-quality density map and accurately estimate the crowd count in complex crowd scenes. The multiscale generator utilizes the fusion features from multiple hierarchical layers to detect people with large-scale variation. The resulting density map produced by the multiscale generator is processed by a discriminator network trained to solve a binary classification task between a poor quality density map and real ground-truth ones. The additional adversarial loss can improve the quality of the density map, which is critical to accurately estimate the crowd counts. The experiments were conducted on multiple datasets with different crowd scenes and densities. The results showed that the proposed method provided better performance compared to current state-of-the-art methods.
Yuan Zhou 0006, Jianxing Yang, Sun-Yuan Kung
IEEE Trans. Cybern.1
2021 Temporal Action Localization Using Long Short-Term Dependency
abstract
Temporal action localization in untrimmed videos is an important but difficult task. Difficulties are encountered in the application of existing methods when modeling the temporal structures of videos. In the present study, we develop a novel method, referred to as the Gemini Network, for effective modeling of temporal structures and achieving high-performance temporal action localization. The significant improvements afforded by the proposed method are due to three major factors. First, temporal dependencies are explicitly distinguished as long-term temporal dependencies and short-term temporal dependencies and are separately captured by two dedicated subnets. Second, a long-range temporal dependency capture module combined with a self-adaptive pooling module is proposed to capture long-term temporal dependency. Third, the proposed method uses auxiliary supervision, with the auxiliary classifier losses affording additional constraints for improving the modeling capability of the network. As a demonstration of its effectiveness, the Gemini Network is used to achieve a state-of-the-art temporal action localization performance on two challenging datasets, namely, THUMOS14 and ActivityNet.
Yuan Zhou 0006, Ruolin Wang, Sun-Yuan Kung
IEEE Trans. Multim.1
2020 A Feature Pair Fusion And Hierarchical Learning Framework For Video Re-Localization
abstract
Video re-localization has become an emerging research topic nowadays but existing methods still have many deficiencies. The existing deficiencies mainly lie in the interference caused by the irrelevant information in the input reference video and the ignorance of the correlation between query and reference video features. Therefore, we present a novel framework named Semantic Relevance Learning Network to address these shortcomings. First, we extract effective proposals from reference video as new inputs to reduce interference from irrelevant video frames. Second, two key components of our proposed model, the Attention-based Fusion Tensor and Semantic Relevance Measurement, jointly explore the intrinsic correlation between video feature pairs and finally get a score as measurement. To better evaluate our proposed model, we reorganize Thumos14 to obtain another new dataset for the video re-localization task. For both ActivityNet and Thumos14, our model achieves the best performance reported so far.
Ruolin Wang, Yuan Zhou 0006
ICIP2
2020 Exploring Highly Efficient Compact Neural Networks For Image Classification
abstract
Group convolution works well with many lightweight convolutional neural networks (CNNs) that can effectively reduce the number of parameters and computational cost. However, feature maps of different groups cannot communicate, which restricts their representation capability. To address this issue, in this work, we propose a novel convolution operation named Hierarchical Group Convolution (HGC) for creating computationally efficient neural networks. Different from standard group convolution which blocks the inter-group information exchange and induces the severe performance degradation, HGC hierarchically fuses the feature maps from each group and leverages the inter-group information effectively. Taking advantage of the proposed operation, we introduce an efficient compact network named HGCNet. Extensive experimental results on image classification task demonstrate that HGCNet obtain significant reduction of computational cost and the number of parameters, while achieving comparable performance over the prior CNN architectures designed for mobile devices.
Xukai Xie, Yuan Zhou 0006, Sun-Yuan Kung
ICIP2
2020 Underwater Image Processing by an Adversarial Network with Feedback Control
Kangming Yan, Yuan Zhou 0006
PRCV (1)2
2020 Depth image super-resolution based on joint sparse coding
Yuan Zhou 0006, Yeda Zhang, Aihua Wang
Pattern Recognit. Lett.2
2020 Adaptive Irregular Graph Construction-Based Salient Object Detection
abstract
Saliency detection represents a vital pre-processing stage of computer vision. Most existing propagation-based salient object detection methods construct a k-regular graph for saliency propagation. Applying a regular graph to a vast smooth region is potentially prone to unnecessary or prolonged propagation errors, leading to the excessive highlighting of the background regions. To mitigate such problems, we substitute the conventional k-regular graph with an adaptive irregular graph for saliency value propagation, thereby avoiding unnecessary iterations over a vast smooth region. We first perform a clustering analysis based on the smoothness, color, and other features of regions. The new graph boosts an adaptive link density by considering the clustering result. In addition, we propose a seeding strategy for the propagation. Based on our experimental studies of six major benchmark datasets, our method performed favorably against the other state-of-the-art methods, both quantitatively and qualitatively.
Yuan Zhou 0006, Shuwei Huo, Chunping Hou, Sun-Yuan Kung
IEEE Trans. Circuits Syst. Video Technol.1
2019 QLFR: A Q-Learning-Based Localization-Free Routing Protocol for Underwater Sensor Networks
abstract
Designing a routing protocol for underwater sensor networks is a great challenge due to characteristics of high energy consumption and high latency. This paper investigates a Q-learning-based localization-free routing protocol (QLFR) to prolong the lifetime as well as reduce the end-to-end delay for underwater sensor networks. Aiming to seek optimal routing policies, Q- value is calculated by jointly considering residual energy and depth information of sensor nodes throughout the routing process. More specifically, we define two cost functions (depth-related cost and energy-related cost) for Q-learning, in order to reduce delay and extend the network lifetime. In addition, a holding time mechanism for packet forwarding is designed according to the priority of forwarding nodes. The key contribution lies in: 1) a novel Q-learning-based routing protocol for UWSNs; 2) a new holding time mechanism for packet forwarding; and 3) a packet- delivery-ratio-based scheme to further reduce unnecessary transmissions. Extensive simulation results demonstrate superiority performance of our routing protocol in terms of reducing end-to-end delay and extending the network lifetime.
Yuan Zhou 0006, Wei Xiang 0001
GLOBECOM1
2019 Rethinking Temporal Structure Modeling Method for Temporal Action Localization
abstract
Temporal action localization in untrimmed videos is an important but difficult task. Difficulties are encountered in the application of existing methods when modeling temporal structures of videos. In the present study, we developed a novel method, referred to as Gemini Network, for effective modeling of temporal structures and achieving high-performance temporal action localization. The significant improvements afforded by the proposed method are attributable to three major factors. First, the developed network utilizes two sub-nets for effective modeling of temporal structures. Second, three parallel feature extraction pipelines are used to prevent interference between the extractions of different stage features. Third, the proposed method utilizes auxiliary supervision, with the auxiliary classifier losses affording additional constraints for improving the modeling capability of the network. As a demonstration of its effectiveness, the Gemini Network was used to achieve state-of-the-art temporal action localization performance on two challenging datasets, namely, THUMOS14 and ActivityNet.
Jianxing Yang, Yuan Zhou 0006, Sumei Li
ICIP3
2019 Dense-Connected Residual Network for Video Super-Resolution
abstract
Recent research has shown that the performances of super-resolution methods can be significantly boosted using deep convolutional neural networks. However, current superresolution methods continue to exhibit relatively low performances for video, partly because they ignore certain crucial inter-frame information from the original low-resolution frame sequence or the hierarchical features of deep networks. In this paper, we propose a novel method for video super resolution named dense-connected residual network (DCRnet) to address the above drawbacks. The DCRnet can preserve the low -frequency contents of motion compensated frames, and facilitate the restoration of high-frequency details by exploiting the hierarchical features from all the convolutional layers. Specifically, we propose a dense-connected residual block (DCRB) as a basic component. The output of one DCRB is the compressed concatenation of all preceding DCRBs features and each residual block features of the current DCRB. Extensive experimentation demonstrates that our method is superior to the current state-of-the-art methods in both quantitative and qualitative metrics.
Xiaoting Du, Yuan Zhou 0006, Yanfang Chen, Yeda Zhang, Jianxing Yang, Dou Jin
ICME2
2019 Unsupervised Global Manifold Alignment for Cross-Scene Hyperspectral Image Classification
Yuan Zhou 0006, Dou Jin
PRCV (2)2
2019 Gemini Network for Temporal Action Localization
Ying Wang 0087, Yuan Zhou 0006
PRCV (2)3
2019 Underwater Image Restoration Using Color-Line Model
abstract
Underwater images typically suffer from low visibility and severe colorcast due to scattering and absorption. In this letter, a novel method is proposed to handle the scattering and absorption problems of light with different wavelengths based on the color-line model. We filter out image patches that exhibit the characteristics of the color-line prior and recover the color line of the patches. Then, the local transmission for each patch is estimated based on the offsets of the color lines along the background-light vector from the origin. We also develop an optimization function to derive the local transmission and to obtain the solution in the underwater environment. Experimental results are presented to show that the proposed method can produce high-quality underwater images with relatively genuine colors, natural appearance, and improved contrast and visibility.
Yuan Zhou 0006, Kangming Yan, Liyang Feng, Wei Xiang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Semi-Supervised Salient Object Detection Using a Linear Feedback Control System Model
abstract
To overcome the challenging problems in saliency detection, we propose a novel semi-supervised classifier which makes good use of a linear feedback control system (LFCS) model by establishing a relationship between control states and salient object detection. First, we develop a boundary homogeneity model to estimate the initial saliency and background likelihoods, which are regarded as the labeled samples in our semi-supervised learning procedure. Then in order to allocate an optimized saliency value to each superpixel, we present an iterative semi-supervised learning framework which integrates multiple saliency cues and image features using an LFCS model. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. This paper also covers comprehensive simulation study based on public datasets, which demonstrates the superiority of the proposed approach.
Yuan Zhou 0006, Shuwei Huo, Wei Xiang 0001, Chunping Hou, Sun-Yuan Kung
IEEE Trans. Cybern.1
2019 Salient Object Detection via Fuzzy Theory and Object-Level Enhancement
abstract
This paper proposes a bottom-up saliency detection method via effective integration of regional saliency measure and object-level information using fuzzy theory. First, we generate an initial saliency map by fusing multiple prior maps. Second, to emphasize the object-level concept of saliency, we further generate many object proposals of the input image. A fuzzy set theory is then applied to measure the objectness score of the object proposals and integrate them into an objectness map. Third, an optimization framework is proposed to effectively fuse various prior saliency cues and object-level information to produce a clean and uniform saliency map as well as to maintain the salient object completeness. Experimental studies in several benchmark datasets confirmed the superiority of the proposed method over state-of-the-art saliency detection methods.
Yuan Zhou 0006, Ailing Mao, Shuwei Huo, Jianjun Lei 0001, Sun-Yuan Kung
IEEE Trans. Multim.1
2019 Semisupervised Learning Based on a Novel Iterative Optimization Model for Saliency Detection
abstract
In this paper, we propose a novel iterative optimization model for bottom-up saliency detection. By exploring bottom-up saliency principles and semisupervised learning approaches, we design a high-performance saliency analysis method for wide ranging scenes. The proposed algorithm consists of two stages: 1) we develop a boundary homogeneity model to characterize the general position and the contour of the salient objects and 2) we propose a novel iterative optimization model, termed gradual saliency optimization, for further performance improvement. Our main contribution falls on the second stage, where we propose an iterative framework with self-repairing mechanisms for refining saliency maps. In this framework, we further develop a more comprehensive optimization function applying a novel semisupervised learning scheme to enhance the traditional saliency measure. More elaborately, the iterative method can gradually improve the output in each iteration and finally converge to high-quality saliency maps. Based on our experiments on four different public data sets, it can be demonstrated that our approach significantly outperforms the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.2
2018 Cross-Layer Design for Network Lifetime Maximization in Underwater Wireless Sensor Networks
abstract
This paper investigates the cross-layer design problem with the goal of maximizing the network lifetime for energy-constrained underwater wireless sensor networks (UWSNs). We first jointly consider link scheduling, transmission power and transmission rate in a proposed optimization problem with the adoption of time division multiple access (TDMA) schedules. Then, we propose an iterative algorithm to solve the optimization problem. It alternates between (1) link scheduling and (2) computation of transmission powers and transmission rates. In fact, the convergence of such iterative algorithm can be mathematically and empirically supported. We evaluate our algorithm for several network topologies. Extensive simulation results demonstrate the superiority of the proposed approach.
Yuan Zhou 0006, Yu Hen Hu, Boyu Wang 0004, Sun-Yuan Kung
ICC2
2018 Super-resolution Imaging Based on Global Interpolation and Structural Similarities
abstract
In this paper, we propose a double dictionary learning method for image super-resolution (SR) reconstruction. Different from existing dictionary learning based super-resolution, we combine both self-similarity and external images to construct a double dictionary learning method. A new optimization model is established using self-similarities and external-similarities as regularization terms. Furthermore, we propose a global interpolation method to reconstruct an accurate initial estimation at the edges. Experimental results show that the proposed algorithm can produce high-quality reconstruction results both perceptually and quantitatively in terms of peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), as compared to existing algorithms.
Yuan Zhou 0006, Shuwei Huo, Sun-Yuan Kung
ICPR1
2018 Heterogeneous image change detection using Deep Canonical Correlation Analysis
abstract
Cross-sensor change detection is nowadays of paramount importance for earth observation applications. Most current change detection techniques are based on homogeneous input images. Due to the detailed and complementary spatial and spectral information, heterogeneous images change detection has become an active research topic. Change detection models need effective feature representations to estimate changes of interest. Although great progress has been made, existing approaches mainly focus on shallow models, which only extracting handcrafted low-level features. To this end, this paper proposes a novel heterogeneous change detection method using deep canonical correlation analysis (DCCA). Specifically, the two heterogeneous images are transformed via a deep neural network, and they are projected in the common latent space in the output layer. Experiments on the commonly used homogenous and heterogeneous image datasets demonstrate the superiority of the proposed method compared with the traditional approaches.
Yuan Zhou 0006, Liyang Feng
ICPR2
2018 Multi-scale Generative Adversarial Networks for Crowd Counting
abstract
We investigate generative adversarial networks as an effective solution to the crowd counting problem. These networks not only learn the mapping from crowd image to corresponding density map, but also learn a loss function to train this mapping. There are many challenges to the task of crowd counting, such as severe occlusions in extremely dense crowd scenes, perspective distortion, and high visual similarity between pedestrians and background elements. To address these problems, we proposed multi-scale generative adversarial network to generate high-quality crowd density maps of arbitrary crowd density scenes. We utilized the adversarial loss from discriminator to improve the quality of the estimated density map, which is critical to accurately predict crowd counts. The proposed multi-scale generator can extract multiple hierarchy features from the crowd image. The results showed that the proposed method provided better performance compared to current state-of-the-art methods.
Jianxing Yang, Yuan Zhou 0006, Sun-Yuan Kung
ICPR2
2018 Blurred Image Region Detection based on Stacked Auto-Encoder
abstract
In this study, we address a fundamental yet challenging problem on detection and classification of blurred regions in partially blurred images. we propose to learn a latent feature representation with stacked auto-encoder (SAE) network to perform blur region detection. Most previous approaches focus on extracting a few blur features in image gradient, Fourier domain, and data-driven local filters. We extract a latent high-level feature representation from such low-level features using the stacked auto-encoder network, thereby improve the accuracy of blur region classification. This high accuracy enables us to successfully separate the clear and blurred regions. Experimental results demonstrate that the proposed method significantly outperforms the state-of-the-arts methods in detecting and classifying blur regions in partially blurred images.
Yuan Zhou 0006, Jianxing Yang, Sun-Yuan Kung
ICPR1
2018 Iterative Feedback Control-Based Salient Object Segmentation
abstract
In this paper, we establish a mathematical model that relates the control states and the saliency values in salient object detection. We show that a linear feedback control system (LFCS) is amenable to saliency detection tasks owing to its functional properties. This inspired us to employ an LFCS to detect salient objects in static images. Based on the novel iteration method, the system gradually converges to an optimized stable state, which is associated with an accurate saliency map. In addition, to initialize the system, we propose a so-called boundary homogeneity based on a priori knowledge of the boundary in order to estimate the background likelihood and indirectly obtain a foreground (saliency) map. The experimental results indicate that such a feedback control model can offer significant improvement in salient object detection performance.
Shuwei Huo, Yuan Zhou 0006, Jianjun Lei 0001, Nam Ling, Chunping Hou
IEEE Trans. Multim.2
2017 Salient object detection via a linear feedback control system
abstract
Linear feedback control systems (LFCS) have been widely applied in signal analysis, filtering, and error correction. Many functional properties of LFCS are amenable to numerous object recognition and detection tasks. In fact, there exists an intimate relationship between control states and salient values. This prompts us to adopt the linear feedback control system to detect salient object in static images. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. In addition, to initialize the system, we propose the so called boundary homogeneity based on a priori knowledge on the boundary to estimate the background likelihood and indirectly depict a foreground (saliency) map. By our experimental results, we demonstrates that such feedback control model can bring about noticable improvement in salient object detection.
Shuwei Huo, Yuan Zhou 0006, Sun-Yuan Kung
ICIP2
2017 Depth-weighted correlation method for visual tracking with occlusion detection
abstract
Despite the significant progress, it remains a challenging task for a tracker to distinguish a target from the background when the target is occluded. In this paper, we propose a new tracking method, named as depth-weighted correlation method(DWCM), to handle heavy occlusion. The proposed method uses depth cues as the weights of candidate objects and applies the framework of spatio-temporal context (STC). We also propose a scale update scheme for DWCM, so as to obtain an appropriate scale for the target. Encouraging experimental results show that the proposed tracker obtains state-of-the-art results and handles occlusion better than competing tracking methods.
Chenghao Li 0004, Yuan Zhou 0006, Bo Cut, Chunping Hou
ICIP2
2017 Label propagation based saliency detection via graph design
abstract
Saliency detection has been widely used as the pre-processing of the computer-vision tasks. Existing propagation based saliency detection methods simply select a k-regular graph for saliency propagation, which usually leads to the mistaken highlighting of the long-range smooth background regions. In this paper, we design a novel graph for label propagation based saliency detection by considering the local consistency and the global symmetry of the image scene and updating the graph model based on smoothness assumption and cluster assumption. Then, we label the reliable seeds and propagate the saliency value through the designed graph. On two widely used large open benchmark data sets, the proposed method significantly outperforms thirteen state-of-the-arts under either quantitative or qualitative evaluation.
Yuan Zhou 0006, Shuwei Huo, Chunping Hou
ICIP2
2017 Joint nonlocal sparse representation for depth map super-resolution
abstract
Depth image super-resolution reconstruction has gained significant popularity due to its practicability. However, conventional depth image super-resolution reconstruction methods access high frequency information either from a high-resolution depth image database or from a high-resolution color image of the same scene, which is limited in specific applications. In this paper, a novel joint nonlocal sparse representation model is proposed, which is able to capture the interdependency of low-resolution depth and intensity information. As a relative new and not well addressed problem, we reconstruct a high-resolution depth image from a single low-resolution depth image with a low-resolution color image as reference. Experiment results demonstrate that the proposed method outperforms many current state-of-the-art depth map super-resolution approaches on both visual effects and objective image quality.
Yeda Zhang, Yuan Zhou 0006, Aihua Wang, Chunping Hou
ICIP2
2017 Semi-supervised saliency classifier based on a linear feedback control system model
abstract
Linear feedback control systems (LFCS) are amenable to numerous object recognition and detection tasks on account of its functional properties in signal filtering and error correction. In fact, there exists an intimate relationship between control states and salient values. Therefore, we propose a novel semi-supervised classifier which makes use of linear feedback control theory to improve saliency detection performance. First, we develop a boundary homogeneity model to estimate the initial saliency and background likelihoods, which may lead to the labeled samples in our semi-supervised learning procedure. Then in order to allocate an optimized saliency value to each superpixel, we present an iterative semi-supervised learning framework which integrates multiple saliency cues and image features using a LCSF model. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. Based on our experiments on public datasets, it can be demonstrated that our approach significantly outperforms the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Sun-Yuan Kung
IJCNN2
2017 Hyperspectral and Multispectral Image Fusion Based on Local Low Rank and Coupled Spectral Unmixing
abstract
Hyperspectral images (HSIs) usually have high spectral and low spatial resolution. Conversely, multispectral images (MSIs) usually have low spectral and high spatial resolution. The fusion of HSI and MSI aims to create spectral images with high spectral and spatial resolution. In this paper, we propose a fusion algorithm by combining linear spectral unmixing with the local low-rank property. By taking advantage of the local low-rank property, we first partition the corresponding spectral image into patches. For each patch pair, we cast the fusion problem as a coupled spectral unmixing problem that extracts the abundance and the endmembers of MSI and HSI, respectively. It then updates the abundance and the endmember through an alternating update algorithm. In fact, the convergence of the alternative update algorithm can be mathematically and empirically supported. We also propose a multiscale postprocessing procedure to combine fusion results obtained under different patch sizes. In experiments on three data sets, the proposed fusion algorithms outperformed state-of-the-art fusion algorithms in both spatial and spectral domains.
Yuan Zhou 0006, Liyang Feng, Chunping Hou, Sun-Yuan Kung
IEEE Trans. Geosci. Remote. Sens.1
2016 An Energy Efficient Clustering Protocol for Lifetime Maximization in Wireless Sensor Networks
abstract
Maximizing the lifetime is an important issue in the design of applications and protocols for wireless sensor networks (WSNs). Clustering sensor nodes is an effective topology control approach helping achieve this goal. In this paper, we present an energy efficiency protocol to prolong the network lifetime based on an improved particle swarm optimization (PSO) algorithm. The protocol takes both energy efficiency and transmission distance into consideration, and the relay nodes are used to balance the heavy consumption of cluster heads. In this way, the network results in better distributed sensors and a well-balanced clustering system enhancing the network's lifetime. We compare the proposed protocol with comparative protocols in different scenarios. Simulation results show that the proposed protocol performs well over other comparative protocols in various scenarios.
Ning Wang 0015, Yuan Zhou 0006, Wei Xiang 0001
GLOBECOM2
2011 Channel Distortion Modeling for Multi-View Video Transmission Over Packet-Switched Networks
abstract
Channel distortion modeling for generic multi-view video transmission remains a unfilled blank, despite that intensive research efforts have been devoted to model traditional 2-D video transmission. This paper aims to fill this blank through developing a recursive distortion model for multi-view video transmission over lossy packet-switched networks. Based on the study on the characteristics of multi-view video coding and the propagating behavior of transmission error due to random frame losses, a recursive mathematical model is derived to estimate the expected channel-induced distortion at both the frame and sequence levels. The model we develop explicitly considers both temporal and inter-view dependencies, induced by motion-compensated and disparity-compensated coding, respectively. The derived model is applicable to all multi-view video encoders using the classical block-based motion-/disparity-compensated prediction framework. Both objective and subjective experimental results are presented to demonstrate that the proposed model is capable of effectively model channel-induced distortion for multi-view video.
Yuan Zhou 0006, Chunping Hou, Wei Xiang 0001, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2010 Modeling of Transmission Distortion for Multi-View Video in Packet Lossy Networks
abstract
In this paper, a mathematical model is proposed to estimate the distortion caused by random packet losses for multi-view video transmission. Based on the study of multi-view video coding, the proposed model takes into account the disparity/motion compensation which relates the channel-induced distortion in the current frame with that in the previous frame or the adjacent view, and allows for any motion-compensated and disparity-compensated concealment method at the decoder. Comparative studies between the modeled and simulated distortion results demonstrates that the proposed model is able to estimate the transmission distortion of multi-view video with high accuracy.
Yuan Zhou 0006, Chunping Hou, Wei Xiang 0001
GLOBECOM1