Xiaotao Wang

dblp:42/10799 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 3 since 2021Computer networks · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Adaptive Optimal Surrounding Control of Multiple Unmanned Surface Vessels via Actor-Critic Reinforcement Learning
abstract
In this article, an optimal surrounding control algorithm is proposed for multiple unmanned surface vessels (USVs), in which actor-critic reinforcement learning (RL) is utilized to optimize the merging process. Specifically, the multiple-USV optimal surrounding control problem is first transformed into the Hamilton-Jacobi-Bellman (HJB) equation, which is difficult to solve due to its nonlinearity. An adaptive actor-critic RL control paradigm is then proposed to obtain the optimal surround strategy, wherein the Bellman residual error is utilized to construct the network update laws. Particularly, a virtual controller representing intermediate transitions and an actual controller operating on a dynamics model are employed as surrounding control solutions for second-order USVs; thus, optimal surrounding control of the USVs is guaranteed. In addition, the stability of the proposed controller is analyzed by means of Lyapunov theory functions. Finally, numerical simulation results demonstrate that the proposed actor-critic RL-based surrounding controller can achieve the surrounding objective while optimizing the evolution process and obtains 9.76% and 20.85% reduction in trajectory length and energy consumption compared with the existing controller.
Renzhi Lu, Xiaotao Wang, Yiyu Ding, Hai-Tao Zhang, Lijun Zhu 0001, Yong He 0003
IEEE Trans. Neural Networks Learn. Syst.2
2024 Learning Real-World Image De-weathering with Imperfect Supervision
abstract
Real-world image de-weathering aims at removing various undesirable weather-related artifacts. Owing to the impossibility of capturing image pairs concurrently, existing real-world de-weathering datasets often exhibit inconsistent illumination, position, and textures between the ground-truth images and the input degraded images, resulting in imperfect supervision. Such non-ideal supervision negatively affects the training process of learning-based de-weathering methods. In this work, we attempt to address the problem with a unified solution for various inconsistencies. Specifically, inspired by information bottleneck theory, we first develop a Consistent Label Constructor (CLC) to generate a pseudo-label as consistent as possible with the input degraded image while removing most weather-related degradation. In particular, multiple adjacent frames of the current input are also fed into CLC to enhance the pseudo-label. Then we combine the original imperfect labels and pseudo-labels to jointly supervise the de-weathering model by the proposed Information Allocation Strategy (IAS). During testing, only the de-weathering model is used for inference. Experiments on two real-world de-weathering datasets show that our method helps existing de-weathering models achieve better performance. Code is available at https://github.com/1180300419/imperfect-deweathering.
Xiaohui Liu 0003, Zhilu Zhang 0001, Xiaohe Wu, Chaoyu Feng, Xiaotao Wang, Wangmeng Zuo
AAAI5
2024 Self-Supervised High Dynamic Range Imaging with Multi-Exposure Images in Dynamic Scenes
abstract
Merging multi-exposure images is a common approach for obtaining high dynamic range (HDR) images, with the primary challenge being the avoidance of ghosting artifacts in dynamic scenes. Recent methods have proposed using deep neural networks for deghosting. However, the methods typically rely on sufficient data with HDR ground-truths, which are difficult and costly to collect. In this work, to eliminate the need for labeled data, we propose SelfHDR, a self-supervised HDR reconstruction method that only requires dynamic multi-exposure images during training. Specifically, SelfHDR learns a reconstruction network under the supervision of two complementary components, which can be constructed from multi-exposure images and focus on HDR color as well as structure, respectively. The color component is estimated from aligned multi-exposure images, while the structure one is generated through a structure-focused network that is supervised by the color component and an input reference (\eg, medium-exposure) image. During testing, the learned reconstruction network is directly deployed to predict an HDR image. Experiments on real-world images demonstrate our SelfHDR achieves superior results against the state-of-the-art self-supervised methods, and comparable performance to supervised ones. Codes are available at https://github.com/cszhilu1998/SelfHDR
Zhilu Zhang 0001, Shuai Liu 0009, Xiaotao Wang, Wangmeng Zuo
ICLR4
2023 Physics-Guided ISO-Dependent Sensor Noise Modeling for Extreme Low-Light Photography
abstract
Although deep neural networks have achieved astonishing performance in many vision tasks, existing learningbased methods are far inferior to the physical model-based solutions in extreme low-light sensor noise modeling. To tap the potential of learning-based sensor noise modeling, we investigate the noise formation in a typical imaging process and propose a novel physics-guided ISO-dependent sensor noise modeling approach. Specifically, we build a normalizing flow-based framework to represent the complex noise characteristics of CMOS camera sensors. Each component of the noise model is dedicated to a particular kind of noise under the guidance of physical models. Moreover, we take into consideration of the ISO dependence in the noise model, which is not completely considered by the existing learning-based methods. For training the proposed noise model, a new dataset is further collected with paired noisy-clean images, as well as flat-field and bias frames covering a wide range of ISO settings. Compared to existing methods, the proposed noise model is equipped with a flexible structure and accurate modeling capabilities, which is beneficial for better denoising performance in extreme low-light scenes. The dataset and code are available at https://github.com/happycaoyue/LLD.
Yue Cao 0009, Ming Liu 0018, Shuai Liu 0009, Xiaotao Wang, Wangmeng Zuo
CVPR4
2023 Spatially Adaptive Self-Supervised Learning for Real-World Image Denoising
abstract
Significant progress has been made in self-supervised image denoising (SSID) in the recent few years. However, most methods focus on dealing with spatially independent noise, and they have little practicality on real-world sRGB images with spatially correlated noise. Although pixel-shuffle downsampling has been suggested for breaking the noise correlation, it breaks the original information of images, which limits the denoising performance. In this paper, we propose a novel perspective to solve this problem, i.e., seeking for spatially adaptive supervision for real-world sRGB image denoising. Specifically, we take into account the respective characteristics of flat and textured regions in noisy images, and construct supervisions for them separately. For flat areas, the supervision can be safely derived from non-adjacent pixels, which are much far from the current pixel for excluding the influence of the noise-correlated ones. And we extend the blind-spot network to a blind-neighborhood network (BNN) for providing supervision on flat areas. For textured regions, the supervision has to be closely related to the content of adjacent pixels. And we present a locally aware network (LAN) to meet the requirement, while LAN itself is selectively supervised with the output of BNN. Combining these two supervisions, a denoising network (e.g., U-Net) can be well-trained. Extensive experiments show that our method performs favorably against state-of-the-art SSID methods on real-world sRGB photographs. The code is available at https://github.com/nagejacob/SpatiallyAdaptiveSSID.
Junyi Li 0005, Zhilu Zhang 0001, Xiaoyu Liu 0006, Chaoyu Feng, Xiaotao Wang, Wangmeng Zuo
CVPR5
2023 CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language Recognition
abstract
The co-occurrence signals (e.g., hand shape, facial expression, and lip pattern) play a critical role in Continuous Sign Language Recognition (CSLR). Compared to RGB data, skeleton data provide a more efficient and concise option, and lay a good foundation for the co-occurrence exploration in CSLR. However, skeleton data are often used as a tool to assist visual grounding and have not attracted sufficient attention. In this paper, we propose a simple yet effective GCN-based approach, named CoSign, to incorporate Co-occurrence Signals and explore the potential of skeleton data in CSLR. Specifically, we propose a group-specific GCN to better exploit the knowledge of each signal and a complementary regularization to prevent complex co-adaptation across signals. Furthermore, we propose a two-stream framework that gradually fuses both static and dynamic information in skeleton data. Experimental results on three public CSLR datasets (PHOENIX14, PHOENIX14-T and CSL-Daily) show that the proposed CoSign achieves competitive performance with recent video-based approaches while reducing the computation cost during training.
Peiqi Jiao, Yuecong Min, Xiaotao Wang, Xilin Chen 0001
ICCV4
2023 Beyond Image Borders: Learning Feature Extrapolation for Unbounded Image Composition
abstract
For improving image composition and aesthetic quality, most existing methods modulate the captured images by striking out redundant content near the image borders. However, such image cropping methods are limited in the range of image views. Some methods have been suggested to extrapolate the images and predict cropping boxes from the extrapolated image. Nonetheless, the synthesized extrapolated regions may be included in the cropped image, making the image composition result not real and potentially with degraded image quality. In this paper, we circumvent this issue by presenting a joint framework for both unbounded recommendation of camera view and image composition (i.e., UNIC). In this way, the cropped image is a sub-image of the image acquired by the predicted camera view, and thus can be guaranteed to be real and consistent in image quality. Specifically, our framework takes the current camera preview frame as input and provides a recommendation for view adjustment, which contains operations unlimited by the image borders, such as zooming in or out and camera movement. To improve the prediction accuracy of view adjustment prediction, we further extend the field of view by feature extrapolation. After one or several times of view adjustments, our method converges and results in both a camera view and a bounding box showing the image composition recommendation. Extensive experiments are conducted on the datasets constructed upon existing image cropping datasets, showing the effectiveness of our UNIC in unbounded recommendation of camera view and image composition. The source code, dataset, and pre-trained models is available at https://github.com/liuxiaoyu1104/UNIC.
Xiaoyu Liu 0006, Ming Liu 0018, Junyi Li 0005, Shuai Liu 0009, Xiaotao Wang, Wangmeng Zuo
ICCV5
2023 Self-supervised Learning to Bring Dual Reversed Rolling Shutter Images Alive
abstract
Modern consumer cameras usually employ the rolling shutter (RS) mechanism, where images are captured by scanning scenes row-by-row, yielding RS distortions for dynamic scenes. To correct RS distortions, existing methods adopt a fully supervised learning manner, where high framerate global shutter (GS) images should be collected as ground-truth supervision. In this paper, we propose a Self-supervised learning framework for Dual reversed RS distortions Correction (SelfDRSC), where a DRSC network can be learned to generate a high framerate GS video only based on dual RS images with reversed distortions. In particular, a bidirectional distortion warping module is proposed for reconstructing dual reversed RS images, and then a self-supervised loss can be deployed to train DRSC network by enhancing the cycle consistency between input and reconstructed dual reversed RS images. Besides start and end RS scanning time, GS images at arbitrary intermediate scanning time can also be supervised in SelfDRSC, thus enabling the learned DRSC network to generate a high framerate GS video. Moreover, a simple yet effective self-distillation strategy is introduced in self-supervised loss for mitigating boundary artifacts in generated GS images. On synthetic dataset, SelfDRSC achieves better or comparable quantitative metrics in comparison to state-of-the-art methods trained in the full supervision manner. On real-world RS cases, our SelfDRSC can produce high framerate GS videos with finer correction textures and better temporary consistency. The source code and trained models are made publicly available at https://github.com/shangwei5/SelfDRSC.
Wei Shang 0001, Dongwei Ren, Chaoyu Feng, Xiaotao Wang, Wangmeng Zuo
ICCV4
2023 Cooperative Target-Surrounding Control of Unmanned Surface Vessels Based on MADDPG
abstract
This article proposes a multi-agent deep reinforce-ment learning algorithm to control a fleet of unmanned surface vessels (USVs) that encircle and capture sea targets. First, a simulation environment for USVs is established based on a dynamic model; two-dimensional control variables are used to control movements in three directions. Second, the multi-agent deep deterministic policy gradient (MADDPG) algorithm is employed to achieve intelligent control of the USVs, using a reward function based on certain prior knowledge. Finally, centralized training and decentralized execution are used to complete the offline learning of multiple agents, and continuous action decisions are made based on the observations of the USV sensors. Simulations demonstrate that the method can capture stationary moving sea targets with any number of multi-agents under disturbances, and exhibits strong robustness and practicality.
Taoman Li, Zihan Gan, Zexing Zhou, Xiaotao Wang, Renzhi Lu
IECON5
2023 HiCLift: a fast and efficient tool for converting chromatin interaction data between genome assemblies
abstract
MOTIVATION: With the continuous effort to improve the quality of human reference genome and the generation of more and more personal genomes, the conversion of genomic coordinates between genome assemblies is critical in many integrative and comparative studies. While tools have been developed for such task for linear genome signals such as ChIP-Seq, no tool exists to convert genome assemblies for chromatin interaction data, despite the importance of three-dimensional genome organization in gene regulation and disease. RESULTS: Here, we present HiCLift, a fast and efficient tool that can convert the genomic coordinates of chromatin contacts such as Hi-C and Micro-C from one assembly to another, including the latest T2T-CHM13 genome. Comparing with the strategy of directly remapping raw reads to a different genome, HiCLift runs on average 42 times faster (hours vs. days), while outputs nearly identical contact matrices. More importantly, as HiCLift does not need to remap the raw reads, it can directly convert human patient sample data, where the raw sequencing reads are sometimes hard to acquire or not available. AVAILABILITY AND IMPLEMENTATION: HiCLift is publicly available at https://github.com/XiaoTaoWang/HiCLift.
Xiaotao Wang
Bioinform.1
2022 Deep Radial Embedding for Visual Sequence Learning
Yuecong Min, Peiqi Jiao, Xiaotao Wang, Xiujuan Chai, Xilin Chen 0001
ECCV (6)4
2022 Kronecker Factorization-Based Multinomial Logistic Regression for Hyperspectral Image Classification
abstract
Multinomial logistic regression (MLR) has become prevailing for supervised learning within hyperspectral images (HSIs) community. It seeks the optimal regressors with the given training data. To better understand HSI data and learn more representative spatial feature, in this letter, a unified framework which combines MLR classifier training and Kronecker factorization (KF)-based feature learning for joint optimization is proposed. It is called Kronecker factorization-based multinomial logistic regression algorithm (KF-MLR). Besides, with Gabor wavelet transform to feed input, two data-oriented strategies are tailored. One is to reshape the feature learning part as a bilinear form instead such that the structure information among different wavelet filters can be exploited as much as possible. The other is to add local regularization term to preserve discriminant information as well. The regressors and feature factor matrices are optimized in iterative fashion. KF-MLR is investigated on several popular HSI datasets. The random experimental results prove it a competitive and promising classifier when compared with other state-of-the-art techniques.
Xiaotao Wang
IEEE Geosci. Remote. Sens. Lett.1
2022 Hyperspectral Image Classification Powered by Khatri-Rao Decomposition-Based Multinomial Logistic Regression
abstract
Multinomial logistic regression (MLR) is of great significance in hyperspectral image (HSI) classification within remote sensing community. It seeks the optimal regressors with logistic loss function. In the past, spatial information has been widely used to improve its classification performance by means of some prepared feature or postprocessing techniques. To better understand HSI data and learn more effective spatial feature, in this paper, a joint optimization framework which combines MLR classifier training with Khatri-Rao decomposition-based feature learning is proposed for HSI classification. It is called Khatri-Rao decomposition-based multinomial logistic regression algorithm (KR-MLR). With Gabor feature as input, KR-MLR customizes two data-oriented strategies. One is to insert a feature learning layer after the initial input and optimize it with classifier concurrently. Moreover, Khatri-Rao decomposition is utilized to convert the problem into tensor space and make it feasible in computation. Another is to add local regularization term to preserve discriminant information as well. The regressors and feature factor matrices are optimized in iterative fashion. The proposed KR-MLR is investigated on four popular HSI data sets. The experimental results show that KR-MLR outperforms other prior arts, proving it a competitive and promising classifier.
Xiaotao Wang
IEEE Trans. Geosci. Remote. Sens.1
2018 Maximum Correntropy Criterion-Based Low-Rank Preserving Projection for Hyperspectral Image Classification
abstract
In this letter, we propose a maximum correntropy criterion-based low-rank preserving projection (MCC-LRPP) for hyperspectral image (HSI) classification, seeking a lowdimensional subspace via low-rank correntropy graph where spectral band structure can be preserved as much as possible. Unlike the sparse and low-rank-based techniques available, MCC-LRPP introduces maximum correntropy criteria (MCC) to model individual band reconstruction error and noise discriminately instead of l2and Frobenius related norms. It is equivalent to a row-weighting regularization problem. It puts more emphasis on bands with less noise and indirectly increase their importance and vice versa. MCC-LRPP enhances band difference and thus preserves their local structure as well as global structure. Indeed, more local structure means more discriminant ability. The experimental results on several popular HSI data sets prove its effectiveness and superiority when compared to other existing dimension reduction means.
Xiaotao Wang
IEEE Geosci. Remote. Sens. Lett.1
2017 Weighted Low-Rank Representation-Based Dimension Reduction for Hyperspectral Image Classification
abstract
A predimension-reduction algorithm that couples weighted low-rank representation (WLRR) with a skinny intrinsic mode functions (IMFs) dictionary is proposed for hyperspectral image (HSI) classification. It seeks a low-rank subspace to solve the performance degradation issue encountered by linear discriminant analysis in a small-sample-size situation. It can also improve the scatter matrix estimation when using a large training set. Unlike those commonly used methods, e.g., the principal component analysis-based ones, WLRR focuses on preserving more structure information. Based on the traditional LRR model, WLRR introduces a local weighted regularization to characterize the correlation between samples such that HSI-specific local structure can be better preserved as well as its global structure. Indeed, more structure information gives more additional discriminant ability. Furthermore, a new discriminant IMFs dictionary is designed to enhance interclass difference via empirical mode decomposition. The proposed method is investigated on several HSI data sets. All experimental results prove it a competitive and promising predimension-reduction means when compared to other traditional techniques.
Xiaotao Wang
IEEE Geosci. Remote. Sens. Lett.1
2016 Context-aware event-driven stereo matching
abstract
Similarity measuring plays as an import role in stereo matching, whether for visual data from standard cameras or for those from novel sensors such as Dynamic Vision Sensors (DVS). Generally speaking, robust feature descriptors contribute to designing a powerful similarity measurement, as demonstrated by classic stereo matching methods. However, the kind and representative ability of feature descriptors for DVS data are so limited that achieving accurate stereo matching on DVS data becomes very challenging. In this paper, a novel feature descriptor is proposed to improve the accuracy for DVS stereo matching. Our feature descriptor can describe the local context or distribution of the DVS data, contributing to constructing an effective similarity measurement for DVS data matching, yielding an accurate stereo matching result. Our method is evaluated by testing our method on groundtruth data and comparing with various standard stereo methods. Experiments demonstrate the efficiency and effectiveness of our method.
Dongqing Zou, Qiang Wang 0023, Xiaotao Wang, Guangqi Shao, Paul K. J. Park
ICIP4
2015 Real-time human body parts localization from dynamic vision sensor
abstract
Dynamic vision sensor (DVS) as a novel type of visual sensors can detect a moving object in a fast and cost effective way by outputting events on edges of the object. This paper proposes a body part localization method using structured output Deep Belief Network (s-DBN) to label the body parts in block of pixels in very fast fashion. Experiments show that our proposed algorithm achieves pixel accuracy 90.13% on body parts localization compared to Deep Belief Network (87.01%) and Random Forests (84.15%) under the same computational cost. For head/hand detection s-DBN has significant better accuracy of 99.3%/87.8% compared to DBN 98.7%/81.7% and RF 97.1%/47.1% under recall rate 99%/90%. Specifically, the process time on a 240×180 sized image is less than 1ms on Intel Core2 2.83GHZ CPU.
Wentao Mao, Qiang Wang 0023, Xiaotao Wang, Shandong Wang, Guangqi Shao, Kyoobin Lee, Paul K. J. Park
ICIP3
2013 Learning a Structured Graphical Model with Boosted Top-Down Features for Ultrasound Image Segmentation
Zhihui Hao, Qiang Wang 0023, Xiaotao Wang, Jung-Bae Kim, Youngkyoo Hwang, Baek Hwan Cho, Won Ki Lee
MICCAI (1)3
2013 Fast ASA Modeling and Texturing Using Subgraph Isomorphism Detection Algorithm of Relational Model
Xiaotao Wang, Pingping Yang
MMM (2)2
2013 An Efficient ML Decoder for Tail-Biting Codes Based on Circular Trap Detection
abstract
Tail-biting codes are efficient coding techniques to eliminate the rate loss in conventional known-tail convolutional codes at a cost of increased complexity in decoders. In addition, tail-biting trellis representation of block codes makes the trellis-based maximum likelihood (ML) decoder desirable for implementation. Circular Viterbi algorithm (CVA) is introduced to decode the tail-biting codes for its decoding efficiency. However, its decoding process suffers from circular traps, which degrade the decoding efficiency. In this paper, we propose an efficient checking rule for the detection of circular traps. Based on this rule, a novel maximum likelihood (ML) decoding algorithm for tail-biting codes is presented. On tail-biting trellis, computational complexity and memory consumption of this decoder are significantly reduced comparing to other available ML decoders, such as the two-phase ML decoder. To further reduce the decoding complexity, we propose a new near-optimal decoding algorithm based on a simplified trap detection strategy. The performance of the above algorithms is validated with simulation.
Xiaotao Wang, Hua Qian, Weidong Xiang, Jing Xu 0001
IEEE Trans. Commun.1
2012 Fractional delay compensation in digital predistortion system
abstract
Delay mismatch between the input and output signals of power amplifiers (PAs) may lead to an erroneous assumption of memory effects. In adaptive digital predistortion (DPD) system, the delay mismatch affects the accuracy of coefficients estimation and degrades performance of the DPD system. In this paper, we reveal the impact of fractional delay mismatch, and analyze the relationship between delay mismatch and memory effects. The fractional delay compensation helps to reduce or eliminate the delay mismatch. Benefits of fractional delay compensation are provided through numerical analysis and experimental results.
Hua Qian, Xiaotao Wang
ICASSP4
2011 An Efficient CVA-Based Decoding Algorithm for Tail-Biting Codes
abstract
Tail-biting convolutional codes (TBCC) provide an efficient method to eliminate the rate loss caused by the known-tail encoding. To simplify the decoder design, circular Viterbi algorithm (CVA) has been proposed by recording and repeating the received block of (soft) symbols beyond the block boundary and continuing Viterbi decoding. However, CVA does not converge in the presence of circular trap. A checking rule is proposed for detecting the circular trap in existing CVA. Based on this rule, an efficient CVA-based decoding algorithm is obtained for tail-biting codes, which exhibits near-optimal performance for both short and long tail-biting codes. This new scheme provides faster convergence speed than the conventional CVA without increasing in complexity and storage space.
Xiaotao Wang, Hua Qian, Jing Xu 0001, Yang Yang 0001
GLOBECOM1