EDBT 2026 Demo / reviewers in the wild / expert
Xiaohe Wu
dblp:20/3663
· DBLP profile ↗
31ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0001-6884-9121ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LK-Road3R: Road point cloud mapping via UAV-based video and deep learning
Jiangchuan Chen, Yunfei Yin, Xiaohe Wu, Abaho G. Gershome, Zejiao Dong |
Expert Syst. Appl. | 4 |
| 2026 | Anomaly detection method for satellite networks based on adaptive federated learning driven by deep reinforcement learning
Yangtao Chen, Xiaohe Wu, Chenxi Cai, Dianying Chen |
Inf. Sci. | 4 |
| 2026 | I2V-Adapter: Fast adapting image pre-trained models for video correspondence
Hannan Lu, Xinyu Zhang 0015, Zhi Tian, Xiaohe Wu, Wangmeng Zuo, Jingdong Wang 0001 |
Pattern Recognit. | 4 |
| 2026 | TCLformer: enhancing multi-scale time series forecasting with temporal decomposition and convolution-enhanced LogSparse self-attention
Xiaohe Wu, Chenxi Cai, Dianying Chen, Yaodi Liu |
J. Supercomput. | 1 |
| 2025 | MC^2: Multi-concept Guidance for Customized Multi-concept GenerationabstractCustomized text-to-image generation, which synthesizes images based on user-specified concepts, has made significant progress in handling individual concepts. However, when extended to multiple concepts, existing methods often struggle with properly integrating different models and avoiding the unintended blending of characteristics from distinct concepts. In this paper, we propose MC2, a novel approach for multi-concept customization that enhances flexibility and fidelity through inference-time optimization. MC2enables the integration of multiple single-concept models with heterogeneous architectures. By adaptively refining attention weights between visual and textual tokens, our method ensures that image regions accurately correspond to their associated concepts while minimizing interference between concepts. Extensive experiments demonstrate that MC2outperforms training-based methods in terms of prompt-reference alignment. Furthermore, MC2can be seamlessly applied to text-to-image generation, providing robust compositional capabilities. To facilitate the evaluation of multi-concept customization, we also introduce a new benchmark, MC++. The code is available at https://github.com/jiangJiaxiu/MC-2. Jiaxiu Jiang, Yabo Zhang, Kailai Feng, Xiaohe Wu, Wenbo Li 0002, Renjing Pei, Wangmeng Zuo |
CVPR | 4 |
| 2025 | Generative Inbetweening through Frame-wise Conditions-Driven Video GenerationabstractGenerative inbetweening aims to generate intermediate frame sequences by utilizing two key frames as input. Although remarkable progress has been made in video generation models, generative inbetweening still faces challenges in maintaining temporal stability due to the ambiguous interpolation path between two key frames. This issue becomes particularly severe when there is a large motion gap between input frames. In this paper, we propose a straight-forward yet highly effective Frame-wise Conditions-driven Video Generation (FCVG) method that significantly enhances the temporal stability of interpolated video frames. Specifically, our FCVG provides an explicit condition for each frame, making it much easier to identify the interpolation path between two input frames and thus ensuring temporally stable production of visually plausible video frames. To achieve this, we suggest extracting matched lines from two input frames that can then be easily interpolated frame by frame, serving as frame-wise conditions seamlessly integrated into existing video generation models. In extensive evaluations covering diverse scenarios such as natural landscapes, complex human poses, camera movements and animations, existing methods often exhibit incoherent transitions across frames. In contrast, our FCVG demonstrates the capability to generate temporally stable videos using both linear and non-linear interpolation curves. Our project page and code are available at https://fcvg-inbetween.github.io/. Dongwei Ren, Qilong Wang 0001, Xiaohe Wu, Wangmeng Zuo |
CVPR | 4 |
| 2025 | DeblurDiff: Real-Word Image Deblurring with Generative Diffusion ModelsabstractDiffusion models have achieved significant progress in image generation and the pre-trained Stable Diffusion (SD) models are helpful for image deblurring by providing clear image priors. However, directly using a blurry image or a pre-deblurred one as a conditional control for SD will either hinder accurate structure extraction or make the results overly dependent on the deblurring network. In this work, we propose a Latent Kernel Prediction Network (LKPN) to achieve robust real-world image deblurring. Specifically, we co-train the LKPN in the latent space with conditional diffusion. The LKPN learns a spatially variant kernel to guide the restoration of sharp images in the latent space. By applying element-wise adaptive convolution (EAC), the learned kernel is utilized to adaptively process the blurry feature, effectively preserving the information of the blurry input. This process thereby more effectively guides the generative process of SD, enhancing both the deblurring efficacy and the quality of detail reconstruction. Moreover, the results at each diffusion step are utilized to iteratively estimate the kernels in LKPN to better restore the sharp latent by EAC in the subsequent step. This iterative refinement enhances the accuracy and robustness of the deblurring process. Extensive experimental results demonstrate that the proposed method outperforms state-of-the-art image deblurring methods on both benchmark and real-world images. Lingshun Kong, Jiawei Zhang 0002, Dongqing Zou, Fu Lee Wang, Jimmy S. J. Ren, Xiaohe Wu, Jiangxin Dong, Jinshan Pan |
NeurIPS | 6 |
| 2025 | S2C-HAR: A Semi-Supervised Human Activity Recognition Framework Based on Contrastive LearningabstractABSTRACT Human activity recognition (HAR) has emerged as a critical element in various domains, such as smart healthcare, smart homes, and intelligent transportation, owing to the rapid advancements in wearable sensing technology and mobile computing. Nevertheless, existing HAR methods predominantly rely on deep supervised learning algorithms, necessitating a substantial supply of high‐quality labeled data, which significantly impacts their accuracy and reliability. Considering the diversity of mobile devices and usage environments, the quest for optimizing recognition performance in deep models while minimizing labeled data usage has become a prominent research area. In this paper, we propose a novel semi‐supervised HAR framework based on contrastive learning named S2C‐HAR, which is capable of generating accurate pseudo‐labels for unlabeled data, thus achieving comparable performance with supervised learning with only a few labels applied. First, a contrastive learning model for HAR (CLHAR) is designed for more general feature representations, which contains a contrastive augmentation transformer pre‐trained exclusively on unlabeled data and fine‐tuned in conjunction with a model‐agnostic classification network. Furthermore, based on the FixMatch technique, unlabeled data with two different perturbations imposed are fed into the CLHAR to produce pseudo‐labels and prediction results, which effectively provides a robust self‐training strategy and improves the quality of pseudo‐labels. To validate the efficacy of our proposed model, we conducted extensive experiments, yielding compelling results. Remarkably, even with only 1% labeled data, our model achieves satisfactory recognition performance, outperforming state‐of‐the‐art methods by approximately 5%. Xue Li 0011, Mingxing Liu, Lanshun Nie, Wenxiao Cheng, Xiaohe Wu, Dechen Zhan |
Concurr. Comput. Pract. Exp. | 5 |
| 2025 | A Two-Stage Boundary-Enhanced Contrastive Learning approach for nested named entity recognition
Yaodi Liu, Rong Tong, Chenxi Cai, Dianying Chen, Xiaohe Wu |
Expert Syst. Appl. | 6 |
| 2025 | Reblurring-Guided Single Image Defocus Deblurring: A Learning Framework with Misaligned Training Pairs
Dongwei Ren, Xinya Shu, Yu Li 0048, Xiaohe Wu, Wangmeng Zuo |
Int. J. Comput. Vis. | 4 |
| 2025 | A discontinuous NER model based on token prediction and contrastive learning to enhance span
Yaodi Liu, Dianying Chen, Chenxi Cai, Xiaohe Wu, Rong Tong |
J. Supercomput. | 5 |
| 2024 | Learning Real-World Image De-weathering with Imperfect SupervisionabstractReal-world image de-weathering aims at removing various undesirable weather-related artifacts. Owing to the impossibility of capturing image pairs concurrently, existing real-world de-weathering datasets often exhibit inconsistent illumination, position, and textures between the ground-truth images and the input degraded images, resulting in imperfect supervision. Such non-ideal supervision negatively affects the training process of learning-based de-weathering methods. In this work, we attempt to address the problem with a unified solution for various inconsistencies. Specifically, inspired by information bottleneck theory, we first develop a Consistent Label Constructor (CLC) to generate a pseudo-label as consistent as possible with the input degraded image while removing most weather-related degradation. In particular, multiple adjacent frames of the current input are also fed into CLC to enhance the pseudo-label. Then we combine the original imperfect labels and pseudo-labels to jointly supervise the de-weathering model by the proposed Information Allocation Strategy (IAS). During testing, only the de-weathering model is used for inference. Experiments on two real-world de-weathering datasets show that our method helps existing de-weathering models achieve better performance. Code is available at https://github.com/1180300419/imperfect-deweathering. Xiaohui Liu 0003, Zhilu Zhang 0001, Xiaohe Wu, Chaoyu Feng, Xiaotao Wang, Wangmeng Zuo |
AAAI | 3 |
| 2024 | Pseudo-ISP: Learning pseudo in-camera signal processing pipeline from a color image denoiser
Yue Cao 0009, Xiaohe Wu, Shuran Qi, Xiao Liu 0040, Zhongqin Wu, Wangmeng Zuo |
Neurocomputing | 2 |
| 2024 | Learning with noisy labels using collaborative sample selection and contrastive semi-supervised learning
Xiaohe Wu, Chao Xu 0003, Yanli Ji, Wangmeng Zuo, Yiwen Guo, Zhaopeng Meng |
Knowl. Based Syst. | 2 |
| 2024 | Corrigendum to "Learning with Noisy Labels Using Collaborative Sample Selection and Contrastive Semi-Supervised Learning" [Knowledge-Based Systems 296 (2024) 111860]
Xiaohe Wu, Chao Xu 0003, Yanli Ji, Wangmeng Zuo, Yiwen Guo, Zhaopeng Meng |
Knowl. Based Syst. | 2 |
| 2024 | Degraded Structure and Hue Guided Auxiliary Learning for low-light image enhancement
Heming Xu, Xiaohe Wu, Wangmeng Zuo |
Knowl. Based Syst. | 4 |
| 2024 | Learning Diverse Tone Styles for Image RetouchingabstractImage retouching, aiming to regenerate the visually pleasing renditions of given images, is a subjective task where the users are with different aesthetic sensations. Most existing methods adopt a deterministic model to learn the retouching style from a specific expert, making it less flexible to meet diverse subjective preferences. Besides, the intrinsic diversity of an expert due to the targeted processing of different images is also deficiently described. To circumvent such issues, we propose to learn diverse image retouching with normalizing flow-based architectures. Unlike current flow-based methods which directly generate the output image, we argue that learning in a one-dimensional style space could 1) disentangle the retouching styles from the image content, 2) lead to a stable style presentation form, and 3) avoid the spatial disharmony effects. For obtaining meaningful image tone style representations, a joint-training pipeline is delicately designed, which is composed of a style encoder, a conditional RetouchNet, and the image tone style normalizing flow (TSFlow) module. In particular, the style encoder predicts the target style representation of an input image, which serves as the conditional information in the RetouchNet for retouching, while the TSFlow maps the style representation vector into a Gaussian distribution in the forward pass. After training, the TSFlow can generate diverse image tone style vectors by sampling from the Gaussian distribution. Extensive experiments on MIT-Adobe FiveK and PPR10K datasets show that our proposed method performs favorably against state-of-the-art methods and is effective in generating diverse results to satisfy different human aesthetic preferences. Source codeterministic and pre-trained models are publicly available at https://github.com/SSRHeart/TSFlow. Haolin Wang 0004, Jiawei Zhang 0002, Ming Liu 0018, Xiaohe Wu, Wangmeng Zuo |
IEEE Trans. Image Process. | 4 |
| 2024 | Invertible network for unpaired low-light image enhancement
Jize Zhang, Haolin Wang 0004, Xiaohe Wu, Wangmeng Zuo |
Vis. Comput. | 3 |
| 2023 | Inferring and Leveraging Parts from Object Shape for Improving Semantic Image SynthesisabstractDespite the progress in semantic image synthesis, it remains a challenging problem to generate photo-realistic parts from input semantic map. Integrating part segmentation map can undoubtedly benefit image synthesis, but is bothersome and inconvenient to be provided by users. To improve part synthesis, this paper presents to infer Parts from Object ShapE (iPOSE) and leverage it for improving semantic image synthesis. However, albeit several part segmentation datasets are available, part annotations are still not provided for many object categories in semantic image synthesis. To circumvent it, we resort to few-shot regime to learn a PartNet for predicting the object part map with the guidance of pre-defined support part maps. PartNet can be readily generalized to handle a new object category when a small number (e.g., 3) of support part maps for this category are provided. Furthermore, part semantic modulation is presented to incorporate both inferred part map and semantic map for image synthesis. Experiments show that our iPOSE not only generates objects with rich part details, but also enables to control the image synthesis flexibly. And our iPOSE performs favorably against the state-of-the-art methods in terms of quantitative and qualitative evaluation. Our code will be publicly available at https://github.com/csyxwei/iPOSE. Yuxiang Wei 0001, Zhilong Ji, Xiaohe Wu, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo |
CVPR | 3 |
| 2023 | Survey on leveraging pre-trained generative adversarial networks for image editing and restorationabstractGenerative adversarial networks (GANs) have drawn enormous attention due to their simple yet effective training mechanism and superior image generation quality. With the ability to generate photorealistic high-resolution (e.g., 1024 × 1024) images, recent GAN models have greatly narrowed the gaps between the generated images and the real ones. Therefore, many recent studies show emerging interest to take advantage of pre-trained GAN models by exploiting the well-disentangled latent space and the learned GAN priors. In this study, we briefly review recent progress on leveraging pre-trained large-scale GAN models from three aspects, i.e., (1) the training of large-scale generative adversarial networks, (2) exploring and understanding the pre-trained GAN models, and (3) leveraging these models for subsequent tasks like image restoration and editing. Ming Liu 0018, Yuxiang Wei 0001, Xiaohe Wu, Wangmeng Zuo, Lei Zhang 0006 |
Sci. China Inf. Sci. | 3 |
| 2023 | On better detecting and leveraging noisy samples for learning with severe label noise
Xiaohe Wu, Chao Xu 0003, Wangmeng Zuo, Zhaopeng Meng |
Pattern Recognit. | 2 |
| 2022 | Unidirectional Video Denoising by Mimicking Backward Recurrent Modules with Look-Ahead Forward Ones
Junyi Li 0005, Xiaohe Wu, Zhenxing Niu, Wangmeng Zuo |
ECCV (18) | 2 |
| 2021 | Research on Authentication and Key Agreement Protocol of Smart Medical Systems Based on Blockchain Technology
Xiaohe Wu, Jianbo Xu, Wei Liang 0005, W. Jian |
ICA3PP (2) | 1 |
| 2020 | Unpaired Learning of Deep Image Denoising
Xiaohe Wu, Ming Liu 0018, Yue Cao 0009, Dongwei Ren, Wangmeng Zuo |
ECCV (4) | 1 |
| 2020 | Learning second-order statistics for place recognition based on robust covariance estimation of CNN features
Zifei Yan, Qilong Wang 0001, Xiaohe Wu, Wangmeng Zuo |
Neurocomputing | 4 |
| 2020 | Remove Cosine Window From Correlation Filter-Based Visual Trackers: When and HowabstractCorrelation filters (CFs) have been continuously advancing the state-of-the-art tracking performance and have been extensively studied in the recent few years. Nonetheless, the existing CF trackers adopt a cosine window to spatially reweight base image to alleviate boundary discontinuity. However, cosine window emphasizes more on the central regions of base image and has the risk of contaminating negative training samples during model learning. On the other hand, spatial regularization deployed in many recent CF trackers plays a similar role as cosine window by enforcing spatial penalty on CF coefficients. Therefore, we in this paper investigate the feasibility to remove cosine window from CF trackers with spatial regularization. When simply removing cosine window, CF with spatial regularization still suffers from small degree of boundary discontinuity. To tackle this issue, binary and Gaussian shaped mask functions are further introduced for eliminating boundary discontinuity while reweighting the estimation error of each training sample, and can be incorporated with multiple CF trackers with spatial regularization. In comparison to the baseline methods with cosine window, our methods are effective in handling boundary discontinuity and sample contamination, thereby benefiting tracking performance. Extensive experiments on four benchmarks show that our methods perform favorably against the state-of-the-art trackers using either handcrafted or deep CNN features. Feng Li 0031, Xiaohe Wu, Wangmeng Zuo, David Zhang 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 2 |
| 2019 | Learning Support Correlation Filters for Visual TrackingabstractFor visual tracking methods based on kernel support vector machines (SVMs), data sampling is usually adopted to reduce the computational cost in training. In addition, budgeting of support vectors is required for computational efficiency. Instead of sampling and budgeting, recently the circulant matrix formed by dense sampling of translated image patches has been utilized in kernel correlation filters for fast tracking. In this paper, we derive an equivalent formulation of a SVM model with the circulant matrix expression and present an efficient alternating optimization method for visual tracking. We incorporate the discrete Fourier transform with the proposed alternating optimization process, and pose the tracking problem as an iterative learning of support correlation filters (SCFs). In the fully-supervision setting, our SCF can find the globally optimal solution with real-time performance. For a given circulant data matrix with$n^2$samples of$n \times n$pixels, the computational complexity of the proposed algorithm is$O(n^2\; \log n)$whereas that of the standard SVM-based approaches is at least$O(n^4)$. In addition, we extend the SCF-based tracking algorithm with multi-channel features, kernel functions, and scale-adaptive approaches to further improve the tracking performance. Experimental results on a large benchmark dataset show that the proposed SCF-based algorithms perform favorably against the state-of-the-art tracking methods in terms of accuracy and speed. Wangmeng Zuo, Xiaohe Wu, Liang Lin 0004, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | VITAL: VIsual Tracking via Adversarial LearningabstractThe tracking-by-detection framework consists of two stages, i.e., drawing samples around the target object in the first stage and classifying each sample as the target object or as background in the second stage. The performance of existing trackers using deep classification networks is limited by two aspects. First, the positive samples in each frame are highly spatially overlapped, and they fail to capture rich appearance variations. Second, there exists extreme class imbalance between positive and negative samples. This paper presents the VITAL algorithm to address these two problems via adversarial learning. To augment positive samples, we use a generative network to randomly generate masks, which are applied to adaptively dropout input features to capture a variety of appearance changes. With the use of adversarial learning, our network identifies the mask that maintains the most robust features of the target objects over a long temporal span. In addition, to handle the issue of class imbalance, we propose a high-order cost sensitive loss to decrease the effect of easy negative samples to facilitate training the classification network. Extensive experiments on benchmark datasets demonstrate that the proposed tracker performs favorably against state-of-the-art approaches. Yibing Song, Chao Ma 0004, Xiaohe Wu, Lijun Gong, Linchao Bao, Wangmeng Zuo, Chunhua Shen, Rynson W. H. Lau, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2018 | Joint Representation and Truncated Inference Learning for Correlation Filter Based Tracking
Yingjie Yao, Xiaohe Wu, Lei Zhang 0036, Shiguang Shan, Wangmeng Zuo |
ECCV (9) | 2 |
| 2018 | F-SVM: Combination of Feature Transformation and SVM Learning via Convex RelaxationabstractThe generalization error bound of the support vector machine (SVM) depends on the ratio of the radius and margin. However, conventional SVM only considers the maximization of the margin but ignores the minimization of the radius, which restricts its performance when applied to joint learning of feature transformation and the SVM classifier. Although several approaches have been proposed to integrate the radius and margin information, most of them either require the form of the transformation matrix to be diagonal, or are nonconvex and computationally expensive. In this paper, we suggest a novel approximation for the radius of the minimum enclosing ball in feature space, and then propose a convex radius-margin-based SVM model for joint learning of feature transformation and the SVM classifier, i.e., F-SVM. A generalized block coordinate descent method is adopted to solve the F-SVM model, where the feature transformation is updated via the gradient descent and the classifier is updated by employing the existing SVM solver. By incorporating with kernel principal component analysis, F-SVM is further extended for joint learning of nonlinear transformation and the classifier. F-SVM can also be incorporated with deep convolutional networks to improve image classification performance. Experiments on the UCI, LFW, MNIST, CIFAR-10, CIFAR-100, and Caltech101 data sets demonstrate the effectiveness of F-SVM. Xiaohe Wu, Wangmeng Zuo, Liang Lin 0004, Wei Jia 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Learning pairwise SVM on deep features for ear recognitionabstractRecently, deep features extracted from Convolutional Neural Networks (CNNs) have been widely adopted in various applications, such as face recognition. Compared with the handcrafted descriptors, deep features have more powerful representation ability which can lead to better performance. Effective feature representations play an important role in ear recognition. While deep features have not been applied to represent the ear images. In this paper, we propose to extract deep features of ear images based on VGG-M Net for solving the ear recognition problem. And due to the lack of training images per person, we propose to use the pairwise SVM for classification firstly. For computational efficiency, Principal Component Analysis (PCA) is exploited to reduce the dimension before classification. Finally, we evaluate our approach on two public ear databases: USTB I and USTB II. The experimental results achieve a promising recognition rate and show superior performance compared with the state-of-the-art methods. Ibrahim Omara, Xiaohe Wu, Wangmeng Zuo |
ICIS | 2 |