VLDB 2026 Research / reviewers in the wild / expert
Xinhua Xu
dblp:98/9114
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0008-0378-5209ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Segmentation and scene understanding · 38% 3D vision · 19% Face, body and person analysis · 16% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
semantic segmentation |
1.7 | 2 | 2025 | MiLNet: Multiplex Interactive Learning Network for RGB-T Semantic Segmentation · IEEE Trans. Image Process. 2025 PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor Scenes · AAAI 2025 |
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | PRIDEV: A Plug-and-Play Refinement for Improved Depth Estimation in Videos · ICRA 2025 |
Computer vision › 3D vision › depth estimation › video depth estimation
monocular video depth estimation |
0.9 | 1 | 2025 | PRIDEV: A Plug-and-Play Refinement for Improved Depth Estimation in Videos · ICRA 2025 |
Computer vision › Vision and language
multimodal fusion |
0.9 | 1 | 2025 | MiLNet: Multiplex Interactive Learning Network for RGB-T Semantic Segmentation · IEEE Trans. Image Process. 2025 |
Computer vision › Segmentation and scene understanding › multimodal segmentation
RGB-D segmentation |
0.9 | 1 | 2025 | PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor Scenes · AAAI 2025 |
Computer vision › Segmentation and scene understanding › semantic segmentation › multimodal semantic segmentation
RGB-T semantic segmentation |
0.9 | 1 | 2025 | MiLNet: Multiplex Interactive Learning Network for RGB-T Semantic Segmentation · IEEE Trans. Image Process. 2025 |
Machine learning › Generative modeling › face synthesis
facial micro-expression generation |
0.8 | 1 | 2024 | Facial Prior Guided Micro-Expression Generation · IEEE Trans. Image Process. 2024 |
Computer vision › Video understanding and tracking › motion analysis › motion modeling
motion model |
0.5 | 1 | 2021 | Facial Prior Based First Order Motion Model for Micro-expression Generation · ACM Multimedia 2021 |
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion model features |
0.3 | 1 | 2025 | PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor Scenes · AAAI 2025 |
Computer vision › Face, body and person analysis › facial expression analysis › facial expression recognition
micro-expression recognition |
0.2 | 1 | 2024 | Facial Prior Guided Micro-Expression Generation · IEEE Trans. Image Process. 2024 |
Methods — techniques the papers use, named apart from their topics
facial prior · 1.3vision foundation model transfer · 0.9temporal depth stabilization module · 0.9pseudo depth aggregation · 0.9inverse hierarchical fusion · 0.9image depth estimation priors · 0.9diffusion model · 0.9asymmetric simulated learning · 0.9landmark estimation · 0.8image animation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor ScenesabstractThe integration of RGB and depth modalities significantly enhances the accuracy of segmenting complex indoor scenes, with depth data from RGB-D cameras playing a crucial role in this improvement. However, collecting an RGB-D dataset is more expensive than an RGB dataset due to the need for specialized depth sensors. Aligning depth and RGB images also poses challenges due to sensor positioning and issues like missing data and noise. In contrast, Pseudo Depth (PD) from high-precision depth estimation algorithms can eliminate the dependence on RGB-D sensors and alignment processes, as well as provide effective depth information and show significant potential in semantic segmentation. Therefore, to explore the practicality of utilizing pseudo depth instead of real depth for semantic segmentation, we design an RGB-PD segmentation pipeline to integrate RGB and pseudo depth and propose a Pseudo Depth Aggregation Module (PDAM) for fully exploiting the informative clues provided by the diverse pseudo depth maps. The PDAM aggregates multiple pseudo depth maps into a single modality, making it easily adaptable to other RGB-D segmentation methods. In addition, the pre-trained diffusion model serves as a strong feature extractor for RGB segmentation tasks, but multi-modal diffusion-based segmentation methods remain unexplored. Therefore, we present a Pseudo Depth Diffusion Model (PDDM) that adopts a large-scale text-image diffusion model as a feature extractor and a simple yet effective fusion strategy to integrate pseudo depth. To verify the applicability of pseudo depth and our PDDM, we perform extensive experiments on the NYUv2 and SUNRGB-D datasets. The experimental results demonstrate that pseudo depth can effectively enhance segmentation performance, and our PDDM achieves state-of-the-art performance, outperforming other methods by +6.98 mIoU on NYUv2 and +2.11 mIoU on SUNRGB-D. Xinhua Xu, Hong Liu 0008, Jianbing Wu |
AAAI | 1 |
| 2025 | PRIDEV: A Plug-and-Play Refinement for Improved Depth Estimation in VideosabstractMonocular video depth estimation is a key challenge in computer vision, highlighting its importance in visual understanding. Monocular depth estimation models trained on single images achieve impressive results on individual frames but often lack temporal consistency when applied to videos, leading to flickering and artifacts. Current video depth estimation methods often rely on additional optical flow or camera poses, which are limited by their accuracy, complex design, and lack robustness. Specially, we propose a plug-and-play method that seamlessly transfers the robustness of image depth estimation to video depth estimation. By leveraging powerful priors from image depth estimation, our method enhances the performance of video depth estimation without requiring additional conditional inputs or extensive pretraining on large and expensive video datasets. We introduce the Temporal Depth Stabilization Module (TDSM), which can seamlessly inflate an image monocular depth estimation model into a video depth estimation model, enabling unified modeling of depth across video sequences and capturing the temporal cues in video. We validate the effectiveness and efficiency of our method across various datasets (e.g., normal and challenging conditions) and different backbones. Extensive experiments demonstrate that our simple and effective method significantly improves monocular depth estimation networks, achieving new state-of-the-art accuracy in both spatial and temporal dimensions. Hong Liu 0008, Jianbing Wu, Xinhua Xu |
ICRA | 4 |
| 2025 | MiLNet: Multiplex Interactive Learning Network for RGB-T Semantic SegmentationabstractSemantic segmentation methods enhance robust and reliable understanding under adverse illumination conditions by integrating complementary information from visible and thermal infrared (RGB-T) images. Existing methods primarily focus on designing various feature fusion modules between different modalities, overlooking that feature learning is the critical aspect of scene understanding. In this paper, we propose a novel module-free Multiplex Interactive Learning Network (MiLNet) for RGB-T semantic segmentation, which adeptly integrates multi-model, multi-modal, and multi-level feature learning, fully exploiting the potential of multiplex feature interaction. Specifically, robust knowledge is transferred from the vision foundation model to our task-specific model to enhance its segmentation performance. In the task-specific model, an asymmetric simulated learning strategy is introduced to facilitate mutual learning of geometric and semantic information between high- and low-level features across modalities. Additionally, an inverse hierarchical fusion strategy based on feature learning pairs is adopted and further refined using multilabel and multiscale supervision. Experimental results on the MFNet and PST900 datasets demonstrate that MiLNet outperforms state-of-the-art methods in terms of mIoU. As a limitation, the model's performance under few-sample conditions could be improved further. The code and results of our method are available at https://github.com/Jinfu-pku/MiLNet. Hong Liu 0008, Xia Li 0005, Jiale Ren, Xinhua Xu |
IEEE Trans. Image Process. | 5 |
| 2024 | Facial Prior Guided Micro-Expression GenerationabstractThis paper focuses on the facial micro-expression (FME) generation task, which has potential application in enlarging digital FME datasets, thereby alleviating the lack of training data with labels in existing micro-expression datasets. Despite obvious progress in the image animation task, FME generation remains challenging because existing image animation methods can hardly encode subtle and short-term facial motion information. To this end, we present a facial-prior-guided FME generation framework that takes advantage of facial priors for facial motion generation. Specifically, we first estimate the geometric locations of action units (AUs) with detected facial landmarks. We further calculate an adaptive weighted prior (AWP) map, which alleviates the estimation error of AUs while efficiently capturing subtle facial motion patterns. To achieve smooth and realistic synthesis results, we use our proposed facial prior module to guide motion representation and generation modules in mainstream image animation frameworks. Extensive experiments on three benchmark datasets consistently show that our proposed facial prior module can be adopted in image animation frameworks and significantly improve their performance on micro-expression generation. Moreover, we use the generation technique to enlarge existing datasets, thereby improving the performance of general action recognition backbones on the FME recognition task. Our code is available at https://github.com/sysu19351158/FPB-FOMM. Xinhua Xu, Youjun Zhao, Yuhang Wen 0001, Zixuan Tang, Mengyuan Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Facial Prior Based First Order Motion Model for Micro-expression GenerationabstractSpotting facial micro-expression from videos finds various potential applications in fields including clinical diagnosis and interrogation, meanwhile this task is still difficult due to the limited scale of training data. To solve this problem, this paper tries to formulate a new task called micro-expression generation and then presents a strong baseline which combines the first order motion model with facial prior knowledge. Given a target face, we intend to drive the face to generate micro-expression videos according to the motion patterns of source videos. Specifically, our new model involves three modules. First, we extract facial prior features from a region focusing module. Second, we estimate facial motion using key points and local affine transformations with a motion prediction module. Third, expression generation module is used to drive the target face to generate videos. We train our model on public CASME II, SAMM and SMIC datasets and then use the model to generate new micro-expression videos for evaluation. Our model achieves the first place in the Facial Micro-Expression Challenge 2021 (MEGC2021), where our superior performance is verified by three experts with Facial Action Coding System certification. Source code is provided in https://github.com/Necolizer/Facial-Prior-Based-FOMM. Youjun Zhao, Yuhang Wen 0001, Zixuan Tang, Xinhua Xu, Mengyuan Liu 0001 |
ACM Multimedia | 5 |