VLDB 2026 Research / reviewers in the wild / expert
Xingxin Xu
dblp:243/3806
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-8968-9331ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Image and video processing · 91% Visual content generation and editing · 9% | |
| Artificial intelligence
2 papers |
Generative modeling · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing
image fusion |
1.8 | 2 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 Conditional Controllable Image Fusion · NeurIPS 2024 |
Image and video processing › image restoration
degradation-aware enhancement |
1.0 | 1 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 |
Image and video processing
image restoration |
1.0 | 1 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 |
Image and video processing › image fusion
multi-modal image fusion |
1.0 | 1 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 |
Machine learning › Generative modeling › generative adversarial network
image-to-image translation |
0.5 | 1 | 2021 | Complementary, Heterogeneous and Adversarial Networks for Image-to-Image Translation · IEEE Trans. Image Process. 2021 |
Visual content generation and editing
image generation |
0.5 | 1 | 2021 | Complementary, Heterogeneous and Adversarial Networks for Image-to-Image Translation · IEEE Trans. Image Process. 2021 |
Image and video processing
image enhancement |
0.3 | 1 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 |
Machine learning › Generative modeling › diffusion model › score-based generative model
denoising diffusion probabilistic model |
0.2 | 1 | 2024 | Conditional Controllable Image Fusion · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | Conditional Controllable Image Fusion · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 1.5conditional generation · 1.5prompt-based encoding · 1.0multi-scale discriminator · 1.0heterogeneous generators · 1.0generative adversarial network · 1.0gated fusion · 1.0cross-modal enhancement · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image FusionabstractImage fusion aims to integrate comprehensive information from images acquired through multiple sources. However, images captured by diverse sensors often encounter various degradations that can negatively affect fusion quality. Traditional fusion methods generally treat image enhancement and fusion as separate processes, overlooking the inherent correlation between them; notably, the dominant regions in one modality of a fused image often indicate areas where the other modality might benefit from enhancement. Inspired by this observation, we introduce the concept of dominant regions for image enhancement and present a Dynamic Relative EnhAnceMent framework for Image Fusion (Dream-IF). This framework quantifies the relative dominance of each modality across different layers and leverages this information to facilitate reciprocal cross-modal enhancement. By integrating the relative dominance derived from image fusion, our approach supports not only image restoration but also a broader range of image enhancement applications. Furthermore, we employ prompt-based encoding to capture degradation-specific details, which dynamically steer the restoration process and promote coordinated enhancement in both multi-modal image fusion and image enhancement scenarios. Extensive experimental results demonstrate that Dream-IF consistently outperforms its counterparts. Xingxin Xu, Bing Cao 0002, Dongdong Li 0004, Qinghua Hu, Pengfei Zhu 0001 |
AAAI | 1 |
| 2024 | Conditional Controllable Image FusionabstractImage fusion aims to integrate complementary information from multiple input images acquired through various sources to synthesize a new fused image. Existing methods usually employ distinct constraint designs tailored to specific scenes, forming fixed fusion paradigms. However, this data-driven fusion approach is challenging to deploy in varying scenarios, especially in rapidly changing environments. To address this issue, we propose a conditional controllable fusion (CCF) framework for general image fusion tasks without specific training. Due to the dynamic differences of different samples, our CCF employs specific fusion constraints for each individual in practice. Given the powerful generative capabilities of the denoising diffusion model, we first inject the specific constraints into the pre-trained DDPM as adaptive fusion conditions. The appropriate conditions are dynamically selected to ensure the fusion process remains responsive to the specific requirements in each reverse diffusion stage. Thus, CCF enables conditionally calibrating the fused images step by step. Extensive experiments validate our effectiveness in general fusion tasks across diverse scenarios against the competing methods without additional training. The code is publicly available. Bing Cao 0002, Xingxin Xu, Pengfei Zhu 0001, Qilong Wang 0001, Qinghua Hu |
NeurIPS | 2 |
| 2023 | A Knowledge-Guided Framework for Fine-Grained Classification of Liver Lesions Based on Multi-Phase CT ImagesabstractAutomatic and accurate differentiation of liver lesions from multi-phase computed tomography imaging is critical for the early detection of liver cancer. Multi-phase data can provide more diagnostic information than single-phase data, and the effective use of multi-phase data can significantly improve diagnostic accuracy. Current fusion methods usually fuse multi-phase information at the image level or feature level, ignoring the specificity of each modality, therefore, the information integration capacity is always limited. In this paper, we propose a Knowledge-guided framework, named MCCNet, which adaptively integrates multi-phase liver lesion information from three different stages to fully utilize and fuse multi-phase liver information. Specifically, 1) a multi-phase self-attention module was designed to adaptively combine and integrate complementary information from three phases using multi-level phase features; 2) a cross-feature interaction module was proposed to further integrate multi-phase fine-grained features from a global perspective; 3) a cross-lesion correlation module was proposed for the first time to imitate the clinical diagnosis process by exploiting inter-lesion correlation in the same patient. By integrating the above three modules into a 3D backbone, we constructed a lesion classification network. The proposed lesion classification network was validated on an in-house dataset containing 3,683 lesions from 2,333 patients in 9 hospitals. Extensive experimental results and evaluations on real-world clinical applications demonstrate the effectiveness of the proposed modules in exploiting and fusing multi-phase information. Xingxin Xu, Qikui Zhu, Hanning Ying, Jiongcheng Li, Xiujun Cai, Shuo Li 0001, Yizhou Yu |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Toward Realistic Face Photo-Sketch Synthesis via Composition-Aided GANsabstractFace photo-sketch synthesis aims at generating a facial sketch/photo conditioned on a given photo/sketch. It covers wide applications including digital entertainment and law enforcement. Precisely depicting face photos/sketches remains challenging due to the restrictions on structural realism and textural consistency. While existing methods achieve compelling results, they mostly yield blurred effects and great deformation over various facial components, leading to the unrealistic feeling of synthesized images. To tackle this challenge, in this article, we propose using facial composition information to help the synthesis of face sketch/photo. Especially, we propose a novel composition-aided generative adversarial network (CA-GAN) for face photo-sketch synthesis. In CA-GAN, we utilize paired inputs, including a face photo/sketch and the corresponding pixelwise face labels for generating a sketch/photo. Next, to focus training on hard-generated components and delicate facial structures, we propose a compositional reconstruction loss. In addition, we employ a perceptual loss function to encourage the synthesized image and real image to be perceptually similar. Finally, we use stacked CA-GANs (SCA-GANs) to further rectify defects and add compelling details. The experimental results show that our method is capable of generating both visually comfortable and identity-preserving face sketches/photos over a wide range of challenging data. In addition, our method significantly decreases the best previous Fréchet inception distance (FID) from 36.2 to 26.2 for sketch synthesis, and from 60.9 to 30.5 for photo synthesis. Besides, we demonstrate that the proposed method is of considerable generalization ability. Jun Yu 0002, Xingxin Xu, Fei Gao 0006, Shengjie Shi, Meng Wang 0001, Dacheng Tao, Qingming Huang |
IEEE Trans. Cybern. | 2 |
| 2021 | Complementary, Heterogeneous and Adversarial Networks for Image-to-Image TranslationabstractImage-to-image translation is to transfer images from a source domain to a target domain. Conditional Generative Adversarial Networks (GANs) have enabled a variety of applications. Initial GANs typically conclude one single generator for generating a target image. Recently, using multiple generators has shown promising results in various tasks. However, generators in these works are typically of homogeneous architectures. In this paper, we argue that heterogeneous generators are complementary to each other and will benefit the generation of images. By heterogeneous, we mean that generators are of different architectures, focus on diverse positions, and perform over multiple scales. To this end, we build two generators by using a deep U-Net and a shallow residual network, respectively. The former concludes a series of down-sampling and up-sampling layers, which typically have large perception field and great spatial locality. In contrast, the residual network has small perceptual fields and works well in characterizing details, especially textures and local patterns. Afterwards, we use a gated fusion network to combine these two generators for producing a final output. The gated fusion unit automatically induces heterogeneous generators to focus on different positions and complement each other. Finally, we propose a novel approach to integrate multi-level and multi-scale features in the discriminator. This multi-layer integration discriminator encourages generators to produce realistic details from coarse to fine scales. We quantitatively and qualitatively evaluate our model on various benchmark datasets. Experimental results demonstrate that our method significantly improves the quality of transferred images, across a variety of image-to-image translation tasks. We have made our code and results publicly available: http://aiart.live/chan/. Fei Gao 0006, Xingxin Xu, Jun Yu 0002, Meimei Shang, Xiang Li 0205, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2020 | Attentive and ensemble 3D dual path networks for pulmonary nodules classification
Hanliang Jiang, Fei Gao 0006, Xingxin Xu, Suguo Zhu |
Neurocomputing | 3 |
| 2019 | Improving Facial Attractiveness Prediction via Co-attention LearningabstractFacial attractiveness prediction has drawn considerable attention from image processing community. Despite the substantial progress achieved by existing works, various challenges remain. One is the lack of accurate representation for facial composition, which is essential for attractiveness evaluation. In this paper, we propose to use pixel-wise labelling masks as the meta information of facial composition, and input them into a network for learning high-level semantic representations. The other challenge is to define to what degree different local properties contribute to facial attractiveness. To tackle this challenge, we employ a co-attention learning mechanism to concurrently characterize the significance of different regions and that of distinct facial components. We conduct experiments on the SCUT-FBP5500 and CelebA datasets. Results show that our co-attention learning mechanism significantly improves the facial attractiveness prediction accuracy. Besides, our method consistently produces appealing results and outperforms previous advanced approaches. Shengjie Shi, Fei Gao 0006, Xuantong Meng, Xingxin Xu, Jingjie Zhu |
ICASSP | 4 |
| 2019 | Cross-Layer Optimization on Charging Strategy for Wireless Sensor Networks Based on Successive Interference Cancellation
Juan Xu 0002, Xingxin Xu, Xu Ding 0001, Lei Shi 0011, Yang Lu 0015 |
WASA | 2 |