EDBT 2026 Demo / reviewers in the wild / expert
Yazhou Xing
dblp:63/7625
· DBLP profile ↗
11ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-2872-0192ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
8 papers |
Image and video processing · 42% Visual content generation and editing · 30% Computational photography and imaging · 19% | |
| Artificial intelligence
8 papers |
Generative modeling · 52% Robot navigation and mapping · 17% Vision and language · 13% |
Topics — the 26 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.9 | 3 | 2025 | SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation · ACM Multimedia 2025 Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners · CVPR 2024 LDM-ISP: Enhancing Neural ISP for Low Light with Latent Diffusion Models · ICRA 2025 |
Computational photography and imaging
image signal processing |
1.4 | 2 | 2025 | LDM-ISP: Enhancing Neural ISP for Low Light with Latent Diffusion Models · ICRA 2025 Invertible Image Signal Processing · CVPR 2021 |
Image and video processing › video processing
temporal consistency |
1.1 | 2 | 2023 | Deep Video Prior for Video Consistency and Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Blind Video Temporal Consistency via Deep Video Prior · NeurIPS 2020 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation · ACM Multimedia 2025 |
Computer vision › Video understanding and tracking
video representation learning |
0.9 | 1 | 2025 | VideoVAE+: Large Motion Video Autoencoding with Cross-Modal Video VAE · ICCV 2025 |
Image and video processing › image enhancement
low-light image enhancement |
0.9 | 1 | 2025 | LDM-ISP: Enhancing Neural ISP for Low Light with Latent Diffusion Models · ICRA 2025 |
Machine learning › Generative modeling
multimodal generation |
0.8 | 1 | 2024 | Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners · CVPR 2024 |
Audio and music processing
sound synthesis |
0.8 | 1 | 2024 | Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners · CVPR 2024 |
Visual content generation and editing
video generation |
0.8 | 1 | 2024 | Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners · CVPR 2024 |
Robotics › Robot navigation and mapping › SLAM › visual SLAM
monocular SLAM |
0.7 | 2 | 2019 | Leveraging Structural Regularity of Atlanta World for Monocular SLAM · ICRA 2019 A Monocular SLAM System Leveraging Structural Regularity in Manhattan World · ICRA 2018 |
Visual content generation and editing › image editing
image compositing |
0.6 | 1 | 2022 | Composite Photograph Harmonization with Complete Background Cues · ACM Multimedia 2022 |
Visual content generation and editing › image editing › human image editing
portrait harmonization |
0.6 | 1 | 2022 | Composite Photograph Harmonization with Complete Background Cues · ACM Multimedia 2022 |
Machine learning › Generative modeling
generative prior |
0.5 | 1 | 2021 | Unsupervised Portrait Shadow Removal via Generative Priors · ACM Multimedia 2021 |
Image and video processing
image restoration |
0.5 | 1 | 2021 | Invertible Image Signal Processing · CVPR 2021 |
Image and video processing › image restoration › shadow removal
portrait shadow removal |
0.5 | 1 | 2021 | Unsupervised Portrait Shadow Removal via Generative Priors · ACM Multimedia 2021 |
Computational photography and imaging › image signal processing
RAW image reconstruction |
0.5 | 1 | 2021 | Invertible Image Signal Processing · CVPR 2021 |
Image and video processing › image restoration
shadow removal |
0.5 | 1 | 2021 | Unsupervised Portrait Shadow Removal via Generative Priors · ACM Multimedia 2021 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.4 | 1 | 2020 | Blind Video Temporal Consistency via Deep Video Prior · NeurIPS 2020 |
Robotics › Robot navigation and mapping › SLAM
visual SLAM |
0.4 | 1 | 2019 | Leveraging Structural Regularity of Atlanta World for Monocular SLAM · ICRA 2019 |
Robotics › Robot navigation and mapping
SLAM |
0.3 | 1 | 2018 | A Monocular SLAM System Leveraging Structural Regularity in Manhattan World · ICRA 2018 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.3 | 1 | 2025 | LDM-ISP: Enhancing Neural ISP for Low Light with Latent Diffusion Models · ICRA 2025 |
Computer vision › Vision and language
cross-modal alignment |
0.2 | 1 | 2024 | Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners · CVPR 2024 |
Visual content generation and editing › image editing › image compositing
image blending |
0.2 | 1 | 2022 | Composite Photograph Harmonization with Complete Background Cues · ACM Multimedia 2022 |
Image and video coding
JPEG compression |
0.1 | 1 | 2021 | Invertible Image Signal Processing · CVPR 2021 |
Computer vision › 3D vision
structural regularity |
0.1 | 1 | 2019 | Leveraging Structural Regularity of Atlanta World for Monocular SLAM · ICRA 2019 |
Computer vision › 3D vision
3d reconstruction |
0.1 | 1 | 2018 | A Monocular SLAM System Leveraging Structural Regularity in Manhattan World · ICRA 2018 |
Methods — techniques the papers use, named apart from their topics
multimodal large language model · 1.7latent diffusion model · 1.7feature modulation · 1.7diffusion model · 1.7latent aligner · 1.5imagebind · 1.5classifier guidance · 1.5iteratively reweighted training · 1.1convolutional network · 1.1cross-modal learning · 0.9VAE · 0.9deep video prior · 0.7layer decomposition · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VideoVAE+: Large Motion Video Autoencoding with Cross-Modal Video VAE
Yazhou Xing, Yingqing He, Jingye Chen, Jiaxin Xie, Xiaowei Chi, Qifeng Chen 0001 |
ICCV | 1 |
| 2025 | LDM-ISP: Enhancing Neural ISP for Low Light with Latent Diffusion ModelsabstractEnhancing a low-light noisy RAW image into a well-exposed and clean sRGB image is a significant challenge for modern digital cameras. Prior approaches have difficulties in recovering fine-grained details and true colors of the scene under extremely low-light environments due to near-to-zero SNR. Meanwhile, diffusion models have shown significant progress towards general domain image generation. In this paper, we propose to leverage the pre-trained latent diffusion model to perform the neural ISP for enhancing extremely low-light images. Specifically, to tailor the pre-trained latent diffusion model to operate on the RAW domain, we train a set of lightweight taming modules to inject the RAW information into the diffusion denoising process via modulating the intermediate features of UNet. We further observe different roles of UNet denoising and decoder reconstruction in the latent diffusion model, which inspires us to decompose the lowlight image enhancement task into latent-space low-frequency content generation and decoding-phase high-frequency detail maintenance. Through extensive experiments on representative datasets, we demonstrate our simple design not only achieves state-of-the-art performance in quantitative evaluations but also shows significant superiority in visual comparisons over strong baselines, which highlight the effectiveness of powerful generative priors for neural ISP under extremely low-light environments. Zhefan Rao, Yazhou Xing, Qifeng Chen 0001 |
ICRA | 3 |
| 2025 | SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video GenerationabstractAudio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring semantic information, such as the classes of sounding sources present in the audio, limiting their ability to generate videos with accurate content and spatial composition. In contrast, we humans can not only naturally identify the semantic categories of sounding sources but also determine their deeply encoded spatial attributes, including locations and movement directions. This useful information can be elucidated by considering specific spatial indicators derived from the inherent physical properties of sound, such as loudness or frequency. As prior methods largely ignore this factor, we present SpA2V, the first framework explicitly exploits these spatial auditory cues from audios to generate videos with high semantic and spatial correspondence. SpA2V decomposes the generation process into two stages: 1) Audio-guided Video Planning: We meticulously adapt a state-of-the-art MLLM for a novel task of harnessing spatial and semantic cues from input audio to construct Video Scene Layouts (VSLs). This serves as an intermediate representation to bridge the gap between the audio and video modalities. 2) Layout-grounded Video Generation: We develop an efficient and effective approach to seamlessly integrate VSLs as conditional guidance into pre-trained diffusion models, enabling VSL-grounded video generation in a training-free manner. Extensive experiments demonstrate that SpA2V excels in generating realistic videos with semantic and spatial alignment to the input audios. Kien T. Pham 0001, Yingqing He, Yazhou Xing, Qifeng Chen 0001, Long Chen 0016 |
ACM Multimedia | 3 |
| 2024 | Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent AlignersabstractVideo and audio content creation serves as the core technique for the movie industry and professional users. Re-cently, existing diffusion-based methods tackle video and audio generation separately, which hinders the technique transfer from academia to industry. In this work, we aim at filling the gap, with a carefully designed optimization-based framework for cross-visual-audio and joint-visual-audio generation. We observe the powerful generation abil-ity of off-the-shelf video or audio generation models. Thus, instead of training the giant models from scratch, we pro-pose to bridge the existing strong models with a shared la-tent representation space. Specifically, we propose a mul-timodality latent aligner with the pre-trained ImageBind model. Our latent aligner shares a similar core as the clas-sifier guidance that guides the diffusion denoising process during inference time. Through carefully designed opti-mization strategy and loss functions, we show the superior performance of our method on joint video-audio generation, visual-steered audio generation, and audio-steered vi-sual generation tasks. The project website can be found at https://yzxing87.github.io/Seeing-and-Hearing/. Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang 0002, Qifeng Chen 0001 |
CVPR | 1 |
| 2023 | Deep Video Prior for Video Consistency and PropagationabstractApplying an image processing algorithm independently to each video frame often leads to temporal inconsistency in the resulting video. To address this issue, we present a novel and general approach for blind video temporal consistency. Our method is only trained on a pair of original and processed videos directly instead of a large dataset. Unlike most previous methods that enforce temporal consistency with optical flow, we show that temporal consistency can be achieved by training a convolutional network on a video with Deep Video Prior (DVP). Moreover, a carefully designed iteratively reweighted training strategy is proposed to address the challenging multimodal inconsistency problem. We demonstrate the effectiveness of our approach on 7 computer vision tasks on videos. Extensive quantitative and perceptual experiments show that our approach obtains superior performance than state-of-the-art methods on blind video temporal consistency. We further extend DVP to video propagation and demonstrate its effectiveness in propagating three different types of information (color, artistic style, and object segmentation). A progressive propagation strategy with pseudo labels is also proposed to enhance DVP's performance on video propagation. Our source codes are publicly available at https://github.com/ChenyangLEI/deep-video-prior. Chenyang Lei, Yazhou Xing, Hao Ouyang, Qifeng Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Composite Photograph Harmonization with Complete Background CuesabstractCompositing portrait photographs or videos to novel backgrounds is an important application in computational photography. Seamless blending along boundaries and globally harmonic colors are two desired properties of the photo-realistic composition of foregrounds and new backgrounds. Existing works are dedicated to either foreground alpha matte generation or after-blending harmonization, leading to sub-optimal background replacement when putting foregrounds and backgrounds together. In this work, we unify the two objectives in a single framework to obtain realistic portrait image composites. Specifically, we investigate the usage of a target background and find that a complete background plays a vital role in both seamlessly blending and harmonization. We develop a network to learn the composition process given an imperfect alpha matte with appearance features extracted from the complete background to adjust color distribution. Our dedicated usage of a complete background enables realistic portrait image composition and also temporally stable results on videos. Extensive quantitative and qualitative experiments on both synthetic and real-world data demonstrate that our method achieves state-of-the-art performance. Yazhou Xing, Yu Li 0003, Xintao Wang 0002, Ye Zhu 0003, Qifeng Chen 0001 |
ACM Multimedia | 1 |
| 2021 | Invertible Image Signal ProcessingabstractUnprocessed RAW data is a highly valuable image format for image editing and computer vision. However, since the file size of RAW data is huge, most users can only get access to processed and compressed sRGB images. To bridge this gap, we design an Invertible Image Signal Processing (InvISP) pipeline, which not only enables rendering visually appealing sRGB images but also allows recovering nearly perfect RAW data. Due to our framework’s inherent reversibility, we can reconstruct realistic RAW data instead of synthesizing RAW data from sRGB images without any memory overhead. We also integrate a differentiable JPEG compression simulator that empowers our framework to re-construct RAW data from JPEG images. Extensive quantitative and qualitative experiments on two DSLR demonstrate that our method obtains much higher quality in both rendered sRGB images and reconstructed RAW data than alternative methods. Yazhou Xing, Zian Qian, Qifeng Chen 0001 |
CVPR | 1 |
| 2021 | Unsupervised Portrait Shadow Removal via Generative PriorsabstractPortrait images often suffer from undesirable shadows cast by casual objects or even the face itself. While existing methods for portrait shadow removal require training on a large-scale synthetic dataset, we propose the first unsupervised method for portrait shadow removal without any training data. Our key idea is to leverage the generative facial priors embedded in the off-the-shelf pretrained StyleGAN2. To achieve this, we formulate the shadow removal task as a layer decomposition problem: a shadowed portrait image is constructed by the blending of a shadow image and a shadow-free image. We propose an effective progressive optimization algorithm to learn the decomposition process. Our approach can also be extended to portrait tattoo removal and watermark removal. Qualitative and quantitative experiments on a real-world portrait shadow dataset demonstrate that our approach achieves comparable performance with supervised shadow removal methods. Our source code is available at https://github.com/YingqingHe/Shadow-Removal-via-Generative-Priors. Yingqing He, Yazhou Xing, Tianjia Zhang, Qifeng Chen 0001 |
ACM Multimedia | 2 |
| 2020 | Blind Video Temporal Consistency via Deep Video PriorabstractApplying image processing algorithms independently to each video frame often leads to temporal inconsistency in the resulting video. To address this issue, we present a novel and general approach for blind video temporal consistency. Our method is only trained on a pair of original and processed videos directly instead of a large dataset. Unlike most previous methods that enforce temporal consistency with optical flow, we show that temporal consistency can be achieved by training a convolutional network on a video with the Deep Video Prior. Moreover, a carefully designed iteratively reweighted training strategy is proposed to address the challenging multimodal inconsistency problem. We demonstrate the effectiveness of our approach on 7 computer vision tasks on videos. Extensive quantitative and perceptual experiments show that our approach obtains superior performance than state-of-the-art methods on blind video temporal consistency. Chenyang Lei, Yazhou Xing, Qifeng Chen 0001 |
NeurIPS | 2 |
| 2019 | Leveraging Structural Regularity of Atlanta World for Monocular SLAMabstractA wide range of man-made environments can be abstracted as the Atlanta world. It consists of a set of Atlanta frames with a common vertical (gravitational) axis and multiple horizontal axes orthogonal to this vertical axis. This paper focuses on leveraging the regularity of Atlanta world for monocular SLAM. First, we robustly cluster image lines. Based on these clusters, we compute the local Atlanta frames in the camera frame by solving polynomial equations. Our method provides the global optimum and satisfies inherent geometric constraints. Second, we define the posterior probabilities to refine the initial clusters and Atlanta frames alternately by the maximum a posteriori estimation. Third, based on multiple local Atlanta frames, we compute the global Atlanta frames in the world frame using Kalman filtering. We optimize rotations by the global alignment and then refine translations and 3D line-based map under the directional constraints. Experiments on both synthesized and real data have demonstrated that our approach outperforms state-of-the-art methods. Haoang Li, Yazhou Xing, Ji Zhao 0001, Jean-Charles Bazin, Zhe Liu 0022, Yun-Hui Liu 0001 |
ICRA | 2 |
| 2018 | A Monocular SLAM System Leveraging Structural Regularity in Manhattan WorldabstractThe structural features in Manhattan world encode useful geometric information of parallelism, orthogonality and/or coplanarity in the scene. By fully exploiting these structural features, we propose a novel monocular SLAM system which provides accurate estimation of camera poses and 3D map. The foremost contribution of the proposed system is a structural feature-based optimization module which contains three novel optimization strategies. First, a rotation optimization strategy using the parallelism and orthogonality of 3D lines is presented. We propose a global binding method to compute an accurate estimation of the absolute rotation of the camera. Then we propose an approach for calculating the relative rotation to further refine the absolute rotation. Second, a translation optimization strategy leveraging coplanarity is proposed. Coplanar features are effectively identified, and we leverage them by a unified model handling both points and lines to calculate the relative translation, and then the optimal absolute translation. Third, a 3D line optimization strategy utilizing parallelism, orthogonality and coplanarity simultaneously is proposed to obtain an accurate 3D map consisting of structural line segments with low computational complexity. Experiments in man-made environments have demonstrated that the proposed system outperforms existing state-of-the-art monocular SLAM systems in terms of accuracy and robustness. Haoang Li, Jian Yao 0002, Jean-Charles Bazin, Xiaohu Lu, Yazhou Xing, Kang Liu 0003 |
ICRA | 5 |