EDBT 2026 Demo / reviewers in the wild / expert
Xiaoqian Lv
dblp:68/10889
· DBLP profile ↗
14ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-8735-1561ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ProsodyTalker: 3D Visual Speech Animation via Prosody DecompositionabstractMost existing 3D visual speech animation methods synthesize lip movements synchronized with speech, which however neglect head poses and therefore degrade the animation realism. The animation of head poses presents two primary challenges: (1) the intricate mapping between speech and head poses remains poorly understood and (2) the absence of 4D face datasets featuring realistic head poses. Inspired by prosody decomposition in speech processing, we discern that head movements correlate with the fundamental frequency (F0) of speech prosody, while lip movements align with the language content. These observations motivate us to propose a novel framework, dubbed ProsodyTalker, that concurrently synthesizes lip and head movements, grounded in the principles of prosody decomposition. The core idea is first to adopt information perturbation to explicitly decompose the speech prosody into pose-related F0 and lip-related language content. Then, an autoregressive content-oriented fusion decoder is employed to enhance lip synchronization in the synthesized facial sequences. To synthesize head poses, we design a transformer-based variational autoencoder to learn a latent distribution of facial sequences and propose an F0-conditioned latent diffusion model to establish a probabilistic mapping from F0 to pose-related latent codes. Furthermore, we contribute a large-scale 4D face dataset containing bunches of variations in identities, head poses and facial motions. Extensive experiments show that our method achieves more realistic animation than state-of-the-art methods. Zonglin Li 0004, Xiaoqian Lv, Qinglin Liu, Quanling Meng, Xin Sun 0003, Shengping Zhang |
AAAI | 2 |
| 2025 | Path-Adaptive Matting for Efficient Inference Under Various Computational Cost ConstraintsabstractIn this paper, we explore a novel image matting task aimed at achieving efficient inference under various computational cost constraints, specifically FLOP limitations, using a single matting network. Existing matting methods which have not explored scalable architectures or path-learning strategies, fail to tackle this challenge. To overcome these limitations, we introduce Path-Adaptive Matting (PAM), a framework that dynamically adjusts network paths based on image contexts and computational cost constraints. We formulate the training of the computational cost-constrained matting network as a bilevel optimization problem, jointly optimizing the matting network and the path estimator. Building on this formalization, we design a path-adaptive matting architecture by incorporating path selection layers and learnable connect layers to estimate optimal paths and perform efficient inference within a unified network. Furthermore, we propose a performance-aware path-learning strategy to generate path labels online by evaluating a few paths sampled from the prior distribution of optimal paths and network estimations, enabling robust and efficient online path learning. Experiments on five image matting datasets demonstrate that the proposed PAM framework achieves competitive performance across a range of computational cost constraints. Qinglin Liu, Zonglin Li 0004, Xiaoqian Lv, Xin Sun 0003, Ru Li 0002, Shengping Zhang |
AAAI | 3 |
| 2024 | Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and DeblurringabstractExisting joint low-light enhancement and deblurring methods learn pixel-wise mappings from paired synthetic data, which results in limited generalization in real-world scenes. While some studies explore the rich generative prior of pre-trained diffusion models, they typically rely on the assumed degradation process and cannot handle unknown real-world degradations well. To address these problems, we propose a novel zero-shot framework, FourierDiff, which embeds Fourier priors into a pre-trained diffusion model to harmoniously handle the joint degradation of luminance and structures. FourierDiff is appealing in its relaxed requirements on paired training data and degradation assumptions. The key zero-shot insight is motivated by image characteristics in the Fourier domain: most luminance information concentrates on amplitudes while structure and content information are closely related to phases. Based on this observation, we decompose the sampled results of the reverse diffusion process in the Fourier domain and take advantage of the amplitude of the generative prior to align the enhanced brightness with the distribution of natural images. To yield a sharp and content-consistent enhanced result, we further design a spatial-frequency alternating optimization strategy to progressively refine the phase of the input. Extensive experiments demonstrate the superior effectiveness of the proposed method, especially in real-world scenes. The code is available at https://github.com/aipixel/FourierDiff. Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Yichen Zheng, Bineng Zhong 0001, Chongyi Li, Liqiang Nie |
CVPR | 1 |
| 2024 | DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-Based Human Video GenerationabstractExisting diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In this paper, we propose a novel framework, DiffPerformer, to synthesize high-fidelity and temporally consistent human video. Without complex architecture modification or costly training, DiffPerformer finetunes a pre-trained diffusion model on a single video of the target character and introduces an implicit video representation as a proxy to learn temporally consistent guidance for the diffusion model. The guidance is encoded into VAE latent space and an iterative optimization loop is constructed between the implicit video representation and the diffusion model, allowing to harness the smooth property of the implicit video representation and the generative capabilities of the diffusion model in a mutually beneficial way. Moreover, we propose 3D-aware human flow as a temporal constraint during the optimization to explicitly model the correspondence between driving poses and human appearance. This alleviates the mis-alignment between driving poses and target performer and therefore maintains the appearance coherence under various motions. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. The code is available at https://github.com/aipixel/DiffPerformer. Chenyang Wang 0002, Zerong Zheng, Tao Yu 0007, Xiaoqian Lv, Bineng Zhong 0001, Shengping Zhang, Liqiang Nie |
CVPR | 4 |
| 2024 | Revisiting Context Aggregation for Image MattingabstractTraditional studies emphasize the significance of context information in improving matting performance. Consequently, deep learning-based matting methods delve into designing pooling or affinity-based context aggregation modules to achieve superior results. However, these modules cannot well handle the context scale shift caused by the difference in image size during training and inference, resulting in matting performance degradation. In this paper, we revisit the context aggregation mechanisms of matting networks and find that a basic encoder-decoder network without any context aggregation modules can actually learn more universal context aggregation, thereby achieving higher matting performance compared to existing methods. Building on this insight, we present AEMatter, a matting network that is straightforward yet very effective. AEMatter adopts a Hybrid-Transformer backbone with appearance-enhanced axis-wise learning (AEAL) blocks to build a basic network with strong context aggregation learning capability. Furthermore, AEMatter leverages a large image training strategy to assist the network in learning context aggregation from data. Extensive experiments on five popular matting datasets demonstrate that the proposed AEMatter outperforms state-of-the-art matting methods by a large margin. The source code is available at https://github.com/aipixel/AEMatter. Qinglin Liu, Xiaoqian Lv, Quanling Meng, Zonglin Li 0004, Xiangyuan Lan, Shuo Yang 0006, Shengping Zhang, Liqiang Nie |
ICML | 2 |
| 2024 | Dual-context aggregation for universal image matting
Qinglin Liu, Xiaoqian Lv, Wei Yu 0002, Changyong Guo, Shengping Zhang |
Multim. Tools Appl. | 2 |
| 2024 | Human Selective MattingabstractExisting human matting methods are incapable of accurately estimating the alpha mattes of arbitrarily selected humans from a group photo. An alternative solution is to apply them to the corresponding cropped image patches. However, this option obtains an inaccurate alpha estimation due to the interference of the body parts of the neighboring humans. In addition, these methods are only trained on finely annotated synthetic data, which causes poor performance in real-world scenarios due to the domain shift. To address these problems, we propose human selective matting (HSMatt), which performs matting for arbitrarily selected humans from a group photo given only a simple bounding box as guidance. Specifically, we design a global–local context network to extract both local and global semantic context features. A human-aware trimap network is then proposed to generate human-aware trimaps for the selected humans, which adopts stacked bidirectional inference modules with intermediate supervision to progressively refine the estimated trimap. Finally, a partially supervised matting network is introduced to estimate the alpha matte, which uses a sample-varying loss to train the network on both the finely annotated synthetic data and coarsely annotated real-world data, resulting in high accuracy and good generalization. To evaluate the proposed HSMatt, we construct the first human selective matting dataset, named HSM-200K, which contains over 200,000 human images with instance-level alpha matte annotations. Experimental results demonstrate that the proposed HSMatt outperforms state-of-the-art methods. Qinglin Liu, Quanling Meng, Xiaoqian Lv, Zonglin Li 0004, Wei Yu 0004, Shengping Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Rectangular-Output Image StitchingabstractImage stitching aims to combine two images with overlapping fields to expand the field-of-view (FoV). However, the stitched images of existing methods are irregular, and need to be processed by rectangling methods, which is time-consuming and prone to be unnatural. In this paper, we propose the first end-to-end framework, Rectangular-output Deep Image Stitching Network (RDISNet), to directly stitch two images into a standard rectangular image while learning color consistency between image pairs and maintaining the authenticity of the content. To further preserve the structure of large objects in the stitched image, we design a dilated BN-RCU block to expand the receptive field of RDISNet for extracting enriched spatial context. Furthermore, we design a novel data synthesis pipeline and build the first rectangular-output deep image stitching dataset (RDIS-D) for jointing image stitching and rectangling. Experimental results demonstrate that RDISNet performs favorably against the state-of-the-art methods. Hongfei Zhou, Yuhe Zhu, Xiaoqian Lv, Qinglin Liu, Shengping Zhang |
ICIP | 3 |
| 2023 | Unsupervised Low-Light Video Enhancement With Spatial-Temporal Co-Attention TransformerabstractExisting low-light video enhancement methods are dominated by Convolution Neural Networks (CNNs) that are trained in a supervised manner. Due to the difficulty of collecting paired dynamic low/normal-light videos in real-world scenes, they are usually trained on synthetic, static, and uniform motion videos, which undermines their generalization to real-world scenes. Additionally, these methods typically suffer from temporal inconsistency (e.g., flickering artifacts and motion blurs) when handling large-scale motions since the local perception property of CNNs limits them to model long-range dependencies in both spatial and temporal domains. To address these problems, we propose the first unsupervised method for low-light video enhancement to our best knowledge, named LightenFormer, which models long-range intra- and inter-frame dependencies with a spatial-temporal co-attention transformer to enhance brightness while maintaining temporal consistency. Specifically, an effective but lightweight S-curve Estimation Network (SCENet) is first proposed to estimate pixel-wise S-shaped non-linear curves (S-curves) to adaptively adjust the dynamic range of an input video. Next, to model the temporal consistency of the video, we present a Spatial-Temporal Refinement Network (STRNet) to refine the enhanced video. The core module of STRNet is a novel Spatial-Temporal Co-attention Transformer (STCAT), which exploits multi-scale self- and cross-attention interactions to capture long-range correlations in both spatial and temporal domains among frames for implicit motion estimation. To achieve unsupervised training, we further propose two non-reference loss functions based on the invertibility of the S-curve and the noise independence among frames. Extensive experiments on the SDSD and LLIV-Phone datasets demonstrate that our LightenFormer outperforms state-of-the-art methods. Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Weigang Zhang, Hongxun Yao, Qingming Huang |
IEEE Trans. Image Process. | 1 |
| 2022 | BacklitNet: A dataset and network for backlit image enhancement
Xiaoqian Lv, Shengping Zhang, Qinglin Liu, Haozhe Xie, Bineng Zhong 0001, Huiyu Zhou 0001 |
Comput. Vis. Image Underst. | 1 |
| 2021 | Is It Easy to Recognize Baby's Age and Gender?
Yang Liu 0119, Ruili He, Xiaoqian Lv, Wei Wang 0107, Xin Sun 0003, Shengping Zhang |
J. Comput. Sci. Technol. | 3 |
| 2021 | Low-light image enhancement via deep Retinex decomposition and bilateral learning
Xiaoqian Lv, Yujing Sun 0004, Jun Zhang 0017, Feng Jiang 0001, Shengping Zhang |
Signal Process. Image Commun. | 1 |
| 2019 | Proposal-Refined Weakly Supervised Object Detection in Underwater Images
Xiaoqian Lv, Qinglin Liu, Jiamin Sun, Shengping Zhang |
ICIG (1) | 1 |
| 2012 | Low-Complexity Compression Method for Hyperspectral Images Based on Distributed Source CodingabstractIn this letter, we propose a low-complexity discrete cosine transform (DCT)-based distributed source coding (DSC) scheme for hyperspectral images. First, the DCT was applied to the hyperspectral images. Then, set-partitioning-based approach was utilized to reorganize DCT coefficients into waveletlike tree structure and extract the sign, refinement, and significance bitplanes. Third, low-density parity-check-based Slepian-Wolf (SW) coder was adopted to implement the DSC strategy. Finally, an auxiliary reconstruction method was employed to improve the reconstruction quality. Experimental results on Airborne Visible/Infrared Imaging Spectrometer data set show that the proposed paradigm significantly outperforms the DSC-based coder in wavelet transform domain (set partitioning in hierarchical tree with SW coding), and its performance is comparable to that of the DSC scheme based on informed quantization at low bit rate. Xuzhou Pan, Rongke Liu, Xiaoqian Lv |
IEEE Geosci. Remote. Sens. Lett. | 3 |