VLDB 2026 Research / reviewers in the wild / expert
Qinglin Lu
dblp:165/1885
· DBLP profile ↗
21ranked-venue papers
3as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Phased One-Step Adversarial Equilibrium for Video Diffusion ModelsabstractVideo diffusion generation suffers from critical sampling efficiency bottlenecks, particularly for large-scale models and long contexts. Existing video acceleration methods, adapted from image-based techniques, lack a single-step distillation ability for large-scale video models and task generalization for conditional downstream tasks. To bridge this gap, we propose the Video Phased Adversarial Equilibrium (V-PAE), a distillation framework that enables high-quality, single-step video generation from large-scale video models. Our approach employs a two-phase process. (i) Stability priming is a warm-up process to align the distributions of real and generated videos. It improves the stability of single-step adversarial distillation in the following process. (ii) Unified adversarial equilibrium is a flexible self-adversarial process that reuses generator parameters for the discriminator backbone. It achieves a co-evolutionary adversarial equilibrium in the Gaussian noise space. For the conditional tasks, we primarily preserve video-image subject consistency, which is caused by semantic degradation and conditional frame collapse during the distillation training in image-to-video (I2V) generation. Comprehensive experiments on VBench-I2V demonstrate that V-PAE outperforms existing acceleration methods by an average of 5.8% in the overall quality score, including semantic alignment, temporal coherence, and frame quality. In addition, our approach reduces the diffusion latency of the large-scale video model (e.g., Wan2.1-I2V-14B) by 100 times, while preserving competitive performance. Jiaxiang Cheng, Xuhua Ren, Hongyi Henry Jin, Tianxiang Zheng 0001, Qinglin Lu |
AAAI | 10 |
| 2025 | Local Conditional Controlling for Text-to-Image Diffusion ModelsabstractDiffusion models have exhibited impressive prowess in the text-to-image task. Recent methods add image-level structure controls, e.g., edge and depth maps, to manipulate the generation process together with text prompts to obtain desired images. This controlling process is globally operated on the entire image, which limits the flexibility of control regions. In this paper, we explore a novel and practical task setting: local control. It focuses on controlling specific local region according to user-defined image conditions, while the remaining regions are only conditioned by the original text prompt. However, it is non-trivial to achieve it. The naive manner of directly adding local conditions may lead to the local control dominance problem, which forces the model to focus on the controlled region and neglect object generation in other regions. To mitigate this problem, we propose Regional Discriminate Loss to update the noised latents, aiming at enhanced object generation in non-control regions. Furthermore, the proposed Focused Token Response suppresses weaker attention scores which lack the strongest response to enhance object distinction and reduce duplication. Lastly, we adopt Feature Mask Constraint to reduce quality degradation in images caused by information differences across the local control region. All proposed strategies are operated at the inference stage. Extensive experiments demonstrate that our method can synthesize high-quality images aligned with the text prompt under local control conditions. Yang Yang 0002, Zekai Luo, Hengjia Li, Zheng Yang 0008, Xiaofei He 0001, Wei Zhao 0019, Qinglin Lu, Wei Liu 0005, Boxi Wu 0001 |
AAAI | 10 |
| 2025 | Concept-Edge Fusion: Background Generation for Product Presentation Based on Text-to-Image Model
Pengfei Deng, Weize Quan, Hanyu Wang 0002, Qinglin Lu, Zhifeng Li 0001, Dong-Ming Yan 0001 |
CVM (2) | 5 |
| 2025 | Sonic: Shifting Focus to Global Audio Perception in Portrait AnimationabstractThe study of talking face generation mainly explores the intricacies of synchronizing facial movements and crafting visually appealing, temporally-coherent animations. However, due to the limited exploration of global audio perception, current approaches predominantly employ auxiliary visual and spatial knowledge to stabilize the movements, which often results in the deterioration of the naturalness and temporal inconsistencies. Considering the essence of audio-driven animation, the audio signal serves as the ideal and unique priors to adjust facial expressions and lip movements, without resorting to interference of any visual signals. Based on this motivation, we propose a novel paradigm, dubbed as Sonic, to shift focus on the exploration of global audio perception. To effectively leverage global audio knowledge, we disentangle it into intra-and inter-clip audio perception and collaborate with both aspects to enhance overall perception. For the intra-clip audio perception, 1). Context-enhanced audio learning, in which long-range intra-clip temporal audio knowledge is extracted to provide facial expression and lip motion priors implicitly expressed as the tone and speed of speech. 2). Motion-decoupled controller, in which the motion of the head and expression movement are disentangled and independently controlled by intra-audio clips. Most importantly, for inter-clip audio perception, as a bridge to connect the intra-clips to achieve the global perception, Time-aware position shift fusion, in which the global inter-clip audio information is considered and fused for long-audio inference via through consecutively time-aware shifted windows. Extensive experiments demonstrate that the novel audio-driven paradigm outperform existing SOTA methodologies in terms of video quality, temporally consistency, lip synchronization precision, and motion diversity. Xiaozhong Ji, Xiaobin Hu, Chuming Lin, Qingdong He, Jiangning Zhang, Donghao Luo 0001, Qin Lin 0003, Qinglin Lu, Chengjie Wang 0001 |
CVPR | 11 |
| 2025 | HunyuanPortrait: Implicit Condition Control for Enhanced Portrait AnimationabstractWe introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the character in the reference image by the facial expression and head pose of the driving videos. In our framework, we utilize pre-trained encoders to achieve the decoupling of portrait motion information and identity in videos. To do so, implicit representation is adopted to encode motion information and is employed as control signals in the animation phase. By leveraging the power of stable video diffusion as the main building block, we carefully design adapter layers to inject control signals into the denoising unet through attention mechanisms. These bring spatial richness of details and temporal consistency. HunyuanPortrait also exhibits strong generalization performance, which can effectively disentangle appearance and motion under different image styles. Our framework outperforms existing methods, demonstrating superior temporal consistency and controllability. Our project is available at HunyuanPortrait. Zunnan Xu, Zhentao Yu, Xiaoyu Jin, Fa-Ting Hong, Xiaozhong Ji, Chengfei Cai, Shiyu Tang, Qin Lin 0003, Xiu Li 0001, Qinglin Lu |
CVPR | 13 |
| 2025 | FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language ModelabstractCurrently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of vision language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-grained editing. To address these issues, we propose FireEdit, an innovative Fine-grained Instruction-based image editing framework that exploits a REgion-aware VLM. FireEdit is designed to accurately comprehend user instructions and ensure effective control over the editing process. Specifically, we enhance the fine-grained visual perception capabilities of the VLM by introducing additional region tokens. Relying solely on the output of the LLM to guide the diffusion model may lead to suboptimal editing results. Therefore, we propose a Time-Aware Target Injection module and a Hybrid Visual Cross Attention module. The former dynamically adjusts the guidance strength at various denoising stages by integrating timestep embeddings with the text embeddings. The latter enhances visual details for image editing, thereby preserving semantic consistency between the edited result and the source image. By combining the VLM enhanced with fine-grained region tokens and the time-dependent diffusion model, FireEdit demonstrates significant advantages in comprehending editing instructions and maintaining high semantic consistency. Extensive experiments indicate that our approach surpasses the state-of-the-art instruction-based image editing methods. Jiahao Li 0001, Zunnan Xu, Yiji Cheng, Fa-Ting Hong, Qin Lin 0003, Qinglin Lu, Xiaodan Liang |
CVPR | 8 |
| 2025 | Audio-Visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationabstractTalking head synthesis is vital for virtual avatars and human-computer interaction. However, most existing methods are typically limited to accepting control from a single primary modality, restricting their practical utility. To this end, we introduce \textbf{ACTalker}, an end-to-end video diffusion framework that supports both multi-signals control and single-signal control for talking head video generation. For multiple control, we design a parallel mamba structure with multiple branches, each utilizing a separate driving signal to control specific facial regions. A gate mechanism is applied across all branches, providing flexible control over video generation. To ensure natural coordination of the controlled video both temporally and spatially, we employ the mamba structure, which enables driving signals to manipulate feature tokens across both dimensions in each branch. Additionally, we introduce a mask-drop strategy that allows each driving signal to independently control its corresponding facial region within the mamba structure, preventing control conflicts. Experimental results demonstrate that our method produces natural-looking facial videos driven by diverse signals and that the mamba layer seamlessly integrates multiple driving modalities without conflict. The project website can be found at https://harlanhong.github.io/publications/actalker/index.html. Fa-Ting Hong, Zunnan Xu, Xiu Li 0001, Qin Lin 0003, Qinglin Lu, Dan Xu 0002 |
ICCV | 7 |
| 2025 | PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and EnhancementabstractDespite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction.
In this paper, we propose PolyVivid, a multi-subject video customization framework that enables flexible and identity-consistent generation. To establish accurate correspondences between subject images and textual entities, we design a VLLM-based text-image fusion module that embeds visual identities into the textual space for precise grounding. To further enhance identity preservation and subject interaction, we propose a 3D-RoPE-based enhancement module that enables structured bidirectional fusion between text and image embeddings. Moreover, we develop an attention-inherited identity injection module to effectively inject fused identity features into the video generation process, mitigating identity drift. Finally, we construct an MLLM-based data pipeline that combines MLLM-based grounding, segmentation, and a clique-based subject consolidation strategy to produce high-quality multi-subject data, effectively enhancing subject distinction and reducing ambiguity in downstream video generation.
Extensive experiments demonstrate that PolyVivid achieves superior performance in identity fidelity, video realism, and subject alignment, outperforming existing open-source and commercial baselines. More comprehensive video results and comparisons are shown on the project page in the supplementary material. Zhentao Yu, Zhengguang Zhou, Jiangning Zhang, Qinglin Lu, Ran Yi 0002 |
NeurIPS | 6 |
| 2025 | Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-TuningabstractRecent advances in multimodal Reward Models (RMs) have shown significant promise in delivering reward signals to align vision models with human preferences. However, current RMs are generally restricted to providing direct responses or engaging in shallow reasoning processes with limited depth, often leading to inaccurate reward signals. We posit that incorporating explicit long chains of thought (CoT) into the reward reasoning process can significantly strengthen their reliability and robustness. Furthermore, we believe that once RMs internalize CoT reasoning, their direct response accuracy can also be improved through implicit reasoning capabilities. To this end, this paper proposes UnifiedReward-Think, the first unified multimodal CoT-based reward model, capable of multi-dimensional, step-by-step long-chain reasoning for both visual understanding and generation reward tasks. Specifically, we adopt an exploration-driven reinforcement fine-tuning approach to elicit and incentivize the model's latent complex reasoning ability: (1) We first use a small amount of image generation preference data to distill the reasoning process of GPT-4o, which is then used for the model's cold start to learn the format and structure of CoT reasoning. (2) Subsequently, by leveraging the model's prior knowledge and generalization capabilities, we prepare large-scale unified multimodal preference data to elicit the model's reasoning process across various vision tasks. During this phase, correct reasoning outputs are retained for rejection sampling to refine the model (3) while incorrect predicted samples are finally used for Group Relative Policy Optimization (GRPO) based reinforcement fine-tuning, enabling the model to explore diverse reasoning paths and optimize for correct and robust solutions. Extensive experiments confirm that incorporating long CoT reasoning significantly enhances the accuracy of reward signals. Notably, after mastering CoT reasoning, the model exhibits implicit reasoning capabilities, allowing it to surpass existing baselines even without explicit reasoning traces. Yuhang Zang, Qinglin Lu, Jiaqi Wang 0003 |
NeurIPS | 5 |
| 2025 | HOMA: Towards Generic Human-Object Interaction in Multimodal Driven Human Animation with Weak ConditionsabstractWhile recent advances in human-object interaction (HOI) video generation showcase promising capabilities for synthesizing coordinated human-object dynamics, existing methods remain constrained by their reliance on meticulously curated motion sequences and actor-specific data, thereby limiting practical scalability and user accessibility. Furthermore, generalization to novel object appearances and interaction scenarios remains understudied. To address these limitations, we propose HOMA, a weakly conditioned multimodal-driven HOI video generation framework that introduces sparse, decoupled motion guidance to enhance controllability and reduce dependency on stringent input conditions. Our approach encodes appearance and motion signals into the dual input space of a multimodal diffusion transformer (MMDiT), fusing them within a shared context space to enable temporally consistent and physically plausible interactions. To optimize learning efficiency and feature injection accuracy, we introduce a parameter-space HOI adapter initialized with pretrained MMDiT weights to preserve prior knowledge while enabling efficient adaptation. Additionally, we design a facial cross-attention adapter for audio-driven lip synchronization, ensuring anatomically accurate speech animation. Extensive experiments demonstrate that HOMA achieves state-of-the-art performance in interaction naturalness and generalization under weak supervision, outperforming existing methods by significant margins. We further illustrate HOMA’s versatility through diverse applications, including text-conditioned generation and interactive object manipulation, facilitated by a user-friendly demo interface. The project page is https://bone-11.github.io/homa-page/. Ziyao Huang 0002, Juan Cao 0001, Yifeng Ma 0006, Zejing Rao, Qin Lin 0003, Qinglin Lu, Fan Tang |
SIGGRAPH Asia | 11 |
| 2025 | LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion ModelsabstractCustomization generation techniques have significantly advanced the synthesis of specific concepts across varied contexts. Multi-concept customization emerges as the challenging task within this domain. Existing approaches often rely on training a fusion matrix of multiple Low-Rank Adaptations (LoRAs) to merge various concepts into a single image. However, we identify this straightforward method faces two major challenges: 1) concept confusion, where the model struggles to preserve distinct individual characteristics, and 2) concept vanishing, where the model fails to generate the intended subjects. To address these issues, we introduce LoRA-Composer, a training-free framework designed for seamlessly integrating multiple LoRAs, thereby enhancing the harmony among different concepts within generated images. LoRA-Composer addresses concept vanishing through concept injection constraints, enhancing visibility via an expanded cross-attention mechanism. To combat concept confusion, concept isolation constraints are introduced, refining the self-attention computation. Furthermore, we propose two inference techniques to accelerate inference speed without performance degradation and enhance the accuracy of the generated region, respectively. Extensive experiments demonstrate that LoRA-Composer significantly outperforms standard baselines, especially in scenarios without image-based conditions such as canny edge or pose estimation. Yang Yang 0002, Chaotian Song, Hengjia Li, Qinglin Lu, Deng Cai 0001, Xiaofei He 0001, Boxi Wu 0001, Wei Liu 0005 |
IEEE Trans. Image Process. | 8 |
| 2023 | GFFT: a Task Graph Based Fast Fourier Transform Optimization FrameworkabstractFast Fourier Transform (FFT) is a widely used mathematical tool in scientific and engineering applications, and optimizing its performance remains a challenging problem. This paper introduces GFFT, a novel task-graph-based FFT optimization framework that leverages modern hardware and software techniques to achieve high-performance computation. GFFT features a tuning model that uses hardware parameters to optimize FFT decomposition, a bi-directional recursive FFT algorithm that avoids strided load in SIMD implementation, and several graph optimizers inspired by deep learning frameworks to enhance performance. In addition, GFFT utilizes task-based parallelism to exploit performance on multi-core processors and provide potential compatibility with heterogeneous systems. Experimental results demonstrate that GFFT outperforms popular FFT frameworks, achieving an average speedup of 1.17x to FFTW and 1.27x to oneMKL on the Intel Xeon processor, 1.18x to AOCL-FFTW on the AMD EPYC processor, and 2.11x to FFTW on the Sunway multi-core processor with a single thread. Additionally, GFFT achieves an average speedup of 11.48x to FFTW and 1.41x to oneMKL on the Intel Xeon processor, 9.87x to AOCL-FFTW on the AMD EPYC processor with 16-threads. Qinglin Lu, Wenjing Ma, Daokun Chen, Fangfang Liu 0004 |
ICPP | 1 |
| 2023 | xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002 |
CCF Trans. High Perform. Comput. | 6 |
| 2023 | Publisher Correction: xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002 |
CCF Trans. High Perform. Comput. | 6 |
| 2023 | An Optimized Framework for Matrix Factorization on the New Sunway Many-core PlatformabstractMatrix factorization functions are used in many areas and often play an important role in the overall performance of the applications. In the LAPACK library, matrix factorization functions are implemented with blocked factorization algorithm, shifting most of the workload to the high-performance Level-3 BLAS functions. But the non-blocked part, the panel factorization, becomes the performance bottleneck, especially for small- and medium-size matrices that are the common cases in many real applications. On the new Sunway many-core platform, the performance bottleneck of panel factorization can be alleviated by keeping the panel in the LDM for the panel factorization. Therefore, we propose a new framework for implementing matrix factorization functions on the new Sunway many-core platform, facilitating the in-LDM panel factorization. The framework provides a template class with wrapper functions, which integrates inter-CPE communication for the Level-1 and Level-2 BLAS functions with flexible interfaces and can accommodate different partitioning schemes. With the framework, writing panel factorization code with data residing in the LDM space can be done with much higher productivity. We implemented three functions ( dgetrf , dgeqrf , and dpotrf ) based on the framework and compared our work with a CPE_BLAS version, which uses the original LAPACK implementation linked with optimized BLAS library that runs on the CPE mesh. Using the most favorable partitioning, the panel factorization part achieves speedup of up to 26.3, 19.1, and 18.2 for the three matrix factorization functions. For the whole function, our implementation is based on a carefully tuned recursion framework, and we added specific optimization to some subroutines used in the factorization functions. Overall, we obtained average speedup of 9.76 on dgetrf , 10.12 on dgeqrf , and 4.16 on dpotrf , compared to the CPE_BLAS version. Based on the current template class, our work can be extended to support more categories of linear algebra functions. Wenjing Ma, Fangfang Liu 0004, Daokun Chen, Qinglin Lu, Hongsen Wang, Xinhui Yuan |
ACM Trans. Archit. Code Optim. | 4 |
| 2022 | Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
Yunlong Tang 0002, Siting Xu, Teng Wang 0007, Qin Lin 0003, Qinglin Lu, Feng Zheng 0001 |
ACCV (2) | 5 |
| 2022 | 2.5 Million-Atom Ab Initio Electronic-Structure Simulation of Complex Metallic Heterostructures with DGDFTabstractOver the past three decades, ab initio electronic structure calculations of large, complex and metallic systems are limited to tens of thousands of atoms in computational accuracy and efficiency on leadership supercomputers. We present a massively parallel discontinuous Galerkin density functional theory (DGDFT) implementation, which adopts adaptive local basis functions to discretize the Kohn-Sham equation, resulting in a block-sparse Hamiltonian matrix. A highly efficient pole expansion and selected inversion (PEXSI) sparse direct solver is implemented in DGDFT to achieve O(N1.5) scaling for quasi two-dimensional systems. DGDFT allows us to compute the electronic structures of complex metallic heterostructures with 2.5 million atoms (17.2 million electrons) using 35.9 million cores on the new Sunway supercomputer. The peak performance of PEXSI can achieve 64 PFLOPS (~5% of theoretical peak), which is un-precedented for sparse direct solvers. This accomplishment paves the way for quantum mechanical simulations into mesoscopic scale for designing next-generation electronic devices. Wei Hu 0006, Hong An, Zhuoqiang Guo, Qingcai Jiang, Xinming Qin, Junshi Chen 0003, Weile Jia, Chao Yang 0001, Zhaolong Luo, Jielan Li, Wentiao Wu, Guangming Tan, Dongning Jia, Qinglin Lu, Yeqi Huang, Liyi Wang, Jinlong Yang 0003 |
SC | 14 |
| 2021 | Overview of Tencent Multi-modal Ads Video UnderstandingabstractMulti-modal Ads Video Understanding Challenge is the first grand challenge aiming to comprehensively understand ads videos. Our challenge includes two tasks: video structuring and multi-label classification. Video structuring asks the participants to accurately predict both the scene boundaries and the multi-label categories of each scene based on a fine-grained and ads-related category hierarchy. This task will advance the foundation of comprehensive ads video understanding, which has a significant impact on many applications in ads, such as video recommendation and user behavior analysis. This paper presents an overview of the video structuring task in our grand challenge, including the background of ads videos, an elaborate description of this task, our proposed dataset, the evaluation protocol, and our baseline model. By ablating the key components of our baseline, we would like to reveal the main challenges of this task and provide useful guidance for future research of this area. Zhenzhi Wang 0001, Liyu Wu, Jiangfeng Xiong, Qinglin Lu |
ACM Multimedia | 5 |
| 2021 | Better Learning Shot Boundary Detection via Multi-taskabstractShot boundary detection (SBD) plays an important role in video understanding, since most recent works take the shot as minimal granularity instead of frames for upstream tasks. However, the large variations of hard-cut and gradual-change transitions within shots significantly limit the performance of SBD. To deal with the variations, we propose a multi-task architecture called Transnet++. Transnet++ disentangles the two types of transition and adopts two separate branches to predict them respectively. Two branches share the same video knowledge space and their results are fused for final prediction. Moreover, we propose a spatial attention module (SAM) to enhance the feature representations which suffers from redundant padding region. Meanwhile, a temporal attention module (TAM) is applied to capture the long-term information of the video for alleviating the over-segmentation problem. Experimental results (91.16% f1-score) on Tencent AVS Dataset demonstrate the effectiveness and superiority of Transnet++ for SBD. Qinglin Lu |
ACM Multimedia | 3 |
| 2015 | Entire Reflective Object Surface Structure Understanding
Qinglin Lu, Olivier Laligant, Eric Fauvet, Anastasia Zakharova |
BMVC | 1 |
| 2015 | Entire reflective object surface structure understanding based on reflection motion estimation
Qinglin Lu, Eric Fauvet, Anastasia Zakharova, Olivier Laligant |
Pattern Recognit. Lett. | 1 |