VLDB 2026 Research / reviewers in the wild / expert
Qin Lin 0003
dblp:88/3408-3
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0008-3462-9735ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sonic: Shifting Focus to Global Audio Perception in Portrait AnimationabstractThe study of talking face generation mainly explores the intricacies of synchronizing facial movements and crafting visually appealing, temporally-coherent animations. However, due to the limited exploration of global audio perception, current approaches predominantly employ auxiliary visual and spatial knowledge to stabilize the movements, which often results in the deterioration of the naturalness and temporal inconsistencies. Considering the essence of audio-driven animation, the audio signal serves as the ideal and unique priors to adjust facial expressions and lip movements, without resorting to interference of any visual signals. Based on this motivation, we propose a novel paradigm, dubbed as Sonic, to shift focus on the exploration of global audio perception. To effectively leverage global audio knowledge, we disentangle it into intra-and inter-clip audio perception and collaborate with both aspects to enhance overall perception. For the intra-clip audio perception, 1). Context-enhanced audio learning, in which long-range intra-clip temporal audio knowledge is extracted to provide facial expression and lip motion priors implicitly expressed as the tone and speed of speech. 2). Motion-decoupled controller, in which the motion of the head and expression movement are disentangled and independently controlled by intra-audio clips. Most importantly, for inter-clip audio perception, as a bridge to connect the intra-clips to achieve the global perception, Time-aware position shift fusion, in which the global inter-clip audio information is considered and fused for long-audio inference via through consecutively time-aware shifted windows. Extensive experiments demonstrate that the novel audio-driven paradigm outperform existing SOTA methodologies in terms of video quality, temporally consistency, lip synchronization precision, and motion diversity. Xiaozhong Ji, Xiaobin Hu, Chuming Lin, Qingdong He, Jiangning Zhang, Donghao Luo 0001, Qin Lin 0003, Qinglin Lu, Chengjie Wang 0001 |
CVPR | 10 |
| 2025 | HunyuanPortrait: Implicit Condition Control for Enhanced Portrait AnimationabstractWe introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the character in the reference image by the facial expression and head pose of the driving videos. In our framework, we utilize pre-trained encoders to achieve the decoupling of portrait motion information and identity in videos. To do so, implicit representation is adopted to encode motion information and is employed as control signals in the animation phase. By leveraging the power of stable video diffusion as the main building block, we carefully design adapter layers to inject control signals into the denoising unet through attention mechanisms. These bring spatial richness of details and temporal consistency. HunyuanPortrait also exhibits strong generalization performance, which can effectively disentangle appearance and motion under different image styles. Our framework outperforms existing methods, demonstrating superior temporal consistency and controllability. Our project is available at HunyuanPortrait. Zunnan Xu, Zhentao Yu, Xiaoyu Jin, Fa-Ting Hong, Xiaozhong Ji, Chengfei Cai, Shiyu Tang, Qin Lin 0003, Xiu Li 0001, Qinglin Lu |
CVPR | 11 |
| 2025 | FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language ModelabstractCurrently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of vision language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-grained editing. To address these issues, we propose FireEdit, an innovative Fine-grained Instruction-based image editing framework that exploits a REgion-aware VLM. FireEdit is designed to accurately comprehend user instructions and ensure effective control over the editing process. Specifically, we enhance the fine-grained visual perception capabilities of the VLM by introducing additional region tokens. Relying solely on the output of the LLM to guide the diffusion model may lead to suboptimal editing results. Therefore, we propose a Time-Aware Target Injection module and a Hybrid Visual Cross Attention module. The former dynamically adjusts the guidance strength at various denoising stages by integrating timestep embeddings with the text embeddings. The latter enhances visual details for image editing, thereby preserving semantic consistency between the edited result and the source image. By combining the VLM enhanced with fine-grained region tokens and the time-dependent diffusion model, FireEdit demonstrates significant advantages in comprehending editing instructions and maintaining high semantic consistency. Extensive experiments indicate that our approach surpasses the state-of-the-art instruction-based image editing methods. Jiahao Li 0001, Zunnan Xu, Yiji Cheng, Fa-Ting Hong, Qin Lin 0003, Qinglin Lu, Xiaodan Liang |
CVPR | 7 |
| 2025 | Audio-Visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationabstractTalking head synthesis is vital for virtual avatars and human-computer interaction. However, most existing methods are typically limited to accepting control from a single primary modality, restricting their practical utility. To this end, we introduce \textbf{ACTalker}, an end-to-end video diffusion framework that supports both multi-signals control and single-signal control for talking head video generation. For multiple control, we design a parallel mamba structure with multiple branches, each utilizing a separate driving signal to control specific facial regions. A gate mechanism is applied across all branches, providing flexible control over video generation. To ensure natural coordination of the controlled video both temporally and spatially, we employ the mamba structure, which enables driving signals to manipulate feature tokens across both dimensions in each branch. Additionally, we introduce a mask-drop strategy that allows each driving signal to independently control its corresponding facial region within the mamba structure, preventing control conflicts. Experimental results demonstrate that our method produces natural-looking facial videos driven by diverse signals and that the mamba layer seamlessly integrates multiple driving modalities without conflict. The project website can be found at https://harlanhong.github.io/publications/actalker/index.html. Fa-Ting Hong, Zunnan Xu, Xiu Li 0001, Qin Lin 0003, Qinglin Lu, Dan Xu 0002 |
ICCV | 6 |
| 2025 | HOMA: Towards Generic Human-Object Interaction in Multimodal Driven Human Animation with Weak ConditionsabstractWhile recent advances in human-object interaction (HOI) video generation showcase promising capabilities for synthesizing coordinated human-object dynamics, existing methods remain constrained by their reliance on meticulously curated motion sequences and actor-specific data, thereby limiting practical scalability and user accessibility. Furthermore, generalization to novel object appearances and interaction scenarios remains understudied. To address these limitations, we propose HOMA, a weakly conditioned multimodal-driven HOI video generation framework that introduces sparse, decoupled motion guidance to enhance controllability and reduce dependency on stringent input conditions. Our approach encodes appearance and motion signals into the dual input space of a multimodal diffusion transformer (MMDiT), fusing them within a shared context space to enable temporally consistent and physically plausible interactions. To optimize learning efficiency and feature injection accuracy, we introduce a parameter-space HOI adapter initialized with pretrained MMDiT weights to preserve prior knowledge while enabling efficient adaptation. Additionally, we design a facial cross-attention adapter for audio-driven lip synchronization, ensuring anatomically accurate speech animation. Extensive experiments demonstrate that HOMA achieves state-of-the-art performance in interaction naturalness and generalization under weak supervision, outperforming existing methods by significant margins. We further illustrate HOMA’s versatility through diverse applications, including text-conditioned generation and interactive object manipulation, facilitated by a user-friendly demo interface. The project page is https://bone-11.github.io/homa-page/. Ziyao Huang 0002, Juan Cao 0001, Yifeng Ma 0006, Zejing Rao, Qin Lin 0003, Qinglin Lu, Fan Tang |
SIGGRAPH Asia | 9 |
| 2022 | Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
Yunlong Tang 0002, Siting Xu, Teng Wang 0007, Qin Lin 0003, Qinglin Lu, Feng Zheng 0001 |
ACCV (2) | 4 |