Wenyu Qin

dblp:118/4771 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
abstract
Current video generation models perform well at single-shot synthesis but struggle with multi-shot videos, facing critical challenges in maintaining character and background consistency across shots and flexibly generating videos of arbitrary length and shot count. To address these limitations, we introduce \textbf{FilmWeaver}, a novel framework designed to generate consistent, multi-shot videos of arbitrary length. First, it employs an autoregressive diffusion paradigm to achieve arbitrary-length video generation. To address the challenge of consistency, our key insight is to decouple the problem into inter-shot consistency and intra-shot coherence. We achieve this through a dual-level cache mechanism: a shot memory caches keyframes from preceding shots to maintain character and scene identity, while a temporal memory retains a history of frames from the current shot to ensure smooth, continuous motion. The proposed framework allows for flexible, multi-round user interaction to create multi-shot videos. Furthermore, due to this decoupled design, our method demonstrates high versatility by supporting downstream tasks such as multi-concept injection and video extension. To facilitate the training of our consistency-aware method, we also developed a comprehensive pipeline to construct a high-quality multi-shot video dataset. Extensive experimental results demonstrate that our method surpasses existing approaches on metrics for both consistency and aesthetic quality, opening up new possibilities for creating more consistent, controllable, and narrative-driven video content.
Xiaokun Liu, Wenyu Qin, Meng Wang 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Shao-Lun Huang
AAAI4
2025 MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
abstract
Multimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by generalizable attention architecture. Advanced methods predominantly focus on language-centric tuning while less exploring multimodal tokens mixed through attention, posing challenges in high-level tasks that require fine-grained cognition and emotion understanding. In this work, we identify the attention deficit disorder problem in multimodal learning, caused by inconsistent cross-modal attention and layer-by-layer decayed attention activation. To address this, we propose a novel attention mechanism, termed MOdular Duplex Attention (MODA), simultaneously conducting the inner-modal refinement and inter-modal interaction. MODA employs a correct-after-align strategy to effectively decouple modality alignment from cross-layer token mixing. In the alignment phase, tokens are mapped to duplex modality spaces based on the basis vectors, enabling the interaction between visual and language modality. Further, the correctness of attention scores is ensured through adaptive masked attention, which enhances the model's flexibility by allowing customizable masking patterns for different modalities. Extensive experiments on 21 benchmark datasets verify the effectiveness of MODA in perception, cognition, and emotion tasks.
Wuyou Xia, Chenxi Zhao 0002, Zhou Yan, Yongjie Zhu, Wenyu Qin, Pengfei Wan 0001, Di Zhang 0026, Jufeng Yang
ICML7
2025 VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
abstract
Understanding and predicting emotions from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis, the intrinsic nature of emotions—characterized by their open-set, dynamic, and context-dependent properties—poses challenge in understanding complex and evolving emotional states with reasonable rationale. To tackle these challenges, we propose a novel affective cues-guided reasoning framework that unifies fundamental attribute perception, expression analysis, and high-level emotional understanding in a stage-wise manner. At the core of our approach is a family of video emotion foundation models (VidEmo), specifically designed for emotion reasoning and instruction-following. These models undergo a two-stage tuning process: first, curriculum emotion learning for injecting emotion knowledge, followed by affective-tree reinforcement learning for emotion reasoning. Moreover, we establish a foundational data infrastructure and introduce a emotion-centric fine-grained dataset (Emo-CFG) consisting of 2.1M diverse instruction-based samples. Emo-CFG includes explainable emotional question-answering, fine-grained captions, and associated rationales, providing essential resources for advancing emotion understanding tasks. Experimental results demonstrate that our approach achieves competitive performance, setting a new milestone across 15 face perception tasks.
Yongjie Zhu, Wenyu Qin, Pengfei Wan 0001, Di Zhang 0026, Jufeng Yang
NeurIPS4
2025 Heterogeneous signcryption scheme from CLC to IBC for IIoT
Wenyu Qin, Zhiwei Chen 0004, Xiaobing Chen, Guanhua Chen 0007, Jian Weng 0001
Peer Peer Netw. Appl.2
2025 An efficient ring signcryption scheme for wireless sensor networks
Yusheng Cui, Wenyu Qin, Zhiwei Chen 0004, Guanhua Chen 0007, Jinsong Shan
Wirel. Networks3
2024 Heterogeneous signcryption scheme with equality test from CLC to PKI for IoV
Wenyu Qin, Zhiwei Chen 0004, Kaijun Sun, Guanhua Chen 0007, Jinsong Shan, Liqing Chen
Comput. Commun.2
2024 BDACD: Blockchain-based decentralized auditing supporting ciphertext deduplication
Yongliang Xu, Wenyu Qin, Jie Zhao 0015, Guanhua Chen 0007, Fugeng Zeng
J. Syst. Archit.3
2024 A blockchain-based auditable deduplication scheme for multi-cloud storage
Yongliang Xu, Wenyu Qin, Jie Zhao 0015, Ge Kan, Fugeng Zeng
Peer Peer Netw. Appl.3
2022 Heterogeneous online/offline signcryption for secure communication in Internet of Things
Huihui Zhu 0001, Wenyu Qin, Jinsong Shan
J. Syst. Archit.3
2022 Secure fuzzy identity-based public verification for cloud storage
Yongliang Xu, Wenyu Qin, Jinsong Shan
J. Syst. Archit.3