VLDB 2026 Research / reviewers in the wild / expert
Zonghong Dai
dblp:264/2890
· DBLP profile ↗
10ranked-venue papers
0as first author
9since 2021 · last 2026
0009-0006-7723-4130ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
Xu Chu 0001, Xinrong Chen, Haochen Li 0001, Zonghong Dai, Hongcheng Fan, Xiaoyue Yuan, Weiping Li 0002, Tong Mo |
DASFAA (6) | 5 |
| 2025 | ReMask-Animate: Refined Character Image Animation Using Mask-Guided AdaptersabstractPose-controlled human video generation is of significant interest and finds extensive applications in areas such as automated advertising and content creation on social media platforms. While existing methods employing pose sequences and reference images for human image animation have exhibited notable performance, they tend to encounter issues such as specific region blurring, background sharpening, and decreased identity consistency. In this paper, we introduce ReMask-Animate, which utilizes masks as additional priors to guide the model's local visual attention to specific areas, thereby alleviating feature confusion between different regions of the image. Three distinct mask-guided adapters are designed for cross-condition regional fusion of hand and face pose features, mitigating feature confusion between the foreground and background, and enhancing the visual consistency of character identity. Moreover, these lightweight adapters introduce minimal computational overhead and can be seamlessly integrated into specific layers of the backbone architecture. Extensive experiments show that our method outperforms state-of-the-art methods on five metrics in public datasets. Additionally, qualitative evaluations highlight a significant improvement in the quality of generated videos, demonstrating our approach's superiority. Xunzhi Xiang, Haiwei Xue, Zonghong Dai, Minglei Li 0001, Ye Yue, Fei Ma 0006, Weijiang Yu, Heng Chang, F. Richard Yu |
AAAI | 3 |
| 2025 | Identity-Preserving Audio-Driven Holistic Human Motion Video GenerationabstractGenerating realistic human motion videos is a pivotal challenge in advancing human-computer interaction. While existing approaches often focus on generating either head or gesture movements from audio, they lack unified control over full-body motion, frequently producing low-resolution and blurred outputs. Additionally, these methods struggle to maintain character identity throughout the generated content. In this paper, we introduce a novel framework that generates photorealistic, personalized human motion videos from audio by decoupling identity features. We integrate both visual features and voice timbre to enhance the preservation of character identity. Our approach follows a four-stage paradigm: (1) frame generation, (2) identity feature customization, (3) audio-motion modeling, and (4) motion-video rendering. Through the collaborative modeling of audio-motion and motion-video stages, our approach effectively maintains the consistency of character identity and background throughout the video, enhancing the realism and coherence of the generated video. Experimental results demonstrate that our framework delivers high-resolution videos with superior fidelity, establishing a new and effective baseline for holistic human motion video generation. Haiwei Xue, Zhensong Zhang, Minglei Li 0001, Zonghong Dai, Zhiyong Wu 0001 |
ICASSP | 4 |
| 2025 | VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweeningabstractWe propose VideoHumanMIB, a novel framework for Video Human Motion In-betweening that enables seamless transitions between different motion video clips, facilitating the generation of longer and more natural digital human videos. While existing video frame interpolation methods work well for similar motions in adjacent frames, they often struggle with complex human movements, resulting in artifacts and unrealistic transitions. To address these challenges, we introduce a two-stage approach: First, we design an Appearance Reconstruction AutoEncoder to decouple appearance and motion information, extracting robust appearance-invariant features. Second, we develop an enhanced diffusion pretrained network that leverages both motion optical flow and human pose as guidance conditions, enabling the model to learn comprehensive latent distributions of possible motions. Rather than operating directly in pixel space, our model works in a learned latent space, allowing it to better capture the underlying motion dynamics. The framework is optimized with a dual-frame constraint loss and a motion flow loss to ensure temporal consistency and natural movement transitions. Extensive experiments demonstrate that our approach generates highly realistic transition sequences that significantly outperform existing methods, particularly in challenging scenarios with large motion variations. The proposed VideoHumanMIB establishes a new baseline for human motion synthesis and enables more natural and controllable digital human animation. Haiwei Xue, Zhensong Zhang, Minglei Li 0001, Zonghong Dai, F. Richard Yu, Fei Ma 0006, Zhiyong Wu 0001 |
IJCAI | 4 |
| 2025 | Human Motion Video Generation: A SurveyabstractHuman motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys in this field focus on individual methods, lacking a comprehensive overview of the entire generative process. This paper addresses this gap by providing an in-depth survey of human motion video generation, encompassing over ten sub-tasks, and detailing the five key phases of the generation process: input, motion planning, motion video generation, refinement, and output. Notably, this is the first survey that discusses the potential of large language models in enhancing human motion video generation. Our survey reviews the latest developments and technological trends in human motion video generation across three primary modalities: vision, text, and audio. By covering over two hundred papers, we offer a thorough overview of the field and highlight milestone works that have driven significant technological breakthroughs. Our goal for this survey is to unveil the prospects of human motion video generation and serve as a valuable resource for advancing the comprehensive applications of digital humans. Haiwei Xue, Xiangyang Luo 0002, Zhanghao Hu, Xin Zhang 0169, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li 0001, Jian Yang 0003, Fei Ma 0006, Zhiyong Wu 0001, Changpeng Yang, Zonghong Dai, F. Richard Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 14 |
| 2025 | Super-NeRF: View-Consistent Detail Generation for NeRF Super-ResolutionabstractThe neural radiance field (NeRF) achieved remarkable success in modeling 3D scenes and synthesizing high-fidelity novel views. However, existing NeRF-based methods focus more on making full use of high-resolution images to generate high-resolution novel views, but less considering the generation of high-resolution details given only low-resolution images. In analogy to the extensive usage of image super-resolution, NeRF super-resolution is an effective way to generate low-resolution-guided high-resolution 3D scenes and holds great potential applications. Up to now, such an important topic is still under-explored. In this article, we propose a NeRF super-resolution method, named Super-NeRF, to generate high-resolution NeRF from only low-resolution inputs. Given multi-view low-resolution images, Super-NeRF constructs a multi-view consistency-controlling super-resolution module to generate various view-consistent high-resolution details for NeRF. Specifically, an optimizable latent code is introduced for each input view to control the generated reasonable high-resolution 2D images satisfying view consistency. The latent codes of each low-resolution image are optimized synergistically with the target Super-NeRF representation to utilize the view consistency constraint inherent in NeRF construction. We verify the effectiveness of Super-NeRF on synthetic, real-world, and even AI-generated NeRFs. Super-NeRF achieves state-of-the-art NeRF super-resolution performance on high-resolution detail generation and cross-view consistency. Yuqi Han, Tao Yu 0007, Xiaohang Yu, Di Xu 0012, Binge Zheng, Zonghong Dai, Changpeng Yang, Yuwang Wang, Qionghai Dai |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion ModelsabstractAudio-driven co-speech human gesture generation has made remarkable advancements recently. However, most previous works only focus on single person audio-driven gesture generation. We aim at solving the problem of conversational co-speech gesture generation that considers multiple participants in a conversation, which is a novel and challenging task due to the difficulty of simultaneously incorporating semantic information and other relevant features from both the primary speaker and the interlocutor. To this end, we propose CoDiffuseGesture, a diffusion model-based approach for speech-driven interaction gesture generation via modeling bilateral conversational intention, emotion, and semantic context. Our method synthesizes appropriate interactive, speech-matched, high-quality gestures for conversational motions through the intention perception module and emotion reasoning module at the sentence level by a pretrained language model. Experimental results demonstrate the promising performance of the proposed method. Haiwei Xue, Zhensong Zhang, Zhiyong Wu 0001, Minglei Li 0001, Zonghong Dai, Helen M. Meng |
ICASSP | 6 |
| 2023 | UnifiedGesture: A Unified Gesture Synthesis Model for Multiple SkeletonsabstractThe automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability across different motion capture standards. In addition, it is a challenging task due to the weak correlation between speech and gestures. To address these problems, we present UnifiedGesture, a novel diffusion model-based speech-driven gesture synthesis approach, trained on multiple gesture datasets with different skeletons. Specifically, we first present a retargeting network to learn latent homeomorphic graphs for different motion capture standards, unifying the representations of various gestures while extending the dataset. We then capture the correlation between speech and gestures based on a diffusion model architecture using cross-local attention and self-attention to generate better speech-matched and realistic gestures. To further align speech and gesture and increase diversity, we incorporate reinforcement learning on the discrete gesture units with a learned reward function. Extensive experiments show that UnifiedGesture outperforms recent approaches on speech-driven gesture generation in terms of CCA, FGD, and human-likeness. Zilin Wang 0002, Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Qiaochu Huang, Songcen Xu, Changpeng Yang, Zonghong Dai |
ACM Multimedia | 11 |
| 2023 | Leveraging the Latent Diffusion Models for Offline Facial Multiple Appropriate Reactions GenerationabstractOffline Multiple Appropriate Facial Reaction Generation (OMAFRG) aims to predict the reaction of different listeners given a speaker, which is useful in the senario of human-computer interaction and social media analysis. In recent years, the Offline Facial Reactions Generation (OFRG) task has been explored in different ways. However, most studies only focus on the deterministic reaction of the listeners. The research of the non-deterministic (i.e. OMAFRG) always lacks of sufficient attention and the results are far from satisfactory. Compared with the deterministic OFRG tasks, the OMAFRG task is closer to the true circumstance but corresponds to higher difficulty for its requirement of modeling stochasticity and context. In this paper, we propose a new model named FRDiff to tackle this issue. Our model is developed based on the diffusion model architecture with some modification to enhance its ability of aggregating the context features. And the inherent property of stochasticity in diffusion model enables our model to generate multiple reactions. We conduct experiments on the datasets provided by the ACM Multimedia REACT2023 and obtain the second place on the board, which demonstrates the effectiveness of our method. Jun Yu 0001, Ji Zhao 0020, Guochen Xie, Fengxin Chen, Minglei Li 0001, Zonghong Dai |
ACM Multimedia | 8 |
| 2020 | Trust Relationship Prediction in Alibaba E-Commerce PlatformabstractThis paper introduces how to infer trust relationships from billion-scale networked data to benefit Alibaba E-Commerce business. To effectively leverage the network correlations between labeled and unlabeled relationships to predict trust relationships, we formalize trust into multiple types and propose a graphical model to incorporate type-based dyadic and triadic correlations, namely eTrust. We also present a fast learning algorithm in order to handle billion-scale networks. Systematically, we evaluate the proposed methods on four different genres of datasets with labeled trust relationships: Alibaba, Epinions, Ciao, and Advogato. Experimental results show that the proposed methods achieve significantly better performance than several comparison methods (+1.7-32.3% by accuracy; p <; <; 0:01, with t-test). Most importantly, when handling the real large networked data with over 1,200,000,000 edges (Ali-large), our method achieves 2,000× speedup to infer trust relationships, comparing with the traditional graph learning algorithms. Finally, we have applied the inferred trust relationships to Alibaba E-commerce platform: Taobao, and achieved 2.75 percent improvement on gross merchandise volume (GMV). Yukuo Cen, Jing Zhang 0001, Gaofei Wang, Yujie Qian, Chuizheng Meng, Zonghong Dai, Hongxia Yang, Jie Tang 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |