VLDB 2026 Research / reviewers in the wild / expert
Hang Zhou 0009
dblp:26/3707-9
· DBLP profile ↗
44ranked-venue papers
5as first author
38since 2021 · last 2026
0000-0002-2616-923XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 5 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 5 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
Yasheng Sun, Hang Zhou 0009, Jiazhi Guan, Quanwei Yang, Kaisiyuan Wang, Borong Liang, Haocheng Feng, Jingdong Wang 0001, Ziwei Liu 0002, Hideki Koike |
Int. J. Comput. Vis. | 3 |
| 2025 | Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion TransformerabstractExisting methodologies for animating portrait images face significant challenges, particularly in handling non-frontal perspectives, rendering dynamic objects around the portrait, and generating immersive, realistic backgrounds. In this paper, we introduce the first application of a pretrained transformer-based video generative model that demonstrates strong generalization capabilities and generates highly dynamic, realistic videos for portrait animation, effectively addressing these challenges. The adoption of a new video backbone model makes previous U-Net-Based methods for identity maintenance, audio conditioning, and video extrapolation inapplicable. To address this limitation, we design an identity reference network consisting of a causal 3D VAE combined with a stacked series of transformer layers, ensuring consistent facial identity across video sequences. Additionally, we investigate various speech audio conditioning and motion frame mechanisms to enable the generation of continuous video driven by speech audio. Our method is validated through experiments on benchmark and newly proposed wild datasets, demonstrating substantial improvements over prior methods in generating realistic portraits characterized by diverse orientations within dynamic and immersive scenes. Jiahao Cui 0003, Yun Zhan, Hanlin Shang, Kaihui Cheng, Shan Mu, Hang Zhou 0009, Jingdong Wang 0001, Siyu Zhu 0001 |
CVPR | 8 |
| 2025 | Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion ModelabstractCurrent digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g., objects) have not been well investigated. Despite human hand synthesis already being an intricate problem, generating objects in contact with hands and their interactions presents an even more challenging task, especially when the objects exhibit obvious variations in size and shape. To tackle these issues, we present a novel video reenactment framework focusing on Human-Object Interaction (HOI) via an adaptive Layout-instructed Diffusion model (Re-HOLD). Our key insight is to employ specialized layout representation for hands and objects, respectively. Such representations enable effective disentanglement of hand modeling and object adaptation to diverse motion sequences. To further improve the quality of the HOI generation, we design an interactive textural enhancement module for both hands and objects by introducing two independent memory banks. We also propose a layout adjustment strategy for the cross-object reenactment scenario to adaptively adjust unreasonable layouts caused by diverse object sizes during inference. Comprehensive qualitative and quantitative evaluations demonstrate that our proposed framework significantly outperforms existing methods. Project page: https://fyycs.github.io/Re-HOLD. Quanwei Yang, Kaisiyuan Wang, Hang Zhou 0009, Haocheng Feng, Errui Ding, Yu Wu 0011, Jingdong Wang 0001 |
CVPR | 4 |
| 2025 | AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion TransformersabstractDespite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic human videos with both accurate lip-sync and delicate co-speech gestures w.r.t. given audio. In this work, we propose AudCast, a generalized audio-driven human video generation framework adopting a cascade Diffusion-Transformers (DiTs) paradigm, which synthesizes holistic human videos based on a reference image and a given audio. 1) Firstly, an audio-conditioned Holistic Human DiT architecture is proposed to directly drive the movements of any human body with vivid gesture dynamics. 2) Then to enhance hand and face details that are well-knownly difficult to handle, a Regional Refinement DiT leverages regional 3D fitting as the bridge to reform the signals, producing the final results. Extensive experiments demonstrate that our framework generates high-fidelity audio-driven holistic human videos with temporal coherence and fine facial and hand details. Resources can be found at https://guanjz20.github.io/projects/AudCast. Jiazhi Guan, Kaisiyuan Wang, Quanwei Yang, Yasheng Sun, Shengyi He, Borong Liang, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Youjian Zhao, Hang Zhou 0009, Ziwei Liu 0002 |
CVPR | 14 |
| 2025 | RQTalker: Speech-driven 3D Facial Animation via Region-aware Vector QuantizationabstractSpeech-driven 3D facial animation has been a long-standing topic due to the complex geometry and motion modeling as well as difficulties in cross-modality learning. Current studies struggle to synthesize human-like lip motions, as they usually represent the movement of the entire face with a compressed global vector, leading to subtle motion loss and thus over-smoothed movements in the local lip region. To cope with this problem, we propose a new speech-driven 3D facial animation framework RQTalker based on the Region-aware Vector Quantization mechanism. The key insight is to first build a region-aware codebook via a self-reconstruction manner, in which each part of the codebook physically corresponds to a facial region with a clear semantic. Our region-aware codebook divides facial movements into local regions for multiple sub-encodings, reducing information loss from compression and improving local facial motion modeling. In addition, we further propose a spatial-temporal Audio-to-Motion Learning Module to produce movements that are spatially more accurate and temporally consistent. Qualitative and quantitative results demonstrate that our method outperforms state-of-the-art approaches. Kaisiyuan Wang, Hang Zhou 0009, Shengyi He, Yu Wu 0001 |
ICASSP | 3 |
| 2025 | GestureHYDRA: Semantic Co-Speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
Quanwei Yang, Luying Huang, Kaisiyuan Wang, Jiazhi Guan, Shengyi He, Fengguo Li, Hang Zhou 0009, Lingyun Yu 0002, Haocheng Feng, Hongtao Xie 0001 |
ICCV | 7 |
| 2025 | Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image AnimationabstractRecent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities.First, we extend the method to produce long-duration videos. To address substantial challenges such as appearance drift and temporal artifacts, we investigate augmentation strategies within the image space of conditional motion frames. Specifically, we introduce a patch-drop technique augmented with Gaussian noise to enhance visual consistency and temporal coherence over long duration.Second, we achieve 4K resolution portrait video generation. To accomplish this, we implement vector quantization of latent codes and apply temporal alignment techniques to maintain coherence across the temporal dimension. By integrating a high-quality decoder, we realize visual synthesis at 4K resolution.Third, we incorporate adjustable semantic textual labels for portrait expressions as conditional inputs. This extends beyond traditional audio cues to improve controllability and increase the diversity of the generated content. To the best of our knowledge, Hallo2, proposed in this paper, is the first method to achieve 4K resolution and generate hour-long, audio-driven portrait image animations enhanced with textual prompts. We have conducted extensive experiments to evaluate our method on publicly available datasets, including HDTF, CelebV, and our introduced ''Wild'' dataset. The experimental results demonstrate that our approach achieves state-of-the-art performance in long-duration portrait video animation, successfully generating rich and controllable content at 4K resolution for duration extending up to tens of minutes. Jiahao Cui 0003, Yao Yao 0008, Hao Zhu 0004, Hanlin Shang, Kaihui Cheng, Hang Zhou 0009, Siyu Zhu 0001, Jingdong Wang 0001 |
ICLR | 7 |
| 2025 | Real-Time Neural Radiance Talking Portrait Synthesis via Audio-Spatial Decomposition
Jiaxiang Tang, Kaisiyuan Wang, Hang Zhou 0009, Xiaokang Chen, Dongliang He, Tianshu Hu, Jingtuo Liu, Ziwei Liu 0002, Jingdong Wang 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | ReSyncer: Rewiring Style-Based Generator for Unified Audio-Visually Synced Facial Performer
Jiazhi Guan, Hang Zhou 0009, Kaisiyuan Wang, Shengyi He, Zhanwang Zhang, Borong Liang, Haocheng Feng, Errui Ding, Jingtuo Liu, Jingdong Wang 0001, Youjian Zhao, Ziwei Liu 0002 |
ECCV (41) | 3 |
| 2024 | 3D-Aware Text-Driven Talking Avatar Generation
Xiuzhe Wu, Yang-Tian Sun, Handi Chen, Hang Zhou 0009, Jingdong Wang 0001, Zhengzhe Liu, Xiaojuan Qi 0001 |
ECCV (88) | 4 |
| 2024 | DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content CreationabstractRecent advances in 3D content creation mostly leverage optimization-based 3D generation via score distillation sampling (SDS).
Though promising results have been exhibited, these methods often suffer from slow per-sample optimization, limiting their practical usage.
In this paper, we propose DreamGaussian, a novel 3D content generation framework that achieves both efficiency and quality simultaneously.
Our key insight is to design a generative 3D Gaussian Splatting model with companioned mesh extraction and texture refinement in UV space.
In contrast to the occupancy pruning used in Neural Radiance Fields, we demonstrate that the progressive densification of 3D Gaussians converges significantly faster for 3D generative tasks.
To further enhance the texture quality and facilitate downstream applications, we introduce an efficient algorithm to convert 3D Gaussians into textured meshes and apply a fine-tuning stage to refine the details.
Extensive experiments demonstrate the superior efficiency and competitive generation quality of our proposed approach.
Notably, DreamGaussian produces high-quality textured meshes in just 2 minutes from a single-view image, achieving approximately 10 times acceleration compared to existing methods. Jiaxiang Tang, Jiawei Ren 0001, Hang Zhou 0009, Ziwei Liu 0002 |
ICLR | 3 |
| 2024 | ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion ModelingabstractAlthough significant progress has been made in human video generation, most previous studies focus on either human facial animation or full-body animation, which cannot be directly applied to produce realistic conversational human videos with frequent hand gestures and various facial movements simultaneously.
To address these limitations, we propose a 2D human video generation framework, named ShowMaker, capable of generating high-fidelity half-body conversational videos via fine-grained diffusion modeling.
We leverage dual-stream diffusion models as the backbone of our framework and carefully design two novel components for crucial local regions (i.e., hands and face) that can be easily integrated into our backbone.
Specifically, to handle the challenging hand generation caused by sparse motion guidance, we propose a novel Key Point-based Fine-grained Hand Modeling module by amplifying positional information from raw hand key points and constructing a corresponding key point-based codebook.
Moreover, to restore richer facial details in generated results, we introduce a Face Recapture module, which extracts facial texture features and global identity features from the aligned human face and integrates them into the diffusion process for face enhancement.
Extensive quantitative and qualitative experiments demonstrate the superior visual quality and temporal consistency of our method. Quanwei Yang, Jiazhi Guan, Kaisiyuan Wang, Lingyun Yu 0002, Wenqing Chu, Hang Zhou 0009, ZhiQiang Feng, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Hongtao Xie 0001 |
NeurIPS | 6 |
| 2024 | TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
Jiazhi Guan, Quanwei Yang, Kaisiyuan Wang, Hang Zhou 0009, Shengyi He, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Hongtao Xie 0001, Youjian Zhao, Ziwei Liu 0002 |
SIGGRAPH Asia | 4 |
| 2024 | ReliTalk: Relightable Talking Portrait Generation from a Single Video
Haonan Qiu, Zhaoxi Chen 0009, Yuming Jiang 0003, Hang Zhou 0009, Xiangyu Fan 0002, Lei Yang 0045, Wayne Wu, Ziwei Liu 0002 |
Int. J. Comput. Vis. | 4 |
| 2023 | Robust Video Portrait Reenactment via Personalized Representation QuantizationabstractWhile progress has been made in the field of portrait reenactment, the problem of how to produce high-fidelity and robust videos remains. Recent studies normally find it challenging to handle rarely seen target poses due to the limitation of source data. This paper proposes the Video Portrait via Non-local Quantization Modeling (VPNQ) framework, which produces pose- and disturbance-robust reenactable video portraits. Our key insight is to learn position-invariant quantized local patch representations and build a mapping between simple driving signals and local textures with non-local spatial-temporal modeling. Specifically, instead of learning a universal quantized codebook, we identify that a personalized one can be trained to preserve desired position-invariant local details better. Then, a simple representation of projected landmarks can be used as sufficient driving signals to avoid 3D rendering. Following, we employ a carefully designed Spatio-Temporal Transformer to predict reasonable and temporally consistent quantized tokens from the driving signal. The predicted codes can be decoded back to robust and high-quality videos. Comprehensive experiments have been conducted to validate the effectiveness of our approach. Kaisiyuan Wang, Changcheng Liang, Hang Zhou 0009, Jiaxiang Tang, Qianyi Wu, Dongliang He, Zhibin Hong, Jingtuo Liu, Errui Ding, Ziwei Liu 0002, Jingdong Wang 0001 |
AAAI | 3 |
| 2023 | StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-Based GeneratorabstractDespite recent advances in syncing lip movements with any audio waves, current methods still struggle to balance generation quality and the model's generalization ability. Previous studies either require long-term data for training or produce a similar movement pattern on all subjects with low quality. In this paper, we propose StyleSync, an effective framework that enables high-fidelity lip synchronization. We identify that a style-based generator would sufficiently enable such a charming property on both one-shot and few-shot scenarios. Specifically, we design a mask-guided spatial information encoding module that preserves the details of the given face. The mouth shapes are accurately modified by audio through modulated convolutions. Moreover, our design also enables personalized lip-sync by introducing style space and generator refinement on only limited frames. Thus the identity and talking style of a target person could be accurately preserved. Extensive experiments demonstrate the effectiveness of our method in producing high-fidelity results on a variety of scenes. Resources can be found at https:/hangz-nju-cuhk.github.io/projects/StyleSync. Jiazhi Guan, Zhanwang Zhang, Hang Zhou 0009, Tianshu Hu, Kaisiyuan Wang, Dongliang He, Haocheng Feng, Jingtuo Liu, Errui Ding, Ziwei Liu 0002, Jingdong Wang 0001 |
CVPR | 3 |
| 2023 | Delicate Textured Mesh Recovery from NeRF via Adaptive Surface RefinementabstractNeural Radiance Fields (NeRF) have constituted a remarkable breakthrough in image-based 3D reconstruction. However, their implicit volumetric representations differ significantly from the widely-adopted polygonal meshes and lack support from common 3D software and hardware, making their rendering and manipulation inefficient. To overcome this limitation, we present a novel framework that generates textured surface meshes from images. Our approach begins by efficiently initializing the geometry and view-dependency decomposed appearance with a NeRF. Subsequently, a coarse mesh is extracted, and an iterative surface refinement algorithm is developed to adaptively adjust both vertex positions and face density based on reprojected rendering errors. We jointly refine the appearance with geometry and bake it into texture images for real-time rendering. Extensive experiments demonstrate that our method achieves superior mesh quality and competitive rendering quality. Jiaxiang Tang, Hang Zhou 0009, Xiaokang Chen, Tianshu Hu, Errui Ding, Jingdong Wang 0001 |
ICCV | 2 |
| 2023 | GoBigger: A Scalable Platform for Cooperative-Competitive Multi-Agent Interactive Simulation
Ming Zhang 0037, Shenghan Zhang, Zhenjie Yang 0001, Lekai Chen, Jinliang Zheng, Chao Yang 0026, Chuming Li, Hang Zhou 0009, Yazhe Niu, Yu Liu 0015 |
ICLR | 8 |
| 2023 | Dual-Modality Co-Learning for Unveiling Deepfake in Spatio-Temporal SpaceabstractThe emergence of photo-realistic deepfakes on a large scale has become a significant societal concern, which has garnered considerable attention from the research community. Several recent studies have identified the critical issue of “temporal inconsistency” resulting from the frame reassembling process of deepfake generation techniques. However, due to the lack of task-specific design, the spatio-temporal modeling of current methods remains insufficient in three critical aspects: 1) inapparent temporal changes are prone to be undermined compared to abundant spatial cues; 2) minor inconsistent regions are often concealed by motions with greater amplitude during downsampling; 3) capturing both transient inconsistencies and persistent motions simultaneously remains a significant challenge. In this paper, we propose a novel Dual-Modality Co-Learning framework tailored for these characteristics, which achieves more effectual deepfake detection with complementary information from RGB and optical flow modalities. In particular, we designed a Multi-Scale Motion Regularization module to encourage the network to equally prioritize both the significant spatial cues and the subtle temporal facial motion cues. Additionally, we developed a Multi-Span Cross-Attention module to effectively integrate the information from both RGB and optical flow modalities and improve the detection accuracy with multi-span predictions. Extensive experiments validate the effectiveness our ideas and demonstrate the superior performance of our approach. Jiazhi Guan, Hang Zhou 0009, Zhizhi Guo, Tianshu Hu, Lirui Deng 0001, Chengbin Quan, Youjian Zhao |
ICMR | 2 |
| 2023 | ReEnFP: Detail-Preserving Face Reconstruction by Encoding Facial PriorsabstractWe address the problem of face modeling, which is still challenging in achieving high-quality reconstruction results efficiently. Neither previous regression-based nor optimization-based frameworks could well balance between the facial reconstruction fidelity and efficiency. We notice that the large amount of in-the-wild facial images contain diverse appearance information, however, their underlying knowledge is not fully exploited for face modeling. To this end, we propose our Reconstruction by Encoding Facial Priors (ReEnFP) pipeline to exploit the potential of unconstrained facial images for further improvement. Our key is to encode generative priors learned by a style-based texture generator on unconstrained data for fast and detail-preserving face reconstruction. With our texture generator pre-trained using a differentiable renderer, faces could be encoded to its latent space as opposed to the time-consuming optimization-based inversion. Our generative prior encoding is further enhanced with a pyramid fusion block for adaptive integration of input spatial information. Extensive experiments show that our method reconstructs photo-realistic facial textures and geometric details with precise identity recovery. Yasheng Sun, Jiangke Lin, Hang Zhou 0009, Dongliang He, Hideki Koike |
WACV | 3 |
| 2023 | Exploiting Visual Context Semantics for Sound Source LocalizationabstractSelf-supervised sound source localization in unconstrained visual scenes is an important task of audio-visual learning. In this paper, we propose a visual reasoning module to explicitly exploit the rich visual context semantics, which alleviates the issue of insufficient utilization of visual information in previous works. The learning objectives are carefully designed to provide stronger supervision signals for the extracted visual semantics while enhancing the audio-visual interactions, which lead to more robust feature representations. Extensive experimental results demonstrate that our approach significantly boosts the localization performances on various datasets, even without initializations pretrained on ImageNet. Moreover, with the visual context exploitation, our framework can accomplish both the audio-visual and purely visual inference, which expands the application scope of the sound source localization task and further raises the competitiveness of our approach. Xinchi Zhou, Dongzhan Zhou, Di Hu 0001, Hang Zhou 0009, Wanli Ouyang |
WACV | 4 |
| 2023 | SeCo: Separating Unknown Musical Visual Sounds with Consistency GuidanceabstractRecent years have witnessed the success of deep learning on the visual sound separation task. However, existing works follow similar settings where the training and testing datasets share the same musical instrument categories, which to some extent limits the versatility of this task. In this work, we focus on a more general and challenging scenario, namely the separation of unknown musical instruments, where the categories in training and testing phases have no direct overlap with each other. To tackle this new setting, we propose the "Separation-with-Consistency" (SeCo) framework, which can accomplish the separation on unknown categories by exploiting the consistency constraints. Furthermore, to capture richer characteristics of the novel melodies, we devise an online matching strategy, which can bring stable enhancements with no cost of extra parameters. Experiments demonstrate that our SeCo framework exhibits strong adaptation ability on the novel musical categories and outperforms the baseline methods by a notable margin. Xinchi Zhou, Dongzhan Zhou, Wanli Ouyang, Hang Zhou 0009, Di Hu 0001 |
WACV | 4 |
| 2022 | Visual Sound Localization in the Wild by Cross-Modal Interference ErasingabstractThe task of audiovisual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real world scenarios, audios are usually contaminated by off screen sound and background noise. They will interfere with the procedure of identifying desired sources and building visual sound connections, making previous studies nonapplicable. In this work, we propose the Interference Eraser (IEr) framework, which tackles the problem of audiovisual sound source localization in the wild. The key idea is to eliminate the interference by redefining and carving discriminative audio representations. Specifically, we observe that the previous practice of learning only a single audio representation is insufficient due to the additive nature of audio signals. We thus extend the audio representation with our Audio Instance Identifier module, which clearly distinguishes sounding instances when audio signals of different volumes are unevenly mixed. Then we erase the influence of the audible but off screen sounds and the silent but visible objects by a Cross modal Referrer module with cross modality distillation. Quantitative and qualitative evaluations demonstrate that our framework achieves superior results on sound localization tasks, especially under real world scenarios. Rui Qian 0001, Hang Zhou 0009, Di Hu 0001, Weiyao Lin, Ziwei Liu 0002, Bolei Zhou, Xiaowei Zhou 0001 |
AAAI | 3 |
| 2022 | SepFusion: Finding Optimal Fusion Structures for Visual Sound SeparationabstractMultiple modalities can provide rich semantic information; and exploiting such information will normally lead to better performance compared with the single-modality counterpart. However, it is not easy to devise an effective cross-modal fusion structure due to the variations of feature dimensions and semantics, especially when the inputs even come from different sensors, as in the field of audio-visual learning. In this work, we propose SepFusion, a novel framework that can smoothly produce optimal fusion structures for visual-sound separation. The framework is composed of two components, namely the model generator and the evaluator. To construct the generator, we devise a lightweight architecture space that can adapt to different input modalities. In this way, we can easily obtain audio-visual fusion structures according to our demands. For the evaluator, we adopt the idea of neural architecture search to select superior networks effectively. This automatic process can significantly save human efforts while achieving competitive performances. Moreover, since our SepFusion provides a series of strong models, we can utilize the model family for broader applications, such as further promoting performance via model assembly, or providing suitable architectures for the separation of certain instrument classes. These potential applications further enhance the competitiveness of our approach. Dongzhan Zhou, Xinchi Zhou, Di Hu 0001, Hang Zhou 0009, Lei Bai 0001, Ziwei Liu 0002, Wanli Ouyang |
AAAI | 4 |
| 2022 | Expressive Talking Head Generation with Granular Audio-Visual ControlabstractGenerating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces realistic. In this paper, we propose the Granularly Controlled Audio-Visual Talking Heads (GC-AVT), which controls lip movements, head poses, and facial expressions of a talking head in a granular manner. Our insight is to decouple the audio-visual driving sources through prior-based pre-processing designs. Detailedly, we disassemble the driving image into three complementary parts including: 1) a cropped mouth that facilitates lip-sync; 2) a masked head that implicitly learns pose; and 3) the upper face which works corporately and complementarily with a time-shifted mouth to contribute the expression. Interestingly, the encoded features from the three sources are integrally balanced through reconstruction training. Extensive experiments show that our method generates expressive faces with not only synced mouth shapes, controllable poses, but precisely animated emotional expressions as well. Borong Liang, Yan Pan 0019, Zhizhi Guo, Hang Zhou 0009, Zhibin Hong, Xiaoguang Han 0001, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
CVPR | 4 |
| 2022 | Learning Hierarchical Cross-Modal Association for Co-Speech Gesture GenerationabstractGenerating speech-consistent body and gesture movements is a long-standing problem in virtual avatar creation. Previous studies often synthesize pose movement in a holistic manner, where poses of all joints are generated simultaneously. Such a straightforward pipeline fails to generate fine-grained co-speech gestures. One observation is that the hierarchical semantics in speech and the hierarchical structures of human gestures can be naturally described into multiple granularities and associated together. To fully utilize the rich connections between speech audio and human gestures, we propose a novel framework named Hierarchical Audio-to-Gesture (HA2G) for co-speech gesture generation. In HA2G, a Hierarchical Audio Learner extracts audio representations across semantic granularities. A Hierarchical Pose Inferer subsequently renders the entire human pose gradually in a hierarchical manner. To enhance the quality of synthesized gestures, we develop a contrastive learning strategy based on audio-text alignment for better audio representations. Extensive experiments and human evaluation demonstrate that the proposed method renders realistic co-speech gestures and out-performs previous methods in a clear margin. Project page: https://alvinliu0.github.io/projects/HA2G. Qianyi Wu, Hang Zhou 0009, Yinghao Xu 0001, Rui Qian 0001, Xiaowei Zhou 0001, Wayne Wu, Bo Dai 0002, Bolei Zhou |
CVPR | 3 |
| 2022 | Few-Shot Head Swapping in the WildabstractThe head swapping task aims at flawlessly placing a source head onto a target body, which is of great importance to various entertainment scenarios. While face swapping has drawn much attention, the task of head swapping has rarely been explored, particularly under the few-shot setting. It is inherently challenging due to its unique needs in head modeling and background blending. In this paper, we present the Head Swapper (HeSer), which achieves few-shot head swapping in the wild through two delicately de-signed modules. Firstly, a Head2Head Aligner is devised to holistically migrate pose and expression information from the target to the source head by examining multi-scale in-formation. Secondly, to tackle the challenges of skin color variations and head-background mismatches in the swapping procedure, a Head2Scene Blender is introduced to si-multaneously modify facial skin color and fill mismatched gaps on the background around the head. Particularly, seamless blending is achieved with the help of a Semantic-Guided Color Reference Creation procedure and a Blending UNet. Extensive experiments demonstrate that the proposed method produces superior head swapping results on a variety of scenes. Changyong Shu, Hemao Wu, Hang Zhou 0009, Jiaming Liu 0003, Zhibin Hong, Changxing Ding, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
CVPR | 3 |
| 2022 | TokenMix: Rethinking Image Mixing for Data Augmentation in Vision Transformers
Jihao Liu, Boxiao Liu, Hang Zhou 0009, Hongsheng Li 0001, Yu Liu 0015 |
ECCV (26) | 3 |
| 2022 | Semantic-Aware Implicit Neural Audio-Driven Video Portrait Generation
Yinghao Xu 0001, Qianyi Wu, Hang Zhou 0009, Wayne Wu, Bolei Zhou |
ECCV (37) | 4 |
| 2022 | StyleSwap: Style-Based Generator Empowers Robust Face Swapping
Hang Zhou 0009, Zhibin Hong, Ziwei Liu 0002, Jiaming Liu 0003, Zhizhi Guo, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
ECCV (14) | 2 |
| 2022 | Delving into Sequential Patches for Deepfake DetectionabstractRecent advances in face forgery techniques produce nearly visually untraceable deepfake videos, which could be leveraged with malicious intentions. As a result, researchers have been devoted to deepfake detection. Previous studies have identified the importance of local low-level cues and temporal information in pursuit to generalize well across deepfake methods, however, they still suffer from robustness problem against post-processings. In this work, we propose the Local- & Temporal-aware Transformer-based Deepfake Detection (LTTD) framework, which adopts a local-to-global learning protocol with a particular focus on the valuable temporal information within local sequences. Specifically, we propose a Local Sequence Transformer (LST), which models the temporal consistency on sequences of restricted spatial regions, where low-level information is hierarchically enhanced with shallow layers of learned 3D filters. Based on the local temporal embeddings, we then achieve the final classification in a global contrastive way. Extensive experiments on popular datasets validate that our approach effectively spots local forgery cues and achieves state-of-the-art performance. Jiazhi Guan, Hang Zhou 0009, Zhibin Hong, Errui Ding, Jingdong Wang 0001, Chengbin Quan, Youjian Zhao |
NeurIPS | 2 |
| 2022 | Audio-Driven Co-Speech Gesture Video GenerationabstractCo-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the image domain remains unsolved. In this work, we formally define and study this challenging problem of audio-driven co-speech gesture video generation, i.e., using a unified framework to generate speaker image sequence driven by speech audio. Our key insight is that the co-speech gestures can be decomposed into common motion patterns and subtle rhythmic dynamics. To this end, we propose a novel framework, Audio-driveN Gesture vIdeo gEneration (ANGIE), to effectively capture the reusable co-speech gesture patterns as well as fine-grained rhythmic movements. To achieve high-fidelity image sequence generation, we leverage an unsupervised motion representation instead of a structural human body prior (e.g., 2D skeletons). Specifically, 1) we propose a vector quantized motion extractor (VQ-Motion Extractor) to summarize common co-speech gesture patterns from implicit motion representation to codebooks. 2) Moreover, a co-speech gesture GPT with motion refinement (Co-Speech GPT) is devised to complement the subtle prosodic motion details. Extensive experiments demonstrate that our framework renders realistic and vivid co-speech gesture video. Demo video and more resources can be found in: https://alvinliu0.github.io/projects/ANGIE Qianyi Wu, Hang Zhou 0009, Yuanqi Du, Wayne Wu, Dahua Lin, Ziwei Liu 0002 |
NeurIPS | 3 |
| 2022 | Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in TransformersabstractPrevious studies have explored generating accurately lip-synced talking faces for arbitrary targets given audio conditions. However, most of them deform or generate the whole facial area, leading to non-realistic results. In this work, we delve into the formulation of altering only the mouth shapes of the target person. This requires masking a large percentage of the original image and seamlessly inpainting it with the aid of audio and reference frames. To this end, we propose the Audio-Visual Context-Aware Transformer (AV-CAT) framework, which produces accurate lip-sync with photo-realistic quality by predicting the masked mouth shapes. Our key insight is to exploit desired contextual information provided in audio and visual modalities thoroughly with delicately designed Transformers. Specifically, we propose a convolution-Transformer hybrid backbone and design an attention-based fusion strategy for filling the masked parts. It uniformly attends to the textural information on the unmasked regions and the reference frame. Then the semantic audio information is involved in enhancing the self-attention computation. Additionally, a refinement network with audio injection improves both image and lip-sync quality. Extensive experiments validate that our model can generate high-fidelity lip-synced results for arbitrary subjects. Yasheng Sun, Hang Zhou 0009, Kaisiyuan Wang, Qianyi Wu, Zhibin Hong, Jingtuo Liu, Errui Ding, Jingdong Wang 0001, Ziwei Liu 0002, Hideki Koike |
SIGGRAPH Asia | 2 |
| 2021 | PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective FeedbackabstractSecond language (L2) English learners often find it difficult to improve their pronunciations due to the lack of expressive and personalized corrective feedback. In this paper, we present Pronunciation Teacher (PTeacher), a Computer-Aided Pronunciation Training (CAPT) system that provides personalized exaggerated audio-visual corrective feedback for mispronunciations. Though the effectiveness of exaggerated feedback has been demonstrated, it is still unclear how to define the appropriate degrees of exaggeration when interacting with individual learners. To fill in this gap, we interview 100 L2 English learners and 22 professional native teachers to understand their needs and experiences. Three critical metrics are proposed for both learners and teachers to identify the best exaggeration levels in both audio and visual modalities. Additionally, we incorporate the personalized dynamic feedback mechanism given the English proficiency of learners. Based on the obtained insights, a comprehensive interactive pronunciation training course is designed to help L2 learners rectify mispronunciations in a more perceptible, understandable, and discriminative manner. Extensive user studies demonstrate that our system significantly promotes the learners’ learning efficiency. Yaohua Bu, Hang Zhou 0009, Jia Jia 0001, Shengqi Chen 0001, Dachuan Shi, Haozhe Wu, Kun Li 0003, Zhiyong Wu 0001, Yuanchun Shi, Xiaobo Lu, Ziwei Liu 0002 |
CHI | 4 |
| 2021 | Audio-Driven Emotional Video PortraitsabstractDespite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important features on natural human faces, is always neglected in their methods. In this work, we present Emotional Video Portraits (EVP), a system for synthesizing high-quality video portraits with vivid emotional dynamics driven by audios. Specifically, we propose the Cross-Reconstructed Emotion Disentanglement technique to decompose speech into two decoupled spaces, i.e., a duration-independent emotion space and a duration- dependent content space. With the disentangled features, dynamic 2D emotional facial landmarks can be deduced. Then we propose the Target-Adaptive Face Synthesis technique to generate the final high-quality video portraits, by bridging the gap between the deduced landmarks and the natural head poses of target videos. Extensive experiments demonstrate the effectiveness of our method both qualitatively and quantitatively.1 Xinya Ji, Hang Zhou 0009, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, Feng Xu 0005 |
CVPR | 2 |
| 2021 | Visually Informed Binaural Audio Generation without Binaural AudiosabstractStereophonic audio, especially binaural audio, plays an essential role in immersive viewing environments. Recent research has explored generating visually guided stereophonic audios supervised by multi-channel audio collections. However, due to the requirement of professional recording devices, existing datasets are limited in scale and variety, which impedes the generalization of supervised methods in real-world scenarios. In this work, we propose PseudoBinaural, an effective pipeline that is free of binaural recordings. The key insight is to carefully build pseudo visual-stereo pairs with mono data for training. Specifically, we leverage spherical harmonic decomposition and head-related impulse response (HRIR) to identify the relationship between spatial locations and received binaural audios. Then in the visual modality, corresponding visual cues of the mono data are manually placed at sound source positions to form the pairs. Compared to fully-supervised paradigms, our binaural-recording-free pipeline shows great stability in cross-dataset evaluation and achieves comparable performance under subjective preference. Moreover, combined with binaural recordings, our method is able to further boost the performance of binaural audio generation under supervised settings1. Xudong Xu, Hang Zhou 0009, Ziwei Liu 0002, Bo Dai 0002, Xiaogang Wang 0001, Dahua Lin |
CVPR | 2 |
| 2021 | Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationabstractWhile accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information such as landmarks and 3D parameters, aiming to generate personalized rhythmic movements. However, the inaccuracy of such estimated information under extreme conditions would lead to degradation problems. In this paper, we propose a clean yet effective framework to generate pose-controllable talking faces. We operate on non-aligned raw face images, using only a single photo as an identity reference. The key is to modularize audio-visual representations by devising an implicit low-dimension pose code. Substantially, both speech content and head pose information lie in a joint non-identity embedding space. While speech content information can be defined by learning the intrinsic synchronization between audio-visual modalities, we identify that a pose code will be complementarily learned in a modulated convolution-based reconstruction framework.Extensive experiments show that our method generates accurately lip-synced talking faces whose poses are controllable by other videos. Moreover, our model has multiple advanced capabilities including extreme view robustness and talking face frontalization.1 Hang Zhou 0009, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang 0001, Ziwei Liu 0002 |
CVPR | 1 |
| 2021 | Speech2Talking-Face: Inferring and Driving a Face with Synchronized Audio-Visual RepresentationabstractWhat can we picture solely from a clip of speech? Previous research has shown the possibility of directly inferring the appearance of a person's face by listening to a voice. However, within human speech lies not only the biometric identity signal but also the identity-irrelevant information such as the talking content. Our goal is to extract as much information from a clip of speech as possible. In particular, we aim at not only inferring the face of a person but also animating it. Our key insight is to synchronize audio and visual representations from two perspectives in a style-based generative framework. Specifically, contrastive learning is leveraged to map both the identity and speech content information within the speech to visual representation spaces. Furthermore, the identity space is strengthened with class centroids. Through curriculum learning, the style-based generator is capable of automatically balancing the information from the two latent spaces. Extensive experiments show that our approach encourages better speech-identity correlation learning while generating vivid faces whose identities are consistent with given speech samples. Moreover, by leveraging the same model, these inferred faces can be driven to talk by the audio. Yasheng Sun, Hang Zhou 0009, Ziwei Liu 0002, Hideki Koike |
IJCAI | 2 |
| 2020 | Rotate-and-Render: Unsupervised Photorealistic Face Rotation From Single-View ImagesabstractThough face rotation has achieved rapid progress in recent years, the lack of high-quality paired training data remains a great hurdle for existing methods. The current generative models heavily rely on datasets with multi-view images of the same person. Thus, their generated results are restricted by the scale and domain of the data source. To overcome these challenges, we propose a novel unsupervised framework that can synthesize photo-realistic rotated faces using only single-view image collections in the wild. Our key insight is that rotating faces in the 3D space back and forth, and re-rendering them to the 2D plane can serve as a strong self-supervision. We leverage the recent advances in 3D face modeling and high-resolution GAN to constitute our building blocks. Since the 3D rotation-and-render on faces can be applied to arbitrary angles without losing details, our approach is extremely suitable for in-the-wild scenarios (i.e. no paired data are available), where existing methods fall short. Extensive experiments demonstrate that our approach has superior synthesis quality as well as identity preservation over the state-of-the-art methods, across a wide range of poses and domains. Furthermore, we validate that our rotate-and-render framework naturally can act as an effective data augmentation engine for boosting modern face recognition systems even on strong baseline models. Hang Zhou 0009, Jihao Liu, Ziwei Liu 0002, Yu Liu 0015, Xiaogang Wang 0001 |
CVPR | 1 |
| 2020 | Discriminability Distillation in Group Representation Learning
Manyuan Zhang, Guanglu Song, Hang Zhou 0009, Yu Liu 0015 |
ECCV (10) | 3 |
| 2020 | Sep-Stereo: Visually Guided Stereophonic Audio Generation by Associating Source Separation
Hang Zhou 0009, Xudong Xu, Dahua Lin, Xiaogang Wang 0001, Ziwei Liu 0002 |
ECCV (12) | 1 |
| 2019 | Talking Face Generation by Adversarially Disentangled Audio-Visual RepresentationabstractTalking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of the talking face regions. Existing works either construct specific face appearance model on specific subjects or model the transformation between lip motion and speech. In this work, we integrate both aspects and enable arbitrary-subject talking face generation by learning disentangled audio-visual representation. We find that the talking face sequence is actually a composition of both subject-related information and speech-related information. These two spaces are then explicitly disentangled through a novel associative-and-adversarial training process. This disentangled representation has an advantage where both audio and video can serve as inputs for generation. Extensive experiments show that the proposed approach generates realistic talking face sequences on arbitrary subjects with much clearer lip motion patterns than previous work. We also demonstrate the learned audio-visual representation is extremely useful for the tasks of automatic lip reading and audio-video retrieval. Hang Zhou 0009, Yu Liu 0015, Ziwei Liu 0002, Ping Luo 0002, Xiaogang Wang 0001 |
AAAI | 1 |
| 2019 | A Graph-Based Framework to Bridge Movies and SynopsesabstractInspired by the remarkable advances in video analytics, research teams are stepping towards a greater ambition - movie understanding. However, compared to those activity videos in conventional datasets, movies are significantly different. Generally, movies are much longer and consist of much richer temporal structures. More importantly, the interactions among characters play a central role in expressing the underlying story. To facilitate the efforts along this direction, we construct a dataset called Movie Synopses Associations (MSA) over 327 movies, which provides a synopsis for each movie, together with annotated associations between synopsis paragraphs and movie segments. On top of this dataset, we develop a framework to perform matching between movie segments and synopsis paragraphs. This framework integrates different aspects of a movie, including event dynamics and character interactions, and allows them to be matched with parsed paragraphs, based on a graph-based formulation. Our study shows that the proposed framework remarkably improves the matching accuracy over conventional feature-based methods. It also reveals the importance of narrative structures and character interactions in movie understanding. Dataset and code are available at: https://ycxioooong.github.io/projects/moviesyn. Qingqiu Huang, Lingfeng Guo, Hang Zhou 0009, Bolei Zhou, Dahua Lin |
ICCV | 4 |
| 2019 | Vision-Infused Deep Audio InpaintingabstractMulti-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, i.e. synthesizing missing audio segments that correspond to their accompanying videos. We identify two key aspects for a successful inpainter: (1) It is desirable to operate on spectrograms instead of raw audios. Recent advances in deep semantic image inpainting could be leveraged to go beyond the limitations of traditional audio inpainting. (2) To synthesize visually indicated audio, a visual-audio joint feature space needs to be learned with synchronization of audio and video. To facilitate a large-scale study, we collect a new multi-modality instrument-playing dataset called MUSIC-Extra-Solo (MUSICES) by enriching MUSIC dataset [51]. Extensive experiments demonstrate that our framework is capable of inpainting realistic and varying audio segments with or without visual contexts. More importantly, our synthesized audio segments are coherent with their video counterparts, showing the effectiveness of our proposed Vision-Infused Audio Inpainter (VIAI). Code, models, dataset and video results are available at https://github.com/Hangz-nju-cuhk/ Vision-Infused-Audio-Inpainter-VIAI. Hang Zhou 0009, Ziwei Liu 0002, Xudong Xu, Ping Luo 0002, Xiaogang Wang 0001 |
ICCV | 1 |