VLDB 2026 Research / reviewers in the wild / expert
Longhao Zhang
dblp:236/7382
· DBLP profile ↗
15ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-2787-3664ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MaTe3D: Mask-Guided Text-based 3D-Aware Portrait Editing
Kangneng Zhou, Daiheng Gao, Xuan Wang 0009, Jie Zhang 0090, Peng Zhang 0080, Xusen Sun, Longhao Zhang, Shiqi Yang 0002, Bang Zhang, Liefeng Bo, Yaxing Wang, Ming-Ming Cheng |
Int. J. Comput. Vis. | 7 |
| 2025 | VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid PriorabstractAudio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, facial motion, head pose generation, and video quality. However, no model has yet led or tied on all these metrics due to the one-to-many mapping between audio and motion. In this paper, we propose VividTalk, a two-stage generic framework that supports generating high-visual quality talking head videos with all the above properties. Specifically, in the first stage, we map the audio to mesh by learning two motions, including non-rigid facial motion and rigid head motion. For facial motion, both blendshape and vertex are adopted as the intermediate representation to maximize the representation ability of the model. For head motion, a novel learnable head pose codebook with a two-phase training mechanism is proposed. In the second stage, we proposed a dual branch motion-vae and a generator to transform the meshes into dense motion and synthesize high-quality video frame-by-frame. Extensive experiments show that the proposed VividTalk can generate high-visual quality talking head videos with lip-sync and realistic enhanced by a large margin, and outperforms previous state-of-the-art works in objective and subjective comparisons. The code will be publicly released upon publication. Xusen Sun, Longhao Zhang, Hao Zhu 0004, Peng Zhang 0080, Bang Zhang, Xinya Ji, Kangneng Zhou, Daiheng Gao, Liefeng Bo, Xun Cao |
3DV | 2 |
| 2025 | INFP: Audio-Driven Interactive Head Generation in Dyadic ConversationsabstractImagine having a conversation with a socially intelligent agent. It can attentively listen to your words and offer visual and linguistic feedback promptly. This seamless interaction allows for multiple rounds of conversation to flow smoothly and naturally. In pursuit of actualizing it, we propose INFP, a novel audio-driven head generation framework for dyadic interaction. Unlike previous head generation works that only focus on single-sided communication, or require manual role assignment and explicit role switching, our model drives the agent portrait dynamically alternates between speaking and listening state, guided by the input dyadic audio. Specifically, INFP comprises a Motion-Based Head Imitation stage and an Audio-Guided Motion Generation stage. The first stage learns to project facial communicative behaviors from real-life conversation videos into a low-dimensional motion latent space, and use the motion latent codes to animate a static image. The second stage learns the mapping from the input dyadic audio to motion latent codes through denoising, leading to the audio-driven head generation in interactive scenarios. To facilitate this line of research, we introduce DyConv, a large scale dataset of rich dyadic conversations collected from the Internet. Extensive experiments and visualizations demonstrate superior performance and effectiveness of our method. Yongming Zhu, Longhao Zhang, Zhengkun Rong, Tianshu Hu, Zhipeng Ge |
CVPR | 2 |
| 2025 | DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid GuidanceabstractWhile recent image-based human animation methods achieve realistic body and facial motion synthesis, critical gaps remain in fine-grained holistic controllability, multi-scale adaptability, and long-term temporal coherence, which leads to their lower expressiveness and robustness. We propose a diffusion transformer (DiT) based framework, DreamActor-M1, with hybrid guidance to overcome these limitations. For motion guidance, our hybrid control signals that integrate implicit facial representations, 3D head spheres, and 3D body skeletons achieve robust control of facial expressions and body movements, while producing expressive and identity-preserving animations. For scale adaptation, to handle various body poses and image scales ranging from portraits to full-body views, we employ a progressive training strategy using data with varying resolutions and scales. For appearance guidance, we integrate motion patterns from sequential frames with complementary visual references, ensuring long-term temporal coherence for unseen regions during complex movements. Experiments demonstrate that our method outperforms the state-of-the-art works, delivering expressive results for portraits, upper-body, and full-body generation with robust long-term consistency. Project Page: https://grisoon.github.io/DreamActor-M1/. Yuxuan Luo 0002, Zhengkun Rong, Longhao Zhang, Tianshu Hu |
ICCV | 4 |
| 2025 | Domain Attention and Confidence-Aware Unsupervised Domain Adaptation Network
Longhao Zhang |
ICIC (21) | 1 |
| 2025 | Textual similarity calculation techniques in the medical field: a retrospective review
Hongzhen Cui, Haoming Ma, Xiaoyue Zhu, Longhao Zhang, Meihua Piao |
Appl. Intell. | 5 |
| 2025 | HGBL: A Fine Granular Hierarchical Multi-Label Text Classification ModelabstractHierarchical multi-label text classification is vital for natural language processing (NLP). However, existing research rarely makes full use of the interaction between labels and text features that are crucial to hierarchical multi-label text classification. To address this issue, a novel model named hierarchy-guided BiLSTM guided contrastive learning classification (HGBL) is proposed, which successfully enhances the interaction between labels and text features by incorporating global context and embedding the idea of contrastive learning into this model. During modeling, Graphormer is adopted to model the dependencies between labels, and the bidirectional recurrent network (BiLSTM) is used to integrate global context including label features. Afterwards, the contrastive learning module embeds hierarchical awareness into the fine-tuned bidirectional encoder representations from transformers (BERT) by training the value of the loss. Experimental results on NYT, WOS and RCV1-V2 datasets show that HGBL exhibits significant competitive advantages compared with 19 competitors in terms of several indicators and can be used effectively for hierarchical multi-label text classification problems. Linlin Dai, Chengxing Liu, Longhao Zhang |
Neural Process. Lett. | 4 |
| 2024 | PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
Longhao Zhang, Zhipeng Ge, Tianshu Hu |
SIGGRAPH Asia | 1 |
| 2023 | One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance FieldabstractTalking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encountered. Recent works instead employ explicit 3D structural representations or implicit neural rendering to improve performance under large pose changes. Nevertheless, the fidelity of identity and expression is not so desirable, especially for novel-view synthesis. In this paper, we propose HiDe-NeRF, which achieves high-fidelity and free-view talking-head synthesis. Drawing on the recently proposed Deformable Neural Radiance Fields, HiDe-NeRF represents the 3D dynamic scene into a canonical appearance field and an implicit deformation field, where the former comprises the canonical source face and the latter models the driving pose and expression. In particular, we improve fidelity from two aspects: (i) to enhance identity expressiveness, we design a generalized appearance module that leverages multi-scale volume features to preserve face shape and details; (ii) to improve expression preciseness, we propose a lightweight deformation module that explicitly decouples the pose and expression to enable precise expression modeling. Extensive experiments demonstrate that our proposed approach can generate better results than previous works. Project page: https://www.waytron.net/hidenerf/ Weichuang Li, Longhao Zhang, Dong Wang 0028, Bin Zhao 0001, Zhigang Wang 0002, Mulin Chen, Bang Zhang, Zhongjian Wang, Liefeng Bo, Xuelong Li 0001 |
CVPR | 2 |
| 2023 | A two-layer BiLSTM model with linear gating for Chinese named entity recognitionabstractChinese named entity recognition (CNER) is one of the most fundamental tasks in natural language processing (NLP), and is key to extracting information from unstructured texts. In recent years, advances in neural network models and pretrained word-level information embedding techniques have played a driving role in the development of NLP. In this context, how to make full use of word vectors to extract information has become one of the research emphases. The diversity of Chinese expressions and the irregular expressions of texts lead to poor recognition results. This paper proposes a two-layer BiLSTM network model with linear gating logic to enhance the model's learning effect of word vectors within sentences and word memory. The aim is to solve the problem of gradient disappearance and improve the model's generalization ability and entity recognition. Through experiments, our model proved effective on three Chinese benchmark datasets: MSRA, the People's Daily Corpus (PRF), and Boson. The precision of NER performs best among similar models. In addition, using the lab-constructed medical dataset of Chinese Drugs for the Heart for testing, our model outperforms the existing BiLSTM model. Finally, statistical analysis of the changes in F1 during training demonstrated faster convergence of our model. Hongzhen Cui, Longhao Zhang |
IJCNN | 2 |
| 2022 | Depth-Aware Generative Adversarial Network for Talking Head Video GenerationabstractTalking head video generation aims to produce a synthetic human face video that contains the identity and pose information respectively from a given source image and a driving video. Existing works for this task heavily rely on 2D representations (e.g. appearance and motion) learned from the input images. However, dense 3D facial geometry (e.g. pixel-wise depth) is extremely important for this task as it is particularly beneficial for us to essentially generate accurate 3D face structures and distinguish noisy information from the possibly cluttered background. Nevertheless, dense 3D geometry annotations are prohibitively costly for videos and are typically not available for this video generation task. In this paper, we introduce a self-supervised face-depth learning method to automatically recover dense 3D facial geometry (i.e. depth) from the face videos without the requirement of any expensive 3D annotation data. Based on the learned dense depth maps, we further propose to leverage them to estimate sparse facial keypoints that capture the critical movement of the human head. In a more dense way, the depth is also utilized to learn 3D-aware cross-modal (i.e. appearance and depth) attention to guide the generation of motion fields for warping source image representations. All these contributions compose a novel depth-aware generative adversarial network (DaGAN) for talking head generation. Extensive experiments conducted demonstrate that our proposed method can generate highly realistic faces, and achieve significant results on the unseen human faces.11https://github.com/harlanhong/CVPR2022-DaGAN Fa-Ting Hong, Longhao Zhang, Li Shen 0005, Dan Xu 0002 |
CVPR | 2 |
| 2022 | AP-GAN: Improving Attribute Preservation in Video Face SwappingabstractFace swapping is a popular subject in face manipulation, which aims to replace the identity of the target face with that of the source face. Existing methods cannot well preserve facial attributes (e.g., pose, expression, skin color, illumination, make-up, occlusion, etc.) of the target face, causing noticeable temporal discontinuity and instability artifacts for video face swapping. In this paper, we propose a lightweight Generative Adversarial Networks based framework named AP-GAN, which can precisely control the attribute of the generated face to be consistent with that of the target face, achieving efficient and high-fidelity video face swapping. Specifically, we derive a U-Net based generator with ID blocks to translate identity and PE blocks to correct pose and expression. Besides, a PE-aware discriminator is designed to help supervise pose and expression of the synthetic face. Furthermore, we propose a discriminator based perceptual loss leveraging multi-scale features of the discriminator to preserve facial attributes like skin color, illumination, make-up and occlusion. AP-GAN is trained on Flickr-Faces-HQ, CelebA-HQ and VGGFace2 and evaluated on FaceForensics++. Extensive experiments and comparisons to the existing state-of-the-art face swapping methods demonstrate the efficacy of our framework. Comprehensive ablation studies are also carried out to isolate the validity of each proposed component and to contrast with other face manipulation approaches. Longhao Zhang, Tian Qiu 0001, Lingqiao Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Adaptive attention augmentor for weakly supervised object localization
Longhao Zhang |
Neurocomputing | 1 |
| 2020 | Multi-task deep learning for fine-grained classification and grading in breast cancer histopathological images
Lingqiao Li, Xipeng Pan, Zhenbing Liu, Yubei He, Zhongming Li, Yong-Xian Fan, Longhao Zhang |
Multim. Tools Appl. | 9 |
| 2019 | Multi-Level Ensemble Network for Scene Recognition
Longhao Zhang, Lingqiao Li, Xipeng Pan |
Multim. Tools Appl. | 1 |