Longhao Zhang

dblp:236/7382 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-2787-3664ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MaTe3D: Mask-Guided Text-based 3D-Aware Portrait Editing
Kangneng Zhou, Daiheng Gao, Xuan Wang 0009, Jie Zhang 0090, Peng Zhang 0080, Xusen Sun, Longhao Zhang, Shiqi Yang 0002, Bang Zhang, Liefeng Bo, Yaxing Wang, Ming-Ming Cheng
Int. J. Comput. Vis.7
2025 VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior
abstract
Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, facial motion, head pose generation, and video quality. However, no model has yet led or tied on all these metrics due to the one-to-many mapping between audio and motion. In this paper, we propose VividTalk, a two-stage generic framework that supports generating high-visual quality talking head videos with all the above properties. Specifically, in the first stage, we map the audio to mesh by learning two motions, including non-rigid facial motion and rigid head motion. For facial motion, both blendshape and vertex are adopted as the intermediate representation to maximize the representation ability of the model. For head motion, a novel learnable head pose codebook with a two-phase training mechanism is proposed. In the second stage, we proposed a dual branch motion-vae and a generator to transform the meshes into dense motion and synthesize high-quality video frame-by-frame. Extensive experiments show that the proposed VividTalk can generate high-visual quality talking head videos with lip-sync and realistic enhanced by a large margin, and outperforms previous state-of-the-art works in objective and subjective comparisons. The code will be publicly released upon publication.
Xusen Sun, Longhao Zhang, Hao Zhu 0004, Peng Zhang 0080, Bang Zhang, Xinya Ji, Kangneng Zhou, Daiheng Gao, Liefeng Bo, Xun Cao
3DV2
2025 INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations
abstract
Imagine having a conversation with a socially intelligent agent. It can attentively listen to your words and offer visual and linguistic feedback promptly. This seamless interaction allows for multiple rounds of conversation to flow smoothly and naturally. In pursuit of actualizing it, we propose INFP, a novel audio-driven head generation framework for dyadic interaction. Unlike previous head generation works that only focus on single-sided communication, or require manual role assignment and explicit role switching, our model drives the agent portrait dynamically alternates between speaking and listening state, guided by the input dyadic audio. Specifically, INFP comprises a Motion-Based Head Imitation stage and an Audio-Guided Motion Generation stage. The first stage learns to project facial communicative behaviors from real-life conversation videos into a low-dimensional motion latent space, and use the motion latent codes to animate a static image. The second stage learns the mapping from the input dyadic audio to motion latent codes through denoising, leading to the audio-driven head generation in interactive scenarios. To facilitate this line of research, we introduce DyConv, a large scale dataset of rich dyadic conversations collected from the Internet. Extensive experiments and visualizations demonstrate superior performance and effectiveness of our method.
Yongming Zhu, Longhao Zhang, Zhengkun Rong, Tianshu Hu, Zhipeng Ge
CVPR2
2025 DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance
abstract
While recent image-based human animation methods achieve realistic body and facial motion synthesis, critical gaps remain in fine-grained holistic controllability, multi-scale adaptability, and long-term temporal coherence, which leads to their lower expressiveness and robustness. We propose a diffusion transformer (DiT) based framework, DreamActor-M1, with hybrid guidance to overcome these limitations. For motion guidance, our hybrid control signals that integrate implicit facial representations, 3D head spheres, and 3D body skeletons achieve robust control of facial expressions and body movements, while producing expressive and identity-preserving animations. For scale adaptation, to handle various body poses and image scales ranging from portraits to full-body views, we employ a progressive training strategy using data with varying resolutions and scales. For appearance guidance, we integrate motion patterns from sequential frames with complementary visual references, ensuring long-term temporal coherence for unseen regions during complex movements. Experiments demonstrate that our method outperforms the state-of-the-art works, delivering expressive results for portraits, upper-body, and full-body generation with robust long-term consistency. Project Page: https://grisoon.github.io/DreamActor-M1/.
Yuxuan Luo 0002, Zhengkun Rong, Longhao Zhang, Tianshu Hu
ICCV4
2025 Domain Attention and Confidence-Aware Unsupervised Domain Adaptation Network
Longhao Zhang
ICIC (21)1
2025 Textual similarity calculation techniques in the medical field: a retrospective review
Hongzhen Cui, Haoming Ma, Xiaoyue Zhu, Longhao Zhang, Meihua Piao
Appl. Intell.5
2025 HGBL: A Fine Granular Hierarchical Multi-Label Text Classification Model
abstract
Hierarchical multi-label text classification is vital for natural language processing (NLP). However, existing research rarely makes full use of the interaction between labels and text features that are crucial to hierarchical multi-label text classification. To address this issue, a novel model named hierarchy-guided BiLSTM guided contrastive learning classification (HGBL) is proposed, which successfully enhances the interaction between labels and text features by incorporating global context and embedding the idea of contrastive learning into this model. During modeling, Graphormer is adopted to model the dependencies between labels, and the bidirectional recurrent network (BiLSTM) is used to integrate global context including label features. Afterwards, the contrastive learning module embeds hierarchical awareness into the fine-tuned bidirectional encoder representations from transformers (BERT) by training the value of the loss. Experimental results on NYT, WOS and RCV1-V2 datasets show that HGBL exhibits significant competitive advantages compared with 19 competitors in terms of several indicators and can be used effectively for hierarchical multi-label text classification problems.
Linlin Dai, Chengxing Liu, Longhao Zhang
Neural Process. Lett.4
2024 PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
Longhao Zhang, Zhipeng Ge, Tianshu Hu
SIGGRAPH Asia1
2023 One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance Field
abstract
Talking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encountered. Recent works instead employ explicit 3D structural representations or implicit neural rendering to improve performance under large pose changes. Nevertheless, the fidelity of identity and expression is not so desirable, especially for novel-view synthesis. In this paper, we propose HiDe-NeRF, which achieves high-fidelity and free-view talking-head synthesis. Drawing on the recently proposed Deformable Neural Radiance Fields, HiDe-NeRF represents the 3D dynamic scene into a canonical appearance field and an implicit deformation field, where the former comprises the canonical source face and the latter models the driving pose and expression. In particular, we improve fidelity from two aspects: (i) to enhance identity expressiveness, we design a generalized appearance module that leverages multi-scale volume features to preserve face shape and details; (ii) to improve expression preciseness, we propose a lightweight deformation module that explicitly decouples the pose and expression to enable precise expression modeling. Extensive experiments demonstrate that our proposed approach can generate better results than previous works. Project page: https://www.waytron.net/hidenerf/
Weichuang Li, Longhao Zhang, Dong Wang 0028, Bin Zhao 0001, Zhigang Wang 0002, Mulin Chen, Bang Zhang, Zhongjian Wang, Liefeng Bo, Xuelong Li 0001
CVPR2
2023 A two-layer BiLSTM model with linear gating for Chinese named entity recognition
abstract
Chinese named entity recognition (CNER) is one of the most fundamental tasks in natural language processing (NLP), and is key to extracting information from unstructured texts. In recent years, advances in neural network models and pretrained word-level information embedding techniques have played a driving role in the development of NLP. In this context, how to make full use of word vectors to extract information has become one of the research emphases. The diversity of Chinese expressions and the irregular expressions of texts lead to poor recognition results. This paper proposes a two-layer BiLSTM network model with linear gating logic to enhance the model's learning effect of word vectors within sentences and word memory. The aim is to solve the problem of gradient disappearance and improve the model's generalization ability and entity recognition. Through experiments, our model proved effective on three Chinese benchmark datasets: MSRA, the People's Daily Corpus (PRF), and Boson. The precision of NER performs best among similar models. In addition, using the lab-constructed medical dataset of Chinese Drugs for the Heart for testing, our model outperforms the existing BiLSTM model. Finally, statistical analysis of the changes in F1 during training demonstrated faster convergence of our model.
Hongzhen Cui, Longhao Zhang
IJCNN2
2022 Depth-Aware Generative Adversarial Network for Talking Head Video Generation
abstract
Talking head video generation aims to produce a synthetic human face video that contains the identity and pose information respectively from a given source image and a driving video. Existing works for this task heavily rely on 2D representations (e.g. appearance and motion) learned from the input images. However, dense 3D facial geometry (e.g. pixel-wise depth) is extremely important for this task as it is particularly beneficial for us to essentially generate accurate 3D face structures and distinguish noisy information from the possibly cluttered background. Nevertheless, dense 3D geometry annotations are prohibitively costly for videos and are typically not available for this video generation task. In this paper, we introduce a self-supervised face-depth learning method to automatically recover dense 3D facial geometry (i.e. depth) from the face videos without the requirement of any expensive 3D annotation data. Based on the learned dense depth maps, we further propose to leverage them to estimate sparse facial keypoints that capture the critical movement of the human head. In a more dense way, the depth is also utilized to learn 3D-aware cross-modal (i.e. appearance and depth) attention to guide the generation of motion fields for warping source image representations. All these contributions compose a novel depth-aware generative adversarial network (DaGAN) for talking head generation. Extensive experiments conducted demonstrate that our proposed method can generate highly realistic faces, and achieve significant results on the unseen human faces.11https://github.com/harlanhong/CVPR2022-DaGAN
Fa-Ting Hong, Longhao Zhang, Li Shen 0005, Dan Xu 0002
CVPR2
2022 AP-GAN: Improving Attribute Preservation in Video Face Swapping
abstract
Face swapping is a popular subject in face manipulation, which aims to replace the identity of the target face with that of the source face. Existing methods cannot well preserve facial attributes (e.g., pose, expression, skin color, illumination, make-up, occlusion, etc.) of the target face, causing noticeable temporal discontinuity and instability artifacts for video face swapping. In this paper, we propose a lightweight Generative Adversarial Networks based framework named AP-GAN, which can precisely control the attribute of the generated face to be consistent with that of the target face, achieving efficient and high-fidelity video face swapping. Specifically, we derive a U-Net based generator with ID blocks to translate identity and PE blocks to correct pose and expression. Besides, a PE-aware discriminator is designed to help supervise pose and expression of the synthetic face. Furthermore, we propose a discriminator based perceptual loss leveraging multi-scale features of the discriminator to preserve facial attributes like skin color, illumination, make-up and occlusion. AP-GAN is trained on Flickr-Faces-HQ, CelebA-HQ and VGGFace2 and evaluated on FaceForensics++. Extensive experiments and comparisons to the existing state-of-the-art face swapping methods demonstrate the efficacy of our framework. Comprehensive ablation studies are also carried out to isolate the validity of each proposed component and to contrast with other face manipulation approaches.
Longhao Zhang, Tian Qiu 0001, Lingqiao Li
IEEE Trans. Circuits Syst. Video Technol.1
2021 Adaptive attention augmentor for weakly supervised object localization
Longhao Zhang
Neurocomputing1
2020 Multi-task deep learning for fine-grained classification and grading in breast cancer histopathological images
Lingqiao Li, Xipeng Pan, Zhenbing Liu, Yubei He, Zhongming Li, Yong-Xian Fan, Longhao Zhang
Multim. Tools Appl.9
2019 Multi-Level Ensemble Network for Scene Recognition
Longhao Zhang, Lingqiao Li, Xipeng Pan
Multim. Tools Appl.1