Zhisheng Wang 0001

dblp:23/3272-1 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0001-7604-3942ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
abstract
Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can achieve human-level auditory perception, particularly in terms of their ability to comprehend latent intentions and implicit emotions in real-world spoken language, remains underexplored. To this end, we introduce the Human-level Perception in Spoken Speech Understanding (HPSU), a new benchmark for fully evaluating the human-level perceptual and understanding capabilities of Speech LLMs. HPSU comprises over 20,000 expert-validated spoken language understanding samples in English and Chinese. It establishes a comprehensive evaluation framework by encompassing a spectrum of tasks, ranging from basic speaker attribute recognition to complex inference of latent intentions and implicit emotions. To address the issues of data scarcity and high cost of manual annotation in real-world scenarios, we developed a semi-automatic annotation process. This process fuses audio, textual, and visual information to enable precise speech understanding and labeling, thus enhancing both annotation efficiency and quality. We systematically evaluate various open-source and proprietary Speech LLMs. The results demonstrate that even top-performing models still fall considerably short of human capabilities in understanding genuine spoken interactions. Consequently, HPSU will be useful for guiding the development of Speech LLMs toward human-level perception and cognition.
Peiji Yang, Yicheng Zhong, Jianxing Yu, Zhisheng Wang 0001, Zihao Gou, Wenqing Chen, Jian Yin 0001
AAAI5
2025 RealPortrait: Realistic Portrait Animation with Diffusion Transformers
abstract
We introduce RealPortrait, a framework based on Diffusion Transformers (DiT), designed to generate highly expressive and visually appealing portrait animations. Given a static portrait image, our method can transfer complex facial expressions and head pose movements extracted from a driving video onto the portrait, transforming it into a lifelike video. Specifically, we exploit the robust spatial-temporal modeling capabilities of DiT, enabling the generation of portrait videos that maintain high-fidelity visual details and ensure temporal coherence. In contrast to conventional image-to-video generation frameworks that necessitate a separate reference network, we incorporate an efficient reference attention within the DiT backbone, thereby obviating the computational overhead and achieving superior reference appearance preservation. Concurrently, we integrate a parallel ControlNet to precisely regulate intricate facial expressions and head poses. Diverging from prior methods that utilize explicit sparse motion representations, such as facial landmarks or 3DMM coefficients, we adopt a dense implicit motion representation as the control guidance. This implicit motion representation excels in capturing nuanced emotional facial expressions and subtle non-rigid dynamics of the lips. To further enhance the generalization capability of the model, we augment the training dataset by incorporating a substantial volume of facial image data through random crop augmentation. This strategy ensures the model's robustness across a wide variety of facial appearances and expressions. Empirical evaluations demonstrate that RealPortrait excels in generating portrait animations with highly-realistic quality and exceptional temporal coherence in appearance retention.
Zejun Yang, Huawei Wei, Zhisheng Wang 0001
AAAI3
2025 Eliciting Implicit Acoustic Styles from Open-domain Instructions to Facilitate Fine-grained Controllable Generation of Speech
abstract
This paper focuses on generating speech with the acoustic style that meets users' needs based on their open-domain instructions.To control the style, early work mostly relies on predefined rules or templates.The control types and formats are fixed in a closed domain, making it hard to meet users' diverse needs.One solution is to resort to instructions in free text to guide the generation.Current work mainly studies the instructions that clearly specify the acoustic styles, such as low pitch and fast speed.However, the instructions are complex, some even vague and abstract, such as "Generate a voice of a woman who is heartbroken due to a breakup."It is hard to infer this implicit style by traditional matching-based methods.To address this problem, we propose a new controllable model.It first utilizes multimodal LLMs with knowledge-augmented techniques to infer the desired speech style from the instructions.The powerful language understanding ability of LLMs can help us better elicit the implicit style factors from the instruction.By using these factors as a control condition, we design a diffusion-based generator adept at finely adjusting speech details, enabling higher flexibility to meet complex users' needs.Next, we verify the output speech from three aspects, i.e., consistency of decoding state, mel-spectrogram, and instruction style.This verified feedback can inversely optimize the generator.Extensive experiments are conducted on three popular datasets.The results show the effectiveness and good controllability of our approach.
Jianxing Yu, Zihao Gou, Zhisheng Wang 0001, Peiji Yang, Wenqing Chen, Jian Yin 0001
EMNLP4
2024 3D Visibility-Aware Generalizable Neural Radiance Fields for Interacting Hands
abstract
Neural radiance fields (NeRFs) are promising 3D representations for scenes, objects, and humans. However, most existing methods require multi-view inputs and per-scene training, which limits their real-life applications. Moreover, current methods focus on single-subject cases, leaving scenes of interacting hands that involve severe inter-hand occlusions and challenging view variations remain unsolved. To tackle these issues, this paper proposes a generalizable visibility-aware NeRF (VA-NeRF) framework for interacting hands. Specifically, given an image of interacting hands as input, our VA-NeRF first obtains a mesh-based representation of hands and extracts their corresponding geometric and textural features. Subsequently, a feature fusion module that exploits the visibility of query points and mesh vertices is introduced to adaptively merge features of both hands, enabling the recovery of features in unseen areas. Additionally, our VA-NeRF is optimized together with a novel discriminator within an adversarial learning paradigm. In contrast to conventional discriminators that predict a single real/fake label for the synthesized image, the proposed discriminator generates a pixel-wise visibility map, providing fine-grained supervision for unseen areas and encouraging the VA-NeRF to improve the visual quality of synthesized images. Experiments on the Interhand2.6M dataset demonstrate that our proposed VA-NeRF outperforms conventional NeRFs significantly. Project Page: https://github.com/XuanHuang0/VANeRF.
Zejun Yang, Zhisheng Wang 0001, Xiaodan Liang
AAAI4
2024 Monocular 3D Hand Mesh Recovery via Dual Noise Estimation
abstract
Current parametric models have made notable progress in 3D hand pose and shape estimation. However, due to the fixed hand topology and complex hand poses, current models are hard to generate meshes that are aligned with the image well. To tackle this issue, we introduce a dual noise estimation method in this paper. Given a single-view image as input, we first adopt a baseline parametric regressor to obtain the coarse hand meshes. We assume the mesh vertices and their image-plane projections are noisy, and can be associated in a unified probabilistic model. We then learn the distributions of noise to refine mesh vertices and their projections. The refined vertices are further utilized to refine camera parameters in a closed-form manner. Consequently, our method obtains well-aligned and high-quality 3D hand meshes. Extensive experiments on the large-scale Interhand2.6M dataset demonstrate that the proposed method not only improves the performance of its baseline by more than 10% but also achieves state-of-the-art performance. Project page: https://github.com/hanhuili/DNE4Hand.
Xiaojian Lin, Zejun Yang, Zhisheng Wang 0001, Xiaodan Liang
AAAI5
2024 ExpCLIP: Bridging Text and Facial Expressions via Semantic Alignment
abstract
The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression templates, which may limit the necessary flexibility for accurately conveying user intent. In this research, we introduce a technique that enables the control of arbitrary styles by leveraging natural language as emotion prompts. This technique presents benefits in terms of both flexibility and user-friendliness. To realize this objective, we initially construct a Text-Expression Alignment Dataset (TEAD), wherein each facial expression is paired with several prompt-like descriptions. We propose an innovative automatic annotation method, supported by CahtGPT, to expedite the dataset construction, thereby eliminating the substantial expense of manual annotation. Following this, we utilize TEAD to train a CLIP-based model, termed ExpCLIP, which encodes text and facial expressions into semantically aligned style embeddings. The embeddings are subsequently integrated into the facial animation generator to yield expressive and controllable facial animations. Given the limited diversity of facial emotions in existing speech-driven facial animation training data, we further introduce an effective Expression Prompt Augmentation (EPA) mechanism to enable the animation generator to support unprecedented richness in style control. Comprehensive experiments illustrate that our method accomplishes expressive facial animation generation and offers enhanced flexibility in effectively conveying the desired style.
Yicheng Zhong, Huawei Wei, Peiji Yang, Zhisheng Wang 0001
AAAI4
2024 Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
Peiji Yang, Yicheng Zhong, Yixuan Zhou 0002, Zhisheng Wang 0001, Zhiyong Wu 0001, Xixin Wu, Helen M. Meng
INTERSPEECH5
2023 Semi-supervised Speech-driven 3D Facial Animation via Cross-modal Encoding
abstract
Existing Speech-driven 3D facial animation methods typically follow the supervised paradigm, involving regression from speech to 3D facial animation. This paradigm faces two major challenges: the high cost of supervision acquisition, and the ambiguity in mapping between speech and lip movements. To address these challenges, this study proposes a novel cross-modal semi-supervised framework, comprising a Speech-to-Image Transcoder and a Face-to-Geometry Regressor. The former jointly learns a common representation space from speech and image domains, enabling the transformation of speech into semantically-consistent facial images. The latter is responsible for reconstructing 3D facial meshes from the transformed images. Both modules require minimal effort to acquire the necessary training data, thereby obviating the dependence on costly supervised data. Furthermore, the joint learning scheme enables the fusion of intricate visual features into speech encoding, thereby facilitating the transformation of subtle speech variations into nuanced lip movements, ultimately enhancing the fidelity of 3D face reconstructions. Consequently, the ambiguity of the direct mapping of speech-to-animation is significantly reduced, leading to coherent and high-fidelity generation of lip motion. Extensive experiments demonstrate that our approach produces competitive results compared to supervised methods.
Peiji Yang, Huawei Wei, Yicheng Zhong, Zhisheng Wang 0001
ICCV4
2015 Online Personalized Recommendation Based on Streaming Implicit User Feedback
Zhisheng Wang 0001, Wei Liu 0061, Jian Yin 0001
APWeb1