Tianhua Qi

dblp:372/7575 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0005-5780-9374ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Speaker-independent speech emotion recognition using group sparse-based adversarial local fisher discriminant analysis
Cheng Lu 0005, Kaifei Zhang, Hailun Lian, Sunan Li, Tianhua Qi, Yuan Zong, Wenming Zheng
Pattern Recognit.5
2025 Enhancing Zero-Shot Emotional Voice Conversion via Speaker Adaptation and Duration Prediction
abstract
Zero-shot Emotional Voice Conversion (EVC) aims to transform a speaker’s emotional state to match a target emotion, even for speakers and emotion categories that were not encountered during training, thereby enhancing the generalization ability of traditional EVC systems. Despite advancements in the field, existing methods often face challenges in preserving speaker identity and ensuring the naturalness of emotional expression, particularly in the context of rhythm modeling. To this end, we propose the Zero-Shot Emotion Voice Conversion (ZSEVC) model, which leverages self-supervised learning for speaker adaptation and duration prediction. To adjust speech rhythm in alignment with the target emotional state, we introduce a rhythm-aware content encoder that captures and refines discrete speech units at a finer granularity. Additionally, a hierarchical emotion fusion scheme is employed to integrate emotional features with content features, enhancing both pronunciation accuracy and emotional expressiveness. Moreover, a residual speaker-emotion fusion module is incorporated to better adapt speaker characteristics to emotional prosodic variation. Experimental results show ZSEVC’s superior performance in terms of naturalness and speaker similarity in zero-shot scenario, successfully generating emotional speeches for unseen emotions and speakers. Speech samples are available at https://wosyoo.github.io/ZSEVC.
Shiyan Wang, Tianhua Qi, Cheng Lu 0005, Zhaojie Luo, Wenming Zheng
ICASSP2
2025 DisenEmo: Learning disentangled emotional representation from facial motion for 3D talking head generation
abstract
Emotional 3D talking head generation synthesizes vivid facial expressions with precise lip synchronization for immersive interactions. We introduce DisenEMO, a novel framework designed to disentangle emotion and content from facial motions, thereby facilitating the synthesis of personalized and expressive audio-driven facial animations. To achieve precise emotional disentanglement, we incorporate an intensity perception constraint which improves the accurate perception of categorized emotion and its intensities, leading to the generation of subtle emotional expressions. To ensure the temporal consistency of facial expressions, we introduce facial dynamic modeling, which refines motion trajectories to better capture emotional nuances. Finally, a motion decoder integrates emotional features with audio features extracted from driving speech, producing 3D talking heads with enhanced emotional expressiveness and realism. Experimental results demonstrate that our method outperforms state-of-the-art approaches. Synthesis samples are available at https://c295bw.github.io/DisenEMO-icip25.github.io/.
Tianhua Qi, Cheng Lu 0005, Wenming Zheng
ICIP2
2025 PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
Tianhua Qi, Shiyan Wang, Cheng Lu 0005, Tengfei Song, Hao Yang 0006, Zhanglin Wu, Wenming Zheng
INTERSPEECH1
2025 Multi-Level Segment Fusion Based on Adaptive Time-Window Selection for Multimodal Personality-Aware Elderly Depression Detection
abstract
Major Depressive Disorder (MDD) is a prevalent and severe psychiatric disorder, and its detection remains challenging due to the complexity and variability of its symptoms. Traditional single-modality methods often fail to capture the full spectrum of depressive cues, which has led to the rise of multimodal methods. The ACM Multimedia 2025 ''Multimodal Personality-Aware Depression Detection Challenge'' (MPDD 2025) aims to advance the development of more accurate depression detection models by incorporating multimodal data. In this paper, we proposed a Multi-Level Segment Fusion Based on Adaptive Time-Window Selection (MSF-ATS) method for the MPDD-Elderly Track. To address the challenge of sparse and transient depressive symptoms, we fuse segment-level classifications to obtain subject-level classifications. An adaptive time-window selection based on mean class variance is employed to choose the window with the smallest variance for more stable detection results. Our method achieved an average score of 0.8576 on the MPDD 2025 official test set, significantly outperforming the baseline score of 0.6675.
Yuyun Liu, Kaifei Zhang, Yinghao Ma, Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong
ACM Multimedia5
2025 Assessing Personality Traits and Interview Performance from Asynchronous Video Interviews
abstract
Asynchronous Video Interviews (AVIs) allow candidates to record responses to predefined questions using digital devices, offering both flexibility and remote accessibility. Assessing personality traits and interview performance via AVIs provides organizations with valuable insights into candidate profiles and facilitates the prediction of future job performance. However, prior benchmark challenges, whose datasets were predominantly sourced from social media, suffer from suboptimal construct and methodological validity, limiting their utility for model development and real-world applications. To address these limitations, we introduce the AVI Grand Challenge at ACM Multimedia 2025, featuring a novel dataset of mock AVIs comprising 3,876 videos from 646 participants in a simulated job application procedure. Interview questions were carefully designed to reflect real-world selection contexts and elicit personality expressions grounded in Trait Activation Theory. Personality traits and job competencies were annotated by trained evaluators and professional recruiters, ensuring both methodological rigor and ecological validity. The solutions and algorithms developed in this challenge are analyzed and summarized in this paper to foster the development of fair, reliable, and AI-driven hiring assessments.
Tianyi Zhang 0013, Tianhua Qi, Antonis Koutsoumpis, Yuan Zong, Wenming Zheng, Janneke K. Oostrom, Djurre Holtrop, Zhaojie Luo, Reinout E. de Vries
ACM Multimedia2
2025 Low-rank joint distribution adaptation for cross-corpus speech emotion recognition
Sunan Li, Cheng Lu 0005, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Yuan Zong
Knowl. Based Syst.5
2024 PAVITS: Exploring Prosody-Aware VITS for End-to-End Emotional Voice Conversion
abstract
In this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of converted audio, we have developed an end-to-end EVC architecture inspired by the high audio quality of VITS. By seamlessly integrating an acoustic converter and vocoder, we effectively address the common issue of mismatch between emotional prosody training and run-time conversion that is prevalent in existing EVC models. To further enhance the emotional naturalness, we introduce an emotion descriptor to model the subtle prosody variations of different speech emotions. Additionally, we propose a prosody predictor, which predicts prosody features from text based on the provided emotion label. Notably, we introduce a prosody alignment loss to establish a connection between latent prosody features from two distinct modalities, ensuring effective training. Experimental results show that the performance of PAVITS is superior to the state-of-the-art EVC methods. Speech Samples are available at https://jeremychee4.github.io/pavits4EVC/.
Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong, Hailun Lian
ICASSP1
2024 Hierarchical Distribution Adaptation for Unsupervised Cross-corpus Speech Emotion Recognition
Cheng Lu 0005, Yuan Zong, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Björn W. Schuller, Wenming Zheng
INTERSPEECH5
2024 Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
Tianhua Qi, Shiyan Wang, Cheng Lu 0005, Yan Zhao 0037, Yuan Zong, Wenming Zheng
INTERSPEECH1
2024 Exploring corpus-invariant emotional acoustic feature for cross-corpus speech emotion recognition
Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Sunan Li, Tianhua Qi, Yuan Zong
Expert Syst. Appl.5