Haoqin Sun

dblp:319/6338 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 15 · 3 first-author · 15 since 2021
YearPublicationVenuePosition
2026 Learning Personalised Human Internal Cognition from External Expressive Behaviours for Real Personality Recognition
abstract
Automatic real personality recognition (RPR) aims to evaluate human real personality traits from their expressive behaviours. However, most existing solutions generally act as external observers to infer observers' personality impressions based on target individuals' expressive behaviours, which significantly deviate from their real personalities and consistently lead to inferior recognition performance. Inspired by the association between real personality and human internal cognition underlying the generation of expressive behaviours, we propose a novel RPR approach that efficiently simulates personalised internal cognition from external short audio-visual behaviours expressed by target individual. The simulated personalised cognition, represented as a set of network weights that enforce the personalised network to reproduce the individual-specific facial reactions, is further encoded as a graph containing two-dimensional node and edge feature matrices, with a novel 2D Graph Neural Network (2D-GNN) proposed for inferring real personality traits from it. To simulate real personality-related cognition, an end-to-end (E2E) strategy is designed to jointly train our cognition simulation, 2D graph construction, and personality recognition modules. Experiments show our approach’s effectiveness in capturing real personality traits with superior computational efficiency.
Xiangyu Kong 0001, Hengde Zhu, Haoqin Sun, Jiayan Gu, Xinyi Ni, Wei Zhang 0243, Shizhe Liu, Siyang Song
AAAI3
2026 TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
abstract
Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generalization, and ethical concerns. We present TTA-Bench, a comprehensive benchmark for evaluating TTA models across functional performance, reliability, and social responsibility. It covers seven dimensions including accuracy, robustness, fairness, and toxicity, and includes 2,999 diverse prompts generated through automated and manual methods. We introduce a unified evaluation protocol that combines objective metrics with over 118,000 human annotations from both experts and general users. Ten state-of-the-art models are benchmarked under this framework, offering detailed insights into their strengths and limitations. TTA-Bench establishes a new standard for holistic evaluation of TTA systems.
Hui Wang 0075, Haoze Liu, Yuhang Jia, Shiwan Zhao, Jiaming Zhou 0001, Haoqin Sun, Hui Bu
AAAI8
2026 DIFFA: Large Language Diffusion Models Can Listen and Understand
abstract
Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, large language diffusion models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context modeling, and robust generation. However, their application to the audio modality remains underexplored. In this work, we introduce DIFFA, the first diffusion-based large audio-language model designed to perform spoken language understanding. DIFFA integrates a frozen diffusion language model with a lightweight dual-adapter architecture that bridges speech understanding and natural language reasoning. We employ a two-stage training pipeline: first, aligning semantic representations via an ASR objective; then, learning instruction-following abilities through synthetic audio-caption pairs automatically generated by prompting LLMs. Despite being trained on only 960 hours of ASR and 127 hours of synthetic instruction data, DIFFA demonstrates competitive performance on major benchmarks, including MMSU, MMAU, and VoiceBench, outperforming several autoregressive open-source baselines. Our results reveal the potential of large language diffusion models for efficient and scalable audio understanding, opening a new direction for speech-driven AI.
Jiaming Zhou 0001, Hongjie Chen 0001, Shiwan Zhao, Jian Kang 0006, Jie Li 0001, Enzhi Wang, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xuelong Li 0001
AAAI8
2026 SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
abstract
Hui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu, Junyang Chen, Yanzhe Zhang, Shiwan Zhao, Jinyu Li, Jiaming Zhou, Haoqin Sun, Yan Lu, Yong Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hui Wang 0075, Jinghua Zhao 0004, Yifan Yang 0005, Shujie Liu 0001, Shiwan Zhao, Jinyu Li 0001, Jiaming Zhou 0001, Haoqin Sun, Yan Lu 0001
ACL (1)10
2026 PD-DDPM: Prior-driven diffusion model for single image dehazing
Haoqin Sun, Jiaxin Gong
Image Vis. Comput.1
2026 DBIDM: Implementing blind image separation through a dual branch interactive diffusion model
Jiaxin Gong, Haoqin Sun
Pattern Recognit. Lett.3
2026 SFFDM: Diffusion model-based spatial-frequency fusion network for remote sensing image dehazing
Haoqin Sun, Zongqin Yue
Signal Process. Image Commun.1
2025 ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
abstract
Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children’s speech remains challenging due to differences in pronunciation, tone, and pace compared to adult speech. In this paper, we introduce a new Mandarin speech dataset focused on children aged 3 to 5, addressing the scarcity of resources in this area. The dataset comprises 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. We provide a comprehensive analysis of speaker demographics, speech duration distribution and geographic coverage. Additionally, we evaluate ASR performance on models trained from scratch, such as Conformer, as well as fine-tuned pre-trained models like HuBERT and Whisper, where fine-tuning demonstrates significant performance improvements. Furthermore, we assess speaker verification (SV) on our dataset, showing that, despite the challenges posed by the unique vocal characteristics of young children, the dataset effectively supports both ASR and SV tasks. This dataset is a valuable contribution to Mandarin child speech research and holds potential for applications in educational technology and child-computer interaction. It will be open-source and freely available for all academic purposes.
Jiaming Zhou 0001, Shiwan Zhao, Jiabei He 0001, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xi Yang 0023, Yequan Wang, Yonghua Lin
ACL (1)5
2025 Emotion-Preserving Prosody Anonymization Network for Voice Privacy Protection
abstract
Balancing emotion preservation and privacy protection in voice anonymization presents a significant challenge, particularly due to the difficulty of effectively handling prosody, a key feature in speech. While preserving prosodic features in anonymized speech enhances emotional expression, it also increases the risk of leaking speaker information. To address this conflict, we propose a lightweight Emotion-Preserving Prosody Anonymization (EPPA) network, which extracts speaker-independent prosodic features to preserve speech emotion while converting them into another speaker’s style for anonymization. By combining EPPA with timbre cloning for anonymization while retaining speech content, we achieve a more balanced voice conversion. Evaluated using the Voice Privacy Challenge (VPC) 2024 metrics, our proposed EPPA, utilizing the closest center distance (CCD) anonymization strategy, demonstrates strong performance across emotional expression, content clarity, and privacy protection, achieving the highest ranking in both average and weighted ranks compared to the six baseline solutions.
Jiabei He 0001, Shiwan Zhao, Jiaming Zhou 0001, Haoqin Sun, Hui Wang 0075
ICASSP4
2025 MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
abstract
The technology for generating music from textual descriptions has seen rapid advancements. However, evaluating text-to-music (TTM) systems remains a significant challenge, primarily due to the difficulty of balancing performance and cost with existing objective and subjective evaluation methods. In this paper, we propose an automatic assessment task for TTM models to align with human perception. To address the TTM evaluation challenges posed by the professional requirements of music evaluation and the complexity of the relationship between text and music, we collect MusicEval, the first generative music assessment dataset. This dataset contains 2,748 music clips generated by 31 advanced and widely used models in response to 384 text prompts, along with 13,740 ratings from 14 music experts. Furthermore, we design a CLAP-based assessment model built on this dataset, and our experimental results validate the feasibility of the proposed task, providing a valuable reference for future development in TTM evaluation. The dataset is available at https://www.aishelltech.com/AISHELL_7A.
Hui Wang 0075, Jinghua Zhao 0004, Shiwan Zhao, Hui Bu, Jiaming Zhou 0001, Haoqin Sun
ICASSP8
2025 Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework
abstract
Multimodal emotion recognition systems rely heavily on the full availability of modalities, suffering significant performance declines when modal data is incomplete. To tackle this issue, we present the Cross-Modal Alignment, Reconstruction, and Refinement (CM-ARR) framework, an innovative approach that sequentially engages in cross-modal alignment, reconstruction, and refinement phases to handle missing modalities and enhance emotion recognition. This framework utilizes unsupervised distribution-based contrastive learning to align heterogeneous modal distributions, reducing discrepancies and modeling semantic uncertainty effectively. The reconstruction phase applies normalizing flow models to transform these aligned distributions and recover missing modalities. The refinement phase employs supervised point-based contrastive learning to disrupt semantic correlations and accentuate emotional traits, thereby enriching the affective content of the reconstructed representations. Extensive experiments confirm the superior performance of CM-ARR. Notably, averaged across six scenarios of missing modalities, CM-ARR achieves absolute improvements of 2.11%/2.12% (WAR/UAR), and 1.71%/1.96% (WAR/UAR), respectively, on IEMOCAP and MSP-IMPROV datasets.
Haoqin Sun, Shiwan Zhao, Shaokai Li, Xiangyu Kong 0001, Xuechen Wang, Jiaming Zhou 0001, Aobo Kong, Wenjia Zeng
ICASSP1
2025 Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
abstract
Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features across these modalities is significant, with most existing approaches adopting a singular alignment strategy. Such a narrow focus not only limits model performance but also fails to address the complexity and ambiguity inherent in emotional expressions. In response, this paper introduces a Multi-Granularity Cross-Modal Alignment (MGCMA) framework, distinguished by its comprehensive approach encompassing distribution-based, instance-based, and token-based alignment modules. This framework enables a multi-level perception of emotional information across modalities. Our experiments on IEMOCAP demonstrate that our proposed method outperforms current state-of-the-art techniques.
Xuechen Wang, Shiwan Zhao, Haoqin Sun, Hui Wang 0075, Jiaming Zhou 0001
ICASSP3
2025 M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
abstract
State-of-the-art models like OpenAI’s Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizing diverse subdialects. In this paper, we propose M2R-Whisper, a novel multi-stage and multi-scale retrieval augmentation approach designed to enhance ASR performance in low-resource settings. Building on the principles of in-context learning (ICL) and retrieval-augmented techniques, our method employs sentence-level ICL in the pre-processing stage to harness contextual information, while integrating token-level k-Nearest Neighbors (kNN) retrieval as a post-processing step to further refine the final output distribution. By synergistically combining sentence-level and token-level retrieval strategies, M2R-Whisper effectively mitigates various types of recognition errors. Experiments conducted on Mandarin and subdialect datasets, including AISHELL-1 and KeSpeech, demonstrate substantial improvements in ASR accuracy, all achieved without any parameter updates.
Jiaming Zhou 0001, Shiwan Zhao, Jiabei He 0001, Hui Wang 0075, Wenjia Zeng, Haoqin Sun, Aobo Kong
ICASSP7
2025 Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
abstract
The kNN-CTC model has proven to be effective for monolingual automatic speech recognition (ASR). However, its direct application to multilingual scenarios like code-switching, presents challenges. Although there is potential for performance improvement, a kNN-CTC model utilizing a single bilingual datastore can inadvertently introduce undesirable noise from the alternative language. To address this, we propose a novel kNN-CTC-based code-switching ASR (CS-ASR) framework that employs dual monolingual datastores and a gated datastore selection mechanism to reduce noise interference. Our method selects the appropriate datastore for decoding each frame, ensuring the injection of language-specific information into the ASR process. We apply this framework to cutting-edge CTC-based models, developing an advanced CS-ASR system. Extensive experiments demonstrate the remarkable effectiveness of our gated datastore mechanism in enhancing the performance of zero-shot Chinese-English CS-ASR.
Jiaming Zhou 0001, Shiwan Zhao, Hui Wang 0075, Tian-Hao Zhang, Haoqin Sun, Xuechen Wang
ICASSP5
2025 RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Haoqin Sun, Jingguang Tian, Jiaming Zhou 0001, Hui Wang 0075, Jiabei He 0001, Shiwan Zhao, Xiangyu Kong 0001, Desheng Hu, Xinkang Xu, Xinhui Hu
INTERSPEECH1
2025 Discrete Audio Representations for Automated Audio Captioning
Jingguang Tian, Haoqin Sun, Xinhui Hu, Xinkang Xu
INTERSPEECH2
2025 FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
abstract
To advance continuous token modeling and temporal-coherence enforcement, we propose FELLE, an autoregressive model that integrates language modeling with token-wise flow matching. By leveraging the autoregressive nature of language models and the generative efficacy of flow matching, FELLE effectively predicts continuous-valued tokens (mel-spectrograms). For each continuous-valued token, FELLE modifies the general prior distribution in flow matching by incorporating information from the previous step, improving coherence and stability. Furthermore, to enhance synthesis quality, FELLE introduces a coarse-to-fine flow-matching mechanism, generating continuous-valued tokens hierarchically, conditioned on the language model's output. Experimental results demonstrate the potential of incorporating flow-matching techniques in autoregressive mel-spectrogram modeling, leading to significant improvements in TTS generation quality, as shown in https://aka.ms/felle.
Hui Wang 0075, Shujie Liu 0001, Lingwei Meng, Jinyu Li 0001, Yifan Yang 0005, Shiwan Zhao, Haiyang Sun 0004, Haoqin Sun, Jiaming Zhou 0001, Yan Lu 0001
ACM Multimedia9
2025 Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs
Aobo Kong, Shiwan Zhao, Qicheng Li, Jiaming Zhou 0001, Haoqin Sun
NLPCC (1)7
2025 StreamMel: Real-Time Zero-Shot Text-to-Speech Via Interleaved Continuous Autoregressive Modeling
abstract
Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at:https://aka.ms/StreamMel.
Hui Wang 0075, Yifan Yang 0005, Shujie Liu 0001, Jinyu Li 0001, Lingwei Meng, Tie-Yan Liu, Jiaming Zhou 0001, Haoqin Sun, Yan Lu 0001
IEEE Signal Process. Lett.8
2024 Fine-Grained Disentangled Representation Learning For Multimodal Emotion Recognition
abstract
Multimodal emotion recognition (MMER) is an active research field that aims to accurately recognize human emotions by fusing multiple perceptual modalities. However, inherent heterogeneity across modalities introduces distribution gaps and information redundancy, posing significant challenges for MMER. In this paper, we propose a novel fine-grained disentangled representation learning (FDRL) framework to address these challenges. Specifically, we design modality-shared and modality-private encoders to project each modality into modality-shared and modality-private subspaces, respectively. In the shared subspace, we introduce a fine-grained alignment component to learn modality-shared representations, thus capturing modal consistency. Subsequently, we tailor a fine-grained disparity component to constrain the private subspaces, thereby learning modality-private representations and enhancing their diversity. Lastly, we introduce a fine-grained predictor component to ensure that the labels of the output representations from the encoders remain unchanged. Experimental results on the IEMOCAP dataset show that FDRL outperforms the state-of-the-art methods, achieving 78.34% and 79.44% on WAR and UAR, respectively.
Haoqin Sun, Shiwan Zhao, Xuechen Wang, Wenjia Zeng
ICASSP1
2024 Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
Haoqin Sun, Shiwan Zhao, Xiangyu Kong 0001, Xuechen Wang, Hui Wang 0075, Jiaming Zhou 0001
INTERSPEECH1
2024 Uncertainty-Aware Mean Opinion Score Prediction
Hui Wang 0075, Shiwan Zhao, Jiaming Zhou 0001, Xiguang Zheng, Haoqin Sun, Xuechen Wang
INTERSPEECH5
2023 Multi-Level Knowledge Distillation for Speech Emotion Recognition in Noisy Conditions
Yang Liu 0262, Haoqin Sun, Qingyue Wang, Zhen Zhao 0006, Xugang Lu, Longbiao Wang
INTERSPEECH2
2023 A Discriminative Feature Representation Method Based on Cascaded Attention Network With Adversarial Strategy for Speech Emotion Recognition
abstract
Currently, speech emotion recognition models still could not show satisfactory performance due to the complexity of emotions. In most of the previous studies, there is a common problem that some of the particular emotions are severely misclassified. In this article, we propose a novel framework integrating cascaded attention network and adversarial joint loss strategy for speech emotion recognition, aiming at discriminating the confusions by emphasizing more on the emotions which are difficult to be correctly classified. First, we extract log-Mels, deltas and delta-deltas of log-Mels as 3D features to effectively reduce the interference of external factors. Next, we introduce a cascaded attention network to extract effective emotional features, where spatiotemporal attention selectively locates the targeted emotional regions from the input features. In these targeted regions, the self attention with head fusion captures the long-distance dependence of temporal features. Finally, an adversarial joint loss strategy is proposed to distinguish the emotional embeddings with high similarity by the generated hard triplets in an adversarial fashion. To evaluate our proposed method, experiments are performed with the IEMOCAP, CASIA, and EMODB corpora. The experimental results demonstrate that our proposed method significantly outperforms the state-of-the-art approaches on all datasets.
Yang Liu 0262, Haoqin Sun, Wenbo Guan, Yuqi Xia, Masashi Unoki, Zhen Zhao 0006
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Discriminative Feature Representation Based on Cascaded Attention Network with Adversarial Joint Loss for Speech Emotion Recognition
Yang Liu 0262, Haoqin Sun, Wenbo Guan, Yuqi Xia, Zhen Zhao 0006
INTERSPEECH2
2022 Multi-modal speech emotion recognition using self-attention mechanism and multi-scale fusion framework
Yang Liu 0262, Haoqin Sun, Wenbo Guan, Yuqi Xia, Zhen Zhao 0006
Speech Commun.2