Xianhong Chen

dblp:209/6969 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TSAD: Trace-Semantic Adaptive Disentanglement for Detecting and Grounding Multi-Modal Media Manipulation
abstract
The proliferation of sophisticated multi-modal misinformation necessitates detectors capable of perceiving both subtle manipulation traces and high-level semantic inconsistencies. However, a pivotal challenge remains: existing methods often project features into a unified latent space, leading to a “semantic crowding” effect where robust semantic signals drown out fragile, low-level forgery artifacts. To resolve this intrinsic dimensionality mismatch, we propose TSAD (Trace-Semantic Adaptive Disentanglement), a novel framework designed to detect and ground multi-modal media manipulation. At its core, TSAD introduces a Trace-Semantic Disentanglement (TSD) module that casts feature extraction as a dual-manifold projection problem, enabling explicit separation of heterogeneous representations. This mechanism orthogonally separates input data into a Structural Anomaly Space, aggregating local neighborhood pixel correlations to capture splicing boundaries, and a Semantic Fidelity Space that preserves global content identity. Furthermore, to handle the varying complexity of forgery patterns, we propose a Granularity-Adaptive Calibration (GAC) module that dynamically learns instance-specific gating weights to fuse these disentangled cues. Extensive experiments on the large-scale DGM4 benchmark validate the effectiveness and robustness of TSAD, showing clear advantages over existing state-of-the-art approaches in both binary and fine-grained evaluation settings.
Hujin Peng, Xianhong Chen, Yixi Tian, Menghan Liang, Xiaofeng Wang 0002, Tan Deng, Wanwei Jiang
ICMR5
2026 CCASNet: Criss-Cross Attention Enhanced Network with Dual-Channel Spatial Modeling for Medical Image Segmentation
Xianhong Chen, Hujin Peng, Junyuan Gong, Zeyun Liu, Tan Deng
MMM (1)3
2026 Global diversity-based entropy minimization for source-free unsupervised domain adaptation of speaker verification
Xianhong Chen, Zhuorui Li, Mao-shen Jia
Speech Commun.1
2026 Multi-Level Interaction for Emotion Recognition From Unaligned Speech and Text
abstract
In multimodal emotion recognition, the diversity and temporal unalignment of speech and text modalities pose significant challenges for effective fusion. To address this issue, Multi-level Interaction for Emotion Recognition from Unaligned Speech and Text (MIUST) is proposed. Inspired by the hierarchical and multi-level integration process of human emotion cognition, the MIUST framework is designed to be consisted of a unimodal emotion recognition module, a multi-level cross-modal fusion module incorporating both coarse-grained feature learning and fine-grained modal fusion (FGMF), and an emotion classification module that synthesizes information from all stages. This multi-level branched fusion architecture more closely mirrors the human emotion understanding process. In addition, the FGMF can achieve cross-granularity fusion of speech and text. It learns emotional correlations between individual speech frames and textual words, thereby eliminating the need for explicit temporal alignment between modalities. Experimental results on the IEMOCAP and MELD datasets demonstrate the effectiveness of MIUST, achieving 77.78% weighted accuracy (WA), 78.79% unweighted accuracy (UA), and 77.82% weighted F1-score (W-F1) on IEMOCAP, outperforming existing state-of-the-art methods. These results validate that MIUST effectively improves multimodal emotion recognition performance by leveraging multi-level and cross-modal feature interaction. Our code is available athttps://github.com/HANLM15/MIUST.
Lingmin Han, Xianhong Chen, Mao-shen Jia, Changchun Bao
IEEE Trans. Affect. Comput.2
2024 Target Speaker Extraction by Directly Exploiting Contextual Information in the Time-Frequency Domain
abstract
In target speaker extraction, many studies rely on the speaker embedding which is obtained from an enrollment of the target speaker and employed as the guidance. However, solely using speaker embedding may not fully utilize the contextual information contained in the enrollment. In this paper, we directly exploit this contextual information in the time-frequency (T-F) domain. Specifically, the T-F representations of the enrollment and the mixed signal are interacted to compute the weighting matrices through an attention mechanism. These weighting matrices reflect the similarity among different frames of the T-F representations and are further employed to obtain the consistent T-F representations of the enrollment. These consistent representations are served as the guidance, allowing for better exploitation of the contextual information. Furthermore, the proposed method achieves the state-of-the-art performance on the benchmark dataset and shows its effectiveness in the complex scenarios.
Changchun Bao, Xianhong Chen
ICASSP4
2024 Coarse-to-Fine Target Speaker Extraction Based on Contextual Information Exploitation
abstract
To address the cocktail party problem, the target speaker extraction (TSE) has received increasing attention recently. Typically, the TSE is explored in two scenarios. The first scenario is a specific one, where the target speaker is present and the signal received by the microphone contains at least two speakers. The second scenario is a universal one, where the target speaker may be present or absent and the received signal may contain one or multiple speakers. Numerous TSE studies utilize the target speaker's embedding to guide the extraction. However, solely utilizing this embedding may not fully leverage the contextual information within the enrollment. To address this limitation, a novel approach that directly exploits the contextual information in the time-frequency (T-F) domain was proposed. This paper improves this approach by integrating our previously proposed coarse-to-fine framework. For the specific scenario, an interaction block is employed to facilitate direct interaction between the T-F representations of the enrollment and received signal. This direct interaction leads to the consistent representation of the enrollment that serves as guidance for the coarse extraction. Afterwards, the T-F representation of the coarsely extracted signal is utilized to guide the refining extraction. The residual representation obtained during the refining extraction increases the extraction precision. Besides, this paper explores an undisturbed universal scenario where the noise and reverberation are not considered. A two-level decision-making scheme is devised to generalize our proposed method for this undisturbed universal scenario. The proposed method achieves high performance and is proven effective for both scenarios.
Changchun Bao, Xianhong Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Monaural Speech Separation Method Based on Recurrent Attention with Parallel Branches
Changchun Bao, Xianhong Chen
INTERSPEECH4
2023 Coarse-to-fine speech separation method in the time-frequency domain
Changchun Bao, Xianhong Chen
Speech Commun.3
2021 Phoneme-Unit-Specific Time-Delay Neural Network for Speaker Verification
abstract
Variations of speech content increase the difficulty of speaker verification. In this paper, to alleviate the negative effect of the variations, phoneme-unit-specific time-delay neural network (PUSTDNN) is proposed and applied to the state-of-the-art x-vector system. It models each phoneme unit with an individual time-delay neural network (TDNN). That is to say, each TDNN mainly deals with a phoneme unit. Compared with handling all phoneme units together, when handling a phoneme unit, a TDNN can extract more discriminative speaker information, thus improving the system performance. Two realizations of the PUSTDNN are proposed. The first one can retain speech temporal information. The second one further combines all the TDNNs in a PUSTDNN into a larger TDNN to reduce computational complexity. To avoid model overfitting, the phoneme units are obtained by clustering phonemes based on the phonetic knowledge and phonetic sparsity degree. The PUSTDNN is also compared with two other techniques, i.e., phonetic vector and multitask. Experiments on the Fisher, NIST SRE10, and VoxCeleb datasets show that the phonetic vector technique is most robust to the phoneme unit recognition accuracy. When the accuracy is high enough, the multitask performs better than the phonetic vector, and the PUSTDNN performs best and can achieve over 10% relative improvement compared with the x-vector baseline.
Xianhong Chen, Changchun Bao
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 THUEE System for NIST SRE19 CTS Challenge
Ruyun Li, Tianyu Liang, Yi Liu 0049, Yangcheng Wu, Can Xu 0003, Xianhong Chen, Weiqiang Zhang 0001, Shouyi Yin, Liang He 0003
INTERSPEECH9
2019 Multi-objective Optimization Training of PLDA for Speaker Verification
abstract
Most current state-of-the-art text-independent speaker verifi-cation systems take probabilistic linear discriminant analysis (PLDA) as their backend classifiers. The parameters of PL-DA are often estimated by maximizing the objective function, which focuses on increasing the value of log-likelihood function, but ignoring the distinction between speakers. In order to better distinguish speakers, we propose a multi-objective optimization training for PLDA. Experiment results show that the proposed method has more than 10% relative performance improvement in both EER and MinDCF on the NIST SRE14 i-vector challenge dataset, and about 20% relative performance improvement in EER on the MCE18 dataset.
Liang He 0003, Xianhong Chen, Can Xu 0003, Jia Liu 0001
ICASSP2
2019 Distance-Dependent Metric Learning
abstract
In most existing metric learning methods, the data pairs are equally treated without considering their diversity. In fact, for pairs with different distances, the main purpose of the metric acted on them has some differences. If the pairs have smaller distance, metric should focus more on expanding the negative pairs, which are more easily misjudged. While if the pairs have larger distance, metric should focus more on shrinking the positive pairs. However, most metric learning methods neglect these differences. In this letter, we propose a distance-dependent metric learning (D2ML) method. It partitions data pairs into different clusters according to the ℓ2distance between them. Each cluster is associated with a Mahalanobis metric that learns the pairs' distance. This not only allows us to make each metric more targeted and adapt to the data diversity flexibly, but also avoids the problem of computing the distance between points assigned to different clusters, which happens in some local metric learning methods. D2ML is further extended to D3ML to embrace the nonlinear capacity of neural network. Experiments on UCI datasets and speaker recognition i-vector machine learning challenge show that the proposed methods are superior to other metric learning methods.
Xianhong Chen, Liang He 0003, Can Xu 0003, Jia Liu 0001
IEEE Signal Process. Lett.1
2018 Local Pairwise Linear Discriminant Analysis for Speaker Verification
abstract
Linear discriminant analysis-probabilistic linear discriminant analysis (LDA-PLDA) is a standard and effective backend in the field of speaker verification. The object of LDA is to perform dimensionality reduction while minimizing within-class covariance and maximizing between-class covariance. For a target class (or speaker), our task is to make a binary decision about whether a test utterance is from a specific target speaker. Generally, the nontarget test utterances that are close to the target speaker are easily misjudged. Inspired by this idea, we propose a local pairwise linear discriminant analysis (LPLDA) algorithm. This new method focuses on maximizing the local pairwise covariance, which represents the local structure between the target class samples and neighboring nontarget class samples, instead of the between-class covariance, which represents the global structure of the data. Experiments on the NIST SRE 2010, 2014, and 2016 database show that, the proposed LPLDA-PLDA backend has significant performance improvements over the LDA-PLDA backend.
Liang He 0003, Xianhong Chen, Can Xu 0003, Jia Liu 0001, Michael T. Johnson
IEEE Signal Process. Lett.2