Xueyuan Chen

dblp:126/2361 · DBLP profile ↗
← Back
27ranked-venue papers
10as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DEALSD: A deep edge assisted line segment detector
Zhongyi Sha, Baojiang Zhong, Xueyuan Chen, Zikai Wang 0007
Expert Syst. Appl.3
2025 Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
abstract
With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features.Each approach has demonstrated strong capabilities in audiorelated processing tasks.However, the performance gap between these two paradigms has not been thoroughly explored.To address this gap, we present a fair comparison of selfsupervised learning (SSL)-based discrete and continuous features under the same experimental settings.We evaluate their performance across six spoken language understandingrelated tasks using both small and large-scale LLMs (Qwen1.5-0.5B and Llama3.1-8B).We further conduct in-depth analyses, including efficient comparison, SSL layer analysis, LLM layer analysis, and robustness comparison.Our findings reveal that continuous features generally outperform discrete tokens in various tasks.Each speech processing method exhibits distinct characteristics and patterns in how it learns and processes speech information.We hope our findings will provide valuable insights to advance spoken language understanding in SpeechLLMs.
Dingdong Wang, Junan Li, Dongchao Yang, Xueyuan Chen, Helen M. Meng
EMNLP5
2025 Integrating Potential Pronunciations for Enhanced Mispronunciation Detection and Diagnosis Ability in LLMs
abstract
Large Language Models (LLMs) have exhibited significant potentials across various tasks. However, how to leverage the power of LLMs in the mispronunciation detection and diagnosis (MDD) task is still under-explored. In this paper, we propose a PP-ATP model, which integrates potential pronunciations covering common mispronunciations into the prompt part of LLMs, to enhance the MDD ability of LLMs in second language (L2) English. Specifically, the proposed PP-ATP model is composed of an audio encoder, an LLM decoder, and an adapter. Taking speech representations from the audio encoder as the audio prompt and reference sentence with canonical and potential pronunciations as text prompt, the LLM decoder is adapted to predict the actual pronunciation in the given L2 speech. Experiments show that our PP-ATP model achieves new state-of-the-art (SOTA) performance in MDD on CU-CHLOE corpus, confirming the effectiveness of potential pronunciation integration.
Minglin Wu, Xueyuan Chen, Helen M. Meng
ICASSP3
2025 ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
abstract
Recent advancements in audio language models have underscored the pivotal role of audio tokenization, which converts audio signals into discrete tokens, thereby facilitating the application of language model architectures to the audio domain. In this study, we introduce ALMTokenizer, a novel low-bitrate and semantically rich audio codec tokenizer for audio language models. Prior methods, such as Encodec, typically encode individual audio frames into discrete tokens without considering the use of context information across frames. Unlike these methods, we introduce a novel query-based compression strategy to capture holistic information with a set of learnable query tokens by explicitly modeling the context information across frames. This design not only enables the codec model to capture more semantic information but also encodes the audio signal with fewer token sequences. Additionally, to enhance the semantic information in audio codec models, we introduce the following: (1) A masked autoencoder (MAE) loss, (2) Vector quantization based on semantic priors, and (3) An autoregressive (AR) prediction loss. As a result, ALMTokenizer achieves competitive reconstruction performance relative to state-of-the-art approaches while operating at a lower bitrate. Within the same audio language model framework, ALMTokenizer outperforms previous tokenizers in audio understanding and generation tasks.[https://dongchaoyang.top/ALMTokenizer/]
Dongchao Yang, Songxiang Liu, Haohan Guo, Jiankun Zhao, Helin Wang, Zeqian Ju, Xueyuan Chen, Xu Tan 0003, Xixin Wu, Helen M. Meng
ICML9
2025 DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
Xueyuan Chen, Dongchao Yang, Minglin Wu, Xixin Wu, Zhiyong Wu 0001, Helen M. Meng
INTERSPEECH1
2025 Molecular graph contrastive learning with line graph
Xueyuan Chen, Shangzhe Li, Ruomei Liu, Bowen Shi 0001, Junran Wu, Ke Xu 0001
Pattern Recognit.1
2024 NC2D: Novel Class Discovery for Node Classification
abstract
Novel Class Discovery (NCD) involves identifying new categories within unlabeled data by utilizing knowledge acquired from previously established categories. However, existing NCD methods often struggle to maintain a balance between the performance of old and new categories. Discovering unlabeled new categories in a class-incremental way is more practical but also more challenging, as it is frequently hindered by either catastrophic forgetting of old categories or an inability to learn new ones. Furthermore, the implementation of NCD on continuously scalable graph-structured data remains an under-explored area. In response to these challenges, we introduce for the first time a more practical NCD scenario for node classification (i.e., NC-NCD), and propose a novel self-training framework with prototype replay and distillation called SWORD, adopted to our NC-NCD setting. Our approach enables the model to cluster unlabeled new category nodes after learning labeled nodes while preserving performance on old categories without reliance on old category nodes. SWORD achieves this by employing a self-training strategy to learn new categories and preventing the forgetting of old categories through the joint use of feature prototypes and knowledge distillation. Extensive experiments on four common benchmarks demonstrate the superiority of SWORD over other state-of-the-art methods.
Xueyuan Chen, Ruomei Liu, Bowen Shi 0001, Junran Wu, Ke Xu 0001
CIKM2
2024 Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
abstract
Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in complex, noisy acoustic environments. To address these challenges, we propose a novel multi-modal framework to utilize visual information, e.g., lip movements, in DSR as extra clues for reconstructing the highly abnormal pronunciations. The multi-modal framework consists of: (i) a multi-modal encoder to extract robust phoneme embeddings from dysarthric speech with auxiliary visual features; (ii) a variance adaptor to infer the normal phoneme duration and pitch contour from the extracted phoneme embeddings; (iii) a speaker encoder to encode the speaker’s voice characteristics; and (iv) a mel-decoder to generate the reconstructed mel-spectrogram based on the extracted phoneme embeddings, prosodic features and speaker embeddings. Both objective and subjective evaluations conducted on the commonly used UASpeech corpus show that our proposed approach can achieve significant improvements over baseline systems in terms of speech intelligibility and naturalness, especially for the speakers with more severe symptoms. Compared with original dysarthric speech, the reconstructed speech achieves 42.1% absolute word error rate reduction for patients with more severe dysarthria levels.1
Xueyuan Chen, Yuejiao Wang, Xixin Wu, Disong Wang, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng
ICASSP1
2024 Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech Synthesis
abstract
The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-supervised style enhancing method with VQ-VAE-based pre-training for expressive audiobook speech synthesis. Firstly, a text style encoder is pre-trained with a large amount of unlabeled text-only data. Secondly, a spectrogram style extractor based on VQ-VAE is pre-trained in a self-supervised manner, with plenty of audio data that covers complex style variations. Then a novel architecture with two encoder-decoder paths is specially designed to model the pronunciation and high-level style expressiveness respectively, with the guidance of the style extractor. Both objective and subjective evaluations demonstrate that our proposed method can effectively improve the naturalness and expressiveness of the synthesized speech in audiobook synthesis especially for the role and out-of-domain scenarios.1
Xueyuan Chen, Xi Wang 0016, Shaofei Zhang, Lei He 0005, Zhiyong Wu 0001, Xixin Wu, Helen M. Meng
ICASSP1
2024 Contrast-Guided Wireframe Parsing
abstract
Existing deep wireframe parsing methods typically focus on the semantic saliency of scene structural lines without paying particular attention to their visual saliency. As a result, these methods often face the challenge of multiple responses to proximate line segments or erroneous responses to non-structural elements like shadows. To address this fundamental issue, a novel Contrast Guidance Module (CGM) is proposed. In the CGM, a low-level image attribute, i.e., the image contrast, is leveraged to measure the visual saliency of structural lines. The CGM augments feature maps in CNN networks, subtly balancing the interpretation of line segments with their contextual significance in the image. This approach not only refines detection accuracy but also enriches the understanding of spatial geometry. Extensive experiments conducted on benchmark datasets have shown that our proposed CGM consistently outperforms the current state-of-the-art methods.
Xueyuan Chen, Baojiang Zhong
ICIP1
2024 Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
abstract
Audio-visual target speech extraction (AV-TSE) is one of the enabling technologies in robotics and many audiovisual applications. One of the challenges of AV-TSE is how to effectively utilize audio-visual synchronization information in the process. AV-HuBERT can be a useful pre-trained model for lip-reading, which has not been adopted by AV-TSE. In this paper, we would like to explore the way to integrate a pretrained AV-HuBERT into our AV-TSE system. We have good reasons to expect an improved performance. To benefit from the inter and intra-modality correlations, we also propose a novel Mask-And-Recover (MAR) strategy for self-supervised learning. The experimental results on the VoxCeleb2 dataset show that our proposed model outperforms the baselines both in terms of subjective and objective metrics, suggesting that the pre-trained AV-HuBERT model provides more informative visual cues for target speech extraction. Furthermore, through a comparative study, we confirm that the proposed Mask-And-Recover strategy is significantly effective.
Xueyuan Chen, Xixin Wu, Haizhou Li 0001, Helen M. Meng
IJCNN2
2024 CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
Xueyuan Chen, Dongchao Yang, Dingdong Wang, Xixin Wu, Zhiyong Wu 0001, Helen M. Meng
INTERSPEECH1
2024 SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
Dongchao Yang, Dingdong Wang, Haohan Guo, Xueyuan Chen, Xixin Wu, Helen M. Meng
INTERSPEECH4
2024 Uncovering Capabilities of Model Pruning in Graph Contrastive Learning
abstract
Graph contrastive learning has achieved great success in pre-training graph neural networks without ground-truth labels. Leading graph contrastive learning follows the classical scheme of contrastive learning, forcing model to identify the essential information from augmented views. However, general augmented views are produced via random corruption or learning, which inevitably leads to semantics alteration. Although domain knowledge guided augmentations alleviate this issue, the generated views are domain specific and undermine the generalization. In this work, motivated by the firm representation ability of sparse model from pruning, we reformulate the problem of graph contrastive learning via contrasting different model versions rather than augmented views. We first theoretically reveal the superiority of model pruning in contrast to data augmentations. In practice, we take original graph as input and dynamically generate a perturbed graph encoder to contrast with the original encoder by pruning its transformation weights. Furthermore, considering the integrity of node embedding in our method, we are capable of developing a local contrastive loss to tackle the hard negative samples that disturb the model training. We extensively validate our method on various benchmarks regarding graph classification via unsupervised and transfer learning. Compared to the state-of-the-art (SOTA) works, better performance can always be obtained by the proposed method.
Junran Wu, Xueyuan Chen, Shangzhe Li
ACM Multimedia2
2024 VRDistill: Vote Refinement Distillation for Efficient Indoor 3D Object Detection
abstract
Recently, indoor 3D object detection has shown impressive progress. However, these improvements have come at the cost of increased memory consumption and longer inference times, making it difficult to apply these methods in practical scenarios. To address this issue, knowledge distillation has emerged as a promising technique for model acceleration. In this paper, we propose the VRDistill framework, the first knowledge distillation framework designed for efficient indoor 3D object detection. Our VRDistill framework includes a refinement module and a soft foreground mask operation to enhance the quality of the distillation. The refinement module utilizes trainable layers to improve the quality of the teacher's votes, while the soft foreground mask operation focuses on foreground votes, further enhancing the distillation performance. Comprehensive experiments on the ScanNet and SUN-RGBD datasets demonstrate the effectiveness and generalization ability of our VRDistill framework.
Ze Yuan, Jinyang Guo 0002, Dakai An, Junran Wu, Xueyuan Chen, Ke Xu 0001
ACM Multimedia7
2024 A unified efficient deep image compression framework and its application on human-centric Task
Xueyuan Chen, Guo Lu
Multim. Tools Appl.1
2024 MPG-LSD: A high-quality line segment detector based on multi-scale perceptual grouping
Baojiang Zhong, Xueyuan Chen, Hangjia Zheng
Pattern Recognit.3
2024 Simultaneous Retrieval of Land Surface Temperature and Soil Moisture Using Multichannel Passive Microwave Data
abstract
Land surface temperature (LST) and soil moisture (SM) are two important parameters in land surface ecosystem at regional and global scale. The accurate acquisition of LST and SM can benefit various fields, including agriculture and climate which are closely related to human life. The independent retrievals of LST and SM from passive microwave observations are mutually restricted and highly dependent on auxiliary data. To solve this problem, a simulations retrieval method of LST and SM was proposed based on the characteristics of multi-frequency and dual-polarization. The simultaneous solution of LST and SM was realized by approximating and correcting the radiative transfer equation (RTE). The performance of the proposed method was evaluated using simulated data, resulting in a root mean square error (RMSE) of approximately 1.63 K and 0.063 m3/m3. This method was further used to retrieve LST and SM from AMSR-E observations. The retrieved LST was compared to the MODIS land surface temperature product under clear-sky, with RMSE of 5.68 K. The retrieved LST was validated using the ISD air temperature under cloudy-sky, with RMSE of 4.29 K. The accuracy of retrieved LST changes with the variation of vegetation. The retrieved SM was evaluated using the CCI soil moisture product and in-situ observations. The result shows that the accuracy ranges from 0.0157 to 0.1115 m3/m3with the change of vegetation. This study gives a feasible method to retrieve LST and SM simultaneously with reasonable accuracy.
Xiao-Jing Han, Na Yao, Pei Leng, Wenjing Han, Xueyuan Chen
IEEE Trans. Geosci. Remote. Sens.6
2024 Forecasting Turning Points in Stock Price by Integrating Chart Similarity and Multipersistence
abstract
Forecasting financial data plays a crucial role in financial market. Relying solely on prices or price trends as prediction targets often leads to a vast of invalid transactions. As a result, researchers have increasingly turned their attention to turning points as the prediction target. Surprisingly, existing methods have largely overlooked the role of technical charts, despite turning points being closely related to the technical charts. Recently, several researchers have attempted to utilize chart information via converting price sequences into images for turning point forecasting, but robustness and convergence problems arise. To address these challenges and enhance the turning point predictions, this article introduces a new method known as MPCNet. Specifically, we first transform the price series into a graph structure using chart similarity to robustly extract valuable information from technical charts. Additionally, we introduce the multipersistence topology tool to accurately predict stock turning points and provide convergence guarantee. Experimental results demonstrate the significant superiority of our proposed model over existing methods. Furthermore, based on additional performance evaluations using real stock data, MPCNet consistently achieves the highest average return during the transaction backtesting period. Meanwhile, we provide empirical validation of robustness and theoretical analysis to confirm its convergence, establishing it as a superior tool for financial forecasting.
Shangzhe Li, Yingke Liu, Xueyuan Chen, Junran Wu, Ke Xu 0001
IEEE Trans. Knowl. Data Eng.3
2023 SEGA: Structural Entropy Guided Anchor View for Graph Contrastive Learning
abstract
In contrastive learning, the choice of "view" controls the information that the representation captures and influences the performance of the model. However, leading graph contrastive learning methods generally produce views via random corruption or learning, which could lead to the loss of essential information and alteration of semantic information. An anchor view that maintains the essential information of input graphs for contrastive learning has been hardly investigated. In this paper, based on the theory of graph information bottleneck, we deduce the definition of this anchor view; put differently, the anchor view with essential information of input graph is supposed to have the minimal structural uncertainty. Furthermore, guided by structural entropy, we implement the anchor view, termed SEGA, for graph contrastive learning. We extensively validate the proposed anchor view on various benchmarks regarding graph classification under unsupervised, semi-supervised, and transfer learning and achieve significant performance boosts compared to the state-of-the-art methods.
Junran Wu, Xueyuan Chen, Bowen Shi 0001, Shangzhe Li, Ke Xu 0001
ICML2
2023 ReGRL: An Informative Graph Representation via Hierarchical Recursive Learning for Legal Case Recommendation
abstract
Legal Case Recommendation (LCR) is to find out the documents that are most similar to the input case from the judicial point of view. Since the legal documents are long texts and have strong legal attributes, the traditional recommendation method based on text similarity is difficult to accurately understand the legal documents, resulting in poor effect of LCR. To address this problem, we propose Recursive Graph Representation Learning (ReGRL) to hierarchically learn the information in the graph, and obtain a more informative graph representation to accurately understand the case. ReGRL captures nodes, edge, and community information at different levels to integrate information of different granularity. To achieve this, ReGRL performs top-down graph decomposition and bottom-up graph encoding in a recursive form, which allows ReGRL to flexibly control the depth of the learning layers and provide accurate case representations for LCR. Experimental results show that ReGRL can not only generate a good representation, but has a better performance compared to text representation methods for LCR. In addition, we also performed different experiments to analyze the principle of ReGRL and verify the effectiveness of recursive procedures.
Xueyuan Chen, Xiao Wei 0002, Hang Yu 0006, Xiangfeng Luo
IJCNN1
2022 Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis
abstract
Naturalness and expressiveness are crucial for audiobook speech synthesis, but now are limited by the averaged global-scale speaking style representation. In this paper, we propose an unsupervised multi-scale context-sensitive text-to-speech model for audiobooks. A multi-scale hierarchical context encoder is specially designed to predict both global-scale context style embedding and local-scale context style embedding from a wider context of input text in a hierarchical manner. Likewise, a multi-scale reference encoder is introduced to extract reference style embeddings at both global and local scales from the reference speech, which is used to guide the prediction of speaking styles. On top of these, a bi-reference attention mechanism is used to align both local-scale reference style embedding sequence and local-scale context style embedding sequence with corresponding phoneme embedding sequence. Both objective and subjective experiment results on a real-world multi-speaker Mandarin novel audio dataset demonstrate the excellent performance of our proposed method over all baselines in terms of naturalness and expressiveness of the synthesized speech.
Xueyuan Chen, Shun Lei, Zhiyong Wu 0001, Weifeng Zhao, Helen M. Meng
COLING1
2022 A Character-Level Span-Based Model for Mandarin Prosodic Structure Prediction
abstract
The accuracy of prosodic structure prediction is crucial to the naturalness of synthesized speech in Mandarin text-to-speech system, but now is limited by widely-used sequence-to-sequence framework and error accumulation from previous word segmentation results. In this paper, we propose a span-based Mandarin prosodic structure prediction model to obtain an optimal prosodic structure tree, which can be converted to corresponding prosodic label sequence. Instead of the prerequisite for word segmentation, rich linguistic features are provided by Chinese character-level BERT and sent to encoder with self-attention architecture. On top of this, span representation and label scoring are used to describe all possible prosodic structure trees, of which each tree has its corresponding score. To find the optimal tree with the highest score for a given sentence, a bottom-up CKYstyle algorithm is further used. The proposed method can predict prosodic labels of different levels at the same time and accomplish the process directly from Chinese characters in an end-to-end manner. Experiment results on two real-world datasets demonstrate the excellent performance of our span-based method over all sequence-to-sequence baseline approaches.
Xueyuan Chen, Changhe Song, Yixuan Zhou 0002, Zhiyong Wu 0001, Changbin Chen, Zhongqin Wu, Helen M. Meng
ICASSP1
2022 Structural Entropy Guided Graph Hierarchical Pooling
abstract
Following the success of convolution on non-Euclidean space, the corresponding pooling approaches have also been validated on various tasks regarding graphs. However, because of the fixed compression ratio and stepwise pooling design, these hierarchical pooling methods still suffer from local structure damage and suboptimal problem. In this work, inspired by structural entropy, we propose a hierarchical pooling approach, SEP, to tackle the two issues. Specifically, without assigning the layer-specific compression ratio, a global optimization algorithm is designed to generate the cluster assignment matrices for pooling at once. Then, we present an illustration of the local structure damage from previous methods in reconstruction of ring and grid synthetic graphs. In addition to SEP, we further design two classification models, SEP-G and SEP-N for graph classification and node classification, respectively. The results show that SEP outperforms state-of-the-art graph pooling methods on graph classification benchmarks and obtains superior performance on node classifications.
Junran Wu, Xueyuan Chen, Ke Xu 0001, Shangzhe Li
ICML2
2022 Price graphs: Utilizing the structural information of financial time series for stock prediction
Junran Wu, Ke Xu 0001, Xueyuan Chen, Shangzhe Li, Jichang Zhao
Inf. Sci.3
2021 Retrieval of Land Surface Temperature and Soil Moisture from Passive Microwave Observations
abstract
Land surface temperature (LST) and soil moisture (SM) are two important parameters in land surface ecosystem at regional and global scale. The accurate acquisition of LST and SM can benefit various fields, including agriculture and climate which are closely related to human life. This study proposed a simultaneous retrieval method of LST and SM based on the approximate and correction of passive microwave radiation transfer equation. Compared to LST and SM in simulated database, the accuracy of retrieved LST is approximately 1.63 K and the accuracy of retrieved SM is about 0.063 m3/m3.
Xiao-Jing Han, Huajun Tang, Zhao-Liang Li, Sibo Duan, Pei Leng, Yongchang Wu, Xueyuan Chen
IGARSS7
2021 Validating static warnings via testing code fragments
abstract
Static analysis is an important approach for finding bugs and vulnerabilities in software. However, inspecting and confirming static warnings are challenging and time-consuming. In this paper, we present a novel solution that automatically generates test cases based on static warnings to validate true and false positives. We designed a syntactic patching algorithm that can generate syntactically valid, semantic preserving executable code fragments from static warnings. We developed a build and testing system to automatically test code fragments using fuzzers, KLEE and Valgrind. We evaluated our techniques using 12 real-world C projects and 1955 warnings from two commercial static analysis tools. We successfully built 68.5% code fragments and generated 1003 test cases. Through automatic testing, we identified 48 true positives and 27 false positives, and 205 likely false positives. We matched 4 CVE and real-world bugs using Helium, and they are only triggered by our tool but not other baseline tools. We found that testing code fragments is scalable and useful; it can trigger bugs that testing entire programs or testing procedures failed to trigger.
Ashwin Kallingal Joshy, Xueyuan Chen, Benjamin Steenhoek, Wei Le
ISSTA2