EDBT 2026 Demo / reviewers in the wild / expert
Chen Xu 0008
dblp:54/1474-8
· DBLP profile ↗
22ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-2408-9340ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WaveEx: Accelerating Flow Matching-based Speech Generation via Wavelet-guided ExtrapolationabstractFlow matching-based generative models offer a principled approach to modeling continuous-time dynamics in speech generation. However, inference is often computationally expensive due to repeated neural network evaluations required by ODE solvers. We propose WaveEx, a training-free and plug-in acceleration framework which replaces portions of ODE integration with wavelet-guided extrapolation. By leveraging the multi-scale structure of latent trajectories, WaveEx predicts future states directly in the frequency domain without additional model evaluations or architectural changes. WaveEx consistently accelerates inference across diverse speech generation tasks. The gains are especially pronounced in tasks like speech synthesis (up to 5.73× speedup) and music generation (2.75×), where flow matching plays a central role in alignment modeling and dense ODE integration. Even in tasks with simpler input-output mappings such as speech enhancement (4.55×) and voice conversion (2.75×), WaveEx still achieves notable acceleration, demonstrating the robustness and generalizability of the approach. These results highlight wavelet-guided extrapolation as a lightweight and broadly applicable alternative to full ODE solving for flow matching-based speech generation. Xiyan Gui, Zhengkun Ge, Yuan Ge 0001, Chang Zou, Zhikang Niu, Qixi Zheng, Chen Xu 0008, Xie Chen 0001, Tong Xiao 0001, Linfeng Zhang 0001 |
AAAI | 9 |
| 2026 | Act in Collusion: Distributed Multi-Target Backdoor Attacks in Federated LearningabstractFederated learning (FL) is widely used in Internet-of-Things (IoT) systems, but its distributed training process also exposes it to backdoor attacks. Existing studies mainly consider single-target or centralized multi-target settings, while coordinated distributed multi-target attacks remain underexplored. In practical IoT scenarios, one adversarial entity may control multiple distributed malicious clients and assign each client distinct triggers and target labels. Under this setting, existing distributed backdoor methods often fail to preserve the effectiveness of all backdoors because malicious updates conflict during aggregation. To address this issue, we propose a Distributed Multi-Target Backdoor Attack (DMBA) for FL. DMBA introduces a Backdoor Replay (BR) mechanism to reduce discrepancies among malicious gradients and a Channel-Frequency Composite Trigger (CFCT) strategy to improve trigger distinguishability and alleviate local interference. Experiments on multiple datasets show that DMBA ensures attack success rates above 80% for all implanted back-doors, whereas some baseline backdoors fall below 50% and may even approach 0. Tao Liu 0038, Dapeng Man, Jiguang Lv, Chen Xu 0008, Weiye Xi, Huanran Wang, Wu Yang 0001 |
IEEE Internet Things J. | 4 |
| 2026 | Decoupling representation learning and classifier for long-tailed adversarial training
Hengheng Xiong, Dapeng Man, Jiguang Lv, Chen Xu 0008, Fanyi Zeng, Yuyan Shi, Mingzhu Lai, Wu Yang 0001 |
Pattern Recognit. | 4 |
| 2025 | Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio EncodersabstractWeiqiao Shan, Yuang Li, Yuhao Zhang, Yingfeng Luo, Chen Xu, Xiaofeng Zhao, Long Meng, Yunfei Lu, Min Zhang, Hao Yang, Tong Xiao, JingBo Zhu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weiqiao Shan, Yuang Li, Yingfeng Luo, Chen Xu 0008, Long Meng, Yunfei Lu, Min Zhang 0042, Hao Yang 0006, Tong Xiao 0001 |
EMNLP | 5 |
| 2025 | A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech TranslationabstractSimultaneous Speech Translation (SimulST) involves generating target language text while continuously processing streaming speech input, presenting significant real-time challenges. Multi-task learning is often employed to enhance SimulST performance but introduces optimization conflicts between primary and auxiliary tasks, potentially compromising overall efficiency. The existing model-level conflict resolution methods are not well-suited for this task which exacerbates inefficiencies and leads to high GPU memory consumption. To address these challenges, we propose a Modular Gradient Conflict Mitigation (MGCM) strategy that detects conflicts at a finer-grained modular level and resolves them utilizing gradient projection. Experimental results demonstrate that MGCM significantly improves SimulST performance, particularly under medium and high latency conditions, achieving a 0.68 BLEU score gain in offline tasks. Additionally, MGCM reduces GPU memory consumption by over 95% compared to other conflict mitigation methods, establishing it as a robust solution for SimulST tasks. Yangfan Du, Jianjin Wang, Yuan Ge 0001, Chen Xu 0008, Tong Xiao 0001, Guocheng Chen |
ICASSP | 5 |
| 2025 | Collapsing Sequence-Level Data-Policy Coverage via Poisoning Attack in Offline Reinforcement LearningabstractOffline reinforcement learning (RL) heavily relies on the coverage of pre-collected data over the target policy’s distribution. Existing studies aim to improve data-policy coverage to mitigate distributional shifts, but overlook security risks from insufficient coverage, and the single-step analysis is not consistent with the multi-step decision-making nature of offline RL. To address this, we introduce the sequence-level concentrability coefficient to quantify coverage, and reveal its exponential amplification on the upper bound of estimation errors through theoretical analysis. Building on this, we propose the Collapsing Sequence-Level Data-Policy Coverage (CSDPC) poisoning attack. Considering the continuous nature of offline RL data, we convert state-action pairs into decision units, and extract representative decision patterns that capture multi-step behavior. We identify rare patterns likely to cause insufficient coverage, and poison them to reduce coverage and exacerbate distributional shifts. Experiments show that poisoning just 1% of the dataset can degrade agent performance by 90%. This finding provides new perspectives for analyzing and safeguarding the security of offline RL. Dapeng Man, Chen Xu 0008, Fanyi Zeng, Tao Liu 0038, Shucheng He, Chaoyang Gao, Wu Yang 0001 |
UAI | 3 |
| 2025 | A lightweight secret-sharing-based defense against model poisoning attacks in privacy-preserving federated learning
Hengheng Xiong, Jiguang Lv, Dapeng Man, Yukun Zhu, Tao Liu 0038, Huanran Wang, Chen Xu 0008, Wu Yang 0001 |
Comput. Commun. | 7 |
| 2025 | FLoV2T: A fine-grained malicious traffic classification method based on federated learning for AIoT
Fanyi Zeng, Chen Xu 0008, Dapeng Man, Junhui Jiang 0001, Wu Yang 0001 |
Comput. Commun. | 2 |
| 2024 | Beyond Traditional Threats: A Persistent Backdoor Attack on Federated LearningabstractBackdoors on federated learning will be diluted by subsequent benign updates. This is reflected in the significant reduction of attack success rate as iterations increase, ultimately failing. We use a new metric to quantify the degree of this weakened backdoor effect, called attack persistence. Given that research to improve this performance has not been widely noted, we propose a Full Combination Backdoor Attack (FCBA) method. It aggregates more combined trigger information for a more complete backdoor pattern in the global model. Trained backdoored global model is more resilient to benign updates, leading to a higher attack success rate on the test set. We test on three datasets and evaluate with two models across various settings. FCBA's persistence outperforms SOTA federated learning backdoor attacks. On GTSRB, post-attack 120 rounds, our attack success rate rose over 50% from baseline. The core code of our method is available at https://github.com/PhD-TaoLiu/FCBA. Tao Liu 0038, Zhu Feng, Zhiqin Yang, Chen Xu 0008, Dapeng Man, Wu Yang 0001 |
AAAI | 5 |
| 2024 | Bridging the Gaps of Both Modality and Language: Synchronous Bilingual CTC for Speech Translation and Speech RecognitionabstractIn this study, we present synchronous bilingual Connectionist Temporal Classification (CTC), an innovative framework that leverages dual CTC to bridge the gaps of both modality and language in the speech translation (ST) task. Utilizing transcript and translation as concurrent objectives for CTC, our model bridges the gap between audio and text as well as between source and target languages. Building upon the recent advances in CTC application, we develop an enhanced variant, BiL-CTC+, that establishes new state-of-the-art performances on the MuST-C ST benchmarks under resource-constrained scenarios. Intriguingly, our method also yields significant improvements in speech recognition performance, revealing the effect of cross-lingual learning on transcription and demonstrating its broad applicability. The source code is available at https://github.com/xuchennlp/S2T. Chen Xu 0008, Erfeng He, Qianqian Dong, Tong Xiao 0001, Dapeng Man, Wu Yang 0001 |
ICASSP | 1 |
| 2024 | Soft Alignment of Modality Space for End-to-End Speech TranslationabstractEnd-to-end Speech Translation (ST) aims to convert speech into target text within a unified model. The inherent differences between speech and text modalities often impede effective cross-modal and cross-lingual transfer. Existing methods typically employ hard alignment (H-Align) of individual speech and text segments, which can degrade textual representations. To address this, we introduce Soft Alignment (S-Align), using adversarial training to align the representation spaces of both modalities. S-Align creates a modality-invariant space while preserving individual modality quality. Experiments on three languages from the MuST-C dataset show S-Align outperforms H-Align across multiple tasks and offers translation capabilities on par with specialized translation models. Kaiqi Kou, Chen Xu 0008, Chunliang Zhang, Tong Xiao 0001 |
ICASSP | 4 |
| 2024 | PolyVoice: Language Models for Speech to Speech TranslationabstractWith the huge success of GPT models in natural language processing, there is a growing interest in applying language modeling approaches to speech tasks.
Currently, the dominant architecture in speech-to-speech translation (S2ST) remains the encoder-decoder paradigm, creating a need to investigate the impact of language modeling approaches in this area.
In this study, we introduce PolyVoice, a language model-based framework designed for S2ST systems. Our framework comprises three decoder-only language models: a translation language model, a duration language model, and a speech synthesis language model.
These language models employ different types of prompts to extract learned information effectively. By utilizing unsupervised semantic units, our framework can transfer semantic information across these models, making it applicable even to unwritten languages.
We evaluate our system on Chinese $\rightarrow$ English and English $\rightarrow$ Spanish language pairs. Experimental results demonstrate that \method outperforms the state-of-the-art encoder-decoder model, producing voice-cloned speech with high translation and audio quality.
Speech samples are available at https://polyvoice.github.io. Qianqian Dong, Zhiying Huang, Qi Tian 0001, Chen Xu 0008, Tom Ko, Yunlong Zhao 0004, Tang Li 0001, Xuxin Cheng, Fengpeng Yue, Ye Bai 0001, Lu Lu 0015, Zejun Ma 0001, Yuping Wang 0005, Mingxuan Wang, Yuxuan Wang 0002 |
ICLR | 4 |
| 2024 | Recent Advances in End-to-End Simultaneous Speech Translation
Yangfan Du, Erfeng He, Yingfeng Luo, Chen Xu 0008, Tong Xiao 0001 |
IJCAI | 6 |
| 2023 | Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataabstractWe present a method for introducing a text encoder into pre-trained end-to-end speech translation systems. It enhances the ability of adapting one modality (i.e., source-language speech) to another (i.e., source-language text). Thus, the speech translation model can learn from both unlabeled and labeled data, especially when the source-language text data is abundant. Beyond this, we present a denoising method to build a robust text encoder that can deal with both normal and noisy text data. Our system sets new state-of-the-arts on the MuST-C En-De, En-Fr, and LibriSpeech En-Fr tasks. Chen Xu 0008, Bojie Hu, Chunliang Zhang, Tong Xiao 0001 |
AAAI | 2 |
| 2023 | CTC-based Non-autoregressive Speech TranslationabstractChen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun, Yuhao Zhang, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao, Anxiang Ma, Jingbo Zhu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Chen Xu 0008, Qingxuan Sun, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao 0001, Anxiang Ma |
ACL (1) | 1 |
| 2023 | Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationabstractSignificant improvements in end-to-end speech translation (ST) have been achieved through the application of multi-task learning.However, the extent to which auxiliary tasks are highly consistent with the ST task, and how much this approach truly helps, have not been thoroughly studied.In this paper, we investigate the consistency between different tasks, considering different times and modules.We find that the textual encoder primarily facilitates cross-modal conversion, but the presence of noise in speech impedes the consistency between text and speech representations.Furthermore, we propose an improved multi-task learning (IMTL) approach for the ST task, which bridges the modal gap by mitigating the difference in length and representation.We conduct experiments on the MuST-C dataset.The results demonstrate that our method attains stateof-the-art results.Moreover, when additional data is used, we achieve the new SOTA result on MuST-C English to Spanish task with 20.8% of the training time required by the current SOTA method. Chen Xu 0008, Tong Xiao 0001, Chunliang Zhang |
EMNLP | 2 |
| 2023 | Recent Advances in Direct Speech-to-text TranslationabstractRecently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech translation aiming to summarize the current state-of-the-art techniques. First, we categorize the existing research work into three directions based on the main challenges --- modeling burden, data scarcity, and application issues. To tackle the problem of modeling burden, two main structures have been proposed, encoder-decoder framework (Transformer and the variants) and multitask frameworks. For the challenge of data scarcity, recent work resorts to many sophisticated techniques, such as data augmentation, pre-training, knowledge distillation, and multilingual modeling. We analyze and summarize the application issues, which include real-time, segmentation, named entity, gender bias, and code-switching. Finally, we discuss some promising directions for future work. Chen Xu 0008, Rong Ye, Qianqian Dong, Chengqi Zhao, Tom Ko, Mingxuan Wang, Tong Xiao 0001 |
IJCAI | 1 |
| 2023 | Information Magnitude Based Dynamic Sub-sampling for Speech-to-text
Chenghao Gao, Kaiqi Kou, Chen Xu 0008, Tong Xiao 0001 |
INTERSPEECH | 4 |
| 2021 | Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation EncodersabstractChen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang, Shen Huang, Qi Ju, Tong Xiao, Jingbo Zhu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Xu 0008, Bojie Hu, Yanyang Li, Shen Huang, Qi Ju 0002, Tong Xiao 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | Chinese Poetry Generation with Metrical Constraints
Yingfeng Luo, Changliang Li, Canan Huang, Chen Xu 0008, Binghao Wei, Tong Xiao 0001 |
NLPCC (1) | 4 |
| 2020 | Dynamic Curriculum Learning for Low-Resource Neural Machine TranslationabstractLarge amounts of data has made neural machine translation (NMT) a big success in recent years.But it is still a challenge if we train these models on small-scale corpora.In this case, the way of using data appears to be more important.Here, we investigate the effective use of training data for low-resource NMT.In particular, we propose a dynamic curriculum learning (DCL) method to reorder training samples in training.Unlike previous work, we do not use a static scoring function for reordering.Instead, the order of training samples is dynamically determined in two ways -loss decline and model competence.This eases training by highlighting easy samples that the current model has enough competence to learn.We test our DCL method in a Transformerbased system.Experimental results show that DCL outperforms several strong baselines on three low-resource machine translation benchmarks and different sized data of WMT'16 En-De. Chen Xu 0008, Bojie Hu, Yufan Jiang, Zeyang Wang, Shen Huang, Qi Ju 0002, Tong Xiao 0001 |
COLING | 1 |
| 2019 | Analysis of Back-Translation Methods for Low-Resource Neural Machine Translation
Nuo Xu 0010, Yinqiao Li, Chen Xu 0008, Yanyang Li, Tong Xiao 0001 |
NLPCC (2) | 3 |