EDBT 2026 Demo / reviewers in the wild / expert
Xuanjun Chen
dblp:277/5190 · also Xuan-Jun Chen
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Generalized Source Tracing for Codec-Based Deepfake SpeechabstractRecent attempts at source tracing for codecbased deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using simulated CoSG data while maintaining strong performance on real CoSG-generated audio remains an open challenge. In this paper, we show that models trained solely on codec-resynthesized data tend to overfit to non-speech regions and struggle to generalize to unseen content. To mitigate these challenges, we introduce the Semantic-Acoustic Source Tracing Network (SASTNet), which jointly leverages Whisper for semantic feature encoding and Wav2vec2 with AudioMAE for acoustic feature encoding. Our proposed SASTNet achieves state-of-theart performance on the CoSG test set of CodecFake+ dataset, demonstrating its effectiveness for reliable source tracing. I-Ming Lin, Xuanjun Chen, Lin Zhang 0054, Hung-yi Lee, Jyh-Shing Roger Jang |
ASRU | 2 |
| 2025 | Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech EnhancementabstractIn multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of full-band and sub-band spectral and spatial features. However, these approaches face limitations in fully modeling complex temporal dependencies, especially in dynamic acoustic environments. To overcome these challenges, we modify the current advanced model McNet by introducing an improved version of Mamba, a state-space model, and further propose MCMamba. MCMamba has been completely reengineered to integrate full-band and narrow-band spatial information with sub-band and full-band spectral features, providing a more comprehensive approach to modeling spatial and spectral information. Our experimental results demonstrate that MCMamba significantly improves the modeling of spatial and spectral features in multichannel speech enhancement, outperforming McNet and achieving very promis- ing performance on the CHiME-3 dataset. Additionally, we find that Mamba performs exceptionally well in modeling spectral information. Wenze Ren, Yi-Cheng Lin, Xuanjun Chen, Rong Chao, Kuo-Hsuan Hung, You-Jin Li, Wen-Yuan Ting, Hsin-Min Wang, Yu Tsao 0001 |
ICASSP | 4 |
| 2025 | Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 TasksabstractMultimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication gaps and facilitating more intuitive interactions. However, the absence of a comprehensive evaluation benchmark poses a significant challenge. We present Dynamic-SUPERB Phase-2, an open and evolving benchmark for the comprehensive evaluation of instruction-based universal speech models. Building upon the first generation, this second version incorporates 125 new tasks contributed collaboratively by the global research community, expanding the benchmark to a total of 180 tasks, making it the largest benchmark for speech and audio evaluation. While the first generation of Dynamic-SUPERB was limited to classification tasks, Dynamic-SUPERB Phase-2 broadens its evaluation capabilities by introducing a wide array of novel and diverse tasks, including regression and sequence generation, across speech, music, and environmental audio. Evaluation results show that no model performed well universally. SALMONN-13B excelled in English ASR and Qwen2-Audio-7B-Instruct showed high accuracy in emotion recognition, but current models still require further innovations to handle a broader range of tasks. We open-source all task data and the evaluation pipeline at https://github.com/dynamic-superb/dynamic-superb. Chien-Yu Huang, Wei-Chih Chen, Shu-Wen Yang, Andy T. Liu, Chen-An Li, Yu-Xiang Lin, Wei-Cheng Tseng, Anuj Diwan, Yi-Jen Shih, Jiatong Shi, Chih-Kai Yang, Xuanjun Chen, Chi-Yuan Hsiao, Puyuan Peng, Shih-Heng Wang, Chun-Yi Kuan, Ke-Han Lu, Kai-Wei Chang 0001, Fabian Ritter Gutierrez |
ICLR | 13 |
| 2025 | Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
Xuanjun Chen, I-Ming Lin, Lin Zhang 0054, Jiawei Du 0003, Hung-yi Lee, Jyh-Shing Roger Jang |
INTERSPEECH | 1 |
| 2024 | Multimodal Transformer Distillation for Audio-Visual SynchronizationabstractAudio-visual synchronization aims to determine whether the mouth movements and speech in the video are synchronized. VocaLiST reaches state-of-the-art performance by incorporating multimodal Transformers to model audio-visual interact information. However, it requires high computing resources, making it impractical for real-world applications. This paper proposed an MTD-VocaLiST model, which is trained by our proposed multimodal Transformer distillation (MTD) loss. MTD loss enables MTDVocaLiST model to deeply mimic the cross-attention distribution and value-relation in the Transformer of VocaLiST. Additionally, we harness uncertainty weighting to fully exploit the interaction information across all layers. Our proposed method is effective in two aspects: From the distillation method perspective, MTD loss outperforms other strong distillation baselines. From the distilled model’s performance perspective: 1) MTDVocaLiST outperforms similar-size SOTA models, SyncNet, and Perfect Match models by 15.65% and 3.35%; 2) MTDVocaLiST reduces the model size of VocaLiST by 83.52%, yet still maintaining similar performance. Xuanjun Chen, Chung-Che Wang, Hung-yi Lee, Jyh-Shing Roger Jang |
ICASSP | 1 |
| 2024 | Neural Codec-based Adversarial Sample Detection for Speaker Verification
Xuanjun Chen, Jiawei Du 0003, Jyh-Shing Roger Jang, Hung-yi Lee |
INTERSPEECH | 1 |
| 2024 | Singing Voice Graph Modeling for SingFake Detection
Xuanjun Chen, Roger Jang, Hung-yi Lee |
INTERSPEECH | 1 |
| 2024 | PointCIM: A Computing-in-Memory Architecture for Accelerating Deep Point Cloud AnalyticsabstractEfficient deep point cloud (PC) analytics is crucial for numerous emerging applications such as autonomous vehicles and augmented and virtual reality. Our roofline model analysis reveals that the “memory wall” bottleneck primarily constrains the execution efficiency of deep PC analytics, providing valuable insight into optimization opportunities. In contrast to previous works, which greatly rely on approximating the original algorithm to fit hardware limitations, the approach presented in this paper is analytical; that is, our approach does not require any modification to the original algorithm, thus preserving its integrity and accuracy. In this paper, we introduce PointCIM, the first deep PC analytics accelerator that leverages computing-in-memory (CIM) optimization opportunities to address memory inefficiency. We identify that existing in-memory methods cannot fully support the distance function required by PC network inference. To address the challenge, we propose computation optimizations, including the Base+Offset mapping and early stopping for bit-serial computation, not only to enable full support for PC network inference in memory, but also to significantly improve hardware efficiency. We design the CIM architecture support for the proposed computation optimizations, including the memristor crossbar architecture, custom peripheral logic, data layout, and pipelined execution. Evaluation results show that the designed accelerator provides an average speedup of 17.1× and an energy reduction of 9.6× compared to the baseline of a typical edge SoC. We also compare PointCIM with several state-of-the-art PC accelerators, yielding up to 10.7× speedup and 4.9× energy savings. Xuanjun Chen, Han-Ping Chen, Chia-Lin Yang |
MICRO | 1 |
| 2024 | DFADD: The Diffusion and Flow-Matching Based Audio Deepfake DatasetabstractMainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, human-level audio synthesis leads to identity misuse and information security issues. Currently, many anti-spoofing models have been developed against deepfake audio. However, the efficacy of current state-of-the-art anti-spoofing models in countering audio synthesized by diffusion and flow-matching based TTS systems remains unknown. In this paper, we proposed the Diffusion and Flow-matching based Audio Deepfake (DFADD) dataset. The DFADD dataset collected the deepfake audio based on advanced diffusion and flowmatching TTS models. Additionally, we reveal that current anti-spoofing models lack sufficient robustness against highly human-like audio generated by diffusion and flow-matching TTS systems. The proposed DFADD dataset addresses this gap and provides a valuable resource for developing more resilient anti-spoofing models. Jiawei Du 0003, I-Ming Lin, I-Hsiang Chiu, Xuanjun Chen, Wenze Ren, Yu Tsao 0001, Hung-yi Lee, Jyh-Shing Roger Jang |
SLT | 4 |
| 2024 | Codec-Superb @ SLT 2024: A Lightweight Benchmark For Neural Audio Codec ModelsabstractNeural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The ideal neural audio codec should maintain content, paralinguistics, speaker characteristics, and audio information even at low bitrates. Recently, numerous advanced neural codec models have been proposed. However, codec models are often tested under varying experimental conditions. As a result, we introduce the Codec-SUPERB challenge at SLT 20241, designed to facilitate fair and lightweight comparisons among existing codec models and inspire advancements in the field. This challenge brings together representative speech applications and objective metrics, and carefully selects license-free datasets, sampling them into small sets to reduce evaluation computation costs. This paper presents the challenge’s rules, datasets, participant systems, results, and findings.1https://codecsuperb.github.io/ Xuanjun Chen, Yi-Cheng Lin, Kai-Wei Chang 0001, Jiawei Du 0003, Ke-Han Lu, Alexander H. Liu, Ho-Lam Chung, Yuan-Kuei Wu, Dongchao Yang, Songxiang Liu, Yi-Chiao Wu, Xu Tan 0003, James R. Glass, Shinji Watanabe 0001, Hung-yi Lee |
SLT | 2 |
| 2023 | Unified Agile Accuracy Assessment in Computing-in-Memory Neural Accelerators by Layerwise Dynamical IsometryabstractDeploying neural networks (NN) on computing-in-memory (CIM) neural accelerators incurs additional hardware factors in the test accuracy, which add substantial extra evaluation overhead. This work takes the first step to quantitatively analyze how information propagates in CIM neural accelerators as well as how additional CIM factors influence that information propagation. From our analysis, we propose a new metric named Unified-QCN that is theoretically linked to the test accuracy according to layerwise dynamical isometry (LDI), providing us with a compass to avoid direct time-consuming simulations. Our method consistently delivers high correlations with the test accuracy for various NN backbones on different datasets. Xuanjun Chen, Cynthia Kuan, Chia-Lin Yang |
DAC | 1 |
| 2022 | Push-Pull: Characterizing the Adversarial Robustness for Audio-Visual Active Speaker DetectionabstractAudio-visual active speaker detection (AVASD) is well-developed, and now is an indispensable front-end for several multi-modal applications. However, to the best of our knowledge, the adversarial robustness of AVASD models hasn't been investigated, not to mention the effective defense against such attacks. In this paper, we are the first to reveal the vulnerability of AVASD models under audio-only, visual-only, and audio-visual adversarial attacks through extensive experiments. What's more, we also propose a novel audio-visual interaction loss (AVIL) for making attackers difficult to find feasible adversarial examples under an allocated attack budget. The loss aims at pushing the inter-class embeddings to be dispersed, namely non-speech and speech clusters, sufficiently disentangled, and pulling the intra-class embeddings as close as possible to keep them compact. Experimental results show the AVIL outperforms the adversarial training by 33.14 mAP (%) under multi-modal attacks. Xuanjun Chen, Helen M. Meng, Hung-yi Lee, Jyh-Shing Roger Jang |
SLT | 1 |
| 2021 | ezGeno: an automatic model selection package for genomic data analysisabstractMOTIVATION: To facilitate the process of tailor-making a deep neural network for exploring the dynamics of genomic DNA, we have developed a hands-on package called ezGeno. ezGeno automates the search process of various parameters and network structures and can be applied to any kind of 1D genomic data. Combinations of multiple abovementioned 1D features are also applicable. RESULTS: For the task of predicting TF binding using genomic sequences as the input, ezGeno can consistently return the best performing set of parameters and network structure, as well as highlight the important segments within the original sequences. For the task of predicting tissue-specific enhancer activity using both sequence and DNase feature data as the input, ezGeno also regularly outperforms the hand-designed models. Furthermore, we demonstrate that ezGeno is superior in efficiency and accuracy compared to the one-layer DeepBind model and AutoKeras, an open-source AutoML package. AVAILABILITY AND IMPLEMENTATION: The ezGeno package can be freely accessed at https://github.com/ailabstw/ezGeno. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun-Liang Lin, Tsung-Ting Hsieh, Yi-An Tung, Xuanjun Chen, Yu-Chun Hsiao, Chia-Lin Yang, Tyng-Luh Liu, Chien-Yu Chen 0001 |
Bioinform. | 4 |