EDBT 2026 Demo / reviewers in the wild / expert
Jun Ma 0018
dblp:91/4845-18
· DBLP profile ↗
18ranked-venue papers
0as first author
12since 2021 · last 2025
0009-0003-8713-0667ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 10 since 2021Artificial intelligence and machine learning · 13 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning DistillationabstractThe rapid advancement of large language models (LLMs) has significantly enhanced their reasoning abilities, enabling increasingly complex tasks. However, these capabilities often diminish in smaller, more computationally efficient models like GPT-2. Recent research shows that reasoning distillation can help small models acquire reasoning capabilities, but most existing methods focus primarily on improving teacher-generated reasoning paths. Our observations reveal that small models can generate high-quality reasoning paths during sampling, even without chain-of-thought prompting, though these paths are often latent due to their low probability under standard decoding strategies. To address this, we propose Self-Enhanced Reasoning Training (SERT), which activates and leverages latent reasoning capabilities in small models through self-training on filtered, self-generated reasoning paths under zero-shot conditions. Experiments using OpenAI’s GPT-3.5 as the teacher model and GPT-2 models as the student models demonstrate that SERT enhances the reasoning abilities of small models, improving their performance in reasoning distillation. Yong Zhang 0058, Zhitao Li 0002, Ming Li 0010, Ning Cheng 0001, Minchuan Chen, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 8 |
| 2024 | EfficientTTS 2: Variational End-to-End Text-to-Speech Synthesis and Voice ConversionabstractRecently, the field of Text-to-Speech (TTS) has been dominated by one-stage text-to-waveform models which have significantly improved speech quality compared to two-stage models. In this work, we propose EfficientTTS 2 (EFTS2), a one-stage high-quality end-to-end TTS framework that is fully differentiable and highly efficient. Our method adopts an adversarial training process, with a differentiable aligner and a hierarchical-VAE-based waveform generator. These design choices free the model from the use of external aligners, invertible structures, and complex training procedures as most previous TTS works have. Moreover, we extend EFTS2 to the voice conversion (VC) task and propose EFTS2-VC, an end-to-end VC model that allows high-quality speech-to-speech conversion. Experimental results suggest that the two proposed models achieve better or at least comparable speech quality compared to baseline models, while also providing faster inference speeds and smaller model sizes. Chenfeng Miao, Qingying Zhu, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Exploring multi-task learning and data augmentation in dementia detection with self-supervised pretrained models
Minchuan Chen, Chenfeng Miao, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 3 |
| 2023 | Improving End-to-End Modeling For Mandarin-English Code-Switching Using Lightweight Switch-Routing Mixture-of-Experts
Fengyun Tan, Chaofeng Feng, Tao Wei 0003, Shuai Gong, Jinqiang Leng, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 7 |
| 2022 | A compact transformer-based GAN vocoder
Chenfeng Miao, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2022 | Towards Efficiently Learning Monotonic Alignments for Attention-based End-to-End Speech Recognition
Chenfeng Miao, Ziyang Zhuang, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 5 |
| 2021 | SEQ-CPC : Sequential Contrastive Predictive Coding for Automatic Speech RecognitionabstractInspired by the contrastive predictive coding (CPC), we propose a feature representation scheme for automatic speech recognition (ASR), which encodes sequential dependency information from raw audio signals. Following the original CPC, for a given frame, mutual information (MI) lower bound is maximized between historical context and future prediction. While computing the MI lower bound, based on original CPC, we develop the sequential CPC (SEQ-CPC), which takes the sequential information between frames into consideration. Since speech frames are not independent events, incorporating sequential information leads to better recognition performance. Experimental results on WSJ corpus show that SEQ-CPC achieves the best performance than CPC and NCE which is the contrastive objective used in wav2vec. Haimei Kang, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 8 |
| 2021 | Improving Neural Text Normalization with Partial Parameter Generator and Pointer-Generator NetworkabstractText Normalization (TN) is an essential part in conversational systems like text-to-speech synthesis (TTS) and automatic speech recognition (ASR). It is a process of transforming non-standard words (NSW) into a representation of how the words are to be spoken. Existing approaches to TN are mainly rule-based or hybrid systems, which require abundant hand-crafted rules. In this paper, we treat TN as a neural machine translation problem and present a pure data-driven TN system using Transformer framework. Partial Parameter Generator (PPG) and Pointer-Generator Network (PGN) are combined in our model to improve accuracy of normalization and act as auxiliary modules to reduce the number of simple errors. The experiments demonstrate that our proposed model reaches remarkable performance on various semiotic classes. Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 4 |
| 2021 | Unsupervised Learning for Multi-Style Speech Synthesis with Limited DataabstractExisting multi-style speech synthesis methods require either style labels or large amounts of unlabeled training data, making data acquisition difficult. In this paper, we present an unsupervised multi-style speech synthesis method that can be trained with limited data. We leverage instance discriminator to guide a style encoder to learn meaningful style representations from a multi-style dataset. Furthermore, we employ information bottleneck to filter out style-irrelevant information in the representations, which can improve speech quality and style similarity. Our method is able to produce desirable speech using a fairly small dataset, where the baseline GST-Tacotron fails. ABX tests show that our model significantly outperforms GST-Tacotron in both emotional speech synthesis task and multi-speaker speech synthesis task. In addition, we demonstrate that our method is able to learn meaningful style features with only 50 training samples per style. Chenfeng Miao, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 4 |
| 2021 | EfficientTTS: An Efficient and High-Quality Text-to-Speech ArchitectureabstractIn this work, we address the Text-to-Speech (TTS) task by proposing a non-autoregressive architecture called EfficientTTS. Unlike the dominant non-autoregressive TTS models, which are trained with the need of external aligners, EfficientTTS optimizes all its parameters with a stable, end-to-end training procedure, allowing for synthesizing high quality speech in a fast and efficient manner. EfficientTTS is motivated by a new monotonic alignment modeling approach, which specifies monotonic constraints to the sequence alignment with almost no increase of computation. By combining EfficientTTS with different feed-forward network structures, we develop a family of TTS models, including both text-to-melspectrogram and text-to-waveform networks. We experimentally show that the proposed models significantly outperform counterpart models such as Tacotron 2 and Glow-TTS in terms of speech quality, training efficiency and synthesis speed, while still producing the speeches of strong robustness and great diversity. In addition, we demonstrate that proposed approach can be easily extended to autoregressive models such as Tacotron 2. Chenfeng Miao, Zhengchen Liu, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
ICML | 5 |
| 2021 | Improving Polyphone Disambiguation for Mandarin Chinese by Combining Mix-Pooling Strategy and Window-Based Attention
Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
Interspeech | 4 |
| 2021 | EfficientSing: A Chinese Singing Voice Synthesis System Using Duration-Free Acoustic Model and HiFi-GAN Vocoder
Zhengchen Liu, Chenfeng Miao, Qingying Zhu, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
Interspeech | 5 |
| 2020 | Flow-TTS: A Non-Autoregressive Network for Text to Speech Based on FlowabstractIn this work, we propose Flow-TTS, a non-autoregressive end-to-end neural TTS model based on generative flow. Unlike other non-autoregressive models, Flow-TTS can achieve high-quality speech generation by using a single feed-forward network. To our knowledge, Flow-TTS is the first TTS model utilizing flow in spectrogram generation network and the first non-autoregssive model which jointly learns the alignment and spectrogram generation through a single network. Experiments on LJSpeech show that the speech quality of Flow-TTS heavily approaches that of human and is even better than that of autoregressive model Tacotron 2 (outperforms Tacotron 2 with a gap of 0.09 in MOS). Meanwhile, the inference speed of Flow-TTS is about 23 times speed-up over Tacotron 2, which is comparable to FastSpeech.1 Chenfeng Miao, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 4 |
| 2020 | Nonparallel Emotional Speech Conversion Using VAE-GAN
Yuexin Cao, Zhengchen Liu, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2020 | Non-Parallel Voice Conversion with Fewer Labeled Data by Conditional Generative Adversarial Networks
Minchuan Chen, Weijian Hou, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 3 |
| 2020 | Improving Replay Detection System with Channel Consistency DenseNeXt for the ASVspoof 2019 Challenge
Junjie Cheng, Yanmei Gu, Huacan Wang, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 5 |
| 2020 | Contextualized Emotion Recognition in Conversation as Sequence TaggingabstractEmotion recognition in conversation (ERC) is an important topic for developing empathetic machines in a variety of areas including social opinion mining, health-care and so on.In this paper, we propose a method to model ERC task as sequence tagging where a Conditional Random Field (CRF) layer is leveraged to learn the emotional consistency in the conversation.We employ LSTM-based encoders that capture self and inter-speaker dependency of interlocutors to generate contextualized utterance representations which are fed into the CRF layer.For capturing long-range global context, we use a multi-layer Transformer encoder to enhance the LSTM-based encoder.Experiments show that our method benefits from modeling the emotional consistency and outperforms the current state-of-theart methods on multiple emotion classification datasets. Jun Ma 0018, Jing Xiao 0006 |
SIGdial | 3 |
| 2019 | Cross-Lingual, Multi-Speaker Text-To-Speech Synthesis Using Neural Speaker Embedding
Mengnan Chen, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 4 |