VLDB 2026 Research / reviewers in the wild / expert
Zejun Ma 0001
dblp:18/10648-1
· DBLP profile ↗
86ranked-venue papers
3as first author
80since 2021 · last 2026
0000-0001-5508-1328ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 62 · 3 first-author · 56 since 2021Artificial intelligence and machine learning · 53 · 2 first-author · 50 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMSearch-R1: Incentivizing LMMs to SearchabstractRobust deployment of large multimodal models (LMMs) in real-world scenarios requires access to external knowledge sources, given the complexity and dynamic nature of real-world information. Existing approaches such as retrieval-augmented generation (RAG) and prompt engineered search agents rely on rigid pipelines, often leading to inefficient or excessive search behaviors. We present MMSearch-R1, the first end-to-end reinforcement learning framework that enables LMMs to perform on-demand, multi-turn search in real-world Internet environments. Our framework integrates both image and text search tools, allowing the model to reason about when and how to invoke them guided by an outcome-based reward with a search penalty. To support training, We collect a multimodal search VQA dataset through a semi-automated pipeline that covers diverse visual and textual knowledge needs and curate a search-balanced subset with both search-required and search-free samples, which proves essential for shaping efficient and on-demand search behavior. Extensive experiments on knowledge-intensive and info-seeking VQA tasks show that our model not only outperforms RAG-based baselines of the same model size, but also matches the performance of a larger RAG-based model while reducing search calls by over 30%. We further analyze key empirical findings to offer actionable insights for advancing research in multimodal search. Wei Li 0119, Bo Li 0080, Zejun Ma 0001, Ziwei Liu 0002 |
ACL (1) | 7 |
| 2025 | LiveCC: Learning Video LLM with Streaming Speech Transcription at ScaleabstractRecent video large language models (Video LLMs) often depend on costly human annotations or proprietary APIs (e.g., GPT-4o) to produce training data, which limits their training at scale. In this paper, we explore large-scale training for Video LLM with cheap automatic speech recognition (ASR) transcripts. Specifically, we propose a novel streaming training approach that densely interleaves the ASR words and video frames according to their timestamps. Compared to previous studies in vision-language representation with ASR, our method naturally fits the streaming characteristics of ASR, thus enabling the model to learn temporally-aligned, fine-grained vision-language modeling. To support the training algorithm, we introduce a data pipeline for YouTube videos and their closed captions (CC), resulting in Live-CC-10M pre-training set and Live-WhisperX-408K high-quality supervised fine-tuning (SFT) set. Remarkably, even without SFT, the pre-trained model LiveCC-7B demonstrates significant improvements in general video QA and exhibits a new capability in real-time video commentary. To evaluate this, we carefully design a new benchmark LiveSports-3K, using LLM-as-a-judge to measure the free-form commentary. Experiments show our final model LiveCC-7B can surpass LLaVA-Video-72B in commentary quality even working in a real-time mode. Meanwhile, it achieves state-of-the-art results at the 7B scale on popular benchmarks such as VideoMME, demonstrating its broad generalizability. All resources of this paper have been released at showlab.github.io/livecc. Joya Chen, Ziyun Zeng, Zejun Ma 0001, Zheng Shou 0001 |
CVPR | 5 |
| 2025 | SCM: Enhancing Large Language Model with Self-Controlled Memory Framework
Xinnian Liang, Jian Yang 0003, Hui Huang 0021, Zhenhe Wu, Shuangzhi Wu, Zejun Ma 0001, Zhoujun Li 0001 |
DASFAA (6) | 7 |
| 2025 | Audio-centric Video Understanding Benchmark without Text ShortcutabstractYudong Yang, Jimin Zhuang, Guangzhi Sun, Changli Tang, Yixuan Li, Peihan Li, Yifan Jiang, Wei Li, Zejun Ma, Chao Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yudong Yang, Jimin Zhuang, Guangzhi Sun, Changli Tang, Peihan Li, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
EMNLP | 9 |
| 2025 | LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal ModelsabstractVisual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image tasks, their applications to multi-image scenarios remains less explored. Additionally, prior LMM research separately tackles different scenarios, leaving it impossible to generalize cross scenarios with new
emerging capabilities. To this end, we introduce LLaVA-Interleave, which simultaneously tackles Multi-image, Multi-frame (video), Multi-view (3D), and Multi-patch (single-image) scenarios in LMMs. To enable these capabilities, we regard the interleaved data format as a general template and compile the M4-Instruct dataset with 1,177.6k samples, spanning 4 primary domains with 14
tasks and 41 datasets. We also curate the LLaVA-Interleave Bench to comprehensively evaluate the multi-image performance of LMMs. Through extensive
experiments, LLaVA-Interleave achieves leading results in multi-image, video,
and 3D benchmarks, while maintaining the performance of single-image tasks.
Besides, our model also exhibits several emerging capabilities, e.g., transferring tasks across different settings and modalities. Feng Li 0040, Renrui Zhang, Hao Zhang 0097, Yuanhan Zhang, Bo Li 0080, Wei Li 0119, Zejun Ma 0001, Chunyuan Li |
ICLR | 7 |
| 2025 | Improving LLM Video Understanding with 16 Frames Per SecondabstractHuman vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted from images sampled at a fixed low frame rate of frame-per-second (FPS) $\leqslant$2, leading to critical visual information loss. In this paper, we introduce F-16, the first multimodal LLM designed for high-frame-rate video understanding. By increasing the frame rate to 16 FPS and compressing visual tokens within each 1-second clip, F-16 efficiently captures dynamic visual features while preserving key semantic information.
Experimental results demonstrate that higher frame rates considerably enhance video understanding across multiple benchmarks, providing a new approach to improving video LLMs beyond scaling model size or training data. F-16 achieves state-of-the-art performance among 7-billion-parameter video LLMs on both general and fine-grained video understanding benchmarks, such as Video-MME and TemporalBench. Furthermore, F-16 excels in complex spatiotemporal tasks, including high-speed sports analysis (*e.g.*, basketball, football, gymnastics, and diving), outperforming SOTA proprietary visual models like GPT-4o and Gemini-1.5-pro.
Additionally, we introduce a novel decoding method for F-16 that enables highly efficient low-frame-rate inference without requiring model retraining. We will release the source code, model checkpoints, and data at [https://github.com/bytedance/F-16](https://github.com/bytedance/F-16). Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
ICML | 7 |
| 2025 | video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language ModelabstractWhile recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in general video understanding. This paper proposes video-SALMONN-o1, the first open-source reasoning-enhanced audio-visual LLM designed for general video understanding tasks. To enhance its reasoning abilities, we develop a reasoning-intensive dataset featuring challenging audio-visual questions with step-by-step solutions. We also propose process direct preference optimization (pDPO), which leverages contrastive step selection to achieve efficient step-level reward modelling tailored for multimodal inputs. Additionally, we introduce RivaBench, the first reasoning-intensive video understanding benchmark, featuring over 4,000 high-quality, expert-curated question-answer pairs across scenarios such as standup comedy, academic presentations, and synthetic video detection. video-SALMONN-o1 achieves 3-8% accuracy improvements over the LLaVA-OneVision baseline across different video reasoning benchmarks. Besides, pDPO achieves 6-8% improvements compared to the supervised fine-tuning model on RivaBench. Enhanced reasoning enables video-SALMONN-o1 zero-shot synthetic video detection capabilities. Guangzhi Sun, Yudong Yang, Jimin Zhuang, Changli Tang, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
ICML | 7 |
| 2025 | ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear AttentionabstractLinear attention mechanisms deliver significant advantages for Large Language Models (LLMs) by providing linear computational complexity, enabling efficient processing of ultra-long sequences (e.g., 1M context). However, existing Sequence Parallelism (SP) methods, essential for distributing these workloads across devices, become the primary performance bottleneck due to substantial communication overhead. In this paper, we introduce ZeCO (Zero Communication Overhead) sequence parallelism for linear attention models, a new SP method designed to overcome these limitations and achieve practically end-to-end near-linear scalability for long sequence training. For example, training a model with a 1M sequence length across 64 devices using ZeCO takes roughly the same time as training with an 16k sequence on a single device. At the heart of ZeCO lies All-Scan, a novel collective communication primitive. All-Scan provides each SP rank with precisely the initial operator state it requires while maintaining a minimal communication footprint, effectively eliminating communication overhead. Theoretically, we prove the optimaity of ZeCO, showing that it introduces only negligible time and space overhead. Empirically, we compare the communication costs of different sequence parallelism strategies and demonstrate that All-Scan achieves the fastest communication in SP scenarios. Specifically, on 256 GPUs with an 8M sequence length, ZeCO achieves a 60\% speedup compared to the current state-of-the-art (SOTA) SP method. We believe ZeCO establishes a clear path toward efficiently training next-generation LLMs on previously intractable sequence lengths. Yuhong Chou, Rui-Jie Zhu 0003, Tianjian Li, Congying Chu, Qian Liu 0033, Jibin Wu, Zejun Ma 0001 |
NeurIPS | 9 |
| 2025 | Robust SuperAlignment: Weak-to-Strong Robustness Generalization for Vision-Language ModelsabstractNumerous well-established studies have demonstrated the superhuman capabilities of modern Vision-Language Models (VLMs) across a wide range of tasks. However, growing is the doubt about the continuing availability of reliable high-quality labeling (supervision) from human annotators, leading to stagnation of the model's performance. To address this challenge, ``superalignment'' employs the so-called weak-to-strong generalization paradigm, where the supervision from a weak model can provide generalizable knowledge for a strong model. While effective in aligning knowledge for clean samples between the strong and weak models, the standard weak-to-strong approach typically fails to capture adversarial robustness, exposing strong VLMs to adversarial attacks. This inability to transfer adversarial robustness is because adversarial samples are normally missing in the superalignment stage. To this end, we are the first to propose the weak-to-strong (adversarial) robustness generalization method to elicit zero-shot robustness in large-scale models by an unsupervised scheme, mitigating the unreliable information source for alignment from two perspectives: alignment re-weighting and source guidance refinement. We analyze settings under which robustness generalization is possible. Extensive experiments across various vision-language benchmarks validate the effectiveness of our method in numerous scenarios, demonstrating its plug-and-play applicability to large-scale VLMs. Junhao Dong 0001, Xinghua Qu, Zejun Ma 0001, Piotr Koniusz, Yew-Soon Ong |
NeurIPS | 4 |
| 2025 | Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency OptimizationabstractLarge Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we introduce a novel test-time iterative optimization framework to address this, employing a closed-loop system where LLMs iteratively refine code based on empirical performance feedback from an execution sandbox. We explore three training strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization~(GRPO). Experiments on our Venus dataset and the APPS benchmark show that SFT and DPO rapidly saturate in efficiency gains. In contrast, GRPO, using reinforcement learning (RL) with execution feedback, continuously optimizes code performance, significantly boosting both pass@1 (from 47% to 62%) and the likelihood of outperforming human submissions in efficiency (from 31% to 45%). Our work demonstrates effective test-time code efficiency improvement and critically reveals the power of RL in teaching LLMs to truly self-improve code efficiency. We released our code and data at https://github.com/Elfsong/Afterburner. Mingzhe Du, Anh Tuan Luu, Yue Liu 0008, Yuhao Qing, Dong Huang 0005, Qian Liu 0033, Zejun Ma 0001, See-Kiong Ng |
NeurIPS | 8 |
| 2025 | General-Reasoner: Advancing LLM Reasoning Across All DomainsabstractReinforcement learning (RL) has recently demonstrated strong potential in enhancing the reasoning capabilities of large language models (LLMs).
Particularly, the "Zero" reinforcement learning introduced by Deepseek-R1-Zero, enables direct RL training of base LLMs without relying on an intermediate supervised fine-tuning stage.
Despite these advancements, current works for LLM reasoning mainly focus on mathematical and coding domains, largely due to data abundance and the ease of answer verification.
This limits the applicability and generalization of such models to broader domains, where questions often have diverse answer representations, and data is more scarce.
In this paper, we propose General-Reasoner, a novel training framework designed to enhance LLM reasoning capabilities across diverse domains. Our key contributions include: (1) constructing a large-scale, high-quality dataset of questions with verifiable answers curated by web crawling, covering a wide range of disciplines; and (2) developing a generative model-based answer verifier, which replaces traditional rule-based verification with the capability of chain-of-thought and context-awareness. We train a series of models and evaluate them on a wide range of datasets covering wide domains like physics, chemistry, finance, electronics etc. Our comprehensive evaluation across these 12 benchmarks (e.g. MMLU-Pro, GPQA, SuperGPQA, TheoremQA, BBEH and MATH AMC) demonstrates that General-Reasoner outperforms existing baseline methods, achieving robust and generalizable reasoning performance while maintaining superior effectiveness in mathematical reasoning tasks. Xueguang Ma, Qian Liu 0033, Dongfu Jiang, Ge Zhang 0009, Zejun Ma 0001, Wenhu Chen |
NeurIPS | 5 |
| 2025 | VisionThink: Smart and Efficient Vision Language Model via Reinforcement LearningabstractRecent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens.
However, we observe that most real-world scenarios do not require such an extensive number of visual tokens. While the performance drops significantly in a small subset of OCR-related tasks, models still perform accurately in most other general VQA tasks with only 1/4 resolution.
Therefore, we propose to dynamically process distinct samples with different resolutions, and present a new paradigm for visual token compression, namely, VisionThink.
It starts with a downsampled image and smartly decides whether it is sufficient for problem solving. Otherwise, the model could output a special token to request the higher-resolution image. Compared to existing Efficient VLM methods that compress tokens using fixed pruning ratios or thresholds, VisionThink autonomously decides whether to compress tokens case by case. As a result, it demonstrates strong fine-grained visual understanding capability on OCR-related tasks, and meanwhile saves substantial visual tokens on simpler tasks.
We adopt reinforcement learning and propose the LLM-as-Judge strategy to successfully apply RL to general VQA tasks. Moreoever, we carefully design a reward function and penalty mechanism to achieve a stable and reasonable image resize call ratio.
Extensive experiments demonstrate the superiority, efficiency, and effectiveness of our method.
All our code and data are open-sourced. Senqiao Yang, Wei Li 0159, Zejun Ma 0001, Bei Yu 0001, Hengshuang Zhao, Jiaya Jia |
NeurIPS | 6 |
| 2024 | SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASRabstractMulti-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among various approaches because of its simplistic architecture and exceptional performance. However, the frequent speaker changes in token-level SOT (t-SOT) present challenges for the autoregressive decoder in effectively utilizing context to predict output sequences. To address this issue, we introduce a masked t-SOT label, which serves as the cornerstone of an auxiliary training loss. Additionally, we utilize a speaker similarity matrix to refine the self-attention mechanism of the decoder. This strategic adjustment enhances contextual relationships within the same speaker’s tokens while minimizing interactions between different speakers’ tokens. We denote our method as speaker-aware SOT (SA-SOT). Experiments on the Librispeech datasets demonstrate that our SA-SOT obtains a relative cpWER reduction ranging from 12.75% to 22.03% on the multi-talker test sets. Furthermore, with more extensive training, our method achieves an impressive cpWER of 3.41%, establishing a new state-of-the-art result on the LibrispeechMix dataset. Zhiyun Fan, Linhao Dong, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001 |
ICASSP | 5 |
| 2024 | Extending Multilingual ASR to New Languages Using Supplementary Encoder and Decoder ComponentsabstractExtending multilingual automatic speech recognition (mASR) systems to new languages poses challenges, particularly when training data for existing languages is limited or unavailable. To tackle this issue, we suggest utilizing supplementary encoder and decoder components. Specifically, we propose appending and fine-tuning a distinct decoder designed for new languages, while preserving the parameters of existing languages to minimize disruption to their performance. Furthermore, we advocate attaching an additional encoder component to enhance acoustic representation learning for new languages, resulting in substantial improvements in word error rate performance. Our experimental findings demonstrate the effectiveness of the proposed methods for the task of extending language support within mASR systems. Yerbolat Khassanov, Tianfeng Chen, Tze Yuang Chong, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001 |
ICASSP | 7 |
| 2024 | Extending Large Language Models for Speech and Audio CaptioningabstractMultimodal large language models (LLMs) have shown promising visual perception abilities by connecting with image encoders, but their performance on auditory tasks has not yet been widely investigated. Meanwhile, automatic speech recognition (ASR) and automatic audio captioning (AAC) are often achieved with separate systems, resulting in incomplete auditory perception abilities. To fill in these gaps, in this paper, we present the first study that achieves both ASR and AAC by connecting an LLM with auditory encoders. A dual auditory encoder structure is proposed, integrating the Whisper encoder for speech and the BEATs encoder for audio events with a high temporal resolution by using a Q-Former at the window level. Experiments for ASR and AAC are performed correspondingly on the widely used LibriSpeech, GigaSpeech, WavCaps, AudioCaps, and Clotho datasets and yield promising results. In particular, state-of-the-art results are achieved on GigaSpeech, AudioCaps and Clotho. Our model is also able to caption speech and audio events simultaneously from clips with mixed speech and background audio events, which is a step towards more complete machine auditory perception. Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Chao Zhang 0031 |
ICASSP | 8 |
| 2024 | Connecting Speech Encoder and Large Language Model for ASRabstractThe impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM. This paper presents a comparative study of three commonly used structures as connectors, including fully connected layers, multi-head cross-attention, and Q-Former. Speech encoders from the Whisper model series as well as LLMs from the Vicuna model series with different model sizes were studied. Experiments were performed on the commonly used LibriSpeech, Common Voice, and GigaSpeech datasets, where the LLMs with Q-Formers demonstrated consistent and considerable word error rate (WER) reductions over LLMs with other connector structures. Q-Former-based LLMs can generalise well to out-of-domain datasets, where 12% relative WER reductions over the Whisper baseline ASR model were achieved on the Eval2000 test set without using any in-domain training data from Switchboard. Moreover, a novel segment-level Q-Former is proposed to enable LLMs to recognise speech segments with a duration exceeding the limitation of the encoders, which results in 17% relative WER reductions over other connector structures on 90-second-long speech data. Wenyi Yu, Changli Tang, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Chao Zhang 0031 |
ICASSP | 8 |
| 2024 | Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisabstractZero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspects: 1) previous works of zero-shot TTS are typically trained with single-sentence prompts, which significantly restricts their performance when the data is relatively sufficient during the inference stage. 2) The prosodic information in prompts is highly coupled with timbre, making it untransferable to each other.
This paper introduces Mega-TTS 2, a generic prompting mechanism for zero-shot TTS, to tackle the aforementioned challenges. Specifically, we design a powerful acoustic autoencoder that separately encodes the prosody and timbre information into the compressed latent space while providing high-quality reconstructions. Then, we propose a multi-reference timbre encoder and a prosody latent language model (P-LLM) to extract useful information from multi-sentence prompts. We further leverage the probabilities derived from multiple P-LLM outputs to produce transferable and controllable prosody.
Experimental results demonstrate that Mega-TTS 2 could not only synthesize identity-preserving speech with a short prompt of an unseen speaker from arbitrary sources but consistently outperform the fine-tuning method when the volume of data ranges from 10 seconds to 5 minutes. Furthermore, our method enables to transfer various speaking styles to the target timbre in a fine-grained and controlled manner. Audio samples can be found in https://boostprompt.github.io/boostprompt/. Ziyue Jiang 0001, Jinglin Liu, Yi Ren 0006, Jinzheng He, Zhenhui Ye, Shengpeng Ji, Qian Yang 0006, Chen Zhang 0020, Pengfei Wei 0001, Xiang Yin 0006, Zejun Ma 0001, Zhou Zhao 0001 |
ICLR | 12 |
| 2024 | PolyVoice: Language Models for Speech to Speech TranslationabstractWith the huge success of GPT models in natural language processing, there is a growing interest in applying language modeling approaches to speech tasks.
Currently, the dominant architecture in speech-to-speech translation (S2ST) remains the encoder-decoder paradigm, creating a need to investigate the impact of language modeling approaches in this area.
In this study, we introduce PolyVoice, a language model-based framework designed for S2ST systems. Our framework comprises three decoder-only language models: a translation language model, a duration language model, and a speech synthesis language model.
These language models employ different types of prompts to extract learned information effectively. By utilizing unsupervised semantic units, our framework can transfer semantic information across these models, making it applicable even to unwritten languages.
We evaluate our system on Chinese $\rightarrow$ English and English $\rightarrow$ Spanish language pairs. Experimental results demonstrate that \method outperforms the state-of-the-art encoder-decoder model, producing voice-cloned speech with high translation and audio quality.
Speech samples are available at https://polyvoice.github.io. Qianqian Dong, Zhiying Huang, Qi Tian 0001, Chen Xu 0008, Tom Ko, Yunlong Zhao 0004, Tang Li 0001, Xuxin Cheng, Fengpeng Yue, Ye Bai 0001, Lu Lu 0015, Zejun Ma 0001, Yuping Wang 0005, Mingxuan Wang, Yuxuan Wang 0002 |
ICLR | 15 |
| 2024 | SALMONN: Towards Generic Hearing Abilities for Large Language ModelsabstractHearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a speech audio language music open neural network, built by integrating a pre-trained text-based large language model (LLM) with speech and audio encoders into a single multimodal model. SALMONN enables the LLM to directly process and understand general audio inputs and achieve competitive performances on a number of speech and audio tasks used in training, such as
automatic speech recognition and translation, auditory-information-based question answering, emotion recognition, speaker verification, and music and audio captioning etc. SALMONN also has a diverse set of emergent abilities unseen in the training, which includes but is not limited to speech translation to untrained languages, speech-based slot filling, spoken-query-based question answering, audio-based storytelling, and speech audio co-reasoning etc. The presence of cross-modal emergent abilities is studied, and a novel few-shot activation tuning approach is proposed to activate such abilities. To our knowledge, SALMONN is the first model of its type and can be regarded as a step towards AI with generic hearing abilities. The source code, model checkpoints and data are available at https://github.com/bytedance/SALMONN. Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Chao Zhang 0031 |
ICLR | 8 |
| 2024 | Real3D-Portrait: One-shot Realistic 3D Talking Portrait SynthesisabstractOne-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talking face animation. Besides, while the existing works mainly focus on synthesizing the head part, it is also vital to generate natural torso and background segments to obtain a realistic talking portrait video. To address these limitations, we present Real3D-Potrait, a framework that (1) improves the one-shot 3D reconstruction power with a large image-to-plane model that distills 3D prior knowledge from a 3D face generative model; (2) facilitates accurate motion-conditioned animation with an efficient motion adapter; (3) synthesizes realistic video with natural torso movement and switchable background using a head-torso-background super-resolution model; and (4) supports one-shot audio-driven talking face generation with a generalizable audio-to-motion model. Extensive experiments show that Real3D-Portrait generalizes well to unseen identities and generates more realistic talking portrait videos compared to previous methods. Video samples are available at https://real3dportrait.github.io. Zhenhui Ye, Tianyun Zhong, Yi Ren 0006, Jiaqi Yang 0008, Weichuang Li, Jiawei Huang 0008, Ziyue Jiang 0001, Jinzheng He, Rongjie Huang 0001, Jinglin Liu, Chen Zhang 0020, Xiang Yin 0006, Zejun Ma 0001, Zhou Zhao 0001 |
ICLR | 13 |
| 2024 | video-SALMONN: Speech-Enhanced Audio-Visual Large Language ModelsabstractSpeech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video processing, which can understand not only visual frame sequences, audio events and music, but speech as well. To obtain fine-grained temporal information required by speech understanding, while keeping efficient for other video elements, this paper proposes a novel multi-resolution causal Q-Former (MRC Q-Former) structure to connect pre-trained audio-visual encoders and the backbone large language model. Moreover, dedicated training approaches including the diversity loss and the unpaired audio-visual mixed training scheme are proposed to avoid frames or modality dominance. On the introduced audio-visual evaluation benchmark, video-SALMONN achieves more than 25% absolute accuracy improvements on the video-QA task and over 30% absolute accuracy improvements on audio-visual QA tasks with human speech. In addition, video-SALMONN demonstrates remarkable video comprehension and reasoning abilities on tasks that are unprecedented by other av-LLMs. Our training code and model checkpoints are available at https://github.com/bytedance/SALMONN/ Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Yuxuan Wang 0002, Chao Zhang 0031 |
ICML | 8 |
| 2024 | Can Large Language Models Understand Spatial Audio?
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001, Yuxuan Wang 0002, Chao Zhang 0031 |
INTERSPEECH | 9 |
| 2024 | MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
Yifei Xin, Zhesong Yu, Bilei Zhu, Lu Lu 0015, Zejun Ma 0001 |
INTERSPEECH | 6 |
| 2024 | BiFSMNv2: Pushing Binary Neural Networks for Keyword Spotting to Real-Network PerformanceabstractDeep neural networks, such as the deep-FSMN, have been widely studied for keyword spotting (KWS) applications while suffering expensive computation and storage. Therefore, network compression technologies such as binarization are studied to deploy KWS models on edge. In this article, we present a strong yet efficient binary neural network for KWS, namely, BiFSMNv2, pushing it to the real-network accuracy performance. First, we present a dual-scale thinnable 1-bit-architecture (DTA) to recover the representation capability of the binarized computation units by dual-scale activation binarization and liberate the speedup potential from an overall architecture perspective. Second, we also construct a frequency-independent distillation (FID) scheme for KWS binarization-aware training, which distills the high- and low-frequency components independently to mitigate the information mismatch between full-precision and binarized representations. Moreover, we propose the learning propagation binarizer (LPB), a general and efficient binarizer that enables the forward and backward propagation of binary KWS networks to be continuously improved through learning. We implement and deploy BiFSMNv2 on ARMv8 real-world hardware with a novel fast bitwise computation kernel (FBCK), which is proposed to fully use registers and increase instruction throughput. Comprehensive experiments show our BiFSMNv2 outperforms the existing binary networks for KWS by convincing margins across different datasets and achieves comparable accuracy with the full-precision networks (only a tiny 1.51% drop on Speech Commands V1-12). We highlight that benefiting from the compact architecture and optimized hardware kernel, BiFSMNv2 can achieve an impressive 25.1× speedup and 20.2× storage-saving on edge hardware. Haotong Qin, Xudong Ma, Yifu Ding 0001, Yang Zhang 0088, Zejun Ma 0001, Jiakai Wang, Jie Luo 0004, Xianglong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Improving Large-Scale Deep Biasing With Phoneme Features and Text-Only Data in Streaming TransducerabstractDeep biasing for the Transducer can improve the recognition performance of rare words or contextual entities, which is essential in practical applications, especially for streaming Automatic Speech Recognition (ASR). However, deep biasing with large-scale rare words remains challenging, as the performance drops significantly when more distractors exist and there are words with similar grapheme sequences in the bias list. In this paper, we combine the phoneme and textual information of rare words in Transducers to distinguish words with similar pronunciation or spelling. Moreover, the introduction of training with text-only data containing more rare words benefits large-scale deep biasing. The experiments on the Librispeech corpus demonstrate that the proposed method achieves state-of-the-art performance on rare word error rate for different scales and levels of bias lists. Jin Qiu, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001 |
ASRU | 6 |
| 2023 | Bytecover3: Accurate Cover Song Identification On Short QueriesabstractDeep learning based methods have become a paradigm for cover song identification (CSI) in recent years, where the ByteCover systems have achieved state-of-the-art results on all the mainstream datasets of CSI. However, with the burgeon of short videos, many real-world applications require matching short music excerpts to full-length music tracks in the database, which is still under-explored and waiting for an industrial-level solution. In this paper, we upgrade the previous ByteCover systems to ByteCover3 that utilizes local features to further improve the identification performance of short music queries. ByteCover3 is designed with a local alignment loss (LAL) module and a two-stage feature retrieval pipeline, allowing the system to perform CSI in a more precise and efficient way. We evaluated ByteCover3 on multiple datasets with different benchmark settings, where ByteCover3 beat all the compared methods including its previous versions. Xingjian Du, Xia Liang, Huidong Liang, Bilei Zhu, Zejun Ma 0001 |
ICASSP | 6 |
| 2023 | Leveraging Phone-Level Linguistic-Acoustic Similarity For Utterance-Level Pronunciation ScoringabstractRecent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding and actual pronunciation of the target phone as the phone-level pronunciation quality representation. In this paper, we propose to use linguistic-acoustic similarity to explicitly measure the deviation of non-native production from its native reference for pronunciation assessment. Specifically, the deviation is first estimated by the cosine similarity between reference phone embedding and corresponding acoustic embedding. Next, a phone-level Goodness of pronunciation (GOP) pre-training stage is introduced to guide this similarity-based learning for better initialization of the aforementioned two embeddings. Finally, a transformer-based hierarchical pronunciation scorer is used to map a sequence of phone embeddings, acoustic embeddings along with their similarity measures to predict the final utterance-level score. Experimental results on the non-native databases suggest that the proposed system significantly outperforms the baselines, where the acoustic and phone embeddings are simply added or concatenated. A further examination shows that the phone embeddings learned in the proposed approach are able to capture linguistic-acoustic attributes of native pronunciation as references. Wei Liu 0147, Kaiqi Fu, Xiaohai Tian, Shuju Shi, Wei Li 0119, Zejun Ma 0001, Tan Lee |
ICASSP | 6 |
| 2023 | An ASR-Free Fluency Scoring Approach with Self-Supervised LearningabstractA typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for the subsequent calculation of fluency-related features or directly modeling speech fluency with an end-to-end approach. This paper describes a novel ASR-free approach for automatic fluency assessment using self-supervised learning (SSL). Specifically, wav2vec2.0 is used to extract frame-level speech features, followed by K-means clustering to assign a pseudo label (cluster index) to each frame. A BLSTM-based model is trained to predict an utterance-level fluency score from frame-level SSL features and the corresponding cluster indexes. Neither speech transcription nor time stamp information is required in the proposed system. It is ASR-free and can potentially avoid the ASR errors effect in practice. Experimental results carried out on non-native English databases show that the proposed approach significantly improves the performance in the "open response" scenario as compared to previous methods and matches the recently reported performance in the "read aloud" scenario. Wei Liu 0147, Kaiqi Fu, Xiaohai Tian, Shuju Shi, Wei Li 0119, Zejun Ma 0001, Tan Lee |
ICASSP | 6 |
| 2023 | Internal Language Model Estimation Based Adaptive Language Model Fusion for Domain AdaptationabstractASR model deployment environment is ever-changing, and the incoming speech can be switched across different domains during a session. This brings a challenge for effective domain adaptation when only target domain text data is available, and our objective is to obtain obviously improved performance on the target domain while the performance on the general domain is less undermined. In this paper, we propose an adaptive LM fusion approach called internal language model estimation based adaptive domain adaptation (ILME-ADA). To realize such an ILME-ADA, an interpolated log-likelihood score is calculated based on the maximum of the scores from the internal LM and the external LM (ELM) respectively. We demonstrate the efficacy of the proposed ILME-ADA method with both RNN-T and LAS modeling frameworks employing neural network and n-gram LMs as ELMs respectively on two domain specific (target) test sets. The proposed method can achieve significantly better performance on the target test sets while it gets minimal performance degradation on the general test set, compared with both shallow and ILME-based LM fusion methods. Rao Ma, Jin Qiu, Yanan Qin, Haihua Xu 0001, Peihao Wu, Zejun Ma 0001 |
ICASSP | 7 |
| 2023 | LiteG2P: A Fast, Light and High Accuracy Model for Grapheme-to-Phoneme ConversionabstractAs a key component of automated speech recognition (ASR) and the front-end in text-to-speech (TTS), grapheme-to-phoneme (G2P) plays the role of converting letters to their corresponding pronunciations. Existing methods are either slow or poor in performance, and are limited in application scenarios, particularly in the process of on-device inference. In this paper, we integrate the advantages of both expert knowledge and connectionist temporal classification (CTC) based neural network and propose a novel method named LiteG2P which is fast, light and theoretically parallel. With the carefully leading design, LiteG2P can be applied both on cloud and on device. Experimental results on the CMU dataset show that the performance of the proposed method is superior to the state-of-the-art CTC based method with 10 times fewer parameters, and even comparable to the state-of-the-art Transformer-based sequence-to-sequence model with less parameters and 33 times less computation. Peisong Huang, Yuxiang Zou, Shichao Liu 0003, Xiang Yin 0006, Zejun Ma 0001 |
ICASSP | 7 |
| 2023 | Virtual Try-On with Pose-Garment Keypoints Guided InpaintingabstractVirtual try-on is an important technology supporting on-line apparel shopping, which provides consumers with a virtual experience to fit garments without physically wearing them. Recently, the image-based virtual try-on has received growing research attention. However, the synthetic results of existing virtual try-on methods usually present distortions in garment shape and lose pattern details. In this paper, we propose a pose-garment keypoints guided inpainting method for the image-based virtual try-on task, which produces high-fidelity try-on images and well preserves the shapes and patterns of the garments. In our method, human pose and garment keypoints are extracted from source images and constructed as graphs to predict the garment keypoints at the target pose. After which, the predicted key-points are used as guide information to predict the target segmentation map and warp the garment image. The try-on image is finally generated with a semantic-conditioned inpainting scheme using the segmentation map and recomposed person image as conditions. To verify the effectiveness of our proposed method, we conduct extensive experiments on the VITON-HD dataset under both paired and unpaired experimental settings. The qualitative and quantitative results show that our method significantly outperforms prior methods at different image resolutions. The codes repository link is https://github.com/lizhi-ntu/KGI. Pengfei Wei 0001, Xiang Yin 0006, Zejun Ma 0001, Alex Chichung Kot |
ICCV | 4 |
| 2023 | AudioQR: Deep Neural Audio Watermarks For QR CodeabstractImage-based quick response (QR) code is frequently used, but creates barriers for the visual impaired people. With the goal of ``AI for good", this paper proposes the AudioQR, a barrier-free QR coding mechanism for the visually impaired population via deep neural audio watermarks. Previous audio watermarking approaches are mainly based on handcrafted pipelines, which is less secure and difficult to apply in large-scale scenarios. In contrast, AudioQR is the first comprehensive end-to-end pipeline that hides watermarks in audio imperceptibly and robustly. To achieve this, we jointly train an encoder and decoder, where the encoder is structured as a concatenation of transposed convolutions and multi-receptive field fusion modules. Moreover, we customize the decoder training with a stochastic data augmentation chain to make the watermarked audio robust towards different audio distortions, such as environment background, room impulse response when playing through the air, music surrounding, and Gaussian noise. Experiment results indicate that AudioQR can efficiently hide arbitrary information into audio without introducing significant perceptible difference. Our code is available at https://github.com/xinghua-qu/AudioQR. Xinghua Qu, Xiang Yin 0006, Pengfei Wei 0001, Lu Lu 0015, Zejun Ma 0001 |
IJCAI | 5 |
| 2023 | Improving Frame-level Classifier for Word Timings with Non-peaky CTC in End-to-End Automatic Speech Recognition
Xianzhao Chen, Yist Y. Lin, Zejun Ma 0001 |
INTERSPEECH | 5 |
| 2023 | Knowledge Distillation Approach for Efficient Internal Language Model Estimation
Haihua Xu 0001, Yerbolat Khassanov, Lu Lu 0015, Zejun Ma 0001, Ji Wu 0002 |
INTERSPEECH | 6 |
| 2023 | GenerTTS: Pronunciation Disentanglement for Timbre and Style Generalization in Cross-Lingual Text-to-Speech
Yahuan Cong, Haopeng Lin, Shichao Liu 0003, Yi Ren 0006, Xiang Yin 0006, Zejun Ma 0001 |
INTERSPEECH | 8 |
| 2023 | Language-specific Boundary Learning for Improving Mandarin-English Code-switching Speech Recognition
Zhiyun Fan, Linhao Dong, Chen Shen 0011, Zhenlin Liang, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001 |
INTERSPEECH | 7 |
| 2023 | Phonetic and Prosody-aware Self-supervised Learning Approach for Non-native Fluency Scoring
Kaiqi Fu, Shaojun Gao, Shuju Shi, Xiaohai Tian, Wei Li 0119, Zejun Ma 0001 |
INTERSPEECH | 6 |
| 2023 | Text-only Domain Adaptation using Unified Speech-Text Representation in Transducer
Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001 |
INTERSPEECH | 5 |
| 2023 | Random Utterance Concatenation Based Data Augmentation for Improving Short-video Speech RecognitionabstractOne of limitations in end-to-end automatic speech recognition (ASR) framework is its performance would be compromised if train-test utterance lengths are mismatched.In this paper, we propose an on-the-fly random utterance concatenation (RUC) based data augmentation method to alleviate train-test utterance length mismatch issue for short-video ASR task.Specifically, we are motivated by observations that our human-transcribed training utterances tend to be much shorter for short-video spontaneous speech (∼3 seconds on average), while our test utterance generated from voice activity detection front-end is much longer (∼10 seconds on average).Such a mismatch can lead to suboptimal performance.Empirically, it's observed the proposed RUC method significantly improves long utterance recognition without performance drop on short one.Overall, it achieves 5.72% word error rate reduction on average for 15 languages and improved robustness to various utterance length. Yist Y. Lin, Haihua Xu 0001, Van Tung Pham, Yerbolat Khassanov, Tze Yuang Chong, Lu Lu 0015, Zejun Ma 0001 |
INTERSPEECH | 9 |
| 2023 | Disentangling the Contribution of Non-native Speech in Automated Pronunciation Assessment
Shuju Shi, Kaiqi Fu, Yiwei Gu, Xiaohai Tian, Shaojun Gao, Wei Li 0119, Zejun Ma 0001 |
INTERSPEECH | 7 |
| 2023 | StyleS2ST: Zero-shot Style Transfer for Direct Speech-to-speech Translation
Yi Ren 0006, Lei Xie 0001, Xiang Yin 0006, Zejun Ma 0001 |
INTERSPEECH | 8 |
| 2023 | S2CD: Self-heuristic Speaker Content Disentanglement for Any-to-Any Voice Conversion
Pengfei Wei 0001, Xiang Yin 0006, Xinghua Qu, Zhiqiang Xu 0003, Zejun Ma 0001 |
INTERSPEECH | 7 |
| 2023 | Emotionally Situated Text-to-Speech Synthesis in User-Agent ConversationabstractConversational Text-to-speech Synthesis (TTS) aims to generate speech with proper style in the user-agent conversation scenario. Although previous works have explored modeling the context in the dialogue history to provide style information for the agent, there are still deficiencies in modeling the role-aware multi-modal context. Moreover, previous works ignore the emotional dependencies between the user and the agent, which includes: 1) agent understands emotional states of users, and 2) agent expresses proper emotion in the generated speech. In this work, we propose an Emotionally Situated Text-to-speech Synthesis (EmoSit-TTS) framework to understand users' semantics and subtle emotional states, and generate speech with proper speaking style and emotional expression in the user-agent conversation. Experiments on the DailyTalk dataset show the superiority of our proposed framework for the user-agent conversational TTS, especially in terms of emotion-aware expressiveness, which outperforms other state-of-the-art methods by 0.69 on MOS. Demos of our proposed framework are available at https://anonydemo.github.io. Yuchen Liu 0003, Shichao Liu 0003, Xiang Yin 0006, Zejun Ma 0001, Qin Jin |
ACM Multimedia | 5 |
| 2023 | Towards Building Voice-based Conversational Recommender Systems: Datasets, Potential Solutions and ProspectsabstractConversational recommender systems (CRSs) have become crucial emerging research topics in the field of RSs, thanks to their natural advantages of explicitly acquiring user preferences via interactive conversations and revealing the reasons behind recommendations. However, the majority of current CRSs are text-based, which is less user-friendly and may pose challenges for certain users, such as those with visual impairments or limited writing and reading abilities. Therefore,for the first time, this paper investigates the potential of voice-based CRS (VCRSs) to revolutionize the way users interact with RSs in a natural, intuitive, convenient, and accessible fashion. To support such studies, we create two VCRSs benchmark datasets in the e-commerce and movie domains, after realizing the lack of such datasets through an exhaustive literature review. Specifically, we first empirically verify the benefits and necessity of creating such datasets. Thereafter, we convert the user-item interactions to text-based conversations through the ChatGPT-driven prompts for generating diverse and natural templates, and then synthesize the corresponding audios via the text-to-speech model. Meanwhile, a number of strategies are delicately designed to ensure the naturalness and high quality of voice conversations. On this basis, we further explore the potential solutions and point out possible directions to build end-to-end VCRSs by seamlessly extracting and integrating voice-based inputs, thus delivering performance-enhanced, self-explainable, and user-friendly VCRSs. Our study aims to establish the foundation and motivate further pioneering research in the emerging field of VCRSs. This aligns with the principles of explainable AI and AI for social good, viz., utilizing technology's potential to create a fair, sustainable, and just world. Our codes and datasets are available on GitHub (https://github.com/hyllll/VCRS ). Xinghua Qu, Zhu Sun 0001, Xiang Yin 0006, Yew-Soon Ong, Lu Lu 0015, Zejun Ma 0001 |
SIGIR | 7 |
| 2023 | Graph contrastive learning with implicit augmentations
Huidong Liang, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Ke Chen 0021, Junbin Gao |
Neural Networks | 4 |
| 2023 | Adaptive Transfer Kernel Learning for Transfer Gaussian Process RegressionabstractTransfer regression is a practical and challenging problem with important applications in various domains, such as engineering design and localization. Capturing the relatedness of different domains is the key of adaptive knowledge transfer. In this paper, we investigate an effective way of explicitly modelling domain relatedness through transfer kernel, a transfer-specified kernel that considers domain information in the covariance calculation. Specifically, we first give the formal definition of transfer kernel, and introduce three basic general forms that well cover existing related works. To cope with the limitations of the basic forms in handling complex real-world data, we further propose two advanced forms. Corresponding instantiations of the two forms are developed, namely${Trk}_{\alpha \beta }$and${Trk}_{\omega }$based on multiple kernel learning and neural networks, respectively. For each instantiation, we present a condition with which the positive semi-definiteness is guaranteed and a semantic meaning is interpreted to the learned domain relatedness. Moreover, the condition can be easily used in the learning ofTrGP$_{\alpha \beta }$andTrGP$_{\omega }$that are the Gaussian process models with the transfer kernels${Trk}_{\alpha \beta }$and${Trk}_{\omega }$respectively. Extensive empirical studies show the effectiveness ofTrGP$_{\alpha \beta }$andTrGP$_{\omega }$on domain relatedness modelling and transfer adaptiveness. Pengfei Wei 0001, Yiping Ke, Yew-Soon Ong, Zejun Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Transfer Kernel Learning for Multi-Source Transfer Gaussian Process RegressionabstractMulti-source transfer regression is a practical and challenging problem where capturing the diverse relatedness of different domains is the key of adaptive knowledge transfer. In this paper, we propose an effective way of explicitly modeling the domain relatedness of each domain pair through transfer kernel learning. Specifically, we first discuss the advantages and disadvantages of existing transfer kernels in handling the multi-source transfer regression problem. To cope with the limitations of the existing transfer kernels, we further propose a novel multi-source transfer kernel$k_{ms}$. The proposed$k_{ms}$assigns a learnable parametric coefficient to model the relatedness of each inter-domain pair, and simultaneously regulates the relatedness of the intra-domain pair to be 1. Moreover, to capture the heterogeneous data characteristics of multiple domains,$k_{ms}$exploits different standard kernels for different domain pairs. We further provide a theorem that not only guarantees the positive semi-definiteness of$k_{ms}$but also conveys a semantic interpretation to the learned domain relatedness. Moreover, the theorem can be easily used in the learning of the corresponding transfer Gaussian process model with$k_{ms}$. Extensive empirical studies show the effectiveness of our proposed method on domain relatedness modelling and transfer performance. Pengfei Wei 0001, Thanh Vinh Vo, Xinghua Qu, Yew-Soon Ong, Zejun Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataabstractDeep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a single model to target multiple sources, they have difficulty generalizing to unseen sources. In this paper, we propose a three-component pipeline to train a universal audio source separator from a large, but weakly-labeled dataset: AudioSet. First, we propose a transformer-based sound event detection system for processing weakly-labeled training data. Second, we devise a query-based audio separation model that leverages this data for model training. Third, we design a latent embedding processor to encode queries that specify audio targets for separation, allowing for zero-shot generalization. Our approach uses a single model for source separation of multiple sound types, and relies solely on weakly-labeled data for training. In addition, the proposed audio separator can be used in a zero-shot setting, learning to separate types of audio sources that were never seen in training. To evaluate the separation performance, we test our model on MUSDB18, while training on the disjoint AudioSet. We further verify the zero-shot performance by conducting another experiment on audio source types that are held-out from training. The model achieves comparable Source-to-Distortion Ratio (SDR) performance to current supervised models in both cases. Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
AAAI | 4 |
| 2022 | Dynamic Transfer Gaussian Process RegressionabstractIn this paper, we work on a challenging dynamic transfer regression problem where domains come in a streaming manner. At each time stage, a new domain emerges and is taken as the target domain while all the domains in previous time stages are taken as source domains. We propose a transfer Gaussian process model GPdk with a novel dynamic transfer kernel DyTK to handle the dynamic transfer regression problem. Specifically, DyTK is with a sequential form to fit the domain stream. To adaptively control the knowledge transfer strength, DyTK is designed to be capable of modeling the inter-domain relatedness of every inter-domain pair. A theorem that ensures DyTK to be positive semi-definite is then proposed. We also theoretically analyze the transfer performance of GPdk by deriving its generalization error bounds. The error bounds further motivate us to propose a parameter reuse strategy to alleviate the scalability issue of GPdk along time. Extensive experiments on both synthetic and real-world datasets show the effectiveness of GPdk in handling dynamic transfer regression problems. Pengfei Wei 0001, Xinghua Qu, Wen Song 0004, Zejun Ma 0001 |
CIKM | 4 |
| 2022 | HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and DetectionabstractAudio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted in this field. However, existing audio transformers require large GPU memories and long training time, meanwhile relying on pretrained vision models to achieve high performance, which limits the model’s scalability in audio tasks. To combat these problems, we introduce HTS-AT: an audio transformer with a hierarchical structure to reduce the model size and training time. It is further combined with a token-semantic module to map final outputs into class featuremaps, thus enabling the model for the audio event detection (i.e. localization in time). We evaluate HTS-AT on three datasets of audio classification where it achieves new state-of-the-art (SOTA) results on AudioSet and ESC50, and equals the SOTA on Speech Command V2. It also achieves better performance in event localization than the previous CNN-based models. Moreover, HTS-AT requires only 35% model parameters and 15% training time of the previous audio transformer. These results demonstrate the high performance and high efficiency of HTS-AT. Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 4 |
| 2022 | Bytecover2: Towards Dimensionality Reduction of Latent Embedding for Efficient Cover Song IdentificationabstractConvolutional neural network (CNN)-based methods have dominated the recent research of cover song identification (CSI). A typical example is the ByteCover system we proposed, which has achieved state-of-the-art results on all the mainstream datasets of CSI. In this paper, we propose an up-graded version of ByteCover, termed ByteCover2, which further improves ByteCover in both identification performance and efficiency. Compared with ByteCover, ByteCover2 is designed with an additional PCA-FC module, which integrates the capability of principal component analysis (PCA) and fully-connected (FC) neural network for dimensionality reduction of the audio embedding, allowing ByteCover2 to perform CSI in a more precise and efficient way. We evaluated ByteCover2 on multiple datasets in different dimension sizes and training settings, where ByteCover2 beat all the compared methods including ByteCover, even with a dimension size of 128, which is 15 times smaller than that of ByteCover. Xingjian Du, Ke Chen 0021, Bilei Zhu, Zejun Ma 0001 |
ICASSP | 5 |
| 2022 | Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge SelectionabstractNowadays, most methods for end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on phrase-level contextual modeling and attention-based relevance modeling, they may suffer from the confusion between similar context-specific phrases, which hurts predictions at the token level. In this work, we focus on mitigating confusion problems with fine-grained contextual knowledge selection (FineCoS). In FineCoS, we introduce fine-grained knowledge to reduce the uncertainty of token predictions. Specifically, we first apply phrase selection to narrow the range of phrase candidates, and then conduct token attention on the tokens in the selected phrase candidates. Moreover, we re-normalize the attention weights of most relevant phrases in inference to obtain more focused phrase-level contextual representations, and inject position information to help model better discriminate phrases or tokens. On LibriSpeech and an in-house 160,000-hour dataset, we explore the proposed methods based on an all-neural biasing method, collaborative decoding (ColDec). The proposed methods further bring at most 6.1% relative word error rate reduction on LibriSpeech and 16.4% relative character error rate reduction on the in-house dataset. Minglun Han, Linhao Dong, Zhenlin Liang, Zejun Ma 0001, Bo Xu 0002 |
ICASSP | 6 |
| 2022 | Improving Pseudo-Label Training For End-To-End Speech Recognition Using Gradient MaskabstractIn the recent trend of semi-supervised speech recognition, both self-supervised representation learning and pseudo-labeling have shown promising results. In this paper, we propose a novel approach to combine their ideas for end-to-end speech recognition model. Without any extra loss function, we utilize the Gradient Mask to optimize the model when training on pseudo-label. This method forces the speech recognition model to predict from the masked in-put to learn strong acoustic representation and make training robust to label noise. In our semi-supervised experiments, the method can improve the model’s performance when training on pseudo-label and our method achieved competitive results comparing with other semi-supervised approaches on the Librispeech 100 hours experiments. Shaoshi Ling, Chen Shen 0011, Zejun Ma 0001 |
ICASSP | 4 |
| 2022 | Language Adaptive Cross-Lingual Speech Representation Learning with Sparse Sharing Sub-NetworksabstractUnsupervised cross-lingual speech representation learning (XLSR) has recently shown promising results in speech recognition by leveraging vast amounts of unlabeled data across multiple languages. However, standard XLSR model suffers from language interference problem due to the lack of language specific modeling ability. In this work, we investigate language adaptive training on XLSR models. More importantly, we propose a novel language adaptive pretraining approach based on sparse sharing sub-networks. It makes room for language specific modeling by pruning out unimportant parameters for each language, without requiring any manually designed language specific component. After pruning, each language only maintains a sparse sub-network, while the sub-networks are partially shared with each other. Experimental results on a downstream multilingual speech recognition task show that our proposed method significantly outperforms baseline XLSR models on both high resource and low resource languages. Besides, our proposed method consistently outperforms other adaptation methods and requires fewer parameters. Yizhou Lu, Mingkun Huang, Xinghua Qu, Pengfei Wei 0001, Zejun Ma 0001 |
ICASSP | 5 |
| 2022 | The Volcspeech System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription ChallengeabstractThis paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to make the clustering-based speaker diarization system enable to handle overlapped speech. Front-end dereverberation and the direction-of-arrival (DOA) estimation are used to improve the accuracy of speaker diarization. Multi-channel combination and overlap detection are applied to reduce the missed speaker error. A modified DOVER-Lap is also proposed to fuse the results from different systems. We achieve the final DER of 5.79% on the Eval set and 7.23% on the Test set, which ranks 4th in the diarization challenge. For Track 2, we develop our system using the Conformer model in a joint CTC-attention architecture. Serialized output training (SOT) is adopted to multi-speaker overlapped speech recognition. We propose a neural front-end module to model multi-channel audio and train the model end-to-end. Various data augmentation methods are utilized to mitigate over-fitting in the multi-channel multi-speaker E2E system. Transformer language model fusion is developed to achieve better performance. The final CER is 19.2% on the Eval set and 20.8% on the Test set, which ranks 2nd in the ASR challenge. Chen Shen 0011, Wenzhi Fan, Shixue Wen, Jun Zhang 0066, Jingsheng Yang, Zejun Ma 0001 |
ICASSP | 9 |
| 2022 | Towards Using Clothes Style Transfer for Scenario-Aware Person Video GenerationabstractClothes style transfer for person video generation is a challenging task, due to drastic variations of intra-person appearance and video scenarios. To tackle this problem, most recent AdaIN-based architectures are proposed to extract clothes and scenario features for generation. However, these approaches suffer from being short of fine-grained details and are prone to distort the origin person. To further improve the generation performance, we propose a novel framework with disentangled multi-branch encoders and a shared decoder. Moreover, to pursue the strong video spatio-temporal consistency, an inner-frame discriminator is delicately designed with input being cross-frame difference. Besides, the proposed frame-work possesses the property of scenario adaptation. Extensive experiments on the TEDXPeople benchmark demonstrate the superiority of our method over state-of-the-art approaches in terms of image quality and video coherence. Jingning Xu, Benlai Tang, Siyuan Bian, Wenyi Guo, Xiang Yin 0006, Zejun Ma 0001 |
ICASSP | 7 |
| 2022 | S3T: Self-Supervised Pre-Training with Swin Transformer For Music ClassificationabstractIn this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive easily accessible unlabeled music data. S3T introduces a momentum-based paradigm, MoCo, with Swin Transformer as its feature extractor to music time-frequency domain. For better music representations learning, S3T contributes a music data augmentation pipeline and two specially designed pre-processors. To our knowledge, S3T is the first method combining the Swin Transformer with a self-supervised learning method for music classification. We evaluate S3T on music genre classification and music tagging tasks with linear classifiers trained on learned representations. Experimental results show that S3T outperforms the previous self-supervised method (CLMR) by 12.5 percents top-1 accuracy and 4.8 percents PR-AUC on two tasks respectively, and also surpasses the task-specific state-of-the-art supervised methods. Besides, S3T shows advances in label efficiency using only 10% labeled data exceeding CLMR on both tasks with 100% labeled data. Chen Zhang 0020, Bilei Zhu, Zejun Ma 0001 |
ICASSP | 4 |
| 2022 | BiFSMN: Binary Neural Network for Keyword SpottingabstractThe deep neural networks, such as the Deep-FSMN, have been widely studied for keyword spotting (KWS) applications. However, computational resources for these networks are significantly constrained since they usually run on-call on edge devices. In this paper, we present BiFSMN, an accurate and extreme-efficient binary neural network for KWS. We first construct a High-frequency Enhancement Distillation scheme for the binarization-aware training, which emphasizes the high-frequency information from the full-precision network's representation that is more crucial for the optimization of the binarized network. Then, to allow the instant and adaptive accuracy-efficiency trade-offs at runtime, we also propose a Thinnable Binarization Architecture to further liberate the acceleration potential of the binarized network from the topology perspective. Moreover, we implement a Fast Bitwise Computation Kernel for BiFSMN on ARMv8 devices which fully utilizes registers and increases instruction throughput to push the limit of deployment efficiency. Extensive experiments show that BiFSMN outperforms existing binarization methods by convincing margins on various datasets and is even comparable with the full-precision counterpart (e.g., less than 3% drop on Speech Commands V1-12). We highlight that benefiting from the thinnable architecture and the optimized 1-bit implementation, BiFSMN can achieve an impressive 22.3x speedup and 15.5x storage-saving on real-world edge hardware. Haotong Qin, Xudong Ma, Yifu Ding 0001, Yang Zhang 0088, Zejun Ma 0001, Jie Luo 0004, Xianglong Liu 0001 |
IJCAI | 7 |
| 2022 | Token-level Speaker Change Detection Using Speaker Difference and Speech Content via Continuous Integrate-and-fire
Zhiyun Fan, Zhenlin Liang, Linhao Dong, Jun Zhang 0066, Zejun Ma 0001, Bo Xu 0002 |
INTERSPEECH | 8 |
| 2022 | Using Fluency Representation Learned from Sequential Raw Features for Improving Non-native Fluency Scoring
Kaiqi Fu, Shaojun Gao, Xiaohai Tian, Wei Li 0119, Zejun Ma 0001 |
INTERSPEECH | 5 |
| 2022 | Bring dialogue-context into RNN-T for streaming ASR
Junfeng Hou, Jinkun Chen, Yufeng Tang, Jun Zhang 0066, Zejun Ma 0001 |
INTERSPEECH | 6 |
| 2022 | Internal Language Model Estimation Through Explicit Context Vector Learning for Attention-based Encoder-decoder ASR
Rao Ma, Haihua Xu 0001, Zejun Ma 0001 |
INTERSPEECH | 5 |
| 2022 | A Transfer and Multi-Task Learning based Approach for MOS Prediction
Xiaohai Tian, Kaiqi Fu, Shaojun Gao, Yiwei Gu, Wei Li 0119, Zejun Ma 0001 |
INTERSPEECH | 7 |
| 2022 | Towards high-fidelity singing voice conversion with acoustic reference and contrastive predictive codingabstractRecently, phonetic posteriorgrams (PPGs) based methods have been quite popular in non-parallel singing voice conversion systems. However, due to the lack of acoustic information in PPGs, style and naturalness of the converted singing voices are still limited. To solve these problems, in this paper, we utilize an acoustic reference encoder to implicitly model singing characteristics. We experiment with different auxiliary features, including mel spectrograms, HuBERT, and the middle hidden feature (PPG-Mid) of pretrained automatic speech recognition (ASR) model, as the input of the reference encoder, and finally find the HuBERT feature is the best choice. In addition, we use contrastive predictive coding (CPC) module to further smooth the voices by predicting future observations in latent space. Experiments show that, compared with the baseline models, our proposed model can significantly improve the naturalness of converted singing voices and the similarity with the target singer. Moreover, our proposed model can also make the speakers with just speech data sing. Benlai Tang, Xiang Yin 0006, Yuan Wan, Yibiao Yu, Zejun Ma 0001 |
INTERSPEECH | 7 |
| 2022 | Importance Prioritized Policy DistillationabstractPolicy distillation (PD) has been widely studied in deep reinforcement learning (RL), while existing PD approaches assume that the demonstration data (i.e., state-action pairs in frames) in a decision making sequence is uniformly distributed. This may bring in unwanted bias since RL is a reward maximizing process instead of simple label matching. Given such an issue, we denote the frame importance as its contribution to the expected reward on a particular frame, and hypothesize that adapting such frame importance could benefit the performance of the distilled student policy. To verify our hypothesis, we analyze why and how frame importance matters in RL settings. Based on the analysis, we propose an importance prioritized PD framework that highlights the training on important frames, so as to learn efficiently. Particularly, the frame importance is measured by the reciprocal of weighted Shannon entropy from a teacher policy's action prescriptions. Experiments on Atari games and policy compression tasks show that capturing the frame importance significantly boosts the performance of the distilled policies. Xinghua Qu, Yew-Soon Ong, Abhishek Gupta 0001, Pengfei Wei 0001, Zhu Sun 0001, Zejun Ma 0001 |
KDD | 6 |
| 2022 | Synthesising Audio Adversarial Examples for Automatic Speech RecognitionabstractAdversarial examples in automatic speech recognition (ASR) are naturally sounded by humans yet capable of fooling well trained ASR models to transcribe incorrectly. Existing audio adversarial examples are typically constructed by adding constrained perturbations on benign audio inputs. Such attacks are therefore generated with an audio dependent assumption. For the first time, we propose the Speech Synthesising based Attack (SSA), a novel threat model that constructs audio adversarial examples entirely from scratch, i.e., without depending on any existing audio to fool cutting-edge ASR models. To this end, we introduce a conditional variational auto-encoder (CVAE) as the speech synthesiser. Meanwhile, an adaptive sign gradient descent algorithm is proposed to solve the adversarial audio synthesis task. Experiments on three datasets (i.e., Audio Mnist, Common Voice, and Librispeech) show that our method could synthesise naturally sounded audio adversarial examples to mislead the start-of-the-art ASR models. Our web-page containing generated audio demos is at https://sites.google.com/view/ssa-asr/home. Xinghua Qu, Pengfei Wei 0001, Mingyong Gao, Zhu Sun 0001, Yew-Soon Ong, Zejun Ma 0001 |
KDD | 6 |
| 2022 | GIO: A Timbre-informed Approach for Pitch Tracking in Highly Noisy EnvironmentsabstractAs one of the fundamental tasks in music and speech signal processing, pitch tracking has been attracting attention for decades. While a human can focus on the voiced pitch even in highly noisy environments, most existing automatic pitch tracking systems show unsatisfactory performance encountering noise. To mimic human auditory, a data-driven model named GIO is proposed in this paper, in which timbre information is introduced to guide pitch tracking. The proposed model takes two inputs: a short audio segment to extract pitch from and a timbre embedding derived from the speaker's or singer's voice. In experiments, we use a music artist classification model to extract timbre embedding vectors. A dual-branch structure and a two-step training method are designed to enable the model to predict voice presence. The experimental results show that the proposed model gains a significant improvement in noise robustness and outperforms existing state-of-the-art methods with fewer parameters. Xiaoheng Sun, Xia Liang, Qiqi He, Bilei Zhu, Zejun Ma 0001 |
ICMR | 5 |
| 2022 | Sequence-Level Speaker Change Detection With Difference-Based Continuous Integrate-and-FireabstractSpeaker change detection is an important task in multi-party interactions such as meetings and conversations. In this paper, we address the speaker change detection task from the perspective of sequence transduction. Specifically, we propose a novel encoder-decoder framework that directly converts the input feature sequence to the speaker identity sequence. The difference-based continuous integrate-and-fire mechanism is designed to support this framework. It detects speaker changes by integrating the speaker difference between the encoder outputs frame-by-frame and transfers encoder outputs to segment-level speaker embeddings according to the detected speaker changes. The whole framework is supervised by the speaker identity sequence, a weaker label than the precise speaker change points. The experiments on the AMI and DIHARD-I corpora show that our sequence-level method consistently outperforms a strong frame-level baseline that uses the precise speaker change labels. Zhiyun Fan, Linhao Dong, Zejun Ma 0001, Bo Xu 0002 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Bytecover: Cover Song Identification Via Multi-Loss TrainingabstractWe present in this paper ByteCover, which is a new feature learning method for cover song identification (CSI). Byte-Cover is built based on the classical ResNet model, and two major improvements are designed to further enhance the capability of the model for CSI. In the first improvement, we introduce the integration of instance normalization (IN) and batch normalization (BN) to build IBN blocks, which are major components of our ResNet-IBN model. With the help of the IBN blocks, our CSI model can learn features that are invariant to the changes of musical attributes such as key, tempo, timbre and genre, while preserving the version information. In the second improvement, we employ the BN-Neck method to allow a multi-loss training and encourage our method to jointly optimize a classification loss and a triplet loss, and by this means, the inter-class discrimination and intra-class compactness of cover songs, can be ensured at the same time. A set of experiments demonstrated the effectiveness and efficiency of ByteCover on multiple datasets, and in the Da-TACOS dataset, ByteCover outperformed the best competitive system by 18.0%. Xingjian Du, Zhesong Yu, Bilei Zhu, Xiaoou Chen, Zejun Ma 0001 |
ICASSP | 5 |
| 2021 | Singing Melody Extraction from Polyphonic Music based on Spectral Correlation ModelingabstractConvolutional neural network (CNN) based methods have achieved state-of-the-art performance for singing melody extraction from polyphonic music. However, most of these methods focus on the learning of local features, while relationships among spectral components locating far apart are often neglected. In this paper, we explore the idea of modeling spectral correlation explicitly for melody extraction. Specifically, we present a spectral correlation module (SCM) that can learn to model the relationships among all frequency bands in a time-frequency representation, thus allowing the encoding of global spectral information into a conventional CNN. Furthermore, we propose to integrate center frequencies with the input feature map of SCM to improve the performance. We implement a light-weight model comprised of SCM blocks to verify the efficacy of our system. Our system achieves a state-of-the-art overall accuracy of 83.5% on the MedleyDB dataset. Xingjian Du, Bilei Zhu, Qiuqiang Kong, Zejun Ma 0001 |
ICASSP | 4 |
| 2021 | An Hrnet-Blstm Model With Two-Stage Training For Singing Melody ExtractionabstractWell-labeled datasets available for melody extraction are scarce, which limits the further advancement of deep learning based methods. To overcome this problem, we propose to use a pitch refinement method to refine the semitone-level pitch sequences decoded from massive melody MIDI files to generate a large number of fundamental frequency (F0) values for model training. Since the refined pitch values used for the first round of training contain errors, a small set of well-labeled data is used for a second round of training. A high-resolution network (HRNet), initially developed for human pose estimation, is introduced for melody extraction. It considers multi-resolution feature learning, making the resulting representation semantically richer. Subsequently, a bidirectional long short-term memory (BLSTM) layer is used to exploit the temporal information of melody. In addition, a new loss function where the unvoiced frames only contribute to voicing detection but not to pitch classification is also proposed to alleviate the class imbalance problem. Experiment results on three public datasets show that the proposed system outperforms four state-of-the-art algorithms in most cases. Yongwei Gao, Xingjian Du, Bilei Zhu, Xiaoheng Sun, Wei Li 0012, Zejun Ma 0001 |
ICASSP | 6 |
| 2021 | Rule-Embedded Network for Audio-Visual Voice Activity Detection in Live Musical Video StreamsabstractDetecting anchor’s voice in live musical streams is an important preprocessing step for music and speech signal processing. Existing approaches to voice activity detection (VAD) primarily rely on audio, however, audio-based VAD is difficult to effectively focus on the target voice in noisy environments. This paper proposes a rule-embedded network to fuse the audio-visual (A-V) inputs for better detection of the target voice. The core role of the rule in the model is to coordinate the relation between the bi-modal information and use visual representations as a mask to filter out the information of non-target sound. Experiments show that: 1) with the help of cross-modal fusion using the proposed rule, the detection results of the A-V branch outperform that of the audio branch in the same model framework; 2) the performance of the bimodal A-V model far outperforms that of audio-only models, indicating that the incorporation of both audio and visual signals is highly beneficial for VAD. To attract more attention to the cross-modal music and audio signal processing, a new live musical video corpus with frame-level labels is introduced. Yuanbo Hou, Bilei Zhu, Zejun Ma 0001, Dick Botteldooren |
ICASSP | 4 |
| 2021 | PPG-Based Singing Voice Conversion with Adversarial Representation LearningabstractSinging voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion works, we propose a novel model to steadily convert songs while keeping their naturalness and intonation. We build an end-to-end architecture, taking phonetic posteriorgrams (PPGs) as inputs and generating mel spectrograms. Specifically, we implement two separate encoders: one encodes PPGs as content, and the other compresses mel spectrograms to supply acoustic and musical information. To improve the performance on timbre and melody, an adversarial singer confusion module and a mel-regressive representation learning module are designed for the model. Objective and subjective experiments are conducted on our private Chinese singing corpus. Comparing with the baselines, our methods can significantly improve the conversion performance in terms of naturalness, melody, and voice similarity. Moreover, our PPG-based method is proved to be robust for noisy sources. Benlai Tang, Xiang Yin 0006, Yuan Wan, Chen Shen 0011, Zejun Ma 0001 |
ICASSP | 7 |
| 2021 | A Chapter-Wise Understanding System for Text-To-Speech in Chinese NovelsabstractIn TTS-based audiobook production, multi-role dubbing and emotional expressions can significantly improve the naturalness of audiobooks. However, it requires manual annotation of original novels with explicit speaker and emotion tags in sentence level, which is extremely time-consuming and costly. In this paper, we propose a chapter-wise understanding system for Chinese novels, to predict speaker and emotion tags automatically based on the chapter-level context. Compared with baselines of each component, our models obtain higher performance. Audiobooks produced by our proposed system along with a multi-speaker emotional TTS system, are proved to achieve comparable quality score to audiobooks made by individual producers. Demos are demonstrated in https://jeffpan.net/icassp/2021/main.html. Xiang Yin 0006, Chenchang Xu, Zejun Ma 0001 |
ICASSP | 6 |
| 2021 | Improving RNN Transducer Modeling for Small-Footprint Keyword SpottingabstractThe recurrent neural network transducer (RNN-T) model has been proved effective for keyword spotting (KWS) recently. However, compared with cross-entropy (CE) or connectionist temporal classification (CTC) based models, the additional prediction network in the RNN-T model increases the model size and computational cost. Besides, since the keyword training data usually only contain the keyword sequence, the prediction network might has over-fitting problems. In this paper, we improve the RNN-T modeling for small-footprint keyword spotting in three aspects. First, to address the over-fitting issue, we explore multi-task training where a CTC loss is added to the encoder. The CTC loss is calculated with both KWS data and ASR data, while the RNN-T loss is calculated with ASR data so that only the encoder is augmented with KWS data. Second, we use the feed-forward neural network to replace the LSTM for prediction network modeling. Thus all possible prediction network outputs could be pre-computed for decoding. Third, we further improve the model with transfer learning, where a model trained with 160 thousand hours of ASR data is used to initialize the KWS model. On a self-collected far-field wake-word testset, the proposed RNN-T system greatly improves the performance comparing with a strong "keyword-filler" baseline. Haitao Yao, Yaming Liu, Zejun Ma 0001 |
ICASSP | 5 |
| 2021 | Emitting Word Timings with HMM-Free End-to-End System in Automatic Speech Recognition
Xianzhao Chen, Zejun Ma 0001, Zongxia Xie |
Interspeech | 5 |
| 2021 | Attention-Based Cross-Modal Fusion for Audio-Visual Voice Activity Detection in Musical Video StreamsabstractMany previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse musical video data, voice activity detection is a necessary step. This paper attempts to detect the speech and singing voices of target performers in musical video streams using audio-visual information. To integrate information of audio and visual modalities, a multi-branch network is proposed to learn audio and image representations, and the representations are fused by attention based on semantic similarity to shape the acoustic representations through the probability of anchor vocalization. Experiments show the proposed audio-visual multi-branch network far outperforms the audio-only model in challenging acoustic environments, indicating the cross-modal information fusion based on semantic correlation is sensible and successful. Yuanbo Hou, Zhesong Yu, Xia Liang, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Dick Botteldooren |
Interspeech | 6 |
| 2021 | HMM-Free Encoder Pre-Training for Streaming RNN TransducerabstractThis work describes an encoder pre-training procedure using frame-wise label to improve the training of streaming recurrent neural network transducer (RNN-T) model.Streaming RNN-T trained from scratch usually performs worse than nonstreaming RNN-T.Although it is common to address this issue through pre-training components of RNN-T with other criteria or frame-wise alignment guidance, the alignment is not easily available in end-to-end manner.In this work, frame-wise alignment, used to pre-train streaming RNN-T's encoder, is generated without using a HMM-based system.Therefore an allneural framework equipping HMM-free encoder pre-training is constructed.This is achieved by expanding the spikes of CTC model to their left/right blank frames, and two expanding strategies are proposed.To our best knowledge, this is the first work to simulate HMM-based frame-wise label using CTC model for pre-training.Experiments conducted on LibriSpeech and MLS English tasks show the proposed pre-training procedure, compared with random initialization, reduces the WER by relatively 5%∼11% and the emission latency by 60 ms.Besides, the method is lexicon-free, so it is friendly to new languages without manually designed lexicon. Jingyu Sun, Yufeng Tang, Junfeng Hou, Jinkun Chen, Jun Zhang 0066, Zejun Ma 0001 |
Interspeech | 7 |
| 2021 | Fine-Grained Prosody Modeling in Neural Speech Synthesis Using ToBI Representation
Yuxiang Zou, Shichao Liu 0003, Xiang Yin 0006, Haopeng Lin, Zejun Ma 0001 |
Interspeech | 7 |
| 2021 | Towards Realistic Visual Dubbing with Heterogeneous SourcesabstractThe task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data sources of videos and audios, thus causing the failure to leverage heterogeneous data sufficiently. In practice, it may be intractable to collect the perfect homologous data in some cases, for example, audio-corrupted or picture-blurry videos. To explore this kind of data and support high-fidelity few-shot visual dubbing, in this paper, we novelly propose a simple yet efficient two-stage framework with a higher flexibility of mining heterogeneous data. Specifically, our two-stage paradigm employs facial landmarks as intermediate prior of latent representations and disentangles the lip movements prediction from the core task of realistic talking head generation. By this means, our method makes it possible to independently utilize the training corpus for two-stage sub-networks using more available heterogeneous data easily acquired. Besides, thanks to the disentanglement, our framework allows a further fine-tuning for a given talking head, thereby leading to better speaker-identity preserving in the final synthesized results. Moreover, the proposed method can also transfer appearance features from others to the target speaker. Extensive experimental results demonstrate the superiority of our proposed method in generating highly realistic videos synchronized with the speech over the state-of-the-art. Tianyi Xie, Liucheng Liao, Benlai Tang, Xiang Yin 0006, Jianfei Yang 0001, Jiali Yao, Yang Zhang 0088, Zejun Ma 0001 |
ACM Multimedia | 10 |
| 2020 | A Unified Sequence-to-Sequence Front-End Model for Mandarin Text-to-Speech SynthesisabstractIn Mandarin text-to-speech (TTS) system, the front-end text processing module significantly influences the intelligibility and naturalness of synthesized speech. Building a typical pipeline-based front-end which consists of multiple individual components requires extensive efforts. In this paper, we proposed a unified sequence-to-sequence front-end model for Mandarin TTS that converts raw texts to linguistic features directly. Compared to the pipeline-based front-end, our unified front-end can achieve comparable performance in polyphone disambiguation and prosody word prediction, and improve intonation phrase prediction by 0.0738 in F1 score. We also implemented the unified front-end with Tacotron and WaveRNN to build a Mandarin TTS system. The synthesized speech by that got a comparable MOS (4.38) with the pipeline-based front-end (4.37) and close to human recordings (4.49). Xiang Yin 0006, Zhiling Zhang, Shichao Liu 0003, Yang Zhang 0088, Zejun Ma 0001, Yuxuan Wang 0002 |
ICASSP | 6 |
| 2020 | A Hybrid Text Normalization System Using Multi-Head Self-Attention For MandarinabstractIn this paper, we propose a hybrid text normalization system using multi-head self-attention. The system combines the advantages of a rule-based model and a neural model for text preprocessing tasks. Previous studies in Mandarin text normalization usually use a set of hand-written rules, which are hard to improve on general cases. The idea of our proposed system is motivated by the neural models from recent studies and has a better performance on our internal news corpus. This paper also includes different attempts to deal with imbalanced pattern distribution of the dataset. Overall, the performance of the system is improved by over 1.9% on sentence-level. This idea can potentially be adopted by different languages with rule-based text normalization systems. Xiang Yin 0006, Chen Li 0042, Shichao Liu 0003, Yang Zhang 0088, Yuxuan Wang 0002, Zejun Ma 0001 |
ICASSP | 8 |
| 2019 | Learning Hierarchical Representations for Expressive Speaking Style in End-to-End Speech SynthesisabstractAlthough Global Style Tokens (GSTs) are a recently-proposed method to uncover expressive factors of variation in speaking style, they are a mixture of style attributes without explicitly considering the factorization of multiple-level speaking styles. In this work, we introduce a hierarchical GST architecture with residuals to Tacotron, which learns multiple-level disentangled representations to model and control different style granularities in synthesized speech. We make hierarchical evaluations conditioned on individual tokens from different GST layers. As the number of layers increases, we tend to observe a coarse to fine style decomposition. For example, the first GST layer learns a good representation of speaker IDs while finer speaking style or emotion variations can be found in higher-level layers. Meanwhile, the proposed model shows good performance of style transfer. Xiaochun An, Yuxuan Wang 0002, Shan Yang 0001, Zejun Ma 0001, Lei Xie 0001 |
ASRU | 4 |
| 2012 | Unsupervised training of subspace gaussian mixture models for conversational telephone speech recognitionabstractThis paper presents our preliminary works on exploring unsupervised training of subspace gaussian mixture models for under-resourced CTS recognition task. The subspace model yields better performance than conventional GMM model, particularly in small or middle-sized training set. As an effective way to save human efforts, unsupervised learning is often applied to automatically transcribe a large amount of speech archives. The additional auto-transcribed data may help to improve model accuracy. In this paper, experiments are carried out on two publicly available English conversational telephone speech corpora. Both GMM and SGMM model in combination with unsupervised learning are examined and compared in this paper. Zejun Ma 0001, Bo Xu 0002 |
ICASSP | 1 |
| 2011 | An Empirical Study of Multilingual Spoken Term Detection
Zejun Ma 0001, Bo Xu 0002 |
INTERSPEECH | 1 |
| 2011 | Fusing Multiple Confidence Measures for Chinese Spoken Term Detection
Zejun Ma 0001, Bo Xu 0002 |
INTERSPEECH | 1 |