Rao Ma

dblp:258/9035 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0002-8552-8992ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Assessment of L2 Oral Proficiency using Speech Large Language Models
Rao Ma, Mengjie Qian 0001, Stefano Bannò, Kate M. Knill, Mark J. F. Gales
INTERSPEECH1
2025 Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
Mengjie Qian 0001, Rao Ma, Stefano Bannò, Kate M. Knill, Mark J. F. Gales
INTERSPEECH2
2024 Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
abstract
Recent developments in large speech foundation models like Whisper have led to their widespread use in many automatic speech recognition (ASR) applications.These systems incorporate 'special tokens' in their vocabulary, such as <|endoftext|>, to guide their language generation process.However, we demonstrate that these tokens can be exploited by adversarial attacks to manipulate the model's behavior.We propose a simple yet effective method to learn a universal acoustic realization of Whisper's <|endoftext|> token, which, when prepended to any speech signal, encourages the model to ignore the speech and only transcribe the special token, effectively 'muting' the model.Our experiments demonstrate that the same, universal 0.64-second adversarial audio segment can successfully mute a target Whisper ASR model for over 97% of speech samples.Moreover, we find that this universal adversarial audio segment often transfers to new datasets and tasks.Overall this work demonstrates the vulnerability of Whisper models to 'muting' adversarial attacks, where such attacks can pose both risks and potential benefits in real-world settings: for example the attack can be used to bypass speech moderation systems, or conversely the attack can also be used to protect private speech data. 1
Vyas Raina, Rao Ma, Charles McGhee, Kate M. Knill, Mark J. F. Gales
EMNLP2
2024 Towards End-to-End Spoken Grammatical Error Correction
abstract
Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded pipeline comprising an ASR system, disfluency removal, and GEC, with the associated concern of propagating errors between these individual modules. In this paper, we introduce an alternative "end-to-end" approach to spoken GEC, exploiting a speech recognition foundation model, Whisper. This foundation model can be used to replace the whole framework or part of it, e.g., ASR and disfluency removal. These end-to-end approaches are compared to more standard cascaded approaches on the data obtained from a free-speaking spoken language assessment test, Linguaskill. Results demonstrate that end-to-end spoken GEC is possible within this architecture, but the lack of available data limits current performance compared to a system using large quantities of text-based GEC data. Conversely, end-to-end disfluency detection and removal, which is easier for the attention-based Whisper to learn, does outperform cascaded approaches. Additionally, the paper discusses the challenges of providing feedback to candidates when using end-to-end systems for spoken GEC.
Stefano Bannò, Rao Ma, Mengjie Qian 0001, Kate M. Knill, Mark J. F. Gales
ICASSP2
2024 Learn and Don't Forget: Adding a New Language to ASR Foundation Models
Mengjie Qian 0001, Rao Ma, Kate M. Knill, Mark J. F. Gales
INTERSPEECH3
2024 Investigating the Emergent Audio Classification Ability of ASR Foundation Models
abstract
Rao Ma, Adian Liusie, Mark Gales, Kate Knill. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Rao Ma, Adian Liusie, Mark J. F. Gales, Kate M. Knill
NAACL-HLT1
2024 Zero-Shot Audio Topic Reranking Using Large Language Models
abstract
Multimodal Video Search by Examples (MVSE) investigates using video clips as the query term for information retrieval, rather than the more traditional text query. This enables far richer search modalities such as images, speaker, content, topic, and emotion. A key element for this process is highly rapid and flexible search to support large archives, which in MVSE is facilitated by representing video attributes with embeddings. This work aims to compensate for any performance loss from this rapid archive search by examining reranking approaches. In particular, zero-shot reranking methods using large language models (LLMs) are investigated as these are applicable to any video archive audio content. Performance is evaluated for topic-based retrieval on a publicly available video archive, the BBC Rewind corpus. Results demonstrate that reranking significantly improves retrieval ranking without requiring any task-specific in-domain training data. Furthermore, three sources of information (ASR transcriptions, automatic summaries and synopses) as input for LLM reranking were compared. To gain a deeper understanding and further insights into the performance differences and limitations of these text sources, we employ a fact-checking approach to analyse the information consistency among them.
Mengjie Qian 0001, Rao Ma, Adian Liusie, Erfan Loweimi, Kate M. Knill, Mark J. F. Gales
SLT2
2023 Internal Language Model Estimation Based Adaptive Language Model Fusion for Domain Adaptation
abstract
ASR model deployment environment is ever-changing, and the incoming speech can be switched across different domains during a session. This brings a challenge for effective domain adaptation when only target domain text data is available, and our objective is to obtain obviously improved performance on the target domain while the performance on the general domain is less undermined. In this paper, we propose an adaptive LM fusion approach called internal language model estimation based adaptive domain adaptation (ILME-ADA). To realize such an ILME-ADA, an interpolated log-likelihood score is calculated based on the maximum of the scores from the internal LM and the external LM (ELM) respectively. We demonstrate the efficacy of the proposed ILME-ADA method with both RNN-T and LAS modeling frameworks employing neural network and n-gram LMs as ELMs respectively on two domain specific (target) test sets. The proposed method can achieve significantly better performance on the target test sets while it gets minimal performance degradation on the general test set, compared with both shallow and ILME-based LM fusion methods.
Rao Ma, Jin Qiu, Yanan Qin, Haihua Xu 0001, Peihao Wu, Zejun Ma 0001
ICASSP1
2023 N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space
abstract
Error correction models form an important part of Automatic Speech Recognition (ASR) post-processing to improve the readability and quality of transcriptions.Most prior works use the 1-best ASR hypothesis as input and therefore can only perform correction by leveraging the context within one sentence.In this work, we propose a novel N-best T5 model for this task, which is fine-tuned from a T5 model and utilizes ASR N-best lists as model input.By transferring knowledge from the pretrained language model and obtaining richer information from the ASR decoding space, the proposed approach outperforms a strong Conformer-Transducer baseline.Another issue with standard error correction is that the generation process is not well-guided.To address this a constrained decoding process, either based on the N-best list or an ASR lattice, is used which allows additional information to be propagated.
Rao Ma, Mark J. F. Gales, Kate M. Knill, Mengjie Qian 0001
INTERSPEECH1
2023 Adapting an Unadaptable ASR System
abstract
As speech recognition model sizes and training data requirements grow, it is increasingly common for systems to only be available via APIs from online service providers rather than having direct access to models themselves.In this scenario it is challenging to adapt systems to a specific target domain.To address this problem we consider the recently released OpenAI Whisper ASR as an example of a large-scale ASR system to assess adaptation methods.An error correction based approach is adopted, as this does not require access to the model, but can be trained from either 1-best or N-best outputs that are normally available via the ASR API.LibriSpeech is used as the primary target domain for adaptation.The generalization ability of the system in two distinct dimensions are then evaluated.First, whether the form of correction model is portable to other speech recognition domains, and secondly whether it can be used for ASR models having a different architecture.
Rao Ma, Mengjie Qian 0001, Mark J. F. Gales, Kate M. Knill
INTERSPEECH1
2022 Internal Language Model Estimation Through Explicit Context Vector Learning for Attention-based Encoder-decoder ASR
Rao Ma, Haihua Xu 0001, Zejun Ma 0001
INTERSPEECH2
2021 AISpeech-SJTU ASR System for the Accented English Speech Recognition Challenge
abstract
This paper describes the AISpeech-SJTU ASR system for the Interspeech-2020 Accented English Speech Recognition Challenge (AESRC). This task is challenging due to the diversity of pronunciation accuracy, intonation speed and pronunciation of some syllables. All participants were restricted to develop their systems based on the speech and text corpora provided by the organizer. To work around the data-scarcity problem, data augmentation was first explored including noise simulation, SpecAugment, speed perturbation and TTS simulation. Moreover, SOTA CNN-transformer-based joint CTC-attention system was built and accent adaptation was proposed to train an accent robust system. Finally, the first-pass recognition hypotheses generated from CTC head were rescored by forward, backward LSTM-LM and the attention head. Our system with the best configuration achieves second place in the challenge, resulting in a word error rate (WER) of 4.00% on dev set and 4.47% WER on test set, while WER on test set of the top-performing, second runner-up and official baseline systems are 4.06%, 4.52%, 8.29%, respectively.
Tian Tan 0002, Yizhou Lu, Rao Ma, Sen Zhu, Yanmin Qian
ICASSP3
2021 AISpeech-SJTU Accent Identification System for the Accented English Speech Recognition Challenge
abstract
This paper describes the AISpeech-SJTU system for the accent identification track of the Interspeech-2020 Accented English Speech Recognition Challenge. In this challenge track, only 160-hour accented English data collected from 8 countries and the auxiliary Librispeech dataset are provided for training. To build an accurate and robust accent identification system, we explore the whole system pipeline in detail. First, we introduce the ASR based phone posteriorgram (PPG) feature to accent identification and verify its efficacy. Then, a novel TTS based approach is carefully designed to augment the very limited accent training data for the first time. Finally, we propose the test time augmentation and embedding fusion schemes to further improve the system performance. Our final system is ranked first in the challenge and outperforms all the other participants by a large margin. The submitted system achieves 83.63% average accuracy on the challenge evaluation data, ahead of the others by more than 10% in absolute terms.
Houjun Huang, Xu Xiang, Yexin Yang, Rao Ma, Yanmin Qian
ICASSP4
2020 Unsupervised Dual Paraphrasing for Two-stage Semantic Parsing
abstract
One daunting problem for semantic parsing is the scarcity of annotation.Aiming to reduce nontrivial human labor, we propose a two-stage semantic parsing framework, where the first stage utilizes an unsupervised paraphrase model to convert an unlabeled natural language utterance into the canonical utterance.The downstream naive semantic parser accepts the intermediate output and returns the target logical form.Furthermore, the entire training process is split into two phases: pre-training and cycle learning.Three tailored self-supervised tasks are introduced throughout training to activate the unsupervised paraphrase model.Experimental results on benchmarks OVERNIGHT and GE-OGRANNO demonstrate that our framework is effective and compatible with supervised training.
Ruisheng Cao, Su Zhu, Chen Liu 0019, Rao Ma, Yanbin Zhao, Lu Chen 0002, Kai Yu 0004
ACL5
2020 Addressing the Polysemy Problem in Language Modeling with Attentional Multi-Sense Embeddings
abstract
Neural network language models have gained considerable popularity due to their promising performance. Distributed word embeddings are utilized to represent semantic information. However, each word is associated with a single vector in the embedding layer, disabling the model from capturing the meanings of polysemous words. In this work, we address this problem by assigning multiple fine-grained sense embeddings to each word in the embedding layers. The proposed model discriminates among different senses of a word with attention mechanism in an unsupervised manner. Experiments demonstrate the benefits of our approach in language modeling and ASR rescoring. Investigations are also made on standard word similarity tasks. The results indicate that our proposed method is efficient in modeling polysemy and therefore obtains better word representations.
Rao Ma, Lesheng Jin, Qi Liu 0018, Lu Chen 0002, Kai Yu 0004
ICASSP1
2020 Neural Lattice Search for Speech Recognition
abstract
To improve the accuracy of automatic speech recognition, a two-pass decoding strategy is widely adopted. The first-pass model generates compact word lattices, which are utilized by the second-pass model to perform rescoring. Currently, the most popular rescoring methods are N-best rescoring and lattice rescoring with long short-term memory language models (LSTMLMs). However, these methods encounter the problem of limited search space or inconsistency between training and evaluation. In this paper, we address these problems with an end-to-end model for accurately extracting the best hypothesis from the word lattice. Our model is composed of a bidirectional LatticeLSTM encoder followed by an attentional LSTM decoder. The model takes word lattice as input and generates the single best hypothesis from the given lattice space. When combined with an LSTMLM, the proposed model yields 9.7% and 7.5% relative WER reduction compared to N-best rescoring methods and lattice rescoring methods within the same amount of decoding time.
Rao Ma, Qi Liu 0018, Lu Chen 0002, Kai Yu 0004
ICASSP1
2020 An Investigation on Different Underlying Quantization Schemes for Pre-trained Language Models
Zihan Zhao 0001, Yuncong Liu, Lu Chen 0002, Qi Liu 0018, Rao Ma, Kai Yu 0004
NLPCC (1)5
2020 Neural Network Language Model Compression With Product Quantization and Soft Binarization
abstract
Large memory consumption of the neural network language models (NN LMs) prohibits their use in many resource-constrained scenarios. Hence, effective NN LM compression approaches that are independent of NN structures are of great interest. However, previous approaches usually achieve a high compression ratio at the cost of obvious performance loss. In this paper, two recently proposed quantization approaches, product quantization (PQ) and soft binarization are effectively combined to address the issue. PQ decomposes word embedding matrices into a Cartesian product of low dimensional subspaces and quantizes each subspace separately. Soft binarization uses a small number of float scalars and the knowledge distillation technique to recover the performance loss during the binarization. Experiments show that the proposed approaches can achieve a high compression ratio, from 70 to over 100, while still maintaining comparable performance to the uncompressed NN LM on both PPL and word error rate criteria.
Kai Yu 0004, Rao Ma, Kaiyu Shi, Qi Liu 0018
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Prior Knowledge Driven Label Embedding for Slot Filling in Natural Language Understanding
abstract
Traditional slot filling in natural language understanding (NLU) predicts a one-hot vector for each word. This form of label representation lacks semantic correlation modeling, which leads to severe data sparsity problem, especially when adapting an NLU model to a new domain. To address this issue, a novel label embedding based slot filling framework is proposed in this article. Here, distributed label embedding is constructed for each slot using prior knowledge. Three encoding methods are investigated to incorporate different kinds of prior knowledge about slots: atomic concepts, slot descriptions, and slot exemplars. The proposed label embeddings tend to share text patterns and reuses data with different slot labels. This makes it useful for adaptive NLU with limited data. Also, since label embedding is independent of NLU model, it is compatible with almost all deep learning based slot filling models. The proposed approaches are evaluated on three datasets. Experiments on single domain and domain adaptation tasks show that label embedding achieves significant performance improvement over traditional one-hot label representation as well as advanced zero-shot approaches.
Su Zhu, Rao Ma, Kai Yu 0004
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Highly Efficient Neural Network Language Model Compression Using Soft Binarization Training
abstract
The long short-term memory language model (LSTM LM) has been widely investigated in large vocabulary continuous speech recognition (LVCSR) task. Despite the excellent performance of LSTM LM, its usage in resource-constrained environments, such as portable devices, is limited due to the high consumption of memory. Binarized language model has been proposed to achieve significant memory reduction at the cost of performance degradation at high compression ratio. In this paper, we propose a soft binarization approach to recover the performance of binarized LSTM LM. Experiments show that the proposed method can achieve a high compression rate of 30 × with almost no performance loss in both language modeling and speech recognition tasks.
Rao Ma, Qi Liu 0018, Kai Yu 0004
ASRU1