Mengjie Qian 0001

dblp:199/7644-1 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-2614-5411ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Assessment of L2 Oral Proficiency using Speech Large Language Models
Rao Ma, Mengjie Qian 0001, Stefano Bannò, Kate M. Knill, Mark J. F. Gales
INTERSPEECH2
2025 Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
Mengjie Qian 0001, Rao Ma, Stefano Bannò, Kate M. Knill, Mark J. F. Gales
INTERSPEECH1
2024 Towards End-to-End Spoken Grammatical Error Correction
abstract
Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded pipeline comprising an ASR system, disfluency removal, and GEC, with the associated concern of propagating errors between these individual modules. In this paper, we introduce an alternative "end-to-end" approach to spoken GEC, exploiting a speech recognition foundation model, Whisper. This foundation model can be used to replace the whole framework or part of it, e.g., ASR and disfluency removal. These end-to-end approaches are compared to more standard cascaded approaches on the data obtained from a free-speaking spoken language assessment test, Linguaskill. Results demonstrate that end-to-end spoken GEC is possible within this architecture, but the lack of available data limits current performance compared to a system using large quantities of text-based GEC data. Conversely, end-to-end disfluency detection and removal, which is easier for the attention-based Whisper to learn, does outperform cascaded approaches. Additionally, the paper discusses the challenges of providing feedback to candidates when using end-to-end systems for spoken GEC.
Stefano Bannò, Rao Ma, Mengjie Qian 0001, Kate M. Knill, Mark J. F. Gales
ICASSP3
2024 On the Usefulness of Speaker Embeddings for Speaker Retrieval in the Wild: A Comparative Study of x-vector and ECAPA-TDNN Models
Erfan Loweimi, Mengjie Qian 0001, Kate M. Knill, Mark J. F. Gales
INTERSPEECH2
2024 Learn and Don't Forget: Adding a New Language to ASR Foundation Models
Mengjie Qian 0001, Rao Ma, Kate M. Knill, Mark J. F. Gales
INTERSPEECH1
2024 Zero-Shot Audio Topic Reranking Using Large Language Models
abstract
Multimodal Video Search by Examples (MVSE) investigates using video clips as the query term for information retrieval, rather than the more traditional text query. This enables far richer search modalities such as images, speaker, content, topic, and emotion. A key element for this process is highly rapid and flexible search to support large archives, which in MVSE is facilitated by representing video attributes with embeddings. This work aims to compensate for any performance loss from this rapid archive search by examining reranking approaches. In particular, zero-shot reranking methods using large language models (LLMs) are investigated as these are applicable to any video archive audio content. Performance is evaluated for topic-based retrieval on a publicly available video archive, the BBC Rewind corpus. Results demonstrate that reranking significantly improves retrieval ranking without requiring any task-specific in-domain training data. Furthermore, three sources of information (ASR transcriptions, automatic summaries and synopses) as input for LLM reranking were compared. To gain a deeper understanding and further insights into the performance differences and limitations of these text sources, we employ a fact-checking approach to analyse the information consistency among them.
Mengjie Qian 0001, Rao Ma, Adian Liusie, Erfan Loweimi, Kate M. Knill, Mark J. F. Gales
SLT1
2024 Multi-modal video search by examples - A video quality impact analysis
abstract
Abstract As the proliferation of video content continues, and many video archives lack suitable metadata, therefore, video retrieval, particularly through example‐based search, has become increasingly crucial. Existing metadata often fails to meet the needs of specific types of searches, especially when videos contain elements from different modalities, such as visual and audio. Consequently, developing video retrieval methods that can handle multi‐modal content is essential. An innovative Multi‐modal Video Search by Examples (MVSE) framework is introduced, employing state‐of‐the‐art techniques in its various components. In designing MVSE, the authors focused on accuracy, efficiency, interactivity, and extensibility, with key components including advanced data processing and a user‐friendly interface aimed at enhancing search effectiveness and user experience. Furthermore, the framework was comprehensively evaluated, assessing individual components, data quality issues, and overall retrieval performance using high‐quality and low‐quality BBC archive videos. The evaluation reveals that: (1) multi‐modal search yields better results than single‐modal search; (2) the quality of video, both visual and audio, has an impact on the query precision. Compared with image query results, audio quality has a greater impact on the query precision (3) a two‐stage search process (i.e. searching by Hamming distance based on hashing, followed by searching by Cosine similarity based on embedding); is effective but increases time overhead; (4) large‐scale video retrieval is not only feasible but also expected to emerge shortly.
Guanfeng Wu, Abbas Haider, Xing Tian, Erfan Loweimi, Chi-Ho Chan, Mengjie Qian 0001, Muhammad Junaid Awan, Ivor T. A. Spence, Rob Cooper, Wing W. Y. Ng, Josef Kittler, Mark J. F. Gales, Hui Wang 0001
IET Comput. Vis.6
2023 N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space
abstract
Error correction models form an important part of Automatic Speech Recognition (ASR) post-processing to improve the readability and quality of transcriptions.Most prior works use the 1-best ASR hypothesis as input and therefore can only perform correction by leveraging the context within one sentence.In this work, we propose a novel N-best T5 model for this task, which is fine-tuned from a T5 model and utilizes ASR N-best lists as model input.By transferring knowledge from the pretrained language model and obtaining richer information from the ASR decoding space, the proposed approach outperforms a strong Conformer-Transducer baseline.Another issue with standard error correction is that the generation process is not well-guided.To address this a constrained decoding process, either based on the N-best list or an ASR lattice, is used which allows additional information to be propagated.
Rao Ma, Mark J. F. Gales, Kate M. Knill, Mengjie Qian 0001
INTERSPEECH4
2023 Adapting an Unadaptable ASR System
abstract
As speech recognition model sizes and training data requirements grow, it is increasingly common for systems to only be available via APIs from online service providers rather than having direct access to models themselves.In this scenario it is challenging to adapt systems to a specific target domain.To address this problem we consider the recently released OpenAI Whisper ASR as an example of a large-scale ASR system to assess adaptation methods.An error correction based approach is adopted, as this does not require access to the model, but can be trained from either 1-best or N-best outputs that are normally available via the ASR API.LibriSpeech is used as the primary target domain for adaptation.The generalization ability of the system in two distinct dimensions are then evaluated.First, whether the form of correction model is portable to other speech recognition domains, and secondly whether it can be used for ASR models having a different architecture.
Rao Ma, Mengjie Qian 0001, Mark J. F. Gales, Kate M. Knill
INTERSPEECH2
2018 Overview of the 2018 Spoken CALL Shared Task
abstract
Contains fulltext : 199312.pdf (Publisher’s version ) (Open Access)
Claudia Baur, Andrew Caines, Cathy Chua, Johanna Gerlach, Mengjie Qian 0001, Manny Rayner, Martin J. Russell, Helmer Strik, Xizi Wei
INTERSPEECH5
2018 Phone Recognition Using a Non-Linear Manifold with Broad Phone Class Dependent DNNs
abstract
Although it is generally accepted that different broad phone classes (BPCs) have different production mechanisms and are better described by different types of features, most automatic speech recognition (ASR) systems use the same features and decision criteria for all phones. Motivated by this observation, this paper proposes a two-level DNN structure, referred to as a BPC-DNN, inspired by the notion of a topological manifold. In the first level, several small separate BPC-dependent DNNs are applied to different broad phonetic classes and in the second level the outputs of these DNNs are fused to obtain senone-dependent posterior probabilities, which can be used for frame level classification or integrated into Viterbi decoding for phone recognition. In a previous paper using this approach we reported improved frame classification accuracy on the TIMIT corpus compared with a conventional DNN. The contribution of the present paper is to demonstrate that this advantage extends to full phone recognition. Our most recent results show that the BPC-DNN achieves reductions in error rate relative to a conventional DNN of 16% and 8% for frame classification and phone recognition, respectively.
Mengjie Qian 0001, Linxue Bai, Peter Jancovic, Martin J. Russell
INTERSPEECH1
2018 The University of Birmingham 2018 Spoken CALL Shared Task Systems
abstract
This paper describes the systems developed by the University of Birmingham for the 2018 CALL Shared Task (ST) challenge. The task is to perform automatic assessment of grammatical and linguistic aspects of English spoken by German-speaking Swiss teenagers. Our developed systems consist of two components, automatic speech recognition (ASR) and text processing (TP). We explore several ways of building a DNN-HMM ASR system using out-of-domain AMI speech corpus plus a limited amount of ST data. In development experiments on the initial ST data, our final ASR system achieved the word-error-rate (WER) of 12.00%, compared to 14.89% for the official ST baseline DNN-HMM system. The WER of 9.28% was achieved on the test set data. For TP component, we first post-process the ASR output to deal with hesitations and then pass this to a template-based grammar, which we expanded from the provided baseline. We also developed a TP system based on machine learning methods, which enables to better accommodate variability of spoken language. We also fused outputs from several systems using a linear logistic regression. Our best system submitted to the challenge achieved F-measure of 0.914, D of 10.764 and D_{full} score of 5.691 on the final test set.
Mengjie Qian 0001, Xizi Wei, Peter Jancovic, Martin J. Russell
INTERSPEECH1