VLDB 2026 Research / reviewers in the wild / expert
Zhengrui Ma
dblp:276/3133
· DBLP profile ↗
14ranked-venue papers
6as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Language Models Are Read/Write Policy-Makers for Simultaneous GenerationabstractSimultaneous generation models write generation results while reading streaming inputs, necessitating a policy-maker to determine the appropriate output timing. Existing simultaneous generation methods generally adopt the traditional encoder-decoder architecture and learn the generation and policy-making capabilities through complex dynamic programming techniques. Although LLMs excel at text generation, they face challenges in taking on the role of policy-makers through traditional training methods, limiting their exploration in simultaneous generation. To overcome these limitations, we propose a novel LLM-driven Simultaneous Generation (LSG) framework, which allows the off-the-shelf LLM to decide the generation timing and produce output concurrently. Specifically, LSG selects the generation policy that minimizes latency as the baseline policy. Referring to the baseline policy, LSG enables the LLM to devise an improved generation policy that better balances latency and generation quality, and writes generation results accordingly. Experiments on simultaneous translation and streaming automatic speech recognition tasks show that our method can achieve state-of-the-art performance utilizing the open-source LLMs and demonstrate practicality in real-world scenarios. Shoutao Guo, Shaolei Zhang 0001, Zhengrui Ma, Yang Feng 0004 |
AAAI | 3 |
| 2025 | LLaMA-Omni: Seamless Speech Interaction with Large Language ModelsabstractModels like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs. To address this, we propose LLaMA-Omni, a novel model architecture designed for low-latency and high-quality speech interaction with LLMs. LLaMA-Omni integrates a pretrained speech encoder, a speech adaptor, an LLM, and a streaming speech decoder. It eliminates the need for speech transcription, and can simultaneously generate text and speech responses directly from speech instructions with extremely low latency. We build our model based on the latest Llama-3.1-8B-Instruct model. To align the model with speech interaction scenarios, we construct a dataset named InstructS2S-200K, which includes 200K speech instructions and corresponding speech responses. Experimental results show that compared to previous speech-language models, LLaMA-Omni provides better responses in both content and style, with a response latency as low as 226ms. Additionally, training LLaMA-Omni takes less than 3 days on just 4 GPUs, paving the way for the efficient development of speech-language models in the future. Qingkai Fang, Shoutao Guo, Zhengrui Ma, Shaolei Zhang 0001, Yang Feng 0004 |
ICLR | 4 |
| 2025 | Overcoming Non-monotonicity in Transducer-based Streaming GenerationabstractStreaming generation models are utilized across fields, with the Transducer architecture being popular in industrial applications. However, its input-synchronous decoding mechanism presents challenges in tasks requiring non-monotonic alignments, such as simultaneous translation. In this research, we address this issue by integrating Transducer’s decoding with the history of input stream via a learnable monotonic attention. Our approach leverages the forward-backward algorithm to infer the posterior probability of alignments between the predictor states and input timestamps, which is then used to estimate the monotonic context representations, thereby avoiding the need to enumerate the exponentially large alignment space during training. Extensive experiments show that our MonoAttn-Transducer effectively handles non-monotonic alignments in streaming scenarios, offering a robust solution for complex generation tasks. Code is available at https://github.com/ictnlp/MonoAttn-Transducer. Zhengrui Ma, Yang Feng 0004, Min Zhang 0005 |
ICML | 1 |
| 2025 | FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech ProcessingabstractThe rapid advancement of Large Language Models (LLMs) has spurred significant progress in Large Speech-Language Models (LSLMs), enhancing their capabilities in both speech understanding and generation. While existing LSLMs often concentrate on augmenting speech generation or tackling a diverse array of short-speech tasks, the efficient processing of long-form speech remains a critical yet underexplored challenge. This gap is primarily attributed to the scarcity of long-speech training datasets and the high computational costs associated with long sequences. To address these limitations, we introduce FastLongSpeech, a novel framework designed to extend LSLM capabilities for efficient long-speech processing without necessitating dedicated long-speech training data. FastLongSpeech incorporates an iterative fusion strategy that can compress excessively long-speech sequences into manageable lengths. To adapt LSLMs for long-speech inputs, it introduces a dynamic compression training approach, which exposes the model to short-speech sequences at varying compression ratios, thereby transferring the capabilities of LSLMs to long-speech tasks. To assess the long-speech capabilities of LSLMs, we develop a long-speech understanding benchmark called LongSpeech-Eval. Experiments show that our method exhibits strong performance in both long-speech and short-speech tasks, while greatly improving inference efficiency. Shoutao Guo, Shaolei Zhang 0001, Qingkai Fang, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
NeurIPS | 4 |
| 2025 | Efficient Speech Language Modeling via Energy Distance in Continuous Latent SpaceabstractWe introduce \emph{SLED}, an alternative approach to speech language modeling by encoding speech waveforms into sequences of continuous latent representations and modeling them autoregressively using an energy distance objective. The energy distance offers an analytical measure of the distributional gap by contrasting simulated and target samples, enabling efficient training to capture the underlying continuous autoregressive distribution. By bypassing reliance on residual vector quantization, SLED avoids discretization errors and eliminates the need for the complicated hierarchical architectures common in existing speech language models. It simplifies the overall modeling pipeline while preserving the richness of speech information and maintaining inference efficiency. Empirical results demonstrate that SLED achieves strong performance in both zero-shot and streaming speech synthesis, showing its potential for broader applications in general-purpose speech language models. Demos and code are available at \url{https://github.com/ictnlp/SLED-TTS}. Zhengrui Ma, Yang Feng 0004, Chenze Shao, Fandong Meng, Jie Zhou 0016, Min Zhang 0005 |
NeurIPS | 1 |
| 2024 | Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?abstractTwo-pass direct speech-to-speech translation (S2ST) models have shown promising results which decompose S2ST into speech-to-text translation (S2TT) and text-to-speech (TTS), yet conduct end-to-end training by sharing the target text representation between S2TT and TTS models.However, the training of these models still requires large-scale parallel speech data comprising triplets, which is extremely challenging to collect.On the other hand, S2TT and TTS have accumulated a large amount of data and numerous pretrained models, which can be used to reduce the reliance on parallel speech data.To this end, we propose a composite S2ST model named ComSpeech, which connects pretrained S2TT and TTS models by introducing a vocabulary adaptor based on connectionist temporal classification (CTC).The vocabulary adaptor is employed to adapt the output text sequence of S2TT to the input text sequence of TTS, which are different due to the use of different vocabularies.In this way, ComSpeech can still be trained end-to-end and only needs a small amount of parallel speech data to finetune.We further propose a novel training method ComSpeech-ZS to eliminate the reliance on parallel speech data by aligning the text representation space of S2TT and TTS.Experimental results on the CVSS dataset show that when the parallel speech data is available, ComSpeech surpasses previous two-pass models like UnitY and Translatotron 2 in both translation quality and decoding speed.When there is no parallel speech data, ComSpeech-ZS lags behind ComSpeech by only 0.7 ASR-BLEU and outperforms the cascaded models. 1 Qingkai Fang, Shaolei Zhang 0001, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
ACL (1) | 3 |
| 2024 | A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any TranslationabstractSimultaneous translation models play a crucial role in facilitating communication.However, existing research primarily focuses on text-totext or speech-to-text models, necessitating additional cascade components to achieve speechto-speech translation.These pipeline methods suffer from error propagation and accumulate delays in each cascade component, resulting in reduced synchronization between the speaker and listener.To overcome these challenges, we propose a novel non-autoregressive generation framework for simultaneous speech translation (NAST-S2x 1 ), which integrates speechto-text and speech-to-speech tasks into a unified end-to-end framework.We develop a nonautoregressive decoder capable of concurrently generating multiple text or acoustic unit tokens upon receiving fixed-length speech chunks.The decoder can generate blank or repeated tokens and employ CTC decoding to dynamically adjust its latency.Experimental results show that NAST-S2x outperforms state-of-theart models in both speech-to-text and speechto-speech tasks.It achieves high-quality simultaneous interpretation within a delay of less than 3 seconds and provides a 28× decoding speedup in offline generation. 2 Zhengrui Ma, Qingkai Fang, Shaolei Zhang 0001, Shoutao Guo, Yang Feng 0004, Min Zhang 0005 |
ACL (1) | 1 |
| 2024 | StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningabstractSimultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication.Beyond accomplishing translation between speech, Simul-S2ST requires a policy to control the model to generate corresponding target speech at the opportune moment within speech inputs, thereby posing a double challenge of translation and policy.In this paper, we propose StreamSpeech, a direct Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning.Adhering to a multi-task learning approach, StreamSpeech can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model.Experiments on CVSS benchmark demonstrate that StreamSpeech achieves state-of-the-art performance in both offline S2ST and Simul-S2ST tasks.Besides, StreamSpeech is able to present high-quality intermediate results (i.e., ASR or translation results) during simultaneous translation process, offering a more comprehensive real-time communication experience 1 . Shaolei Zhang 0001, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
ACL (1) | 4 |
| 2023 | Non-autoregressive Streaming Transformer for Simultaneous TranslationabstractSimultaneous machine translation (SiMT) models are trained to strike a balance between latency and translation quality.However, training these models to achieve high quality while maintaining low latency often leads to a tendency for aggressive anticipation.We argue that such issue stems from the autoregressive architecture upon which most existing SiMT models are built.To address those issues, we propose non-autoregressive streaming Transformer (NAST) which comprises a unidirectional encoder and a non-autoregressive decoder with intra-chunk parallelism.We enable NAST to generate the blank token or repetitive tokens to adjust its READ/WRITE strategy flexibly, and train it to maximize the nonmonotonic latent alignment with an alignmentbased latency loss.Experiments on various SiMT benchmarks demonstrate that NAST outperforms previous strong autoregressive SiMT baselines. Zhengrui Ma, Shaolei Zhang 0001, Shoutao Guo, Chenze Shao, Min Zhang 0005, Yang Feng 0004 |
EMNLP | 1 |
| 2023 | Fuzzy Alignments in Directed Acyclic Graph for Non-Autoregressive Machine Translation
Zhengrui Ma, Chenze Shao, Shangtong Gui, Min Zhang 0005, Yang Feng 0004 |
ICLR | 1 |
| 2023 | Non-autoregressive Machine Translation with Probabilistic Context-free GrammarabstractNon-autoregressive Transformer(NAT) significantly accelerates the inference of neural machine translation. However, conventional NAT models suffer from limited expression power and performance degradation compared to autoregressive (AT) models due to the assumption of conditional independence among target tokens. To address these limitations, we propose a novel approach called PCFG-NAT, which leverages a specially designed Probabilistic Context-Free Grammar (PCFG) to enhance the ability of NAT models to capture complex dependencies among output tokens. Experimental results on major machine translation benchmarks demonstrate that PCFG-NAT further narrows the gap in translation quality between NAT and AT models. Moreover, PCFG-NAT facilitates a deeper understanding of the generated sentences, addressing the lack of satisfactory explainability in neural machine translation. Code is publicly available at https://github.com/ictnlp/PCFG-NAT. Shangtong Gui, Chenze Shao, Zhengrui Ma, Xishan Zhang, Yunji Chen, Yang Feng 0004 |
NeurIPS | 3 |
| 2023 | Beyond MLE: Convex Learning for Text GenerationabstractMaximum likelihood estimation (MLE) is a statistical method used to estimate the parameters of a probability distribution that best explain the observed data. In the context of text generation, MLE is often used to train generative language models, which can then be used to generate new text. However, we argue that MLE is not always necessary and optimal, especially for closed-ended text generation tasks like machine translation. In these tasks, the goal of model is to generate the most appropriate response, which does not necessarily require it to estimate the entire data distribution with MLE. To this end, we propose a novel class of training objectives based on convex functions, which enables text generation models to focus on highly probable outputs without having to estimate the entire data distribution. We investigate the theoretical properties of the optimal predicted distribution when applying convex functions to the loss, demonstrating that convex functions can sharpen the optimal distribution, thereby enabling the model to better capture outputs with high probabilities. Experiments on various text generation tasks and models show the effectiveness of our approach. It enables autoregressive models to bridge the gap between greedy and beam search, and facilitates the learning of non-autoregressive models with a maximum improvement of 9+ BLEU points. Moreover, our approach also exhibits significant impact on large language models (LLMs), substantially enhancing their generative capability on various tasks. Source code is available at \url{https://github.com/ictnlp/Convex-Learning}. Chenze Shao, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
NeurIPS | 2 |
| 2022 | Prediction Difference Regularization against Perturbation for Neural Machine TranslationabstractRegularization methods applying input perturbation have drawn considerable attention and have been frequently explored for NMT tasks in recent years.Despite their simplicity and effectiveness, we argue that these methods are limited by the under-fitting of training data.In this paper, we utilize prediction difference for ground-truth tokens to analyze the fitting of token-level samples and find that underfitting is almost as common as over-fitting.We introduce prediction difference regularization (PD-R), a simple and effective method that can reduce over-fitting and under-fitting at the same time.For all token-level samples, PD-R minimizes the prediction difference between the original pass and the input-perturbed pass, making the model less sensitive to small input changes, thus more robust to both perturbations and under-fitted training data.Experiments on three widely used WMT translation tasks show that our approach can significantly improve over existing perturbation regularization methods.On WMT16 En-De task, our model achieves 1.80 SacreBLEU improvement over vanilla transformer. Dengji Guo, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
ACL (1) | 2 |
| 2020 | Towards Clustering-friendly Representations: Subspace Clustering via Graph FilteringabstractFinding a suitable data representation for a specific task has been shown to be crucial in many applications. The success of subspace clustering depends on the assumption that the data can be separated into different subspaces. However, this simple assumption does not always hold since the raw data might not be separable into subspaces. To recover the "clustering-friendly" representation and facilitate the subsequent clustering, we propose a graph filtering approach by which a smooth representation is achieved. Specifically, it injects graph similarity into data features by applying a low-pass filter to extract useful data representations for clustering. Extensive experiments on image and document clustering datasets demonstrate that our method improves upon state-of-the-art subspace clustering techniques. Especially, its comparable performance with deep learning methods emphasizes the effectiveness of the simple graph filtering scheme for many real-world applications. An ablation study shows that graph filtering can remove noise, preserve structure in the image, and increase the separability of classes. Zhengrui Ma, Zhao Kang 0001, Guangchun Luo, Ling Tian, Wenyu Chen 0001 |
ACM Multimedia | 1 |