VLDB 2026 Research / reviewers in the wild / expert
Shoutao Guo
dblp:331/5767
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2025
0009-0006-1662-7504ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Machine translation · 37% Speech recognition and synthesis · 33% Deep learning architectures and training · 9% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
simultaneous machine translation |
3.7 | 5 | 2025 | Large Language Models Are Read/Write Policy-Makers for Simultaneous Generation · AAAI 2025 StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning · ACL (1) 2024 Decoder-only Streaming Transformer for Simultaneous Translation · ACL (1) 2024 |
Natural language and speech › Machine translation › speech translation
speech-to-speech translation |
1.5 | 2 | 2024 | StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning · ACL (1) 2024 A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation · ACL (1) 2024 |
Machine learning › Deep learning architectures and training
transformer |
1.4 | 2 | 2024 | Decoder-only Streaming Transformer for Simultaneous Translation · ACL (1) 2024 Non-autoregressive Streaming Transformer for Simultaneous Translation · EMNLP 2023 |
Natural language and speech › Speech recognition and synthesis
speech language model |
1.1 | 2 | 2025 | LLaMA-Omni: Seamless Speech Interaction with Large Language Models · ICLR 2025 LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis · ACL (1) 2025 |
Natural language and speech › Speech recognition and synthesis
speech interaction |
0.9 | 1 | 2025 | LLaMA-Omni: Seamless Speech Interaction with Large Language Models · ICLR 2025 |
Natural language and speech › Speech recognition and synthesis › speech language model
speech large language model |
0.9 | 1 | 2025 | FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing · NeurIPS 2025 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.9 | 1 | 2025 | LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis · ACL (1) 2025 |
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems |
0.9 | 1 | 2025 | LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis · ACL (1) 2025 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › real-time speech recognition
streaming speech recognition |
0.9 | 1 | 2025 | Large Language Models Are Read/Write Policy-Makers for Simultaneous Generation · AAAI 2025 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
non-autoregressive speech generation |
0.8 | 1 | 2024 | A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation · ACL (1) 2024 |
Natural language and speech › Machine translation › simultaneous machine translation
simultaneous speech translation |
0.8 | 1 | 2024 | A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation · ACL (1) 2024 |
Machine learning › Generative modeling › transformer-based generative model
non-autoregressive transformer |
0.7 | 1 | 2023 | Non-autoregressive Streaming Transformer for Simultaneous Translation · EMNLP 2023 |
Machine learning › Reinforcement learning
policy learning |
0.7 | 1 | 2023 | Learning Optimal Policy for Simultaneous Machine Translation via Binary Search · ACL (1) 2023 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
speech encoder · 1.7large language model · 1.7non-autoregressive decoding · 1.4speech adaptor · 0.9policy-making · 0.9iterative fusion · 0.9dynamic programming · 0.9dynamic compression training · 0.9autoregressive streaming speech decoder · 0.9streaming self-attention · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Language Models Are Read/Write Policy-Makers for Simultaneous GenerationabstractSimultaneous generation models write generation results while reading streaming inputs, necessitating a policy-maker to determine the appropriate output timing. Existing simultaneous generation methods generally adopt the traditional encoder-decoder architecture and learn the generation and policy-making capabilities through complex dynamic programming techniques. Although LLMs excel at text generation, they face challenges in taking on the role of policy-makers through traditional training methods, limiting their exploration in simultaneous generation. To overcome these limitations, we propose a novel LLM-driven Simultaneous Generation (LSG) framework, which allows the off-the-shelf LLM to decide the generation timing and produce output concurrently. Specifically, LSG selects the generation policy that minimizes latency as the baseline policy. Referring to the baseline policy, LSG enables the LLM to devise an improved generation policy that better balances latency and generation quality, and writes generation results accordingly. Experiments on simultaneous translation and streaming automatic speech recognition tasks show that our method can achieve state-of-the-art performance utilizing the open-source LLMs and demonstrate practicality in real-world scenarios. Shoutao Guo, Shaolei Zhang 0001, Zhengrui Ma, Yang Feng 0004 |
AAAI | 1 |
| 2025 | LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech SynthesisabstractReal-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction.Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language models (LLMs).In this paper, we introduce LLaMA-Omni 2, a series of speech language models (SpeechLMs) ranging from 0.5B to 14B parameters, capable of achieving highquality real-time speech interaction.LLaMA-Omni 2 is built upon the Qwen2.5 series models, integrating a speech encoder and an autoregressive streaming speech decoder.Despite being trained on only 200K multi-turn speech dialogue samples, LLaMA-Omni 2 demonstrates strong performance on several spoken question answering and speech instruction following benchmarks, surpassing previous state-of-theart SpeechLMs like GLM-4-Voice, which was trained on millions of hours of speech data. 1 Qingkai Fang, Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 3 |
| 2025 | LLaMA-Omni: Seamless Speech Interaction with Large Language ModelsabstractModels like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs. To address this, we propose LLaMA-Omni, a novel model architecture designed for low-latency and high-quality speech interaction with LLMs. LLaMA-Omni integrates a pretrained speech encoder, a speech adaptor, an LLM, and a streaming speech decoder. It eliminates the need for speech transcription, and can simultaneously generate text and speech responses directly from speech instructions with extremely low latency. We build our model based on the latest Llama-3.1-8B-Instruct model. To align the model with speech interaction scenarios, we construct a dataset named InstructS2S-200K, which includes 200K speech instructions and corresponding speech responses. Experimental results show that compared to previous speech-language models, LLaMA-Omni provides better responses in both content and style, with a response latency as low as 226ms. Additionally, training LLaMA-Omni takes less than 3 days on just 4 GPUs, paving the way for the efficient development of speech-language models in the future. Qingkai Fang, Shoutao Guo, Zhengrui Ma, Shaolei Zhang 0001, Yang Feng 0004 |
ICLR | 2 |
| 2025 | FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech ProcessingabstractThe rapid advancement of Large Language Models (LLMs) has spurred significant progress in Large Speech-Language Models (LSLMs), enhancing their capabilities in both speech understanding and generation. While existing LSLMs often concentrate on augmenting speech generation or tackling a diverse array of short-speech tasks, the efficient processing of long-form speech remains a critical yet underexplored challenge. This gap is primarily attributed to the scarcity of long-speech training datasets and the high computational costs associated with long sequences. To address these limitations, we introduce FastLongSpeech, a novel framework designed to extend LSLM capabilities for efficient long-speech processing without necessitating dedicated long-speech training data. FastLongSpeech incorporates an iterative fusion strategy that can compress excessively long-speech sequences into manageable lengths. To adapt LSLMs for long-speech inputs, it introduces a dynamic compression training approach, which exposes the model to short-speech sequences at varying compression ratios, thereby transferring the capabilities of LSLMs to long-speech tasks. To assess the long-speech capabilities of LSLMs, we develop a long-speech understanding benchmark called LongSpeech-Eval. Experiments show that our method exhibits strong performance in both long-speech and short-speech tasks, while greatly improving inference efficiency. Shoutao Guo, Shaolei Zhang 0001, Qingkai Fang, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
NeurIPS | 1 |
| 2024 | Decoder-only Streaming Transformer for Simultaneous TranslationabstractSimultaneous Machine Translation (SiMT) generates translation while reading source tokens, essentially producing the target prefix based on the source prefix.To achieve good performance, it leverages the relationship between source and target prefixes to exact a policy to guide the generation of translations.Although existing SiMT methods primarily focus on the Encoder-Decoder architecture, we explore the potential of Decoder-only architecture, owing to its superior performance in various tasks and its inherent compatibility with SiMT.However, directly applying the Decoder-only architecture to SiMT poses challenges in terms of training and inference.To alleviate the above problems, we propose the first Decoder-only SiMT model, named Decoder-only Streaming Transformer (DST).Specifically, DST separately encodes the positions of the source and target prefixes, ensuring that the position of the target prefix remains unaffected by the expansion of the source prefix.Furthermore, we propose a Streaming Self-Attention (SSA) mechanism tailored for the Decoder-only architecture.It is capable of obtaining translation policy by assessing the sufficiency of input source information and integrating with the soft-attention mechanism to generate translations.Experiments demonstrate that our approach achieves state-of-theart performance on three translation tasks 1 . Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 1 |
| 2024 | A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any TranslationabstractSimultaneous translation models play a crucial role in facilitating communication.However, existing research primarily focuses on text-totext or speech-to-text models, necessitating additional cascade components to achieve speechto-speech translation.These pipeline methods suffer from error propagation and accumulate delays in each cascade component, resulting in reduced synchronization between the speaker and listener.To overcome these challenges, we propose a novel non-autoregressive generation framework for simultaneous speech translation (NAST-S2x 1 ), which integrates speechto-text and speech-to-speech tasks into a unified end-to-end framework.We develop a nonautoregressive decoder capable of concurrently generating multiple text or acoustic unit tokens upon receiving fixed-length speech chunks.The decoder can generate blank or repeated tokens and employ CTC decoding to dynamically adjust its latency.Experimental results show that NAST-S2x outperforms state-of-theart models in both speech-to-text and speechto-speech tasks.It achieves high-quality simultaneous interpretation within a delay of less than 3 seconds and provides a 28× decoding speedup in offline generation. 2 Zhengrui Ma, Qingkai Fang, Shaolei Zhang 0001, Shoutao Guo, Yang Feng 0004, Min Zhang 0005 |
ACL (1) | 4 |
| 2024 | StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningabstractSimultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication.Beyond accomplishing translation between speech, Simul-S2ST requires a policy to control the model to generate corresponding target speech at the opportune moment within speech inputs, thereby posing a double challenge of translation and policy.In this paper, we propose StreamSpeech, a direct Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning.Adhering to a multi-task learning approach, StreamSpeech can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model.Experiments on CVSS benchmark demonstrate that StreamSpeech achieves state-of-the-art performance in both offline S2ST and Simul-S2ST tasks.Besides, StreamSpeech is able to present high-quality intermediate results (i.e., ASR or translation results) during simultaneous translation process, offering a more comprehensive real-time communication experience 1 . Shaolei Zhang 0001, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
ACL (1) | 3 |
| 2024 | Glancing Future for Simultaneous Machine TranslationabstractSimultaneous machine translation (SiMT) outputs translation while reading the source sentence. Unlike conventional sequence-to-sequence (seq2seq) training, existing SiMT methods adopt the prefix-to-prefix (prefix2prefix) training, where the model predicts target tokens based on partial source tokens. However, the prefix2prefix training diminishes the ability of the model to capture global information and introduces forced predictions due to the absence of essential source information. Consequently, it is crucial to bridge the gap between the prefix2prefix training and seq2seq training to enhance the translation capability of the SiMT model. In this paper, we propose a novel method that glances future in curriculum learning to achieve the transition from the seq2seq training to prefix2prefix training. Specifically, we gradually reduce the available source information from the whole sentence to the prefix corresponding to that latency. Our method is applicable to a wide range of SiMT methods and experiments demonstrate that our method outperforms strong baselines1. Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ICASSP | 1 |
| 2023 | Learning Optimal Policy for Simultaneous Machine Translation via Binary SearchabstractSimultaneous machine translation (SiMT) starts to output translation while reading the source sentence and needs a precise policy to decide when to output the generated translation.Therefore, the policy determines the number of source tokens read during the translation of each target token.However, it is difficult to learn a precise translation policy to achieve good latency-quality trade-offs, because there is no golden policy corresponding to parallel sentences as explicit supervision.In this paper, we present a new method for constructing the optimal policy online via binary search.By employing explicit supervision, our approach enables the SiMT model to learn the optimal policy, which can guide the model in completing the translation during inference.Experiments on four translation tasks show that our method can exceed strong baselines across all latency scenarios 1 Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 1 |
| 2023 | Non-autoregressive Streaming Transformer for Simultaneous TranslationabstractSimultaneous machine translation (SiMT) models are trained to strike a balance between latency and translation quality.However, training these models to achieve high quality while maintaining low latency often leads to a tendency for aggressive anticipation.We argue that such issue stems from the autoregressive architecture upon which most existing SiMT models are built.To address those issues, we propose non-autoregressive streaming Transformer (NAST) which comprises a unidirectional encoder and a non-autoregressive decoder with intra-chunk parallelism.We enable NAST to generate the blank token or repetitive tokens to adjust its READ/WRITE strategy flexibly, and train it to maximize the nonmonotonic latent alignment with an alignmentbased latency loss.Experiments on various SiMT benchmarks demonstrate that NAST outperforms previous strong autoregressive SiMT baselines. Zhengrui Ma, Shaolei Zhang 0001, Shoutao Guo, Chenze Shao, Min Zhang 0005, Yang Feng 0004 |
EMNLP | 3 |