Shoutao Guo

dblp:331/5767 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2025
0009-0006-1662-7504ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Machine translation · 37% Speech recognition and synthesis · 33% Deep learning architectures and training · 9%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
simultaneous machine translation
3.752025
Large Language Models Are Read/Write Policy-Makers for Simultaneous Generation · AAAI 2025
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning · ACL (1) 2024
Decoder-only Streaming Transformer for Simultaneous Translation · ACL (1) 2024
Natural language and speech › Machine translation › speech translation
speech-to-speech translation
1.522024
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning · ACL (1) 2024
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation · ACL (1) 2024
Machine learning › Deep learning architectures and training
transformer
1.422024
Decoder-only Streaming Transformer for Simultaneous Translation · ACL (1) 2024
Non-autoregressive Streaming Transformer for Simultaneous Translation · EMNLP 2023
Natural language and speech › Speech recognition and synthesis
speech language model
1.122025
LLaMA-Omni: Seamless Speech Interaction with Large Language Models · ICLR 2025
LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis · ACL (1) 2025
Natural language and speech › Speech recognition and synthesis
speech interaction
0.912025
LLaMA-Omni: Seamless Speech Interaction with Large Language Models · ICLR 2025
Natural language and speech › Speech recognition and synthesis › speech language model
speech large language model
0.912025
FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing · NeurIPS 2025
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.912025
LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems
0.912025
LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis · ACL (1) 2025
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › real-time speech recognition
streaming speech recognition
0.912025
Large Language Models Are Read/Write Policy-Makers for Simultaneous Generation · AAAI 2025
Natural language and speech › Speech recognition and synthesis › speech synthesis
non-autoregressive speech generation
0.812024
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation · ACL (1) 2024
Natural language and speech › Machine translation › simultaneous machine translation
simultaneous speech translation
0.812024
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation · ACL (1) 2024
Machine learning › Generative modeling › transformer-based generative model
non-autoregressive transformer
0.712023
Non-autoregressive Streaming Transformer for Simultaneous Translation · EMNLP 2023
Machine learning › Reinforcement learning
policy learning
0.712023
Learning Optimal Policy for Simultaneous Machine Translation via Binary Search · ACL (1) 2023
Natural language and speech › Language models and text generation
large language model
0.312025
FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

speech encoder · 1.7large language model · 1.7non-autoregressive decoding · 1.4speech adaptor · 0.9policy-making · 0.9iterative fusion · 0.9dynamic programming · 0.9dynamic compression training · 0.9autoregressive streaming speech decoder · 0.9streaming self-attention · 0.8
YearPublicationVenuePosition
2025 Large Language Models Are Read/Write Policy-Makers for Simultaneous Generation
abstract
Simultaneous generation models write generation results while reading streaming inputs, necessitating a policy-maker to determine the appropriate output timing. Existing simultaneous generation methods generally adopt the traditional encoder-decoder architecture and learn the generation and policy-making capabilities through complex dynamic programming techniques. Although LLMs excel at text generation, they face challenges in taking on the role of policy-makers through traditional training methods, limiting their exploration in simultaneous generation. To overcome these limitations, we propose a novel LLM-driven Simultaneous Generation (LSG) framework, which allows the off-the-shelf LLM to decide the generation timing and produce output concurrently. Specifically, LSG selects the generation policy that minimizes latency as the baseline policy. Referring to the baseline policy, LSG enables the LLM to devise an improved generation policy that better balances latency and generation quality, and writes generation results accordingly. Experiments on simultaneous translation and streaming automatic speech recognition tasks show that our method can achieve state-of-the-art performance utilizing the open-source LLMs and demonstrate practicality in real-world scenarios.
Shoutao Guo, Shaolei Zhang 0001, Zhengrui Ma, Yang Feng 0004
AAAI1
2025 LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
abstract
Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction.Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language models (LLMs).In this paper, we introduce LLaMA-Omni 2, a series of speech language models (SpeechLMs) ranging from 0.5B to 14B parameters, capable of achieving highquality real-time speech interaction.LLaMA-Omni 2 is built upon the Qwen2.5 series models, integrating a speech encoder and an autoregressive streaming speech decoder.Despite being trained on only 200K multi-turn speech dialogue samples, LLaMA-Omni 2 demonstrates strong performance on several spoken question answering and speech instruction following benchmarks, surpassing previous state-of-theart SpeechLMs like GLM-4-Voice, which was trained on millions of hours of speech data. 1
Qingkai Fang, Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004
ACL (1)3
2025 LLaMA-Omni: Seamless Speech Interaction with Large Language Models
abstract
Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs. To address this, we propose LLaMA-Omni, a novel model architecture designed for low-latency and high-quality speech interaction with LLMs. LLaMA-Omni integrates a pretrained speech encoder, a speech adaptor, an LLM, and a streaming speech decoder. It eliminates the need for speech transcription, and can simultaneously generate text and speech responses directly from speech instructions with extremely low latency. We build our model based on the latest Llama-3.1-8B-Instruct model. To align the model with speech interaction scenarios, we construct a dataset named InstructS2S-200K, which includes 200K speech instructions and corresponding speech responses. Experimental results show that compared to previous speech-language models, LLaMA-Omni provides better responses in both content and style, with a response latency as low as 226ms. Additionally, training LLaMA-Omni takes less than 3 days on just 4 GPUs, paving the way for the efficient development of speech-language models in the future.
Qingkai Fang, Shoutao Guo, Zhengrui Ma, Shaolei Zhang 0001, Yang Feng 0004
ICLR2
2025 FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing
abstract
The rapid advancement of Large Language Models (LLMs) has spurred significant progress in Large Speech-Language Models (LSLMs), enhancing their capabilities in both speech understanding and generation. While existing LSLMs often concentrate on augmenting speech generation or tackling a diverse array of short-speech tasks, the efficient processing of long-form speech remains a critical yet underexplored challenge. This gap is primarily attributed to the scarcity of long-speech training datasets and the high computational costs associated with long sequences. To address these limitations, we introduce FastLongSpeech, a novel framework designed to extend LSLM capabilities for efficient long-speech processing without necessitating dedicated long-speech training data. FastLongSpeech incorporates an iterative fusion strategy that can compress excessively long-speech sequences into manageable lengths. To adapt LSLMs for long-speech inputs, it introduces a dynamic compression training approach, which exposes the model to short-speech sequences at varying compression ratios, thereby transferring the capabilities of LSLMs to long-speech tasks. To assess the long-speech capabilities of LSLMs, we develop a long-speech understanding benchmark called LongSpeech-Eval. Experiments show that our method exhibits strong performance in both long-speech and short-speech tasks, while greatly improving inference efficiency.
Shoutao Guo, Shaolei Zhang 0001, Qingkai Fang, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004
NeurIPS1
2024 Decoder-only Streaming Transformer for Simultaneous Translation
abstract
Simultaneous Machine Translation (SiMT) generates translation while reading source tokens, essentially producing the target prefix based on the source prefix.To achieve good performance, it leverages the relationship between source and target prefixes to exact a policy to guide the generation of translations.Although existing SiMT methods primarily focus on the Encoder-Decoder architecture, we explore the potential of Decoder-only architecture, owing to its superior performance in various tasks and its inherent compatibility with SiMT.However, directly applying the Decoder-only architecture to SiMT poses challenges in terms of training and inference.To alleviate the above problems, we propose the first Decoder-only SiMT model, named Decoder-only Streaming Transformer (DST).Specifically, DST separately encodes the positions of the source and target prefixes, ensuring that the position of the target prefix remains unaffected by the expansion of the source prefix.Furthermore, we propose a Streaming Self-Attention (SSA) mechanism tailored for the Decoder-only architecture.It is capable of obtaining translation policy by assessing the sufficiency of input source information and integrating with the soft-attention mechanism to generate translations.Experiments demonstrate that our approach achieves state-of-theart performance on three translation tasks 1 .
Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004
ACL (1)1
2024 A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation
abstract
Simultaneous translation models play a crucial role in facilitating communication.However, existing research primarily focuses on text-totext or speech-to-text models, necessitating additional cascade components to achieve speechto-speech translation.These pipeline methods suffer from error propagation and accumulate delays in each cascade component, resulting in reduced synchronization between the speaker and listener.To overcome these challenges, we propose a novel non-autoregressive generation framework for simultaneous speech translation (NAST-S2x 1 ), which integrates speechto-text and speech-to-speech tasks into a unified end-to-end framework.We develop a nonautoregressive decoder capable of concurrently generating multiple text or acoustic unit tokens upon receiving fixed-length speech chunks.The decoder can generate blank or repeated tokens and employ CTC decoding to dynamically adjust its latency.Experimental results show that NAST-S2x outperforms state-of-theart models in both speech-to-text and speechto-speech tasks.It achieves high-quality simultaneous interpretation within a delay of less than 3 seconds and provides a 28× decoding speedup in offline generation. 2
Zhengrui Ma, Qingkai Fang, Shaolei Zhang 0001, Shoutao Guo, Yang Feng 0004, Min Zhang 0005
ACL (1)4
2024 StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
abstract
Simultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication.Beyond accomplishing translation between speech, Simul-S2ST requires a policy to control the model to generate corresponding target speech at the opportune moment within speech inputs, thereby posing a double challenge of translation and policy.In this paper, we propose StreamSpeech, a direct Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning.Adhering to a multi-task learning approach, StreamSpeech can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model.Experiments on CVSS benchmark demonstrate that StreamSpeech achieves state-of-the-art performance in both offline S2ST and Simul-S2ST tasks.Besides, StreamSpeech is able to present high-quality intermediate results (i.e., ASR or translation results) during simultaneous translation process, offering a more comprehensive real-time communication experience 1 .
Shaolei Zhang 0001, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004
ACL (1)3
2024 Glancing Future for Simultaneous Machine Translation
abstract
Simultaneous machine translation (SiMT) outputs translation while reading the source sentence. Unlike conventional sequence-to-sequence (seq2seq) training, existing SiMT methods adopt the prefix-to-prefix (prefix2prefix) training, where the model predicts target tokens based on partial source tokens. However, the prefix2prefix training diminishes the ability of the model to capture global information and introduces forced predictions due to the absence of essential source information. Consequently, it is crucial to bridge the gap between the prefix2prefix training and seq2seq training to enhance the translation capability of the SiMT model. In this paper, we propose a novel method that glances future in curriculum learning to achieve the transition from the seq2seq training to prefix2prefix training. Specifically, we gradually reduce the available source information from the whole sentence to the prefix corresponding to that latency. Our method is applicable to a wide range of SiMT methods and experiments demonstrate that our method outperforms strong baselines1.
Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004
ICASSP1
2023 Learning Optimal Policy for Simultaneous Machine Translation via Binary Search
abstract
Simultaneous machine translation (SiMT) starts to output translation while reading the source sentence and needs a precise policy to decide when to output the generated translation.Therefore, the policy determines the number of source tokens read during the translation of each target token.However, it is difficult to learn a precise translation policy to achieve good latency-quality trade-offs, because there is no golden policy corresponding to parallel sentences as explicit supervision.In this paper, we present a new method for constructing the optimal policy online via binary search.By employing explicit supervision, our approach enables the SiMT model to learn the optimal policy, which can guide the model in completing the translation during inference.Experiments on four translation tasks show that our method can exceed strong baselines across all latency scenarios 1
Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004
ACL (1)1
2023 Non-autoregressive Streaming Transformer for Simultaneous Translation
abstract
Simultaneous machine translation (SiMT) models are trained to strike a balance between latency and translation quality.However, training these models to achieve high quality while maintaining low latency often leads to a tendency for aggressive anticipation.We argue that such issue stems from the autoregressive architecture upon which most existing SiMT models are built.To address those issues, we propose non-autoregressive streaming Transformer (NAST) which comprises a unidirectional encoder and a non-autoregressive decoder with intra-chunk parallelism.We enable NAST to generate the blank token or repetitive tokens to adjust its READ/WRITE strategy flexibly, and train it to maximize the nonmonotonic latent alignment with an alignmentbased latency loss.Experiments on various SiMT benchmarks demonstrate that NAST outperforms previous strong autoregressive SiMT baselines.
Zhengrui Ma, Shaolei Zhang 0001, Shoutao Guo, Chenze Shao, Min Zhang 0005, Yang Feng 0004
EMNLP3