EDBT 2026 Demo / reviewers in the wild / expert
Shaolei Zhang 0001
dblp:11/7954-1
· DBLP profile ↗
23ranked-venue papers
11as first author
22since 2021 · last 2025
0000-0002-7254-9380ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 11 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Language Models Are Read/Write Policy-Makers for Simultaneous GenerationabstractSimultaneous generation models write generation results while reading streaming inputs, necessitating a policy-maker to determine the appropriate output timing. Existing simultaneous generation methods generally adopt the traditional encoder-decoder architecture and learn the generation and policy-making capabilities through complex dynamic programming techniques. Although LLMs excel at text generation, they face challenges in taking on the role of policy-makers through traditional training methods, limiting their exploration in simultaneous generation. To overcome these limitations, we propose a novel LLM-driven Simultaneous Generation (LSG) framework, which allows the off-the-shelf LLM to decide the generation timing and produce output concurrently. Specifically, LSG selects the generation policy that minimizes latency as the baseline policy. Referring to the baseline policy, LSG enables the LLM to devise an improved generation policy that better balances latency and generation quality, and writes generation results accordingly. Experiments on simultaneous translation and streaming automatic speech recognition tasks show that our method can achieve state-of-the-art performance utilizing the open-source LLMs and demonstrate practicality in real-world scenarios. Shoutao Guo, Shaolei Zhang 0001, Zhengrui Ma, Yang Feng 0004 |
AAAI | 2 |
| 2025 | LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech SynthesisabstractReal-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction.Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language models (LLMs).In this paper, we introduce LLaMA-Omni 2, a series of speech language models (SpeechLMs) ranging from 0.5B to 14B parameters, capable of achieving highquality real-time speech interaction.LLaMA-Omni 2 is built upon the Qwen2.5 series models, integrating a speech encoder and an autoregressive streaming speech decoder.Despite being trained on only 200K multi-turn speech dialogue samples, LLaMA-Omni 2 demonstrates strong performance on several spoken question answering and speech instruction following benchmarks, surpassing previous state-of-theart SpeechLMs like GLM-4-Voice, which was trained on millions of hours of speech data. 1 Qingkai Fang, Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 4 |
| 2025 | AlignX: Advancing Multilingual Large Language Models with Multilingual Representation AlignmentabstractMultilingual large language models (LLMs) possess impressive multilingual understanding and generation capabilities.However, their performance and cross-lingual alignment often lag for non-dominant languages.A common solution is to fine-tune LLMs on largescale and more balanced multilingual corpora, but such approaches often lead to imprecise alignment and suboptimal knowledge transfer, struggling with limited improvements across languages.In this paper, we propose AlignX to bridge the multilingual performance gap, which is a two-stage representationlevel framework for enhancing multilingual performance of pre-trained LLMs.In the first stage, we align multilingual representations with multilingual semantic alignment and language feature integration.In the second stage, we stimulate the multilingual capability of LLMs via multilingual instruction fine-tuning.Experimental results on several pre-trained LLMs demonstrate that our approach enhances LLMs' multilingual general and cross-lingual generation capability.Further analysis indicates that AlignX brings the multilingual representations closer and improves the cross-lingual alignment.1 Mengyu Bu, Shaolei Zhang 0001, Zhongjun He, Hua Wu 0003, Yang Feng 0004 |
EMNLP | 2 |
| 2025 | IG-Pruning: Input-Guided Block Pruning for Large Language ModelsabstractWith the growing computational demands of large language models (LLMs), efficient inference has become increasingly critical for practical deployment. Depth pruning has emerged as a promising approach for reducing the computational costs of large language models by removing transformer layers. However, existing methods typically rely on fixed block masks, which can lead to suboptimal performance across different tasks and inputs. In this paper, we propose IG-Pruning, a novel input-aware block-wise pruning method that dynamically selects layer masks at inference time. Our approach consists of two stages: (1) Discovering diverse mask candidates through semantic clustering and L0 optimization, and (2) Implementing efficient dynamic pruning without the need for extensive training. Experimental results demonstrate that our method consistently outperforms state-of-the-art static depth pruning methods, making it particularly suitable for resource-constrained deployment scenarios. Kangyu Qiao, Shaolei Zhang 0001, Yang Feng 0004 |
EMNLP | 2 |
| 2025 | LLaMA-Omni: Seamless Speech Interaction with Large Language ModelsabstractModels like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs. To address this, we propose LLaMA-Omni, a novel model architecture designed for low-latency and high-quality speech interaction with LLMs. LLaMA-Omni integrates a pretrained speech encoder, a speech adaptor, an LLM, and a streaming speech decoder. It eliminates the need for speech transcription, and can simultaneously generate text and speech responses directly from speech instructions with extremely low latency. We build our model based on the latest Llama-3.1-8B-Instruct model. To align the model with speech interaction scenarios, we construct a dataset named InstructS2S-200K, which includes 200K speech instructions and corresponding speech responses. Experimental results show that compared to previous speech-language models, LLaMA-Omni provides better responses in both content and style, with a response latency as low as 226ms. Additionally, training LLaMA-Omni takes less than 3 days on just 4 GPUs, paving the way for the efficient development of speech-language models in the future. Qingkai Fang, Shoutao Guo, Zhengrui Ma, Shaolei Zhang 0001, Yang Feng 0004 |
ICLR | 5 |
| 2025 | LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision TokenabstractThe advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and integrate them and textual instructions into the context of large language models (LLMs), where large-scale parameters and numerous context tokens (predominantly vision tokens) result in substantial computational overhead. Previous efforts towards efficient LMMs always focus on replacing the LLM backbone with smaller models, while neglecting the crucial issue of token quantity. In this paper, we introduce LLaVA-Mini, an efficient LMM with minimal vision tokens. To achieve a high compression ratio of vision tokens while preserving visual information, we first analyze how LMMs understand vision tokens and find that most vision tokens only play a crucial role in the early layers of LLM backbone, where they mainly fuse visual information into text tokens. Building on this finding, LLaVA-Mini introduces modality pre-fusion to fuse visual information into text tokens in advance, thereby facilitating the extreme compression of vision tokens fed to LLM backbone into one token. LLaVA-Mini is a unified large multimodal model that can support the understanding of images, high-resolution images, and videos in an efficient manner. Experiments across 11 image-based and 7 video-based benchmarks demonstrate that LLaVA-Mini outperforms LLaVA-v1.5 with just 1 vision token instead of 576. Efficiency analyses reveal that LLaVA-Mini can reduce FLOPs by 77%, deliver low-latency responses within 40 milliseconds, and process over 10,000 frames of video on the GPU hardware with 24GB of memory. Shaolei Zhang 0001, Qingkai Fang, Yang Feng 0004 |
ICLR | 1 |
| 2025 | FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech ProcessingabstractThe rapid advancement of Large Language Models (LLMs) has spurred significant progress in Large Speech-Language Models (LSLMs), enhancing their capabilities in both speech understanding and generation. While existing LSLMs often concentrate on augmenting speech generation or tackling a diverse array of short-speech tasks, the efficient processing of long-form speech remains a critical yet underexplored challenge. This gap is primarily attributed to the scarcity of long-speech training datasets and the high computational costs associated with long sequences. To address these limitations, we introduce FastLongSpeech, a novel framework designed to extend LSLM capabilities for efficient long-speech processing without necessitating dedicated long-speech training data. FastLongSpeech incorporates an iterative fusion strategy that can compress excessively long-speech sequences into manageable lengths. To adapt LSLMs for long-speech inputs, it introduces a dynamic compression training approach, which exposes the model to short-speech sequences at varying compression ratios, thereby transferring the capabilities of LSLMs to long-speech tasks. To assess the long-speech capabilities of LSLMs, we develop a long-speech understanding benchmark called LongSpeech-Eval. Experiments show that our method exhibits strong performance in both long-speech and short-speech tasks, while greatly improving inference efficiency. Shoutao Guo, Shaolei Zhang 0001, Qingkai Fang, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
NeurIPS | 2 |
| 2024 | Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?abstractTwo-pass direct speech-to-speech translation (S2ST) models have shown promising results which decompose S2ST into speech-to-text translation (S2TT) and text-to-speech (TTS), yet conduct end-to-end training by sharing the target text representation between S2TT and TTS models.However, the training of these models still requires large-scale parallel speech data comprising triplets, which is extremely challenging to collect.On the other hand, S2TT and TTS have accumulated a large amount of data and numerous pretrained models, which can be used to reduce the reliance on parallel speech data.To this end, we propose a composite S2ST model named ComSpeech, which connects pretrained S2TT and TTS models by introducing a vocabulary adaptor based on connectionist temporal classification (CTC).The vocabulary adaptor is employed to adapt the output text sequence of S2TT to the input text sequence of TTS, which are different due to the use of different vocabularies.In this way, ComSpeech can still be trained end-to-end and only needs a small amount of parallel speech data to finetune.We further propose a novel training method ComSpeech-ZS to eliminate the reliance on parallel speech data by aligning the text representation space of S2TT and TTS.Experimental results on the CVSS dataset show that when the parallel speech data is available, ComSpeech surpasses previous two-pass models like UnitY and Translatotron 2 in both translation quality and decoding speed.When there is no parallel speech data, ComSpeech-ZS lags behind ComSpeech by only 0.7 ASR-BLEU and outperforms the cascaded models. 1 Qingkai Fang, Shaolei Zhang 0001, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
ACL (1) | 2 |
| 2024 | Decoder-only Streaming Transformer for Simultaneous TranslationabstractSimultaneous Machine Translation (SiMT) generates translation while reading source tokens, essentially producing the target prefix based on the source prefix.To achieve good performance, it leverages the relationship between source and target prefixes to exact a policy to guide the generation of translations.Although existing SiMT methods primarily focus on the Encoder-Decoder architecture, we explore the potential of Decoder-only architecture, owing to its superior performance in various tasks and its inherent compatibility with SiMT.However, directly applying the Decoder-only architecture to SiMT poses challenges in terms of training and inference.To alleviate the above problems, we propose the first Decoder-only SiMT model, named Decoder-only Streaming Transformer (DST).Specifically, DST separately encodes the positions of the source and target prefixes, ensuring that the position of the target prefix remains unaffected by the expansion of the source prefix.Furthermore, we propose a Streaming Self-Attention (SSA) mechanism tailored for the Decoder-only architecture.It is capable of obtaining translation policy by assessing the sufficiency of input source information and integrating with the soft-attention mechanism to generate translations.Experiments demonstrate that our approach achieves state-of-theart performance on three translation tasks 1 . Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 2 |
| 2024 | A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any TranslationabstractSimultaneous translation models play a crucial role in facilitating communication.However, existing research primarily focuses on text-totext or speech-to-text models, necessitating additional cascade components to achieve speechto-speech translation.These pipeline methods suffer from error propagation and accumulate delays in each cascade component, resulting in reduced synchronization between the speaker and listener.To overcome these challenges, we propose a novel non-autoregressive generation framework for simultaneous speech translation (NAST-S2x 1 ), which integrates speechto-text and speech-to-speech tasks into a unified end-to-end framework.We develop a nonautoregressive decoder capable of concurrently generating multiple text or acoustic unit tokens upon receiving fixed-length speech chunks.The decoder can generate blank or repeated tokens and employ CTC decoding to dynamically adjust its latency.Experimental results show that NAST-S2x outperforms state-of-theart models in both speech-to-text and speechto-speech tasks.It achieves high-quality simultaneous interpretation within a delay of less than 3 seconds and provides a 28× decoding speedup in offline generation. 2 Zhengrui Ma, Qingkai Fang, Shaolei Zhang 0001, Shoutao Guo, Yang Feng 0004, Min Zhang 0005 |
ACL (1) | 3 |
| 2024 | StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningabstractSimultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication.Beyond accomplishing translation between speech, Simul-S2ST requires a policy to control the model to generate corresponding target speech at the opportune moment within speech inputs, thereby posing a double challenge of translation and policy.In this paper, we propose StreamSpeech, a direct Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning.Adhering to a multi-task learning approach, StreamSpeech can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model.Experiments on CVSS benchmark demonstrate that StreamSpeech achieves state-of-the-art performance in both offline S2ST and Simul-S2ST tasks.Besides, StreamSpeech is able to present high-quality intermediate results (i.e., ASR or translation results) during simultaneous translation process, offering a more comprehensive real-time communication experience 1 . Shaolei Zhang 0001, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
ACL (1) | 1 |
| 2024 | TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceabstractLarge Language Models (LLMs) sometimes suffer from producing hallucinations, especially LLMs may generate untruthful responses despite knowing the correct knowledge.Activating the truthfulness within LLM is the key to fully unlocking LLM's knowledge potential.In this paper, we propose TruthX, an inference-time intervention method to activate the truthfulness of LLM by identifying and editing the features within LLM's internal representations that govern the truthfulness.TruthX employs an auto-encoder to map LLM's representations into semantic and truthful latent spaces respectively, and applies contrastive learning to identify a truthful editing direction within the truthful space.During inference, by editing LLM's internal representations in truthful space, TruthX effectively enhances the truthfulness of LLM.Experiments show that TruthX improves the truthfulness of 13 advanced LLMs by an average of 20% on Truth-fulQA benchmark.Further analyses suggest that TruthX can control LLM to produce truthful or hallucinatory responses via editing only one vector in LLM's internal representations 1 . Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 1 |
| 2024 | Glancing Future for Simultaneous Machine TranslationabstractSimultaneous machine translation (SiMT) outputs translation while reading the source sentence. Unlike conventional sequence-to-sequence (seq2seq) training, existing SiMT methods adopt the prefix-to-prefix (prefix2prefix) training, where the model predicts target tokens based on partial source tokens. However, the prefix2prefix training diminishes the ability of the model to capture global information and introduces forced predictions due to the absence of essential source information. Consequently, it is crucial to bridge the gap between the prefix2prefix training and seq2seq training to enhance the translation capability of the SiMT model. In this paper, we propose a novel method that glances future in curriculum learning to achieve the transition from the seq2seq training to prefix2prefix training. Specifically, we gradually reduce the available source information from the whole sentence to the prefix corresponding to that latency. Our method is applicable to a wide range of SiMT methods and experiments demonstrate that our method outperforms strong baselines1. Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ICASSP | 2 |
| 2023 | Learning Optimal Policy for Simultaneous Machine Translation via Binary SearchabstractSimultaneous machine translation (SiMT) starts to output translation while reading the source sentence and needs a precise policy to decide when to output the generated translation.Therefore, the policy determines the number of source tokens read during the translation of each target token.However, it is difficult to learn a precise translation policy to achieve good latency-quality trade-offs, because there is no golden policy corresponding to parallel sentences as explicit supervision.In this paper, we present a new method for constructing the optimal policy online via binary search.By employing explicit supervision, our approach enables the SiMT model to learn the optimal policy, which can guide the model in completing the translation during inference.Experiments on four translation tasks show that our method can exceed strong baselines across all latency scenarios 1 Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 2 |
| 2023 | Non-autoregressive Streaming Transformer for Simultaneous TranslationabstractSimultaneous machine translation (SiMT) models are trained to strike a balance between latency and translation quality.However, training these models to achieve high quality while maintaining low latency often leads to a tendency for aggressive anticipation.We argue that such issue stems from the autoregressive architecture upon which most existing SiMT models are built.To address those issues, we propose non-autoregressive streaming Transformer (NAST) which comprises a unidirectional encoder and a non-autoregressive decoder with intra-chunk parallelism.We enable NAST to generate the blank token or repetitive tokens to adjust its READ/WRITE strategy flexibly, and train it to maximize the nonmonotonic latent alignment with an alignmentbased latency loss.Experiments on various SiMT benchmarks demonstrate that NAST outperforms previous strong autoregressive SiMT baselines. Zhengrui Ma, Shaolei Zhang 0001, Shoutao Guo, Chenze Shao, Min Zhang 0005, Yang Feng 0004 |
EMNLP | 2 |
| 2023 | Hidden Markov Transformer for Simultaneous Machine Translation
Shaolei Zhang 0001, Yang Feng 0004 |
ICLR | 1 |
| 2023 | Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationabstractSimultaneous sequence generation is a pivotal task for real-time scenarios, such as streaming speech recognition, simultaneous machine translation and simultaneous speech translation, where the target sequence is generated while receiving the source sequence. The crux of achieving high-quality generation with low latency lies in identifying the optimal moments for generating, accomplished by learning a mapping between the source and target sequences. However, existing methods often rely on task-specific heuristics for different sequence types, limiting the model’s capacity to adaptively learn the source-target mapping and hindering the exploration of multi-task learning for various simultaneous tasks. In this paper, we propose a unified segment-to-segment framework (Seg2Seg) for simultaneous sequence generation, which learns the mapping in an adaptive and unified manner. During the process of simultaneous generation, the model alternates between waiting for a source segment and generating a target segment, making the segment serve as the natural bridge between the source and target. To accomplish this, Seg2Seg introduces a latent segment as the pivot between source to target and explores all potential source-target mappings via the proposed expectation training, thereby learning the optimal moments for generating. Experiments on multiple simultaneous generation tasks demonstrate that Seg2Seg achieves state-of-the-art performance and exhibits better generality across various tasks. Shaolei Zhang 0001, Yang Feng 0004 |
NeurIPS | 1 |
| 2022 | Modeling Dual Read/Write Paths for Simultaneous Machine TranslationabstractSimultaneous machine translation (SiMT) outputs translation while reading source sentence and hence requires a policy to decide whether to wait for the next source word (READ) or generate a target word (WRITE), the actions of which form a read/write path.Although the read/write path is essential to SiMT performance, no direct supervision is given to the path in the existing methods.In this paper, we propose a method of dual-path SiMT which introduces duality constraints to direct the read/write path.According to duality constraints, the read/write path in source-totarget and target-to-source SiMT models can be mapped to each other.As a result, the two SiMT models can be optimized jointly by forcing their read/write paths to satisfy the mapping.Experiments on En↔Vi and De↔En tasks show that our method can outperform strong baselines under all latency. Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 1 |
| 2022 | Reducing Position Bias in Simultaneous Machine Translation with Length-Aware FrameworkabstractSimultaneous machine translation (SiMT) starts translating while receiving the streaming source inputs, and hence the source sentence is always incomplete during translating.Different from the full-sentence MT using the conventional seq-to-seq architecture, SiMT often applies prefix-to-prefix architecture, which forces each target word to only align with a partial source prefix to adapt to the incomplete source in streaming inputs.However, the source words in the front positions are always illusoryly considered more important since they appear in more prefixes, resulting in position bias, which makes the model pay more attention on the front source positions in testing.In this paper, we first analyze the phenomenon of position bias in SiMT, and develop a Length-Aware Framework to reduce the position bias by bridging the structural gap between SiMT and fullsentence MT.Specifically, given the streaming inputs, we first predict the full-sentence length and then fill the future source position with positional encoding, thereby turning the streaming inputs into a pseudo full-sentence.The proposed framework can be integrated into most existing SiMT methods to further improve performance.Experiments on two representative SiMT methods, including the state-of-theart adaptive policy, show that our method successfully reduces the position bias and thereby achieves better SiMT performance. Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 1 |
| 2022 | Information-Transport-based Policy for Simultaneous TranslationabstractSimultaneous translation (ST) outputs translation while receiving the source inputs, and hence requires a policy to determine whether to translate a target token or wait for the next source token.The major challenge of ST is that each target token can only be translated based on the current received source tokens, where the received source information will directly affect the translation quality.So naturally, how much source information is received for the translation of the current target token is supposed to be the pivotal evidence for the ST policy to decide between translating and waiting.In this paper, we treat the translation as information transport from source to target and accordingly propose an Information-Transport-based Simultaneous Translation (ITST).ITST quantifies the transported information weight from each source token to the current target token, and then decides whether to translate the target token according to its accumulated received information.Experiments on both text-to-text ST and speech-to-text ST (a.k.a., streaming speech translation) tasks show that ITST outperforms strong baselines and achieves state-of-the-art performance 1 . Shaolei Zhang 0001, Yang Feng 0004 |
EMNLP | 1 |
| 2021 | Future-Guided Incremental Transformer for Simultaneous TranslationabstractSimultaneous translation (ST) starts translations synchronously while reading source sentences, and is used in many online scenarios. The previous wait-k policy is concise and achieved good results in ST. However, wait-k policy faces two weaknesses: low training speed caused by the recalculation of hidden states and lack of future source information to guide training. For the low training speed, we propose an incremental Transformer with an average embedding layer (AEL) to accelerate the speed of calculation of the hidden states during training. For future-guided training, we propose a conventional Transformer as the teacher of the incremental Transformer, and try to invisibly embed some future information in the model through knowledge distillation. We conducted experiments on Chinese-English and German-English simultaneous translation tasks and compared with the wait-k policy to evaluate the proposed method. Our method can effectively increase the training speed by about 28 times on average at different k and implicitly embed some predictive abilities in the model, achieving better translation quality than wait-k baseline. Shaolei Zhang 0001, Yang Feng 0004, Liangyou Li |
AAAI | 1 |
| 2021 | Universal Simultaneous Machine Translation with Mixture-of-Experts Wait-k PolicyabstractSimultaneous machine translation (SiMT) generates translation before reading the entire source sentence and hence it has to trade off between translation quality and latency.To fulfill the requirements of different translation quality and latency in practical applications, the previous methods usually need to train multiple SiMT models for different latency levels, resulting in large computational costs.In this paper, we propose a universal SiMT model with Mixture-of-Experts Wait-k Policy to achieve the best translation quality under arbitrary latency with only one trained model.Specifically, our method employs multi-head attention to accomplish the mixture of experts where each head is treated as a wait-k expert with its own waiting words number, and given a test latency and source inputs, the weights of the experts are accordingly adjusted to produce the best translation.Experiments on three datasets show that our method outperforms all the strong baselines under different latency, including the state-of-the-art adaptive policy. Shaolei Zhang 0001, Yang Feng 0004 |
EMNLP (1) | 1 |
| 2019 | Opinion Knowledge Injection Network for Aspect Extraction
Shaolei Zhang 0001, Kai Shuang |
ICONIP (2) | 1 |