Qingkai Fang

dblp:301/3107 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0001-8575-591XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 8 first-author · 16 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
abstract
Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction.Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language models (LLMs).In this paper, we introduce LLaMA-Omni 2, a series of speech language models (SpeechLMs) ranging from 0.5B to 14B parameters, capable of achieving highquality real-time speech interaction.LLaMA-Omni 2 is built upon the Qwen2.5 series models, integrating a speech encoder and an autoregressive streaming speech decoder.Despite being trained on only 200K multi-turn speech dialogue samples, LLaMA-Omni 2 demonstrates strong performance on several spoken question answering and speech instruction following benchmarks, surpassing previous state-of-theart SpeechLMs like GLM-4-Voice, which was trained on millions of hours of speech data. 1
Qingkai Fang, Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004
ACL (1)1
2025 LLaMA-Omni: Seamless Speech Interaction with Large Language Models
abstract
Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs. To address this, we propose LLaMA-Omni, a novel model architecture designed for low-latency and high-quality speech interaction with LLMs. LLaMA-Omni integrates a pretrained speech encoder, a speech adaptor, an LLM, and a streaming speech decoder. It eliminates the need for speech transcription, and can simultaneously generate text and speech responses directly from speech instructions with extremely low latency. We build our model based on the latest Llama-3.1-8B-Instruct model. To align the model with speech interaction scenarios, we construct a dataset named InstructS2S-200K, which includes 200K speech instructions and corresponding speech responses. Experimental results show that compared to previous speech-language models, LLaMA-Omni provides better responses in both content and style, with a response latency as low as 226ms. Additionally, training LLaMA-Omni takes less than 3 days on just 4 GPUs, paving the way for the efficient development of speech-language models in the future.
Qingkai Fang, Shoutao Guo, Zhengrui Ma, Shaolei Zhang 0001, Yang Feng 0004
ICLR1
2025 LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
abstract
The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and integrate them and textual instructions into the context of large language models (LLMs), where large-scale parameters and numerous context tokens (predominantly vision tokens) result in substantial computational overhead. Previous efforts towards efficient LMMs always focus on replacing the LLM backbone with smaller models, while neglecting the crucial issue of token quantity. In this paper, we introduce LLaVA-Mini, an efficient LMM with minimal vision tokens. To achieve a high compression ratio of vision tokens while preserving visual information, we first analyze how LMMs understand vision tokens and find that most vision tokens only play a crucial role in the early layers of LLM backbone, where they mainly fuse visual information into text tokens. Building on this finding, LLaVA-Mini introduces modality pre-fusion to fuse visual information into text tokens in advance, thereby facilitating the extreme compression of vision tokens fed to LLM backbone into one token. LLaVA-Mini is a unified large multimodal model that can support the understanding of images, high-resolution images, and videos in an efficient manner. Experiments across 11 image-based and 7 video-based benchmarks demonstrate that LLaVA-Mini outperforms LLaVA-v1.5 with just 1 vision token instead of 576. Efficiency analyses reveal that LLaVA-Mini can reduce FLOPs by 77%, deliver low-latency responses within 40 milliseconds, and process over 10,000 frames of video on the GPU hardware with 24GB of memory.
Shaolei Zhang 0001, Qingkai Fang, Yang Feng 0004
ICLR2
2025 FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing
abstract
The rapid advancement of Large Language Models (LLMs) has spurred significant progress in Large Speech-Language Models (LSLMs), enhancing their capabilities in both speech understanding and generation. While existing LSLMs often concentrate on augmenting speech generation or tackling a diverse array of short-speech tasks, the efficient processing of long-form speech remains a critical yet underexplored challenge. This gap is primarily attributed to the scarcity of long-speech training datasets and the high computational costs associated with long sequences. To address these limitations, we introduce FastLongSpeech, a novel framework designed to extend LSLM capabilities for efficient long-speech processing without necessitating dedicated long-speech training data. FastLongSpeech incorporates an iterative fusion strategy that can compress excessively long-speech sequences into manageable lengths. To adapt LSLMs for long-speech inputs, it introduces a dynamic compression training approach, which exposes the model to short-speech sequences at varying compression ratios, thereby transferring the capabilities of LSLMs to long-speech tasks. To assess the long-speech capabilities of LSLMs, we develop a long-speech understanding benchmark called LongSpeech-Eval. Experiments show that our method exhibits strong performance in both long-speech and short-speech tasks, while greatly improving inference efficiency.
Shoutao Guo, Shaolei Zhang 0001, Qingkai Fang, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004
NeurIPS3
2024 Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
abstract
Two-pass direct speech-to-speech translation (S2ST) models have shown promising results which decompose S2ST into speech-to-text translation (S2TT) and text-to-speech (TTS), yet conduct end-to-end training by sharing the target text representation between S2TT and TTS models.However, the training of these models still requires large-scale parallel speech data comprising triplets, which is extremely challenging to collect.On the other hand, S2TT and TTS have accumulated a large amount of data and numerous pretrained models, which can be used to reduce the reliance on parallel speech data.To this end, we propose a composite S2ST model named ComSpeech, which connects pretrained S2TT and TTS models by introducing a vocabulary adaptor based on connectionist temporal classification (CTC).The vocabulary adaptor is employed to adapt the output text sequence of S2TT to the input text sequence of TTS, which are different due to the use of different vocabularies.In this way, ComSpeech can still be trained end-to-end and only needs a small amount of parallel speech data to finetune.We further propose a novel training method ComSpeech-ZS to eliminate the reliance on parallel speech data by aligning the text representation space of S2TT and TTS.Experimental results on the CVSS dataset show that when the parallel speech data is available, ComSpeech surpasses previous two-pass models like UnitY and Translatotron 2 in both translation quality and decoding speed.When there is no parallel speech data, ComSpeech-ZS lags behind ComSpeech by only 0.7 ASR-BLEU and outperforms the cascaded models. 1
Qingkai Fang, Shaolei Zhang 0001, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004
ACL (1)1
2024 A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation
abstract
Simultaneous translation models play a crucial role in facilitating communication.However, existing research primarily focuses on text-totext or speech-to-text models, necessitating additional cascade components to achieve speechto-speech translation.These pipeline methods suffer from error propagation and accumulate delays in each cascade component, resulting in reduced synchronization between the speaker and listener.To overcome these challenges, we propose a novel non-autoregressive generation framework for simultaneous speech translation (NAST-S2x 1 ), which integrates speechto-text and speech-to-speech tasks into a unified end-to-end framework.We develop a nonautoregressive decoder capable of concurrently generating multiple text or acoustic unit tokens upon receiving fixed-length speech chunks.The decoder can generate blank or repeated tokens and employ CTC decoding to dynamically adjust its latency.Experimental results show that NAST-S2x outperforms state-of-theart models in both speech-to-text and speechto-speech tasks.It achieves high-quality simultaneous interpretation within a delay of less than 3 seconds and provides a 28× decoding speedup in offline generation. 2
Zhengrui Ma, Qingkai Fang, Shaolei Zhang 0001, Shoutao Guo, Yang Feng 0004, Min Zhang 0005
ACL (1)2
2024 StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
abstract
Simultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication.Beyond accomplishing translation between speech, Simul-S2ST requires a policy to control the model to generate corresponding target speech at the opportune moment within speech inputs, thereby posing a double challenge of translation and policy.In this paper, we propose StreamSpeech, a direct Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning.Adhering to a multi-task learning approach, StreamSpeech can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model.Experiments on CVSS benchmark demonstrate that StreamSpeech achieves state-of-the-art performance in both offline S2ST and Simul-S2ST tasks.Besides, StreamSpeech is able to present high-quality intermediate results (i.e., ASR or translation results) during simultaneous translation process, offering a more comprehensive real-time communication experience 1 .
Shaolei Zhang 0001, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004
ACL (1)2
2023 Back Translation for Speech-to-text Translation Without Transcripts
abstract
The success of end-to-end speech-to-text translation (ST) is often achieved by utilizing source transcripts, e.g., by pre-training with automatic speech recognition (ASR) and machine translation (MT) tasks, or by introducing additional ASR and MT data.Unfortunately, transcripts are only sometimes available since numerous unwritten languages exist worldwide.In this paper, we aim to utilize large amounts of targetside monolingual data to enhance ST without transcripts.Motivated by the remarkable success of back translation in MT, we develop a back translation algorithm for ST (BT4ST) to synthesize pseudo ST data from monolingual target data.To ease the challenges posed by short-to-long generation and one-to-many mapping, we introduce self-supervised discrete units and achieve back translation by cascading a target-to-unit model and a unit-to-speech model.With our synthetic ST data, we achieve an average boost of 2.3 BLEU on MuST-C En→De, En→Fr, and En→Es datasets.More experiments show that our method is especially effective in low-resource scenarios.12
Qingkai Fang, Yang Feng 0004
ACL (1)1
2023 Understanding and Bridging the Modality Gap for Speech Translation
abstract
How to achieve better end-to-end speech translation (ST) by leveraging (text) machine translation (MT) data?Among various existing techniques, multi-task learning is one of the effective ways to share knowledge between ST and MT in which additional MT data can help to learn source-to-target mapping.However, due to the differences between speech and text, there is always a gap between ST and MT.In this paper, we first aim to understand this modality gap from the target-side representation differences, and link the modality gap to another well-known problem in neural machine translation: exposure bias.We find that the modality gap is relatively small during training except for some difficult cases, but keeps increasing during inference due to the cascading effect.To address these problems, we propose the Cross-modal Regularization with Scheduled Sampling (CRESS) method.Specifically, we regularize the output predictions of ST and MT, whose target-side contexts are derived by sampling between ground truth words and self-generated words with a varying probability.Furthermore, we introduce token-level adaptive training which assigns different training weights to target tokens to handle difficult cases with large modality gaps.Experiments and analysis show that our approach effectively bridges the modality gap, and achieves promising results in all eight directions of the MuST-C dataset. 1
Qingkai Fang, Yang Feng 0004
ACL (1)1
2023 CMOT: Cross-modal Mixup via Optimal Transport for Speech Translation
abstract
End-to-end speech translation (ST) is the task of translating speech signals in the source language into text in the target language.As a cross-modal task, end-to-end ST is difficult to train with limited data.Existing methods often try to transfer knowledge from machine translation (MT), but their performances are restricted by the modality gap between speech and text.In this paper, we propose Cross-modal Mixup via Optimal Transport (CMOT) to overcome the modality gap.We find the alignment between speech and text sequences via optimal transport and then mix up the sequences from different modalities at a token level using the alignment.Experiments on the MuST-C ST benchmark demonstrate that CMOT achieves an average BLEU of 30.0 in 8 translation directions, outperforming previous methods.Further analysis shows CMOT can adaptively find the alignment between modalities, which helps alleviate the modality gap between speech and text.
Qingkai Fang, Yang Feng 0004
ACL (1)2
2023 Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine Translation
abstract
Multimodal machine translation (MMT) simultaneously takes the source sentence and a relevant image as input for translation.Since there is no paired image available for the input sentence in most cases, recent studies suggest utilizing powerful text-to-image generation models to provide image inputs.Nevertheless, synthetic images generated by these models often follow different distributions compared to authentic images.Consequently, using authentic images for training and synthetic images for inference can introduce a distribution shift, resulting in performance degradation during inference.To tackle this challenge, in this paper, we feed synthetic and authentic images to the MMT model, respectively.Then we minimize the gap between the synthetic and authentic images by drawing close the input image representations of the Transformer Encoder and the output distributions of the Transformer Decoder.Therefore, we mitigate the distribution disparity introduced by the synthetic images during inference, thereby freeing the authentic images from the inference process.Experimental results show that our approach achieves state-ofthe-art performance on the Multi30K En-De and En-Fr datasets, while remaining independent of authentic images during inference.
Wenyu Guo, Qingkai Fang, Dong Yu 0003, Yang Feng 0004
EMNLP2
2023 DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translation
abstract
Direct speech-to-speech translation (S2ST) translates speech from one language into another using a single model. However, due to the presence of linguistic and acoustic diversity, the target speech follows a complex multimodal distribution, posing challenges to achieving both high-quality translations and fast decoding speeds for S2ST models. In this paper, we propose DASpeech, a non-autoregressive direct S2ST model which realizes both fast and high-quality S2ST. To better capture the complex distribution of the target speech, DASpeech adopts the two-pass architecture to decompose the generation process into two steps, where a linguistic decoder first generates the target text, and an acoustic decoder then generates the target speech based on the hidden states of the linguistic decoder. Specifically, we use the decoder of DA-Transformer as the linguistic decoder, and use FastSpeech 2 as the acoustic decoder. DA-Transformer models translations with a directed acyclic graph (DAG). To consider all potential paths in the DAG during training, we calculate the expected hidden states for each target token via dynamic programming, and feed them into the acoustic decoder to predict the target mel-spectrogram. During inference, we select the most probable path and take hidden states on that path as input to the acoustic decoder. Experiments on the CVSS Fr$\rightarrow$En benchmark demonstrate that DASpeech can achieve comparable or even better performance than the state-of-the-art S2ST model Translatotron 2, while preserving up to 18.53$\times$ speedup compared to the autoregressive baseline. Compared with the previous non-autoregressive S2ST model, DASpeech does not rely on knowledge distillation and iterative decoding, achieving significant improvements in both translation quality and decoding speed. Furthermore, DASpeech shows the ability to preserve the speaker's voice of the source speech during translation.
Qingkai Fang, Yang Feng 0004
NeurIPS1
2022 Neural Machine Translation with Phrase-Level Universal Visual Representations
abstract
Multimodal machine translation (MMT) aims to improve neural machine translation (NMT) with additional visual information, but most existing MMT methods require paired input of source sentence and image, which makes them suffer from shortage of sentence-image pairs.In this paper, we propose a phrase-level retrieval-based method for MMT to get visual information for the source input from existing sentence-image data sets so that MMT can break the limitation of paired sentence-image input.Our method performs retrieval at the phrase level and hence learns visual information from pairs of source phrase and grounded region, which can mitigate data sparsity.Furthermore, our method employs the conditional variational auto-encoder to learn visual representations which can filter redundant visual information and only retain visual information related to the phrase.Experiments show that the proposed method significantly outperforms strong baselines on multiple MMT datasets, especially when the textual context is limited.
Qingkai Fang, Yang Feng 0004
ACL (1)1
2022 STEMM: Self-learning with Speech-text Manifold Mixup for Speech Translation
abstract
How to learn a better speech representation for end-to-end speech-to-text translation (ST) with limited labeled data?Existing techniques often attempt to transfer powerful machine translation (MT) capabilities to ST, but neglect the representation discrepancy across modalities.In this paper, we propose the Speech-TExt Manifold Mixup (STEMM) method to calibrate such discrepancy.Specifically, we mix up the representation sequences of different modalities, and take both unimodal speech sequences and multimodal mixed sequences as input to the translation model in parallel, and regularize their output predictions with a selflearning framework.Experiments on MuST-C speech translation benchmark and further analysis show that our method effectively alleviates the cross-modal representation discrepancy, and achieves significant improvements over a strong baseline on eight translation directions.* indicates corresponding authors.
Qingkai Fang, Rong Ye, Lei Li 0005, Yang Feng 0004, Mingxuan Wang
ACL (1)1
2022 Low-resource Neural Machine Translation with Cross-modal Alignment
abstract
How to achieve neural machine translation with limited parallel data?Existing techniques often rely on large-scale monolingual corpora, which is impractical for some low-resource languages.In this paper, we turn to connect several low-resource languages to a particular high-resource one by additional visual modality.Specifically, we propose a cross-modal contrastive learning method to learn a shared space for all languages, where both a coarsegrained sentence-level objective and a finegrained token-level one are introduced.Experimental results and further analysis show that our method can effectively learn the crossmodal and cross-lingual alignment with a small amount of image-text pairs and achieves significant improvements over the text-only baseline under both zero-shot and few-shot scenarios.Our code could be found at https: //github.com/ictnlp/LNMT-CA.
Qingkai Fang, Yang Feng 0004
EMNLP2
2021 Geometric Object 3D Reconstruction from Single Line Drawing Image Based on a Network for Classification and Sketch Extraction
Zhuoying Wang, Qingkai Fang, Yongtao Wang
ICDAR (1)2