Biao Fu

dblp:144/8117 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-4532-7432ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021
YearPublicationVenuePosition
2026 Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture
abstract
Simultaneous speech translation (SimulST) produces translations incrementally while processing partial speech input. Although large language models (LLMs) have shown strong capabilities in offline translation tasks, applying them to SimulST poses notable challenges. Existing LLM-based SimulST approaches either incur significant computational overhead due to repeated encoding of bidirectional speech encoder, or they depend on a fixed read/write policy, limiting the efficiency and performance. In this work, we introduce Efficient and Adaptive Simultaneous Speech Translation (EASiST) with fully unidirectional architecture, including both speech encoder and LLM. EASiST includes a multi-latency data curation strategy to generate semantically aligned SimulST training samples and redefines SimulST as an interleaved generation task with explicit read/write tokens. To facilitate adaptive inference, we incorporate a lightweight policy head that dynamically predicts read/write actions. Additionally, we employ a multi-stage training strategy to align speech-text modalities and optimize both translation and policy behavior. Experiments on both in-domain (MuST-C) and out-of-domain (Europarl-ST) En-De and En-Es datasets demonstrate that EASiST offers superior latency-quality trade-offs compared to several strong baselines.
Biao Fu, Donglei Yu, Minpeng Liao, Chengxi Li 0014, Xinjie Chen, Yidong Chen 0001, Kai Fan 0002, Xiaodong Shi
AAAI1
2026 From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization
abstract
Xinjie Chen, Minpeng Liao, Guoxin Chen, Chengxi Li, Biao Fu, Kai Fan, Xinggao Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinjie Chen, Minpeng Liao, Guoxin Chen, Chengxi Li 0014, Biao Fu, Kai Fan 0002, Xinggao Liu
ACL (1)5
2026 Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
abstract
Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators’ ability to adapt reasoning effort based on difficulty, we propose TwT (Translation with Thought), a resource-rational framework that learns to modulate inference between intuitive and deliberate reasoning. TwT is trained in two stages: (1) supervised fine-tuning on difficulty-aware long chain-of-though traces distilled from DeepSeek-R1 and rewritten by GPT-4o to reflect human-like reasoning economy, and (2) reinforcement learning with a hybrid reward to optimize translation quality and reasoning efficiency. Evaluated on 15 benchmarks spanning in-domain and out-of-domain settings, as well as 3 seen and 59 unseen languages, with ablations across three backbone models, TwT-7B and TwT-14B outperform much larger SOTA reasoning models in translation quality, while reducing token usage by 32–60%. These results confirm that aligning translation behavior with cognitive principles enables robust generalization, high translation quality, and efficient reasoning in MDMT.
Yongshi Ye, Biao Fu, Chongxuan Huang, Yidong Chen 0001, Xiaodong Shi
ACL (1)2
2025 From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment
abstract
Large language models (LLMs) have demonstrated remarkable multilingual capabilities, however, how to evaluate cross-lingual alignment remains underexplored.Existing alignment benchmarks primarily focus on sentence embeddings, but prior research has shown that neural models tend to induce a nonsmooth representation space, which impact of semantic alignment evaluation on low-resource languages.Inspired by neuroscientific findings that similar information activates overlapping neuronal regions, we propose a novel Neuron State-Based Cross-Lingual Alignment (NeuronXA) to assess the cross-lingual a lignment capabilities of LLMs, which offers a more semantically grounded approach to assess cross-lingual alignment.We evaluate NeuronXA on several prominent multilingual LLMs (LLaMA, Qwen, Mistral, GLM, and OLMo) across two transfer tasks and three multilingual benchmarks.The results demonstrate that with only 100 parallel sentence pairs, Neu-ronXA achieves a Pearson correlation of 0.9556 with downstream tasks performance and 0.8524 with transferability.These findings demonstrate NeuronXA's effectiveness in assessing both cross-lingual alignment and transferability, even with a small dataset.This highlights its potential to advance cross-lingual alignment research and to improve the semantic understanding of multilingual LLMs.
Chongxuan Huang, Yongshi Ye, Biao Fu, Qifeng Su, Xiaodong Shi
ACL (1)3
2025 Towards Simultaneous Sign Language Production: A Future-Context-Aware Approach
abstract
Sign Language Production (SLP) has achieved promising progress in offline settings, where full input text is available before generation. However, such methods are unsuitable for real-time applications requiring low latency. In this work, we introduce Simultaneous Sign Language Production (SimulSLP), a new task that generates sign pose sequences incrementally from streaming text input. We first formalize the SimulSLP task and adapt the Average Token Delay metric to quantify latency. Then, we benchmark this task using three strong baselines from offline SLP—an end-to-end system and two cascaded pipelines with neural and dictionary-based Gloss-to-Pose modules—under a wait-k policy. However, all baselines suffer from a mismatch between full-sequence training and partial-input inference. To mitigate this, we propose a Future-Context-Aware Inference (FCAI) strategy. FCAI enhances partial input representations by predicting a small number of future tokens using a large language model. Before decoding, speculative features from the predicted tokens are discarded to ensure alignment with the observed input. Experiments on PHOENIX2014T show that FCAI significantly improves the quality-latency trade-off, especially in low-latency settings, offering a promising step toward SimulSLP.
Biao Fu, Xiaodong Shi, Yidong Chen 0001
IEEE Signal Process. Lett.1
2024 Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal Alignment
abstract
Sign language translation (SLT) aims to convert continuous sign language videos into textual sentences. As a typical multi-modal task, there exists an inherent modality gap between sign language videos and spoken language text, which makes the cross-modal alignment between visual and textual modalities crucial. However, previous studies tend to rely on an intermediate sign gloss representation to help alleviate the cross-modal problem thereby neglecting the alignment across modalities that may lead to compromised results. To address this issue, we propose a novel framework based on Conditional Variational autoencoder for SLT (CV-SLT) that facilitates direct and sufficient cross-modal alignment between sign language videos and spoken language text. Specifically, our CV-SLT consists of two paths with two Kullback-Leibler (KL) divergences to regularize the outputs of the encoder and decoder, respectively. In the prior path, the model solely relies on visual information to predict the target text; whereas in the posterior path, it simultaneously encodes visual information and textual knowledge to reconstruct the target text. The first KL divergence optimizes the conditional variational autoencoder and regularizes the encoder outputs, while the second KL divergence performs a self-distillation from the posterior path to the prior path, ensuring the consistency of decoder outputs.We further enhance the integration of textual information to the posterior path by employing a shared Attention Residual Gaussian Distribution (ARGD), which considers the textual information in the posterior path as a residual component relative to the prior path. Extensive experiments conducted on public datasets demonstrate the effectiveness of our framework, achieving new state-of-the-art results while significantly alleviating the cross-modal representation discrepancy. The code and models are available at https://github.com/rzhao-zhsq/CV-SLT.
Biao Fu, Jinsong Su, Yidong Chen 0001
AAAI3
2024 Layer-Wise Representation Fusion for Compositional Generalization
abstract
Existing neural models are demonstrated to struggle with compositional generalization (CG), i.e., the ability to systematically generalize to unseen compositions of seen components. A key reason for failure on CG is that the syntactic and semantic representations of sequences in both the uppermost layer of the encoder and decoder are entangled. However, previous work concentrates on separating the learning of syntax and semantics instead of exploring the reasons behind the representation entanglement (RE) problem to solve it. We explain why it exists by analyzing the representation evolving mechanism from the bottom to the top of the Transformer layers. We find that the ``shallow'' residual connections within each layer fail to fuse previous layers' information effectively, leading to information forgetting between layers and further the RE problems. Inspired by this, we propose LRF, a novel Layer-wise Representation Fusion framework for CG, which learns to fuse previous layers' information back into the encoding and decoding process effectively through introducing a fuse-attention module at each encoder and decoder layer. LRF achieves promising results on two realistic benchmarks, empirically demonstrating the effectiveness of our proposal. Codes are available at https://github.com/thinkaboutzero/LRF.
Yafang Zheng, Shuangtao Li, Zhaohong Lai, Biao Fu, Yidong Chen 0001, Xiaodong Shi
AAAI7
2024 Adaptive Simultaneous Sign Language Translation with Confident Translation Length Estimation
abstract
Traditional non-simultaneous Sign Language Translation (SLT) methods, while effective for pre-recorded videos, face challenges in real-time scenarios due to inherent inference delays. The emerging field of simultaneous SLT aims to address this issue by progressively translating incrementally received sign video. However, the sole existing work in simultaneous SLT adopts a fixed gloss-based policy, which suffer from limitations in boundary prediction and contextual comprehension. In this paper, we delve deeper into this area and propose an adaptive policy for simultaneous SLT. Our approach introduces the concept of “confident translation length”, denoting maximum accurate translation achievable from current input. An estimator measures this length for streaming sign video, enabling the model to make informed decisions on whether to wait for more input or proceed with translation. To train the estimator, we construct a training data of confident translation length based on the longest common prefix between translations of partial and complete inputs. Furthermore, we incorporate adaptive training, utilizing pseudo prefix pairs, to refine the offline translation model for optimal performance in simultaneous scenarios. Experimental results on PHOENIX2014T and CSL-Daily demonstrate the superiority of our adaptive policy over existing methods, particularly excelling in situations requiring extremely low latency.
Biao Fu, Ruiquan Zhang, Xiaodong Shi, Jinsong Su, Yidong Chen 0001
LREC/COLING2
2024 Improving Non-Autoregressive Sign Language Translation with Random Ordering Progressive Prediction Pretraining
abstract
Recently, the Non-AutoRegressive (NAR) decoding mechanism, effectively reducing the inference latency of text generation, has been applied to Sign Language Translation (SLT). Typically, the current best NAR SLT model using a Curriculum-based Non-autoregressive Decoder (CND) outperforms AutoRegressive (AR) baselines in speed and performance. Although it has been proven that AutoRegressive Pre-trained Language Models (AR-PLMs) further boost the performance of AR SLT models, combining NAR Pretrained Language Models (NAR-PLMs) with NAR SLT model remains challenge due to (1) existing NAR-PLMs’ inability to model token dependencies between decoder layers, crucial for NAR SLT models using CND; (2) the modality gap between the decoder’s inputs of the NAR-PLMs and NAR SLT models. To address these, we propose a Random Ordering Progressive Prediction Pre-training task for NAR SLT models using CND, enabling the decoder to predict target sequences in diverse orderings and enhancing the modeling of target token dependencies between layers. Moreover, we propose a CTC-enhanced Soft Copy method to incorporate target-side information in the decoder’s inputs, alleviating the modality gap. Experimental results on PHOENIX-2014T and CSL-Daily demonstrate that our model consistently outperforms all strong baselines and achieves competitive performance with AR SLT models equipped with AR-PLMs.
Pei Yu, Changhao Lai, Biao Fu, Yidong Chen 0001
ECAI6
2024 Multi-Level Cross-Modal Alignment for Speech Relation Extraction
abstract
Liang Zhang, Zhen Yang, Biao Fu, Ziyao Lu, Liangying Shao, Shiyu Liu, Fandong Meng, Jie Zhou, Xiaoli Wang, Jinsong Su. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Biao Fu, Ziyao Lu, Liangying Shao, Fandong Meng, Jie Zhou 0016, Xiaoli Wang 0002, Jinsong Su
EMNLP3
2024 An Explicit Multi-Modal Fusion Method for Sign Language Translation
abstract
Sign Language Translation (SLT) aims to convert sign language videos into corresponding spoken text sequences. However, the inherent modality gap between sign language video and text hinders the development of SLT. Motivated by the linguistic consistency between gloss1and text, we propose EMF-SLT, an Explicit Multi-modal Fusion method for Sign Language Translation to mitigate the modality gap with the help of gloss. Specifically, EMF-SLT first leverages a vector quantizer and a fusion module to align and fuse sign language and gloss features, respectively, resulting in more informative multi-modal features for the decoder. Then, a multi-task mutual learning framework is introduced to regularize the output predictions from different modalities, which ensures the consistency of outputs across modalities and encourages different modalities to learn from each other. Experiments on two SLT benchmarks and further analyses show that our method achieves significant improvements over the baselines and effectively alleviates the modality gap.
Biao Fu, Pei Yu, Xiaodong Shi, Yidong Chen 0001
ICASSP2
2023 Exploring Self-Distillation Based Relational Reasoning Training for Document-Level Relation Extraction
abstract
Document-level relation extraction (RE) aims to extract relational triples from a document. One of its primary challenges is to predict implicit relations between entities, which are not explicitly expressed in the document but can usually be extracted through relational reasoning. Previous methods mainly implicitly model relational reasoning through the interaction among entities or entity pairs. However, they suffer from two deficiencies: 1) they often consider only one reasoning pattern, of which coverage on relational triples is limited; 2) they do not explicitly model the process of relational reasoning. In this paper, to deal with the first problem, we propose a document-level RE model with a reasoning module that contains a core unit, the reasoning multi-head self-attention unit. This unit is a variant of the conventional multi-head self-attention and utilizes four attention heads to model four common reasoning patterns, respectively, which can cover more relational triples than previous methods. Then, to address the second issue, we propose a self-distillation training framework, which contains two branches sharing parameters. In the first branch, we first randomly mask some entity pair feature vectors in the document, and then train our reasoning module to infer their relations by exploiting the feature information of other related entity pairs. By doing so, we can explicitly model the process of relational reasoning. However, because the additional masking operation is not used during testing, it causes an input gap between training and testing scenarios, which would hurt the model performance. To reduce this gap, we perform conventional supervised training without masking operation in the second branch and utilize Kullback-Leibler divergence loss to minimize the difference between the predictions of the two branches. Finally, we conduct comprehensive experiments on three benchmark datasets, of which experimental results demonstrate that our model consistently outperforms all competitive baselines. Our source code is available at https://github.com/DeepLearnXMU/DocRE-SD
Jinsong Su, Zijun Min, Zhongjian Miao, Qingguo Hu, Biao Fu, Xiaodong Shi, Yidong Chen 0001
AAAI6
2023 Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and Inference
abstract
A popular approach to streaming speech translation is to employ a single offline model with a wait-k policy to support different latency requirements, which is simpler than training multiple online models with different latency constraints.However, there is a mismatch problem in using a model trained with complete utterances for streaming inference with partial input.We demonstrate that speech representations extracted at the end of a streaming input are significantly different from those extracted from a complete utterance.To address this issue, we propose a new approach called Future-Aware Streaming Translation (FAST) that adapts an offline ST model for streaming input.FAST includes a Future-Aware Inference (FAI) strategy that incorporates future context through a trainable masked embedding, and a Future-Aware Distillation (FAD) framework that transfers future context from an approximation of full speech to streaming input.Our experiments on the MuST-C EnDe, EnEs, and EnFr benchmarks show that FAST achieves better trade-offs between translation quality and latency than strong baselines.Extensive analyses suggest that our methods effectively alleviate the aforementioned mismatch problem between offline training and online inference.1
Biao Fu, Minpeng Liao, Kai Fan 0002, Zhongqiang Huang, Boxing Chen, Yidong Chen 0001, Xiaodong Shi
EMNLP1
2023 A Token-Level Contrastive Framework for Sign Language Translation
abstract
Sign Language Translation (SLT) is a promising technology to bridge the communication gap between the deaf and the hearing people. Recently, researchers have adopted Neural Machine Translation (NMT) methods, which usually require large-scale corpus for training, to achieve SLT. However, the publicly available SLT corpus is very limited, which causes the collapse of the token representations and the inaccuracy of the generated tokens. To alleviate this issue, we propose Con-SLT, a novel token-level Contrastive learning framework for Sign Language Translation , which learns effective token representations by incorporating token-level contrastive learning into the SLT decoding process. Concretely, ConSLT treats each token and its counterpart generated by different dropout masks as positive pairs during decoding, and then randomly samples K tokens in the vocabulary that are not in the current sentence to construct negative examples. We conduct comprehensive experiments on two benchmarks (PHOENIX14T and CSL-Daily) for both end-to-end and cascaded settings. The experimental results demonstrate that ConSLT can achieve better translation quality than the strong baselines1.
Biao Fu, Peigen Ye, Pei Yu, Xiaodong Shi, Yidong Chen 0001
ICASSP1
2023 Efficient Sign Language Translation with a Curriculum-based Non-autoregressive Decoder
abstract
Most existing studies on Sign Language Translation (SLT) employ AutoRegressive Decoding Mechanism (AR-DM) to generate target sentences. However, the main disadvantage of the AR-DM is high inference latency. To address this problem, we introduce Non-AutoRegressive Decoding Mechanism (NAR-DM) into SLT, which generates the whole sentence at once. Meanwhile, to improve its decoding ability, we integrate the advantages of curriculum learning and NAR-DM and propose a Curriculum-based NAR Decoder (CND). Specifically, the lower layers of the CND are expected to predict simple tokens that could be predicted correctly using source-side information solely. Meanwhile, the upper layers could predict complex tokens based on the lower layers' predictions. Therefore, our CND significantly reduces the model's inference latency while maintaining its competitive performance. Moreover, to further boost the performance of our CND, we propose a mutual learning framework, containing two decoders, i.e., an AR decoder and our CND. We jointly train the two decoders and minimize the KL divergence between their outputs, which enables our CND to learn the forward sequential knowledge from the strengthened AR decoder. Experimental results on PHOENIX2014T and CSL-Daily demonstrate that our model consistently outperforms all competitive baselines and achieves 7.92/8.02× speed-up compared to the AR SLT model respectively. Our source code is available at https://github.com/yp20000921/CND.
Pei Yu, Biao Fu, Yidong Chen 0001
IJCAI3
2022 Continuous Prompt Enhanced Biomedical Entity Normalization
Zhaohong Lai, Biao Fu, Shangfei Wei, Xiaodong Shi
NLPCC (2)2
2021 CTRD: A Chinese Theme-Rheme Discourse Dataset
Biao Fu, Yiqi Tong, Dawei Tian, Yidong Chen 0001, Xiaodong Shi
NLPCC (1)1