Yuhang Guo 0001

dblp:74/10083-1 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
14since 2021 · last 2026
0009-0005-6343-285XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
abstract
Through reinforcement learning (RL) with outcome correctness rewards, large reasoning models (LRMs) with scaled inference computation have demonstrated substantial success on complex reasoning tasks. However, the one-sided reward, focused solely on final correctness, limits its ability to provide detailed supervision over internal reasoning process. This deficiency leads to suboptimal internal reasoning quality, manifesting as issues like over-thinking, under-thinking, redundant-thinking, and disordered-thinking. Inspired by the recent progress in LRM self-rewarding, we introduce self-rewriting framework, where a model rewrites its own reasoning texts, and subsequently learns from the rewritten reasoning to improve the internal thought process quality. For algorithm design, we propose a selective rewriting approach wherein only "simple" samples, defined by the model's consistent correctness, are rewritten, thereby preserving all original reward signals of GRPO. For practical implementation, we compile rewriting and vanilla generation within one single batch, maintaining the scalability of the RL algorithm and introducing only 10% overhead. Extensive experiments on diverse tasks with different model sizes validate the effectiveness of self-rewriting. In terms of the accuracy-length tradeoff, the self-rewriting approach achieves improved accuracy (+0.6) with substantially shorter reasoning (-46%) even without explicit instructions in rewriting prompts to reduce reasoning length, outperforming existing strong baselines. In terms of internal reasoning quality, self-rewriting achieves significantly higher scores (+7.2) under the LLM-as-a-judge metric, successfully mitigating internal reasoning flaws.
Jiashu Yao, Heyan Huang, Shuang Zeng, Chuwei Luo, WangJie You, Jie Tang 0001, Yuhang Guo 0001, Yangyang Kang
AAAI8
2026 Mem²Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
abstract
Zihao Cheng, Zeming Liu, Yingyu Shan, Xinyi Wang, Xiangrong Zhu, Yunpu Ma, Hongru Wang, Yuhang Guo, Wei Lin, Yunhong Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zeming Liu, Yingyu Shan, Xiangrong Zhu 0002, Yunpu Ma, Hongru Wang 0003, Yuhang Guo 0001, Yunhong Wang 0001
ACL (1)8
2026 PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception
abstract
Embodied Action Sequence Planning focuses on the capability of embodied agents to implement action planning via environmental perception.This technology enables diverse intelligent assistance for real-world scenarios such as home and office environments.To address the limitations of existing embodied agents in meeting the requirement for proactivity and achieving joint understanding of visual and audio information, this study investigates the ability of embodied agents to proactively provide assistance through action sequence planning based on joint understanding of vision and audio perception without explicit human instructions.Correspondingly, we propose PEAP, the first multimodal proactive embodied action sequence planning dataset.We evaluate the performance of multiple Large Language Models on the PEAP dataset.The results demonstrate that these models still exhibit significant deficiencies on this task particularly lacking accurate environmental perception capabilities.Furthermore, ablation experiment and replacement experiment further corroborate that the joint understanding of multimodal information can significantly improve the models' performance on proactive embodied action sequence planning task.Our dataset and code are publicly available 1 .
Tianwei Lan, Zeming Liu, Zhaoxin Fan, Haifeng Wang 0001, Yuhang Guo 0001
ACL (1)6
2026 Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
abstract
Yanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang, Wenbo Yu, Dawei Song, Jie Tang, Yuhang Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang, Jie Tang 0001, Yuhang Guo 0001
ACL (1)8
2026 Controllable timbre cloning and style replication with reference speech examples for multimodal human-computer interaction
Tianwei Lan, Yuhang Guo 0001, Mengyuan Deng, Jing Wang 0037, Wenwu Wang 0001, Chong Feng 0001
Neurocomputing2
2026 Towards multi-language repository-level code generation: From-scratch to guided tasks
Silin Li, Zeming Liu, Yuhang Guo 0001, Yuanfang Guo, Yunhong Wang 0001, Haifeng Wang 0001
Neurocomputing5
2025 ReFF: Reinforcing Format Faithfulness in Language Models Across Varied Tasks
abstract
Following formatting instructions to generate well-structured content is a fundamental yet often unmet capability for large language models (LLMs). To study this capability, which we refer to as format faithfulness, we present FormatBench, a comprehensive format-related benchmark. Compared to previous format-related benchmarks, FormatBench involves a greater variety of tasks in terms of application scenes (traditional NLP tasks, creative works, autonomous agency tasks), human-LLM interaction styles (single-turn instruction, multi-turn chat), and format types (inclusion, wrapping, length, coding). Moreover, each task in FormatBench is attached with a format checker program. Extensive experiments on the benchmark reveal that state-of-the-art open- and closed-source LLMs still suffer from severe deficiency in format faithfulness. By virtue of the decidable nature of formats, we propose to Reinforce Format Faithfulness (ReFF) to help LLMs generate formatted output as instructed without compromising general quality. Without any annotated data, ReFF can substantially improve the format faithfulness rate (e.g., from 21.6% in original LLaMA3 to 95.0% on caption segmentation task), while keep the general quality comparable (e.g., from 47.3 to 46.4 in F1 scores). Combined with labeled training data, ReFF can simultaneously improve both format faithfulness (e.g., from 21.6% in original LLaMA3 to 75.5%) and general quality (e.g., from 47.3 to 61.6 in F1 scores). We further offer an interpretability analysis to explain how ReFF improves both format faithfulness and general quality.
Jiashu Yao, Heyan Huang, Zeming Liu, Haoyu Wen, Boao Qian, Yuhang Guo 0001
AAAI7
2025 HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
abstract
Large language models (LLMs) have the potential to revolutionize smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately, which is extremely beneficial for building a smarter home environment.While recent studies have explored integrating LLMs into smart home systems, they primarily focus on handling straightforward, valid single-device operation instructions.However, real-world scenarios are far more complex and often involve users issuing invalid instructions or controlling multiple devices simultaneously.These have two main challenges: LLMs must accurately identify and rectify errors in user instructions and execute multiple user instructions perfectly.To address these challenges and advance the development of LLM-based smart home assistants, we introduce HomeBench, the first smart home dataset with valid and invalid instructions across single and multiple devices in this paper.We have experimental results on 13 distinct LLMs; e.g., GPT-4o achieves only a 0.0% success rate in the scenario of invalid multi-device instructions, revealing that the existing stateof-the-art LLMs still cannot perform well in this situation even with the help of in-context learning, retrieval-augmented generation, and fine-tuning.Our code and dataset are publicly available at the link 1 .
Silin Li, Yuhang Guo 0001, Jiashu Yao, Zeming Liu, Haifeng Wang 0001
ACL (1)2
2025 PRIM: Towards Practical In-Image Multilingual Machine Translation
abstract
In-Image Machine Translation (IIMT) aims to translate images containing texts from one language to another. Current research of end-to-end IIMT mainly conducts on synthetic data, with simple background, single font, fixed text position, and bilingual translation, which can not fully reflect real world, causing a significant gap between the research and practical conditions. To facilitate research of IIMT in real-world scenarios, we explore Practical In-Image Multilingual Machine Translation (IIMMT). In order to convince the lack of publicly available data, we annotate the PRIM dataset, which contains real-world captured one-line text images with complex background, various fonts, diverse text positions, and supports multilingual translation directions. We propose an end-to-end model VisTrans to handle the challenge of practical conditions in PRIM, which processes visual text and background information in the image separately, ensuring the capability of multilingual translation while improving the visual quality. Experimental results indicate the VisTrans achieves a better translation quality and visual effect compared to other models. The code and dataset are available at: https://github.com/BITHLP/PRIM.
Yanzhi Tian, Zeming Liu, Chong Feng 0001, Heyan Huang, Yuhang Guo 0001
EMNLP7
2024 TED-EL: A Corpus for Speech Entity Linking
abstract
Speech entity linking amis to recognize mentions from speech and link them to entities in knowledge bases. Previous work on entity linking mainly focuses on visual context and text context. In contrast, speech entity linking focuses on audio context. In this paper, we first propose the speech entity linking task. To facilitate the study of this task, we propose the first speech entity linking dataset, TED-EL. Our corpus is a high-quality, human-annotated, audio, text, and mention-entity pair parallel dataset derived from Technology, Entertainment, Design (TED) talks and includes a wide range of entity types (24 types). Based on TED-EL, we designed two types of models: ranking-based and generative speech entity linking models. We conducted experiments on the TED-EL dataset for both types of models. The results show that the ranking-based models outperform the generative models, achieving an F1 score of 60.68%.
Silin Li, Ruoyu Song 0002, Tianwei Lan, Zeming Liu, Yuhang Guo 0001
LREC/COLING5
2024 FAME: Towards Factual Multi-Task Model Editing
abstract
Large language models (LLMs) embed extensive knowledge and utilize it to perform exceptionally well across various tasks.Nevertheless, outdated knowledge or factual errors within LLMs can lead to misleading or incorrect responses, causing significant issues in practical applications.To rectify the fatal flaw without the necessity for costly model retraining, various model editing approaches have been proposed to correct inaccurate knowledge within LLMs in a cost-efficient way.To evaluate these model editing methods, previous work introduced a series of datasets.However, most of the previous datasets only contain fabricated data in a single format, which diverges from realworld model editing scenarios, raising doubts about their usability in practice.To facilitate the application of model editing in real-world scenarios, we propose the challenge of practicality.To resolve such challenges and effectively enhance the capabilities of LLMs, we present FAME, an factual, comprehensive, and multi-task dataset, which is designed to enhance the practicality of model editing.We then propose SKEME, a model editing method that uses a novel caching mechanism to ensure synchronization with the real world.The experiments demonstrate that SKEME performs excellently across various tasks and scenarios, confirming its practicality.
Zeng Li 0002, Yingyu Shan, Zeming Liu, Jiashu Yao, Yuhang Guo 0001
EMNLP5
2023 Rethinking the Reasonability of the Test Set for Simultaneous Machine Translation
abstract
Simultaneous machine translation (SimulMT) models start translation before the end of the source sentence, making the translation monotonically aligned with the source sentence. However, the general full-sentence translation test set is acquired by offline translation of the entire source sentence, which is not designed for SimulMT evaluation, making us rethink whether this will underestimate the performance of SimulMT models. In this paper, we manually annotate a monotonic test set based on the MuST-C English-Chinese test set, denoted as SiMuST-C. Our human evaluation confirms the acceptability of our annotated test set. Evaluations on three different SimulMT models verify that the underestimation problem can be alleviated on our test set. Further experiments show that finetuning on an automatically extracted monotonic training set improves SimulMT models by up to 3 BLEU points.
Mengge Liu, Wen Zhang 0015, Xiang Li 0104, Jian Luan 0001, Bin Wang 0004, Yuhang Guo 0001, Shuoying Chen
ICASSP6
2022 Overview of the NLPCC2022 Shared Task on Speech Entity Linking
Ruoyu Song 0002, Xiaoyu Tian, Yuhang Guo 0001
NLPCC (2)4
2021 Multi-relational EHR representation learning with infusing information of Diagnosis and Medication
abstract
Medical concept embedding which aims at learning interpretable low-dimensional representations of medical codes has become one of the key technologies to enable the machine (deep) learning models to imitate the doctor’s cognitive reasoning process in a variety of clinical tasks. Most existing works focus on leveraging the medical ontology to get the representations but remains ineffective in dealing with 1) the inconsistency between the knowledge of the medical ontology and the observations in health records, and 2) the deficiency of discovering the relations among multi-types of medical concepts. To address these challenges, this paper proposes MrER(Multi-relational EHR representation learning method). It’s a heterogeneous graph convolutional network with a self-adaptive adjacency matrix, to infer the multi-relations among different types of medical concepts and align them in the same subspace for the complex knowledge inference. Moreover, an temporal convolutional network is introduced to capture the dependency patterns in the sequence of medical records. The entire framework is trained in an end-to-end fashion. The experimental results show that MrER achieves competitive performance advantages in sequential diagnosis prediction task in comparison with state-of-the-art methods and the learned embeddings have good interpretability regarding the relationship between medical codes.
Yuhang Guo 0001, Hao Wu 0066, Jingxiu Li, Xin Li 0033
COMPSAC2
2017 A Parallel Recurrent Neural Network for Language Modeling with POS Tags
Chao Su 0002, Heyan Huang, Shumin Shi, Yuhang Guo 0001, Hao Wu 0066
PACLIC4
2017 Incorporating target language semantic roles into a string-to-tree translation model
abstract
The string-to-tree model is one of the most successful syntax-based statistical machine translation (SMT) models. It models the grammaticality of the output via target-side syntax. However, it does not use any semantic information and tends to produce translations containing semantic role confusions and error chunk sequences. In this paper, we propose two methods to use semantic roles to improve the performance of the string-to-tree translation model: (1) adding role labels in the syntax tree; (2) constructing a semantic role tree, and then incorporating the syntax information into it. We then perform string-to-tree machine translation using the newly generated trees. Our methods enable the system to train and choose better translation rules using semantic information. Our experiments showed significant improvements over the state-of-the-art string-to-tree translation system on both spoken and news corpora, and the two proposed methods surpass the phrase-based system on large-scale training data.
Chao Su 0002, Yuhang Guo 0001, Heyan Huang, Shumin Shi, Chong Feng 0001
Frontiers Inf. Technol. Electron. Eng.2
2013 Microblog Entity Linking by Leveraging Extra Posts
abstract
Linking name mentions in microblog posts to a knowledge base, namely microblog entity linking, is useful for text mining tasks on microblog.Entity linking in long text has been well studied in previous works.However few work has focused on short text such as microblog post.Microblog posts are short and noisy.Previous method can extract few features from the post context.In this paper we propose to use extra posts for the microblog entity linking task.Experimental results show that our proposed method significantly improves the linking accuracy over traditional methods by 8.3% and 7.5% respectively.
Yuhang Guo 0001, Bing Qin 0001, Ting Liu 0001, Sheng Li 0003
EMNLP1
2013 Improving Candidate Generation for Entity Linking
Yuhang Guo 0001, Bing Qin 0001, Yuqin Li, Ting Liu 0001, Sheng Li 0003
NLDB1
2011 A Graph-based Method for Entity Linking
Yuhang Guo 0001, Wanxiang Che, Ting Liu 0001, Sheng Li 0003
IJCNLP1