EDBT 2026 Demo / reviewers in the wild / expert
Yifan Song 0002
dblp:66/7929-2
· DBLP profile ↗
17ranked-venue papers
2as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | P-Aligner: Pre-Aligning LLMs via Principled Instruction Synthesis
Feifan Song 0001, Bofei Gao, Yifan Song 0002, Weimin Xiong, Yuyang Song, Tianyu Liu 0001, Houfeng Wang |
WWW | 3 |
| 2025 | ISR: Self-Refining Referring Expressions for Entity GroundingabstractEntity grounding, a crucial task in constructing multimodal knowledge graphs, aims to align entities from knowledge graphs with their corresponding images.Unlike conventional visual grounding tasks that use referring expressions (REs) as inputs, entity grounding relies solely on entity names and types, presenting a significant challenge.To address this, we introduce a novel Iterative Self-Refinement (ISR) scheme to enhance the multimodal large language model's capability to generate high quality REs for the given entities as explicit contextual clues.This training scheme, inspired by human learning dynamics and human annotation processes, enables the MLLM to iteratively generate and refine REs by learning from successes and failures, guided by outcome rewards from a visual grounding model.This iterative cycle of self-refinement avoids overfitting to fixed annotations and fosters continued improvement in referring expression generation.Extensive experiments demonstrate that our methods surpasses other methods in entity grounding, highlighting its effectiveness, robustness and potential for broader applications 1 . Zhuocheng Yu, Bingchan Zhao, Yifan Song 0002, Sujian Li, Zhonghui He |
ACL (1) | 3 |
| 2025 | Hierarchical Memory Organization for Wikipedia GenerationabstractEugene J. Yu, Dawei Zhu, Yifan Song, Xiangyu Wong, Jiebin Zhang, Wenxuan Shi, Xiaoguang Li, Qun Liu, Sujian Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Eugene J. Yu, Yifan Song 0002, Xiangyu Wong, Jiebin Zhang, Qun Liu 0001, Sujian Li |
ACL (1) | 3 |
| 2025 | Exploring Fine-Grained Human Motion Video CaptioningabstractDetailed descriptions of human motion are crucial for effective fitness training, which highlights the importance of research in fine-grained human motion video captioning. Existing video captioning models often fail to capture the nuanced semantics of videos, resulting in the generated descriptions that are coarse and lack details, especially when depicting human motions. To benchmark the Body Fitness Training scenario, in this paper, we construct a fine-grained human motion video captioning dataset named BoFiT and design a state-of-the-art baseline model named BoFiT-Gen (Body Fitness Training Text Generation). BoFiT-Gen makes use of computer vision techniques to extract angular representations of human motions from videos and LLMs to generate fine-grained descriptions of human motions via prompting. Results show that BoFiT-Gen outperforms previous methods on comprehensive metrics. We aim for this dataset to serve as a useful evaluation set for visio-linguistic models and drive further progress in this field. Our dataset is released at https://github.com/colmon46/bofit. Bingchan Zhao, Zhuocheng Yu, Tongchen Yang, Yifan Song 0002, Mingyu Jin, Sujian Li |
COLING | 5 |
| 2025 | VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward ModelsabstractVision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench’s effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson’s r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io. Lei Li 0039, Yuancheng Wei, Zhihui Xie 0002, Xuqing Yang, Yifan Song 0002, Peiyi Wang, Chenxin An, Tianyu Liu 0001, Sujian Li, Bill Y. Lin, Lingpeng Kong, Qi Liu 0049 |
CVPR | 5 |
| 2025 | Harnessing Webpage UIs for Text-Rich Visual UnderstandingabstractText-rich visual understanding—the ability to interpret both textual content and visual elements within a scene—is crucial for multimodal large language models (MLLMs) to effectively interact with structured environments. We propose leveraging webpage UIs as a naturally structured and diverse data source to enhance MLLMs’ capabilities in this area. Existing approaches, such as rule-based extraction, multimodal model captioning, and rigid HTML parsing, are hindered by issues like noise, hallucinations, and limited generalization. To overcome these challenges, we introduce MultiUI, a dataset of 7.3 million samples spanning various UI types and tasks, structured using enhanced accessibility trees and task taxonomies. By scaling multimodal instructions from web UIs through LLMs, our dataset enhances generalization beyond web domains, significantly improving performance in document understanding, GUI comprehension, grounding, and advanced agent tasks. This demonstrates the potential of structured web data to elevate MLLMs’ proficiency in processing text-rich visual environments and generalizing across domains. Junpeng Liu 0001, Tianyue Ou, Yifan Song 0002, Yuxiao Qu, Wai Lam, Chenyan Xiong, Wenhu Chen, Graham Neubig, Xiang Yue |
ICLR | 3 |
| 2025 | The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-DeterminismabstractYifan Song, Guoyin Wang, Sujian Li, Bill Yuchen Lin. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yifan Song 0002, Sujian Li, Bill Y. Lin |
NAACL (Long Papers) | 1 |
| 2024 | Trial and Error: Exploration-Based Trajectory Optimization of LLM AgentsabstractLarge Language Models (LLMs) have become integral components in various autonomous agent systems.In this study, we present an exploration-based trajectory optimization approach, referred to as ETO.This learning method is designed to enhance the performance of open LLM agents.Contrary to previous studies that exclusively train on successful expert trajectories, our method allows agents to learn from their exploration failures.This leads to improved performance through an iterative optimization framework.During the exploration phase, the agent interacts with the environment while completing given tasks, gathering failure trajectories to create contrastive trajectory pairs.In the subsequent training phase, the agent utilizes these trajectory preference pairs to update its policy using contrastive learning methods like DPO (Rafailov et al., 2023).This iterative cycle of exploration and training fosters continued improvement in the agents.Our experiments on three complex tasks demonstrate that ETO consistently surpasses baseline performance by a large margin.Furthermore, an examination of task-solving efficiency and potential in scenarios lacking expert trajectory underscores the effectiveness of our approach.1 Yifan Song 0002, Da Yin, Xiang Yue, Sujian Li, Bill Y. Lin |
ACL (1) | 1 |
| 2024 | Watch Every Step! LLM Agent Learning via Iterative Step-level Process RefinementabstractLarge language model agents have exhibited exceptional performance across a range of complex interactive tasks.Recent approaches have utilized tuning with expert trajectories to enhance agent performance, yet they primarily concentrate on outcome rewards, which may lead to errors or suboptimal actions due to the absence of process supervision signals.In this paper, we introduce the Iterative step-level Process Refinement (IPR) framework, which provides detailed step-by-step guidance to enhance agent training.Specifically, we adopt the Monte Carlo method to estimate step-level rewards.During each iteration, the agent explores along the expert trajectory and generates new actions.These actions are then evaluated against the corresponding step of expert trajectory using step-level rewards.Such comparison helps identify discrepancies, yielding contrastive action pairs that serve as training data for the agent.Our experiments on three complex agent tasks demonstrate that our framework outperforms a variety of strong baselines.Moreover, our analytical findings highlight the effectiveness of IPR in augmenting action efficiency and its applicability to diverse models † . Weimin Xiong, Yifan Song 0002, Xiutian Zhao, Cheng Li 0040, Wei Peng 0011, Sujian Li |
EMNLP | 2 |
| 2024 | LongEmbed: Extending Embedding Models for Long Context RetrievalabstractEmbedding models play a pivotal role in modern NLP applications such as document retrieval.However, existing embedding models are limited to encoding short documents of typically 512 tokens, restrained from application scenarios requiring long inputs.This paper explores context window extension of existing embedding models, pushing their input length to a maximum of 32,768.We begin by evaluating the performance of existing embedding models using our newly constructed LONGEM-BED benchmark, which includes two synthetic and four real-world tasks, featuring documents of varying lengths and dispersed target information.The benchmarking results highlight huge opportunities for enhancement in current models.Via comprehensive experiments, we demonstrate that training-free context window extension strategies can effectively increase the input length of these models by several folds.Moreover, comparison of models using Absolute Position Encoding (APE) and Rotary Position Encoding (RoPE) reveals the superiority of RoPE-based embedding models in context window extension, offering empirical guidance for future models.Our benchmark, code and trained models will be released to advance the research in long context embedding models. Liang Wang 0046, Nan Yang 0002, Yifan Song 0002, Furu Wei, Sujian Li |
EMNLP | 4 |
| 2024 | PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise TrainingabstractLarge Language Models (LLMs) are trained with a pre-defined context length, restricting their use in scenarios requiring long inputs. Previous efforts for adapting LLMs to a longer length usually requires fine-tuning with this target length (Full-length fine-tuning), suffering intensive training cost. To decouple train length from target length for efficient context window extension, we propose Positional Skip-wisE (PoSE) training that smartly simulates long inputs using a fixed context window. This is achieved by first dividing the original context window into several chunks, then designing distinct skipping bias terms to manipulate the position indices of each chunk. These bias terms and the lengths of each chunk are altered for every training example, allowing the model to adapt to all positions within target length. Experimental results show that PoSE greatly reduces memory and time overhead compared with Full-length fine-tuning, with minimal impact on performance. Leveraging this advantage, we have successfully extended the LLaMA model to 128k tokens using a 2k training context window. Furthermore, we empirically confirm that PoSE is compatible with all RoPE-based LLMs and position interpolation strategies. Notably, our method can potentially support infinite length, limited only by memory usage in inference. With ongoing progress for efficient inference, we believe PoSE can further scale the context window beyond 128k. Nan Yang 0002, Liang Wang 0046, Yifan Song 0002, Furu Wei, Sujian Li |
ICLR | 4 |
| 2024 | CoUDA: Coherence Evaluation via Unified Data AugmentationabstractDawei Zhu, Wenhao Wu, Yifan Song, Fangwei Zhu, Ziqiang Cao, Sujian Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yifan Song 0002, Fangwei Zhu, Ziqiang Cao, Sujian Li |
NAACL-HLT | 3 |
| 2023 | Rationale-Enhanced Language Models are Better Continual Relation LearnersabstractContinual relation extraction (CRE) aims to solve the problem of catastrophic forgetting when learning a sequence of newly emerging relations.Recent CRE studies have found that catastrophic forgetting arises from the model's lack of robustness against future analogous relations.To address the issue, we introduce rationale, i.e., the explanations of relation classification results generated by large language models (LLM), into CRE task.Specifically, we design the multi-task rationale tuning strategy to help the model learn current relations robustly.We also conduct contrastive rationale replay to further distinguish analogous relations.Experimental results on two standard benchmarks demonstrate that our method outperforms the state-of-the-art CRE models.Our code is available at https://github.com/WeiminXiong/ RationaleCL Weimin Xiong, Yifan Song 0002, Peiyi Wang, Sujian Li |
EMNLP | 2 |
| 2023 | DocRED-FE: A Document-Level Fine-Grained Entity and Relation Extraction DatasetabstractJoint entity and relation extraction (JERE) is one of the most important tasks in information extraction. However, most existing works focus on sentence-level coarse-grained JERE, which have limitations in real-world scenarios. In this paper, we construct a large-scale document-level fine-grained JERE dataset DocRED-FE, which improves DocRED with Fine-Grained Entity Type. Specifically, we redesign a hierarchical entity type schema including 11 coarse-grained types and 119 fine-grained types, and then re-annotate DocRED manually according to this schema. Through comprehensive experiments we find that: (1) DocRED-FE is challenging to existing JERE models; (2) Our fine-grained entity types promote relation classification. We make DocRED-FE with instruction and the code for our baselines publicly available at https://github.com/PKU-TANGENT/DOCRED-FE. Weimin Xiong, Yifan Song 0002, Sujian Li |
ICASSP | 3 |
| 2022 | ConFiguRe: Exploring Discourse-level Chinese Figures of SpeechabstractFigures of speech, such as metaphor and irony, are ubiquitous in literature works and colloquial conversations. This poses great challenge for natural language understanding since figures of speech usually deviate from their ostensible meanings to express deeper semantic implications. Previous research lays emphasis on the literary aspect of figures and seldom provide a comprehensive exploration from a view of computational linguistics. In this paper, we first propose the concept of figurative unit, which is the carrier of a figure. Then we select 12 types of figures commonly used in Chinese, and build a Chinese corpus for Contextualized Figure Recognition (ConFiguRe). Different from previous token-level or sentence-level counterparts, ConFiguRe aims at extracting a figurative unit from discourse-level context, and classifying the figurative unit into the right figure type. On ConFiguRe, three tasks, i.e., figure extraction, figure type classification and figure recognition, are designed and the state-of-the-art techniques are utilized to implement the benchmarks. We conduct thorough experiments and show that all three tasks are challenging for existing models, thus requiring further research. Our dataset and code are publicly available at https://github.com/pku-tangent/ConFiguRe. Qiusi Zhan, Zhejian Zhou, Yifan Song 0002, Jiebin Zhang, Sujian Li |
COLING | 4 |
| 2022 | Learning Robust Representations for Continual Relation Extraction via Adversarial Class AugmentationabstractContinual relation extraction (CRE) aims to continually learn new relations from a classincremental data stream.CRE model usually suffers from catastrophic forgetting problem, i.e., the performance of old relations seriously degrades when the model learns new relations.Most previous work attributes catastrophic forgetting to the corruption of the learned representations as new relations come, with an implicit assumption that the CRE models have adequately learned the old relations.In this paper, through empirical studies we argue that this assumption may not hold, and an important reason for catastrophic forgetting is that the learned representations do not have good robustness against the appearance of analogous relations in the subsequent learning process.To address this issue, we encourage the model to learn more precise and robust representations through a simple yet effective adversarial class augmentation mechanism (ACA), which is easy to implement and model-agnostic.Experimental results show that ACA can consistently improve the performance of state-of-theart CRE models on two popular benchmarks. Peiyi Wang, Yifan Song 0002, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Sujian Li, Zhifang Sui |
EMNLP | 2 |
| 2022 | Robust Fine-tuning via Perturbation and Interpolation from In-batch InstancesabstractFine-tuning pretrained language models (PLMs) on downstream tasks has become common practice in natural language processing. However, most of the PLMs are vulnerable, e.g., they are brittle under adversarial attacks or imbalanced data, which hinders the application of the PLMs on some downstream tasks, especially in safe-critical scenarios. In this paper, we propose a simple yet effective fine-tuning method called Match-Tuning to force the PLMs to be more robust. For each instance in a batch, we involve other instances in the same batch to interact with it. To be specific, regarding the instances with other labels as a perturbation, Match-Tuning makes the model more robust to noise at the beginning of training. While nearing the end, Match-Tuning focuses more on performing an interpolation among the instances with the same label for better generalization. Extensive experiments on various tasks in GLUE benchmark show that Match-Tuning consistently outperforms the vanilla fine-tuning by 1.64 scores. Moreover, Match-Tuning exhibits remarkable robustness to adversarial attacks and data imbalance. Shoujie Tong, Qingxiu Dong, Damai Dai, Yifan Song 0002, Tianyu Liu 0001, Baobao Chang, Zhifang Sui |
IJCAI | 4 |