VLDB 2026 Research / reviewers in the wild / expert
Chenhan Yuan
dblp:239/5838
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0001-9667-0460ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ELAINE-medLLM: Lightweight English Japanese Chinese Trilingual Large Language Model for Bio-medical DomainabstractWe propose ELAINE (EngLish-jApanese-chINesE)-medLLM, a trilingual (English, Japanese, Chinese) large language model adapted for the bio-medical domain based on Llama-3-8B. The training dataset was carefully curated in terms of volume and diversity to adapt to the biomedical domain and endow trilingual capability while preserving the knowledge and abilities of the base model. The training follows 2-stage paths: continued pre-training and supervised fine-tuning (SFT). Our results demonstrate that ELAINE-medLLM exhibits superior trilingual capabilities compared to existing bilingual or multilingual medical LLMs without severely sacrificing the base model’s capability. Ken Yano, Zheheng Luo, Jimin Huang, Qianqian Xie, Masaki Asada, Chenhan Yuan, Kailai Yang, Makoto Miwa, Sophia Ananiadou, Jun'ichi Tsujii |
COLING | 6 |
| 2025 | VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark ModelsabstractMing Cheng, Jiaying Gong, Chenhan Yuan, William A Ingram, Edward Fox, Hoda Eldardiry. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jiaying Gong, Chenhan Yuan, William A. Ingram, Edward A. Fox, Hoda Eldardiry |
NAACL (Long Papers) | 3 |
| 2025 | CAST: Corpus-Aware Self-similarity Enhanced Topic modellingabstractYanan Ma, Chenghao Xiao, Chenhan Yuan, Sabine N Van Der Veer, Lamiece Hassan, Chenghua Lin, Goran Nenadic. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chenghao Xiao, Chenhan Yuan, Sabine van der Veer, Lamiece Hassan, Goran Nenadic |
NAACL (Long Papers) | 3 |
| 2025 | CARE: Decoding-Time Safety Alignment via Rollback and Introspection InterventionabstractAs large language models (LLMs) are increasingly deployed in real-world applications, ensuring the safety of their outputs during decoding has become a critical challenge. However, existing decoding-time interventions, such as Contrastive Decoding, often force a severe trade-off between safety and response quality. In this work, we propose **CARE**, a novel framework for decoding-time safety alignment that integrates three key components: (1) a guard model for real-time safety monitoring, enabling detection of potentially unsafe content; (2) a rollback mechanism with a token buffer to correct unsafe outputs efficiently at an earlier stage without disrupting the user experience; and (3) a novel introspection-based intervention strategy, where the model generates self-reflective critiques of its previous outputs and incorporates these reflections into the context to guide subsequent decoding steps. The framework achieves a superior safety-quality trade-off by using its guard model for precise interventions, its rollback mechanism for timely corrections, and our novel introspection method for effective self-correction. Experimental results demonstrate that our framework achieves a superior balance of safety, quality, and efficiency, attaining a **low harmful response rate** and **minimal disruption to the user experience** while **maintaining high response quality**. Xiaomeng Hu, Fei Huang 0002, Chenhan Yuan, Junyang Lin, Tsung-Yi Ho |
NeurIPS | 3 |
| 2025 | Resource Allocation in Wideband Cooperative ISAC SystemsabstractThis paper investigates the resource allocation problem for multi-user wideband cooperative integrated sensing and communication (ISAC) networks based on orthogonal frequency-division multiplexing (OFDM) waveforms. In order to balance sensing and communication performance with limited spectrum resources in this wideband cell-free system, we aim to maximize the sum rate, encompassing both communication and radar rates, while adhering to constraints related to access point (AP) power and spectrum resources. We utilize alternate optimization (AO) methods to optimize power and spectrum resources separately. For power optimization, we employ the fractional programming (FP) algorithm to convert the problem into a convex one, which can be quickly solved by the primal-dual subgradient (PDS) method. As for subcarrier allocation optimization, we derive its closed-form solution. Simulation results indicate that the communication and sensing performance of the cell-free ISAC system outperforms that of the conventional centralized ISAC system. Chenhan Yuan, Boshi Wang, Zhiyuan Yu 0007, Cunhua Pan, Hong Ren |
VTC2025-Spring | 1 |
| 2025 | A Reinforcement Learning Framework for N-Ary Document-Level Relation ExtractionabstractKnowledge Bases (KBs) have become more complex because some facts in KBs include more than two entities. The construction and completion of these KBs require a new relation extraction task to retrieve complex facts from the text. To address this issue, we present a new N-ary Document-Level relation extraction task that involves extracting relations that 1) include an arbitrary number of entities, and 2) can span multiple sentences within a document. This new task requires inferring relation labels and entity completeness, i.e., whether the entities in the document are (insufficient to describe the relation. We propose a reinforcement learning-based relation classifier training framework that can adapt most existing binary document-level relation extractors to this task. Extensive experimental evaluation demonstrates that our proposed framework is effective in reducing the impact of noise introduced by distant supervision or unrelated sentences in the document. Chenhan Yuan, Ryan Rossi, Andrew Katz, Hoda Eldardiry |
IEEE Trans. Big Data | 1 |
| 2024 | Predicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language ModelabstractTransformer-based large language models (LLMs) exhibit limitations such as generating unsafe responses, unreliable reasoning, etc. Existing inference intervention approaches attempt to mitigate these issues by finetuning additional models to produce calibration signals (such as rewards) that guide the LLM's decoding process.However, this solution introduces substantial time and space overhead due to the separate models required.This work proposes NOn-disruptive parameters insertion (Otter), inserting extra parameters into the transformer architecture to predict calibration signals along with the original LLM output.Otter offers state-of-the-art performance on multiple demanding tasks while saving up to 86.5% extra space and 98.5% extra time.Furthermore, Otter seamlessly integrates with existing inference engines, requiring only a oneline code change, and the original model response remains accessible after the parameter insertion. Chenhan Yuan, Fei Huang 0005, Ru Peng, Keming Lu, Bowen Yu 0002, Chang Zhou 0005, Jingren Zhou 0001 |
EMNLP | 1 |
| 2024 | Dólares or Dollars? Unraveling the Bilingual Prowess of Financial LLMs Between Spanish and EnglishabstractDespite Spanish's pivotal role in the global finance industry, a pronounced gap exists in Spanish financial natural language processing (NLP) and application studies compared to English, especially in the era of large language models (LLMs).To bridge this gap, we unveil Toisón de Oro, the first bilingual framework that establishes instruction datasets, finetuned LLMs, and evaluation benchmark for financial LLMs in Spanish joint with English.We construct a rigorously curated bilingual instruction dataset including over 144K Spanish and English samples from 15 datasets covering 7 tasks.Harnessing this, we introduce FinMA-ES, an LLM designed for bilingual financial applications.We evaluate our model and existing LLMs using FLARE-ES, the first comprehensive bilingual evaluation benchmark with 21 datasets covering 9 tasks.The FLARE-ES benchmark results Xiao Zhang 0060, Ruoyu Xiang, Chenhan Yuan, Duanyu Feng, Weiguang Han, Alejandro Lopez-Lira, Xiao-Yang Liu, Meikang Qiu, Sophia Ananiadou, Min Peng 0002, Jimin Huang, Qianqian Xie |
KDD | 3 |
| 2024 | FinBen: A Holistic Financial Benchmark for Large Language ModelsabstractLLMs have transformed NLP and shown promise in various fields, yet their potential in finance is underexplored due to a lack of comprehensive benchmarks, the rapid development of LLMs, and the complexity of financial tasks. In this paper, we introduce FinBen, the first extensive open-source evaluation benchmark, including 42 datasets spanning 24 financial tasks, covering eight critical aspects: information extraction (IE), textual analysis, question answering (QA), text generation, risk management, forecasting, decision-making, and bilingual (English and Spanish). FinBen offers several key innovations: a broader range of tasks and datasets, the first evaluation of stock trading, novel agent and Retrieval-Augmented Generation (RAG) evaluation, and two novel datasets for regulations and stock trading. Our evaluation of 21 representative LLMs, including GPT-4, ChatGPT, and the latest Gemini, reveals several key findings: While LLMs excel in IE and textual analysis, they struggle with advanced reasoning and complex tasks like text generation and forecasting. GPT-4 excels in IE and stock trading, while Gemini is better at text generation and forecasting. Instruction-tuned LLMs improve textual analysis but offer limited benefits for complex tasks such as QA. FinBen has been used to host the first financial LLMs shared task at the FinNLP-AgentScen workshop during IJCAI-2024, attracting 12 teams. Their novel solutions outperformed GPT-4, showcasing FinBen's potential to drive innovations in financial LLMs. All datasets and code are publicly available for the research community, with results shared and updated regularly on the Open Financial LLM Leaderboard. Qianqian Xie, Weiguang Han, Ruoyu Xiang, Xiao Zhang 0060, Yueru He, Mengxi Xiao, Yongfu Dai, Duanyu Feng, Yijing Xu, Haoqiang Kang, Ziyan Kuang, Chenhan Yuan, Kailai Yang, Zheheng Luo, Zhiwei Liu 0003, Guojun Xiong, Zhiyang Deng, Yuechen Jiang, Zhiyuan Yao 0001, Haohang Li, Yangyang Yu, Gang Hu 0003, Xiao-Yang Liu, Alejandro Lopez-Lira, Benyou Wang, Yanzhao Lai, Min Peng 0002, Sophia Ananiadou, Jimin Huang |
NeurIPS | 14 |
| 2024 | Back to the Future: Towards Explainable Temporal Reasoning with Large Language ModelsabstractTemporal reasoning is a crucial natural language processing (NLP) task, providing a nuanced understanding of time-sensitive contexts within textual data. Although recent advancements in Large Language Models (LLMs) have demonstrated their potential in temporal reasoning, the predominant focus has been on tasks such as temporal expression detection, normalization, and temporal relation extraction. These tasks are primarily designed for the extraction of direct and past temporal cues from given contexts and to engage in simple reasoning processes. A significant gap remains when considering complex reasoning tasks such as event forecasting, which requires multi-step temporal reasoning on events and prediction on the future timestamp. Another notable limitation of existing methods is their incapability to illustrate their reasoning process for explaining their prediction, hindering explainability. In this paper, we introduce the first task of explainable temporal reasoning, to predict an event's occurrence at a future timestamp based on context which requires multiple reasoning over multiple events, and subsequently provide a clear explanation for their prediction. Our task offers a comprehensive evaluation of both the LLMs' complex temporal reasoning ability, the future event prediction ability, and explainability-a critical attribute for AI applications. To support this task, we present the first instruction-tuning dataset of explainable temporal reasoning (ExpTime) with 26k derived from the temporal knowledge graph datasets, using a novel knowledge-graph-instructed-generation strategy. Based on the dataset, we propose the first open-source LLM series TimeLlaMA based on the foundation LLM LlaMA2, with the ability of instruction following for explainable temporal reasoning. We compare the performance of our method and a variety of LLMs, where our method achieves the state-of-the-art performance of temporal prediction and explanation generation. We also explore the impact of instruction tuning and different training sizes of instruction-tuning data, highlighting LLM's capabilities and limitations in complex temporal prediction and explanation generation. Chenhan Yuan, Qianqian Xie, Jimin Huang, Sophia Ananiadou |
WWW | 1 |
| 2024 | Temporal relation extraction with contrastive prototypical samplingabstractTemporal relation extraction aims to infer the temporal order of either two events in the document. Because of the nature of events in real life, severe imbalanced temporal relation classes exist in the temporal relation extraction task. Even though various methods have been proposed to improve the overall performance, the accuracy of these methods on minority temporal classes is limited. In this work, we present a contrastive prototypical learning architecture to address this problem, which explicitly models the spatial similarity between instances in the embedding space so that instances from minority classes can be distinguished from the large classes. To make it compatible with current temporal relation extraction settings, we propose a novel sampling memory queue-based method so that the architecture can be applied to a limited batch size scenario. We further design a context encoding layer that incorporates both contextualized information and linguistic features such as tense information and dependency. Our extensive experiments on TimeBank-Dense, TDDiscourse, and MATRES datasets demonstrate that our model can significantly improve the performance of minority relation classes and, therefore increase the overall learning ability. Chenhan Yuan, Qianqian Xie, Sophia Ananiadou |
Knowl. Based Syst. | 1 |
| 2022 | Clustering-based Unsupervised Generative Relation ExtractionabstractExisting unsupervised relation extraction methods work by extracting sentence features and using these features as inputs to train a generative model. This model is then used to cluster similar relations. However, these methods do not consider correlations between sentences with the same entity pair during training, which can negatively impact model performance. To address this issue, we propose a Clustering-based Unsupervised generative Relation Extraction (CURE) framework that leverages an Encoder-Decoder architecture to train a relation extractor as the encoder. Given multiple sentences with the same entity pair as inputs, CURE is deployed by predicting the shortest path between entity pairs on the dependency graph of one of the sentences. After that, we extract the relation information using the encoder. Then, entity pairs that share the same relation are clustered based on their corresponding relation information. Each cluster is labeled based on the words in the shortest paths corresponding to the entity pairs in each cluster. Experimental results demonstrate the effectiveness of CURE compared to state-of-the-art models across all benchmark datasets. Chenhan Yuan, Ryan Rossi, Andrew Katz, Hoda Eldardiry |
IEEE Big Data | 1 |
| 2021 | Unsupervised Relation Extraction: A Variational Autoencoder ApproachabstractUnsupervised relation extraction works by clustering entity pairs that have the same relations in the text. Some existing variational autoencoder (VAE)-based approaches train the relation extraction model as an encoder that generates relation classifications.A decoder is trained along with the encoder to reconstruct the encoder input based on the encodergenerated relation classifications.These classifications are a latent variable so they are required to follow a pre-defined prior distribution which results in unstable training.We propose a VAE-based unsupervised relation extraction technique that overcomes this limitation by using the classifications as an intermediate variable instead of a latent variable.Specifically, classifications are conditioned on sentence input, while the latent variable is conditioned on both the classifications and the sentence input.This allows our model to connect the decoder with the encoder without putting restrictions on the classification distribution; which improves training stability.Our approach is evaluated on the NYT dataset and outperforms state-of-the-art methods. Chenhan Yuan, Hoda Eldardiry |
EMNLP (1) | 1 |