Zhiyu Chen 0002

dblp:71/1661-2 · also Zhiyu Zoey Chen · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0006-6028-4836ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 RescueLens: LLM-Powered Triage and Action on Volunteer Feedback for Food Rescue
abstract
Food rescue organizations simultaneously tackle food insecurity and waste by working with volunteers to redistribute food from donors who have excess to recipients who need it. Volunteer feedback allows food rescue organizations to identify issues early and ensure volunteer satisfaction. However, food rescue organizations monitor feedback manually, which can be cumbersome and labor-intensive, making it difficult to prioritize which issues are most important. In this work, we investigate how large language models (LLMs) assist food rescue organizers in understanding and taking action based on volunteer experiences. We work with 412 Food Rescue, a large food rescue organization based in Pittsburgh, Pennsylvania, to design RescueLens, an LLM-powered tool that automatically categorizes volunteer feedback, suggests donors and recipients to follow up with, and updates volunteer directions based on feedback. We evaluate the performance of RescueLens on an annotated dataset, and show that it can recover 96% of volunteer issues at 71% precision. Moreover, by ranking donors and recipients according to their rates of volunteer issues, RescueLens allows organizers to focus on 0.5% of donors responsible for more than 30% of volunteer issues. RescueLens is now deployed at 412 Food Rescue and through semi-structured interviews with organizers, we find that RescueLens streamlines the feedback process so organizers better allocate their time.
Naveen Raman 0001, Jingwu Tang, Zhiyu Chen 0002, Zheyuan Shi, Sean Hudson, Ameesh Kapoor, Fei Fang 0001
AAAI3
2025 Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning
abstract
Instruction Fine-Tuning (IFT) significantly enhances the zero-shot capabilities of pretrained Large Language Models (LLMs). While coding data is known to boost LLM reasoning abilities during pretraining, its role in activating internal reasoning capacities during IFT remains understudied. This paper investigates a key question: How does coding data impact LLMs' reasoning capacities during IFT stage? To explore this, we thoroughly examine the impact of coding data across different coding data proportions, model families, sizes, and reasoning domains, from various perspectives. Specifically, we create three IFT datasets with increasing coding data proportions, fine-tune six LLM backbones across different families and scales on these datasets, evaluate the tuned models' performance across twelve tasks in three reasoning domains, and analyze the outcomes from three broad-to-granular perspectives: overall, domain-level, and task-specific. Our holistic analysis provides valuable insights into each perspective. First, coding data tuning enhances the overall reasoning capabilities of LLMs across different model families and scales. Moreover, while the impact of coding data varies by domain, it shows consistent trends within each domain across different model families and scales. Additionally, coding data generally provides comparable task-specific benefits across model families, with optimal proportions in IFT datasets being task-dependent.
Xinlu Zhang, Zhiyu Chen 0002, Xi Ye 0003, Xianjun Yang, Lichang Chen, William Yang Wang, Linda R. Petzold
AAAI2
2025 Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
abstract
Agentic Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) by enabling dynamic, multi-step reasoning and information retrieval.However, these systems often exhibit sub-optimal search behaviors like over-search (retrieving redundant information) and under-search (failing to initiate retrieval for necessary information), which hinder efficiency and reliability.This work formally defines and quantifies these behaviors, revealing their prevalence across multiple QA datasets and agentic RAG systems (e.g., one model could have avoided searching in 27.7% of its search steps).Furthermore, we demonstrate a crucial link between these inefficiencies and the models' uncertainty regarding their own knowledge boundaries, where response accuracy correlates with model's uncertainty or confidence in its search decisions.To address this, we propose β-GRPO, a reinforcement learning-based training method that incorporates confidence threshold to reward high-certainty search decisions.Experiments on seven QA benchmarks show that β-GRPO enable a 3B model with better agentic RAG ability, outperforming other strong baselines with a 4% higher average exact match score, with lower over-search and under-search rate 1 .
Xinlu Zhang, Xinya Du, Zhiyu Chen 0002
EMNLP5
2025 LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
abstract
Shuo Yan, Ruochen Li, Ziming Luo, Zimu Wang, Daoyang Li, Liqiang Jing, Kaiyu He, Peilin Wu, Juntong Ni, George Michalopoulos, Yue Zhang, Ziyang Zhang, Mian Zhang, Zhiyu Chen, Xinya Du. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ziming Luo, Daoyang Li, Liqiang Jing, Kaiyu He, Juntong Ni, George Michalopoulos, Zhiyu Chen 0002, Xinya Du
EMNLP14
2025 CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
abstract
Mian Zhang, Xianjun Yang, Xinlu Zhang, Travis Labrum, Jamie C. Chiu, Shaun M. Eack, Fei Fang, William Yang Wang, Zhiyu Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xianjun Yang, Xinlu Zhang, Travis Labrum, Jamie C. Chiu, Shaun M. Eack, Fei Fang 0001, William Yang Wang, Zhiyu Chen 0002
NAACL (Long Papers)9
2024 Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
abstract
Zekun Li, Zhiyu Zoey Chen, Mike Ross, Patrick Huber, Seungwhan Moon, Zhaojiang Lin, Luna Dong, Adithya Sagar, Xifeng Yan, Paul A. Crook. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zekun Li 0001, Zhiyu Chen 0002, Mike Ross, Patrick Huber, Seungwhan Moon, Zhaojiang Lin, Xin Dong 0001, Adithya Sagar, Xifeng Yan, Paul A. Crook
ACL (1)2
2024 PATIENT-ψ: Using Large Language Models to Simulate Patients for Training Mental Health Professionals
abstract
Ruiyi Wang, Stephanie Milani, Jamie C. Chiu, Jiayin Zhi, Shaun M. Eack, Travis Labrum, Samuel M Murphy, Nev Jones, Kate V Hardy, Hong Shen, Fei Fang, Zhiyu Chen. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Ruiyi Wang, Stephanie Milani, Jamie C. Chiu, Jiayin Zhi, Shaun M. Eack, Travis Labrum, Samuel M. Murphy, Nev Jones, Kate Hardy, Hong Shen 0004, Fei Fang 0001, Zhiyu Chen 0002
EMNLP12
2023 Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling
abstract
Health conditions among patients in intensive care units (ICUs) are monitored via electronic health records (EHRs), composed of numerical time series and lengthy clinical note sequences, both taken at $\textit{irregular}$ time intervals. Dealing with such irregularity in every modality, and integrating irregularity into multimodal representations to improve medical predictions, is a challenging problem. Our method first addresses irregularity in each single modality by (1) modeling irregular time series by dynamically incorporating hand-crafted imputation embeddings into learned interpolation embeddings via a gating mechanism, and (2) casting a series of clinical note representations as multivariate irregular time series and tackling irregularity via a time attention mechanism. We further integrate irregularity in multimodal fusion with an interleaved attention mechanism across temporal steps. To the best of our knowledge, this is the first work to thoroughly model irregularity in multimodalities for improving medical predictions. Our proposed methods for two medical prediction tasks consistently outperforms state-of-the-art (SOTA) baselines in each single modality and multimodal fusion scenarios. Specifically, we observe relative improvements of 6.5%, 3.6%, and 4.3% in F1 for time series, clinical notes, and multimodal fusion, respectively. These results demonstrate the effectiveness of our methods and the importance of considering irregularity in multimodal EHRs.
Xinlu Zhang, Zhiyu Chen 0002, Xifeng Yan, Linda R. Petzold
ICML3
2023 Robust NLP for Finance (RobustFin)
abstract
Natural language processing (NLP) technologies have been widely applied in business domains such as e-commerce and customer service, but their adoption in the financial sector has been constrained by industry-specific performance standards and regulatory restrictions. This challenge has created new opportunities for core research in related areas. Recent advancements in NLP, such as the advent of large language models, has encouraged adoption in the finance sector. However, compared to other domains, finance has stricter requirements for robustness, explainability, and generalizability. Given this background, we propose to organize the first Robust NLP for Finance (RobustFin) workshop at KDD '23 to encourage the study of and research on robustness and explainability technologies with regard to financial NLP. The goal of the workshop is to extend the applications of NLP in finance, while motivating further research in robust NLP.
Sameena Shah, Xiaodan Zhu 0001, Gerard de Melo, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Zhiyu Chen 0002
KDD8
2022 Inductive Relation Prediction by BERT
abstract
Relation prediction in knowledge graphs is dominated by embedding based methods which mainly focus on the transductive setting. Unfortunately, they are not able to handle inductive learning where unseen entities and relations are present and cannot take advantage of prior knowledge. Furthermore, their inference process is not easily explainable. In this work, we propose an all-in-one solution, called BERTRL (BERT-based Relational Learning), which leverages pre-trained language model and fine-tunes it by taking relation instances and their possible reasoning paths as training samples. BERTRL outperforms the SOTAs in 15 out of 18 cases in both inductive and transductive settings. Meanwhile, it demonstrates strong generalization capability in few-shot learning and is explainable. The data and code can be found at https://github.com/zhw12/BERTRL.
Hanwen Zha, Zhiyu Chen 0002, Xifeng Yan
AAAI2
2022 ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering
abstract
With the recent advance in large pre-trained language models, researchers have achieved record performances in NLP tasks that mostly focus on language pattern matching.The community is experiencing the shift of the challenge from how to model language to the imitation of complex reasoning abilities like human beings.In this work, we investigate the application domain of finance that involves realworld, complex numerical reasoning.We propose a new large-scale dataset, CONVFINQA, aiming to study the chain of numerical reasoning in conversational question answering.Our dataset poses great challenge in modeling longrange, complex numerical reasoning paths in real-world conversations.We conduct comprehensive experiments and analyses with both the neural symbolic methods and the promptingbased methods, to provide insights into the reasoning mechanisms of these two divisions.We believe our new dataset should serve as a valuable resource to push forward the exploration of real-world, complex reasoning tasks as the next research focus.Our dataset and code is publicly available 1 .
Zhiyu Chen 0002, Charese Smiley, Sameena Shah, William Yang Wang
EMNLP1
2021 FinQA: A Dataset of Numerical Reasoning over Financial Data
abstract
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, William Yang Wang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Zhiyu Chen 0002, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao 'Kenneth' Huang, Bryan R. Routledge, William Yang Wang
EMNLP (1)1
2020 Logical Natural Language Generation from Open-Domain Tables
abstract
Neural natural language generation (NLG) models have recently shown remarkable progress in fluency and coherence.However, existing studies on neural NLG are primarily focused on surface-level realizations with limited emphasis on logical inference, an important aspect of human thinking and language.In this paper, we suggest a new NLG task where a model is tasked with generating natural language statements that can be logically entailed by the facts in an open-domain semi-structured table.To facilitate the study of the proposed logical NLG problem, we use the existing Tab-Fact dataset (Chen et al., 2019) featured with a wide range of logical/symbolic inferences as our testbed, and propose new automatic metrics to evaluate the fidelity of generation models w.r.t.logical inference.The new task poses challenges to the existing monotonic generation frameworks due to the mismatch between sequence order and logical order.In our experiments, we comprehensively survey different generation architectures (LSTM, Transformer, Pre-Trained LM) trained with different algorithms (RL, Adversarial Training, Coarse-to-Fine) on the dataset and made following observations: 1) Pre-Trained LM can significantly boost both the fluency and logical fidelity metrics, 2) RL and Adversarial Training are trading fluency for fidelity, 3) Coarse-to-Fine generation can help partially alleviate the fidelity issue while maintaining high language fluency.
Wenhu Chen, Jianshu Chen, Yu Su 0001, Zhiyu Chen 0002, William Yang Wang
ACL4
2020 Few-Shot NLG with Pre-Trained Language Model
abstract
Neural-based end-to-end approaches to natural language generation (NLG) from structured data or knowledge are data-hungry, making their adoption for real-world applications difficult with limited data. In this work, we propose the new task of few-shot natural language generation. Motivated by how humans tend to summarize tabular data, we propose a simple yet effective approach and show that it not only demonstrates strong performance but also provides good generalization across domains. The design of the model architecture is based on two aspects: content selection from input data and language modeling to compose coherent sentences, which can be acquired from prior knowledge. With just 200 training examples, across multiple domains, we show that our approach achieves very reasonable performances and outperforms the strongest baseline by an average of over 8.0 BLEU points improvement. Our code and data can be found at https://github.com/czyssrs/Few-Shot-NLG
Zhiyu Chen 0002, Harini Eavani, Wenhu Chen, Yinyin Liu, William Yang Wang
ACL1
2019 Global Textual Relation Embedding for Relational Understanding
abstract
Pre-trained embeddings such as word embeddings and sentence embeddings are fundamental tools facilitating a wide range of downstream NLP tasks.In this work, we investigate how to learn a general-purpose embedding of textual relations, defined as the shortest dependency path between entities.Textual relation embedding provides a level of knowledge between word/phrase level and sentence level, and we show that it can facilitate downstream tasks requiring relational understanding of the text.To learn such an embedding, we create the largest distant supervision dataset by linking the entire English ClueWeb09 corpus to Freebase.We use global co-occurrence statistics between textual and knowledge base relations as the supervision signal to train the embedding.Evaluation on two relational understanding tasks demonstrates the usefulness of the learned textual relation embedding.
Zhiyu Chen 0002, Hanwen Zha, Honglei Liu 0001, Wenhu Chen, Xifeng Yan, Yu Su 0001
ACL (1)1
2017 IMAP: An iterative method for aligning protein-protein interaction networks
abstract
Biological network alignment benefits the evolutionary and comparative biology by providing regions of topological and functional similarity between different species. However, most existing network aligners follow heuristic methods and only capture the static information that based purely on the original isolated networks, while there also exists valuable interactive information hidden in the resulted alignment that provides additional signals for further improvement. In this paper, we propose an iterative method IMAP to improve the quality of existing network aligners. IMAP starts from an imperfect seed alignment generated by any aligner, and then iteratively refines it by capturing interactive information hidden in current alignment until convergence. Within each iteration, we calculate the likelihood of pairwise alignment using supervised learning techniques, hence heuristic functions are no longer required. Furthermore, we extend IMAP to start from multiple seed aligners to combine their individual advantages. Comprehensive experiments indicate that IMAP improves existing network aligners significantly in terms of node correctness, topology conservation and biological similarity. Therefore, IMAP can benefit the subsequent cross-species biology researches by providing high-quality alignment between PPI networks.
Xuezhi Cao, Zhiyu Chen 0002, Yong Yu 0001
BIBM2