Yining Ye

dblp:274/2421 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 82% Multi-agent systems · 18%
Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 50% Empirical software engineering · 50%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
LLM agents
0.912025
Rational Decision-Making Agent with Learning Internal Utility Judgment · ICLR 2025
Knowledge, reasoning and agents › Multi-agent systems › multi-agent learning
utility learning
0.912025
Rational Decision-Making Agent with Learning Internal Utility Judgment · ICLR 2025
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
function calling
0.812024
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs · ICLR 2024
Natural language and speech › Language models and text generation
instruction tuning
0.812024
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs · ICLR 2024
Natural language and speech › Language models and text generation › LLM agents
tool use
0.812024
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs · ICLR 2024
Empirical software engineering › mining software repositories
github
0.312025
Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub · ACL (1) 2025
Software maintenance and evolution
software ecosystems
0.312025
Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

large language model prompting · 1.7pairwise comparison · 0.9iterative experience exploration · 0.9elo-based utility learning · 0.9neural API retrieval · 0.8depth-first search · 0.8
YearPublicationVenuePosition
2025 Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub
abstract
Bohan Lyu, Xin Cong, Heyang Yu, Pan Yang, Cheng Qian, Zihe Wang, Yujia Qin, Yining Ye, Yaxi Lu, Chen Qian, Zhong Zhang, Yukun Yan, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Bohan Lyu 0001, Xin Cong, Heyang Yu, Pan Yang 0022, Cheng Qian 0008, Yujia Qin, Yining Ye, Yaxi Lu, Zhong Zhang 0004, Yukun Yan, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)8
2025 Rational Decision-Making Agent with Learning Internal Utility Judgment
abstract
With remarkable advancements, large language models (LLMs) have attracted significant efforts to develop LLM-based agents capable of executing intricate multi-step decision-making tasks. Existing approaches predominantly build upon the external performance measure to guide the decision-making process but the reliance on the external performance measure as prior is problematic in real-world scenarios, where such prior may be unavailable, flawed, or even erroneous. For genuine autonomous decision-making for LLM-based agents, it is imperative to develop rationality from their posterior experiences to judge the utility of each decision independently. In this work, we propose RaDAgent (Rational Decision-Making Agent), which fosters the development of its rationality through an iterative framework involving Experience Exploration and Utility Learning. Within this framework, Elo-based Utility Learning is devised to assign Elo scores to individual decision steps to judge their utilities via pairwise comparisons. Consequently, these Elo scores guide the decision-making process to derive optimal outcomes. Experimental results on the Game of 24, WebShop, ToolBench and RestBench datasets demonstrate RaDAgent’s superiority over baselines, achieving about 7.8% improvement on average. Besides, RaDAgent also can reduce costs (ChatGPT API calls), highlighting its effectiveness and efficiency.
Yining Ye, Xin Cong, Shizuo Tian, Yujia Qin, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR1
2024 ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
abstract
Despite the advancements of open-source large language models (LLMs), e.g., LLaMA, they remain significantly limited in tool-use capabilities, i.e., using external tools (APIs) to fulfill human instructions. The reason is that current instruction tuning largely focuses on basic language tasks but ignores the tool-use domain. This is in contrast to the excellent tool-use capabilities of state-of-the-art (SOTA) closed-source LLMs, e.g., ChatGPT. To bridge this gap, we introduce ToolLLM, a general tool-use framework encompassing data construction, model training, and evaluation. We first present ToolBench, an instruction-tuning dataset for tool use, which is constructed automatically using ChatGPT. Specifically, the construction can be divided into three stages: (i) API collection: we collect 16,464 real-world RESTful APIs spanning 49 categories from RapidAPI Hub; (ii) instruction generation: we prompt ChatGPT to generate diverse instructions involving these APIs, covering both single-tool and multi-tool scenarios; (iii) solution path annotation: we use ChatGPT to search for a valid solution path (chain of API calls) for each instruction. To enhance the reasoning capabilities of LLMs, we develop a novel depth-first search-based decision tree algorithm. It enables LLMs to evaluate multiple reasoning traces and expand the search space. Moreover, to evaluate the tool-use capabilities of LLMs, we develop an automatic evaluator: ToolEval. Based on ToolBench, we fine-tune LLaMA to obtain an LLM ToolLLaMA, and equip it with a neural API retriever to recommend appropriate APIs for each instruction. Experiments show that ToolLLaMA demonstrates a remarkable ability to execute complex instructions and generalize to unseen APIs, and exhibits comparable performance to ChatGPT. Our ToolLLaMA also demonstrates strong zero-shot generalization ability in an out-of-distribution tool-use dataset: APIBench.
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin 0001, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou 0016, Mark Gerstein, Dahai Li, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR3
2021 Country Image in COVID-19 Pandemic: A Case Study of China
abstract
Country image has a profound influence on international relations and economic development. In the worldwide outbreak of COVID-19, countries and their people display different reactions, resulting in diverse perceived images among foreign public. Therefore, in this article, we take China as a specific and typical case and investigate its image with aspect-based sentiment analysis on a large-scale Twitter dataset. To our knowledge, this is the first study to explore country image in such a fine-grained way. To perform the analysis, we first build a manually-labeled Twitter dataset with aspect-level sentiment annotations. Afterward, we conduct the aspect-based sentiment analysis with BERT to explore the image of China. We discover an overall sentiment change from non-negative to negative in the general public, and explain it with the increasing mentions of negative ideology-related aspects and decreasing mentions of non-negative fact-based aspects. Further investigations into different groups of Twitter users, including U.S. Congress members, English media, and social bots, reveal different patterns in their attitudes toward China. This article provides a deeper understanding of the changing image of China in COVID-19 pandemic. Our research also demonstrates how aspect-based sentiment analysis can be applied in social science researches to deliver valuable insights.
Fanchao Qi, Yining Ye, Zhiyuan Liu 0001, Maosong Sun 0001, Jianbin Jin
IEEE Trans. Big Data4