Yueru He

dblp:371/1115 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
0009-0003-0514-8266ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 54% Multi-agent systems · 26% Trustworthy machine learning · 9%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Computational finance and economics · 100%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 15 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
1.532026
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application · ACL (1) 2026
Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation · SIGIR 2026
FinBen: A Holistic Financial Benchmark for Large Language Models · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems
agent-based simulation
1.012026
When Agents Trade: Live Multi-Market Trading Arena for LLM Agents · WWW 2026
Natural language and speech › Language models and text generation › large language model evaluation › truthfulness evaluation
factuality verification
1.012026
All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection · ACL (1) 2026
Knowledge, reasoning and agents › Multi-agent systems
trading agents
1.012026
When Agents Trade: Live Multi-Market Trading Arena for LLM Agents · WWW 2026
Computational finance and economics › financial data analysis
financial text analysis
1.012026
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application · ACL (1) 2026
Recommender systems › domain-specific recommendation › service recommendation
financial recommendation
1.012026
Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation · SIGIR 2026
Recommender systems › domain-specific recommendation
stock recommendation
1.012026
Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation · SIGIR 2026
Natural language and speech › Language models and text generation › large language model evaluation
LLM agent evaluation
0.912025
INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent · ACL (1) 2025
Computational finance and economics
financial decision-making
0.912025
INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › document analysis
financial text analysis
0.812024
FinBen: A Holistic Financial Benchmark for Large Language Models · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems
0.812024
FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making · NeurIPS 2024
Computational finance and economics
portfolio management
0.812024
FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making · NeurIPS 2024
Natural language and speech › Language models and text generation
LLM agents
0.312026
When Agents Trade: Live Multi-Market Trading Arena for LLM Agents · WWW 2026
Natural language and speech › Language models and text generation
multimodal language model
0.312026
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
financial question answering
0.212024
FinBen: A Holistic Financial Benchmark for Large Language Models · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

large language model · 3.3benchmark construction · 3.0multilingual evaluation · 2.0longitudinal benchmarking · 2.0large language model agents · 2.0conversational recommendation · 2.0agent · 1.7counterfactual reasoning · 1.0verbal reinforcement · 0.8retrieval-augmented generation · 0.8multi-agent system · 0.8instruction tuning · 0.8
YearPublicationVenuePosition
2026 All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection
abstract
Yuechen Jiang, Zhiwei Liu, Yupeng Cao, Yueru He, Ziyang Xu, Chen Xu, Zhiyang Deng, Prayag Tiwari, Xi Chen, Alejandro Lopez-Lira, Jimin Huang, Junichi Tsujii, Sophia Ananiadou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuechen Jiang, Zhiwei Liu 0003, Yupeng Cao, Yueru He, Zhiyang Deng, Prayag Tiwari, Xi Chen 0003, Alejandro Lopez-Lira, Jimin Huang, Jun'ichi Tsujii, Sophia Ananiadou
ACL (1)4
2026 MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
abstract
Xueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang, Yueru He, Yang Ren, Mingyang Jiang, Vincent Jim Zhang, Yuqing Guo, Jeff Zhao, Huan He, Yi Han, Yun Feng, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Xiaoyu Wang, Penglei Gao, Shengyuan Lin, Keyi Wang, Shanshan Yang, Yilun Zhao, Zhiwei Liu, Peng Lu, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen, Junichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xueqing Peng, Lingfei Qian, Yan Wang 0015, Ruoyu Xiang, Yueru He, Mingyang Jiang, Vincent Jim Zhang, Jeff Zhao, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Penglei Gao, Shengyuan Lin, Yilun Zhao 0001, Zhiwei Liu 0003, Peng Lu 0006, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen 0002, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E. Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen 0003, Jun'ichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie
ACL (1)5
2026 Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation
abstract
Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or short-sighted under market volatility and may conflict with a user's long-term goals. Treating what users chose as the sole ground truth, therefore, conflates behavioral imitation with decision quality. We introduce Conv-FinRe, a conversational and longitudinal benchmark for stock recommendation that evaluates LLMs beyond behavior matching. Given an onboarding interview, step-wise market context, and advisory dialogues, models must generate rankings over a fixed investment horizon. Crucially, Conv-FinRe provides multi-view references that distinguish descriptive behavior from normative utility grounded in investor-specific risk preferences, enabling diagnosis of whether an LLM follows rational analysis, mimics user noise, or is driven by market momentum. We build the benchmark from real market data and human decision trajectories, instantiate controlled advisory conversations, and evaluate a suite of state-of-the-art LLMs. Results reveal a persistent tension between rational decision quality and behavioral alignment: models that perform well on utility-based ranking often fail to match user choices, whereas behaviorally aligned models can overfit short-term noise. The dataset is publicly released on Hugging Face. https://huggingface.co/collections/TheFinAI/conv-finre, and the codebase is available on GitHub. https://github.com/The-FinAI/Conv-FinRe.
Yan Wang 0015, Lingfei Qian, Yueru He, Xueqing Peng, Dongji Feng, Zhuohan Xie, Vincent Jim Zhang, Fengran Mo, Jimin Huang, Yankai Chen 0001, Jian-Yun Nie
SIGIR4
2026 When Agents Trade: Live Multi-Market Trading Arena for LLM Agents
Lingfei Qian, Xueqing Peng, Hanley Smith, Yueru He, Haohang Li, Yupeng Cao, Yangyang Yu, Guojun Xiong, Peng Lu 0006, Yan Wang 0015, Vincent Jim Zhang, Alejandro Lopez-Lira, Jimin Huang, Jian-Yun Nie, Sophia Ananiadou
WWW5
2025 INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent
abstract
Haohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji, Zhiyang Deng, Yueru He, Yuechen Jiang, Zining Zhu, K.p. Subbalakshmi, Jimin Huang, Lingfei Qian, Xueqing Peng, Jordan W. Suchow, Qianqian Xie. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Haohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji, Zhiyang Deng, Yueru He, Yuechen Jiang, Zining Zhu 0001, K. P. Subbalakshmi, Jimin Huang, Lingfei Qian, Xueqing Peng, Jordan W. Suchow, Qianqian Xie
ACL (1)6
2024 FinBen: A Holistic Financial Benchmark for Large Language Models
abstract
LLMs have transformed NLP and shown promise in various fields, yet their potential in finance is underexplored due to a lack of comprehensive benchmarks, the rapid development of LLMs, and the complexity of financial tasks. In this paper, we introduce FinBen, the first extensive open-source evaluation benchmark, including 42 datasets spanning 24 financial tasks, covering eight critical aspects: information extraction (IE), textual analysis, question answering (QA), text generation, risk management, forecasting, decision-making, and bilingual (English and Spanish). FinBen offers several key innovations: a broader range of tasks and datasets, the first evaluation of stock trading, novel agent and Retrieval-Augmented Generation (RAG) evaluation, and two novel datasets for regulations and stock trading. Our evaluation of 21 representative LLMs, including GPT-4, ChatGPT, and the latest Gemini, reveals several key findings: While LLMs excel in IE and textual analysis, they struggle with advanced reasoning and complex tasks like text generation and forecasting. GPT-4 excels in IE and stock trading, while Gemini is better at text generation and forecasting. Instruction-tuned LLMs improve textual analysis but offer limited benefits for complex tasks such as QA. FinBen has been used to host the first financial LLMs shared task at the FinNLP-AgentScen workshop during IJCAI-2024, attracting 12 teams. Their novel solutions outperformed GPT-4, showcasing FinBen's potential to drive innovations in financial LLMs. All datasets and code are publicly available for the research community, with results shared and updated regularly on the Open Financial LLM Leaderboard.
Qianqian Xie, Weiguang Han, Ruoyu Xiang, Xiao Zhang 0060, Yueru He, Mengxi Xiao, Yongfu Dai, Duanyu Feng, Yijing Xu, Haoqiang Kang, Ziyan Kuang, Chenhan Yuan, Kailai Yang, Zheheng Luo, Zhiwei Liu 0003, Guojun Xiong, Zhiyang Deng, Yuechen Jiang, Zhiyuan Yao 0001, Haohang Li, Yangyang Yu, Gang Hu 0003, Xiao-Yang Liu, Alejandro Lopez-Lira, Benyou Wang, Yanzhao Lai, Min Peng 0002, Sophia Ananiadou, Jimin Huang
NeurIPS6
2024 FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making
abstract
Large language models (LLMs) have demonstrated notable potential in conducting complex tasks and are increasingly utilized in various financial applications. However, high-quality sequential financial investment decision-making remains challenging. These tasks require multiple interactions with a volatile environment for every decision, demanding sufficient intelligence to maximize returns and manage risks. Although LLMs have been used to develop agent systems that surpass human teams and yield impressive investment returns, opportunities to enhance multi-source information synthesis and optimize decision-making outcomes through timely experience refinement remain unexplored. Here, we introduce FinCon, an LLM-based multi-agent framework tailored for diverse financial tasks. Inspired by effective real-world investment firm organizational structures, FinCon utilizes a manager-analyst communication hierarchy. This structure allows for synchronized cross-functional agent collaboration towards unified goals through natural language interactions and equips each agent with greater memory capacity than humans. Additionally, a risk-control component in FinCon enhances decision quality by episodically initiating a self-critiquing mechanism to update systematic investment beliefs. The conceptualized beliefs serve as verbal reinforcement for the future agent’s behavior and can be selectively propagated to the appropriate node that requires knowledge updates. This feature significantly improves performance while reducing unnecessary peer-to-peer communication costs. Moreover, FinCon demonstrates strong generalization capabilities in various financial tasks, including stock trading and portfolio management.
Yangyang Yu, Zhiyuan Yao 0001, Haohang Li, Zhiyang Deng, Yuechen Jiang, Yupeng Cao, Jordan W. Suchow, Zhenyu Cui, Zhaozhuo Xu, K. P. Subbalakshmi, Guojun Xiong, Yueru He, Jimin Huang, Qianqian Xie
NeurIPS15