VLDB 2026 Research / reviewers in the wild / expert
Yueru He
dblp:371/1115
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0009-0003-0514-8266ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 54% Multi-agent systems · 26% Trustworthy machine learning · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Computational finance and economics · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 15 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model evaluation |
1.5 | 3 | 2026 | MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application · ACL (1) 2026 Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation · SIGIR 2026 FinBen: A Holistic Financial Benchmark for Large Language Models · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems
agent-based simulation |
1.0 | 1 | 2026 | When Agents Trade: Live Multi-Market Trading Arena for LLM Agents · WWW 2026 |
Natural language and speech › Language models and text generation › large language model evaluation › truthfulness evaluation
factuality verification |
1.0 | 1 | 2026 | All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection · ACL (1) 2026 |
Knowledge, reasoning and agents › Multi-agent systems
trading agents |
1.0 | 1 | 2026 | When Agents Trade: Live Multi-Market Trading Arena for LLM Agents · WWW 2026 |
Computational finance and economics › financial data analysis
financial text analysis |
1.0 | 1 | 2026 | MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application · ACL (1) 2026 |
Recommender systems › domain-specific recommendation › service recommendation
financial recommendation |
1.0 | 1 | 2026 | Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation · SIGIR 2026 |
Recommender systems › domain-specific recommendation
stock recommendation |
1.0 | 1 | 2026 | Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation · SIGIR 2026 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM agent evaluation |
0.9 | 1 | 2025 | INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent · ACL (1) 2025 |
Computational finance and economics
financial decision-making |
0.9 | 1 | 2025 | INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis › document analysis
financial text analysis |
0.8 | 1 | 2024 | FinBen: A Holistic Financial Benchmark for Large Language Models · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems |
0.8 | 1 | 2024 | FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making · NeurIPS 2024 |
Computational finance and economics
portfolio management |
0.8 | 1 | 2024 | FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making · NeurIPS 2024 |
Natural language and speech › Language models and text generation
LLM agents |
0.3 | 1 | 2026 | When Agents Trade: Live Multi-Market Trading Arena for LLM Agents · WWW 2026 |
Natural language and speech › Language models and text generation
multimodal language model |
0.3 | 1 | 2026 | MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application · ACL (1) 2026 |
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
financial question answering |
0.2 | 1 | 2024 | FinBen: A Holistic Financial Benchmark for Large Language Models · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
large language model · 3.3benchmark construction · 3.0multilingual evaluation · 2.0longitudinal benchmarking · 2.0large language model agents · 2.0conversational recommendation · 2.0agent · 1.7counterfactual reasoning · 1.0verbal reinforcement · 0.8retrieval-augmented generation · 0.8multi-agent system · 0.8instruction tuning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation DetectionabstractYuechen Jiang, Zhiwei Liu, Yupeng Cao, Yueru He, Ziyang Xu, Chen Xu, Zhiyang Deng, Prayag Tiwari, Xi Chen, Alejandro Lopez-Lira, Jimin Huang, Junichi Tsujii, Sophia Ananiadou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuechen Jiang, Zhiwei Liu 0003, Yupeng Cao, Yueru He, Zhiyang Deng, Prayag Tiwari, Xi Chen 0003, Alejandro Lopez-Lira, Jimin Huang, Jun'ichi Tsujii, Sophia Ananiadou |
ACL (1) | 4 |
| 2026 | MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial ApplicationabstractXueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang, Yueru He, Yang Ren, Mingyang Jiang, Vincent Jim Zhang, Yuqing Guo, Jeff Zhao, Huan He, Yi Han, Yun Feng, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Xiaoyu Wang, Penglei Gao, Shengyuan Lin, Keyi Wang, Shanshan Yang, Yilun Zhao, Zhiwei Liu, Peng Lu, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen, Junichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xueqing Peng, Lingfei Qian, Yan Wang 0015, Ruoyu Xiang, Yueru He, Mingyang Jiang, Vincent Jim Zhang, Jeff Zhao, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Penglei Gao, Shengyuan Lin, Yilun Zhao 0001, Zhiwei Liu 0003, Peng Lu 0006, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen 0002, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E. Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen 0003, Jun'ichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie |
ACL (1) | 5 |
| 2026 | Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial RecommendationabstractMost recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or short-sighted under market volatility and may conflict with a user's long-term goals. Treating what users chose as the sole ground truth, therefore, conflates behavioral imitation with decision quality. We introduce Conv-FinRe, a conversational and longitudinal benchmark for stock recommendation that evaluates LLMs beyond behavior matching. Given an onboarding interview, step-wise market context, and advisory dialogues, models must generate rankings over a fixed investment horizon. Crucially, Conv-FinRe provides multi-view references that distinguish descriptive behavior from normative utility grounded in investor-specific risk preferences, enabling diagnosis of whether an LLM follows rational analysis, mimics user noise, or is driven by market momentum. We build the benchmark from real market data and human decision trajectories, instantiate controlled advisory conversations, and evaluate a suite of state-of-the-art LLMs. Results reveal a persistent tension between rational decision quality and behavioral alignment: models that perform well on utility-based ranking often fail to match user choices, whereas behaviorally aligned models can overfit short-term noise. The dataset is publicly released on Hugging Face. https://huggingface.co/collections/TheFinAI/conv-finre, and the codebase is available on GitHub. https://github.com/The-FinAI/Conv-FinRe. Yan Wang 0015, Lingfei Qian, Yueru He, Xueqing Peng, Dongji Feng, Zhuohan Xie, Vincent Jim Zhang, Fengran Mo, Jimin Huang, Yankai Chen 0001, Jian-Yun Nie |
SIGIR | 4 |
| 2026 | When Agents Trade: Live Multi-Market Trading Arena for LLM Agents
Lingfei Qian, Xueqing Peng, Hanley Smith, Yueru He, Haohang Li, Yupeng Cao, Yangyang Yu, Guojun Xiong, Peng Lu 0006, Yan Wang 0015, Vincent Jim Zhang, Alejandro Lopez-Lira, Jimin Huang, Jian-Yun Nie, Sophia Ananiadou |
WWW | 5 |
| 2025 | INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based AgentabstractHaohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji, Zhiyang Deng, Yueru He, Yuechen Jiang, Zining Zhu, K.p. Subbalakshmi, Jimin Huang, Lingfei Qian, Xueqing Peng, Jordan W. Suchow, Qianqian Xie. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji, Zhiyang Deng, Yueru He, Yuechen Jiang, Zining Zhu 0001, K. P. Subbalakshmi, Jimin Huang, Lingfei Qian, Xueqing Peng, Jordan W. Suchow, Qianqian Xie |
ACL (1) | 6 |
| 2024 | FinBen: A Holistic Financial Benchmark for Large Language ModelsabstractLLMs have transformed NLP and shown promise in various fields, yet their potential in finance is underexplored due to a lack of comprehensive benchmarks, the rapid development of LLMs, and the complexity of financial tasks. In this paper, we introduce FinBen, the first extensive open-source evaluation benchmark, including 42 datasets spanning 24 financial tasks, covering eight critical aspects: information extraction (IE), textual analysis, question answering (QA), text generation, risk management, forecasting, decision-making, and bilingual (English and Spanish). FinBen offers several key innovations: a broader range of tasks and datasets, the first evaluation of stock trading, novel agent and Retrieval-Augmented Generation (RAG) evaluation, and two novel datasets for regulations and stock trading. Our evaluation of 21 representative LLMs, including GPT-4, ChatGPT, and the latest Gemini, reveals several key findings: While LLMs excel in IE and textual analysis, they struggle with advanced reasoning and complex tasks like text generation and forecasting. GPT-4 excels in IE and stock trading, while Gemini is better at text generation and forecasting. Instruction-tuned LLMs improve textual analysis but offer limited benefits for complex tasks such as QA. FinBen has been used to host the first financial LLMs shared task at the FinNLP-AgentScen workshop during IJCAI-2024, attracting 12 teams. Their novel solutions outperformed GPT-4, showcasing FinBen's potential to drive innovations in financial LLMs. All datasets and code are publicly available for the research community, with results shared and updated regularly on the Open Financial LLM Leaderboard. Qianqian Xie, Weiguang Han, Ruoyu Xiang, Xiao Zhang 0060, Yueru He, Mengxi Xiao, Yongfu Dai, Duanyu Feng, Yijing Xu, Haoqiang Kang, Ziyan Kuang, Chenhan Yuan, Kailai Yang, Zheheng Luo, Zhiwei Liu 0003, Guojun Xiong, Zhiyang Deng, Yuechen Jiang, Zhiyuan Yao 0001, Haohang Li, Yangyang Yu, Gang Hu 0003, Xiao-Yang Liu, Alejandro Lopez-Lira, Benyou Wang, Yanzhao Lai, Min Peng 0002, Sophia Ananiadou, Jimin Huang |
NeurIPS | 6 |
| 2024 | FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision MakingabstractLarge language models (LLMs) have demonstrated notable potential in conducting complex tasks and are increasingly utilized in various financial applications. However, high-quality sequential financial investment decision-making remains challenging. These tasks require multiple interactions with a volatile environment for every decision, demanding sufficient intelligence to maximize returns and manage risks. Although LLMs have been used to develop agent systems that surpass human teams and yield impressive investment returns, opportunities to enhance multi-source information synthesis and optimize decision-making outcomes through timely experience refinement remain unexplored. Here, we introduce FinCon, an LLM-based multi-agent framework tailored for diverse financial tasks. Inspired by effective real-world investment firm organizational structures, FinCon utilizes a manager-analyst communication hierarchy. This structure allows for synchronized cross-functional agent collaboration towards unified goals through natural language interactions and equips each agent with greater memory capacity than humans. Additionally, a risk-control component in FinCon enhances decision quality by episodically initiating a self-critiquing mechanism to update systematic investment beliefs. The conceptualized beliefs serve as verbal reinforcement for the future agent’s behavior and can be selectively propagated to the appropriate node that requires knowledge updates. This feature significantly improves performance while reducing unnecessary peer-to-peer communication costs. Moreover, FinCon demonstrates strong generalization capabilities in various financial tasks, including stock trading and portfolio management. Yangyang Yu, Zhiyuan Yao 0001, Haohang Li, Zhiyang Deng, Yuechen Jiang, Yupeng Cao, Jordan W. Suchow, Zhenyu Cui, Zhaozhuo Xu, K. P. Subbalakshmi, Guojun Xiong, Yueru He, Jimin Huang, Qianqian Xie |
NeurIPS | 15 |