Qiyuan Ge

dblp:426/6938 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0008-0569-1855ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › evaluation › online evaluation
a/b testing
1.012026
Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing · KDD (1) 2026
Information retrieval
evaluation
1.012026
Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing · KDD (1) 2026
Human-AI interaction
LLM-based agents
1.012026
Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing · KDD (1) 2026
Information retrieval › evaluation
user simulation
0.312026
Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing · KDD (1) 2026

Methods — techniques the papers use, named apart from their topics

memory retrieval · 2.0large language model · 2.0multimodal perception · 1.0multi-modal perception · 1.0
YearPublicationVenuePosition
2026 Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing
abstract
diningIn recommender systems, online A/B testing is a crucial method for evaluating the performance of different models. However, conducting online A/B testing often presents significant challenges, including substantial economic costs, user experience degradation, and considerable time requirements. With the Large Language Models' powerful capacity, LLM-based agent shows great potential to replace traditional online A/B testing. Nonetheless, current agents fail to simulate the perception process and interaction patterns, due to the lack of real environments and visual perception capability. To address these challenges, we introduce a multi-modal user agent for A/B testing (A/B Agent). Specifically, we construct a recommendation sandbox environment for A/B testing, enabling multimodal and multi-page interactions that align with real user behavior on online platforms. The designed agent leverages multimodal information perception, fine-grained user preferences, and integrates profiles, action memory retrieval, and a fatigue system to simulate complex human decision-making. We validated the potential of the agent as an alternative to traditional A/B testing from three perspectives: model, data, and features. Furthermore, we found that the data generated by A/B Agent can effectively enhance the capabilities of recommendation models. Our code is publicly available at https://github.com/Applied-Machine-Learning-Lab/ABAgent. © 2026 Owner/Author.
Wenlin Zhang 0001, Xiangyang Li 0004, Qiyuan Ge, Kuicai Dong, Pengyue Jia, Xiaopeng Li 0014, Zijian Zhang 0009, Maolin Wang 0001, Yichao Wang 0002, Huifeng Guo, Ruiming Tang, Xiangyu Zhao 0001
KDD (1)3