EDBT 2026 Demo / reviewers in the wild / expert
Qiyuan Ge
dblp:426/6938
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0008-0569-1855ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › evaluation › online evaluation
a/b testing |
1.0 | 1 | 2026 | Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing · KDD (1) 2026 |
Information retrieval
evaluation |
1.0 | 1 | 2026 | Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing · KDD (1) 2026 |
Human-AI interaction
LLM-based agents |
1.0 | 1 | 2026 | Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing · KDD (1) 2026 |
Information retrieval › evaluation
user simulation |
0.3 | 1 | 2026 | Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing · KDD (1) 2026 |
Methods — techniques the papers use, named apart from their topics
memory retrieval · 2.0large language model · 2.0multimodal perception · 1.0multi-modal perception · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B TestingabstractdiningIn recommender systems, online A/B testing is a crucial method for evaluating the performance of different models. However, conducting online A/B testing often presents significant challenges, including substantial economic costs, user experience degradation, and considerable time requirements. With the Large Language Models' powerful capacity, LLM-based agent shows great potential to replace traditional online A/B testing. Nonetheless, current agents fail to simulate the perception process and interaction patterns, due to the lack of real environments and visual perception capability. To address these challenges, we introduce a multi-modal user agent for A/B testing (A/B Agent). Specifically, we construct a recommendation sandbox environment for A/B testing, enabling multimodal and multi-page interactions that align with real user behavior on online platforms. The designed agent leverages multimodal information perception, fine-grained user preferences, and integrates profiles, action memory retrieval, and a fatigue system to simulate complex human decision-making. We validated the potential of the agent as an alternative to traditional A/B testing from three perspectives: model, data, and features. Furthermore, we found that the data generated by A/B Agent can effectively enhance the capabilities of recommendation models. Our code is publicly available at https://github.com/Applied-Machine-Learning-Lab/ABAgent. © 2026 Owner/Author. Wenlin Zhang 0001, Xiangyang Li 0004, Qiyuan Ge, Kuicai Dong, Pengyue Jia, Xiaopeng Li 0014, Zijian Zhang 0009, Maolin Wang 0001, Yichao Wang 0002, Huifeng Guo, Ruiming Tang, Xiangyu Zhao 0001 |
KDD (1) | 3 |