VLDB 2026 Research / reviewers in the wild / expert
Yimin Shi 0001
dblp:57/3424-1
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0003-3375-4602ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Question answering and dialogue systems · 67% Language models and text generation · 33% | |
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 50% Data integration and cleaning · 38% Data mining · 12% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
dialogue evaluation |
0.9 | 1 | 2025 | A Framework for Evaluating AI Agents in Open-Ended Conversations via Scripted Simulation · KDD (2) 2025 |
Data integration and cleaning
entity matching |
0.9 | 1 | 2025 | ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries · Proc. VLDB Endow. 2025 |
Recommender systems › user modeling
user representation learning |
0.9 | 1 | 2025 | You Are What You Bought: Generating Customer Personas for E-commerce Applications · SIGIR 2025 |
Data mining › clustering › user clustering
customer segmentation |
0.3 | 1 | 2025 | You Are What You Bought: Generating Customer Personas for E-commerce Applications · SIGIR 2025 |
Recommender systems › e-commerce recommendation
product recommendation |
0.3 | 1 | 2025 | You Are What You Bought: Generating Customer Personas for E-commerce Applications · SIGIR 2025 |
Methods — techniques the papers use, named apart from their topics
submodular optimization · 1.7approximation algorithm · 1.7scripted simulation · 0.9deep learning · 0.9AI agent simulation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Framework for Evaluating AI Agents in Open-Ended Conversations via Scripted SimulationabstractTraditional evaluations of conversational AI have primarily focused on "closed-ended" interactions, where a human user queries the AI system, such as in customer support.However, many advanced real-world applications-such as job interviews, podcast hosting, and legal or healthcare intake discussions-require "open-ended" interactions in which the AI must take initiative by formulating questions to fully understand the human user's story.As AI assumes broader roles that demand greater autonomy, evaluating its performance in open-ended conversations becomes significantly more complex.This paper introduces a novel framework for rigorously assessing AI agents in such open-ended interviews.In this framework, a secondary AI agent (Agent B) simulates a human interviewee by strictly following a structured script that defines topic strengths and weaknesses, along with guidelines for when to hint at or deviate from these topics.Meanwhile, the primary AI agent (Agent A) engages in dynamic questioning to uncover the script's underlying facts and narrative.By comparing the final conversation transcript to the original script, we assess Agent A using metrics such as completeness, consistency, and investigative depth.This approach not only establishes a new benchmark for open-ended conversational skills but also provides insight into how effectively an AI agent can detect and navigate strategic diversions in scripted behavior. Clarice Wang, Yimin Shi 0001, Xiaokui Xiao |
KDD (2) | 2 |
| 2025 | You Are What You Bought: Generating Customer Personas for E-commerce ApplicationsabstractIn e-commerce, user representations are essential for various applications. Existing methods often use deep learning techniques to convert customer behaviors into implicit embeddings. However, these embeddings are difficult to understand and integrate with external knowledge, limiting the effectiveness of applications such as customer segmentation, search navigation, and product recommendations. To address this, our paper introduces the concept of the customer persona. Condensed from a customer's numerous purchasing histories, a customer persona provides a multi-faceted and human-readable characterization of specific purchase behaviors and preferences, such as Busy Parents or Bargain Hunters. Yimin Shi 0001, Shiqi Zhang 0004, Haixun Wang, Xiaokui Xiao |
SIGIR | 1 |
| 2025 | ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification QueriesabstractRecently, large language models (LLMs) have demonstrated remarkable capabilities in understanding and generating natural language content, attracting widespread attention in both industry and academia. An increasing number of services offer LLMs for various tasks via APIs. Different LLMs demonstrate expertise in different domains of queries (e.g., text classification queries). Meanwhile, LLMs of different scales, complexities, and performance are priced diversely. Driven by this, several researchers are investigating strategies for selecting an ensemble of LLMs, aiming to decrease overall usage costs while enhancing performance. However, to our best knowledge, none of the existing works addresses the problem, how to find an LLM ensemble subject to a cost budget, which maximizes the ensemble performance with guarantees. In this paper, we formalize the performance of an ensemble of models (LLMs) using the notion of correctness probability, which we formally define. We develop an approach for aggregating responses from multiple LLMs to enhance ensemble performance. Building on this, we formulate the Optimal Ensemble Selection (OES) problem of selecting a set of LLMs subject to a cost budget that maximizes the overall correctness probability. We show that the correctness probability function is non-decreasing and non-submodular and provide evidence that the OES problem is likely to be NP-hard. By leveraging a submodular function that upper bounds correctness probability, we develop an algorithm, ThriftLLM, and prove that it achieves an instance-dependent approximation guarantee with high probability. Our framework functions as a data processing system that selects appropriate LLM operators to deliver high-quality results under budget constraints. It achieves state-of-the-art performance for text classification and entity matching queries on multiple real-world datasets against various baselines in our extensive experimental evaluation, while using a relatively lower cost budget, strongly supporting the effectiveness and superiority of our method. Keke Huang, Yimin Shi 0001, Dujian Ding, Yifei Li 0008, Laks V. S. Lakshmanan, Xiaokui Xiao |
Proc. VLDB Endow. | 2 |
| 2022 | An Energy-efficient and Privacy-aware Decomposition Framework for Edge-assisted Federated LearningabstractDeep Learning (DL) is an essential technology for modern intelligent sensor network and interactive multimedia applications, having problems with user data privacy when training on a central cloud. While Federated Learning (FL) motivates to preserve user privacy, it also causes new problems of lower user terminal usability and training efficiency, which caused substantial energy consumption. This article proposes a novel energy-efficient and privacy-aware decomposition framework to improve user-side FL efficiency under pre-defined privacy requirements with the assistance of Mobile Edge Computing (MEC) and Software Decomposition. It takes the propagation of each neural layer as the migrating unit and considers the tradeoff relationship between privacy and efficiency. We also propose an online scheduling algorithm to optimize the framework’s training performance. Furthermore, we summarize eight privacy-sensitive information classes on which existing privacy attacks base and design configurable privacy preservation mechanisms for each class. Simulations and experiments prove the effectiveness of our framework and algorithm in FL efficiency improvement and the effects of different privacy constraints on the overall training efficiency. Yimin Shi 0001, Haihan Duan, Lei Yang 0024, Wei Cai 0002 |
ACM Trans. Sens. Networks | 1 |
| 2020 | Edge-Assisted Federated Learning: An Empirical Study from Software Decomposition Perspective
Yimin Shi 0001, Haihan Duan, Yuanfang Chi, Keke Gai, Wei Cai 0002 |
ICA3PP (2) | 1 |