Yimin Shi 0001

dblp:57/3424-1 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0003-3375-4602ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Question answering and dialogue systems · 67% Language models and text generation · 33%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 50% Data integration and cleaning · 38% Data mining · 12%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
dialogue evaluation
0.912025
A Framework for Evaluating AI Agents in Open-Ended Conversations via Scripted Simulation · KDD (2) 2025
Data integration and cleaning
entity matching
0.912025
ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries · Proc. VLDB Endow. 2025
Recommender systems › user modeling
user representation learning
0.912025
You Are What You Bought: Generating Customer Personas for E-commerce Applications · SIGIR 2025
Data mining › clustering › user clustering
customer segmentation
0.312025
You Are What You Bought: Generating Customer Personas for E-commerce Applications · SIGIR 2025
Recommender systems › e-commerce recommendation
product recommendation
0.312025
You Are What You Bought: Generating Customer Personas for E-commerce Applications · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

submodular optimization · 1.7approximation algorithm · 1.7scripted simulation · 0.9deep learning · 0.9AI agent simulation · 0.9
YearPublicationVenuePosition
2025 A Framework for Evaluating AI Agents in Open-Ended Conversations via Scripted Simulation
abstract
Traditional evaluations of conversational AI have primarily focused on "closed-ended" interactions, where a human user queries the AI system, such as in customer support.However, many advanced real-world applications-such as job interviews, podcast hosting, and legal or healthcare intake discussions-require "open-ended" interactions in which the AI must take initiative by formulating questions to fully understand the human user's story.As AI assumes broader roles that demand greater autonomy, evaluating its performance in open-ended conversations becomes significantly more complex.This paper introduces a novel framework for rigorously assessing AI agents in such open-ended interviews.In this framework, a secondary AI agent (Agent B) simulates a human interviewee by strictly following a structured script that defines topic strengths and weaknesses, along with guidelines for when to hint at or deviate from these topics.Meanwhile, the primary AI agent (Agent A) engages in dynamic questioning to uncover the script's underlying facts and narrative.By comparing the final conversation transcript to the original script, we assess Agent A using metrics such as completeness, consistency, and investigative depth.This approach not only establishes a new benchmark for open-ended conversational skills but also provides insight into how effectively an AI agent can detect and navigate strategic diversions in scripted behavior.
Clarice Wang, Yimin Shi 0001, Xiaokui Xiao
KDD (2)2
2025 You Are What You Bought: Generating Customer Personas for E-commerce Applications
abstract
In e-commerce, user representations are essential for various applications. Existing methods often use deep learning techniques to convert customer behaviors into implicit embeddings. However, these embeddings are difficult to understand and integrate with external knowledge, limiting the effectiveness of applications such as customer segmentation, search navigation, and product recommendations. To address this, our paper introduces the concept of the customer persona. Condensed from a customer's numerous purchasing histories, a customer persona provides a multi-faceted and human-readable characterization of specific purchase behaviors and preferences, such as Busy Parents or Bargain Hunters.
Yimin Shi 0001, Shiqi Zhang 0004, Haixun Wang, Xiaokui Xiao
SIGIR1
2025 ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries
abstract
Recently, large language models (LLMs) have demonstrated remarkable capabilities in understanding and generating natural language content, attracting widespread attention in both industry and academia. An increasing number of services offer LLMs for various tasks via APIs. Different LLMs demonstrate expertise in different domains of queries (e.g., text classification queries). Meanwhile, LLMs of different scales, complexities, and performance are priced diversely. Driven by this, several researchers are investigating strategies for selecting an ensemble of LLMs, aiming to decrease overall usage costs while enhancing performance. However, to our best knowledge, none of the existing works addresses the problem, how to find an LLM ensemble subject to a cost budget, which maximizes the ensemble performance with guarantees. In this paper, we formalize the performance of an ensemble of models (LLMs) using the notion of correctness probability, which we formally define. We develop an approach for aggregating responses from multiple LLMs to enhance ensemble performance. Building on this, we formulate the Optimal Ensemble Selection (OES) problem of selecting a set of LLMs subject to a cost budget that maximizes the overall correctness probability. We show that the correctness probability function is non-decreasing and non-submodular and provide evidence that the OES problem is likely to be NP-hard. By leveraging a submodular function that upper bounds correctness probability, we develop an algorithm, ThriftLLM, and prove that it achieves an instance-dependent approximation guarantee with high probability. Our framework functions as a data processing system that selects appropriate LLM operators to deliver high-quality results under budget constraints. It achieves state-of-the-art performance for text classification and entity matching queries on multiple real-world datasets against various baselines in our extensive experimental evaluation, while using a relatively lower cost budget, strongly supporting the effectiveness and superiority of our method.
Keke Huang, Yimin Shi 0001, Dujian Ding, Yifei Li 0008, Laks V. S. Lakshmanan, Xiaokui Xiao
Proc. VLDB Endow.2
2022 An Energy-efficient and Privacy-aware Decomposition Framework for Edge-assisted Federated Learning
abstract
Deep Learning (DL) is an essential technology for modern intelligent sensor network and interactive multimedia applications, having problems with user data privacy when training on a central cloud. While Federated Learning (FL) motivates to preserve user privacy, it also causes new problems of lower user terminal usability and training efficiency, which caused substantial energy consumption. This article proposes a novel energy-efficient and privacy-aware decomposition framework to improve user-side FL efficiency under pre-defined privacy requirements with the assistance of Mobile Edge Computing (MEC) and Software Decomposition. It takes the propagation of each neural layer as the migrating unit and considers the tradeoff relationship between privacy and efficiency. We also propose an online scheduling algorithm to optimize the framework’s training performance. Furthermore, we summarize eight privacy-sensitive information classes on which existing privacy attacks base and design configurable privacy preservation mechanisms for each class. Simulations and experiments prove the effectiveness of our framework and algorithm in FL efficiency improvement and the effects of different privacy constraints on the overall training efficiency.
Yimin Shi 0001, Haihan Duan, Lei Yang 0024, Wei Cai 0002
ACM Trans. Sens. Networks1
2020 Edge-Assisted Federated Learning: An Empirical Study from Software Decomposition Perspective
Yimin Shi 0001, Haihan Duan, Yuanfang Chi, Keke Gai, Wei Cai 0002
ICA3PP (2)1