VLDB 2026 Research / reviewers in the wild / expert
Shilei Ji
dblp:291/4397
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0009-0001-4443-0294ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SageCopilot: An LLM-Empowered Autonomous Agent for Data Science as a ServiceabstractWhile the field of natural language to SQL(NL2SQL) has made significant advancements in translating natural language instructions into executable SQL scripts for data querying and processing, achieving full automation within the broader data science pipeline–encompassing data querying, analysis, visualization, and reporting–remains a complex challenge. This study introduces SageCopilot, an advanced, industry-grade system that automates the data science pipeline by integrating Large Language Models (LLMs), Autonomous Agents (AutoAgents), and Language User Interfaces (LUIs). Designed with a two-phase architecture, SageCopilot uses an offline phase to generate high-quality demonstrations supporting In-Context Learning (ICL), which powers the online phase to transform user inputs into executable scripts for database queries, analysis, and visualization tasks. Leveraging specialized components such as NL2SQL, Text2Analyze, and Text2Viz, as well as chain-of-thought prompting for multi-turn interactions, SageCopilot achieves superior end-to-end automation. Rigorous experimentation with real-world datasets demonstrates the system's ability to minimize human intervention while ensuring correctness and user-friendly operation. Yuan Liao 0003, Jiang Bian 0003, Yuhui Yun, Jiaming Chu, Yuchen Li 0006, Xuhong Li 0002, Shilei Ji, Haoyi Xiong |
IEEE Trans. Serv. Comput. | 10 |
| 2023 | Distributed and deep vertical federated learning with big dataabstractSummary In recent years, data are typically distributed in multiple organizations while the data security is becoming increasingly important. Federated learning (FL), which enables multiple parties to collaboratively train a model without exchanging the raw data, has attracted more and more attention. Based on the distribution of data, FL can be realized in three scenarios, that is, horizontal, vertical, and hybrid. In this article, we propose to combine distributed machine learning techniques with vertical FL and propose a distributed vertical federated learning (DVFL) approach. The DVFL approach exploits a fully distributed architecture within each party in order to accelerate the training process. In addition, we exploit homomorphic encryption to protect the data against honest‐but‐curious participants. We conduct extensive experimentation in a large‐scale cluster environment and a cloud environment in order to show the efficiency and scalability of our proposed approach. The experiments demonstrate the good scalability of our approach and the significant efficiency advantage (up to 6.8 times with a single server and 15.1 times with multiple servers in terms of the training time) compared with baseline frameworks. Ji Liu 0003, Xuehai Zhou, Lei Mo, Shilei Ji, Yuan Liao 0003, Qin Gu, Dejing Dou |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | Cross-model consensus of explanations and beyond for image classification models: an empirical study
Xuhong Li 0002, Haoyi Xiong, Siyu Huang, Shilei Ji, Dejing Dou |
Mach. Learn. | 4 |
| 2023 | Data Placement for Multi-Tenant Data Federation on the CloudabstractDue to privacy concerns of users and law enforcement in data security and privacy, it becomes more and more difficult to share data among organizations. Data federation brings new opportunities to the data-related cooperation among organizations by providing abstract data interfaces. With the development of cloud computing, organizations store data on the cloud to achieve elasticity and scalability for data processing. The existing data placement approaches generally only consider one aspect, which is either execution time or monetary cost, and do not consider data partitioning for hard constraints. In this paper, we propose an approach to enable data processing on the cloud with the data from different organizations. The approach consists of a data federation platform namedFedCubeand a Lyapunov-based data placement algorithm.FedCubeenables data processing on the cloud. We use the data placement algorithm to create a plan in order to partition and store data on the cloud so as to achieve multiple objectives while satisfying the constraints based on a multi-objective cost model. The cost model is composed of two objectives, i.e., reducing monetary cost and execution time. We present an experimental evaluation to show our proposed algorithm significantly reduces the total cost (up to 69.8%) compared with existing approaches. Ji Liu 0003, Lei Mo, Jingbo Zhou 0003, Shilei Ji, Haoyi Xiong, Dejing Dou |
IEEE Trans. Cloud Comput. | 5 |
| 2023 | $\mathcal {AFCS}:$AFCS: Aggregation-Free Spatial-Temporal Mobile Community SensingabstractWhile spatial-temporal environment monitoring has become an indispensable way to collect data for enabling smart cities and intelligent transportation applications, the cost to deploy, operate and maintain a sensor network with sensors and massive communication infrastructure is too high to bear. Compared to the infrastructure-based sensing approach, community sensing, or namely mobile crowdsensing, that leverage community members' mobile devices to collect data becomes a feasible way to scale up the spatial-temporal coverage of the sensing system. However, a community sensing system would need to aggregate sensors and location data from community members and thus would raise concerns on privacy and data security In this paper, we present a novel community sensing paradigm AFCS -Sensor and Location Data Aggregation-Free Community Sensing, which is designed to obtain the environment information (e.g., spatial-temporal distributions of air pollution, temperature, and bike-shares) in each subarea of the target area, without aggregating sensor and location data collected by community members. AFCS proposes to orchestrate with the Trusted Execution Environments (TEEs) of every community member's mobile device to cover the communication, computation and storage with spatial-temporal data. Further, AFCS proposes a novel Decentralized Spatial-Temporal Compressive Sensing framework based on Parallelized Stochastic Gradient Descent. Through learning the latent structure of the spatial-temporal data via decentralized optimization, AFCS approximates the value of the sensor data in each subarea (both covered and uncovered) for each sensing cycle using the sensor data locally stored in every member's TEE instance. Experiments based on real-world datasets and the Virtual Mobile Infrastructure (VMI) with TEE emulations demonstrate that AFCS exhibits low approximation error (i.e., less than 0:2°C in city-wide temperature sensing, 10 units of PM2.5 index in urban air pollution sensing, and 2 bikes in city-wide bike sharing prediction) and performs comparably to (sometimes better than) state-of-the-art algorithms based on the data aggregation and centralized computation Jiang Bian 0003, Haoyi Xiong, Zhiyuan Wang 0003, Jingbo Zhou 0003, Shilei Ji, Hongyang Chen 0001, Daqing Zhang 0001, Dejing Dou |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Feynman: Federated Learning-Based Advertising for Ecosystems-Oriented Mobile Apps RecommendationabstractWhile recommender systems have been ubiquitously used in digital marketing and online business development, the conversions of online advertising for mobile apps installation and activation sometimes are far from satisfactory, due to the lack of feedback from App-related activities, leading to a poor record of Return on Investment (RoI). Though the advertisers, e.g., App operators and App Store, are granted to log users’ app-related activities such as installation, activation, usages, and preferences per the agreement, they usually limit the access to such data from advertisement publishers, due to the privacy concerns. To improve conversions of online advertising under privacy controls, we proposeFeynman—afederated learning-based advertising platform for ecosystems-orientedmobileapps recommendation.Feynmanaims at improving the RoI of mobile app recommendation from an ecosystem's perspective, i.e., per investment in advertising an app (Goal. 1) increasing the number of new installs/users of the app, and then (Goal. 2) increasing the number of new active users (preferably with frequent in-app purchase activities). Incorporating with a federated computing platform,Feynmanleverages users’ records stored in advertisers to refine the pool of targeting users for ads distribution, and jointly builds the predictive models for users’ purchase activities forecasting using features from the Ads publisher and the advertiser. With refined target pools and more accurate models,Feynmanhas successfully helped several mobile apps in China by attracting more than 100 million users to further enlarge their user populations and revenues from in-app purchases. Note that rather than proposing new techniques for federated learning, the design ofFeynmandedicates to show its promising performance in the industrial practices of advertising using federated computing and privacy protected strategies. In three cases that we report in this paper,Feynmanoutperforms the state-of-the-art plans in terms of several key measurements, including Click-Through Rates (CTR), Conversion Rate (CVR), Cost per Action (CPA), and Non-targeting User Hit-Rates (NTHR). Jiang Bian 0003, Jizhou Huang, Shilei Ji, Yuan Liao 0003, Xuhong Li 0002, Qingzhong Wang, Jingbo Zhou 0003, Dejing Dou, Yaqing Wang 0002, Haoyi Xiong |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | From distributed machine learning to federated learning: a survey
Ji Liu 0003, Jizhou Huang, Yang Zhou 0001, Xuhong Li 0002, Shilei Ji, Haoyi Xiong, Dejing Dou |
Knowl. Inf. Syst. | 5 |