VLDB 2026 Research / reviewers in the wild / expert
Yuan Liao 0003
dblp:30/3498-3
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0008-4826-7886ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SageCopilot: An LLM-Empowered Autonomous Agent for Data Science as a ServiceabstractWhile the field of natural language to SQL(NL2SQL) has made significant advancements in translating natural language instructions into executable SQL scripts for data querying and processing, achieving full automation within the broader data science pipeline–encompassing data querying, analysis, visualization, and reporting–remains a complex challenge. This study introduces SageCopilot, an advanced, industry-grade system that automates the data science pipeline by integrating Large Language Models (LLMs), Autonomous Agents (AutoAgents), and Language User Interfaces (LUIs). Designed with a two-phase architecture, SageCopilot uses an offline phase to generate high-quality demonstrations supporting In-Context Learning (ICL), which powers the online phase to transform user inputs into executable scripts for database queries, analysis, and visualization tasks. Leveraging specialized components such as NL2SQL, Text2Analyze, and Text2Viz, as well as chain-of-thought prompting for multi-turn interactions, SageCopilot achieves superior end-to-end automation. Rigorous experimentation with real-world datasets demonstrates the system's ability to minimize human intervention while ensuring correctness and user-friendly operation. Yuan Liao 0003, Jiang Bian 0003, Yuhui Yun, Jiaming Chu, Yuchen Li 0006, Xuhong Li 0002, Shilei Ji, Haoyi Xiong |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Distributed and deep vertical federated learning with big dataabstractSummary In recent years, data are typically distributed in multiple organizations while the data security is becoming increasingly important. Federated learning (FL), which enables multiple parties to collaboratively train a model without exchanging the raw data, has attracted more and more attention. Based on the distribution of data, FL can be realized in three scenarios, that is, horizontal, vertical, and hybrid. In this article, we propose to combine distributed machine learning techniques with vertical FL and propose a distributed vertical federated learning (DVFL) approach. The DVFL approach exploits a fully distributed architecture within each party in order to accelerate the training process. In addition, we exploit homomorphic encryption to protect the data against honest‐but‐curious participants. We conduct extensive experimentation in a large‐scale cluster environment and a cloud environment in order to show the efficiency and scalability of our proposed approach. The experiments demonstrate the good scalability of our approach and the significant efficiency advantage (up to 6.8 times with a single server and 15.1 times with multiple servers in terms of the training time) compared with baseline frameworks. Ji Liu 0003, Xuehai Zhou, Lei Mo, Shilei Ji, Yuan Liao 0003, Qin Gu, Dejing Dou |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Feynman: Federated Learning-Based Advertising for Ecosystems-Oriented Mobile Apps RecommendationabstractWhile recommender systems have been ubiquitously used in digital marketing and online business development, the conversions of online advertising for mobile apps installation and activation sometimes are far from satisfactory, due to the lack of feedback from App-related activities, leading to a poor record of Return on Investment (RoI). Though the advertisers, e.g., App operators and App Store, are granted to log users’ app-related activities such as installation, activation, usages, and preferences per the agreement, they usually limit the access to such data from advertisement publishers, due to the privacy concerns. To improve conversions of online advertising under privacy controls, we proposeFeynman—afederated learning-based advertising platform for ecosystems-orientedmobileapps recommendation.Feynmanaims at improving the RoI of mobile app recommendation from an ecosystem's perspective, i.e., per investment in advertising an app (Goal. 1) increasing the number of new installs/users of the app, and then (Goal. 2) increasing the number of new active users (preferably with frequent in-app purchase activities). Incorporating with a federated computing platform,Feynmanleverages users’ records stored in advertisers to refine the pool of targeting users for ads distribution, and jointly builds the predictive models for users’ purchase activities forecasting using features from the Ads publisher and the advertiser. With refined target pools and more accurate models,Feynmanhas successfully helped several mobile apps in China by attracting more than 100 million users to further enlarge their user populations and revenues from in-app purchases. Note that rather than proposing new techniques for federated learning, the design ofFeynmandedicates to show its promising performance in the industrial practices of advertising using federated computing and privacy protected strategies. In three cases that we report in this paper,Feynmanoutperforms the state-of-the-art plans in terms of several key measurements, including Click-Through Rates (CTR), Conversion Rate (CVR), Cost per Action (CPA), and Non-targeting User Hit-Rates (NTHR). Jiang Bian 0003, Jizhou Huang, Shilei Ji, Yuan Liao 0003, Xuhong Li 0002, Qingzhong Wang, Jingbo Zhou 0003, Dejing Dou, Yaqing Wang 0002, Haoyi Xiong |
IEEE Trans. Serv. Comput. | 4 |