Liu Yang 0008

dblp:354/0279 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-4393-1791ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MFS: An Efficient Model Family Serving System for LLMs
abstract
LLM serving providers typically offer a suite of structurally similar models, known as model families, such as the open-source Llama2 series featuring 7B, 13B, and 70B models. While numerous optimizations for LLM serving have been proposed, the potential for leveraging synergies between models within the same family has not been thoroughly explored. This paper introduces MFS, an innovative multi-tiered LLM model family serving system to exploit the structural similarities and parameter redundancies across different scales of models within a family. By utilizing a novel fine-tuning technique called Knowledge Precipitation, MFS restructures the largest model in a family to encapsulate smaller models within its architecture, enabling a unified multi-tiered serving pipeline. Based on the multi-tiered model, MFS realizes a highly parallelized tiered-level batching approach, significantly enhancing system efficiency. It also enables the sharing of intermediate features and KV-cache between models and facilitates multi-level sampling techniques during the inference phase. Experimental results demonstrate that MFS achieves substantial improvements over existing methods, including a 56.1% reduction in end-to-end token generation latency and a 47.8% decrease in GPU memory footprint without compromising the quality of generated content.
Yunxuan Zhang, Hao Wang 0116, Han Tian, Liu Yang 0008, Xudong Liao, Wenxue Li 0004, Ping Yin, Bowen Liu 0002, Kai Chen 0005
EuroSys4
2024 Efficient Decentralized Federated Singular Vector Decomposition
Di Chai, Junxue Zhang 0001, Liu Yang 0008, Yilun Jin, Leye Wang, Kai Chen 0005, Qiang Yang 0001
USENIX ATC3
2024 Federated Meta Embedding Concept Stock Recommendation
abstract
Mining relevant stocks given a trending topic/concept in capital markets is an application with significant economic and societal impacts. Previous concept stock recommendation system mines concept stocks only from public social media like financial news. On stock forums, investors discuss emerging concepts and stocks by using forum comments, which are unneglectable resources to capture trending concept stocks accurately and timely. However, the comment data from a single forum is insufficient to build a high-quality recommendation system. The forums are data silos protected by privacy regulations, and their comments are still underutilized. In this paper, we propose a federated concept stock recommendation baseline and an optimized method that both leverage the private forum comments and public social media without compromising privacy regulations. Our baseline,i.e., Federated Meta Embedding (FedME), is built upon the federated learning framework and learns a concept-stock embedding jointly from private and public data. Our optimized method, Federated Graph Meta Embedding (FedGME), improves FedME by using a graph to combine two sources of embeddings and additional human experts' concept-stock knowledge. Empirically, the experiments on two concept stock datasets show that FedME and FedGME substantially improve the performance of recommendation. Our methods provide practical guidance on privacy-preserving FinTech applications.
Zhuoyi Peng, Yi Yang 0042, Liu Yang 0008, Kai Chen 0005
IEEE Trans. Big Data3
2024 A Survey for Federated Learning Evaluations: Goals and Measures
abstract
Evaluation is a systematic approach to assessing how well a system achieves its intended purpose. Federated learning (FL) is a novel paradigm for privacy-preserving machine learning that allows multiple parties to collaboratively train models without sharing sensitive data. However, evaluating FL is challenging due to its interdisciplinary nature and diverse goals, such as utility, efficiency, and security. In this survey, we first review the major evaluation goals adopted in the existing studies and then explore the evaluation metrics used for each goal. We also introduceFedEval, an open-source platform that provides a standardized and comprehensive evaluation framework for FL algorithms in terms of their utility, efficiency, and security. Finally, we discuss several challenges and future research directions for FL evaluation.
Di Chai, Leye Wang, Liu Yang 0008, Junxue Zhang 0001, Kai Chen 0005, Qiang Yang 0001
IEEE Trans. Knowl. Data Eng.3
2024 High-Performance Hardware Acceleration Architecture for Cross-Silo Federated Learning
abstract
Cross-silo federated learning (FL) adopts various cryptographic operations to preserve data privacy, which introduces significant performance overhead. In this paper, we identify nine widely-used cryptographic operations and design an efficient hardware architecture to accelerate them. However, directly offloading them on hardware statically leads to (1) inadequate hardware acceleration due to the limited resources allocated to each operation; (2) insufficient resource utilization, since different operations are used at different times. To address these challenges, we propose FLASH, a high-performance hardware acceleration architecture for cross-silo FL systems. At its heart, FLASH extracts two basic operators—modular exponentiation and multiplication—behind the nine cryptographic operations and implements them as highly-performant engines to achieve adequate acceleration. Furthermore, it leverages a dataflow scheduling scheme to dynamically compose different cryptographic operations based on these basic engines to obtain sufficient resource utilization. We have implemented a fully-functional FLASH prototype with Xilinx VU13P FPGA and integrated it with FATE, the most widely-adopted cross-silo FL framework. Experimental results show that, for the nine cryptographic operations, FLASH achieves up to$14.0\times$and$3.4\times$acceleration over CPU and GPU, translating to up to$6.8\times$and$2.0\times$speedup for realistic FL applications, respectively. We finally evaluate the FLASH design as an ASIC, and it achieves$23.6\times$performance improvement upon the FPGA prototype.
Junxue Zhang 0001, Xiaodian Cheng, Liu Yang 0008, Jinbin Hu 0001, Han Tian, Kai Chen 0005
IEEE Trans. Parallel Distributed Syst.3
2023 Globally Consistent Federated Graph Autoencoder for Non-IID Graphs
abstract
Graph neural networks (GNNs) have been applied successfully in many machine learning tasks due to their advantages in utilizing neighboring information. Recently, with the global enactment of privacy protection regulations, federated GNNs have gained increasing attention in academia and industry. However, the graphs owned by different participants could be non-independently-and-identically distributed (non-IID), leading to the deterioration of federated GNNs' accuracy. In this paper, we propose a globally consistent federated graph autoencoder (GCFGAE) to overcome the non-IID problem in unsupervised federated graph learning via three innovations. First, by integrating federated learning with split learning, we train a unique global model instead of FedAvg-styled global and local models, yielding results consistent with that of the centralized GAE. Second, we design a collaborative computation mechanism considering overlapping vertices to reduce communication overhead during forward propagation. Third, we develop a layer-wise and block-wise gradient computation strategy to reduce the space and communication complexity during backward propagation. Experiments on real-world datasets demonstrate that GCFGAE achieves not only higher accuracy but also around 500 times lower communication overhead and 1000 times smaller space overhead than existing federated GNN models.
Kun Guo 0003, Yutong Fang, Wenyu He, Liu Yang 0008, Kai Chen 0005, Ximeng Liu, Wenzhong Guo
IJCAI7
2023 FLASH: Towards a High-performance Hardware Acceleration Architecture for Cross-silo Federated Learning
Junxue Zhang 0001, Xiaodian Cheng, Liu Yang 0008, Jinbin Hu 0001, Kai Chen 0005
NSDI4
2022 Addressing Network Bottlenecks with Divide-and-Shuffle Synchronization for Distributed DNN Training
abstract
Bulk synchronous parallel (BSP) is the de-facto paradigm for distributed DNN training in today’s production clusters. However, due to the global synchronization nature, its performance can be significantly influenced by network bottlenecks caused by either static topology heterogeneity or dynamic bandwidth contentions. Existing solutions, either system-level optimizations strengthening BSP (e.g., Ring or Hierarchical All-reduce) or algorithmic optimizations replacing BSP (e.g., ASP or SSP, which relax the global barriers), do not completely solve the problem, as they may still suffer from communication inefficiency or risk convergence inaccuracy.In this paper, we present a novel divide-and-shuffle synchronization (DS-Sync) to realize communication efficiency without sacrificing convergence accuracy for distributed DNN training. At its heart, by taking into account the network bottlenecks, DS-Sync improves communication efficiency by dividing workers into non-overlap groups to synchronize independently in a bottleneck-free manner. Meanwhile, it maintains convergence accuracy by iteratively shuffling workers among different groups to ensure a global consensus. We theoretically prove that DS-Sync converges properly in non-convex and smooth conditions like DNN. We further implement DS-Sync and integrate it with PyTorch, and our testbed experiments show that DS-Sync can achieve up to 94% improvements on the end-to-end training time with existing solutions while maintaining the same accuracy.
Weiyan Wang, Cengguang Zhang, Liu Yang 0008, Kai Chen 0005, Kun Tan 0002
INFOCOM3
2022 Practical Lossless Federated Singular Vector Decomposition over Billion-Scale Data
abstract
With the enactment of privacy-preserving regulations, e.g., GDPR, federated SVD is proposed to enable SVD-based applications over different data sources without revealing the original data. However, many SVD-based applications cannot be well supported by existing federated SVD solutions. The crux is that these solutions, adopting either differential privacy (DP) or homomorphic encryption (HE), suffer from accuracy loss caused by unremovable noise or degraded efficiency due to inflated data.
Di Chai, Leye Wang, Junxue Zhang 0001, Liu Yang 0008, Shuowei Cai, Kai Chen 0005, Qiang Yang 0001
KDD4
2022 Improving Availability of Vertical Federated Learning: Relaxing Inference on Non-overlapping Data
abstract
Vertical Federated Learning (VFL) enables multiple parties to collaboratively train a machine learning model over vertically distributed datasets without data privacy leakage. However, there is a limitation of the current VFL solutions: current VFL models fail to conduct inference on non-overlapping samples during inference. This limitation seriously damages the VFL model’s availability because, in practice, overlapping samples may only take up a small portion of the whole data at each party which means a large part of inference tasks will fail. In this article, we propose a novel VFL framework which enables federated inference on non-overlapping data. Our framework regards the distributed features as privileged information which is available in the training period but disappears during inference. We distill the knowledge of such privileged features and transfer them to the parties’ local model which only processes local features. Furthermore, we adopt Oblivious Transfer (OT) to preserve data ID privacy during training and inference. Empirically, we evaluate the model on the real-world dataset collected from Criteo and Taobao. Besides, we also provide a security analysis of the proposed framework.
Zhenghang Ren, Liu Yang 0008, Kai Chen 0005
ACM Trans. Intell. Syst. Technol.2
2020 Exploring Clustering of Bandits for Online Recommendation System
abstract
Cluster-of-bandit policy leverages contextual bandits in a collaborative filtering manner and aids personalized services in the online recommendation system (RecSys). When facing insufficient observations, the cluster-of-bandit policy could achieve more outstanding performance because of knowledge sharing. Cluster-of-bandit policy aims to maximize the cumulative feedback, e.g., clicks, from users. Nevertheless, in the way of their goal exist two kinds of uncertainties. First, cluster-of-bandit algorithms make recommendations according to their uncertain estimation of user interests. Second, cluster-of-bandit algorithms transfer relevant knowledge upon uncertain and noisy user clusters. Existing algorithms only consider the first one, while leaving the latter one untouched. To address the two challenges together, in this paper, we propose the ClexB policy for online RecSys. On the one hand, ClexB estimates user clustering more accurately and with less uncertainty via explorable-clustering. On the other hand, ClexB also exploits and explores user interests by sharing information within and among user clusters. In summary, ClexB explores knowledge transfer and further aids the inferences about user interests. Besides, we provide extensive empirical experiments on both the synthetic and real-world datasets and regret analysis, further consolidating the superiority of ClexB.
Liu Yang 0008, Bo Liu 0015, Leyu Lin, Feng Xia 0006, Kai Chen 0005, Qiang Yang 0001
RecSys1