VLDB 2026 Research / reviewers in the wild / expert
Xingjian Lu
dblp:120/4838
· DBLP profile ↗
26ranked-venue papers
6as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | POP-FL: Towards Efficient Federated Learning on Edge Using Parallel Over-ParameterizationabstractFederated Learning (FL) is a promising paradigm for mining massive data while respecting users' privacy. However, the deployment of FL on resource-constrained edge devices remains elusive due to its high resource demand. In this paper, unlike existing works that use expensive dense models, we propose to utilize dynamic sparse training in FL and design a novel sparse-to-sparse FL framework, named as POP-FL. The framework can reduce both computation and communication overheads while maintaining the performance of the global model. Specifically, POP-FL partitions massive clients into groups and performs parallel parameter exploration, i.e.,Parallel Over-Parameterization, over the collaboration between these groups. This exploration can greatly improve the expressibility and generalizability of sparse training in FL (especially for extreme sparsity levels) through reliably covering sufficient parameters and dynamically updating the global sparse network's structure during the training process. Experimental results show that compared with existing sparse-to-sparse training methods in both iid and non-iid data distribution, POP-FL achieves the best inference accuracy on various representative networks. Xingjian Lu, Haikun Zheng, Wenyan Liu 0001, Yuhui Jiang, Hongyue Wu |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | Joint Client Selection and Bandwidth Allocation of Wireless Federated Learning by Deep Reinforcement LearningabstractFederated Learning (FL) is a promising paradigm for massive data mining service while protecting users’ privacy. In wireless federated learning networks (WFLNs), limited communication resources and heterogeneity of user devices have essential impacts on training efficiency of FL, hence it is critical to select clients and allocate network bandwidths among them in each learning round to improve the training efficiency. In this article, we formulate the joint client selection and bandwidth allocation optimization problem as a MDP process and design a FL framework CSBWA to solve it. CSBWA relies on DRL-based REINFORCE algorithm to automatically perform effective policy based on observed information, e.g., client states, historical bandwidths, and feedback rewards. It is able to achieve lower time cost and energy consumption with long-term FL performance guarantee by jointly optimizing the client selection and bandwidth allocation. Experimental results show the effectiveness of CSBWA in reducing time cost and energy consumption while guaranteeing model performance of wireless federated learning compared with existing state-of-art methods. Xingjian Lu, Yuhui Jiang, Haikun Zheng |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | Integrating Staleness and Shapley Value Consistency for Efficient K-Asynchronous Federated LearningabstractIn the big data era, Federated Learning (FL), which allows multiple participants to collaboratively train a global model without sharing their raw data, emerges as a promising solution to address the challenges of isolated data silos and privacy protection. Federated learning has two main communication strategies: synchronous and asynchronous. Synchronous FL ensures stable convergence but may encounter model quality degradation and server crash risks. Asynchronous FL avoids the straggler effect and supports more participants, but unstable convergence and non-IID data could affect the model performance. In this paper, inspired by real-world FL scenarios, we propose a highly efficient K-Asynchronous FL framework, KFLBSV, which addresses the limitations of synchronous and asynchronous strategies to some extent, leading to improved model performance and convergence speed. The framework allows clients to upload updates multiple times within the same round instead of blocking after each upload, thereby enhancing training efficiency. To ensure the stability and performance of the global model, we introduce a novel aggregation method. By approximating Shapley value to assess model consistency and balancing client contribution frequency and model staleness, we allocate weights more accurately to each participating client. We extensively conducted experiments on benchmark datasets using three distinct models, and the results show that KFLBSV outperforms existing algorithms in terms of both model performance and convergence speed. Yuhui Jiang, Xingjian Lu, Ying Li 0001 |
IEEE Big Data | 2 |
| 2023 | ASPFL: Efficient Personalized Federated Learning for Edge Based on Adaptive Sparse TrainingabstractOne of the primary challenges in cloud-edge environments is efficiently utilizing significant amounts of data on edge devices for machine learning tasks, enabling adaptation to increasingly complex computing and service scenarios. Federated Learning (FL) is a machine learning paradigm that enables collaborative training of models involving multiple data warehouses in a privacy-preserving manner. However, classical federated learning has poor convergence on highly heterogeneous data, which limits its performance of global model on each edge device. The emergence of Personalized Federated Learning (PFL) effectively alleviates data heterogeneity, but learning a personalized model may incur greater overheads. In this paper, we propose an efficient FL framework named as ASPFL, which uses dynamic sparse training for personalized federated learning to maintain model performance while reducing computational and communication overheads in cloud-edge environments. By adaptively allocating the dynamic sparsity from a global perspective to explore sparse network structure during training, ASPFL improves the independent parameter exploration process of local sparse training to adapt to various heterogeneous situations and solves the Non-IID challenge of FL. The abundant experimental results show that ASPFL outperforms state-of-the-art methods in performance, overheads, and convergence speed in PFL. Yuhui Jiang, Xingjian Lu, Haikun Zheng |
ICWS | 2 |
| 2023 | ACORN: Adaptive Compression-Reconstruction for Video Services in 5G-U Industrial IoTabstractIoT devices are enabled to capture and upload videos with increasing bitrates. Massive IIoT is eager for effective video processing techniques to satisfy the requirements of real-time video services. With the emergence of 5G-unlicensed (5G-U), ultra-low latency video applications become possible. However, existing encoding standards for video services in Web 2.0, such as H.265, are not naturally designed for IIoT video streaming, leading to bandwidth pressure where 5G-U coexists with various other wireless signals. To tackle this problem and to support low-latency video utilization by IIoT video sources, we propose an Adaptive Compression-Reconstruction framework named ACORN, which is based on compressed sensing and recent advances in deep learning. At end nodes, we compress multiple sequential video frames into a single frame to reduce video volume. We design a QoE-aware parameter selection mechanism to deal with volatile network environments during compression. With learnable gated convolution layers and channel-wise soft-thresholding operators, ACORN also builds a real-time reconstruction module. Experimental results reveal that video analytics can be conducted on compressed frames. The reconstruction algorithm in ACORN is with $1-4 \mathrm{~dB}$ improvements. Moreover, both the encoding time cost and the encoded video volume are reduced by more than $4 \times$ under the ACORN framework. Jiale Lei, Peihao Yang, Linghe Kong, Yehan Ma, Xingjian Lu, Deyu Lin, Guihai Chen, E. Zhao |
MSN | 5 |
| 2022 | Hybrid differential privacy based federated learning for Internet of Things
Wenyan Liu 0001, Junhong Cheng, Xiaoling Wang 0004, Xingjian Lu, Jianwei Yin |
J. Syst. Archit. | 4 |
| 2021 | Large-scale Fake Click Detection for E-commerce Recommendation SystemsabstractWith the development of e-commerce platforms, e-commerce recommendation systems are playing an increasingly important role for the purpose of product recommendation. As a new attack model against e-commerce recommendation systems, the "Ride Item's Coattails" attack creates fake click information to establish the deceptive correlation between popular products and low-quality products in order to mislead the recommendation system of e-commerce platform to boost the sales of low-quality products. This attack is characterized by high concealment and strong destructiveness, which can cause great damage to e-commerce recommendation systems, and adversely affect the usability of the e-commerce platform and users' shopping experience. It is therefore of great practical significance to study how to quickly and effectively identify the false click information and the corresponding "Ride Item's Coattails" attack to better safeguard e-commerce recommendation systems. At present, there is no previously reported relevant research work conducted specifically for addressing the detection of the "Ride Item's Coattails" attack. In this work, we carried out pioneering work in analyzing and summarizing the characteristics of the false click information produced by attackers on the target products in the "Ride Item's Coattails" attack and designed a set of attack detection techniques suitable for e-commerce recommendation systems. Experimental results on real e-commerce datasets show that our proposed techniques can quickly and effectively detect the large-scale fake click information as well as the associated "Ride Item's Coattails" attack in e-commerce recommendation systems. Jingdong Li, Zhao Li 0007, Ji Zhang 0001, Xiaoling Wang 0004, Xingjian Lu, Jingren Zhou 0001 |
ICDE | 6 |
| 2021 | Incorporating Network Structure with Node Information for Semi-supervised Anomaly Detection on Attributed Graphs
Bofeng Chen, Jingdong Li, Xingjian Lu, Chaofeng Sha |
WISE (1) | 3 |
| 2020 | WFApprox: Approximate Window Functions Processing
Chunbo Lin, Jingdong Li, Xingjian Lu |
DASFAA (1) | 4 |
| 2020 | POEM: Position Order Enhanced Model for Session-based Recommendation ServiceabstractSession-based recommendation, which aims to predict the next action of an anonymous user base on the interaction information in a session, plays a crucial role in many online services. Recent works solve the problem with the latest deep learning techniques and have achieved good performance on some datasets. However, they have some shortcomings that affect their practical application value: a) the drift process of users' interests in the browsing is not well explored; b) the association between a user's current interests and general preferences in the session is not adequately considered. They mostly assume that the last interaction has a significant impact on the next interaction, which makes them work well only in limited scenarios and specific datasets. To address these limitations, we propose a session-based recommendation model called POEM, which explicitly considers the impact of interaction order relationships on recommendations by emphasizing position attributes in the session. Specifically, POEM models the macro and micro importance of each item in the session, the influence of user interaction order on the item-level collaboration, and the session-level collaboration reflected in the user interest drift process, respectively. Extensive experiments of the effectiveness, efficiency, and universality on three real-world datasets show that our method outperforms various state-of-the-art session-based recommendation methods consistently. Mingyou Sun, Jiahao Yuan 0002, Zihan Song 0001, Xingjian Lu, Xiaoling Wang 0004 |
ICWS | 5 |
| 2020 | AQapprox: Aggregation Queries Approximation with Distribution-Aware Online Sampling
Xingjian Lu |
WISE (2) | 3 |
| 2020 | Bulk Savings for Bulk Transfers: Minimizing the Energy-Cost for Geo-Distributed Data CentersabstractWith the fast proliferation of cloud computing, major cloud service providers, e.g., Amazon, Google, Facebook, etc., have been deploying more and more geographically distributed data centers to provide customers with better reliability and quality of services. A basic demand in such a geo-distributed data center system is to transfer bulk volumes of data from one data center to another. Geographic distribution and large delay-tolerance of such inter-data-center bulk data transfers provide cloud service providers opportunities to optimize the operating cost. Most existing studies on inter-data-center bulk data transfers focus on minimizing the network bandwidth cost. However, the energy-cost of the bulk data transfers, which also accounts for a large proportion of operating cost in the data centers, still remains unexplored. This is an important problem, especially in the multi-electricity-market environment, where the electricity price exhibits both spatial and temporal diversities. In this paper, we systematically study the problem of how to route and schedule inter-data-center bulk data transfers to minimize the energy-cost for geo-distributed data centers. We model this problem as a min-cost multi-commodity flow problem and develop an efficient two-stage optimization method to solve it. Extensive evaluations with real-life inter-data-center network and electricity prices show that our method brings significant energy-cost savings over existing bulk data transfer methods. Xingjian Lu, Fanxin Kong, Xue (Steve) Liu, Jianwei Yin, Qiao Xiang, Huiqun Yu |
IEEE Trans. Cloud Comput. | 1 |
| 2019 | ADMM-Based Decentralized Electric Vehicle Charging with Trip Duration LimitsabstractWith the large-scale deployment of Electric Vehicles (EVs), the unbalanced distribution of charging needs and random charging behaviors cause charging stations (CSs) congestion. This degrades EV drivers' quality of experience by extending charging waiting time and increasing charging fee. Thus, EV owners are facing a critical issue on how to decrease the cost of charging, which consists of two parts: charging duration and charging fee. A great deal of existing work is confined to finding CSs to optimize the two parts individually. However, it still remains unexplored how to jointly minimize charging duration and charging fee under an overall time limit (i.e., deadline) of a scheduled trip. The problem is the focus of this paper. First, we formulate this problem as a 0-1 Integer Linear Programming problem and show its NP-Hardness. Then, we propose an efficient distributed algorithm based on the Alternating Direction Method of Multipliers (ADMM). The algorithm decomposes the original problem into sub-problems that can be solved locally and in parallel between charging stations and the global coordinator. Finally, we carry out extensive simulations based on real-life transport network data, and the results show that the proposed approach brings significant cost savings over existing ones. Gaoqi He, Zhifu Chai, Xingjian Lu, Fanxin Kong, Bin Sheng 0001 |
RTSS | 3 |
| 2019 | Physical-barrier detection based collective motion analysis
Gaoqi He, Dongxu Jiang, Yubo Yuan 0001, Xingjian Lu |
Frontiers Comput. Sci. | 5 |
| 2019 | Distributed Data Center Bandwidth Allocation for Cloud-Based StreamingabstractCloud-based video streaming systems such as YouTube and Netflix are usually supported by the content delivery networks and data centers that can consume many megawatts of power. Most existing work independently studies the issues of improving quality of experience (QoE) for viewers and reducing the cost and emissions associated with the enormous energy usage of data centers. By contrast, this paper addresses them both, and jointly optimizes the QoE, the energy cost and emissions by intelligently allocating data center bandwidth among different client groups. Specially, we propose a distributed algorithm to achieve the optimal bandwidth allocation, given the prediction of future workload. The algorithm novelly decomposes the optimization process into separate ones, which are solved iteratively across data centers and clients. Further, the algorithm has robust performance guarantee in terms of the variance of the prediction error. We demonstrate its convergence and robustness by both proofs using theoretical analysis and validation based on trace-driven simulations. The results further show that the proposed algorithm converges very fast and achieves much better QoE-cost balance than existing approaches. Fanxin Kong, Xingjian Lu, Xue (Steve) Liu |
IEEE Trans. Sustain. Comput. | 2 |
| 2017 | A double-region learning algorithm for counting the number of pedestrians in subway surveillance videos
Gaoqi He, Dongxu Jiang, Xingjian Lu, Yubo Yuan 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2016 | JTangCMS: An efficient monitoring system for cloud platforms
Xingjian Lu, Jianwei Yin, Naixue Xiong, Shuiguang Deng, Gaoqi He, Huiqun Yu |
Inf. Sci. | 1 |
| 2016 | Shadow obstacle model for realistic corner-turning behavior in crowd simulationabstractThis paper describes a novel model known as the shadow obstacle model to generate a realistic corner-turning behavior in crowd simulation. The motivation for this model comes from the observation that people tend to choose a safer route rather than a shorter one when turning a corner. To calculate a safer route, an optimization method is proposed to generate the corner-turning rule that maximizes the viewing range for the agents. By combining psychological and physical forces together, a full crowd simulation framework is established to provide a more realistic crowd simulation. We demonstrate that our model produces a more realistic corner-turning behavior by comparison with real data obtained from the experiments. Finally, we perform parameter analysis to show the believability of our model through a series of experiments. Gaoqi He, Zhen Liu 0002, Xingjian Lu |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2015 | Geographical Job Scheduling in Data Centers with Heterogeneous Demands and ServersabstractThe fast proliferation of cloud computing promotes the rapid development of large-scale commercial data centers. Tens or even hundreds of geographically distributed data centers have been deployed for better reliability and quality of services. This brings huge energy consumption for data centers. Previous research has proved that the geographical load balancing technique can achieve significant energy cost savings for geographically distributed data centers. However, existing methods for geographical load balancing often assume data centers with homogeneous servers, and workloads with single-dimension or uniform resource demands. This is an over-simplification in reality, especially when modern data centers are typically constructed from a variety of server classes. In this paper, we systematically study the problem of job scheduling for geographically distributed data centers to embrace the heterogeneity of underlying platforms and workloads. We develop a novel distributed algorithm to solve the problem efficiently based on the alternating direction method of multipliers. Extensive evaluations based on real-life data center topology, traffic traces, and electricity price data show high efficiency and efficacy of our method. Xingjian Lu, Fanxin Kong, Jianwei Yin, Xue (Steve) Liu, Huiqun Yu, Guisheng Fan |
CLOUD | 1 |
| 2015 | Distributed Optimal Datacenter Bandwidth Allocation for Dynamic Adaptive Video StreamingabstractVideo streaming systems such as YouTube and Netflix are usually supported by the content delivery networks and datacenters that can consume many megawatts of power. Most existing works independently study the issues of improving quality of experience (QoE) for viewers and reducing the cost and emissions associated with the enormous energy usage of datacenters. By contrast, this paper addresses them both, and jointly optimizes the QoE, the energy cost and emissions by intelligently allocating datacenter bandwidth among different client groups. Specially, we propose a distributed algorithm for achieving the optimal bandwidth allocation. The algorithm novelly decomposes the optimization process into separate ones, which are solved iteratively across datacenters and clients. We demonstrate its convergence by both theoretical proof and experimental validation. The experimental results show that the proposed algorithm converges very fast and achieves much better QoE-cost balance than existing approaches. Fanxin Kong, Xingjian Lu, Mingyuan Xia 0001, Xue (Steve) Liu, Haibing Guan |
ACM Multimedia | 2 |
| 2015 | BURSE: A Bursty and Self-Similar Workload Generator for Cloud ComputingabstractAs two of the most important characteristics of workloads, burstiness and self-similarity are gaining more and more attention. Workload generation, which is a key technique for performance analysis and simulations, has also attracted an increasing interest in cloud community in recent years. Though a large number of methods for synthetically generating bursty or self-similar workloads have been proposed in the literature, none of them can deal with workload generation with both of the two characteristics. In this paper, a configurable and intelligible synthetic generator (BURSE) is proposed for bursty and self-similar workloads in cloud computing based on a superposition of two-state Markov Modulated Poisson Processes (MMPP2s). The proposed generator can produce workloads with both specified intension of burstiness and self-similarity. Detailed experimental evaluation demonstrates the accuracy, robustness and good applicability of BURSE. Jianwei Yin, Xingjian Lu, Xinkui Zhao, Hanwei Chen, Xue (Steve) Liu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | System resource utilization analysis and prediction for cloud based applications under bursty workloads
Jianwei Yin, Xingjian Lu, Hanwei Chen, Xinkui Zhao, Naixue Xiong |
Inf. Sci. | 2 |
| 2013 | A synthetic bursty workload generation method for web 2.0 benchmarkabstractAs one of most important characteristics of Web-based systems'workloads, burstiness is gaining more and more attentions. And synthetically generating bursty workloads is a key technique for performance analysis. In this paper, a configurable and intelligible synthetic bursty workload generation method for Web 2.0 benchmark Olio has been proposed based on 2-state Markovian arrival processes (MAP2). By comparing the actual value of index of dispersion for counts (IDC) estimated from system logs with the target value deduced from MAP2 model, we show that our method is more accurate than related work. Jianwei Yin, Hanwei Chen, Xingjian Lu, Xinkui Zhao |
CLUSTER | 3 |
| 2013 | Distance-aware virtual cluster performance optimization: A hadoop case studyabstractCloud computing and big data are becoming two important developing trends in information technology area. However, data-intensive computing has some challenges to work well on virtual machines in cloud computing for virtualized resource competition and complex network communication. Network becomes one of the most notorious bottlenecks, which highlights strategies to lower communication and transmission cost in virtual cluster. In this paper, we present a novel cluster performance optimization strategy named vClusterOpt. vClusterOpt finds out centralized subgraphs of node graph and choose node with the shortest logical distance as kernel node of the subgraph to reduce inter-machine communication and transmission cost under virtual cluster. To calculate logical distance accurately, we define two kinds of logical distance: Logical Communication Distance(LCD) and Logical Transmission Distance(LTD). VM with the shortest LCD with others is used as the communication kernel node who has the most information communication stress, while VM with the shortest LTD is treated as transmission kernel node who has the most data transmission stress. We choose benchmarks running on Hadoop as the represent of data-intensive computing service to demonstrate effectiveness of our approach. Experiments show that an average of 20% performance improvement can get by our distance-aware virtual cluster optimization strategy. Xinkui Zhao, Jianwei Yin, Zuoning Chen, Xingjian Lu |
CLUSTER | 4 |
| 2013 | An Approach for Bursty and Self-similar Workload Generation
Xingjian Lu, Jianwei Yin, Hanwei Chen, Xinkui Zhao |
WISE (2) | 1 |
| 2012 | An Efficient Data Dissemination Approach for Cloud Monitoring
Xingjian Lu, Jianwei Yin, Ying Li 0001, Shuiguang Deng, Mingfa Zhu |
ICSOC | 1 |