VLDB 2026 Research / reviewers in the wild / expert
Sijing Duan
dblp:217/1056
· DBLP profile ↗
18ranked-venue papers
9as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 15 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Cellular User Re-Identification Risks With Networking Behaviors Analysis and ModelingabstractMobile network operators (e.g., China Mobile, Verizon) are significant for providing communication services and collecting massive amounts of data. However, operators are increasingly concerned about customer data breaches involving third-party application providers (e.g., Tencent, Apple, Netflix). This concern is particularly aggravated when anonymous datasets shared with third-party providers or publicly released can be linked to user data compromised in breaches, leading to severe re-identification attacks and privacy threats. However, comprehensive methods for identifying such privacy risks on a large scale are lacking due to limited networking behavioral data. To address this, we aim to measure the re-identification privacy risk associated with sharing or releasing cellular traces amidst data breaches. Based on the analysis of key privacyimpacting features in traffic usage and base station association data, we propose a novel re-identification method, SURE, which learns similarities between cellular traces to classify if traces belong to the same user. Extensive experiments on a largescale dataset of 10,000 users over four months demonstrate SURE's superior performance, with AUC scores exceeding 0.9. Our findings reveal significant re-identification risks in data sharing/release, influenced by data scale and user attributes, corroborated by a public dataset. Sijing Duan, Feng Lyu 0001, Yi Ding 0011, Xiaohao He, Yaoxue Zhang |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Fraudulent Delivery Detection with Multimodal Courier Behavior Data in Last-Mile DeliveryabstractThe rapid growth of e-commerce has made last-mile delivery a critical service in daily life. Despite regulations mandating doorstep delivery, the pressure of penalties for delays can lead to fraudulent delivery behaviors, where couriers may report package receipt without actually deliver the package to assigned locations. Existing studies on fraud behavior detection focus on exploring user (courier) behaviors for fraud behavior detection. However, due to the inaccuracy of GPS positioning and the variability of user behavior patterns caused by dynamic environmental factors, relying solely on behavior data remains insufficient for detecting fraudulent deliveries. In this paper, we present a Multimodal Fraudulent Delivery Detection framework (MFDD), which integrates heterogeneous data from multiple agents (courier-side and user-side)-including couriers' physical behavior, digital behavior, and conversations containing customer feedback-for detecting fraudulent deliveries in the last-mile delivery. We employ attention mechanisms to extract features from each modality and use cross-modal fusion to capture complex and varied relationships between multimodal data. To further mitigate modality imbalance during training, we introduce a dynamic gradient-modulation strategy that balances learning across all modalities. We implement and evaluate MFDD on real-world, human-annotated data, achieving a 9.6% improvement in precision and a 5.8% increase in accuracy over the state-of-the-art methods. We also deploy the model in the production environment of JD Logistics, and results show that compared to existing methods, MFDD improves accuracy by 15.3%, reducing estimated annual costs by over 18.5 million CNY. Sijing Duan, Shuxin Zhong, Zhiqing Hong, Weijian Zuo, Desheng Zhang 0002, Yi Ding 0011 |
CIKM | 2 |
| 2025 | Auto-UIT: Automating UAV Inspection Trajectory by Recognizing Pylon Structure from 3D Point CloudabstractUAV-assisted inspection is critical for modern power grid maintenance, enhancing efficiency and safety in remote areas. However, automatically designing UAV inspection trajectories is challenging due to the cluttered inspection environments, small inspection targets, and pervasive obstacles. We propose Auto-UIT, a novel method for generating inspection trajectories in noisy, sparse, and complex 3D point cloud. Auto-UIT has three core techniques: (1) A local structure-enhanced pylon segmentation, which accurately segments pylons, power lines, and surroundings in noisy point cloud for effective inspection target identification and trajectory planning. (2) A 3D fingerprint-based pylon type recognition that compensates for point cloud sparsity to complete missing inspection targets based on the pylon type. (3) An adaptive trajectory generation that samples positions in response to diverse pylon orientations and pervasive environmental obstacles, ensuring UAV operational safety. Our experiments on a real-world dataset across four distinct areas demonstrate that Auto-UIT outperforms existing baseline methods in all three tasks. Furthermore, a four-month deployment in a power grid inspection system—covering a 270 km2 primary mountainous area—yielded an expert first-review acceptance rate of 91.86% for the generated trajectories, and reduced design time by an average of 88.19% compared to manual methods, significantly improving inspection efficiency. Feng Lyu 0001, Lijuan He, Mingliu Liu, Sijing Duan, Hao Wu 0067, Jieyu Zhou, Yi Ding 0011, Zaixun Ling |
MobiCom | 4 |
| 2025 | Demo: UAV Trajectory Generation from Sparse and Noisy 3D Point CloudsabstractUAV-assisted inspection is critical for modern power grid maintenance, enhancing efficiency and safety in remote areas. However, automatically designing UAV inspection trajectories is challenging due to the cluttered inspection environments, small inspection targets, and pervasive obstacles. We propose a novel method for generating inspection trajectories in noisy, sparse, and complex 3D point cloud. It has three core techniques: (1) A local structure-enhanced pylon segmentation, which accurately segments pylons, power lines, and surroundings in noisy point cloud for effective inspection target identification and trajectory planning. (2) A 3D fingerprint-based pylon type recognition that compensates for point cloud sparsity to complete missing inspection targets based on the pylon type. (3) An adaptive trajectory generation that samples positions in response to diverse pylon orientations and pervasive environmental obstacles, ensuring UAV operational safety. A four-month deployment in a power grid inspection system—covering a 270 km2 primary mountainous area—yielded an expert first-review acceptance rate of 91.86% for the generated trajectories, and reduced design time by an average of 88.19% compared to manual methods, significantly improving inspection efficiency. Demo video and dataset are available at https://ljhe006.github.io/autouit/. Lijuan He, Feng Lyu 0001, Mingliu Liu, Hao Wu 0067, Sijing Duan, Jieyu Zhou, Yi Ding 0011, Zaixun Ling |
MobiCom | 5 |
| 2025 | NC-Load: On-Demand Program Loading and Running for Computing Sharing Among IoT DevicesabstractThe number of Internet of Things (IoT) devices has increased rapidly in recent years, but lack effective methods to integrate their computational power. In this article, we propose NC-Load, which couples IoT devices into a multiprocessor system, allowing process scheduling across different devices to share their computing power and improve overall throughput. Specifically, NC-Load consists of three key designs, i.e., remote page fault (RPF), lightweight program cropping, and identical memory layout migration, contributing to three merits compared to existing systems: 1) high storage efficiency: the target device launches the program with a locally stored lightweight icon and leverages RPFs to retrieve the required code/data from the source device; 2) on-demand memory loading: only the required memory portions are transmitted when scheduling programs across different devices, which ensures quick recovery of the program; and 3) consistent memory layout: to ensure consistency of addresses after program offloading, the virtual memory area layout of the source device is migrated to the target device. We implement NC-Load on Linux 6.1 and conduct performance evaluation using unmodified programs and the N-Queens cases. The results demonstrate that NC-Load can achieve superior performance in terms of storage efficiency, program performance, memory usage, and throughput. Yanhao Dong, Sijing Duan, Feng Lyu 0001, Yongmin Zhang, Ju Ren 0001, Yaoxue Zhang |
IEEE Internet Things J. | 2 |
| 2025 | MoCo: Urban User Mobile Contact Detection Based on Cellular Signaling TraceabstractMobile contact exhibits user co-traveling events within the same transportation tool, which is crucial for resident profiling, face-to-face interaction detection, etc. In this paper, we investigate urban user mobile contact detection with cellular signaling traces, which is cost-efficient to enable large-scale detection. Specifically, we develop a data collection platform to collect substantial user signaling traces, covering different types of road scenarios within a city. With the collected traces, we perform systematic data analysis to reveal several technical challenges, which are sparsity of signaling trajectory, remote base station noise, and fuzzy matching difficulties. To address challenges, we propose a mobile contact detection method namedMoCo. InMoCoframework, we first conduct data denoising to remove the noise from remote base stations. Then, we devise a spatio-temporal filter to eliminate unlikely mobile contact traces in both spatial and temporal domains, reducing the computational overhead. Finally, we design a detection network that integrates the submodules of data alignment, feature encoder, spatio-temporal representation learner, and user mobile contact detector. Extensive evaluation results demonstrate the superiority ofMoCoin comparison with state-of-the-art baselines. Robust experiments show thatMoCocan work efficiently in different transportation modes and urban densities. Sijing Duan, Feng Lyu 0001, Huali Lu, Peng Yang 0004, Huaqing Wu, Yaoxue Zhang, Xuemin Shen |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Flexible and Effective Cellular Traffic Data Synthesis with Large Language ModelabstractCellular traffic data hold significant potential for applications such as network planning, traffic prediction, mobility modeling, and personalized recommendations. However, limited data accessibility hinders more open data-driven research. Previous studies have explored data synthesis, while exhibiting flexible limitations in supporting conditional traffic synthesis, and are vulnerable to multidimensional data modeling. In this paper, we present LLMCell, a flexible and effective framework that leverages the arbitrary conditioning and contextual understanding capabilities of the large language model (LLM) to generate high-quality synthetic cellular traffic data. The LLMCell comprises three key components: i) a textual encoder for converting raw cellular traffic data into textual representations, ii) a generative model learner to fine-tune pre-trained LLM based on encoded textual representation for cellular traffic generation, and iii) a synthetic data sampling module for final synthetic data sampling and textual-to-data transformation. Experiments conducted on a large-scale dataset demonstrate the superior fidelity and utility of LLMCell over state-of-the-art baselines, and the synthetic data can effectively preserve user privacy. We release our synthetic dataset to the public to benefit future research in the wireless network community1. Sijing Duan, Feng Lyu 0001, Jinfeng Cen, Ju Ren 0001, Peng Yang 0004, Yaoxue Zhang |
GLOBECOM | 1 |
| 2024 | MOTO: Mobility-Aware Online Task Offloading With Adaptive Load Balancing in Small-Cell MECabstractMobile edge computing is a promising computing paradigm enabling mobile devices to offload computation-intensive tasks to nearby edge servers. However, within small-cell networks, the user mobilities can result in uneven spatio-temporal loads, which have not been well studied by considering adaptive load balancing, thus limiting the system performance. Motivated by the data analytics and observations on a real-world user association dataset in a large-scale WiFi system, in this paper, we investigate the mobility-aware online task offloading problem with adaptive load balancing to minimize the total computation costs. However, the problem is intractable directly without prior knowledge of future user mobility behaviors and spatio-temporal computation loads of edge servers. To tackle this challenge, we transform and decompose the original task offloading optimization problem into two sub-problems, i.e., task offloading control (ToC) and server grouping (SeG). Then, we devise an online control scheme, namedMOTO(i.e.,Mobility-awareOnlineTaskOffloading), which consists of two components, i.e., Long Short Term Memory based algorithm and Dueling Double DQN based algorithm, to efficiently solve theToCandSeGsub-problems, respectively. Extensive trace-driven experiments are carried out and the results demonstrate the effectiveness ofMOTOin reducing computational costs of mobile devices and achieving load balancing when compared to the state-of-the-art benchmarks. Sijing Duan, Feng Lyu 0001, Huaqing Wu, Wenxiong Chen, Huali Lu, Xuemin Shen |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Dynamic RRH-BBU Mapping for C-RAN: A Data-Driven ApproachabstractThe increasing network traffic and dynamic user connections have posed challenges for cellular operators in reducing operating costs while ensuring the quality of service (QoS) for users. Cloud radio access network (C-RAN) addresses these issues by separating baseband units (BBUs) and remote radio heads (RRHs), creating a centralized BBU pool. To optimize C-RAN performance, the key is to dynamically assigning RRHs to BBUs, which is challenging due to cost and QoS constraints. In this paper, we propose a data-driven RRH-BBU mapping scheme (KC-A3C) with deep reinforcement learning (DRL) to improve the performance of large-scale C-RANs. First, we analyze a dataset from a cellular operator containing approximately 26,652 active base stations and use the features of the dataset to construct an RRH popularity metric to cluster RRHs. Second, we model the RRH-BBU mapping as a Markov decision process and use the synchronous Advantage Actor-Critic (A3C) algorithm to find the optimal mapping scheme with the highest long-term gain in a dynamic environment, considering resource utilization, RRH migration, and BBU load balancing. Evaluations using real-world datasets show that our proposed scheme outperforms baseline methods. Fan Wu 0014, Jie Gao 0002, Sijing Duan, Feng Lyu 0001, Huaqing Wu, Yaoxue Zhang, Xuemin Shen |
GLOBECOM | 4 |
| 2023 | VeLP: Vehicle Loading Plan Learning from Human Behavior in Nationwide Logistics SystemabstractFor a nationwide logistics transportation system, it is critical to make the vehicle loading plans (i.e., given many packages, deciding vehicle types and numbers) at each sorting and distribution center. This task is currently completed by dispatchers at each center in many logistics companies and consumes a lot of workloads for dispatchers. Existing works formulate such an issue as a cargo loading problem and solve it by combinatorial optimization methods. However, it cannot work in some real-world nationwide applications due to the lack of accurate cargo volume information and effective model design under complicated impact factors as well as temporal correlation. In this paper, we explore a new opportunity to utilize large-scale route and human behavior data (i.e., dispatchers' decision process on planning vehicles) to generate vehicle loading plans (i.e., plans). Specifically, we collect a five-month nationwide operational dataset from JD Logistics in China and comprehensively analyze human behaviors. Based on the data-driven analytics insights, we design a Vehicle Loading Plan learning model, named VeLP, which consists of a pattern mining module and a deep temporal cross neural network, to learn the human behaviors on regular and irregular routes, respectively. Extensive experiments demonstrate the superiority of VeLP, which achieves performance improvement by 35.8% and 50% for trunk and branch routes compared with baselines, respectively. Besides, we deployed VeLP in JDL and applied it in about 400 routes, reducing the time by approximately 20% in creating plans. It saves significant human workload and improves operational efficiency for the logistics company. Sijing Duan, Feng Lyu 0001, Xin Zhu 0007, Yi Ding 0011, Haotian Wang 0008, Desheng Zhang 0002, Yaoxue Zhang, Ju Ren 0001 |
Proc. VLDB Endow. | 1 |
| 2022 | Mobility-Aware Computation Offloading with Adaptive Load Balancing in Small-Cell MECabstractMobile edge computing (MEC) is a promising computing paradigm enabling mobile devices to offload computation-intensive tasks to nearby edge servers for fast processing. In this paper, we investigate the computing task offloading in small-cell MEC systems. Considering the unevenly distributed mobile users, it is critical to balance the computing load among edge servers to better utilize the computing resources. To this end, we formulate a joint task offloading control and load balancing problem to minimize the average computational cost of users. The formulated problem is a mixed-integer nonlinear optimization problem and is intractable with system scale. To solve the problem in real time, we propose a reinforcement learning-based grouping and task offloading control (RLGTC) scheme. Specifically, we first decompose the problem into two sub-problems with the Tammer method, i.e., the task offloading control (ToC) and server grouping (SeG) sub-problems. Then, we devise two algorithms based on the Kalman Filter technique and reinforcement learning with Dueling Double DQN to solve them, respectively. Extensive data-driven experiments demonstrate the effectiveness of the RLGTC scheme in achieving load balancing and reducing UEs’ computational costs compared to the state-of-the-art benchmarks. Feng Lyu 0001, Huaqing Wu, Sijing Duan, Fan Wu 0014, Yaoxue Zhang, Xuemin Shen |
ICC | 4 |
| 2022 | Dynamic Pricing Scheme for Edge Computing Services: A Two-layer Reinforcement Learning ApproachabstractEdge computing servers (ECSs) have been widely deployed in large-scale mobile edge computing (MEC) systems, which can provide nearby computing services by charging users a price. Service pricing schemes can regulate user task offloading and affect the total revenue of service providers. Investigating how to maximize the revenue of service provider and improve the utilization of edge computing resources becomes crucial while is challenging, considering the users mobility and the uncertainty of users service requests. In this paper, we model the dynamic pricing process of ECS as a Markov decision process and propose a dynamic pricing approach based on Dueling Double Deep Q Network (D3QN) by using the current load conditions and user characteristics, the goal of which is to maximize the revenue of service provider. In addition, considering more ECSs in the MEC system, with the dynamic variations of ECSs loads and the different arrival rate of user tasks, we propose a joint scheduling approach based on D3QN (called RLJS) to collectively improve the total service revenue of service providers. Specifically, we first use a data-driven method to group the ECSs and then devise a D3QN-based task scheduling scheme to distribute tasks among ECS groups by considering the load and price conditions in real time. Simulation results demonstrate the efficacy of RLJS in improving the total revenue of the system provider and reducing the user delays. Feng Lyu 0001, Xinyao Cai, Fan Wu 0014, Huali Lu, Sijing Duan, Ju Ren 0001 |
IWQoS | 5 |
| 2022 | Deep action: A mobile action recognition framework using edge offloading
Heguo Zhang, Sijing Duan, Yunzhen Luo, Fucheng Jia |
Peer-to-Peer Netw. Appl. | 3 |
| 2022 | Multitype Highway Mobility Analytics for Efficient Learning Model Design: A Case of Station Traffic PredictionabstractThe provincial highway transportation system supports substantial cross-city transitions of people and logistics, where the prediction tasks in terms of station/road traffic, urban transitions, and individual traveling are crucial for boosting data intelligence. However, to achieve efficient prediction model design, the predictability analytics with data is the basis, but has not been sufficiently investigated in the existing literature yet. To bridge this gap, in this paper, we study one large-scale dataset collected from one provincial highway transportation system, which contains totally 21,685,765 vehicles and 351,766,743 transaction records, and conduct a comprehensive mobility analytics on its predictable performance. We first investigate the station traffic by mining its spatio-temporal correlations, then examine the multi-type urban transition flows (i.e., people flows and logistics) by demystifying the difference and similarity between the two types of behaviors, and finally analyze the uncertainty of individual traveling behaviors in terms of the destination and arriving time. After that, in accordance with the analytical findings, we cast a case study of data-driven model design for station traffic prediction. Specifically, a novel learning model is devised, named STAR, i.e., Spatio-Temporal Attention based pRediction model, which consists of station outflow/inflow temporal embedding components and spatio-temporal attention blocks to push the limit of prediction capability. Extensive experiments corroborate the efficacy of the proposed STAR. Sijing Duan, Feng Lyu 0001, Ju Ren 0001, Peng Yang 0004, Desheng Zhang 0002, Yaoxue Zhang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Optimizing Federated Learning on Device Heterogeneity with A Sampling StrategyabstractFederated learning (FL) is a novel machine learning that performs distributed training locally on devices and aggregating the local models into a global one. The limited network bandwidth and the tremendous amount of model data that need to be transported bring up expensive communication cost. Meanwhile, heterogeneity in the devices’ local datasets and computation power exerts a huge influence on the performance of FL. To address these issues, we provide an empirical and mathematical analysis of device heterogeneity on the performance of model convergence and quality, then propose a holistic design to efficiently sample devices. Furthermore, we design a dynamic strategy to further speed up convergence and propose the FedAgg algorithm to alleviate the deviation caused by device heterogeneity. With extensive experiments performed in PyTorch, we show that the number of communication rounds required in FL can be reduced by up to 52% on the MNIST dataset, 32% on CIFAR-10, and 28% on FashionMNIST in comparison to the Federated Averaging algorithm. Xiaohui Xu, Sijing Duan, Yunzhen Luo |
IWQoS | 2 |
| 2020 | JointRec: A Deep-Learning-Based Joint Cloud Video Recommendation Framework for Mobile IoTabstractIn the era of Internet of Things (IoT), watching videos on mobile devices has been a popular application in our daily life. How to recommend videos to users is one of the most concerned problem for Internet video service providers (IVSPs). In order to provide better recommendation service to users, they deploy cloud servers in a geo-distributed manner. Each server is responsible for analyzing a local area of user data. Therefore, these cloud servers form information islands and the characteristics of data present nonindependent and identically distribution (non-i.i.d). In this scenario, it is difficult to provide accurate video recommendation service to the minority of users in each area. To tackle this issue, we propose JointRec, a deep learning-based joint cloud video recommendation framework. JointRec integrates the JointCloud architecture into mobile IoT and achieves federated training among distributed cloud servers. Specifically, we first design a dual-convolutional probabilistic matrix factorization (Dual-CPMF) model to conduct video recommendation. Based on this model, each cloud can recommend videos by exploiting the user's profiles and description of videos that users rate, thereby providing more accurate video recommendation services. Then, we present a federated recommendation algorithm which enables each cloud to share their weights and train a model cooperatively. Furthermore, considering the heavy communication costs in the process of federated training, we combine low-rank matrix factorization and 8-bit quantization method to reduce uplink communication costs and network bandwidth. We validate the proposed approach on the real-world data set, and the experimental results indicate the effectiveness of our proposed approach. Sijing Duan, Lingxiang Li, Yaoxue Zhang |
IEEE Internet Things J. | 1 |
| 2018 | Resource allocation optimisation for delay-sensitive traffic in energy harvesting cloud radio access networkabstractIn this study, the authors study a sustainable resource allocation scheme for delay‐sensitive applications in an energy harvesting (EH)‐cloud radio access network (CRAN). The authors formulate an optimisation problem to maximise the user equipment (UE) utility and provide them with strong delay‐guarantee by jointly considering the stochastic EH process, dynamic wireless channel state. By using the Lyapunov stochastic network optimisation technique combined with virtual queues, the authors decompose the formulated problem into four sub‐problems, including channel allocation, data dropping, UE request scheduling and energy management. Based on the solutions of these sub‐problems, a UE optimal resource allocation algorithm is proposed to maximise UE utility while guaranteeing the delay bound and the sustainability of remote radio heads. Furthermore, this algorithm does not require any prior statistical information of the system, e.g. EH process and channel state. Both theoretical analyses and simulation results demonstrate that the proposed algorithm can achieve close‐to‐optimal UE utility, bounded data buffer, delay‐guarantee, and required battery capacity for the operation of the CRAN. Sijing Duan, Zhigang Chen 0001 |
IET Commun. | 1 |
| 2018 | Resource allocation for hybrid energy powered cloud radio access network with battery leakageabstractThis study proposes a resource allocation policy for hybrid energy powered cloud radio access network with battery leakage. To optimise the network utility, the authors first formulate a network utility maximisation problem while jointly considering multiple random processes, which include energy harvesting, data arrival, wireless channel condition, and grid energy prices. To tackle this problem, they exploit the Lyapunov optimisation technique to develop an online dynamic resource allocation framework, which contains four subproblems, i.e. data admission, hybrid energy management, power allocation and route scheduling. Based on the solutions of these subproblems, a network utility optimisation resource allocation (ORA) algorithm is proposed. Specifically, the ORA algorithm only needs to track the current system states without requiring a prior knowledge about channel and energy conditions. Theoretical performance analyses and simulation results verify that the proposed algorithm can achieve close‐to‐optimal utility with bounded data buffer and battery capacity. Sijing Duan, Ju Ren 0001, Yaoxue Zhang |
IET Commun. | 1 |