VLDB 2026 Research / reviewers in the wild / expert
Dian Shen
dblp:139/4309
· DBLP profile ↗
68ranked-venue papers
8as first author
56since 2021 · last 2026
0000-0003-0422-5285ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 26 · 3 first-author · 25 since 2021Systems, architecture and hardware · 16 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 10 · 10 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compressing LLM Knowledge into Graph Representations for Text-attributed Graphs LearningabstractText-attributed graphs (TAGs) require jointly modeling relational structure and node-level text.Existing GNN-LLM approaches perform by incorporating large language models at inference time for processing the text attributes, resulting in costly deployment.More fundamentally, LLM knowledge is typically used in a sample-wise manner, leading to inefficient utilization across graph instances.In this work, we study how interactions with LLM embedding spaces affect graph representations, and show that projecting into the LLM space can learn better GNNs.That is to say, the knowledge encoded in LLM embeddings can be compressed into graph representations.Based on this insight, we propose a framework that internalizes LLM knowledge within graph models and supports inference-efficient TAG learning.Our framework employs a hierarchical Proxy-Purifier module with distribution-level regularization, using LLM embeddings only as training-time guidance.With this module, the model operates TAGs without invoking LLMs, achieving high efficiency as standard GNNs without LLMs.Notably, experiments on five popular TAG tasks further demonstrate that our method can also achieve consistent performance gains, in comparison to existing GNN-LLM approaches. Runhuai Chen, Dian Shen, Kaihong Huang, Beilun Wang |
ACL (1) | 2 |
| 2026 | LiGen: Active Lipid Generation via a Molecular Language ModelabstractLipid nanoparticles (LNPs) can deliver cargos to both tumor and immune cells, playing a crucial role in biomedicine.Traditional approaches rely on experimental screening and expert knowledge, which can be costly and time-consuming.Recent methods based on language models have accelerated this process using deep learning.Although these methods can retrieve molecules for fusion or rank candidates from existing libraries, they are still limited by the scope of known formulations.In this work, we propose LiGen to generate lipid molecules efficiently and actively, facilitating the discovery of high-performing LNP formulations.We first train a lipid-specific molecular language model, LiCore, to learn hidden representations of lipid molecules.We then explore the learned latent space to generate improved candidate formulations.This process is guided by a trained predictor, which evaluates delivery efficiency and provides directional signals.In reconstruction task, LiCore achieves near-perfect reconstruction performance output with a low invalid ratio on both the LNP-Virtual900k and LNP-Exp12k datasets.The predictor consistently improves ranking-oriented metrics across multiple cell lines, with our method outperforming the best baselines by an average of 4.1%, 10.8%, and 8.1% in Top-50, Top-10, and Top-5 identification accuracy, respectively.Guided by predictor, LiGen generates novel lipid candidates that achieve a 30.7% relative improvement over baseline methods in predicted delivery efficiency, with some candidates exceeding 50% improvement. Ying Zhan 0001, Xiuqi Tang, Yan Zhang 0100, Xiao Tan 0005, Dian Shen, Beilun Wang |
ACL (1) | 5 |
| 2026 | FreeProbe: Towards Low-Overhead Runtime Observability for High-Performance User-Space SystemsabstractObservability, especially for fine-grained internal metrics, is essential for optimizing high-performance user-space systems. Non-intrusive runtime instrumentation provides such visibility without code modification, rebuilding, or redeployment, making it attractive for production environments. However, in microsecond- or sub-microsecond-scale code paths, existing approaches incur prohibitive overhead due to costly context switches. Even with sampling, this overhead remains because sampling are performed inside the instrumentation logic and cannot eliminate the inherent interception cost. Bin Yang 0027, Xinshu Wang, Beilun Wang, Dian Shen |
APNet | 5 |
| 2026 | eCCA: Deploying In-Network Congestion Control Algorithms in Practice with eBPF
Xinshu Wang, Bin Yang 0027, Jiantao Cheng, Dian Shen |
APNet | 5 |
| 2026 | Decentralizing Compressed Sensing for Federated Learning with Hardware-Software Codesign
Fan Sun 0005, Fang Dong 0001, Dian Shen |
INFOCOM | 3 |
| 2026 | NGSim: A High-Fidelity and Efficient Simulator for Optimizing Network Function Graphs
Bin Yang 0027, Dian Shen, Jianrui Liu, Beilun Wang |
INFOCOM | 2 |
| 2026 | PolicyCache: Intra-flow Learning in Congestion Control
Han Tian, Xudong Liao, Decang Sun, Wenxue Li 0004, Bin Huang 0024, Senbo Fu, Junxue Zhang 0001, Dian Shen, Kai Chen 0005 |
NSDI | 11 |
| 2026 | Identification of Influential Node Group in Attributed Graph through Explaining Graph Neural NetworkabstractIdentification of influential groups of nodes in attributed graphs has applications in a wide range of real-world problems, for instance, collecting important proceedings in citation networks, or identifying essential genes for diagnosing disease in Protein-Protein Interaction networks. Previous approaches for influence maximization manipulated on the graph structure, despite their proliferation, neglect the node attribute information containing additional knowledge. In this work, we introduce Global Graph UNderstanding (GGUN), a perturbation-based framework leveraging the explanatory power of Graph Neural Networks. It takes into account the entire graph structure and node attributes simultaneously and fuses knowledge through GNN layers. Following the perturbation-based explanation, GGUN fills the gap between Deep Neural Network gradient-based feature importance analysis and discrete structure in the graph, which is formulated as a combinatorial optimization problem. Moreover, GGUN obtains an efficient solution by relaxing the infeasible combinatorial optimization problem with performance guaranteed. Evaluations of synthetic and real-world datasets show that GGUN outperforms baselines on both quantitative metrics and human-intelligible analysis. Xiao Tan 0005, Tongtong Su, Yan Zhang 0100, Binghui Xu, Dian Shen, Meng Wang 0009, Beilun Wang |
WWW | 6 |
| 2026 | TolerStore: Tolerating Malicious Nodes in Decentralized Storage NetworkabstractA decentralized storage network (DSN) collects idle storage resources from Internet nodes for low-cost rental to users, and its scale has grown exponentially. In a DSN, most users rely on a centralized third-party service provider (SP) to process userside data and interact with decentralized storage nodes (SNs), making users suffer from a single point of failure. Additionally, since both SP and SNs may offer malicious services, enabling fault tolerance, data confidentiality, and availability guarantee with public verifiability is crucial in the presence of such threats. In this paper, we propose TolerStore, a completely decentralized service framework for DSNs with decentralized SPs and SNs, which can tolerate Byzantine SPs and malicious SNs. To the best of our knowledge, TolerStore is the first to develop a blockchain with multiple SPs for privacy-aware data processing in DSNs, which is formally proven to ensure Byzantine fault tolerance, data confidentiality, public verification, and data availability. Furthermore, we propose an optimized Byzantine Fault-Tolerant consensus with an adaptive leader rotation, incorporating homomorphic fingerprints to verify privacy-aware data processing with enhanced performance. We implement a TolerStore prototype over Hyperledger Fabric, and extensive experiments show that it tolerates [$\frac{N-1}{3}$] Byzantine SPs and 50% malicious SNs with up to 99.99% data availability. Wanning Bao, Liangmin Wang 0001, Haiqin Wu, Dian Shen, Boris Düdder |
IEEE Trans. Computers | 4 |
| 2026 | Corrigendum: DESIGN: Online Device Selection and Edge Association for Federated Synergy Learning-enabled AIoTabstractThis is a corrigendum for the article “DESIGN: Online Device Selection and Edge Association for Federated Synergy Learning-enabled AIoT” published in ACM Trans. Intell. Syst. Technol. 15, 5, Article 104 (November 2024), 28 pages. Shucun Fu, Fang Dong 0001, Dian Shen, Runze Chen 0001, Jiangshan Hao |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2026 | Popularity-Aware Layer-Wise Caching and Function Scheduling for Dynamic Workflow at the EdgeabstractServerless Edge Computing (SEC) has emerged as a promising paradigm for delivering low-latency, resource-efficient services for edge-native applications, which are implemented as dependent functions, forming Directed Acyclic Graph (DAG) workflows. Unfortunately, the application's performance is hindered by the notorious issue of cold startup, especially in resource-constrained SEC environments. Layer- wise container caching has been proven to be an effective startup acceleration solution in SEC, due to its fine granularity and flexibility. However, due to the dynamic nature of call graphs and the skewness in function popularity in DAG workflows, as well as the heterogeneity of container layer cold start time and edge computing environments, the performance of existing layer- wise caching mechanisms degrades significantly. To solve this problem, we propose an efficient DAG workflow deployment method in SEC to minimize the application completion time (ACT) in the long term. We model the problem as a joint optimization of container layer- wise caching and function scheduling, which is a Time-coupled Integer Nonlinear Programming (TINLP) problem. To solve it, we first convert it to an Integer Linear Programming (ILP) problem and propose an online algorithm with theoretical performance guarantees. Extensive experiments demonstrate that our method achieves up to$2.92\times$speedup in ACT reduction. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | Joint Optimization of DNN Model Caching and Request Routing in Mobile Edge ComputingabstractMobile edge computing (MEC) can pre-cache deep neural networks (DNNs) near end-users, providing low-latency services and improving users’ quality of experience (QoE). However, caching all DNN models at capacity-limited edge servers is difficult, and the impact of model loading time on QoE remains underexplored. We explore dynamic DNNs by disassembling a complete DNN model into interrelated submodels to enable fine-grained joint optimization of submodel caching and request routing to balance inference precision and loading latency. In this paper, we study the joint dynamic model caching and request routing problem in MEC networks, aiming to maximize user request inference precision under constraints of server resources, latency, and model loading time. We propose CoCaR, an offline algorithm based on linear programming and random rounding that optimizes joint decisions with a provable performance bound. Furthermore, we develop an online extension, CoCaROL, to adapt to dynamic and unpredictable request patterns. The simulation results demonstrate that CoCaR improves the average inference precision for user requests by 40.1% over state-of-the-art baselines. In addition, CoCaR-OL achieves an improvement of at least 32.3% in users’ QoE over competitive baselines. Shuting Qiu, Fang Dong 0001, Siyu Tan, Ruiting Zhou, Dian Shen, Patrick P. C. Lee, Qilin Fan |
IEEE Trans. Netw. | 5 |
| 2025 | Enhancing Network Traffic Prediction by Integrating Graph Transformer with a Temporal Model
Xiucheng Sun, Runqun Xiong, Dian Shen, Junzhou Luo |
APNet | 3 |
| 2025 | HiSpTRSV: Exploring Tile-Level Parallelism for SpTRSV Acceleration on FPGAsabstractSparse Triangular Solve (SpTRSV) is a critical level2 kernel in sparse Basic Linear Algebra Subprograms (BLAS). While Field-Programmable Gate Array (FPGA) accelerators for SpTRSV focus on optimizing individual tiles, they overlook intertile parallelism. Designing an inter-tile parallelism accelerator poses challenges, including constructing fine-grained dependency graph, handling communication overhead, and balancing workloads. HiSpTRSV addresses these challenges through dependency graph parsing, tile-based highly parallel algorithm, filtering mechanisms, and bidirectional matching with modular indexing. Experiments show that HiSpTRSV outperforms the state-of-the-art SpTRSV accelerator in terms of a 34.3% performance improvement. HiSpTRSV achieves a $3.58 \times$ speedup and $9.59 \times$ higher energy efficiency compared to GPUs. Fan Sun 0005, Fang Dong 0001, Dian Shen |
DAC | 3 |
| 2025 | Information-Agnostic Model Poisoning Attacks Against Byzantine-Robust Federated Learning
Yan Zhang 0100, Yueyao Chen, Xiao Tan 0005, Dian Shen, Meng Wang 0009, Beilun Wang |
DASFAA (4) | 4 |
| 2025 | eNetSTL: Towards an In-kernel Library for High-Performance eBPF-based Network FunctionsabstractUsing extended Berkeley Packet Filter (eBPF) to implement networking functions (NFs) has been a promising trend for modern network infrastructure. In this paper, we endeavor to implement 35 representative NFs with eBPF, but encounter inherent problems of either incomplete functionality or performance degradation of up to 49.2%. Conventional solutions like modifying the eBPF infrastructure or implementing functions directly in the kernel can lead to intrusive and unstable modifications. Bin Yang 0027, Dian Shen, Junxue Zhang 0001, Lunqi Zhao, Beilun Wang, Guyue Liu, Kai Chen 0005 |
EuroSys | 2 |
| 2025 | Bi-perspective Splitting Defense: Achieving Clean-Seed-Free Backdoor SecurityabstractBackdoor attacks have seriously threatened deep neural networks (DNNs) by embedding concealed vulnerabilities through data poisoning. To counteract these attacks, training benign models from poisoned data garnered considerable interest from researchers. High-performing defenses often rely on additional clean subsets/seeds, which is untenable due to increasing privacy concerns and data scarcity. In the absence of additional clean subsets/seeds, defenders resort to complex feature extraction and analysis, resulting in excessive overhead and compromised performance. To address these challenges, we identify the key lies in sufficient utilization of both the easier-to-obtain target labels and clean hard samples. In this work, we propose a Bi-perspective Splitting Defense (BSD). BSD distinguishes clean samples using both semantic and loss statistics characteristics through open set recognition-based splitting (OSS) and altruistic model-based data splitting (ALS) respectively. Through extensive experiments on benchmark datasets and against representative attacks, we empirically demonstrate that BSD surpasses existing defenses by over 20% in average Defense Effectiveness Rating (DER), achieving clean data-free backdoor security. Yangyang Shen, Xiao Tan 0005, Dian Shen, Meng Wang 0009, Beilun Wang |
ICML | 3 |
| 2025 | CoCaR: Enabling Efficient Dynamic DNN-Based Model Caching and Request Routing in MEC
Shuting Qiu, Fang Dong 0001, Siyu Tan, Dian Shen, Ruiting Zhou, Qilin Fan |
INFOCOM | 4 |
| 2025 | HyperCom: Enabling High Performance and Composable Data Structures for Software Network Functions with eBPF
Bin Yang 0027, Dian Shen, Lunqi Zhao, Jianrui Liu, Jiantao Cheng, Beilun Wang |
INFOCOM | 2 |
| 2025 | NoTeNet: Normalized Mutual Information-Driven Tuning-free Dynamic Dependence Network Inference Method for Multimodal DataabstractDynamic Dependence Network (DDN) inference is crucial for understanding evolving relationships in multimodal time series web data, with broad applications in fields like medical and financial network analysis. The inherent dynamic nature, temporal continuity, and heterogeneous data sources in multimodal time series data pose three fundamental challenges: computational efficiency, prediction stability and robustness, and modality quality disparity. Previous methods, generally lacking utilization of multiple modalities, either struggle with computational efficiency due to the time-intensive manual hyperparameter tuning, or compromise prediction stability and robustness by neglecting temporal coherence. To address these challenges, we propose a Normalized mutual information-driven Tuning-free Dynamic Dependence Network inference method for multimodal data, namely NoTeNet. NoTeNet provides a promising paradigm that can integrate two different data modalities to enhance prediction accuracy. It uses normalized mutual information transforms noisy auxiliary data into relationship matrices and employs a kernel function for smooth temporal estimation. Additionally, NoTeNet significantly reduces the need for manual hyperparameter adjustments, offering a tuning-free approach with theoretical guarantees. On various synthetic datasets and real-world data, NoTeNet demonstrates superior prediction accuracy and efficiency without the need for hyperparameter tuning, making it potential for a wide range of web data applications. Xiao Tan 0005, Yangyang Shen, Yan Zhang 0100, Jingwen Shao, Dian Shen, Meng Wang 0009, Beilun Wang |
WWW | 5 |
| 2025 | ADPTD: Adaptive Data Partition With Unbiased Task Dispatching for Video Analytics at the EdgeabstractRecently, edge-assisted methods have been proposed as a promising technique to deliver fast and accurate on-device video analytics by partitioning frame data and dispatching them to edge servers for parallel execution. However, the data partition (DP) reduces the detection latency but decreases accuracy since objects may cross the boundaries of adjacent blocks. The effect of DP on the accuracy and latency depends on multiple vital parameters (e.g., target size, density, network, and computing resources) in an unknown and time-varying fashion. Moreover, these parameters are determined by the application scenarios and edge environment, which are uncertain and heterogeneous at the edge. Hence, how to partition frames to strike a balance between accuracy and latency is a nontrivial and intractable problem. To this end, we propose an online learning-based device-edge–cloud collaboration framework, ADPTD, to guide DP at the edge. We propose an optimal task dispatching algorithm (OTD) to minimize detection latency. Then, we propose a multiarmed bandit-based algorithm to pick a DP strategy and invoke OTD to dispatch tasks in each time slot. Theoretical analysis reveals that ADPTD achieves sublinear regret. Extensive experimental results show that ADPTD outperforms the state-of-the-art methods, achieving a latency reduction of up to$2.53\times $and improving accuracy by up to 49.4%. Zhaowu Huang, Fang Dong 0001, Haopeng Zhu, Mengyang Liu, Dian Shen, Ruiting Zhou, Xiaolin Guo, Baijun Chen |
IEEE Internet Things J. | 5 |
| 2025 | Hourglass: An Adaptive Range Filter with Lightweight Hybrid EncodingabstractRange filters can check whether a queried range is non-empty within a key set, with no false negatives and a low false positive rate. However, existing range filters fail to address recurring false positives in skewed or adversarial queries. In this paper, we propose Hourglass, an adaptive range filter that defends against recurring false positives through lightweight hybrid encoding and semi-sorted adaptivity. Hourglass partitions keys into prefixes, stored in a semi-sorted cuckoo filter, and suffixes, encoded using hybrid encoding schemes based on their sparsity. By preserving the order of fingerprints, the semi-sorted cuckoo filter improves space efficiency. Additionally, Hourglass introduces a new adaptivity strategy that updates fingerprints without violating the semi-sorting order. Further, Hourglass introduces a correlation-aware space allocation model to optimize space across varying key-query correlation degrees. The evaluations show that Hourglass outperforms state-of-the-art range filters under adversarial workloads, achieving a 9.8-35.4X lower false positive rate. Moreover, they demonstrate that Hourglass delivers robust performance on both synthetic and real-world datasets, as well as under varying key-query correlation degrees. Rong Gu 0001, Meng Li 0010, Haipeng Dai 0001, Baohan Wang, Dian Shen |
Proc. ACM Manag. Data | 7 |
| 2025 | HPST-GT: Full-Link Delivery Time Estimation Via Heterogeneous Periodic Spatial-Temporal Graph TransformerabstractA warehouse-distribution integration (WDI) e-commerce platform is an approach that combines warehousing and distribution processes, which is increasingly adopted in industry to enhance business efficiency. In the WDI e-commerce, one of the most important problems is to estimate the full-link delivery time for decision-making. Traditional methods designed for separate warehouse-distribution models struggle to address challenges in integrated systems. The difficulties stem from two main factors: (i) the contextual influence exerted by neighboring units within heterogeneous delivery networks, and (ii) the uncertainty in delivery times caused by dynamic and periodic temporal factors such as fluctuations in online sales volumes and the varying characteristics of different delivery units (e.g., warehouses and sorting centers). To address these challenges, we propose a novel full-link delivery time estimation framework calledHeterogeneousPeriodicSpatial-TemporalGraphTransformer (HPST-GT). First, we develop heterogeneous graph transformers to capture the hierarchical and diverse information of the warehouse-distribution network. Next, we design spatial-temporal transformers based on heterogeneous features to analyze the correlation between spatial and temporal information. Finally, we create a heterogeneous spatial-temporal graph prediction module to estimate full-link delivery time. Our method, evaluated on a one-month dataset from a leading e-commerce platform, surpasses current benchmarks across multiple performance metrics. Shuai Wang 0008, Hai Wang 0019, Li Lin 0011, Xiaohui Zhao 0006, Tian He 0001, Dian Shen, Wei Xi 0003 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | Enhancing Link Performance for Mobile LoRa NetworksabstractLoRa, as a typical representative of Low Power Wide Area Networks (LPWAN), has been widely used to connect massive IoT devices. However, in mobile applications, there is significant packet loss in LoRa transmission due to link performance degradation. Existing studies take little account of end-devices' movement, particularly when the movement pattern is unknown. We propose LMLoRa to enhance theLink Performance forMobile LoRa networks in general scenarios for both single-gateway and multi-gateway applications. The key observation is that, due to LoRa's unique feature, repeating the original packet content enables the use of smaller, more energy-saving transmission parameters, which not only enhances link performance but also reduces energy consumption. Technically, we propose a link performance estimation model based on packet content repetition for both single-gateway and multi-gateway mobile networks. Then, we propose the corresponding channel frequency selection model to avoid transmission collisions. Finally, we design low-overhead communication mechanisms to operate the system. To evaluate the performance of LMLoRa in various scenarios, we design and implement real-world testbeds and a simulation platform for both single-gateway and multi-gateway scenarios. Extensive results show that LMLoRa improves packet delivery ratio by an average of 33.4% to 69.2% compared with the state-of-the-art. Ciyuan Chen, Zhuqing Xu, Runqun Xiong, Dian Shen, Weizheng Wang 0001, Junzhou Luo, Xiaohua Jia |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Multi-Dimensional Training Optimization for Efficient Federated Synergy LearningabstractEdge learning (EL) is an end-to-edge collaborative learning paradigm enabling devices to participate in model training and data analysis, opening countless opportunities for edge intelligence. As a promising EL framework, federated synergy learning (FSyL) mitigates the computation and communication overhead on resource-constrained devices by offloading partial model layers to the edge server for synergistic training. Nevertheless, due to the system and statistical heterogeneity, naively using existing FSyL methods is significantly time-consuming and causes accuracy degradation. Motivated by this issue, this paper introduces a novel FSyL framework that integrates multi-dimensional training optimization and formulates the edge learning cost minimization (ELCM) problem. To tackle the ELCM efficiently, we designOL-MG, anOnLineModel Splitting and Resource ProvisioningGame. Specifically, we first reformulate and decompose the original ELCM based on data quality evaluation. Then, given a model splitting decision, we determine the optimal resource provisioning in Sub-problem1, based on which optimal model splitting in Sub-problem2 is modeled as a potential game. Subsequently, we introduce a decentralized algorithm to find a Nash equilibrium (NE) solution. Furthermore, we further extendOL-MGto support a budget-aware multi-edge scenario. Extensive experiments demonstrate that the proposed mechanism significantly outperforms state-of-the-art methods in cost-saving and accuracy improvement. Shucun Fu, Fang Dong 0001, Runze Chen 0001, Dian Shen, Jinghui Zhang 0001, Qiang He 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Resource-Efficient DNN Inference With Early Exiting in Serverless Edge ComputingabstractServerless Edge Computing (SEC) has gained widespread adoption in improving resource utilization due to its triggered event-driven model. However, deploying deep neural network (DNN) inference services directly in SEC leads to resource inefficiencies, which stem from two key factors. First, existing methods adopt model-wise function encapsulation, which requires the entire DNN model to occupy memory throughout its execution lifecycle. This increases both memory footprint and occupancy time. Second, uniform DNN inference for diversity input leads to redundant computations and additional inference time. To this end, we propose REDI, a novel framework that leverages fine-grained block-wise function encapsulation and progressive inference to provide resource-efficient DNN inference while ensuring latency requirements. REDI enables the release of memory from already inferred shallow networks and allows each request to exit early based on input data complexity, eliminating redundant computations. To fully unleash the potential, REDI jointly considers resource heterogeneity, data diversity, and environment dynamics to investigate the block-wise function placement problem. We introduce an uncertainty-aware online learning-driven algorithm with bounded regret. Finally, we conduct extensive trace-driven experiments to evaluate our methods, demonstrating that REDI achieves a significant speedup of up to$6.52\times$in terms of resource usage cost compared to state-of-the-art methods. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Jinghui Zhang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | DiG-In-GNN: Discriminative Feature Guided GNN-Based Fraud Detector against Inconsistencies in Multi-Relation Fraud GraphabstractFraud detection on multi-relation graphs aims to identify fraudsters in graphs. Graph Neural Network (GNN) models leverage graph structures to pass messages from neighbors to the target nodes, thereby enriching the representations of those target nodes. However, feature and structural inconsistency in the graph, owing to fraudsters' camouflage behaviors, diminish the suspiciousness of fraud nodes which hinders the effectiveness of GNN-based models. In this work, we propose DiG-In-GNN, Discriminative Feature Guided GNN against Inconsistency, to dig into graphs for fraudsters. Specifically, we use multi-scale contrastive learning from the perspective of the neighborhood subgraph where the target node is located to generate guidance nodes to cope with the feature inconsistency. Then, guided by the guidance nodes, we conduct fine-grained neighbor selection through reinforcement learning for each neighbor node to precisely filter nodes that can enhance the message passing and therefore alleviate structural inconsistency. Finally, the two modules are integrated together to obtain discriminable representations of the nodes. Experiments on three fraud detection datasets demonstrate the superiority of the proposed method DiG-In-GNN, which obtains up to 20.73% improvement over previous state-of-the-art methods. Our code can be found at https://github.com/GraphBerry/DiG-In-GNN. Jinghui Zhang 0001, Zhengjia Xu, Dingyang Lyu 0001, Dian Shen, Jiahui Jin 0001, Fang Dong 0001 |
AAAI | 5 |
| 2024 | Factor Model-Based Large Covariance Estimation from Streaming Data Using a Knowledge-Based Sketch MatrixabstractCovariance matrix estimation is an important problem in statistics, with wide applications in finance, neuroscience, meteorology, oceanography, and other fields. However, when the data are high-dimensional and constantly generated and updated in a streaming fashion, the covariance matrix estimation faces huge challenges, including the curse of dimensionality and limited memory space. The existing methods either assume sparsity, ignoring any possible common factor among the variables, or obtain poor performance in recovering the covariance matrix directly from sketched data. To address these issues, we propose a novel method - KEEF: Knowledge-based Time and Memory Efficient Covariance Estimator in Factor Model and its extended variation. Our method leverages historical data to train a knowledge-based sketch matrix, which is used to accelerate the factor analysis of streaming data and directly estimates the covariance matrix from the sketched data. We provide theoretical guarantees, showing the advantages of our method in terms of time and space complexity, as well as accuracy. We conduct extensive experiments on synthetic and real-world data, comparing KEEF with several state-of-the-art methods, demonstrating the superior performance of our method. Xiao Tan 0005, Hao Qian 0003, Jun Zhou 0011, Peibo Duan, Dian Shen, Meng Wang 0009, Beilun Wang |
CIKM | 6 |
| 2024 | Achieving Low Queueing Latency in Time-Slotted LoRa NetworksabstractLoRa, as a Low-Power Wide Area Networks (LP-WAN) technology, is extensively employed for connecting Internet of Things (IoT) applications. LoRa time-slotted networks have gained popularity due to their high channel utilization and robust anti-interference capability. However, the queueing latency of end-devices (EDs) in these networks is often overlooked in the time-slot-scheduled LoRa network, leading to data obsolescence and insufficient notification time. Existing research mainly focuses on reducing transmission delay and avoiding collisions in LoRa networks, while neglecting the importance of ensuring low queueing latency for EDs. In this paper, we propose a semidefinite relaxation (SDR)-based channel scheduler called Q-MAC to achieve low queueing latency in time-slotted LoRa networks. The core idea is to allocate time slots and channels effectively for packets while avoiding collisions. To accomplish this, we formulate an optimization model to minimize latency and packet collisions. This model is a multivariable-coupled non-convex integer problem, we transform the model into a Quadratically Constrained Quadratic Programming (QCQP) problem. Subsequently, we employ the SDR and heuristic algorithms to obtain feasible solutions. Simulation results demonstrate that Q-MAC can significantly reduce queueing latency, achieving an average improvement of 8.57 × compared to existing approaches. Ciyuan Chen, Junzhou Luo, Dian Shen, Zhuqing Xu, Runqun Xiong |
CSCWD | 3 |
| 2024 | Hynify: A High-throughput and Unified Accelerator for Multi-Mode Nonparametric StatisticsabstractNonparametric statistics methods are a class of robust and potent machine learning operators, which are widely used in various domains such as finance, medicine, and computer science. Such methods deliver an accurate estimation without an assumed data distribution. Moreover, they can handle discrete data with various data sources. Despite their desirable features, the calculation of large-scale nonparametric statistics is both compute- and memory-intensive, and the performance overhead hinders them from widespread usage. Kaihong Huang, Dian Shen, Juntao Yang, Beilun Wang |
DAC | 2 |
| 2024 | Large Covariance Estimation from Streaming Data with Knowledge-Based Sketch Matrix
Xiao Tan 0005, Meng Wang 0009, Dian Shen, Weitong Chen 0001, Beilun Wang |
DASFAA (5) | 4 |
| 2024 | Rendering Super Resolution Video Streaming Efficiently with in-Network ComputingabstractEmerging live video streaming applications, e.g., Ultra High Definition videos and interactive video streaming, have put forward new demands for ultra-high bandwidth and reduced delay to match the desired quality of experience. Since current on-device Super-Resolution (SR) approaches are hindered by the limited end-device capabilities, we are motivated to take advantage of Mobile Edge Computing (MEC) and Computing in the Network technologies, such that SR videos can be processed on a more powerful infrastructure by integrating the resources from end-devices through the edge and all the way to the cloud. However, the integration of SR and MEC is non-trivial due to the challenges introduced by the features of SR tasks, and the heterogeneous nature of MEC resources. In this paper, we endeavor to explore and solve these challenges by presenting AVSA, which renders SR live video streaming efficiently with in-network computing. AVSA can adaptively allocate SR work-loads under heterogeneous resources and yield a cost-effective workload allocation with theoretical performance guarantees. Simulation results show that, compared with the state-of-the-art methods, our method achieves up to 13.06× speedup in terms of end-to-end latency. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Baijun Chen, Daheng Yin |
HPCC | 3 |
| 2024 | SkewCache: Skewed Layer-wise Caching for Function Chains in Serverless Edge ComputingabstractIn serverless edge computing (SEC), traditional monolithic applications are encapsulated in multiple dependent functions, in which event-driven service provisioning introduces the cold start problem. Existing works adopt uniform container caching for each function to mitigate the cold start problem, which remains with relatively low efficiency due to ignoring the skewed invocation frequency and the heterogeneity of cold-start behaviors of functions. In this paper, we propose SkewCache, an efficient layer-wise container caching framework for frequency-skewed function chains in SEC. Based on the container’s layer structure, SkewCache enables fine-grained and balanced container caching on multiple edge servers, which takes into account both the invocation frequency skewness and the cold start latency of each function. The problem is modeled as an integer nonlinear programming (INLP) to minimize the application completion time (ACT). To solve the INLP, we first convert it to an equivalent integer linear programming (ILP) form. Then, we propose an approximation algorithm to solve the ILP with a guaranteed approximation ratio. To evaluate the performance of the proposed algorithm, we conduct intensive simulations and the results show that our algorithms outperform baselines, achieving up to 1.94× speedup in terms of ACT reduction. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Haodong Tian |
HPCC | 3 |
| 2024 | Adaptive Group Personalization for Federated Mutual Transfer LearningabstractMutual transfer learning aims to improve prediction with knowledge from related domains. Recently, federated learning is applied in this field to address the communication and privacy concerns. However, previous clustered federated learning (CFL) solutions lack theoretical guarantee of learnability recovery and require time-consuming hyper-parameter tuning, while centralized mutual transfer learning methods lack adaptability to concept drifts. In this paper, we propose the Adaptive Group Personalization method (AdaGrP) to overcome these challenges. We adaptively decide the recovery threshold with a nonparametric method, adaptive threshold correction, for tuning-free solution with relaxed condition. Theoretical results guarantee the perfect learnability recovery with the corrected threshold. Empirical results show AdaGrP achieves 16.9% average improvement in learnability structure recovery compared with state-of-the-art CFL baselines. Haoqing Xu, Dian Shen, Meng Wang 0009, Beilun Wang |
ICML | 2 |
| 2024 | Lmlora: Enhancing Link Performance for Mobile Lora NetworksabstractLoRa, as a typical representative of Low Power Wide Area Networks (LPWAN), has been widely used to connect massive IoT devices. However, in mobile applications, there is massive packet loss in LoRa transmission due to link performance degradation, especially when LoRa end-devices move far from the gateway or into obstructed areas. Existing studies take little account of end-device movement, particularly when the movement pattern is unknown. We propose LMLoRa to enhance the Link Performance for Mobile LoRa networks in general scenarios. The key observation is that repeating the original packet content enhances link performance and allows smaller and more energy-efficient transmission parameter selections. Technically, LMLoRa proposes a link performance estimation model for mobile LoRa networks based on packet content repetition. Second, we exploit key hardware features of LoRa to obtain much continuous RSSI information for link quality prediction. Additionally, LMLoRa develops a channel frequency allocation policy to mitigate transmission collisions. Finally, LMLoRa designs a communication mechanism to assist the estimation model and work the whole system with low communication overhead. We design and implement LMLoRa in complex realworld environments, results show that LMLoRa enhances packet reception rate by 33.4% and energy efficiency by 14.4% on average compared with the state-of-the-art. Ciyuan Chen, Zhuqing Xu, Xiaohua Jia, Jingkai Lin, Runqun Xiong, Dian Shen, Xirui Dong, Junzhou Luo |
ICNP | 6 |
| 2024 | FasMe: Fast and Sample-efficient Meta Estimator for Precision Matrix Learning in Small Sample SettingsabstractPrecision matrix estimation is a ubiquitous task featuring numerous applications such as rare disease diagnosis and neural connectivity exploration. However, this task becomes challenging in small sample settings, where the number of samples is significantly less than the number of dimensions, leading to unreliable estimates. Previous approaches either fail to perform well in small sample settings or suffer from inefficient estimation processes, even when incorporating meta-learning techniques.
To this end, we propose a novel approach FasMe for Fast and Sample-efficient Meta Precision Matrix Learning, which first extracts meta-knowledge through a multi-task learning diagram. Then, meta-knowledge constraints are applied using a maximum determinant matrix completion algorithm for the novel task. As a result, we reduce the sample size requirements to $O(\log p/K)$ per meta-training task and $O(\log\vert \mathcal{G}\vert)$ for the meta-testing task. Moreover, the hereby proposed model only needs $O(p \log\epsilon^{-1})$ time and $O(p)$ memory for converging to an $\epsilon$-accurate solution. On multiple synthetic and biomedical datasets, FasMe is at least ten times faster than the four baselines while promoting prediction accuracy in small sample settings. Xiao Tan 0005, Yangyang Shen, Dian Shen, Meng Wang 0009, Peibo Duan, Beilun Wang |
NeurIPS | 4 |
| 2024 | Privacy-preserving model splitting and quality-aware device association for federated edge learningabstractAbstract Federated edge learning (FEEL) provides a promising device‐edge collaborative learning paradigm, which enables edge devices to parallel participate in model co‐creation while preserving user privacy, opening countless opportunities to enable edge intelligence. With the growing demand for intelligent services, extensive FEEL deployment is inevitable. Nevertheless, existing FL schemes neglect two unique features (i.e., resource heterogeneity and data heterogeneity) in real‐world edge learning and thus may negatively affect the training efficiency and accuracy. Specifically, (1) heterogeneous and limited device resources cause massive laggards, which bring intolerable training delay; (2) heterogeneous data distribution causes device quality divergence, bringing severe training accuracy degradation. This article proposes a split‐based FEEL framework and an adaptive model splitting and quality‐aware device association scheme (MSDA) to tackle the aforementioned challenges. MSDA contains two levels: at the model splitting level, according to device capability and model structure, an adaptive splitting mechanism is proposed to provide a low‐latency and privacy‐preserving model splitting strategy for each device and guide subsequent device association. At the device association level, each device is simulated as a player with a quality weight in the potential game. Then a quality‐aware decentralized device association mechanism is designed to ensure that more high‐quality devices upload local updates before the deadline with the help of the edge server. Finally, experimental results demonstrate that MSDA yields significant improvements, achieving up to 3.1 training speedup and 39% accuracy improvement compared to state‐of‐the‐art methods. Shucun Fu, Fang Dong 0001, Dian Shen, Tianyang Lu |
Softw. Pract. Exp. | 3 |
| 2024 | DESIGN: Online Device Selection and Edge Association for Federated Synergy Learning-enabled AIoTabstractThe artificial intelligence of things (AIoT) is an emerging technology that enables numerous AIoT devices to participate in big data analytics and machine learning (ML) model training, providing various customized intelligent services for industry manufacturing. Federated learning (FL) empowers AIoT applications with privacy-preserving distributed model training without sharing raw data. However, due to IoT devices’ limited computing and memory resources, existing FL approaches for AIoT applications cannot support efficient large-scale model training. Federated synergy learning (FSyL) is a promising collaborative paradigm that alleviates the computation and communication overhead on resource-constrained AIoT devices via offloading part of the ML model to the edge server for end-to-edge collaborative training. Existing FSyL works neither efficiently address the inter-round device selection to improve model diversity nor determine the intra-round edge association to reduce the training cost, which hinders the applications of FSyL-enable AIoT. Motivated by this issue, this article first investigates the bottlenecks of executing FSyL in AIoT. It builds an optimization model of joint inter-round device selection and intra-round edge association for balancing model diversity and training cost. To tackle the intractable coupling problem, we present a framework named Online DEvice SelectIon and EdGe AssociatioN for Cost-Diversity Tradeoffs FSyL (DESIGN). First, the edge association subproblem is extracted from the original problem, and game theory determines the optimal association decision for an arbitrary device selection. Then, based on the optimal association decision, device selection is modeled as a combinatorial multi-armed bandit (CMAB) problem. Finally, we propose an online mechanism to obtain joint DESIGN decisions. The performance of DESIGN is theoretically analyzed and experimentally evaluated on real-world datasets. The results show that DESIGN can achieve up to \(84.3\%\) in cost-saving with an accuracy improvement of \(23.6\%\) compared with the state-of-the-art. Shucun Fu, Fang Dong 0001, Dian Shen, Runze Chen 0001, Jiangshan Hao |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | eMPTCP: A Framework to Fully Extend Multipath TCPabstractMPTCP provides the basic multipath support for network applications to deliver high throughput and robust communication. However, the original MPTCP is designed with limited extensibility. Various research works have tried to extend MPTCP to attain better performance or richer functionalities. These existing approaches either modify the kernel implementation of MPTCP, which involve considerable engineering efforts and may accidentally introduce safety issues, or control MPTCP via userspace tools, which suffer from restricted functionality support. To address this issue, we propose eMPTCP, an easy-to-use framework to fully extend MPTCP without safety risks. Internally, eMPTCP has a modular and pluggable model which allows operators to specify a comprehensive MPTCP extension as a chain of sub-policies. eMPTCP further enforces the policies through packet header manipulations. To ensure safety, eMPTCP is implemented using eBPF. Despite the stringent constraints of eBPF, we show that it is possible to implement an elaborated framework for a fully extensible MPTCP. Through verifying MPTCP in a number of real-world cases and extensive experiments, we show that eMPTCP is able to support a wide range of MPTCP extensions, while the overhead of eMPTCP operations in the kernel is in the scale of nanosecond, and the extra processing time accounts for only about 0.63% of flows’ transmission time. Dian Shen, Bin Yang 0027, Junxue Zhang 0001, Fang Dong 0001, John C. S. Lui |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Joint Optimization of Device Selection and Resource Allocation for Multiple Federations in Federated Edge LearningabstractFederated edge learning (FEEL) is a promising collaborative paradigm, which employs edge devices (EDs) to train machine learning models for a federation. It opens countless opportunities to enable edge intelligence. The increasingly diversified demands for intelligent services are driving the deployment of various federations at the edge. Existing works on FEEL focus on a single federation and ignore inter-federation device competition and intra-device resource allocation, which hinders the applications of FEEL. To address this issue, this article first investigates the bottlenecks of executing multiple federations and builds a joint optimization model as a two-stage Stackelberg game involving device selection and resource allocation. To tackle the problem efficiently, we present a game-theoretical approach namedDeviceSelection andResourceAllocation forMultipleFederationsGame (DSRAMF-G). First, following the arbitrary device selection of leaders (i.e., federations), the time cost minimization of followers (i.e., EDs) is modeled as a convex problem to obtain the optimal resource allocation. Then, based on followers’ optimal responses, device selection is modeled as a congestion game. We prove the existence of the Nash equilibrium and propose a decentralized mechanism. Finally, extensive experiments show that DSRAMF-G significantly outperforms the state-of-the-art methods, achieving up to 5.9x training speedup and 2.8x resource-savings. Shucun Fu, Fang Dong 0001, Dian Shen, Jinghui Zhang 0001, Zhaowu Huang, Qiang He 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Take CARE: Improving Inherent Robustness of Spiking Neural Networks with Channel-wise Activation Recalibration ModuleabstractSpiking Neural Networks (SNNs) are considered the next generation of deep neural networks for their computation efficiency and biological plausibility. Still, SNN models can be fooled with adversarial perturbations and noises. There is an urgent need for building a robust SNN model that can be deployed in safety-critical domains. Recent works successfully proposed some defense methods inspired by those designed for traditional deep neural network models. However, these methods neglect the inherent robustness of SNN models, which has been proven by previous studies. In this paper, we dedicate ourselves to improving the inherent robustness of SNN without additional training. To do that, we unveil that the success of most attacks relies on obfuscating the model activation. Inspired by this phenomenon, we propose a spiking neural network framework Channel-wise Activation Recalibration (CARE) to improve SNN inherent robustness, which is named CARENet. By analyzing the model activation pattern, we prove that the CARE module has a strong capability of activation preservation. We evaluate our method on three benchmarks. Under diverse attacks, including hybrid attacks using multiple attacks, our method shows significant accuracy gains compared to baselines. Furthermore, our framework achieves competitive performance on natural benchmarks. Yan Zhang 0100, Dian Shen, Meng Wang 0009, Beilun Wang |
ICDM | 3 |
| 2023 | WAEVSR: Enabling Collaborative Live Video Super-Resolution in Wide-Area MEC EnvironmentabstractLive video streaming is increasingly popular for its rich content and real-time interactions, but its demand for bandwidth has put a heavy burden on backbone networks. To save bandwidth, recent studies have proposed neural-enhanced live video streaming that deploys deep neural networks (DNNs) for video super-resolution (VSR) on end devices or nearby edge devices to enhance video quality by taking low-resolution frames as input and producing high-resolution output frames. In this solution, the high computational demands of high-quality VSR DNNs make them difficult to support on single end or edge device, necessitating the use of distributed resources in edge facilities. However, the distributed deployment of high-quality VSR DNNs for low-latency inference remains challenging due to the inherent data dependencies of VSR DNNs and the heterogeneity and dynamics of edge facilities. In this paper, we present WAEVSR, a novel collaborative neural-enhanced live video super-resolution system that enables effective leverage of distributed resources to maximize the latency-bounded quality in wide-area MEC environments. WAEVSR consists of two key components: 1) It deploys a parallel-friendly video super-resolution DNN among edge devices, 2) with an inference controller based on the variable-size sliding window to balance the latency and quality of distributed inference in the heterogeneous and dynamics MEC environment. Prototype-based evaluation shows that WAEVSR can achieve 2.5 × lower end-to-end latency than traditional super-resolution serving with a 0.01 drop in SSIM score. The case study also demonstrates its higher stability on latency than vanilla distributed MEC deployment. Daheng Yin, Fang Dong 0001, Baijun Chen, Dian Shen, Ruiting Zhou, Xiaolin Guo, Zhaowu Huang |
IWQoS | 4 |
| 2023 | Joint Quality Evaluation, Model Splitting and Resource Provisioning for Split Edge LearningabstractEdge learning (EL) is an end-edge collaborative learning paradigm that enables numerous edge devices to participate in model training and data analysis, opening countless opportunities to enable edge intelligence. As is a promising EL approach, split edge learning (SPEL) alleviates the computation and communication overhead on resource-constrained devices via offloading part of the machine learning (ML) model to the edge server for cooperative training. Nevertheless, due to the system and statistical heterogeneity of the edge environment, naively using existing SPEL methods brings significantly time-consuming and accuracy degradation. Specifically, system heterogeneity causes intolerable time costs in each training round, while statistical heterogeneity further results in weight divergence and more training rounds to achieve global convergence. Motivated by this issue, this paper designs an efficient SPEL scheme to minimize the total time cost of participating devices. Specifically, we propose a novel SPEL framework and formulate the edge learning cost minimization (ELCM) problem that involves jointly optimizing model splitting and resource provisioning. We design OL-MG, i.e., OnLine Model Splitting and Resource Provisioning Game scheme, to solve the ELCM problem. In OL-MG, we first transform and decompose the original ELCM into two subproblems based on data quality evaluation. Second, we determine the optimal resource provisioning of Sub-problem1 with a given model splitting decision, based on which optimal model splitting of Sub-problem2 is modeled as a potential game. Then, we propose a decentralized algorithm to find a Nash equilibrium (NE) solution for the ELCM problem. Experimental results from both hardware prototype and simulation demonstrate that OL-MG outperforms the state-of-the-art methods, achieving up to 3.1x training cost savings and 40% accuracy improvement. Shucun Fu, Fang Dong 0001, Dian Shen, Qiang He 0001 |
SECON | 3 |
| 2023 | Enabling large-scale low-power LoRa data transmission via multiple mobile LoRa gateways
Ciyuan Chen, Junzhou Luo, Zhuqing Xu, Runqun Xiong, Dian Shen, Zhimeng Yin 0001 |
Comput. Networks | 5 |
| 2023 | DeepMetricCorr: Fast flow correlation for data center networks with deep metric learning
Zunyi Liu, Dian Shen, Jiaang Bao, Fang Dong 0001, Jiong You |
Comput. Networks | 2 |
| 2023 | PADP-FedMeta: A personalized and adaptive differentially private federated meta learning mechanism for AIoT
Fang Dong 0001, Xinghua Ge, Qinya Li, Jinghui Zhang 0001, Dian Shen, Xiao Liu 0004, Gang Li 0009, Fan Wu 0006, Junzhou Luo |
J. Syst. Archit. | 5 |
| 2023 | Multi-Exit DNN Inference Acceleration Based on Multi-Dimensional Optimization for Edge IntelligenceabstractEdge intelligence, as a prospective paradigm for accelerating DNN inference, is mostly implemented by model partitioning which inevitably incurs the large transmission overhead of DNN's intermediate data. A popular solution introduces multi-exit DNNs to reduce latency by enabling early exits. However, existing work ignores the correlation between exit settings and synergistic inference, causing incoordination of device-to-edge. To address this issue, this paper first investigates the bottlenecks of executing multi-exit DNNs in edge computing and builds a novel model for inference acceleration with exit selection, model partition, and resource allocation. To tackle the intractable coupling subproblems, we propose a Multi-exit DNN inference Acceleration framework based on Multi-dimensional Optimization (MAMO). In MAMO, the exit selection subproblem is first extracted from the original problem. Then, bidirectional dynamic programming is employed to determine the optimal exit setting for an arbitrary multi-exit DNN. Finally, based on the optimal exit setting, a DRL-based policy is developed to learn joint decisions of model partition and resource allocation. We deploy MAMO on a real-world testbed and evaluate its performance in various scenarios. Extensive experiments show that it can adapt to heterogeneous tasks and dynamic networks, and accelerate DNN inference by up to 13.7x compared with the state-of-the-art. Fang Dong 0001, Huitian Wang, Dian Shen, Zhaowu Huang, Qiang He 0001, Jinghui Zhang 0001, Liangsheng Wen |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Enabling Distributed and Optimal RDMA Resource Sharing in Large-Scale Data Center Networks: Modeling, Analysis, and ImplementationabstractRemote Direct Memory Access (RDMA) suffers from unfairness issues and performance degradation when multiple applications share RDMA network resources. Hence, an efficient resource scheduling mechanism is urged to optimally allocates RDMA resources among applications. However, traditional Network Utility Maximization (NUM) based solutions are inadequate for RDMA due to three challenges: 1) The standard NUM-oriented algorithm cannot deal with coupling variables introduced by multiple dependent RDMA operations; 2) The stringent constraint of RDMA on-board resources complicates the standard NUM by bringing extra optimization dimensions; 3) Naively applying traditional algorithms for NUM suffers from scalability issues in solving a large-scale RDMA resource scheduling problem. In this paper, we present how to optimally share the RDMA resources in large-scale data center networks with a distributed manner. First, we propose Distributed RDMA NUM (DRUM) to model the RDMA resource scheduling problem as a new variation of the NUM problem. Second, we present distributed algorithms to efficiently solve the large-scale, interdependent RDMA resource sharing problem for different RDMA use cases. Through theoretical analysis, the convergence and parallelism of proposed algorithms are guaranteed. Finally, we implement the algorithms as a kernel-level indirection module in the real-world RDMA environment, so as to provide end-to-end resource sharing and performance guarantee. Through extensive evaluations by large-scale simulations and testbed experiments, we show that our method significantly improves applications’ performance under resource contention, achieving$1.7-3.1\times $higher throughput, and in a dynamic context, the largest performance improvement reaches 98.1% and 64.1% in terms of latency and throughput, respectively. Dian Shen, Junzhou Luo, Fang Dong 0001, Xiaolin Guo, Ciyuan Chen, John C. S. Lui |
IEEE/ACM Trans. Netw. | 1 |
| 2022 | Last-mile Matters: Mitigating the Tail Latency of Virtualized Networks with Multipath Data PlaneabstractVirtualized network has become the cornerstone of today's large-scale cloud data centers. In particular, the data plane of virtualized network, consisting of virtual switch, virtual router and other software network functionalities, performs all network packets processing of virtual machines (VMs). However, current virtualized data plane solutions incur drastic performance interference with co-resident VMs, and thus suffer from unpredictable network performance, especially in terms of tail latency. In this work, we show that the performance issue stems from the fact that CPU plays a dual role of both communication and computation in virtualized networks. A number of virtual network components and their complex packets processing create an undue burden on the hosts' CPUs and in turn cause the mutual performance interference among VMs and networks. To address this issue, we present a multipath data plane solution, where the traffic of VMs can be adaptively and seamlessly offloaded to the adjacent hosts. At the core of this design is to optimize the VM traffic allocation among multiple paths. We formulate the VM multipath traffic allocation problem with coupled variables of computing and network resources, which were only considered as mutually independent in prior researches. Then we present a distributed algorithm to efficiently solve the large-scale, interdependent global optimization problem, with convergence and optimality guarantees. Through extensive simulations and real-world testbed experiments, we show that our solution delivers consistent performance improvement (up to$6.7\times$improvement in aggregate throughput and$21.4\times$reduction in tail latency, respectively) in the dynamic cloud system. Dian Shen, Yi Zhai 0004, Fang Dong 0001, Junzhou Luo |
CLUSTER | 1 |
| 2022 | Towards the Full Extensibility of Multipath TCP with eMPTCPabstractMPTCP provides the basic multipath support for network applications to deliver high throughput and robust communication. However, the original MPTCP is designed with limited extensibility. Various research works have tried to extend MPTCP to attain better performance or richer functionalities. These existing approaches either modify the kernel implementation of MPTCP, which involve considerable engineering efforts and may accidentally introduce security issues, or control MPTCP via user-space tools, which suffer from restricted functionality support. To address this issue, we propose eMPTCP, an easy-to-use framework to fully extend MPTCP without security risks. Internally, eMPTCP has a modular and pluggable model which allows operators to specify a comprehensive MPTCP extension as a chain of sub-policies. eMPTCP further enforces the policies through packet header manipulations. To ensure safety, eMPTCP is implemented using eBPF. Despite the stringent constraints of eBPF, we show that it is possible to implement an elaborated framework for a fully extensible MPTCP. Through verifying MPTCP in a number of real-world cases and extensive experiments, we show that eMPTCP is able to support a wide range of MPTCP extensions, while the overhead of eMPTCP operations in the kernel is in the scale of nanosecond, and the extra processing time accounts for only about 0.63% of flows' transmission time. Bin Yang 0027, Dian Shen, Junxue Zhang 0001, Fang Dong 0001, Junzhou Luo, John C. S. Lui |
ICNP | 2 |
| 2022 | Enabling Latency-Sensitive DNN Inference via Joint Optimization of Model Surgery and Resource Allocation in Heterogeneous EdgeabstractNowadays, edge computing is widely adopted to resolve the emerging deep neural networks (DNNs)-driven intelligence scenarios with the requirement of low-latency and high-accuracy, which includes heterogeneous end devices and DNNs. In such scenarios, the influx of data and computation into a shared edge server incurs prohibitive latency. Thus, we exploit the advantage of Multi-exit DNNs (ME-DNNs) that tasks can exit early at appropriate depths to save inference time. However, naively using ME-DNNs in the heterogeneous edge still fails to deliver fast inference due to improper model surgery and resource allocation. Zhaowu Huang, Fang Dong 0001, Dian Shen, Huitian Wang, Xiaolin Guo, Shucun Fu |
ICPP | 3 |
| 2022 | LoRaDrone: Enabling Low-Power LoRa Data Transmission via a Mobile ApproachabstractLow-Power Wide Area Networks (LPWANs) are widely used to connect large-scale Internet of Things (IoT) applications. Long Range (LoRa) is a promising LPWAN technology sensitive to energy consumption, since LoRa nodes are generally battery-powered, and the battery life will influence the lifetime of the LoRa network. In practice, the battery life of LoRa nodes is short in many scenarios, due to the long transmission distance form the gateway leading to high energy consumption. Existing techniques for energy-efficient data transmission mainly focus on static gateways, and will consume huge energy of remote nodes. In this paper, we propose to integrate LoRa with mobility to minimize the energy consumption of nodes by effectively shortening the transmission distances, and design the first mobile LoRa data transmission system called LoRaDrone by leveraging the unmanned aerial vehicle (UAV) gateway flying close to nodes. Specifically, we present a low-power communication mechanism and a dynamic channel allocation policy to minimize the energy consumed in sensing and communicating with the UAV gateway, while considering the distinctive LoRa parallel reception and complex transmission collisions. Then, an optimal speed scheduling strategy is designed to ensure the reliability of data transmission, and minimize the energy consumption of the UAV. Evaluations on various scales verify the effectiveness of LoRaDrone under different nodes' distributions and UAV paths. Compared with the baselines, the energy consumption of nodes using LoRaDrone is at most reduced by$\mathbf{70.37}\times$at 5000 nodes. Ciyuan Chen, Junzhou Luo, Zhuqing Xu, Runqun Xiong, Zhimeng Yin 0001, Jingkai Lin, Dian Shen |
MSN | 7 |
| 2022 | Exploiting the Computational Path Diversity with In-network Computing for MECabstractWith Computing in the Network technologies, Mobile Edge Computing (MEC) has expanded the resource distribution and tightly integrated computing-network capabilities from the end-devices, through the edge, to the cloud infrastructure, including at points in between. Thus, edge computing is able to deliver a more collaborative processing, better service responding to the increasing application needs in low latency processing. In the presence of integrated computing-network resources and their increased capacity, current proximity-to-data methods in edge computing lead to sub-optimal performance in terms of processing latency. Addressing this issue, this paper presents a Low-latency Adaptive Workload Allocation framework (LAWA) to harness the growing in-network computing resources to deliver low latency processing capabilities for emerging latency-constrained applications. LAWA defines an application by its computational source and destination. Considering the diversity of computing and network resources, we try to find an optimal computational path and its workload allocation. We model the problem as a mixed integer programming problem. To solve this problem, we propose the computational pathfinding and workload allocation algorithms with optimality guarantees. Experimental results show that, comparing with the state-of-the-art methods, our method achieves up to 8.04× speedup, in terms of end-to-end latency. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Zhenyang Ni, Yulong Jiang, Daheng Yin |
SECON | 3 |
| 2021 | Towards Tunable RDMA Parameter Selection at Runtime for Datacenter ApplicationsabstractBecause of the low-latency and high-throughput benefits of RDMA, an increasing number of collaborative applications in datacenters are re-designed with RDMA to boost the performance. Among various low-level hardware primitives provided by RDMA, exposed as parameters of APIs, the application designers select and hardcode them to exploit all the performance benefits of RDMA. However, with the dynamic nature of datacenter application, the hardcoded and fixed parameter selection fails to take full advantages of RDMA capabilities, which can cause up to 35% throughput performance loss. To address this issue, we present a tunable RDMA parameter selection framework, which allows parameter tuning at runtime, adaptive to the dynamic application and server status. To attain the native RDMA performance, we use a lightweight decision tree to reduce the overhead of RDMA parameter selection. Finally, we implement the tunable RDMA parameter selection framework with native RDMA API to provide a more abstract API. To demonstrate the effectiveness of our method, we implement a key-value service based on the abstract API. Experiment results show that our implementation has only a very small overhead compared with the native RDMA, while the optimized key-value service achieves 112% more throughput than Pilaf and 66% more throughput than FaRM. Fang Dong 0001, Dian Shen, Chengtian Zhang, Jinghui Zhang 0001, Junzhou Luo |
CSCWD | 3 |
| 2021 | A CPU Load-awared Virtual Router Placement Strategy in Cloud NetworkabstractWith the increment of the scale of users and networks, network virtualization technology has been widely used by service providers in cloud networks to address elastic network demands. As a fundamental network virtualization component realizing cross-tenant traffic routing, the placement of virtual routers has become a considerable factor influencing the network performance. And through in-depth experiments, we also found that CPU load augments may well incur a significant degradation in network throughput performance, revealing the problems of the placement of virtual routers in existing cloud network modes: 1) Ignore the bandwidth loss caused by the CPU load variation. 2) Lack of theoretical support for the optimal scheme. Based on these above, we propose a CPU load-aware virtual router placement strategy, which balances the computing load situation of each virtual router, and adopts branch and bound algorithm and convex optimization to achieve the approximate optimal placement within 0.1% error. We have evaluated our strategy in our cloud testbed, and find a 20% improvement in terms of cross-tenant throughput compared with the worst case of existing strategies, Fang Dong 0001, Dian Shen, Yi Zhai 0004, Ciyuan Chen |
CSCWD | 3 |
| 2021 | Enabling Low Latency Edge Intelligence based on Multi-exit DNNs in the WildabstractIn recent years, deep neural networks (DNNs) have witnessed a booming of artificial intelligence Internet of Things applications with stringent demands across high accuracy and low latency. A widely adopted solution is to process such computation-intensive DNNs inference tasks with edge computing. Nevertheless, existing edge-based DNN processing methods still cannot achieve acceptable performance due to the intensive transmission data and unnecessary computation. To address the above limitations, we take the advantage of Multi-exit DNNs (ME-DNNs) that allows the tasks to exit early at different depths of the DNN during inference, based on the input complexity. However, naively deploying ME-DNNs in edge still fails to deliver fast and consistent inference in the wild environment. Specifically, 1) at the model-level, unsuitable exit settings will increase additional computational overhead and will lead to excessive queuing delay; 2) at the computation-level, it is hard to sustain high performance consistently in the dynamic edge computing environment. In this paper, we present a Low Latency Edge Intelligence Scheme based on Multi-Exit DNNs (LEIME) to tackle the aforementioned problem. At the model-level, we propose an exit setting algorithm to automatically build optimal ME-DNNs with lower time complexity; At the computation-level, we present a distributed offloading mechanism to fine-tune the task dispatching at runtime to sustain high performance in the dynamic environment, which has the property of close-to-optimal performance guarantee. Finally, we implement a prototype system and extensively evaluate it through testbed and large-scale simulation experiments. Experimental results demonstrate that LEIME significantly improves applications' performance, achieving 1.1–18.7 × speedup in different situations. Zhaowu Huang, Fang Dong 0001, Dian Shen, Junxue Zhang 0001, Huitian Wang, Guangxing Cai, Qiang He 0001 |
ICDCS | 3 |
| 2020 | Distributed and Optimal RDMA Resource Scheduling in Shared Data Center NetworksabstractRemote Direct Memory Access (RDMA) suffers from unfairness issues and performance degradation when multiple applications share RDMA network resources. Hence, an efficient resource scheduling mechanism is urged to optimally allocates RDMA resources among applications. However, traditional Network Utility Maximization (NUM) based solutions are inadequate for RDMA due to three challenges: 1) The standard NUM-oriented algorithm cannot deal with coupling variables introduced by multiple dependent RDMA operations; 2) The stringent constraint of RDMA on-board resources complicates the standard NUM by bringing extra optimization dimensions; 3) Naively applying traditional algorithms for NUM suffers from scalability and convergence issues in solving a large-scale RDMA resource scheduling problem. Dian Shen, Junzhou Luo, Fang Dong 0001, Xiaolin Guo, John C. S. Lui |
INFOCOM | 1 |
| 2020 | Facilitating Application-Aware Bandwidth Allocation in the Cloud with One-Step-Ahead Traffic InformationabstractBandwidth allocation to virtual machines (VMs) has a significant impact on the performance of communication-intensive big data applications hosted in VMs. It is crucial to accurately determine how much bandwidth to be reserved for VMs and when to adjust it. Past approaches typically resort to predicting the long-term network demands of applications for bandwidth allocation. However, lacking of prediction accuracy, these methods lead to the unpredictable application performance. Recently, it is conceded that the network demands of applications can only be accurately derived right before each of their execution phases. Hence, it is challenging to timely allocate the bandwidth to VMs with limited information. In this paper, we design and implement AppBag, an Application-aware Bandwidth guarantee framework, which allocates the accurate bandwidth to VMs with one-step-ahead traffic information. We propose an algorithm to allocate the bandwidth to VMs and map them onto feasible hosts. To reduce the overhead when adjusting the allocation, an efficient Lazy Migration (LM) algorithm is proposed with bounded performance. We conduct extensive evaluations using real-world applications, showing that AppBag can handle the bandwidth requests at run-time, while reducing the execution time of applications by 47.3 percent and the global traffic by 36.7 percent, compared to the state-of-the-art methods. Dian Shen, Junzhou Luo, Fang Dong 0001, Jiahui Jin 0001, Junxue Zhang 0001, Jun Shen 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2019 | MBECN: Enabling ECN with Micro-burst Traffic in Multi-queue Data CenterabstractModern multi-queue data centers often use the standard Explicit Congestion Notification (ECN) scheme to achieve high network performance. However, one substantial drawback of this approach is that micro-burst traffic can cause the instantaneous queue length to exceed the ECN's threshold, resulting in numerous mismarkings. After enduring too many mismarkings, senders may overreact, leading to severe throughput loss. As a solution to this dilemma, we propose our own adaptationthe Micro-burst ECN (MBECN) scheme-to mitigate mismarking. MBECN finds a more appropriate threshold baseline for each queue to absorb micro-bursts, based on steady-state analysis and an ideal generalized processor sharing (GPS) model. By adopting a queue-occupation-based dynamically adjusting algorithm, MBECN effectively handles packet backlog without hurting latency. Through testbed experiments, we find that MBECN improves throughput by ~20% and reduces flow completion time (FCT) by ~40%. Using large scale simulations, we find that throughput can be improved by 1.5~2.4× with DCTCP and 1.26~1.35× with ECN*. We also measure network delay and find that latency only increases by 7.36%. Kexi Kang, Jinghui Zhang 0001, Jiahui Jin 0001, Dian Shen, Junzhou Luo, Zhiang Wu 0001 |
CLUSTER | 4 |
| 2019 | QAECN: Dynamically Tuning ECN Threshold with Micro-burst in Multi-queue Data CentersabstractPacket loss is a common problem in data center networks. The factors causing packet loss are various. Among them, micro-burst is the most important reason. Some previous works have studied the causes and influence o f micro-burst in single queue data center. However, through simulations and experiments, we find that micro-burst could bring m ore serious performance degradation in multi-queue data centers. The micro-burst traffic could cause E CN marking ratio rising from 4% to 22%, and cause throughput loss by up to 40%. Through observing queue length, we find that the standard E CN, which adopts immutable threshold, is not suitable for micro-burst traffic because micro-burst could trigger spurious congestion signals frequently, especially in DCTCP. In this paper, we not only show how much influence the micro-burst brings, but also propose Queue-length Aware ECN (QA-ECN) scheme to mitigate micro-burst. Finally, the simulations and experiments show that QAECN could reduce ECN marking ratio to 2.5%. In addition, the throughput and flow completion time could be improved by up to 22.9% and 34.1%, respectively. Kexi Kang, Jinghui Zhang 0001, Jiahui Jin 0001, Dian Shen, Runqun Xiong, Junzhou Luo |
CSCWD | 4 |
| 2019 | ParaNF: Enabling Delay-Balanced Network Function Parallelism in NFVabstractIn Network Function Virtualization (NFV), multiple network functions cooperate to provide various network services. To reduce the end-to-end latency through a chain of network functions, research hotspots have turned to complete NF parallelism frameworks. However, several issues remain in them such as the manual dependency analysis on NFs and the excessive parallelism for NFs. Therefore, in this paper, we present ParaNF, an effective delay-balanced NF parallelism framework. ParaNF mainly consists of two logical components. First, the ParaNF orchestrator conducts a dynamic dependency analysis to find out which NFs can be parallelized and then conducts a delay-balanced NF parallelism optimization strategy. Second, the ParaNF infrastructure performs light-weight, dynamic packet copying and merging guided by an efficient label mechanism to support high-performance NF parallelism. We implement a ParaNF prototype with DPDK. Our evaluations show that ParaNF not only realizes the line-speed packet processing, but also achieves significant reduction in latency by up to 47% than the traditional SFC and 35% than OpenBox. Junzhou Luo, Fang Dong 0001, Dian Shen |
CSCWD | 4 |
| 2019 | Rendering differential performance preference through intelligent network edge in cloud data centersabstractSummary Sharing the network infrastructure, the performance of emerging distributed applications and services in data centers is directly impacted by the network. As these applications are becoming more and more demanding, it is challenging to satisfy their requirements of low latency, high throughput, and low packet loss rate simultaneously. Prior approaches typically resort to flow control or scheduling mechanisms, prioritizing flows according to their demands. However, none of the methods can solely satisfy the various demands of data center applications. Addressing this challenge, we propose tasch, a preference aware flow scheduling mechanism equipped in the software network edge (ie, end‐host networking). This mechanism utilizes multiple separate queues for flows with different preferences, which guarantees low packet delay for latency‐sensitive flows and provides bandwidth guarantees for throughput‐sensitive flows. A coordinating algorithm is presented to share the network resource among multiple queues with pareto‐optimality. tasch is implemented as a thin and plugable kernel module in Linux based hypervisors, which lies between the complicated physical network and tenants VMs. Subsequently, based on the flow traces of real‐world applications, extensive experiments were conducted to verify the effectiveness of network management mechanism. Dian Shen, Yidan Gao, Xiaolin Guo, Runqun Xiong |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Virtual network fault diagnosis mechanism based on fault injectionabstractDiagnosing faults in virtual networks is always a popular research area. Existing researches primarily focus on diagnosing faults in physical networks, while they could not identify the faults introduced by virtual networks. Besides, the high complexity of algorithms and the requirement for modifying hardware may limit their scope of use. To address these drawbacks, in this paper, we propose a novel approach to diagnose faults in virtual networks. The rational of our approach is that the faults can be identified when located in the packet traces, with the knowledge that the possible known faults that can happen in that location. To achieve this goal, we apply packet marking, fault injection and machine learning techniques to provide precise fault diagnosis. Experimental results show that our approach can efficiently identify 73% of the faults while for virtual network-specific faults, our approach can diagnose 86% of them. Our system can also support real-time or near real-time fault analysis. Fang Dong 0001, Dian Shen, Runqun Xiong, Jiahui Jin 0001 |
CSCWD | 3 |
| 2017 | Enabling application-aware flexible graph partition mechanism for parallel graph processing systemsabstractSummary With the emerging of the large‐scale graph data,Pregel‐like graph parallel processing systems have been an essential tool to efficiently process the graph data. The first step to use thePregel‐like systems is to partition the graph into multiple blocks and distribute them on multiple machines. The partition strategy plays a significant role in determining the performance because a good partition could both ensure load balance and optimize network communication overhead, and vice versa. However, existing partition strategies fail to meet the requirements because they suffer from the following drawbacks: (1) they ignore the application features and (2) they ignore the multi‐application feature in productive environment. To overcome those drawbacks, we proposed thesuperblockpartition strategy, which utilizes theatomic blocksgenerated by pre‐processing of the original graph and could be constructed and re‐constructed dynamically according to the submitted applications in real time. The hash‐based and clustering‐based pre‐partition methods are covered in details. The application feature extraction method and heuristicsuperblockpartition algorithm are proposed to construct the superblocks. Experimental results show that thesuperblockpartition strategy could boost the graph processing performance and its partition efficiency also outperforms the hash‐based and topology optimal partition strategy. Copyright © 2016 John Wiley & Sons, Ltd. Fang Dong 0001, Junxue Zhang 0001, Junzhou Luo, Dian Shen, Jiahui Jin 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | Towards a fast and secure design for enterprise-oriented cloud storage systemsabstractSummary With the rapid development of information technology, enormous volumes of data are being generated by enterprises at all times. The management and storage of these large‐scale data have always been challenging enterprises. As these data are usually shared among users in a collaborative manner, secure data access and access performance are 2 key concerns for data storage of enterprises. However, current solutions fail to meet the requirements of enterprises since they suffer from the following drawbacks: (1) they do not support fine‐grained access control and cannot meet the strict secure data access requirements of enterprises, and (2) they suffer from the unpredictable access latency. Thus in this paper, we propose Frostor, an enterprise‐oriented cloud storage system, which addresses the secure data access issue through a user account and IP‐based fine‐grained access control mechanism, and guarantees the access performance via a two‐level performance optimization mechanism. We further implement Frostor and deploy it on the testbed environment in a real data center. Extensive evaluations have shown that Frostor implements fine‐grained access control, while achieving a significant reduction (≥60%) on access latency. Fang Dong 0001, Dian Shen, Zhuqing Xu, Junzhou Luo |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | AppBag: Application-Aware Bandwidth Allocation for Virtual Machines in Cloud EnvironmentabstractIt is challenging to allocate the network bandwidth to virtual machines(VMs) hosting communication-intensive applications. Due to the temporal and spatial variability of the hosted applications, it is crucial how much bandwidth to be reserved for each VM and when to adjust it. Prior approaches typically resort to predicting the applications' network demands, according to which the VMs are placed once for all or periodically migrated. However, recent works conceded that the network demands of applications can only be accurately derived right before each execution phase. In this paper, we propose AppBag, an Application-aware Bandwidth guarantee framework which allocates the bandwidth to VMs using only one-stepahead information. An efficient VM migration algorithm is then proposed to adjust the bandwidth allocation and corresponding VM placement, subjected to the network demands variation in future execution phases. We further implement AppBag with OpenStack and deploy it on the testbed environment in our data center. Extensive evaluations using popular applications show that AppBag can handle the bandwidth requests at run-time while improving applications' performance and reducing the global traffic in the data center fabric. Dian Shen, Junzhou Luo, Fang Dong 0001, Junxue Zhang 0001 |
ICPP | 1 |
| 2015 | Stochastic modeling of dynamic right-sizing for energy-efficiency in cloud data centers
Dian Shen, Junzhou Luo, Fang Dong 0001, Wei Wang 0089, Guoqing Jin, Weidong Li 0001 |
Future Gener. Comput. Syst. | 1 |
| 2014 | Game theory based dynamic resource allocation for hybrid environment with cloud and big data applicationabstractVirtualization based cloud and big data applications have been widely adopted in various fields. Because deploying the big data applications on the cloud will cause obvious performance degradation, the cloud and big data applications are provided with fixed resource separately. However, the traditional fixed resource allocation mechanism has two drawbacks: (1) low resource utility and (2) unresponsiveness to the performance degradation. To address these drawbacks, the cloud and big data hybrid environment is designed, where fair resource allocation is used to ensure fairness between cloud and big data applications while virtual machine migration is used to make each virtual machine in cloud application reach its own satisfactory. Herein, game theory is used to model the conflict and negotiation between cloud and big data applications. Firstly, the Nash Equilibrium is used to discover the best strategy for both applications. Secondly, as for virtual machine migration, we use Nash Bargaining game to present the situation where virtual machines compete for more resources allocation while their minimal demand is ensured. Finally, experiments are carried out to prove that the hybrid environment outperforms the traditional method both in resource utility and application performance. Junxue Zhang 0001, Fang Dong 0001, Dian Shen, Junzhou Luo |
SMC | 3 |