VLDB 2026 Research / reviewers in the wild / expert
Yingwen Chen 0001
dblp:66/2378-1 · also YingWen Chen 0001
· DBLP profile ↗
50ranked-venue papers
6as first author
34since 2021 · last 2025
0000-0002-8171-1861ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 21 · 3 first-author · 13 since 2021Systems, architecture and hardware · 13 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DecFLLM: A Privacy-Preserving Fine-Tuning Framework for Federated Large Language Models via Adapter Decomposition
Maojiang Wang, Silong Chen, Xu Yang 0028, Zhenyu Qiu, Yingwen Chen 0001, Yuchuan Luo, Shaojing Fu |
ICA3PP (4) | 5 |
| 2025 | Dual-Path Model for Pulmonary Artery SegmentationabstractThe pulmonary artery (PA) is a multi-level vascular system composed of the main pulmonary artery (mPA) and the branch pulmonary arteries (bPA). Accurate segmentation of the PA is of significant importance for the diagnosis of diseases such as pulmonary embolism. However, since the vascular features of the mPA and the bPA are very different, the existing holistic PA segmentation methods will cause the network to pay too much attention to the mPA with significant features and ignore the bPA with weak features, resulting in the imbalance of PA segmentation. To address this issue, we propose a Dual-Path Pulmonary Artery Segmentation Model, which employs two separate paths to learn the features of the mPA and bPA, thereby enhancing the network’s ability to learn features of the bPA. Additionally, we have designed a Skeleton-Optimized Feature Learning Mechanism that optimizes the topological structure of the PA through skeletal guidance, reducing fragmentation and false positives. We conducted experiments on the public dataset provided by PARSE2022, and the results demonstrate that our model indeed improves the segmentation accuracy of the bPA while maintaining the segmentation effectiveness of the mPA. Furthermore, the Skeleton-Optimized Feature Learning Mechanism plays a significant role in reducing vascular fragmentation and false positives. Yingwen Chen 0001, Zhengbo Zhang |
ICASSP | 2 |
| 2025 | Can LLMs only talk? Experimental studies on task scheduling with Large Language ModelsabstractLarge Language Models (LLMs) have emerged as a disruptive technology for Natural Language Processing (NLP), achieving success in NLP-related generative applications. However, the potential capability of LLMs in other domains remains largely unexplored. To explore the potential of task scheduling with LLMs, we model a typical task scheduling scenario in cloud computing and transfer scheduling problems as natural language prompts. Afterward, the knowledge and reasoning abilities of LLMs are enabled to generate scheduling decisions. Six well-known and open-source LLMs are integrated into our framework to perform experimental studies, and the results are evaluated from multiple perspectives and compared with each other. Besides, traditional heuristic algorithms and a basic Reinforcement Learning (RL) method are all performed for comparison. Our results demonstrate: 1) compared to most heuristic methods, the decisions made by LLMs achieve better scheduling performance; 2) compared to the basic RL method, LLMs exhibit better generalization on various workload patterns; 3) the larger parameter size of the LLMs has, the better scheduling performance it achieves. To the best of our knowledge, our experimental study is the first exploration to apply LLMs in task scheduling. Our findings highlight the promising potential of LLMs as a novel approach to task scheduling, offering new avenues for research and practice. Mengjuan Li, Zhengguang Chen, Huan Zhou 0006, Yingwen Chen 0001, Baokang Zhao, Xue Ouyang 0003, Jinshu Su |
ICCCN | 5 |
| 2025 | PPAGAN: A Privacy-Preserving Self-attention GAN Framework for Image Synthesis
Maojiang Wang, Yuchuan Luo, Yingwen Chen 0001, Xu Yang 0028, Shaojing Fu |
ICIC (4) | 3 |
| 2025 | Does One-shot Give the Best Shot? Mitigating Model Inconsistency in One-shot Federated LearningabstractTurning the multi-round vanilla Federated Learning into one-shot FL (OFL) significantly reduces the communication burden and makes a big leap toward practical deployment. However, this work empirically and theoretically unravels that existing OFL falls into a garbage (inconsistent one-shot local models) in and garbage (degraded global model) out pitfall. The inconsistency manifests as divergent feature representations and sample predictions. This work presents a novel OFL framework FAFI that enhances the one-shot training on the client side to essentially overcome inferior local uploading. Specifically, unsupervised feature alignment and category-wise prototype learning are adopted for clients’ local training to be consistent in representing local samples. On this basis, FAFI uses informativeness-aware feature fusion and prototype aggregation for global inference. Extensive experiments on three datasets demonstrate the effectiveness of FAFI, which facilitates superior performance compared with 11 OFL baselines (+10.86% accuracy). Code available at https://github.com/zenghui9977/FAFI_ICML25 Wenke Huang 0003, Tongqing Zhou, Guancheng Wan, Yingwen Chen 0001, Zhiping Cai |
ICML | 6 |
| 2025 | A Closer Look at IPv6 IP-ID Behavior in the Wild
Fengyuan Huang, Zhenzhong Yang, Bingnan Hou, Yingwen Chen 0001, Zhiping Cai |
PAM | 5 |
| 2025 | Memory-efficient programmable packet parsing for multi-tenant terabit networks
Xuetan Cheng, Yingwen Chen 0001, Lailong Luo, Deke Guo |
Comput. Networks | 2 |
| 2025 | GENDN: A Geospatially Enhanced NDN Framework for Location-Related Pub/Sub Services in NTN-Enabled IoTabstractLeveraging satellites and aerials vehicle, nonterrestrial network (NTN)-enabled IoT networks enhance coverage and reliability, enabling global data connections in remote and underserved regions. A key application within these networks is the location-related publish/subscribe service (LPSS), which is geospatial location sensitive, real time, and energy efficient, supporting disaster early warning and environmental monitoring. We demonstrate that, compared to IP technology, named data networking (NDN) is more suited to supporting LPSS. However, current NTN-enabled IoT networks lack mechanisms to utilize geospatial characteristics effectively. Additionally, interactions between IoT devices and aerial vehicles or satellites face challenges, such as low bandwidth, high latency, and intermittent connectivity, which hinder the efficiency of LPSS. We propose geospatially enhanced NDN (GENDN), an adapted NDN framework for supporting LPSS. GENDN incorporates Geohash encoding in content names, allowing flexible use of geospatial characteristics in data subscription. GENDN enhances request aggregation, enabling a single Interest packet (I-pkt) to subscribe to all data in adjacent areas without sequential matching and retrieval. Simulation experiments demonstrate that, compared to traditional NDN, GENDN: 1) effectively leverages geospatial data characteristics, increasing the hit rate of I-pkts in LPSS; 2) reduces the PIT size and network communication overhead, enhancing real-time performance and energy efficiency; and 3) shows potential for large-scale deployment in NTN-enabled IoT environments. Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002, Gaofeng Lv |
IEEE Internet Things J. | 1 |
| 2025 | SMCCA: A sharded multi-task collaborative consensus algorithm for unmanned vehicle networks
Yongming Fu, Yingwen Chen 0001, Mengyuan Zhu, Huan Zhou 0006, Jiachao Wang, Jinshu Su |
J. Syst. Archit. | 3 |
| 2025 | Noise-Robust Federated Learning via Interclient Co-DistillationabstractFederated learning (FL) is a new learning paradigm that enables multiple clients to collaboratively train a high-performance model while preserving user privacy. However, the effectiveness of FL heavily relies on the availability of accurately labeled data, which can be challenging to obtain in real-world scenarios. To address this issue and robustly train shared models using distributed noisy labeled data, we propose FedDQ, a noise-robust FL framework that utilizes co-distillation and quality-aware aggregation techniques. FedDQ incorporates two key features: a noise-adaptive training strategy and an efficient label-correcting mechanism. The noise-adaptive training strategy relies on the estimation of labels' noise levels to dynamically adjust clients' training engagement, which mitigates the impact of wrong labels while efficiently exploring features from clean data. In addition, FedDQ designs a two-head network and employs it for co-distillation. The co-distillation strategy facilitates knowledge transfer among clients to share the representational capabilities. Besides, FedDQ enhances label correction to rectify improper labels through co-filtering and label correction. The experimental results demonstrate the effectiveness of FedDQ in improving model performance and handling noisy data challenges in FL settings. On the CIFAR-100 dataset with noisy labels, FedDQ exhibits a notable improvement of up to 32.4% compared to the baseline method. Liang Gao 0001, Li Li 0064, Yingwen Chen 0001, Shaojing Fu, Dongsheng Wang 0004, Siwei Wang 0001, Cheng-Zhong Xu 0001, Ming Xu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | FooDog: Empower TSN for Efficient PolicingabstractTime-Sensitive Networking (TSN) is an emerging real-time Ethernet technology that provides deterministic communication for time-sensitive (TS) traffic. At its core, TSN utilizes Per-Stream Filtering and Policing (PSFP) gates to mitigate the disruption of unavoidable frame drift. However, as first identified in this work, the naive PSFP gate design results in heavy memory usage, which hinders normal switching functions. This work proposes an efficient PSFP gate design called FooDog. FooDog employs a two-stage structure and a dual-engine policing mechanism to realize memory-efficient, logic-compact, and fast policing while maintaining minimal latency and jitter for TS traffic. Results on FPGA prototypes show that FooDog consumes only hundreds of kilobits of memory, reducing on-chip memory overheads by more than 90% compared to the unoptimized PSFP gate design. Additionally, it maintains end-to-end latency in the microsecond range and jitter below 150 nanoseconds under abnormal traffic conditions, comparable to typical TSN performance without anomalies. Xuyan Jiang, Xiangrui Yang 0002, Tongqing Zhou, Wenfei Wu, Wenwen Fu, Wei Quan 0004, Yingwen Chen 0001, Yihao Jiao, Zhigang Sun 0002 |
IEEE Trans. Netw. | 7 |
| 2025 | Magneto: Load-Balanced Key-Value Service for Write-Intensive WorkloadsabstractHigh-performance key-value (KV) storage is critical for the cloud, providing KV services to various cloud applications. A key challenge for KV services is that workloads of cloud applications are often write-intensive and exhibit highly-skewed characteristics, which result in load imbalance among storage servers thus lowering system performance. To address this problem, prior works set up an in-switch write-back caching mechanism, which adopts a centralized controller to balance writes. However, the limited bandwidth between the switches and the controller is a severe bottleneck for achieving high performance. In this paper, we present Magneto, a novel key-value service architecture for the cloud. At the core of Magneto is a delayed-write mechanism in the switch data plane to absorb frequent write queries for hot items. It effectively balances the load under write-intensive workloads without involving the controller. Magneto also designs a reliability mechanism to ensure system reliability during switch state transitions and failures. We implement a prototype using an FPGA-integrated switch, which has high packet-processing performance and contains enough memory to provide both the in-switch cache and the write buffer. Extensive evaluation shows that Magneto can achieve 8.4x system throughput gains compared to baseline systems when handling a skewed workload consisting of 70% reads and 30% writes. Moreover, it can reduce the load on back-end servers up to 56% in total. Yuanhang Gao, Yingwen Chen 0001, Xiangrui Yang 0002, Huan Zhou 0006, Shihua Tang |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | An Efficient Replication-Based Aggregation Verification and Correctness Assurance Scheme for Federated LearningabstractFederated learning(FL), enabling multiple clients collaboratively to train a model via a parameter server, is an effective approach to address the issue of data silos. However, due to the self-interest and laziness of servers, they may not correctly aggregate the global model parameters, which will cause the final model trained to deviate from the training goal. In the existing proposals, the cryptography-based verification scheme involves heavy computation overheads. On the other hand, the replication-based verification method, relying on a dual-server architecture, can ensure the correctness of aggregation and reduce computation overheads, but incur at least twice the communication cost as that of the task itself. To address these issues, we propose a novel replication-based aggregation scheme for FL, which enables efficient verification and stronger correctness assurance. The scheme employs a main-secondary server architecture, which allows the secondary servers to partakes in aggregation tasks at a predetermined probability, consequently mitigating the validation overhead. Moreover, we resort to the game theory and design a Learning Contract to impose penalties on dishonest servers, enforcing rational servers to correctly compute global model parameters. Under the use of Betrayal Contract to prevent collusion among servers, we further design a training game to efficiently verify global model parameters and ensure their correctness. Finally, we analyze the correctness of the proposed scheme and demonstrate that the computational overhead of our scheme is$\frac{{n + 1}}{{2n}}$of the previous replication-based validation scheme, obtaining a significant reduction in communication cost, where$n$means the training rounds. Experimental results further validate our deduction. Shihong Wu, Yuchuan Luo, Shaojing Fu, Yingwen Chen 0001, Ming Xu 0002 |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | A Geohash-based Naming Scheme for NDN Optimization
Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002 |
APNet | 2 |
| 2024 | DeInfer: A GPU resource allocation algorithm with spatial sharing for near-deterministic inferring tasksabstractFor the applications of artificial intelligence, training models with GPUs are widely noted, while inferring requirements are somehow neglected. In some scenario, it is quite important to finish the deep learning inference (DLI) task and get the response in time, e.g., anomaly detection in AIOps or QoE (Quality of Experience) assurance for customers. However, the challenges of GPU inferring are as follows: 1) the interference among inferring tasks on the shared GPU are not well studied and even not considered, which may cause a surge in the inference latency due to the hardware contention; 2) the deadline miss rate caused by the arrival rate is not clearly considered, which often exhibits significant fluctuations in real-world cases. Therefore, the interference among tasks and arrival rates of tasks should be well designed to decrease the deadline miss rate, when sharing GPU resources. To tackle the issue, we propose the algorithm, DeInfer, through following manners: 1) we identify the key factors that lead to interference, conduct a systematic study and develop a highly accurate interference prediction algorithm based on the random forest algorithm, achieving a four times improvement compared to the state-of-the-art interference prediction algorithms; 2) we utilize the queue theory to model the randomness of the arrival process and put forward a GPU resource allocation algorithm, which reduces the deadline miss rate by an average of over 30%. Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002, Yanfei Yin |
ICPP | 1 |
| 2024 | A Distributed Framework for Subgraph Isomorphism Leveraging CPU and GPU Heterogeneous ComputingabstractSubgraph isomorphism enumerates all embeddings in a data graph that are identical to a query graph. It is a well-known NP-hard problem widely used in various domains, such as bioinformatics, chem-informatics, and social network analysis. Recent works are focused on using GPUs for subgraph isomorphism. Due to the massive scale of intermediate results, current GPU implementations face challenges in scaling across multiple nodes due to high communication costs. The computational power of CPUs is not fully utilized in this process. We present a distributed framework for subgraph isomorphism that leverages CPU and GPU heterogeneous computing. It eliminates the intermediate results on GPU and significantly reduces communication overhead during the load-balancing process. The experiments indicate that our algorithm can be extended to multiple nodes with an almost linear efficiency improvement. Furthermore, our method also significantly outperforms other existing works on GPUs. It can reach an improvement of up to 21 × compared to the state-of-the-art implementation CuTS in the distributed environment. Chen Chen 0016, Li Shen 0007, Yingwen Chen 0001 |
ICPP | 3 |
| 2024 | GeoNDN: Naming the localized data with the Geohash-based Scheme for NDN optimizationabstractNamed Data Networking (NDN) is an emerging paradigm for future networks that facilitates content retrieval by managing data directly through their names. However, in Non-Terrestrial Networks (NTN), including satellites and UAVs, data often have strong geospatial characteristics due to dynamic topologies and geographic dependencies. The current NDN protocol struggles to leverage these geospatial characteristics due to the mismatch between two-dimensional geographical coordinates and one-dimensional data names, resulting in inefficient data retrieval. We propose GeoNDN, a Geohash-based naming scheme that effectively utilizes geospatial characteristics during data retrieval. GeoNDN enhances the aggregation of Interest packets when consumers retrieve data from nearby areas, allowing all data from a target area to be obtained with a single Interest packet instead of sequentially matching each piece of data. Simulation experiments tailored to NTN wireless scenarios demonstrate that GeoNDN reduces the Pending Interest Table (PIT) scale by approximately 25%, decreases network communication overhead by 21%, and increases the Interest packet hit rate by about 12% compared to the traditional NDN naming. The experimental results also show that GeoNDN’s optimization effects are proportional to request density, highlighting its potential for large-scale deployment. Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002, Gaofeng Lv |
ISPA | 1 |
| 2024 | Optimizing In-network Caching for Key-Value Stores under Write-intensive WorkloadsabstractKey-value stores are critical for modern data centers, yet they struggle with performance degradation under write-intensive workloads due to frequent cache invalidations. Traditional read-optimized caching mechanisms fail to maintain efficiency as frequent updates lead to obsolete cached data. As a result, queries sent by clients will suffer from long queuing delays or even be dropped and overall system performance will be degraded severely. This paper presents a novel key-value stores architecture, named Freq-absorb, which addresses this issue through a delayed-write mechanism delaying writes commitment to backend servers. By buffering write queries within the switch for a specific period, Freq-absorb reduces the impact of cache invalidations, improves switch hit rates, and enhances overall system performance. We implement a prototype using an FPGA-based switch that acts as the ToR (Top of Rack) switch to achieve cache and the appended write buffer. Our experimental results show that compared to the read cache mechanism, our approach can reduce switch miss from 99% to 48.4% when handling write-intensive workloads, significantly enhancing performance. Yuanhang Gao, Yingwen Chen 0001, Xiangrui Yang 0002, Huan Zhou 0006, Shihua Tang, Yuanfeng Chen |
ISPA | 2 |
| 2024 | DRL-Tomo: a deep reinforcement learning-based approach to augmented data generation for network tomographyabstractAbstract Accurate and current comprehension of network status is crucial for efficient network management. Nevertheless, direct network measurement strategies entail substantial traffic overhead and demand intricate coordination among network entities, making them impractical. Network tomography, an indirect measurement approach, utilizes insights garnered from measured parts to deduce characteristics of the entire network. Past studies frequently depend on acquiring challenging-to-access information, such as the complete network topology or support from specialized protocols. Unfortunately, these constraints pose challenges in non-cooperative scenarios where obtaining such information is difficult. Recent endeavors pursue emancipating tomography from dependence on copious information, striving to predict unmeasured path performance using limited data. Nevertheless, the disparity between the measured data and actual performance has hindered the accuracy. In response, we introduce an innovative tomography framework named DRL-Tomo, designed to alleviate potential biases. DRL-Tomo initiates by generating augmented data through deep reinforcement learning, gradually approximating the genuine performance of unmeasured paths. Subsequently, a neural network model is trained using this augmented data, enabling precise inferences. Our experiments, encompassing both real-world and synthetic datasets, vividly demonstrate DRL-Tomo’s remarkable enhancement. Specifically, it achieves a substantial 10%–67% improvement in path delay prediction and an impressive 30%–98% enhancement in path loss rate prediction. Changsheng Hou, Bingnan Hou, Xionglve Li, Tongqing Zhou, Yingwen Chen 0001, Zhiping Cai |
Comput. J. | 5 |
| 2024 | A Sketch Framework for Fast, Accurate and Fine-Grained Analysis of Application TrafficabstractAbstract Nowadays, with the continuous increase in internet traffic, the demand for real-time and high-speed traffic analysis has grown significantly. However, existing traffic analysis technologies are either limited by specific applications or data, unable to expand for widespread implementation, or in offline mode are unable to keep up with dynamic adjustments required in certain network management scenarios. A promising approach is to utilize sketch technology to enhance real-time traffic analysis. Unfortunately, existing technologies suffer from defects, such as overly coarse-grained statistics that cannot perform precise application-level traffic analysis, and irreversibility, which cannot support real-time queries in a friendly way. To achieve real-time fine-grained application traffic analysis in general scenarios, we propose AppSketch, a real-time network traffic measurement tool. AppSketch adopts a one-pass approach to classify and label the application information of each packet in the network flows. It then hashes the flow, identified with the application tag, into a carefully designed multiple-key sketch, for gathering application-specific statistics. We conducted extensive experiments using a real-world network traffic dataset collected on a university campus. The results showed that AppSketch achieved high accuracy while requiring less update time than other alternatives. Moreover, AppSketch occupies limited memory ($ {\leq }$64KB), making it suitable for online network devices. Changsheng Hou, Chunbo Jia, Bingnan Hou, Tongqing Zhou, Yingwen Chen 0001, Zhiping Cai |
Comput. J. | 5 |
| 2023 | 6Search: A reinforcement learning-based traceroute approach for efficient IPv6 topology discovery
Ning Liu 0015, Chunbo Jia, Bingnan Hou, Changsheng Hou, Yingwen Chen 0001, Zhiping Cai |
Comput. Networks | 5 |
| 2023 | P2Ride: Practical and Privacy-Preserving Ride-Matching Scheme for RidesharingabstractAs a popular instance of sharing economy, ridesharing has been widely adopted in recent years. To use the convenient ridesharing service, riders and drivers have to share with the service provider their private trip information, which impedes users from freely enjoying the benefits of ridesharing. However, existing studies in ridesharing mainly focus on the optimization of rider-driver matching but ignore the protection of privacy of users. In this paper, we propose P2Ride, a Practical and Privacy-preserving Ride-matching scheme for ridesharing, which enables the service provider to efficiently match drivers with appropriate riders without learning the privacy of both drivers and riders. In P2Ride, we first convert the complex ride-matching computation into equality testing by leveraging overlapping partition systems, and then achieve the privacy-preserving ride-matching by designing a novel non-interactive private equality testing protocol. We prove the security of the proposed P2Ride theoretically. Moreover, a prototype of the P2Ride is implemented, and the experiment results over a real-world dataset demonstrate that the proposed P2Ride can achieve both high ride-matching accuracy and practical efficiency. Yuchuan Luo, Shaojing Fu, Xiaohua Jia, Ming Xu 0002, Yingwen Chen 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | FedDC: Federated Learning with Non-IID Data via Local Drift Decoupling and CorrectionabstractFederated learning (FL) allows multiple clients to collectively train a high-performance global model without sharing their private data. However, the key challenge in federated learning is that the clients have significant statistical heterogeneity among their local data distributions, which would cause inconsistent optimized local models on the clientside. To address this fundamental dilemma, we propose a novel federated learning algorithm with local drift decoupling and correction (FedDC). Our FedDC only introduces lightweight modifications in the local training phase, in which each client utilizes an auxiliary local drift variable to track the gap between the local model parameter and the global model parameters. The key idea of FedDC is to utilize this learned local drift variable to bridge the gap, i.e., conducting consistency in parameter-level. The experiment results and analysis demonstrate that FedDC yields expediting convergence and better performance on various image classification tasks, robust in partial participation settings, non-iid data, and heterogeneous clients. Liang Gao 0001, Huazhu Fu, Li Li 0064, Yingwen Chen 0001, Ming Xu 0002, Cheng-Zhong Xu 0001 |
CVPR | 4 |
| 2022 | Tereis: A Package-Based Scheduling in Deep Learning SystemsabstractDeep learning (DL) systems are typically used to accelerate training DL jobs. Training DL models requires feeding mass input data. It takes a long time to transfer training data from the storage nodes to the compute nodes. However, the computational resources of the GPUs are idle during the data transmission period, which results in a waste of computing resources. In DL systems, a large number of short-term jobs are queuing longer than their own execution times. Meanwhile, many multi-GPU jobs are suffering a long-queuing time due to not enough free GPU. To the best of our knowledge, no studies try to use the idle computation resources of GPU in the data transmission period.We propose Tereis, a package-based scheduler to make full use of GPU. Tereis predicts a DL job’s execution time and data transmission time, then safely package two jobs on the same GPU. One of the packaged jobs will be completed before the other job ends transferring data, which is what ‘safely’ means. This ‘safe’ Packaging does not cause GPU contention. Tereis also designs multi-level queues to prevent starvation. We implemented Tereis in python on the actual cluster and evaluated its performance. The experiment results show that Tereis decreases the average waiting time by 3.2× to10.5×, and has an improvement of 18% to 42% on makespan compared to other methods. Furthermore, we have made large-scale simulations to explore the sensitivities of Tereis. We find that Tereis performs better in the scenario where the job set has a distribution of large dataset size and jobs are submitted frequently. Chen Chen 0016, Yingwen Chen 0001, Jianchen Han |
ICPADS | 3 |
| 2022 | PickyMan: A Preemptive Scheduler for Deep Learning Jobs on GPU ClustersabstractDeep learning (DL) jobs normally run on GPU clusters. Some DL jobs need to be scheduled preemptively to avoid long waiting times. However, preempting a DL job is time-consuming, which consists of suspending and resuming. Suspending needs to complete the training process of the current epoch, and resuming needs to reload the model and the training data. The existing schedulers almost do not consider the overhead of preempting jobs; thus, they may preempt jobs with large time loss, increasing the waiting time and the makespan.In this paper, we present PickyMan, a preemptive scheduler to minimize the overhead of preempting jobs to reduce the average waiting time and the makespan. PickyMan has some innovations. (1) Predict execution time using network traffic and database. It predicts the execution time of a DL job by profiling the network traffic from storage nodes to computation nodes and using a database, without the requirement of allocating extra resources from the cluster. It can use profiled information of only four jobs to predict the execution times for other same-model jobs, and most of the predicted errors are less than 10%. (2) Modeling the overhead of preemption. It builds a model to predict the time loss of job suspensions and resumptions with an average error of less than 5%. (3)We abstract the problem of choosing the appropriate jobs for preemption as one of finding an ordered division of the set of running jobs and solve it quickly with a greedy algorithm. By conducting experiments on the small-scale actual cluster and making large-scale simulations, PickyMan reduces the average waiting time by 10%–92% and further reduces the makespan by up to 14%, compared to existing methods. Chen Chen 0016, Yingwen Chen 0001, Zhaoyun Chen, Jianchen Han, Guangtao Xue |
IPCCC | 2 |
| 2022 | Approximate Shortest Distance Queries with Advanced Graph Analytics over Large-scale Encrypted GraphsabstractUnderstanding graph characteristics is of great importance for graph analytics. Among the many properties, shortest path distance is the fundamental and widely used one. With the advent of cloud computing, it is a natural choice for the data owners to host their massive graphs on the cloud and outsource the shortest distance querying service to it. However, the new paradigm brings serious security concerns as graph data and shortest distance queries may contain sensitive information of data owners and users. In this paper, we propose a novel scheme to support privacy-preserving approximate shortest distance queries with advanced graph analytics over large-scale encrypted graphs, which enables an untrusted cloud to answer shortest distance queries as well as advanced graph metrics (e.g., node centrality) without knowing the content of queries and the sensitive information of outsourced graphs. Compared with the state-of-the-art solutions, our design can support not only efficient and accurate shortest distance approximation, but also advanced graph analytics. We prove that our scheme is secure under the chosen-plaintext model. Experimental results over real-world datasets show that our scheme achieves high approximation accuracy with practical efficiency. Yuchuan Luo, Dongsheng Wang 0004, Shaojing Fu, Ming Xu 0002, Yingwen Chen 0001 |
MSN | 5 |
| 2022 | Collusion-Tolerant Data Aggregation Method for Smart Grid
Yingwen Chen 0001, Kaiyu Cai, Dongsheng Wang 0004, Yuchuan Luo, Guangtao Xue |
WASA (1) | 2 |
| 2022 | Privacy-preserving WiFi Fingerprint Localization Based on Spatial Linear Correlation
Xu Yang 0028, Yuchuan Luo, Ming Xu 0001, Shaojing Fu, Yingwen Chen 0001 |
WASA (1) | 5 |
| 2022 | FGFL: A blockchain-based fair incentive governor for Federated Learning
Liang Gao 0001, Li Li 0064, Yingwen Chen 0001, Cheng-Zhong Xu 0001, Ming Xu 0002 |
J. Parallel Distributed Comput. | 3 |
| 2021 | FIFL: A Fair Incentive Mechanism for Federated LearningabstractFederated learning is a novel machine learning framework that enables multiple devices to collaboratively train high-performance models while preserving data privacy. Federated learning is a kind of crowdsourcing computing, where a task publisher shares profit with workers to utilize their data and computing resources. Intuitively, devices have no interest to participate in training without rewards that match their expended resources. In addition, guarding against malicious workers is also essential because they may upload meaningless updates to get undeserving rewards or damage the global model. In order to effectively solve these problems, we propose FIFL, a fair incentive mechanism for federated learning. FIFL rewards workers fairly to attract reliable and efficient ones while punishing and eliminating the malicious ones based on a dynamic real-time worker assessment mechanism. We evaluate the effectiveness of FIFL through theoretical analysis and comprehensive experiments. The evaluation results show that FIFL fairly distributes rewards according to workers’ behaviour and quality. FIFL increases the system revenue by 0.2% to 3.4% in reliable federations compared with baselines. In the unreliable scenario containing attackers which destroy the model’s performance, the system revenue of FIFL outperforms the baselines by more than 46.7%. Liang Gao 0001, Li Li 0064, Yingwen Chen 0001, Wenli Zheng, Cheng-Zhong Xu 0001, Ming Xu 0002 |
ICPP | 3 |
| 2021 | A Hybrid Framework for Class-Imbalanced Classification
Lailong Luo, Yingwen Chen 0001, Junxu Xia, Deke Guo |
WASA (1) | 3 |
| 2021 | EMPC: Energy-Minimization Path Construction for data collection and wireless charging in WRSN
Ping Zhong 0002, Aikun Xu, Shigeng Zhang, Yiming Zhang 0003, Yingwen Chen 0001 |
Pervasive Mob. Comput. | 5 |
| 2021 | A Mobile-assisted Edge Computing Framework for Emerging IoT ApplicationsabstractEdge computing (EC) is a promising paradigm for providing ultra-low latency experience for IoT applications at the network edge, through pre-caching required services in fixed edge nodes. However, the supply-demand mismatch can arise while meeting the peak period of some specific service requests. The mismatch between capacity provision and user demands can be fatal to the delay-sensitive user requests of emerging IoT applications and will be further exacerbated due to the long service provisioning cycle. To tackle this problem, we propose the mobile-assisted edge computing framework to improve the QoS of fixed edge nodes by exploiting mobile edge nodes. Furthermore, we devise a CRI (Credible, Reciprocal, and Incentive) auction mechanism to stimulate mobile edge nodes to participate in the services for user requests. The advantages of our mobile-assisted edge computing framework include higher task completion rate, profit maximization, and computational efficiency. Meanwhile, the theoretical analysis and experimental results guarantee the desirable economic properties of our CRI auction mechanism. Deke Guo, Siyuan Gu, Lailong Luo, Xueshan Luo, Yingwen Chen 0001 |
ACM Trans. Sens. Networks | 6 |
| 2021 | A Blockchain-Based Medical Data Sharing Mechanism with Attribute-Based Access Control and Privacy ProtectionabstractThe rapid development of wearable sensors and the 5G network empowers traditional medical treatment with the ability to collect patients’ information remotely for monitoring and diagnosing purposes. Meanwhile, the health‐related mobile apps and devices also generate a large amount of medical data, which is critical for promoting disease research and diagnosis. However, medical data is too sensitive to share, which is also a common issue for IoT (Internet of Things) data. The traditional centralized cloud‐based medical data sharing schemes have to rely on a single trusted third party. Therefore, the schemes suffer from single‐point failure and lack of privacy protection and access control for the data. Blockchain is an emerging technique to provide an approach for managing data in a decentralized manner. Especially, the blockchain‐based smart contract technique enables the programmability for participants to access the data. All the interactions are authenticated and recorded by the other participants of the blockchain network, which is tamper resistant. In this paper, we leverage the K‐anonymity and searchable encryption techniques and propose a blockchain‐based privacy‐preserving scheme for medical data sharing among medical institutions and data users. To be specific, the consortium blockchain, Hyperledger Fabric, is adopted to allow data users to search for encrypted medical data records. The smart contract, i.e., the chaincode, implements the attribute‐based access control mechanisms to guarantee that the data can only be accessed by the user with proper attributes. The K‐anonymity and searchable encryption ensure that the medical data is shared without privacy leaking, i.e., figuring out an individual patient from queries. We implement a prototype system using the chaincode of Hyperledger Fabric. From the functional perspective, security analysis shows that the proposed scheme satisfies security goals and precedes others. From the performance perspective, we conduct experiments by simulating different numbers of medical institutions. The experimental results demonstrate that the scalability and performance of our scheme are practical. Yingwen Chen 0001, Linghang Meng, Huan Zhou 0006, Guangtao Xue |
Wirel. Commun. Mob. Comput. | 1 |
| 2019 | Building Trustful Crowdsensing Service on the Edge
Biao Yu, Yingwen Chen 0001, Shaojing Fu, Wanrong Yu |
WASA | 2 |
| 2018 | The Fusion of VMs and Processes: A System Perspective of cKernelabstractVirtual machines (VMs) and processes are two important abstractions for cloud virtualization, where VMs usually install a complete operating system (OS) executing user processes. Although existing in different layers in the virtualization hierarchy, VMs and processes have overlapped functionalities. For example, they are both intended to provide execution abstraction (e.g., physical/virtual memory address space), and share similar objectives of isolation, cooperation and scheduling. However, neither of them could provide the benefits of the other: VMs provide higher isolation, security and portability, while processes are more efficient, flexible and easier to schedule and cooperate. Currently, this heavyweight architecture degrades both efficiency and security of cloud services. There are two trends for cloud virtualization: the first is to enhance processes to achieve VM-like security, and the second is to reduce VMs to achieve process-like flexibility. Based on these observations, our vision is that in the near future VMs and processes might be fused into one new abstraction for cloud virtualization that embraces the best of both, providing VM-level isolation and security while preserving process-level efficiency and flexibility. We describe a reference implementation, dubbed cKernel (customized kernel), for the new abstraction. Essentially, cKernel enhances the exokernel architecture by (i) adopting the LibOS paradigm to assemble isolated, smallest possible "execution environments", and (ii) following the the "core-shell" model to dynamically add traditional process features to the environments. Yiming Zhang 0003, Dongsheng Li 0001, Yingwen Chen 0001, Ping Zhong 0002, Yongqiang Xiong, Huaimin Wang 0001 |
ICDCS | 5 |
| 2017 | An Adaptive MAC Protocol for Wireless Rechargeable Sensor Networks
Ping Zhong 0002, Shuaihua Ma, Jianliang Gao, Yingwen Chen 0001 |
WASA | 5 |
| 2015 | Privacy-Preserving Public Auditing Together with Efficient User Revocation in the Mobile Environments
Feng Chen 0015, Yuchuan Luo, Yingwen Chen 0001 |
WASA | 4 |
| 2014 | Empirical Study on Spatial and Temporal Features for Vehicular Wireless Communications
Yingwen Chen 0001, Ming Xu 0002, Pei Li 0001 |
WASA | 1 |
| 2014 | Throughput Prediction-Based Rate Adaptation for Real-Time Video Streaming over UAVs Networks
Tongqing Zhou, Ming Xu 0002, Yingwen Chen 0001 |
WASA | 4 |
| 2013 | Roadside Infrastructure Placement for Information Dissemination in Urban ITS Based on a Probabilistic Model
Bo Xie 0005, Geming Xia, Yingwen Chen 0001, Ming Xu 0002 |
NPC | 3 |
| 2013 | A traffic light extension to Cell Transmission Model for estimating urban traffic jamabstractUrban traffic congestions have become a financial and societal burden in many cities. Efficient traffic management solutions mitigating such congestions require a reliable modeling and estimation of traffic jams. In urban traffic, the modeling challenges are related to flow collisions and gridlocks created at intersections. In this paper, we propose a traffic light extended Cell Transmission Model (CTM), where the influence of flow collisions and gridlocks are modeled by a single CK parameter. Our approach only requires adapting CK for each intersection type/geometry instead of a complex mathematical formulae proposed in related works. We formalize the description of our urban CTM, and evaluate its capability to model traffic volumes and jams against the microscopic traffic simulator SUMO (Simulation of Urban MObility). Results show that the CK parameter is able to closely reproduce the impact of collisions and gridlocks on traffic jam, making the proposed urban CTM suitable to predict traffic congestions in urban environments. Bo Xie 0005, Ming Xu 0002, Jérôme Härri, Yingwen Chen 0001 |
PIMRC | 4 |
| 2013 | Local Information Storage Protocol for Urban Vehicular Networks
Bo Xie 0005, Yingwen Chen 0001, Ming Xu 0002, Yuangang Wang |
WASA | 2 |
| 2012 | An Empirical View on Opportunistic Forwarding Relay Selection Using Contact Records
Jianbin Jia, Yingwen Chen 0001, Ming Xu 0002 |
NPC | 2 |
| 2012 | Möbius-deBruijn: The product of Möbius cube and deBruijn digraph
Deke Guo, Guiming Zhu, Hai Jin 0001, Panlong Yang, Yingwen Chen 0001, Xianqing Yi, Junxian Liu |
Inf. Process. Lett. | 5 |
| 2011 | Network-Leading Association Scheme in IEEE 802.11 Wireless Mesh NetworksabstractThe association policy in current IEEE 802.11 networks usually considers Received Signal Strength Indication (RSSI) to be the only metric to capture access link quality. However, when a Mesh Client (MC) in IEEE 802.11-based Wireless Mesh Network (WMN) needs to be associated with the most appropriate Mesh Access Point (MAP), the quality of both the access link and the routing path in mesh backhaul should be considered. To take into account this requirement, most existing approaches rely on MAPs to provide more information for MCs such as traffic load, routing metric, airtime cost, etc. These solutions inevitably need to modify the IEEE 802.11 standard or wireless interface drivers on both MAPs and MCs, which is not feasible or flexible in actual deployment. In this paper, a network-leading association scheme is proposed for IEEE 802.11 WMNs. It is completely operated by MAPs and does not require any modification on MCs. It also adapts to the dynamic network environment and always selects the best MAP to serve an MC. Simulation results on ns-3 platform indicate that the network-leading association scheme remarkably improves the performance of IEEE 802.11 WMNs as compared with the existing approaches. Xudong Wang 0001, Ming Xu 0002, Yingwen Chen 0001 |
ICC | 4 |
| 2010 | POCOSIM: A Power Control and Scheduling Scheme in Multi-Rate Wireless Mesh Networks
Weihuang Li, Yingwen Chen 0001, Ming Xu 0002 |
UIC | 3 |
| 2006 | In-Network Data Processing forWireless Sensor NetworksabstractIn wireless sensor networks, energy is the most crucial resource. In-network data processing is a common technique in which an intermediate proxy node is chosen to house a possibly complicated data transformation function to consolidate the sensor data streams from the source nodes, en route to the sink node. We investigate into the placement problem of the proxy. We formulate and solve the energy minimization problem analytically, based on an ENergy- Efficient Rate-Governed Yardstick (ENERGY). An optimal solution is derived based on complete network topology information. Taking into account realistic sensor network constraints that only neighboring network connectivity is known to a node, we develop an approximate but effective solution, ENERGY . We evaluate the performance of ENERGY, which performs well even in low-density networks and for queries requesting from data sources at a distance. Yingwen Chen 0001, Hong Va Leong, Ming Xu 0002, Jiannong Cao 0001, Keith C. C. Chan, Alvin Chan Toong Shoon |
MDM | 1 |
| 2006 | An Anti-void Geographic Routing Algorithm for Wireless Sensor Networks
Ming Xu 0002, Yingwen Chen 0001, Wanrong Yu |
MSN | 2 |
| 2004 | A Resource Reservation Protocol for Mobile Cellular Networks
Ming Xu 0002, Zhijiao Zhang, Yingwen Chen 0001 |
ISPA | 3 |