VLDB 2026 Research / reviewers in the wild / expert
Wenda Tang
dblp:188/4498
· DBLP profile ↗
23ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0001-6684-4642ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 7 first-author · 11 since 2021Computer networks · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FractalGPU: Fair and Elastic GPU Sharing for General-Purpose Computing
Kaicheng Guo, Lingyun Yang, Wenda Tang, Pengwei Du, Qian Da, Zhengwei Qi |
ICDCS | 5 |
| 2026 | Improving the Serverless Function Cache Efficiency With FlameabstractFunction caching is one of the fundamental techniques in FaaS platforms to alleviate coldstart overhead. However, as cache instances consume significant cloud resources (e.g., memory), it is challenging to balance function performance and cache cost. Current systems use simple and rudderless cache polices with a “local cache control” design, which ignores function characteristics such as workload skewness from hot functions and results in either cache contentions or cache resource waste.In this paper, inspired by software-defined networks, we proposeFlame, an efficient cache system to manage cached functions with hotspot-aware instance scheduling and cache allocation. It consists of a two-layer design. Firstly, by decoupling the cache control plane from worker nodes and introducing a centralized cache controller,Flamecan schedule functions from a global view of the cluster’s status, thereby reducing inter-node workload skew. Second,Flamedivides the prior monolithic cache pool within each node into multiple partitions and dynamically assigns them to different hot functions, thereby further mitigating intra-node cache contention. Experimental results from realworld workloads show thatFlamecan reduce cache resource usage by 36% on average while improving function performance by nearly 7× compared to the state-of-the-art method. Wenda Tang, Laiping Zhao, Keqiu Li, Jie Wu 0001 |
IEEE Trans. Computers | 2 |
| 2025 | Metis: A Non-Clairvoyant, Workflow-Aware OS Scheduler for Serverless ApplicationsabstractServerless workflows introduce unique challenges for modern cluster schedulers, as they consist of highly concurrent and ephemeral functions with unpredictable execution patterns. Through analysis of workloads derived from production serverless trace characteristics, we observe that existing OS-level schedulers, such as Linux CFS, lack workflow-level awareness and make scheduling decisions solely at the function level, which can result in bottlenecks within workflows and prolonged Workflow Completion Times (WCTs). We present Metis, anon-clairvoyant, workflow-aware OS scheduler designed specifically for serverless workflow workloads, which aims to reduce average WCTs by treating workflows as first-class scheduling entities. Metis implements Workflow-Aware Least-Attained Service (WLAS), a non-clairvoyant scheduling algorithm that leverages workflow-level virtual clocks and critical path estimation to reduce WCTs and ensure fairness. By utilizing eBPF to hook into existing OS primitives, Metis achieves practical deployment with minimal kernel modifications. Extensive synthetic trace-driven simulations demonstrate that Metis reduces average WCTs by 31.3% and the 95th percentile by 14.3% compared to state-of-the-art function-centric scheduling approaches. Real-system experiments across diverse workflow patterns show improvements ranging from 47.2% to 58.2% in average end-to-end latency and the 99th percentile by up to 73.9% over baseline schedulers. Wenda Tang, Jie Wu 0001 |
SoCC | 1 |
| 2025 | Hierarchical-Caching-Driven Distributed Architecture for Accelerating Model Training
Jinbin Hu 0001, Wenda Tang, Jin Wang 0001 |
ICA3PP (8) | 2 |
| 2025 | Origami: Efficient ML-Driven Metadata Load Balancing for Distributed File SystemsabstractModern distributed file systems (DFSs) rely on metadata server clusters to manage large-scale files and achieve scalability. However, the hierarchical namespace structure and dynamic user workloads pose severe challenges for efficient metadata partitioning and load balancing. Existing approaches primarily focus on identifying and redistributing hot metadata to address imbalances. While these load-balancing strategies offer potential benefits, they often reduce metadata locality, ultimately failing to improve the end-to-end job completion time—a key metric prioritized by users. Although recent research reveals that learning-based approaches are effective in predicting hotspots, they have been shown to be less effective in improving metadata performance. We revisit metadata load balancing strategies and propose a learning-based metadata load balance framework Origami, which focuses on minimizing end-to-end job completion time rather than equalizing loads. Origami first decomposes the overhead of metadata operations and assesses the impact of migration decisions on user requests, allowing us to compute the benefits of migration decisions for job completion time when future requests are known. Subsequently, Origami propose the Meta-OPT algorithm to determine near-optimal migration decisions. Finally, we implemented OrigamiFS, on which we collected statistical data to train and validate ML-models capable of predicting migration benefits. By predicting the benefits of migration decisions and employing Meta-OPT to quickly explore nearly optimal migration decisions, Origami makes a better trade-off between load balancing and namespace locality. Our evaluation shows that compared to state-of-the-art methods, Origami increases aggregated metadata throughput by 1.12-2.51 × across three real-world workloads, and enhances end-to-end throughput by 1.11-2.02 ×. Yiduo Wang 0002, Wenda Tang, Linghang Meng, Liang Li 0016, Jie Wu 0001 |
ICPP | 2 |
| 2025 | Leave No One Behind: Fair and Efficient Tiered Memory Management for Multi-ApplicationsabstractThe emergence of byte-addressable memory technologies, such as CXL-attached memory, has catalyzed extensive research into tiered memory management. Existing tiering solutions optimize system-wide performance by migrating frequently accessed (“hot”) data to fast-tier memory via page migration, which serves as the de facto mechanism in modern OS. However, these strategies often fail in the multi-tenant environment, where diverse workloads interfere with each other. For example, latency-critical workloads co-located with throughput-oriented ones may face the “cold page dilemma,” where critical pages are misclassified as “cold” and migrated to slower tiers, leading to significant performance degradation. Moreover, current methods often assume negligible migration overhead, which becomes problematic in multi-core systems handling write-intensive workloads that incur substantial costs. This paper proposes Vulcan, a workload-aware tiered memory management framework that targets fair and efficient tiering in multi-tenant environments. Vulcan introduces four key innovations: (1) workload-dependent migration mechanism, which decouples page migration from the OS kernel to enhance operational flexibility for multi-workloads; (2) QoS-aware fair resource partitioning, which dynamically optimizes fast memory distribution through per-workload fast tier hit ratios and fairness-oriented allocation policies; (3) per-thread page table replication, which minimizes TLB coherence overhead during migration; and (4) biased page migration policy, which optimizes efficiency by considering both access characteristics (read-intensive vs. write-intensive) and thread-level page ownership (private vs. shared). We evaluated Vulcan using multiple representative cloud applications with realistic working sets in co-location scenarios. Vulcan improves performance by 12.4% on average and achieves a 75.3% improvement in fairness compared to existing state-of-the-art memory tiering solutions. Wenda Tang, Yiduo Wang 0002, Yanwen Wang 0002, Jie Wu 0001 |
ICPP | 1 |
| 2025 | Nip it in the Bud: Unsupervised KPI Incipient Fault Detection via Dynamic Latent Feature EnsemblingabstractMonitoring Key Performance Indicator (KPI) trends in a timely manner enables early detection of performance degradation, which is critical to maintaining cloud service reliability. However, incipient anomalies that subtly precede KPI degradation are notoriously difficult to detect due to interference from noise and the high-dimensional, correlated nature of monitoring data. In modern production environments, KPIs are recorded as massive multivariate time series (MTS), where intricate temporal and inter-metric dependencies further obscure early signals of system instability. Unfortunately, existing fault detection methods either rely on oversimplified statistical assumptions that ignore these dependencies or depend on largescale supervised training, which is rarely feasible in dynamic, label-scarce settings. We introduce Heimdallr, an unsupervised detection framework tailored to uncover early-stage anomalies that causally affect KPIs. The core design of Heimdallr is a KPI-oriented monitoring and attribution mechanism that models and partitions the latent space based on KPI behavior. This design enables not only the early detection of anomalies but also causal attribution to specific KPI shifts, thereby enhancing system observability and resilience. Heimdallr is built upon two key innovations: Dynamic-inner Related Component Analysis (DiRCA), a latent structure modeling technique that captures dynamic temporal dependencies across metrics, enabling interpretable representations of underlying system behavior; KPI Feature Ensemble Monitoring Network (KFEMNet), a three-layer hierarchical architecture that, following DiRCA's decomposition, extracts fine-grained deviations and detects incipient anomalies with high sensitivity. Extensive experiments on both synthetic and real-world datasets convincingly demonstrate that Heimdallr consistently outperforms existing state-of-the-art methods, achieving higher early detection accuracy and lower false alarm rates, while maintaining low overhead and high interpretability suitable for production deployment. Yanwen Wang 0002, Wenda Tang, Jie Wu 0001 |
SRDS | 2 |
| 2024 | PheScale: Leveraging Transformer Models for Proactive VM Auto-scaling
Yanqin Zheng, Changjian Wang 0001, Wenda Tang, Tianxiang Ai |
ADMA (1) | 5 |
| 2024 | Yggdrasil: Reducing Network I/O Tax with (CXL-Based) Distributed Shared MemoryabstractIn communication-intensive applications that run on hosts with high-speed network hardware, a common challenge arises from the significant burden placed on the native socket system within the OS. Researchers have devoted considerable effort to optimizing the kernel networking stack and moving the TCP/IP stack to user-space. In this paper, we describe a novel socket replacement solution, Yggdrasil, a CXL-based user-space high-performance socket system. Yggdrasil is fully compatible with Linux socket, making it a drop-in replacement for existing applications without the need for code modifications. In order to optimize performance, Yggdrasil employs CXL-based distributed shared memory (DSM) for inter-host communication whenever it is available. In cases where DSM is not accessible, Yggdrasil transparently switches back to Linux socket for communication. A key element in achieving isolation in Yggdrasil involves a trusted user-space monitoring daemon responsible for managing control plane operations like connection setup and access control. Within the data plane of Yggdrasil, a peer-to-peer model is adopted for communication between processes. To bridge the semantic gap between socket and DSM, we exploit several techniques to ensure compatibility and performance, including (1) transparent dynamic fast/slow data path navigation, (2) decentralized CXL memory management, (3) lock-free queue based QoS-aware dynamic data polling, and (4) semantics-aware memory page migration. By evaluating Yggdrasil on both emulated and real CXL hardware, we show that Yggdrasil outperforms Linux socket in Memcached throughput by 8.2 × and reduces latency by 24 ∼ 320 × in a micro benchmark across different message sizes. Wenda Tang, Tianxiang Ai |
ICPP | 1 |
| 2024 | PheCon: Fine-Grained VM Consolidation with Nimble Resource Defragmentation in Public Cloud PlatformsabstractResource fragmentation is inevitable due to the unknown and fluctuating sequence of requests to create or delete virtual machines (VMs) in cloud platforms. VM consolidations can be effective in addressing resource fragmentation issues to ensure better utilization of multi-dimension resources. However, current VM consolidation solutions primarily focus on optimizing the utilization of resources at the level of physical machines (PMs) and often require migrating all VMs of the target PM to other PMs, which can result in unnecessary and inappropriate VM migrations. In this paper, we propose a novel nimble fine-grained VM consolidation algorithm, PheCon, which focuses on fine-grained consolidation by considering VM flavors as the unit of consolidation. It attempts to aggregate the resource fragments from PMs and gather resources for additional VM allocations with specific flavors. In contrast to the state-of-the-art method, PheCon does not attempt to release PMs, but instead focuses on utilizing resource fragmentation to increase the number of additional VM allocations. Besides, to further reduce the resource fragmentation on PMs, PheCon leverages a hierarchical swapping method that enables placing VMs into PMs with insufficient free resources by swapping a part of smaller VMs to other PMs. In addition, to improve generalizability, PheCon takes NUMA systems into consideration, determining both the target PM and NUMA nodes for VM consolidation. Comprehensive evaluation using simulation and our production cloud datasets shows that PheCon could reduce the number of VM migrations by 35% on average compared to the state-of-the-art method. Jiazhen Zhu, Wenda Tang, Nan Gong, Tianxiang Ai |
ICPP | 2 |
| 2023 | Smart decision for device selection in D2D-assisted multi-path video transmission networkabstractAbstract Watching the live video on a bus/train/tram has become an important pattern for people to enjoy their travel time. Due to the complex environment changes and obstacles caused by the rapid movement, the network state of the user device is often unstable. A large number of interruptions and resolution reductions seriously affect the user's watching experience. Through direct transmission between devices, Device‐to‐Device (D2D) can effectively improve video quality. Accurate decision for helper device selection is the key to optimize video streaming services based on D2D technology. However, the existing methods through D2D rarely consider the states and the intention of the helper devices when selecting cooperative devices, such as the watching status, the state of charge, the state of cache, and the connection quality. This is highly probable to reduce the efficiency of D2D transmission and the cooperation enthusiasm of helper devices. In addition, the existing methods seldom consider the multiple network interfaces in the device. In view of these challenges, we propose a dynamic smart decision method for device selection in D2D‐assisted multi‐path video transmission network (named DS‐DAMP). Specifically, this method takes the current states and the resource constraints into consideration to select the appropriate helper devices for video cooperative transmission, so as to improve the cooperation satisfaction of helper devices and optimize the video quality of the requesting device. Besides, we make full use of multiple network interfaces to realize multi‐path parallel transmission between devices. A large number of experiments have proved the good performance of our method in terms of video quality and network utility. Xuan Zhao 0005, Bowen Liu 0002, Xutong Jiang, Wenda Tang, Wan-Chun Dou |
Expert Syst. J. Knowl. Eng. | 4 |
| 2022 | Demeter: QoS-aware CPU scheduling to reduce power consumption of multiple black-box workloadsabstractEnergy consumption in cloud data centers has become an increasingly important contributor to greenhouse gas emissions and operation costs. To reduce energy-related costs and improve environmental sustainability, most modern data centers consolidate Virtual Machine (VM) workloads belonging to different application classes, some being latency-critical (LC) and others being more tolerant to performance changes, known as best-effort (BE). However, in public cloud scenarios, the real classes of applications are often opaque to data center operators. The heterogeneous applications from different cloud tenants are usually consolidated onto the same hosts to improve energy efficiency, but it is not trivial to guarantee decent performance isolation among colocated workloads. We tackle the above challenges by introducing Demeter, a QoS-aware power management controller for heterogeneous black-box workloads in public clouds. Demeter is designed to work without offline profiling or prior knowledge about black-box workloads. Through the correlation analysis between network throughput and CPU resource utilization, Demeter automatically classifies black-box workloads as either LC or BE. By provisioning differentiated CPU management strategies (including dynamic core allocation and frequency scaling) to LC and BE workloads, Demeter achieves considerable power savings together with a minimum impact on the performance of all workloads. We discuss the design and implementation of Demeter in this work, and conduct extensive experimental evaluations to reveal its effectiveness. Our results show that Demeter not only meets the performance demand of all workloads, but also responds quickly to dynamic load changes in our cloud environment. In addition, Demeter saves an average of 10.6% power consumption than state of the art mechanisms. Wenda Tang, Yutao Ke, Senbo Fu, Hongliang Jiang |
SoCC | 1 |
| 2022 | Themis: Fair Memory Subsystem Resource Sharing with Differentiated QoS in Public CloudsabstractTo reduce the increasing cost of building and operating cloud data centers, cloud providers are seeking various mechanisms to achieve higher resource effectiveness. For example, cloud operators are leveraging dynamic resource management techniques to consolidate a higher density of application workloads into commodity physical servers to maximize server resource utilization. However, higher workload density is a major source of performance interference problems in multi-tenant clouds. Existing performance isolation techniques such as dedicated CPU cores for specific workloads are not enough as there are still common resource (e.g., last-level cache and memory bandwidth in memory subsystem) on the processor that are shared among all CPUs on the same NUMA node. While prior work has proposed a variety of resource partitioning techniques, it still remains unexplored to characterize the impact of memory subsystem resource partitioning for the consolidated workloads with different priorities and investigate software support to dynamically manage memory subsystem resource sharing in a real-time manner. To bridge the gap, we propose Themis, a feedback-based controller that enables a priority-aware and fairness-aware memory subsystem resource management strategy to guarantee the performance of high-priority workloads while maintaining fairness across all colocated workloads in high-density clouds. Themis is evaluated with multiple typical cloud applications in our data center environment. The results show that Themis improves the performance of various workloads by up to 3.15%, and fairness by more than 70% in memory subsystem resource allocation compared to existing state-of-the-art work. Wenda Tang, Senbo Fu, Yutao Ke |
ICPP | 1 |
| 2021 | A WiFi-aware method for mobile data offloading with deadline constraintsabstractSummary With an increasing number of public WiFi hotspots have been constructed in metropolitan areas, citizens can leverage these WiFi hotspots to surf the Internet with their smart devices almost everywhere. Compared to cellular networks, most people prefer using WiFi because this can not only save money by reducing the cellular data traffic but also save energy to prolong the battery lifetime of mobile devices. Mobile offloading techniques give the opportunity for smart devices to offload cellular data traffic to WiFi networks whenever they are connected WiFi networks. Although offloading to WiFi networks can cause communication delay, many delay‐tolerant applications still prefer doing so in order to save the cellular data traffic. Existing offloading approaches only take the application delay tolerance into consideration and therefore easy to induce some tasks missing deadlines. In this paper, we propose a novel WiFi‐aware method for mobile data offloading through WiFi networks with deadline constraints. Specifically, it takes the WiFi distribution in city areas into consideration while making offloading decision. The simulation results on real‐world data have shown that our method improves the percentage of tasks that can meet their deadlines, and reduces the average task completion time by comparing with other outstanding methods. Wenda Tang, Chaobing Wu, Lianyong Qi, Xuyun Zhang, Xiaolong Xu 0001, Wan-Chun Dou |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Blockchain-based Mobility-aware Offloading mechanism for Fog computing services
Wan-Chun Dou, Wenda Tang, Bowen Liu 0002, Xiaolong Xu 0001, Qiang Ni |
Comput. Commun. | 2 |
| 2020 | An insurance theory based optimal cyber-insurance contract against moral hazard
Wan-Chun Dou, Wenda Tang, Xiaotong Wu, Lianyong Qi, Xiaolong Xu 0001, Xuyun Zhang, Chunhua Hu 0001 |
Inf. Sci. | 2 |
| 2020 | TrCMP: A dependable app usage inference design for user behavior analysis through cyber-physical parameters
Xuan Zhao 0005, Md. Zakirul Alam Bhuiyan, Lianyong Qi, Hongli Nie, Wenda Tang, Wan-Chun Dou |
J. Syst. Archit. | 5 |
| 2020 | Data security over wireless transmission for enterprise multimedia security with fountain codes
Hongli Nie, Xutong Jiang, Wenda Tang, Wan-Chun Dou |
Multim. Tools Appl. | 3 |
| 2019 | AnalogMUSIC: A Concurrent Beam Training Scheme for Multiple Users in mmWave SystemsabstractIn this paper, a multi-user concurrent beam training scheme, named AnalogMUSIC, is designed for mmWave systems with a single radio frequency (RF) chain. In AnalogMUSIC, a one-RF chain based multiple signal classification (MUSIC) algorithm is developed for accurate direction estimation, where the training beams are switched in a time-division manner. To pursue a short training time, the training beams are switched at rough directions predefined by a coarse- resolution codebook. We prove that the traditional MUSIC algorithm can be achieved in the frequency domain. For orthogonal frequency division multiple access (OFDMA) systems, the multi-user training signals can be decoupled in the frequency domain. Thus, AnalogMUSIC achieves the concurrent beam training for multiple users by allocating different subcarriers to different users. To further adapt the number of switched training beams to various channel conditions, two coarse-resolution codebooks and a beam mapping mechanism are designed. Extensive simulations show that AnalogMUSIC can effectively improve beam training accuracy and substantially reduce delay overhead, compared with the existing schemes. Bangzhao Zhai, Wenda Tang, Aimin Tang, Mingzeng Dai, Xudong Wang 0001 |
GLOBECOM | 2 |
| 2019 | A heuristic line piloting method to disclose malicious taxicab driver's privacy over GPS big data
Wan-Chun Dou, Wenda Tang, Shui Yu 0001, Kim-Kwang Raymond Choo |
Inf. Sci. | 2 |
| 2019 | An offloading method using decentralized P2P-enabled mobile edge servers in edge computing
Wenda Tang, Xuan Zhao 0005, Wajid Rafique, Lianyong Qi, Wan-Chun Dou, Qiang Ni |
J. Syst. Archit. | 1 |
| 2018 | A Deadline-Aware Coflow Scheduling Approach for Big Data ApplicationsabstractMany datacenters usually process complex jobs such as MapReduce jobs. From a network perspective, most of these jobs trigger multiple parallel data flows, which comprise acoflowgroup semantically. When to schedule the jobs in datacenter or across multiple datacenters, most of current job schedulers have not considered the underlying network traffic load, which is suboptimal for jobs completion times. We present a new deadline-aware coflow scheduling approach called DCS, which takes the underlying network traffic load into consideration while guaranteeing high percentage of coflows that meet their deadlines. DCS aims to alleviate the network congestion in datacenters whose network worload are unbalanced, and it includes two stages for coflow scheduling: Firstly, it generates the task placement proposal by considering the underlying network workload. Secondly, it makes scheduling decision by estimating both task's execution time and transmission waiting time under the previous task placement proposal. The real-world data based simulation results have shown that DCS outperforms all existing solutions on reducing the percentage of coflows that miss their deadlines. Wenda Tang, Duanchao Li, Taigui Huang, Wan-Chun Dou, Shui Yu 0001 |
ICC | 1 |
| 2017 | Model parameters of molecular evolution explain genomic correlationsabstractOne long-standing research focus in evolutionary genomics is trying to resolve how biological variables (expression, essentiality, protein-protein interaction, structural stability, etc.) determine the rate of protein evolution. While these studies have considerably deepened our understanding of molecular evolution, many issues remain unsolved. In this opinion article, after having a brief survey of literatures, we establish relationships between model parameters of molecular evolution and genomic variables, based on which, most-observed genomic correlations and confounds can be explained by model parameter combinations under different conditions, which include the strength of stabilizing selection, mutational variance, expression sufficiency, gene pleiotropy, as well as the effective population size. We suggest that the problem to discern biological variable(s) that may determine the rate of protein evolution can be tackled at two levels. The first level, as discussed here, is to demonstrate how the model of molecular evolution can predict potential genomic correlations under various conditions. And the second level is to estimate genome-wide variations of model parameters (or combinations) that help to identify canonical biological variables that may underlie the rate variation among genes that ranges up to at least three magnitudes. Wenda Tang |
Briefings Bioinform. | 2 |