EDBT 2026 Demo / reviewers in the wild / expert
Wenli Zheng
dblp:147/1003
· DBLP profile ↗
46ranked-venue papers
13as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 7 first-author · 13 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Selective Diffusion Distillation for Real-World High-Scale Image Super-ResolutionabstractHigh-scale image super-resolution (SR) has become increasingly important with the rapid growth of mobile devices and high-resolution displays. However, current SR methods primarily focus on lower scales and generalize poorly to high-scale scenarios due to severe information loss and complex real-world degradations. In this paper, we propose a novel Selective Diffusion Distillation (SDD) framework for real-world high-scale SR, which distills reliable knowledge from a low-scale diffusion teacher to a high-scale student. Specifically, considering severe information loss in high-scale inputs, directly distilling from low-scale models may result in feature misalignment. To address this, we introduce a Degradation-aware Metric Learning (DML) approach to align feature distributions across different degradation levels. In addition, since the diffusion-based teacher may hallucinate artifacts in ambiguous regions, blindly imitating these unreliable outputs can degrade the student’s fidelity. To tackle this, we propose a Region-aware Selective Distillation (RSD) strategy to filter out uncertain predictions and adaptively supervise only on reliable areas. To evaluate the effectiveness of our method, we introduce Real-UltraSR, a new real-world benchmark that contains diverse high-scale LR-HR pairs, including x8, x10, x12, and x14. Extensive experiments demonstrate that our SDD framework achieves state-of-the-art performance across multiple benchmarks. Wenli Zheng, Huiyuan Fu, Xin Wang 0001, Huadong Ma |
AAAI | 1 |
| 2026 | Learning Continuous Degradation for Real-World Arbitrary-Scale Video Super-ResolutionabstractArbitrary-scale video super-resolution (VSR) aims to enhance video resolution at continuous scales and has attracted increasing attention in recent years. However, existing methods typically rely on fixed degradation modes, such as bicubic downsampling, which often fail to handle the complex degradations of real-world videos. Current real-world datasets only cover limited scales (e.g., ×2, ×4) and are insufficient to capture the diverse degradations required for arbitrary-scale VSR. To address this, we present RealArbVSR, the first real-world VSR dataset with both integer and decimal scale factors, providing a wider range of degradation levels. Moreover, to generate continuous degradations beyond the collected scales, we propose the Continuous Degradation Generation Network (CDGN), which synthesizes realistic LR videos with arbitrary degradations. Specifically, we design a Scale-aware Degradation Module (SDM) to adaptively learn scale-specific degradations and an Implicit Filter Module (IFM) that represents spatial-temporal features as a continuous feature domain for arbitrary-scale LR frame generation. Extensive experiments demonstrate that our CDGN trained on RealArbVSR produces high-fidelity LR videos with arbitrary degradations and significantly enhances the performance of VSR models in real-world scenarios. The RealArbVSR dataset and source code will be publicly released for further research. Wenli Zheng, Huiyuan Fu, Chuanming Wang, Enyuan Zhang, Hengming Mao, Heng Zhang 0042, Huadong Ma |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | NeurDORA: Neural-Aided Decentralized Offloading Based on Resource AuctionabstractMobile Edge Computing (MEC) is a key solution to overcome vehicles' limited on-board computation capabilities when dealing with data-intensive tasks.Existing studies usually assume guaranteed resource availability at base stations (BSes), or restrict a user to a single BS, overlooking the resource-competition failures due to multiple BS accesses.This paper proposes a decentralized offloading algorithm (NeurDORA) for multi-user-multi-BS (MUMB) scenarios, ensuring efficient and fair allocation of BS resources.With Neur-DORA, each vehicle predicts its chance of successfully securing communication and computation resources at each candidate BS, and then competes for offloading opportunities in an iterative auction process that converges to Nash Equilibrium.Simulation results show that NeurDORA reduces offloading failures by 35-53% and achieves a 12-26% reduction in average offloading latency. Rentian Wei, Wenli Zheng |
CF | 2 |
| 2025 | EncloPC: Power Control for Server Enclosures with Heterogeneous Workloads at Power Peaks
Wenli Zheng |
ICA3PP (1) | 1 |
| 2025 | Exploiting Virtual Energy Storage to Improve Power Oversubscription in Colocation Data CentersabstractFor a colocation data center with power oversubscription, its power capacity can be inadequate when the power loads of all the tenants peak simultaneously, due to lack of collaboration. In this paper, we point out that such simultaneous power peaks can be common and regular challenges rather than occasional events. To address the problem, we propose to exploit the virtual energy storage device (vESD) as an incentive mechanism in the pricing of rent rates and energy costs, to guide the power management of their tenants and eliminate the regular simultaneous power peaks. In addition, we design a dynamic allocation and pricing strategy to support the tenants to flexibly rent vESD for their respective cost optimization, and the fairness is guaranteed with Nash Equilibrium. The experiments based on real-world data center traces show that using vESD can improve the utilization of power infrastructures and save the provisioning costs by 32%. Wenli Zheng |
ICCCN | 1 |
| 2025 | Recommendation-Expert Framework for Fast and Adaptive Scheduling in Computing Power NetworkabstractEfficient task scheduling in heterogeneous Computing Power Networks (CPNs) is challenging due to the diversity of resource capabilities and task characteristics. Existing reinforcement learning (RL)-based methods suffer from large action spaces and limited generalization, while heuristic or optimizationbased approaches lack adaptability. In this paper, we propose the Recommendation-Expert (RE) scheduling framework, which combines a recommendation network with multiple lightweight expert networks. The recommendation network encodes the system state and selects the most appropriate expert to make a scheduling decision, enabling fast and context-aware task placement. We evaluate RE on two simulation platforms using real-world traces. Experimental results show that RE respectively reduces energy consumption and response time by up to$\mathbf{4 2. 0 6 \%}$and 37.95 %, and improves task throughput by up to 93.63 % compared to state-of-the-art baselines, or reduces scheduling latency by as much as 94.04 % with comparable task execution performance. Wenli Zheng |
ICCD | 2 |
| 2025 | Towards Fine-Grained Scalability for Stateful Stream Processing SystemsabstractDynamic scaling is critical to stream processing engines, as their long-running nature demands adaptive resource management. Existing scaling approaches easily cause performance degradation due to coarse-grained synchronization and inefficient state migration, resulting in system halt or high processing latency. In this paper, we propose DRRS, an on-the-fly scaling method that reduces performance overhead at the system level with three key innovations: (i) fine-grained scaling signals coupled with a re-routing mechanism that significantly mitigates propagation delay, (ii) a sophisticated record-scheduling mechanism that substantially reduces processing suspension, and (iii) subscale division, a mechanism that partitions migrating states into independent subsets, thereby reducing dependency-related overhead to enable finer-grained control and better runtime adaptability during scaling. DRRS is implemented on Apache Flink and, when compared to state-of-the-art approaches, reduces peak and average latencies by up to 81.1% and 95.5% respectively, while achieving a 72%-86% reduction in scaling duration, without disruption in non-scaling periods. Yunfan Qing, Wenli Zheng |
ICDE | 2 |
| 2025 | EvRAW: Event-guided Structural and Color Modeling for RAW-to-sRGB Image ReconstructionabstractEvent-based image reconstruction has achieved remarkable progress, benefiting from the high temporal resolution and high dynamic range of event cameras. However, most event-based methods focus on enhancing sRGB image quality, neglecting the potential of leveraging event data for RAW-to-sRGB conversion. Due to the limitations of camera sensors, images processed through standard ISP pipelines often suffer from motion blur and color distortion in dynamic scenes. In contrast, RAW images preserve uncompressed scene information, integrating event signals at this stage enables finer texture recovery and more accurate color correction. To tackle these challenges, we propose EvRAW, a novel event-assisted RAW-to-sRGB image reconstruction network that integrates event signals to promote high-fidelity sRGB image reconstruction. Specifically, we introduce a Motion-guided Structural Enhancement (MSE) module that extracts motion patterns from event streams and aggregates dynamic features to restore fine textures. Additionally, we propose an Adaptive Color Correction (ACC) module that performs region-wise gamma correction and channel-wise color decoding to enhance color fidelity under complex lighting conditions. To evaluate performance in challenging real-world scenarios, we collect a pixel-aligned RAW-Event dataset specifically for this task. Extensive experiments demonstrate that EvRAW achieves state-of-the-art performance in RAW-to-sRGB reconstruction on both synthetic and real-world datasets. Wenli Zheng, Huiyuan Fu, Xicong Wang, Hao Kang, Chuanming Wang, Jin Liu 0024, Heng Zhang 0042, Huadong Ma |
ACM Multimedia | 1 |
| 2025 | CLMTR: a generic framework for contrastive multi-modal trajectory representation learning
Anqi Liang, Bin Yao 0002, Jiong Xie, Wenli Zheng, Yanyan Shen, Qiqi Ge |
GeoInformatica | 4 |
| 2024 | Dynamic heterogeneous attributed network embeddingabstractInformation networks generally exhibit three characteristics, namely dynamicity, heterogeneity, and node attribute diversity. However, most existing network embedding approaches only consider two of the three when embedding each node into low-dimensional space. Adding to such an existing approach a technique of processing the remaining characteristic can easily cause incompatibility. One solution to process the three characteristics together is to treat the dynamic heterogeneous attributed network (DHAN) as a temporal sequence of heterogeneous attributed network (HAN) snapshots. For example, existing graph convolutional networks (GCNs)-based DHAN embedding approaches embed the HAN snapshots to get static representations offline, and then dynamically capture temporal dependencies between adjacent snapshots online to maintain fresh representations of the DHAN. However, those approaches encounter the convergence problem when stacking multiple convolutional layers to capture more topological information. Some other existing approaches dynamically update the representations of HAN snapshots online, neglecting the efficiency requirement of online scenarios and the temporal dependencies between snapshots. To address the two issues, we propose a new framework called Dynamic Heterogeneous Attributed Network Embedding (DHANE), consisting of a static model MGAT and a dynamic model NICE. MGAT captures more topological information while maintaining GCN convergence by performing metagraph-based attention in each convolutional layer. NICE preserves network freshness while reducing the computational load of the update by only examining network changes and updating their embedding representations. Extensive experiments show that DHANE achieves up to 27× speedup and 9.1-26.4% higher accuracy on several real dynamic heterogeneous attributed networks for online classification. Hongbo Li 0003, Wenli Zheng, Feilong Tang 0001, Yitong Song 0001, Bin Yao 0002, Yanmin Zhu 0006 |
Inf. Sci. | 2 |
| 2024 | Bayesian-Driven Automated Scaling in Stream Computing With Multiple QoS TargetsabstractStream processing systems commonly work with auto-scaling to ensure resource efficiency and quality of service (QoS). Existing auto-scaling solutions lack accuracy in resource allocation because they rely on static QoS-resource models that fail to account for high workload variability and use indirect metrics with much distractive information. Moreover, different types of QoS metrics present different characteristics and thus need individual auto-scaling methods. In this paper, we propose a versatile auto-scaling solution for operator-level parallelism configuration, called AuTraScale+, to meet the throughput, processing-time latency, and event-time latency targets. AuTraScale+ follows the Bayesian optimization framework to make scaling decisions. First, it uses the Gaussian process model to eliminate the negative influence of uncertain factors on the performance model accuracy. Second, it leverages the expected improvement-based (EI-based) acquisition function to search and recommend the optimal configuration quickly. Besides, to make a more accurate scaling decision when the new model is not ready, AuTraScale+ proposes a transfer learning algorithm to estimate the benefits of all configurations at a new rate based on existing models and then recommend the optimal one. We implement and evaluate AuTraScale+ on the Flink platform. The experimental results on three representative workloads demonstrate that compared with the state-of-the-art methods, AuTraScale+ can reduce 66.6% and 36.7% resource consumption, respectively, in the scale-down and scale-up scenarios while achieving their throughput and processing-time latency targets. Compared with other methods of optimizing event-time latency, AuTraScale+ saves 26.9% of resources on average. Liang Zhang 0027, Wenli Zheng, Kuangyu Zheng, Hongzi Zhu, Chao Li 0009, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Kronos: towards bus contention-aware job scheduling in warehouse scale computers
Shang Zhao 0003, Quan Chen 0002, Shanpei Chen, Tao Ma 0006, Yong Yang 0013, Wenli Zheng, Minyi Guo |
Frontiers Comput. Sci. | 8 |
| 2023 | Few-shot time-series anomaly detection with unsupervised domain adaptation
Hongbo Li 0003, Wenli Zheng, Feilong Tang 0001, Yanmin Zhu 0006, Jielong Huang |
Inf. Sci. | 2 |
| 2022 | FaaSFlow: enable efficient workflow execution for function-as-a-serviceabstractServerless computing (Function-as-a-Service) provides fine-grain resource sharing by running functions (or Lambdas) in containers. Data-dependent functions are required to be invoked following a pre-defined logic, which is known as serverless workflows. However, our investigation shows that the traditional master-worker based workflow execution architecture performs poorly in serverless context. One significant overhead results from the master-side workflow schedule pattern, with which the functions are triggered in the master node and assigned to worker nodes for execution. Besides, the data movement between workers also reduces the throughput. Zijun Li 0001, Yushi Liu 0003, Linsong Guo, Quan Chen 0002, Jiagan Cheng, Wenli Zheng, Minyi Guo |
ASPLOS | 6 |
| 2022 | LoADPart: Load-Aware Dynamic Partition of Deep Neural Networks for Edge OffloadingabstractThe emerging edge computing technique provides support for the computation tasks that are delay-sensitive and compute-intensive, such as deep neural network inference, by offloading them from a user-end device to an edge server for fast execution. The increasing offloaded tasks on an edge server are gradually facing the contention of both the network and computation resources. The existing offloading approaches often partition the deep neural network at a place where the amount of data transmission is small to save network resource, but rarely consider the problem caused by computation resource shortage on the edge server. In this paper, we design LoADPart, a deep neural network offloading system. LoADPart can dynamically and jointly analyze both the available network bandwidth and the computation load of the edge server, and make proper decisions of deep neural network partition with a light-weighted algorithm, to minimize the end-to-end inference latency. We implement LoADPart for MindSpore, a deep learning framework supporting edge AI, and compare it with state-of-the-art solutions in the experiments on 6 deep neural networks. The results show that under the variation of server computation load, LoADPart can reduce the end-to-end latency by 14.2% on average and up to 32.3% in some specific cases. Hongzhou Liu, Wenli Zheng, Li Li 0064, Minyi Guo |
ICDCS | 2 |
| 2022 | CSC: Collaborative System Configuration for I/O-Intensive Applications in Multi-Tenant CloudsabstractI/O-intensive applications are important workloads of public clouds. Multiple cloud applications co-run on the same physical machine in different virtual machines (VMs), and the shared resources (e.g., disk bandwidth) are often isolated for fairness. Our investigation shows that the performance of an I/O-intensive application is impacted by both disk bandwidth allocation and the page cache settings in the guest operating system. However, none of prior work considers adjusting the page cache settings for better performance, when the disk bandwidth allocation is adjusted. We therefore propose CSC, a system that collaboratively identifies the appropriate disk bandwidth allocation and page cache settings in the guest operating system of each VM. CSC aims to improve the system-wide I/O throughput of the physical machine, while also improve the I/O throughput of each individual I/O-intensive application in VMs. CSC comprises an online disk bandwidth allocator and an adaptive dirty page setting optimizer. The bandwidth allocator monitors the disk bandwidth utilization and re-allocates some bandwidth from free VMs to busy VMs periodically. After the re-allocation, the opti-mizer identifies the appropriate dirty page settings in the guest operating system of the VMs using Bayesian Optimization. The experimental results show that CSC improves the performance of I/O-intensive applications by 9.5 % on average (up to 17.29 %) when 5 VMs are co-located while fairness is guaranteed. Haowei Huang, Pu Pang, Quan Chen 0002, Jieru Zhao, Wenli Zheng, Minyi Guo |
IPDPS | 5 |
| 2022 | Performance optimization for cloud computing systems in the microservice era: state-of-the-art and research opportunities
Xiaofeng Hou, Lu Zhang 0049, Chao Li 0009, Wenli Zheng, Minyi Guo |
Frontiers Comput. Sci. | 5 |
| 2022 | Reliability and Incentive of Performance Assessment for Decentralized Clouds
Jiuchen Shi, Xiaoqing Cai, Wenli Zheng, Quan Chen 0002, Deze Zeng, Tatsuhiro Tsuchiya, Minyi Guo |
J. Comput. Sci. Technol. | 3 |
| 2022 | Integrated Power Anomaly Defense: Towards Oversubscription-Safe Data CentersabstractEnergy storage devices (e.g., batteries) are critical components for high-availability data center infrastructure today. Without resilient energy management of these devices, existing power-hungry data centers are largely unguarded targets for cyber criminals. Particularly for some of today's scale-out data centers, power infrastructure oversubscription unavoidably taxes the data center's backup energy resources (i.e., UPS), leaving very little room for dealing with power emergency. As a result, an attacker could manipulate the computing system to generate peak power demand and disrupt power-constrained server racks. This article aims at protecting data centers from malicious loads that seek to drain precious energy backup, overload server racks and compromise workload performance. We term such load as Elusive Power Peak (EPP) and demonstrate its basic three-phase attacking model. To defend against EPP, we propose IPAD, a remediation solution build on integrated software and hardware mechanisms. IPAD not only increases the attacking cost considerably by hiding vulnerable server racks from visible power peaks, but also strengthens the last line of defense against hidden power spikes with fine-grained power control strategy. We show that IPAD can effectively raise the bar of power-related attack, with reasonable design overhead. Xiaofeng Hou, Chao Li 0009, Jinghang Yang, Wenli Zheng, Xiaoyao Liang, Minyi Guo |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | Exploiting big.LITTLE Batteries for Software Defined Management on Mobile DevicesabstractBattery service time is a critical constraint on the availability and functionality of mobile devices. Equipping larger batteries may mitigate such deficiency yet it raises the challenges on thermal limit and physical size. Changing the battery chemistry is another solution, which however usually benefits the energy efficiency of only part of the applications, depending on their software behaviors. To address the challenges of battery energy efficiency and heat dissipation in limited physical space, we propose CAPMAN, a management framework that jointly optimizes thecooling andactivepowermanagement in a smartphone, a typical mobile device, equipped with a hybrid battery pack. We establish the framework with three components. First, we abstract the correlation among the batteries, devices and software into a finite state machine model, whose state transitions can be triggered by actions like system calls and user activities. Second, we propose a battery scheduling algorithm that determines the more suitable battery for cooling/active power use, with respect to the dynamic software behaviors and their impact on the hardware states, based on a Markov decision process (MDP). Third, we design a facility for joint cooling and active power management by coordinating TECs and batteries. With the three major designs, CAPMAN realizes software defined management that schedules heterogeneous batteries and TEC cooling in a timely manner. In addition, CAPMAN provides an online algorithm with a proved O(1.05)-competitiveness performance. With a pair of big.LITTLE batteries, we prototype CAPMAN on multiple popular smartphones and a PYNQ development board. The evaluation with real-world workloads shows that compared to the current mainstream, CAPMAN can achieve 114 percent longer battery service time under skewed loads; compared to the state-of-the-practice baselines, CAPMAN shows 55 percent performance gain and 53 percent less energy use on average. Those results approve that big.LITTLE batteries with sophisticated software defined management is an effective way to prolong the battery service times on mobile devices. Zichen Xu 0001, Wenli Zheng, Yuhao Wang 0001, Minyi Guo |
IEEE Trans. Mob. Comput. | 3 |
| 2021 | CHARM: Collaborative Host and Accelerator Resource Management for GPU DatacentersabstractEmerging latency-critical (LC) services often have both CPU and GPU stages (e.g. DNN-assisted services) and require short response latency. Co-locating best-effort (BE) applications on the both CPU side and GPU side with the LC service improves resource utilization. However, resource contention often results in the QoS violation of LC services. We therefore present CHARM, a collaborative host-accelerator resource management system. CHARM ensures the required QoS target of DNN-assisted LC services, while maximizing the resource utilization of both the host and accelerator. CHARM is comprised of a BE-aware QoS target allocator, a unified heterogeneous resource manager, and a collaborative accelerator-side QoS compensator. The QoS target allocator determines the time limit of an LC service running on the host side and the accelerator side. The resource manager allocates the shared resources on both host side and accelerator side. The QoS compensator allocates more resources to the LC service to speed up its execution, if it runs slower than expected. Experimental results on an Nvidia GPU RTX 2080Ti show that CHARM improves the resource utilization by 43.2%, while ensuring the required QoS target compared with state-of-the-art solutions. Wei Zhang 0149, Kaihua Fu, Ningxin Zheng, Quan Chen 0002, Chao Li 0009, Wenli Zheng, Minyi Guo |
ICCD | 6 |
| 2021 | A Fast Domain Adaptation Network for Image Super-Resolution
Wenli Zheng, Fenghai Li |
ICIG (3) | 3 |
| 2021 | FIFL: A Fair Incentive Mechanism for Federated LearningabstractFederated learning is a novel machine learning framework that enables multiple devices to collaboratively train high-performance models while preserving data privacy. Federated learning is a kind of crowdsourcing computing, where a task publisher shares profit with workers to utilize their data and computing resources. Intuitively, devices have no interest to participate in training without rewards that match their expended resources. In addition, guarding against malicious workers is also essential because they may upload meaningless updates to get undeserving rewards or damage the global model. In order to effectively solve these problems, we propose FIFL, a fair incentive mechanism for federated learning. FIFL rewards workers fairly to attract reliable and efficient ones while punishing and eliminating the malicious ones based on a dynamic real-time worker assessment mechanism. We evaluate the effectiveness of FIFL through theoretical analysis and comprehensive experiments. The evaluation results show that FIFL fairly distributes rewards according to workers’ behaviour and quality. FIFL increases the system revenue by 0.2% to 3.4% in reliable federations compared with baselines. In the unreliable scenario containing attackers which destroy the model’s performance, the system revenue of FIFL outperforms the baselines by more than 46.7%. Liang Gao 0001, Li Li 0064, Yingwen Chen 0001, Wenli Zheng, Cheng-Zhong Xu 0001, Ming Xu 0002 |
ICPP | 4 |
| 2021 | Dubhe: Towards Data Unbiasedness with Homomorphic Encryption in Federated Learning Client SelectionabstractFederated learning (FL) is a distributed machine learning paradigm that allows clients to collaboratively train a model over their own local data. FL promises the privacy of clients and its security can be strengthened by cryptographic methods such as additively homomorphic encryption (HE). However, the efficiency of FL could seriously suffer from the statistical heterogeneity in both the data distribution discrepancy among clients and the global distribution skewness. We mathematically demonstrate the cause of performance degradation in FL and examine the performance of FL over various datasets. To tackle the statistical heterogeneity problem, we propose a pluggable system-level client selection method named Dubhe, which allows clients to proactively participate in training, meanwhile preserving their privacy with the assistance of HE. Experimental results show that Dubhe is comparable with the optimal greedy method on the classification accuracy, with negligible encryption and communication overhead. Shulai Zhang, Quan Chen 0002, Wenli Zheng, Jingwen Leng, Minyi Guo |
ICPP | 4 |
| 2021 | SmartDistance: A Mobile-based Positioning System for Automatically Monitoring Social DistanceabstractCoronavirus disease 2019 (COVID-19) has resulted in an ongoing pandemic. Since COVID-19 spreads mainly via close contact among people, social distancing has become an effective manner to slow down the spread. However, completely forbidding close contact can also lead to unacceptable damage to the society. Thus, a system that can effectively monitor people's social distance and generate corresponding alerts when a high infection probability is detected is in urgent need. In this paper, we propose SmartDistance, a smartphone based software framework that monitors people's interaction in an effective manner, and generates a reminder whenever the infection probability is high. Specifically, SmartDistance dynamically senses both the relative distance and orientation during social interaction with a well-designed relative positioning system. In addition, it recognizes different events (e.g., speaking, coughing) and determines the infection space through a droplet transmission model. With event recognition and relative positioning, SmartDistance effectively detects risky social interaction, generates an alert immediately, and records the relevant data for close contact reporting. We prototype SmartDistance on different Android smartphones, and the evaluation shows it reduces the false positive rate from 33% to 1% and the false negative rate from 5% to 3% in infection risk detection. Li Li 0064, Wenli Zheng, Cheng-Zhong Xu 0001 |
INFOCOM | 3 |
| 2021 | QoS-Aware and Resource Efficient Microservice Deployment in Cloud-Edge ContinuumabstractUser-facing services are now evolving towards the microservice architecture where a service is built by connecting multiple microservice stages. While an entire service is heavy, the microservice architecture shows the opportunity to only offload some microservice stages to the edge devices that are close to the end users. However, emerging techniques often result in the violation of Quality-of-Service (QoS) of microservice-based services in cloud-edge continuum, as they do not consider the communication overhead or the resource contention between microservices.We propose Nautilus, a runtime system that effectively deploys microservice-based user-facing services in cloud-edge continuum. It ensures the QoS of microservice-based user-facing services while minimizing the required computational resources. Nautilus is comprised of a communication-aware microservice mapper, a contention-aware resource manager and a load-aware microservice scheduler. The mapper divides the microservice graph into multiple partitions based on the communication overhead and maps the partitions to the nodes. On each node, the resource manager determines the optimal resource allocation for its microservices based on reinforcement learning that may capture the complex contention behaviors. The microservice scheduler monitors the QoS of the entire service, and migrates microservices from busy nodes to idle ones at runtime. Our experimental results show that Nautilus reduces the computational resource usage by 23.9% and the network bandwidth usage by 53.4%, while achieving the required 99%-ile latency. Kaihua Fu, Wei Zhang 0149, Quan Chen 0002, Deze Zeng, Xin Peng 0001, Wenli Zheng, Minyi Guo |
IPDPS | 6 |
| 2021 | AuTraScale: An Automated and Transfer Learning Solution for Streaming System Auto-ScalingabstractThe complexity and variability of streaming data have brought a great challenge to the elasticity of the data processing systems. Streaming systems, such as Flink and Storm, need to adapt to the changes of workload with auto-scaling to meet the QoS requirements while saving resources. However, the accuracy of classical models (such as a queueing model) for QoS prediction decreases with the increase of the complexity and variability of streaming data and the resource interference. On the other hand, the indirect metrics used to optimize QoS may not accurately guide resource adjustment. Those problems can easily lead to waste of resources or QoS violation in practice. To solve the above problems, we propose AuTraScale, an automated and transfer learning auto-scaling solution, to determine the appropriate parallelism and resource allocation that meet the latency and throughput targets. AuTraScale uses Bayesian optimization to adapt to the complex relationship between resources and QoS, minimizing the impact of resource interference on the prediction accuracy, and a new metric that measures the performance of operators for accurate optimization. Even when the input data rate changes, it can quickly adjust the parallelism of each operator in response, with a transfer learning algorithm. We have implemented and evaluated AuTraScale on a Flink platform. The experimental results show that, compared with the state-of-the-art method like DRS and DS2, AuTraScale can reduce 66.6% and 36.7% resource consumption respectively in the scale-down and scale-up scenarios while ensuring QoS requirements, and save 13.5% resource on average when the input data rate changes. Liang Zhang 0027, Wenli Zheng, Chao Li 0009, Minyi Guo |
IPDPS | 2 |
| 2020 | CAPMAN: Cooling and Active Power Management in big.LITTLE Battery Supported DevicesabstractModern smartphone is far from being ubiquitous due to limited energy capacity. Recent research suggests that heterogeneous batteries may expose power saving opportunities that fit dynamic software patterns. Yet it is still a challenge on thermal and power management for a hybrid battery pack in a smartphone. To address this challenge, we propose a system framework, called CAPMAN , which supports joint optimization of cooling and active power management in smartphones. The framework consists of three major techniques: 1) A Markov decision process (MDP) technique that models battery types, and cooling/active power use from state and action nodes; 2) A structural similarity approximation that speeds up the convergence of MDP computation, providing battery scheduling decisions; 3) A TEC and battery management facility to realize the cooling and active power management. In addition, CAPMAN provides an online algorithm with a proved worst-case ${\mathbf{O}}\left( {\frac{1}{{1 - \rho }}} \right)$-competitiveness performance, whereρis the factor of discount. We have prototyped CAPMAN with popular smartphones and heterogeneous batteries, and evaluated them with real-world workloads. Results show that CAPMAN can achieve 114% longer service time under skewed loads, compared to the original phone. Compared to the state-of-the-practice baselines, CAPMAN shows 55% performance gain and 53% less energy use on average. As such, CAPMAN approves that big.LITTLE batteries with a careful system design is an effective way to prolong smartphone service times. Zichen Xu 0001, Wenli Zheng, Yuhao Wang 0001 |
ICDCS | 3 |
| 2020 | OVERSEE: Outsourcing Verification to Enable Resource Sharing in Edge EnvironmentabstractMulti-tenant (or colocation) data centers are good solutions to support edge computing, since each enterprise or organization usually has limited servers at an edge site. When any data center tenant faces a burst of workload, renting resources from the other tenants in the same data center can provide the required resources while keeping the merits of edge computing, but its challenges of reliability and performance are daunting. In this paper, we propose OVERSEE, an outsourcing verification mechanism that enables resource sharing in multi-tenant data centers, fully exploiting the benefits of edge computing. OVERSEE addresses the above two challenges by making skillful use of Intel SGX suite. OVERSEE consists of two sub schemes, the Report-Proof mechanism and Sampling-Challenging mechanism. The Report-Proof mechanism guarantees a task outsourced by a tenant can be executed correctly, i.e., completely and without modification, in the operating environment provided by another tenant. The Sampling-Challenging mechanism can be used to verify that sufficient computing capacity is provided to achieve the required QoS according to the resource lease agreement between the tenants. The theoretical analysis shows the effectiveness of OVERSEE and the experimental results show that it brings minimal overhead. Xiaoqing Cai, Jiuchen Shi, Wenli Zheng, Quan Chen 0002, Chao Li 0009, Jingwen Leng, Minyi Guo |
ICPP | 5 |
| 2020 | Sturgeon: Preference-aware Co-location for Improving Utilization of Power Constrained ComputersabstractLarge-scale datacenters often host latency-sensitive services that have stringent Quality-of-Service requirement and experience diurnal load pattern. Co-locating best-effort applications that have no QoS requirement with latency-sensitive services has been widely used to improve the resource utilization with careful shared resource management. However, existing co-location techniques tend to result in the power overload problem on power constrained computers due to the ignorance of the power consumption. To this end, we propose Sturgeon, a runtime system proactively manages resources between colocated applications in a power constrained environment, to ensure the QoS of latency-sensitive services while maximizing the resource utilization. Our investigation shows that, at a given load, there are multiple feasible resource configurations to meet both QoS requirement and power budget, while one of them yields the maximum throughput of best-effort applications. To find such a configuration, we establish models to accurately predict the performance and power consumption of the colocated applications. Sturgeon monitors the QoS periodically in order to eliminate the potential QoS violation caused by the unpredictable interference. The experimental results show that Sturgeon improves the throughput of best-effort applications by 24.96% compared to the state-of-the-art technique, while guaranteeing the 95%-ile latency within the QoS target. Pu Pang, Quan Chen 0002, Deze Zeng, Chao Li 0009, Jingwen Leng, Wenli Zheng, Minyi Guo |
IPDPS | 6 |
| 2020 | Predicting and reining in application-level slowdown on spatial multitasking GPUs
Mengze Wei, Wenyi Zhao, Quan Chen 0002, Jingwen Leng, Chao Li 0009, Wenli Zheng, Minyi Guo |
J. Parallel Distributed Comput. | 7 |
| 2019 | POSTER: Precise Capacity Planning for Database Public CloudsabstractDatabase platform-as-a-service (dbPaaS) is developing rapidly and a large number of databases have been migrated to run on the Clouds for the low cost and flexibility. Emerging Clouds rely on the tenants to provide the resource specification for their database workloads. However, they tend to over-estimate the resource requirement of their databases, resulting in the unnecessarily high cost and low Cloud utilization. A methodology that automatically suggests the "just-enough" resource specification that fulfills the performance requirement of every database workload is profitable. To this end, we propose URSA, a capacity planning system for dbPaaS Clouds. Our real system experimental results show that URSA can accurately plan the capacity for dbPaaS. Ningxin Zheng, Quan Chen 0002, Yong Yang 0013, Wenli Zheng, Minyi Guo |
PACT | 5 |
| 2019 | When Power Oversubscription Meets Traffic Flood Attack: Re-Thinking Data Center Peak Load ManagementabstractThe state-of-the-art techniques on data center peak power management are too optimistic; they overestimate their benefits in a potentially insecure operating environment. Especially in data centers that oversubscribe power infrastructure, it is likely that unexpected traffics can violate power budget before an effective network DoS attack is observed. In this work, we take the first to investigate the joint effect of power throttling and traffic flooding. We characterize a special operating region in which DoS attacks can provoke undesirable power peaks without exhibiting network traffic anomalies. In this region, an attacker can trigger power emergency by sending normal traffics throughout the Internet. We term this new type of threat as DOPE (Denial of Power and Energy). We show that existing technologies are insufficient for eliminating DOPE without negative performance effects on legitimate users. To enhance data center resiliency, we propose a request-aware power management framework called Anti-DOPE. The key feature of Anti-DOPE is bridging the gap between network traffic controlling and server power management. Specifically, it pre-processes of incoming requests to isolate malicious power attacks on the network load balancer side and then post-processes of compute node performance to minimize the collateral damage it may cause. Anti-DOPE is orthogonal to prior power management schemes and requires minute system modification. Using Alibaba container trace we show that Anti-DOPE allows 44% shorter average response time. It also improves the 90th percentile tail latency by 68.1% compared to the other power controlling methods. Xiaofeng Hou, Mingyu Liang, Chao Li 0009, Wenli Zheng, Quan Chen 0002, Minyi Guo |
ICPP | 4 |
| 2019 | Avalon: towards QoS awareness and improved utilization through multi-resource management in datacentersabstractExisting techniques for improving datacenter utilization while guaranteeing the QoS are based on the assumption that queries have similar behaviors. However, user queries in emerging compute demanding services demonstrate significantly diverse behavior and require adaptive parallelism. Our study shows that the end-to-end latency of the compute demanding query is determined together by the system-wide load, its workload, its parallelism, contention on shared cache, and memory bandwidth. When hosting such new services, the current cross-query resource allocation results in either severe QoS violation or significant resource under-utilization. Quan Chen 0002, Zhenning Wang, Jingwen Leng, Chao Li 0009, Wenli Zheng, Minyi Guo |
ICS | 5 |
| 2019 | Themis: Predicting and Reining in Application-Level Slowdown on Spatial Multitasking GPUsabstractPredicting performance degradation of a GPU application when it is co-located with other applications on a spatial multitasking GPU without prior application knowledge is essential in public Clouds. Prior work mainly targets CPU co-location, and is inaccurate and/or inefficient for predicting performance of applications at co-location on spatial multitasking GPUs. Our investigation shows that hardware event statistics caused by co-located applications, which can be collected with negligible overhead, strongly correlate with their slowdowns. Based on this observation, we present Themis, an online slowdown predictor that can precisely and efficiently predict application slowdown without prior application knowledge. We first train a precise slowdown model offline using hardware event statistics collected from representative co-locations. When new applications co-run, Themis collects event statistics and predicts their slowdowns simultaneously. Our evaluation shows that Themis has negligible runtime overhead and can precisely predict application-level slowdown with prediction error smaller than 9.5%. Based on Themis, we also implement an SM allocation engine to rein in application slowdown at co-location. Case studies show that the engine successfully enforces fair sharing and QoS. Wenyi Zhao, Quan Chen 0002, Jingwen Leng, Chao Li 0009, Wenli Zheng, Li Li 0012, Minyi Guo |
IPDPS | 7 |
| 2019 | SprintCon: Controllable and Efficient Computational Sprinting for Data Center ServersabstractComputational sprinting is an effective mechanism to temporarily boost the performance of data center servers. However, given the great effect on performance improvement, how to make the sprinting process controllable and how to maximize the sprinting efficiency have not been well discussed yet. Those can be significant problems for a data center when computational sprinting is needed for more than a few minutes, since it requires the support of energy storage, whose capacity is limited. The control and efficiency of sprinting not only involve how fast to run servers and how to allocate resources to corunning workloads, but also the impact on power overload, and how to handle the overload with circuit breakers and energy storage to ensure power safety. Different workloads can impact sprinting in different ways, and hence efficient sprinting requires workload-specific strategies. In this paper, we propose SprintCon to realize controllable and efficient computational sprinting for data center servers. SprintCon mainly consists of a power load allocator and two different power controllers. The allocator analyzes how to divide the power load to different power sources. The server power controller adapts the CPU cores that process batch workloads, to improve the efficiency in terms of computing, energy and cost. The UPS power controller dynamically adjusts the discharge rate of UPS energy storage to satisfy the time-varying power demand of interactive workloads, and ensure power safety. The experiment results show that compared to state-of-theart solutions, SprintCon can achieve 6-56% better computing performance and up to 87% less demand of energy storage. Wenli Zheng, Chao Li 0009, Bin Yao 0002, Minyi Guo |
IPDPS | 1 |
| 2018 | Power Grab in Aggressively Provisioned Data Centers: What is the Risk and What Can Be Done About ItabstractAggressively provisioned data centers achieve great cost savings by over-committing the very expensive power distribution infrastructure. However, existing proposals for managing load power demand in such a data center are largely utilization-driven, overlooking power-related interferences among users. An important observation is that some tasks can impact existing power budget management framework and disrupt normal operation by taking away the precious public power capacity. This vulnerability exposes data centers to a new type of risk that we call power grab, which is essentially hostile power resource competition. It could worsen the performance-utilization tradeoff in a power-constrained computing environment. Anticipating a growing case for power-oriented com-petition, we propose CFP, a resilient power capacity management frame-work for improving the fairness and service quality in scale-out data centers. Our solution features a market-based power re-source allocation and billing scheme that involves users in the loop. It allows the data center to bypass the formidable task of identifying malicious users and defend against power grab with reward and punishment incentives. We build a proof-of-concept system and also evaluate our design with realistic Google cluster traces. Compared to prior arts, CFP can increase the average performance-cost ratio by 1.8X. It can boost the total throughput in an APDC by 15% under severe power contention. Our design allows scale-out data centers to safely exploit the benefits that power over-subscription may provide, with minor overhead. Xiaofeng Hou, Luoyao Hao, Chao Li 0009, Quan Chen 0002, Wenli Zheng, Minyi Guo |
ICCD | 5 |
| 2017 | PowerNetS: Coordinating Data Center Network With Servers and Cooling for Power OptimizationabstractRecently, a lot of research efforts have been made to optimize the large amounts of energy consumed by different devices in data centers, including servers, cooling, and the data center network (DCN). Unfortunately, current research addresses these devices mostly in a separate manner, leading to inferior optimization results. This paper proposes PowerNetS, a power optimization framework that coordinates servers and DCN, as well as cooling, for minimized power consumption of a data center. PowerNetS leverages workload correlation analysis for more energy savings during server and traffic consolidations. More importantly, PowerNetS tries to change the DCN topology during server consolidation, in order to have more intra-server traffic and shorter flows that go through fewer switches. For example, two virtual machines previously located on two different servers can now be migrated to the same server, so that the flow between them no longer needs to use switches, which allows more devices to sleep for energy savings without network performance degradation. PowerNetS has been implemented on a physical testbed with 6 servers and 10 virtual switches that are configured using a production 48-port OpenFlow switch. Our evaluation with Wikipedia, Yahoo!, and IBM traces shows that PowerNetS can save up to 51.6% of energy by coordinating servers and DCN, which is 44.3% and 15.8% more than two state-of-the-art baselines, respectively. By further coordinating with cooling to utilize different cooling efficiencies at different locations within a data center, PowerNetS can achieve 8.8%-14.6% additional energy savings. Kuangyu Zheng, Wenli Zheng, Li Li 0064 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2017 | Hybrid Energy Storage with Supercapacitor for Cost-Efficient Data Center Power Shaving and CappingabstractRecent studies have proposed to dynamically reshape the power demand curve of a data center (i.e., power shaving) with energy storage devices, particularly uninterruptible power supply (UPS) batteries. Power shaving can be used to limit the peak power demand in a data center, in order to reduce both the power infrastructure investment (i.e., cap-ex) and the electricity bills (i.e., op-ex). However, power shaving requires the UPS batteries to be frequently charged/discharged, which is known to compromise the battery lifetime and availability. This paper presents a detailed quantitative study that explores different options to integrate supercapacitor (SC) with batteries for cost-efficient energy storage. Compared with batteries, SC allows more charge/discharge cycles and has a higher power density, which are desirable for fast power shaving. However, SC also has undesirable characteristics (e.g., relatively high self-discharging rate and cost). Therefore, we quantitatively compare three possible energy storage options (i.e., Battery-only, SC-only, and Battery+SC) in detail, with different SC self-discharging rate assumptions. SC options (SC-only and Battery+SC) are shown to be more cost-efficient designs, saving the energy storage cost by 34 percent, on average, compared with Battery-only. For a 10 MW data center in a 10-year period, the savings can be converted to $3 M in total cost of ownership (TCO) reduction by allowing more servers to be deployed. In addition, we also propose the integration of energy storage with dynamic voltage and frequency scaling (DVFS) to cap the peak power demand (i.e., power capping). Specifically, we comparatively studyfour power capping algorithms and discuss their applicable scenarios. Finally, we introduce our proof-of-concept SC physical testbed and present preliminary hardware testing results. Wenli Zheng |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | TECfan: Coordinating Thermoelectric Cooler, Fan, and DVFS for CMP Energy OptimizationabstractThe cooling needs of modern processors are dominantly provided by cooling fans, which can only offer global cooling capability. Even for cooling down one single hot unit (i.e., local hot spot), the fan needs to run at a high speed level, which consumes a lot of cooling power. Fortunately, emerging technologies, such as thermoelectric cooler (TEC), offer effective local cooling, which can be integrated with the fan to improve the overall cooling efficiency. However, relying on only TEC and fan may not optimize the total energy consumption of a chip multiprocessor (CMP), because the CMP core power states impact both computing and cooling power consumption. Therefore, for optimizing CMP energy, it is necessary to intelligently manage the processor power states and coordinate it with the cooling system. In this paper, we propose TECfan, a hierarchical runtime optimization framework that integrates TEC, fan, and DVFS for the overall energy efficiency of CMP. TECfan coordinates TEC and fan for efficient cooling, and also exploits DVFS to adapt the computing power consumption and execution time. Specifically, we first formulate CMP energy optimization with temperature constraint as a nonlinear optimization problem. Since there are no known polynomial-time algorithms for such a problem, solving it online is prohibitive. Hence, a novel heuristic algorithm is designed to solve it with acceptable time overheads. Our experiment results show that TECfan leads to 29% less energy consumption for medium workload compared to a state-of-the-art solution and 27%less overall compared to fan-based cooling. Wenli Zheng |
IPDPS | 1 |
| 2015 | Data Center Sprinting: Enabling Computational Sprinting at the Data Center LevelabstractMicroprocessors may need to keep most of their cores off in the era of dark silicon due to thermal constraints. Recent studies have proposed Computational Sprinting, which allows a chip to temporarily exceed its power and thermal limits by turning on all its cores for a short time period, such that its computing performance is boosted for bursty computation demands. However, conducting sprinting in a data center faces new challenges due to power and thermal constraints at the data center level, which are exacerbated by recently proposed power infrastructure under-provisioning and reliance on renewable energy, as well as the increasing server density. In this paper, we propose Data Center Sprinting, a methodology that enables a data center to temporarily boost its computing performance by turning on more cores in the era of dark silicon, in order to handle occasional workload bursts. We demonstrate the feasibility of this approach by analyzing the tripping characteristics of data center circuit breakers and the discharging characteristics of energy storage devices, in order to realize safe sprinting without causing undesired server overheating or shutdown. We evaluate a prototype of Data Center Sprinting on a hardware testbed and in data enter-level simulations. The experimental results show that our solution can improve the average computing performance of a data center by a factor of 1.62 to 2.45 for 5 to 30 minutes. Wenli Zheng |
ICDCS | 1 |
| 2015 | BYY harmony learning of log-normal mixtures with automated model selection
Wenli Zheng, Zhijie Ren, Jinwen Ma |
Neurocomputing | 1 |
| 2015 | TE-Shave: Reducing Data Center Capital and Operating Expenses with Thermal Energy StorageabstractPower shaving has recently been proposed to dynamically shave the power peaks of a data center with energy storage devices (ESD), such that more servers can be safely hosted. In addition to the reduction of capital investment (cap-ex), power shaving also helps cut the electricity bills (op-ex) of a data center by reducing the high utility tariffs related to peak power. However, existing work on power shaving focuses exclusively on electrical ESDs (e.g., UPS batteries) to shave the server-side power demand. In this paper, we propose TE-Shave, a generalized power shaving framework that exploits both UPS batteries and a new knob, thermal energy storage (TES) tanks equipped in many data centers. Specifically, TE-Shave utilizes stored cold water or ice to manipulate the cooling power, which accounts for 30-40 percent of the total power cost of a data center. Our extensive evaluation with real-world workload traces shows that TE-Shave saves cap-ex and op-ex up to $2,668/day and $825/day, respectively, for a data center with 17,920 servers. Even for future data centers that are projected to have more efficient cooling and thus a smaller portion of cooling power, e.g., a quarter of today's level, TE-Shave still leads to 28 percent more savings than existing work that focuses only on the server-side power. TE-Shave is also coordinated with traditional TES solutions for further reduced op-ex, and integrated with processor throttling to cap the power draw (i.e., power capping). Our hardware testbed results show that TE-Shave can improve the system performance up to 23 percent. Wenli Zheng |
IEEE Trans. Computers | 1 |
| 2014 | Exploiting thermal energy storage to reduce data center capital and operating expensesabstractPower shaving has recently been proposed to dynamically shave the power peaks of a data center with energy storage devices (ESD), such that more servers can be safely hosted. In addition to the reduction of capital investment (cap-ex), power shaving also helps cut the electricity bills (op-ex) of a data center by reducing the high utility tariffs related to peak power. However, existing work on power shaving focuses exclusively on electrical ESDs (e.g., UPS batteries) to shave the server-side power demand. In this paper, we propose TE-Shave, a generalized power shaving framework that exploits both UPS batteries and a new knob, thermal energy storage (TES) tanks equipped in many data centers. Specifically, TE-Shave utilizes stored cold water or ice to manipulate the cooling power, which accounts for 30-40% of the total power cost of a data center. Our extensive evaluation with real-world workload traces shows that TE-Shave saves cap-ex and op-ex up to $2,668/day and $825/day, respectively, for a data center with 17,920 servers. Even for future data centers that are projected to have more efficient cooling and thus a smaller portion of cooling power, e.g., a quarter of today's level, TE-Shave still leads to 28% more savings than existing work that focuses only on the server-side power. TE-Shave is also coordinated with traditional TES solutions for further reduced op-ex. Wenli Zheng |
HPCA | 1 |
| 2014 | Diagonal Log-Normal Generalized RBF Neural Network for Stock Price Prediction
Wenli Zheng, Jinwen Ma |
ISNN | 1 |
| 2014 | Reducing the expenses of geo-distributed data centers with portable containerized modules
Marco Brocanelli, Wenli Zheng |
Perform. Evaluation | 2 |