Di Wang 0003

dblp:18/5410-3 · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-8003-9738ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 4 first-author · 6 since 2021Computer networks · 9 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 8 · 4 first-authorDatabases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MPulse: A Programmable and Autonomic Fault Detection System via Hierarchical Liveness Exchange
Di Wang 0003, Haifeng Zhou, Zhengyan Zhou, Jiayu Luo
INFOCOM1
2026 Holm: A DPU-Based Robust Host-Side Latency Monitoring with Low-Overhead and Selective Full-Coverage
Haifeng Zhou, Di Wang 0003, Wenbin Zhang 0011, Dianxing Tang, Zhengyan Zhou, Chunming Wu 0001
INFOCOM2
2026 Achieving Precise Host Congestion Mitigation by DPU Offloading
Haifeng Zhou, Di Wang 0003, Dianxing Tang, Zhengyan Zhou, Chunming Wu 0001
INFOCOM2
2026 PSM: Timely and Resource-Efficient Sketch Migration in Network Measurement
Hongyan Liu 0001, Xiang Chen 0017, Zhengyan Zhou, Di Wang 0003, Chunming Wu 0001
IWQoS6
2026 MSRTrack: LLM-Powered Object Tracking with Motion and Semantic Reasoning
abstract
State-of-the-art object trackers primarily model appearance relations between the image template and the search region with Siamese networks. However, this well-established approach has a limited ability to leverage both motion and semantic cues of the target object, leading to increasing errors in challenging scenarios like drastic appearance changes and similar-looking distractors. To address the above weaknesses, we propose a novel tracking framework with Motion and Semantic Reasoning (MSRTrack), integrating short-term motion modeling and distinctive semantic features for robust tracking across diverse conditions. Powered by vision large language models (VLLMs) and the Segment Anything Model 2 (SAM2), MSRTrack identifies unique semantic attributes of the target, exploits motion cues across consecutive frames, and complements appearance-based trackers with strong semantic and dynamic reasoning capabilities. Unlike previous vision language tracking (VLT) methods that rely on broad captioning, MSRTrack automatically focuses on a concise set of key semantic attributes of the target, substantially improving target lost recovery and distractor rejection. MSRTrack achieves state-of-the-art performance across multiple tracking benchmarks, with 2.2% improvement on the LaSOT dataset, 9.5% improvement on the VastTrack dataset, and 1.4% on the TNL2K dataset.
Di Wang 0003, José M. F. Moura
WACV2
2026 FedMT: Multitask Federated Learning With Competitive GPU Resource Sharing
abstract
Federated learning (FL) nowadays involves heterogeneous compound learning tasks as cognitive applications’ complexity increases. For example, a self-driving system hosts multiple tasks simultaneously (e.g., detection, classification, segmentation, etc.) and expects FL to retain life-long intelligence involvement. However, our analysis demonstrates that, when deploying compound FL models for multiple training tasks on a GPU, certain issues arise: As different tasks’ skewed data distributions and corresponding models cause highly imbalanced learning workloads, current GPU scheduling methods lack effective resource allocations; Therefore, existing FL schemes, only focusing on heterogeneous data distribution but runtime computing, cannot practically achieve optimally synchronized federation. To address these issues, we propose a full-stack FL optimization scheme to tackle both intra-device GPU scheduling and inter-device FL coordination for multi-task training. Specifically, our works illustrate two key insights in this research domain: Competitive resource sharing is beneficial for parallel model executions, and the proposed concept of “virtual resource” could effectively characterize and guide the practical per-task resource utilization and allocation; Additionally, architectural-level coordination improves FL performance by aligning task workloads with GPU utilization. Our experiments demonstrate that the FL performance could be significantly escalated. Specifically, we observed a 2.16×–2.38× increase in intra-device GPU training throughput and a 2.53×–2.80× boost in inter-device FL coordination efficiency across diverse multi-task scenarios.
Fuxun Yu, Di Wang 0003, Minjia Zhang, Ang Li 0005, Zhi Tian, Xiang Chen 0010
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Toward Sensor-In-the-Loop LLM Agent: Benchmarks and Implications
abstract
This paper explores sensor-informed personal agents that can take advantage of sensor hints on wearables to enhance the personal agent's response. We demonstrate that such a sensor-in-the-loop AI agent design can be easily integrated into existing LLM agents by building a prototype named WellMax based on existing well-developed techniques such as structured prompt templates and few-shot prompting. The head-to-head comparison with a non-sensor-informed agent across five use scenarios demonstrates that this sensor-in-the-loop design can effectively improve users' needs and their overall experience. The deep-dive into agents' replies and participants' feedback further reveals that sensor-in-the-loop agents not only provide more contextually relevant responses but also exhibit a better understanding of user priorities and situational nuances. In addition, we conduct two case studies to examine the potential pitfalls and distill key insights from this sensor-in-the-loop agent. We hope this work can spawn new ideas for building more intelligent, empathetic, and effective AI-driven personal assistants.
Zhiwei Ren, Minjia Zhang, Di Wang 0003, Xiaoran Fan, Longfei Shangguan
SenSys4
2025 FlowTracker: A refined and versatile data plane measurement approach
Chunming Wu 0001, Zhengyan Zhou, Di Wang 0003, Dezhang Kong, Muhammad Khurram Khan, Xuan Liu 0006
J. Netw. Comput. Appl.4
2025 Rethinking Latency-Aware DNN Design With GPU Tail Effect Analysis
abstract
As the size of Deep Neural Networks (DNNs) continues to grow, their runtime latency also scales. While model pruning and Neural Architecture Search (NAS) can effectively reduce the computation workload, their effectiveness fails to consistently translate into runtime latency reduction. In this paper, we identify the root cause behind the mismatch between workload reduction and latency reduction is GPU tail effect – a classic system issue caused by resource under-utilization in the last processing wave of the GPU. We conduct detailed DNN workload characterization and demonstrate the prevalence of GPU tail effect across different DNN architectures, and meanwhile reveal that the unique deep structure and the light-weight layer workload of DNNs exacerbate the tail effect for DNN inference. We then propose a tail-awareness design space enhancement and DNN optimization algorithm to optimize existing NAS and pruning designs and achieve better runtime latency and model accuracy performance. Extensive experiments show 11%-27% latency reduction over SOTA DNN pruning and NAS methods.
Fuxun Yu, Longfei Shangguan, Di Wang 0003, Dimitrios Stamoulis, Rishi Madhok, Nikolaos Karianakis, Ang Li 0005, Yiran Chen 0001, Xiang Chen 0010
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Forecasting Graph-Based Time-Dependent Data with Graph Sequence Attention
abstract
Forecasting graph-based, time-dependent data has broad practical applications but presents challenges. Effective models must capture both spatial and temporal dependencies in the data, while also incorporating auxiliary information to enhance prediction accuracy. In this article, we identify limitations in current state-of-the-art models regarding temporal dependency handling. To overcome this, we introduce GSA-Forecaster, a new deep learning model designed for forecasting in graph-based, time-dependent contexts. GSA-Forecaster utilizes graph sequence attention, a new attention mechanism proposed in this article, to effectively manage temporal dependencies. GSA-Forecaster integrates the data’s graph structure directly into its architecture, addressing spatial dependencies. Additionally, it incorporates auxiliary information to refine its predictions further. We validate its performance using real-world graph-based, time-dependent datasets, where it demonstrates superior effectiveness compared to existing state-of-the-art models.
Yang Li 0183, Di Wang 0003, José M. F. Moura
ACM Trans. Knowl. Discov. Data2
2024 FlexPDD: Enabling Proportional Delay Differentiation Service on Programmable Switches
abstract
Quality-of-Service (QoS) guarantees are crucial for meeting the diverse performance requirements of applications in packet networks. The Proportional Delay Differentiation (PDD) model offers relative service differentiation based on the delay requirements of different traffic classes. However, implementing PDD on current hardware switches faces challenges due to the lack of inherent queuing behavior description in switch ASICs. This paper introduces FlexPDD, a dynamic and adaptive packet prioritization mechanism designed to implement the PDD model on programmable switches. FlexPDD leverages the flexibility of programmable switch to adjust the mapping between packet classes and output queues dynamically, ensuring precise control over delay differentiation. Our implementation of FlexPDD on a Barefoot Tofino switch and an NS3 simulator demonstrates its feasibility and effectiveness. The results indicate that FlexPDD successfully maintains approximate delay differentiation among service classes proportional to their delay weights, highlighting its potential as a practical solution for achieving advanced service differentiation in modern network infrastructures.
Dezhang Kong, Zhengyan Zhou, Di Wang 0003, Shuangxi Chen, Chunming Wu 0001
GLOBECOM4
2024 Public Opinion Evolution in Cyberspace: A Case Analysis of Pelosi's Visit to Taiwan
abstract
The dynamics of public opinion on social media affects people’s feeling and minds about international affairs and leads to the reconstruction of societal states for international conflicts. In this article, we analyze the topics’ evolution on social media during the Pelosi visit. Such kind of analysis should help the related departments sense and beware the situation effectively and efficiently, and may provide technical supports for proper policy making and responses. To facilitate this purpose, a new method is proposed and an abbreviated large-graph clustering (ALGC) algorithm has been designed to generate documents and topic representation for alleviating the overhead of high computational complexity of large graphs by reducing the dimensionality of the attention matrix and adjacency matrix. The evolution pattern of topics is also analyzed in and between different time periods. Experiment results show that the proposed method performs well, achieving a high clustering accuracy with lower computational cost. The dataset used in this article is also released for public analysis.
Tao Chen 0023, Baoyu Zhang, Xiao Wang 0002, Weishan Zhang, Chitin Hon, Di Wang 0003, Long Chen 0001, Qiang Li 0060, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.6
2024 A Dual Neural Network for Defect Detection With Highly Imbalanced Data in 3-D Printing
abstract
Digital light processing (DLP) is a popular additive manufacturing technology that uses light irradiation to fabricate 3-D devices via a projector to achieve laser-sensitive resin curing. However, the performance and reliability of DLP can be affected by internal defects such as printing errors and the accumulation of residual stress. Existing defect detection methods rely on monitoring the printed parts, which leads to resource wastage and struggles to effectively handle imbalanced defect data. In this article, we propose a defect detection method called dual neural network, which involves detecting defects in materials before the printing process to prevent resource wastage and serious consequences. Specifically, to handle the highly imbalanced class distribution problem in online DLP defect detection, dual neural network utilizes a domain learner and balance learner to effectively balance the information of the minority class and learn the generalization knowledge from the imbalanced defect dataset. Experimental results demonstrate the effectiveness of our proposed method, which has also been applied to real-world production equipment successfully.
Fang Wang 0033, Gang Xiong 0001, Qihang Fang, Zhen Shen 0004, Di Wang 0003, Xisong Dong, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.5
2024 Knowledge Graph-Based Reinforcement Federated Learning for Chinese Question and Answering
abstract
Knowledge question and answering (Q&A) is widely used. However, most existing semantic parsing methods in Q&A usually use cascading, which can incur error accumulation. In addition, using only one institution’s Q&A data definitely will limit the Q&A performance, while data privacy prevents sharing between institutions. This article proposes a knowledge graph-based reinforcement federated learning (KGRFL)-based Q&A approach to address these challenges. We design an end-to-end multitask semantic parsing model [MSP-bidirectional and auto-regressive transformers (BART)] that identifies question categories while converting questions into SPARQL statements to improve semantic parsing. Meanwhile, a reinforcement learning (RL)-based model fusion strategy is proposed to improve the effectiveness of federated learning, which enables multi-institution joint modeling and data privacy protection using cross-domain knowledge. In particular, it also reduces the negative impact of low-quality clients on the global model. Furthermore, a prompt learning-based entity disambiguation method is proposed to address the semantic ambiguity problem because of joint modeling. The experiments show that the proposed method performs well on different datasets. The Q&A results of the proposed approach outperform the approach of using only a single institution. Experiments also demonstrate that the proposed approach is resilient to security attacks, which is required for real applications.
Liang Xu 0009, Tao Chen 0023, Zhaoxiang Hou, Weishan Zhang, Chitin Hon, Xiao Wang 0002, Di Wang 0003, Long Chen 0001, Wenyin Zhu, Yunlong Tian, Huansheng Ning, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.7
2023 Cost-effective On-device Continual Learning over Memory Hierarchy with Miro
abstract
Continual learning (CL) trains NN models incrementally from a continuous stream of tasks. To remember previously learned knowledge, prior studies store old samples over a memory hierarchy and replay them when new tasks arrive. Edge devices that adopt CL to preserve data privacy are typically energy-sensitive and thus require high model accuracy while not compromising energy efficiency, i.e., cost-effectiveness. Our work is the first to explore the design space of hierarchical memory replay-based CL to gain insights into achieving cost-effectiveness on edge devices. We present Miro, a novel system runtime that carefully integrates our insights into the CL framework by enabling it to dynamically configure the CL system based on resource states for the best cost-effectiveness. To reach this goal, Miro also performs online profiling on parameters with clear accuracy-energy trade-offs and adapts to optimal values with low overhead. Extensive evaluations show that Miro significantly outperforms baseline systems we build for comparison, consistently achieving higher cost-effectiveness.
Suyeon Jeong, Minjia Zhang, Di Wang 0003, Myeongjae Jeon
MobiCom4
2023 RFT: Toward Highly Reliable Flow Data Transmission in Network Measurement
abstract
How to satisfy the latency and reliability requirements of flow data transfer is an essential problem. To address this problem, we propose RFT, a framework that aims to satisfy the user-specified latency and reliability requirements of flow data transfer, especially in the situation where the network resources are insufficient. Firstly, we formulate the problem of satisfying the user-specified latency and reliability requirements of data transfer via mixed integer linear programming (MILP), and a heuristic algorithm is then designed to solve it in a polynomialtime. Secondly, to satisfy these requirements under insufficient network resources, we proposed a greedy-based algorithm used to select the minimum number of links added to the network, which can be deployed with low cost, especially in production networks such as data centers. Finally, we have implemented RFT on a 64$\times$100 Gbps Intel Barefoot Tofino switch. Our experimental results indicate that RFT satisfies the user-specified latency and reliability requirements in all test cases at acceptable costs, even when the network resources are insufficient.
Xiang Chen 0017, Di Wang 0003, Zhengyan Zhou, Wenhai Wang, Chunming Wu 0001, Haifeng Zhou
SECON3
2022 CarM: hierarchical episodic memory for continual learning
abstract
Continual Learning (CL) is an emerging machine learning paradigm in mobile or IoT devices that learns from a continuous stream of tasks. To avoid forgetting of knowledge of the previous tasks, episodic memory (EM) methods exploit a subset of the past samples while learning from new data. Despite the promising results, prior studies are mostly simulation-based and unfortunately do not promise to meet an insatiable demand for both EM capacity and system efficiency in practical system setups. We propose CarM, the first CL framework that meets the demand by a novel hierarchical EM management strategy. CarM has EM on high-speed RAMs for system efficiency and exploits the abundant storage to preserve past experiences and alleviate the forgetting by allowing CL to efficiently migrate samples between memory and storage. Extensive evaluations show that our method significantly outperforms popular CL methods while providing high training efficiency.
Soobee Lee, Minindu Weerakoon, Minjia Zhang, Di Wang 0003, Myeongjae Jeon
DAC5
2022 SC-UDA: Style and Content Gaps aware Unsupervised Domain Adaptation for Object Detection
abstract
Current state-of-the-art object detectors can have significant performance drop when deployed in the wild due to domain gaps with training data. Unsupervised Domain Adaptation (UDA) is a promising approach to adapt detectors for new domains/environments without any expensive label cost. Previous mainstream UDA works for object detection usually focused on image-level and/or feature-level adaptation by using adversarial learning methods. In this work, we show that such adversarial-based methods can only reduce domain style gap, but cannot address the domain content gap that is also important for object detectors. To overcome this limitation, we propose the SC-UDA framework to concurrently reduce both gaps: We propose fine-grained domain style transfer to reduce the style gaps with finer image details preserved for detecting small objects; Then we leverage the pseudo label-based self-training to reduce content gaps; To address pseudo label error accumulation during self-training, novel optimizations are proposed, including uncertainty-based pseudo labeling and imbalanced mini-batch sampling strategy. Experiment results show that our approach consistently outperforms prior state-of-the-art methods (up to 8.6%, 2.7% and 2.5% mAP on three UDA benchmarks).
Fuxun Yu, Di Wang 0003, Yinpeng Chen, Nikolaos Karianakis, Pei Yu, Dimitrios Lymberopoulos, Sidi Lu, Weisong Shi, Xiang Chen 0010
WACV2
2022 AntiDoteX: Attention-Based Dynamic Optimization for Neural Network Runtime Efficiency
abstract
Deep neural networks (DNNs) achieved great cognitive performance at the expense of a considerable computation workload. To relieve the computational burden, many optimization works are developed to reduce the model redundancy by identifying and removing insignificant model components, such as weight sparsity and filter pruning methods. However, these works only evaluate model components’ static significance with parameter information, ignoring their dynamic interaction with external inputs. Specifically, due to the difference in per-input features, the model components’ significance can dynamically change and, thus, the static methods can only achieve suboptimal performance. Focusing on this aspect, we propose a dynamic DNN optimization framework in this work. Based on the neural network attention mechanism, we propose a comprehensive dynamic optimization framework, including 1) testing-phase dynamic feature map pruning; 2) training-phase optimization by training with targeted dropout; and 3) deployment-phase one-for-all (OFA) model adaptability enhancement. By providing a holistic dynamic testing, training, and deployment co-optimization framework, our work has the following benefits: first, it can accurately identify and aggressively remove per-input feature redundancy by considering the model-input interaction and involving the channel/column-wise pruning flexibility; meanwhile, the training-testing co-optimization favors the dynamic pruning and helps maintain the model accuracy even with a very high feature pruning ratio. Finally, the deployment enhancement provides one unified OFA model to support full-spectrum feature sparsity ratios. The unified model can be dynamically reconfigured to meet different resource budgets without any retraining cost, and thus provide significant deployment flexibility. Extensive experiments show that our method could bring 37.4%–54.5% floating-point operations reduction with negligible accuracy drop on various test benchmarks. Meanwhile, the OFA deployment optimization enables us to use one model to support at most ten different resource constraints without any retraining cost.
Fuxun Yu, Dimitrios Stamoulis, Di Wang 0003, Yanzhi Wang 0001, Xiang Chen 0010
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 Automated Runtime-Aware Scheduling for Multi-Tenant DNN Inference on GPU
abstract
With the fast development of deep neural networks (DNNs), many real-world applications are adopting multiple models to conduct compound tasks, such as co-running classification, detection, and segmentation models on autonomous vehicles. Such multi-tenant DNN inference cases greatly exacerbate the computational complexity and call for comprehensive collaboration for graph-level operator scheduling, runtime-level resource awareness, as well as hardware scheduler support. However, the current scheduling support for such multi-tenant inference is still relatively backward. In this work, we propose a resource-aware scheduling framework for efficient multi-tenant DNN inference on GPU, which automatically coordinates DNN computing in different execution levels. Leveraging the unified scheduling intermediate representation and the automated ML-based searching algorithm, optimal schedules could be generated to wisely adjust model concurrency and interleave DNN model operators, maintaining a continuously balanced resource utilization across the entire inference process, and eventually improving the runtime efficiency. Experiments show that we could consistently achieve$1.3\times\sim 1.7\times$speed-up, comparing to regular DNN runtime libraries (e.g., CuDNN, TVM) and particular concurrent scheduling methods (e.g., NVIDIA Multi-Stream).
Fuxun Yu, Shawn Bray, Di Wang 0003, Longfei Shangguan, Xulong Tang, Xiang Chen 0010
ICCAD3
2021 Fed2: Feature-Aligned Federated Learning
abstract
Federated learning learns from scattered data by fusing collaborative models from local nodes. However, conventional coordinate-based model averaging by FedAvg ignored the random information encoded per parameter and may suffer from structural feature misalignment. In this work, we propose Fed2, a feature-aligned federated learning framework to resolve this issue by establishing a firm structure-feature alignment across the collaborative models. Fed2 is composed of two major designs: First, we design a feature-oriented model structure adaptation method to ensure explicit feature allocation in different neural network structures. Applying the structure adaptation to collaborative models, matchable structures with similar feature information can be initialized at the very early training stage. During the federated learning process, we then propose a feature paired averaging scheme to guarantee aligned feature distribution and maintain no feature fusion conflicts under either IID or non-IID scenarios. Eventually, Fed2 could effectively enhance the federated learning convergence performance under extensive homo- and heterogeneous settings, providing excellent convergence speed, accuracy, and computation/communication efficiency.
Fuxun Yu, Weishan Zhang, Zhuwei Qin, Di Wang 0003, Zhi Tian, Xiang Chen 0010
KDD5
2021 REIN the RobuTS: Robust DNN-Based Image Recognition in Autonomous Driving Systems
abstract
In recent years, the neural network (NN) has shown its great potential in image recognition tasks of autonomous driving systems, such as traffic sign recognition, pedestrian detection, etc. However, theoretically well-trained NNs usually fail their performance when facing real-world scenarios. For example, adverse real-world conditions, e.g., bad weather and lighting conditions, can introduce different physical variations and cause considerable accuracy degradation. As for now, the generalization capability of NNs is still one of the most critical challenges for the autonomous driving system. To facilitate the robust image recognition tasks, in this work, we build the RobuTS dataset: a comprehensive Robust Traffic Sign Recognition dataset, which includes images with different environmental variations, e.g., rain, fog, darkening, and blurring. Then to enhance the NN's generalization capability, we propose two generalization-enhanced training schemes: 1) REIN for robust training without data in adverse scenarios and 2) Self-Teaching (ST) for robust training with unlabeled adverse data. The great advantages of such two training schemes are they are data-free (REIN) and label-free (ST), thus effectively reducing the huge human efforts/cost of on-road driving data collection, as well as the expensive manual data annotation. We conduct extensive experiments to validate our methods' performance on both classification and detection tasks. For classification tasks, our proposed training algorithms could consistently improve model performance by +15%-25% (REIN) and +16%-30% (ST) in all adverse scenarios of our RobuTS datasets. For detection tasks, our ST could also improve the detector's performance by +10.1 mean average precision (mAP) on Foggy-Cityscapes, outperforming previous state-of-the-art works by +2.2 mAP.
Fuxun Yu, Zhuwei Qin, Di Wang 0003, Xiang Chen 0010
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 AntiDote: Attention-based Dynamic Optimization for Neural Network Runtime Efficiency
abstract
Convolutional Neural Networks (CNNs) achieved great cognitive performance at the expense of considerable computation load. To relieve the computation load, many optimization works are developed to reduce the model redundancy by identifying and removing insignificant model components, such as weight sparsity and filter pruning. However, these works only evaluate model components’ static significance with internal parameter information, ignoring their dynamic interaction with external inputs. With per-input feature activation, the model component significance can dynamically change, and thus the static methods can only achieve sub-optimal results. Therefore, we propose a dynamic CNN optimization framework in this work. Based on the neural network attention mechanism, we propose a comprehensive dynamic optimization framework including (1) testing-phase channel and column feature map pruning, as well as (2) training-phase optimization by targeted dropout. Such a dynamic optimization framework has several benefits: (1) First, it can accurately identify and aggressively remove per-input feature redundancy with considering the model-input interaction; (2) Meanwhile, it can maximally remove the feature map redundancy in various dimensions thanks to the multi-dimension flexibility; (3) The training-testing co-optimization favors the dynamic pruning and helps maintain the model accuracy even with very high feature pruning ratio. Extensive experiments show that our method could bring 37.4%∼54.5% FLOPs reduction with negligible accuracy drop on various of test networks.
Fuxun Yu, Di Wang 0003, Yanzhi Wang 0001, Xiang Chen 0010
DATE3
2020 DC-CNN: Computational Flow Redefinition for Efficient CNN through Structural Decoupling
abstract
Recently Convolutional Neural Networks (CNNs) are widely applied into novel intelligent applications and systems. However, the CNN computation performance is significantly hindered by its computation flow, which computes the model structure sequentially by layers with massive convolution operations. Such a layer-wise sequential computation flow can cause certain performance issues, such as resource under-utilization, huge memory overhead, etc. To solve these problems, we propose a novel CNN structural decoupling method, which could decouple CNN models into "critical paths" and eliminate the inter-layer data dependency. Based on this method, we redefine the CNN computation flow into parallel and cascade computing paradigms, which can significantly enhance the CNN computation performance with both multi-core and single-core CPU processors. Experiments show that, our DC-CNN framework could reduce 24% to 33% latency on multi-core CPUs for CIFAR and ImageNet. On small-capacity mobile platforms, cascade computing could reduce the latency by average 24% on ImageNet and 42% on CIFAR10. Meanwhile, the memory reduction could also reach average 21% and 64%, respectively.
Fuxun Yu, Zhuwei Qin, Di Wang 0003, Ping Xu 0002, Zhi Tian, Xiang Chen 0010
DATE3
2020 Exploring the Design Space of Efficient Deep Neural Networks
abstract
This paper gives an overview of our ongoing work on the design space exploration of efficient deep neural networks (DNNs), specifically on the novel optimization perspectives that past work have mainly overlooked. We cover two complementary aspects of efficient DNN design: (1) static architecture design efficiency and (2) dynamic model execution efficiency. In the static architecture design, one of the major challenges of NAS is the low search efficiency. Different with current mainstream efficient search algorithm optimization, we identify the new perspective in efficient search space design. In the dynamic model execution, current major optimization methods still target at the model structure redundancy, e.g., weight/filter pruning, connection pruning, etc. We instead identify the new dimension of DNN feature map redundancy. By showcasing such new perspectives, further advantages could be potentially attained by integrating both current optimizations and our new perspectives.
Fuxun Yu, Dimitrios Stamoulis, Di Wang 0003, Dimitrios Lymberopoulos, Xiang Chen 0010
SEC3
2019 Single-Path NAS: Designing Hardware-Efficient ConvNets in Less Than 4 Hours
Dimitrios Stamoulis, Ruizhou Ding, Di Wang 0003, Dimitrios Lymberopoulos, Bodhi Priyantha, Jie Liu 0001, Diana Marculescu
ECML/PKDD (2)3
2018 Representing and Recommending Shopping Baskets with Complementarity, Compatibility and Loyalty
abstract
We study the problem of representing and recommending products for grocery shopping. We carefully investigate grocery transaction data and observe three important patterns: products within the same basket complement each other in terms of functionality (complementarity); users tend to purchase products that match their preferences (compatibility); and a significant fraction of users repeatedly purchase the same products over time (loyalty). Unlike conventional e-commerce settings, complementarity and loyalty are particularly predominant in the grocery shopping domain. This motivates a new representation learning approach to leverage complementarity and compatibility holistically, as well as a new recommendation approach to explicitly account for users' 'must-buy' purchases in addition to their overall preferences and needs. Doing so not only improves product classification and recommendation performance on both public and proprietary transaction data covering various grocery store types, but also reveals interesting findings about the relationships between preferences, necessity, and loyalty in consumer purchases.
Mengting Wan, Di Wang 0003, Jie Liu 0001, Paul N. Bennett, Julian J. McAuley
CIKM2
2017 Rain or Shine? - Making Sense of Cloudy Reliability Data
abstract
Cloud datacenters must ensure high availability for the hosted applications and failures can be the bane of datacenter operators. Understanding the what, when and why of failures can help tremendously to mitigate their occurrence and impact. Failures can, however, depend on numerous spatial and temporal factors spanning hardware, workloads, support facilities, and even the environment. One has to rely on failure data from the field to quantify the influence of these factors on failures. Towards this goal, we collect failures data along with many parameters that might influence failures from two large production datacenters with very diverse characteristics. We show that multiple factors simultaneously affect failures, and these factors may interact in non-trivial ways. This makes conventional approaches that study aggregate characteristics or single parameter influences, rather inaccurate. Instead, we build a multi-factor analysis framework to systematically identify influencing factors, quantify their relative impact, and help in more accurate decision making for failure mitigation. We demonstrate this approach for three important decisions: spare capacity provisioning, comparing the reliability of hardware for vendor selection, and quantifying flexibility in datacenter climate control for cost-reliability trade-offs.
Iyswarya Narayanan, Bikash Sharma, Di Wang 0003, Sriram Govindan, Laura Caulfield, Anand Sivasubramaniam, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid
ICDCS3
2017 Optimal Peak Shaving Using Batteries at Datacenters: Characterizing the Risks and Benefits
abstract
A datacenter's power consumption is a major contributor to its operational expenditures (op-ex) and one-time capital expenditures (cap-ex). The recurring electricity cost is often in large determined by datacenter peak-demand under peak-based pricing which is employed by major electric utility providers. There is a growing interest in reducing a datacenter's electricity costs by using throttling techniques and/or energy storage devices (batteries) which are readily available at most datacenters as a backup energy source. A datacenter's power-demand uncertainty makes this a challenging problem, which is largely neglected in existing work, by assuming perfect predictability of power demand. We model this inherent uncertainty as a Markov chain and also evaluate the risk of over/under charging batteries as a result of the randomness in power demand. We design an online optimization framework for peak shaving which considers Conditional Value at Risk and allows for navigating cost-risk trade-offs of datacenters based on their energy infrastructure and workload characteristics. We show that this framework offers significantly higher (up to 2X) cost-savings with small risks of over/under charging batteries, compared to existing stochastic optimization techniques. This framework leverages Markov Decision Processes to perform online dynamic peak shaving, considering battery degradation costs under peak-based pricing.
Neda Nasiriani, George Kesidis, Di Wang 0003
MASCOTS3
2017 Modeling Consumer Preferences and Price Sensitivities from Large-Scale Grocery Shopping Transaction Logs
abstract
In order to match shoppers with desired products and provide personalized promotions, whether in online or offline shopping worlds, it is critical to model both consumer preferences and price sensitivities simultaneously. Personalized preferences have been thoroughly studied in the field of recommender systems, though price (and price sensitivity) has received relatively little attention. At the same time, price sensitivity has been richly explored in the area of economics, though typically not in the context of developing scalable, working systems to generate recommendations. In this study, we seek to bridge the gap between large-scale recommender systems and established consumer theories from economics, and propose a nested feature-based matrix factorization framework to model both preferences and price sensitivities. Quantitative and qualitative results indicate the proposed personalized, interpretable and scalable framework is capable of providing satisfying recommendations (on two datasets of grocery transactions) and can be applied to obtain economic insights into consumer behavior.
Mengting Wan, Di Wang 0003, Matthew Taddy, Justin Rao, Jie Liu 0001, Dimitrios Lymberopoulos, Julian J. McAuley
WWW2
2016 SizeCap: Efficiently handling power surges in fuel cell powered data centers
abstract
Fuel cells are a promising power source for future data centers, offering high energy efficiency, low greenhouse gas emissions, and high reliability. However, due to mechanical limitations related to fuel delivery, fuel cells are slow to adjust to sudden increases in data center power demands, which can result in temporary power shortfalls. To mitigate the impact of power shortfalls, prior work has proposed to either perform power capping by throttling the servers, or to leverage energy storage devices (ESDs) that can temporarily provide enough power to make up for the shortfall while the fuel cells ramp up power generation. Both approaches have disadvantages: power capping conservatively limits server performance and can lead to service level agreement (SLA) violations, while ESD-only solutions must significantly overprovision the energy storage device capacity to tolerate the shortfalls caused by the worst-case (i.e., largest) power surges, which greatly increases the total cost of ownership (TCO). We propose SizeCap, the first ESD sizing framework for fuel cell powered data centers, which coordinates ESD sizing with power capping to enable a cost-effective solution to power shortfalls in data centers. SizeCap sizes the ESD just large enough to cover the majority of power surges, but not the worst-case surges that occur infrequently, to greatly reduce TCO. It then uses the smaller capacity ESD in conjunction with power capping to cover the power shortfalls caused by the worst-case power surges. As part of our new flexible framework, we propose multiple power capping policies with different degrees of awareness of fuel cell and workload behavior, and evaluate their impact on workload performance and ESD size. Using traces from Microsoft's production data center systems, we demonstrate that SizeCap significantly reduces the ESD size (by 85%ofor a workload with infrequent yet large power surges, and by 50% for a workload with frequent power surges) without violating any SLAs.
Yang Li 0183, Di Wang 0003, Saugata Ghose, Jie Liu 0001, Sriram Govindan, Sean James, Eric Peterson, John Siegler, Rachata Ausavarungnirun, Onur Mutlu
HPCA2
2016 SSD Failures in Datacenters: What, When and Why?
abstract
Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, which cannot really help understand how SSD failures can manifest and impact production systems, in order to take appropriate actions. Besides failure data, a detailed characterization requires wide spectrum of data about factors influencing SSD failures, right from provisioning (what models' where and when deployed' etc.) to the operational ones (workloads, read-write intensities, write amplification, etc.). We analyze over half a million SSDs that span multiple generations spread across several datacenters which host a wide range of workloads over nearly 3 years. By studying the diverse set of factors on SSD failures, and their symptoms, our work provides the first look at the what, when and why characteristics of SSD failures in production datacenters.
Iyswarya Narayanan, Di Wang 0003, Myeongjae Jeon, Bikash Sharma, Laura Caulfield, Anand Sivasubramaniam, Ben Cutler, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid
SIGMETRICS2
2016 SSD Failures in Datacenters: What? When? and Why?
abstract
Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, and such information cannot really help understand how SSD failures can manifest and impact the operation of production systems, in order to take appropriate remedial measures. Besides actual failure data and the symptoms exhibited by SSDs before failing, a detailed characterization effort requires wide set of data about factors influencing SSD failures, right from provisioning factors to the operational ones. This paper presents an extensive SSD failure characterization by analyzing a wide spectrum of data from over half a million SSDs that span multiple generations spread across several datacenters which host a wide spectrum of workloads over nearly 3 years. By studying the diverse set of design, provisioning and operational factors on failures, and their symptoms, our work provides the first comprehensive analysis of the what, when and why characteristics of SSD failures in production datacenters.
Iyswarya Narayanan, Di Wang 0003, Myeongjae Jeon, Bikash Sharma, Laura Caulfield, Anand Sivasubramaniam, Ben Cutler, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid
SYSTOR2
2014 Underprovisioning backup power infrastructure for datacenters
abstract
While there has been prior work to underprovision the power distribution infrastructure for a datacenter to save costs, the ability to underprovision the backup power infrastructure, which contributes significantly to capital costs, is little explored. There are two main components in the backup infrastructure - Diesel Generators (DGs) and UPS units - which can both be underprovisioned (or even removed) in terms of their power and/or energy capacities. However, embarking on such underprovisioning mandates studying several ramifications - the resulting cost savings, the lower availability, and the performance and state loss consequences on individual applications - concurrently. This paper presents the first such study, considering cost, availability, performance and application consequences of underprovisioning the backup power infrastructure. We present a framework to quantify the cost of backup capacity that is provisioned, and implement techniques leveraging existing software and hardware mechanisms to provide as seamless an operation as possible for an application within the provisioned backup capacity during a power outage. We evaluate the cost-performance-availability trade-offs for different levels of backup underprovisioning for applications with diverse reliance on the backup infrastructure. Our results show that one may be able to completely do away with DGs, compensating for it with additional UPS energy capacities, to significantly cut costs and still be able to handle power outages lasting as high as 40 minutes (which constitute bulk of the outages). Further, we can push the limits of outage duration that can be handled in a cost-effective manner, if applications are willing to tolerate degraded performance during the outage. Our evaluations also show that different applications react differently to the outage handling mechanisms, and that the efficacy of the mechanisms is sensitive to the outage duration. The insights from this paper can spur new opportunities for future work on backup power infrastructure optimization.
Di Wang 0003, Sriram Govindan, Anand Sivasubramaniam, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib
ASPLOS1
2013 Virtualizing power distribution in datacenters
abstract
Power infrastructure contributes to a significant portion of datacenter expenditures. Overbooking this infrastructure for a high percentile of the needs is becoming more attractive than for occasional peaks. There exist several computing knobs to cap the power draw within such under-provisioned capacity. Recently, batteries and other energy storage devices have been proposed to provide a complementary alternative to these knobs, which when decentralized (or hierarchically placed), can temporarily take the load to suppress power peaks propagating up the hierarchy. With aggressive under-provisioning, the power hierarchy becomes as central a datacenter resource as other computing resources, making it imperative to carefully allocate, isolate and manage this resource (including batteries), across applications. Towards this goal, we present vPower, a software system to virtualize power distribution. vPower includes mechanisms and policies to provide a virtual power hierarchy for each application. It leverages traditional computing knobs as well as batteries, to apportion and manage the infrastructure between co-existing applications in the hierarchy. vPower allows applications to specify their power needs, performs admission control and placement, dynamically monitors power usage, and enforces allocations for fairness and system efficiency. Using several datacenter applications, and a 2-level power hierarchy prototype containing batteries at both levels, we demonstrate the effectiveness of vPower when working in an under-provisioned power infrastructure, using the right computing knobs and the right batteries at the right time. Results show over 50% improved system utilization and scale-out for vPower's over-booking, and between 12-28% better application performance than traditional power-capping control knobs. It also ensures isolation between applications competing for power.
Di Wang 0003, Chuangang Ren, Anand Sivasubramaniam
ISCA1
2013 ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumption
abstract
Peak power management of datacenters has tremendous cost implications. While numerous mechanisms have been proposed to cap power consumption, real datacenter power consumption data is scarce. To address this gap, we collect power demands at multiple spatial and fine-grained temporal resolutions from the load of geo-distributed datacenters of Microsoft over 6 months. We conduct aggregate analysis of this data, to study its statistical properties. With workload characterization a key ingredient for systems design and evaluation, we note the importance of better abstractions for capturing power demands, in the form of peaks and valleys. We identify and characterize attributes for peaks and valleys, and important correlations across these attributes that can influence the choice and effectiveness of different power capping techniques. With the wide scope of exploitability of such characteristics for power provisioning and optimizations, we illustrate its benefits with two specific case studies.
Di Wang 0003, Chuangang Ren, Sriram Govindan, Anand Sivasubramaniam, Bhuvan Urgaonkar, Aman Kansal, Kushagra Vaid
SIGMETRICS1
2013 Aggressive Datacenter Power Provisioning with Batteries
abstract
Datacenters spend $10--25 per watt in provisioning their power infrastructure, regardless of the watts actually consumed. Since peak power needs arise rarely, provisioning power infrastructure for them can be expensive. One can, thus, aggressively underprovision infrastructure assuming that simultaneous peak draw across all equipment will happen rarely. The resulting nonzero probability of emergency events where power needs exceed provisioned capacity, however small, mandates graceful reaction mechanisms to cap the power draw instead of leaving it to disruptive circuit breakers/fuses. Existing strategies for power capping use temporal knobs local to a server that throttle the rate of execution (using power modes), and/or spatial knobs that redirect/migrate excess load to regions of the datacenter with more power headroom. We show these mechanisms to have performance degrading ramifications, and propose an entirely orthogonal solution that leverages existing UPS batteries to temporarily augment the utility supply during emergencies. We build an experimental prototype to demonstrate such power capping on a cluster of 8 servers, each with an individual battery, and implement several online heuristics in the context of different datacenter workloads to evaluate their effectiveness in handling power emergencies. We show that our battery-based solution can: (i) handle emergencies of short durations on its own, (ii) supplement existing reaction mechanisms to enhance their efficacy for longer emergencies, and (iii) create more slack for shifting applications temporarily to nonpeak durations.
Sriram Govindan, Di Wang 0003, Anand Sivasubramaniam, Bhuvan Urgaonkar
ACM Trans. Comput. Syst.2
2012 Leveraging stored energy for handling power emergencies in aggressively provisioned datacenters
abstract
Datacenters spend $10-25 per watt in provisioning their power infrastructure, regardless of the watts actually consumed. Since peak power needs arise rarely, provisioning power infrastructure for them can be expensive. One can, thus, aggressively under-provision infrastructure assuming that simultaneous peak draw across all equipment will happen rarely. The resulting non-zero probability of emergency events where power needs exceed provisioned capacity, however small, mandates graceful reaction mechanisms to cap the power draw instead of leaving it to disruptive circuit breakers/fuses. Existing strategies for power capping use temporal knobs local to a server that throttle the rate of execution (using power modes), and/or spatial knobs that redirect/migrate excess load to regions of the datacenter with more power headroom. We show these mechanisms to have performance degrading ramifications, and propose an entirely orthogonal solution that leverages existing UPS batteries to temporarily augment the utility supply during emergencies. We build an experimental prototype to demonstrate such power capping on a cluster of 8 servers, each with an individual battery, and implement several online heuristics in the context of different datacenter workloads to evaluate their effectiveness in handling power emergencies. We show that: (i) our battery-based solution can handle emergencies of short duration on its own, (ii) supplement existing reaction mechanisms to enhance their efficacy for longer emergencies, and (iii) battery even provide feasible options when other knobs do not suffice.
Sriram Govindan, Di Wang 0003, Anand Sivasubramaniam, Bhuvan Urgaonkar
ASPLOS2
2012 Carbon-Aware Energy Capacity Planning for Datacenters
abstract
Datacenters are facing increasing pressure to cap their carbon footprints at low cost. Recent work has shown the significant environmental benefits of using renewable energy for datacenters by supply-following techniques (workload scheduling, geographical load balancing, etc.) However, all such prior work has only considered on-site renewable generation when numerous other options also exist, which may be superior to on-site renewables for many datacenters. Alternative ways for datacenters to incorporate renewable energy into their overall energy portfolio include: construction of or investment into off-site renewable farms at locations with more abundant renewable energy potential, indirect purchase of renewable energy through buying renewable energy certificates (RECs), purchase of renewable energy products such as power purchase agreements (PPAs) or through third-party renewable providers. We propose a general, optimization-based framework to minimize datacenter costs in the presence of different carbon footprint reduction goals, renewable energy characteristics, policies, utility tariff, and energy storage devices (ESDs). We expect that our work can help datacenter operators make informed decisions about sustainable, renewable-energy-powered IT system design.
Chuangang Ren, Di Wang 0003, Bhuvan Urgaonkar, Anand Sivasubramaniam
MASCOTS2
2012 Energy storage in datacenters: what, where, and how much?
abstract
Energy storage - in the form of UPS units - in a datacenter has been primarily used to fail-over to diesel generators upon power outages. There has been recent interest in using these Energy Storage Devices (ESDs) for demand-response (DR) to either shift peak demand away from high tariff periods, or to shave demand allowing aggressive under-provisioning of the power infrastructure. All such prior work has only considered a single/specific type of ESD (typically re-chargeable lead-acid batteries), and has only employed them at a single level of the power delivery network. Continuing technological advances have provided us a plethora of competitive ESD options ranging from ultra-capacitors, to different kinds of batteries, flywheels and even compressed air-based storage. These ESDs offer very different trade-offs between their power and energy costs, densities, lifetimes, and energy efficiency, among other factors, suggesting that employing hybrid combinations of these may allow more effective DR than with a single technology. Furthermore, ESDs can be placed at different, and possibly multiple, levels of the power delivery hierarchy with different associated trade-offs. To our knowledge, no prior work has studied the extensive design space involving multiple ESD technology provisioning and placement options. This paper intends to fill this critical void, by presenting a theoretical framework for capturing important characteristics of different ESD technologies, the trade-offs of placing them at different levels of the power hierarchy, and quantifying the resulting cost-benefit trade-offs as a function of workload properties.
Di Wang 0003, Chuangang Ren, Anand Sivasubramaniam, Bhuvan Urgaonkar, Hosam K. Fathy
SIGMETRICS1