EDBT 2026 Demo / reviewers in the wild / expert
Sa Wang
dblp:42/8496
· DBLP profile ↗
37ranked-venue papers
4as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RRAtention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context InferenceabstractThe quadratic complexity of attention mechanisms poses a critical bottleneck for large language models processing long contexts. While dynamic sparse attention methods offer input-adaptive efficiency, they face fundamental trade-offs: requiring preprocessing, lacking global evaluation, violating query independence, or incurring high computational overhead. We present RRAttention, a novel dynamic sparse attention method that simultaneously achieves all desirable properties through a head round-robin (RR) sampling strategy. By rotating query sampling positions across attention heads within each stride, RRAttention maintains query independence while enabling efficient global pattern discovery with stride-level aggregation. Our method reduces complexity from O(L^2) to O(L^2/S^2) and employs adaptive Top-\tau selection for optimal sparsity. Extensive experiments on natural language understanding (HELMET) and multimodal video comprehension (Video-MME) demonstrate that RRAttention recovers over 99% of full attention performance while computing only half of the attention blocks, achieving 2.4\times speedup at 128K context length and outperforming existing dynamic sparse attention methods. The code is available at https://github.com/PaddlePaddle/PaddleFleet (see ‘Research/RRAttention‘). Siran Liu, Guoxia Wang, Sa Wang, Jinle Zeng, Haoyang Xie, Siyu Lou, Jiabin Yang, Dianhai Yu |
ACL (1) | 3 |
| 2026 | RaidenSwap: A Multi-Swap Remote System for Multi-core ApplicationsabstractKernel-based remote memory systems are gaining traction in datacenters due to their significant improvement in memory utilization and their ability to transparently provide applications with unlimited memory capacity. However, the high degree of parallelism in contemporary applications leads to a significant demand for remote access throughput, which mismatches with the state-of-the-art kernel swap path due to its inherently limited parallelism. As a result, it cannot scale up the multi-core applications. We dive into the implementation of the swap path and identify the root cause behind it - significant lock contentions and inefficient swap tasks offloading. Kefan Liu, Ke Liu 0004, Xu Zhang 0033, Ning Liu 0031, Sa Wang, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005 |
EuroSys | 7 |
| 2026 | TraceRTL: Agile Performance Evaluation for Microarchitecture ExplorationabstractWhile agile chip development methodologies have accelerated RTL design and simulation, performance evaluation remains constrained by three challenges: (1) inefficient feature prototyping caused by the tight coupling between functional correctness and performance evaluation, particularly for large-scale, error-prone microarchitectures; (2) limited workloads due to incomplete peripheral/software environments or unavailable source code; and (3) time-consuming warm-up phases in sampling-based simulation, required to mitigate cold-start effects. To address these challenges, we propose TRACERTL, an agile, trace-driven performance evaluation methodology that decouples the functional and performance components of CPU RTL designs. It introduces three techniques: (1) a trace-driven performance exploration framework that bypasses full functional correctness while preserving performance accuracy; (2) a trace transformation technique, TraceBridge, that replays traces across different formats and instruction sets; and (3) a fast warm-up strategy, TraceDedup, that eliminates redundant traces and efficiently initializes microarchitectural states. Using TRACERTL, we develop the first trace-driven RTL CPU derived from XiangShan, a high-performance out-of-order RISC-V processor. TRACERTL achieves performance accuracies of 99.87% and 99.86% on SPECint2017 and SPECfp2017, respectively. With TraceBridge, we evaluate x86-based Google workload traces on a RISC-V RTL CPU and reveal distinct memory-bound behavior. TraceDedup further accelerates warm-up phases in sampling-based simulations by$\text{1. 5} \times$to$\text{1 1. 8} \times$. Zifei Zhang 0001, Yinan Xu 0001, Sa Wang, Dan Tang 0002, Yungang Bao |
HPCA | 3 |
| 2026 | TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile VerificationabstractVerification is a critical process for ensuring the correctness of modern processors. The increasing complexity of processor designs and the emergence of new instruction set architectures (ISAs) like RISC-V have created demands for more agile and efficient verification methodologies, particularly regarding verification efficiency and faster coverage convergence. While simulation-based approaches now attempt to incorporate advanced software testing techniques such as fuzzing to improve coverage, they face significant limitations when applied to processor verification, notably poor performance and inadequate test case quality. Hardware-accelerated solutions using FPGA or ASIC platforms have tried to address these issues, yet they struggle with challenges including host-FPGA communication overhead, inefficient test pattern generation, and suboptimal implementation of the entire multi-step verification process. In this paper, we present TurboFuzz, an end-to-end hardwareaccelerated verification framework that implements the entire Test Generation-Simulation-Coverage Feedback loop on a single FPGA for modern processor verification. TurboFuzz enhances test quality through optimized test case (seed) control flow, efficient inter-seed scheduling, and hybrid fuzzer integration, thereby improving coverage and execution efficiency. Additionally, it employs a feedback-driven generation mechanism to accelerate coverage convergence. Experimental results show that TurboFuzz achieves up to 2.23× more coverage collection than software-based fuzzers within the same time budget, and up to 571× performance speedup when detecting real-world issues, while maintaining full visibility and debugging capabilities with moderate area overhead. Xueqi Li 0001, Sa Wang, David Boland, Yungang Bao, Kan Shi |
HPCA | 4 |
| 2026 | Democratizing and Accelerating Hardware Verification with Software-Native Optimization
Yunlong Xie, Zhicheng Yao, Fangyuan Song, Junyue Wang, Haojin Tang, Yinan Xu 0001, Ziyuan Gao, Duan Yu, Jiayi Rao, Junyu Yue, Yunqi Lu, Zechen Yang, Xu An, Qi Ge, Jiuyue Ma, Jian-Yi Meng, Kan Shi, Dan Tang 0002, Sa Wang, Yungang Bao |
ISCA | 27 |
| 2026 | SketchPlan: Full-Visibility Sketch-Based Telemetry with Limited Programmable Switch Coverage
Jinbo Sun, Haifeng Sun 0004, Jintao He, Qun Huang 0001, Sa Wang, Yungang Bao |
IWQoS | 5 |
| 2025 | Poby: SmartNIC-accelerated Image Provisioning for Coldstart in Clouds
Zihao Chang, Haifeng Sun 0004, Yunlong Xie, Kan Shi, Ninghui Sun, Yungang Bao, Sa Wang |
USENIX ATC | 8 |
| 2025 | AEPSO: An adaptive learning particle swarm optimization for solving the hyperparameters of dynamic periodic regulation grey model
Gang Hu 0002, Sa Wang, Bin Shu, Guo Wei 0004 |
Expert Syst. Appl. | 2 |
| 2025 | Particle swarm optimization for hybrid mutant slime mold: An efficient algorithm for solving the hyperparameters of adaptive Grey-Markov modified model
Gang Hu 0002, Sa Wang, Jiulong Zhang, Essam H. Houssein |
Inf. Sci. | 2 |
| 2025 | ShuffleInfer: Disaggregate LLM Inference for Mixed Downstream WorkloadsabstractTransformer-based large language model (LLM) inference serving is now the backbone of many cloud services. LLM inference consists of a prefill phase and a decode phase. However, existing LLM deployment practices often overlook the distinct characteristics of these phases, leading to significant interference. To mitigate interference, our insight is to carefully schedule and group inference requests based on their characteristics. We realize this idea in ShuffleInfer through three pillars. First, it partitions prompts into fixed-size chunks so that the accelerator always runs close to its computation-saturated limit. Second, it disaggregates prefill and decode instances so each can run independently. Finally, it uses a smart two-level scheduling algorithm augmented with predicted resource usage to avoid decode scheduling hotspots. Results show that ShuffleInfer improves time-to-first-token (TTFT), job completion time (JCT), and inference efficiency in terms of performance per dollar by a large margin, e.g., it uses 38% less resources all the while lowering average TTFT and average JCT by 97% and 47%, respectively. Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Chenxi Wang 0005, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan |
ACM Trans. Archit. Code Optim. | 9 |
| 2024 | HAPPIES: a History-Aware Efficient Cloud Resource Overcommitment SystemabstractImproving resource utilization in datacenters is vital for reducing costs for cloud service providers (CSPs). Increasing resource utilization must be balanced with maintaining quality of service (QoS) for latency-critical applications. In cloud environments, users often request excessive resources for applications to ensure QoS. To address this issue, CSPs use resource overcommitment - offering users resources that exceed the actual capacity of physical infrastructure. However, if not properly managed, such strategies may result in performance degradation or even request failure. Therefore, to achieve optimal resource utilization while maintaining QoS to applications, it is critical to implement a fine-grained overcommitment strategy.We propose HAPPIES, a History-aware management system with a precise prediction for machine resource demand. HAPPIES uses historical usage to extract resource characteristics and build application portraits that describe their resource demands. Compared to the existing strategy, this is a more aggressive overcommitment strategy that achieves higher resource utilization. We simulated experiments on 3,021 nodes and deployed over 14,000 applications on them. Results show that HAPPIES significantly outperforms Kubernetes Least Request and Peak Oracle in load balancing. Not only does it reduce the number of nodes experiencing high utilization, but it also decreases the peak usage of the most heavily utilized nodes. Therefore, HAPPIES scheduling reduces the risk of a machine being used beyond capacity. Ziwei Huang 0003, Shibo Tang, Zihao Chang, Qichao Lu, Jian Ouyang, Wenbin Lv, Zhicheng Yao, Yungang Bao, Sa Wang |
CCGrid | 10 |
| 2024 | INS: Identifying and Mitigating Performance Interference in Clouds via Interference-Sensitive PathsabstractIdentifying and managing performance interference in clouds has long been a critical and challenging task for cloud providers. They keep seeking useful performance indicators from underlying systems to monitor cloud applications accurately. However, state-of-the-art indicators are either sensitive to limited applications and resource contention or are unrobust to the continually changing production environments. There still lacks a practical and efficient indicator for production environments. Ziwei Huang 0003, Mengyao Xie, Shibo Tang, Zihao Chang, Zhicheng Yao, Yungang Bao, Sa Wang |
SoCC | 7 |
| 2024 | PathFuzz: Broadening Fuzzing Horizons with Footprint Memory for CPUsabstractCoverage metrics have been widely adopted to quantify the completeness of hardware verification. Recently, coverage-guided fuzzing has emerged as a popular method for automatically creating test inputs toward higher verification coverage reach. However, we observe that its effectiveness on CPUs is hindered by limited sources of seed corpus and efficiency of mutations. To broaden the fuzzing horizons, this paper proposes the PathFuzz framework incorporating an efficient input format for fuzzing CPUs, the footprint memory, with seed corpus from real-world large-scale programs. Experiments demonstrate that using PathFuzz reaches over 95% verification coverage with four long-standing bugs newly identified in two well-known open-source CPU designs. Yinan Xu 0001, Sa Wang, Dan Tang 0002, Ninghui Sun, Yungang Bao |
DAC | 2 |
| 2024 | Aceso: Efficient Parallel DNN Training through Iterative Bottleneck AlleviationabstractMany parallel mechanisms, including data parallelism, tensor parallelism, and pipeline parallelism, have been proposed and combined together to support training increasingly large deep neural networks (DNN) on massive GPU devices. Given a DNN model and GPU cluster, finding the optimal configuration by combining these parallelism mechanisms is an NP-hard problem. Widely adopted mathematical programming approaches have been proposed to search in a configuration subspace, but they are still too costly when scaling to large models over numerous devices. Youshan Miao, Xiaoxiang Shi, Saeed Maleki, Fan Yang 0024, Yungang Bao, Sa Wang |
EuroSys | 8 |
| 2024 | LazyCAT: Efficient Fine-Grained Cache Partitioning with Two BoundariesabstractIntel CAT is a widely available cache partitioning technique in commercial hardware but falls short in partitioning granularity. We propose LazyCAT, a fine-grained, on-demand, and easy-to-use cache partitioning technique, which not only can ensure the QoS of High-Priority (HP) applications but also yield the under-utilized cache sets to other Best-Effort (BE) applications for better resource efficiency. LazyCAT retains the easy-to-use philosophy of CAT and introduces a new soft LLC partitioning boundary, which is lower than the original CAT partitioning boundary (hard boundary). LazyCAT detects and selects under-utilized cache sets in HP applications during profiling, and specifies them to the soft boundary at runtime dynamically, yielding the cache blocks between these two boundaries for other applications. Meanwhile, LazyCAT provides users with a simple software interface to guide the set-level space allocation according to their needs. Experimental results show that LazyCAT exhibits substantial performance improvements (up to 12.2%) for BE applications with less than 3% performance degradation of HP applications. Chuanqi Zhang, Xueqi Li 0001, Ninghui Sun, Yungang Bao, Sa Wang |
HPCC | 5 |
| 2024 | PP-Stream: Toward High-Performance Privacy-Preserving Neural Network Inference via Distributed Stream ProcessingabstractPrivacy preservation is critical for neural network inference, which often involves collaborative execution of different parties to make predictions on sensitive data based on sensitive neural network models. However, the expensive cryptographic operations of privacy preservation also pose performance chal-lenges to neural network inference. We address this performance-security tension by designing PP-Stream, a distributed stream processing system for high-performance privacy-preserving neural network inference. PP-Stream adopts hybrid privacy-preserving mechanisms for linear and non-linear operations of neural network inference. It treats inference data as real-time data streams, and parallelizes the inference operations across multiple pipelined stages that are executed by multiple servers and threads. It also solves the load-balanced resource allocation across servers and threads as an optimization problem. We prototype PP-Stream and show via testbed experiments that it achieves low inference latencies on various neural network models. Qingxiu Liu, Qun Huang 0001, Xiang Chen 0017, Sa Wang, Shujie Han 0001, Patrick P. C. Lee |
ICDE | 4 |
| 2024 | Tentacles: A Middleware with Multi-Network Communication Reliability for Vehicle-Infrastructure Cooperative Autonomous DrivingabstractVehicle-Infrastructure Cooperative Autonomous Driving (CAD) is a new paradigm of autonomous driving, which relies on the cooperation between intelligent roads and autonomous vehicles. This paradigm has been shown to be safer and more efficient compared to the on-vehicle-only autonomous driving paradigm. Our real-world deployment data indicate that the effectiveness of Vehicle-Infrastructure CAD is constrained by the reliability and performance of commercial communication networks. This paper targets this exact problem and proposes Tentacles, a middleware to achieve high communication reliability between intelligent roads and autonomous vehicles, in the context of Vehicle-Infrastructure CAD. Specifically, Tentacles dynamically matches Vehicle-Infrastructure CAD applications and the underlying communication technologies based on varying communication performance and quality needs. Evaluation results confirm that Tentacles reduces deadline violations by more than 88%, significantly improving the reliability of Vehicle-Infrastructure CAD systems. Tianze Wu, Sa Wang, Yungang Bao, Weisong Shi |
VTC Fall | 2 |
| 2024 | Panoptic Segmentation with Convex Object RepresentationabstractAbstract The accurate representation of objects holds pivotal significance in the realm of panoptic segmentation. Presently, prevalent object representation methodologies, including box-based, keypoint-based and query-based techniques, encounter a challenge known as the ‘representation confusion’ issue in specific scenarios, often resulting in the mislabeling of instances. In response, this paper introduces Convex Object Representation (COR), a straightforward yet highly effective approach to address this problem. COR leverages a CNN-based Euclidean Distance Transform to convert the target instance into a convex heatmap. Simultaneously, it offers a parallel embedding method for encoding the object. Subsequently, COR characterizes objects based on the distinctive embedding vectors of their convex vertices. This paper seamlessly integrates COR into a state-of-the-art query-based panoptic segmentation framework. Experimental findings validate that COR successfully mitigates the representation confusion predicament, enhancing segmentation accuracy. The COR-augmented methods exhibit notable improvements of +1.3 and +0.7 points in PQ on the Cityscapes validation and MS COCO panoptic 2017 validation datasets, respectively. Zhicheng Yao, Sa Wang, Jinbin Zhu, Yungang Bao |
Comput. J. | 2 |
| 2024 | Function Interaction Risks in Robot Apps: Analysis and Policy-Based SolutionabstractRobot apps are becoming more automated, complex and diverse. An app usually consists of many functions, interacting with each other and the environment. This allows robots to conduct various tasks. However, it also opens a new door for cyber attacks: adversaries can leverage these interactions to threaten the safety of robot operations. Unfortunately, this issue is rarely explored in past works. We present thefirstsystematic investigation about the function interactions in common robot apps. First, we disclose the potential risks and damages caused by malicious interactions. We introduce a comprehensive graph to model the function interactions in robot apps by analyzing 3,100 packages from the Robot Operating System (ROS) platform. From this graph, we identify and categorize three types of interaction risks. Second, we propose novel methodologies to detect and mitigate these risks and protect the operations of robot apps. We introduce security policies for each type of risks, and design coordination nodes to enforce the policies and regulate the interactions. We conduct extensive experiments on 110 robot apps from the ROS platform and two complex apps (Baidu Apollo and Autoware) widely adopted in industry. Evaluation results showed our methodologies can correctly identify and mitigate all potential risks. Yuan Xu 0033, Yungang Bao, Sa Wang, Tianwei Zhang 0004 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | Skadi: Building a Distributed Runtime for Data Systems in Disaggregated Data CentersabstractData-intensive systems are the backbone of today's computing and are responsible for shaping data centers. Over the years, cloud providers have relied on three principles to maintain cost-effective data systems: use disaggregation to decouple scaling, use domain-specific computing to battle waning laws, and use serverless to lower costs. Although they work well individually, they fail to work in harmony: an issue amplified by emerging data system workloads. Cunchen Hu, Chenxi Wang 0005, Sa Wang, Ninghui Sun, Yungang Bao, Jieru Zhao, Sanidhya Kashyap, Pengfei Zuo, Xusheng Chen, Liangliang Xu, Yizhou Shan |
HotOS | 3 |
| 2023 | Functional Verification for Agile Processor Development: A Case for Workflow Integration
Yinan Xu 0001, Kaifan Wang, Huaqiang Wang, Linjuan Zhang, Zifei Zhang 0001, Dan Tang 0002, Sa Wang, Kan Shi, Ninghui Sun, Yungang Bao |
J. Comput. Sci. Technol. | 10 |
| 2022 | Towards Developing High Performance RISC-V Processors Using Agile MethodologyabstractWhile research has shown that the agile chip design methodology is promising to sustain the scaling of computing performance in a more efficient way, it is still of limited usage in actual applications due to two major obstacles: 1) Lack of tool-chain and developing framework supporting agile chip design, especially for large-scale modern processors. 2) The conventional verification methods are less agile and become a major bottleneck of the entire process. To tackle both issues, we propose MINJIE, an open-source platform supporting agile processor development flow. MINJIE integrates a broad set of tools for logic design, functional verification, performance modelling, pre-silicon validation and debugging for better development efficiency of state-of-the-art processor designs. We demonstrate the usage and effectiveness of MINJIE by building two generations of an open-source superscalar out-of-order RISC-V processor code-named XIANGSHAN using agile methodologies. We quantify the performance of XIANGSHAN using SPEC CPU2006 benchmarks and demonstrate that XIANGSHAN achieves industry-competitive performance. Yinan Xu 0001, Dan Tang 0002, Guokai Chen, Lingrui Gou, Qianruo Li, Zuojun Li, Jiazhan Tan, Huaqiang Wang, Huizhe Wang, Kaifan Wang, Chuanqi Zhang, Fawang Zhang, Linjuan Zhang, Zifei Zhang 0001, Yaoyang Zhou, Yike Zhou, Jiangrui Zou, Ye Cai 0001, Dandan Huan, Zusong Li, Jiye Zhao, Qiyuan Quan, Xingwu Liu, Sa Wang, Kan Shi, Ninghui Sun, Yungang Bao |
MICRO | 34 |
| 2022 | NfvInsight: A Framework for Automatically Deploying and Benchmarking VNF Chains
Tianni Xu, Haifeng Sun 0004, Xiao-Ming Zhou, Xiufeng Sui, Sa Wang, Qun Huang 0001, Yungang Bao |
J. Comput. Sci. Technol. | 6 |
| 2021 | SEER: A Time Prediction Model for CNNs from GPU Kernel's ViewabstractWith the deepening of research and increasing size of data sets, deep neural networks have become larger and larger. To reduce the training time of large neural networks, researchers propose to optimize neural networks from different levels. When performing optimizations, prior knowledge about execution time of each part of the network can help avoid repeatedly time-consuming testing and profiling process. However it is quite challenging to build an accurate iteration time prediction model, due to opaque underlying implementation of network operators and complex architecture of accelerators. In this paper, we propose SEER, an iteration time prediction model for CNNs, targeting on GPU platforms. We propose to categorize convolution kernels into three different types: Compute-bound, DRAM-bound and Under-utilized, then we build performance model for each type respectively. We combined analytical models and learning-based models to make the performance model accurate and in line with GPU execution model. Experimental results show that, our model achieves 14.71% prediction error on convolution kernels and up to 1.79% prediction error for the overall computation time in one iteration of common CNNs. Besides, when used for selecting the best convolution algorithm, our model shows 7.14% lower error rate than cuDNN's official algorithm picker. Sa Wang, Yungang Bao |
PACT | 2 |
| 2021 | Omegaflow: a high-performance dependency-based architectureabstractThis paper investigates how to better track and deliver dependency in dependency-based cores to exploit instruction-level parallelism (ILP) as much as possible. To this end, we first propose an analytical performance model for the state-of-art dependency-based core, Forwardflow, and figure out two vital factors affecting its upper bound of performance. Then we propose Omegaflow,a dependency-based architecture adopting three new techniques, which respond to the discovered factors. Experimental results show that Omegaflow improves IPC by 24.6% compared to the state-of-the-art design, approaching the performance of the OoO architecture with an ideal scheduler (94.4%) without increasing the clock cycle and consumes only 8.82% more energy than Forwardflow. Yaoyang Zhou, Chuanqi Zhang, Yinan Xu 0001, Huizhe Wang, Sa Wang, Ninghui Sun, Yungang Bao |
ICS | 6 |
| 2021 | Towards Practical Cloud Offloading for Low-cost Ground Vehicle WorkloadsabstractLow-cost Ground Vehicles (LGVs) have been widely adopted to conduct various tasks in our daily life. However, the limited on-board battery capacity and computation resources prevent LGVs from taking more complex and intelligent workloads. A promising approach is to offload the computation from local LGVs to remote servers. However, current cloud-robotic research and platforms are still at a very early stage. Compared to other systems and devices, optimizing LGV workload offloading faces more challenges, such as the uncertainty of environments and the mobility feature of devices.In this paper, we explore the opportunities of optimizing cloud offloading of LGV workloads from the perspectives of performance, energy efficiency and network robustness. We first build an analytical model to reveal the computation role and impact of each function in LGV workloads. Then we propose several optimization strategies (fine-grained migration, cloud acceleration, real-time monitoring and adjustment) to accelerate workload computation, reduce on-board energy consumption, and increase the network robustness. We implement an end-to-end cloud-robotic framework with such strategies to achieve dynamic and adaptive offloading. Evaluations on physical LGVs show that our strategies can significantly reduce the total energy consumption by 2.12× and mission completion time by 2.53×, and maintain strong robust ness under poor network quality. Yuan Xu 0033, Tianwei Zhang 0004, Jimin Han, Sa Wang, Yungang Bao |
IPDPS | 4 |
| 2021 | PR-Sketch: Monitoring Per-key Aggregation of Streaming Data with Nearly Full AccuracyabstractComputing per-key aggregation is indispensable in streaming data analysis formulated as two phases, an update phase and a recovery phase. As the size and speed of data streams rise, accurate per-key information is useful in many applications like anomaly detection, attack prevention, and online diagnosis. Even though many algorithms have been proposed for per-key aggregation in stream processing, their accuracy guarantees only cover a small portion of keys. In this paper, we aim to achieve nearly full accuracy with limited resource usage. We follow the line of sketch-based techniques. We observe that existing methods suffer from high errors for most keys. The reason is that they track keys by complicated mechanism in the update phase and simply calculate per-key aggregation from some specific counter in the recovery phase. Therefore, we present PR-Sketch, a novel sketching design to address the two limitations. PR-Sketch builds linear equations between counter values and per-key aggregations to improve accuracy, and records keys in the recovery phase to reduce resource usage in the update phase. We also provide an extension called fast PR-Sketch to improve processing rate further. We derive space complexity, time complexity, and guaranteed error probability for both PR-Sketch and fast PR-Sketch. We conduct trace-driven experiments under 100K keys and 1M items to compare our algorithms with multiple state-of-the-art methods. Results demonstrate the resource efficiency and nearly full accuracy of our algorithms. Siyuan Sheng, Qun Huang 0001, Sa Wang, Yungang Bao |
Proc. VLDB Endow. | 3 |
| 2020 | A Software Stack for Composable Cloud Robotics System
Tianwei Zhang 0004, Sa Wang, Yungang Bao |
ICA3PP (2) | 3 |
| 2020 | Leveraging 3D blendshape for facial expression recognition using CNN
Sa Wang, Zhengxin Cheng, Xiaoming Deng 0001, Liang Chang 0001, Fuqing Duan, Ke Lu 0002 |
Sci. China Inf. Sci. | 1 |
| 2020 | A Case for Adaptive Resource Management in Alibaba Datacenter Using Neural Networks
Sa Wang, Yan-Hai Zhu, Shan-Pei Chen, Tianze Wu, Wen-Jie Li, Xusheng Zhan, Haiyang Ding, Weisong Shi, Yungang Bao |
J. Comput. Sci. Technol. | 1 |
| 2019 | QoSMT: supporting precise performance control for simultaneous multithreading architectureabstractSimultaneous multithreading (SMT) technology improves CPU throughput, but also causes unpredictable performance fluctuations for co-located workloads. Although recent major SMT processors have adopted some techniques to promote hardware support for quality-of-service (QoS), achieving both precise performance control and high throughput on SMT architectures is still a challenging open problem. Yaoyang Zhou, Xusheng Zhan, Huizhe Wang, Sa Wang, Ningmei Yu, Ninghui Sun, Yungang Bao |
ICS | 7 |
| 2019 | Who limits the resource efficiency of my datacenter: an analysis of Alibaba datacenter tracesabstractCloud platform provides great flexibility and cost-efficiency for end-users and cloud operators. However, low resource utilization in modern datacenters brings huge wastes of hardware resources and infrastructure investment. To improve resource utilization, a straightforward way is co-locating different workloads on the same hardware. To figure out the resource efficiency and understand the key characteristics of workloads in co-located cluster, we analyze an 8-day trace from Alibaba's production trace. We reveal three key findings as follows. First, memory becomes the new bottleneck and limits the resource efficiency in Alibaba's datacenter. Second, in order to protect latency-critical applications, batch-processing applications are treated as second-class citizens and restricted to utilize limited resources. Third, more than 90% of latency-critical applications are written in Java applications. Massive self-contained JVMs further complicate resource management and limit the resource efficiency in datacenters. Zihao Chang, Sa Wang, Haiyang Ding, Yihui Feng, Yungang Bao |
IWQoS | 3 |
| 2019 | Practices of backuping homomorphically encrypted databases
Sa Wang, Yiwen Shao, Yungang Bao |
Frontiers Comput. Sci. | 1 |
| 2017 | Labeled von Neumann Architecture for Software-Defined Cloud
Yungang Bao, Sa Wang |
J. Comput. Sci. Technol. | 2 |
| 2016 | Hug the Elephant: Migrating a Legacy Data Analytics Application to Hadoop EcosystemabstractBig data applications that rely on relational databases gradually expose limitations on scalability and performance. In recent years, Hadoop ecosystem has been widely adopted as an evolving solution. This paper presents the migration of a legacy data analytics application in a provincial data center. The target platform follows "no one size fits all" method. Considering different workloads, data storage is hybrid with distributed file system (HDFS) and distributed NoSQL database. Beyond the architecture re-design, we focus on the problem of data model transformation from relational database to NoSQL database. We propose a query-aware approach to free developers from tedious manual work. The approach generates query-specific views (NoView) for NoSQL and re-structures the views to align with NoSQL's data model. Our results show that the migrated application achieves high scalability and high performance. We believe that our practice provides valuable insights (such as NoSQL data modeling methodology), and the techniques can be easily applied to other similar migrations. Jie Liu 0008, Sa Wang, Lijie Xu, Jixin Ren, Dan Ye 0004, Jun Wei 0001, Tao Huang 0001 |
ICSME | 3 |
| 2015 | VMon: Monitoring and Quantifying Virtual Machine Interference via Hardware Performance CounterabstractVirtualization greatly improves resource utilization in IaaS platforms, but it also introduces potential interference between virtual machines (VMs). For example, VMs may suffer from performance degradation, when they are located in one host and compete for sharing physical resources. Thus, how to efficiently monitor and quantify the VMs interference becomes a key challenge for IaaS providers. In this paper, we present Vmon, a system to transparently monitor and quantify the interference between VMs with the hardware performance counters (HPCs). By collecting the HPCs of different VMs and exploring the LLC miss rates within HPCs, Vmon analyzes the relationship between the LLC miss rates and VM performance degradation to predict the interference between different resource-intensive VMs, and mitigate the VMs interference. The experimental results show that Vmon predicts the performance degradation in the accuracy of more than 90% with less than 10% performance overhead. Sa Wang, Wenbo Zhang 0006, Tao Wang 0030, Chunyang Ye, Tao Huang 0001 |
COMPSAC | 1 |
| 2011 | Bench4Q: A QoS-Oriented E-Commerce BenchmarkabstractE-commerce systems are typically QoS-sensitive, so QoS-oriented tunings of e-commerce servers are very important for such systems. However, existing e-commerce benchmarks are insufficient for supporting QoS-oriented tunings, because some critical QoS features of e-commerce systems cannot be precisely evaluated by them. One example of these features is the integrality of service, which is usually expressed as a session, provided to customers. This paper presents a QoS-oriented e-commerce benchmark, which is named Bench4Q and is an extension of TPC-W supporting QoS-oriented tuning of e-commerce servers. The main features of Bench4Q include: (1) supporting session-based metrics analysis and (2) simulating QoS-sensitive load for QoS-oriented capacity analysis. We illustrate the promising benefits of these features for QoS-oriented tuning of an e-commerce server by a series of Bench4Q benchmarking on a typical e-commerce server. Wenbo Zhang 0006, Sa Wang, Wei Wang 0049, Hua Zhong 0007 |
COMPSAC | 2 |