EDBT 2026 Demo / reviewers in the wild / expert
Zuoning Chen
dblp:17/9580
· DBLP profile ↗
36ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0003-1975-5414ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 2 since 2021Security and privacy · 3 · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Approx-L: An Error-Balanced Approximate Floating-Point Divider with Multi-Level Linear Compensation
Shangshang Yao, Huidong Ji, Zuoning Chen |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | DyTopK: Accelerating Top-K SpMV on Embedded FPGAs via Dynamic Floating-Point QuantizationabstractTop-K sparse matrix–vector multiplication is the computational backbone of modern recommendation systems and graph neural networks (GNNs). While shifting these workloads to the edge offers the potential for ultralow latency and enhanced privacy, it presents significant architectural challenges. Recent HBM-based field-programmable gate array (FPGA) accelerators deployed in data centers have shown promise, but their designs are ill-suited for power-constrained embedded FPGAs, which suffer from the limited bandwidth of standard DDR interfaces and restricted on-chip logic resources. In this article, we propose DyTopK, a hardware–software codesigned accelerator tailored specifically for embedded edge devices. Unlike prior approaches that rely on coarse-grained blockwise quantization or complex filtering, DyTopK introduces a novel Dynamic FP4/FP8 Quantization scheme. This mechanism adapts the numerical format at the granularity of individual nonzero elements. To maximize the utility of restricted off-chip bandwidth, we further propose a bandwidth-maximal sparse matrix compression format named packet CSR (PCSR), perfectly aligning with the 64-bit data bus width of standard DDR4 interfaces. Complementing these data-centric optimizations, we design a high-efficiency hardware architecture featuring a scalable processing element array, allowing for flexible expansion based on available logic resources. Experimental results demonstrate that DyTopK achieves significant improvements, delivering$12.4\times $speedup in effective bandwidth utilization and$8.6\times $higher computational throughput compared to state-of-the-art CPU baselines. Furthermore, it outperforms prior FPGA implementations by$2.3\times $in energy efficiency while maintaining negligible accuracy loss (less than 0.5% recall drop) at K = 100. Shangshang Yao, Zuoning Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Behaviour-diverse automatic penetration testing: a coverage-based deep reinforcement learning approach
Yizhou Yang, Longde Chen, Lanning Wang, Haohuan Fu, Xin Liu 0081, Zuoning Chen |
Frontiers Comput. Sci. | 7 |
| 2025 | UKFaaS: Lightweight, High-Performance and Secure FaaS Communication With UnikernelabstractUnikernel is a promising runtime for serverless computing with its lightweight and isolated architecture. It offers a secure and efficient environment for applications. However, famous serverless frameworks like Knative have introduced heavyweight component sidecars to assist function instance deployment in a non-intrusive manner. But the sidecar not only hinders the throughput of unikernel function services but also consumes excessive memory resources. Moreover, the intricate network communication pathways among various services pose significant challenges for deploying unikernels in production serverless environments. Although shared-memory based communication on the same server can solve the communication bottleneck of unikernel-based function instances. The situation where malicious programs on the server make the shared memory untrustworthy limits the deployment of such technologies.We propose UKFaaS, a lightweight and high-performance serverless framework. UKFaaS leverages the advantages of customized operating systems through unikernel and it non-intrusively integrates sidecar functionality into the unikernel, avoiding the overhead of sidecar request forwarding. Additionally, UKFaaS innovatively implements data communication between unikernels in the same server to eliminate VM-Exit bottlenecks in RPC (remote process call) based on VMFUNC without relying on memory sharing. The preliminary experimental results indicate that UKFaaS can realize 1.8×-3.5× request throughput per second (RPS) compared with the advanced serverless system FaasFlow, UaaF and Nightcore in the Google online boutique microservice benchmark. Zhenqian Chen, Yuchun Zhan, Xinkui Zhao, Muyu Yang, Siwei Tan, Lufei Zhang, Liqiang Lu, Jianwei Yin, Zuoning Chen |
IEEE Trans. Computers | 10 |
| 2025 | AdaptHM: A Fully Adaptive Data Migration Strategy for Hybrid Memory SystemsabstractData migration strategies (DMS) improve the overall performance of hybrid memory systems by migrating frequently accessed (hot) data to faster memory. However, designing an efficient DMS is challenging since the key metrics of DMS -hot data selection, migration granularity, and migration frequency -are sensitive to access patterns of workloads. Most existing strategies focus on only one of these metrics and often overlook the crucial impact of access patterns, resulting in sub-optimal performance and unnecessary migration traffic. In this paper, we propose AdaptHM, a fully access-pattern-aware Adaptive data migration strategy for Hybrid Memory systems. AdaptHM achieves adaptability on all three metrics through its unique multi-level data framework. First, AdaptHM adopts a group-level competition policy to select hot blocks, which responds faster to access patterns than threshold-based policies. Second, AdaptHM enables segment-level dynamic migration granularity by decoupling migration from remapping, which shows better access pattern resilience than existing schemes with fixed-size global migration granularity. Third, AdaptHM adjusts the migration frequency at set-level by periodically assessing the migration benefit, avoiding unnecessary migrations. Experimental results demonstrate that AdaptHM improves the performance by an average of 12.78% and reduces energy consumption by up to 37.24% compared to the state-of-the-art scheme. Zhouxuan Peng, Dan Feng 0001, Jianxi Chen, Yachun Liu, Jinlei Hu, Jintong Zhang, Tianyu Wan, Zuoning Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2024 | Scaling Up Memory Disaggregated Applications with SMARTabstractRecent developments in RDMA networks are leading to the trend of memory disaggregation. However, the performance of each compute node is still limited by the network, especially when it needs to perform a large number of concurrent fine-grained remote accesses. According to our evaluations, existing IOPS-bound disaggregated applications do not scale well beyond 32 cores, and therefore do not take full advantage of today's many-core machines. Feng Ren, Kang Chen 0001, Huaxia Xia, Zuoning Chen, Yongwei Wu 0001 |
ASPLOS (1) | 5 |
| 2023 | HadaFS: A File System Bridging the Local and Shared Burst Buffer for Exascale Supercomputers
Xiaobin He, Bin Yang 0043, Shupeng Shi, Dexun Chen, Wei Xue 0003, Zuoning Chen |
FAST | 10 |
| 2023 | Scalability and efficiency challenges for the exascale supercomputing system: practice of a parallel supporting environment on the Sunway exascale prototype systemabstractWith the continuous improvement of supercomputer performance and the integration of artificial intelligence with traditional scientific computing, the scale of applications is gradually increasing, from millions to tens of millions of computing cores, which raises great challenges to achieve high scalability and efficiency of parallel applications on super-large-scale systems. Taking the Sunway exascale prototype system as an example, in this paper we first analyze the challenges of high scalability and high efficiency for parallel applications in the exascale era. To overcome these challenges, the optimization technologies used in the parallel supporting environment software on the Sunway exascale prototype system are highlighted, including the parallel operating system, input/output (I/O) optimization technology, ultra-large-scale parallel debugging technology, 10-million-core parallel algorithm, and mixed-precision method. Parallel operating systems and I/O optimization technology mainly support large-scale system scaling, while the ultra-large-scale parallel debugging technology, 10-million-core parallel algorithm, and mixed-precision method mainly enhance the efficiency of large-scale applications. Finally, the contributions to various applications running on the Sunway exascale prototype system are introduced, verifying the effectiveness of the parallel supporting environment design. Xiaobin He, Xin Chen 0023, Xin Liu 0081, Dexun Chen, Yuling Yang, Yunlong Feng, Longde Chen, Xiaona Diao, Zuoning Chen |
Frontiers Inf. Technol. Electron. Eng. | 11 |
| 2022 | Haica: A High Performance Computing & Artificial Intelligence Fused Computing Architecture
Zhengbo Chen, Fang Zheng 0007, Zuoning Chen |
ICA3PP | 5 |
| 2022 | DFBuffer: High-performance data forwarding software optimized for single-process I/O scenariosabstractMost supercomputers adopt a data forwarding architecture to achieve storage scalability. However, it results in a significant reduction in single-process bandwidth compared to direct file system access. Moreover, considering that a majority of applications uses only a single process for writing and reading data, the low single-process performance also leads to a time overhead for these applications. This paper proposes an userspace forwarding mechanism DFBUFFER with two performance optimization methods: user-space multi-thread request processing and data write buffer in a unit of file. The client of DFBUFFER is embedded in the application as a library reducing the software overhead, and the server implements multi-thread I/O request processing to improve bandwidth efficiency. The data write buffer can asynchronously handle write requests, which accelerates the write bandwidth of compute nodes. We evaluate DFBUFFER on the Sunway exascale prototype system. The results indicate that in the regular mode of DFBUFFER, both the write and read latency are reduced, and the write bandwidth and large-block read bandwidth of single-process are increased by 1.8 times and 2.8 times respectively. The DFBUFFER buffer mode increases the write bandwidth of a single process by 0.8 times over the regular mode. Although the performance advantage of the regular mode of DFBUFFER gradually weakens with the increase of concurrent processes, the DFBUFFER buffer mode has the effect of improving the write bandwidth, the 64-IO-processes application is increased by 0.2 times. Xiaobin He, Bin Yang 0043, Zuoning Chen |
ICPADS | 6 |
| 2022 | DSSA: Dual-Side Sparse Systolic Array Architecture for Accelerating Convolutional Neural Network TrainingabstractEver-growing CNN size incurs a significant amount of redundancy in model parameters, which in turn, puts considerable burden on hardware. Unstructured pruning is widely used to reduce model sparsity. While, the irregularity introduced by unstructured pruning makes it difficult to accelerate sparse CNNs on systolic array. To address this issue, a variety of accelerators have been proposed. SIGMA, the state-of-the-art sparse GEMM accelerator, achieves significant speedup over systolic array. However, SIGMA suffers from two disadvantages: 1) it only supports one-side sparsity, leaving potential for further performance gains; 2) SIGMA improves utilization of large-sized systolic arrays at the cost of extra overhead. Zhengbo Chen, Fang Zheng 0007, Zuoning Chen |
ICPP | 5 |
| 2022 | SeqDLM: A Sequencer-Based Distributed Lock Manager for Efficient Shared File Access in a Parallel File SystemabstractDistributed locks are used to guarantee the distributed client-cache coherence in parallel file systems. However, they lead to poor performance in the case of parallel writes under high-contention workloads. We analyze the distributed lock manager and find out that lock conflict resolution is the root cause of the poor performance, which involves frequent lock revocations and slow data flushing from client caches to data servers. We design a distributed lock manager named SeqDLM by exploiting the sequencer mechanism. SeqDLM mitigates the lock conflict resolution overhead using early grant and early revocation while keeping the same semantics as traditional distributed locks. To evaluate SeqDLM, we have implemented a parallel file system called ccPFS using both SeqDLM and traditional distributed locks. Evaluations on 96 nodes show SeqDLM outperforms the traditional distributed locks by up to$\boldsymbol{10.3}\times$for high-contention parallel writes on a shared file with multiple stripes. Shaonan Ma, Kang Chen 0001, Teng Ma 0006, Xin Liu 0081, Dexun Chen, Yongwei Wu 0001, Zuoning Chen |
SC | 8 |
| 2022 | Evaluating performance of AI operators using roofline model
Zhengbo Chen, Fang Zheng 0007, Rujun Sun, Zuoning Chen |
Appl. Intell. | 6 |
| 2022 | Path Sensitive Fuzzing for Native ApplicationsabstractCoverage-guided fuzzing is a widely used and effective solution to find software vulnerabilities. Tracking code coverage and utilizing it to guide fuzzing are crucial to coverage-guided fuzzers. However, tracking full and accurate path coverage is infeasible in practice due to the high instrumentation overhead. Popular fuzzers (e.g., AFL) often usecoarsecoverage information, e.g., edge hit counts stored in a compact bitmap, to achieve highly efficient greybox testing. Such inaccuracy and incompleteness in coverage introduce serious limitations to fuzzers. First, it causespath collisions, which prevent fuzzers from discovering potential paths that lead to new crashes. More importantly, it prevents fuzzers from making wise decisions on fuzzing strategies. In this article, we propose a coverage sensitive fuzzing solution CollAFL. It mitigates path collisions by providing more accurate coverage information, while still preserving low instrumentation overhead. It also utilizes the coverage information to apply three new fuzzing strategies, promoting the speed of discovering new paths and vulnerabilities. We implemented two variants of this solution, namely CollAFL (based on AFL) and CollAFL-bin (based on AFL-dyninst), to test applications with and without source code respectively, and evaluated them on 24 popular applications. The results showed that path collisions are common, i.e., up to 75 percent of edges could collide with others in some applications. But our solutions CollAFL and CollAFL-bin could reduce the edge collision ratio to nearly zero. Moreover, armed with the three fuzzing strategies, they outperform their counterparts (i.e., AFL and AFL-dyninst) in terms of both code coverage and vulnerability discovery. On average, CollAFL covered 20 percent more program paths, and found 320 percent more unique crashes and 260 percent more bugs than AFL in 200 hours. Moreover, CollAFL-bin covered 15 percent more paths, and found 200 percent more unique crashes and 150 percent more vulnerabilities than AFL-dyninst, showing that the proposed solution also works for binary application fuzzing. In total, CollAFL found 157 new security bugs with 95 new CVEs assigned. Shuitao Gan, Chao Zhang 0008, Xiaojun Qin, Xuwen Tu, Zhongyu Pei, Zuoning Chen |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2022 | AMT: asynchronous in-place matrix transpose mechanism for sunway many-core processor
Zhengbo Chen, Fang Zheng 0007, Zuoning Chen |
J. Supercomput. | 6 |
| 2022 | Jdebug: A Fast, Non-Intrusive and Scalable Fault Locating Tool for Ten-Million-Scale Parallel ApplicationsabstractThis article presents Jdebug, a fast, non-intrusive and scalable fault locating tool for extreme-scale parallel applications. Large-scale debugging has drawn more attention with the increasing scale of supercomputers and applications. To eliminate program intrusion caused by traditional instrumentation or interception during debugging information acquisition, we introduce the out-of-band management into large-scale debugging. We propose a rapid information gathering scheme that separates user and debugging traffic to solve scalability problem and to eliminate program interference during merging data. Observations of Program Counters (PC) and performance characteristics in suspended applications find abnormalities and help locate abnormal threads caused by software errors or hardware failures effectively. Evaluation shows that Jdebug collects PCs of over 20 million cores on the new Sunway supercomputer within 1.97 seconds, and can locate the abnormal threads in 1.4 seconds with an accuracy of 92.5%. In the running test of three fundamental benchmarks (HPL, HPCG, Graph500) and seventeen real-world applications, Jdebug quickly and accurately locates abnormal threads to help find scalability errors and hardware failures including memory access failures, communication failures, and execution component failures, which validates its effectiveness. Dajia Peng, Yunlong Feng, Xin Liu 0081, Wei Xue 0003, Dexun Chen, Jiawei Song, Zuoning Chen |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2020 | GREYONE: Data Flow Sensitive Fuzzing
Shuitao Gan, Chao Zhang 0008, Peng Chen 0034, Bodong Zhao, Xiaojun Qin, Zuoning Chen |
USENIX Security Symposium | 7 |
| 2020 | Lessons Learned from Optimizing the Sunway Storage System for Higher Application I/O Performance
Zuoning Chen, Wei Xue 0003, Bin Yang 0043 |
J. Comput. Sci. Technol. | 3 |
| 2018 | CollAFL: Path Sensitive FuzzingabstractCoverage-guided fuzzing is a widely used and effective solution to find software vulnerabilities. Tracking code coverage and utilizing it to guide fuzzing are crucial to coverage-guided fuzzers. However, tracking full and accurate path coverage is infeasible in practice due to the high instrumentation overhead. Popular fuzzers (e.g., AFL) often use coarse coverage information, e.g., edge hit counts stored in a compact bitmap, to achieve highly efficient greybox testing. Such inaccuracy and incompleteness in coverage introduce serious limitations to fuzzers. First, it causes path collisions, which prevent fuzzers from discovering potential paths that lead to new crashes. More importantly, it prevents fuzzers from making wise decisions on fuzzing strategies. In this paper, we propose a coverage sensitive fuzzing solution CollAFL. It mitigates path collisions by providing more accurate coverage information, while still preserving low instrumentation overhead. It also utilizes the coverage information to apply three new fuzzing strategies, promoting the speed of discovering new paths and vulnerabilities. We implemented a prototype of CollAFL based on the popular fuzzer AFL and evaluated it on 24 popular applications. The results showed that path collisions are common, i.e., up to 75% of edges could collide with others in some applications, and CollAFL could reduce the edge collision ratio to nearly zero. Moreover, armed with the three fuzzing strategies, CollAFL outperforms AFL in terms of both code coverage and vulnerability discovery. On average, CollAFL covered 20% more program paths, found 320% more unique crashes and 260% more bugs than AFL in 200 hours. In total, CollAFL found 157 new security bugs with 95 new CVEs assigned. Shuitao Gan, Chao Zhang 0008, Xiaojun Qin, Xuwen Tu, Zhongyu Pei, Zuoning Chen |
IEEE Symposium on Security and Privacy | 7 |
| 2018 | A Large-Scale Study of Failures on Petascale Supercomputers
Rui-Tao Liu, Zuoning Chen |
J. Comput. Sci. Technol. | 2 |
| 2018 | Post-exascale supercomputing: research opportunities aboundabstractExascale supercomputing refers to the scientific research efforts and activities to build and use supercomputers that can perform scientific computing at the speed of exaflops, or 10 18 floating-point (64-bit) operations per second.Exascale supercomputing is a major milestone in surpassing the state-of-the-art standard of petascale supercomputing, i.e., 10 15 floating-point operations per second, established a decade ago in 2008.Exascale scientific computing research is already in full bloom worldwide.The USA leads this research and development direction, with federal government funding starting in as early as 2008.Japan, Europe, India, and China soon followed suit.The most recent focus on the field is in Europe, with 1.4 billion euros budgeted for building pre-exascale supercomputers by 2020, and an additional 2.7 billion euros proposed for building an exascale supercomputer by 2023 (Feldman, 2018).It is expected that multiple exascale supercomputers will become operational in the USA, Europe, and Asia by 2020-2024, supporting cutting edge research in many scientific fields.In this context, the Chinese Academy of Engineering (CAE) organized a special issue of "Postexascale Zuoning Chen, Jack J. Dongarra, Zhiwei Xu 0002 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2018 | A Simulation Analysis of Redundancy and Reliability in Primary Storage DeduplicationabstractDeduplication has been widely used to improve storage efficiency in modern primary and secondary storage systems, yet how deduplication fundamentally affects storage system reliability remains debatable. This paper aims to analyze and compare storage system reliability with and without deduplication in primary workloads using public file system snapshots from two research groups. We first study the redundancy characteristics of the file system snapshots. We then propose a trace-driven, deduplication-aware simulation framework to analyze data loss in both chunk and file levels due to sector errors and whole-disk failures. Compared to without deduplication, our analysis shows that deduplication consistently reduces the damage of sector errors due to intra-file redundancy elimination, but potentially increases the damages of whole-disk failures if the highly referenced chunks are not carefully placed on disk. To improve reliability, we examine a deliberate copy technique that stores and repairs first the most referenced chunks in a small dedicated physical area (e.g., 1 percent of the physical capacity), and demonstrate its effectiveness through our simulation framework. Min Fu 0002, Shujie Han 0001, Patrick P. C. Lee, Dan Feng 0001, Zuoning Chen |
IEEE Trans. Computers | 5 |
| 2017 | SimMon: a toolkit for simulation of monitoring mechanisms in cloud computing environmentabstractSummary Monitoring is a precondition for intelligent management in cloud computing environment, such as dynamic resource allocation. Typically, working as an auxiliary tool, a monitoring system is expected to incur the least additional resource usage, thus the strategies to improve the efficiency of monitoring mechanisms become significant. Yet the scale and monitoring requirement of different data centers vary, we cannot determine whether a monitoring mechanism would work well in a new data center before it serves the data center. To evaluate monitoring mechanisms, we proposeSimMon, a toolkit for simulating monitoring mechanisms in cloud computing environments.SimMonis designed to simulate the topologies, actions, and strategies in data collection, dissemination, storage, and management processes.SimMonprovides a controllable and repeatable way to evaluate monitoring mechanisms. In this paper, we describe the requirements analysis, design, implementation, and evaluation ofSimMon. We simulate several different monitoring systems and compare their cost on time and resource to evaluate the efficiency ofSimMon. We reproduce two usage scenarios from former literatures to demonstrate the effectiveness ofSimMonon monitoring mechanisms simulation and evaluation. We build a real‐world working environment to validate the capability ofSimMonon mimicking the characteristics of cloud monitoring systems. Copyright © 2016 John Wiley & Sons, Ltd. Xinkui Zhao, Jianwei Yin, Chen Zhi, Zuoning Chen |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | Evolution of Cloud Operating System: From Technology to Ecosystem
Zuoning Chen, Kang Chen 0001, Jinlei Jiang, Lufei Zhang, Song Wu 0001, Zhengwei Qi, Chunming Hu, Yongwei Wu 0001, Yuzhong Sun, Aobing Sun, Zilu Kang |
J. Comput. Sci. Technol. | 1 |
| 2017 | CloudScout: A Non-Intrusive Approach to Service Dependency DiscoveryabstractNowadays, numerous enterprises are migrating their applications into cloud computing environments. Typically, the applications are composed of several dependent service components that span many hosts and network devices. In light of this, exploring the dependency between service components can be beneficial for achieving fast network application response time. Moreover, it is significant to consolidate service components according to resource constraints, service dependency, and network structure. However, it is a tedious task to discover the dependency among service components without expert knowledge of the running application. In this paper, we propose CloudScout, a non-intrusive approach that is capable of automatically discovering dependent service components. CloudScout analyzes the correlation among service components based on the time-series information from system monitoring logs. We address two key challenges in CloudScout: service distance calculation and dependent service clustering. We conduct experiments on five applications with 290 service components that span 20 physical hosts across two data centers. The experimental results demonstrate that CloudScout can successfully discover the dependency among service components and facilitate reducing the network latency of network applications and distributed applications. Jianwei Yin, Xinkui Zhao, Chen Zhi, Zuoning Chen, Zhaohui Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2016 | MonValley: An Unified Monitoring and Management Framework for Cloud ServicesabstractMonitoring is the cornerstone for cloud service management, so it is significant for a cloud monitoring tool to support customization for specific monitoring requirements, to indicate the correlation between target services, and to guide adaptive service management. Unfortunately, traditional monitoring tools are always developed independently with service management platforms and are provided as self-contained softwares, which limit their capabilities on addressing these requirements. In this paper, we propose MonValley, an unified monitoring and management framework for cloud services. It consists of four components: 1) a high-level language for practitioners to express monitoring specifications on the services from all three layers, 2) a compiler to translate the expressive program into an executable program, 3) an execution engine to execute the executable program, 4) a runtime to provide supports on basic functionalities, such as data transmission. MonValley provides an innovative approach to facilitate the integrated monitoring and adaptive management for cloud services. Xinkui Zhao, Jianwei Yin, Chen Zhi, Zuoning Chen |
ICWS | 4 |
| 2016 | vSpec: workload-adaptive operating system specialization for virtual machines in cloud computing
Xinkui Zhao, Jianwei Yin, Zuoning Chen |
Sci. China Inf. Sci. | 3 |
| 2016 | Reducing Fragmentation for In-line Deduplication Backup Storage via Exploiting Backup History and Cache KnowledgeabstractIn backup systems, the chunks of each backup are physically scattered after deduplication, which causes a challenging fragmentation problem. We observe that the fragmentation comes into sparse and out-of-order containers. The sparse container decreases restore performance and garbage collection efficiency, while the out-of-order container decreases restore performance if the restore cache is small. In order to reduce the fragmentation, we propose History-Aware Rewriting algorithm (HAR) and Cache-Aware Filter (CAF). HAR exploits historical information in backup systems to accurately identify and reduce sparse containers, and CAF exploits restore cache knowledge to identify the out-of-order containers that hurt restore performance. CAF efficiently complements HAR in datasets where out-of-order containers are dominant. To reduce the metadata overhead of the garbage collection, we further propose a Container-Marker Algorithm (CMA) to identify valid containers instead of valid chunks. Our extensive experimental results from real-world datasets show HAR significantly improves the restore performance by 2.84-175.36 × at a cost of only rewriting 0.5-2.03 percent data. Min Fu 0002, Dan Feng 0001, Yu Hua 0001, Xubin He, Zuoning Chen, Jingning Liu, Wen Xia, Fangting Huang, Qing Liu 0007 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2015 | Can Cloud Service Get His Family? A Step Towards Service Family DetectingabstractIn cloud computing environment, an application is always composed of several service components. A collection of service components is called a service family, and we name the cloud service components as service family members. In this paper, we propose a solution named Icebreaker to assemble service components belonging to the same application without sniffing tenants' privacy. Icebreaker characterizes each service component with basic resource consuming information and proposes a new distance calculating algorithm named iEntropy to distinct service components. We adaptively adopt Affinity Propagation (AP) clustering algorithm and maximum Silhouette index to identify the number of service family and assemble the service family members. Experiments are conducted on RUBiS, Hadoop and ApacheBench clusters with 169 VMs. Evaluation results show that Icebreaker can get 96.45% accuracy. Xinkui Zhao, Jianwei Yin, Chen Zhi, Pengxiang Lin, Zuoning Chen |
CLUSTER | 5 |
| 2015 | monBench: A Database Performance Benchmark for Cloud Monitoring SystemabstractMonitoring system provides a clear insight into the state and performance of service components in cloud computing platforms. It collects metrics from dispersed sensors and stores them in databases for show and future query. To choose the most suitable database for a monitoring system is complex, since the performance requirement on cloud monitoring system in different data centers differs widely. In this paper, we propose a benchmark named monBench to evaluate the performance of databases in cloud monitoring systems. Monbench extracts structural data from real-world monitoring logs to construct the benchmarking workload. Several performance metrics, such as throughput and query time, are evaluated by monBench. Xinkui Zhao, Jianwei Yin, Chen Zhi, Pengxiang Lin, Shichun Feng, Zuoning Chen |
CLUSTER | 7 |
| 2015 | Design Tradeoffs for Data Deduplication Performance in Backup Workloads
Min Fu 0002, Dan Feng 0001, Yu Hua 0001, Xubin He, Zuoning Chen, Wen Xia, Yujuan Tan |
FAST | 5 |
| 2015 | SimMon: A Toolkit for Simulating Monitoring Mechanism in Cloud Computing Environments
Xinkui Zhao, Jianwei Yin, Pengxiang Lin, Chen Zhi, Shichun Feng, Zuoning Chen |
ICSOC | 7 |
| 2014 | Accelerating Restore and Garbage Collection in Deduplication-based Backup Systems via Exploiting Historical Information
Min Fu 0002, Dan Feng 0001, Yu Hua 0001, Xubin He, Zuoning Chen, Wen Xia, Fangting Huang, Qing Liu 0007 |
USENIX ATC | 5 |
| 2013 | Workload Classification Model for Specializing Virtual Machine Operating SystemabstractThere is growing demand on strategies to help cloud computing utilize its scale adaptiveness and cost effectiveness advantages. Previous operating systems(OS) are designed to suit all, leading to that virtual machines with different workloads use indiscriminate processing platform. However, there are conflicts between generality and performance, limited resource utilization and low processing efficiency of common OS penalize system performance. Therefore, we design four kinds of OS optimization strategies corresponding to four primary classes of workloads: CPU-Intensive, Memory-Intensive, I/O-Intensive and Network-Intensive. In this paper, we propose a Feedback-Based Workload Classification(FBWC) model which contains metrics collector, data preprocessor, Training Set Refresh Support Vector Machine(TSRSVM) classifier, decision maker and operating system tuner to classify workloads into appropriate class. TSRSVM combines support vectors of origin training set and correctly classified testing set together as new training set to get higher classification accuracy and efficiency. Comprehensive experiments compared with K Nearest Neighbors(KNN) and SVM demonstrate effectiveness of FBWC model and TSRSVM classification algorithm. Performance comparison between common virtual machine and the tuned one shows high degree performance improvement by OS specialization. Xinkui Zhao, Jianwei Yin, Zuoning Chen |
IEEE CLOUD | 3 |
| 2013 | Distance-aware virtual cluster performance optimization: A hadoop case studyabstractCloud computing and big data are becoming two important developing trends in information technology area. However, data-intensive computing has some challenges to work well on virtual machines in cloud computing for virtualized resource competition and complex network communication. Network becomes one of the most notorious bottlenecks, which highlights strategies to lower communication and transmission cost in virtual cluster. In this paper, we present a novel cluster performance optimization strategy named vClusterOpt. vClusterOpt finds out centralized subgraphs of node graph and choose node with the shortest logical distance as kernel node of the subgraph to reduce inter-machine communication and transmission cost under virtual cluster. To calculate logical distance accurately, we define two kinds of logical distance: Logical Communication Distance(LCD) and Logical Transmission Distance(LTD). VM with the shortest LCD with others is used as the communication kernel node who has the most information communication stress, while VM with the shortest LTD is treated as transmission kernel node who has the most data transmission stress. We choose benchmarks running on Hadoop as the represent of data-intensive computing service to demonstrate effectiveness of our approach. Experiments show that an average of 20% performance improvement can get by our distance-aware virtual cluster optimization strategy. Xinkui Zhao, Jianwei Yin, Zuoning Chen, Xingjian Lu |
CLUSTER | 3 |
| 2009 | A Virtualized Self-Adaptive Parallel Programming Framework for Heterogeneous High Productivity ComputersabstractThis paper proposed a Virtualized Self-Adaptive Heterogeneous High Productivity Computers Parallel Programming Framework (VAPPF), which is composed of Virtualization-Based Runtime System (VRTS) and Virtualized Adaptive Parallel Programming Model (VAPPM). Virtualization-Based Runtime System is composed of Node-Level Virtual Machine Monitor (NVMM) and System-Level Virtual Infrastructure (SVI). VAPPM program model is not only compatible with conventional data parallel, but also support task parallel. Moreover, with the concept of Domains and virtualized process Locale, Virtualization-Based Runtime System can map between computation and processors according to system-level resources view and performance model. By conceal the hardware details through both runtime system level and programming model level by virtualization, the framework provides programmers a middle-level view independent of hardware details. Programmers can do their programming and debugging works on this middle-level view, and then, the runtime system map it into specific hardware environment. By this way, programming can be relatively separated from specific hardware architectures, this model realized an efficient work division between programmers and systems, and can help to improve the system’s programmability, scalability, portability, robustness, performance, and productivity. Zuoning Chen, Ninghui Sun, Fenbin Qi, Chaoqun Dong, Laiwang Cheng |
ISPA | 2 |