EDBT 2026 Demo / reviewers in the wild / expert
Peng Wang 0037
dblp:95/4442-37
· DBLP profile ↗
29ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0002-4008-1963ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A general lightweight and adaptive cache space allocation scheme
Ke Liu 0014, Hua Wang 0008, Yajun Tan, Peng Wang 0037, Yuanzhang Wang, Ke Zhou 0001, Quan Fu |
Future Gener. Comput. Syst. | 4 |
| 2026 | HyTorC: Hybrid Address Translation for SSDs Supporting CompressionabstractHigh-capacity solid-state drives (SSDs) with expected capacities of one PByte and more will address cloud storage and archiving environments previously dominated by magnetic disks. These new applications are very cost-sensitive, so unnecessary overhead must be reduced as much as possible without sacrificing the performance advantages of SSDs. An expensive component within a scale-up SSD is the on-device memory to map host logical page numbers to flash pages. This article, therefore, proposes HyTorC, which builds on the idea of S-FTL to represent sequentially stored logical pages using bitmaps instead of providing one entry per logical page. HyTorC extends this idea by introducing, for the first time in an FTL, compression of these bitmaps using run-length encoding, and by investigating the effects of background scrubbing to realign randomly written pages into contiguous runs. This background scrubbing allows HyTorC to keep large portions of the mapping table in block mapping mode, further reducing the memory footprint. HyTorC retains the flexibility of the page mapping scheme and supports compression of logical blocks within the SSD, allowing multiple compressed logical pages to be stored within a single physical page. HyTorC is fully implemented in an open-channel SSD. Our tests show that HyTorC can reduce memory consumption by an average of 98.9% over standard page mapping, 95.6% over the DFTL scheme, 87.3% over S-FTL, and 56.5% over the learned index-based approach LeaFTL for the Alibaba Cloud block traces, and by 98.6% over standard page mapping, 95.5% over the DFTL scheme, 84.1% over S-FTL, and 61.5% over LeaFTL for the Microsoft Research Cambridge traces. HyTorC focuses on the memory footprint of the FTL and not on performance. However, the performance evaluation shows that HyTorC achieves similar performance compared to page mapping and FTLs based on learned indexes. Yu Zhang 0294, Renhai Chen, Gong Zhang 0001, Peng Wang 0037, Xin Yao 0008, Keji Huang, André Brinkmann |
ACM Trans. Storage | 4 |
| 2025 | NeuVSA: A Unified and Efficient Accelerator for Neural Vector SearchabstractNeural Vector Search (NVS) has exhibited superior search quality over traditional key-based strategies for information retrieval tasks. An effective NVS architecture requires high recall, low latency, and high throughput to enhance user experience and cost-efficiency. However, implementing NVS on existing neural network accelerators and vector search accelerators is sub-optimal due to the separation between the embedding stage and vector search stage at both algorithm and architecture levels. Fortunately, we unveil that Product Quantization (PQ) opens up an opportunity to break separation. However, existing PQ algorithms and accelerators still focus on either the embedding stage or the vector search stage, rather than both simultaneously. Simply combining existing solutions still follows the beaten track of separation and suffers from insufficient parallelization, frequent data access conflicts, and the absence of scheduling, thus failing to reach optimal recall, latency, and throughput. To this end, we propose a unified and efficient NVS accelerator dubbed NeuVSA based on algorithm and architecture co-design philosophy. Specifically, on the algorithm level, we propose a learned PQ-based unified NVS algorithm that consolidates two separate stages into the same computing and memory access paradigm. It integrates an end-to-end joint training strategy to learn the optimal codebook and index for enhanced recall and reduced PQ complexity, thus achieving smoother acceleration. On the architecture level, we customize a homogeneous NVS accelerator based on the unified NVS algorithm. Each sub-accelerator is optimized to exploit all parallelism exposed by unified NVS, incorporating a structured index assignment strategy and an elastic on-chip buffer to alleviate buffer conflicts for reduced latency. All sub-accelerators are coordinated using a hardware-aware scheduling strategy for boosted throughput. Experimental results show that the joint training strategy improves recall by 4.6% over the separated strategy and accuracy by 43.5% over LUT-NN. NeuVSA achieves $2.82 \times$ to $416.17 \times$ lower latency over CPU, GPU, DFX+ANNA, and PQA+ANNA, and up to $49.60 \times$ and $10.57 \times$ higher average throughput over CPU and GPU, respectively. NeuVSA also reduces chip area by 65.2% over PQA+ANNA. Ziming Yuan, Wen Li 0013, Jie Zhang 0048, Shengwen Liang, Ying Wang 0001, Cheng Liu 0008, Huawei Li 0001, Xiaowei Li 0001, Jiafeng Guo, Peng Wang 0037, Renhai Chen, Gong Zhang 0001 |
HPCA | 11 |
| 2025 | ADAPT: Dynamic Grouping and Cross-Group Aggregation for GC-Efficient Log-Structured Storage in SSD ArraysabstractLog-structured storage (LSS) has been widely adopted in SSD-based array architectures due to its high I/O performance and flash-friendly design. However, LSS requires optimized data placement to curb the write amplification (WA) resulting from its garbage collection (GC) mechanisms. We reveal that existing data placement strategies suffer from high padding overhead and redundant GC migrations under production workloads due to the granularity mismatch between log-structured storage management and array-level organization. To tackle these inefficiencies, we propose ADAPT, an access-density-aware data placement strategy to decrease padding overhead and minimize WA. ADAPT separates user-written blocks into groups with an adaptive threshold by considering workload-level access density and block-level popularity. Additionally, cross-group dynamic aggregation exploits unused space in cold groups to optimize real-time block padding for hot data. Moreover, we deploy a proactive demotion placement algorithm to coordinate GC processes and eliminate redundant migrations. Experimental evaluations demonstrate that compared with state-of-the-art approaches, ADAPT lowers WA by 21.8%-46.3% and decreases padding traffic by 40%-72.1% under production workloads, while delivering high throughput and maintaining manageable memory overhead under heavy workloads. Ruisong Zhou, Peng Wang 0037, Chunhua Li 0002, Ke Zhou 0001 |
ICPP | 2 |
| 2025 | HLN-Tree: A memory-efficient B+-Tree with huge leaf nodes and locality predictorsabstractKey-value stores in Cloud environments can contain more than 2 45 unique elements and be larger than 100 PByte. B + -Trees are well suited for these larger-than-memory datasets and seamlessly index data stored on thousands of secondary storage devices. Unfortunately, it is often uneconomical to even store all inner tree nodes in memory for these dataset sizes. Therefore, lookup performance is affected by the additional IOs for reading inner nodes. This number of inner nodes can be reduced by increasing the size of leaf nodes. We propose HLN-Trees, which support huge leaf nodes without increasing the IO sizes for individual index operations. They partition leaf nodes in arrays of independent subnodes and combine ideas from BD-trees with rebalancing, learning key deviations, and storing locality predictors. HLN-Trees have been initially designed for uniform random key distributions and support arbitrary key distributions through an additional layer of hashing in leaf nodes. HLN-Trees decrease the number of inner nodes by up to 256× for uniform random key distributions and by 16× to 64× for arbitrary ones compared to B + -Trees, while keeping their performance at the same level even at high concurrency levels. We show analytically and through real-world and synthetic benchmarks that HLN-Trees also outperform state-of-the-art learned indexes for secondary storage. André Brinkmann, Reza Salkhordeh, Florian Wiegert, Peng Wang 0037, Xin Yao 0008, Renhai Chen, Keji Huang, Gong Zhang 0001 |
ACM Trans. Storage | 4 |
| 2024 | Palantir: Hierarchical Similarity Detection for Post-Deduplication Delta CompressionabstractDeduplication compresses backup data by identifying and removing duplicate blocks. However, deduplication cannot detect when two blocks are very similar, which opens up opportunities for further data reduction using delta compression. Most existing works find similar blocks by characterizing each block by a set of features and matching similar blocks using coarse-grained super-features. If two blocks share a super-feature, delta compression only needs to store their delta for the new block. Hongming Huang, Peng Wang 0037, Hong Xu 0001, Chun Jason Xue, André Brinkmann |
ASPLOS (2) | 2 |
| 2024 | MV-Adapter: Multimodal Video Transfer Learning for Video Text RetrievalabstractState-of-the-art video-text retrieval (VTR) methods typically involve fully fine-tuning a pre-trained model (e.g. CLIP) on specific datasets. However, this can result in significant storage costs in practical applications as a separate model per task must be stored. To address this issue, we present our pioneering work that enables parameter-efficient VTR using a pre-trained model, with only a small number of tunable parameters during training. Towards this goal, we propose a new method dubbed Multimodal Video Adapter (MV-Adapter) for efficiently transferring the knowledge in the pre-trained CLIP from image-text to video-text. Specifically, MV-Adapter utilizes bottleneck structures in both video and text branches, along with two novel components. The first is a Temporal Adaptation Module that is incorporated in the video branch to introduce global and local temporal contexts. We also train weights calibrations to adjust to dynamic variations across frames. The second is Cross Modality Tying that generates weights for video/text branches through sharing cross modality factors, for better aligning between modalities. Thanks to above innovations, MV-Adapter can achieve comparable or better performance than standard full fine-tuning with negligible parameters overhead. Notably, MV-Adapter consistently outperforms various competing methods in V2T/T2V tasks with large margins on five widely used VTR benchmarks (MSR-VTT, MSVD, LSMDC, DiDemo, and ActivityNet). Codes will be released. Xiaojie Jin 0004, Weibo Gong, Xueqing Deng, Peng Wang 0037, Zhao Zhang 0001, Xiaohui Shen, Jiashi Feng |
CVPR | 6 |
| 2024 | HyperDB: a Novel Key Value Store for Reducing Background Traffic in Heterogeneous SSD StorageabstractLog-structured merge tree (LSM-tree) has been widely adopted by modern key-value stores. Deploying LSM-tree across heterogeneous SSD storage which combines the fast but expensive NVMe storage tier with the slow but economical SATA storage tier has emerged as the optimal choice for maximizing cost-effectiveness. However, existing studies typically focus on optimizing the performance of individual storage layers, thereby impeding the full utilization potential of both storage layers. We notice that they tend to over-rely on one storage layer and underutilize the other. In this paper, we present HyperDB, a novel hybrid key-value store designed to enhance the overall performance of both layers via deploying tailored data structures in different media. Especially, HyperDB devises a zone-based data layout for NVMe SSDs to reduce migration overhead, while also implementing a semi-sorted table on the SATA storage layer to minimize merge overhead. Furthermore, we propose a preemptive compaction method at the block-granularity level to further alleviate resource consumption caused by background compaction. Experimental results show that HyperDB achieves 2.25 × faster on average throughput and a 60.3% reduction in background task traffic, compared to the standard use of RocksDB in data centers today. Ruisong Zhou, Yuzhan Zhang, Chunhua Li 0002, Ke Zhou 0001, Peng Wang 0037, Gong Zhang 0001, Ji Zhang 0010 |
ICPP | 5 |
| 2024 | CGHit: A Content-Oriented Generative-Hit Framework for Content Delivery NetworksabstractThe service provided by content delivery networks (CDNs) may overlook content locality, leaving the potential to improve performance. In this study, we explore the feasibility of leveraging generated data as a replacement for fetching data in missing scenarios based on content locality. Due to sufficient local computing resources and reliable generation efficiency, we propose a content-oriented generative-hit framework (CGHit) for CDNs. CGHit utilizes idle computing resources on edge nodes to generate requested data based on similar or related cached data, achieving hits. Extensive experiments in a real-world system demonstrate that CGHit reduces the average access latency by half. In addition, experiments conducted on a simulator confirm that CGHit can enhance current caching algorithms, leading to lower latency and reduced bandwidth usage. Peng Wang 0037, Yu Liu 0040, Ke Liu 0014, Ke Zhou 0001, Zhihai Huang |
NAS | 1 |
| 2024 | SLAP: Segmented Reuse-Time-Label Based Admission Policy for Content Delivery Network Cachingabstract‘‘Learned” admission policies have shown promise in improving Content Delivery Network (CDN) cache performance and lowering operational costs. Unfortunately, existing learned policies are optimized with a few fixed cache sizes while in reality, cache sizes often vary over time in an unpredictable manner. As a result, existing solutions cannot provide consistent benefits in production settings. We present SLAP , a learned CDN cache admission approach based on segmented object reuse time prediction. SLAP predicts an object’s reuse time range using the Long-Short-Term-Memory model and admits objects that will be reused (before eviction) given the current cache size. SLAP decouples model training from cache size, allowing it to adapt to arbitrary sizes. The key to our solution is a novel segmented labeling scheme that makes SLAP without requiring precise prediction on object reuse time. To further make SLAP a practical and efficient solution, we propose aggressive reusing of computation and training on sampled traces to optimize model training, and a specialized predictor architecture that overlaps prediction computation with miss object fetching to optimize model inference. Our experiments using production CDN traces show that SLAP achieves significantly lower write traffic (38%-59%), longer SSDs lifetime (104%-178%), a consistently higher hit rate (3.2%-11.7%), and requires no effort to adapt to changing cache sizes, outperforming existing policies. Ke Liu 0014, Hua Wang 0008, Ke Zhou 0001, Peng Wang 0037, Ji Zhang 0010 |
ACM Trans. Archit. Code Optim. | 5 |
| 2024 | $\varepsilon$ɛ-LAP: A Lightweight and Adaptive Cache Partitioning Scheme With Prudent Resizing Decisions for Content Delivery NetworksabstractAs dependence on Content Delivery Networks (CDNs) increases, there is a growing need for innovative solutions to optimize cache performance amid increasing traffic and complicated cache-sharing workloads. Allocating exclusive resources to applications in CDNs boosts the overall cache hit ratio (OHR), enhancing efficiency. However, the traditional method of creating the miss ratio curve (MRC) is unsuitable for CDNs due to the diverse sizes of items and the vast number of applications, leading to high computational overhead and performance inconsistency. To tackle this issue, we propose alightweight andadaptive cachepartitioning scheme called$\varepsilon$-LAP. This scheme uses a corresponding shadow cache for each partition and sorts them based on the average hit numbers on the granularity unit in the shadow caches. During partition resizing,$\varepsilon$-LAP transfers storage capacity, measured in units of granularity, from the$(N-k+1)$-th ($k\leq \frac{N}{2}$) partition to the$k$-th partition. A learning threshold parameter, i.e.,$\varepsilon$, is also introduced to prudently determine when to resize partitions, improving caching efficiency. This can eliminate about 96.8% of unnecessary partition resizing without compromising performance.$\varepsilon$-LAP, when deployed inPicCloudatTencent, improved OHR by 9.34% and reduced the average user access latency by 12.5 ms. Experimental results show that$\varepsilon$-LAP outperforms other cache partitioning schemes in terms of both OHR and access latency, and it effectively adapts to workload variations. Peng Wang 0037, Yu Liu 0040, Zhelong Zhao, Ke Liu 0014, Ke Zhou 0001, Zhihai Huang |
IEEE Trans. Cloud Comput. | 1 |
| 2024 | Beyond Belady to Attain a Seemingly Unattainable Byte Miss Ratio for Content Delivery NetworksabstractReducing the byte miss ratio (BMR) in the Content Delivery Network (CDN) caches can help providers save on the cost of paying for traffic. When evicting objects or files of different sizes in the caches of CDNs, it is no longer sufficient to pursue an optimal object miss ratio (OMR) by approximating Belady to ensure an optimal BMR. Our experimental observations suggest that there are multiple request sequence windows. In these windows, a replacement policy prioritizes the eviction of objects with large sizes and ultimately evicts the object with the longest reuse distance, lowering the BMR without increasing the OMR. To accurately capture those windows, we monitor the changes in OMR and BMR using a deep reinforcement learning (RL) model and then implement a BMR-friendly replacement algorithm in these windows. Based on this policy, we propose a Belady and Size Eviction (LRU-BaSE) algorithm that reduces BMR while maintaining OMR. To make LRU-BaSE efficient and practical, we address the feedback delay problem of RL with a two-pronged approach. On the one hand, we shorten the LRU-base decision region based on the observation that the rear section of the cache queue contains most of the eviction candidates. On the other hand, the request distribution on CDNs makes it feasible to divide the learning region into multiple sub-regions that are each learned with reduced time and increased accuracy. In real CDN systems, LRU-BaSE outperforms LRU by reducing “backing to OS” traffic and access latency by 30.05% and 17.07%, respectively, on average. In simulator tests, LRU-BaSE outperforms state-of-the-art cache replacement policies. On average, LRU-BaSE's BMR is 0.63% and 0.33% less than that of Belady and Practical Flow-based Offline Optimal (PFOO), respectively. In addition, compared to Learning Relaxed Belady (LRB), LRU-BaSE can yield relatively stable performance when facing workload drift. Peng Wang 0037, Hong Jiang 0001, Yu Liu 0040, Zhelong Zhao, Ke Zhou 0001, Zhihai Huang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | Smart Cache Insertion and Promotion Policy for Content Delivery NetworksabstractImproving hit rates can be achieved by enhancing cache replacement algorithms with the identification of zero-reuse objects (ZROs) and inserting them at the end of the cache queue. Note that the promotion policy needs to achieve a similar task as the above insertion policy since the hit object may immediately become a ZRO (called P-ZRO) that is not suitable for placement at the front of the queue. However, existing studies have yet to consider P-ZROs, and current insertion algorithms struggle to simultaneously identify both ZROs and P-ZROs. To address these issues, we propose integrating the insertion and promotion policies. We do this by treating hit objects as special missing objects and employing reinforcement learning to create a unified model for both policies, where the learning function recognizes the relationship between performance changes and the emergence of ZROs and P-ZROs. Our proposed solution is a smart cache insertion and promotion policy (SCIP) that dynamically adjusts the insertion position using a bimodal insertion policy for both missing and hit objects, guided by the model. Extensive experiments demonstrate that SCIP significantly improves overall performance in real-world content delivery network systems and outperforms state-of-the-art insertion policies in terms of miss ratios in the simulator. In addition, deploying SCIP on optimal cache replacement algorithms can further decrease their miss ratios. Peng Wang 0037, Yu Liu 0040, Zhelong Zhao, Ke Zhou 0001, Zhihai Huang, Yanxiong Chen |
ICPP | 1 |
| 2023 | Bottleneck-Aware Non-Clairvoyant Coflow Scheduling With FaiabstractCoflow scheduling is critical to data-parallel applications in data centers. While schemes like Varys can achieve optimal performance, they require a priori information about coflows which is hard to obtain in practice. Existing non-clairvoyant solutions like Aalo generalize least attained service (LAS) scheduling discipline to address this issue. However, they fail to identify the bottleneck flows in a coflow and tend to allocate excessive bandwidth to the non-bottleneck flows, leading to bandwidth wastage and inferior overall performance. To this end, we present Fai that strives to improve the overall coflow performance by accelerating the bottleneck flows without priori knowledge. Fai employs bottleneck-aware scheduling. It adopts loose coordination to update coflow priority and flow rates based on total bytes sent. In addition, Fai detects bottleneck flows based on a flow’s rate and bytes sent, and de-allocates bandwidth for other flows to match the bottleneck rate without affecting the coflow completion time (CCT). The saved bandwidth is then distributed among coflows according to their priority to improve overall performance. Testbed evaluation on a 40-node cluster shows that Fai improves average (P95) CCT by 1.73× (3.43×), compared to Aalo. Large-scale trace-driven simulations also show that Fai outperforms Aalo substantially. Libin Liu 0001, Chengxi Gao, Peng Wang 0037, Hongming Huang, Jiamin Li 0002, Hong Xu 0001, Wei Zhang 0049 |
IEEE Trans. Cloud Comput. | 3 |
| 2022 | SuperCDC: A Hybrid Design of High-Performance Content-Defined Chunking for Fast DeduplicationabstractContent-Defined Chunking (CDC) has been widely applied in data deduplication systems in the past since it can detect much more redundant data than Fixed-Size Chunking (FSC). CDC approach becomes faster and faster to match the improvement of high performance storage systems, and there are two main kinds of acceleration mechanisms: calculation-efficient acceleration and stream-informed acceleration. We observe the opportunity to combine the benefits of these two mechanisms to chase a faster speed, and the challenges of memory overhead and deduplication ratio loss, which are caused by stream histories and the configuration of minimum/maximum chunk size, respectively. Motivated by these observations, we proposed SuperCDC with several corresponding techniques, including hybridizing calculation-efficient processing with a stream-informed design, memory-efficient structure for tracking stream history, and Min-Max chunking for improving deduplication ratio. Evaluations suggest that SuperCDC achieves an up to 4.94× faster chunking speed and 6.23% higher deduplication ratio while saving 99.58% memory space on stream histories compared with the state-of-the-art Gear-based RapidCDC. Binzhaoshuo Wan, Lifeng Pu, Xiangyu Zou, Peng Wang 0037, Wen Xia |
ICCD | 5 |
| 2022 | Adaptive Size-Aware Cache Insertion Policy for Content Delivery NetworksabstractContent delivery networks (CDNs) are large distributed cache systems that deliver objects with inconsistent sizes. The zero-reuse objects that are not reused in a time window but still loaded and evicted in the cache waste cache resources and result in degradation of object hit ratio (OHR) in CDNs. Although prohibiting these objects from entering the cache is a viable solution, the variable workloads and various object sizes in CDNs make the determination of zero-reuse difficult, resulting in an increased risk of bandwidth overhead in the data center by the misjudgment. To alleviate this problem, we propose to use the insertion policy to give each object at least one chance to be hit. Meanwhile, we find that the distribution of zero-reuse objects correlates with their sizes through data analysis. As a result, we propose an adaptive size-aware cache insertion policy (ASC-IP) for the OHR improvement and design an adaptive scheme to dynamically adjust the size threshold used to determine the zero-reuse objects, adapting the mutative access patterns with negligible overhead. We have deployed ASC-IP in TDC of Company-T and ASC-IP can improve the OHR by 9.6% and reduce the user access latency by 7.14ms on average and reduce the back-to-source bandwidth by 8.75Gbps. In addition, on Twitter, Wikipedia, and a real-world Trace-T, we show that ASC-IP outperforms state-of-the-art cache algorithms working on CDNs and can upgrade LRU-based replacement algorithms with negligible overheads. Peng Wang 0037, Yu Liu 0040, Zhelong Zhao, Ke Zhou 0001, Zhihai Huang, Yanxiong Chen |
ICCD | 1 |
| 2022 | A Lightweight and Adaptive Cache Partitioning Scheme for Content Delivery NetworksabstractAllocating exclusive resources for different applications in content delivery networks (CDNs) allows for a higher overall hit ratio. The cache partitioning schemes on Last-Level Cache (LLC) are promising solutions that dynamically split cache sizes into partitions corresponding to threads by the miss ratio curve (MRC). Nonetheless, due to the sheer number of applications and various item sizes in CDNs, partitioning via MRC will cause high computational overheads and performance fluctuations. As a result, in this paper, we propose a lightweight and adaptive cache partitioning scheme (LAP) for CDNs. LAP establishes a shadow cache for each partition, where the size of the partition and its shadow cache is equal to the size of the integral cache. The average number of hits on the granularity unit in the shadow caches, where the size of the granularity equals the size of the probable largest item, is used to sort N partitions in decreasing order. When resizing partitions, LAP transfers a capacity of the size of granularity from the (N – k + 1)-th $\left( {k \leq \frac{N}{2}} \right)$ partition into the k-th partition. Meanwhile,we provide a threshold that neglects partition resizing and improves partitioning efficiency. This lightweight scheme can enhance resource utilization by progressively adapting to workload variations. We have deployed LAP in PicCloud of Company-T and LAP can improve the OHR by 9.34% and reduce the average user access latency by 12.5ms. Then, we verify LAP in the public trace from Akamai and the real trace from PicCloud. Experimental results demonstrate that LAP outperforms other cache partitioning schemes and tackles the performance cliff problem with little overhead. Peng Wang 0037, Zhelong Zhao, Yu Liu 0040, Ke Zhou 0001, Zhihai Huang, Yanxiong Chen |
ICCD | 1 |
| 2022 | A Focused Garbage Collection Approach for Primary Deduplicated Storage with Low Memory OverheadabstractSince one chunk could be shared by many files after data deduplication, Garbage Collection (GC) is an essential but complex task to reclaim stale chunks in large-scale primary deduplication systems. Traditional Mark&Sweep is a widely used approach but suffers from the increasingly traversing time and huge memory overhead of Liveness Array (i.e., a data structure reflects the liveness of alive chunks) in the Mark phase. This paper proposes a new method named Focused Garbage Collection (FGC) to accelerate the Mark phase for primary deduplication storage significantly. Specifically, we design a global Austere Reference Graph with low memory cost that efficiently represents files’ reference relationships (i.e., sharing chunks after deduplication) by considering the deduplication characteristics of workloads in primary systems. Austere Reference Graph helps FGC focus on the deleted files and their correlative files to quickly mark stale chunks, while traditional approaches need to traverse all files. Consequently, FGC’s traversing time and Liveness Array size will be greatly reduced in the Mark phase. Evaluation results show that compared with traditional Mark&Sweep, FGC decreases the time consumption in the Mark phase 1.3×-7.34× in a stand-alone primary deduplication system and 128×-256× network traffic reduction for the Mark phase while only introducing < 0.05% extra memory overhead for the reference graph. Jingsong Yuan, Xiangyu Zou, Zhichao Cao 0002, Wen Xia, Peng Wang 0037, Li Chen 0008 |
ICCD | 7 |
| 2022 | ScaleFlux: Efficient Stateful Scaling in NFVabstractNetwork function virtualization (NFV) enables elastic scaling to middlebox deployment and management. Therefore, efficient stateful scaling is an important task because operators often need to shift traffic and the associated flow states across VNF instances to deal with time-varying loads. Existing NFV scaling methods, however, typically focus on one aspect of the scaling pipeline and does not offer an end-to-end scaling framework. This article presents ScaleFlux, a complete stateful scaling system that efficiently reduces flow-level latency and achieves near-optimal resource usage. ScaleFlux (1) monitors traffic load for each VNF instance and adopts a queue-based mechanism to detect load burstiness timely, (2) deploys a flow bandwidth predictor to predict flow bandwidth time-series with the ABCNN-LSTM model, and (3) schedules the necessary flow and state migration using the simulated annealing algorithm to achieve both flow-level latency guarantee and resource usage minimization. Testbed evaluation with a five-machine cluster shows that ScaleFlux reduces flow completion time by at least 8.7× for all the workloads and achieves near-optimal CPU usage during scaling. Libin Liu 0001, Hong Xu 0001, Zhixiong Niu, Jingzong Li, Wei Zhang 0049, Peng Wang 0037, Jiamin Li 0002, Chun Jason Xue, Cong Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | vPipe: A Virtualized Acceleration System for Achieving Efficient and Scalable Pipeline Parallel DNN TrainingabstractThe increasing computational complexity of DNNs achieved unprecedented successes in various areas such as machine vision and natural language processing (NLP), e.g., the recent advanced Transformer has billions of parameters. However, as large-scale DNNs significantly exceed GPU's physical memory limit, they cannot be trained by conventional methods such as data parallelism. Pipeline parallelism that partitions a large DNN into small subnets and trains them on different GPUs is a plausible solution. Unfortunately, the layer partitioning and memory management in existing pipeline parallel systems are fixed during training, making them easily impeded by out-of-memory errors and the GPU under-utilization. These drawbacks amplify when performing neural architecture search (NAS) such as the evolved Transformer, where different network architectures of Transformer needed to be trained repeatedly. vPipe is the first system that transparently provides dynamic layer partitioning and memory management for pipeline parallelism. vPipe has two unique contributions, including (1) an online algorithm for searching a near-optimal layer partitioning and memory management plan, and (2) a live layer migration protocol for re-balancing the layer distribution across a training pipeline. vPipe improved the training throughput of two notable baselines (Pipedream and GPipe) by 61.4-463.4 percent and 24.8-291.3 percent on various large DNNs and training settings. Shixiong Zhao, Fanxin Li, Xusheng Chen, Xiuxian Guan, Jianyu Jiang, Dong Huang 0005, Yuhao Qing, Sen Wang 0004, Peng Wang 0037, Gong Zhang 0001, Cheng Li 0001, Ping Luo 0002, Heming Cui |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2021 | HR-NAS: Searching Efficient High-Resolution Neural Architectures With Lightweight TransformersabstractHigh-resolution representations (HR) are essential for dense prediction tasks such as segmentation, detection, and pose estimation. Learning HR representations is typically ignored in previous Neural Architecture Search (NAS) methods that focus on image classification. This work proposes a novel NAS method, called HR-NAS, which is able to find efficient and accurate networks for different tasks, by effectively encoding multiscale contextual information while maintaining high-resolution representations. In HR-NAS, we renovate the NAS search space as well as its searching strategy. To better encode multiscale image contexts in the search space of HR-NAS, we first carefully design a lightweight transformer, whose computational complexity can be dynamically changed with respect to different objective functions and computation budgets. To maintain high-resolution representations of the learned networks, HR-NAS adopts a multi-branch architecture that provides convolutional encoding of multiple feature resolutions, inspired by HRNet [73]. Last, we proposed an efficient fine-grained search strategy to train HR-NAS, which effectively explores the search space, and finds optimal architectures given various tasks and computation resources. As shown in Fig. 1 (a), HR-NAS is capable of achieving state-of-the-art trade-offs between performance and FLOPs for three dense prediction tasks and an image classification task, given only small computational budgets. For example, HR-NAS surpasses SqueezeNAS [63] that is specially designed for semantic segmentation while improving efficiency by 45.9%. Code is available at https://github.com/dingmyu/HR-NAS. Mingyu Ding, Xiaochen Lian, Peng Wang 0037, Xiaojie Jin 0004, Zhiwu Lu 0001, Ping Luo 0002 |
CVPR | 4 |
| 2020 | Label-Attended Hashing for Multi-Label Image RetrievalabstractFor the multi-label image retrieval, the existing hashing algorithms neglect the dependency between objects and thus fail to capture the attention information in the feature extraction, which affects the precision of hash codes. To address this problem, we explore the inter-dependency between objects through their co-occurrence correlation from the label set and adopt Multi-modal Factorized Bilinear (MFB) pooling component so that the image representation learning can capture this attention information. We propose a Label-Attended Hashing (LAH) algorithm which enables an end-to-end hash model with inter-dependency feature extraction. LAH first combines Convolutional Neural Network (CNN) and Graph Convolution Network (GCN) to separately generate the image representation and label co-occurrence embeddings, then adopts MFB to fuse these two modal vectors, finally learns the hash function with a Cauchy distribution based loss function via back propagation. Extensive experiments on public multi-label datasets demonstrate that (1) LAH can achieve the state-of-the-art retrieval results and (2) the usage of co-occurrence relationship and MFB not only promotes the precision of hash codes but also accelerates the hash learning. GitHub address: https://github.com/IDSM-AI/LAH. Yanzhao Xie, Yu Liu 0040, Yangtao Wang, Lianli Gao, Peng Wang 0037, Ke Zhou 0001 |
IJCAI | 5 |
| 2019 | Luopan: Sampling-Based Load Balancing in Data Center NetworksabstractData center networks demand high-performance, robust, and practical data plane load balancing protocols. Despite progress, existing work falls short of meeting these requirements. We design, analyze, and evaluate Luopan, a novel sampling based load balancing protocol that overcomes these challenges. Luopan operates at flowcell granularity similar to Presto. It periodically samples a few paths for each destination switch and directs flowcells to the least congested one. By being congestion-aware, Luopan improves flow completion time (FCT), and is more robust to topological asymmetries compared to Presto. The sampling approach simplifies the protocol and makes it much more scalable for implementation in large-scale networks compared to existing congestion-aware schemes. We provide analysis to show that Luopan's periodic sampling has the same asymptotic behavior as instantaneous sampling: taking 2 random samples provides exponential improvements over 1 sample. We conduct comprehensive packet-level simulations with production workloads. The results show that Luopan consistently outperforms state-of-the-art schemes in large-scale topologies. Compared to Presto, Luopan with 2 samples improves the 99.9%ile FCT of mice flows by up to 35 percent, and average FCT of medium and elephant flows by up to 30 percent. Luopan also performs significantly better than Local Sampling with large asymmetry. Peng Wang 0037, George Trimponias, Hong Xu 0001, Yanhui Geng |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Kuijia: Traffic Rescaling in Software-Defined Data Center WANsabstractNetwork faults like link or switch failures can cause heavy congestion and packet loss. Traffic engineering systems need a lot of time to detect and react to such faults, which results in significant recovery times. Recent work either preinstalls a lot of backup paths in the switches to ensure fast rerouting or proactively prereserves bandwidth to achieve fault resiliency. Our idea agilely reacts to failures in the data plane while eliminating the preinstallation of backup paths. We propose Kuijia, a robust traffic engineering system for data center WANs, which relies on a novel failover mechanism in the data plane called rate rescaling. The victim flows on failed tunnels are rescaled to the remaining tunnels and enter lower priority queues to avoid performance impairment of aboriginal flows. Real system experiments show that Kuijia is effective in handling network faults and significantly outperforms the conventional rescaling method. Che Zhang, Hong Xu 0001, Libin Liu 0001, Zhixiong Niu, Peng Wang 0037 |
Secur. Commun. Networks | 5 |
| 2018 | CoCloud: Enabling Efficient Cross-Cloud File Collaboration Based on Inefficient Web APIsabstractCloud storage services such as Dropbox have been widely used for file collaboration among multiple users. However, this desirable functionality is yet restricted to the “walled-garden” of each service. At present, the only feasible approach to cross-cloud file collaboration seems to be using web APIs, whose performance is known to be highly unstable and unpredictable. Now that using inefficient web APIs is inevitable, in this paper we attempt to achieve sound user-perceived performance for cross-cloud file collaboration. This attempt is enabled by two key observations from real-world measurements. First, for each cloud, we are always able to deploy one or several nearby (client) proxies which can efficiently access the web APIs. Second, during file collaboration, significant similarity exists among different versions of a file. This can be exploited to substantially reduce inter-proxy traffic and thus shorten the data sync time. Guided by the observations, we design and implement an open-source prototype system called CoCloud. Currently, it supports file collaboration among four popular cloud storage services in the US and China. Its performance is well acceptable to users under representative workloads, even approaching or exceeding that of intra-cloud collaboration in many cases. Jinlong E, Yong Cui 0001, Peng Wang 0037, Zhenhua Li 0001, Chaokun Zhang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | CoCloud: Enabling efficient cross-cloud file collaboration based on inefficient web APIsabstractCloud storage services such as Dropbox have been widely used for file collaboration among multiple users. However, this desirable functionality is yet restricted to the “walled-garden” of each service. At present, the only effective approach to cross-cloud file collaboration seems to be using web APIs, whose performance is known to be highly unstable and unpredictable. Now that using inefficient web APIs is inevitable, in this paper we attempt to achieve sound user-perceived performance for cross-cloud file collaboration. This attempt is enabled by two key observations from real-world measurements. First, for each cloud, we are always able to deploy one or several nearby (client) proxies which can efficiently access the web APIs. Second, during file collaboration, significant similarity exists among different versions of a file. This can be exploited to substantially reduce inter-proxy traffic and thus shorten the data sync time. Guided by the observations, we design and implement an open-source prototype system called CoCloud. Currently, it supports file collaboration among four popular cloud storage services in the US and China. Its performance is well acceptable to users under representative workloads, even approaching or exceeding intra-cloud performance in many cases. Jinlong E, Yong Cui 0001, Peng Wang 0037, Zhenhua Li 0001, Chaokun Zhang |
INFOCOM | 3 |
| 2017 | Expeditus: Congestion-Aware Load Balancing in Clos Data Center NetworksabstractData center networks often use multi-rooted Clos topologies to provide a large number of equal cost paths between two hosts. Thus, load balancing traffic among the paths is important for high performance and low latency. However, it is well known that ECMP-the de facto load balancing scheme-performs poorly in data center networks. The main culprit of ECMP's problems is its congestion agnostic nature, which fundamentally limits its ability to deal with network dynamics. We propose Expeditus, a novel distributed congestion-aware load balancing protocol for general 3-tier Clos networks. The complex 3-tier Clos topologies present significant scalability challenges that make a simple per-path feedback approach infeasible. Expeditus addresses the challenges by using simple local information collection, where a switch only monitors its egress and ingress link loads. It further employs a novel two-stage path selection mechanism to aggregate relevant information across switches and make path selection decisions. Testbed evaluation on Emulab and large-scale ns-3 simulations demonstrate that, Expeditus outperforms ECMP by up to 45% in tail flow completion times (FCT) for mice flows, and by up to 38% in mean FCT for elephant flows in 3-tier Clos networks. Peng Wang 0037, Hong Xu 0001, Zhixiong Niu, Dongsu Han, Yongqiang Xiong |
IEEE/ACM Trans. Netw. | 1 |
| 2016 | Expeditus: Congestion-aware Load Balancing in Clos Data Center NetworksabstractData center networks often use multi-rooted Clos topologies to provide a large number of equal cost paths between two hosts. Thus, load balancing traffic among the paths is important for high performance and low latency. However, it is well known that ECMP---the de facto load balancing scheme---performs poorly in data center networks. The main culprit of ECMP's problems is its congestion agnostic nature, which fundamentally limits its ability to deal with network dynamics. Peng Wang 0037, Hong Xu 0001, Zhixiong Niu, Dongsu Han, Yongqiang Xiong |
SoCC | 1 |
| 2016 | Luopan: Sampling based load balancing in data center networksabstractData center networks demand high-performance, robust, and practical data plane load balancing protocols. Despite progress, existing work falls short of satisfying these requirements. We design and evaluate Luopan, a novel sampling based load balancing protocol that overcomes these challenges. Luopan operates at flowcell granularity similar to Presto. It periodically samples a few paths to each destination switch and directs flowcells to the least congested one. By being congestion-aware, Luopan improves flow completion time (FCT), and is more robust to topological asymmetries compared to Presto. The sampling approach simplifies the protocol and makes it much more scalable for implementation in large-scale networks compared to existing congestion-aware schemes. We conduct comprehensive packet-level simulations with a production workload. The results show that Luopan consistently outperforms state-of-the-art schemes in large-scale symmetric and asymmetric topologies. Compared to Presto, Luopan with 2 samples improves the 99%ile FCT of mice flows by up to 45%, and average FCT of medium flows by ~20%. Peng Wang 0037, George Trimponias, Hong Xu 0001, Hongyuan Liu 0003, Yanhui Geng |
ICNP | 1 |