VLDB 2026 Research / reviewers in the wild / expert
Bingzhe Li
dblp:167/9945
· DBLP profile ↗
52ranked-venue papers
17as first author
34since 2021 · last 2026
0000-0002-5815-9706ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 17 first-author · 29 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAGMA: A Multi-Graph based Agentic Memory Architecture for AI AgentsabstractMemory-Augmented Generation (MAG) extends Large Language Models with external memory to support long-context reasoning, but existing approaches largely rely on semantic similarity over monolithic memory stores, entangling temporal, causal, and entity information.This design limits interpretability and alignment between query intent and retrieved evidence, leading to suboptimal reasoning accuracy.In this paper, we propose MAGMA, a multi-graph agentic memory architecture that represents each memory item across orthogonal semantic, temporal, causal, and entity graphs.MAGMA formulates retrieval as policy-guided traversal over these relational views, enabling query-adaptive selection and structured context construction.By decoupling memory representation from retrieval logic, MAGMA provides transparent reasoning paths and fine-grained control over retrieval.Experiments on LoCoMo and LongMemEval demonstrate that MAGMA consistently outperforms state-of-the-art agentic memory systems in long-horizon reasoning tasks. Dongming Jiang, Guanpeng Li, Bingzhe Li |
ACL (1) | 4 |
| 2026 | VeloxGNN: Efficient Out-of-Core GNN Training with Delayed Gradient PropagationabstractTraining Graph Neural Networks (GNNs) on large-scale data is essential in various applications, e.g., transportation, and molecular biology. As graph sizes increasingly surpass main memory capacities, the out-of-core (OOC) GNN training system (OOC-based) has been proposed, a scheme that storing graphs on external storage, such as SSD or HDD, sequentially loading and processing smaller partitions. However, existing OOC-based systems face the key challenges: excessive data migration between storage and memory, as well as reduced model accuracy. In this paper, we theoretically and empirically analyze the limitations of state-of-the-art OOC-based systems and identify opportunities for optimization. Guided by our theoretical insights, we propose VeloxGNN, a novel system to improve data migration efficiency while maintaining high model accuracy. First, we introduce a novel algorithm, named Delayed Gradient Propagation (DGP), which specifically designed for OOC-based system. DGP leverages both historical node embeddings and unbiased gradients to achieve two key objectives simultaneously: minimizing the data migration (by ensuring that dataset is read at most once) while maintaining model accuracy. To support DGP, we then propose system-level optimizations: dynamic memory management, DGP-aware loading order, and a new graph partitioning method that separates labeled and unlabeled data. Experimental results show that VeloxGNN achieves memory-based accuracy while reducing training time by 17.7% to 73.3% across various datasets and GNN models, outperforming state-of-the-art methods. This result highlights VeloxGNN's potential for efficient and scalable GNN training on large-scale graph data. Tsun-Yu Yang, Zhaoyan Shen, Ming-Chang Yang, Bingzhe Li |
HPCA | 5 |
| 2025 | AdaCM^2: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory ReductionabstractThe advancements in large language models (LLMs) have propelled the improvement of video understanding tasks by incorporating LLMs with visual models. However, most existing LLM-based models (e.g., VideoLLaMA, VideoChat) are constrained to processing short-duration videos. Recent attempts to understand long-term videos by extracting and compressing visual features into a fixed memory size. Nevertheless, those methods leverage only visual modality to merge video tokens and overlook the correlation between visual and textual queries, leading to difficulties in effectively handling complex question-answering tasks. To address the challenges of long videos and complex prompts, we propose AdaCM2, which, for the first time, introduces an adaptive cross-modality memory reduction approach to video-text alignment in an auto-regressive manner on video streams. Our extensive experiments on various video understanding tasks, such as video captioning, video question answering, and video classification, demonstrate that AdaCM2achieves state-of-the-art performance across multiple datasets while significantly reducing memory usage. Notably, it achieves a 4.5% improvement across multiple tasks in the LVU dataset with a GPU memory consumption reduction of up to 65%. Yuanbin Man, Ying Huang 0008, Chengming Zhang 0006, Bingzhe Li, Wei Niu 0002, Miao Yin |
CVPR | 4 |
| 2025 | Oasis: An Out-of-core Approximate Graph System via All-Distances Sketches
Tsun-Yu Yang, Yizou Chen, Bingzhe Li, Ming-Chang Yang |
FAST | 4 |
| 2025 | Pixel-DNA: Increasing Robustness of Approximate DNA Storage for Images by Using Hierarchical DeduplicationabstractDNA is a high-density long-lasting storage medium that has the potential to keep up with the ever-increasing data storage needs. However, DNA storage is incredibly error-prone. The more nucleotides stored in the DNA strand the greater the chance of data loss. Moreover, DNA storage faces a specific phenomenon called error propagation due to its encoding from binary to nucleotide, resulting in an even higher error rate. When the majority of relevant data is stored inside the DNA strand error propagation is limited to the single strand. Reducing the overall data stored in DNA allows for more error correction and error protection. In this paper, we propose an isolation and deduplication scheme called Pixel-DNA for images. PixelDNA uses specific properties of DNA and DNA strands to implement multiple levels of deduplication while maximizing the data isolation. Also, a data segmentation scheme is applied to further increase the density of DNA storage. Based on the experimental results, Pixel-DNA allows for increased robustness and recoverability of images while increasing the density of DNA storage capacity. Alex Sensintaffar, David Hung-Chang Du, Bingzhe Li |
ICCD | 3 |
| 2025 | A Deep Dive into Protocol Design: How to Improve IPFS Performance without Sacrificing DecentralizationabstractThe InterPlanetary File System (IPFS) is a prominent decentralized storage solution; however, it struggles with performance issues. To address this challenge, the IPFS team has patched a series of centralized components, resulting in improved performance while giving more significant roles to specific entities. Balancing speed and decentralization has always posed a complex dilemma for storage systems. In this paper, we conduct a series of experiments and analyses to identify the performance advantages and constraints associated with the IPFS decentralized protocol. Based on thorough analysis, we propose a novel scheme named XIPFS to optimize IPFS. This scheme includes facilitating parallel block exchange across multiple nodes, refining content Publication strategies, and improving node selection algorithms for content routing. Our goal is to maximize the benefits of decentralized multi-source parallel downloading while minimizing the negative impact of decentralized indexing on execution time. Compared to previous approaches, the proposed optimizations are lightweight and fully compatible with the decentralized protocol, enabling autonomous execution and utility realization at each node. Experimental results demonstrate that XIPFS significantly improves node performance without compromising decentralization. Zhaoyan Shen, Mengying Zhao, Dongxiao Yu, Bingzhe Li |
ICDE | 5 |
| 2025 | You Only Spectralize Once: Taking a Spectral Detour to Accelerate Graph Neural NetworkabstractTraining Graph Neural Networks (GNNs) often relies on repeated, irregular, and expensive message-passing operations over all nodes (e.g., $N$), leading to high computational overhead. To alleviate this inefficiency, we revisit the GNNs training from a spectral perspective. In many real-world graphs, node features and embeddings exhibit sparse representation in the Graph Fourier domain. This inherent spectral sparsity aligns well with the principles of Compressed Sensing, which posits that signals sparse in one transform domain can be accurately reconstructed from a significantly reduced number of measurements. This observation motivates the design of a more efficient GNNs that operates predominantly in compressed spectral subspace. Thus, we propose You Only Spectralize Once (YOSO), a GNN training scheme that performs single Graph Fourier Transformation to project features onto a learnable orthonormal Fourier basis, retaining only $M$ spectral coefficients ($M \ll N$). The entire GNN computation is then carried out in reduced spectral domain. Final full-graph embeddings are recovered only at output layer by solving a bounded $\ell_{2,1}$-regularized optimization problem. Theoretically, drawing upon Compressed Sensing theory, we prove stable recovery throughout training by showing that the projection onto our learnable Fourier basis can satisfy the Restricted Isometry Property when $M=\mathcal{O}(k \log N)$ for $k$-row-sparse spectra, acting as the measurement process. Empirically, YOSO achieves an average 74\% reduction in training time across five benchmark datasets compared to state-of-the-art methods, while maintaining competitive accuracy. Zhichun Guo, Guanpeng Li, Bingzhe Li |
NeurIPS | 4 |
| 2025 | Advancing Archival Data Storage: The Promises and Challenges of DNA Storage SystemabstractAs the volume of data is rapidly produced every day, there is a need for the storage media to keep up with the growth rate of digital data created. Despite emerging storage solutions that have been proposed such as Solid State Drive with quad-level cells or penta-level cells, Shingled Magnetic Recording, Linear Tape-Open, and so on, these technologies still fall short of meeting the demand for preserving huge amounts of available data. Moreover, current storage solutions have a limited lifespan, often lasting just a few years. To ensure long-term preservation, data must be continuously migrated to new storage drives. Therefore, there is a need for alternative storage technologies that not only offer high storage capacity but also long persistency. In contrast to existing storage devices, Synthetic Deoxyribonucleic Acid (DNA) storage emerges as a promising candidate for archival data storage, offering both high-density storage capacity and the potential for long-term data preservation. In this article, we will introduce DNA storage, discuss the capabilities of DNA storage based on the current biotechnologies, discuss possible improvements in DNA storage, and explore further improvements with future technologies. Currently, the limitations of DNA storage are due to its weaknesses including high error rates, long access latency, and so on. In this article, we will focus on possible DNA storage research issues based on its relevant bio and computer technologies. Also, we will provide potential solutions and forward-looking predictions about the development and the future of DNA storage. We will discuss DNA storage from the following five perspectives: (1) We will describe the basic background of DNA storage including the basic technologies of read/write DNA storage, data access processes such as Polymerase Chain Reaction-based random access, encoding schemes from digital data to DNA, and required DNA storage format. (2) We will describe the issues of DNA storage based on the current technologies including bio-constraints during the encoding process such as avoiding long homopolymers and containing certain GC contents, different types of errors in synthesis and sequencing processes, low practical capacity with the current technologies, slow read and write performance, and low encoding density for random accesses. (3) Based on the previously mentioned issues, we will summarize the current solutions for each issue, and also give and discuss the potential solutions based on the future technologies. (4) From a system perspective, we will discuss how the DNA storage system will look if the DNA storage becomes commercialized and is widely equipped in archive systems. Some questions will be discussed, including: (i) How do we efficiently index data in DNA storage? (ii) What is a good storage hierarchical storage system with DNA storage? (iii) What will DNA storage be like with the development of technology? (5) Finally, we will provide a comparison with other competitive technologies. Alex Sensintaffar, Yixun Wei, Li Ou, David Hung-Chang Du, Bingzhe Li |
ACM Trans. Storage | 5 |
| 2024 | An Encoding Scheme to Enlarge Practical DNA Storage Capacity by Reducing Primer-Payload CollisionsabstractDeoxyribonucleic Acid (DNA), with its ultra-high storage density and long durability, is a promising long-term archival storage medium and is attracting much attention today. A DNA storage system encodes and stores digital data with synthetic DNA sequences and decodes DNA sequences back to digital data via sequencing. Many encoding schemes have been proposed to enlarge DNA storage capacity by increasing DNA encoding density. However, only increasing encoding density is insufficient because enhancing DNA storage capacity is a multifaceted problem. Yixun Wei, Bingzhe Li, David Hung-Chang Du |
ASPLOS (2) | 2 |
| 2024 | Grafu: Unleashing the Full Potential of Future Value Computation for Out-of-core Synchronous Graph ProcessingabstractAs graphs exponentially grow recently, out-of-core graph systems have been invented to process large-scale graphs by keeping massive data in storage. Among them, many systems process the graphs iteration-by-iteration and provide synchronous semantics that allows easy programmability by forcing the computation dependency of vertex values between iterations. On the other hand, although future value computation is an effective IO optimization for out-of-core graph systems by computing vertex values of future iterations in advance, it is challenging to take full advantage of future value computation while guaranteeing iteration-based dependency. In fact, based on our investigation, even state-of-the-art work along this direction has a wide gap from optimality in IO reduction and further requires substantial overhead in computation as well as extra memory consumption. Tsun-Yu Yang, Cale England, Bingzhe Li, Ming-Chang Yang |
ASPLOS (2) | 4 |
| 2024 | Celeritas: Out-of-Core Based Unsupervised Graph Neural Network via Cross-Layer Computing 2024abstractGraph neural networks (GNN) one of the most popular neural network models, are extensively applied in graph-related fields, including drug discovery, recommendation systems, etc. Unsupervised graph learning as one type of GNN plays a crucial role in various graph-related missions like node classification and edge prediction. However, with the increasing size of real-world graph datasets, processing such massive graphs in host memory becomes impractical, and GNN training demands a substantial storage volume to accommodate the vast amount of graph data. Consequently, GNN training results in significant I/O migration between the host and storage. Although state-of-the-art frameworks have made strides in mitigating I/O overhead by considering embedding locality, their GNN frameworks still suffer from long training times. In this paper, we propose a fully out-of-core framework, called Celeritas, which speeds up the unsupervised GNN training on a single machine by co-designing the GNN algorithm and storage systems. First, based on the theoretical analysis, we propose a new partial combination operation to enable the embedding updates across GNN layers. This cross-layer computing achieves future computation for the embedding stored in memory to save data migration. Second, due to the dependency between embedding and edges, we consider their data locality together. Based on the cross-layer computing property, we propose a new loading order to fully utilize the data stored in the main memory to save I/O. Finally, a new sampling scheme called two-level sampling is proposed associated with a new partition algorithm to further reduce data migration and computation overhead while maintaining similar training accuracy. The real system experiments indicate that the proposed Celeritas can reduce the total training time of different G NN models from 44.76 % to 73.85 % compared to state-of-art schemes for different graph datasets. Tsun-Yu Yang, Ming-Chang Yang, Zhaoyan Shen, Bingzhe Li |
HPCA | 5 |
| 2024 | An Access Pattern-aware Hybrid Learning-based and Conventional Mapping for Solid-State DrivesabstractLearning-based address mapping strategy has been proposed as an alternative to the conventional one-to-one mapping strategy for logical to physical address translation in flash-based Solid-State Drives (SSDs). This strategy employs a machine learning (ML)-based index to replace the memory-intensive mappings, thereby saving memory space. However, current learning-based mapping strategies mainly focus on the sequential characteristics of I/O workloads, making them less effective in scenarios with frequent updates. In this work, we introduce an access pattern-aware hybrid learning-based and conventional mapping scheme for SSDs, called APH-FTL. This approach features a machine learning-driven data classifier that employs distinct address mapping strategies for data requests with varying access patterns. Additionally, we develop a benefit-guided space regulation strategy to manage in-device memory space effectively and provide a solution to the challenge of mapping policies transformation. We evaluate APH-FTL through simulations using various real-world workloads. Experimental results demonstrate that APH-FTL significantly improves performance, reducing response time by up to 76.59% and write amplification (WA) by up to 82.97% compared with state-of-the-art schemes. Xiaosu Guo, Jie Wang 0128, Zhaoyan Shen, Dongxiao Yu, Zhiping Jia, Bingzhe Li |
ICCAD | 7 |
| 2024 | ASHL: An Adaptive Multi-Stage Distributed Deep Learning Training Scheme for Heterogeneous EnvironmentsabstractWith the increment of data sets and models sizes, distributed deep learning has been proposed to accelerate training and improve the accuracy of DNN models. The parameter server framework is a popular collaborative architecture for data-parallel training, which works well for homogeneous environments by properly aggregating the computation/communication capabilities of different workers. However, in heterogeneous environments, the resources of different workers vary a lot. Some stragglers may seriously limit the whole speed, which impacts the overall training process. In this paper, we propose an adaptive multi-stage distributed deep learning training framework, named ASHL, for heterogeneous environments. First, a profiling scheme is proposed to capture the capabilities of each worker to reasonably plan the training and communication tasks on each worker, and lay the foundation for the formal training. Second, a hybrid-mode training scheme (i.e., coarse-grained and fined-grained training) is proposed to balance the model accuracy and training speed. The coarse-grained training scheme (named AHL) adopts an asynchronous communication strategy, which involves less frequent communications. Its main goal is to make the model quickly converge to a certain level. The fine-grained training stage (named SHL) uses a semi-asynchronous communication strategy and adopts a high communication frequency. Its main goal is to improve the model convergence effect. Finally, a compression-based communication scheme is proposed to further increase the communication efficiency of the training process. Our experimental results show that ASHL reduces the overall training time by more than 35% to converge to the same degree and has better generalization ability compared with state-of-the-art schemes like ADSP. Zhaoyan Shen, Qingxiang Tang, Tianren Zhou, Yuhao Zhang 0006, Zhiping Jia, Dongxiao Yu, Zhiyong Zhang 0006, Bingzhe Li |
IEEE Trans. Computers | 8 |
| 2024 | A Semantic-Integrated LSM-Tree-Based Key-Value Storage Engine for Blockchain SystemsabstractBlockchain systems play an important role in distributed ledgers, database systems, etc. As more and more blocks are mined, the storage burden of blockchain system is significantly increased. The current blockchain system uniformly transforms all its data into key-value (KV) items and stores them to the underlying Log-Structure Merged tree (LSM-tree) storage engine ignoring the software semantics. Consequently, it not only aggravates the write amplification effect of the storage engine, but also increases the redundancy of data query steps, resulting in the performance bottleneck of blockchain system. In this paper, we propose a semantic-integrated LSM-tree based Key-Value storage engine for blockchain systems, called Block-LSM, which significantly improves the data synchronization and data query efficiency of blockchain system. Specifically, we first design a shared prefix scheme to transform blockchain data into ordered KV pairs to alleviate the key range overlaps of different levels in the underlying LSM-tree based storage engine. Moreover, we propose to maintain several semantic-orientated memory buffers to isolate different kinds of blockchain data, and implement memory buffer space management strategy to further improve memory efficiency. To save space overhead, Block-LSM further aggregates multiple blocks into a group and assigns the same prefix to all KV items from the same block group. We also reduce step redundancy in transaction queries by modifying the body data storage format. Finally, we implement Block-LSM in a real blockchain environment and conduct a series of comparative experiments with the typical blockchain system Ethereum. The evaluation results show that Block-LSM significantly reduces up to 7.56× storage write amplification and increases throughput by 8.64× compared with the original Ethereum design. In terms of data lookups (i.e. transaction and account lookup), Block-LSM improves the throughput by 50% compared to the original Ethereum design. Yuhao Zhang 0006, Xiaojun Cai, Zhiping Jia, Zhaoyan Shen, Yi Wang 0003, Zili Shao, Bingzhe Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2023 | Reinforcement Learning-Assisted Management for Convertible SSDsabstractConvertible SSDs, which allow flash cells to convert between different types of flash cells (e.g., SLC/MLC/TLC/QLC), are designed for achieving both high performance and high density. However, previous designs with two types of flash cells encounter a performance cliff degradation once the flash cells of single bit mode are consumed. In this work, we propose a novel level-based convertible SSD (e.g., including SLC-MLC-QLC), named RL-cSSD, that adopts an intermediate layer (e.g., MLC) as a performance cushion. A reinforcement learning-assisted device management scheme is designed to coordinate the data allocation, garbage collection and flash conversion processes considering both the SSD internal status and workload patterns. We evaluated RL-cSSD with various real-world workloads based on simulation. The experimental results show that the proposed RL-cSSD provides 72.98% higher performance on average compared with state-of-the-art schemes. Zhiping Jia, Mengying Zhao, Zhaoyan Shen, Bingzhe Li |
DAC | 6 |
| 2023 | K8sES: Optimizing Kubernetes with Enhanced Storage Service-Level ObjectivesabstractKubernetes (k8s) is a system for managing containerized applications across multiple hosts. It offers automatic deployment, maintenance, scaling, and resource management for applications. Applications in k8s usually have different storage requirements in the form of service-level objectives (SLOs). However, the current k8s storage management has several limitations which cause explicit performance and cost overhead. K8s administrators have to configure storage in advance manually, and users must know configurations and capabilities of provided storage. Users' storage SLOs can be easily violated in k8s.In this paper, we design and implement k8s Enhanced Storage (k8sES) which efficiently supports applications with various storage SLOs along with all other requirements in the Kubernetes environment. We design and incorporate storage scheduling as part of the node scheduling process in k8s. Applications will be scheduled onto the correct nodes and storage without intervention from either users or administrators. Proper storage resources will be dynamically carved based on users' storage SLOs. In addition, we provide a tool to monitor the I/O activities of both applications and storage devices in k8sES. The evaluation shows that k8sES can better meet users' storage SLOs along with other requirements. Also, k8sES can achieve higher resource utilization efficiency with overhead similar to that of the current k8s. Hao Wen 0001, Zhichao Cao 0002, Bingzhe Li, David Hung-Chang Du, Ayman Abouelwafa, Doug Voigt, Shiyong Liu, Jim Diehl, Fenggang Wu |
ICCD | 3 |
| 2023 | DP-DNA: A Digital Pattern-Aware DNA Encoding Scheme to Improve Encoding Density of DNA StorageabstractWith the rapid increase of available digital data, Deoxyribonucleic Acid (DNA) storage is identified as such a promising candidate due to its long persistency and high areal density, especially for archival storage systems. However, due to biochemical constraints, currently the encoding densities of various DNA storage systems are much less than this upper bound. In this paper, we propose a new Digital Pattern-aware DNA encoding scheme, called DP-DNA, which satisfies the DNA biochemical constraints and efficiently stores digital data in DNA storage with high encoding density. To satisfy the biochemical constraints, our proposed scheme is based on several rotation codes. DP-DNA first analyzes the patterns of each short binary sequence, which will be encoded to a DNA strand, and then selects an appropriate code for encoding the target binary sequence to achieve a high encoding density. An additional encoding field is added to the DNA encoding format, which can distinguish the encoding scheme used for each DNA strand, and thus we can decode DNA data back to its original digital data. Moreover, a new 2bit-code with the highest encoding density (i.e., 2bits/nt) is proposed to add to the pool of code candidates to further increase the encoding density. In addition, a variable-length scheme is applied to increase the feasibility of using 2bit-code scheme. Finally, the experimental results indicate that the proposed DP-DNA achieves 5.9% - 103.5% higher encoding density than the existing encoding schemes with various datasets. Bingzhe Li, Li Ou, Bo Yuan 0001, David Hung-Chang Du |
MASCOTS | 1 |
| 2023 | Hardware-aware neural architecture search for stochastic computing-based neural networks on tiny devices
Yuhong Song, Edwin H.-M. Sha, Qingfeng Zhuge, Rui Xu 0013, Xiaowei Xu 0004, Bingzhe Li, Lei Yang 0018 |
J. Syst. Archit. | 6 |
| 2023 | ChainKV: A Semantics-Aware Key-Value Store for Ethereum SystemabstractThe Log-Structure Merged tree (LSM-tree) based key-value (KV) store has been widely adopted as the storage engine for blockchain systems, such as Ethereum, in which blockchain data are uniformly transformed into randomly distributed KV items for persistence. However, blockchain semantics are ignored during this process, making the blockchain storage suffer from heavy read/write amplification problems. Moreover, as the Ethereum network scales up, tremendous data further exacerbates its storage burden. Until now, most studies have focused on sharding, data archiving, decentralized distributed storage, etc., to mitigate the burden of the storage layer. However, the incompatibility between Ethereum semantics and the characteristics of the storage engine is ignored. In this paper, we present ChainKV, a new semantics-aware storage paradigm to improve the storage management performance for the Ethereum system. Firstly, based on Ethereum blockchain semantics, ChainKV separately stores different types of data in multiple storage zones in the KV store to mitigate the read/write amplification problem. Secondly, following the mechanism of the verification process in the authenticated data structure (ADS), a new ADS data transformer is proposed to exploit the data locality when persisting ADS. Moreover, a new space gaming caching policy is adopted to coordinate the cache space management for two independent storage zones. Finally, we propose an optional lightweight node crash recovery mechanism to eliminate functional redundancy between the Ethereum protocol and the storage engine. The experimental results indicate that ChainKV outperforms the prior Ethereum systems by up to 1.99× and 4.20× for synchronization and query operations, respectively Bingzhe Li, Xiaojun Cai, Zhiping Jia, Lei Ju 0001, Zili Shao, Zhaoyan Shen |
Proc. ACM Manag. Data | 2 |
| 2023 | A Multiagent Reinforcement Learning-Assisted Cache Cleaning Scheme for DM-SMRabstractTo support nonsequential writes, persistent cache (PC) is constructed in drive managed SMR (DM-SMR) drive. However, PC cleaning introduces drastic performance degradation and enlarges tail latencies. In this article, we propose to utilize reinforcement learning (RL) to mitigate the long-tail latency of PC cleaning. Our scheme uses the lightweight$Q$-learning method to monitor and learn the idle time of I/O workloads, based on which PC cleaning is intelligently guided, thus maximally exploit idle time between requests and hiding tail latency from normal requests. In addition, a multiagent RL scheme with clustering algorithm is adopted to further mitigate the tail latencies and adapt to variable workloads. We emulate a DM-SMR drive inside a Linux device driver to implement our proposed scheme. According to the experimental results, our scheme can effectively reduce the tail latency by 59.45% at the 99.9th percentile and the average latency by 48.75% compared with a typical shingled magnetic recording (SMR) design. Zhaoyan Shen, Yungang Pan, Yuhao Zhang 0006, Zhiping Jia, Xiaojun Cai, Bingzhe Li, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | BSC: Block-based Stochastic Computing to Enable Accurate and Efficient TinyMLabstractAlong with the progress of AI democratization, machine learning (ML) has been successfully applied to edge applications, such as smart phones and automated driving. Nowadays, more applications require ML on tiny devices with extremely limited resources, like implantable cardioverter de-fibrillator (ICD), which is known as TinyML. Unlike ML on the edge, TinyML with a limited energy supply has higher demands on low-power execution. Stochastic computing (SC) using bitstreams for data representation is promising for TinyML since it can perform the fundamental ML operations using simple logical gates, instead of the complicated binary adder and multiplier. However, SC commonly suffers from low accuracy for ML tasks due to low data precision and inaccuracy of arithmetic units. Increasing the length of the bitstream in the existing works can mitigate the precision issue but incur higher latency. In this work, we propose a novel SC architecture, namely Block-based Stochastic Computing (BSC). BSC divides inputs into blocks, such that the latency can be reduced by exploiting high data parallelism. Moreover, optimized arithmetic units and output revision (OUR) scheme are proposed to improve accuracy. On top of it, a global optimization approach is devised to determine the number of blocks, which can make a better latency-power trade-off. Experimental results show that BSC can outperform the existing designs in achieving over 10% higher accuracy on ML tasks and over$6\times$power reduction. Yuhong Song, Edwin H.-M. Sha, Qingfeng Zhuge, Rui Xu 0013, Yongzhuo Zhang, Bingzhe Li, Lei Yang 0018 |
ASP-DAC | 6 |
| 2022 | Work-in-Progress: ExpCache: Online-Learning based Cache Replacement Policy for Non-Volatile MemoryabstractAs emerging memory technologies (e.g., non-volatile memory (NVM)) coming out and machine learning algorithms successfully applying to different fields, the potentials of cache replacement policy for NVM-based systems with the integration of machine learning algorithms are worthy of being exploited to improve the performance of computer systems. In this work, we proposed a machine learning based cache replacement algorithm, named ExpCache, to improve the system performance with NVM as the main memory. By considering the non-volatility characteristic of the NVM devices, we split the whole NVM into two caches, including a read cache and a write cache, for retaining different types of requests. The pages in each cache are managed by both LRU and LFU policies for balancing the recency and frequency of workloads. The online Expert machine learning algorithm is responsible for selecting a proper policy to evict a page from one of the caches based on the access patterns of workloads. In experimental results, the proposed ExpCache outperforms previous studies in terms of hit ratio and the number of dirty pages written back to storage. Jinfeng Yang, Bingzhe Li, Zhaoyan Shen, David Hung-Chang Du, David J. Lilja |
CASES | 2 |
| 2022 | Re-LSM: A ReRAM-Based Processing-in-Memory Framework for LSM-Based Key-Value StoreabstractLog-structured merge (LSM) tree based key-value (KV) stores organize writes into hierarchical batches for high-speed writing. However, the notorious compaction process of LSM-tree severely hurts system performance. It not only involves huge I/O operations but also consumes tremendous computation and memory resources. In this paper, first we find that when compaction happens in the high levels (i.e., L0, L1) of the LSM-tree, it may saturate all system computation and memory resources, and eventually stall the whole system. Based on this observation, we present Re-LSM, a ReRAM-based Processing-in-Memory (PIM) framework for LSM-based Key-Value Store. Specifically, in Re-LSM, we propose to offload certain computation and memory-intensive tasks in the high levels of the LSM-tree to the ReRAM-based PIM space. A high parallel ReRAM compaction accelerator is designed by decomposing the three-phased compaction into basic logic operating units. Evaluation results based on db_bench and YCSB show that Re-LSM achieves 2.2× improvement on the throughput of random writes compared to RocksDB, and the ReRAM-based compaction accelerator speedups the CPU-based implementation by 64.3× and saves 25.5× energy. Zhaoyan Shen, Yiheng Tong, Zhiping Jia, Lei Ju 0001, Jiezhi Chen, Bingzhe Li |
ICCAD | 7 |
| 2022 | HL-DNA: A Hybrid Lossy/Lossless Encoding Scheme to Enhance DNA Storage Density and Robustness for ImagesabstractWith the storage's demand for high density and long-term preservation, Deoxyribonucleic Acid (DNA) has become a promising candidate to satisfy the requirement of archival storage for rapidly increased digital volume. However, due to the biochemical constraints, DNA storage faces critical issues of low practical capacity and robustness. In this paper, we target image applications and propose to apply approximation to DNA storage to improve the overall encoding density and robustness of DNA storage by using a hybrid lossy and lossless encoding scheme (called HL-DNA). Several lossy and lossless encoding schemes (lossy and lossless codes) are proposed and used to encode incoming binary sequences. These two types of codes are coordinated to balance the encoding density and errors. The lossless codes are used to limit the errors and the lossy codes are used to improve the encoding density. Moreover, the introduced approximation and newly proposed hybrid encoding schemes in one DNA strand can improve the robustness of DNA storage. Finally, the experimental results indicate that the proposed HL-DNA improves the encoding density of DNA storage and makes it much close to the ideal case. Also, HL-DNA achieves higher robustness to the injected errors than other DNA storage codes. David Hung-Chang Du, Li Ou, Bingzhe Li |
ICCD | 4 |
| 2022 | Machine Learning-based Adaptive Migration Algorithm for Hybrid Storage SystemsabstractHybrid storage systems are prevalent in most large-scale enterprise storage systems since they balance storage performance, storage capacity and cost. The goal of such systems is to serve the majority of the I/O requests from high-performance devices and store less frequently used data in low-performance devices. A large data migration volume between tiers can cause a huge overhead in practical hybrid storage systems. Therefore, how to balance the trade-off between the migration cost and potential performance gain is a challenging and critical issue in hybrid storage systems. In this paper, we focused on the data migration problem of hybrid storage systems with two classes of storage devices. A machine learning-based migration algorithm called K-Means assisted Support Vector Machine (K-SVM) migration algorithm is proposed. This algorithm is capable of more precisely classifying and efficiently migrating data between performance and capacity tiers. Moreover, this K-SVM migration algorithm involves a K-Means clustering algorithm to dynamically select a proper training dataset such that the proposed algorithm can significantly reduce the volume of migrating data. Finally, the real implementation results indicate that the ML-based algorithm reduces the migration data volume by about 40% and achieves 70% lower latency than other algorithms. Milan Shetti, Bingzhe Li, David Hung-Chang Du |
NAS | 2 |
| 2022 | A Survey of Blockchain Data Management SystemsabstractBlockchain has been widely deployed in various fields, such as finance, education, and public services. Blockchain has decentralized mechanisms with persistency and auditability and runs as an immutable distributed ledger, where transactions are jointly performed through cryptocurrency-based consensus algorithms by worldwide distributed nodes. There have been many survey papers reviewing the blockchain technologies from different perspectives, e.g., digital currencies, consensus algorithms, and smart contracts. However, none of them have focused on the blockchain data management systems. To fill in this gap, we have conducted a comprehensive survey on the data management systems, based on three typical types of blockchain, i.e., standard blockchain, hybrid blockchain, and DAG ( Directed Acyclic Graph )-based blockchain. We categorize their data management mechanisms into three layers: blockchain architecture, blockchain data structure, and blockchain storage engine, where block architecture indicates how to record transactions on a distributed ledger, blockchain data structure refers to the internal structure of each block, and blockchain storage engine specifies the storage form of data on the blockchain system. For each layer, the works advancing the state-of-the-art are discussed together with technical challenges. Furthermore, we lay out several possible future research directions for the blockchain data management systems. Bingzhe Li, Wanli Chang 0001, Zhiping Jia, Zhaoyan Shen, Zili Shao |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | Reinforcement Learning-Assisted Cache Cleaning to Mitigate Long-Tail Latency in DM-SMRabstractDM-SMR adopts Persistent Cache (PC) to accommodate non-sequential write operations. However, the PC cleaning process induces severe long-tail latency. In this paper, we propose to mitigate the tail latency of PC cleaning by using Reinforcement Learning (RL). Specifically, a real-time lightweight Q-learning model is built to analyze the idle window of I/O workloads, based on which PC cleaning is judiciously scheduled, thereby maximally utilizing the I/O idle window and effectively hiding the tail latency from regular requests. We implement our technique inside a Linux device driver with an emulated SMR drive. Experimental results show that our technique can reduce the tail latency by 57.65% at 99.9th percentile and the average response time by 46.11% compared to a typical SMR design. Yungang Pan, Zhiping Jia, Zhaoyan Shen, Bingzhe Li, Wanli Chang 0001, Zili Shao |
DAC | 4 |
| 2021 | Block-LSM: An Ether-aware Block-ordered LSM-tree based Key-Value Storage EngineabstractEthereum as one of the largest blockchain systems plays an important role in the distributed ledger, database systems, etc. As more and more blocks are mined, the storage burden of Ethereum is significantly increased. The current Ethereum system uniformly transforms all its data into key-value (KV) items and stores them to the underlying Log-Structure Merged tree (LSM-tree) storage engine ignoring the software semantics. Consequently, it not only exacerbates the write amplification effect of the storage engine but also hurts the performance of Ethereum. In this paper, we proposed a new Ethereum-aware storage model called Block-LSM, which significantly improves the data synchronization of the Ethereum system. Specifically, we first design a shared prefix scheme to transform Ethereum data into ordered KV pairs to alleviate the key range overlaps of different levels in the underlying LSM-tree based storage engine. Moreover, we propose to maintain several semantic-orientated memory buffers to isolate different kinds of Ethereum data. To save space overhead, Block-LSM further aggregates multiple blocks into a group and assigns the same prefix to all KV items from the same block group. Finally, we implement Block-LSM in the real Ethereum environment and conduct a series of experiments. The evaluation results show that Block-LSM significantly reduces up to 3.7× storage write amplification and increases throughput by 3× compared with the original Ethereum design. Bingzhe Li, Xiaojun Cai, Zhiping Jia, Zhaoyan Shen, Yi Wang 0003, Zili Shao |
ICCD | 2 |
| 2021 | EFM: Elastic Flash Management to Enhance Performance of Hybrid Flash MemoryabstractNAND-based flash memory has become a prevalent storage media due to its low access latency and high performance. By setting up different incremental step pulse programming (ISPP) values and threshold voltages, the tradeoffs between lifetime and access latency in NAND-based flash memory can be exploited. The existing studies that exploit the tradeoffs by using heuristic algorithms do not consider the dynamically changed access latency due to wearing-out, resulting in low access performance. In this paper, we proposed a new Elastic Flash Management scheme, called EFM, to manage data in hybrid flash memory, which consists of multiple physical regions with different read/write latencies according to their ISPP values and threshold voltages. EFM includes a Long-Term Classifier (LT-Classifier) and a Short-Term Classifier (ST-Classifier) to accurately track dynamically changed workloads by considering current quantitative differences of read/write latencies and workload access patterns. Moreover, a reduced effective wearing management is proposed to prolong the lifetime of flash memory by scheduling write-intensive workloads to the region with a reduced threshold voltage and the lowest write cost. Experimental results indicate that EFM reduces the average read/write latencies by about 54% - 296% and obtain 17.7% lifetime improvement on average compared to the existing studies. Bingzhe Li, Bo Yuan 0001, David Hung-Chang Du |
ICCD | 1 |
| 2021 | WAS-Deletion: Workload-Aware Secure Deletion Scheme for Solid-State DrivesabstractDue to the intrinsic properties of Solid-State Drives (SSDs), invalid data remain in SSDs before erased by a garbage collection process, which increases the risk of being attacked by adversaries. Previous studies use erase and cryptography based schemes to purposely delete target data but face extremely large overhead. In this paper, we propose a Workload-Aware Secure Deletion scheme, called WAS-Deletion, to reduce the overhead of secure deletion by three major components. First, the WAS-Deletion scheme efficiently splits invalid and valid data into different blocks based on workload characteristics. Second, the WAS-Deletion scheme uses a new encryption allocation scheme, making the encryption follow the same direction as the write on multiple blocks and vertically encrypts pages with the same key in one block. Finally, a new adaptive scheduling scheme can dynamically change the configurations of different regions to further reduce secure deletion overhead based on the current workload. The experimental results indicate that the newly proposed WAS-Deletion scheme can reduce the secure deletion cost by about 1.2x to 12.9x compared to previous studies. Bingzhe Li, David Hung-Chang Du |
ICCD | 1 |
| 2021 | IMG-DNA: approximate DNA storage for imagesabstractDeoxyribonucleic Acid (DNA) as a storage medium with high density and long-term preservation properties can satisfy the requirement of archival storage for rapidly increased digital volume. The read and write processes of DNA storage are error-prone. Images widely used in social media have the properties of fault tolerance which are well fitted to the DNA storage. However, prior work simply investigated the feasibility of DNA storage storing different types of data and simply store images in DNA storage, which did not fully investigate the fault-tolerant potential of images in the DNA storage. In this paper, we proposed a new image-based DNA system called IMG-DNA, which can efficiently store images in DNA storage with improved DNA storage robustness. First, a new DNA architecture is proposed to fit JPEG-based images and improve the image's robustness in DNA storage. Moreover, barriers inserted in DNA sequences efficiently prevent error propagation in images of DNA storage. The experimental results indicate that the proposed IMG-DNA achieves much higher fault-tolerant than prior work. Bingzhe Li, Li Ou, David Hung-Chang Du |
SYSTOR | 1 |
| 2021 | HeuristicDB: a hybrid storage database system using a non-volatile memory block deviceabstractHybrid storage systems are widely used in big data fields to balance system performance and cost. However, due to a poor understanding of the characteristics of database block requests, past studies in this area cannot fully utilize the performance gain from emerging storage devices. This study presents a hybrid storage database system, called HeuristicDB, which uses an emerging non-volatile memory (NVM) block device as an extension of the database buffer pool. To consider the unique performance behaviors of NVM block devices and the block-level characteristics of database requests, a set of heuristic rules that associate database (block) requests with the appropriate quality of service for the purpose of caching priority are proposed. Using online analytical processing (OLAP) and online transactional processing (OLTP) benchmarks, both trace-based examination and system implementation on MySQL are carried out to evaluate the effectiveness of the proposed design. The experimental results indicate that HeuristicDB provides up to 75% higher performance and migrates 18X fewer data between storage and the NVM block device than existing systems. Jinfeng Yang, Bingzhe Li, David J. Lilja |
SYSTOR | 2 |
| 2021 | TrackLace: Data Management for Interlaced Magnetic RecordingabstractInterlaced Magnetic Recording (IMR) is a promising technology which achieves higher data density and lower write amplification (WA) than Shingled Magnetic Recording (SMR). In IMR, top tracks and bottom tracks are interlaced so each bottom track is partially overlapped with two adjacent top tracks. Top tracks can be updated without any WA, but bottom track updates require reading and rewriting of affected valid data on the two neighboring top tracks. There are few published studies discussing WA in IMR drives. We propose TrackLace to reduce WA for IMR. TrackLace consists of three techniques: Z-Alloc allocates user data to the tracks in alternating directions and spreads unallocated tracks among allocated tracks; Top-Buffer opportunistically utilizes unallocated top tracks to buffer bottom track updates; and Block-Swap progressively swaps bottom track hot data with top track cold data during high space utilization. To further optimize TrackLace performance, we propose a virtual frame design that can keep the relocated block (due to Top-Buffer or Block-Swap) close to its original location and an adaptive buffering mechanism that can avoid unnecessary redirections depending on the write locality. Evaluations show that TrackLace can reduce WA by 45 percent and lower average latency by 31percent compared with baseline schemes. Fenggang Wu, Bingzhe Li, Baoquan Zhang, Zhichao Cao 0002, Jim Diehl, Hao Wen 0001, David Hung-Chang Du |
IEEE Trans. Computers | 2 |
| 2021 | FluidSMR: Adaptive Management for Hybrid SMR DrivesabstractHybrid Shingled Magnetic Recording (H-SMR) drives are the most recently developed SMR drives, which allow dynamic conversion of the recording format between Conventional Magnetic Recording (CMR) and SMR on a single disk drive. We identify the unique opportunities of H-SMR drives to manage the tradeoffs between performance and capacity, including the possibility of adjusting the SMR area capacity based on storage usage and the flexibility of dynamic data swapping between the CMR area and SMR area. We design and implement FluidSMR, an adaptive management scheme for hybrid SMR Drives, to fully utilize H-SMR drives under different workloads and capacity usages. FluidSMR has a two-phase allocation scheme to support a growing usage of the H-SMR drive. The scheme can intelligently determine the sizes of the CMR and the SMR space in an H-SMR drive based on the dynamic changing of workloads. Moreover, FluidSMR uses a cache in the CMR region, managed by a proposed loop-back log policy, to reduce the overhead of updates to the SMR region. Evaluations using enterprise traces demonstrate that FluidSMR outperforms baseline schemes in various workloads by decreasing the average I/O latency and effectively reducing/controlling the performance impact of the format conversion between CMR and SMR. Fenggang Wu, Bingzhe Li, David Hung-Chang Du |
ACM Trans. Storage | 2 |
| 2020 | Can We Store the Whole World's Data in DNA Storage?
Bingzhe Li, Nae Young Song, Li Ou, David Hung-Chang Du |
HotStorage | 1 |
| 2019 | Energy-Efficient Convolutional Neural Networks with Deterministic Bit-Stream ProcessingabstractStochastic computing (SC) has been used for low-cost and low power implementation of neural networks. Inherent inaccuracy and long latency of processing random bit-streams have made prior SC-based implementations inefficient compared to conventional fixed-point designs. Random or pseudo-random bitstreams often need to be processed for a very long time to produce acceptable results. This long latency leads to a significantly higher energy consumption than binary design counterparts. Low-discrepancy sequences have been recently used for fast-converging deterministic computation with stochastic constructs. In this work, we propose a low-cost, low-latency, and energy-efficient implementation of convolutional neural networks based on low-discrepancy deterministic bit-streams. Experimental results show a significant reduction in the energy consumption compared to previous random bitstream-based implementations and to the optimized fixed-point design with no quality degradation. S. Rasoul Faraji, M. Hassan Najafi, Bingzhe Li, David J. Lilja, Kia Bazargan |
DATE | 3 |
| 2019 | Sliding Look-Back Window Assisted Data Chunk Rewriting for Improving Deduplication Restore Performance
Zhichao Cao 0002, Shiyong Liu, Fenggang Wu, Bingzhe Li, David Hung-Chang Du |
FAST | 5 |
| 2019 | TASecure: Temperature-Aware Secure Deletion Scheme for Solid State DrivesabstractWith the increasing concerns of security, the secure deletion for SSDs becomes very costly due to its out-of-place update (i.e., an update is performed in a new location leaving the old data un-touched). Some previous studies used a combined erase-based and cryptography-based method to find a heuristic deletion scheme. However, the deletion overhead is still large since a key may be associated with too many pages. Therefore, how to reduce the secure deletion overhead further is becoming an interesting research issue. In this paper, a temperature-aware secure deletion scheme (TASecure) is proposed. Data with the same temperature are stored in the same region in order to reduce the overhead of secure deletion. The temperature of a data is referred to the frequency of the data is updated. Moreover, the regions based on their temperature are associated with different numbers of keys and are applied with different secure deletion schemes. Finally, the experimental results indicate that our proposed secure deletion scheme can reduce the secure deletion cost about 1.2x to 19x compared to previous studies. Bingzhe Li, David Hung-Chang Du |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | Low Cost Hybrid Spin-CMOS Compressor for Stochastic Neural NetworksabstractWith expansion of neural network (NN) applications lowering their hardware implementation cost becomes an urgent task especially in back-end applications where the power-supply is limited. Stochastic computing (SC) is a promising solution to realize low-cost hardware designs. Implementation of matrix multiplication has been a bottleneck in previous stochastic neural networks (SC-NNs). In this paper, we introduce spintronic components into the design of SC-NNs. A novel spin-CMOS matrix multiplier is proposed in which the stochastic multiplications are performed by CMOS AND gates while the sum of products is implemented by spintronic compressor gates. The experimental results indicate that compared to the conventional binary implementations the proposed hybrid spin-CMOS architecture can achieve over 125x, 4.5x and 43x; reduction in terms of power, energy and area consumptions, respectively. Moreover, compared to previous CMOS-based SC-NNs, our design saves the power by 3.1x - 7.3x, reduces energy consumption by 3.1x - 7.3x and decreases area by 1.4x - 7.6x while maintaining similar recognition rates. Bingzhe Li, Jiaxi Hu, M. Hassan Najafi, Steven J. Koester, David J. Lilja |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | ZoneAlloy: Elastic Data and Space Management for Hybrid SMR Drives
Fenggang Wu, Bingzhe Li, Zhichao Cao 0002, Baoquan Zhang, Ming-Hong Yang, Hao Wen 0001, David Hung-Chang Du |
HotStorage | 2 |
| 2019 | HAML-SSD: A Hardware Accelerated Hotness-Aware Machine Learning based SSD ManagementabstractSolid state drive (SSD) as a fast storage device has been playing an important role across many applications from mobile computing to large distributed systems in recent years. However, the performance of the SSD can be degraded tremendously due to the intrinsic properties of NAND-based flash memory including limited erase cycles and asymmetric write and erase operations. Previous works separated hot/cold data into different blocks in order to improve SSD performance. “Hotness” is typically defined as the cumulative update frequencies of pages. However, we believe that an additional new parameter, average update time interval, should also be considered into the “hotness” definition associated with the update frequency. Moreover, to adaptively classify hot/cold data, a machine learning algorithm is applied to better accommodate the dynamically changed I/O access patterns of traces. In this paper, a machine learning (ML) based SSD management called HAML-SSD is proposed. The purpose of applying the ML algorithm is to dynamically cluster the data with similar “hotness” based on a new definition of “hotness”. Thus, a two-dimension clustering algorithm is used for storing the pages categorized into the same cluster within the same block. Moreover, to obtain reasonable training time, a specific hardware component called HAML-unit is designed in the SSD. Finally, the experimental results indicate that the HAML-SSD decreases the response time around 26.3% - 57.7% compared to previous works with the evaluation of real traces. Bingzhe Li, Chunhua Deng, Jinfeng Yang, David J. Lilja, Bo Yuan 0001, David Hung-Chang Du |
ICCAD | 1 |
| 2019 | Low-Cost Stochastic Hybrid Multiplier for Quantized Neural NetworksabstractWith increased interests of neural networks, hardware implementations of neural networks have been investigated. Researchers pursue low hardware cost by using different technologies such as stochastic computing (SC) and quantization. More specifically, the quantization is able to reduce total number of trained weights and results in low hardware cost. SC aims to lower hardware costs substantially by using simple gates instead of complex arithmetic operations. However, the advantages of both quantization and SC in neural networks are not well investigated. In this article, we propose a new stochastic multiplier with simple CMOS transistors called the stochastic hybrid multiplier for quantized neural networks. The new design uses the characteristic of quantized weights and tremendously reduces the hardware cost of neural networks. Experimental results indicate that our stochastic design achieves about 7.7x energy reduction compared to its counterpart binary implementation while maintaining slightly higher recognition error rates than the binary implementation. Compared to previous stochastic neural network implementations, our work derives at least 4x, 9x, and 10x reduction in terms of area, power, and energy, respectively. Bingzhe Li, M. Hassan Najafi, David J. Lilja |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2019 | Neural Network Classifiers Using a Hardware-Based Approximate Activation Function with a Hybrid Stochastic MultiplierabstractNeural networks are becoming prevalent in many areas, such as pattern recognition and medical diagnosis. Stochastic computing is one potential solution for neural networks implemented in low-power back-end devices such as solar-powered devices and Internet of Things (IoT) devices. In this article, we investigate a new architecture of stochastic neural networks with a hardware-oriented approximate activation function. The newly proposed approximate activation function can be hidden in the proposed architecture and thus reduce the whole hardware cost. Additionally, to further reduce the hardware cost of the stochastic implementation, a new hybrid stochastic multiplier is proposed. It contains OR gates and a binary parallel counter, which aims to reduce the number of inputs of the binary parallel counter. The experimental results indicate the newly proposed approximate architecture without hybrid stochastic multipliers achieves more than 25%, 60%, and 3x reduction compared to previous stochastic neural networks, and more than 30x, 30x, and 52% reduction compared to conventional binary neural networks, in terms of area, power, and energy, respectively, while maintaining the similar error rates compared to the conventional neural networks. Furthermore, the stochastic implementation with hybrid stochastic multipliers further reduces area about 18% to 80%, power from 15% to 113.1%, and energy about 15% to 131%, respectively. Bingzhe Li, Yaobin Qin, Bo Yuan 0001, David J. Lilja |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2019 | NetStorage: A synchronized trace-driven replayer for network-storage system evaluation
Bingzhe Li, Hao Wen 0001, Farnaz Toussi, Clark Anderson, Bernard A. King-Smith, David J. Lilja, David Hung-Chang Du |
Perform. Evaluation | 1 |
| 2018 | Data Management Design for Interlaced Magnetic Recording
Fenggang Wu, Baoquan Zhang, Zhichao Cao 0002, Hao Wen 0001, Bingzhe Li, Jim Diehl, David Hung-Chang Du |
HotStorage | 5 |
| 2018 | Tier-Code: An XOR-Based RAID-6 Code with Improved Write and Degraded-Mode Read PerformanceabstractThe RAID-6 configuration is more tolerant of disk failures than other RAID levels because of its ability to tolerate two disk failures. However, previous RAID-6 codes suffer from two major overheads - the time of encoding or decoding processes plus the need to access multiple blocks when updating parities or recovering failed blocks. For example, the PS and Reed-Solomon codes do not have optimal computation complexity, while P-code, X-code and RDP-code must access multiple blocks to update parities during write operations. This work proposes a new XOR- based RAID-6 code, called Tier-code, which not only achieves the optimal parity computation complexity, but also increases the write and degraded-mode read performance compared to previous codes. It uses two tiers of coding, one at the block level and the other at the chunk level. Experimental results of software testing, simulation and ASIC synthesis for this new hierarchical code demonstrate that Tier-code can outperform the previous RAID-6 codes in both write performance and degraded-mode read performance while maintaining the optimal computation complexity in both hardware and software implementations. Bingzhe Li, Soheil Mohajer, Weikang Qian, David J. Lilja |
NAS | 1 |
| 2017 | Neural Network Classifiers Using Stochastic Computing with a Hardware-Oriented Approximate Activation FunctionabstractNeural networks are becoming prevalent in many areas, such as pattern recognition and medical diagnosis. Stochastic computing is one potential solution for neural networks implemented in low-power back-end devices such as solar-powered devices and Internet-of-things (IoT) devices. In this paper, we investigate a new architecture of stochastic neural networks with a hardware-oriented approximate activation function. The new proposed approximate activation function can be omitted while keeping the functionality well. Thus, it reduces the stochastic implementation complexity and hardware costs. Moreover, the new architecture significantly improves recognition error rates compared to previous stochastic neural networks with sigmoid function. Three classical types of neural networks are explored, multiple layer perceptron (MLP), restricted Boltzmann machine (RBM) and convolutional neural networks (CNN). The experimental results indicate the new proposed architecture achieves more than 25%, 60% and 3× reduction than previous stochastic neural networks, and more than 30×, 30× and 52% reduction than conventional binary neural networks, in terms of area, power and energy, respectively, while maintaining the similar error rates compared to the conventional neural networks. Bingzhe Li, Yaobin Qin, Bo Yuan 0001, David J. Lilja |
ICCD | 1 |
| 2017 | Kinetic Action: Performance Analysis of Integrated Key-Value Storage Devices vs. LevelDB ServersabstractWith the rise of cloud storage and many data intensive applications, there is an unprecedented growth in the volume of unstructured data. In response, key-value object storage is becoming more popular for the ease with which it can store, manage, and retrieve large amounts of this data. Seagate recently launched Kinetic direct-access-over-Ethernet hard drives which incorporate a LevelDB key-value store inside each drive. In this work, we evaluate these drives using micro as well as macro benchmarks to help understand the performance limits, trade-offs, and implications of replacing traditional hard drives with Kinetic drives in data centers and high performance systems. We perform in-depth throughput and latency benchmarking of these Kinetic drives (each acting as a tiny independent server) from a client machine connected to them via Ethernet. We compare these results to a SATA-based and a faster SAS-based traditional server running LevelDB. Our sample Kinetic drives are CPU-bound, but they still average sequential write throughput of 63 MB/sec and sequential read throughput of 78 MB/sec for 1 MB value sizes. They also demonstrate unique Kinetic features including direct disk-to-disk data transfer. Our macro benchmarking using the Yahoo Cloud Serving Benchmark (YCSB) shows that mid-range LevelDB servers outperform the Kinetic drives for several workloads; however, this is not always the case. For larger value sizes, even these first generation sample Kinetic drives outperform a full server for several different workloads. Manas Minglani, Jim Diehl, Bingzhe Li, Dongchul Park, David J. Lilja, David Hung-Chang Du |
ICPADS | 4 |
| 2017 | TraceRAR: An I/O Performance Evaluation Tool for Replaying, Analyzing, and Regenerating TracesabstractAdopting a new technology, such as a new storage system, is a complicated process because the supporting ecosystems also have to be changed. As a result, any new technology requires exhaustive performance evaluation to justify the cost of switching. However, synthetic workloads or benchmarks typically cannot completely characterize the actual workload. On the other hand, the time and effort required to obtain an appropriate trace can be prohibitive. This work presents a block-level performance measurement tool for storage systems combined with a trace re- player, a trace characteristics analyzer, and a trace re-generator. This new tool is compatible with several different platforms, including Linux and AIX. The purpose of the tool is to evaluate system performance when executing a given application, and to help users determine which system best fits their specific application. Additionally, the trace analyzer can provide details about the characteristics of a given trace. Using the trace analysis results, the re- generator can produce arbitrarily long I/O traces to improve the accuracy of the performance evaluation. The tool also can be used to determine whether a particular system can be adapted to a specific application, and to make comparisons between systems. Bingzhe Li, Farnaz Toussi, Clark Anderson, David J. Lilja, David Hung-Chang Du |
NAS | 1 |
| 2016 | Using Stochastic Computing to Reduce the Hardware Requirements for a Restricted Boltzmann Machine ClassifierabstractArtificial neural networks are powerful computational systems with interconnected neurons. Generally, these networks have a very large number of computation nodes which forces the designer to use software-based implementations. However, the software based implementations are offline and not suitable for portable or real-time applications. Experiments show that compared with the software based implementations, FPGA-based systems can greatly speed up the computation time, making them suitable for real-time situations and portable applications. However, the FPGA implementation of neural networks with a large number of nodes is still a challenging task. Bingzhe Li, M. Hassan Najafi, David J. Lilja |
FPGA | 1 |
| 2016 | Ps-Code: A New Code for Improved Degraded Mode Read and Write Performance of RAID SystemsabstractThe newer storage systems often have to deal with multiple disk failures. Thus, RAID (Redundant Array of Independent Disk) is becoming increasingly important for two primary benefits: (1) high read performance and (2) robust fault tolerance. In addition, the recent RAID codes achieve better IO performance by using only XORs (optimal computational complexity) to compute parities. However, this achieved optimal computational complexity limits the further IO performance improvement. In this paper, we are proposing a new code called PS-code (Parity Stripe code) that employs improved Cauchy Reed Solomon codes for computing the parity to further improve write and degraded mode read performance. Furthermore, we extend the novelty by reducing the number of writes required for updating the parity. We simulated the PS-code in the DiskSim environment and devised a new metric called multi- block access complexity to perform an improved evaluation of the performance of the PS-code on workloads that represent real life scenario such as multi-block updates. Also, the experimental results demonstrate that the PS-code, on average, achieves 66%-86% better write performance and achieves 8.9%-23.6% higher degraded-mode read performance compared to previous works including P-code, RDP code, X-code, and HV-code. Finally, the comparison between the vertical and horizontal codes demonstrates that the vertical codes have better read performance than the horizontal codes in most cases. Bingzhe Li, Manas Minglani, David J. Lilja |
NAS | 1 |
| 2015 | An FPGA implementation of a Restricted Boltzmann Machine classifier using stochastic bit streamsabstractArtificial neural networks (ANNs) usually require a very large number of computation nodes and can be implemented either in software or directly in hardware, such as FPGAs. Software-based approaches are offline and not suitable for real-time applications, but they support a large number of nodes. FPGA-based implementations, in contrast, can greatly speedup the computation time. However, resource limitations in an FPGA restrict the maximum number of computation nodes in hardware-based approaches. This work exploits stochastic bit streams to implement the Restricted Boltzmann Machine (RBM) handwritten digit recognition application completely on an FPGA. Exploiting this approach saves a large number of hardware resources making the FPGA-based implementation of large ANNs feasible. Bingzhe Li, M. Hassan Najafi, David J. Lilja |
ASAP | 1 |