VLDB 2026 Research / reviewers in the wild / expert
Tianyu Wang 0009
dblp:35/8397-9
· DBLP profile ↗
42ranked-venue papers
8as first author
38since 2021 · last 2026
0000-0001-7030-1990ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 6 first-author · 31 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EEMU: A FEMU-based Accurate, Parametric Ethernet-SSD EmulatorabstractEthernet-SSDs incorporate built-in Ethernet connectivity, allowing them to function as standalone network-attached storage devices. This architecture enables efficient and scalable disaggregated storage by providing direct network access without host dependencies. Understanding their internal architecture and fine-grained performance interactions is critical for advancing this emerging design, yet is difficult to study on proprietary hardware. In this paper, we present EEMU, a configurable and accurate Ethernet-SSD emulator built on the FEMU framework. EEMU reproduces the end-to-end NVMeoF I/O translation path and models latency across three key components: the Network Interface, the NVMeoF Target Module, and the AXI Bus. It further employs a fine-grained parametric latency model that decomposes end-to-end latency into component-level contributions, enabling systematic exploration of architectural trade-offs via controlled parameterization. Experiments show that EEMU closely matches both component-level behaviors and end-to-end performance across diverse Ethernet-SSD configurations. Xikun Jiang, Tianyu Wang 0009, Zhaoyan Shen, Zili Shao |
DATE | 3 |
| 2026 | An Efficient Cloud Storage Model with Compacted Metadata Management for Performance Monitoring Timeseries Systems
Tianyu Wang 0009, Zili Shao |
FAST | 2 |
| 2026 | Region-Based Collaborative Caching With Joint Latency and Lifetime Optimization for Hybrid SMR-Flash StorageabstractShingled Magnetic Recording (SMR) disks have been improved significantly over traditional Hard Disk Drives (HDDs) in storage capacity and cost by a shingled-like structure that overlaps with adjacent tracks. However, the overall system performance is severely sacrificed due to the large portion of Read-Merge-Write (RMW) operations triggered by the non-sequential nature of SMR disks. To handle such performance degradation, Persistent Cache (PC) and built-in NAND flash cache are applied to absorb non-sequential writes. However, when the cache is full, triggering write-back operations extends the I/O response time. Additionally, the erase of NAND flash also compromises its lifetime.This paper aims to manage the NAND flash region and the original SMR band region uniformly instead of only adopting NAND flash as a cache layer. To achieve this, we propose a Region-based Co-optimized strategy named Multi-Regional Collaborative Management (MCM). Our approach segments I/O requests using a fast and efficient Bloom filter alongside a Band-based partitioning module. Additionally, the NAND flash memory region is managed based on the characteristics of the segmented data, enabling maximum control over frequent data updates in the NAND flash memory. The NAND flash cache will not serve as the first-level cache as before. Meanwhile, through Region-aware Wear-leveling (WL) and garbage collection (GC) strategies, the lifetime of NAND flash memory is extended, and cleaning efficiency is improved. In addition, further analysis of the impact of the wear-leveling parameters on RMWs and NAND flash lifetime under different workloads. According to the experimental results, compared with the typical Skylight (baseline), our scheme reduces the average response time and write amplification by 74.56% and 98.99%, respectively. At the same time, our approach also outperforms some state-of-the-art solutions, such as MU-RMW, RMW-F, and FC. Zhengang Chen, Zhi-Ping Shi 0002, Tianyu Wang 0009 |
IEEE Trans. Computers | 5 |
| 2026 | Resolving Gray Code Dilemma With Bidirectional Programming for Efficient QLC SSDsabstractQLC NAND flash is widely adopted in modern storage systems. By trading off read/write performance for storage density through a “time-for-space" approach, QLC enables ultra-high storage capacity. To mitigate performance degradation, Gray code and the two-step programming (TSP) algorithm are used. However, Gray code also has limitations: multiple Gray codes incur circuit overhead, while a single Gray code causes extra I/O latency overhead. This paradox seems unsolvable at first glance, requiring an innovative solution that maintains I/O performance without additional circuit overhead. A promising solution lies in selecting an appropriate Gray code and preventing hot data placement on slow physical pages. This paper proposes BDP, a novel Bi-Directional Programming scheme that adopts a single Gray code to fit both traditional (forward) and reverse programming directions based on TSP. The objective of BDP is to resolve the inherent contradiction between I/O performance preservation and implementation overhead. BDP optimizes the system performance through hardware/software co-design. At the hardware level, a fixed Gray code is employed to avoid additional circuit complexity. At the software level, two strategies (i.e. hotness-aware data allocation and background data migration) are proposed to further mitigate the misplacement of hot data on slow pages in QLC SSDs. The experimental results demonstrate that BDP significantly reduces the allocation of hot data to slow pages and enhances overall I/O performance compared to representative schemes. Yi Wang 0003, Shaoqi Li, Yongbiao Zhu, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | PCD-ORAM: A Path-Aware and Cross-Layer Design to Enhance Data Locality in Oblivious RAM
Yi Wang 0003, Zhencheng Wang, Weixuan 'Vincent' Chen, Xianhua Wang, Chenlin Ma, Tianyu Wang 0009, Rui Mao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | Registry: Enhancing Vertex Reusability for GCN Inference on Hybrid Stacked Memory
Zhaoyu Zhong, Jiaxian Chen, Yunhao Dong, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001, Yi Wang 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Groundhog: Accelerating Spatio-Temporal Data Analytics With Fine-Grained In-Storage ProcessingabstractWith the rapid growth of mobile devices and applications, a prodigious number of spatio-temporal data are generated constantly. To process these data for applications like traffic forecasting, existing spatio-temporal systems rely on the move-data-to-computation paradigm. However, this approach incurs significant data movement overhead between hosts and storage devices, particularly when a spatio-temporal query is executed on a non-preferred data layout or when the query has a small result size due to its inherent nature. To address this issue, this work introduces Groundhog, an efficient in-storage computing technique designed specifically for spatio-temporal queries, aimed at reducing unnecessary data movement and computations. Groundhog introduces three key designs for efficient in-storage computing: (i) a self-contained and segment-based storage model, which is lightweight for in-storage computing and enables fine-grained pruning for spatio-temporal queries; (ii) a set of fine-grained techniques to optimize query processing inside storage devices for spatio-temporal queries; and (iii) an in-storage-computing-aware query planner, which offloads spatio-temporal queries in a fine-grained manner using a cost-based approach. We implemented Groundhog on real hardware and demonstrated how to apply fine-grained techniques to accelerate various spatio-temporal queries. Extensive experiments conducted on real-world datasets demonstrate that Groundhog achieves significant performance improvements, with latency reductions of up to$81\%$for widely used spatio-temporal queries compared to host computing solutions. Tianyu Wang 0009, Zizhan Chen, Zili Shao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2026 | AKV: Agile Read-Efficiently Key-Value OLTP Engine for Non-Volatile MemoryabstractNon-volatile memory (NVM), as an emergingstor age technology, offers several advantageous features for OLTP engines, including byte-addressability, high capacity, low energy consumption, and data persistence across power failures. Despite these benefits, the current mainstream OLTP engines still commonly adopt a hybrid architecture that deeply couples DRAM with NVM, which results in a complex system architecture and high recovery costs. In this paper, we aim to construct a highly available, stable, and recoverable OLTP engine that guarantees ACID properties through anagile system architecture. We introduce AKV (Agile Key-Value), an NVM-only OLTP storage engine designed to provide effective space utilization, high throughput, and fast failure recovery. AKV addresses the challenges of NVM space management, write redundancy, and concurrency control with two novel techniques: dual-version concurrency control and circular dual-version storage. Experimental results demonstrate that AKV achieves higher throughput (up to 69.7%) and faster recovery (up to 54×) compared to existing storage engines in most scenarios of the TPC-C benchmarks. Additionally, the codebase of AKV (4k+ lines) is more concise than that of SOTA OLTP engines like Zen (8k+ lines) and Falcon (11k+ lines). In addition, this study innovatively proposes a read abort optimization strategy based on dynamic version changes. The experimental results show that this strategy can significantly reduce the transaction abort rate of AKV in specific workload scenarios while maintaining stable system throughput, achieving a maximum reduction of up to 73% in the abort count. Jianbin Qin, Tianyu Wang 0009, Yuxing Chen 0003, Anqun Pan, Rui Mao 0001, Yu-Xuan Qiu, Makoto Onizuka, Chuan Xiao 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Unifying Two Operators with One PIM: Leveraging Hybrid Bonding for Efficient LLM Inference
Jiaxian Chen, Yuxuan Qi, Kaoyi Sun, Zhiliang Lin, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003 |
APPT | 5 |
| 2025 | Move Less, Retrieve Fast: A Retrieval-in-Memory Architecture for Language ModelsabstractRetrieval-augmented language models (RALMs) have attracted widespread attention for addressing the limitations of traditional large language models. However, challenges involved in retrieval, including substantial data movement and irregular access patterns, seriously impact the efficiency and deployment of RALMs. The emerging 3D-stacked processing-in-memory (PIM) architecture, characterized by its high memory bandwidth and near-data computing capabilities, presents a promising solution for efficient retrieval. To support large-scale retrieval in RALMs, the PIM architecture should be carefully designed with joint software and hardware optimization. This paper presents Rimast, a retrieval-in-memory architecture for fast retrieval in RALMs. The objective is to minimize data movement and improve overall performance through hardwaresoftware co-design. At the hardware level, a hierarchical PIM architecture with a retrieval-in-memory dataflow is designed to reduce unnecessary data transfer. At the software level, skew-free data mapping and adaptive offloading strategies are proposed to address the irregular access patterns associated with retrieval in RALMs. We demonstrate the effectiveness of the proposed Rimast using extensive experiments. The experimental results demonstrate that Rimast effectively reduces data movement, achieving average speedups of $273 \times 55 \times$, and $2.41 \times$ over CPUs, GPUs, and prior art accelerators, respectively. Jiaxian Chen, Yuxuan Qi, Jianan Yuan, Kaoyi Sun, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003 |
DAC | 5 |
| 2025 | Anchor First, Accelerate Next: Revolutionizing GNNs with PIM by Harnessing Stationary DataabstractSubstantial data movement caused by irregular graph topologies hinders the efficient processing of graph neural networks (GNNs). Although the emerging near-bank processing-in-memory (PIM) architecture offers a promising solution to reduce data transfer between memory and computing units, cross-bank communication remains a critical challenge, limiting the benefits of PIM architectures. Our findings indicate that only $35.6 \%$ of the data can stay stationary within PIM units on average, with the rest requiring movement due to graph dependencies. This situation worsens as the number of PIM units increases, reducing the ratio to $18.7 \%$. In this paper, we argue that to fully leverage PIM architectures, systems must maximize stationary data and minimize the movement of non-stationary data. Following this principle, we propose Anchor, a scalable PIM architecture that exploits stationary data for GNNs through a hardware-software co-design approach. To maximize stationary data, we introduce the graph partitioning algorithm Mastav, which carefully allocates vertices and edges to preserve data locality. To minimize the movement of non-stationary data, we employ a two-step strategy. First, a customized dataflow ensures that non-stationary data is accessed and distributed exactly once. Second, an optimized communication mechanism reduces redundant data transfers through critical paths. Our extensive experiments demonstrate that Anchor significantly reduces processing latency and data movement compared to representative schemes. Jiaxian Chen, Yuxuan Qi, Yongbiao Zhu, Jianan Yuan, Kaoyi Sun, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003 |
DAC | 6 |
| 2025 | Expanding Logical Space Freely: A Memory-efficient Mapping Table Design for Compressional SSDsabstractCompressional SSDs can be a mixed blessing. While they offer users expanded logical space beyond the physical capacity, they complicate the Flash Translation Layer (FTL) design by requiring a larger Logical-to-Physical (L2P) address mapping, which places a heavier burden on the limited in-SSD memory, leading to a degraded I/O performance. In this paper, we aim to reduce the memory footprint of the L2P mapping table in compressional SSDs by proposing a novel N-to-1 L2P mapping table design that consolidates multiple logical entries into a single entry. This approach eliminates the duplication of physical page numbers when a physical page contains several compressed logical pages. To accommodate the dynamic compression ratios of real-world workloads, we introduce promotion and demotion that enable mapping table entries to migrate between pages with different compression ratios. Additionally, to address the issue of partial invalidation-where some compressed logical pages within a physical page are invalid due to the N-to-1 mapping-we present a compression-aware garbage collection algorithm aimed at minimizing the number of copy operations for partially invalid physical pages. We have implemented our design in MQsim, a widely used SSD simulator, and have conducted a series of experiments to evaluate the effectiveness of the proposed techniques. The results demonstrate that our approach significantly reduces the mapping table size in compressional SSDs, leading to an improved mapping table cache hit ratio and a reduced I/O latency compared to traditional compressional SSDs. Zixuan Huang 0011, Tianyu Wang 0009, Kecheng Huang, Zelin Du, Zili Shao |
DAC | 2 |
| 2025 | MiniWear: Minimizing Flash Wear via Hybrid Persistent Cache for Extended EF-SMR LifetimeabstractAs the huge discrepancy between traffic and capacity persists, the lifetime of flash in EF-SMR systems faces a grave issue. EF-SMR systems combine NAND flash with Shingled Magnetic Recording (SMR) disks to achieve both low cost and high performance. However, previous research has primarily focused on issues such as write amplification and tail-latency in EF-SMR disks, overlooking the critical issue of flash lifetime. Studying the durability of EF-SMR systems is essential for developing future high-performance, low-cost storage solutions.This paper presents MiniWear, a hybrid persistent cache (PC) design aimed at extending the lifetime of EF-SMR systems. MiniWear adopts a hybrid medium persistent cache and proposes a customized scheduling strategy to reduce flash wear without impacting the EF-SMR system performance. At the hardware level, the hybrid PC of EF-SMR, composed of flash and SMR disk, is organized into Flash-PC and SMR-PC. At the software level, a fine-grained scheduling strategy is proposed to better manage PC resources. Additionally, we introduce a proactive balancing strategy to address PC resource idleness. Experimental results show that, compared to existing methods, MiniWear can reduce flash wear by up to 66.67%. Chenlin Ma, Kaoyi Sun, Yuxuan Qi, Jiaxian Chen, Xiaochuan Zheng, Tianyu Wang 0009, Yi Wang 0003 |
DAC | 6 |
| 2025 | Dancer: Dynamic Compression and Quantization Architecture for Deep Graph Convolutional NetworkabstractGraph Convolutional Networks (GCNs) have been widely applied in fields such as social network analysis and recommendation systems. Recently, deep GCNs have emerged, enabling the exploration of deeper hidden information. Compared to traditional shallow GCNs, deep GCNs feature significantly more layers, leading to considerable computational and data movement challenges. Processing-In-Memory (PIM) offers a promising solution for efficiently handling GCNs by enabling near-data computation, thus reducing data transfer between processing units and memory. However, previous work mainly focused on shallow GCNs and has shown limited performance with deep GCNs. In this paper, we present Dancer, an innovative PIM-based GCN accelerator. Dancer optimizes data movement during the inference process, significantly improving efficiency and reducing energy consumption. Specifically, we introduce a novel compressed graph storage architecture and a dynamic quantization technique to minimize data transfers at each layer of the GCN. Additionally, through a detailed analysis of weight dynamics changes, we propose a sparsity propagation strategy to further alleviate the computational and data transfer burden between layers. Experimental results demonstrate that, compared to current state-of-the-art methods, Dancer achieves 3.7× speedup, 7.6× energy efficiency, and reduces of 9.6× DRAM access on average. Yunhao Dong, Zhaoyu Zhong, Yi Wang 0003, Chenlin Ma, Tianyu Wang 0009 |
DATE | 5 |
| 2025 | A Practical Learning-Based FTL for Memory Constrained Mobile Flash StorageabstractThe rapidly growing mobile market is pushing flash storage manufacturers to expand capacity into the terabyte range. However, this presents a significant challenge for mobile storage management: more logical-to-physical page mappings are desired to be efficiently managed and cached while the available caching space is extremely limited. This motivates us to shift toward a new learning-based paradigm: rather than maintaining mappings for individual pages, the learning-based approach can represent mapping relationships for a set of continuous pages. However, to construct linear models, existing methods that either consume the already-limited memory space or reuse flash garbage collection demonstrate poor model construction capabilities or significantly degrade flash performance, making them impractical for real-world use. In this paper, we propose LFTL, a practical, learning-based on-demand flash translation layer design for flash management in mobile devices. In contrast to prior work that centered around gathering sufficient mappings for linear model construction, our key insight is that linear patterns can be extracted and refined by leveraging the orderly, LPA-aligned write stream typical of mobile devices. By doing this, highly accurate linear models can be constructed regardless of the constraints of mobile device's cache limitation. We have implemented a fully functional prototype of LFTL based on FEMU. Our evaluation results show that LFTL is more adaptable to memory-constrained storage devices than state-of-the-art learning-based approaches. Zelin Du, Kecheng Huang, Tianyu Wang 0009, Xin Yao 0008, Renhai Chen, Zili Shao |
DATE | 3 |
| 2025 | One Gray Code Fits All: Optimizing Access Time with Bi-Directional Programming for QLC SSDsabstractGray code, a voltage-level-to-data-bit translation scheme, is widely used in QLC SSDs. However, it causes the four data bits in QLC to exhibit significantly different read and write performance with up to 8 × latency variation, severely impacting the worst-case performance of QLC SSDs. This paper presents BDP, a novel Bi-Directional Programming scheme. Based on a fixed Gray code, BDP combines both the normal (forward) and reverse programming directions to enable runtime programming direction arbitration. Experimental results show that BDP can effectively improve the read and write performance of SSD compared to representative schemes. Shaoqi Li, Tianyu Wang 0009, Yongbiao Zhu, Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao |
DATE | 2 |
| 2025 | EF-IMR: Embedded Flash with Interlaced Magnetic Recording TechnologyabstractInterlaced Magnetic Recording (IMR), a technology that improves storage density through track overlap, introduces significant latency due to Read-Modify-Write (RMW) operations. Writing to overlapped tracks affects underlying tracks, requiring additional I/O operations to read, back up, and rewrite them, resulting in significant head movement latency. We propose EF-IMR, a new architecture that ensures crash consistency in IMR while minimizing RMW latency and head movement. EF-IMR reduces head movement during RMW operations and decreases redundant RMW operations. Evaluations under real-world, intensive I/O workloads show that EF-IMR reduces RMW latency by 20.11 % and head movement latency by 89.37% compared to existing methods. Chenlin Ma, Xiaochuan Zheng, Kaoyi Sun, Tianyu Wang 0009, Yi Wang 0003 |
DATE | 4 |
| 2025 | A Storage Model with Fine-Grained In-Storage Query Processing for Spatio-Temporal DataabstractMassive spatio-temporal data are continuously generated by various moving objects. To process these data for applications such as traffic forecasting, existing spatio-temporal systems all employ the move-data-to-computation paradigm. However, this approach suffers from significant data movement overhead between hosts and drives. To address this issue, this work introduces Groundhog, an efficient in-storage computing technique designed specifically for spatio-temporal queries, aimed at reducing unnecessary data movement and computations. Groundhog introduces three key designs for efficient in-storage computing: (i) a self-contained and segment-based storage model, which is lightweight for in-storage computing and enables fine-grained pruning for spatio-temporal queries; (ii) a set of fine-grained techniques to optimize query processing inside storage devices for spatio-temporal queries; and (iii) an in-storage-computing-aware query planner, which offloads spatio-temporal queries in a fine-grained manner using a cost-based approach. We implemented Groundhog on a real hardware board. Extensive experiments conducted on real-world datasets demonstrate that Groundhog achieves significant performance improvements, with latency reductions of up to 81 % for widely used spatio-temporal queries compared to host computing solutions. Tianyu Wang 0009, Zizhan Chen, Zili Shao |
ICDE | 2 |
| 2025 | CAMO: A High-Performance CIM-based Lightweight CNN Accelerator for Mobile DevicesabstractDigital Compute-in-Memory (CIM) macros revolutionize the Von Neumann architecture by significantly reducing data movements between CPU and memory. However, when dealing with lightweight CNNs with various convolution types, existing GEMM (general matrix multiplication)-oriented solutions suffer from underutilization and large activation traffic, leading to unsatisfied energy and area efficiencies. In this context, we propose CAMO, in which the key contributions are: (1) A novel convolution mapping mechanism suitable for CIM macros, that maximizes data reuse and reduces activation traffic. (2) A convolution-capable CIM macro, that also supports small-scale GEMM. (3) A CIM-based architecture that supports multiple computing modes. The experimental results show that CAMO achieves up to 31.49× performance speedup and 77.5% activation traffic reduction compared to the baseline architecture. Xin Ju 0005, Renyu Yang, Mei Wen, Junzhong Shen, Tianyu Wang 0009, Zhaoyan Shen, Zili Shao |
ISCAS | 5 |
| 2025 | SmartCache: Context-aware Semantic Cache for Efficient Multi-turn LLM InferenceabstractLarge Language Models (LLMs) for multi-turn conversations suffer from inefficiency: semantically similar queries across different user sessions trigger redundant computation and duplicate memory-intensive Key-Value (KV) caches. Existing optimizations such as prefix caching overlook semantic similarities, while typical semantic caches either ignore conversational context or are not integrated with low-level KV cache management.
We propose SmartCache, a system-algorithm co-design framework that tackles this inefficiency by exploiting semantic query similarity across sessions. SmartCache leverages a Semantic Forest structure to hierarchically index conversational turns, enabling efficient retrieval and reuse of responses only when both the semantic query and conversational context match.
To maintain accuracy during topic shifts, it leverages internal LLM attention scores—computed during standard prefill—to dynamically detect context changes with minimal computational overhead. Importantly, this semantic understanding is co-designed alongside the memory system: a novel two-level mapping enables transparent cross-session KV cache sharing for semantically equivalent states, complemented by a semantics-aware eviction policy that significantly improves memory utilization. This holistic approach significantly reduces redundant computations and optimizes GPU memory utilization.
The evaluation demonstrates SmartCache's effectiveness across multiple benchmarks. On the CoQA and SQuAD datasets, SmartCache reduces KV cache memory usage by up to $59.1\%$ compared to prefix caching and $56.0\%$ over semantic caching, while cutting Time-to-First-Token (TTFT) by $78.0\%$ and $71.7\%$, respectively. It improves answer quality metrics, achieving $39.9\%$ higher F1 and $39.1\%$ higher ROUGE-L for Qwen-2.5-1.5B on CoQA. The Semantic-aware Tiered Eviction Policy (STEP) outperforms LRU/LFU by $29.9\%$ in reuse distance under skewed workloads. Chengye Yu, Tianyu Wang 0009, Zili Shao, Song Jiang 0001 |
NeurIPS | 2 |
| 2025 | A Q-Learning-Based Display Energy Optimization Scheme for Android SystemsabstractMobile devices have gained immense popularity in recent years and have become an integral part of people’s daily lives. This surge in usage presents new challenges for energy conservation, particularly concerning the screens of mobile phones. Not only are screens being used for longer durations, but also their adjustment needs to be dynamically conducted during runtime. Existing approaches for display energy optimization mainly focus on the display content while disregarding user interactions. This limitation prevents the full exploitation of opportunities to reduce screen energy consumption during runtime. Detection-based display energy optimization approaches incorporate user interactions to some extent. However, these approaches require additional hardware and offer relatively coarse-grained control. To address these challenges, we propose QLEO, a Q-learning-based runtime display energy optimization scheme for Android systems. QLEO takes both the display and user interaction features into account. By dynamically classifying the user state into active and idle states based on these features, we enable adaptive screen brightness adjustments according to the user’s current state. We formulate display energy optimization as a constrained dynamic optimization problem and propose an integrated Q-learning model to effectively solve it. QLEO is designed as a module within the hardware abstraction layer, capable of obtaining display and user interaction features from the upper Android framework, and directly controlling the screen brightness through the underlying Linux kernel. We have implemented QLEO on real hardware and released the source code for public access. The experimental results demonstrate that QLEO achieves significant energy reduction without compromising the user experience. Zizhan Chen, Tianyu Wang 0009, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | MAP-SIM: A DNN-Specific Mapping Optimization Framework for Shared-Memory CPU-Systolic Array ArchitecturesabstractAs performance demands continue to rise, Shared-Memory Heterogeneous Systems (SMHSs) have been widely adopted for their ability to enable efficient communication and data sharing between different heterogeneous cores. However, existing SMHS face challenges in uneven workload distribution among heterogeneous cores and suboptimal mapping schemes, preventing them from fully leveraging their architectural advantages. To address these issues, this paper proposes a mapping-aware framework for modeling SMHSs called MAP-SIM. By performing performance modeling for CPUs and Systolic Arrays (SAs), and considering rational schemes for the partition and mapping of computational tasks, MAP-SIM aims to evaluate and optimize the computational performance of heterogeneous multicore architectures. The experimental results show that compared to previous work, MAP-SIM can increase simulation speed by 14 to 67 times and can also enhance the computational performance of SMHS by 1.4 to 4.4 times. Mei Wen, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008, Tianyu Wang 0009, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Leanor: A Learning-Based Accelerator for Efficient Approximate Nearest Neighbor Search via Reduced Memory AccessabstractApproximate Nearest Neighbor Search (ANNS) is a classical problem in data science. ANNS is both computationally-intensive and memory-intensive. As a typical implementation of ANNS, Inverted File with Product Quantization (IVFPQ) has the properties of high precision and rapid processing. However, the traversal of non-nearest neighbor vectors in IVFPQ leads to redundant memory accesses. This significantly impacts retrieval efficiency. A promising approach involves the utilization of learned indexes, leveraging insights from data distribution to optimize search efficiency. Existing learned indexes are primarily customized for low-dimensional data. How to tackle ANNS in high-dimensional vectors is a challenging issue. Yi Wang 0003, Jianan Yuan, Jiaxian Chen, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001 |
DAC | 5 |
| 2024 | PipeSSD: A Lock-free Pipelined SSD Firmware Design for Multi-core ArchitectureabstractModern SSD firmware is continuously optimized for higher parallelism to match the growing frontend PCIe bandwidth with more backend flash channels. Although a multi-core microprocessor is typically adopted to concurrently process independent NVMe requests from multiple NVMe queues, the existing one-to-many thread-request mapping model with each thread serving one or more incoming I/O requests has poor scalability due to severe lock contention problem, especially in cache management. Zelin Du, Shaoqi Li, Zixuan Huang 0011, Jin Xue, Kecheng Huang, Tianyu Wang 0009, Zili Shao |
DAC | 6 |
| 2024 | TPGraph: A Highly-scalable Time-partitioned Graph Model for Tracing BlockchainabstractThe continuously increasing volume of blockchain data presents significant challenges to blockchain traceability. Current tracing approaches, which rely on heuristic analytics directly applied to blockchain data, exhibit degraded tracking accuracy due to the involvement of only partial data. On the other hand, incorporating all blockchain data for analytics is not scalable given the unprecedented data volumes. Interestingly, we observe that despite the vast amount of blockchain data, blockchain tracing primarily focuses on time-related transaction spaces rather than all transactions. This insight motivates us to rethink the blockchain tracing problem to achieve both accuracy and scalability. Xiangao Chen, Tianyu Wang 0009, Kecheng Huang, Zili Shao |
SYSTOR | 2 |
| 2024 | TwinPilots: A New Computing Paradigm for GPU-CPU Parallel LLM InferenceabstractWhen trained Large Language Models (LLMs) become available, it is desirable to carry out LLM inferences at the user end with limited resources. A common belief on LLM inference is that GPU is essentially the only meaningful processor as almost all computation is tensor multiplication that GPU excels in. However, this belief and its practice are challenged by the fact that GPU has insufficient memory and runs at a much slower speed due to constantly waiting for data to be loaded from the CPU memory via a slow PCIe bus. This makes the CPU a processor with meaningful computing power that can be leveraged to accelerate the inference. Chengye Yu, Tianyu Wang 0009, Zili Shao, Linjie Zhu, Song Jiang 0001 |
SYSTOR | 2 |
| 2024 | A Bloom-Filter-Based Unique Address Checking Approach for DAG-Based Blockchain SystemsabstractWinternitz one-time signature (WOTS) is a quantum-resistant signature mechanism that has been widely used in direct acyclic graph (DAG)-based blockchain systems. However, it needs to generate a unique private/public key pair for each transaction, which severely limits blockchain performance by performing a time-consuming unique address checking process when generating transactions. This article proposes a bloom-filter-based approach, called ABACUS, to optimize the unique address checking process in WOTS. In ABACUS, we separate the large address space into multiple small subspaces and apply bloom filters to perform uniqueness checking for all addresses in one subspace. Specifically, we propose a two-level address space mechanism to strike a balance between the checking efficiency and the memory/storage space overhead of the bloom filter design. A bucket-based scalable bloom filter (SBF) design is proposed to match the growth of used addresses with efficient I/O access by storing all sub-bloom-filters together in one bucket. To further reduce disk I/Os, ABACUS incorporates an in-memory write buffer and a read-only cache. To reclaim wasted addresses, we also propose a layered address recycling mechanism by efficiently checking and reusing unused addresses that were identified as used due to false positives of bloom filters. We have implemented a fully functional prototype of ABACUS and integrated it into IOTA, a widely used DAG-based blockchain system. A series of experiments have been conducted on a private IOTA system. Experimental results show that ABACUS can significantly reduce the transaction generation time by up to four orders of magnitude, and achieve up to$3\times $boost on the system throughput with more than 80% space-saving. We have released the open-source code of ABACUS for public access. Tianyu Wang 0009, Zizhan Chen, Zhaoyan Shen, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | NICE: A Nonintrusive In-Storage-Computing Framework for Embedded ApplicationsabstractEmbedded machine learning applications face challenges related to massive data movement and high computational intensity, exacerbated by the limited performance of mobile devices. Computational storage devices (CSDs) pose huge potential for accelerating both data-intensive and computation-intensive embedded machine learning tasks by effectively reducing data movement and leveraging built-in accelerators. However, existing in-storage-computing (ISC) frameworks either require invasive customization of existing host driver layers or necessitate complex device firmware modifications, hindering the widespread deployment of CSDs. In addition, the lack of file semantics and the constrained internal resources within CSD implicitly compromise system performance and impact normal read/write performance. In this article, we aim to provide a nonintrusive in-storage-computing framework for embedded applications, named NICE. This framework includes an easy-to-use ISC programming interface that bypasses the kernel stack and requires no modification to the host NVMe driver, which is achieved through a novel hyper-addressing-based programming library and a file-aware page data layout within the CSD. In addition, we incorporate a lightweight kernel with coroutine-based command scheduling and several FPGA-based accelerators within the storage device firmware to enhance the performance of embedded machine learning applications while ensuring that the normal I/O performance remains unaffected. NICE is implemented on real CSD hardware integrated with ARM and FPGA. Experimental results demonstrate that our NICE framework can achieve an average latency performance improvement of$43.5\times $($9.32\times $) compared to CPU-(GPU-) based embedded machine learning solutions using the state-of-the-art NVIDIA Jetson NX platform, with$27.5\times $($4.3\times $) higher energy efficiency. NICE also has$34.2\times $less software and I/O performance overheads than state-of-the-art ISC frameworks. Tianyu Wang 0009, Yongbiao Zhu, Shaoqi Li, Jin Xue, Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Lightning Talk: Model, Framework and Integration for In-Storage Computing with Computational SSDsabstractIn-storage computing with computational SSDs is emerging as one effective solution for I/O bottlenecks in big data applications such as AI learning model training. Specifically, with in-SSD computing, computation can be pushed down to SSDs and the volume of the output data that will be transferred back to the host can be greatly reduced. However, there are several fundamental issues for applications to fully exploit in-SSD computing with simple and efficient function offloading. In this paper, we present three challenges for in-SSD computing, namely, data model, programming framework, and storage/computing integration, and discuss possible research directions. Tianyu Wang 0009, Jin Xue, Zelin Du, Yaotian Cui, Zili Shao |
DAC | 1 |
| 2023 | Region-based Flash Caching with Joint Latency and Lifetime Optimization in Hybrid SMR Storage SystemsabstractThe frequent Read-Modify-Write operations (RMWs) in Shingled Magnetic Recording (SMR) disks severely degrade the random write performance of the system. Although the adoption of persistent cache (PC) and built-in NAND flash cache alleviates some of the RMWs, when the cache is full, the triggered write-back operations still prolong I/O response time and the erasure of NAND flash also sacrifices its lifetime. In this paper, we propose a Region-based Co-optimized strategy named Multi-Regional Collaborative Management (MCM) to optimize the average response time by separately managing sequential/random and hot/cold data and extend the NAND flash lifetime by a region-aware wear-leveling strategy. The experimental results show that our MCM reduces 71 % of the average response time and 96% of RMWs on average compared with the Skylight (baseline). For the comparison with the state-of-art flash-based cache (FC) approach, we can still save the average response time and flash erase operations by 17.2 % and 33.32 %, respectively. Zhengang Chen, Zhi-Ping Shi 0002, Tianyu Wang 0009 |
DATE | 5 |
| 2023 | SoftSSD: enabling rapid flash firmware prototyping for solid-state drivesabstractRecently, solid-state drives (SSDs) have been used in a wide range of emerging data processing systems. Essentially, an SSD is a complex embedded system that involves both hardware and software design. For the latter, firmware modules such as the flash translation layer (FTL) orchestrate internal operations and flash management, and are crucial to the overall input/output performance of an SSD. Despite the rapid development of new SSD features in the market, the research of flash firmware has been mostly based on simulations due to the lack of a realistic and extensible SSD development platform. In this paper, we propose SoftSSD, a software-oriented SSD development platform for rapid flash firmware prototyping. The core of SoftSSD is a novel framework with an event-driven programming model. With the programming model, new FTL algorithms can be implemented and integrated into a full-featured flash firmware in a straightforward way. The resulting flash firmware can be deployed and evaluated on a hardware development board, which can be connected to a host system via peripheral component interconnect express and serve as a normal non-volatile memory express SSD. Different from existing hardware-oriented development platforms, SoftSSD implements the majority of SSD components (e.g., host interface controller) in software, so that data flows and internal states that were once confined in the hardware can now be examined with a software debugger, providing the observability and extensibility that are critical to the rapid prototyping and research of flash firmware. We describe the programming model and hardware design of SoftSSD. We also perform experiments with real application workloads on a prototype board to demonstrate the performance and usefulness of SoftSSD, and release the open-source code of SoftSSD for public access. Jin Xue, Renhai Chen, Tianyu Wang 0009, Zili Shao |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2023 | MSA: A Novel App Development Framework for Transparent Multiscreen Support on Android AppsabstractMultidisplay Android systems are emerging and require new display management. The current Android design exposes and offloads the multiscreen display management to app developers. As a consequence, apps cannot utilize multiscreen without redevelopment. This article proposes MSA, a novel Android app development framework for transparent multiscreen support. MSA cooperates with existing Android system services to map the single-screen views provided by apps to multiple screens and maps input events backward correspondingly. Following the new framework, the app development of single-screen Android systems and multiscreen Android systems are unified. Thus, both existing apps and future app development can directly utilize multiple screens with single-screen-based techniques, such as multiwindow and resizable-activity, instead of redevelopment through multiscreen-dedicated new methods such as presentation classes. We have implemented an MSA prototype with real hardware and released the source code for public access. Experimental results show that using MSA, without any modifications, existing apps can directly run and fully exploit multiple screens with better performance and less overhead compared with the state-of-the-art multiscreen Android system. Zizhan Chen, Tianyu Wang 0009, Jin Xue, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | MCMQ: Simulation Framework for Scalable Multi-Core Flash Firmware of Multi-Queue SSDsabstractSolid-state drives (SSDs) have been used in a wide range of emerging data processing systems. To fully utilize the massive internal parallelism delivered by SSDs, manufacturers begin to utilize high-performance multi-core microprocessors in scalable flash firmware to process I/O requests concurrently. Designing scalable multi-core flash firmwares requires simulation tools that can model the features of a multi-core environment. However, existing SSD simulators assume a single-threading execution model and are not capable of modelling overheads incurred by multi-threading firmware execution such as lock contentions. In this paper, we propose MCMQ, a novel framework for simulating scalable multi-core flash firmware. The framework is based on an emulated multi-core RISC processor and supports executing multiple I/O traces in parallel through a multi-queue interface. Experiment results show the effectiveness of the proposed framework. We have released the open-source code of MCMQ for public access. Jin Xue, Tianyu Wang 0009, Zili Shao |
DATE | 2 |
| 2022 | CNN Acceleration with Joint Optimization of Practical PIM and GPU on Embedded DevicesabstractThe lightweight AI models are proposed for embedded devices to mitigate computation and memory requirements. However, it still faces the memory wall problem (i.e., data movement between computing units and memory). Many processing-in-memory (PIM) based accelerators have been proposed to conquer this problem. Whereas, in practical deployment, there are still many challenges that need to be addressed.In this paper, we propose an embedded GPU-PIM integrated framework, with a fabricated HBM-PIM device from Samsung, to accelerate lightweight CNN inference. Considering various characteristics of different layers in lightweight CNNs, we first propose a fine-grained roofline model to guide the layer offloading to PIM cores. To fully utilize the computation and bandwidth resources of both GPU and PIM, we model the task allocation into a multi-dimensional knapsack problem (MDKP), which is addressed by a dynamic programming algorithm. Meanwhile, considering some practical issues such as PIM mode switch and memory address interleaving during the offloading, we further introduce a PIM-aware memory layout optimization strategy to reduce unnecessary inter-channel memory accesses and PIM mode switches. Experiment results show that our proposed framework achieves averagely the 5.9X and 37.9% inference performance boosts compared to the original GPU-only approach and the GPU-HBM approach, respectively. Tianyu Wang 0009, Zhaoyan Shen, Zili Shao |
ICCD | 1 |
| 2022 | TagTree: Global Tagging Index with Efficient Querying for Time Series DatabasesabstractModern time series databases come with a tag-based query interface that allows users to select time series, which are essentially sequences of timestamped data values, based on a set of specific tags. A tagging index is an important component that can efficiently provide such tag-based services. However, existing methods store tag information in external databases or time-partitioned data structures, which has a negative impact on query performance. In this paper, we present a novel abstraction for efficient queries of tag information in time series databases: a hybrid tagging index that manages all tags in one place. By managing tag information globally in a single disk-based data structure, we can fundamentally relieve memory pressure and eliminate I/O overhead of duplicate metadata from existing methods. Furthermore, the tagging index is internally partitioned by time to support time range based queries and data retention which are essential to time series databases. We implement the proposed tagging index as a standalone module which can be integrated with time series storage engines. Experiments on the TSBS benchmark show our proposed method can significantly speed up queries by on average 84.0% and 87.2% compared to Prometheus (using a time-partitioned segment method) and Graphite (using an external database for tag management), respectively. Jin Xue, Tianyu Wang 0009, Zili Shao |
IPDPS | 3 |
| 2022 | Co-mining: a processing-in-memory assisted framework for memory-intensive PoW accelerationabstractRecently, HBM (High Bandwidth Memory) and PIM (Processing in Memory) integrated technology such as Samsung function-in-memory DRAM opens a new door for memory-intensive PoW acceleration by jointly exploiting GPU, PIM and HBM. In this paper, we for the first time propose a GPU/PIM Co-Mining framework to accelerate memory intensive PoW by fully exploiting HBM-PIM's bandwidth and coordinately scheduling mining tasks in both GPU and PIM. Specifically, we first design a linear programming model to intelligently guide the GPU/PIM task scheduling. An extended finite-state-machine model is designed for the GPU memory controller to switch PIM working mode (compute/memory mode) accordingly. Finally, considering the speed difference between intra-/inter-channel memory accesses, a hybrid memory access method is proposed to minimize inter-channel data movements. We evaluate Co-Mining based on Samsung's HBM2-based function-in-memory architecture. The experimental results show that it can achieve up to 38.5% hashrate improvement compared with the method by directly integrating PIM into PoW acceleration with GPU. Tianyu Wang 0009, Zhaoyan Shen, Zili Shao |
LCTES | 1 |
| 2022 | Understanding Characteristics and System Implications of DAG-Based Blockchain in IoT EnvironmentsabstractBlockchain is starting to be deployed in the Internet of Things (IoT) to enable autonomous device-to-device transactions. However, traditional block-based blockchain techniques, such as Bitcoin and Ethereum, are not suitable for IoT environments due to their low throughput, high computation overhead, and costly transaction fee. To satisfy the requirements of IoT environments, directed-acyclic-graph (DAG)-based approaches, aiming to provide cheap blockchain services with low latency and high throughput, are emerging. This article presents a set of comprehensive experimental studies on IOTA, a representative DAG-based blockchain. We aim to exhibit its unique characteristics mainly from three aspects: 1) performance; 2) security; and 3) system robustness. We have developed a series of benchmark tools and judiciously selected typical configurations to perform experimental examinations with a real private IOTA network. Our studies reveal several interesting findings: 1) the throughput of IOTA is higher than the traditional block-based blockchain but far less than the reported thousands of transactions per second (TPS) in its whitepaper, even with scaling-up configurations; 2) the database query heavily impacts the performance of IOTA, even more than its mining [i.e., Proof of Work (PoW)] process; and 3) the system robustness of IOTA is closely related to the frequency of the incoming transactions while the milestone sent by the centralized coordinator has little effect on the system robustness. We make our benchmark tools public and expect our works can inspire system architects, application designers, and practitioners with new optimization directions and potential application cases for further exploration. Tianyu Wang 0009, Qian Wang 0042, Zhaoyan Shen, Zhiping Jia, Zili Shao |
IEEE Internet Things J. | 1 |
| 2021 | LiteIndex: Memory-Efficient Schema-Agnostic Indexing for JSON documents in SQLiteabstractSQLite with JSON (JavaScript Object Notation) format is widely adopted for local data storage in mobile applications such as Twitter and Instagram. With more data are generated and stored, it becomes vitally important to efficiently index and search JSON records in SQLite. However, current methods in SQLite either require full text search (that incurs big memory usage and long query latency) or indexing based on expression (that needs to be manually created by specifying search keys). On the other hand, existing JSON automatic indexing techniques, mainly focusing on big data and cloud environments, depend on a colossal tree structure that cannot be applied in memory-constrained mobile devices. Siqi Shang, Qihong Wu, Tianyu Wang 0009, Zili Shao |
ASP-DAC | 3 |
| 2020 | ABACUS: Address-partitioned Bloom filter on Address Checking for UniquenesS in IoT BlockchainabstractDAG-based blockchain systems have been deployed to enable trustworthy peer-to-peer transactions for IoT devices. Unique address checking, as a key part of transaction generation for privacy and security protection in DAG-based blockchain systems, incurs big latency overhead and degrades system throughput. Tianyu Wang 0009, Qun Ma, Zhaoyan Shen, Zili Shao |
ICCAD | 1 |
| 2020 | A Highly Parallelized PIM-Based Accelerator for Transaction-Based Blockchain in IoT EnvironmentabstractBlockchain has gained a lot of attention from both academia and industry. However, traditional standard blockchains, such as Bitcoin and Ethereum, suffer from low throughput, high computation overhead, and large transaction fee, which is not suitable for Internet of Things (IoT) transactions. Recently, transaction-based approaches, such as Tangle structure which is based on a directed acyclic graph (DAG), have emerged to solve blockchain scalability issues for IoT environment. In transaction-based blockchain, for a transaction, namely, a node, to be attached to the Tangle, it needs to verify two other transactions. However, with the Tangle expanding, this attaching process consumes huge computational resources and energy, which severely limits the performance of the transaction-based blockchain. In this article, we present Re-Tangle, a highly parallelized processing-in-memory (PIM)-based accelerator for transaction-based blockchain. Re-Tangle is composed of a random walking module, a transaction validation module, and a PoW module, to improve the Tangle system performance. These modules transfer Tangle functions, such as fast exponentiation and modular, into ReRAM-based logic analog computation units. In the random walking module, Re-Tangle maintains an exponentiation table to reduce its design complexity and improve its computation efficiency. In the transaction validation module, Re-Tangle further proposes a highly parallel modular unit to accelerate the validation of different tags in a transaction. In the PoW module, we decompose the Curl hash function into basic logic OR, AND, SHIFT, and XOR operations, and map these logic operations to ReRAM crossbars in parallel to accelerate the working process. The experimental results show that Re-Tangle distinguishes itself from other architectures with significant performance improvement and energy saving. The throughput of Re-Tangle is about 22.4× and 2.38× higher compared with CPU and GPU, respectively, and the energy consumption of Re-Tangle is 83.5× and 5.77× less for equal workload. Qian Wang 0042, Zhiping Jia, Tianyu Wang 0009, Zhaoyan Shen, Mengying Zhao, Renhai Chen, Zili Shao |
IEEE Internet Things J. | 3 |
| 2019 | Re-Tangle: A ReRAM-based Processing-in-Memory Architecture for Transaction-based BlockchainabstractBlockchain has gained a lot of attentions from both academic and industry. Transaction-based approaches such like Tangle structure, which is based on a DAG (Directed Acyclic Graph), are emerging to solve the blockchain scalability issues for IoT environment. In transaction-based blockchain, for a transaction, namely a node, to be attached to the Tangle, it needs to verify two other transactions. However, with the Tangle expanding, this attaching process consumes huge computational resources and energy, which severely limits the performance of the transaction-based blockchain. In this paper, we present Re-Tangle, a novel transaction-based blockchain acceleration architecture that explores the opportunity of performing massive parallel operations with low hardware and energy cost. Re-Tangle consists of a random walking module and a transaction validation module, which transfer Tangle functions into ReRAM-based logic analog computation units. In the random walking module, Re-Tangle maintains a exponentiation translator to reduce its design complexity and improve its computation efficiency for exponentiation. In the transaction validation module, Re-Tangle further proposes a highly parallel modular unit to accelerate the validation of different tags in a transaction. The experience results show that Re-Tangle distinguishes itself from other architectures, with significant performance improvement and energy saving. The throughput of Re-Tangle is about 19.4× and 2.13× higher compared with CPU and GPU, respectively, and the energy consumption of Re-Tangle is 63.35 × and 4.92 × less. Qian Wang 0042, Tianyu Wang 0009, Zhaoyan Shen, Zhiping Jia, Mengying Zhao, Zili Shao |
ICCAD | 2 |
| 2018 | H2-RAID: A Novel Hybrid RAID Architecture Towards High Reliability
Tianyu Wang 0009, Zhiyong Zhang 0006, Mengying Zhao, Zhiping Jia, Jianping Yang, Yang Wu 0003 |
ICA3PP (4) | 1 |