VLDB 2026 Research / reviewers in the wild / expert
Ao Ren
dblp:185/5754
· DBLP profile ↗
67ranked-venue papers
5as first author
48since 2021 · last 2026
0000-0002-2322-8038ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 51 · 4 first-author · 35 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D2 Prune: Sparsifying Large Language Models via Dual Taylor Expansion and Attention Distribution AwarenessabstractLarge language models (LLMs) face significant deployment challenges due to their massive computational demands. While pruning offers a promising compression solution, existing methods suffer from two critical limitations: (1) They neglect activation distribution shifts between calibration data and test data, resulting in inaccurate error estimations; (2) Overlooking the long-tail distribution characteristics of activations in the attention module. To address these limitations, this paper proposes a novel pruning method, D²Prune. First, we propose a dual Taylor expansion-based method that jointly models weight and activation perturbations for precise error estimation, leading to precise pruning mask selection and weight updating and facilitating error minimization during pruning. Second, we propose an attention-aware dynamic update strategy that preserves the long-tail attention pattern by jointly minimizing the KL divergence of attention distributions and the reconstruction error. Extensive experiments show that D²Prune consistently outperforms SOTA methods across various LLMs (e.g., OPT-125M, LLaMA2/3, Qwen3). Moreover, the dynamic attention update mechanism also generalizes well to ViT-based vision models like DeiT, achieving superior accuracy on ImageNet-1K. Lang Xiong, Ning Liu 0007, Ao Ren, Yuheng Bai, Haining Fang, Binyan Zhang, Yujuan Tan, Duo Liu 0002 |
AAAI | 3 |
| 2026 | HitKV: Activation Frequency Knows Which Tokens Are ImportantabstractThe demand for long-context processing in large language models (LLMs) continues to escalate alongside rapid advancements in their capabilities. However, the intermediate attention keys and values (KV cache) employed to avoid re-computations, also grow linearly with sequence length, far exceeding the memory capacity of consumer-grade GPUs. Consequently, many studies have proposed KV cache compression methods that evict unimportant tokens based on variant attention scoring strategies. These methods typically retain the KV pairs of the top-k scoring tokens under a fixed memory budget. However, they still face several limitations. First, they disregard the activation frequency of tokens, specifically the count of times tokens achieve top-k scores in the attention distribution of following tokens. The methods based on variant attention scores may incorrectly evict some high-activation-frequency yet low final-scoring tokens. Second, the activation frequency exhibits different distribution patterns across layers and tasks. Neglecting these differences negatively impacts model performance and task adaptability. Our analysis of the actual token activation frequency and its unique characteristics across layers and task types reveals potential opportunities to address these issues. In this paper, we propose HitKV, which employs hit rates to directly characterize token activation frequencies, enabling adaptive layer-aware and task-aware KV cache eviction under the uniform memory allocation strategies. Also, HitKV can be easily integrated into layer-specific memory allocation methods. Experimental results demonstrate that HitKV maintains model performance with preserving only 3% of the KV cache, achieves high-quality generation outputs in long-text generation tasks, and delivers 4× throughput improvement over baselines. Sanle Zhao, Yujuan Tan, Jing Yu 0026, Zhuoxin Bai, Zongjie Wang, Ao Ren |
AAAI | 8 |
| 2026 | RefineDedup: efficient deduplication for mobile systems via application-wise learning
Wei Li 0322, Xianzhang Chen, Xingjie Zhou, Duo Liu 0002, Yujuan Tan, Ao Ren, Kan Zhong, Lei Qiao 0002 |
Sci. China Inf. Sci. | 6 |
| 2026 | LiPRA: Lightweight pruning rate allocation for LLMs via global sensitivity measurement
Haining Fang, Ning Liu 0007, Lang Xiong, Zhenyu Wang 0002, Xianzhang Chen, Ao Ren, Yujuan Tan |
Neurocomputing | 7 |
| 2026 | SAHChain: A Hybrid Storage Blockchain System Supporting Semantic Expressiveness and Retrieval
Chaoxia Qin, Duo Liu 0002, Bing Guo 0003, Yujuan Tan, Ao Ren, Kan Zhong, Liang Liang 0002 |
IEEE Trans. Computers | 5 |
| 2026 | Latency Optimization in Hybrid Memory System for GNNsabstractGraph Neural Networks (GNNs) require high-capacity, low-latency memory systems to process large graphs. A hierarchical hybrid memory architecture combining high-capacity Non-Volatile Memory (NVM) and low-latency DRAM offers a promising solution. However, the inherent sparsity of graph data results in poor locality for GNN memory requests, leading to low DRAM cache hit rates and numerous misses, which significantly impairs the hybrid memory system’s performance. A critical issue is that DRAM misses in serial access mode incur substantial latency. While parallel access mode can mitigate this for misses, it introduces long-tail latency and wastes bandwidth for DRAM hits. In this paper, we focus on addressing these issues from two aspects: increasing the cache hit rate and decreasing the miss latency. We mainly propose two predictors: a future data access predictor that enables accurate prefetching to DRAM, thereby improving cache hit rates, and a data location predictor that determines whether data resides in DRAM or NVM, optimizing the choice between serial and parallel access modes to reduce miss latency. By integrating these predictors, we achieve efficient data access in both DRAM and NVM. Our experiments show a 49.5% reduction in memory delay and a 38.1% increase in memory bandwidth utilization compared to baseline. Zhaoyang Zeng, Yujuan Tan, Wei Chen 0101, Zhuoxin Bai, Ao Ren, Duo Liu 0002, Xianzhang Chen |
IEEE Trans. Computers | 6 |
| 2025 | RAN: Accelerating Data Repair with Available Nodes in Erasure-Coded StorageabstractDistributed storage systems ensure data availability through fault-tolerant mechanisms, with erasure coding widely adopted for its low storage overhead. However, erasure coding generates significant repair traffic during data recovery, severely degrading performance. Recent repair algorithms aim to alleviate network bottlenecks at congested nodes, but they primarily address downlink bottlenecks while neglecting uplink constraints, which fundamentally limit repair efficiency. Furthermore, these algorithms lack a systematic approach for handling diverse failure scenarios, complicating recover implementation. In this paper, we propose RAN, an aggregation-based repair algorithm that alleviates both uplink and downlink bottlenecks by optimizing bandwidth utilization across all available nodes and aggregating network transfers via programmable network devices. Additionally, RAN systematically maximizes repair performance across diverse failure scenarios through a unified procedure. Experiments on Amazon EC2 show that RAN improves repair throughput by up to$\mathbf{6 8. 9 \%}$for degraded read and$\mathbf{2 6 6. 6 \%}$for full-node recovery compared to state-of-the-art algorithms. Canghai Yang, Kan Zhong, Yujuan Tan, Ao Ren, Duo Liu 0002 |
CLUSTER | 4 |
| 2025 | LIO-DPC: Accurate and Fast LiDAR-Inertial Odometry with Dynamic Pose ChainabstractLiDAR-inertial odometry is widely used in robotics navigation, autonomous driving, and drone operation to provide precise, low-latency motion estimation. Filter-based methods are fast but suffer from significant cumulative errors. Graph optimization methods reduce cumulative errors through loop closure detection but are computationally expensive. In this work, we propose LIO-DPC, a framework that combines the benefits of the filter-based approach and graph-based approach. First, we propose a dynamic pose chain optimization method. It generates an initial pose chain using the fast filter. This is followed by applying computationally efficient local graph optimization to a set of local pose chains to generate refined relative poses, which are then used to update the motion estimation. Second, we propose a loop sparsification approach to select representative loops that are both temporally and spatially proximate, to reduce the computational complexity in graph optimization and minimize loop errors. Extensive experiments demonstrate that LIO-DPC achieves real-time performance and outperforms state-of-the-art methods in accuracy. Yuexin Mu, Ao Ren, Duo Liu 0002, Zihao Zhang 0002, Haojie Lu, Longyi Zhou, Huachen Tan, Kan Zhong, Yujuan Tan, Chaoxia Qin |
DAC | 2 |
| 2025 | CoSF: A Co-Optimization Framework for Operator Splitting and Fusion
Wei Li 0322, Ao Ren, Qingqiu Lan, Haining Fang, Zhenyu Wang 0002, Yujuan Tan, Kan Zhong, Duo Liu 0002 |
Euro-Par (1) | 2 |
| 2025 | Cocache: An Accurate and Low-Overhead Dynamic Caching Method for GNNs
Zhaoyang Zeng, Yujuan Tan, Zhuoxin Bai, Kan Zhong, Duo Liu 0002, Ao Ren |
Euro-Par (2) | 8 |
| 2025 | MPNAS: Multimodal Sentiment Analysis Pruning via Neural Architecture SearchabstractWith the rapid development of social media, sentiment analysis from multimodal posts has garnered significant attention in recent years. However, the substantial size of these models impedes their deployment on resource-constrained embedded devices. Although pruning has been extensively studied to reduce the size of unimodal models, specific challenges remain for Multimodal Sentiment Analysis (MSA) models. First, existing techniques prune fixed original models into sparse models, while our findings indicate that different model architectures of identical size yield varying performance outcomes. Second, prior studies fail to explore the unique characteristics of MSA models, resulting in suboptimal pruning performance. To address these challenges, we propose MPNAS, a unified pruning framework via Neural Architecture Search (NAS) for MSA models. Specifically, we formulate pruning as a NAS problem and analyze MSA model characteristics to guide the subnet search. We conduct an initial coarse-grained NAS on the original model, expanding the search space slightly to identify suitable subnets that enhance pruning rates and accuracy. Subsequently, we refine coarse-grained subnets in a fine-grained NAS stage, where MSA model characteristics guide the search process. Extensive experiments on three representative datasets demonstrate the superiority of our approach over existing methods. Binyan Zhang, Ao Ren, Zihao Zhang 0002, Moming Duan, Duo Liu 0002, Yujuan Tan, Kan Zhong |
ICASSP | 2 |
| 2025 | CAST: An Efficient Framework for Schedules Performance Prediction Based on Compact ASTsabstractWith the advances of deep learning, efficient model inference is crucial. Deep learning compilers optimize inference by decomposing models into subgraphs and searching schedules for them, whose evaluation relies on accurate cost models. Existing methods suffer from high transformation overheads or limited prediction accuracy caused by insufficient structural representation of subgraphs and schedules. To address these limitations, we propose CAST, a framework that predicts schedule performance based on Abstract Syntax Trees (ASTs). CAST proposes AST classification based on structural similarity and class-specific cost models. Experiments show CAST achieves significantly reduced prediction errors and up to$13 \times$higher efficiency than prior methods. Qingqiu Lan, Ao Ren, Zhenyu Wang 0002, Wei Li 0322, Hongbin Zhu, Yujuan Tan, Duo Liu 0002, Kan Zhong, Chaoxia Qin |
ICCD | 2 |
| 2025 | DualSpar: A Dual-Granularity Memory Framework with Adaptive Sparsity for Efficient LLM InferenceabstractThe block-based inference engine, powered by noncontiguous key-value (KV) cache management, has emerged as a new paradigm for large language model (LLM) inference due to its efficient memory utilization. However, in large-batch, longcontext workloads, the substantial demand for KV cache remains a major bottleneck in inference performance. Existing research leverages the sparsity of the attention mechanisms by removing non-critical tokens to limit the KV cache size. However, we observe that in block-based inference engines, current sparse methods require defragmentation after token removal to maintain tensor continuity, incurring significant overhead. Additionally, we find a correlation between request length and sparsity potential, yet existing methods apply a uniform sparsity strategy at batch level without dynamic adjustment. To address these, we propose DualSpar, a novel sparse KV cache framework with dual-granularity memory and adaptive sparsity strategy. First, it binds token importance to KV cache block granularity, achieving low-overhead pre-consolidation. Second, it incorporates system load and request length into sparsity decision, fully exploiting the sparse potential of different requests while reducing the queuing latency in large-batch processing. Evaluations show that DualSpar achieves up to a 3.16× throughput improvement, a 3.75× faster time-to-first-token (TTFT), and an 87.2% reduction in defragmentation overhead while maintaining high accuracy. Yujuan Tan, Zhuoxin Bai, Sanle Zhao, Yujiao Wang, Zongjie Wang, Ao Ren, Kan Zhong |
ICCD | 7 |
| 2025 | Co-GNN: A Co-optimization Framework for Memory and Computation in Sampling-Based GNN Training
Yan Gan, Yujuan Tan, Yujiao Wang, Zongjie Wang, Duo Liu 0002, Ao Ren, Kan Zhong, Chaoxia Qin, Mingrui Qiang |
ICIC (21) | 7 |
| 2025 | RobTrack: A Robust 3D Multi-object Tracking Method for Edge Devices
Mingrui Qiang, Ao Ren, Yujuan Tan, Jing Yu 0026, Zhuoxin Bai, Duo Liu 0002, Kan Zhong, Chaoxia Qin |
ICIC (5) | 2 |
| 2025 | FASP: A Fast and Accurate Framework for Schedule Performance EvaluationabstractWith the widespread application of deep neural networks, improving inference efficiency has become increasingly critical. To speed up the inference, deep learning compilers search for high-performance schedules for the DNN tensor programs. During the process, cost models have been extensively studied to evaluate the performance of the schedules, such that high-performance ones can be efficiently obtained. However, existing methods suffer from either high overhead or low accuracy of performance evaluation, both of which limit the efficiency of the final schedule. To address these issues, we propose FASP, a fast and accurate framework for schedule performance evaluation, based on Abstract Syntax Trees (ASTs). First, we propose a redundancy-aware ASTs reduction method to generate compact ASTs for more accurate feature extraction. Second, we propose a feature extraction method based on compact ASTs, which extracts features by accounting for computation nodes, loop nodes, and their structural relationships. Third, we propose a composition-similarity-driven ASTs classification method and a class-specific cost model architecture for more accurate performance evaluation. FASP overcomes the limitations of prior methods by significantly reducing evaluation errors. Experiments show its excellent performance in both single-model and cross-model evaluation, with errors ranging from 6 % to$\mathbf{1 3 \%}$. Moreover, FASP can obtain high-performance schedules with$13 \times$lower latency. Qingqiu Lan, Ao Ren, Zhenyu Wang 0002, Wei Li 0322, Hongbin Zhu, Yujuan Tan, Duo Liu 0002, Kan Zhong, Chaoxia Qin |
ICPADS | 2 |
| 2025 | PIM-IoT: Enabling hierarchical, heterogeneous, and agile Processing-in-Memory in IoT systems
Kan Zhong, Qiao Li 0001, Ao Ren, Yujuan Tan, Xianzhang Chen, Linbo Long, Duo Liu 0002 |
Future Gener. Comput. Syst. | 3 |
| 2025 | GNNBoost: Accelerating sampling-based GNN training on large scale graph by optimizing data preparation
Yujuan Tan, Yan Gan, Zhaoyang Zeng, Zhuoxin Bai, Lei Qiao 0002, Duo Liu 0002, Kan Zhong, Ao Ren |
J. Syst. Archit. | 8 |
| 2025 | LAShards: Low-Overhead and Self-Adaptive MRC Construction for Non-Stack AlgorithmsabstractShared cache systems have become increasingly crucial, especially in cloud services, where the Miss Ratio Curve (MRC) is a widely used tool for evaluating cache performance. The MRC depicts the relationship between the cache miss ratio and cache size, indicating how cache performance trends with varying cache sizes. Recent advancements have enabled efficient MRC construction for stack replacement policies. For non-stack policies, miniature simulation downsizes the actual cache size and data stream through spatially hashed sampling, providing a general method for MRC construction. However, this approach still faces significant challenges. Firstly, constructing an MRC requires numerous mini-caches to obtain miss ratios, consuming significant cache resources, leading to tremendous memory and computing overhead. Secondly, it cannot adapt to the dynamic I/O workloads, resulting in less precise MRC.To address these issues, we propose LAShards, a low-overhead and self-adaptive MRC construction method for non-stack replacement policies. The key idea behind LAShards is to exploit the locality and burstiness in access patterns. It can statically reduce memory usage and dynamically adapt to workloads. Compared to previous works, LAShards can save up to 20× of memory resources, and increase throughput by up to 10×. Sanle Zhao, Yujuan Tan, Zhaoyang Zeng, Jing Yu 0026, Zhuoxin Bai, Ao Ren, Xianzhang Chen, Duo Liu 0002 |
IEEE Trans. Computers | 6 |
| 2025 | DSAV: A Deep Sparse Acceleration Framework for Voxel-Based 3-D Object DetectionabstractVoxel-based 3-D object detection has been widely applied in robotics, virtual reality, and autonomous driving. However, inefficiency in the voxelization and backbone-network computation, which are the main components of the voxel-based models, prevents efficient 3-D object detection. First, due to the high sparsity and irregularity of the point cloud, the voxelization process usually requires generalized platforms, such as CPUs, and causes low voxelization speed. Second, the voxel-based models contain considerable transposed convolutional layers, and existing accelerators introduce considerable additional hardware to support both the convolution and transposed convolution operations. Nonetheless, this strategy incurs significant hardware costs. Besides, transposed convolutions result in various patterns of sparse feature maps, and pruning as a representative model compression technique, results in sparse weight matrices. The two types of sparsity impose challenges in accelerating the voxel-based models, including activation-weight matching efficiency, low partial-sum accumulation efficiency, and workload imbalance issues. In this work, we propose DSAV, a 3-D object detection accelerator to address these obstacles. Specifically, we first propose a hash-based voxelizer for efficient voxelization, by storing and indexing voxels hierarchically. Then, we collaboratively design the transposed convolution acceleration method, structured pruning method, and accelerator architecture for the voxel-based models. As a result, the accelerator can fully leverage the sparsity lies in both feature maps and weight matrices. Experimental results show that the proposed accelerator can outperform the prior studies by$19{\times } \sim 19.8{\times }$faster in voxelization and$4.29{\times } \sim 38.01\times $faster in backbone inference. Finally, the accelerator achieves$4.61{\times } \sim 31.63{\times }$speedups than its counterparts in 3-D object detection tasks. Haining Fang, Yujuan Tan, Ao Ren, ZhiYong Qin, Duo Liu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | CMCache: An Adaptive Cross-Level Data Placement Method for Multilevel CacheabstractMultilevel cache systems enhance I/O performance by optimizing data placement across various cache levels from a global perspective. However, existing methods often struggle to place data at the optimal cache level promptly due to their reliance on historical access patterns and inflexible placement strategies. These methods face two main challenges: 1) for already cached data with sufficient access history, existing approaches only optimize movement between adjacent cache levels, potentially delaying data arrival at its globally optimal cache level and leading to unnecessary bandwidth consumption and increased latency and 2) for newly entered data without access history, current methods cannot accurately predict their future hotness and simply place them at a fixed cache level (i.e., first or final level), overlooking future accesses of new data and potentially resulting in high cache miss rates or cache pollution. To address these issues, we propose CMCache, an adaptive cross-level data placement method for multilevel cache. CMCache applies distinct placement strategies for cached and new data to reach the optimal level timely, considering their different characteristics. It also logically divides cache space into two sections to manage cached and new data separately, dynamically adjusting section sizes based on access patterns. This approach significantly improves data placement efficiency, achieving up to an 89% reduction in miss rates and a 79% decrease in average response times compared to existing methods. Zhaoyang Zeng, Yujuan Tan, Zhulin Ma, Sanle Zhao, Duo Liu 0002, Xianzhang Chen, Ao Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2024 | VIFA: An Efficient Visible and Infrared Image Fusion Architecture for Multi-task Applications via Continual Learning
Jiaxing Shi, Ao Ren, ZhiYong Qin, Zhenyu Wang 0002, Yujuan Tan, Duo Liu 0002 |
ACCV (8) | 2 |
| 2024 | Rethinking Literary Plagiarism in LLMs through the Lens of Copyright Laws
Huachen Tan, Moming Duan, Duo Liu 0002, Haojie Lu, Yuexin Mu, Longyi Zhou, Ao Ren, Yujuan Tan, Kan Zhong |
ACML | 7 |
| 2024 | FinerDedup: Sifting Fingerprints for Efficient Data Deduplication on Mobile DevicesabstractData deduplication is promised to extend the lifetime and capacity of storage on mobile devices. However, existing data deduplication works show high memory consumption and indexing costs for maintaining a fingerprint for each data block, especially when the duplicate ratio of data blocks on mobile systems is about 10% to 30%. In this paper, we propose a novel approach called FinerDedup to optimize the memory costs and retrieval efficiency of data deduplication. FinerDedup drastically reduces the number of fingerprints by screening out the duplicate data blocks via random forest and Bloom filter. We implement FinerDedup on real mobile devices with Android 10 and evaluate it with real workloads. Extensive experimental results show that FinerDedup can reduce 85% of fingerprints and 20% of I/O latency over the widely-used DmDedup. Xianzhang Chen, Xingjie Zhou, Wei Li 0322, Duo Liu 0002, Yujuan Tan, Ao Ren |
DAC | 7 |
| 2024 | RACI: A Resource-Aware Cooperative Inference Framework on Heterogeneous Edge DevicesabstractCooperative inference for deep neural networks (DNNs) across edge devices has received increasing attention, due to the benefits of low latency, low power consumption, and privacy preservation. Cooperative inference partitions a DNN model into multiple segments, which will then be allocated to distributed devices for parallel inference. Nonetheless, prior works fail to comprehensively study the impact of layer configurations, dynamic network bandwidths, and heterogeneous device capabilities on the inference speed, resulting in suboptimal inference performance. In this work, we conduct a comprehensive analysis of these key factors and figure out the limitations of conventional transfer-based and redundant computation-based methods. Based on the analysis, we first propose a latency prediction agent that accounts for the layer configurations, network bandwidths, and device computing capabilities, aiming to quickly evaluate the inference latency. Furthermore, we propose RACI, a resource-aware cooperative DNNs inference framework on heterogeneous edge devices. It co-trains a model agent for model partition and a workload agent for workload allocation to generate co-optimized model partition and workload allocation strategies, leading to high cooperation inference acceleration. Experimental results demonstrate that RACI outperforms the state-of-the-art approaches by 1.1× -5.2× in terms of inference speedup for three representative DNN models. Zhenyu Wang 0002, Ao Ren, Duo Liu 0002, Haining Fang, Jiaxing Shi, Yujuan Tan, Xianzhang Chen |
ICCAD | 2 |
| 2024 | DPC: DPU-accelerated High-Performance File System ClientabstractTo achieve efficient file access to the file system backend, file system clients employ various intricate optimization techniques, such as local data/metadata caching and direct data access. However, these techniques impose a significant load on the host CPU, posing substantial challenges to the valuable CPU resources. Kan Zhong, Zhiwang Yu, Qiao Li 0001, Xianqiang Luo, Linbo Long, Yujuan Tan, Ao Ren, Duo Liu 0002 |
ICPP | 7 |
| 2024 | An FPGA-based kNN Seach Accelerator for point cloud registrationabstractPoint cloud registration assumes a crucial role in several fields, including 3D reconstruction and pose estimation. The prevalent technique for point cloud registration is Iterative Closest Point (ICP). Nonetheless, the k-nearest neighbor (kNN) search, an essential component of the ICP process, often falls short in meeting real-time requirements due to its substantial time consumption. Consequently, extensive efforts have been dedicated to accelerating the kNN search within the ICP framework. This study proposes an FPGA-based kNN Search Accelerator, leveraging an improved LSH approach to expedite point cloud access and search. Experimental results underscore its superiority with a 120x and 15x speed-up compared to CPU and GPU implementations of the kNN process, respectively. Remarkably, the kNN search completes in a mere 0.64 ms, surpassing the performance of prior works. Chengliang Wang 0002, Zhetong Huang, Ao Ren |
ISCAS | 3 |
| 2024 | CEIU: Consistent and Efficient Incremental Update mechanism for mobile systems on flash storage
Ruiqing Lei, Xianzhang Chen, Duo Liu 0002, Chunlin Song, Yujuan Tan, Ao Ren |
J. Syst. Archit. | 6 |
| 2024 | BGS: Accelerate GNN training on multiple GPUs
Yujuan Tan, Zhuoxin Bai, Duo Liu 0002, Zhaoyang Zeng, Yan Gan, Ao Ren, Xianzhang Chen, Kan Zhong |
J. Syst. Archit. | 6 |
| 2024 | Optimizing the Performance of Consistency-Aware Deduplication Using Persistent MemoryabstractBlock-level data deduplication is a widely-used technology for saving storage space by filtering the data blocks with the same hash value. However, existing block-level data deduplication approaches either ignore the data consistency of deduplication or suffer severe performance degradation for providing consistency guarantees. In this paper, we propose Consistency-Aware Deduplication (CADedup+) to achieve high-performance block-level data deduplication with data consistency. The main idea of CADedup+ is to achieve an efficient journaling mechanism for deduplication by taking advantage of persistent memory (PM), such as byte-addressability and near-DRAM access latency. To balance the trade-offs between performance and consistency requirements in data deduplication, we carefully design three modes of journaling mechanism, i.e., writeback mode, ordered mode, and journal mode, for CADedup+. We properly place the deduplication metadata of CADedup+ onto the DRAM-PM hybrid memory architecture to minimize PM costs according to the features of metadata updates. The deduplication metadata on PM is managed by a set of metadata transactions and updated with the help of the efficient hardware atomic operations provided by CPU. We implement CADedup+ in the generic block layer in Linux kernel 4.9.0. We conduct extensive experiments on Intel Optane PMEM to evaluate CADedup+ with typical benchmarks. Experimental results show that CADedup+ can reduce 63%-70% write volume and 50%-60% I/O latency over Dmdedup, a widely-used open-source block-level data deduplication system, while ensuring deduplication consistency. Chunlin Song, Xianzhang Chen, Duo Liu 0002, Yujuan Tan, Ao Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | FreePrune: An Automatic Pruning Framework Across Various Granularities Based on Training-Free EvaluationabstractNetwork pruning is an effective technique that reduces the computational costs of networks while maintaining accuracy. However, pruning requires expert knowledge and hyperparameter tuning, such as determining the pruning rate for each layer. Automatic pruning methods address this challenge by proposing an effective training-free metric to quickly evaluate the pruned network without fine-tuning. However, most existing automatic pruning methods only investigate a certain pruning granularity, and it remains unclear whether metrics benefit automatic pruning at different granularities. Neural architecture search also studies training-free metrics to accelerate network generation. Nevertheless, whether they apply to pruning needs further investigation. In this study, we first systematically analyze various advanced training-free metrics for various granularities in pruning, and then we investigate the correlation between the training-free metric score and the after-fine-tuned model accuracy. Based on the analysis, we proposed FreePrune score, a more general metric compatible with all pruning granularities. Aiming at generating high-quality pruned networks and unleashing the power of FreePrune score, we further propose FreePrune, an automatic framework that can rapidly generate and evaluate the candidate networks, leading to a final pruned network with both high accuracy and pruning rate. Experiments show that our method achieves high correlation on various pruning granularities and comprehensively improves the accuracy. Ning Liu 0007, Haining Fang, Qiu Lin, Yujuan Tan, Xianzhang Chen, Duo Liu 0002, Kan Zhong, Ao Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2023 | IFHE: Intermediate-Feature Heterogeneity Enhancement for Image Synthesis in Data-Free Knowledge DistillationabstractData-free knowledge distillation (DFKD) explores training a compact student network only by a pre-trained teacher without real data. Prevailing DFKD methods mainly consist of image synthesis and knowledge distillation. The synthesized images are crucial to enhance the student network performance. However, the images synthesized by existing methods cause high homogeneity on intermediate features, incurring undesired distillation performance. To address this problem, we propose the Intermediate-Feature Heterogeneity Enhancement (IFHE) method, which effectively enhances the heterogeneity of synthesized images by minimizing the loss between intermediate features and pre-set labels of the synthesized images Our IFHE outperforms the SOTA results on CIFAR-10/100 datasets of representative networks. Ning Liu 0007, Ao Ren, Duo Liu 0002 |
DAC | 3 |
| 2023 | Optimizing the Performance of NDP Operations by Retrieving File Semantics in StorageabstractIn-storage Near-Data Processing (NDP) architectures can reduce data movement between the host and the storage device by offloading computing tasks to the storage. This encourages many studies on building NDP applications, such as recommendation systems and databases, on computational SSDs. However, in the data path of existing NDP architectures, an NDP application has to find out the address of the requested file data by calling the I/O stacks of the kernel on the host, which incurs large overhead for transferring data between the host and the computational SSD. In this paper, we present File Semantics Retriever (FSR) to optimize the data path of NDP architectures by locating and fetching the requested file data directly in the computational SSD. The key idea is to recognize the file system layout and the metadata structures in the storage with the collaboration of a user-space library and a handler in the firmware of the computational SSD. We implement a prototype of FSR and evaluate it on the Cosmos plus OpenSSD, a widely-used computational SSD platform. The experimental results show that FSR outperforms existing NDP architectures in both benchmarks and real-world NDP applications. Xianzhang Chen, Jiapin Wang, Duo Liu 0002, Yujuan Tan, Ao Ren |
DAC | 7 |
| 2023 | HBP: Hierarchically Balanced Pruning and Accelerator Co-Design for Efficient DNN InferenceabstractWeight pruning is studied to accelerate DNN inference by reducing the parameters and computations. Irregular pruning achieves high sparsity while incurring low computation parallelism and imbalanced workloads. The coarse-grained structured pruning sacrifices sparsity for higher parallelism. To strike a better balance, we propose Hierarchically Balanced Pruning by applying fine-grained but structured adjustments based on irregular pruning. Besides, it partitions the weight matrix into hierarchical blocks and constrains the sparsity of the blocks for balanced workloads. Furthermore, an accelerator is proposed to unleash the power of the pruning method. Experimental results show our method achieves 1.1×-6 higher sparsity than prior studies, and the accelerator achieves 1.2×-13× speedup and 3.3× energy efficiency improvement than its counterparts. Ao Ren, Yuhao Wang 0002, Tao Zhang 0032, Jiaxing Shi, Duo Liu 0002, Xianzhang Chen, Yujuan Tan, Yuan Xie 0001 |
DAC | 1 |
| 2023 | An Efficient Scheduling Algorithm for Multi-mode Tasks on Near-Data Processing SSDs
Xianzhang Chen, Duo Liu 0002, Yujuan Tan, Ao Ren |
ICA3PP (7) | 6 |
| 2023 | Re-compact: Structured Pruning and SpMM Kernel Co-design for Accelerating DNNs on GPUsabstractPruning algorithms and sparse matrix-matrix multiplication (SpMM) kernels have been widely studied to accelerate DNN inference on GPUs. However, unstructured pruning spoils the regularity of data layout and incurs undesired speedup performance. Prior structured pruning methods remove weights coarsely and cause significant accuracy loss. In this work, we co-design Re-compact Pruning and SpMM kernel to achieve high acceleration while maintaining accuracy. Re-compact Pruning compacts both the sparse rows and columns with similar sparse patterns into dense blocks, on which the vector pruning is performed. An SpMM kernel is designed to fully leverage the regularity of the pruned matrix. It contains a Tile Padding operation to balance the workloads, a Tile Sorting operation to enable loop unrolling during compilation, and a Merge Write-back strategy to reduce the memory accesses. Experimental results show that our Re-compact Pruning and SpMM kernel outperform its counterparts by 4.53 ×, 4.44 ×, and 3.24 × for accelerating ResNet50, NMT, and Transformer, respectively. Ao Ren, Xianzhang Chen, Qiu Lin, Yujuan Tan, Duo Liu 0002 |
ICCD | 2 |
| 2023 | Data-Quality-Driven Federated Learning for Optimizing Communication CostsabstractFederated Learning (FL) is a distributed machine learning approach that allows mobile devices to train a global model cooperatively, without uploading privacy-sensitive data to the cloud. To improve the accuracy of the model, the model needs to be updated frequently. However, FL system under mobile edge-end has to adapt to limited communication bandwidth. At the same time, the property of statistical heterogeneity in FL means that we cannot blindly reduce the number of clients. We found that the accuracy of the global model depends greatly on the clients whose data is more similar and balanced. In this paper, we first define the "data quality" of clients to appraise the impact of data on a client to the accuracy of the global model. Then, based on the data quality, we design a client selection to optimize the communication costs of FL by screening out the clients that determine the accuracy of the global model. To the authors’ best knowledge, this is the first paper to save the costs of FL by assessing data quality of clients. Experimental results show that on imbalanced SVHN, the communication cost of our algorithm is reduced by 56% compared with vanilla FL. Compared with vanilla FL, requires all clients to participate in training, our algorithm shows -1.58% and +3.19% and -0.01% of average accuracy on the imbalanced CIFAR10, imbalanced FMNIST and imbalanced SVHN datasets, respectively. In other words, our algorithm can reduce communication overhead with negligible degradation of accuracy. Xuehong Fan, Nanzhong Wu, Shukan Liu, Xianzhang Chen, Duo Liu 0002, Yujuan Tan, Ao Ren |
ICPADS | 7 |
| 2023 | RadarSSD: A Computational Storage for Radar Signal ProcessingabstractRadar signals contain a multitude of small data items with multidimensional characteristics and various types of errors. It is challenging to store and recognize radar signals in real-time. Traditional computer architectures require data to be moved from storage to the host for processing, resulting in a "storage wall" problem. This problem is caused by low storage bandwidth, long I/O stacks, and excessive data transfers, which significantly reduce the efficiency of radar signal recognition. In this paper, we propose RadarSSD address these challenges by utilizing near-data processing (NDP) architecture to recognize radar signals within the solid-state drive (SSD), through which the high overhead of data movements can be avoided. To support efficient data I/O operations, we design a stripe-like data layout for storing radar signals taking advantage of their time sequential feature. We present a task slicing mechanism to reduce I/O blocking from in-storage data processing, and a dedicated interface for providing highly-efficient direct SSD access. We implement RadarSSD in a real computational SSD platform. Extensive experimental results show that RadarSSD can reduce power consumption while improving I/O and recognize performance, with a maximum improvement of 12.4 ×, 11.5 ×, and 4.1 × recognize speed compared to the systems that manage radar signals using MySQL, MongoDB, and Ext4. Xianzhang Chen, Duo Liu 0002, Ao Ren, Zhaoyang Zeng, Yujuan Tan |
ICPP | 4 |
| 2023 | FedMDS: An Efficient Model Discrepancy-Aware Semi-Asynchronous Clustered Federated Learning FrameworkabstractFederated learning (FL) is an emerging distributed machine learning paradigm that protects privacy and tackles the problem of isolated data islands. At present, there are two main communication strategies of FL: synchronous FL and asynchronous FL. The advantages of synchronous FL are the high precision and easy convergence of the model. However, this synchronous communication strategy has the risk of the straggler effect. Asynchronous FL has a natural advantage in mitigating the straggler effect, but there are threats of model quality degradation and server crash. In this paper, we propose a model discrepancy-aware semi-asynchronous clustered FL framework,FedMDS, which alleviates the straggler effect by 1) a clustered strategy based on the delay and direction of the model update and 2) a synchronous trigger mechanism that limits the model staleness.FedMDSleverages the clustered algorithm to reschedule the clients. Each group of clients performs asynchronous updates until the synchronous update mechanism based on the model discrepancy is triggered. We evaluateFedMDSbased on four typical federated datasets in a non-IID setting and compareFedMDSto the baselines. The experimental results show thatFedMDSsignificantly improves average test accuracy by more than$+9.2\%$on the four datasets compared toTA-FedAvg. In particular,FedMDSimproves absolute Top-1 test accuracy by$+37.6\%$on FEMNIST compared toTA-FedAvg. The frequency of the average synchronization waiting time ofFedMDSis significantly lower than that ofTA-FedAvgon all datasets. Moreover,FedMDScan improve the accuracy and alleviate the straggler effect. Yu Zhang 0184, Duo Liu 0002, Moming Duan, Xianzhang Chen, Ao Ren, Yujuan Tan, Chengliang Wang 0002 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | VEA: An FPGA-Based Voxel Encoding Accelerator for 3D Object Detection with LiDARabstractVoxel-based 3D object detection methods have been applied in various applications such as autonomous driving, robot navigation, and Augmented Reality. However, the sparse and unstructured characteristics of the point cloud and voxels prevent high-performance voxel encoding and usually require generalized platforms, such as CPUs. In this paper, an FPGAbased Voxel Encoding Accelerator (VEA) is proposed, which contains a generalized voxel generator and a feature extender. The generalized voxel generator decouples the point storage and voxel information storage, leading to high-speed voxelization and low memory consumption. The feature extender can efficiently extract the geometric information of the voxels and extend the features of the points. Based on the proposed VEA, an FPGA-based 3D object detection accelerator is implemented, and experimental results show that the proposed VEA can outperform prior studies by 19× faster in voxelization and 1.3×~ 9.6× faster in object detection. Ao Ren, Yujuan Tan, Zhetong Huang, Chengliang Wang 0002, Xianzhang Chen, Duo Liu 0002 |
ICCD | 2 |
| 2022 | CADedup: High-performance Consistency-aware Deduplication Based on Persistent MemoryabstractBlock-level data deduplication is prevalent in various-scaled storage systems for saving storage space and improving I/O performance by reducing write operations. However, data deduplication induces additional metadata of blocks, leading to I/O amplification. Furthermore, to ensure the correctness of deduplicated user data, data deduplication systems need to guarantee crash consistency. In this paper, we propose CADedup, to achieve high performance while ensuring crash consistency by using persistent memory. By taking advantage of the byte-addressability and near-DRAM latency of persistent memory, we design an efficient journaling mechanism to manage the deduplication metadata of CADedup. Additionally, we adopt a hybrid storage architecture of DRAM and persistent memory to minimize space costs. We implement CADedup through the device-mapper interface in the Linux kernel. We conduct extensive experiments on Intel Optane PMEM to evaluate CAD-edup with widely-used benchmarks. Experimental results show that compared with the no-deduplication system, CADedup can achieve up to 1×-3× improvement in many workloads of server storage and has negligible throughput drop in the worst case. Chunlin Song, Xianzhang Chen, Duo Liu 0002, Xiaoliu Feng, Yujuan Tan, Ao Ren |
ICCD | 8 |
| 2022 | Measuring Data Reconstruction Defenses in Collaborative Inference SystemsabstractThe collaborative inference systems are designed to speed up the prediction processes in edge-cloud scenarios, where the local devices and the cloud system work together to run a complex deep-learning model. However, those edge-cloud collaborative inference systems are vulnerable to emerging reconstruction attacks, where malicious cloud service providers are able to recover the edge-side users’ private data. To defend against such attacks, several defense countermeasures have been recently introduced. Unfortunately, little is known about the robustness of those defense countermeasures. In this paper, we take the first step towards measuring the robustness of those state-of-the-art defenses with respect to reconstruction attacks. Specifically, we show that the latent privacy features are still retained in the obfuscated representations. Motivated by such an observation, we design a technology called Sensitive Feature Distillation (SFD) to restore sensitive information from the protected feature representations. Our experiments show that SFD can break through defense mechanisms in model partitioning scenarios, demonstrating the inadequacy of existing defense mechanisms as a privacy-preserving technique against reconstruction attacks. We hope our findings inspire further work in improving the robustness of defense mechanisms against reconstruction attacks for collaborative inference systems. Mengda Yang, Juan Wang 0006, Hongxin Hu, Ao Ren, Xiaoyang Xu 0001, Wenzhe Yi |
NeurIPS | 5 |
| 2022 | Federated learning with workload-aware client scheduling in heterogeneous systems
Duo Liu 0002, Moming Duan, Yu Zhang 0184, Ao Ren, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002 |
Neural Networks | 5 |
| 2022 | FRL: Fast and Reconfigurable Accelerator for Distributed Sound Source LocalizationabstractSound source localization (SSL) has been widely applied in industrial and civil fields. And with the development of wearable devices and the Internet of Things (IoT), it is attractive to deploy the SSL system onto embedded and portable devices. However, the software-based SSL system causes excessive response delay and is often affected by environmental noise. To overcome this obstacle, we propose the fast and precise localization (FPL) algorithm for distributed SSL systems. It combines the benefits of both time difference of arrival (TDOA) and steered response power (SRP) methods, and thus it is able to localize sound sources fast and precisely. To further improve the localization speed, we propose the fast and reconfigurable localization (FRL) accelerator, which is an algorithm-hardware co-designed SSL accelerator. It adopts multiple distributed localization nodes for higher localization precision and higher robustness to environmental interference, and it can be configured into either the fast or precise mode to adapt to various environments. Experimental evaluations show that our proposed FPL algorithm can achieve high localization speed and precision, and the field-programmable gate array (FPGA)-based FRL accelerator outperforms the software implementation by$48.6\times $and outperforms the prior FPGA-based SSL accelerators by$20\times \sim 838.2\times $. Chengliang Wang 0002, Heping Liu, Zhihai Zhang, Xianzhang Chen, Yujuan Tan, Duo Liu 0002, Ao Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2022 | SENTunnel: Fast Path for Sensor Data Access on Automotive Embedded Systems
Rongwei Zheng, Xianzhang Chen, Duo Liu 0002, Jiapin Wang, Ao Ren, Chengliang Wang 0002, Yujuan Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Flexible Clustered Federated Learning for Client-Level Data Distribution ShiftabstractFederated Learning (FL) enables the multiple participating devices to collaboratively contribute to a global neural network model while keeping the training data locally. Unlike the centralized training setting, the non-IID, imbalanced (statistical heterogeneity) and distribution shifted training data of FL is distributed in the federated network, which will increase the divergences between the local models and the global model, further degrading performance. In this paper, we propose a flexible clustered federated learning (CFL) framework named FlexCFL, in which we 1) group the training of clients based on the similarities between the clients’ optimization directions for lower training divergence; 2) implement an efficient newcomer device cold start mechanism for framework scalability and practicality; 3) flexibly migrate clients to meet the challenge of client-level data distribution shift. FlexCFL can achieve improvements by dividing joint optimization into groups of sub-optimization and can strike a balance between accuracy and communication efficiency in the distribution shift environment. The convergence and complexity are analyzed to demonstrate the efficiency of FlexCFL. We also evaluate FlexCFL on several open datasets and made comparisons with related CFL frameworks. The results show that FlexCFL can significantly improve absolute test accuracy by$+10.6\%$on FEMNIST compared withFedAvg,$+3.5\%$on FashionMNIST compared withFedProx,$+8.4\%$on MNIST compared withFeSEM,$+4.7\%$on Sentiment140 compare withIFCA. The experiment results show that FlexCFL is also communication efficient in the distribution shift environment. Moming Duan, Duo Liu 0002, Xinyuan Ji, Yu Wu 0016, Liang Liang 0002, Xianzhang Chen, Yujuan Tan, Ao Ren |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2021 | FedSAE: A Novel Self-Adaptive Federated Learning Framework in Heterogeneous SystemsabstractFederated Learning (FL) is a novel distributed machine learning which allows thousands of edge devices to train model locally without uploading data concentrically to the server. But since real federated settings are resource-constrained, FL is encountered with systems heterogeneity which causes a lot of stragglers directly and then leads to significantly accuracy reduction indirectly. To solve the problems caused by systems heterogeneity, we introduce a novel self-adaptive federated framework FedSAE which adjusts the training task of devices automatically and selects participants actively to alleviate the performance degradation. In this work, we 1) propose FedSAE which leverages the complete information of devices' historical training tasks to predict the affordable training workloads for each device. In this way, FedSAE can estimate the reliability of each device and self-adaptively adjust the amount of training load per client in each round. 2)combine our framework with Active Learning to self-adaptively select participants. Then the framework accelerates the convergence of the global model. In our framework, the server evaluates devices' value of training based on their training loss. Then the server selects those clients with bigger value for the global model to reduce communication overhead. The experimental result indicates that in a highly heterogeneous system, FedSAE converges faster than FedAvg, the vanilla FL framework. Furthermore, FedSAE outperforms than FedAvg on several federated datasets - FedSAE improves test accuracy by 26.7% and reduces stragglers by 90.3% on average. Moming Duan, Duo Liu 0002, Yu Zhang 0184, Ao Ren, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002 |
IJCNN | 5 |
| 2021 | CSAFL: A Clustered Semi-Asynchronous Federated Learning FrameworkabstractFederated learning (FL) is an emerging distributed machine learning paradigm that protects privacy and tackles the problem of isolated data islands. At present, there are two main communication strategies of FL: synchronous FL and asynchronous FL. The advantages of synchronous FL are that the model has high precision and fast convergence speed. However, this synchronous communication strategy has the risk that the central server waits too long for the devices, namely, the straggler effect which has a negative impact on some time-critical applications. Asynchronous FL has a natural advantage in mitigating the straggler effect, but there are threats of model quality degradation and server crash. Therefore, we combine the advantages of these two strategies to propose a clustered semi-asynchronous federated learning (CSAFL) framework. We evaluate CSAFL based on four imbalanced federated datasets in a non-IID setting and compare CSAFL to the baseline methods. The experimental results show that CSAFL significantly improves test accuracy by more than +5% on the four datasets compared to TA-FedAvg. In particular, CSAFL improves absolute test accuracy by +34.4% on non-IID FEMNIST compared to TA-FedAvg. Yu Zhang 0184, Moming Duan, Duo Liu 0002, Ao Ren, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002 |
IJCNN | 5 |
| 2020 | DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural NetworksabstractThe rapidly growing parameter volume of deep neural networks (DNNs) hinders the artificial intelligence applications on resource constrained devices, such as mobile and wearable devices. Neural network pruning, as one of the mainstream model compression techniques, is under extensive study to reduce the model size and thus the amount of computation. And thereby, the state-of-the-art DNNs are able to be deployed on those devices with high runtime energy efficiency. In contrast to irregular pruning that incurs high index storage and decoding overhead, structured pruning techniques have been proposed as the promising solutions. However, prior studies on structured pruning tackle the problem mainly from the perspective of facilitating hardware implementation, without diving into the deep to analyze the characteristics of sparse neural networks. The neglect on the study of sparse neural networks causes inefficient trade-off between regularity and pruning ratio. Consequently, the potential of structurally pruning neural networks is not sufficiently mined.In this work, we examine the structural characteristics of the irregularly pruned weight matrices, such as the diverse redundancy of different rows, the sensitivity of different rows to pruning, and the position characteristics of retained weights. By leveraging the gained insights as a guidance, we first propose the novel block-max weight masking (BMWM) method, which can effectively retain the salient weights while imposing high regularity to the weight matrix. As a further optimization, we propose a density-adaptive regular-block (DARB) pruning that can effectively take advantage of the intrinsic characteristics of neural networks, and thereby outperform prior structured pruning work with high pruning ratio and decoding efficiency. Our experimental results show that DARB can achieve 13× to 25× pruning ratio, which are 2.8× to 4.3× improvements than the state-of-the-art counterparts on multiple neural network models and tasks. Moreover, DARB can achieve 14.3× decoding efficiency than block pruning with higher pruning ratio. Ao Ren, Tao Zhang 0032, Yuhao Wang 0002, Sheng Lin 0001, Peiyan Dong, Yen-Kuang Chen, Yuan Xie 0001, Yanzhi Wang 0001 |
AAAI | 1 |
| 2019 | ADMM-NN: An Algorithm-Hardware Co-Design Framework of DNNs Using Alternating Direction Methods of MultipliersabstractModel compression is an important technique to facilitate efficient embedded and hardware implementations of deep neural networks (DNNs), a number of prior works are dedicated to model compression techniques. The target is to simultaneously reduce the model storage size and accelerate the computation, with minor effect on accuracy. Two important categories of DNN model compression techniques are weight pruning and weight quantization. The former leverages the redundancy in the number of weights, whereas the latter leverages the redundancy in bit representation of weights. These two sources of redundancy can be combined, thereby leading to a higher degree of DNN model compression. However, a systematic framework of joint weight pruning and quantization of DNNs is lacking, thereby limiting the available model compression ratio. Moreover, the computation reduction, energy efficiency improvement, and hardware performance overhead need to be accounted besides simply model size reduction, and the hardware performance overhead resulted from weight pruning method needs to be taken into consideration. To address these limitations, we present ADMM-NN, the first algorithm-hardware co-optimization framework of DNNs using Alternating Direction Method of Multipliers (ADMM), a powerful technique to solve non-convex optimization problems with possibly combinatorial constraints. The first part of ADMM-NN is a systematic, joint framework of DNN weight pruning and quantization using ADMM. It can be understood as a smart regularization technique with regularization target dynamically updated in each ADMM iteration, thereby resulting in higher performance in model compression than the state-of-the-art. The second part is hardware-aware DNN optimizations to facilitate hardware-level implementations. We perform ADMM-based weight pruning and quantization considering (i) the computation reduction and energy efficiency improvement, and (ii) the hardware performance overhead due to irregular sparsity. The first requirement prioritizes the convolutional layer compression over fully-connected layers, while the latter requires a concept of the break-even pruning ratio, defined as the minimum pruning ratio of a specific layer that results in no hardware performance degradation. Without accuracy loss, ADMM-NN achieves 85× and 24× pruning on LeNet-5 and AlexNet models, respectively, --- significantly higher than the state-of-the-art. The improvements become more significant when focusing on computation reduction. Combining weight pruning and quantization, we achieve 1,910× and 231× reductions in overall model size on these two benchmarks, when focusing on data storage. Highly promising results are also observed on other representative DNNs such as VGGNet and ResNet-50. We release codes and models at https://github.com/yeshaokai/admm-nn. Ao Ren, Tianyun Zhang, Shaokai Ye, Wenyao Xu, Xuehai Qian, Xue Lin 0001, Yanzhi Wang 0001 |
ASPLOS | 1 |
| 2019 | A Majority Logic Synthesis Framework for Adiabatic Quantum-Flux-Parametron Superconducting CircuitsabstractAdiabatic Quantum-Flux-Parametron (AQFP) logic is an adiabatic superconductor logic that has been proposed as alternative to CMOS logic with extremely high energy efficiency. In AQFP technology, majority-based gates have the same area as two-input AND/OR gates while offering more complex logic. Therefore, majority-based logic (MAJ) is more preferred than and-or-inverter-based logic (AOI) to implement logic functions in AQFP for higher energy efficiency. In this paper, we propose a majority gates synthesis framework for AQFP circuits that is capable of converting any AOI netlist to its corresponding MAJ netlist by mapping all feasible three-input sub- netlists to corresponding MAJ based implementations. In addition, the proposed tool can insert the optimal amount of buffers and splitters for equivalent delay as required in the AQFP technology. Experimental results suggest that the proposed method can reduce delay and area by up to 60.00% and 60.98%, respectively. Ruizhe Cai, Olivia Chen, Ao Ren, Ning Liu 0007, Caiwen Ding, Nobuyuki Yoshikawa, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | A Buffer and Splitter Insertion Framework for Adiabatic Quantum-Flux-Parametron Superconducting CircuitsabstractAdiabatic Quantum-Flux-Parametron (AQFP) logic is an adiabatic superconductor logic that has been proposed as alternative to CMOS logic with extremely high energy efficiency. In AQFP technology, gates are driven by AC-power, which also serves as clock signal to synchronize the outputs of all gates in the same clock phase. As a matter of fact, AQFP circuits may require huge amount of buffers and splitters to be inserted to allow inputs to any gate having equal delay. Existing buffer and splitter insertion method does not deliver optimization, which could lead to huge space and delay overhead. A better automated buffer and splitter framework is imminent for more efficient AQFP circuits design. In this paper, we propose an automated buffer and splitter insertion method that is capable of adding optimized amount of buffers and splitters to any given gate-level netlist to achieve equal delay for all gates. The proposed method achieve equal delay by inserting buffers and splitters with any library limitation on the size of splitters. Experimental results suggest that the proposed method can deliver better results compared with the existing method, with up-to 40.84% less in size and 3.13% less in delay when splitter fan-out size is limited to four. Ruizhe Cai, Olivia Chen, Ao Ren, Ning Liu 0007, Nobuyuki Yoshikawa, Yanzhi Wang 0001 |
ICCD | 3 |
| 2019 | A stochastic-computing based deep learning framework using adiabatic quantum-flux-parametron superconducting technologyabstractThe Adiabatic Quantum-Flux-Parametron (AQFP) superconducting technology has been recently developed, which achieves the highest energy efficiency among superconducting logic families, potentially 104--105 gain compared with state-of-the-art CMOS. In 2016, the successful fabrication and testing of AQFP-based circuits with the scale of 83,000 JJs have demonstrated the scalability and potential of implementing large-scale systems using AQFP. As a result, it will be promising for AQFP in high-performance computing and deep space applications, with Deep Neural Network (DNN) inference acceleration as an important example. Ruizhe Cai, Ao Ren, Olivia Chen, Ning Liu 0007, Caiwen Ding, Xuehai Qian, Jie Han 0001, Wenhui Luo, Nobuyuki Yoshikawa, Yanzhi Wang 0001 |
ISCA | 2 |
| 2019 | Normalization and dropout for stochastic computing-based deep convolutional neural networks
Ji Li 0006, Zhe Li 0001, Ao Ren, Caiwen Ding, Jeffrey T. Draper, Shahin Nazarian, Qinru Qiu, Bo Yuan 0001, Yanzhi Wang 0001 |
Integr. | 4 |
| 2019 | HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural NetworksabstractDeep convolutional neural networks (DCNNs) are one of the most promising deep learning techniques and have been recognized as the dominant approach for almost all recognition and detection tasks. The computation of DCNNs is memory intensive due to large feature maps and neuron connections, and the performance highly depends on the capability of hardware resources. With the recent trend of wearable devices and Internet of Things, it becomes desirable to integrate the DCNNs onto embedded and portable devices that require low power and energy consumptions and small hardware footprints. Recently stochastic computing (SC)-DCNN demonstrated that SC as a low-cost substitute to binary-based computing radically simplifies the hardware implementation of arithmetic units and has the potential to satisfy the stringent power requirements in embedded devices. In SC, many arithmetic operations that are resource-consuming in binary designs can be implemented with very simple hardware logic, alleviating the extensive computational complexity. It offers a colossal design space for integration and optimization due to its reduced area and soft error resiliency. In this paper, we present HEIF, a highly efficient SC-based inference framework of the large-scale DCNNs, with broad applications including (but not limited to) LeNet-5 and AlexNet, that achieves high energy efficiency and low area/hardware cost. Compared to SC-DCNN, HEIF features: 1) the first (to the best of our knowledge) SC-based rectified linear unit activation function to catch up with the recent advances in software models and mitigate degradation in application-level accuracy; 2) the redesigned approximate parallel counter and optimized stochastic multiplication using transmission gates and inverse mirror adders; and 3) the new optimization of weight storage using clustering. Most importantly, to achieve maximum energy efficiency while maintaining acceptable accuracy, HEIF considers holistic optimizations on cascade connection of function blocks in DCNN, pipelining technique, and bit-stream length reduction. Experimental results show that in large-scale applications HEIF outperforms previous SC-DCNN by the throughput of 4.1×, by area efficiency of up to 6.5×, and achieves up to 5.6× energy improvement. Zhe Li 0001, Ji Li 0006, Ao Ren, Ruizhe Cai, Caiwen Ding, Xuehai Qian, Jeffrey T. Draper, Bo Yuan 0001, Jian Tang 0008, Qinru Qiu, Yanzhi Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | VIBNN: Hardware Acceleration of Bayesian Neural NetworksabstractBayesian Neural Networks (BNNs) have been proposed to address the problem of model uncertainty in training and inference. By introducing weights associated with conditioned probability distributions, BNNs are capable of resolving the overfitting issue commonly seen in conventional neural networks and allow for small-data training, through the variational inference process. Frequent usage of Gaussian random variables in this process requires a properly optimized Gaussian Random Number Generator (GRNG). The high hardware cost of conventional GRNG makes the hardware implementation of BNNs challenging. In this paper, we propose VIBNN, an FPGA-based hardware accelerator design for variational inference on BNNs. We explore the design space for massive amount of Gaussian variable sampling tasks in BNNs. Specifically, we introduce two high performance Gaussian (pseudo) random number generators: 1) the RAM-based Linear Feedback Gaussian Random Number Generator (RLF-GRNG), which is inspired by the properties of binomial distribution and linear feedback logics; and 2) the Bayesian Neural Network-oriented Wallace Gaussian Random Number Generator. To achieve high scalability and efficient memory access, we propose a deep pipelined accelerator architecture with fast execution and good hardware utilization. Experimental results demonstrate that the proposed VIBNN implementations on an FPGA can achieve throughput of 321,543.4 Images/s and energy efficiency upto 52,694.8 Images/J while maintaining similar accuracy as its software counterpart. Ruizhe Cai, Ao Ren, Ning Liu 0007, Caiwen Ding, Luhao Wang, Xuehai Qian, Massoud Pedram, Yanzhi Wang 0001 |
ASPLOS | 2 |
| 2018 | Structured Weight Matrices-Based Hardware Accelerators in Deep Neural Networks: FPGAs and ASICsabstractBoth industry and academia have extensively investigated hardware accelerations. To address the demands in increasing computational capability and memory requirement, in this work, we propose the structured weight matrices (SWM)-based compression technique for both Field Programmable Gate Array (FPGA) and application-specific integrated circuit (ASIC) implementations. In the algorithm part, the SWM-based framework adopts block-circulant matrices to achieve a fine-grained tradeoff between accuracy and compression ratio. The SWM-based technique can reduce computational complexity from O(n2) to O(nlog n) and storage complexity from O(n2) to O(n) for each layer and both training and inference phases. For FPGA implementations on deep convolutional neural networks (DCNNs), we achieve at least 152X and 72X improvement in performance and energy efficiency, respectively using the SWM-based framework, compared with the baseline of IBM TrueNorth processor under same accuracy constraints using the data set of MNIST, SVHN, and CIFAR-10. For FPGA implementations on long short term memory (LSTM) networks, the proposed SWM-based LSTM can achieve up to 21X enhancement in performance and 33.5X gains in energy efficiency compared with the ESE accelerator. For ASIC implementations, the proposed SWM-based ASIC design exhibits impressive advantages in terms of power, throughput, and energy efficiency. Experimental results indicate that this method is greatly suitable for applying DNNs onto both FPGAs and mobile/IoT devices. Caiwen Ding, Ao Ren, Geng Yuan, Ning Liu 0007, Bo Yuan 0001, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | Towards acceleration of deep convolutional neural networks using stochastic computingabstractIn recent years, Deep Convolutional Neural Network (DCNN) has become the dominant approach for almost all recognition and detection tasks and outperformed humans on certain tasks. Nevertheless, the high power consumptions and complex topologies have hindered the widespread deployment of DCNNs, particularly in wearable devices and embedded systems with limited area and power budget. This paper presents a fully parallel and scalable hardware-based DCNN design using Stochastic Computing (SC), which leverages the energy-accuracy trade-off through optimizing SC components in different layers. We first conduct a detailed investigation of the Approximate Parallel Counter (APC) based neuron and multiplexer-based neuron using SC, and analyze the impacts of various design parameters, such as bit stream length and input number, on the energy/power/area/accuracy of the neuron cell. Then, from an architecture perspective, the influence of inaccuracy of neurons in different layers on the overall DCNN accuracy (i.e., software accuracy of the entire DCNN) is studied. Accordingly, a structure optimization method is proposed for a general DCNN architecture, in which neurons in different layers are implemented with optimized SC components, so as to reduce the area, power, and energy of the DCNN while maintaining the overall network performance in terms of accuracy. Experimental results show that the proposed approach can find a satisfactory DCNN configuration, which achieves 55X, 151X, and 2X improvement in terms of area, power and energy, respectively, while the error is increased by 2.86%, compared with the conventional binary ASIC implementation. Ji Li 0006, Ao Ren, Zhe Li 0001, Caiwen Ding, Bo Yuan 0001, Qinru Qiu, Yanzhi Wang 0001 |
ASP-DAC | 2 |
| 2017 | Algorithm-hardware co-optimization of the memristor-based framework for solving SOCP and homogeneous QCQP problemsabstractA memristor crossbar, which is constructed with memristor devices, has the unique ability to change and memorize the state of each of its memristor elements. It also has other highly desirable features such as high density, low power operation and excellent scalability. Hence the memristor crossbar technology can potentially be utilized for developing low-complexity and high-scalability solution frameworks for solving a large class of convex optimization problems, which involve extensive matrix operations and have critical applications in multiple disciplines. This paper, as the first attempt towards this direction, proposes a novel memristor crossbar-based framework for solving two important convex optimization problems, i.e., second-order cone programming (SOCP) and homogeneous quadratically constrained quadratic programming (QCQP) problems. In this paper, the alternating direction method of multipliers (ADMM) is adopted. It splits the SOCP and homogeneous QCQP problems into sub-problems that involve the solution of linear systems, which could be effectively solved using the memristor crossbar in O(1) time complexity. The proposed algorithm is an iterative procedure that iterates a constant number of times. Therefore, algorithms to solve SOCP and homogeneous QCQP problems have pseudo-O(N) complexity, which is a significant reduction compared to the state-of-the-art software solvers (O(N3.5)-O(N4)). Ao Ren, Sijia Liu 0001, Ruizhe Cai, Wujie Wen, Pramod K. Varshney, Yanzhi Wang 0001 |
ASP-DAC | 1 |
| 2017 | SC-DCNN: Highly-Scalable Deep Convolutional Neural Network using Stochastic ComputingabstractWith the recent advance of wearable devices and Internet of Things (IoTs), it becomes attractive to implement the Deep Convolutional Neural Networks (DCNNs) in embedded and portable systems. Currently, executing the software-based DCNNs requires high-performance servers, restricting the widespread deployment on embedded and mobile IoT devices. To overcome this obstacle, considerable research efforts have been made to develop highly-parallel and specialized DCNN accelerators using GPGPUs, FPGAs or ASICs. Ao Ren, Zhe Li 0001, Caiwen Ding, Qinru Qiu, Yanzhi Wang 0001, Ji Li 0006, Xuehai Qian, Bo Yuan 0001 |
ASPLOS | 1 |
| 2017 | Structural design optimization for deep convolutional neural networks using stochastic computingabstractDeep Convolutional Neural Networks (DCNNs) have been demonstrated as effective models for understanding image content. The computation behind DCNNs highly relies on the capability of hardware resources due to the deep structure. DCNNs have been implemented on different large-scale computing platforms. However, there is a trend that DCNNs have been embedded into light-weight local systems, which requires low power/energy consumptions and small hardware footprints. Stochastic Computing (SC) radically simplifies the hardware implementation of arithmetic units and has the potential to satisfy the small low-power needs of DCNNs. Local connectivities and down-sampling operations have made DCNNs more complex to be implemented using SC. In this paper, eight feature extraction designs for DCNNs using SC in two groups are explored and optimized in detail from the perspective of calculation precision, where we permute two SC implementations for inner-product calculation, two down-sampling schemes, and two structures of DCNN neurons. We evaluate the network in aspects of network accuracy and hardware performance for each DCNN using one feature extraction design out of eight. Through exploration and optimization, the accuracies of SC-based DCNNs are guaranteed compared with software implementations on CPU/GPU/binary-based ASIC synthesis, while area, power, and energy are significantly reduced by up to 776x, 190x, and 32835x. Zhe Li 0001, Ao Ren, Ji Li 0006, Qinru Qiu, Bo Yuan 0001, Jeffrey T. Draper, Yanzhi Wang 0001 |
DATE | 2 |
| 2017 | Softmax Regression Design for Stochastic Computing Based Deep Convolutional Neural NetworksabstractRecently, Deep Convolutional Neural Networks (DCNNs) have made tremendous advances, achieving close to or even better accuracy than human-level perception in various tasks. Stochastic Computing (SC), as an alternate to the conventional binary computing paradigm, has the potential to enable massively parallel and highly scalable hardware implementations of DCNNs. In this paper, we design and optimize the SC based Softmax Regression function. Experiment results show that compared with a binary SR, the proposed SC-SR under longer bit stream can reach the same level of accuracy with the improvement of 295X, 62X, 2617X in terms of power, area and energy, respectively. Binary SR is suggested for future DCNNs with short bit stream length input whereas SC-SR is recommended for longer bit stream. Ji Li 0006, Zhe Li 0001, Caiwen Ding, Ao Ren, Bo Yuan 0001, Qinru Qiu, Jeffrey T. Draper, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2017 | Ultra-fast robust compressive sensing based on memristor crossbarsabstractIn this paper, we propose a new approach for robust compressive sensing (CS) using memristor crossbars that are constructed by recently invented memristor devices. The exciting features of a memristor crossbar, such as high density, low power and great scalability, make it a promising candidate to perform large-scale matrix operations. To apply memristor crossbars to solve a robust CS problem, the alternating directions method of multipliers (ADMM) is employed to split the original problem into subproblems that involve the solution of systems of linear equations. A system of linear equations can then be solved using memristor crossbars with astonishing O(1) time complexity. We also study the impact of hardware variations on the memristor crossbar based CS solver from both theoretical and practical points of view. The resulting overall complexity is given by O(n), which achieves O(n2.5) speed-up compared to the state-of-the-art software approach. Numerical results are provided to illustrate the effectiveness of the proposed CS solver. Sijia Liu 0001, Ao Ren, Yanzhi Wang 0001, Pramod K. Varshney |
ICASSP | 2 |
| 2017 | Deep reinforcement learning: Framework, applications, and embedded implementations: Invited paperabstractThe recent breakthroughs of deep reinforcement learning (DRL) technique in Alpha Go and playing Atari have set a good example in handling large state and actions spaces of complicated control problems. The DRL technique is comprised of (i) an offline deep neural network (DNN) construction phase, which derives the correlation between each state-action pair of the system and its value function, and (ii) an online deep Q-learning phase, which adaptively derives the optimal action and updates value estimates. In this paper, we first present the general DRL framework, which can be widely utilized in many applications with different optimization objectives. This is followed by the introduction of three specific applications: the cloud computing resource allocation problem, the residential smart grid task scheduling problem, and building HVAC system optimal control problem. The effectiveness of the DRL technique in these three cyber-physical applications have been validated. Finally, this paper investigates the stochastic computing-based hardware implementations of the DRL framework, which consumes a significant improvement in area efficiency and power consumption compared with binary-based implementation counterparts. Hongjia Li 0003, Tianshu Wei, Ao Ren, Qi Zhu 0002, Yanzhi Wang 0001 |
ICCAD | 3 |
| 2017 | Hardware Acceleration of Bayesian Neural Networks Using RAM Based Linear Feedback Gaussian Random Number GeneratorsabstractBayesian neural networks (BNNs) have been proposed to address the problem of model uncertainty in training. By introducing weights associated with conditioned probability distributions, BNN is capable to resolve overfitting issues commonly seen in conventional neural networks. Frequent usage of Gaussian random variables requires a properly optimized Gaussian Random Number Generator (GRNG). The high hardware cost of conventional GRNG makes the hardware realization of BNN challenging. In this paper, a new hardware acceleration architecture for variational inference in BNNs is proposed to facilitate the applicability of BNN in larger-scale applications. In addition, the proposed implementation introduced the RAM based Linear Feedback based GRNG (RLF-GRNG) for effective weight sampling in BNNs. The RAM based Linear Feedback method can effectively utilize RAM resources for parallel Gaussian random number generation while requiring limited and sharable control logic. Implementation on an Altera Cyclone V FPGA suggests that the RLF-GRNG utilizes much less RAM resources compared to other GRNG methods. Experiments results show that the proposed hardware implementation of a BNN can still attain similar accuracy compared to software implementation. Ruizhe Cai, Ao Ren, Luhao Wang, Massoud Pedram, Yanzhi Wang 0001 |
ICCD | 2 |
| 2017 | Hardware-driven nonlinear activation for stochastic computing based deep convolutional neural networksabstractRecently, Deep Convolutional Neural Networks (DCNNs) have made unprecedented progress, achieving the accuracy close to, or even better than human-level perception in various tasks. There is a timely need to map the latest software DCNNs to application-specific hardware, in order to achieve orders of magnitude improvement in performance, energy efficiency and compactness. Stochastic Computing (SC), as a low-cost alternative to the conventional binary computing paradigm, has the potential to enable massively parallel and highly scalable hardware implementation of DCNNs. One major challenge in SC based DCNNs is designing accurate nonlinear activation functions, which have a significant impact on the network-level accuracy but cannot be implemented accurately by existing SC computing blocks. In this paper, we design and optimize SC based neurons, and we propose highly accurate activation designs for the three most frequently used activation functions in software DCNNs, i.e, hyperbolic tangent, logistic, and rectified linear units. Experimental results on LeNet-5 using MNIST dataset demonstrate that compared with a binary ASIC hardware DCNN, the DCNN with the proposed SC neurons can achieve up to 61X, 151X, and 2X improvement in terms of area, power, and energy, respectively, at the cost of small precision degradation. In addition, the SC approach achieves up to 21X and 41X of the area, 41X and 72X of the power, and 198200X and 96443X of the energy, compared with CPU and GPU approaches, respectively, while the error is increased by less than 3.07%. ReLU activation is suggested for future SC based DCNNs considering its superior performance under a small bit stream length. Ji Li 0006, Zhe Li 0001, Caiwen Ding, Ao Ren, Qinru Qiu, Jeffrey T. Draper, Yanzhi Wang 0001 |
IJCNN | 5 |
| 2016 | DSCNN: Hardware-oriented optimization for Stochastic Computing based Deep Convolutional Neural NetworksabstractDeep Convolutional Neural Networks (DCNN), a branch of Deep Neural Networks which use the deep graph with multiple processing layers, enables the convolutional model to finely abstract the high-level features behind an image. Large-scale applications using DCNN mainly operate in high-performance server clusters, GPUs or FPGA clusters; it is restricted to extend the applications onto mobile/wearable devices and Internet-of-Things (IoT) entities due to high power/energy consumption. Stochastic Computing is a promising method to overcome this shortcoming used in specific hardware-based systems. Many complex arithmetic operations can be implemented with very simple hardware logic in the SC framework, which alleviates the extensive computation complexity. The exploration of network-wise optimization and the revision of network structure with respect to stochastic computing based hardware design have not been discussed in previous work. In this paper, we investigate Deep Stochastic Convolutional Neural Network (DSCNN) for DCNN using stochastic computing. The essential calculation components using SC are designed and evaluated. We propose a joint optimization method to collaborate components guaranteeing a high calculation accuracy in each stage of the network. The structure of original DSCNN is revised to accommodate SC hardware design's simplicity. Experimental Results show that as opposed to software inspired feature extraction block in DSCNN, an optimized hardware oriented feature extraction block achieves as higher as 59.27% calculation precision. And the optimized DSCNN can achieve only 3.48% network test error rate compared to 27.83% for baseline DSCNN using software inspired feature extraction block. Zhe Li 0001, Ao Ren, Ji Li 0006, Qinru Qiu, Yanzhi Wang 0001, Bo Yuan 0001 |
ICCD | 2 |