EDBT 2026 Demo / reviewers in the wild / expert
Yi Zou 0001
dblp:19/2439-1
· DBLP profile ↗
42ranked-venue papers
0as first author
21since 2021 · last 2026
0000-0002-4382-4670ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 9 since 2021Security and privacy · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI ModelsabstractStochastic computing (SC) offers hardware simplicity but suffers from low throughput, while high-throughput Digital Computing-in-Memory (DCIM) is bottlenecked by costly adder logic for matrix-vector multiplication (MVM). To address this trade-off, this paper introduces a digital stochastic CIM (DS-CIM) architecture that achieves both high accuracy and efficiency. We implement signed multiply-accumulation (MAC) in a compact, unsigned OR-based circuit by modifying the data representation. Throughput is enhanced by replicating this low-cost circuit 64 times with only a 1× area increase. Our core strategy, a shared Pseudo Random Number Generator (PRNG) with 2D partitioning, enables single-cycle mutually exclusive activation to eliminate OR-gate collisions. We also resolve the 1s saturation issue via stochastic process analysis and data remapping, significantly improving accuracy and resilience to input sparsity. Our high-accuracy DS-CIM1 variant achieves 94.45% accuracy for INT8 ResNet18 on CIFAR-10 with a root-mean-squared error (RMSE) of just 0.74%. Meanwhile, our high-efficiency DS-CIM2 variant attains an energy efficiency of 3566.1 TOPS/W and an area efficiency of 363.7 TOPS/mm2, while maintaining a low RMSE of 3.81%. The DS-CIM capability with larger models is further demonstrated through experiments with INT8 ResNet50 on ImageNet and the FP8 LLaMA-7B model. Kunming Shao, Jiangnan Yu, Zhipeng Liao, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 6 |
| 2026 | CaPGNN: Optimizing parallel graph neural network training with joint caching and resource-aware graph partitioning
Xianfeng Song, Yi Zou 0001 |
Neurocomputing | 2 |
| 2026 | Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-Fly Aligned-Mantissa Bitwidth PredictionabstractFP8 low-precision formats have gained significant adoption in transformer inference and training. However, existing digital compute-in-memory (DCIM) architectures face challenges in supporting variable FP8 aligned-mantissa bitwidths, as unified alignment strategies and fixed-precision multiply accumulate (MAC) units struggle to handle input data with diverse distributions. This work presents a flexible FP8 DCIM accelerator with three innovations: 1) a dynamic shift-aware bitwidth prediction (DSBP) with on-the-fly input prediction that adaptively adjusts weight (2/4/6/8b) and input ($2\sim 12$b) aligned-mantissa precision; 2) a FIFO-based input alignment unit (FIAU) replacing complex barrel shifters with pointer-based control; and 3) a precision-scalable INT MAC array achieving flexible weight precision with minimal overhead. Implemented in 28-nm CMOS with a$64~\times ~96$CIM array, the design achieves 20.4 TFLOPS/W for fixed E5M7, demonstrating$2.8\times $higher FP8 efficiency than previous work while supporting all FP8 formats. Results on Llama-7b show that the DSBP achieves higher efficiency than fixed bitwidth mode at the same accuracy level on both BoolQ and Winogrande datasets, with configurable parameters enabling flexible accuracy–efficiency tradeoffs. Kunming Shao, Zhipeng Liao, Xijie Huang, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | Edge-Accelerated Fall Detection Via Mmwave Biosensing and Quantized CNN: A Soft-Hardware Co-Design ApproachabstractWe present a real-time fall detection system using millimeter-Wave (mmWave) radar and FPGA-accelerated deep learning. The system captures human motion via commercial mmWave radar and extracts Doppler features through a signal processing pipeline comprising Fast Fourier Transform, constant false alarm rate detection, and clustering. A lightweight Convolutional Neural Network (CNN) is designed and trained on Doppler maps for binary fall classification, achieving 99.67 % accuracy. To support efficient deployment, we apply dynamic fixed-point quantization to compress the CNN from 32-bit floating point to 16-bit fixed point, significantly reducing memory and computation cost with minimal accuracy loss. A hardware-software co-design framework is implemented on FPGA, leveraging loop tiling, ping-pong cache, and parallel MAC units for acceleration. The proposed system achieves 21.93 Giga Operation Per Second throughput and 9.2 ms inference latency, with total end-to-end processing time of 51 ms per frame. It enables real-time fall detection for one and two targets with 97 % and 95 % accuracy, respectively, under a 4.2 W power budget. Our design offers a low-power, privacy-preserving, and deployable solution for intelligent elderly care applications. Zufeng Liang, Jiake Tian, Yi Zou 0001 |
BIBM | 3 |
| 2025 | OA-LAMA: An Outlier-Adaptive LLM Inference Accelerator with Memory-Aligned Mixed-Precision Group QuantizationabstractLarge language models (LLMs) face significant deployment challenges due to their substantial memory and computational demands. While low-precision quantization offers a promising solution, the presence of activation outliers severely degrades model accuracy. Existing approaches either compromise hardware efficiency through misaligned memory access or sacrifice quantization granularity through rigid bit-width allocation, particularly when handling non-uniform tensor distributions across and within layers. This paper presents a hardware-software co-designed framework resulting in an outlier-adaptive LLM inference accelerator with memory-aligned mixed-precision group quantization, named OA-LAMA. The framework comprises three key innovations: First, an outlier-adaptive memory-aligned mixed-precision group (OAMAG) format with a novel outlier reordering technique is proposed to preserve accuracy while maintaining DRAM-aligned memory access. Second, a distribution-aware group allocation strategy is proposed to address inter-layer outlier ratio variance. Finally, we design the OA-LAMA hardware architecture with a three-level accumulation architecture and timing-balanced processing elements to support the OAMAG format efficiently. Evaluations demonstrate that OA-LAMA achieves better accuracy than state-of-the-art 4-bit quantization methods while delivering 1.21–3.09× performance improvement and 1.35–2.47× energy efficiency gains over leading LLM accelerators. OA-LAMA establishes new Pareto frontiers in accuracy-efficiency co-optimization for LLM inference. OA-LAMA is open-sourced at https://github.com/CLab-HKUST-GZ/ICCAD25_OA-LAMA.git. Huangxu Chen, Yingbo Hao, Yi Zou 0001 |
ICCAD | 3 |
| 2025 | Optimization of Graph Neural Networks Training Using Graph ReorderingabstractGraph neural networks (GNNs) are specifically designed for graph-structured data and have gained significant attention. However, training GNN on large-scale graphs remains challenging due to the iterative aggregation of high-dimensional features and high computational complexity. Graph sparsity often leads to inefficient memory access and prolonged training times. To address these issues, we propose the Maximum Common Neighbor Graph Reordering (MaCoN-GR) algorithm, which optimizes the memory layout by reducing the physical distance between nodes and their neighbors. We first evaluate MaCoN-GR through its performance on Sparse Matrix-Vector Multiplication (SpMV), a core operation in GNN training. Experimental results demonstrate that MaCoN-GR achieves speedups ranging from 1.10× to 1.35×, indicating that our method effectively enhances memory access efficiency and reduces latency. Building upon this result, we further apply MaCoN-GR to accelerate GNN training. Within the Topology-Oriented Sampling (TOS) model, GraphSAGE, MaCoN-GR achieves up to 1.08× speedup on large-scale datasets, along with notable improvements in accuracy. In the Feature-Oriented Sampling (FOS) framework, we evaluate both MaCoN-GR and METIS partitioning on FOSGNN, achieving speedups between 1.12× and 2.34×, while also improving convergence and model accuracy. These results highlight the effectiveness of MaCoN-GR in optimizing both memory performance and GNN training efficiency. Guohua Wen, Yi Zou 0001, Xianfeng Song |
IJCNN | 2 |
| 2025 | A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight CombinationabstractDeploying mixed-precision neural networks on edge devices is friendly to hardware resources and power consumption. To support fully mixed-precision neural network inference, it is necessary to design flexible hardware accelerators for continuous varying precision operations. However, the previous works have issues on hardware utilization and overhead of reconfigurable logic. In this paper, we propose an efficient accelerator for 2 ∼ 8-bit precision scaling with serial activation input and parallel weight preloaded. First, we set two loading modes for the weight operands and decompose the weight into the corresponding bitwidths, which extends the weight precision support efficiently. Then, to improve hardware utilization of low-precision operations, we design the architecture that performs bit-serial MAC operation with systolic dataflow, and the partial sums are combined spatially. Furthermore, we designed an efficient carry save adder tree supporting both signed and unsigned number summation across rows. The experiment result shows that the proposed accelerator, synthesized with TSMC 28nm CMOS technology, achieves peak throughput of 4.09TOPS and peak energy efficiency of 68.94TOPS/W at 2/2-bit operations. Kunming Shao, Fengshi Tian, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001 |
ISCAS | 6 |
| 2025 | DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM ComputationabstractRetrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge retrieval but faces challenges on edge devices due to high storage, energy, and latency demands. Computing-in-Memory (CIM) offers a promising solution by storing document embeddings in CIM macros and enabling in-situ parallel retrievals but is constrained by either low memory density or limited computational accuracy. To address these challenges, we present DIRC-RAG, a novel edge RAG acceleration architecture leveraging Digital In-ReRAM Computation (DIRC). DIRC integrates a high-density multi-level ReRAM subarray with an SRAM cell, utilizing SRAM and differential sensing for robust ReRAM readout and digital multiply-accumulate (MAC) operations. By storing all document embeddings within the CIM macro, DIRC achieves ultra-low-power, single-cycle data loading, substantially reducing both energy consumption and latency compared to off-chip DRAM. A query-stationary (QS) dataflow is supported for RAG tasks, minimizing on-chip data movement and reducing SRAM buffer requirements. We introduce error optimization for the DIRC ReRAM-SRAM cell by extracting the bit-wise spatial error distribution of the ReRAM subarray and applying targeted bit-wise data remapping. An error detection circuit is also implemented to enhance readout resilience against device-and circuit-level variations.Simulation results demonstrate that DIRC-RAG under TSMC 40nm process achieves an on-chip non-volatile memory density of 5.18Mb/mm2and a throughput of 131 TOPS. It delivers a 4MB retrieval latency of 5.6μs/query and an energy consumption of 0.956μJ/query, while maintaining the retrieval precision. Kunming Shao, Zhipeng Liao, Jiangnan Yu, Xijie Huang, Jingyu He, Fengshi Tian, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui |
ISLPED | 9 |
| 2025 | mmDigit: A Real-Time Digit Recognition Framework in Air-Writing Using FMCW RadarabstractMillimeter-wave (mmWave) radar sensors show significant promise in noncontact human-computer interaction. Using air-writing as a substitute for conventional input devices, such as keyboards and mice has become a pivotal topic in contemporary research. In response to the absence of a dataset in existing studies about air-writing digits and the insufficient exploration of real-time recognition in edge devices, we propose a real-time air-writing digit recognition framework based on mmWave radar, termed mmDigit. Initially, we use mmWave radar equipped with frequency-modulated continuous wave (FMCW) technology to collect digital echo data and design a data processing pipeline to track and reconstruct digital trajectory images. These images are subsequently fed into a lightweight neural network, which is only 6.9K in parameter size, for exploring the images’ quality and the recognition and cross-user capabilities of small-scale air-writing datasets. To enhance mmDigit’s performance, we implement a transfer learning strategy to accommodate a broader range of digit writing styles and habits, achieving a recognition accuracy of 99.14% and a cross-user capability of 94.13%. Additionally, applying a knowledge distillation strategy enables the lightweight network to extract and learn deep-layer features, thereby improving the cross-user recognition accuracy to 96.22%. Jiake Tian, Yi Zou 0001, Jiale Lai, Fangming Liu |
IEEE Internet Things J. | 2 |
| 2024 | CSIFA: A Configurable SRAM-based In-Memory FFT AcceleratorabstractIn the era of Artificial Intelligent (AI) and big data, the demand for signal processing algorithms for large datasets has grown across various fields, notably in the application of the Fast Fourier Transform (FFT). This algorithm is crucial in computation and storage-heavy applications like image and radar signal processing, where it has been integral for decades. In-memory computing (IMC), unlike traditional computing models, integrates computation and storage in the same unit, reducing data transfer time and energy costs. This paper presents CSIFA, a hardware accelerator for up to 1024-point FFT algorithm, utilizing SRAM and some other digital logic to enhance efficiency. Our evaluation indicates that CSIFA performing FFT is able to achieve an overall throughput 15MB/s, computational power 6.65mW, and 115.3GOPS, showcasing its high throughput and energy efficiency for edge application scenarios. Yiyang Lin, Yi Zou 0001 |
ASAP | 2 |
| 2024 | Memory Access Acceleration Through Architecture Design for Edge SoCsabstractThe design of the SoC architecture plays a pivotal role in defining the computational power and performance of Artificial Intelligence of Things (AIoT) products. The CPU serves as the core controller of the SoC, thereby forming the fundamental aspect of its architectural design. This paper presents an SoC architecture featuring a CPU passthrough strategy, which significantly diminishes CPU memory access latency by up to 71% and boosts CPU performance by a minimum of 8% at specific frequencies. Yi Zou 0001, Guohua Wen |
ASAP | 2 |
| 2024 | SoK: Rowhammer on Commodity Operating SystemsabstractRowhammer has drawn much attention from both academia and industry in the past years as rowhammer exploitation poses severe consequences to system security. Since the first comprehensive study of rowhammer in 2014, a number of rowhammer attacks have been demonstrated against dynamic random access memory (DRAM)-based commodity systems to break software confidentiality, integrity and availability. Accordingly, numerous software defenses have been proposed to mitigate rowhammer attacks on commodity systems of either legacy (e.g., DDR3) or recent DRAM (e.g., DDR4). Besides, multiple hardware defenses (e.g., Target Row Refresh) from the industry have been deployed into recent DRAM to eliminate rowhammer, which we categorize as production defenses. Zhi Zhang 0001, Decheng Chen, Jiahao Qi, Yueqiang Cheng, Shijie Jiang, Yiyang Lin, Yansong Gao 0001, Surya Nepal, Yi Zou 0001, Jiliang Zhang 0002, Yang Xiang 0001 |
AsiaCCS | 9 |
| 2024 | GALA-GNN: Optimization for Training GNNs of Large-scale Graphs on Heterogeneous PlatformsabstractGraph Neural Networks (GNNs) demonstrate remarkable learning efficacy on graph-structured data across various real-world domains. However, due to large-scale graph datasets, it is impractical to train the complete graph on a single GPU. Subgraph-level parallel training has proven to be an effective approach for GNN training on large-scale graphs. Nonetheless, this approach presents certain issues: 1. Due to the absence of neighboring vertices at the boundaries, gradient estimation during the training process may exhibit systemic bias, potentially compromising training accuracy; 2. Existing training frameworks often neglect to consider computational capability differences among trainers when allocating the workloads. We propose GALA-GNN to alleviate the aforementioned issues: 1. We introduce a heuristic edge partitioning algorithm, Neighbor Expansion Simulated Annealing, i.e., NESA, which minimizes the number of vertices with missing neighborhoods in subgraphs. Additionally, it allows for the division of the original graph into subgraphs with uneven workloads by setting a threshold manually; 2. We present Computation-Aware Subgraph Enlarge, which not only prevents significant loss in training accuracy but also reduces idle time for high-computation trainers during the training process, thereby enhancing overall computational resource utilization. We compare our proposed GALA-GNN with the current state-of-the-art GNN training framework, DGL. On the medium-scale dataset Flickr, GALA-GNN achieves up to 7.63x speedup, and on the large-scale dataset ogbn-products, it achieves up to 4.84x speedup, without causing significant loss in training accuracy. Yi Zou 0001, Xianfeng Song, Yiqiu Liu |
ISPA | 2 |
| 2024 | 3D Detection Under Foggy Weather: The Effective Multi-Sensor Fusion of LiDAR and RadarabstractAdverse weather conditions are a significant challenge for fully automated driving systems in autonomous vehicles. While radar sensor integration can alleviate some of the adverse effects of weather, the inherent sparsity and altitude discrepancies in radar data remain significant concerns. In response to these challenges, we introduce an innovative sensor fusion method for LiDAR and radar data under foggy weather conditions. This method effectively augments radar data in voxel space. It facilitates the synchronization and complementation of Li-DAR and radar features in the bird's eye view space, thereby achieving high-precision 3D object detection. The effectiveness of this method has been extensively validated through extensive experimentation on the Oxford Radar RobotCar dataset under foggy weather conditions. Jiake Tian, Dacheng Li, Yi Zou 0001, Jiale Lai, Zufeng Liang |
SECON | 3 |
| 2024 | mmHPE: Human Pose Estimation Based on Point Cloud from Millimeter-wave RadarabstractRehabilitation therapy involving repetitive exer-cises targeting specific human joints under the supervision of a doctor is crucial for patients with movement disorders. Yet the cost of commuting and the demand for medical resources are inconvenient for patients. Human-computer interaction can provide remote rehabilitation guidance for patients at home through human pose estimation technology, but privacy concerns with optical sensors and the cost and discomfort of wearable sensors have hindered progress in this field. To address the above challenges, we propose mmHPE, an innovative 3D human pose estimation framework that uses millimeter-wave (mmWave) radar sensors. It initially manipulates the raw data captured from radar sensors to generate a spatiotemporal sequence point cloud dataset. Afterward, we create a Convolutional Neural Network (CNN) that is linked to a Bidirectional Long Short-Term Memory (Bi-LSTM) network. Moreover, a multi-head attention mechanism is employed to boost the network's performance and to accurately estimate the locations of human skeletons. Ultimately, the 21 points with the corre-sponding human pose position are successfully reconstructed. We investigate the mmHPE framework's feasibility and cross-domain stability in different home environments in real-world scenarios. This innovation proffers patients a convenient and privacy-conscious solution for their rehabilitation training req-uisites. Jiale Lai, Jiake Tian, Yi Zou 0001, Xianfeng Song, Fangming Liu, Dacheng Li |
SMC | 3 |
| 2024 | A Novel Pre-Equalization Scheme for Visible Light Communications with Direct Learning ArchitectureabstractVisible light communication is receiving increasing attention as a promising indoor wireless access technology with broad prospects. In this paper, we propose a novel pre-equalization algorithm with direct learning architecture for compensating the nonlinear distortions of visible light communication systems. Particularly, we introduce a deep reinforcement learning based design to directly optimize the parameters by interacting with the environment without any auxiliary networks. We demonstrate that the proposed approach outperforms the deep learning based pre-equalizer with indirect learning in terms of convergence speed, nonlinear compensation and noise immunity. Moreover, we find that the main contribution to the performance gain of the proposed method over the deep learning based preequalizer comes from the hidden layer. Huantian Huang, Yi Zou 0001 |
WCNC | 3 |
| 2024 | Reverse Backdoor Distillation: Towards Online Backdoor Attack Detection for Deep Neural Network ModelsabstractThe backdoor attack on deep neural network models implants malicious data patterns in a model to induce attacker-desirable behaviors. Existing defense methods fall into the online and offline categories, in which the offline models achieve state-of-the-art detection rates but are restricted by heavy computation overhead. In contrast, their more deployable online counterparts lack the means to detect source-specific backdoors with large sizes. This work proposes a new online backdoor detection method—Reverse Backdoor Distillation (RBD) to handle issues associated with source-specific and source-agnostic backdoor attacks. RBD, designed with the novel perspective of distilling instead of erasing backdoor knowledge, is a complementary backdoor detection methodology that can be used in conjunction with other online backdoor defenses. Considering the fact that trigger data will cause overwhelming neuron activation while clean data will not, RBD distills backdoor attack pattern knowledge from a suspicious model to create a shadow model, which is subsequently deployed online along with the original model in scope to predict a backdoor attack. We extensively evaluate RBD on several datasets (MNIST, GTSRB, CIFAR-10) with diverse model architectures and trigger patterns. RBD outperforms online benchmarks in all experimental settings. Notably, RBD demonstrates superior capability in detecting source-specific attacks, where comparison methods fail, underscoring the effectiveness of our proposed technique. Moreover, RBD achieves a computational savings of at least 97%. Zeming Yao, Hangtao Zhang, Yicheng Guo, Xin Tian 0015, Wei Peng 0011, Yi Zou 0001, Leo Yu Zhang, Chao Chen 0015 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | DyLaClass: Dynamic Labeling Based Classification for Optimal Sparse Matrix Format Selection in Accelerating SpMVabstractSparse matrix-vector multiplication (SpMV) is crucial in many scientific and engineering applications, particularly concerning the effectiveness of different sparse matrix storage formats for various architectures, no single format excels across all hardware. Previous research has focused on trying different algorithms to build predictors for the best format, yet it overlooked how to address the issue of the best format changing in the same hardware environment and how to reduce prediction overhead rather than merely considering the overhead in building predictors. This paper proposes a novel classification algorithm for optimizing sparse matrix storage formats, DyLaClass, based on dynamic labeling and flexible feature selection. Particularly, we introduce mixed labels and features with strong correlations, allowing us to achieve ultra-high prediction accuracy with minimal feature inputs, significantly reducing feature extraction overhead. For the first time, we propose the concept of the most suitable storage format rather than the best storage format, which can stably predict changes in the best format for the same matrix across multiple SpMV executions. We further demonstrate the proposed method on the University of Florida’s public sparse matrix collection dataset. Experimental results show that compared to existing work, our method achieves up to 91% classification accuracy. Using two different hardware platforms for verification, the proposed method outperforms existing methods by 1.26 to 1.43 times. Most importantly, the stability of the proposed prediction model is 25.5% higher than previous methods, greatly increasing the feasibility of the model in practical field applications. Yi Zou 0001, Xianfeng Song, Fangming Liu, Quan Xue |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Implicit Hammer: Cross-Privilege-Boundary Rowhammer Through Implicit AccessesabstractRowhammer is a hardware vulnerability in DRAM memory, where repeated access to hammer rows can induce bit flips in neighboringvictim rows. Rowhammer attacks have enabled privilege escalation, sandbox escape, cryptographic key disclosures, etc. A key requirement ofallexisting rowhammer attacks is that an attacker must have access to at least part of an exploitable hammer row. We term such rowhammer attacks as Explicit Hammer. Recently, several proposals leverage the spatial proximity between the accessed hammer rows and the location of the victim rows for a defense against rowhammer. These all aim to deny the attacker's permission to access hammer rows near sensitive data, thus defeating explicit hammer-based attacks. In this paper, we question the core assumption underlying these defenses. We present Implicit Hammer, a confused-deputy attack that causes accesses to hammer rows that the attacker is not allowed to access. It is a paradigm shift in rowhammer attacks since it crosses privilege boundary to stealthily rowhammer an inaccessible row by implicit DRAM accesses. Such accesses are achieved by abusing inherent features of modern hardware and/or software. We propose a generic model to rigorously formalize the necessary conditions to initiate implicit hammer and explicit hammer, respectively. Compared to explicit hammer, implicit hammer can defeat the advanced software-only defenses, stealthy in hiding itself and hard to be mitigated. To demonstrate the practicality of implicit hammer, we have created two implicit hammer's instances, called PThammer and SyscallHammer. Zhi Zhang 0001, Yueqiang Cheng, Wenhao Wang 0001, Yansong Gao 0001, Dongxi Liu, Surya Nepal, Anmin Fu, Yi Zou 0001 |
IEEE Trans. Dependable Secur. Comput. | 10 |
| 2022 | A Real-time Digit Gesture Recognition System Based on mmWave RadarabstractGesture communication is one of the most general communication methods in the world, with the obvious advantage of exchanging information without worrying about the borderline of different languages. Therefore, establishing a cost-effective way of capturing and understanding human gestures has long been a popular research topic regarding human-machine interaction, particularly in emerging scenarios such as smart cities, etc. In this paper, we propose a system based on a commercially available mmWave radar to recognize digits represented by the travel path of the human hand using a specially designed convolutional neural network (CNN) algorithm. We illustrate the proposed system is capable of recording the path of the moving hand in real-time at the cost of 1 transmitter, 2 receivers, and 2.78 GHz bandwidth from the mmWave radar. Our experimental results show that an average prediction accuracy of 98.8% is achieved in a validation test based on a 7:3 ratio split from existing dataset and an average prediction accuracy of 95.3% in generalization test using fresh data. Chun Yuan 0010, Youxuan Zhong, Jiake Tian, Yi Zou 0001 |
ICMLA | 4 |
| 2021 | Detecting Hardware-Assisted Virtualization With Inconspicuous FeaturesabstractRecent years have witnessed the proliferation of the deployment of virtualization techniques. Virtualization is designed to be transparent, that is, unprivileged users should not be able to detect whether a system is virtualized. Such detection can result in serious security threats such as evading virtual machine (VM)-based malware dynamic analysis and exploiting vulnerabilities for cross-VM attacks. The traditional software-based virtualization leaves numerous artifacts/fingerprints, which can be exploited without much effort to detect the virtualization. In contrast, current mainstream hardware-assisted virtualization significantly enhances the virtualization transparency, making itself more transparent and difficult to be detected. Nonetheless, we showcase three new identified low-level inconspicuous features, which can be leveraged by an unprivileged adversary to effectively and stealthily detect the hardware-assisted virtualization. All three features come from the chipset fingerprints, rather than the traces of software-based virtualization implementations (e.g., Xen or KVM). The identified features include i) Translation-Lookaside Buffer (TLB) stores an extra layer of address translations; ii) Last-Level Cache (LLC) caches one more layer of page-table entries; and iii) Level-1 Data (L1D) Cache is unstable. Based on the above features, we develop three corresponding virtualization detection techniques, which are then comprehensively evaluated on three native environments and three popular cloud providers: i) Amazon Elastic Compute Cloud, ii) Google Compute Engine and iii) Microsoft Azure. Experimental results validate that these three adversarial detection techniques are effective (with no false positive) and stealthy (without triggering suspicious system events, e.g., VM-exit) in detecting the above commodity virtualized environments. Zhi Zhang 0001, Yueqiang Cheng, Yansong Gao 0001, Surya Nepal, Dongxi Liu, Yi Zou 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2013 | Accelerator-rich CMPs: From concept to real hardwareabstractApplication-specific accelerators provide 10-100× improvement in power efficiency over general-purpose processors. The accelerator-rich architectures are especially promising. This work discusses a prototype of accelerator-rich CMPs (PARC). During our development of PARC in real hardware, we encountered a set of technical challenges and proposed corresponding solutions. First, we provided system IPs that serve a sea of accelerators to transfer data between userspace and accelerator memories without cache overhead. Second, we designed a dedicated interconnect between accelerators and memories to enable memory sharing. Third, we implemented an accelerator manager to virtualize accelerator resources for users. Finally, we developed an automated flow with a number of IP templates and customizable interfaces to a C-based synthesis flow to enable rapid design and update of PARC. We implemented PARC in a Virtex-6 FPGA chip with integration of platform-specific peripherals and booting of unmodified Linux. Experimental results show that PARC can fully exploit the energy benefits of accelerators at little system overhead. Yuting Chen 0003, Jason Cong, Mohammad Ali Ghodrat, Muhuan Huang, Chunyue Liu, Bingjun Xiao, Yi Zou 0001 |
ICCD | 7 |
| 2012 | Platform characterization for Domain-Specific ComputingabstractWe believe that by adapting architectures to fit the requirements of a given application domain, we can significantly improve the efficiency of computation. To validate the idea for our application domain, we evaluate a wide spectrum of commodity computing platforms to quantify the potential benefits of heterogeneity and customization for the domain-specific applications. In particular, we choose medical imaging as the application domain for investigation, and study the application performance and energy efficiency across a diverse set of commodity hardware platforms, such as general-purpose multi-core CPUs, massive parallel many-core GPUs, low-power mobile CPUs and fine-grain customizable FPGAs. This study leads to a number of interesting observations that can be used to guide further development of domain-specific architectures. Alex Bui, Kwang-Ting Cheng, Jason Cong, Luminita A. Vese, Yi-Chu Wang, Yi Zou 0001 |
ASP-DAC | 7 |
| 2012 | Optimizing memory hierarchy allocation with loop transformations for high-level synthesisabstractFor the majority of computation-intensive application systems, off-chip memory bandwidth is a critical bottleneck for both performance and power consumption. The efficient utilization of limited on-chip memory resources plays a vital role in reducing the off-chip memory accesses. This paper presents an efficient approach for optimizing the on-chip memory allocation by loop transformations in the imperfectly nested loops. We analytically model the on-chip buffer size and off-chip bandwidth after affine loop transformation, loop fusion/distribution and code motion. Branch-and-bound and knapsack reuse techniques are proposed to reduce the computation complexity in finding optimal solutions. Experimental results show that our scheme can save 40% of on-chip memory size with the same bandwidth consumption compared to the previous approaches. Jason Cong, Peng Zhang 0007, Yi Zou 0001 |
DAC | 3 |
| 2012 | Combining module selection and replication for throughput-driven streaming programsabstractStreaming processing is widely adopted in many data-intensive applications in various domains. FPGAs are commonly used to realize these applications since they can exploit inherent data parallelism and pipelining in the applications to achieve a better performance. In this paper we investigate the design space exploration problem (DSE) when mapping streaming applications onto FPGAs. Previous works narrowly focus on using techniques like replication or module selection to meet the throughput target. We propose to combine these two techniques together to guide the design space exploration. A formal formulation and solution to this combined problem is presented in this paper. Our objective is to optimize the total area cost subject to the throughput constraint. In particular, we are able to handle the feedback loops in the streaming programs, which, to the best of our knowledge, has never been discussed in previous work. Our methodology is evaluated with high-level synthesis tools, and we demonstrate our workflow on a set of benchmarks that vary from module kernel design such as FFT to large designs such as an MPEG-4 decoder. Jason Cong, Muhuan Huang, Bin Liu 0006, Peng Zhang 0007, Yi Zou 0001 |
DATE | 5 |
| 2012 | FPGA-accelerated 3D reconstruction using compressive sensingabstractThe radiation dose associated with computerized tomography (CT) is significant. Optimization-based iterative reconstruction approaches, e.g., compressive sensing provide ways to reduce the radiation exposure, without sacrificing image quality. However, the computational requirement such algorithms is much higher than that of the conventional Filtered Back Projection (FBP) reconstruction algorithm. This paper describes an FPGA implementation of one important iterative kernel called EM, which is the major computation kernel of a recent EM+TV reconstruction algorithm. We show that a hybrid approach (CPU+GPU+FPGA) can deliver a better performance and energy efficiency than GPU-only solutions, providing 13X boost of throughput than a dual-core CPU implementation. Jason Cong, Ming Yan 0006, Yi Zou 0001 |
FPGA | 4 |
| 2012 | Mapping a data-flow programming model onto heterogeneous platformsabstractIn this paper we explore mapping of a high-level macro data-flow programming model called Concurrent Collections (CnC) onto heterogeneous platforms in order to achieve high performance and low energy consumption while preserving the ease of use of data-flow programming. Modern computing platforms are becoming increasingly heterogeneous in order to improve energy efficiency. This trend is clearly seen across a diverse spectrum of platforms, from small-scale embedded SOCs to large-scale super-computers. However, programming these heterogeneous platforms poses a serious challenge for application developers. We have designed a software flow for converting high-level CnC programs to the Habanero-C language. CnC programs have a clear separation between the application description, the implementation of each of the application components and the abstraction of hardware platform, making it an excellent programming model for domain experts. Domain experts can later employ the help of a tuning expert (either a compiler or a person) to tune their applications with minimal effort. We also extend the Habanero-C runtime system to support work-stealing across heterogeneous computing devices and introduce task affinity for these heterogeneous components to allow users to fine tune the runtime scheduling decisions. We demonstrate a working example that maps a pipeline of medical image-processing algorithms onto a prototype heterogeneous platform that includes CPUs, GPUs and FPGAs. For the medical imaging domain, where obtaining fast and accurate results is a critical step in diagnosis and treatment of patients, we show that our model offers up to 17.72X speedup and an estimated usage of 0.52X of the power used by CPUs alone, when using accelerators (GPUs and FPGAs) and CPUs. Alina Simion Sbîrlea, Yi Zou 0001, Zoran Budimlic, Jason Cong, Vivek Sarkar |
LCTES | 2 |
| 2011 | Domain-specific processor with 3D integration for medical image processingabstractThe growth of 3D technology had led to opportunities for stacked multiprocessor-accelerator computing platforms with high-bandwidth and low-latency TSV connections between them, resulting in high computing performance and better energy efficiency. This work evaluates the performance and energy benefits of such an advanced architecture and addresses associated design problems. To better utilize the reconfigurable hardware resource and to explore the opportunity of kernel sharing across applications, we propose to use a dedicated domain-specific computing platform. In particular, we have chosen medical image processing as the domain in this work to accelerate due to its growing for real-time processing demand yet inadequete performance on conventional computing architectures. A design flow is proposed in this work for the 3D multiprocessor-accelerator platform and a number of methods are applied to optimize the average performance of all the applications in the targeted domain under area and bandwidth constraints. Experiments show that the applications in this domain can gain a 7.4× speed-up and 18.8× energy savings on average running on our platform using CMP cores and domain-specific accelerators as compared to their counterparts coded in CPU only. Jason Cong, Karthik Gururaj, Muhuan Huang, Bingjun Xiao, Yi Zou 0001 |
ASAP | 6 |
| 2011 | A reuse-aware prefetching scheme for scratchpad memoryabstractScratchpad memory (SPM) has been utilized as prefetch buffer in embedded systems and parallel architectures to hide memory access latency. However, the impact of reuse pattern on SPM prefetching has not been fully investigated. In this paper we quantify the impact of reuse on SPM prefetching efficiency and propose a reuse-aware SPM prefetching (RASP) scheme. The average performance and energy improvements are 15.9% and 22.0% over cache prefetching, 12.9% and 31.2% over prefetch-only SPM management, 18.5% and 10% over DRDU [1] with SPM prefetching support. Jason Cong, Hui Huang 0001, Chunyue Liu, Yi Zou 0001 |
DAC | 4 |
| 2011 | Resolving implicit barrier synchronizations in FPGA HLS (abstract only)abstractExisting C-to-FPGA high-level synthesis (HLS) tools are good at exploiting instruction-level and loop-level parallelism, but not efficient at exploiting task-level parallelism (or requires extensive manual re-write). In this paper, we present an automated flow and architecture template to map task-level data-model (coarse-grain task graphs) onto FPGAs. We automatically generate multiple communicating FSMDs (finite-state machine with datapath) based on the architecture template to manage the required synchronization and communication. The key architectural components in our approach include a progress table that traces task execution, an automatic memory duplication scheme that enables a larger parallelism and provides isolation between different tasks, and a memory lifetime analysis and reuse scheme to reduce the on-chip memory consumption. Through data-flow driven execution, the presented flow can reduce the total latency by 30%, with moderate area overhead, when compared with the RTL implementation using a single hierarchical FSMD generated by a state-of-art HLS tool. Jason Cong, Yi Zou 0001 |
FPGA | 2 |
| 2011 | Accelerating Fluid Registration Algorithm on Multi-FPGA PlatformsabstractIn the clinical applications, medical image registrations on the images taken from different times and/or through different modalities are needed in order to have an objective clinical assessment of the patient. Viscous fluid registration is a powerful PDE-based method that can register large deformations in the imaging process. This paper presents our implementation of the fluid registration algorithm on a multi-FPGA platform Convey HC-1. We obtain a 35X speedup versus single-threaded software on a CPU. The implementation is achieved using a high-level synthesis (HLS) tool, with additional source-code level optimizations including fixed-point conversion, tiling, prefetching, data-reuse, and streaming across modules using a ghost zone (time-tiling) approach. The experience of this case study also identifies further automation steps needed by existing HLS software. Jason Cong, Muhuan Huang, Yi Zou 0001 |
FPL | 3 |
| 2011 | Combined loop transformation and hierarchy allocation for data reuse optimizationabstractExternal memory bandwidth is a crucial bottleneck in the majority of computation-intensive applications for both performance and power consumption. Data reuse is an important technique for reducing the external memory access by utilizing the memory hierarchy. Loop transformation for data locality and memory hierarchy allocation are two major steps in data reuse optimization flow. But they were carried out independently. This paper presents a combined approach which optimizes loop transformation and memory hierarchy allocation simultaneously to achieve global optimal results on external memory bandwidth and on-chip data reuse buffer size. We develop an efficient and optimal solution to the combined problem by decomposing the solution space into two subspaces with linear and nonlinear constraints respectively. We show that we can significantly prune the solution space without losing its optimality. Experimental results show that our scheme can save up to 31% of on-chip memory size compared to the separated two-step method when the memory hierarchy allocation problem is not trivial. Also, run-time complexity is acceptable for the practical cases. Jason Cong, Peng Zhang 0007, Yi Zou 0001 |
ICCAD | 3 |
| 2011 | An energy-efficient adaptive hybrid cache
Jason Cong, Karthik Gururaj, Hui Huang 0001, Chunyue Liu, Glenn Reinman, Yi Zou 0001 |
ISLPED | 6 |
| 2011 | Automatic memory partitioning and scheduling for throughput and power optimizationabstractMemory bottleneck has become a limiting factor in satisfying the explosive demands on performance and cost in modern embedded system design. Selected computation kernels for acceleration are usually captured by nest loops, which are optimized by state-of-the-art techniques like loop tiling and loop pipelining. However, memory bandwidth bottlenecks prevent designs from reaching optimal throughput with respect to available parallelism. In this paper we present an automatic memory partitioning technique which can efficiently improve throughput and reduce energy consumption of pipelined loop kernels for given throughput constraints and platform requirements. Also, our proposed algorithm can handle general array access beyond affine array references. Our partition scheme consists of two steps. The first step considers cycle accurate scheduling information to meet the hard constraints on memory bandwidth requirements specifically for synchronized hardware designs. An ILP formulation is proposed to solve the memory partitioning and scheduling problem optimally for small designs, followed by a heuristic algorithm which is more scalable and equally effective for solving large scale problems. Experimental results show an average 6× throughput improvement on a set of real-world designs with moderate area increase (about 45% on average), given that less resource sharing opportunities exist with higher throughput in optimized designs. The second step further partitions the memory banks for reducing the dynamic power consumption of the final design. In contrast to previous approaches, our technique can statically compute memory access frequencies in polynomial time with little or no profiling. Experimental results show about 30% power reduction on the same set of benchmarks. Jason Cong, Wei Jiang 0035, Bin Liu 0006, Yi Zou 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2010 | A Comparative Study on the Architecture Templates for Dynamic Nested LoopsabstractLoops are the most typical constructs that the FCCM community tries to accelerate. Many loop constructs have dynamic loop bounds and may face load balancing issues in their parallel realizations. In this paper we discuss different architecture templates for these dynamic nested loops, and show the benefits and trade-offs between various implementation strategies. We show that a statically scheduled architecture may perform better if the overhead in banking conflicts overwrites the benefits in load balancing by dynamic scheduled architectures. Jason Cong, Yi Zou 0001 |
FCCM | 2 |
| 2010 | Accelerating Monte Carlo based SSTA using FPGAabstractMonte Carlo based SSTA serves as the golden standard against alternative SSTA algorithms, but it is seldom used in practice due to its high computation time. In this paper, we accelerate Monte Carlo based SSTA using the FPGA platform. A simple dataflow pipeline technique will not work well due to the excessive usage of FPGA logic slices. We leverage the recently proposed pattern matching method to identify common circuit structures, and further use a mathematical programming based formulation to explore the trade-off between performance and logic slices consumption. The proposed design provides two orders of magnitude speedup compared to the CPU-based implementation. Jason Cong, Karthik Gururaj, Wei Jiang 0035, Bin Liu 0006, Kirill Minkovich, Yi Zou 0001 |
FPGA | 7 |
| 2009 | Evaluation of Static Analysis Techniques for Fixed-Point Precision OptimizationabstractPrecision analysis and optimization is very important when transforming a floating-point algorithm into fixed-point hardware implementations. The core analysis techniques are either based on dynamic analysis or static analysis. We believe in static error analysis, as it is the only technique that can guarantee the desired worst-case accuracy. In this paper we study various underlying arithmetic candidates that can be used in static error analysis and compare their computed sensitivities. The approaches studied include Affine Arithmetic(AA), General Interval Arithmetic (GIA) and Automatic Differentiation (Symbolic Arithmetic). Our study shows that symbolic method is preferred for expressions with higher order cancellation. For programs without strong cancellation, any method works fairly well and GIA slightly outperforms others. We also study the impact of program transformations on these arithmetics. Jason Cong, Karthik Gururaj, Bin Liu 0006, Chunyue Liu, Zhiru Zhang, Yi Zou 0001 |
FCCM | 7 |
| 2009 | Revisiting bitwidth optimizationsabstractThis paper revisits the classical bitwidth optimization problem for fixed-point designs. Our approach also starts from static analysis (range analysis and precision analysis) techniques. We first point out that AA-based precision analysis, which is widely used in many previous works, are not correctly applied for some corner cases. We propose to use interval coefficients to model the sensitivity (the mathematic form is similar to classical generalized interval arithmetic) to correct those issues. Based on the analysis, we use a convex optimization problem to optimize the area costs. In the formulation, we also allow different bitwidths for the multiple uses of the same signal. We get an additional 2% to 3% area reduction compared with previous AA-based analysis and optimizations. We also conduct some optimality study to see how much the gap between the current methods and the possible optimal point is. Jason Cong, Karthik Gururaj, Bin Liu 0006, Chunyue Liu, Yi Zou 0001, Zhiru Zhang |
FPGA | 5 |
| 2009 | Automatic memory partitioning and scheduling for throughput and power optimizationabstractHardware acceleration is crucial in modern embedded system design to meet the explosive demands on performance and cost. Selected computation kernels for acceleration are usually captured by nest loops, which are optimized by state-of-the-art techniques like loop tiling and loop pipelining. However, memory bandwidth bottlenecks prevent designs to reach optimal throughput with respect to available parallelism. In this paper we present an automatic memory partitioning technique which can efficiently improve throughput and reduce energy consumption of pipelined loop kernels for given throughput constraints and platform requirement. Our partition scheme consists of two steps, the first step considers cycle accurate scheduling information to meet the hard constraints on memory bandwidth requirements specifically for synchronized hardware designs. Experimental results show an average 6X throughput improvement on a set of real world designs with moderate area increase (about 45% on average), given that less resource sharing opportunities exist with higher throughput in optimized designs. The second step further partitions the memory banks for reducing the dynamic power consumption of the final design. In contrast with previous approaches, our technique can statically compute memory access frequencies in polynomial time with little to none profiling. Experimental results show about 30% power reduction on the same set of benchmarks. Jason Cong, Wei Jiang 0035, Bin Liu 0006, Yi Zou 0001 |
ICCAD | 4 |
| 2009 | Parallel multi-level analytical global placement on graphics processing unitsabstractGPU platforms are becoming increasingly attractive for implementing accelerators because they feature a larger number of cores with improved programmability. In this paper, we describe our implementation of a state-of-the-art academic multi-level analytical placer mPL [8] on Nvidia's massively parallel GT200 series platforms. We detail our efforts on performance tuning and optimizations. When compared to software implementation on Intel's recent generation Xeon CPU, the speed of the global placement part of mPL is 15X faster on average using a Tesla C1060 card, with comparable WL. (less than 1% WL degradation on average) Jason Cong, Yi Zou 0001 |
ICCAD | 2 |
| 2009 | FPGA-Based Hardware Acceleration of Lithographic Aerial Image SimulationabstractLithography simulation, an essential step in design for manufacturability (DFM), is still far from computationally efficient. Most leading companies use large clusters of server computers to achieve acceptable turn-around time. Thus coprocessor acceleration is very attractive for obtaining increased computational performance with a reduced power consumption. This article describes the implementation of a customized accelerator on FPGA using a polygon-based simulation model. An application-specific memory partitioning scheme is designed to meet the bandwidth requirements for a large number of processing elements. Deep loop pipelining and ping-pong buffer based function block pipelining are also implemented in our design. Initial results show a 15X speedup versus the software implementation running on a microprocessor, and more speedup is expected via further performance tuning. The implementation also leverages state-of-art C-to-RTL synthesis tools. At the same time, we also identify the need for manual architecture-level exploration for parallel implementations. Moreover, we implement the algorithm on NVIDIA GPUs using the CUDA programming environment, and provide some useful comparisons for different kinds of accelerators. Jason Cong, Yi Zou 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2008 | Lithographic aerial image simulation with FPGA-based hardwareaccelerationabstractLithography simulation, as an essential step in design for manufacturability (DFM), is still far from computationally efficient. Most leading companies use large clusters of server computers to achieve acceptable turn-around time. Thus co-processor acceleration is very attractive for obtaining increased computational performance with reduced power consumption. This paper describes an implementation of a customized accelerator on FPGA using a polygon-based simulation model. An application-specific memory partitioning scheme is designed to meet the bandwidth requirements for a large number of processing elements. Deep loop pipelining and ping-pong buffer based function block pipelining are also implemented in our design. Initial results show a 15X speedup versus the software implementation running on a microprocessor, and more speedup is expected via further performance tuning. The implementation also leverages state-of-art C-to-RTL synthesis tools. At the same time, we also identified the need for manual architecture-level exploration for parallel implementations Jason Cong, Yi Zou 0001 |
FPGA | 2 |