EDBT 2026 Demo / reviewers in the wild / expert
Lei Ju 0001
dblp:97/3791-1
· DBLP profile ↗
96ranked-venue papers
7as first author
48since 2021 · last 2026
0000-0001-6186-5399ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 69 · 6 first-author · 34 since 2021Software engineering, systems software and programming languages · 13 · 2 first-author · 6 since 2021Security and privacy · 12 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neura: A Unified Framework for Hierarchical and Adaptive CGRAsabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a promising solution for energy-efficient acceleration across multiple application domains. Yet, CGRAs face significant scalability challenges that hinder their widespread adoption, stemming from three main concerns: (1) Mapping Scalability — existing mapping algorithms struggle to find feasible and optimal solutions as the design complexity grows; (2) Architectural Limitations — rigid mapping granularity and memory access restrict flexibility and performance; and (3) Dynamic Multi-Kernel Support — dynamic and simultaneous execution of multiple kernels are not thoroughly explored, limiting the applicability of CGRAs in complex multi-kernel scenarios. Cheng Tan 0002, Miaomiao Jiang, Ruihong Yin, Yanghui Ou, Lei Ju 0001, Jeff Zhang 0001 |
ASPLOS (2) | 7 |
| 2026 | A Framework for Developing and Optimizing Fully Homomorphic Encryption Programs on GPUsabstractIn sensitive domains such as healthcare and finance, machine learning increasingly employs Fully Homomorphic Encryption (FHE) to secure both user data and models. Although FHE's intrinsic parallelism naturally aligns with GPU architectures, optimizing GPU kernels alone remains insufficient for efficient end-to-end FHE application development. The inherent complexity of FHE schemes and intricate GPU-specific details impede developers from focusing on high-level program logic. Additionally, FHE's high memory requirements, fine-grained memory operations, and redundant computations introduce further optimization challenges, resulting in inefficiencies even when GPU kernels are individually optimized. This paper introduces EasyFHE, a framework designed to simplify the development and optimization of GPU-accelerated FHE applications. Similar to PyTorch, EasyFHE provides high-level interfaces for defining computational logic while automatically handling low-level tasks, such as implementation selection and memory management. Furthermore, it incorporates an optimization framework that systematically addresses performance bottlenecks by applying tailored optimization passes during the lowering from high-level FHE programs to GPU kernels. Compared to state-of-the-art open-source GPU FHE libraries, EasyFHE uniquely supports FHE programs with memory requirements exceeding typical GPU capacities, achieving an average speedup of 2.88× with a peak of 4.39×. Jianyu Zhao 0004, Xueyu Wu 0001, Guang Fan 0001, Mingzhe Zhang 0005, Shoumeng Yan, Lei Ju 0001, Zhuoran Ji |
ASPLOS (2) | 6 |
| 2026 | Conflux: A High-Performance Keyword Private Retrieval System for Dynamic DatasetsabstractHomomorphic Encryption (HE)-based Private Information Retrieval (PIR) allows clients to retrieve plaintext records from untrusted servers without revealing query content. While promising in theory, existing solutions fall short in practice due to two fundamental limitations: (1) poor support for dynamic datasets, which limits applicability in real-world, evolving workloads; and (2) excessive I/O overhead from full-database scans. These bottlenecks prevent current designs from bridging the gap between cryptographic privacy and system efficiency. In this paper, we introduce Conflux, an efficient keyword PIR system for dynamic data environments through a protocolarchitecture co-design approach. At the protocol level, Conflux employs a novel two-phase retrieval mechanism, consisting of an oblivious filtering phase followed by a precise retrieval phase. This design natively supports efficient online insertions, deletions, and updates, while maintaining near-optimal computational complexity. At the system level, Conflux adopts a heterogeneous accelerator architecture that tightly couples computational storage devices and incorporates software-hardware co-optimization techniques to mitigate I/O bottlenecks. Experimental results show that Conflux reduces query processing time by up to$2.64 \times$compared to the state-of-the-art methods, while retaining full support for dynamic datasets. Zhaoyan Shen, Lei Ju 0001 |
HPCA | 5 |
| 2026 | Pipelonk: Accelerating End-to-End Zero-Knowledge Proof Generation on GPUs for PLONK-Based ProtocolsabstractZero-knowledge proofs (ZKPs) are cryptographic protocols that allow verification of statements without disclosing the underlying information. Among them, PLONK-based ZKPs are particularly notable for offering succinct, non-interactive proofs of knowledge with a universal trusted setup, leading to widespread adoption in blockchain and cryptocurrency applications. Nonetheless, their broader deployment is hindered by long proof-generation times and substantial memory demands. While GPUs can accelerate these computations, their limited memory capacity introduces significant challenges for efficient end-to-end proof generation. Zhiyuan Zhang 0008, Yanxin Cai, Wenhao Yin, Xueyu Wu 0001, Yi Wang 0003, Lei Ju 0001, Zhuoran Ji |
PPoPP | 6 |
| 2026 | Path-Sensitive Abstract Interpretation for WCET EstimationabstractWorst-Case Execution Time (WCET) analysis provides an upper bound on a program’s execution time and is fundamental to the design and verification of real-time systems. Accurate modeling of cache behavior is critical for WCET estimation, as cache-miss latency is typically orders of magnitude larger than cache-hit latency. Since cache behavior is path dependent, existing methods commonly use abstract interpretation to estimate cache behaviors without enumerating all paths. However, conventional abstract interpretation is context-agnostic—adopting the most conservative case across paths—and thus may produce an overestimated WCET bound. To bridge the gap between scalability and accuracy, we propose a path-sensitive abstract-interpretation-based cache analysis that maintains a set of cache states drawn from critical execution paths to derive context-aware cache behavior. This path-sensitive cache analysis integrates seamlessly into standard WCET frameworks, resulting in tight yet provably sound WCET bounds. Experiments show that our approach improves WCET accuracy by an average of 24.83% without sacrificing scalability. Shangshang Xiao, Mengxia Sun, Wei Zhang 0173, Naijun Zhan, Lei Ju 0001 |
Proc. ACM Program. Lang. | 5 |
| 2026 | PIRacle: A Fast and Scalable Private Information Retrieval System for Key-Value Stores
Zhaoyan Shen, Yi Wang 0003, Lei Ju 0001 |
IEEE Trans. Computers | 5 |
| 2026 | A ReRAM-Based Processing-in-Memory Framework for LSM-Based Key-Value StoreabstractLog-structured merge (LSM) tree-based key-value (KV) stores organize writes into hierarchical batches to optimize write performance. However, the notorious compaction process and multi-level query mechanism of LSM-tree severely hurt system performance. Our preliminary experiments show that (1) When compaction occurs in the L0 and L1 of the LSM-tree, it may saturate system computation and memory resources, ultimately causing the entire system to stall, and (2) large number of iterative retrievals across multiple levels is usually required to locate the queried data, while redundant key range overlap in L0 further increases the overhead. Based on these observations, we introduce Re-LSM+, a ReRAM-based Processing-in-Memory framework for LSM-based Key-Value Stores. In Re-LSM+, we offload compaction tasks from the higher levels of the LSM-tree to the PIM processing part. A highly parallel ReRAM compaction accelerator is designed by breaking down the three-phase compaction process into basic logic operations. Additionally, we design an index table and a multi-layer Bloom filter for different levels to improve the query efficiency of the LSM-tree. Evaluation results from db_bench show that Re-LSM+ achieves a 2.37× improvement in random write throughput compared to RocksDB. Furthermore, the ReRAM-based compaction accelerator achieves a 68.16× speedup over the CPU-based implementation and reduces energy consumption to 25.5×. Yuhao Zhang 0006, Zhaoyan Shen, Dongxiao Yu, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2025 | Accelerating Number Theoretic Transform with Multi-GPU Systems for Efficient Zero Knowledge ProofabstractZero-knowledge proofs validate statements without revealing any information, pivotal for applications such as verifiable outsourcing and digital currencies. However, their broad adoption is limited by the prolonged proof generation times, mainly due to two operations: Multi-Scalar Multiplication (MSM) and Number Theoretic Transform (NTT). While MSM has been efficiently accelerated using multi-GPU systems, NTT has not, due to the high inter-GPU communication overhead incurred by its permutation data access pattern. Zhuoran Ji, Jianyu Zhao 0004, Peimin Gao, Xiangkai Yin, Lei Ju 0001 |
ASPLOS (1) | 5 |
| 2025 | TensorNTT: Architecture-Aware Optimizations for Number-Theoretic Transform on Tensor Core Unit
Xiangkai Yin, Shuoyu Wang, Zimeng Zhou, Lei Ju 0001, Zhuoran Ji |
IEEE Big Data | 5 |
| 2025 | Co-Prime: A Co-design Framework for Privacy Preserving Machine Learning on FPGAabstractIn enormous privacy-sensitive machine learning application domains with collaborative data acquisition from multiple participants, secure multi-party computation (MPC) becomes a promising solution for privacy-preserving machine learning (PPML). Secret sharing protocols is a prevalent MPC strategy, where frequent data distribution and recombination are applied to uphold the confidentiality of participants' data. A key challenge for practical deployment of secret sharing protocols in PPML is the massive and unbalanced computation and communication workloads occurred in various linear and non-linear stages of machine learning. The imbalance could be further amplified when powerful hardware accelerators are designed to reduce the computation latency. In this work, we propose Co-Prime, an FPGA-based 3PC framework for efficient PPML without assistance from a secure third party. Co-Prime integrates protocol and hardware co-optimizations to mitigate the communication bottlenecks in secret sharing schemes. Particularly, Co-Prime proposes a novel protocol conversion technique that seamlessly converts data formats to adaptively adopt preferred protocols in various stages of PPML. Accelerator-friendly MPC primitives and system-level design space exploration schemes are designed to achieve latency hiding through overlapping computation and network communication. Finally, it enables direct interaction with data streams via network communication modules on FPGAs to further reduce the network communication overhead. Experimental results demonstrate significant performance improvements over existing privacy-preserving machine learning frameworks, with 2-18x speedup in inference latency across various LAN/WAN environments and neural network models. Jiming Xu, Lei Ju 0001, Wei Zhang 0173 |
CCS | 5 |
| 2025 | FPGA-TrustZone: Security Extension of TrustZone to FPGA for SoC-FPGA Heterogeneous ArchitectureabstractTo address the growing security issues faced by ARM-based mobile devices today, TrustZone was adopted to provide a trusted execution environment (TEE) to protect sensitive data. Such TrustZone-based models have been proven to be effective, but they target CPU architectures and do not work for the security of widely used heterogeneous computing platforms such as FPGAs. To solve this issue, we propose a comprehensive SoC-FPGA security framework, FPGA-TrustZone, to support FPGA TEE by extending the security of ARM TrustZone. Experiments on real SoC-FPGA hardware development boards show that FPGA-TrustZone provides high security with low performance overhead. Xindong Fan, Shuchen Wang, Lei Ju 0001, Zimeng Zhou |
DAC | 5 |
| 2025 | zkVC: Fast Zero-Knowledge Proof for Private and Verifiable ComputingabstractIn the context of cloud computing, services are held on cloud servers, where the clients send their data to the server and obtain the results returned by server. However, the computation, data and results are prone to tampering due to the vulnerabilities on the server side. Thus, verifying the integrity of computation is important in the client-server setting. The cryptographic method known as Zero-Knowledge Proof (ZKP) is renowned for facilitating private and verifiable computing. ZKP allows the client to validate that the results from the server are computed correctly without violating the privacy of the server’s intellectual property. Zero-Knowledge Succinct NonInteractive Argument of Knowledge (zkSNARKs), in particular, has been widely applied in various applications like blockchain and verifiable machine learning. Despite their popularity, existing zkSNARKs approaches remain highly computationally intensive. For instance, even basic operations like matrix multiplication require an extensive number of constraints, resulting in significant overhead. In addressing this challenge, we introduce $z k V C$, which optimizes the ZKP computation for matrix multiplication, enabling rapid proof generation on the server side and efficient verification on the client side. zkVC integrates optimized ZKP modules, such as Constraint-reduced Polynomial Circuit (CRPC) and Prefix-Sum Query (PSQ), collectively yielding a more than $\mathbf{1 2}$-fold increase in proof speed over prior methods. The code is available at https://github.com/UCF-Lou-Lab-PET/zkformer. Yancheng Zhang, Mengxin Zheng, Jingtong Hu, Lei Ju 0001, Yan Solihin, Qian Lou |
DAC | 6 |
| 2025 | When to Skip: Sparsity-Aware Acceleration for Intermittent Neural Network InferenceabstractDeep neural networks (DNNs) are increasingly being utilized on battery-less IoT devices to facilitate intelligent edge computing. These devices, which lack batteries, often encounter frequent power interruptions and operate under the intermittent computing paradigm. Under this paradigm, system states are periodically backed up before power failures, allowing for the resumption of execution from the latest backup to maintain progress. However, the shift on the applications from traditional control-based embedded tasks to memory and computation-intensive DNN tasks presents new challenges for intermittent computing. Particularly, the significant size of intermediate results during DNN inference results in significant backup and resumption overhead. Our research indicates that exploiting the sparsity of DNN weights and inputs in intermittent systems can effectively mitigate both backup and resumption overhead without compromising inference accuracy. In this study, we introduce a novel sparse-aware acceleration framework for intermittent DNN inference. This framework selectively skips redundant computations during DNN inference, reducing both intermittent system backups and computation overhead. Experimental results on an MSP430 device show that our approach achieves up to$2.98 \times$speedup and an average of$2.56 \times$latency reduction compared to state-of-the-art baselines, while maintaining high inference accuracy. Wei Zhang 0173, Lei Ju 0001 |
HPCC | 4 |
| 2025 | Adaptive ML-KEM: A Configurable HW-SW Architecture for Post-Quantum CryptographyabstractThis paper presents a hardware/software (HW/SW) co-designed Module-Lattice-based Key-Encapsulation Mechanism (ML-KEM) accelerator featuring an ARM processor for runtime scheduling and a reconfigurable FPGA for efficient cryptographic kernel execution. Our design supports seamless, bitstream-free switching among ML-KEM's three security levels (1/3/5), enabling real-time key generation, encapsulation, and decapsulation. We propose a Preprocessed Multi-path Delay Commutator NTT (PMDC-NTT) architecture, achieving 15.4% logic reduction and 14.2% latency improvement over conventional MDC-NTT. The proposed HW/SW co-design achieves over$6.7 \times$speedup through pipelined task parallelism, demonstrating significant advantages in both performance and flexibility over prior works. Lei Ju 0001, Zimeng Zhou |
ICCD | 4 |
| 2025 | SmartPIR: A Private Information Retrieval System using Computational Storage Devices
Honghui You, Lei Ju 0001, Zhaoyan Shen |
MICRO | 5 |
| 2025 | Demo: Real-Time Inference on GPU-Based Heterogeneous SoCs with GPU Cache Locking
Kehao Ma, Wei Zhang 0173, Mengying Zhao, Lei Ju 0001 |
RTCSA | 4 |
| 2025 | DAHE: Parameter-Adaptive and Memory Efficient FPGA Acceleration of Homomorphic EncryptionabstractWhile homomorphic encryption (HE) has been well-recognized as a promising data privacy protection technique, there are many challenges to the real-world deployment of HE applications. In this work, we propose a design flow for parameter-adaptive and memory-efficient FPGA acceleration of homomorphic encryption. In the framework, we explore the correlations between HE parameter selection to meet various design objectives and the huge design space due to underlying FPGA hardware resource allocation. Particularly, we demonstrate that adaptive management of the FPGA memory hierarchy is crucial to supporting diverse cryptosystem parameter selection for application-level security, accuracy, and performance requirements. We propose a resource-efficient and flexible micro-architectural design for HE operations, where data access patterns in various pipeline execution stages are optimized for high memory bandwidth utilization. Furthermore, a memory-aware performance model is built for automatic design space exploration for cryptosystem parameter selection and hardware resource provisioning. Experimental results show 1.50X and 1.16X speedup for the NTT and Rotation operations w.r.t. the state-of-the-art FPGA implementation. Meanwhile, the proposed framework generates flexible and high-performance accelerator code for real HE application kernels with different cryptosystem parameters on a wide range of FPGA devices. Yilan Zhu, Honghui You, Wei Zhang 0173, Jiming Xu, Qian Lou, Shoumeng Yan, Lei Ju 0001 |
IEEE Trans. Computers | 7 |
| 2025 | WCET Estimation for CNN Inference on FPGA SoC With Multi-DPU EnginesabstractThe Deep Learning Processor Unit (DPU) released in the official Xilinx Vitis AI toolchain stands as a commercial off-the-shelf solution tailored for accelerating convolutional neural network (CNN) inference on Xilinx FPGA devices. While most FPGA accelerator focus on high performance and energy-efficiency, analyzing the worst-case execution time (WCET) bound is essential for using CNN accelerations in real-time embedded systems design. In this work, we show that in a multi-DPU environment, the observed worst-case inference time for a CNN inference task could become 3X larger w.r.t. the best case inference time, which prompts the prominent importance of a static timing analysis for FPGA-based CNN inference. We propose, to the best of the authors’ knowledge, the first static timing analysis framework for CNN inference in a multi-DPU environment. The proposed framework introduces a generalized timing behavior model for shared bus arbitration and memory access contention between parallel running DPU engines. Additionally, it incorporates a fine-grained memory access contention analysis that takes into account the characteristics of deep learning applications. For a single-DPU environment, the analysis result is 27% tighter in average compared with the state-of-the-art results. Furthermore, our proposed method produces relatively tight estimated results in the multi-DPU environment. Wei Zhang 0173, Yunlong Yu 0004, Nan Guan, Naijun Zhan, Lei Ju 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2024 | Accelerating Multi-Scalar Multiplication for Efficient Zero Knowledge Proofs with Multi-GPU SystemsabstractZero-knowledge proof is a cryptographic primitive that allows for the validation of statements without disclosing any sensitive information, foundational in applications like verifiable outsourcing and digital currency. However, the extensive proof generation time limits its widespread adoption. Even with GPU acceleration, proof generation can still take minutes, with Multi-Scalar Multiplication (MSM) accounting for about 78.2% of the workload. To address this, we present DistMSM, a novel MSM algorithm tailored for distributed multi-GPU systems. At the algorithmic level, DistMSM adapts Pippenger's algorithm for multi-GPU setups, effectively identifying and addressing bottlenecks that emerge during scaling. At the GPU kernel level, DistMSM introduces an elliptic curve arithmetic kernel tailored for contemporary GPU architectures. It optimizes register pressure with two innovative techniques and leverages tensor cores for specific big integer multiplications. Compared to state-of-the-art MSM implementations, DistMSM offers an average 6.39× speedup across various elliptic curves and GPU counts. An MSM task that previously took seconds on a single GPU can now be completed in mere tens of milliseconds. It showcases the substantial potential and efficiency of distributed multi-GPU systems in ZKP acceleration. Zhuoran Ji, Zhiyuan Zhang 0008, Jiming Xu, Lei Ju 0001 |
ASPLOS (3) | 4 |
| 2024 | FHE-CGRA: Enable Efficient Acceleration of Fully Homomorphic Encryption on CGRAsabstractFully Homomorphic Encryption (FHE) is an attractive privacy-preserving technique that allows computation directly on encrypted data without decryption. However, it incurs significant performance and memory costs due to intensive computations. In this work, we investigate the execution of FHE-enabled machine learning (ML) applications. We show that the runtime hardware reconfigurability of the underlying execution units of homomorphic operations is highly desirable for efficient hardware resource utilization during FHE-ML execution, due to the changing FHE encryption variants across different ML stages (e.g., the multiplicative level of the ciphertext) and corresponding optimal execution unit design. Based on the observation, we propose FHE-CGRA, a coarse-grained re-configurable architecture (CGRA) acceleration framework with an MLIR-based compiler toolchain for end-to-end homomorphic applications. The experiment shows that FHE-CGRA achieves up-to 8.15× speedup against a conventional CGRA baseline for accelerating the inference of FHE-encrypted convolution neural network (FHE-CNN) models, and up-to 16.48× power efficiency w.r.t. the state-of-the-art FPGA-based FHE-CNN accelerator design. Miaomiao Jiang, Yilan Zhu, Honghui You, Cheng Tan 0002, Zhaoying Li 0004, Jiming Xu, Lei Ju 0001 |
DAC | 7 |
| 2024 | Cache-aware Task Decomposition for Efficient Intermittent Computing SystemsabstractEnergy harvesting offers a scalable and cost-effective power solution for IoT devices, but it introduces the challenge of frequent and unpredictable power failures due to the unstable environment. To address this, intermittent computing has been proposed, which periodically backs up the system state to non-volatile memory (NVM), enabling robust and sustainable computing even in the face of unreliable power supplies. In modern processors, write back cache is extensively utilized to enhance system performance. However, it poses a challenge during backup operations as it buffers updates to memory, potentially leading to inconsistent system states. One solution is to adopt a write-through cache, which avoids the inconsistency issue but incurs increased memory access latency for each write reference. Some existing work enforces a cache flushing before backups to maintain a consistent system state, resulting in significant backup overhead. In this paper, we point out that although cache delays updates to the main memory, it may preserve a recoverable system state in the main memory. Leveraging this characteristic, we propose a cache-aware task decomposition method that divides an application into multiple tasks, ensuring that no dirty cache lines are evicted during their execution. Furthermore, the cache-aware task decomposition maintains an unchanged memory state during the execution of each task, enabling us to parallelize the backup process with task execution and effectively hide the backup latency. Experimental results with different power traces demonstrate the effectiveness of the proposed system. Wei Zhang 0173, Mengying Zhao, Zimeng Zhou, Lei Ju 0001 |
DAC | 5 |
| 2024 | Freshness-aware Data Backup for Batteryless Sensing SystemsabstractBatteryless sensing systems rely on energy harvested from the environment to execute. However, as the harvested energy is generally weak and unstable, the system may experience frequent power failures during processing and sensing. To make forward progress across power outages, the system backs up the system state from static random access memory (SRAM) to non-volatile memory (NVM) before power failures and then restores it upon reboot. Moreover, to avoid losing the collected data, existing approaches save all the collected data from SRAM to NVM before system-off. The data saving and the frequent system reboots consume a lot of energy and time and thus cause a long blocking time. However, the data stored in SRAM can be retained for a short period even after the system is turned off, as the data retention voltage of SRAM is lower than the minimum operating voltage of the microcontroller unit (MCU). In this paper, we leverage the SRAM data retention capability to retain data with a short lifetime on SRAM, while only save data with a long lifetime to NVM. Consequently, the backup overhead is significantly reduced. However, a design challenge is to decide the turn-off voltage to minimize the blocking time. Specifically, turning off at a higher voltage leaves more energy for a longer retention time and results in lower data saving overhead. But this may also cause more on-offs, leading to more system states saving and system states restoring overhead. To address this challenge, the paper proposes a method to adaptively compute the optimal turn-off voltage. Experimental results show that the proposed method can significantly reduce the blocking time caused by data saving and system reboots. The system can collect more data and exhibits improved responsiveness in sensing the environment. Yunlong Yu 0004, Wei Zhang 0173, Songran Liu, Mingsong Lv, Nan Guan, Lei Ju 0001 |
HPCC | 7 |
| 2024 | SoTimer: A Software-based Timekeeper for Energy Harvesting SystemsabstractThe proliferation of IoT devices has made energy harvesting systems an appealing solution for power numerous IoT devices without the limitations imposed by battery life constraints. Despite its potential, energy harvesting systems often encounter frequent power failures leading to interruptions of tasks due to the generally weak energy output they rely on. In modern computer systems, maintaining a consistent power flow is crucial for tasks such as synchronization and system stability. However, in energy harvesting systems, maintaining this consistency is challenging as the system is hard to track time during power failures. Current approaches rely on additional hardware and exploit physical phenomena to estimate power-off duration, but these methods are susceptible to environmental variations like temperature, humidity, and radiation, and are not easily applicable to commercial-off-the-shelf (COTS) micro-controllers. In this study, we introduce SoTimer, a software-based timekeeping solution designed to mitigate the impact of environmental changes without the need for extra hardware. The core concept of SoTimer lies in the stability of charging power between adjacent power-on and power-off periods, as evidenced by extensive measurements under different energy resources. Building on this insight, we propose a machine learning algorithm to predict power-off times based on power conditions observed during power-on phases. Furthermore, we introduce optimization techniques tailored for MCUs with limited computational capabilities to ensure efficient inference. Experimental results demonstrate that SoTimer achieves a high inference accuracy and a negligible runtime overhead. Yunlong Yu 0004, Wei Zhang 0173, Lei Ju 0001 |
HPCC | 4 |
| 2024 | ICED: An Integrated CGRA Framework Enabling DVFS-Aware AccelerationabstractCoarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. However, the relationships of voltage and frequency with the utilization of CGRA resources and the dynamic management of them are not well explored, leading to inefficient designs. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non-performance-constraining kernels. This paper proposes ICED - an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICED proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICED is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICED improves average utilization by$\mathbf{2}.\mathbf{3}\times$and energy-efficiency by$\mathbf{1}.\mathbf{32}\times$over a conventional CGRA. With streaming applications, ICED can achieve up to$\mathbf{1}.\mathbf{26}\times$energy-efficiency compared with a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput. Cheng Tan 0002, Miaomiao Jiang, Deepak Patil, Yanghui Ou, Zhaoying Li 0004, Lei Ju 0001, Tulika Mitra, Antonino Tumeo, Jeff Zhang 0001 |
MICRO | 6 |
| 2024 | A Compiler-Like Framework for Optimizing Cryptographic Big Integer Multiplication on GPUsabstractWith the growth of digital data and rising security concerns, techniques for privacy-preserving computation have become increasingly essential. Big integer multiplication, pivotal for these applications, is compute-intensive but poses challenges for GPU acceleration due to its complexity and the need for application-specific tailored implementations. This paper presents IMCompiler, a compiler-like framework that automatically gen-erates optimized GPU kernels for integer multiplications used in cryptosystems. It features a frontend-IR-backend structure, where the Intermediate Representation (IR) employs a segmented integer multiplication algorithm to decouple architecture-specific optimizations from high-level parameters. The frontend can then easily translate integer multiplication with various high-level parameters into the IR, while the backend focuses on fine-tuning a single GPU kernel for each device, enabling automatic code generation. Moreover, we introduce a computation diagram to facilitate the analysis of parallelization strategies, inspiring many optimizations, including two-dimensional parallelization, tailored caching strategy, index transposing, and lazy carrying. Experiments show that IMCompiler achieves a 4.47× speedup compared to the widely used baseline and 1.42 × over Nvidia's official library. The speedup will be even higher for larger integers and higher-capacity GPUs. Zhuoran Ji, Jianyu Zhao 0004, Jiming Xu, Shoumeng Yan, Lei Ju 0001 |
MICRO | 6 |
| 2024 | POSTER: Accelerating High-Precision Integer Multiplication used in Cryptosystems with GPUsabstractHigh-precision integer multiplication is crucial in privacy-preserving computational techniques but poses acceleration challenges on GPUs due to its complexity and the diverse bit lengths in cryptosystems. This paper introduces GIM, an efficient high-precision integer multiplication algorithm accelerated with GPUs. It employs a novel segmented integer multiplication algorithm that separates implementation details from bit length, facilitating code optimizations. We also present a computation diagram to analyze parallelization strategies, leading to a series of enhancements. Experiments demonstrate that this approach achieves a 4.47× speedup over the commonly used baseline. Zhuoran Ji, Jiming Xu, Lei Ju 0001 |
PPoPP | 4 |
| 2024 | Cache Behavior Analysis with SP-Relative Addressing for WCET Estimation
Shangshang Xiao, Mengxia Sun, Wei Zhang 0173, Naijun Zhan, Lei Ju 0001 |
SETTA | 5 |
| 2024 | AVA: Inconspicuous Attribute Variation-based Adversarial Attack bypassing DeepFake DetectionabstractWhile DeepFake applications are becoming popular in recent years, their abuses pose a serious privacy threat. Unfortunately, most related detection algorithms to mitigate the abuse issues are inherently vulnerable to adversarial attacks because they are built atop DNN-based classification models, and the literature has demonstrated that they could be bypassed by introducing pixel-level perturbations. Though corresponding mitigation has been proposed, we have identified a new attribute-variation-based adversarial attack (AVA) that perturbs the latent space via a combination of Gaussian prior and semantic discriminator to bypass such mitigation. It perturbs the semantics in the attribute space of DeepFake images, which are inconspicuous to human beings (e.g., mouth open) but can result in substantial differences in DeepFake detection. We evaluate our proposed AVA attack on nine state-of-the-art DeepFake detection algorithms and applications. The empirical results demonstrate that AVA attack defeats the state-of-the-art black box attacks against DeepFake detectors and achieves more than a 95% success rate on two commercial DeepFake detectors. Moreover, our human study indicates that AVA-generated DeepFake images are often imperceptible to humans, which presents huge security and privacy concerns. Xiangtao Meng, Li Wang 0120, Shanqing Guo, Lei Ju 0001, Qingchuan Zhao |
SP | 4 |
| 2024 | A GPU-Based Privacy-Preserving Machine Learning Acceleration SchemeabstractAs the application of artificial intelligence expands, privacy-preserving machine learning has become a critical research focus. Secret sharing, as a commonly used privacy-preserving technique, has broad application prospects due to its lightweight ciphertext computation characteristics. While secret sharing offers lightweight computation, most schemes are implemented on CPU platforms, leaving room for exploration on GPUs. This paper proposes a GPU-based acceleration scheme for privacy-preserving machine learning, utilizing the ABY3 secret sharing protocol and a 32-bit integer ring, while supporting regularization techniques. Experimental results on standard neural networks, such as VGG-16 and AlexNet, demonstrate a significant performance improvement. The proposed approach achieves a 95× improvement over the CPU-based Falcon scheme and a 3.5× improvement over the GPU-based CryptGPU scheme during privacy training, while reducing communication volume by over 50% in both inference and training phases. Zengrui Huang, Zhiyong Zhang 0006, Wei Zhang 0173, Lei Ju 0001 |
TrustCom | 5 |
| 2024 | CommanderUAP: a practical and transferable universal adversarial attacks on speech recognition modelsabstractAbstract Most of the adversarial attacks against speech recognition systems focus on specific adversarial perturbations, which are generated by adversaries for each normal example to achieve the attack. Universal adversarial perturbations (UAPs), which are independent of the examples, have recently received wide attention for their enhanced real-time applicability and expanded threat range. However, most of the UAP research concentrates on the image domain, and less on speech. In this paper, we propose a staged perturbation generation method that constructs CommanderUAP, which achieves a high success rate of universal adversarial attack against speech recognition models. Moreover, we apply some methods from model training to improve the generalization in attack and we control the imperceptibility of the perturbation in both time and frequency domains. In specific scenarios, CommanderUAP can also transfer attack some commercial speech recognition APIs. Jinxiao Zhao, Lei Ju 0001 |
Cybersecur. | 5 |
| 2024 | HMC-FHE: A Heterogeneous Near Data Processing Framework for Homomorphic EncryptionabstractFully homomorphic encryption (FHE) offers a promising solution to ensure data privacy by enabling computations directly on encrypted data. However, its notorious performance degradation severely limits the practical application, due to the explosion of both the ciphertext volume and computation. In this article, leveraging the diversity of computing power and memory bandwidth requirements of FHE operations, we present HMC-FHE, a robust acceleration framework that combines both GPU and hybrid memory cube (HMC) processing engines to accelerate FHE applications cooperatively. HMC-FHE incorporates four key hardware/software co-design techniques: 1) a fine-grained kernel offloading mechanism to efficiently offload FHE operations to relevant processing engines; 2) a ciphertext partitioning scheme to minimize data transfer across decentralized HMC processing engines; 3) an FHE operation pipeline scheme to facilitate pipelined execution between GPU and HMC engines; and 4) a kernel tuning scheme to guarantee the parallelism of GPU and HMC engines. We demonstrate that the GPU-HMC architecture with proper resource management serves as a promising acceleration scheme for memory-intensive FHE operations. Compared with the state-of-the-art GPU-based acceleration scheme, the proposed framework achieves up to$2.65\times $performance gains and reduces$1.81\times $energy consumption with the same peak computation capacity. Zhining Cao, Zhaoyan Shen, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Static Scheduling of Weight Programming for DNN Acceleration with Resource Constrained PIMabstractMost existing architectural studies on ReRAM-based processing-in-memory (PIM) DNN accelerators assume that all weights of the DNN can be mapped to the crossbar at once. However, these studies are over-idealized. ReRAM crossbar resources for calculation are limited because of technological limitations, so multiple weight mapping procedures are required during the inference process. In this article, we propose a static scheduling framework which generates the mapping between DNN weights and ReRAM cells with minimum runtime weight programming cost. We first build a ReRAM crossbar programming latency model by simultaneously considering the DNN weight patterns, ReRAM programming operations, and PIM architecture characteristics. Then, the model is used in the searching process to obtain an optimized weight-to-OU mapping table with minimum online programming latency. Finally, an OU scheduler is used to coordinate the activation sequences of OUs in the crossbars to perform the inference computation correctly. Evaluation results show the proposed framework significantly reduces the weight programming overhead and the overall inference latency for various DNN models with different input datasets. Xin Gao 0012, Yuhao Zhang 0006, Zhaoyan Shen, Lei Ju 0001 |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2023 | Accelerating DNN Inference with Heterogeneous Multi-DPU EnginesabstractThe Deep Learning Processor (DPU) programmable engine released by the official Xilinx Vitis AI toolchain has become one of the commercial off-the-shelf (COTS) solutions for Convolutional Neural Networks (CNNs) inference on Xilinx FPGAs. While modern FPGA devices generally have enough hardware resources to accommodate multi-DPUs simultaneously, the Xilinx toolchain currently only supports the deployment of multiple homogeneous DPUs engines that running independent inference tasks (task-level parallelism). In this work, we demonstrate that deployment of multiple heterogeneous DPU engines makes better resource efficiency for a given FPGA device. Moreover, we show that pipelined execution of a CNN inference task over heterogeneous multi-DPU engines may further improve overall inference throughput with carefully designed CNN layers-to-DPU mapping and scheduling. Finally, for a given CNN model and an FPGA device, we propose a comprehensive framework that automatically determines the optimal heterogeneous DPU deployment, and adaptively chooses the execution scheme between task-level and pipelined parallelism. Compared with the state-of-the-art solution with homogeneous multi-DPU engines and network-level parallelism, the proposed framework shows an average improvement of 13% (up-to 19%) and 6.6% (up-to 10%) on the Xilinx Zynq UltraScale+ MPSoC ZCU104 and ZCU102 platforms, respectively. Zelin Du, Wei Zhang 0173, Zimeng Zhou, Zili Shao, Lei Ju 0001 |
DAC | 5 |
| 2023 | Work or Sleep: Freshness-Aware Energy Scheduling for Wireless Powered Communication Networks with Interference ConsiderationabstractThis paper explores how to schedule energy to optimize the information freshness in wireless powered communication networks (WPCNs) when considering channel interference among adjacent sensor nodes. We introduce Age of Information (AoI) to quantitatively evaluate the information freshness and formulate the AoI optimization problem. Unlike prior works focusing on system optimization for WPCNs while ignoring channel interference in energy transfer, this work reveals situations where channel interference among adjacent sensor nodes cannot be neglected and explores optimizing information freshness with interference consideration. To take the phenomena into account, we propose an energy scheduling solution to detect the channel interference and then judiciously determine the energy and time allocation for individual sensor nodes to improve the AoI performance as well as the system throughput. We implement a multi-node WPCN testbed to validate the functional correctness of the proposed solution, and extensive experiments have demonstrated the effectiveness of the proposed solution. The experimental results show that the proposed solution can reduce the average AoI by 54.7% and the average throughput by 49.8% on average compared to the state-of-the-art solutions. Lei Ju 0001, Chun Jason Xue, Mingliang Zhou 0001, Wei Zhang 0173, Zimeng Zhou |
DAC | 2 |
| 2023 | FxHENN: FPGA-based acceleration framework for homomorphic encrypted CNN inferenceabstractFully homomorphic encryption (FHE) is a promising data privacy solution for machine learning, which allows the inference to be performed with encrypted data. However, it typically leads to 5-6 orders of magnitude higher computation and storage overhead. This paper proposes the first full-fledged FPGA acceleration framework for FHE-based convolution neural network (HE-CNN) inference. We then design parameterized HE operation modules with intra- and inter- HE-CNN layer resource management based on FPGA high-level synthesis (HLS) design flow. With sophisticated resource and performance modeling of the HE operation modules, the proposed FxHENN framework automatically performs design space exploration to determine the optimized resource provisioning and generates the accelerator circuit for a given HE-CNN model on a target FPGA device. Compared with the state-of-the-art CPU-based HE-CNN inference solution, FxHENN achieves up to 13.49X speedup of inference latency, and 1187.12X energy efficiency. Meanwhile, given this is the first attempt in the literature on FPGA acceleration of fullfledged non-interactive HE-CNN inference, our results obtained on low-power FPGA devices demonstrate HE-CNN inference for edge and embedded computing is practical. Yilan Zhu, Lei Ju 0001, Shanqing Guo |
HPCA | 3 |
| 2023 | Runtime Row/Column Activation Pruning for ReRAM-based Processing-in-Memory DNN AcceleratorsabstractResistive random access memory (ReRAM)-based processing-in-memory (PIM) DNN accelerators have shown great potential in improving model efficiency and saving energy. To further improve memory and computation efficiency, model weight sparsity has been widely explored in ReRAM-based accelerator designs. However, these optimized accelerators rarely touched the model activation sparsity. In this paper, we observe that there exist plenty sparse rows/columns in the DNN model activation matrix which have negligible effect on accuracy, termed as insensitive rows/columns. Pruning them has little impact on model accuracy but would have significant potential to improve the performance and energy efficiency of DNN accelerators. Therefore, we propose a new ReRAM-based PIM accelerator, named as RapPIM, to take advantage of the model activation sparsity. In RapPIM, we first propose an insensitive activation rows/columns pruning method to search and prune the insensitive rows/columns. Then, we present an activation low-bits skipping strategy and a forward propagation delay hiding strategy to further improve model performance and minimize the latency of activation pruning on forward propagation. Our evaluations with several well-known DNN models show that the RapPIM achieves up to 2.40× speedup and 44.82% power reduction compared with the state-of-the-art ReRAM-based accelerator. Xikun Jiang, Zhaoyan Shen, Siqing Sun, Ping Yin, Zhiping Jia, Lei Ju 0001, Zhiyong Zhang 0006, Dongxiao Yu |
ICCAD | 6 |
| 2023 | Towards the universal defense for query-based audio adversarial attacks on speech recognition systemabstractAbstract Recently, studies show that deep learning-based automatic speech recognition (ASR) systems are vulnerable to adversarial examples (AEs), which add a small amount of noise to the original audio examples. These AE attacks pose new challenges to deep learning security and have raised significant concerns about deploying ASR systems and devices. The existing defense methods are either limited in application or only defend on results, but not on process. In this work, we propose a novel method to infer the adversary intent and discover audio adversarial examples based on the AEs generation process. The insight of this method is based on the observation: many existing audio AE attacks utilize query-based methods, which means the adversary must send continuous and similar queries to target ASR models during the audio AE generation process. Inspired by this observation, We propose a memory mechanism by adopting audio fingerprint technology to analyze the similarity of the current query with a certain length of memory query. Thus, we can identify when a sequence of queries appears to be suspectable to generate audio AEs. Through extensive evaluation on four state-of-the-art audio AE attacks, we demonstrate that on average our defense identify the adversary’s intent with over $$90\%$$ 90 % accuracy. With careful regard for robustness evaluations, we also analyze our proposed defense and its strength to withstand two adaptive attacks. Finally, our scheme is available out-of-the-box and directly compatible with any ensemble of ASR defense models to uncover audio AE attacks effectively without model retraining. Lei Ju 0001 |
Cybersecur. | 4 |
| 2023 | Towards the transferable audio adversarial attack via ensemble methodsabstractAbstract In recent years, deep learning (DL) models have achieved significant progress in many domains, such as autonomous driving, facial recognition, and speech recognition. However, the vulnerability of deep learning models to adversarial attacks has raised serious concerns in the community because of their insufficient robustness and generalization. Also, transferable attacks have become a prominent method for black-box attacks. In this work, we explore the potential factors that impact adversarial examples (AEs) transferability in DL-based speech recognition. We also discuss the vulnerability of different DL systems and the irregular nature of decision boundaries. Our results show a remarkable difference in the transferability of AEs between speech and images, with the data relevance being low in images but opposite in speech recognition. Motivated by dropout-based ensemble approaches, we propose random gradient ensembles and dynamic gradient-weighted ensembles, and we evaluate the impact of ensembles on the transferability of AEs. The results show that the AEs created by both approaches are valid for transfer to the black box API. Lei Ju 0001 |
Cybersecur. | 4 |
| 2023 | PRAP-PIM: A weight pattern reusing aware pruning method for ReRAM-based PIM DNN acceleratorsabstractResistive Random-Access Memory (ReRAM) based Processing-in-Memory (PIM) frameworks are proposed to accelerate the working process of DNN models by eliminating the data movement between the computing and memory units. To further mitigate the space and energy consumption, DNN model weight sparsity and weight pattern repetition are exploited to optimize these ReRAM-based accelerators. However, most of these works only focus on one aspect of this software/hardware co-design framework and optimize them individually, which makes the design far from optimal. In this paper, we propose PRAP-PIM, which jointly exploits the weight sparsity and weight pattern repetition by using a weight pattern reusing aware pruning method. By relaxing the weight pattern reusing precondition, we propose a similarity-based weight pattern reusing method that can achieve a higher weight pattern reusing ratio. Experimental results show that PRAP-PIM achieves 1.64× performance improvement and 1.51× energy efficiency improvement in popular deep learning benchmarks, compared with the state-of-the-art ReRAM-based DNN accelerators. Zhaoyan Shen, Jinhao Wu, Xikun Jiang, Yuhao Zhang 0006, Lei Ju 0001, Zhiping Jia |
High Confid. Comput. | 5 |
| 2023 | ChainKV: A Semantics-Aware Key-Value Store for Ethereum SystemabstractThe Log-Structure Merged tree (LSM-tree) based key-value (KV) store has been widely adopted as the storage engine for blockchain systems, such as Ethereum, in which blockchain data are uniformly transformed into randomly distributed KV items for persistence. However, blockchain semantics are ignored during this process, making the blockchain storage suffer from heavy read/write amplification problems. Moreover, as the Ethereum network scales up, tremendous data further exacerbates its storage burden. Until now, most studies have focused on sharding, data archiving, decentralized distributed storage, etc., to mitigate the burden of the storage layer. However, the incompatibility between Ethereum semantics and the characteristics of the storage engine is ignored. In this paper, we present ChainKV, a new semantics-aware storage paradigm to improve the storage management performance for the Ethereum system. Firstly, based on Ethereum blockchain semantics, ChainKV separately stores different types of data in multiple storage zones in the KV store to mitigate the read/write amplification problem. Secondly, following the mechanism of the verification process in the authenticated data structure (ADS), a new ADS data transformer is proposed to exploit the data locality when persisting ADS. Moreover, a new space gaming caching policy is adopted to coordinate the cache space management for two independent storage zones. Finally, we propose an optional lightweight node crash recovery mechanism to eliminate functional redundancy between the Ethereum protocol and the storage engine. The experimental results indicate that ChainKV outperforms the prior Ethereum systems by up to 1.99× and 4.20× for synchronization and query operations, respectively Bingzhe Li, Xiaojun Cai, Zhiping Jia, Lei Ju 0001, Zili Shao, Zhaoyan Shen |
Proc. ACM Manag. Data | 5 |
| 2023 | A Comprehensive Memory Management Framework for CPU-FPGA Heterogenous SoCsabstractEfficient utilization of restrained memory resources is of paramount importance in CPU-FPGA heterogeneous multiprocessor system-on-chip (HMPSoC)-based system design for memory-intensive applications. State-of-the-art high level synthesis (HLS) tools rely on the system programmers to manually determine the data placement within the complex memory hierarchy. Different data placement policies may lead to different system performance, and finding an optimal data placement policy is a nontrivial problem. For instance, we show counter-intuitive results that traditional frequency and locality-based data placement strategy designed for CPU architecture leads to nonoptimal system performance in CPU-FPGA HMPSoCs. In this work, we first propose an automatic data placement framework for field programmable gate array (FPGA) kernels to determine whether each array object should be accessed via the on-chip BRAM, shared CPU L2-cache, or DDR memory to achieve the optimal performance. Moreover, we find that when the CPU kernel and the FPGA kernel are executed in parallel, memory contentions may degrade the performance and the optimal data placement policy designed for the FPGA kernel alone will not achieve the optimal overall system performance. In this article, we proposed to use cache partitioning to alleviate the impact brought by memory contentions. We extend the framework designed for FPGA by adding the cross-layer memory contentions analysis to automatically generate an optimal data placement policy and cache partitioning mechanism for the parallel executing kernels. The proposed data placement framework can be seamlessly integrated with the commercial Vivado HLS. The experimental results on the Zedboard platform show an average$1.5\times $performance speedup for FPGA kernels compared with a greedy-based allocation strategy. When FPGA kernels and CPU kernels are executed in parallel, the FPGA kernel and the CPU kernel have a performance speedup of$1.62\times $and$1.10\times $on average, respectively. Zelin Du, Qianling Zhang, Mao Lin, Shiqing Li, Xin Li 0137, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Adaptive Task-Based Intermittent Computing System With Parallel State BackupabstractEnergy harvesting promises to power billions of Internet of Things devices without being restricted by battery life. Since the energy harvester generally outputs weak and unstable energy, the system may suffer frequent and unpredictable power failures, thus falling into cyclically reboots without forward progress. The task-based intermittent computing system which periodically backs up system states into nonvolatile memory (NVM) is proposed to solve the nonprogress problem, with the nontrivial cost of frequent backups. How to reduce the backup overhead becomes a major research problem for intermittent computing. This article, for the first time, proposes to parallelize state backup and program execution with asynchronous direct memory access (DMA) to hide the backup latency into the program’s execution. But, straightforwardly executing the state backup and the program in parallel may cause an inconsistent system state. In specific, the system state may be modified by the program during backup, and therefore may be backed up incorrectly and further cause the system to deliver an incorrect computation result. We make a deep analysis on the system behavior and observe that, although the system state may be backed up incorrectly, the incorrect backup will be covered by the subsequent correct backups soon as the backup operations are performed frequently. In addition, only a small part of variables among all the program states may cause incorrect computation result. So, in this article, we aggressively allow incorrect backups to occur and propose a backup error detection method and a fault-tolerant backup management to guarantee the correctness of the system’s execution. To augment the parallel backup method, an adaptive execution method is further proposed to reduce the number of backups and balance the ratio between task execution time and backup latency. We design a run-time system to implement the proposed approach, and experimental results conducted on an STM32F7-based platform show that the proposed method can achieve a$2.6\times $average speedup. Wei Zhang 0173, Qianling Zhang, Mingsong Lv, Songran Liu, Zimeng Zhou, Qiulin Chen, Nan Guan, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | Optimizing Worst Case Data Freshness in RF-Powered Networked Embedded SystemsabstractMaintaining real-time data freshness plays a critical role in ensuring system correctness and optimizing the system performance in networked embedded systems (NESs). To quantitatively measure the freshness of the collected real-time data, the concept of Age of Information (AoI) has been extensively studied in recent years. This article explores how to minimize the worst case AoI of real-time data in radio-frequency (RF)-powered NESs. In such systems, one hybrid access point (HAP) transfers wireless power to a set of distributed sensor nodes, and in the meantime, receives the information from these sensor nodes. We utilize the metric of AoI to measure the data freshness and present a comprehensive analysis of the worst case AoI of the real-time data in the target system. Based on the analysis, an optimal energy schedule solution is designed to judiciously determine individual sensor nodes’ energy and time allocation to minimize the worst case AoI. Considering the varying importance of different information and sensor nodes in the target system, we further propose the optimal time and energy allocation scheme for minimizing the weighted worst case AoI. A multinode RF-powered NES testbed is implemented to validate the functional correctness of our solutions. The results show that our solutions significantly outperform the state-of-the-art solutions, reducing the worst case AoI and weighted worst case AoI by 69.3% and 75.1% on average, respectively. Zimeng Zhou, Chenchen Fu, Chun Jason Xue, Song Han 0002, Wei Zhang 0173, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Multi-Party Sequential Data Publishing Under Differential PrivacyabstractGiven a set of local sequential datasets held by multiple parties, we study the problem of publishing a synthetic dataset that preserves approximate sequentiality information of the integrated dataset while satisfying differential privacy for each local dataset. The existing solutions for publishing differentially private sequential data in the centralized setting mostly adopt tree-based approaches. Such approaches rely on different tree structures that encode sequential data's statistical information. The construction of a tree structure is normally done by recursively splitting nodes whose noisyscores(e.g., entropy or count) are larger than a given threshold. However, extending similar ideas to the multi-party setting is challenging. First, the comparison between noisy scores and a given threshold needs to be done in a distributed manner without letting the parties know the noisy scores, while satisfying differential privacy for each local dataset. Second, in the multi-party setting the large number of node splitting decisions incurs prohibitive computation costs. In addressing the above challenges, we presentDPST, a distributed prediction suffix tree construction solution. In DPST, we first introduce a novel node splitting decision method that calculates the comparison result under encryption with substantially improved efficiency. Then we present a novel batch-based tree construction approach to reduce computation costs. In order to achieve high parallel performance without incurring any extra communication cost, we introduce theconjunctionandslidemethods to ensure that each batch contains a stable number of carefully arrangeddecision tasks. To further reduce communication and computation costs, we propose a prefix-based pre-pruning method to reduce the number of nodes that need to be judged whether to split by an interactive protocol. Extensive experiments on real datasets demonstrate that our DPST solution offers desirable data utility with low computation and communication costs. Peng Tang 0002, Rui Chen 0012, Sen Su, Shanqing Guo, Lei Ju 0001, Gaoyuan Liu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Precise and scalable shared cache contention analysis for WCET estimationabstractWorst-Case Execution Time (WCET) analysis for real-time tasks must precisely predict cache hit/miss of memory accesses. While bringing great performance benefits, multi-core processors significantly complicate the cache analysis problem due to the shared cache contentions among different cores. Existing methods pessimistically consider that memory references of parallel executing tasks will contend with each other as long as they are mapped to the same cache line. However, in reality, numerous shared cache contentions are mutually exclusive, due to the partial orders among the programs executed in parallel. The presence of shared cache contentions greatly exacerbates the computational complexity of the WCET computation, as finding the longest path needs exploring an exponentially large partial ordering space. In this paper, we propose a quantitative method with O(n2) time complexity to precisely estimate the worst-case extra execution time (WCEET) caused by shared cache contentions. The proposed method can be easily integrated into the abstract-interpretation based WCET estimation framework. Experiments with MRTC benchmarks show that our method can averagely tighten the WCET estimation by 13% without sacrificing the analysis efficiency. Wei Zhang 0173, Mingsong Lv, Wanli Chang 0001, Lei Ju 0001 |
DAC | 4 |
| 2022 | coxHE: A software-hardware co-design framework for FPGA acceleration of homomorphic computationabstractData privacy becomes a crucial concern in the AI and big data era. Fully homomorphic encryption (FHE) is a promising data privacy protection technique where the entire computation is performed on encrypted data. However, the dramatic increase of the computation workload restrains the usage of FHE for the real-world applications. In this paper, we propose an FPFA accelerator design framework for CKKS-based HE. While the KeySwitch operations are the primary performance bottleneck of FHE computation, we propose a low latency design of KeySwitch module with reduced intra-operation data dependency. Compared with the state-of-the-art FPGA based key-switch implementation that is based on Verilog, the proposed high-level synthesis (HLS) based design reduces the operation latency by 40%. Furthermore, we propose an automated design space exploration framework which generates optimal encryption parameters and accelerators for a given application kernel and the target FPGA device. Experimental results for a set of real HE application kernels on different FPGA devices show that our HLS-based flexible design framework produces substantially better accelerator design compared with a fixed-parameter HE accelerator in terms of security, approximation error, and overall performance. Mingqin Han, Yilan Zhu, Qian Lou, Zimeng Zhou, Shanqing Guo, Lei Ju 0001 |
DATE | 6 |
| 2022 | Re-LSM: A ReRAM-Based Processing-in-Memory Framework for LSM-Based Key-Value StoreabstractLog-structured merge (LSM) tree based key-value (KV) stores organize writes into hierarchical batches for high-speed writing. However, the notorious compaction process of LSM-tree severely hurts system performance. It not only involves huge I/O operations but also consumes tremendous computation and memory resources. In this paper, first we find that when compaction happens in the high levels (i.e., L0, L1) of the LSM-tree, it may saturate all system computation and memory resources, and eventually stall the whole system. Based on this observation, we present Re-LSM, a ReRAM-based Processing-in-Memory (PIM) framework for LSM-based Key-Value Store. Specifically, in Re-LSM, we propose to offload certain computation and memory-intensive tasks in the high levels of the LSM-tree to the ReRAM-based PIM space. A high parallel ReRAM compaction accelerator is designed by decomposing the three-phased compaction into basic logic operating units. Evaluation results based on db_bench and YCSB show that Re-LSM achieves 2.2× improvement on the throughput of random writes compared to RocksDB, and the ReRAM-based compaction accelerator speedups the CPU-based implementation by 64.3× and saves 25.5× energy. Zhaoyan Shen, Yiheng Tong, Zhiping Jia, Lei Ju 0001, Jiezhi Chen, Bingzhe Li |
ICCAD | 5 |
| 2021 | Differentially Private Publication of Multi-Party Sequential DataabstractGiven a set of local sequential datasets held by multiple parties, we study the problem of publishing a synthetic dataset that preserves approximate sequentiality information of the integrated dataset while satisfying differential privacy for each local dataset. The existing solutions for publishing differentially private sequential data in the centralized setting mostly adopt tree-based approaches. Such approaches rely on different tree structures that encode sequential data's statistical information. The construction of a tree structure is normally done by recursively splitting nodes whose noisy scores (e.g., entropy or count) are larger than a given threshold. However, extending similar ideas to the multi-party setting is challenging. First, the comparison between noisy scores and a given threshold needs to be done in a distributed manner without letting the parties know the noisy scores, while satisfying differential privacy for each local dataset. Second, in the multi-party setting the large number of node splitting decisions incurs prohibitive computation costs. In addressing the above challenges, we present DPST, a distributed prediction suffix tree construction solution. In DPST, we first introduce a novel node splitting decision method that calculates the comparison result under encryption with substantially improved efficiency. Then we present a novel batch-based tree construction approach to reduce the computation costs. In order to achieve high parallel performance without incurring any extra communication cost, we introduce the conjunction and slide methods to ensure that each batch contains a stable number of carefully arranged decision tasks. Extensive experiments on real datasets demonstrate that our DPST solution offers desirable data utility with low computation and communication costs. Peng Tang 0002, Rui Chen 0012, Sen Su, Shanqing Guo, Lei Ju 0001, Gaoyuan Liu |
ICDE | 5 |
| 2020 | Scope-Aware Useful Cache Block Calculation for Cache-Related Pre-Emption Delay Analysis With Set-Associative Data CachesabstractTiming analysis of real-time systems must consider cache-related pre-emption delay (CRPD) costs when pre-emptive scheduling is used. While most previous work on CRPD analysis only considers instruction caches, the CRPD incurred on data caches is actually more significant. The state-of-the-art CRPD analysis methods are based on useful cache block (UCB) calculation. Unfortunately, as shown in this article, directly extending the existing UCB calculation techniques from instruction caches to data caches will lead to both unsoundness and significant imprecision. To solve these problems, we develop a new UCB calculation technique for data caches, which redefines the analysis unit (to address the unsoundness in the existing method) and precisely captures the dynamic cache access behavior by taking the temporal scopes of memory blocks into consideration. The experimental results show that our new technique yields substantially tighter CRPD estimations comparing with the state-of-the-art. Wei Zhang 0173, Nan Guan, Lei Ju 0001, Yue Tang 0001, Weichen Liu 0001, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Applying Multiple Level Cell to Non-volatile FPGAsabstractStatic random access memory– (SRAM) based field programmable gate arrays (FPGAs) are currently facing challenges of limited capacity and high leakage power. To solve this problem, non-volatile memory (NVM) is proposed as the alternative to build non-volatile FPGAs (NVFPGAs). Even though the feasibility of NVFPGA has been confirmed, the utilization of multiple level cells (MLCs) has not been fully exploited yet. In this article, we study architecture of MLC-based NVFPGAs, and propose five cluster structures. To give detailed comparisons and extensive discussions, we conduct experiments for area, performance and leakage power evaluation. Based on explorations of the characteristics of MLC-based NVFPGAs, we further present MLC-aware timing-driven packing method to improve delay. In critical paths, our proposed method reduces the overhead of the additional delay in slow MLC cells. Experiments show that, compared to SRAM-based FPGAs, the proposed architecture with the proposed CAD flow can reduce the area, critical path delay and leakage power by 31%, 10%, and 95%, respectively. Mengying Zhao, Lei Ju 0001, Zhiping Jia, Jingtong Hu, Chun Jason Xue |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2019 | Automatic data placement for CPU-FPGA heterogeneous multiprocessor System-on-ChipsabstractEfficient utilization of restrained memory resources is of paramount importance in CPU-FPGA heterogeneous multiprocessor system-on-chip (HMPSoC) based system design for memory-intensive applications. State-of-the-art high level synthesis (HLS) tools rely on the system programmers to manually determine the data placement within the complex memory hierarchy. In this paper, we propose an automatic data placement framework which can be seamlessly integrated with the commercial Vivado HLS. We first show counter-intuitive results that traditional frequency and locality based data placement strategy designed for CPU architecture leads to non-optimal system performance in CPU-FPGA HMPSoCs. Built on top of our memory latency analysis model, the proposed integer linear programming (ILP) based framework determines whether each array object should be access via the on-chip BRAM, shared CPU L2-cache, or DDR memory directly. Experimental results on the Zedboard platform show an average 1.39X performance speedup compared with a greedy-based allocation strategy. Shiqing Li, Yixun Wei, Lei Ju 0001 |
DATE | 3 |
| 2019 | EMC: Energy-Aware Morphable Cache Design for Non-Volatile ProcessorsabstractWearable, implantable and Internet of Things devices are attracting increasing attention from both research and industry fields. Energy harvesting is a promising alternative of battery to power these embedded systems. However, the intrinsic instability of energy harvesting systems leads to potential frequent power interruptions. In traditional volatile processor, all the status will be lost at power failures and the program needs to re-start after power resumes. In order to survive the power failures and enable accumulative execution, non-volatile processor (NVP) is proposed to back up volatile information before power depletion and recover the system status after power resumes. Non-volatile memory (NVM) is typically attached for cache and main memory backup. There are researches working on optimization of the backup. However, little of them involve multiple level cell (MLC) NVM. In this work, we first discuss the benefit of applying MLC NVM for cache backup and the architecture of morphable hybrid cache, and then propose a three-stage energy-aware cache management strategy to improve the system performance and energy utilization while guaranteeing successful backups. Backup-aware cache replacement policies are also developed for backup optimization. Evaluation shows that the proposed EMC scheme can achieve 10.6 percent performance improvement and simultaneous 25.2 percent energy reduction when compared with the single level cell (SLC) based hybrid cache. Weining Song, Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Zhiping Jia |
IEEE Trans. Computers | 4 |
| 2018 | Set variation-aware shared LLC management for CPU-GPU heterogeneous architectureabstractHeterogeneous CPU-GPU multiprocessor systems-on-chip (HMPSoC) becomes a popular architecture choice for high performance embedded systems, where shared last-level cache (LLC) management becomes a critical design consideration. We observe that within a sampling period, CPU and GPU may have distinct access behaviors over various LLC sets. In this work, we propose a light-weighted and fined-grained cache management policy to cope with the CPU-GPU access behavior variation among cache sets. In particular, CPU and GPU requests are prioritized disparately in each LLC set during cache block insertion and promotion, based on the per-core utility behaviors and a per-set CPU-GPU miss counter. Experimental results show that our LLC management scheme outperforms the two state-of-the-art schemes TAP-RRIP and LSP by 12.6% and 10.01%, respectively. Zhaoying Li 0004, Lei Ju 0001, Hongjun Dai, Mengying Zhao, Zhiping Jia |
DATE | 2 |
| 2018 | Position prediction system based on spatio-temporal regularity of object mobility
Xin Li 0137, Chongsheng Yu, Lei Ju 0001, Lei Dou, Yuqing Sun 0001 |
Inf. Syst. | 3 |
| 2018 | NVM-Based FPGA Block RAM With Adaptive SLC-MLC ConversionabstractThe capacity of SRAM-based FPGA block RAM (BRAM) is restrained by the low density and high leakage power of the current CMOS technology. In this paper, we propose a nonvolatile memory (NVM)-based BRAM architecture which enables flexible conversions between single-level cell (SLC) and multilevel cell (MLC) states. We show that despite the high per-access latency and power consumption, MLC-based BRAM blocks reduce the routing cost between logic units and on-chip data storages, which potentially leads to a smaller critical path delay and power consumption. Therefore, we propose an NVM BRAM architecture and an EDA framework which adaptively packs data into SLC- or MLC-state BRAMs during FPGA design flow in order to achieve better system performance. This paper illustrates that a simple memory device replacement from SRAM to NVM leads to nonoptimal system performance. On the other hand, compared with operating all NVM BRAM blocks in the SLC state with better per-access latency and power consumption, the proposed hybrid SLC-MLC architecture and design flow improves the critical path delay by 18.51%, with a system power reduction of 25.83% at the same time. Moreover, compared with the traditional “fast” SRAM-based BRAM blocks under the same BRAM area constraint, our hybrid NVM BRAM architecture improves the critical path delay by 8.55% on average, with an average system power reduction of 54.34% at the same time. Lei Ju 0001, Xiaojin Sui, Shiqing Li, Mengying Zhao, Chun Jason Xue, Jingtong Hu, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Analyzing Data Cache Related Preemption Delay With Multiple PreemptionsabstractTiming analysis of real-time tasks under preemptive scheduling must take cache-related preemption delay (CRPD) into account. Typically, a task may be preempted more than once during the execution in each period. To bound the total CRPD of${k}$preemptions, existing CRPD analysis techniques estimate the CRPD at each program point, and use the sum of the${k}$-largest CRPD among all program points as the total CRPD upper bound. In this paper, we disclose that the above-mentioned approach, although works well for instruction caches, leads to significant overestimation when dealing with data caches. This is because on data caches, the CRPD of preemptions at different program points may have correlations, and the total CRPD of multiple preemptions is in general smaller than the simple sum of the worst-case CRPD of each preemption. To address this problem, we propose a new technique to efficiently explore the correlation among the CRPD of different preemptions, and thus more precisely calculate the total CRPD. Experiments with benchmark programs show that the proposed technique leads to substantially tighter total CRPD estimation with multiple preemptions comparing with the state-of-the-art. Wei Zhang 0173, Nan Guan, Lei Ju 0001, Weichen Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Shared Last-Level Cache Management and Memory Scheduling for GPGPUs with Hybrid Main MemoryabstractMemory intensive workloads become increasingly popular on general purpose graphics processing units (GPGPUs), and impose great challenges on the GPGPU memory subsystem design. On the other hand, with the recent development of non-volatile memory (NVM) technologies, hybrid memory combining both DRAM and NVM achieves high performance, low power, and high density simultaneously, which provides a promising main memory design for GPGPUs. In this article, we explore the shared last-level cache management for GPGPUs with consideration of the underlying hybrid main memory. To improve the overall memory subsystem performance, we exploit the characteristics of both the asymmetric read/write latency of the hybrid main memory architecture, as well as the memory coalescing feature of GPGPUs. In particular, to reduce the average cost of L2 cache misses, we prioritize cache blocks from DRAM or NVM based on observations that operations to NVM part of main memory have a large impact on the system performance. Furthermore, the cache management scheme also integrates the GPU memory coalescing and cache bypassing techniques to improve the overall system performance. To minimize the impact of memory divergence behaviors among simultaneously executed groups of threads, we propose a hybrid main memory and warp aware memory scheduling mechanism for GPGPUs. Experimental results show that in the context of a hybrid main memory system, our proposed L2 cache management policy and memory scheduling mechanism improve performance by 15.69% on average for memory intensive benchmarks, whereas the maximum gain can be up to 29% and achieve an average memory subsystem energy reduction of 21.27%. Chuanqi Zang, Lei Ju 0001, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | Cooperative DVFS for energy-efficient HEVC decoding on embedded CPU-GPU architectureabstractThe next generation video coding standard High Efficiency Video Coding (HEVC) provides better compression rate for high resolution videos, at the cost of substantially higher computational complexity. While some latest off-the-shelf consumer electronics support HEVC via ASIC solutions, software implementation of real-time HEVC remains an open challenge for resource-constraint embedded systems. In this work, we present an HEVC decoder design on a low-power embedded heterogeneous multiprocessor System-on-Chip (HMPSoC) with CPU and GPU. Our analysis shows that the massive parallel architecture of GPU leads to a relatively smooth fluctuation on the processing time between video frames. Moreover, the dynamic workload of each frame has a monotonic correlation with a particular coding parameter that can be obtained at decoding time. Based on these observations, we propose an application-specific userspace CPU-GPU DVFS scheme which effectively saves the energy consumption for HEVC decoding. Furthermore, given our accurate workload prediction, only a small frame buffer is required to ensure real-time video decoding. Fan Gong, Lei Ju 0001, Deshan Zhang, Mengying Zhao, Zhiping Jia |
DAC | 2 |
| 2017 | Maximizing Forward Progress with Cache-aware Backup for Self-powered Non-volatile ProcessorsabstractEnergy harvesting is replacing battery to power embedded systems such as Internet of Things and wearable devices. Unstable energy supply brings challenges to energy harvesting powered system, resulting in frequent interruptions. Non-volatile processor is proposed to back up volatile logics before energy depletion and recover the system status after energy resumes. The backup efficiency of memory content significantly affects program performance. There are existing researches focusing on backup optimizations, but they did not fully consider cache behaviors. In this paper, we introduce cache persistence analysis into memory backup for self-powered non-volatile processors. The evaluation shows that the proposed cache-aware backup delivers on average 45.6% improvement in forward progress, achieving 40.2% and 12.7% higher system performance compared with instant and cache-unaware backup. Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Zhiping Jia |
DAC | 3 |
| 2017 | Shared last-level cache management for GPGPUs with hybrid main memoryabstractMemory intensive workloads become increasingly popular on general purpose graphics processing units (GPGPUs), and impose great challenges on the GPGPU memory subsystem design. On the other hand, with the recent development of non-volatile memory (NVM) technologies, hybrid memory combining both DRAM and NVM achieves high performance, low power and high density simultaneously, which provides a promising main memory design for GPGPUs. In this work, we explore the shared last-level cache management for GPGPUs with consideration of the underlying hybrid main memory. In order to improve the overall memory subsystem performance, we exploit the characteristics of both the asymmetric read/write latency of the hybrid main memory architecture, as well as the memory coalescing feature of GPGPU. In particular, to reduce the average cost of L2 cache misses, we prioritize cache blocks from DRAM or NVM based on observation that operations to NVM part of main memory have large impact on the system performance. Furthermore, the cache management scheme also integrates the GPU memory coalescing and cache bypassing techniques to improve the overall cache hit ratio. Experimental results show that in the context of a hybrid main memory system, our proposed L2 cache management policy improves performance against the traditional LRU policy and a state-of-the-art GPU cache strategy EABP [20] by up to 27.76% and 14%, respectively. Xiaojun Cai, Lei Ju 0001, Chuanqi Zang, Mengying Zhao, Zhiping Jia |
DATE | 3 |
| 2017 | Design Exploration for Multiple Level Cell Based Non-Volatile FPGAsabstractStatic random access memory (SRAM) based field programmable gate arrays (FPGAs) are currently facing challenges of limited capacity and high leakage power. To solve this problem, non-volatile memory (NVM) is proposed as the alternative to build non-volatile FPGAs (NVFPGAs). Even though the feasibility of NVFPGA has been confirmed, the utilization of multiple level cells (MLC) has not been fully exploited yet. In this paper, we study architecture of MLC based NVFPGAs, and propose five cluster structures, as well as the corresponding working mode supported by MLC based clusters. To give detailed comparisons and extensive discussions, we conduct experiments for area, performance and leakage power evaluation. Experiments show that, compared to SRAM based FPGAs, the proposed architecture can reduce the area, latency and leakage power by 32.66% and 7.45%, and 96.13%, respectively. Mengying Zhao, Lei Ju 0001, Zhiping Jia, Chun Jason Xue, Jingtong Hu |
ICCD | 3 |
| 2017 | Unified nvTCAM and sTCAM architecture for improving packet matching performanceabstractSoftware-Defined Networking (SDN) allows controlling applications to install fine-grained forwarding policies in the underlying switches. Ternary Content Addressable Memory (TCAM) enables fast lookups in hardware switches with flexible wildcard rule patterns. However, the performance of packet processing is severely constrained by the capacity of TCAM, which aggravates the processing burden and latency issues. In this paper, we propose a hybrid TCAM architecture which consists of NVM-based TCAM (nvTCAM) and SRAM-based TCAM (sTCAM), utilizing nvTCAM to cache the most popular rules to improve cache-hit-ratio while relying on a very small-size sTCAM to handle cache-miss traffic to effectively decrease update latency. Considering the special rule dependency, we present an efficient Rule Migration Replacement (RMR) policy to make full utilization of both nvTCAM and sTCAM to obtain better performance. Experimental results show that the proposed architecture outperforms current TCAM architectures. Xianzhong Ding, Zhiyong Zhang 0006, Zhiping Jia, Lei Ju 0001, Mengying Zhao, Huawei Huang |
LCTES | 4 |
| 2017 | Scope-Aware Useful Cache Block Analysis for Data Cache Related Preemption DelayabstractStatic timing analysis is crucial for design of realtime systems. While the worst-case execution time of a task is typically computed or measured in a single task environment, the presence of caches imposes additional cache related preemption delay (CRPD) cost to the lower priority tasks in a preemptive multi-tasking system. In this work, we show that existing instruction CRPD analysis techniques cannot be straightforwardly extended for safe and precise data CRPD analysis. In order to capture the dynamic behavior of the data memory references, we introduce the notion of temporal scopes into the abstract cache state (ACS) to capture the data memory blocks that must or may reside in the cache during certain time intervals of program execution. Based on the improved ACS representation, we present a temporal scope aware useful cache block (UCB) calculation for safe and tight estimation of the data CRPD cost. Experimental results show that the proposed technique leads to substantially tighter CRPD estimation, and is applicable to programs with complex data reference patterns. Wei Zhang 0173, Fan Gong, Lei Ju 0001, Nan Guan, Zhiping Jia |
RTAS | 3 |
| 2017 | Energy-aware morphable cache management for self-powered non-volatile processorsabstractWearable, implantable and Internet of Things devices are attracting increasing attention from both research and industry. Energy harvesting is a promising alternative of battery to power these embedded systems. However, the intrinsic instability of energy harvesting systems leads to potential frequent power interruptions. In order to survive the power failures, non-volatile processor (NVP) is proposed to back up volatile information before power depletion and recover the system status after power resumes. Non-volatile memory (NVM) is typically attached for cache and main memory backup. There are researches working on optimization of the backup, however, little of them involve multiple level cell (MLC) NVM. In this work, we first discuss the benefit of applying MLC NVM for cache backup, and then propose a three-stage energy-aware cache management strategy to improve the system performance and energy utilization while guaranteeing successful backups. Evaluation shows that the proposed scheme can achieve 16.7% energy reduction with comparative performance with the single level cell (SLC) based hybrid cache. Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Xin Li 0001, Zhiping Jia |
RTCSA | 3 |
| 2017 | Memory-Aware Embedded Control Systems DesignabstractControl applications are often implemented on highly cost-sensitive and resource-constrained embedded platforms, such as microcontrollers with a small on-chip memory. Typically, control algorithms are designed using model-based approaches, where the details of the implementation platform are completely ignored. As a result, optimizations that integrate platform-level characteristics into the control algorithms design are largely missing. With the emergence of cyber-physical systems (CPS)-oriented thinking, there has lately been a strong interest in co-design of control algorithms and their implementation platforms, leading to work on networked control systems and computation-aware control algorithms design. However, there has so far been no work on integrating the characteristics of a memory architecture into the design of control algorithms. In this paper we, for the first time, show that accounting for the impact of on-chip memory (or cache) reuse on the performance of control applications motivates new techniques for control algorithms design. This leads to significant improvement in quality of control for given resource availability, or more efficient implementations of embedded control applications. We believe that this paper opens up a variety of possibilities for memory-related optimizations of embedded control systems, that will be pursued by researchers working on computer-aided design for CPS. Wanli Chang 0001, Dip Goswami, Samarjit Chakraborty, Lei Ju 0001, Chun Jason Xue, Sidharta Andalam |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | SLA-aware energy-efficient scheduling scheme for Hadoop YARN
Xiaojun Cai, Feng Li 0014, Lei Ju 0001, Zhiping Jia |
J. Supercomput. | 4 |
| 2016 | Write-back aware shared last-level cache management for hybrid main memoryabstractHybrid main memory with both DRAM and emerging non-volatile memory (NVM) becomes a promising solution for high performance and energy-efficient embedded systems. Cache plays an important role and highly affects the number of write backs to NVM and DRAM blocks. However, existing cache policies fail to fully address the significant asymmetry between NVM operations (especially writes) and DRAM operations, leading to non-optimal system designs. We propose a write-back aware last-level cache management scheme for the hybrid main memory, which improves the cache hit ratio of NVM memory blocks and minimizes write-backs to NVM. Experimental results show that our proposed framework leads to better performance and energy saving compared with the state-of-the-art cache management scheme for hybrid main memory architecture. Deshan Zhang, Lei Ju 0001, Mengying Zhao, Xiang Gao 0012, Zhiping Jia |
DAC | 2 |
| 2016 | Unified DRAM and NVM hybrid buffer cache architecture for reducing journaling overhead
Zhiyong Zhang 0006, Lei Ju 0001, Zhiping Jia |
DATE | 2 |
| 2016 | Energy efficient task allocation for hybrid main memory architecture
Xiaojun Cai, Lei Ju 0001, Xin Li 0002, Zhiyong Zhang 0006, Zhiping Jia |
J. Syst. Archit. | 2 |
| 2016 | Data aggregation framework for energy-efficient WirelessHART networks
Feng Li 0014, Lei Ju 0001, Zhiping Jia |
J. Syst. Archit. | 2 |
| 2015 | A three-stage-write scheme with flip-bit for PCM main memoryabstractPhase-change memory (PCM) is a nonvolatile memory which suffers slow write performance and limited write endurance. Besides, writing a one to a PCM cell needs longer time but less electrical current than writing a zero. In traditional PCM schemes, zeros and ones in a word are written at the same time and word write time has to be the time to write a one, thus incurring time waste. In this paper, we propose a three-stage write scheme with flip-bit for PCM main memory to reduce the number of changed bits and write latency. In our scheme, write operation is divided into comparison, write-0 and write-1 stages. In the comparison stage, new data and old data are compared and the new data is re-encoded by a flip-bit to minimize changed bits. Then the flip-bit and re-encoded data are written to PCM cells in an accelerating manner. All zero bits and one bits are written separately in later two stages to avoid the time waste in traditional write. Our scheme shrinks time consumption and reduces bit changes caused by write operation over other existing schemes. The experimental results show that this scheme decreases 43.5% bit changes, 16.6% write time and 34.6% write energy consumption on average. Xin Li 0002, Lei Ju 0001, Zhiping Jia |
ASP-DAC | 3 |
| 2015 | Approximation-aware scheduling on heterogeneous multi-core architecturesabstractThe high performance demand of embedded systems along with restrictive thermal design power (TDP) constraint have lead to the emergence of the heterogenous multi-core architectures, where cores with the same instruction-set architecture but different power-performance characteristics provide new opportunities for energy-efficient computing. Heterogeneity introduces challenges in scheduling the tasks to the appropriate cores and selecting the frequency assignment of each core. In this paper, we introduce an approximation-aware scheduling framework for soft real-time tasks on the heterogeneous multi-core architectures. We consider multiple versions of a task obtained by introducing approximation in the computation to provide different levels of quality of service (QoS) versus performance tradeoffs. The additional choice of approximation allows us more flexibility in meeting the performance and TDP constraints while maximizing QoS per unit of energy. Cheng Tan 0002, Thannirmalai Somu Muthukaruppan, Tulika Mitra, Lei Ju 0001 |
ASP-DAC | 4 |
| 2015 | Managing hybrid on-chip scratchpad and cache memories for multi-tasking embedded systemsabstractOn-chip memory management is essential in design of high performance and energy-efficient embedded systems. While many off-the-shelf embedded processors employ a hybrid on-chip SRAM architecture including both scratchpad memories (SPMs) and caches, many existing work on SPM management ignore the synergy between caches and SPMs. In this work, we propose a static SPM allocation strategy for the hybrid on-chip memory architecture in a multi-tasking environment, which minimizes the overall access latency and energy consumption of the instruction memory subsystem. We capture cache conflict misses via a fine-grained temporal cache behavior model. An integer linear programming (ILP) based formulation is proposed to generate an function-level SPM allocation scheme, where both intra- and inter-task cache interference as well as access frequency are captured for an optimal memory subsystem design. Compared with the state-of-the-art static SPM allocation strategy in a multitasking environment, experimental results show that our SPM management scheme achieves 30.51% further improvement in instruction memory subsystem performance, and up to 34.92% in terms of energy saving. Zimeng Zhou, Lei Ju 0001, Zhiping Jia, Xin Li 0002 |
ASP-DAC | 2 |
| 2015 | Reducing Journaling Overhead with Hybrid Buffer Cache
Zhiyong Zhang 0006, Lei Ju 0001, Zhiping Jia |
ICA3PP (4) | 2 |
| 2015 | Hybrid scratchpad and cache memory management for energy-efficient parallel HEVC encodingabstractThe next-generation video coding standard High Efficiency Video Coding (HEVC) provides better compression rates for high resolution videos compared with H.264, at the cost of significantly increased needs for computation power and memory bandwidth. Therefore, memory subsystem optimization is of paramount importance to support HEVC on resource and energy constrained embedded consumer electronics. In this paper, we present a hybrid on-chip memory architecture with both caches and scratchpad memories (SPMs) for parallel HEVC encoding. A run-time prediction algorithm is proposed to effectively identify the most-frequently accessed memory regions in the search window(s) for processing individual coding tree units (CTUs). Depending on their intra- and inter-core reuses, these regions are loaded into the private or shared SPMs for guaranteed on-chip memory accesses. On the other hand, a relatively small hardware-controlled cache is used for the rest of data accesses. Moreover, an adaptive power gating scheme is proposed to power off SPM sectors with expired load windows to further reduce the on-chip leakage power. Compared with the state-of-the-art solution, experimental results show that our proposed memory management framework supports high speed parallel HEVC processing with substantially smaller on-chip memory size, which achieves up to 76.23% on-chip leakage energy savings, and 33.31% energy saving for the overall memory subsystem. Lei Ju 0001, Zhiping Jia |
ICCD | 2 |
| 2015 | A Novel OpenFlow-Based DDoS Flooding Attack Detection and Response Mechanism in Software-Defined NetworkingabstractSoftware-Defined Networking (SDN) and OpenFlow have brought a promising architecture for the future networks. However, there are still a lot of security challenges to SDN. To protect SDN from the Distributed denial-of-service (DDoS) flooding attack, this paper extends the flow entry counters and adds a mark action of OpenFlow, then proposes an entropy-based distributed attack detection model, a novel IP traceback and source filtering response mechanism in SDN with OpenFlow-based Deterministic Packet Marking. It achieves detecting the attack at the destination and filtering the malicious traffic at the source and can be easily implemented in SDN controller program, software or programmable switch, such as Open vSwitch and NetFPGA. The experimental results show that this scheme can detect the attack quickly, achieve a high detection accuracy with a low false positive rate, shield the victim from attack traffic and also avoid the attacker consuming resource and bandwidth on the intermediate links. Rui Wang 0075, Zhiyong Zhang 0006, Lei Ju 0001, Zhiping Jia |
Int. J. Inf. Secur. Priv. | 3 |
| 2015 | Instruction Cache Locking Using Temporal Reuse ProfileabstractThe performance of most embedded systems is critically dependent on the average memory access latency. Improving the cache hit rate can have significant positive impact on the performance of an application. Modern embedded processors often feature cache locking mechanisms that allow memory blocks to be locked in the cache under software control. Cache locking was primarily designed to offer timing predictability for hard real-time applications. Hence, prior techniques focus on employing cache locking to improve the worst-case execution time. However, cache locking can be quite effective in improving the average-case execution time of general embedded applications as well. In this paper, we explore static instruction cache locking to improve the average-case program performance. We introduce temporal reuse profile (TRP) to accurately and efficiently model the cost and benefit of locking memory blocks in the cache. We consider two locking mechanisms, line locking and way locking. For each locking mechanism, we propose a branch-and-bound algorithm and a heuristic approach that use the TRP to determine the most beneficial memory blocks to be locked in the cache. Experimental results show that the heuristic approach achieves close to the results of branch-and-bound algorithm and can improve the performance by 12% on average for 4 KB cache across a suite of real-world benchmarks. Moreover, our heuristic provides significant improvement compared to the state-of-the-art locking algorithm both in terms of performance and efficiency. Yun Liang 0001, Tulika Mitra, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | An Improved Energy-Efficient Scheduling for Precedence Constrained Tasks in Multiprocessor Clusters
Xin Li 0002, Yanheng Zhao, Yibin Li 0002, Lei Ju 0001, Zhiping Jia |
ICA3PP (1) | 4 |
| 2014 | Reliable and Energy Efficient Routing Algorithm for WirelessHART
Feng Li 0014, Lei Ju 0001, Zhiping Jia, Zhaopeng Zhang |
ICA3PP (1) | 3 |
| 2014 | Energy efficient real-time task scheduling for embedded systems with hybrid main memoryabstractAvailable energy is the most critical limitation on the performance of embedded systems along with the increasing sophistication. Phase Change Memory (PCM), with high density and low idle power has recently been extensively studied as a promising alternative main memory of DRAM. In this paper, a hybrid PCM/DRAM main memory is utilized to leverage the low power of PCM and high performance of DRAM. We reconsider the real-time task scheduling problem of hybrid PCM/DRAM-based embedded systems. To maximize energy saving, two static scheduling algorithms under Rate-Monotonic(RM) and Earliest Deadline First (EDF) are proposed while guaranteeing the real-time constraints of all tasks. Since the actual execution time is much shorter than the worst-case execution time in real environment, we propose two dynamic mechanisms to optimize the energy consumption of our static solutions, so as to fully use the slack time produced by completed tasks. All the proposed algorithms minimize the number of task migrations from PCM to DRAM and ensure each task instance can be migrated at most once. Experimental results show our real-time scheduling algorithms reduce 25.7% to 47.2% of energy consumption on average. Zhiyong Zhang 0006, Lei Ju 0001, Zhiping Jia |
RTCSA | 3 |
| 2014 | On-demand gateway broadcast scheme for connecting mobile ad hoc networks to the InternetabstractGateway discovery algorithm is a fundamental protocol for interconnecting mobile ad hoc network (MANET) with the Internet. In most existing schemes, each gateway node broadcasts gateway advertisements to announce its presence. The decision of when to emit advertisements can influence the performance of the network. Traditional gateway discovery schemes adopt the method of periodically emitting advertisements with a time interval. However, this method does not fulfill the actual needs of the source nodes. This paper proposes a novel adaptive scheme for gateway discovery, in which the gateway broadcasts advertisements only on-demand instead of periodic emission. In order to obtain the network's actual demands for gateway advertisement, routes to the gateway are monitored. In particular, if any route is predicted to be broken, the source node requires fresh gateway advertisements to update routes, and then the gateway will be triggered to broadcast to fulfill such demands. We study the performance of on-demand gateway discovery scheme by a comparison approach. The results show that the proposed adaptive gateway discovery scheme greatly outperforms the conventional solutions: it is capable of achieving higher packet delivery ratio and lower end-to-end delay, while minimizing the routing overhead. Huaqiang Xu, Lei Ju 0001, Chongxian Guo, Zhiping Jia |
SMARTCOMP | 2 |
| 2014 | A High-Performance Distributed Certificate Revocation Scheme for Mobile Ad Hoc NetworksabstractMobile ad hoc networks (MANETs) are wireless networks which have a wide range applications due to their dynamic topologies and easy to deployment. However, such networks are also more vulnerable to attacks compared with traditional wireless networks. Certificate revocation is an effective mechanism for providing network security services. Existing schemes are not well suited for MANETs because of incurring much overhead or bring low accuracy on certificate revocation. Therefore, we propose a high-performance distributed certificate revocation scheme in which certificates of malicious nodes will be revoked quickly and accurately. Certificate revocation is the result of the collaborative effect of multiple accusations. For diluting damages to networks, one accusation is enough to limit the accusation function of the accused node. To enhance the accuracy of certificate revocation, our scheme requires nodes just accepting those accusations in which trust levels of accuser nodes are not less than accused nodes'. To guarantee the rapidity, we restore accusation functions of the falsely accused nodes after revoking certificates of all malicious nodes who ever accused them. Moreover, we design one mechanism to reward nodes who ever accused those malicious nodes, and in return, accusations made by them will accelerate the certificate revocation processes of other malicious nodes. Simulation results demonstrate the effectiveness and efficiency of our scheme in certificate revocation. In addition, our scheme achieves a great improvement of just limiting accusation functions of malicious nodes. Chongxian Guo, Huaqiang Xu, Lei Ju 0001, Zhiping Jia, Jihai Xu |
TrustCom | 3 |
| 2014 | High Performance FPGA Implementation of Elliptic Curve Cryptography over Binary FieldsabstractIn this paper, we propose a high performance hardware implementation architecture of elliptic curve scalar multiplication over binary fields. The proposed architecture is based on the Montgomery ladder method and uses polynomial basis for finite field (FF) arithmetic. A single Karatsuba multiplier runs with no idle cycle significantly increases the performance of FF multiplication while spending small amount of hardware resources, and other FF operations performed in parallel with the FF multiplier. The optimized circuits lead to a lesser area requirement compared to other high performance implementations. An implementation for the National Institute of Standards and Technology (NIST) recommended curve with degree 163 is shown, the proposed design can reach 121 MHz with 10,417 slices when implemented on Xilinx Virtex-4 XC4VLX200 FPGA device, the total time required for one elliptic curve scalar multiplication is 9.0 μs. Lei Ju 0001, Xiaojun Cai, Zhiping Jia, Zhiyong Zhang 0006 |
TrustCom | 2 |
| 2013 | Binarization-Based Human Detection for Compact FPGA Implementation
Shuai Xie, Yibin Li 0002, Zhiping Jia, Lei Ju 0001 |
APPT | 4 |
| 2013 | Binarization based implementation for real-time human detectionabstractHardware implementation of human detection is a challenging task for embedded designs. This paper presents a real-time image-based field-programmable gate array (FPGA) implementation of human detection. Our implementation is based on the histograms of oriented gradients (HOG) feature and linear support vector machine (SVM) classifier. The novelty of this work is that we replace normalization process of HOG with a modified binarization process. Therefore, during classification process with SVM classifier, all multiplication operations are replaced by addition operations. All these modifications result in reduction of hardware resource. Experimental evaluation reveals that 293 fps can be achieved on a low-end Xilinx Spartan-3e FPGA. Moreover, a detection accuracy of 1.97% miss rate and 1% false positive rate is achieved. For further demonstration, a prototype system is developed with an OV7670 camera device. Restricted to the speed of camera, a detection rate of 30 fps is achieved. Shuai Xie, Yibin Li 0002, Zhiping Jia, Lei Ju 0001 |
FPL | 4 |
| 2013 | Trust prediction and trust-based source routing in mobile ad hoc networks
Hui Xia 0001, Zhiping Jia, Xin Li 0002, Lei Ju 0001, Edwin H.-M. Sha |
Ad Hoc Networks | 4 |
| 2013 | Impact of trust model on on-demand multi-path routing in mobile ad hoc networks
Hui Xia 0001, Zhiping Jia, Lei Ju 0001, Xin Li 0002, Edwin H.-M. Sha |
Comput. Commun. | 3 |
| 2012 | Tenant Onboarding in Evolving Multi-tenant Software-as-a-Service SystemsabstractA multi-tenant software as a service (SaaS) system has to meet the needs of several tenant organizations, which connect to the system to utilize its services. To leverage economies of scale through re-use, a SaaS vendor would, in general, like to drive commonality amongst the requirements across tenants. However, many tenants will also come with some custom requirements that may be a pre-requisite for them to adopt the SaaS system. These requirements then need to be addressed by evolving the SaaS system in a controlled manner, while still supporting the requirements of existing tenants. In this paper, we focus on functional variability amongst tenants in a multi-tenant SaaS and develop a framework to help evolve such systems systematically. We adopt an intuitive formal model of services that is easily amenable to tenant requirement analysis and provides a robust way to support multiple tenant on boarding, which is modeled as a bi-objective optimization problem that attempts to maximize vendor profit and tenant functional commonality. We perform a substantial case study of a multi-tenant blog server to demonstrate the benefits of our proposed approach. Lei Ju 0001, Bikram Sengupta, Abhik Roychoudhury |
ICWS | 1 |
| 2012 | The Research and Application of a Specific Instruction Processor for SMS4abstractSMS4 is a block cipher used in the Chinese National Standard for Wired Authentication and Privacy Infrastructure (WAPI). Like other cryptosystems, the SMS4 contains data-intensive computation, whose throughput makes a substantial contribution to the overall system performance. Design and implementation of the SMS4 cryptosystem to meet the real-time requirements of applications are challenge problems, especially for embedded network systems with limited computation resources. In this work, we present a systematic design approach of application-specific instruction-set processor (ASIP) for the SMS4 cryptographic algorithm, which exploits and compromises between the flexibility of software execution and the performance of the application-specific integrated circuit (ASIC) based implementation of the SMS4 cryptosystem. We identify and perform a design space exploration of the custom instructions found for the SMS4 algorithm, and extend the instruction set architecture (ISA) of a standard 32-bit RISC processor to accommodate them. We employ the Electronic System Level (ESL) methodology in the development of the proposed ASIP using the Xilinx Virtex5 LX110T FPGA platform. Results show that compared to the original RISC ISA, our ASIP for SMS4 achieves 2.93 times performance improvement and 52.4% less program memory utilization, with only 28.2% more resource required. Zhenzhou Li, Feng Li 0014, Zhiping Jia, Lei Ju 0001, Renhai Chen |
TrustCom | 4 |
| 2012 | Performance debugging of Esterel specifications
Lei Ju 0001, Bach Khoa Huynh, Abhik Roychoudhury, Samarjit Chakraborty |
Real Time Syst. | 1 |
| 2011 | Scope-Aware Data Cache Analysis for WCET EstimationabstractCaches are widely used in modern computer systems to bridge the increasing gap between processor speed and memory access time. On the other hand, presence of caches, especially data caches, complicates the static worst case execution time (WCET) analysis. Access pattern analysis (e.g., cache miss equations) are applicable to only a specific class of programs, where all array accesses must have predictable access patterns. Abstract interpretation-based methods (must/persistence analysis) determines possible cache conflicts based on coarse-grained memory access information from address analysis, which usually leads to significantly pessimistic estimation. In this paper, we first present a refined persistence analysis method which fixes the potential underestimation problem in the original persistence analysis. Based on our new persistence analysis, we propose a framework to combine access pattern analysis and abstract interpretation for accurate data cache analysis. We capture the dynamic behavior of a memory access by computing its temporal scope (the loop iterations where a given memory block is accessed for a given data reference) during address analysis. Temporal scopes as well as loop hierarchy structure (the static scopes) are integrated and utilized to achieve a more precise abstract cache state modeling. Experimental results shows that our proposed analysis obtains up to 74% reduction in the WCET estimates compared to existing data cache analysis. Bach Khoa Huynh, Lei Ju 0001, Abhik Roychoudhury |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2011 | Multicast Trusted Routing with QoS Multi-constraints in Wireless Ad Hoc NetworksabstractThe wireless ad hoc networks has been attracting increasing attention of researchers owing to its good performance and special applications. Multicast routing with QoS multi-constraints problem in this network has drawn wide spread attention from researchers who have been using different methods to solve it. One solution to this problem is to merge paths in accordance with a swarm intelligence algorithm after finding paths from source node to destination nodes in order to obtain a multicast tree that satisfies QoS multi-constraints. However, the security of routing paths is not guaranteed. To overcome this shortcoming, in this paper, we introduce the concept of trust into multicast routing problem which is made as another QoS constraint, and finally we propose a multicast trusted routing algorithm with QoS multi-constraints based on a modified ant colony algorithm. In this new algorithm, ants move constantly on the network to find an optimal constrained multicast security tree. Simulation results indicate that our algorithm can quickly find the feasible solution to solve constrained multicast security routing issues. Compared to CSTMAN, the packet delivery ratio of new algorithm is improved. Hui Xia 0001, Zhiping Jia, Lei Ju 0001, Youqin Zhu |
TrustCom | 3 |
| 2010 | Timing analysis of esterel programs on general-purpose multiprocessorsabstractSynchronous languages like Esterel have gained wide popularity in certain domains such as avionics. However, platform-specific timing analysis of code generated from Esterel-like specifications have mostly been neglected so far. The growing volume of electronics and software in domains like automotive, calls for formal-specification based code generation to replace manually written and optimized code. Such cost-sensitive domains require tight estimation of timing properties of the generated code. Towards this goal, we propose a scheme for generating C code from Esterel specifications for a multiprocessor platform, followed by timing analysis of the generated code. Due to dependencies across program fragments mapped onto different processors, traditional Worst-Case Execution Time (WCET) analysis techniques for sequential programs cannot applied be to this setting. Our proposed timing analysis technique is tailored to capture such inter-processor code dependencies. Our main novelty stems from how we detect and remove infeasible paths arising from a multiprocessor implementation during our timing analysis. We apply our timing analysis on a number of standard Esterel benchmarks, which show that performing the proposed inter-processor infeasible path elimination may lead to up to 14.3% tighter estimation of the WCRT, thereby leading to resource over-dimensioning and poor design. Lei Ju 0001, Bach Khoa Huynh, Abhik Roychoudhury, Samarjit Chakraborty |
DAC | 1 |
| 2009 | Context-sensitive timing analysis of Esterel programsabstractTraditionally, synchronous languages, such as Esterel, have been compiled into hardware, where timing analysis is relatively easy. When compiled into software -- e.g., into sequential C code -- very conservative estimation techniques have been used, where the focus has only been on obtaining safe timing estimates and not on the cost of the implementation. While this was acceptable in avionics, efficient implementations and hence tight timing estimates are needed in more cost-sensitive application domains. Lately, a number of advances in Worst-Case Execution Time (WCET) analysis techniques, coupled with the growing use of software in domains such as automotives, have led to a considerable interest in timing analysis of code generated from Esterel specifications. In this paper we propose techniques to obtain tight estimates on the processing time of input events by sequential C code generated from Esterel programs. Execution of an Esterel program -- as in all other synchronous languages -- is logically made up of a sequence of clock ticks. In reality, they take non-zero time which depends on the generated C code as well as the underlying hardware platform on which this code is executed. Apart from exploiting the specific structure of this C code to obtain tight WCET estimates, we capture program-level contexts across ticks in order to obtain tight estimates on response times of events whose processing spans across multiple clock ticks. Such tighter estimates immediately translate into more cost-effective implementations. Our experimental results with realistic case studies show 30% reduction in timing estimates when program level context information is taken into account. Lei Ju 0001, Bach Khoa Huynh, Samarjit Chakraborty, Abhik Roychoudhury |
DAC | 1 |
| 2008 | Schedulability Analysis of MSC-based System ModelsabstractMessage sequence charts (MSCs) are widely used for describing interaction scenarios between the components of a distributed system. Consequently, worst-case response time estimation and schedulability analysis of MSC-based specifications form natural building blocks for designing distributed real-time systems. However, currently there exists a large gap between the timing and quantitative performance analysis techniques that exist in the real-time systems literature, and the modeling/specification techniques that are advocated by the formal methods community. As a result, although a number of schedulability analysis techniques are known for a variety of task graph-based models, it is not clear if they can be used to effectively analyze standard specification formalisms such as MSCs. In this paper we make an attempt to bridge this gap by proposing a schedulability analysis technique for MSC-based system specifications. We show that compared to existing timing analysis techniques for distributed real-time systems, our proposed analysis gives tighter results, which immediately translate to better system design and improved resource dimensioning. We illustrate the details of our analysis using a setup from the automotive electronics domain, which consist of two real-life application programs (that are naturally modeled using MSCs) running on a platform consisting of multiple electronic control units (ECUs) connected via a FlexRay bus. Lei Ju 0001, Abhik Roychoudhury, Samarjit Chakraborty |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2007 | Accounting for cache-related preemption delay in dynamic priority schedulability analysis
Lei Ju 0001, Samarjit Chakraborty, Abhik Roychoudhury |
DATE | 1 |