EDBT 2026 Demo / reviewers in the wild / expert
Lu Peng 0001
dblp:16/6199-1
· DBLP profile ↗
67ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-3545-286XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 48 · 3 first-author · 11 since 2021Software engineering, systems software and programming languages · 6Artificial intelligence and machine learning · 5 · 5 since 2021Security and privacy · 5Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantum Probabilistic Label Refining: Enhancing Label Quality for Robust Image ClassificationabstractLearning with softmax cross-entropy on one-hot labels often leads to overconfident predictions and poor robustness under noise or perturbations. Label smoothing mitigates this by redistributing some confidence uniformly, but treats all samples equally, ignoring intra-class variability. We propose a hybrid quantum–classical framework that leverages quantum non-determinism to refine data labels into probabilistic ones, offering more nuanced, human-like uncertainty representations than label smoothing or Bayesian approaches. A variational quantum circuit (VQC) encodes inputs into multi-qubit quantum states, using entanglement and superposition to capture subtle feature correlations. Measurement via the Born rule extracts probabilistic soft labels that reflect input-specific uncertainty. These labels are then used to train a classical convolutional neural network (CNN) with soft-target cross-entropy loss. On MNIST and Fashion-MNIST, our method improves robustness—achieving up to 50% higher accuracy under noise—while maintaining competitive clean-data accuracy. It also enhances model calibration and interpretability, as CNN outputs better reflect quantum-derived uncertainty. This work introduces Quantum Probabilistic Label Refining, which bridges quantum measurement and classical deep learning to enable robust training via refined, correlation-aware labels, without architectural changes or adversarial techniques. Fang Qi, Lu Peng 0001, Zhengming Ding |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | PruningQC: Boosting the Quantum Computation Fidelity by Pruning Redundant Gates
Fang Qi, Yongshan Ding 0001, Victor Bankston, Ji Liu 0007, Lu Peng 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | NANI: Energy-efficient Neuron-Aware hardware Noise Injection for adversarial defense using undervolting
Qiyu Wan, Jing Wang 0055, Mingsong Chen 0001, Lu Peng 0001, Xin Fu 0001 |
J. Syst. Archit. | 5 |
| 2025 | Regional Weather Variable Predictions by Machine Learning With Near-Surface Observational and Atmospheric Numerical DataabstractAccurate and timely regional weather prediction is vital for sectors dependent on weather-related decisions. Traditional prediction methods, based on atmospheric equations, often struggle with coarse temporal resolutions and inaccuracies. This article presents a novel machine learning (ML) model, called Micro-Macro (MiMa), that integrates both near-surface observational data from Kentucky Mesonet stations (collected every 5 min, known as Micro data) and hourly atmospheric numerical outputs (termed as Macro data) for fine-resolution weather forecasting. The MiMa model employs an encoder-decoder transformer structure, with two encoders for processing multivariate data from both datasets and a decoder for forecasting weather variables over short time horizons. Each instance of the MiMa model, called a modelet, predicts the values of a specific weather parameter at an individual mesonet station. The approach is extended with Regional MiMa (Re-MiMa) modelets, which are designed to predict weather variables at ungauged locations by training on multivariate data from a few representative stations in a region, tagged with their elevations. Re-MiMa can provide highly accurate predictions across an entire region, even in areas without observational stations. Experimental results show that MiMa significantly outperforms current models, with Re-MiMa offering precise short-term forecasts for ungauged locations, marking a significant advancement in weather forecasting accuracy and applicability. Yihe Zhang 0001, Bryce Turney, Purushottam Sigdel, Xu Yuan 0001, Eric Rappin, Adrian Lago, Sytske K. Kimball, Li Chen 0019, Paul J. Darby, Lu Peng 0001, Sercan Aygün, Yazhou Tu, M. Hassan Najafi, Nian-Feng Tzeng |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2024 | A New Routing Strategy to Improve Success Rates of Quantum ComputersabstractIn the current noisy intermediate-scale quantum (NISQ) Era, Quantum Computing faces significant challenges due to noise, which severely restricts the application of computing complex algorithms. Superconducting quantum chips, one of the pioneer quantum computation technologies, introduce additional noise when moving qubits to adjacent locations for operation on designated two-qubit gates. The current compilers rely on decision models that either count the swap gates or multiply the gate errors when choosing swap paths at the routing stage. Our research has unveiled the overlooked situations for error propagations through the circuit, leading to accumulations that may affect the final output. Fang Qi, Xin Fu 0001, Xu Yuan 0001, Nian-Feng Tzeng, Lu Peng 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2024 | Soft Error Resilience Analysis of LSTM NetworksabstractLong Short-Term Memory (LSTM) deep neural networks are diverse in the tasks they can accomplish, such as image captioning and speech recognition. However, they remain susceptible to transient faults when deployed in environments with high-energy particles or radiation. It remains unknown how the potential transient faults will impact LSTM models. Therefore, we investigate the resilience of the weights and biases of these networks through four implementations of the original LSTM network. Based on the observations made through the fault injection of these networks, we propose an effective method of fault mitigation through Hamming encoding of selected weights and biases in a given network. Christopher P. Vasquez, Travis LeCompte, Xu Yuan 0001, Nian-Feng Tzeng, Lu Peng 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2024 | Federated Text-driven Prompt Generation for Vision-Language ModelsabstractPrompt learning for vision-language models, e.g., CoOp, has shown great success in adapting CLIP to different downstream tasks, making it a promising solution for federated learning due to computational reasons. Existing prompt learning techniques replace hand-crafted text prompts with learned vectors that offer improvements on seen classes, but struggle to generalize to unseen classes. Our work addresses this challenge by proposing Federated Text-driven Prompt Generation (FedTPG), which learns a unified prompt generation network across multiple remote clients in a scalable manner. The prompt generation network is conditioned on task-related text input, thus is context-aware, making it suitable to generalize for both seen and unseen classes. Our comprehensive empirical evaluations on nine diverse image classification datasets show that our method is superior to existing federated prompt learning methods, achieving better overall generalization on both seen and unseen classes, as well as datasets. Chaithanya Kumar Mummadi, Madan Ravi Ganesh, Lu Peng 0001, Wan-Yi Lin |
ICLR | 6 |
| 2023 | Stochastic Computing for Reliable Memristive In-Memory ComputationabstractIn-Memory Computing (IMC) is a promising computing paradigm to accelerate Big Data applications. It reduces the data movement between memory and processing units, and provides massive parallelism. Memristive technology is one of the promising technologies for IMC. This emerging technology, however, is still in evolution, facing practical challenges. Memristive memories are prone to softerror while storing the data and during computations. The traditional binary encoding commonly used in memristive IMC is highly sensitive to soft-errors, which makes developing reliable memristive IMC more challenging. Stochastic Computing (SC) is a re-emerging computing paradigm that is highly robust against soft-errors as any bit flip leads to only a least significant bit error. In this work, we study SC as a solution to increase the reliability of memristive IMC. We investigate how and to what extent SC may address or improve the reliability issues of current memristive technology, and memristive IMC. We also evaluate the characteristics yielded by memristive stochastic IMC and compare them with those of the traditional reliability techniques. Mohsen Riahi Alam, M. Hassan Najafi, Nima Taherinejad, Mohsen Imani, Lu Peng 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | Graph Neural Network Assisted Quantum Compilation for Qubit AllocationabstractQuantum computers in the current noisy intermediate-scale quantum (NISQ) era face two major limitations - size and error vulnerability. Although quantum error correction (QEC) methods exist, they are not applicable at the current size of computers, requiring thousands of qubits, while NISQ systems have nearly one hundred at most. One common approach to improve reliability is to adjust the compilation process to create a more reliable final circuit, where the two most critical compilation decisions are the qubit allocation and qubit routing problems. We focus on solving the qubit allocation problem and identifying initial layouts that result in a reduction of error. To identify these layouts, we combine reinforcement learning with a graph neural network (GNN)-based Q-network to process the mesh topology of the quantum computer, known as the backend, and make mapping decisions, creating a Graph Neural Network Assisted Quantum Compilation (GNAQC) strategy. We train the architecture using a set of four backends and six circuits and find that GNAQC improves output fidelity by roughly 12.7% over pre-existing allocation methods. Travis LeCompte, Fang Qi, Xu Yuan 0001, Nian-Feng Tzeng, M. Hassan Najafi, Lu Peng 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2023 | MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision TransformerabstractPrecise crop yield prediction provides valuable information for agricultural planning and decision-making processes. However, timely predicting crop yields remains challenging as crop growth is sensitive to growing season weather variation and climate change. In this work, we develop a deep learning-based solution, namely Multi-Modal Spatial-Temporal Vision Transformer (MMST-ViT), for predicting crop yields at the county level across the United States, by considering the effects of short-term meteorological variations during the growing season and the long-term climate change on crops. Specifically, our MMST-ViT consists of a Multi-Modal Transformer, a Spatial Transformer, and a Temporal Transformer. The Multi-Modal Transformer leverages both visual remote sensing data and short-term meteorological data for modeling the effect of growing season weather variations on crop growth. The Spatial Transformer learns the high-resolution spatial dependency among counties for accurate agricultural tracking. The Temporal Transformer captures the long-range temporal dependency for learning the impact of long-term climate change on crops. Meanwhile, we also devise a novel multi-modal contrastive learning technique to pre-train our model without extensive human supervision. Hence, our MMST-ViT captures the impacts of both short-term weather variations and long-term climate change on crops by leveraging both satellite images and meteorological data. We have conducted extensive experiments on over 200 counties in the United States, with the experimental results exhibiting that our MMST-ViT outperforms its counterparts under three performance metrics of interest. Our dataset and code are available at https://github.com/fudong03/MMST-ViT. Fudong Lin, Summer Crawford, Kaleb Guillot, Yihe Zhang 0001, Xu Yuan 0001, Li Chen 0019, Shelby Williams, Robert Minvielle, Xiangming Xiao, Drew Gholson, Nicolas Ashwell, Tri Setiyono, Brenda Tubana, Lu Peng 0001, Magdy A. Bayoumi, Nian-Feng Tzeng |
ICCV | 15 |
| 2023 | Boosting Performance and QoS for Concurrent GPU B+trees by Combining-Based SynchronizationabstractConcurrent B+trees have been widely used in many systems. With the scale of data requests increasing exponentially, the systems are facing tremendous performance pressure. GPU has shown its potential to accelerate concurrent B+trees performance. When many concurrent requests are processed, the conflicts should be detected and resolved. Prior methods guarantee the correctness of concurrent GPU B+trees through lock-based or software transactional memory (STM)-based approaches. However, these methods complicate the request processing logic, increase the number of memory accesses and bring execution path divergence. They lead to performance degradation and variance in response time increasing. Moreover, previous methods do not guarantee linearizability among concurrent requests. Chuanlei Zhao, Lu Peng 0001, Yuzhe Lin, Fengzhe Zhang, Yunping Lu |
PPoPP | 3 |
| 2022 | GeauxTrace: A Scalable Privacy-Protecting Contact Tracing App Design Using BlockchainabstractContact tracing is the approach to identifying physical contact between human beings using a variety of data such as personal details and locations to discover the potential infection of diseases. Since the outbreak of the COVID-19 pandemic, contact tracing has been used extensively to quarantine the people at risk to stop the spread. Moreover, the data collected during contact tracing are typical spatiotemporal data, which can be used to study the disease and discover the spread pattern. However, both traditional labor-intensive and modern digital-based approaches have limitations in terms of cost and privacy concerns. In this paper, we proposed GeauxTrace, a Blockchain-based privacy-protecting contact tracing platform, which separates private data from proof of contact. Sensitive data collected by the front-end app via Bluetooth-based methods are stored locally, and only the proofs of contacts are uploaded onto the immutable private blockchain, which forms a global contact graph at the backend. Our approach not only enables multi-hop risky users to be notified but also reveals the infection patterns via the global graph, which could help study diseases and assist the policymaker. Our implementation shows the feasibility of the proposed platform in real-world scenarios and achieves the performance of 20-30 user requests per second. Fang Qi, John Ner, Tianqing Feng, Brian T. Cunningham, Lu Peng 0001 |
BDCAT | 6 |
| 2022 | Cascade Variational Auto-Encoder for Hierarchical DisentanglementabstractWhile deep generative models pave the way for many emerging applications, decreased interpretability for larger model sizes and complexities hinders their generalizability to wide domains such as economy, security, healthcare, etc. Considering this obstacle, a common practice is to learn interpretable representations through latent feature disentanglement, aiming for exposing a set of mutually independent factors of data variations. However, existing methods either fail to catch the trade-off between the synthetic data quality and model interpretability, or consider the first-order feature disentangling only, overlooking the fact that a subset of salient features can carry decomposable semantic meanings and hence be of high-order in nature. Hence, we in this paper propose a novel generative modeling paradigm by introducing a Bayesian network-based regularize on a cascade Variational Auto-Encoder (VAE). Specifically, this regularizer guides the learner to discover a representation space that comprises both first-order disentangled features and high-order salient features, with the feature interplay captured by the Bayesian structure. Experiments demonstrate that this regularizer gives us free control over the representation space and can guide the learner to discover decomposable semantic meanings by capturing the interplay among independent factors. Meanwhile, we benchmark extensive experiments on six widely-used vision datasets, and the results exhibit that our approach outperforms the state-of-the-art VAE competitors in terms of the trade-off between the synthetic data quality and model interpretability. Although our design is framed in the VAE regime, it in effect is generic and can be better amenable to both GANs and VAEs in terms of letting them concurrently enjoy both high model interpretability and high synthesis quality. Fudong Lin, Xu Yuan 0001, Lu Peng 0001, Nian-Feng Tzeng |
CIKM | 3 |
| 2022 | High performance GPU concurrent B+treeabstractConcurrent B+trees have been widely used in many systems from file systems to databases. With the volume of data requests expanding exponentially, the systems are facing tremendous performance pressure. GPUs have shown their potential to accelerate the concurrent B+trees operations with their high volume of parallel computing resources and large memory bandwidth. In concurrent B+tree, the conflicts should be detected and resolved when multiple concurrent requests are traversing and operating on the tree. However, conflict detection and handling in concurrent B+tree complicates the request processing logic, increases the number of memory accesses and leads to execution path divergence. That leads to performance degradation and increased response time variance. Chuanlei Zhao, Lu Peng 0001, Yuzhe Lin, Fengzhe Zhang, Jinhu Jiang |
PPoPP | 3 |
| 2022 | Protecting Synchronization Mechanisms of Parallel Big Data Kernels via LoggingabstractWith the growing effort to reduce power consumption in machines, fault tolerance becomes more of a concern. This holds particularly for large-scale computing, where execution failures due to soft faults waste excessive time and resources. These large-scale applications are normally parallel in nature and rely on control structures tailored specifically for parallel computing, such as locks and barriers. While there are many studies on resilient software, to our knowledge none of them focus on protecting these parallel control structures. In this work, we present a method of ensuring the correct operation of both locks and barriers in parallel applications. Our method tracks the memory locations used within parallel sections and detects a violation of the control structures. Upon detecting any violation, the violating thread is rolled back to the beginning of the structure and reattempts it, similar to rollback mechanisms in transactional memory systems. We test the method on representative samples of the BigDataBench kernels and find it exhibits a mean error reduction of 93.6% for basic mutex locks and barriers with a mean 6.55% execution time overhead at 64 threads. Additionally, we provide a comparison to transactional memory methods and demonstrate up to a mean 57.5% execution time overhead reduction. Travis LeCompte, Lu Peng 0001, Xu Yuan 0001, Nian-Feng Tzeng |
IEEE Trans. Computers | 2 |
| 2021 | GPU-Assisted Memory ExpansionabstractRecent graphic processing units (GPUs) often come with large on-board physical memory to accelerate diverse parallel program executions on big datasets with regular access patterns, including machine learning (ML) and data mining (DM). Such a GPU may underutilize its physical memory during lengthy ML model training or DM, making it possible to lend otherwise unused GPU memory to applications executed concurrently on the host machine. This work explores an effective approach that lets memory-intensive applications run on the host machine CPU with its memory expanded dynamically onto available GPU on-board DRAM, called GPU-assisted memory expansion (GAME). Targeting computer systems equipped with the recent GPUs, our GAME approach permits speedy executions on CPU with large memory footprints by harvesting unused GPU on-board memory on-demand for swapping, far surpassing competitive GPU executions. Implemented in user space, our GAME prototype lets GPU memory house swapped-out memory pages transparently, without code modifications for high usability and portability. The evaluation of NAS-NPB benchmark applications demonstrates that GAME expedites monotasking (or multitasking) executions considerably by up to 2.1× (or 3.1×), when memory footprints exceed the CPU DRAM size and an equipped GPU has unused VDRAM available for swapping use. Pisacha Srinuan, Purushottam Sigdel, Xu Yuan 0001, Lu Peng 0001, Paul J. Darby III, Christopher Aucoin, Nian-Feng Tzeng |
NAS | 4 |
| 2021 | Precise Weather Parameter Predictions for Target Regions via Neural Networks
Yihe Zhang 0001, Xu Yuan 0001, Sytske K. Kimball, Eric Rappin, Li Chen 0019, Paul J. Darby III, Tom Johnsten, Lu Peng 0001, Boisy Pitre, David M. Bourrie, Nian-Feng Tzeng |
ECML/PKDD (5) | 8 |
| 2020 | BPU: A Blockchain Processing Unit for Accelerated Smart Contract ExecutionabstractModern blockchains use smart contracts to implement automatic and decentralized programs, which are the foundations of Decentralized Applications (DApp). The poor performance on general purpose computers has become the bottleneck that limits the blockchain and smart contracts from being widely used. In this paper, we present BPU, a high-performance modularized blockchain processing unit. BPU aims at bringing performance and flexibility to the blockchain and DApp processing. Our design achieves significant speedup compared against the software implementation on an Intel CPU. Lu Peng 0001 |
DAC | 2 |
| 2020 | ATT: A Fault-Tolerant ReRAM Accelerator for Attention-based Neural NetworksabstractCrossbar-based resistive RAM has been widely used in deep learning accelerator designs because it largely eliminates weight movement between memory and processing units. The high-density storage and low leakage power make it a good fit for edge/IoT devices. However, existing ReRAM designs for traditional neural networks cannot support Attention-based Neural Networks, which are stacked with encoders and decoders instead of convolutional layers or fully connected layers. In addition to matrix-matrix multiplications in traditional neural networks, an encoder or a decoder also includes the attention mechanism, the layer normalization and the gaussian error linear unit. These new characteristics make the data flow far more complicated than that of a convolutional layer. Faulty ReRAM devices are additional obstacles when mapping weights that severely degrade computation accuracy. Existing hardware redundancy strategies that are unaware of application characteristics usually result in inefficient designs. In this work, we analyze the data flow of these attention-based neural networks and propose a ReRAM-based accelerator with a dedicated pipeline design for Attention-based Neural Networks. When considering cells with hard faults in crossbars, we further propose NuXG, a non-uniform redundancy strategy, to meet accuracy requirements and save energy consumption by decreasing the redundancy ratio. Finally, we evaluate results and demonstrate that the proposed can achieve more than two times improved performance over existing redundancy schemes in both power efficiency and throughput for Attention-based Neural Networks. Moreover, it also significantly outperforms an NVIDIA GPU. Haoqiang Guo, Lu Peng 0001, Jian Zhang 0004, Travis LeCompte |
ICCD | 2 |
| 2020 | Robust Cache-Aware Quantum Processor LayoutabstractQuantum computation has taken over as one of the largest current research areas in computer architecture and information theory. With the potential to make a large number of factorization-based encryption methods obsolete, companies and governments around the globe are racing to build the first large-scale quantum computer. Currently, most quantum computers are noisy intermediate-scale quantum (NISQ), using a relatively small collection of unreliable qubits. While error correction methods exist, they require a large number of ancilla qubits to protect the data qubits which is not practical for use on current NISQ machines. However, following the Dowling-Neven Law, available qubits on a superconducting chip are growing at an exponential rate similar to Moore's Law. Looking toward larger scale quantum machines, we examine a method to increase usable qubit density of quantum machines implementing error correction by using quantum caches that utilize simpler error correction codes. Alternatively, this also allows for the design of reliable systems while meeting the performance and qubit requirements for quantum algorithms. We modify the Qiskit quantum simulation library to work with caches and investigate the effects of region size and topology on the swap characteristics of algorithm execution. We also present our results and discuss recommended topologies for each algorithm. Lastly, we present mix scale-out simulations to examine the impact of cache on future large-scale machines. The default central cache topology gains a maximum performance increase of 2.15 times compared to the worst topology, which creates a robust cache-aware quantum processor layout. Travis LeCompte, Fang Qi, Lu Peng 0001 |
SRDS | 3 |
| 2020 | Computer comparisons in the presence of performance variation
Samuel Irving, Bin Li 0008, Shaoming Chen, Lu Peng 0001, Lide Duan |
Frontiers Comput. Sci. | 4 |
| 2020 | Architectural Support for NVRAM Persistence in GPUsabstractNon-volatile Random Access Memories (NVRAM) have emerged in recent years to bridge the performance gap between the main memory and external storage devices, such as Solid State Drives (SSD). In addition to higher storage density, NVRAM provides byte-addressability, higher bandwidth, near-DRAM latency, and easier access compared to block devices such as traditional SSDs. This enables new programming paradigms taking advantage of durability and larger memory footprint. With the range and size of GPU workloads expanding, NVRAM will present itself as a promising addition to GPU's memory hierarchy. To utilize the non-volatility of NVRAMs, programs should allow durable stores, maintaining consistency through a power loss event. This is usually done through a logging mechanism that works in tandem with a transaction execution layer which can consist of a transactional memory or a locking mechanism. Together, this results in a transaction processing system that preserves the ACID properties. GPUs are designed with high throughput in mind, leveraging high degrees of parallelism. Transactional memory proposals enable fine-grained transactions at the GPU thread-level. However, with lower write bandwidths compared to that of DRAMs, using NVRAM as-is may yield sub-optimal overall system performance when threads experience long latency. To address this problem, we propose using Helper Warps to move persistence out of the critical path of transaction execution, alleviating the impact of latencies. Our mechanism achieves a speedup of 4.4 and 1.5 under bandwidth limits of 1.6 GB/s and 12 GB/s and is projected to maintain speed advantage even when NVRAM bandwidth gets as high as hundreds of GB/s in certain cases. Due to the speedup, our proposed method also results in reduction in overall energy consumption. Sui Chen, Lei Liu 0037, Lu Peng 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | A High Throughput B+tree for SIMD ArchitecturesabstractB+tree is one of the most important data structures and has been widely used in different fields. With the increase of concurrent queries and data-scale in storage, designing an efficient B+tree structure has become critical. Due to abundant computation resources, SIMD architectures provide potential opportunities to achieve high query throughput for B+tree. However, prior methods cannot achieve satisfactory performance results due to low resource utilization and poor memory performance. In this paper, we first identify the gaps between B+tree and SIMD architectures. Concurrent B+tree queries involve many global memory accesses and different divergences, which mismatch with SIMD architecture features. Based on this observation, we propose Harmonia, a novel B+tree structure to bridge the gaps. In Harmonia, a B+tree structure is divided into a key region and a child region. The key region stores the nodeswith its keys in a breadth-first order. The child region is organized as a prefix-sum array, which only stores each node's first child index in the key region. Since the prefix-sum child region is small and the children's index can be retrieved through index computations, most of it can be stored in on-chip caches, which can achieve good cache locality. To make it more efficient, Harmonia also includes two optimizations: partially-sorted aggregation and narrowed thread-group traversal, which can mitigate memory and execution divergence and improve resource utilization. Evaluations on a 28-core INTEL CPU show that Harmonia can achieve up to 207 million queries per second, which is about 1.7X faster than that of CPU-based HB+Tree, a recent state-of-the-art solution. And on a Volta TITAN V GPU, it can achieve up to 3.6 billion queries per second, which is about 3.4X faster than that of GPU-based HB+Tree. Zhaofeng Yan, Yuzhe Lin, Chuanlei Zhao, Lu Peng 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2019 | Fooling AI with AI: An Accelerator for Adversarial Attacks on Deep Learning Visual ClassificationabstractRecent studies identify that Deep learning Neural Networks (DNNs) are vulnerable to subtle perturbations, which are not perceptible to the human visual system but can fool the DNN models and lead to wrong outputs. These algorithms are the first efforts to move forward to secure deep learning by providing an avenue to train future defense networks. We propose the first hardware accelerator for adversarial attacks based on memristor crossbar arrays. Our design significantly improves the throughput of a visual adversarial perturbation system, which can further improve the robustness and security of future deep learning systems. Based on the algorithm uniqueness, we propose four implementations for the adversarial attack accelerator (A^3) to improve the throughput, energy efficiency, and computational efficiency. Haoqiang Guo, Lu Peng 0001, Jian Zhang 0004, Fang Qi, Lide Duan |
ASAP | 2 |
| 2019 | Efficient GPU NVRAM Persistence with Helper WarpsabstractNon-volatile Random-Access Memories (NVRAM) have emerged in recent years to bridge the performance gap between the main memory and external storage devices. To utilize the non-volatility of NVRAMs, programs should allow durable stores, meaning consistency must be maintained during a power loss event. GPUs are designed with high throughput, leveraging high degrees of parallelism. However, with lower NVRAM write bandwidths compared to that of DRAMs, using NVRAM as is may yield suboptimal overall system performance. To address this problem, we propose using Helper Warps to move persistence out of the critical path of transaction execution, alleviating the impact of latencies. Our mechanism achieves a speedup of 4.4 and 1.5 under bandwidth limits of 1.6 GB/s and 12 GB/s and is projected to maintain speed advantage even when NVRAM bandwidth gets as high as hundreds of GB/s in certain cases. Sui Chen, Faen Zhang, Lei Liu 0037, Lu Peng 0001 |
DAC | 4 |
| 2019 | Harmonia: a high throughput B+tree for GPUsabstractB+tree is one of the most important data structures and has been widely used in different fields. With the increase of concurrent queries and data-scale in storage, designing an efficient B+tree structure has become critical. Due to abundant computation resources, GPUs provide potential opportunities to achieve high query throughput for B+tree. However, prior methods cannot achieve satisfactory performance results due to low resource utilization and poor memory performance. Zhaofeng Yan, Yuzhe Lin, Lu Peng 0001 |
PPoPP | 3 |
| 2019 | Long Short-Term Memory Network Design for Analog ComputingabstractWe present an analog-integrated circuit implementation of long short-term memory network, which is compatible with digital CMOS technology. We have used multiple-input floating gate MOSFETs as both the front-end to obtain converted analog signals and the differential pairs in proposed analog multipliers. Analog crossbar is built by the analog multiplier processing matrix and bitwise multiplications. We have shown that using current signals as internal transmission signals can largely reduce computation delay, compared to the digital implementation. We also have introduced analog blocks to work as activation functions for the algorithm. In the back-end of our design, we have used current comparators to achieve the output to be readable to external digital systems. We have designed the LSTM network with the matrix size of 16 × 16 in TSMC 180nm CMOS technology. The post-layout simulations show that the latency of one computing cycle is 1.19ns without memory, and power dissipation of the single analog LSTM computing core with 2 kilobytes SRAM at 200MHz is 460.3mW. The overhead of power dissipation due to SRAM access is 8.3%, in which the computing of each LSTM layer requires one computing cycle. The energy efficiency is 0.95TOP/s/W. Ashok Srivastava, Lu Peng 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2019 | Hierarchical Hybrid Memory Management in OS for Tiered Memory SystemsabstractThe emerging hybrid DRAM-NVM architecture is challenging the existing memory management mechanism at the level of the architecture and operating system. In this paper, we introduce Memos, a memory management framework which can hierarchically schedule memory resources over the entire memory hierarchy including cache, channels, and main memory comprising DRAM and NVM simultaneously. Powered by our newly designed kernel-level monitoring module that samples the memory patterns by combining TLB monitoring with page walks, and page migration engine, Memos can dynamically optimize the data placement in the memory hierarchy in response to the memory access pattern, current resource utilization, and memory medium features. Our experimental results show that Memos can achieve high memory utilization, improving system throughput by around 20.0 percent; reduce the memory energy consumption by up to 82.5 percent; and improve the NVM lifetime by up to 34X. Lei Liu 0037, Shengjie Yang, Lu Peng 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | qSwitch: Dynamical Off-Chip Bandwidth Allocation Between Local and Remote AccessesabstractMultisocket computer systems are popular in workstations and servers. However, they suffer from the relatively low bandwidth of intersocket communication especially for massive parallel workloads that generate many intersocket requests for synchronizations and remote memory accesses. Intersocket traffic puts pressure on the underlying network connecting all processors with a limited bandwidth confined by pin resources. Given this constraint, we propose to dynamically increase the intersocket bandwidth by sacrificing off-chip memory bandwidth when systems have heavy intersocket communication but few off-chip memory accesses. Our design increases the physical bandwidth for intersocket communication via switching the function of pins from off-chip memory accesses to intersocket communication and can achieve an average performance speedup of 1.28 in geocentric mean for selected parallel multithreaded benchmarks. Shaoming Chen, Lu Peng 0001, Samuel Irving, Ashok Srivastava |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Carpool: a bufferless on-chip network supporting adaptive multicast and hotspot alleviationabstractModern chip multiprocessors (CMPs) employ on-chip networks to enable communication between the individual cores. Operations such as coherence and synchronization generate a significant amount of the on-chip network traffic, and often create network requests that have one-to-many (i.e., a core multicasting a message to several cores) or many-to-one (i.e., several cores sending the same message to a common hotspot destination core) flows. As the number of cores in a CMP increases, one-to-many and many-to-one flows result in greater congestion on the network. To alleviate this congestion, prior work provides hardware support for efficient one-to-many and many-to-one flows in buffered on-chip networks. Unfortunately, this hardware support cannot be used in bufferless on-chip networks, which are shown to have lower hardware complexity and higher energy efficiency than buffered networks, and thus are likely a good fit for large-scale CMPs. Xi-Yue Xiang, Saugata Ghose, Lu Peng 0001, Onur Mutlu, Nian-Feng Tzeng |
ICS | 4 |
| 2017 | Accelerating GPU Hardware Transactional Memory with Snapshot IsolationabstractSnapshot Isolation (SI) is an established model in the database community, which permits write-read conflicts to pass and aborts transactions only on write-write conflicts. With the Write Skew anomaly correctly eliminated, SI can reduce the occurrence of aborts, save the work done by transactions, and greatly benefit long transactions involving complex data structures. Sui Chen, Lu Peng 0001, Samuel Irving |
ISCA | 2 |
| 2017 | A novel switchable pin method for regulating power in chip-multiprocessor
Ashok Srivastava, Lu Peng 0001, Shaoming Chen, Saraju P. Mohanty |
Integr. | 3 |
| 2017 | Soft error resilience of Big Data kernels through algorithmic approaches
Travis LeCompte, Walker Legrand, Sui Chen, Lu Peng 0001 |
J. Supercomput. | 4 |
| 2017 | Exploring Energy-Efficient Cache Design in Emerging Mobile PlatformsabstractMobile devices are quickly becoming the most widely used processors in consumer devices. Since their major power supply is battery, energy-efficient computing is highly desired. In this article, we focus on energy-efficient cache design in emerging mobile platforms. We observe that more than 40% of L2 cache accesses are OS kernel accesses in interactive smartphone applications. Such frequent kernel accesses cause serious interferences between the user and kernel blocks in the L2 cache, leading to unnecessary block replacements and high L2 cache miss rate. We first propose to statically partition the L2 cache into two separate segments, which can be accessed only by the user code and kernel code, respectively. Meanwhile, the overall size of the two segments is shrunk, which reduces the energy consumption while still maintaining the similar cache miss rate. We then find completely different access behaviors between the two separated kernel and user segments and explore the multi-retention STT-RAM-based user and kernel segments to obtain higher energy savings in this static partition-based cache design. Finally, we propose to dynamically partition the L2 cache into the user and kernel segments to minimize overall cache size. We also integrate the short-retention STT-RAM into this dynamic partition-based cache design for maximal energy savings. The experimental results show that our static technique reduces cache energy consumption by 75% with 2% performance loss, and our dynamic technique further shows strong capability to reduce cache energy consumption by 85% with only 3% performance loss. Kaige Yan, Lu Peng 0001, Mingsong Chen 0001, Xin Fu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2017 | Using Switchable Pins to Increase Off-Chip Bandwidth in Chip-MultiprocessorsabstractOff-chip memory bandwidth has been considered as one of the major limiting factors of processor performance, especially for multi-cores and many-cores. Conventional processor design allocates a large portion of off-chip pins to deliver power, leaving a small number of pins for processor signal communication. We observe that a processor requires much less power during memory intensive stages than is available. This is due to the fact that the frequencies of processor cores waiting for data to be fetched from offchip memories can be scaled down in order to save power without degrading performance. Motivated by this observation, we propose a dynamic pin switching technique to alleviate this bandwidth limitation. This technique is introduced to dynamically exploit surplus power delivery pins to provide extra bandwidth during memory intensive program phases, thereby significantly boosting performance. This work is extended to compare two approaches for increasing off chip bandwidths using switchable pins. Additionally, it shows significant performance improvements for memory intensive workloads on a memory subsystem using Phase Change Memory. Shaoming Chen, Samuel Irving, Lu Peng 0001, Ying Zhang 0016, Ashok Srivastava |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Efficient GPU hardware transactional memory through early conflict resolutionabstractIt has been proposed that Transactional Memory be added to Graphics Processing Units (GPUs) in recent years. One proposed hardware design, Warp TM, can scale to 1000s of concurrent transactions. As a programming method that can atomicize an arbitrary number of memory access locations and greatly reduce the efforts to program parallel applications, transactional memory handles the complexity of inter-thread synchronization. However, when thousands of transactions run concurrently on a GPU, conflicts and resource contentions arise, causing performance loss. In this paper, we identify and analyze the cause of conflicts and contentions and propose two enhancements that try to resolve conflicts early: (1) Early-Abort global conflict resolution that allows conflicts to be detected before they reach the Commit Units so that contention in the Commit Units is reduced and (2) Pause-and-Go execution scheme that reduces the chance of conflict and the performance penalty of re-executing long transactions. These two enhancements are enabled by a single hardware modification. Our evaluation shows the combination of the two enhancements greatly improves overall execution speed while reducing energy consumption. Sui Chen, Lu Peng 0001 |
HPCA | 2 |
| 2016 | Parallelizing image feature extraction algorithms on multi-core platforms
Yunping Lu, Haibo Chen 0001, Lu Peng 0001 |
J. Parallel Distributed Comput. | 6 |
| 2016 | Soft error resilience in Big Data kernels through modular analysis
Sui Chen, Greg Bronevetsky, Lu Peng 0001, Bin Li 0008, Xin Fu 0001 |
J. Supercomput. | 3 |
| 2016 | Performance Analysis of Multimedia Retrieval Workloads Running on MulticoresabstractMultimedia data has become a major data type in the Big Data era. The explosive volume of such data and the increasing real-time requirement to retrieve useful information from it have put significant pressure in processing such data in a timely fashion. However, while prior efforts have done in-depth analysis on architectural characteristics of traditional multimedia processing and text-based retrieval algorithms, there has been no systematic study towards the emerging multimedia retrieval applications. This may impede the architecture design and system evaluation of these applications. In this paper, we make the first attempt to construct a multimedia retrieval benchmark suite (MMRBench for short) that can be used to evaluate architectures and system designs for multimedia retrieval applications. MMRBench covers modern multimedia retrieval algorithms with different versions (sequential, parallel and distributed). MMRBench also provides a series of flexible interfaces as well as certain automation tools. With such a flexible design, the algorithms in MMRBench can be used both in individual kernel-level evaluation and in integration to form a complete multimedia data retrieval infrastructure for full system evaluation. Furthermore, we use performance counters to analyze a set of architecture characteristics of multimedia retrieval algorithms in MMRBench, including the characteristics of core level, chip level and inter-chip level. The study shows that micro-architecture design in current processor is inefficient (both in performance and power) for these multimedia retrieval workloads, especially in core resources and memory systems. We then derive some insights into the architecture design and system evaluation for such multimedia retrieval algorithms. Yunping Lu, Xin Wang 0019, Haibo Chen 0001, Lu Peng 0001, Wenyun Zhao |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2015 | Precise computer comparisons via statistical resampling methodsabstractPerformance variability, stemming from non-deterministic hardware and software behaviors or deterministic behaviors such as measurement bias, is a well-known phenomenon of computer systems which increases the difficulty of comparing computer performance metrics. Conventional methods use various measures (such as geometric mean) to quantify the performance of different benchmarks to compare computers without considering variability. This may lead to wrong conclusions. In this paper, we propose three resampling methods for performance evaluation and comparison: a randomization test for a general performance comparison between two computers, bootstrapping confidence estimation, and an empirical distribution and five-number-summary for performance evaluation. The results show that 1) the randomization test substantially improves our chance to identify the difference between performance comparisons when the difference is not large; 2) bootstrapping confidence estimation provides an accurate confidence interval for the performance comparison measure (e.g. ratio of geometric means); and 3) when the difference is very small, a single test is often not enough to reveal the nature of the computer performance and a five-number-summary to summarize computer performance. We illustrate the results and conclusion through detailed Monte Carlo simulation studies and real examples. Results show that our methods are precise and robust even when two computers have very similar performance metrics. Bin Li 0008, Shaoming Chen, Lu Peng 0001 |
ISPASS | 3 |
| 2015 | NBTI alleviation on FinFET-made GPUs by utilizing device heterogeneity
Ying Zhang 0016, Sui Chen, Lu Peng 0001, Shaoming Chen |
Integr. | 3 |
| 2015 | A framework for evaluating comprehensive fault resilience mechanisms in numerical programs
Sui Chen, Greg Bronevetsky, Bin Li 0008, Marc Casas, Lu Peng 0001 |
J. Supercomput. | 5 |
| 2014 | QoS management on heterogeneous architecture for parallel applicationsabstractQuality of service (QoS) management is widely employed to provide differentiable performance to programs with distinctive priorities on conventional chip multi-processor (CMP) platforms. Recently, heterogeneous architecture integrating diverse processor cores on the same silicon has been proposed to better serve various application domains and it is expected to be an important design paradigm of future processors. Therefore, the QoS management on emerging heterogeneous systems will be of great significance. On the other hand, parallel applications are becoming increasingly important in modern computing community in order to explore the benefit of thread-level parallelism on CMPs. However, considering the diverse characteristics of thread synchronization, data sharing, and parallelization pattern, governing the execution of multiple parallel programs with different performance requirements becomes a complicated yet significant problem. In this paper, we study QoS management for parallel applications running on heterogeneous CMP systems. We comprehensively assess a series of task-to-core mapping policies on a real heterogeneous hardware (QuickIA) by characterizing their impacts on performance of individual applications. Our evaluation results show that the proposed QoS policies are effective to improve the performance of programs with highest priority while striking good tradeoff with system fairness. Ying Zhang 0016, Li Zhao 0002, Ramesh Illikkal, Ravi R. Iyer 0001, Andrew Herdrich, Lu Peng 0001 |
ICCD | 6 |
| 2014 | Increasing off-chip bandwidth in multi-core processors with switchable pinsabstractOff-chip memory bandwidth has been considered as one of the major limiting factors to processor performance, especially for multi-cores and many-cores. Conventional processor design allocates a large portion of off-chip pins to deliver power, leaving a small number of pins for processor signal communication. We observed that the processor requires much less power than that can be supplied during memory intensive stages. This is due to the fact that the frequencies of processor cores waiting for data to be fetched from off-chip memories can be scaled down in order to save power without degrading performance. In this work, motivated by this observation, we propose a dynamic pin switch technique to alleviate the bandwidth limitation issue. The technique is introduced to dynamically exploit the surplus pins for power delivery in the memory intensive phases and uses them to provide extra bandwidth for the program executions, thus significantly boosting the performance. Shaoming Chen, Ying Zhang 0016, Lu Peng 0001, Jesse Ardonne, Samuel Irving, Ashok Srivastava |
ISCA | 4 |
| 2014 | Comprehensive and Efficient Design Parameter Selection for Soft Error Resilient Processors via Universal RulesabstractSoft errors have been significantly degrading the reliability of current processors whose feature sizes and supply voltages are fast scaling down. In this paper, we propose two effective approaches to characterize processor reliability against soft errors at presilicon stage. By utilizing a rule search strategy named Patient Rule Induction Method (PRIM), we are capable of generating a set of selective rules on key design parameters. These rules quantify the design space subregion with the lowest effective soft error rate (SER), thus providing useful guidelines in designing reliable processors. Furthermore, we also propose to use Classification and Regression Trees (CART) to partition the design space into a number of small subregions each being associated with a representative SER value. This gives the processor designer a global view of the SER distribution, enabling a comprehensive analysis over the entire design space. More importantly, both approaches generate “universal” models whose effectiveness is validated with a set of test programs unseen to training. Compared to traditional application-specific design space studies, our models’ cross-program capability can save great training effort in the era of multithreading. Finally, a case study on multiprocessors is performed to simultaneously balance multiple design metrics, including reliability, performance, and power. Lide Duan, Ying Zhang 0016, Bin Li 0008, Lu Peng 0001 |
IEEE Trans. Computers | 4 |
| 2013 | Optimization of Electricity and Server Maintenance Costs in Hybrid Cooling Data CentersabstractThe electricity cost of data centers dominated by server power and cooling power is growing rapidly. To tackle this problem, inlet air with moderate temperature and server consolidation are widely adopted. However, the benefit of these two methods is limited due to conventional air cooling systems ineffectiveness caused by re-circulation and low heat capacity. To address this problem, hybrid air and liquid cooling, as a practical and inexpensive approach, has been introduced. In this paper, we quantitatively analyze the impact of server consolidation and temperature of cooling water on the total electricity and server maintenance costs in hybrid cooling data centers. To minimize the total costs, we proposed to maintain sweet temperature and ASTT (available sleeping time threshold) by which a joint cost optimization can be satisfied. By using real world traces, the potential savings of sweet temperature and ASTT are estimated to be average 18% of the total cost while 99% requests are satisfied compared to a strategy which only reduces electricity cost. Shaoming Chen, Lu Peng 0001 |
IEEE CLOUD | 3 |
| 2013 | Lighting the dark silicon by exploiting heterogeneity on future processorsabstractAs we embrace the deep submicron era, dark silicon caused by the failure of Dennard scaling impedes us from attaining commensurate performance benefit from the increased number of transistors. To alleviate the dark silicon and effectively leverage the advantage of decreased feature size, we consider a set of design paradigms by exploiting heterogeneity in the processor manufacturing. We conduct a thorough investigation on these design patterns from different evaluation perspectives including performance, energy-efficiency, and cost-efficiency. Our observations can provide insightful guidance to the design of future processors in the presence of dark silicon. Ying Zhang 0016, Lu Peng 0001, Xin Fu 0001 |
DAC | 2 |
| 2013 | Predicting Architectural Vulnerability on Multithreaded Processors under Resource Contention and SharingabstractArchitectural vulnerability factor (AVF) characterizes a processor's vulnerability to soft errors. Interthread resource contention and sharing on a multithreaded processor (e.g., SMT, CMP) shows nonuniform impact on a program's AVF when it is co-scheduled with different programs. However, measuring the AVF is extremely expensive in terms of hardware and computation. This paper proposes a scalable two-level predictive mechanism capable of predicting a program's AVF on a SMT/CMP architecture from easily measured metrics. Essentially, the first-level model correlates the AVF in a contention-free environment with important performance metrics and the processor configuration, while the second-level model captures the interthread resource contention and sharing via processor structures' occupancies. By utilizing the proposed scheme, we can accurately estimate any unseen program's soft error vulnerability under resource contention and sharing with any other program(s), on an arbitrarily configured multithreaded processor. In practice, the proposed model can be used to find soft error resilient thread-to-core scheduling for multithreaded processors. Lide Duan, Lu Peng 0001, Bin Li 0008 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2011 | Tree structured analysis on GPU power studyabstractGraphics Processing Units (GPUs) have emerged as a promising platform for parallel computation. With a large number of processor cores and abundant memory bandwidth, GPUs deliver substantial computation power. While providing high computation performance, a GPU consumes high power and needs sufficient power supplies and cooling systems. It is essential to institute an efficient mechanism for evaluating and understanding the power consumption when running real applications on high-end GPUs. In this paper, we present a high-level GPU power consumption model using sophisticated tree-based random forest methods which correlate and predict the power consumption using a set of performance variables. We demonstrate that this statistical model not only predicts the GPU runtime power consumption more accurately than existing regression based approaches, but more importantly, it provides sufficient insights into understanding the correlation of the GPU power consumption with individual performance metrics. We use a GPU simulator that can collect more runtime performance metrics than hardware counters. We measure the power consumption of a wide-range of CUDA kernels on an experimental system with GTX 280 GPU to collect statistical samples for power analysis. The proposed method is applicable to other GPUs as well. Jianmin Chen, Bin Li 0008, Ying Zhang 0016, Lu Peng 0001, Jih-Kwon Peir |
ICCD | 4 |
| 2011 | Universal rules guided design parameter selection for soft error resilient processorsabstractHigh-performance processors suffer from soft error vulnerability due to the increasing on-chip transistor density, shrinking processor feature size, lower threshold voltage, etc. In this paper, we propose to use a rule search strategy, i.e. Patient Rule Induction Method (PRIM), to optimize processor soft error robustness. By exploring a huge microarchitectural design space on the Architectural Vulnerability Factor (AVF), we are capable of generating a set of selective rules on key design parameters. Applying these rules at early design stage effectively identifies the configurations that are inherently reliable to soft errors. Furthermore, we propose a generic approach capable of generating a set of “universal” rules that achieves the optimization of the output variable for different programs in execution. The effectiveness of the universal rule set is validated on programs that are not used in training. This cross-program capability is very useful in the era of multi-threading. Finally, the proposed scheme is extended to multiprocessors where multiple design metrics including reliability, performance and power are balanced. Our proposed methodology is able to generate quantitative and universal solutions for both uniprocessors and multiprocessors. Lide Duan, Ying Zhang 0016, Bin Li 0008, Lu Peng 0001 |
ISPASS | 4 |
| 2011 | Performance and Power Analysis of ATI GPU: A Statistical ApproachabstractWe present a comprehensive study on the performance and power consumption of a recent ATI GPU. By employing a rigorous statistical model to analyze execution behaviors of representative general-purpose GPU (GPGPU) applications, we conduct insightful investigations on the target GPU architecture. Our results demonstrate that the GPU execution throughput and the power dissipation are dependent on different architectural variables. Furthermore, we design a set of micro-benchmarks to study the power consumption features of different function units on the GPU. Based on those results, we derive instructive principles that can guide the design of power-efficient high performance computing systems. Ying Zhang 0016, Bin Li 0008, Lu Peng 0001 |
NAS | 4 |
| 2010 | Weak execution ordering - exploiting iterative methods on many-core GPUsabstractOn NVIDIA's many-core GPUs, there is no synchronization function among parallel thread blocks. When fine-granularity of data communication and synchronization is required for large-scale parallel programs executed by multiple thread blocks, frequent host synchronization are necessary, and they incur a significant overhead. We investigate a class of applications which uses a chaotic version of iterative methods to obtain numerical solutions for partial differential equations (PDE). Such a fast PDE solver is parallelized on GPUs with multiple thread blocks. In this parallel implementation, although frequent data communication is needed between adjacent thread blocks, a precise order of the data communication is not necessary. Separate communication threads are used for periodically exchanging the boundary values with adjacent thread blocks through the global memory. Since a precise order of the data communication is not required, the computation and the communication threads can be overlapped to alleviate the communication overhead. Performance measurements of two popular applications, Poisson image editing from computer graphics and shape from shading from computer vision, on Tesla C1060 show that a speedup of 4–5 times is achievable for both applications in comparison with the solution using host synchronization. Jianmin Chen, Feiqi Su, Jih-Kwon Peir, Jeff Ho, Lu Peng 0001 |
ISPASS | 6 |
| 2010 | Expediating IP lookups with reduced power via TBM and SST supernode caching
Ying Zhang 0016, Lu Peng 0001, Wencheng Lu, Lide Duan, Suresh Rai |
Comput. Commun. | 2 |
| 2010 | A Host-Based Intrusion Detection System Using Architectural Features to Improve Sophisticated Denial-of-Service Attack DetectionsabstractApplication features like port numbers are used by Network-based Intrusion Detection Systems (NIDSs) to detect attacks coming from networks. System calls and the operating system related information are used by Host-based Intrusion Detection Systems (HIDSs) to detect intrusions toward a host. However, the relationship between hardware architecture events and Denial-of-Service (DoS) attacks has not been well revealed. When increasingly sophisticated intrusions emerge, some attacks are able to bypass both the application and the operating system level feature monitors. Therefore, a more effective solution is required to enhance existing HIDSs. In this article, the authors identify the following hardware architecture features: Instruction Count, Cache Miss, Bus Traffic and integrate them into a HIDS framework based on a modern statistical Gradient Boosting Trees model. Through the integration of application, operating system and architecture level features, the proposed HIDS demonstrates a significant improvement of the detection rate in terms of sophisticated DoS intrusions. Li Yang 0001, Lu Peng 0001, Bin Li 0008 |
Int. J. Inf. Secur. Priv. | 3 |
| 2010 | Efficient Microarchitectural Vulnerabilities Prediction Using Boosted Regression Trees and Patient Rule InductionsabstractThe shrinking processor feature size, lower threshold voltage, and increasing clock frequency make modern processors highly vulnerable to transient faults. Architectural Vulnerability Factor (AVF) reflects the possibility that a transient fault eventually causes a visible error in the program output, and it indicates a system's susceptibility to transient faults. Therefore, the awareness of the AVF, especially at early design stage, is greatly helpful to achieve a trade-off between system performance and reliability. However, tracking the AVF during program execution is extremely costly, which makes accurate AVF prediction extraordinarily attractive to computer architects. In this paper, we propose to use Boosted Regression Trees (BRT), a nonparametric tree-based predictive modeling scheme, to identify the correlation across workloads, execution phases, and processor configurations between a key processor structure's AVF and various performance metrics. The proposed method not only makes an accurate prediction but also quantitatively illustrates individual performance variable's importance to the AVF. A quantitative comparison between our model and conventional linear regression is performed in terms of model stability, showing that our model is more stable when the model size varies. Moreover, to reduce the prediction complexity, we also utilize a technique named Patient Rule Induction Method (PRIM) to extract some simple selecting rules on important metrics. Applying these rules during runtime can fast identify execution intervals with a relatively high AVF. A case study that enables PRIM-based ROB redundancy has been performed to demonstrate a possible application of the trained PRIM rules. Bin Li 0008, Lide Duan, Lu Peng 0001 |
IEEE Trans. Computers | 3 |
| 2009 | A case study: Using architectural features to improve sophisticated denial-of-service attack detectionsabstractApplication features such as port numbers are used by network-based intrusion detection systems (NIDSs) to detect attacks coming from networks. System calls and the operating system related information are used by host-based intrusion detection systems (HIDSs) to detect intrusions towards a host. However, the relationship between hardware architecture events and denial-of-service (DoS) attacks has not been well revealed. When increasingly sophisticated intrusions emerge, some attacks are able to bypass both the application and the operating system level feature monitors. Therefore, a more effective solution is required to enhance existing HIDSs. In this paper, we identify the following hardware architecture features: instruction count, cache miss, bus traffic and integrate them into a novel HIDS framework based on a modern statistical gradient boosting trees model. Through the integration of application, operating system and architecture level features, our proposed HIDS demonstrates a significant improvement of the detection rate in terms of sophisticated DoS intrusions. Li Yang 0001, Lu Peng 0001, Bin Li 0008, Alma Cemerlic |
CICS | 3 |
| 2009 | Versatile prediction and fast estimation of Architectural Vulnerability Factor from processor performance metricsabstractThe shrinking processor feature size, lower threshold voltage and increasing clock frequency make modern processors highly vulnerable to transient faults. Architectural vulnerability factor (AVF) reflects the possibility that a transient fault eventually causes a visible error in the program output, and it indicates a system's susceptibility to transient faults. Therefore, the awareness of the AVF especially at early design stage is greatly helpful to achieve a trade-off between system performance and reliability. However, tracking the AVF during program execution is extremely costly, which makes accurate AVF prediction extraordinarily attractive to computer architects. In this paper, we propose to use boosted regression trees, a nonparametric tree-based predictive modeling scheme, to identify the correlation across workloads, execution phases and processor configurations between a key processor structure's AVF and various performance metrics. The proposed method not only makes an accurate prediction but quantitatively illustrates individual performance variable's importance to the AVF. Moreover, to reduce the prediction complexity, we also utilize a technique named patient rule induction method to extract some simple selecting rules on important metrics. Applying these rules during run time can fast identify execution intervals with a relatively high AVF. Lide Duan, Bin Li 0008, Lu Peng 0001 |
HPCA | 3 |
| 2009 | Accurate and efficient processor performance prediction via regression tree based modeling
Bin Li 0008, Lu Peng 0001, Balachandran Ramadass |
J. Syst. Archit. | 2 |
| 2008 | Efficient mart-aided modeling for microarchitecture design space exploration and performance predictionabstractComputer architects usually evaluate new designs by cycle-accurate processor simulation. This approach provides detailed insight into processor performance, power consumption and complexity. However, only configurations in a subspace can be simulated in practice due to long simulation time and limited resource, leading to suboptimal conclusions which might not be applied in a larger design space. In this paper, we propose an automated performance prediction approach which employs state-of-the-art techniques from experiment design, machine learning and data mining. Our method not only produces highly accurate estimations for unsampled points in the design space, but also provides interpretation tools that help investigators to understand performance bottlenecks. According to our experiments, by sampling only 0.02% of the full design space with about 15 millions points, the median percentage errors, based on 5000 independent test points, range from 0.32% to 3.12% in 12 benchmarks. Even for the worst-case performance, the percentage errors are within 7% for 10 out of 12 benchmarks. In addition, the proposed model can also help architects to find important design parameters and performance bottlenecks. Bin Li 0008, Lu Peng 0001, Balachandran Ramadass |
SIGMETRICS | 2 |
| 2008 | SecCMP: Enhancing Critical Secrets Protection in Chip-MultiprocessorsabstractSecurity has been considered as an important issue in processor design. Most of the existing designs of security handling assume the chip as a single secure unit. However, such assumption is vulnerable to exposure resulted from a central failure point. In this article, we propose a secure Chip-Multiprocessor architecture (SecCMP) to handle security related problems such as key protection and core authentication in multi-core systems. Matching the nature of multi-core systems, a distributed threshold secret sharing scheme is employed to protect critical secrets. A critical secret (e.g., encryption key) is divided into multiple shares and distributed among multiple cores instead of being kept a single copy in one core that is sensitive to exposure. The proposed SecCMP can not only enhance the security and fault-tolerance in secret protection but also support core authentication. SecCMP is designed to be an efficient and secure architecture for CMPs. Li Yang 0001, Lu Peng 0001, Balachandran Ramadass |
Int. J. Inf. Secur. Priv. | 2 |
| 2008 | Memory hierarchy performance measurement of commercial dual-core desktop processors
Lu Peng 0001, Jih-Kwon Peir, Tribuvan K. Prakash, Carl Staelin, Yen-Kuang Chen, David M. Koppelman |
J. Syst. Archit. | 1 |
| 2007 | Power Efficient IP Lookup with Supernode CachingabstractIn this paper, we propose a novel supernode caching scheme to reduce IP lookup latencies and energy consumption in network processors. In stead of using an expensive TCAM based scheme, we implement a set associative SRAM based cache. We organize the IP routing table as a supernode tree (a tree bitmap structure) [5]. We add a small supernode cache in-between the processor and the low level memory containing the IP routing table in a tree structure. The supernode cache stores recently visited supernodes of the longest matched prefixes in the IP routing tree. A supernode hitting in the cache reduces the number of accesses to the low level memory, leading to a fast IP lookup. According to our simulations, up to 72% memory accesses can be avoided by a 128KB supernode cache for the selected three trace files. Average supernode cache miss ratio is as low as 4%. Compared to a TCAM with the same size, 77% of energy consumption can be reduced. Lu Peng 0001, Wencheng Lu, Lide Duan |
GLOBECOM | 1 |
| 2007 | Memory Performance and Scalability of Intel's and AMD's Dual-Core Processors: A Case StudyabstractAs chip multiprocessor (CMP) has become the mainstream in processor architectures, Intel and AMD have introduced their dual-core processors to the PC market. In this paper, performance studies on an Intel Core 2 Duo, an Intel Pentium D and an AMD Athlon 64times2 processor are reported. According to the design specifications, key derivations exist in the critical memory hierarchy architecture among these dual-core processors. In addition to the overall execution time and throughput measurement using both multiprogrammed and multi-threaded workloads, this paper provides detailed analysis on the memory hierarchy performance and on the performance scalability between single and dual cores. Our results indicate that for the best performance and scalability, it is important to have (1) fast cache-to-cache communication, (2) large L2 or shared capacity, (3) fast L2 to core latency, and (4) fair cache resource sharing. Three dual-core processors that we studied have shown benefits of some of these factors, but not all of them. Core 2 Duo has the best performance for most of the workloads because of its microarchitecture features such as shared L2 cache. Pentium D shows the worst performance in many aspects due to its technology-remap of Pentium 4. Lu Peng 0001, Jih-Kwon Peir, Tribuvan K. Prakash, Yen-Kuang Chen, David M. Koppelman |
IPCCC | 1 |
| 2006 | Coterminous locality and coterminous group data prefetching on chip-multiprocessorsabstractDue to shared cache contentions and interconnect delays, data prefetching is more critical in alleviating penalties from increasing memory latencies and demands on chip-multiprocessors (CMPs). Through deep analysis of SPEC2000 applications, we find that a part of the nearby data memory references often exhibit highly-repeated patterns with long, but equal block reuse distance. These references can form a coterminous group (CG). Coterminous locality is introduced as that when a member in a CG is referenced, the remaining members will likely be referenced in the near future. Based on the coterminous locality behavior, we implement a novel CG data prefetcher on CMPs. Performance evaluations show that the proposed prefetcher can accurately cover up to 40-50% of the total misses, and result in 50-60% of potential performance improvement for several selected workload mixes Xudong Shi 0003, Jih-Kwon Peir, Lu Peng 0001, Yen-Kuang Chen, Victor W. Lee, B. Liang |
IPDPS | 4 |
| 2004 | Signature Buffer: Bridging Performance Gap between Registers and CachesabstractData communications between producer instructions and consumer instructions through memory incur extra delays that degrade processor performance. We introduce a new storage media with a novel addressing mechanism to avoid address calculations. Instead of a memory address, each load and store is assigned a signature for accessing the new storage. A signature consists of the color of the base register along with its displacement value. A unique color is assigned to a register whenever the register is updated. When two memory instructions have the same signature, they address to the same memory location. This memory signature can be formed early in the processor pipeline. A small signature buffer, addressed by the memory signature, can be established to permit stores and loads bypassing normal memory hierarchy for fast data communication. Performance evaluations based on an Alpha 21264-like pipeline using SPEC2000 integer benchmarks show that an IPC (instruction-per-cycle) improvement of 13-18% is possible using a small 8-entry signature buffer. Lu Peng 0001, Jih-Kwon Peir, Konrad Lai |
HPCA | 1 |
| 2003 | Address-free memory access based on program syntax correlation of loads and storesabstractAn increasing cache latency in next-generation processors incurs profound performance impacts in spite of advanced out-of-order execution techniques. One way to circumvent this cache latency problem is to predict load values at the onset of pipeline execution by exploiting either the load value locality or the address correlation of stores and loads. In this paper, we describe a new load value speculation mechanism based on the program syntax correlation of stores and loads. We establish a symbolic cache (SC) , which is accessed in early pipeline stages to achieve a zero-cycle load. Instead of using memory addresses, the SC is accessed by the encoding bits of base register ID plus the displacement directly from the instruction code. Performance evaluations using SPEC95 and SPEC2000 integer programs on SimpleScalar simulation tools show that the SC achieves higher prediction accuracy in comparison with other load value speculation methods, especially when hardware resources are limited. Lu Peng 0001, Jih-Kwon Peir, Qianrong Ma, Konrad Lai |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Symbolic Cache: Fast Memory Access Based on Program Syntax Correlation of Loads and StoresabstractAn increasing cache latency in next-generation processors incurs profound performance impacts in spite of advanced out-of-order execution techniques. One way to circumvent this cache latency problem is to predict the load values at the onset of pipeline execution by exploiting either the load value locality or the address correlation of stores and loads. We describe a new load value speculation mechanism based on the program syntax correlation of stores and loads. We establish a symbolic cache, which is accessed by the content of memory load and store instructions in early pipeline stages to achieve a zero-cycle load. The performance evaluation using SPEC95 and SPEC2000 integer programs with SimpleScalar tools shows that the symbolic cache provides higher accuracy than both the memory renaming and the value prediction scheme, especially when hardware resources are limited. Qianrong Ma, Jih-Kwon Peir, Lu Peng 0001, Konrad Lai |
ICCD | 3 |