VLDB 2026 Research / reviewers in the wild / expert
Henry Duwe
dblp:136/7930 · also Henry J. Duwe III
· DBLP profile ↗
28ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0003-2310-7399ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 4 since 2021Computer networks · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Microarchitecture Evaluation Framework for Transient Execution Attack Vulnerability: Metrics, Fuzzing, and Sensitivity Analysis
Jordan McGhee, Nayra Lujano, Aiden Peterson, Henry Duwe, Akhilesh Tyagi, Berk Gülmezoglu |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | PAIL: Predictable and Adaptive Intermittent Lifecycling for Robust Coordination Between Batteryless SystemsabstractBatteryless sensor nodes, powered solely by energy harvesting, are a promising alternative to battery-powered sensor nodes. However, energy harvesting rates being very low, unreliable, and time-varying, nodes cannot sustain a continuous operation, making them intermittently powered. As a result, these nodes incur unpredictable wakeup times due to continuously varying off-times. To perform tasks like distributed sensing, time synchronization, and communication for intermittently powered nodes, timely execution and robust coordination is vital. To achieve robust coordination, nodes must guarantee to be ON at a coordinated target time regardless of harvesting variations. To ensure robust coordination of on-times, we propose PAIL, a novel hardware-software approach where the hardware component enforces a constant off-time, with small variations. The software component dynamically corrects for those variations. We fabricated a PAIL sensor node and validated its ability to discover neighboring nodes. Using extensive simulations calibrated from experimental measurements, we observe that PAIL maintains nearly 99.3% coordination at steady state. In the context of pairwise communication, we show that leveraging PAIL coordination, nodes improve packet tail latency by 99%. Vishak Narayanan, Mahmoud Gshash, Vishal Deep, Mathew L. Wymore, Daji Qiao, Nathan M. Neihart, Henry Duwe |
PerCom | 7 |
| 2024 | HANNA: Harvesting-Aware Neural Network Architecture Search for Batteryless Intermittent DevicesabstractBatteryless energy-harvesting IoT devices offer maintenance-free and sustainable deployment solutions to enable pervasive intelligence even in remote areas. To develop intelligent and autonomous applications, we need to make decisions locally, often using deep neural networks (DNNs) because communication of sensed data to the cloud is expensive in terms of latency and energy. Further to enable useful applications guaranteeing timeliness while achieving desired performance such as high accuracy is imperative. However, when making decisions locally, guaranteeing timeliness alongside high accuracy is challenging due to the intertwined design of neural network models and harvested energy constraints. In this work, we propose HANNA that uses energy harvesting aware one-shot differentiable Neural Architecture Search (NAS) to produces adaptive DNN inference to maximize accuracy while maintaining timelines. Compared to other state-of-the-art approaches HANNA is able to improve the average inference accuracy by 10% to 44% while reducing the neural network search cost significantly over a range of datasets, latency bounds and harvesting deployment scenario’s. Rohit Sahu, Vishal Deep, Henry Duwe |
IPCCC | 3 |
| 2024 | Lure: A simulator for networks of batteryless intermittent nodes
Mathew L. Wymore, Rohit Sahu, Thomas Ruminski, Vishal Deep, Morgan Ambourn, Gregory Ling, Vishak Narayanan, William Asiedu, Daji Qiao, Henry Duwe |
Perform. Evaluation | 10 |
| 2023 | Case Studies in Applying Design Thinking to Course Design in Computer EngineeringabstractThis Innovative Practice Full Paper describes case studies from an instructional design process based on design thinking, illustrating tools used during stages of design. Instructional teams investigated the potential relevance of design thinking in engineering course design in electrical and computer engineering. Two teams of educators used a design thinking process in the redesign of two computer engineering courses, one in embedded systems and one in computer organization and architecture. The process of applying design thinking methods and tools was led by a facilitator with expertise in design thinking and electrical and computer engineering. The process leveraged specific tools and collaboration. This paper presents examples from each course, focusing on the design thinking tools used by the instructors and team members, highlighting what design thinking looks like when applied in this setting, and giving specific examples. The purpose is to suggest strategies and provide information and guidance for educators to use tools in their own course design efforts. Diane T. Rover, Henry Duwe, Phillip H. Jones, Nick Fila, Mani Mina |
FIE | 2 |
| 2023 | BOBBER A Prototyping Platform for Batteryless Intermittent AcceleratorsabstractBatteryless systems offer promising platforms to support pervasive, near-sensor intelligence in a sustainable manner. These systems solely rely on ambient energy sources that often provide limited power. One common approach to designing batteryless systems is using intermittent execution---a node banks energy into a capacitive store until a threshold voltage is met and the digital components turn on and consume the banked energy until the energy is depleted and they die. The limited amount of available energy demands the development of application- and domain-specific accelerators to achieve energy efficiency and timeliness. Given the extremely close relationship between volatile state and intermittent behavior, performing actual system prototyping has been critical for demonstrating feasibility of intermittent systems. However, no prototyping platform exists for intermittent accelerators. This paper introduces BOBBER, the first implementation of an intermittent FPGA-based accelerator prototyping platform. We demonstrate BOBBER in the optimization and evaluation of a neural network accelerator powered solely by RF energy harvesting. Vishak Narayanan, Rohit Sahu, Jidong Sun, Henry Duwe |
FPGA | 4 |
| 2022 | Defining and Supporting a Debugging Mindset in Computer Engineering CoursesabstractWhile it is commonly held that debugging is a critical activity for engineers, particularly computer engineers, it is rarely a core component in engineering curriculum. Often it is either considered an innate skill or one to be developed indirectly through coursework. We argue that debugging is more authentically a constellation of mindsets or intrinsic beliefs, values, and dispositions that orient behavior. We define the debugging mindset, describe a series of activities to support the development of such a mindset, and offer evidence of a debugging mindset among students in two computer engineering project courses. Henry Duwe, Diane T. Rover, Phillip H. Jones, Nick Fila, Mani Mina |
FIE | 1 |
| 2022 | Toward a Shared Sense of Time for a Network of Batteryless, Intermittently-powered NodesabstractWireless sensor nodes powered solely by energy-harvesting show promise in enabling truly pervasive, long-duration sensing by avoiding the fragility, cost, and maintenance limitations of batteries. Unfortunately, since the amount of energy harvested is often significantly less than the active consumption of the device, these devices operate intermittently with limited control of when they are on and how long they are off. Such uncontrollable intermittency first poses a challenge for traditional time synchronization error metrics, since nodes cannot reliably communicate time values at known intervals, resulting in the illusion that nodes are out-of-sync. Second, long duration off-times can exceed the inherent timing limit of persistent clocks that intermittent nodes rely on to measure off-times. Bursty groups of long off-times can cause traditional time synchronization mechanisms to re-converge slowly, incurring significant periods of high error in which nodes are effectively out-of-sync. In this paper, we define the meaning of a shared sense of time for intermittently-powered nodes, and propose two intermittency-aware synchronization error metrics. We then propose an intermittency-resilient time synchronization mechanism, called Levee, that exhibits more rapid re-convergence after losing time and a 2.12× reduction in maximum time synchronization error for a 48-hour period. Vishal Deep, Mathew L. Wymore, Daji Qiao, Henry Duwe |
IPCCC | 4 |
| 2022 | A Tale of Two IntermittenciesabstractTwo classes of intermittency have emerged in the batteryless intermittent research community: hard intermittency and soft inter-mittency. While conceptually similar, these two intermittencies represent very different approaches to the intermittency problem. In this position paper, we examine these two intermittencies in detail. We discuss the tradeoffs and evaluate the performance potential of the two intermittencies in the context of communication between intermittent nodes. Finally, we conclude that both types of intermittency have merits under different conditions and application requirements, and we argue for greater understanding of how these two disparate classes of intermittencies may interact and coexist within a single network. Mathew L. Wymore, Henry Duwe |
SenSys | 2 |
| 2021 | Constrained Conservative State Symbolic Co-analysis for Ultra-low-power Embedded SystemsabstractSymbolic simulation and symbolic execution techniques have long been used for verifying designs and testing software. Recently, using symbolic hardware-software co-analysis to characterize unused hardware resources across all possible executions of an application running on a processor has been leveraged to enable application-specific analysis and optimization techniques. Like other symbolic simulation techniques, symbolic hardware-software co-analysis does not scale well to complex applications, due to an explosion in the number of execution paths that must be analyzed to characterize all possible executions of an application. To overcome this issue, prior work proposed a scalable approach by maintaining conservative states of the system at previously-visited locations in the application. However, this approach can be too pessimistic in determining the exercisable subset of resources of a hardware design. In this paper, we propose a technique for performing symbolic co-analysis of an application on a processor's netlist by identifying, propagating, and imposing constraints from the software level onto the gate-level simulation. This produces a more precise, less pessimistic estimate of the gates that an application can exercise when executing on a processor, while guaranteeing coverage of all possible gates that the application can exercise. This also reduces the simulation time of the analysis, significantly, by eliminating the need to explore many simulation paths in the application. Compared to the state-of-art analysis based on conservative states, our constrained approach reduces the number of gates identified as exercisable by up to 34.98%, 11.52% on average, and analysis runtime by up to 84.61%, 43.83% on average. Shashank Hegde, Subhash Sethumurugan, Hari Cherupalli, Henry Duwe, John Sartori |
ASP-DAC | 4 |
| 2021 | Learning and Professional Development Through Integrated Reflective Activities in Electrical and Computer Engineering CoursesabstractThis Research-to-Practice Full Paper describes the implementation of integrated reflective activities in two computer engineering courses. Reflective activities contribute to student learning and professional development. Instructional team members have been examining the need and opportunities to deepen learning by integrating reflective activities into problem-solving experiences. We implemented reflective activities using a coordinated framework for a modified Kolbian cycle. The framework consists of reflection-for-action, reflection-in-action, reflection-on-action, and composted reflections. Reflection-for-action takes place before the experience and involves thinking about and planning future actions. Reflection-in-action takes place during the experience while actively problem-solving. Reflection-on-action takes place after the problem-solving experience. Composting involves revisiting past experiences and reflections to inform future planning. We describe the reflective activities in the context of the coordinated framework, including strategies to support reflection and increase the likelihood of engagement and success. We conclude with an analysis of the activities using the CPREE framework for reflection pathways. Diane T. Rover, Henry Duwe, Mani Mina, Nick Fila, Phillip H. Jones, Lindsey S. Sleeth |
FIE | 2 |
| 2021 | DENNI: Distributed Neural Network Inference on Severely Resource Constrained Edge DevicesabstractPervasive intelligence promises to revolutionize society from Industrial Internet of Things (IIoT), to smart infrastructure and homes, to personal health monitoring. Unfortunately, many edge devices that are pervasively embedded into infrastructure or implanted into humans are severely resource-constrained. As performing computations at the edge becomes increasingly important to meet latency deadlines and retain sensitive data locally, severe resource constraints present a challenge because many algorithms are too large to fit on a single edge device. In this paper, we focus on distributing inference for neural networks (NNs) with convolution and fully connected layers across multiple edge nodes. In order to improve efficiency on severely resource-constrained edge nodes for diverse NN architectures we present an end-to-end, automated approach, DENNI, that optimizes NN distribution with minimal nodes while meeting memory constraints. When targeting a network of edge nodes with 256KB of non-volatile memory connected with Bluetooth Low Energy, DENNI successfully distributes NN inference for a variety of machine learning algorithms across multiple edge nodes where other, static approaches cannot. Rohit Sahu, Ryan Toepfer, Matthew D. Sinclair, Henry Duwe |
IPCCC | 4 |
| 2021 | Experimental Study of Lifecycle Management Protocols for Batteryless Intermittent CommunicationabstractBatteryless energy-harvesting sensor nodes can operate indefinitely, but if the harvesting rate is too low, they must operate intermittently. Intermittent operation imposes various challenges upon the system. One of the least-studied is communication–if nodes are unpowered for long, unpredictable periods of time, how can they reliably communicate with each other? In prior work, we proposed the concept of lifecycle management protocols (LMPs) to mitigate this issue and enable wireless communication directly between intermittent sensor nodes using active radios. In this paper, we propose a design framework for a class of LMPs. We then provide analytical models for the delay and throughput of two-node communication using this framework. Finally, we implement this framework on hardware and validate our models in an experimental setting. To the best of our knowledge, this is the first design framework for, and implementation of, protocols for enabling and improving general-purpose communication between intermittent sensor nodes using active radios. Vishal Deep, Mathew L. Wymore, Alexis A. Aurandt, Vishak Narayanan, Shen Fu, Henry Duwe, Daji Qiao |
MASS | 6 |
| 2020 | Lifecycle Management Protocols for Batteryless, Intermittent Sensor NodesabstractNodes in batteryless sensor networks operate intermittently, making tasks such as node-to-node communication and coordinated computation extremely challenging. Adding to this challenge, a node typically has little control over its intermittency. Therefore, in this paper, we introduce a new class of protocols, which we call lifecycle management protocols (LMPs), to better control and manage the intermittency of batteryless nodes. These protocols may be designed and optimized for a particular task; here, we propose and evaluate a set of LMPs designed to enable direct communication between intermittent batteryless sensor nodes with active radios. Mathew L. Wymore, Vishal Deep, Vishak Narayanan, Henry Duwe, Daji Qiao |
IPCCC | 4 |
| 2020 | HARC: A Heterogeneous Array of Redundant Persistent Clocks for Batteryless, Intermittently-Powered SystemsabstractBatteryless sensing devices powered solely by ambient energy sources are expected to operate in an intermittent manner, since they do not have a predictable, or even continuous, energy supply. When such an intermittent system is powered off, it cannot keep track of time using conventional means. However, a continuous sense of time is critical for any system running real-time or time-sensitive applications. In this paper, we present HARC (Heterogeneous Array of Redundant Persistent Clocks), a novel solution to the problem of timekeeping for batteryless, intermittently-powered systems. HARC uses a heterogeneous, redundant array of capacitor-based persistent clocks that each decay in parallel, but at different rates, to provide variation-resilient high accuracy over a wide range of power off-times. We demonstrate the feasibility and effectiveness of HARC using experimental evaluations on a HARC prototype, and trace-based simulations of HARC-supported communication directly between two devices intermittently-powered by RF harvesting. Vishal Deep, Vishak Narayanan, Mathew L. Wymore, Daji Qiao, Henry Duwe |
RTSS | 5 |
| 2019 | An Adaptive Memory Management Strategy Towards Energy Efficient Machine Inference in Event-Driven Neuromorphic AcceleratorsabstractSpiking neural networks are viable alternatives to classical neural networks for edge processing in low-power embedded and IoT devices. To reap their benefits, neuromorphic network accelerators that tend to support deep networks still have to expend great effort in fetching synaptic states from a large remote memory. Since local computation in these networks is event-driven, memory becomes the major part of the system's energy consumption. In this paper, we explore various opportunities of data reuse that can help mitigate the redundant traffic for retrieval of neuron meta-data and post-synaptic weights. We describe CyNAPSE, a baseline neural processing unit and its accompanying software simulation as a general template for exploration on various levels. We then investigate the memory access patterns of three spiking neural network benchmarks that have significantly different topology and activity. With a detailed study of locality in memory traffic, we establish the factors that hinder conventional cache management philosophies from working efficiently for these applications. To that end, we propose and evaluate a domain-specific management policy that takes advantage of the forward visibility of events in a queue-based event-driven simulation framework. Subsequently, we propose network-adaptive enhancements to make it robust to network variations. As a result, we achieve 13-44% reduction in system power consumption and 8-23% improvement over conventional replacement policies. Saunak Saha, Henry Duwe, Joseph Zambreno |
ASAP | 2 |
| 2017 | Determining Application-specific Peak Power and Energy Requirements for Ultra-low Power ProcessorsabstractMany emerging applications such as IoT, wearables, implantables, and sensor networks are power- and energy-constrained. These applications rely on ultra-low-power processors that have rapidly become the most abundant type of processor manufactured today. In the ultra-low-power embedded systems used by these applications, peak power and energy requirements are the primary factors that determine critical system characteristics, such as size, weight, cost, and lifetime. While the power and energy requirements of these systems tend to be application-specific, conventional techniques for rating peak power and energy cannot accurately bound the power and energy requirements of an application running on a processor, leading to over-provisioning that increases system size and weight. In this paper, we present an automated technique that performs hardware-software co-analysis of the application and ultra-low-power processor in an embedded system to determine application-specific peak power and energy requirements. Our technique provides more accurate, tighter bounds than conventional techniques for determining peak power and energy requirements, reporting 15% lower peak power and 17% lower peak energy, on average, than a conventional approach based on profiling and guardbanding. Compared to an aggressive stressmark-based approach, our technique reports power and energy bounds that are 26% and 26% lower, respectively, on average. Also, unlike conventional approaches, our technique reports guaranteed bounds on peak power and energy independent of an application's input set. Tighter bounds on peak power and energy can be exploited to reduce system size, weight, and cost. Hari Cherupalli, Henry Duwe, Weidong Ye, Rakesh Kumar 0002, John Sartori |
ASPLOS | 2 |
| 2017 | Enabling Effective Module-Oblivious Power Gating for Embedded ProcessorsabstractThe increasingly-stringent power and energy requirements of emerging embedded applications have led to a strong recent interest in aggressive power gating techniques. Conventional techniques for aggressive power gating perform module-based power gating in processors, where power domains correspond to RTL modules. We observe that there can be significant power benefits from module-oblivious power gating, where power domains can include an arbitrary set of gates, possibly from multiple RTL modules. However, since it is not possible to infer the activity of module-oblivious power domains from software alone, conventional software-based power management techniques cannot be applied for module-oblivious power gating in processors. Also, since module-oblivious domains are not encapsulated with a well-defined port list and functionality like RTL modules, hardware-based management of module-oblivious domains is prohibitively expensive. In this paper, we present a technique for low-cost management of moduleoblivious power domains in embedded processors. The technique involves symbolic simulation-based co-analysis of a processor's hardware design and a software binary to derive profitable and safe power gating decisions for a given set of module-oblivious domains when the software binary is run on the processor. Our technique is automated, does not require programmer intervention, and incurs low management overhead. We demonstrate that module-oblivious power gating based on our technique reduces leakage energy by 2× with respect to state-of-the-art aggressive module-based power gating for a common embedded processor. Hari Cherupalli, Henry Duwe, Weidong Ye, Rakesh Kumar 0002, John Sartori |
HPCA | 2 |
| 2017 | Bespoke Processors for Applications with Ultra-low Area and Power ConstraintsabstractA large number of emerging applications such as implantables, wearables, printed electronics, and IoT have ultra-low area and power constraints. These applications rely on ultra-low-power general purpose microcontrollers and microprocessors, making them the most abundant type of processor produced and used today. While general purpose processors have several advantages, such as amortized development cost across many applications, they are significantly over-provisioned for many area- and power-constrained systems, which tend to run only one or a small number of applications over their lifetime. In this paper, we make a case for bespoke processor design, an automated approach that tailors a general purpose processor IP to a target application by removing all gates from the design that can never be used by the application. Since removed gates are never used by an application, bespoke processors can achieve significantly lower area and power than their general purpose counterparts without any performance degradation. Also, gate removal can expose additional timing slack that can be exploited to increase area and power savings or performance of a bespoke design. Bespoke processor design reduces area and power by 62% and 50%, on average, while exploiting exposed timing slack improves average power savings to 65%. Hari Cherupalli, Henry Duwe, Weidong Ye, Rakesh Kumar 0002, John Sartori |
ISCA | 2 |
| 2017 | Software-based gate-level information flow security for IoT systemsabstractThe growing movement to connect literally everything to the internet (internet of things or IoT) through ultra-low-power embedded microprocessors poses a critical challenge for information security. Gate-level tracking of information flows has been proposed to guarantee information flow security in computer systems. However, such solutions rely on non-commodity, secure-by-design processors. In this work, we observe that the need for secure-by-design processors arises because previous works on gate-level information flow tracking assume no knowledge of the application running in a system. Since IoT systems typically run a single application over and over for the lifetime of the system, we see a unique opportunity to provide application-specific gate-level information flow security for IoT systems. We develop a gate-level symbolic analysis framework that uses knowledge of the application running in a system to efficiently identify all possible information flow security vulnerabilities for the system. We leverage this information to provide security guarantees on commodity processors. We also show that security vulnerabilities identified by our analysis framework can be eliminated through software modifications at 15% energy overhead, on average, obviating the need for secure-by-design hardware. Our framework also allows us to identify and eliminate only the vulnerabilities that an application is prone to, reducing the cost of information flow security by 3.3× compared to a software-based approach that assumes no application knowledge. Hari Cherupalli, Henry Duwe, Weidong Ye, Rakesh Kumar 0002, John Sartori |
MICRO | 2 |
| 2017 | Determining Application-Specific Peak Power and Energy Requirements for Ultra-Low-Power ProcessorsabstractMany emerging applications such as the Internet of Things, wearables, implantables, and sensor networks are constrained by power and energy. These applications rely on ultra-low-power processors that have rapidly become the most abundant type of processor manufactured today. In the ultra-low-power embedded systems used by these applications, peak power and energy requirements are the primary factors that determine critical system characteristics, such as size, weight, cost, and lifetime. While the power and energy requirements of these systems tend to be application specific, conventional techniques for rating peak power and energy cannot accurately bound the power and energy requirements of an application running on a processor, leading to overprovisioning that increases system size and weight. In this article, we present an automated technique that performs hardware–software coanalysis of the application and ultra-low-power processor in an embedded system to determine application-specific peak power and energy requirements. Our technique provides more accurate, tighter bounds than conventional techniques for determining peak power and energy requirements. Also, unlike conventional approaches, our technique reports guaranteed bounds on peak power and energy independent of an application’s input set. Tighter bounds on peak power and energy can be exploited to reduce system size, weight, and cost. Hari Cherupalli, Henry Duwe, Weidong Ye, Rakesh Kumar 0002, John Sartori |
ACM Trans. Comput. Syst. | 2 |
| 2016 | Approximate bitcoin miningabstractBitcoin is the most popular cryptocurrency today. A bedrock of the Bitcoin framework is mining, a computation intensive process that is used to verify Bitcoin transactions for profit. We observe that mining is inherently error tolerant due to its embarrassingly parallel and probabilistic nature. We exploit this inherent tolerance to inaccuracy by proposing approximate mining circuits that trade off reliability with area and delay. These circuits can then be operated at Better Than Worst-Case (BTWC) to enable further gains. Our results show that approximation has the potential to increase mining profits by 30%. Matthew Vilim, Henry Duwe, Rakesh Kumar 0002 |
DAC | 2 |
| 2016 | Rescuing Uncorrectable Fault Patterns in On-Chip Memories through Error Pattern TransformationabstractVoltage scaling can effectively reduce processor power, but also reduces the reliability of the SRAM cells in on-chip memories. Therefore, it is often accompanied by the use of an error correcting code (ECC). To enable reliable and efficient memory operation at low voltages, ECCs for on-chip memories must provide both high error coverage and low correction latency. In this paper, we propose error pattern transformation, a novel low-latency error correction technique that allows on-chip memories to be scaled to voltages lower than what has been previously possible. Our technique relies on the observation that the number of on-chip memory errors that many ECCs can correct differs widely depending on the error patterns in the logical words they protect. We propose adaptively rearranging the logical bit to physical bit mapping per word according to the BIST-detectable fault pattern in the physical word. The adaptive logical bit to physical bit mapping transforms many uncorrectable error patterns in the logical words into correctable error patterns and, therefore, improving on-chip memory reliability. This reduces the minimum voltage at which on-chip memory can run by 70mV over the best low-latency ECC baseline, leading to a 25.7% core-wide power reduction for an ARM Cortex-A7-like core. Energy per instruction is reduced by 15.7% compared to the best baseline. Henry Duwe, Xun Jian 0002, Daniel Ruelas-Petrisko, Rakesh Kumar 0002 |
ISCA | 1 |
| 2016 | Bit Serializing a Microprocessor for Ultra-low-powerabstractMany emerging sensor applications are powered by energy harvesters that impose strict power constraints. These applications often do not require high performance or energy efficiency. We explore a technique for minimizing power of a microprocessor for power constrained applications: bit serial computing. Bit serial computing promises power benefits up to the data width for fully bit serializable logic. We perform a best-effort bit serialization of the openMSP430 microprocessor without making instruction set architecture (ISA) modifications. Although it is very challenging to serialize much of the logic in the microprocessor, we show that power benefits of serialization exceed 42% when the serial and parallel designs synthesized for their maximum operating frequency are running at a low duty cycle. Benefits are expected to be higher when ISA modifications are allowed. Matthew Tomei, Henry Duwe, Nam Sung Kim, Rakesh Kumar 0002 |
ISLPED | 2 |
| 2015 | Correction prediction: Reducing error correction latency for on-chip memoriesabstractThe reliability of on-chip memories (e.g., caches) determines their minimum operating voltage (Vmin) and, therefore, the power these memories consume. A strong error correction mechanism can be used to tolerate the increasing memory cell failure rate as supply voltage is reduced. However, strong error correction often incurs a high latency relative to the on-chip memory access time. We propose correction prediction where a fast mechanism predicts the result of strong error correction to hide the long latency of correction. Subsequent pipeline stages execute using the predicted values while the long latency strong error correction attempts to verify the correctness of the predicted values in parallel. We present a simple correction prediction implementation, CP, which uses a fast, but weak error correction mechanism as the correction predictor. Our evaluations for a 32KB 4-way set associative SRAM L1 cache show that the proposed implementation, CP, reduces the average cache access latency by 38%-52% compared to using a strong error correction scheme alone. This reduces the energy of a 2-issue in-order core by 16%-21%. Henry Duwe, Xun Jian 0002, Rakesh Kumar 0002 |
HPCA | 1 |
| 2014 | Better-Than-Worst-Case Design: Progress and Opportunities
Jason Cong, Henry Duwe, Rakesh Kumar 0002 |
J. Comput. Sci. Technol. | 2 |
| 2013 | Low-power, low-storage-overhead chipkill correct via multi-line error correctionabstractDue to their large memory capacities, many modern servers require chipkill correct, an advanced type of memory error detection and correction, to meet their reliability requirements. However, existing chipkill-correct solutions incur high power or storage overheads, or both because they use dedicated error-correction resources per codeword to perform error correction. This requires high overhead for correction and results in high overhead for error detection. We propose a novel chipkill-correct solution, multi-line error correction, that uses resources shared across multiple lines in memory for error correction to reduce the overhead of both error detection and correction. Our evaluations show that the proposed solution reduces memory power by a mean of 27%, and up to 38% with respect to commercial solutions, at a cost of 0.4% increase in storage overhead and minimal impact on reliability. Xun Jian 0002, Henry Duwe, John Sartori, Vilas Sridharan, Rakesh Kumar 0002 |
SC | 2 |
| 2013 | Modular Design of High-Throughput, Low-Latency Sorting UnitsabstractHigh-throughput and low-latency sorting is a key requirement in many applications that deal with large amounts of data. This paper presents efficient techniques for designing high-throughput, low-latency sorting units. Our sorting architectures utilize modular design techniques that hierarchically construct large sorting units from smaller building blocks. The sorting units are optimized for situations in which only the M largest numbers from N inputs are needed, because this situation commonly occurs in many applications for scientific computing, data mining, network processing, digital signal processing, and high-energy physics. We utilize our proposed techniques to design parameterized, pipelined, and modular sorting units. A detailed analysis of these sorting units indicates that as the number of inputs increases their resource requirements scale linearly, their latencies scale logarithmically, and their frequencies remain almost constant. When synthesized to a 65-nm TSMC technology, a pipelined 256-to-4 sorting unit with 19 stages can perform more than 2.7 billion sorts per second with a latency of about 7 ns per sort. We also propose iterative sorting techniques, in which a small sorting unit is used several times to find the largest values. Amin Farmahini Farahani, Henry Duwe, Michael J. Schulte, Katherine Compton |
IEEE Trans. Computers | 2 |