EDBT 2026 Demo / reviewers in the wild / expert
Jeffrey T. Draper
dblp:87/3834 · also Jeff Draper, Jeffrey Draper
· DBLP profile ↗
52ranked-venue papers
4as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 51 · 4 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Hardware reliability and fault tolerance · 28% Parallel and multicore computing · 17% Hardware accelerators and domain-specific architectures · 16% |
Topics — the 29 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
transactional memory |
0.4 | 3 | 2016 | Improving Utilization of Hardware Signatures in Transactional Memory · IEEE Trans. Parallel Distributed Syst. 2013 In-network traffic regulation for Transactional Memory · HPCA 2013 A Filtering Mechanism to Reduce Network Bandwidth Utilization of Transaction Execution · ACM Trans. Archit. Code Optim. 2016 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2019 | HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Emerging computing paradigms › approximate and stochastic computing
stochastic computing |
0.4 | 1 | 2019 | HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
stochastic computing accelerator |
0.4 | 1 | 2019 | HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Electronic design automation
hardware verification and test |
0.2 | 1 | 2016 | Accelerating soft-error-rate (SER) estimation in the presence of single event transients · DAC 2016 |
Memory systems › memory bandwidth management
memory bandwidth reduction |
0.2 | 1 | 2016 | A Filtering Mechanism to Reduce Network Bandwidth Utilization of Transaction Execution · ACM Trans. Archit. Code Optim. 2016 |
Hardware reliability and fault tolerance › radiation effects
radiation-induced faults |
0.2 | 1 | 2016 | SEU Mitigation and Validation of the LEON3 Soft Processor Using Triple Modular Redundancy for Space Processing · FPGA 2016 |
Hardware reliability and fault tolerance › soft errors
single-event upset |
0.2 | 1 | 2016 | SEU Mitigation and Validation of the LEON3 Soft Processor Using Triple Modular Redundancy for Space Processing · FPGA 2016 |
Hardware reliability and fault tolerance
soft errors |
0.2 | 1 | 2016 | Accelerating soft-error-rate (SER) estimation in the presence of single event transients · DAC 2016 |
Hardware reliability and fault tolerance › soft errors
soft error rate estimation |
0.2 | 1 | 2016 | Accelerating soft-error-rate (SER) estimation in the presence of single event transients · DAC 2016 |
Hardware reliability and fault tolerance › redundancy › modular redundancy
triple modular redundancy |
0.2 | 1 | 2016 | SEU Mitigation and Validation of the LEON3 Soft Processor Using Triple Modular Redundancy for Space Processing · FPGA 2016 |
Parallel and multicore computing › transactional memory
hardware transactional memory |
0.2 | 2 | 2016 | In-network traffic regulation for Transactional Memory · HPCA 2013 A Filtering Mechanism to Reduce Network Bandwidth Utilization of Transaction Execution · ACM Trans. Archit. Code Optim. 2016 |
Parallel and multicore computing › transactional memory
conflict detection |
0.2 | 1 | 2013 | Improving Utilization of Hardware Signatures in Transactional Memory · IEEE Trans. Parallel Distributed Syst. 2013 |
Electronic design automation › hardware security
hardware signatures |
0.2 | 1 | 2013 | Improving Utilization of Hardware Signatures in Transactional Memory · IEEE Trans. Parallel Distributed Syst. 2013 |
Interconnection networks and networks-on-chip
traffic management |
0.2 | 1 | 2013 | In-network traffic regulation for Transactional Memory · HPCA 2013 |
Energy-efficient computing › energy-efficient machine learning
energy-efficient inference |
0.1 | 1 | 2019 | HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Integrated circuit design › digital circuit design
combinational logic |
0.1 | 1 | 2016 | Accelerating soft-error-rate (SER) estimation in the presence of single event transients · DAC 2016 |
Hardware reliability and fault tolerance › soft errors
single event transient |
0.1 | 1 | 2016 | Accelerating soft-error-rate (SER) estimation in the presence of single event transients · DAC 2016 |
Reconfigurable computing and FPGAs › FPGA-based processor implementation
soft-core processor |
0.1 | 1 | 2016 | SEU Mitigation and Validation of the LEON3 Soft Processor Using Triple Modular Redundancy for Space Processing · FPGA 2016 |
Reconfigurable computing and FPGAs › FPGA architecture
SRAM-based FPGA |
0.1 | 1 | 2016 | SEU Mitigation and Validation of the LEON3 Soft Processor Using Triple Modular Redundancy for Space Processing · FPGA 2016 |
Memory systems › cache coherence
directory-based coherence |
0.0 | 1 | 2013 | In-network traffic regulation for Transactional Memory · HPCA 2013 |
Processor architecture and microarchitecture
multicore design |
0.0 | 1 | 2013 | Improving Utilization of Hardware Signatures in Transactional Memory · IEEE Trans. Parallel Distributed Syst. 2013 |
Memory systems
memory bandwidth |
0.0 | 1 | 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999 |
Memory systems
processing-in-memory |
0.0 | 1 | 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999 |
Memory systems › memory interface
processor-memory interface |
0.0 | 1 | 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999 |
Interconnection networks and networks-on-chip
network interface |
0.0 | 1 | 1997 | A Bus-Efficient Low-Latency Network Interface for the PDSS Multicomputer · HPDC 1997 |
Electronic design automation › physical design
routing |
0.0 | 1 | 1997 | A Bus-Efficient Low-Latency Network Interface for the PDSS Multicomputer · HPDC 1997 |
Hardware accelerators and domain-specific architectures
irregular application acceleration |
0.0 | 1 | 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999 |
Memory systems
cache coherence |
0.0 | 1 | 1997 | A Bus-Efficient Low-Latency Network Interface for the PDSS Multicomputer · HPDC 1997 |
Methods — techniques the papers use, named apart from their topics
weight clustering · 0.4stochastic computing · 0.4pipelining · 0.4approximate parallel counter · 0.4top-down memoization · 0.2orbit failure rate estimation · 0.2heavy ion radiation testing · 0.2fault injection · 0.2conflict prediction · 0.2SET propagation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Normalization and dropout for stochastic computing-based deep convolutional neural networks
Ji Li 0006, Zhe Li 0001, Ao Ren, Caiwen Ding, Jeffrey T. Draper, Shahin Nazarian, Qinru Qiu, Bo Yuan 0001, Yanzhi Wang 0001 |
Integr. | 6 |
| 2019 | HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural NetworksabstractDeep convolutional neural networks (DCNNs) are one of the most promising deep learning techniques and have been recognized as the dominant approach for almost all recognition and detection tasks. The computation of DCNNs is memory intensive due to large feature maps and neuron connections, and the performance highly depends on the capability of hardware resources. With the recent trend of wearable devices and Internet of Things, it becomes desirable to integrate the DCNNs onto embedded and portable devices that require low power and energy consumptions and small hardware footprints. Recently stochastic computing (SC)-DCNN demonstrated that SC as a low-cost substitute to binary-based computing radically simplifies the hardware implementation of arithmetic units and has the potential to satisfy the stringent power requirements in embedded devices. In SC, many arithmetic operations that are resource-consuming in binary designs can be implemented with very simple hardware logic, alleviating the extensive computational complexity. It offers a colossal design space for integration and optimization due to its reduced area and soft error resiliency. In this paper, we present HEIF, a highly efficient SC-based inference framework of the large-scale DCNNs, with broad applications including (but not limited to) LeNet-5 and AlexNet, that achieves high energy efficiency and low area/hardware cost. Compared to SC-DCNN, HEIF features: 1) the first (to the best of our knowledge) SC-based rectified linear unit activation function to catch up with the recent advances in software models and mitigate degradation in application-level accuracy; 2) the redesigned approximate parallel counter and optimized stochastic multiplication using transmission gates and inverse mirror adders; and 3) the new optimization of weight storage using clustering. Most importantly, to achieve maximum energy efficiency while maintaining acceptable accuracy, HEIF considers holistic optimizations on cascade connection of function blocks in DCNN, pipelining technique, and bit-stream length reduction. Experimental results show that in large-scale applications HEIF outperforms previous SC-DCNN by the throughput of 4.1×, by area efficiency of up to 6.5×, and achieves up to 5.6× energy improvement. Zhe Li 0001, Ji Li 0006, Ao Ren, Ruizhe Cai, Caiwen Ding, Xuehai Qian, Jeffrey T. Draper, Bo Yuan 0001, Jian Tang 0008, Qinru Qiu, Yanzhi Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2017 | Structural design optimization for deep convolutional neural networks using stochastic computingabstractDeep Convolutional Neural Networks (DCNNs) have been demonstrated as effective models for understanding image content. The computation behind DCNNs highly relies on the capability of hardware resources due to the deep structure. DCNNs have been implemented on different large-scale computing platforms. However, there is a trend that DCNNs have been embedded into light-weight local systems, which requires low power/energy consumptions and small hardware footprints. Stochastic Computing (SC) radically simplifies the hardware implementation of arithmetic units and has the potential to satisfy the small low-power needs of DCNNs. Local connectivities and down-sampling operations have made DCNNs more complex to be implemented using SC. In this paper, eight feature extraction designs for DCNNs using SC in two groups are explored and optimized in detail from the perspective of calculation precision, where we permute two SC implementations for inner-product calculation, two down-sampling schemes, and two structures of DCNN neurons. We evaluate the network in aspects of network accuracy and hardware performance for each DCNN using one feature extraction design out of eight. Through exploration and optimization, the accuracies of SC-based DCNNs are guaranteed compared with software implementations on CPU/GPU/binary-based ASIC synthesis, while area, power, and energy are significantly reduced by up to 776x, 190x, and 32835x. Zhe Li 0001, Ao Ren, Ji Li 0006, Qinru Qiu, Bo Yuan 0001, Jeffrey T. Draper, Yanzhi Wang 0001 |
DATE | 6 |
| 2017 | Deadline-Aware Joint Optimization of Sleep Transistor and Supply Voltage for FinFET Based Embedded SystemsabstractLeakage power consumption has recently become a great concern for modern embedded systems. FinFET technologies, power gating, and near- and super-threshold regimes can significantly reduce the power consumption. However, there lacks a comprehensive analysis of jointly applying the aforementioned power saving techniques. In this paper, we investigate the application of power gating to FinFET circuits operating in near- and super-threshold voltage regimes for embedded system applications. A joint optimization algorithm is proposed to determine the width/length, position and threshold type of the sleep transistor together with the operating voltage constrained to a certain deadline, and with the goal of minimizing energy per operation. Experimental results demonstrate that the proposed algorithm achieves up to 99.9% energy reductions when compared to the near-threshold approach without power gating and 95.3% when compared to deadline-free optimization. Huimei Cheng, Ji Li 0006, Jeffrey T. Draper, Shahin Nazarian, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | Softmax Regression Design for Stochastic Computing Based Deep Convolutional Neural NetworksabstractRecently, Deep Convolutional Neural Networks (DCNNs) have made tremendous advances, achieving close to or even better accuracy than human-level perception in various tasks. Stochastic Computing (SC), as an alternate to the conventional binary computing paradigm, has the potential to enable massively parallel and highly scalable hardware implementations of DCNNs. In this paper, we design and optimize the SC based Softmax Regression function. Experiment results show that compared with a binary SR, the proposed SC-SR under longer bit stream can reach the same level of accuracy with the improvement of 295X, 62X, 2617X in terms of power, area and energy, respectively. Binary SR is suggested for future DCNNs with short bit stream length input whereas SC-SR is recommended for longer bit stream. Ji Li 0006, Zhe Li 0001, Caiwen Ding, Ao Ren, Bo Yuan 0001, Qinru Qiu, Jeffrey T. Draper, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 8 |
| 2017 | Hardware-driven nonlinear activation for stochastic computing based deep convolutional neural networksabstractRecently, Deep Convolutional Neural Networks (DCNNs) have made unprecedented progress, achieving the accuracy close to, or even better than human-level perception in various tasks. There is a timely need to map the latest software DCNNs to application-specific hardware, in order to achieve orders of magnitude improvement in performance, energy efficiency and compactness. Stochastic Computing (SC), as a low-cost alternative to the conventional binary computing paradigm, has the potential to enable massively parallel and highly scalable hardware implementation of DCNNs. One major challenge in SC based DCNNs is designing accurate nonlinear activation functions, which have a significant impact on the network-level accuracy but cannot be implemented accurately by existing SC computing blocks. In this paper, we design and optimize SC based neurons, and we propose highly accurate activation designs for the three most frequently used activation functions in software DCNNs, i.e, hyperbolic tangent, logistic, and rectified linear units. Experimental results on LeNet-5 using MNIST dataset demonstrate that compared with a binary ASIC hardware DCNN, the DCNN with the proposed SC neurons can achieve up to 61X, 151X, and 2X improvement in terms of area, power, and energy, respectively, at the cost of small precision degradation. In addition, the SC approach achieves up to 21X and 41X of the area, 41X and 72X of the power, and 198200X and 96443X of the energy, compared with CPU and GPU approaches, respectively, while the error is increased by less than 3.07%. ReLU activation is suggested for future SC based DCNNs considering its superior performance under a small bit stream length. Ji Li 0006, Zhe Li 0001, Caiwen Ding, Ao Ren, Qinru Qiu, Jeffrey T. Draper, Yanzhi Wang 0001 |
IJCNN | 7 |
| 2017 | Accelerated Soft-Error-Rate (SER) Estimation for Combinational and Sequential CircuitsabstractRadiation-induced soft errors have posed an increasing reliability challenge to combinational and sequential circuits in advanced CMOS technologies. Therefore, it is imperative to devise fast, accurate and scalable soft error rate (SER) estimation methods as part of cost-effective robust circuit design. This paper presents an efficient SER estimation framework for combinational and sequential circuits, which considers single-event transients (SETs) in combinational logic and multiple cell upsets (MCUs) in sequential elements. A novel top-down memoization algorithm is proposed to accelerate the propagation of SETs, and a general schematic and layout co-simulation approach is proposed to model the MCUs for redundant sequential storage structures. The feedback in sequential logic is analyzed with an efficient time frame expansion method. Experimental results on various ISCAS85 combinational benchmark circuits demonstrate that the proposed approach achieves up to 560.2X times speedup with less than 3% difference in terms of SER results compared with the baseline algorithm. The average runtime of the proposed framework on a variety of ISCAS89 benchmark circuits is 7.20s, and the runtime is 119.23s for the largest benchmark circuit with more than 3,000 flip-flops and 17,000 gates. Ji Li 0006, Jeffrey T. Draper |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | Accelerating soft-error-rate (SER) estimation in the presence of single event transientsabstractRadiation-induced soft errors have posed an ever increasing reliability challenge as device dimensions keep shrinking in advanced CMOS technology. Therefore, it is imperative to devise fast and accurate soft error rate (SER) estimation methods. Previous works mainly focus on improving the accuracy of the SER results, whereas the speed improvement is limited to partitioning and parallel processing. This paper presents an efficient SER estimation framework for combinational logic circuits in the presence of single-event transients (SETs). A novel top-down memoization algorithm is proposed to accelerate the propagation of SETs. Experimental results of a variety of benchmark circuits demonstrate that the proposed approach achieves up to 560.2X times speedup with less than 3% difference in terms of SER results compared with the baseline algorithm. Ji Li 0006, Jeffrey T. Draper |
DAC | 2 |
| 2016 | SEU Mitigation and Validation of the LEON3 Soft Processor Using Triple Modular Redundancy for Space ProcessingabstractProcessors are an essential component in most satellite payload electronics and handle a variety of functions including command handling and data processing. There is growing interest in implementing soft processors on commercial FPGAs within satellites. Commercial FPGAs offer reconfigurability, large logic density, and I/O bandwidth; however, they are sensitive to ionizing radiation and systems developed for space must implement single-event upset mitigation to operate reliably. This paper investigates the improvements in reliability of a LEON3 soft processor operating on a SRAM-based FPGA when using triple-modular redundancy and other processor-specific mitigation techniques. The improvements in reliability provided by these techniques are validated with both fault injection and heavy ion radiation tests. The fault injection experiments indicate an improvement of 51× and the radiation testing results demonstrate an average improvement of 10×. Orbit failure rate estimations were computed and suggest that the TMR LEON3 processor has a mean-time to failure of over 76 years in a geosynchronous orbit. Michael J. Wirthlin, Andrew M. Keller, Chase McCloskey, Parker Ridd, David S. Lee, Jeffrey T. Draper |
FPGA | 6 |
| 2016 | A Filtering Mechanism to Reduce Network Bandwidth Utilization of Transaction ExecutionabstractHardware Transactional Memory (HTM) relies heavily on the on-chip network for intertransaction communication. However, the network bandwidth utilization of transactions has been largely neglected in HTM designs. In this work, we propose a cost model to analyze network bandwidth in transaction execution. The cost model identifies a set of key factors that can be optimized through system design to reduce the communication cost of HTM. Based on the model and network traffic characterization of a representative HTM design, we identify a huge source of superfluous traffic due to failed requests in transaction conflicts. As observed in a spectrum of workloads, 39% of the transactional requests fail due to conflicts, which renders 58% of the transactional network traffic futile. To combat this pathology, a novel in-network filtering mechanism is proposed. The on-chip router is augmented to predict conflicts among transactions and proactively filter out those requests that have a high probability to fail. Experimental results show the proposed mechanism reduces total network traffic by 24% on average for a set of high-contention TM applications, thereby reducing energy consumption by an average of 24%. Meanwhile, the contention in the coherence directory is reduced by 68%, on average. These improvements are achieved with only 5% area added to a conventional on-chip router design. Lihang Zhao, Lizhong Chen, Woojin Choi, Jeffrey T. Draper |
ACM Trans. Archit. Code Optim. | 4 |
| 2014 | Consolidated conflict detection for hardware transactional memoryabstractHardware Transactional Memory (HTM) promises to ease multithreaded parallel programming with uncompromised performance. Microprocessors supporting HTM implement a conflict detection mechanism to detect data access conflicts between transactions. Understanding the on-chip network bandwidth utilization of such mechanisms is important as the energy and latency cost of routing packets across the chip is growing alarmingly. We investigate the communication characteristics of a typical conflict detection mechanism. A variety of traffic overheads are identified, which accounts for a combined 56% of the total transactional traffic in a wide spectrum of applications. To combat this problem, we propose C2D (Consolidated Conflict Detection), a novel micro-architectural technique to consolidate conflict detection to a logically central (but physically distributed) agent to reduce the bandwidth utilization of conflict detection. Full system evaluation shows that the proposed technique, if applied to conventional eager conflict detection, can reduce 35% of the traffic and hence 27% of the network energy. The consolidated eager conflict detection generates less traffic than a lazy conflict detection scheme thereby closing the gap between bandwidth utilization of eager and lazy conflict detection. Lihang Zhao, Jeffrey T. Draper |
PACT | 2 |
| 2014 | Mitigating the Mismatch between the Coherence Protocol and Conflict Detection in Hardware Transactional MemoryabstractHardware Transactional Memory (HTM) usually piggybacks onto the cache coherence protocol to detect data access conflicts between transactions. We identify an intrinsic mismatch between the typical coherence scheme and transaction execution, which causes a sizable amount of unnecessary transaction aborts. This pathological behavior is called false aborting and increases the amount of wasted computation and on-chip communication. For the TM applications we studied, 41% of the transactional write requests incur false aborting. To combat false aborting, we propose Predictive Unicast and Notification (PUNO), a novel hardware mechanism to 1) replace the inefficient coherence multicast with a unicast scheme to prevent transactions from being disrupted unnecessarily and 2) restrain transaction polling through proactive notification. PUNO reduces transaction aborts by 61% and network traffic by 32% in workloads representative of future TM applications with a VLSI implementation area overhead of 0.41%. Lihang Zhao, Lizhong Chen, Jeffrey T. Draper |
IPDPS | 3 |
| 2014 | Optimal techniques for assigning inter-tier signals to 3D-vias with path control in a 3DICabstract3-dimensional integrated circuit (3DIC) technology is one of the most promising solution to meet the ever-increasing demands for higher device integration and energy-efficiency, while remaining cost-effective, for the current semiconductor industry. In a 3DIC, signals between the tiers are interconnected using top metal layer bondpoints (micro-bumps) or through-silicon-vias (TSVs). These vertical connections are called 3Dvias. Similar to I/O signals, inter-tier signals assigned to 3Dvias influence the standard cell placement in a tier, making this assignment critical for an efficient 3DIC layout. Unlike I/O signals, these inter-tier signals are very large in number and manual assignment (like assigning I/O signals to pins) is impractical and calls for automated techniques. This paper introduces several techniques that provide an optimum assignment and path control with an ability to control critical path lengths for interconnecting inter-tier signals. Ten 3DICs of two designs (5 each) are built using the proposed techniques and an architecture-driving manual assignment. The proposed techniques successfully automate the assignment process, and results show that they achieve up to 9.4% lower total wire length and 10.4% shorter average net length compared to a manual method, and up to 12% lower total and average net length compared to prior work. Gopi Neela, Jeffrey T. Draper |
ISCAS | 2 |
| 2013 | Multiobjective Optimization of Cost, Performance and Thermal Reliability in 3DICsabstractIn this paper we propose a fast and efficient method of multiobjective optimization for 3DIC building block placement. The objectives are cost, performance and thermal reliability. Our proposed method is based on a Quasi-Newton analytical optimization method. Our approach also uses a scalarization method of Compromise Programming, in which the weighted distance of the objectives from their minimum points is optimized. A fast 3DIC thermal map model is used in the optimization algorithm to eliminate the thermal analysis bottleneck during the optimization iterations. In comparison with previous multiobjective optimizations for sample 3DIC configurations, our method reduces the peak temperature by 4.3% and total wire length by 5.7% while it is more than 17x faster in the optimization runtime, on average. Fatemeh Kashfi, Jeffrey T. Draper |
DSD | 2 |
| 2013 | An asymmetric adaptive-precision energy-efficient 3DIC multiplierabstractDecades of research in optimizing multipliers for power, speed, and area efficiency, and the continuing push for further enhancing multipliers reflects their importance. Energy-efficient computing requires optimization of every single logic component and hence, an energy-efficient 64-bit multiplier is proposed. This design exploits the highly frequent occurrence of low-precision operands, by dynamically adapting to asymmetric precision to save energy. Depending on input operands, it functions in three modes: 32x32, 64x64, or asymmetric 32x64/64x32. Results show that the asymmetric multiplier uses up to 42.4% less switching energy, and is overall up to 33.4% more energy-efficient than the baseline design. A 3-dimensional integrated circuit (3DIC) version of the asymmetric multiplier is also introduced. It is designed using a "design for 3D" methodology to mitigate some of the challenges of 3DIC. The average interconnect length and chip footprint of the 3DIC are 17.3% and 46.5% smaller than the traditional IC implementation. Gopi Neela, Jeffrey T. Draper |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | In-network traffic regulation for Transactional MemoryabstractHardware Transactional Memory (HTM) promises to simplify parallel programming on shared-memory chip multiprocessors by providing atomic execution of code blocks. Concurrently, Networks-On-Chip (NOCs) have emerged as an efficient on-chip communication infrastructure but have been largely neglected in HTM designs. In this work, we explore the interaction between the HTM paradigm and NOCs. In the process, we find a huge source of unnecessary network traffic incurred by transactional requests that are unsuccessful. This problem is identified as false forwarding that adversely affects network performance and energy efficiency. Surprisingly, 39% (up to 79% for a specific workload) of the transactional requests have incurred false forwarding over a wide spectrum of workloads. To combat this problem, we propose TMNOC, a novel approach that exploits the co-design of HTM and NOCs to mitigate false forwarding. Transactional requests that have a high probability to fail are filtered out in-network as early as possible to save energy and improve concurrency in the memory system. Experimental results show that our design reduces total network traffic by 20% on average (up to 40%) for a set of high-contention benchmarks representative of future TM workloads, thereby reducing energy consumption by an average of 24% (up to 39%). Meanwhile, the contention in the coherence directory is reduced by 66% on average. These improvements are achieved with only 5% area overhead added to a conventional on-chip router design. Lihang Zhao, Woojin Choi, Lizhong Chen, Jeffrey T. Draper |
HPCA | 4 |
| 2013 | Logic-on-logic partitioning techniques for 3-dimensional integrated circuitsabstractDiminishing returns from transistor scaling, increasing interconnect delay, and the need for high device density and high energy efficiency are pushing the semiconductor industry in the direction of 3-dimensional integrated circuits (3DICs). A homogeneously integrated, logic-on-logic stacked 3DIC has the potential to be a cost-effective solution for these challenges as well as for chip security. However, this emerging integration platform currently suffers from a lack of standard CAD tool support and a 3DIC design flow to build efficient 3DICs. The work presented in this paper proposes new design partitioning techniques to smartly split any given design across various layers to build a 3DIC. These techniques reduce the search space for optimal partitioning by several times depending on the design. Further, a “design for 3D” approach is introduced, which can be used to build powerful custom 3DICs. Finally, as an example, two different 3DICs of a floating point unit are implemented using the proposed partitioning methods and the design flow. Results show an increase in speed of up to 7% and 41.5% reduction in chip footprint for even this small design. Gopi Neela, Jeffrey T. Draper |
ISCAS | 2 |
| 2013 | Implementation of hybrid version management in hardware transactional memoryabstractTransactional Memory (TM) has been proposed as an alternative to locks to simplify parallel programming. While most research has focused on the architectural support of TM, it is at least equally important to investigate the hardware implementation of the key mechanisms of TM to facilitate its deployment in commercial processors. In this paper, we present a light-weight and generic implementation of the key structures to support hybrid version management in Hardware Transactional Memory (HTM). Full-system simulation demonstrates the performance advantage of our design (up to 40% improvement). Synthesis results show that the hardware structures incur an area overhead of only 0.4% and a power overhead of 1.65% to the Sun Rock processor design. Lihang Zhao, Jeffrey T. Draper |
ISCAS | 2 |
| 2013 | Improving Utilization of Hardware Signatures in Transactional MemoryabstractThe transactional memory (TM) paradigm promises to increase programmer productivity by making it easier to write correct parallel programs. In fulfilling this goal, a TM system should maximize its performance with limited hardware resources. Conflict detection is an essential element for maintaining correctness among concurrent transactions in a TM system. Hardware signatures have been proposed as an area-efficient method for detecting conflicts. However, signatures can degrade TM performance by falsely declaring conflicts. Hence, improving the accuracy of signatures within a given hardware budget is a crucial issue for TM to be adopted as a mainstream programming model. In this paper, we propose a simple and effective signature design, the unified signature. Instead of using separate read- and write-signatures, we implement a single signature to track all read- and write-accesses. By merging read- and write-signatures, a unified signature can effectively enlarge the signature coverage without additional overhead. Within the constraints of a given hardware budget, a TM system with a unified signature outperforms a baseline system with the same-sized traditional signatures by reducing the number of falsely detected conflicts. Even though the unified signature scheme incurs read-read dependencies, we show that these false dependencies do not negate the benefit of unified signatures and can effectively be filtered out. A TM system with a 2-Kbit unified signature with a helper signature scheme achieves speedups of 15 percent over baseline TM with 33 percent less area and 49 percent less power. Woojin Choi, Jeffrey T. Draper |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | Mileage-based contention management in transactional memoryabstractIn Transactional Memory (TM), a conflict occurs when a memory block is accessed concurrently by two or more transactions and at least one of them is a write access. The management of conflicts significantly impacts TM performance. There are two alternative approaches for managing conflicts: Reactive Contention Management (RCM) [1] and Proactive Contention Management (PCM) [2]. Previous contention management schemes treat all transactions with no weights, and make a decision based on the information provided by the running transaction instance. Woojin Choi, Lihang Zhao, Jeffrey T. Draper |
PACT | 3 |
| 2012 | TMNOC: a case of HTM and NoC co-design for increased energy efficiency and concurrencyabstractHardware Transactional Memory (HTM) designs must implement conflict detection to guarantee the correctness of transaction execution. A conflict occurs when more than one transaction access the same data and at least one of them attempts to modify the data. The corresponding conflict detection mechanism usually works at a cacheline level that fits naturally into the cache coherence protocol. Thus, the inter-transaction communication for conflict detection is usually mapped onto the coherence communication controlled by the directory-based coherence protocols. In this paper, we identify inefficiency introduced by such mappings. The net effect of such inefficiency is excessive on-chip network traffic that consumes substantial dynamic power as packets are switched over the routers and links. We present TMNOC, a HTM and Network-on-Chip (NoC) co-design to improve network energy efficiency. The on-chip network, instead of a passive communication substrate, proactively filters out transactional requests that waste energy yet having no contribution to the progress of transactions. Experiment results show that TMNOC reduces energy consumption of the on-chip network by 14.5% on average (up to 38%) across a wide range of transaction applications. Lihang Zhao, Woojin Choi, Jeffrey T. Draper |
PACT | 3 |
| 2012 | SEL-TM: Selective Eager-Lazy Management for Improved Concurrency in Transactional MemoryabstractHardware Transactional Memory (HTM) systems implement version management and conflict detection in hardware to guarantee that each transaction is atomic and executes in isolation. In general, HTM implementations fall into two categories, namely, eager systems and lazy systems. Lazy systems have been shown to exploit more concurrency from potentially conflicting transactions. However, lazy systems manage a transaction's entire write set lazily, which gives rise to two main disadvantages: (a) a complex cache protocol and implementation are required to maintain the speculative modifications, and, (b) the latency of committing the entire write set often leads to severe performance degradation of the whole system. It is observed in a wide range of workloads that more than 55% of the transaction aborts are due to conflicts on only three memory blocks. Thus we argue that an eager HTM system can achieve the same level of concurrency as lazy systems by managing only a small portion of a transaction's write set lazily. In this paper, we present Selective-Eager-Lazy HTM (SEL-TM), a new HTM implementation to adopt complementary version management schemes within a transaction whose write set is divided into eagerly- and lazily-managed memory addresses at runtime. An intelligent hardware scheme is designed to select the memory addresses for lazy management as well as determining whether each dynamic instance of a transaction benefits from hybrid management. Experimental results using the STAMP benchmarks show that, on average, SEL-TM improves performance by 14% over an eager system and 22% over a lazy system. The speedup demonstrates that our design is capable of harvesting the concurrency benefit of lazy version management while avoiding some of the performance penalties in lazy HTMs. Lihang Zhao, Woojin Choi, Jeffrey T. Draper |
IPDPS | 3 |
| 2011 | Unified Signatures for Improving Performance in Transactional MemoryabstractTransactional Memory (TM) promises to increase programmer productivity by making it easier to write correct parallel programs. In fulfilling this goal, a TM system should maximize its performance with limited hardware resources. Conflict detection is an essential element for maintaining correctness among concurrent transactions in a TM system. Hardware signatures have been proposed as an area-efficient method for detecting conflicts. However, signatures can degrade TM performance by falsely declaring conflicts. Hence, increasing the quality of signatures within a given hardware budget is a crucial issue for TM to be adopted as a mainstream programming model. In this paper, we propose a simple and effective signature design, unified signature. Instead of using separate read- and write-signatures, as is often done in TM systems, we implement a single signature to track all read- and write-accesses. By merging read- and write-signatures, a unified signature can effectively enlarge the signature size without additional overhead. Within the constraints of a given hardware budget, a TM system with a unified signature outperforms a baseline system with the same hardware budget by reducing the number of falsely detected conflicts. Even though the unified signature scheme incurs read-after-read dependencies, we show that these false dependencies do not negate the benefit of unified signatures for practical signature sizes. A TM system with 2K-bit unified signatures achieves average speedups of 22% over baseline TM systems. Woojin Choi, Jeffrey T. Draper |
IPDPS | 2 |
| 2010 | Cubic Ring Networks: A Polymorphic Topology for Network-on-ChipabstractAs chip multiprocessors transition from multi-core to many-core, on-chip network power is increasingly becoming a key barrier to scalability. Studies have shown that on-chip networks can consume up to 36% of the total chip power, while analysis of network traffic reveals that for extended periods of execution time, network load is well below the network capacity in many applications. In recent studies, researchers have proposed to exploit this temporal variability in network traffic to dynamically turn off links, buffers and segments of the on-chip routers. In this work, we make the case for a polymorphic topology, called Cubic Ring (cRing), that allows dynamically turning off over 30% of resources in a 2D network (and more in higher dimensional networks), with less than 5% increase in average distance. As a result, cRing networks provide an elegant way to trade off network bandwidth for lower (static) power. A complete formalism for the proposed cRing topologies and the associated routing algorithm is presented, along with evaluation under synthetic workloads. Bilal Zafar 0002, Jeffrey T. Draper, Timothy M. Pinkston |
ICPP | 2 |
| 2010 | Locality-aware adaptive grain signatures for Transactional MemoriesabstractTransactional Memory (TM) has attracted considerable attention because it promises to increase programmer productivity by making it easier to write correct parallel programs. To maintain correctness in the face of concurrency, detecting conflicts among simultaneously running transactions is an essential element. Hardware signatures have been proposed as an area-efficient mechanism for conflict detection. A signature can summarize an unbounded amount of addresses and misses no conflicts, but could falsely declare conflicts even when no true conflict exists (false positives) due to aliasing and occupancy. Previous signature designs assume that false positives are destructive to performance and attempt to reduce the total number of false positives. In this paper, we show that some false positives can be helpful to performance by triggering the early abortion of a transaction which would encounter a true conflict later anyway. Based on this observation, we propose an adaptive grain signature to improve performance by dynamically changing the range of address keys based on the history. With the use of adaptive grain signatures, we can increase the number of performance-friendly false positives as well as decrease the number of performance-destructive false positives. Woojin Choi, Jeffrey T. Draper |
IPDPS | 2 |
| 2010 | Implementation of adaptive grain signatures for transactional memoriesabstractHardware signatures for Transactional Memory (TM) systems have been proposed as an efficient mechanism for conflict detection, an essential element in TM for maintaining correctness. A signature misses no conflicts, but could falsely declare conflicts even when no true conflict exists (false positives). In this paper, we show that some false positives can be helpful to the performance by triggering the early abortion of a transaction which would encounter a true conflict later anyway. We propose an adaptive grain signature to improve TM performance by dynamically changing the range of address keys based on the history. With architecture-level simulation and Verilog HDL implementation, we demonstrate that a TM system with our design frequently outperforms baseline TM systems, with marginal area overhead. Woojin Choi, Young Hoon Kang, Taek-Jun Kwon, Jeffrey T. Draper |
ISCAS | 4 |
| 2010 | A single-event upset hardening technique for high speed MOS Current Mode LogicabstractIn this paper, we introduce a SEU-hard MOS Current Mode Logic (MCML) sequential element that is used in high speed communication systems. We have implemented latches and flip-flops in 65 nm technology and show that the critical charge needed to upset the sensitive nodes in these circuits is increased with this proposed design. Simulation has been conducted for clock rates of 0.5, 1, 2 and 4 GHz. The results show the critical charge increases more than 5 times (>440%) with this design while the delay (32 ps) is acceptable for GHz operations. Mahta Haghi, Jeffrey T. Draper |
ISCAS | 2 |
| 2010 | Fault-Tolerant Flow Control in On-chip NetworksabstractScaling of interconnects exacerbates the already challenging reliability of on-chip networks. Although many researchers have provided various fault handling techniques in chip multi-processors (CMPs), the fault-tolerance of the interconnection network is yet to adequately evolve. As an end-to-end recovery approach delays fault detection and complicates recovery to a consistent global state in such a system, a link-level retransmission is endorsed for recovery, making a higher-level protocol simple. In this paper, we introduce a fault-tolerant flow control scheme for soft error handling in on-chip networks. The fault-tolerant flow control recovers errors at a link-level by requesting retransmission and ensures an error-free transmission on a flit-basis with incorporation of dynamic packet fragmentation. Dynamic packet fragmentation is adopted as a part of fault-tolerant flow control to disengage flits from the fault-containment and recover the faulty flit transmission. Thus, the proposed router provides a high level of dependability at the link-level for both datapath and control planes. In simulation with injected faults, the proposed router is observed to perform well, gracefully degrading while exhibiting 97% error coverage in datapath elements. The proposed router has been implemented using a TSMC 45 nm standard cell library. As compared to a router which employs triple modular redundancy (TMR) in datapath elements, the proposed router takes 58% less area and consumes 40% less energy per packet on average. Young Hoon Kang, Taek-Jun Kwon, Jeffrey T. Draper |
NOCS | 3 |
| 2009 | The effect of design parameters on single-event upset sensitivity of MOS current mode logicabstractIn this paper, we describe and discuss the effects of design parameters such as transistor size, output voltage swing and bias current on radiation sensitivity of MOS current mode logic (MCML) type sequential elements that are used in high-speed communication systems. We have implemented latches and flip-flops in 90 nm technology and show how single-event upset can be mitigated just by adjusting particular design factors at the same clock frequency. It is shown that the critical charge needed to upset the logic state of a sequential element increases up to 5 times by increasing the bias current at the cost of more power and up to 2 times by increasing output voltage swing at the cost of more area. The effect of changing operation frequency from 500MHz to 4GHz on single-event upset is also investigated. For frequencies higher than 2 GHz, critical charge improves 1.3 times. Mahta Haghi, Jeffrey T. Draper |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | Multicast routing with dynamic packet fragmentationabstractNetworks-on-Chip (NoCs) become a critical design factor as chip multiprocessors (CMPs) and systems on a chip (SoCs) scale up with technology. With fundamental benefits of high bandwidth and scalability in on-chip networks, a newly added multicast capability can further enhance the performance by reducing the network load and facilitate coherence protocols of many-core CMPs [10]. This paper proposes a novel multicast router with dynamic packet fragmentation in on-chip networks. Packet fragmentation is performed to avoid deadlock in blocking situations, releasing the hold of an output virtual channel (VC) and allowing another packet to use the freed VC. From circuit simulation of the design implemented with IBM 90nm technology, the proposed router reduces latency by 38.6% and consumes 9% less energy than a unicast baseline router at the baseline saturation. Young Hoon Kang, Jeff Sondeen, Jeffrey T. Draper |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | Dynamic packet fragmentation for increased virtual channel utilization in on-chip routersabstractConventional packet-switched on-chip routers provide good resource sharing while minimizing latencies through various techniques. A virtual channel (VC) is allocated on a per-packet basis and held until the entire packet exits the VC buffer. This sometimes leads to inefficient use of VCs at high network loads. A blocked packet can affect adjacent routers, resulting in a congestion propagation effect. In such a scenario, VC buffers may be empty although they are regarded as fully occupied by a blocked packet. This paper proposes a dynamic packet fragmentation technique which releases empty VC buffers by fragmenting packets and allowing other packets to use the freed VC buffers. Thus, fragmentation increases VC utilization. Simulation experiments show performance improvement in terms of latency and throughput up to 20% and 7.5%, respectively. Young Hoon Kang, Taek-Jun Kwon, Jeffrey T. Draper |
NOCS | 3 |
| 2007 | Performance Evaluation of Probe-Send Fault-tolerant Network-on-chip RouterabstractWith increasing reliability concerns for current and next generation VLSI technologies, fault-tolerance is fast becoming an integral part of system-on-chip and multi-core architectures. Another trend for such architectures is network-on-chip (NoC) becoming a standard for on-chip global communication. In an earlier work, a generic fault-tolerant routing algorithm in the context of NoCs has been presented. The proposed routing algorithm works in two phases, namely path exploration (PE) and normal communication. This paper presents fundamental insights into various novel PE approaches, their feasibility and performance trade-offs for k-ary 2-cube NoCs. The dependence of the normal communication phase on the probability of finding paths and their quality in the first phase emphasizes the PE's significance. One major contribution of this work is the investigation of application of constrained randomness to PE for optimizing the quality of paths. Another contribution is the proposed use of merging of traffic to reduce the reconfiguration time by a large amount (73.8% on an average). Sumit D. Mediratta, Jeffrey T. Draper |
ASAP | 2 |
| 2007 | Critical charge and set pulse widths for combinational logic in commercial 90nm cmos technologyabstractThis work presents an efficient hybrid simulation approach, developed for accurate characterization of single-event transients (SETs) in combinational logic. Using this approach, we show that charges as small as 3.5fC can introduce transients in commercial 90nm CMOS technology, hence increasing the likelihood of SET-induced soft errors. SET pulse-widths as large as 942ps are predicted at an LET (Linear Energy Transfer) of 60MeV-cm2/mg. Process-corner variations are shown to modulate SET pulse-widths by up-to 75%. The results suggest that selection of mitigation techniques for SET radiation-hardened circuits cannot exclusively rely on baseline process analyses, as they might grossly underestimate the true SET risk to the design. Riaz Naseer, Jeffrey T. Draper, Younes Boulghassoul, Sandeepan DasGupta, Art Witulski |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Characterization of a Fault-tolerant NoC RouterabstractWith increasing reliability concerns for current and next generation VLSI technologies, fault-tolerance is fast becoming an integral part of system-on-chip (SoC) and multi-core architectures. Another concern for these architectures is increasing global wire lengths with associated issues leading to network-on-chips (NoC) becoming standard for on-chip global communication. We recognize these issues and present an on-chip generic fault-tolerant routing algorithm. The microarchitecture of a NoC router implementing the proposed routing algorithm for a k-ary 2-cube topology is provided. The proposed router works in two phases. In the first phase, the network is explored for an existing path between source-destination pairs after reset or during system reconfiguration after fault detection. Existing paths are cached and used in the second phase of data communication during normal system operation. The presented router architecture also proposes a concept of dynamic multiplexing of virtual channels on physical channels to efficiently utilize physical channel bandwidth. The above approaches complement each other and when combined together, result in an efficiently realizable high-performance NoC fault-tolerant router. An implementation characterization of this k-ary 2-cube torus router in terms of area, power and critical path delay in IBM Cu-08 technology is presented, along with bandwidth and latency characterization for relevant cases. Sumit D. Mediratta, Jeffrey T. Draper |
ISCAS | 2 |
| 2007 | Critical Charge Characterization for Soft Error Rate Modeling in 90nm SRAMabstractDue to continuous technology scaling, the reduction of nodal capacitances and the lowering of power supply voltages result in an ever decreasing minimal charge capable of upsetting the logic state of memory circuits. In this paper the authors investigate the critical charge (Qcrit) required to upset a 6T SRAM cell designed in a commercial 90nm process. The authors characterize Qcritusing different current models and show that there are significant differences in Qcritvalues depending on which models are used. Discrepancies in critical charge characterization are shown to result in under-predictions of the SRAM's associated soft error rate as large as two orders of magnitude. For accurate Qcritcalculation, it is critical that 3D device simulation is used to calibrate the current pulse modeling heavy ion strikes on the circuit, since the stimuli characteristics are technology feature size dependant. Current models with very fast characteristic timing parameters are shown to result in conservative soft error rate predictions; and can assertively be used to model ion strikes when 3D simulation data is not available. Riaz Naseer, Younes Boulghassoul, Jeffrey T. Draper, Sandeepan DasGupta, Art Witulski |
ISCAS | 3 |
| 2006 | 2 Gbps SerDes design based on IBM Cu-11 (130nm) standard cell technologyabstractThis paper introduces a standard cell based design for a Serializer and Deserializer (SerDes) communication link. The proposed design is area, power and design time efficient as compared to conventional SerDes Designs, making it very attractive for modest budget multi-core and multi-processor ASICs with wide communication buses that are difficult to accommodate within the pin count of commonly available packaging. The design employs a “Statistical Random Sampling Technique ” to observe and adjust the synchronization and serialization signals at start up rather than using a resource-heavy PLL or DLL based frequency multiplier/synthesizer and clock data recovery circuits. The serialization and deserialization logic is based on standard cell technology that makes the design highly portable. Multiple serial lines are bundled with a strobe that is used as a reference signal for deserialization. Data-to-strobe timing skew is compensated by adjusting the launch times of strobe and data symbols at the sender side. The edges of the strobe are set within the eye of data symbols to have maximum timing margin, which makes the design inherently tolerant of jitter. Power consumption of the proposed SerDes design is 30 mW per serial link targeted to IBM Cu-11(130 nm) Technology, nearly a 2.5x improvement over the conventional design with a 60 % less area requirement. Rashed Zafar Bhatti, Monty Denneau, Jeffrey T. Draper |
ACM Great Lakes Symposium on VLSI | 3 |
| 2006 | A double-data rate (DDR) processing-in-memory (PIM) device with wideword floating-point capabilityabstractThe data-intensive architecture (DIVA) system incorporates processing-in-memory (PIM) chips as smart-memory coprocessors to a microprocessor. This architecture exploits inherent memory bandwidth both on chip and across the system to target several classes of bandwidth-limited applications. A recently developed PIM chip in TSMC 0.18/spl mu/m technology incorporates a DDR SDRAM interface for its inclusion in commodity systems, such as the HP zx6000 workstation used on this project. Each PIM chip includes eight single-precision floating-point units (FPU) in the wideword pipeline, enabling significant speedups in the target system. This paper focuses on the integration of new subcomponents into the PIM chip design, system integration, and measured system results, demonstrating the significant GFLOP/W feature offered by PIM computing. Tim Barrett, Sumit D. Mediratta, Taek-Jun Kwon, Ravinder Singh, Sachit Chandra, Jeff Sondeen, Jeffrey T. Draper |
ISCAS | 7 |
| 2006 | Phase measurement and adjustment of digital signals using random sampling techniqueabstractThis paper introduces a technique to measure and adjust the relative phase of on-chip high speed digital signals using a random sampling technique of inferential statistics. The proposed technique as applied to timing uncertainty mitigation in the signaling of a digital system is presented as an example; the relative phase information is used to minimize the timing skew. The proposed circuit captures the state of the signals under measurement simultaneously at random instants of time and gathers a large sample data to estimate the relative phase between the signals. By carefully premeditating the sample size, the accuracy and confidence of the result can be set to a level as high as desired. Accurately sensed value of relative phase enables the correction circuit to reduce the maximum correction error, less than half the maximum delay resolution unit available for adjustment. A pure standard cell based circuit design approach is used that reduces the overall design time and circuit complexity. The test results of the proposed circuit manifest a very close correlation to the simulated and theoretically expected results. The random sampling unit (RSU) circuit proposed for phase measurement in this paper occupies 3350 (mum)2area in 130nm technology, which is an order of magnitude smaller than what is required for its analog equivalent in the same technology Rashed Zafar Bhatti, Monty Denneau, Jeffrey T. Draper |
ISCAS | 3 |
| 2006 | DF-DICE: a scalable solution for soft error tolerant circuit designabstractThe delay filtered dual interlocked storage cell (DF-DICE) offers a scalable solution in different radiation environments for soft error mitigation. The area and speed performance for five different single event transient thresholds have been evaluated. The results show that the cost of soft error mitigation is minimal for terrestrial environments (overall area penalty less than 14% and speed penalty within 6% for flip-flop based typical designs) while it is larger for space environments (overall area penalty up to 30% and speed penalty up to 13% for flip-flop based typical designs). The logic of a conventional application specific integrated circuit (ASIC) can easily be converted to a soft-error tolerant design by replacing the existing storage elements with the respective DF-DICE elements Riaz Naseer, Jeffrey T. Draper |
ISCAS | 2 |
| 2005 | Performance Analysis of User-Level PIM Communication in the Data IntensiVe Architecture (DIVA) System
Sumit D. Mediratta, Jeffrey T. Draper |
HiPC | 2 |
| 2003 | Voltage-pulse driven harmonic resonant rail drivers for low-power applicationsabstractWe describe a new design technique for efficient harmonic resonant rail drivers. The proposed circuit implementation is coupled to a standard pulse source and uses only discrete passive components and no external dc power supply. It can thus be externally tuned to minimize the consumed power in the target IC. A new design technique based on current-fed voltage pulse-forming network theory is proposed to find the value of each discrete component for a target frequency and a given load capacitance. The proposed circuit topology can be used to generate any desired periodic 50% duty-cycle waveform by superimposing multiple harmonics of the desired waveform, however, this paper focuses on the generation of trapezoidal-wave clock signals. We have tested the driver with a capacitive load between 38.3 and 97.8 pF with clock frequency ranging between 0.8 and 15 MHz. The overall power dissipation for our second-order harmonic rail driver is 19% of fC/sub L/V/sup 2/ at 15 MHz and 97.8 pF load. Joong-Seok Moon, William C. Athas, Sigfrid D. Soli, Jeffrey T. Draper, Peter A. Beerel |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2002 | Implementation of a 32-bit RISC Processor for the Data-Intensive Architecture Processing-In-Memory ChipabstractThe Data-Intensive Architecture(DIVA) system employs Processing-In-Memory(PIM) chips as smart-memory coprocessors to a micorprocessor. This architecture exploits inherent memory bandwidth both on chip and across the system to target several classes of bandwidth- limited applications, including multimedia applications and pointer-based and sparse-matrix computations. The DIVA project is building a prototype workstation-class system using PIM chips in place of standard DRAMs to demonstrate these concepts. We have recently completed initial testing of the rst version of the prototype PIM device. A key component of this architecture is the scalar processor that coordinates all activ-ity within a PIM node. Since such a component is present in each PIM node,we exploit parallelism to achieve significant speedups rather than relying on costly, high-performance processor design. The resulting scalar processor is then an in-order 32-bit RISC microcontroller that is extremely area-efficient. This paper details the design and implementation of this scalar processor in TSMC 0.18cm technology. In conjunction with other publications, this paper demonstrates that impressive gains can be achieved with very little "smart" logic added to memory devices. Jeffrey T. Draper, Jeff Sondeen, Sumit D. Mediratta, Ihn Kim |
ASAP | 1 |
| 2002 | The architecture of the DIVA processing-in-memory chipabstractThe DIVA (Data IntensiVe Architecture) system incorporates a collection of Processing-In-Memory (PIM) chips as smart-memory co-processors to a conventional microprocessor. We have recently fabricated prototype DIVA PIMs. These chips represent the first smart-memory devices designed to support virtual addressing and capable of executing multiple threads of control. In this paper, we describe the prototype PIM architecture. We emphasize three unique features of DIVA PIMs, namely, the memory interface to the host processor, the 256-bit wide datapaths for exploiting on-chip bandwidth, and the address translation unit. We present detailed simulation results on eight benchmark applications. When just a single PIM chip is used, we achieve an average speedup of 3.3X over host-only execution, due to lower memory stall times and increased fine-grain parallelism. These 1-PIM results suggest that a PIM-based architecture with many such chips yields significantly higher performance than a multiprocessor of a similar scale and at a much reduced hardware cost. Jeffrey T. Draper, Jacqueline Chame, Mary W. Hall, Craig S. Steele, Tim Barrett, Jeff LaCoss, John J. Granacki, Chun Chen 0002, Chang Woo Kang, Ihn Kim Gokhan |
ICS | 1 |
| 2000 | A Zener-diode-activated ESD protection circuit for sub-micron CMOS processesabstractA Zener-diode-activated electrostatic discharge (ESD) protection circuit is implemented in a 0.5 /spl mu/m CMOS process. This ESD circuit uses a substrate p-n-p transistor with its base connected to a Zener diode to discharge the electrostatic energy. The Zener diode implementation utilizes a silicide block capability to avoid short circuits in the active area. Its performance has been tested by the human body model and a high-speed ESD test. Its latchup-free characteristic makes it an ideal circuit for ESD protection of I/O pads, ESD clamping between power rails, poly-antenna effect protection, and overshoot attenuation. Louis Luh, John Choma Jr., Jeffrey T. Draper |
ISCAS | 3 |
| 2000 | Performance optimization for high-order continuous-time ΣΔ modulators with extra loop delayabstractExtra loop delay in a continuous-time modulator can cause a stability problem, especially when the modulator uses a high-order single-loop architecture and operates at a high sampling rate. This paper investigates the impact of extra loop delay on the performance of high-order (single-loop) modulators. A solution to compensate this delay and thereby optimize performance is proposed in this paper. The circuit architecture for this solution is also presented to facilitate a practical realization. Louis Luh, John Choma Jr., Jeffrey T. Draper |
ISCAS | 3 |
| 1999 | Area-Efficient Area Pad Design for High Pin-Count ChipsabstractThis paper presents an area pad layout method to efficiently reduce the space required for interconnection pads and pad drivers. Unlike peripheral pads, area pads use only the top metal layer and therefore allow active circuitry to be laid out underneath. With identical functional elements grouped together, a group of pad drivers share the same well and can be placed tightly together. The use of silicided diffusion reduces the well contact to diffusion contact spacing requirement. By taking advantage of this spacing requirement and using serpentine gate layout, a driver's size can be effectively reduced without reducing the driving capacity. An embedded multicomputer router interface chip has been implemented using these techniques and has achieved 554 pads in a 9 mm/spl times/6 mm chip with a 0.8 /spl mu/m single-poly 3-metal N-well CMOS process. Louis Luh, John Choma Jr., Jeffrey T. Draper |
Great Lakes Symposium on VLSI | 3 |
| 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive ArchitectureabstractProcessing-in-memory (PIM) chips that integrate processor logic into memory devices offer a new opportunity for bridging the growing gap between processor and memory speeds, especially for applications with high memory-bandwidth requirements.The Data-IntensiVe Architecture (DIVA) system combines PIM memories with one or more external host processors and a PIM-to-PIM interconnect.DIVA increases memory bandwidth through two mechanisms: (1) performing selected computation in memory, reducing the quantity of data transferred across the processor-memory interface; and (2) providing communication mechanisms called parcels for moving both data and computation throughout memory, further bypassing the processor-memory bus.DIVA uniquely supports acceleration of important irregular applications, including sparse-matrix and pointer-based computations.In this paper, we focus on several aspects of DIVA designed to effectively support such computations at very high performance levels: (1) the memory model and parcel definitions; (2) the PIM-to-PIM interconnect; and, (3) requirements for the processor-to-memory interface.We demonstrate the potential of PIMbased architectures in accelerating the performance of three irregular computations, sparse conjugate gradient, a natural-join database operation and an object-oriented database query. Mary W. Hall, Peter M. Kogge, Jefferey G. Koller, Pedro C. Diniz, Jacqueline Chame, Jeffrey T. Draper, Jeff LaCoss, John J. Granacki, Jay B. Brockman, Apoorv Srivastava, William C. Athas, Vincent W. Freeh, Joonseok Park |
SC | 6 |
| 1998 | A Continuous-Time Switched-Current Sigma-Delta Modulator with Reduced Loop DelayabstractA novel architecture for a second-order continuous-time switched-current /spl Sigma//spl Delta/ modulator is presented. The loop delay is reduced by predicting the states of the second integrator and feeding the predicted states to the comparator. The predicted states are generated by summing three scaled current mode signals. A gain-manager is used to accurately control the integrator gain to generate the predicted states and stabilize the system. A newly designed high-speed current-mode comparator is capable of summing the three scaled current inputs and comparing them. With a 50 MHz sampling rate, it has achieved 60 dB dynamic range (10-bit) at 1 MHz. The modulator has been fabricated in a 2 /spl mu/m CMOS process with an active area of 0.37 mm/sup 2/. The power dissipation is 16.6 mW from a 5 V single power supply. Louis Luh, John Choma Jr., Jeffrey T. Draper |
Great Lakes Symposium on VLSI | 3 |
| 1997 | A Bus-Efficient Low-Latency Network Interface for the PDSS MulticomputerabstractThe Packaging-Driven Scalable Systems multicomputer (PDSS) project uses several innovative interconnect and routing techniques to construct a low-latency, high-bandwidth (1.3 GB/s) multicomputer network. The PDSS network interface provides a low-latency interface between the network and the processing nodes that allows unprivileged code to initiate network operations while maintaining a high level of protection. The interface design exploits processor-bus cache coherence protocols to deliver very-low-latency cache-to-cache communications between processing nodes. Network operations include a variety of transfers of cache-line-sized packets, including remote read and write, and a distributed barrier-synchronization mechanism. Despite performance-limiting flaws, the initial single-chip implementation of the network router and interface achieves gigabit/s bandwidth and microsecond cache-to-cache latencies between nodes using commodity processor and memory components. Craig S. Steele, Jeffrey T. Draper, Jefferey G. Koller, C. LaCour |
HPDC | 2 |
| 1994 | A Comprehensive Analytical Model for Wormhole Routng in Multicomputer Systems
Jeffrey T. Draper, Joydeep Ghosh |
J. Parallel Distributed Comput. | 1 |
| 1994 | The M-Cache: A Message-Handling Mechanism for Multicomputer Systems
Jeffrey T. Draper, Joydeep Ghosh |
Parallel Comput. | 1 |
| 1993 | Performance Evaluation of a Parallel I/O Subsystem for Hypercube Multicomputers
Joydeep Ghosh, Kelvin D. Goveas, Jeffrey T. Draper |
J. Parallel Distributed Comput. | 3 |