Donald M. Chiarulli

dblp:26/1312 · DBLP profile ↗
← Back
30ranked-venue papers
3as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 1 first-authorSoftware engineering, systems software and programming languages · 5Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorArtificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Emerging computing paradigms · 30% Parallel and multicore computing · 20% Electronic design automation · 11%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
neuromorphic computing
0.312017
Group Scissor: Scaling Neuromorphic Computing Design to Large Neural Networks · DAC 2017
Parallel and multicore computing › many-core systems
many-core computing
0.112009
Massively parallel processing: it's déjà vu all over again · DAC 2009
Parallel and multicore computing › parallel architecture
massively parallel processing
0.112009
Massively parallel processing: it's déjà vu all over again · DAC 2009
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.112017
Group Scissor: Scaling Neuromorphic Computing Design to Large Neural Networks · DAC 2017
Interconnection networks and networks-on-chip
die-to-die interconnect
0.112007
Lightweight Error Correction Coding for System-Level Interconnects · IEEE Trans. Computers 2007
Storage systems › data representation › data encoding
error correction coding
0.112007
Lightweight Error Correction Coding for System-Level Interconnects · IEEE Trans. Computers 2007
Coding theory › error-correcting codes › block codes
nonlinear block code
0.112007
Lightweight Error Correction Coding for System-Level Interconnects · IEEE Trans. Computers 2007
Performance modeling and evaluation
simulation
0.012003
System simulation of mixed-signal multi-domain microsystems with piecewise linear models · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003
Processor architecture and microarchitecture
multicore design
0.012009
Massively parallel processing: it's déjà vu all over again · DAC 2009
Memory systems
memory access patterns
0.011997
Predicting Multiprocessor Memory Access Patterns with Learning Models · ICML 1997
Electronic design automation
system-level simulation
0.022002
A fast optical propagation technique for modeling micro-optical systems · DAC 2002
A CAD Tool for Optical MEMS · DAC 1999
Performance modeling and evaluation
workload characterization
0.011997
Predicting Multiprocessor Memory Access Patterns with Learning Models · ICML 1997
Interconnection networks and networks-on-chip › arbitration
bus arbitration
0.011994
Optoelectronic buses for high-performance computing · Proc. IEEE 1994
Interconnection networks and networks-on-chip
optical interconnection networks
0.011994
Optoelectronic buses for high-performance computing · Proc. IEEE 1994
Electronic design automation
hardware verification and test
0.011990
Timing Verification Using HDTV · DAC 1990
Electronic design automation › hardware verification and test
timing verification
0.011990
Timing Verification Using HDTV · DAC 1990
Integrated circuit design
optoelectronic integrated circuits
0.011997
Computer-Aided Design of Free-Space Opto-Electronic Systems · DAC 1997
Processor architecture and microarchitecture
instruction set architecture
0.011984
A High Performance Factoring Machine · ISCA 1984
Processor architecture and microarchitecture › arithmetic unit
arithmetic logic unit
0.011984
A High Performance Factoring Machine · ISCA 1984

Methods — techniques the papers use, named apart from their topics

structural pruning · 0.3low-rank approximation · 0.3multi-bit differential signaling · 0.2hierarchical error control coding · 0.2piecewise linear modeling · 0.0modified nodal analysis · 0.0discrete-event simulation · 0.0angular frequency optical propagation · 0.0learning models · 0.0gaussian beam propagation · 0.0
YearPublicationVenuePosition
2017 Group Scissor: Scaling Neuromorphic Computing Design to Large Neural Networks
abstract
Synapse crossbar is an elementary structure in neuromorphic computing systems (NCS). However, the limited size of crossbars and heavy routing congestion impede the NCS implementation of large neural networks. In this paper, we propose a two-step framework (namely, group scissor) to scale NCS designs to large neural networks. The first step rank clipping integrates low-rank approximation into the training to reduce total crossbar area. The second step is group connection deletion, which structurally prunes connections to reduce routing congestion between crossbars. Tested on convolutional neural networks of LeNet on MNIST database and ConvNet on CIFAR-10 database, our experiments show significant reduction of crossbar and routing area in NCS designs. Without accuracy loss, rank clipping reduces the total crossbar area to 13.62% or 51.81% in the NCS design of LeNet or ConvNet, respectively. The following group connection deletion further decreases the routing area of LeNet or ConvNet to 8.1% or 52.06%, respectively.
Yandan Wang, Wei Wen 0003, Beiye Liu, Donald M. Chiarulli, Hai Li 0001
DAC4
2016 A Simplified Phase Model for Simulation of Oscillator-Based Computing Systems
abstract
Building oscillator-based computing systems with emerging nano-device technologies has become a promising solution for unconventional computing tasks like computer vision and pattern recognition. However, simulation and analysis of these computing systems is both time and compute intensive due to the nonlinearity of new devices and the complex behavior of coupled oscillators. In order to speed up the simulation of coupled oscillator systems, we propose a simplified phase model to perform phase and frequency synchronization prediction based on a synthesis of earlier models. Our model can predict the frequency-locking behavior with several orders of magnitude speedup compared to direct evaluation, enabling the effective and efficient simulation of the large numbers of oscillators required for practical computing systems. We demonstrate the oscillator-based computing paradigm with three applications, pattern matching, convolution, and image segmentation. The simulation with these models are respectively sped up by factors of 780, 300, and 1120 in our tests.
Yan Fang 0002, Victor V. Yashin, Brandon B. Jennings, Donald M. Chiarulli, Steven P. Levitan
ACM J. Emerg. Technol. Comput. Syst.4
2014 Video analytics using beyond CMOS devices
abstract
The human vision system understands and interprets complex scenes for a variety of visual tasks in real-time while consuming less than 20 Watts of power. The holistic design of artificial vision systems that will approach and eventually exceed the capabilities of human vision systems is a grand challenge. The design of such a system needs advances in multiple disciplines. This paper focuses on advances needed in the computational fabric and provides an overview of a new-genre of architectures inspired by advances in both the understanding of the visual cortex and the emergence of devices with new mechanisms for state computations.
Narayanan Vijaykrishnan, Suman Datta, Gert Cauwenberghs, Donald M. Chiarulli, Steven P. Levitan, H.-S. Philip Wong
DATE4
2014 Feature-aided multiple-hypothesis tracking and classification of biological cells
Stefano Coraluppi, Craig Carthel, Samuel J. Dickerson, Donald M. Chiarulli, Steven P. Levitan
FUSION4
2014 Modeling oscillator arrays for video analytic applications
abstract
Weakly coupled oscillators have shown promise as a computational building block that exploits the power of emerging low power high density nano-devices such as spin torque oscillators and vanadium oxide oscillators. In this paper, we develop a new analytic phase model as well as a circuit simulation to show how clusters and arrays of coupled oscillators can be inserted into three different stages of an image processing pipeline and provide comparable image recognition performance to traditional methods.
Yan Fang 0002, Victor V. Yashin, Andrew J. Seel, Brandon B. Jennings, Reggie Barnett, Donald M. Chiarulli, Steven P. Levitan
ICCAD6
2013 Associative processing with coupled oscillators
abstract
We discuss the opportunities of performing associative processing based on the phase locking of coupled oscillators. The use of coupled oscillators, rather than Boolean logic, provides for implementations using emerging nano-technology such as magnetic spin torque oscillators and resonant body transistor oscillators, which each have the potential of lower energy operations and higher density scaling than traditional CMOS solutions.
Steven P. Levitan, Yan Fang 0002, John A. Carpenter, Chet N. Gnegy, Natalie S. Janosik, Soyo Awosika-Olumo, Donald M. Chiarulli, György Csaba, Wolfgang Porod
ISLPED7
2009 Massively parallel processing: it's déjà vu all over again
abstract
In this paper we will identify those aspects of the concurrent computing landscape that have changed since the 1980's and how those changes might impact the efficacy of parallel computing as we move from single- to multi- to many- and to massive numbers of processing cores.
Steven P. Levitan, Donald M. Chiarulli
DAC2
2007 Non-Linear Circuit Simulation using MATLAB
Steven P. Levitan, Jose A. Martinez, Donald M. Chiarulli
FDL3
2007 Lightweight Error Correction Coding for System-Level Interconnects
abstract
"Lightweight hierarchical error control coding (LHECC)" is a new class of nonlinear block codes that is designed to increase noise immunity and decrease error rate for high-performance chip-to-chip and on-chip interconnects. LHECC is designed such that its corresponding encoder and decoder logic may be tightly integrated into compact, high-speed, and low-latency I/O interfaces. LHECC operates over a new channel technology called multi-bit differential signaling (MBDS). MBDS channels utilize a physical-layer channel code called "N choose M (nCm)" encoding, where each channel is restricted to a symbol set such that half of the bits in each symbol are set to one. These symbol sets have properties that are utilized by LHECC to achieve error correction capability while requiring low or zero relative information overhead. In addition, these codes may be designed such that the latency and size of the corresponding decoders are tightly bounded. The effectiveness of these codes is demonstrated by modeling error behavior of MBDS interconnects over a range of transmission rates and noise characteristics
Jason D. Bakos, Donald M. Chiarulli, Steven P. Levitan
IEEE Trans. Computers2
2006 Nonlinear model order reduction using remainder functions
abstract
This paper describes a novel approach to the problem of model order reduction (MOR) of very large nonlinear systems. We consider the behavior of a dynamic nonlinear system as having two fundamental characteristics: a global behavioral "envelope" that describes major transformations to the state of the system under external stimuli and a local behavior that describes small perturbation responses. The nonlinear low order envelope function is generated by using the remainders from the coalescence of projection bases taken through a space-state sample. A behavioral model can then be expressed as the superposition of these two descriptions, operating according to the input stimuli and the current state value. The global behavior describes major transformations to the state of the system under external stimuli and the local behavior describes small perturbation responses. Local effects are captured by regions through a set of linear projections to a reduced state-space while global effects are captured by examining the non-commonalty among these projections. These "remainders" are used to build a modulation function that will generate the required dynamic changes in the common linear projection. The advantage of the envelope representation for strongly nonlinear systems is that it simplifies the complexity of the model into a two-part problem. Depending on the complexity or cost of the behavioral separation procedure, it can be repeated recursively
Jose A. Martinez, Steven P. Levitan, Donald M. Chiarulli
DATE3
2004 An Application of Parallel Discrete Event Simulation Algorithms to Mixed Domain System Simulation
abstract
We present our system-level co-simulation environment for mixed domain microsystems. The environment provides synchronization and co-simulation between the Chatoyant MOEMS (micro-electro mechanical systems) simulator and ModelTech ModelSim. By using shared memory IPC (inter-process communication) and PDES (parallel discrete event simulation) techniques, we achieve two orders of magnitude speedup over standard pipe/socket communication.
D. K. Reed, Steven P. Levitan, J. Boles, Jose A. Martinez, Donald M. Chiarulli
DATE5
2003 System simulation of mixed-signal multi-domain microsystems with piecewise linear models
abstract
We present a component-based multi-level mixed-signal design and simulation environment for microsystems spanning the domains of electronics, mechanics, and optics. The environment provides a solution to the problem of accurate modeling and simulation of multi-domain devices at the system level. This is achieved by partitioning the system into components that are modeled by analytic expressions. These expressions are reduced via linearization into regions of operation for each element of the component and solved with modified nodal analysis in the frequency domain, which guarantees convergence. Feedback among components is managed by a discrete event simulator sending composite signals between components. For electrical, and mechanical components, interaction is via physical connectivity while optical signals are modeled using complex scalar wavefronts, providing the accuracy necessary to model micro-optical components. Simulation speed vs. simulation accuracy can be tuned by controlling the granularity of the regions of operation of the devices, sample density of the optical wavefronts, or the time steps of the discrete event simulator. The methodology is specifically optimized for loosely coupled systems of complex components such as are found in multi-domain microsystems.
Steven P. Levitan, Jose A. Martinez, Timothy P. Kurzweg, Abhijit Davare, Mark Kahrs, Michael Bails, Donald M. Chiarulli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2002 A fast optical propagation technique for modeling micro-optical systems
abstract
As designers become more aggressive in introducing optical components to micro-systems, rigorous optical models are required for system-level simulation tools. Common optical modeling techniques and approximations are not valid for most optical micro-systems, and those techniques that provide accurate simulation are computationally slow. In this paper, we introduce an angular frequency optical propagation technique that greatly reduces computation time while achieving the accuracy of a full scalar formulation. We present simulations of a diffractive optical MEM grating light valve to show the advantages of this optical propagation method and the integration of the technique into a system-level multi-domain CAD tool.
Timothy P. Kurzweg, Steven P. Levitan, Jose A. Martinez, Mark Kahrs, Donald M. Chiarulli
DAC5
1999 A CAD Tool for Optical MEMS
abstract
Chatoyant models free-space opto-electronic components and systems and performs simulations and analyses that allow designers to make informed system level trade-offs.Recently, the use of MEM bulk and surface micro-machining technology has enabled the fabrication of micro-optical-mechanical systems.This paper presents our models for diffractive optics and new analysis techniques which extend Chatoyant to support optical MEMS design.We show these features in the simulation of two optical MEM systems.
Timothy P. Kurzweg, Steven P. Levitan, Philippe J. Marchand, Jose A. Martinez, Kurt R. Prough, Donald M. Chiarulli
DAC6
1998 Reconfigurable Processor Architectures Exploiting High Bandwidth Optical Channels
abstract
There is growing interest in studying the possibility of reconfigurable architectures as replacements for general purpose computing for certain application domains. Reconfigurable systems can take advantage of deep computational pipelines, perform concurrent execution and are inherently data flow in nature. Furthermore, these systems have the capability of 'on the fly' reconfiguration of all or portions of the hardware to represent all the functionality required to complete the execution of an application. However, these architectures suffer from slow run time reconfiguration (RTR) due to the fact that the configuration memory resides off-chip and hence requires high access latency. This disadvantage limits the system performance and the application domain in which reconfigurable systems could prove effective. To overcome slow RTR, recent approaches include on-chip configuration memory to cache the next possible configurations. This approach trades off die area for fast RTR which diminishes the processing power of the reconfigurable processor. The high cost of adding configuration cache, up to 50% of the die area, would considerably increase the number of hardware reconfigurations required compared to architectures without on-chip cache. This paper presents an alternative reconfigurable architecture which overcomes these limitations by exploiting high bandwidth optical channels. We develop a performance model to analyze and compare the performance of cache based RTR architectures, optical based RTR architectures and hybrid optical-cache based RTR architectures.
Majd F. Sakr, Steven P. Levitan, C. Lee Giles, Donald M. Chiarulli
FCCM4
1997 Computer-Aided Design of Free-Space Opto-Electronic Systems
abstract
This paper presents a system capable of static and dynamic simulationsof heterogeneous opto-electronic systems. It is capable ofmodeling Gaussian optical signal propagation with mechanicaltolerancing at the system level. We present results which demonstratethe system's ability to predict the effects of various componentparameters, such as detector geometry, and system levelparameters, such as alignment tolerances, on system performance.
Steven P. Levitan, Philippe J. Marchand, Timothy P. Kurzweg, M. A. Rempel, Donald M. Chiarulli, C. Fan, F. B. McCormick
DAC5
1997 Predicting Multiprocessor Memory Access Patterns with Learning Models
Majd F. Sakr, Steven P. Levitan, Donald M. Chiarulli, Bill G. Horne, C. Lee Giles
ICML3
1997 Identification of Defective CMOS Devices Using Correlation and Regression Analysis of Frequency Domain Transient Signal Data
abstract
Transient signal analysis is a digital device testing method that is based on the analysis of voltage transients at multiple test points and on I/sub DD/ switching transients on the supply rails. We show that it is possible to identify defective devices by analyzing the transient signals produced at test points on paths not sensitized from the defect site. The small signal variations produced at these test points are analyzed in the frequency domain. Correlation analysis shows a high degree of correlation in these signals across the outputs of defect-free devices. We use regression analysis to show the absence of correlation across the outputs of bridging and open drain defective devices.
James F. Plusquellic, Donald M. Chiarulli, Steven P. Levitan
ITC2
1996 Digital Integrated Circuit Testing using Transient Signal Analysis
abstract
A novel approach to testing CMOS digital circuits is presented that is based on an analysis of I/sub DD/ switching transients on the supply rails and voltage transients at selected test points. We present simulation and hardware experiments which show distinguishable characteristics between the transient waveforms of defective and non-defective devices. These variations are shown to exist for CMOS open drain and bridging defects, located both on and off of a sensitized path.
James F. Plusquellic, Donald M. Chiarulli, Steven P. Levitan
ITC2
1994 Dynamic Reconfiguration of Optically Interconnected Networks with Time-Division Multiplexing
Chunming Qiao, Rami G. Melhem, Donald M. Chiarulli, Steven P. Levitan
J. Parallel Distributed Comput.3
1994 Optoelectronic buses for high-performance computing
abstract
Modern computer buses are typically organized by the three functions of data transfer, addressing, and arbitration/control. In this paper we present a fiber-based bus design which provides optical solutions for each of these functions. The design includes an all-optical addressing system, based on coincident pulse addressing, which eliminates the latency contribution and bandwidth limitation associated with electronic address decoding. The control system uses time-of-flight relationships between a priority chain and a feedback waveguide to implement fully distributed asynchronous and self-timed bus arbitration.>
Donald M. Chiarulli, Steven P. Levitan, Rami G. Melhem, Manoj Bidnurkar, Robert Ditmore, Gregory Gravenstreter, Zicheng Guo, Chungming Qiao, Majd F. Sakr, James P. Teza
Proc. IEEE1
1993 Optical Computing and Interconnection Systems - Guest Editors' Introduction
Rami G. Melhem, Donald M. Chiarulli
J. Parallel Distributed Comput.2
1991 Multicasting in Optical Bus Connected Processors Using Coincident Pulse Techniques
Chunming Qiao, Rami G. Melhem, Donald M. Chiarulli, Steven P. Levitan
ICPP (1)3
1991 Pipelined Communications in Optically Interconnected Arrays
Zicheng Guo, Rami G. Melhem, Richard W. Hall, Donald M. Chiarulli, Steven P. Levitan
J. Parallel Distributed Comput.4
1990 Timing Verification Using HDTV
abstract
In this paper, we provide an overview of a system designed for verifying the consistency of timing specifications for digital circuits. The utility of the system comes from the need to verify that existing digital components will interact correctly when placed together in a system. The system can also be used in the case of verifying specifications of unimplemented components. 1
Alan R. Martello, Steven P. Levitan, Donald M. Chiarulli
DAC3
1990 Optical Bus Control for Distributed Multiprocessors
Donald M. Chiarulli, Steven P. Levitan, Rami G. Melhem
J. Parallel Distributed Comput.1
1989 Space Multiplexing of Waveguides in Optically Interconnected Multiprocessor Systems
abstract
Optical waveguides allow for enhanced bandwidth, loosened loading constraints and large physical distribution of computing resources. Moreover, optics enjoy a unique property that is not shared with electronics, namely the unidirectional propagation of signals. It is this property that is exploited in this paper to increase the effective bandwidth of optical buses. Specifically, a space-multiplexing technique for pipelined messages on optical buses is introduced and analysed. It is shown that pipelined buses support arbitrary routeing permutations in synchronous systems with only linear hardware complexity. Further, a bus arbitration protocol which extends the technique to asynchronous systems is presented. The pipelining of control and data signals represents a significant departure from the conventional exclusive access discipline which characterises bus-interconnected multiprocessors. By relaxing the exclusive access requirement, space multiplexing can support the design of large-scale, distributed, tightly coupled multiprocessor systems.
Rami G. Melhem, Donald M. Chiarulli, Steven P. Levitan
Comput. J.2
1986 Parallel Processing of Quadtrees on a Horizontally Reconfigurable Architecture Computing System
Donald M. Chiarulli, S. Sitharama Iyengar
ICPP2
1985 DRAFT: A dynamically reconfigurable processor for integer arithmetic
abstract
A special computer for high-precision arithmetic features an ALU that is dynamically reconfigurable under program control. The 256-bit ALU consists of 8 32-Mt slices each of which has its own ALU operation code in each microinstruction. The slices can remain logically separated from each other, or he dynamically connected to either or both of their neighbors under control of a segment control code that is part of each microinstruction. The micro-assembly language designed for the machine includes special features to assist in the control of the segmentation, data addressing, and control sequencing. Estimations of the times required to execute arithmetic operations on the machine show that it will be exceptionally fast for problems in computational number theory and factoring of integers.
Donald M. Chiarulli, Walter G. Rudd, Duncan A. Buell
IEEE Symposium on Computer Arithmetic1
1984 A High Performance Factoring Machine
abstract
Factoring, primality testing, and other problems of current interest [1, 2, 3, 4, 5, for example] in experimental number theory require machines with different architectures from those for any other application. Algorithms for such problems are small, but their execution requires high-speed arithmetic on very long integers interspersed with shorter precision computing that can conveniently be done by several processors acting in parallel. In this paper, we describe the preliminary design of a processor specifically for computational number theoretic problems. The primary component is a 256-bit ALU which is dynamically reconfigurable to provide for parallel independent operations on groups of 32-bit subwords within the 256-bit word.
Walter G. Rudd, Duncan A. Buell, Donald M. Chiarulli
ISCA3