Ralph K. Cavin III

dblp:15/240 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
0since 2021 · last 2018
0000-0002-5810-5660ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 11 · 4 first-authorArtificial intelligence and machine learning · 2Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
15 papers
Emerging computing paradigms · 43% Memory systems · 20% Interconnection networks and networks-on-chip · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 77% Computational science and engineering · 23%

Topics — the 30 heaviest of 51, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
beyond-CMOS computing
0.432012
Science and Engineering Beyond Moore's Law · Proc. IEEE 2012
Prolog to the Section on Science and Engineering Beyond Moore's Law · Proc. IEEE 2012
Nanoelectronics Research for Beyond CMOS Information Processing · Proc. IEEE 2010
Memory systems
DRAM
0.112010
Memory Devices: Energy-Space-Time Tradeoffs · Proc. IEEE 2010
Memory systems
emerging memory technologies
0.112010
Memory Devices: Energy-Space-Time Tradeoffs · Proc. IEEE 2010
Interconnection networks and networks-on-chip
optical interconnection networks
0.112010
Device and Architecture Outlook for Beyond CMOS Switches · Proc. IEEE 2010
Emerging computing paradigms
quantum computer architecture
0.112010
Device and Architecture Outlook for Beyond CMOS Switches · Proc. IEEE 2010
Memory systems › non-volatile memory
resistive memory
0.112010
Memory Devices: Energy-Space-Time Tradeoffs · Proc. IEEE 2010
Emerging computing paradigms
bio-inspired computing
0.012012
Science and Engineering Beyond Moore's Law · Proc. IEEE 2012
Integrated circuit design
technology scaling
0.012012
Prolog to the Section on Science and Engineering Beyond Moore's Law · Proc. IEEE 2012
Integrated circuit design
analog and mixed-signal circuits
0.012003
Low-power design methodology for an on-chip bus with adaptive bandwidth capability · DAC 2003
Processor architecture and microarchitecture › general-purpose processor architecture
von neumann architecture
0.012003
Limits to binary logic switch scaling - a gedanken model · Proc. IEEE 2003
Electronic design automation › physical design › clock network synthesis
clock skew optimization
0.022000
Integrated parametric timing optimization of digital systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Timing constraints for wave-pipelined systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Integrated circuit design › technology scaling
CMOS scaling
0.012010
Nanoelectronics Research for Beyond CMOS Information Processing · Proc. IEEE 2010
Performance modeling and evaluation
system-level analysis
0.012010
Memory Devices: Energy-Space-Time Tradeoffs · Proc. IEEE 2010
Electronic design automation › logic synthesis › sequential circuit optimization
retiming
0.012000
Integrated parametric timing optimization of digital systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Electronic design automation › physical design
timing optimization
0.012000
Integrated parametric timing optimization of digital systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Integrated circuit design
digital circuit design
0.011994
Timing constraints for wave-pipelined systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Electronic design automation
timing analysis
0.011994
Timing constraints for wave-pipelined systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Embedded and real-time systems › timing constraints
timing constraint specification
0.011994
Timing constraints for wave-pipelined systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Processor architecture and microarchitecture › pipelining › pipeline design
wave pipelining
0.011994
Timing constraints for wave-pipelined systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Integrated circuit design
ASIC design
0.011990
Design of integrated circuits: directions and challenges · Proc. IEEE 1990
Integrated circuit design › ASIC design
ASIC design methodology
0.011990
Design of integrated circuits: directions and challenges · Proc. IEEE 1990
Interconnection networks and networks-on-chip › routing algorithms
fault-tolerant routing
0.011989
Hamiltonian Cycles in the Shuffle-Exchange Network · IEEE Trans. Computers 1989
Interconnection networks and networks-on-chip
hamiltonian cycle
0.011989
Hamiltonian Cycles in the Shuffle-Exchange Network · IEEE Trans. Computers 1989
Interconnection networks and networks-on-chip › switching network › multistage interconnection network
shuffle-exchange network
0.011989
Hamiltonian Cycles in the Shuffle-Exchange Network · IEEE Trans. Computers 1989
Image and video processing
feature extraction
0.011988
Bit-level concurrency in real-time geometric feature extractions · CVPR 1988
Image and video processing › feature extraction
geometric feature extraction
0.011988
Bit-level concurrency in real-time geometric feature extractions · CVPR 1988
Hardware accelerators and domain-specific architectures
image processing accelerator
0.011988
Exploiting Bit Level Concurrency in Real-Time Geometric Feature Extractions · ISCA 1988
Integrated circuit design
VLSI design
0.011988
Exploiting Bit Level Concurrency in Real-Time Geometric Feature Extractions · ISCA 1988
Integrated circuit design › semiconductor device fabrication
CMOS technology
0.011986
A perspective on CMOS technology trends · Proc. IEEE 1986
Integrated circuit design › low-power circuit design
low-power CMOS design
0.011986
A perspective on CMOS technology trends · Proc. IEEE 1986

Methods — techniques the papers use, named apart from their topics

energy-space-time tradeoff analysis · 0.1statistical analysis of address streams · 0.0gedanken model · 0.0circuit-level power analysis · 0.0linear programming · 0.0mixed integer linear programming · 0.0delay modeling · 0.0tree structure mapping · 0.0pipelined processor design · 0.0combinatorial analysis · 0.0stochastic imbedding · 0.0power series expansion · 0.0circular coverage function · 0.0bessel function elimination · 0.0policy iteration · 0.0markov decision theory · 0.0
YearPublicationVenuePosition
2018 Nurturing the Growth of a National Infrastructure in Emerging Technologies
abstract
A fundamental responsibility of the leadership of a sovereign government is to create policies that enable the overall wellbeing of its citizens. This responsibility has many dimensions but certainly wealth generation through industrial competitiveness is a key component. Today a nation’s economic strength rests not only on ownership of saleable natural resources but increasingly on the technical capabilities and innovation capacities of its citizens. This is particularly true in an environment where technologies are constantly emerging to serve as the foundation for new industries. How does a relatively small nation-state, but with significant resources and aspirations, compete and excel in this environment? This article briefly describes an approach taken by the Abu Dhabi Emirate of the United Arab Emirates (UAE) to foster national competitiveness in selected emerging technologies in which the nation had previously had very little presence.
Rafic Z. Makki, Victor V. Zhirnov, Ralph K. Cavin III, Sami Issa, Marco Iansiti
Proc. IEEE3
2012 Prolog to the Section on Science and Engineering Beyond Moore's Law
abstract
One of the remarkable technical achievements of the past 40 years has been the advances in complexity of the integrated circuit, driven primarily by an exponential rate of reduction in transistor feature sizes, thereby enabling the creation of electronic systems with steadily increasing functionality. The societal benefits in terms of economic growth from Moore's law scaling depend on a knowledge of learning curve cost reductions enabled by Moore's law. It might be possible to increase design efficiency to reduce nonrecurring engineering costs or to develop more cost-effective packaging technologies. The basic physics upon which a transistor operates might be altered 'to utilize other physical phenomena, such as electron spin, magnetic dipoles, photons, etc., to develop new classes of devices that would preserve Moore's law benefits.
Ralph K. Cavin III, Paolo Lugli, Victor V. Zhirnov
Proc. IEEE1
2012 Science and Engineering Beyond Moore's Law
abstract
In this paper, the historical effects and benefits of Moore's law for semiconductor technologies are reviewed, and it is offered that the rapid learning curve obtained to the benefit of society by feature size scaling might be continued in several different ways. The problem is that as features approach the range of a few nanometers, electron-based devices depart radically from the ideal switch and, in fact, become very leaky in the off state. It is argued that there are some short-term solutions involving more highly parallel manufacturing, increased design efficiency, and lower cost packaging technologies that could continue the steep learning curve for cost reductions that have historically been achieved via Moore's Law scaling. Another alternative might be to increase chip functionality by integrating devices that offer broadened chip functionality including, e.g., sensors, energy sources, oscillators, etc. A third alternative would be to invent an entirely new information processing state variable based on different physics, using electron spin, magnetic dipoles, photons, etc., to improve the performance and reduce switching energy for devices whose smallest features are on the order of a few nanometers. Each of these alternatives is being actively explored and an overview of each strategy and progress to date is given in the paper. A final alternative offered in the paper is to learn from information processing examples in nature, specifically in living systems. An E.coli cell of about one cubic micrometer volume is shown to be an incredibly powerful and energy-efficient information processor relative to the performance of an end-of-scaling silicon processor of the same volume. The paper concludes by pointing out some of the crucial differences between E.coli information processing and conventional approaches with the hope technologies can be invented using the hints offered by biosystems.
Ralph K. Cavin III, Paolo Lugli, Victor V. Zhirnov
Proc. IEEE1
2010 Device and Architecture Outlook for Beyond CMOS Switches
abstract
Sooner or later, fundamental limitations destine complementary metal-oxide-semiconductor (CMOS) scaling to a conclusion. A number of unique switches have been proposed as replacements, many of which do not even use electron charge as the state variable. Instead, these nanoscale structures pass tokens in the spin, excitonic, photonic, magnetic, quantum, or even heat domains. Emergent physical behaviors and idiosyncrasies of these novel switches can complement the execution of specific algorithms or workloads by enabling quite unique architectures. Ultimately, exploiting these unusual responses will extend throughput in high-performance computing. Alternative tokens also require new transport mechanisms to replace the conventional chip wire interconnect schemes of charge-based computing. New intrinsic limits to scaling in post-CMOS technologies are likely to be bounded ultimately by thermodynamic entropy and Shannon noise.
Kerry Bernstein, Ralph K. Cavin III, Wolfgang Porod, Alan C. Seabaugh, Jeff Welser
Proc. IEEE2
2010 Nanoelectronics Research for Beyond CMOS Information Processing
abstract
This special issue presents a variety of invited papers covering nanostructures and related materials proposed to extend CMOS scaling to its ultimate limit and enable a variety of new logic and memory devices.
George Bourianoff, Michel Brillouët, Ralph K. Cavin III, Toshiro Hiramoto, James A. Hutchby, Adrian M. Ionescu, Ken Uchida
Proc. IEEE3
2010 Regional, National, and International Nanoelectronics Research Programs: Topical Concentration and Gaps
abstract
This paper will outline some results obtained by an international working group on nanoelectronics, which collects data from major publicly funded programs in Europe, Japan, and the United States on long-term nanoelectronics research. It maps these programs and projects onto a set of research directions that are expected to drive nanoelectronics for the long term. The purpose is to identify those research topics attracting a lot of attention and those important topics that seem less attractive. This paper will give examples of interregional collaborative programs and identify sources of funding specifically provided to support international collaborations.
Michel Brillouët, George Bourianoff, Ralph K. Cavin III, Toshiro Hiramoto, James A. Hutchby, Adrian M. Ionescu, Ken Uchida
Proc. IEEE3
2010 Memory Devices: Energy-Space-Time Tradeoffs
abstract
Many memory candidates based on beyond complementary metal-oxide-semiconductor (CMOS) nanoelectronics have been proposed, but no clear successor has yet been identified. In this paper, we offer a methodology for system-level analysis and address the relationship of the maximum performance of a given memory device type to device physics. The method is illustrated for the classical dynamic RAM (DRAM) device and for the emerging memory device known as the resistive RAM (ReRAM).
Victor V. Zhirnov, Ralph K. Cavin III, Stephan Menzel, Eike Linn, Sebastian Schmelzer, Dennis Bräuhaus, Christina Schindler, Rainer Waser
Proc. IEEE2
2004 A hybrid current/voltage mode on-chip signaling scheme with adaptive bandwidth capability
abstract
This brief describes an adaptive bandwidth bus architecture based on hybrid current/voltage mode repeaters for long global RC interconnect static busses that achieves high-data rates while minimizing the static power dissipation associated with current-mode (CM) signaling. An experimental adaptive bandwidth bus test chip fabricated in AMI 1.6-/spl mu/m Bulk CMOS indicates a reduction in power dissipation of approximately 62% over CM sensing and an increase in maximum data rate of 40% over voltage-mode signaling.
Rizwan Bashirullah, Wentai Liu, Ralph K. Cavin III, Dale Edwards
IEEE Trans. Very Large Scale Integr. Syst.3
2003 Low-power design methodology for an on-chip bus with adaptive bandwidth capability
abstract
This paper describes a low-power design methodology for a bus architecture based on hybrid current/voltage mode signaling for deep sub-micrometer on-chip interconnects that achieves high data transmission rates while minimizing the number of repeaters by nearly 1/3. The technique uses low-impedance current-mode sensing to increase the data throughput and minimizes the static power dissipation inherent to current-mode signaling by adaptively changing the interconnection bandwidth given a change in input signal activity. Since bandwidth is related to power dissipation, the adaptive bus attains energy efficient data transmission by expending minimum power required to support the bus signal activity.The design method is based on statistical analysis of address streams extracted for typical benchmark programs using a microprocessor time-based simulator in combination with circuit-level power analysis. Simulation results indicate improvements in power dissipation of up to 65% and 40% over current and voltage mode signaling schemes, respectively.
Rizwan Bashirullah, Wentai Liu, Ralph K. Cavin III
DAC3
2003 Limits to binary logic switch scaling - a gedanken model
abstract
In this paper we consider device scaling and speed limitations on irreversible von Neumann computing that are derived from the requirement of "least energy computation." We consider computational systems whose material realizations utilize electrons and energy barriers to represent and manipulate their binary representations of state.
Victor V. Zhirnov, Ralph K. Cavin III, James A. Hutchby, George Bourianoff
Proc. IEEE2
2003 Current-mode signaling in deep submicrometer global interconnects
abstract
This paper addresses propagation delay and power dissipation for current mode signaling in deep submicrometer global interconnects. Based on the effective lumped element resistance and capacitance approximation of distributed RC lines, simple yet accurate closed-form expressions of delay and power dissipation are presented. A new closed-form solution of delay under step input excitation is first developed, exhibiting an accuracy that is within 5% of SPICE simulations for a wide range of parameters. The usefulness of this solution is that resistive load termination for current mode signaling is accurately modeled. This model is then extended to a generalized delay formulation for ramp inputs with arbitrary rise time. Using these expressions, the optimum-line width that minimizes the total delay for current mode circuits is found. Additionally, a new power-dissipation model for current-mode signaling is developed to understand the design tradeoffs between current and voltage sensing. Based on the results and derived formulations, a comparison between voltage and current mode repeater insertion for long global deep submicrometer interconnects is presented.
Rizwan Bashirullah, Wentai Liu, Ralph K. Cavin III
IEEE Trans. Very Large Scale Integr. Syst.3
2000 Integrated parametric timing optimization of digital systems
abstract
Clock skew optimization is a timing technique to improve system performance by employing scheduled skews at flip-flops. The integrated framework presented here includes a new linear programming (LP) formulation for the clock skew optimization problem. In this work, we use the concept of a global time frame, instead of a local one, to find a set of optimal skews to minimize system cycle time. The framework provides a firm theoretical foundation for scheduling skews into existing designs. Furthermore, we extend the LP formulation to accommodate retiming in the optimization process. Our framework allows for concurrent timing optimization of a design by retiming the circuit and scheduling clock skews at flip-flops. It is shown that this optimization can be formulated as a mixed-integer linear program and significantly reduce the clock period.
Hong-Yean Hsieh, Wentai Liu, Ralph K. Cavin III
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1995 Concurrent timing optimization of latch-based digital systems
abstract
Many techniques have been proposed to optimize digital system timing. Each technique can be advantageous in particular applications, however they are most often applied individually rather than concurrently. The framework presented here allows for concurrent timing optimization using retiming, intentional clock skew, and wave pipelining for latch-based designed systems with single or multi-phase clocking. This optimization is formulated as a mixed integer linear program. Our integrated framework also includes a new optimization technique called resynchronization which allows for the insertion of latches in the shortest paths and thus avoids race conditions. Our work has been applied to several designs and is able to significantly reduce the clock period.
Hong-Yean Hsieh, Wentai Liu, Ralph K. Cavin III, C. Thomas Gray
ICCD3
1995 High Speed, Fine Resolution Pattern Generation Using the Matched Delay Technique
abstract
This paper presents an architecture for generating a high-speed data pattern with precise edge placement (resolution) by using the matched delay technique. The technique involves passing clock and data signals through arrays of matched delay elements in such a way that the data rate and resolution of the generated data stream are controlled by the difference of these matched delays. This difference can be made much smaller than an absolute gate delay. Since the resolution of conventional designs is determined by these absolute delays, the matched delay technique yields a much finer resolution as well as higher speeds than traditional methods. The matched delay technique lends itself to high-precision and high-speed applications such as fast network interfaces or test pattern generators. This paper also describes a matched delay data generator submitted for fabrication in a MOSIS 1.2 /spl mu/m CMOS technology. Simulations indicate that data signals with on-chip bit rates of 833 Mb/s and resolutions of 100 ps can be generated.
Gary C. Moyer, Mark A. Clements, Wentai Liu, Toby Schaffer, Ralph K. Cavin III
ISCAS5
1994 Circuit delay calculation considering data dependent delays
C. Thomas Gray, Wentai Liu, Ralph K. Cavin III, Hong-Yean Hsieh
Integr.3
1994 Timing constraints for wave-pipelined systems
abstract
Wave-pipelining is a timing methodology used in digital systems to achieve maximal rate operation. Using this technique, new data are applied to the inputs of a combinational block before the previous outputs are available, thus effectively pipelining the combinational logic and maximizing the utilization of the logic without inserting registers. This paper presents a timing constraint formulation for the correct clocking of wave-pipelined systems. Both single- and multiple-stage systems including feedback are considered. Based on the formulation of this paper, several important new results are presented relating to performance limits of wave-pipelined circuits. These results include the specification of distinct and disjoint regions of valid operation dependent on the clock period, intentional clock skew, and the global clock latency. Also, implications and motivations for the use of accurate delay models and exact timing analysis in the determination of combinational logic delays are given, and an analogous relationship between the multi-stage system and the single-stage system in terms of performance limits is shown. The minimum clock period is obtained by clock skew optimization formulated as a linear program. In addition, important special cases are examined and their relative performance limits are analyzed.>
C. Thomas Gray, Wentai Liu, Ralph K. Cavin III
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1993 A bit-serial VLSI architecture for generating moments in real-time
abstract
In computer vision and image processing, the high degree of parallelism and pipelining of algorithms is often obstructed by the raster-scan I/O constraint and the information growing property of multiresolution structures. The approach of formulating algorithms in the pyramid structure as a binary tree structure, and mapping the binary tree structure into a linear pipelined array of 2logN levels for N*N images using a first-in, first-out technique (FIFO) to emulate the tree connections is proposed. It turns out that several geometric feature extraction algorithms such as moment generation can be represented in this scheme so that the inherent information growing of the algorithms enables the exploitation of bit-level concurrency in the architectural design. Consequently, the design of pipelined processor at each level is significantly simplified using bit-serial arithmetic, and this VLSI architecture is capable of generating moments concurrently in real-time.>
Wentai Liu, Su-Shing Chen, Ralph K. Cavin III
IEEE Trans. Syst. Man Cybern.3
1990 The design of a high-performance scalable architecture for image processing applications
abstract
The authors present the organization of an interleaved wrap-around memory system for a partitionable parallel/pipeline architecture with P pipes of L processors each. The architecture is designed to efficiently support real-time image processing and computer vision algorithms, especially those requiring global data operations. The interleaved memory system makes the architecture highly scalable in that L and P can be chosen to optimize performance for particular problems and reconfigurable in that, once L and P are fixed, problems of any size can still be mapped onto the architecture. The authors demonstrate techniques and methods for mapping computational structures to the architecture by considering the case of the 1-D butterfly network (1DBN). Since many other computational structures can be mapped to 1DBN, this gives a firm application base for the architecture. The authors also demonstrate methods for scheduling and controlling the memory system.>
C. Thomas Gray, Wentai Liu, Thomas A. Hughes, Ralph K. Cavin III
ASAP4
1990 P3A: a partitionable parallel/pipeline architecture for real-time image processing
abstract
A high-performance partitionable parallel/pipeline architecture (P/sup 3/A) that is capable of real-time image processing is discussed. The architecture consists of P disjoint pipes of L processors each, connected together through a novel wraparound memory. Many different problem classes, including shuffle-exchange, butterfly, and tree algorithms, can be easily mapped into P/sup 3/A. The power of the architecture lies in its ability to exploit both the spatial and temporal aspects of concurrency balancing parallelism and pipelining.>
C. Thomas Gray, Wentai Liu, Thomas A. Hughes, Ralph K. Cavin III, Su-Shing Chen
ICPR (2)4
1990 Design of integrated circuits: directions and challenges
abstract
Several of the dimensions of IC CAE technology are discussed, focusing on two design styles: custom design, used for commodity products such as DRAMs, microprocessors, etc., where large volume production is planned and area reduction and performance maximization can be expected to return large dividends; and the design of application-specific integrated circuits (ASICs), utilizing either very regular, prepatterned silicon arrays customized at the interconnect level or predesigned, parameterized libraries of cells that are usually arranged in rows and interconnected. The design research that will be required in order to attain the objectives of highly automated design systems and shorter product design cycles for integrated circuits are outlined. Metrics for design system performance are discussed.
Ralph K. Cavin III, Jeffrey L. Hilbert
Proc. IEEE1
1989 The Semiconductor Research Corporation: Cooperative research
abstract
The SRC (Semiconductor Research Corporation) was formed in 1982 to conduct generic, cooperative university research in the field of integrated circuits. An overview is provided of the methodologies used by the SRC for the identification of pacing integrated-circuit technologies, for research program planning and management, and for the transfer of research results to members. Several case studies are developed that illustrate the SRC approach to the conduct of research and that give a perspective on the broad spectrum of research results being produced. The SRC has found that the process of defining generic research goals, followed by the development and implementation of research plans to achieve the stated goals, provides effective focus and metrics for measuring research progress. It is the SRC's experience that focused university research can provide substantial contributions to the advancement of semiconductor technology as well as an additional work force to enhance the industry, university, and government technical infrastructure of the United States.>
Ralph K. Cavin III, Larry W. Sumney, Robert M. Burger
Proc. IEEE1
1989 Hamiltonian Cycles in the Shuffle-Exchange Network
abstract
The usefulness of the shuffle-exchange network in parallel processing applications is well established. The optimal embedding of a shuffle-exchange network of a given size depends on the number of cycles of the shuffle permutation of that size. The cost of one method of adding fault-tolerance through reconfigurability depends upon the number of such cycles, and the manner in which they can be connected to form larger cycles. An exact equation for the number of cycles of a shuffle of size 2/sup W/ is presented. That result is used to demonstrate that it is always possible to form a Hamiltonian cycle on all processors in a shuffle-exchange connected array. From this, it is apparent that there are a large number of ways of sharing spare processors among the members of many cycles. Thus, redundancy can be supplied in any strength, from adding on a spare processor, to adding one for each cycle of the shuffle.>
Wentai Liu, Thomas H. Hildebrandt, Ralph K. Cavin III
IEEE Trans. Computers3
1988 Bit-level concurrency in real-time geometric feature extractions
abstract
An efficient mapping from a tree structure into a pipelined array of 2log N states is presented for processing an N*N image. In the proposed mapping structure the identification of the information growing property inherent in feature-extraction algorithms allows bit-level concurrency to be exploited in the architectural design. Accordingly, the design of each staged pipelined processor is simplified.>
Wentai Liu, Tong-Fei Yeh, William E. Batchelor, Ralph K. Cavin III
CVPR4
1988 Exploiting Bit Level Concurrency in Real-Time Geometric Feature Extractions
abstract
Characteristics and constraints of real-time geometric-feature extraction are discussed. Extracting geometric features from a digital image can be characterized as a computation-intensive task in the environment of a real-time automated vision system. Such tasks require algorithms with a high degree of parallelism and pipelining under the raster-scan I/O constraint. Using the divide-and-conquer technique, many feature extractions have been formulated as a pyramid structure and then mapped into a binary tree. An efficient mapping from a tree structure into a pipelined array of 2logN stages is presented for processing an N*N image. In the proposed mapping structure, the identification of the information growing property allows the exploitation of bit-level concurrency in the architecture design. Accordingly, the design of each staged pipelined processor is simplified containing only bit-serial arithmetic. A single VLSI chip that can generate (p+1)(q+1) moments concurrently in real-time applications is described. This chip has a hardware complexity of O(pq(p+q)log/sup 2/N) units, where p, q stand for the orthogonal orders of the moment. This hardware complexity is better than the O(pq(p+q)/sup 2/log/sup 2/N) units required by the other methods. A single VLSI chip to generate ten moments for a (512*512*8)/pixel image in real time is presented.>
Wentai Liu, Tong-Fei Yeh, William E. Batchelor, Ralph K. Cavin III
ISCA4
1988 Rasterization theory, architectures, and implementations for a class of two-dimensional problems
Wentai Liu, Ralph K. Cavin III
Integr.2
1986 A perspective on CMOS technology trends
abstract
Integrated circuit technology continues to evolve at a rapid pace, driven by the requirements of new applications for electronics of higher performance at ever lower cost. The attributes of CMOS technology in a ULSI environment are an ideal match to these requirements; thus CMOS is becoming the ubiquitous integrated circuit technology. The main feature of CMOS is the existence of complementary n- and p-channel transistors, which results in circuit configurations with virtually zero steady-state current, and consequently low power dissipation. Although CMOS is conceptually a circuit technology, this has implications for fabrication and layout, as well as for functional partitioning in the circuit design environment and test. These considerations are reviewed with special attention to those areas, such as test, latchup, and the design environment, where technical problems are substantial.
William C. Holton, Ralph K. Cavin III
Proc. IEEE2
1984 Introduction to the SRC design sciences program
Ralph K. Cavin III
DAC1
1983 Microelectronic architectures and devices for signal and symbol processing
Ralph K. Cavin III, Noel R. Strader
Integr.1
1982 Analysis of error-gradient adaptive linear estimators for a class of stationary dependent processes
abstract
In many applications, the training data to be processed by an adaptive linear estimator can be assumed to have a finite correlation length. An exact analysis for this class of problems that yields the coefficient bias, coefficient correlation matrix, and mean square estimation error is obtained via a stochastic imbedding procedure. A power series expansion in the gain parameter is used to obtain simplified expressions of order one for the above statistical moments. These new expressions are shown to contain the terms that would result from an analysis based upon the assumption of independent training data plus additional terms arising from data correlation. Algorithm convergence properties are studied by identifying the appropriate matrix eigenvalues from the first-order theory.
Stephen K. Jones, Ralph K. Cavin III, William M. Reed
IEEE Trans. Inf. Theory2
1980 An Efficient Computational Procedure for the Evaluation of the M/M/I Transient State Occupancy Probabilities
abstract
In this note a procedure is given for the numerical evaluation of theM/M/1queue transient state occupancy probabilities which arise in the analysis of dynamic buffer behavior in store and forward networks. The procedure uses the circular coverage function of radar and communication theory to eliminate an infinite series of modified Bessel functions. Computational savings are illustrated by several numerical examples.
Stephen K. Jones, Ralph K. Cavin III, Donald A. Johnston
IEEE Trans. Commun.2
1977 On the Maximum Average Throughput Rate for a Tandem Node Network
abstract
The problem of regulating inputs into a tandem store-and-forward node pair with finite queue limits in order to maximize the average message completion rate is considered from the viewpoint of Markov Decision Theory. A variation of the policy iteration procedure due to Howard [3] is used to compute optimum control policies and message completion rates for various configurations of system parameters. Comparative data are provided on the performance of the optimum control law relative to simpler local control laws which require no information sharing.
Stephen K. Jones, Ralph K. Cavin III, J. N. Holyoak
IEEE Trans. Commun.2