José G. Delgado-Frias

dblp:97/790 · DBLP profile ↗
← Back
50ranked-venue papers
10as first author
3since 2021 · last 2022
0000-0002-7026-9991ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 9 first-author · 3 since 2021Computer networks · 6Artificial intelligence and machine learning · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 ACHS Optimizations on 3D Interconnect Arrangements
abstract
The Asymmetric Crosstalk Harnessed Signaling (ACHS) scheme [1]–[4] provides a significant improvement to the original Crosstalk Harnessed Signaling (CHS) technique [5]. This paper presents a study of design optimizations applied on 3D multilayer interconnect arrangements [2] [3] [6], in which alterations on the encoding matrices or the distribution of wires in the channel translate into a modification of the bit encoding regions and their associated performances. While ACHS has been presented as a vastly superior signaling option than the originally proposed CHS due to the elimination of the highly sensitive common encoding eigenmode [1], further encoding optimizations on the critical (worst performing) data bits and the strobe inside the ACHS implementation can further improve the bus performance. This work presents a full set of performance comparisons involving data and strobe bits, on a ACHS-encoded bus with channels adjacently routed in 3- and 4-layer 3D interconnect arrangements.
Daniel Iparraguirre, José G. Delgado-Frias
ISCAS2
2022 Asymmetric Crosstalk Harness Signaling for Common Eigenmode Elimination
abstract
This paper presents a novel scheme based on the Crosstalk Harnessed Signaling (CHS) technique for parallel high-speed interfaces. The proposed scheme, called Asymmetric Crosstalk Harnessed Signaling (ACHS), provides a robust and consistent eye opening for all the transmitted bits in the interface. The Hadamard matrix employed for CHS encoding has been modified to a non-square format in order to either eliminate the common eigenmode or turn it differential. This in turn results in a signaling scheme that converts a binary data array to a slightly larger signal/interconnect array. The resulting signaling and routing overhead translates into a comparable eye opening across all the bits inside the encoded bus, overcoming the crosstalk sensitivity associated to the common mode present in the original CHS scheme. Both CHS and the proposed ACHS are implemented on a 16-bit source-synchronous bus wired through 3-dimensional interconnect arrangements in a multi-stripline stackup, in order to compare performances in very aggressive crosstalk environments. Simulation results show a consistent eye opening across all data bits when ACHS is applied, rendering a fully functional bus with only 11% routing overhead, against a practically inoperable bus in the CHS case, due to the eye collapse for the common-mode encoded bit.
Daniel Iparraguirre, José G. Delgado-Frias, Howard Heck
IEEE Trans. Computers2
2021 A Crosstalk-Harnessed Signaling Enhancement that Eliminates Common-Mode Encoding
abstract
This paper presents an encoding variation on the Crosstalk-Harnessed Signaling (CHS) technique, aimed towards eliminating the common-mode eigenvector included in the Hadamard matrix for CHS encoding/decoding. This is accomplished by either removing the eigenvector or making it differential by concatenating matrices in a diagonal fashion, or by adding columns to the matrix; this results in a non-square matrix that delivers a larger number of signals to be routed in the interface. Simulation results show the signaling/routing overhead translates in a significantly higher performance for high-speed parallel interfaces with very high routing integration levels.
Daniel Iparraguirre, José G. Delgado-Frias, Howard Heck
ISCAS2
2020 Online Firmware Functional Validation Scheme Using Colored Petri Net Model
abstract
Firmware functional validation suffers from a series of performance limitations in practice, which in turn heavily relies on manual effort and becomes a major bottleneck of product time cycle. The requirement of repetitive run-time firmware execution for the validation environment demands novel techniques to accelerate the validation process. We propose an online firmware functional validation scheme utilizing the colored Petri net (CPN) model which can be generated automatically from the firmware source code. With simulation runs on the generated CPN models at run-time, the firmware execution path is presented and, if an error occurs, the location of error can be identified. An integrated validation tool has been designed and implemented to show the proposed validation methodology's potential and effectiveness. This tool is used in the validation of the universal serial bus (USB) initialization in unified extensible firmware interface (UEFI).
Rongyang Liu, José G. Delgado-Frias, Doug Boyce, Yi Qian 0003, Rahul Khanna
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Full-VDD and near-threshold performance of 8T FinFET SRAM cells
Michael A. Turi, José G. Delgado-Frias
Integr.2
2016 Autonomous management of a recursive area hierarchy for large scale wireless sensor networks using multiple parents
Johnathan Vee Cree, José G. Delgado-Frias
Ad Hoc Networks2
2015 Near-threshold CNTFET SRAM cell design with removed metallic CNT tolerance
abstract
We report a study of power supply reduction to near-threshold for an 8-transistor CNTFET SRAM cell. Voltage at near-threshold has an impact on delays, energy, energy-delay product, leakage current, and static noise margin. In addition, we have incorporated a removed metallic CNT approach to deal with non-semiconductor CNTs. In this study we investigate how to enhance SRAM performance by means of two techniques: Gated Power Supply and Word-line Boosting. Using the gated power supply technique, power saving is over 5X, while the average delay is increased by 3.5X as compared to 0.9V Vdd. On the other hand, word-line boosting technique (where Read and write word lines are boosted with additional 100mV) helps to improve write and read delays that are faster by 3.8X and 1.7X, respectively at Vdd=0.4V. Lowest Energy delay product (EDP) for gated power supply and word-line boosting is at 0.5V and 0.4V, respectively. EDP compared to nominal Vdd of 0.9V is lowered by 38% and 56.9% respectively for the mentioned techniques.
José G. Delgado-Frias, Michael A. Turi
ISCAS1
2013 CNTFET 8T SRAM cell performance with near-threshold power supply scaling
abstract
In this study, we present a Carbon Nanotube (CNT) FET based 8T SRAM cell and its performance in near threshold voltage region. Metallic CNTs (M-CNTs) grow alongside semiconductor CNTs in current synthesis process, but they can be removed using novel techniques. This in turn creates open circuit and degrades the performance and functionality of SRAM cells. In this paper we apply a removed metallic CNT tolerant approach. The near threshold performance of the 8T SRAM cells with the tolerant approach is simulated and optimized to obtain best performance under the supply voltage from 0.4V to 0.9V. An evaluation of energy delay product and SNM shows a favorable tradeoff for the 0.6V power supply. The energy savings for cells with 0.6V power supply are 56.5% and 10.0% for average and worst case, respectively, compared with 0.9V; on the other hand, it has about 58% longer max delay and 25.8% lower static noise margin. The average and worst case values of EDP for 0.6V is 34.0% and 27.8% lower than that of 0.9V, with only 0.09% more invalid cell.
José G. Delgado-Frias
ISCAS2
2012 SRAM leakage in CMOS, FinFET and CNTFET technologies: leakage in 8t and 6t sram cells
abstract
An in-depth study of the static power consumption in 6T and 8T SRAM cell designs based on 32nm CMOS, FinFET and CNTFET technologies is presented. In addition to the inverter leakage currents, memory cells that are not active when write or read operations occur draw current from/to the bus drivers increasing the total standby power consumption. The FinFET schemes yield substantially lower write (1023.5 pA) and read (522.5 pA) leakage currents in 8T cells, which are 10.4% and 4.4% of the amount in CMOS 8T cells. A CNTFET 6T cell consumes 1.9% and 2.8% of the leakage current drawn by a CMOS 6T cell for write and read.
Michael A. Turi, José G. Delgado-Frias
ACM Great Lakes Symposium on VLSI3
2010 FPGA schemes for minimizing the power-throughput trade-off in executing the Advanced Encryption Standard algorithm
Jason Van Dyken, José G. Delgado-Frias
J. Syst. Archit.2
2009 IP Routing table compaction and sampling schemes to enhance TCAM cache performance
Ruirui Guo, José G. Delgado-Frias
J. Syst. Archit.2
2008 High-performance low-power AND and Sense-Amp address decoders with selective precharging
abstract
This paper presents and evaluates two novel address decoding schemes that use selective precharging, the Sense-Amp and the AND decoders, in comparison to the conventional NOR decoder. Simulations for all three designs are performed using 65nm CMOS technology and the delays of all three decoders are set to 120ps for a common base comparison. The most selective AND decoder performs best and dissipates between 0.17% and 43.17% (29.29% on average) and the selective Sense-Amp decoder dissipates between 28.81% and 48.33% (39.96% on average) of the energy dissipated by the nonselective conventional decoder.
Michael A. Turi, José G. Delgado-Frias
ISCAS2
2008 Performance analysis of multipath transmission over 802.11-based multihop ad hoc networks: a cross-layer perspective
abstract
Data transmission in ad hoc networks involves interactions between medium access control (MAC)-layer protocols and data forwarding along network-layer paths. These interactions have been shown to have a significant impact on the performance of a system. This impact on multipath data transmission over multihop IEEE 802.11 MAC-based ad hoc networks is assessed; analysis is from a cross-layer perspective. Both MAC layer protocols and network-layer data forwarding are taken into account in the system models. The frame service time at source in a 802.11 MAC-based multipath data transmission system under unsaturated conditions is studied. Analytical models are developed for two packet generation schemes (round robin and batch) with a Poisson frame arrival process. Moreover, an analytical model is developed to investigate the throughput of a multipath transmission system in 802.11-based multihop wireless networks. Two methods are proposed to estimate the impact of cross-layer interactions on the frame service time in such a system. Two bounds of the system throughput are obtained based on these estimation methods. These models are validated by means of simulation under various scenarios.
José G. Delgado-Frias, Krishnamoorthy Sivakumar
IET Commun.2
2008 A Medium-Grain Reconfigurable Architecture for DSP: VLSI Design, Benchmark Mapping, and Performance
abstract
Reconfigurable hardware has become a well-accepted option for implementing digital signal processing (DSP). Traditional devices such as field-programmable gate arrays offer good fine-grain flexibility. More recent coarse-grain reconfigurable architectures are optimized for word-length computations. We have developed a medium-grain reconfigurable architecture that combines the advantages of both approaches. Modules such as multipliers and adders are mapped onto blocks of 4-bit cells. Each cell contains a matrix of lookup tables that either implement mathematics functions or a random-access memory. A hierarchical interconnection network supports data transfer within and between modules. We have created software tools that allow users to map algorithms onto the reconfigurable platform. This paper analyzes the implementation of several common benchmarks, ranging from floating-point arithmetic to a radix-4 fast Fourier transform. The results are compared to contemporary DSP hardware.
Mitchell J. Myjak, José G. Delgado-Frias
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Redundant Array of Independent Fabrics - An Architecture for Next Generation Network
abstract
As the next generation network begins to incorporate the Internet, telecommunication and TV services, it becomes one of the most critical infrastructures for our society. Routers construct the skeleton of the network. Their kernel, the structure and configuration (scheduler) of the fabric, dominates the networks' performance, scalability, reliability and cost. Based on previous research, we proposed an interleaved architecture of multistage switching fabrics, which will meet the requirements for next generation routers. In this paper, we first assess its performance with a theoretical model which complements our previous simulation results. Moreover, the interleaved fabrics show great tolerance against internal hardware failures. Based on these properties, we propose the architecture of RAIF (redundant array of independent fabrics) for next generation network, which could get better performance and fault tolerance as RAID.
Rongsen He, José G. Delgado-Frias
GLOBECOM2
2007 MARS: Misbehavior Detection in Ad Hoc Networks
abstract
To detect misbehavior on data and mitigate adverse effects, we propose and evaluate a MultipAth routing single path transmission (MARS) scheme. The MARS combines multipath routing, single path data transmission, and end-to-end feedback mechanism together to provide more comprehensive protection against misbehavior from individual or cooperating misbehaving nodes. The MARS scheme and its enhancement E- MARS are evaluated by means of simulation under various adverse scenarios. The simulation results show that the MARS and E-MARS schemes provide better network performance and considerable protection to data transmission than some DSR-based transmission systems at the expense of moderate overhead. Compared to the DSR-based schemes, the proposed schemes deliver up to 45% more data with 20% misbehaving nodes under individual misbehavior, and up to 28% more data with 40% misbehaving nodes under colluded misbehavior.
José G. Delgado-Frias
GLOBECOM2
2007 Hardened by Design Techniques for Implementing Multiple-Bit Upset Tolerant Static Memories
abstract
We present a novel MBU-tolerant design, which utilizes layout-based interleaving and multiple-node disruption tolerant memory latches. This approach protects against grazing incidence particle strikes, which produce disruptions with the widest possible spatial separation. Advantages with respect to size, complexity, and MBU tolerance are realized when this approach is compared to existing solutions.
Daniel R. Blum, José G. Delgado-Frias
ISCAS2
2007 Preface
Laurence T. Yang, José G. Delgado-Frias, Mohammed Y. Niamat, Dimitrios Soudris, Srinivasa Vemuru
Integr.2
2007 Fault Tolerant Interleaved Switching Fabrics For Scalable High-Performance Routers
abstract
Scalable high-performance routers and switches are required to provide a larger number of ports, higher throughput, and good reliability. Most of today's routers and switches are implemented using single crossbar as the switched fabric. The single crossbar complexity increases at O(N2) in terms of crosspoint number, which might become unacceptable for scalability as the port number (N) increases. A delta class self-routing multistage interconnection network (MIN) with the complexity of O(N times log2N) has been widely used in the asynchronous transfer mode switches. However, the reduction of the crosspoint number results in considerable internal blocking. A number of scalable methods have been proposed to solve this problem. One of them uses more stages with recirculation architecture to reroute the deflected packets, which greatly increase the latency. In this paper, we propose an interleaved multistage switching fabrics architecture and assess its throughput with an analytical model and simulations. We compare this novel scheme with some previous parallel architectures and show its benefits. From extensive simulations under different traffic patterns and fault models, our interleaved architecture achieves better performance than its counterpart of single panel fabric. Our interleaved scheme achieves speedups (over the single panel fabric) of 3.4 and 2.25 under uniform and hot-spot traffic patterns, respectively, at maximum load (p = 1). Moreover, the interleaved fabrics show great tolerance against internal hardware failures.
Rongsen He, José G. Delgado-Frias
IEEE Trans. Parallel Distributed Syst.2
2006 Interleaved Multistage Switching Fabrics for Scalable High Performance Routers
abstract
As the Internet grows exponentially, scalable high performance routers and switches on backbone are required to provide a large number of ports, higher throughput, lower delay latency and good reliability. At present, most of these routers and switches are implemented on single crossbar as the switched backplane fabric. But the complexity of the single crossbar is increased with O(N2) in terms of crosspoint number, which is unacceptable for scalability when N becomes large. A delta class self-routing multistage interconnection network with the complexity of O(Ntimeslog2N) has been widely used in the ATM switches. However, the reduction of the crosspoint number results in the serious internal blocking. To solve this problem, quite a few scalable methods have been proposed. One of them, more stages with recirculation architecture is used to reroute the deflected packets, which increase the latency a lot. In this paper, we first bring out the multiple-panel MIN switching fabrics with interleaved recirculation. We also show how to correctly choose the recirculation points to reroute the cells, compared with the wrong connections of former publication. From the simulation under different traffic patterns, this new interleaved architecture, which is insensitive to congestion, could achieve better performance than its counterpart of single panel fabric.
Rongsen He, José G. Delgado-Frias
GLOBECOM2
2006 Superpipelined reconfigurable hardware for DSP
abstract
Reconfigurable hardware offers a number of advantages over custom integrated circuits, including low development cost, high flexibility, and high adaptability to changing requirements. However, this alternative does incur some reduction in performance, especially for computationally intensive tasks such as digital signal processing. Recent developments in both research and industry have aimed to reduce this gap. This paper introduces a novel reconfigurable architecture that pipelines computations at the bit level. The architecture includes a number of features to improve performance, including medium-grain cells, hierarchical interconnections, and minimal clocking overhead. Circuit simulations demonstrate that the basic cell runs at 1.5 GHz in a modest 180-nm technology. At this speed, we estimate that the device could compute a 256-point fast Fourier transform in 829 ns
Mitchell J. Myjak, José G. Delgado-Frias
ISCAS2
2006 A mesochronous pipeline scheme for high performance low power digital systems
abstract
A mesochronous pipeline architecture is described in this paper. Significant performance gains are possible with mesochronous pipeline over conventional pipeline architecture. The clock period in conventional pipeline scheme is proportional to the maximum stage delay while in mesochronous pipelining it is proportional to the maximum delay difference, which means higher clock speeds are possible in the proposed scheme. Also, the clock distribution network is simple and load on it is less in mesochronous approach resulting in significant power savings. An 8/spl times/8-bit multiplier using carry-save adder technique has been implemented in conventional and mesochronous pipeline approach using TSMC 180 nm (drawn length 200 nm). The over all power dissipation in mesochronous approach is less than 50% of the power dissipation in conventional approach. In conventional approach, the power dissipation in clock network and pipeline registers is close to 80% of total power dissipation, while in mesochronous approach the logic dissipates more power.
Suryanarayana Tatapudi, José G. Delgado-Frias
ISCAS2
2006 Performance Analysis of Multipath Data Transmission in Multihop Ad Hoc Networks
abstract
Data transmission in ad hoc networks involves interactions between MAC-layer protocol and data forwarding along network-layer paths. These interactions have been shown to have significant effect on system throughput and source queue characteristics. In this paper, multipath data transmission is studied and analyzed based on the DCF in IEEE 802.11 MAC protocols. Analytical models are developed to demonstrate the frame service time and the queue characteristics in the source station under unsaturated conditions for two frame arrival processes (Poisson and deterministic) and two packet generation processes (round robin and batch). The throughput in the multipath multihop system is also investigated and a bound of it is derived. These models are all validated by means of simulation under various scenarios
José G. Delgado-Frias
SECON2
2006 Multipath Routing Based Secure Data Transmission in Ad Hoc Networks
abstract
The specific characteristics of mobile ad hoc networks (MANETs) make cooperation among all nodes and secure transmission important issues in its research. Misbehaving nodes with different intentions and capabilities would conduct various types of misbehavior in the networks. In this paper, we present and evaluate a scheme, in which multipath routing combined with feedback mechanism are used to tackle misbehaviors on data delivery formed by one or more misbehaving nodes in an ad hoc network. Data and control packets are transmitted through two node-disjoint paths. The source is notified of suspected misconduct of intermediate nodes through feedback mechanism. A simple derivation of this scheme is also discussed. The proposed scheme and the derivation are compared with the single path routing protocol DSR by means of simulation implemented at normal and adverse scenarios. The simulation results show that the proposed scheme and the derivation provide considerable protection in ad hoc networks at the expense of moderate overhead introduced by multipath routing. In a network with up to 40% misbehaving nodes, the proposed scheme and the derivation result in around 17% in data receive rate over the single path DSR
José G. Delgado-Frias
WiMob2
2004 Pipelined Multipliers for Reconfigurable Hardware
abstract
Summary form only given. Reconfigurable devices used in digital signal processing applications must handle large amounts of data in vector form. Most signal processing algorithms use multiplication extensively; thus, the hardware must support this operation to achieve high performance. However, mapping a multiplier on traditional fine-grain devices produces a complex structure whose performance is limited by the routing overhead. In this paper, we present a novel pipelined multiplier structure suitable for medium-grain and coarse-grain reconfigurable cell arrays. We first implement an unsigned n-bit multiplier using m-bit cells. Then, we show how the same structure can work with two's-complement data with small changes to the configuration. The structure requires [n/m]/sup 2/ cells, but can execute vector operations in a pipelined fashion. We also discuss the benefits of using a hierarchical design for large multipliers.
Mitchell J. Myjak, José G. Delgado-Frias
IPDPS2
2001 A VLSI wrapped wave front arbiter for crossbar switches
abstract
Article Share on A VLSI wrapped wave front arbiter for crossbar switches Authors: José G. Delgado-Frias Univ. of Virginia, Charlottesville, VA Univ. of Virginia, Charlottesville, VAView Profile , Girish B. Ratanpal Univ. of Virginia, Charlottesville, VA Univ. of Virginia, Charlottesville, VAView Profile Authors Info & Claims GLSVLSI '01: Proceedings of the 11th Great Lakes symposium on VLSIMarch 2001 Pages 85–88https://doi.org/10.1145/368122.369605Online:01 March 2001Publication History 2citation319DownloadsMetricsTotal Citations2Total Downloads319Last 12 Months7Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
José G. Delgado-Frias, Girish B. Ratanpal
ACM Great Lakes Symposium on VLSI1
2000 A wave-pipelined router architecture using ternary associative memory
abstract
In this paper a wave-pipelining scheme is used to increase the performance of a router architecture. Wave-pipelining has a potential of significantly reducing clock cycle time and power. The design approach considered in this paper allows the propagation of data from stage to stage to occur without the use of intermediate latches. Control signals are used to ensure that intermixing of data waves does not occur. The results of the study show that wave-pipelining helps to reduce the clock period.
José G. Delgado-Frias, Jabulani Nyathi, Laxmi N. Bhuyan
ACM Great Lakes Symposium on VLSI1
2000 A wave-pipelined CMOS associate router for communication switches
abstract
A wave-pipelining approach is used to improve the performance of a VLSI router. Wave-pipelining has the potential of significantly reducing clock cycle time and silicon real estate. The design approach considered in this paper allows data propagation between stages to occur without the use of intermediate latches. Control signals are designed to ensure that intermixing of data waves does not occur. This study's results show that using wave-pipelining reduces the clock period. The circuit delays become the limiting factor, preventing further clock cycle time reduction.
José G. Delgado-Frias, Jabulani Nyathi
ISCAS1
2000 Elementary function generators for neural-network emulators
abstract
Piece-wise first- and second-order approximations are employed to design commonly used elementary function generators for neural-network emulators. Three novel schemes are proposed for the first-order approximations. The first scheme requires one multiplication, one addition, and a 28-byte lookup table. The second scheme requires one addition, a 14-byte lookup table, and no multiplication. The third scheme needs a 14-byte lookup table, no multiplication, and no addition. A second-order approximation approach provides better function precision; it requires more hardware and involves the computation of one multiplication and two additions and access to a 28-byte lookup table. We consider bit serial implementations of the schemes to reduce the hardware cost. The maximum delay for the four schemes ranges from 24- to 32-bit serial machine cycles; the second-order approximation approach has the largest delay. The proposed approach can be applied to compute other elementary function with proper considerations.
Stamatis Vassiliadis, José G. Delgado-Frias
IEEE Trans. Neural Networks Learn. Syst.3
1999 A neuro-emulator with embedded capabilities for generalized learning
Valentine C. Aikens II, José G. Delgado-Frias, Gerald G. Pechanek, Stamatis Vassiliadis
J. Syst. Archit.2
1998 A VLSI Self-Compacting Buffer for DAMQ Communication Switches
abstract
This paper describes a novel VLSI CMOS implementation of a self-compacting buffer (SCB) for the dynamically allocated multi-queue (DAMQ) switch architecture. The SCB is a scheme that dynamically allocates data regions within the input buffer for each output channel. The proposed implementation provides a high-performance solution to buffered communication switches that are required in interconnection networks. This performance comes from not only the DAMQ approach but also the pipelined implementation and novel circuitry. The major components of the SCB are described in detail in this paper. The system has the capability of performing a read, a write, or a simultaneous read/write operation per cycle due to its pipelined architecture.
José G. Delgado-Frias, Richard Diaz
Great Lakes Symposium on VLSI1
1998 A VLSI High-Performance Encoder with Priority Lookahead
abstract
In this paper we introduce a VLSI priority encoder that uses a novel priority lookahead scheme to reduce the delay for the worst case operation of the circuit, while maintaining a very low transistor count. The encoder's topmost input request has the highest priority; this priority descends linearly. Two design approaches for the priority encoder are presented, one without a priority lookahead scheme and one with a priority lookahead scheme. For an N-bit encoder, the circuit with the priority lookahead scheme requires only 1.094 times the number of transistors of the circuit without the priority lookahead scheme. Having a 32-bit encoder as an example, the circuit with the priority lookahead scheme is 2.59 times faster than the circuit without the priority lookahead. The worst case operation delay is 4.4 ns for this lookahead encoder, using a 1-/spl mu/m scalable CMOS technology. The proposed lookahead scheme can be extended to larger encoders.
José G. Delgado-Frias, Jabulani Nyathi
Great Lakes Symposium on VLSI1
1998 A Dictionary Machine Emulation on a VLSI Computing Tree System
abstract
In this paper, we propose a dictionary machine emulation using a novel VLSI tree structure that operates on the dictionary using a blocking technique. We show that dictionary machine operations can be performed through the implementation of a number of processing and communication tasks overlapped on a simple structure. By manipulating the key-records bit serially, and storing them in an external memory rather than within the layers of the structure, we show that the size of the dictionary is limited only by the capacity of the external memory. This structure, which consists of multiple units, can be implemented in VLSI onto a single-chip. The key advantage of our structure is that it provides a means of implementing a high speed and low cost dictionary machine with virtually unlimited capacity; thus, eliminating the need for multiple chips should the dictionary expand. We have that an exhaustive search on a 2048 key-record dictionary can be performed in 29.78 /spl mu/s.
Adger E. Harvin III, José G. Delgado-Frias
Great Lakes Symposium on VLSI2
1998 A Clustering and Genetic Scheme for Large Tsp Optimization Problems
abstract
A novel hybrid genetic scheme to solve the traveling salesman problem is presented in this paper. The proposed scheme has two major steps: clustering and global optimization. The clustering step is used to obtain an initial solution that is used as seed. The global optimization uses this seed to produce a better result using genetic operators. A number of benchmarks have been used to evaluate the proposed scheme. The solutions to these benchmarks which include 105-, 318-, 1060-, 2152-, 3038-, and 5934-city TSPs are extremely close to the best known solutions. This scheme is the first genetic approach that deals with large TSP benchmarks.
Chien-Ying Lu, José G. Delgado-Frias
Cybern. Syst.2
1998 Executing tree routing algorithms on a high-performance pattern associative router
abstract
In this paper a novel programmable approach to execute implicit routing algorithms is presented. The proposed router is based on an associative scheme that uses the attributes of the routing algorithm and the interconnection network topology. In this approach routing algorithms are mapped (or programmed) onto a set of bit-patterns that are matched in parallel. To show the applicability of this router, we have selected oblivious and fault-tolerant routing algorithms for ten different tree interconnection network topologies; however, the proposed scheme is flexible enough to accommodate other network topologies and routing algorithms. For the studied topologies, the number of required bit-patterns is of the same order as the topology degree. The proposed organization requires only one comparison and one read delays. This in turn yields a high-speed port assignment that is comparable to single topology routers (non-flexible routers). In the context of flexible router schemes, the proposed approach not only is one of the fastest but also requires a very small amount of hardware for its implementation.
Douglas H. Summerville, José G. Delgado-Frias, Stamatis Vassiliadis
J. Syst. Archit.2
1996 A VLSI Interconnection Network Router Using a D-CAM with Hidden Refresh
abstract
A VLSI implementation of a programmable router scheme for parallel interconnection network architectures is presented in this paper. The router executes routing algorithms in 1.5 clock cycles, this being the fastest approach for flexible routers. To further increase throughput, the router operation has been made pipelined, achieving 1 routing decision per cycle. The implementation is based on a content addressable memory (CAM) that supports per entry unique bit masking. This programmable CAM requires few entries; this in turn makes it possible to implement a dynamic approach in order to reduce the transistor count. We have provided circuitry and arranged timing to achieve refreshing of the stored data in a hidden fashion. In addition to the CAM, we have incorporated a fast priority scheme that allows only one entry to be selected and a memory that stores the port assignment. The number of required CAM entries is extremely small; it is of the same order as the output ports.
José G. Delgado-Frias, Jabulani Nyathi, Chester L. Miller, Douglas H. Summerville
Great Lakes Symposium on VLSI1
1996 Software Metrics and Microcode: A Case Study
abstract
In this paper, we report the findings of an investigation undertaken at IBM to determine whether or not existing software metrics are applicable to the microcode of large computer systems. As part of this investigation, we calculated several metrics from the microcode developed for the IBM 4381 and IBM 9370 computer systems, and used them as predictive parameters for a number of existing error prediction models. The microcode used in this case study exceeds 1.2 million lines of code written in 12 languages and comprises the microcode for the IBM ES/4381 and IBM ES/9370 computer systems. Our results suggest that only a few of the existing metrics are linearly independent, and that none of the metrics examined can be used in a regression model as a reliable error predictor.
George Triantafyllos, Stamatis Vassiliadis, José G. Delgado-Frias
J. Softw. Maintenance Res. Pract.3
1996 Sigmoid Generators for Neural Computing Using Piecewise Approximations
abstract
A piecewise second order approximation scheme is proposed for computing the sigmoid function. The scheme provides high performance with low implementation cost; thus, it is suitable for hardwired cost effective neural emulators. It is shown that an implementation of the sigmoid generator outperforms, in both precision and speed, existing schemes using a bit serial pipelined implementation. The proposed generator requires one multiplication, no look-up table and no addition. It has been estimated that the sigmoid output is generated with a maximum computation delay of 21 bit serial machine cycles representing a speedup of 1.57 to 2.23 over other proposals.
Stamatis Vassiliadis, José G. Delgado-Frias
IEEE Trans. Computers3
1996 A Flexible Bit-Pattern Associative Router for Interconnection Networks
abstract
A programmable associative approach to execute implicit routing algorithms is presented. Algorithms are mapped onto a set of bit-patterns that are matched in parallel. We have studied and mapped a large number of routing algorithms for a wide range of interconnection network topologies. Here we report three cases that illustrate the capabilities of the router scheme. For the studied topologies, the number of required bit-patterns is of the same order as the topology degree. The proposed approach is one of the fastest routers and requires a very small amount of hardware.
Douglas H. Summerville, José G. Delgado-Frias, Stamatis Vassiliadis
IEEE Trans. Parallel Distributed Syst.2
1995 The multi-associative branch target buffer: a cost effective BTB mechanism
Weili Chu, Stamatis Vassiliadis, José G. Delgado-Frias
Microprocess. Microprogramming3
1994 A VLSI CAM-based flexible oblivious router for multiprocessor interconnection networks
abstract
A VLSI implementation of a flexible router scheme for parallel interconnection network architectures is presented in this paper. The router implements implicit oblivious routing algorithms in 1.5 clock cycles, this being the fastest approach for flexible routers. To further increase performance, the router operation has been made pipelined with a throughput of 1 routing decision per cycle. The implementation is based on a combination of a content addressable memory that supports per entry unique bit masking, a fast priority scheme that allows only one entry to be selected, and a memory that stores the port assignment. The number of required CAM entries is extremely small; it is of the same order as the output ports (or node degree).>
José G. Delgado-Frias, Rovy Sze, Douglas H. Summerville, Valentine C. Aikens II
Great Lakes Symposium on VLSI1
1994 Design and evaluation of a DAMQ multiprocessor network with self-compacting buffers
abstract
The paper describes a new approach to implement dynamically allocated multi-queue (DAMQ) switching elements using a technique called "self-compacting buffers". This technique is efficient in that the amount of hardware required to manage the buffers is relatively small; it offers high performance since it is an implementation of a DAMQ. The paper describes the self-compacting buffer architecture in detail, and compares it against a competing DAMQ switch design. It presents extensive simulation results comparing the performance of a self-compacting buffer switch against an ideal switch including several examples of k-ary n-cubes and delta networks. In addition, simulation results show how the performance of an entire network can be quickly and accurately approximated by simulating just a single switching element.>
Joonho Park, Brian W. O'Krafka, Stamatis Vassiliadis, José G. Delgado-Frias
SC4
1994 An investigation of binary CLA and ripple CMOS adder designs
Chuan-Jen Chang, Stamatis Vassiliadis, José G. Delgado-Frias
Microprocess. Microprogramming3
1993 A massively parallel diagonal-fold array processor
abstract
Image processing for multimedia workstations is a computationally intensive task typically requiring special purpose hardware, for example a nearest neighbor mesh parallel machine organization. One type of nearest neighbor mesh computer consists of a K /spl times/ K square array of Processor Elements (PEs) where each PE is connected to the North, South, East, and West PEs only. In a torus configuration, there are a total of 2K/sup 2/ PE interfaces. Under the assumption of SIMD operation with unidirectional message and data transfers between the PEs, it is possible to reconfigure the array by placing the symmetric PEs together and share the north-south wires with the east-west wires, thereby reducing the wiring complexity in half, i.e. K/sup 2/ PE interfaces without affecting performance. This new machine organization is termed the Diagonal-Fold Mesh Array Processor, providing equivalent performance to a nearest neighbor mesh with half the wiring complexity for unidirectional data transferring algorithms.>
Gerald G. Pechanek, José G. Delgado-Frias, Stamatis Vassiliadis
ASAP2
1992 Digital neural emulators using tree accumulation and communication structures
abstract
Three digital artificial neural network processors suitable for the emulation of fully interconnected neural networks are proposed. The processors use N(2) multipliers and an arrangement of tree structures that provide the communication and accumulation function either individually or in a combined manner using communicating adder trees. The performance for the emulation of an N-neuron network for all processors is achieved in 2log(2)N+C time units, where C is a constant equal to the multiplication, neuron activation, and internal fixed delays. The feasibility and characteristics of the proposed configurations to emulate single and/or multiple neural networks simultaneously are discussed, and a comparison with recently proposed neurocomputer architectures is reported.
Gerald G. Pechanek, Stamatis Vassiliadis, José G. Delgado-Frias
IEEE Trans. Neural Networks3
1991 MPU: A N-Tuple Matching Processor
abstract
A novel matching processor (MPU) for dataflow architectures is proposed. This processor is capable of matching dataflow nodes requiring n-tuple inputs and reducing dataflow processor load. The main features of this processor include: n-tuple matching, associative matching, maskable match fields, storage reclamation capabilities, maskable ALU operations, and a programmable MPU architecture.>
Robert H. Payne, José G. Delgado-Frias
ICCD2
1991 AI in multimedia (panel session)
abstract
In this panel session, the following topics are discussed: artificial intelligence in business; artificial intelligence in multimedia; neural networks as a tool for artificial intelligence: software engineering for knowledge-based systems: and artificial intelligence as a solution for software engineering.>
Nikolaos G. Bourbakis, Robin Williams 0001, Forouzan Golshani, Myron Flickner, Ted Laliotis, Sukhan Lee 0001, José G. Delgado-Frias, Dan W. Hammerstrom, Cris Koutsougeras, Gerald G. Pechanek, Benjamin W. Wah, John Yen, Farokh B. Bastani, Tom Cooper, Karan Harbison-Briggs, Rudy Lauber, Alun D. Preece, Imran A. Zualkernan, Wei-Tek Tsai, Daniel E. Cooke, Martin Feather, Stephen Fickas, N. Minsky, Peter G. Selfridge, Douglas Smith
ICTAI7
1991 SPIN: a sequential pipelined neurocomputer
abstract
A novel digital network architecture, the sequential pipelined neurocomputer (SPIN), is proposed. The SPIN processor emulates neural networks, producing high performance with minimal hardware by sequentially processing each neuron in the modeled completely connected network with a pipelined physical neuron structure. In addition to describing SPIN, performance equations are estimated for the ring systolic, the recurrent systolic array, and the neuromimetic neurocomputer architectures, three previously reported schemes for the emulation of neural networks, and a comparison with the SPIN architecture is reported.>
Stamatis Vassiliadis, Gerald G. Pechanek, José G. Delgado-Frias
ICTAI3
1988 BVE: a wafer-scale engine for differential equation computation
abstract
With the advent of specialized VLSI and WSI hardware components the finite difference algorithms for solving differential equations become more attractive. This paper presents a novel computer architecture dedicated to compute boundary value problems. Improvement in speed and reliability can be obtained by means of WSI technology; however, a fault/defect tolerant scheme must be designed. Here we present a time redundancy approach which uses all the available computational resources; this is there is no spare or idle processors.
José G. Delgado-Frias, D. M. Green
ICS1
1988 Parallel architectures for AI semantic network processing
José G. Delgado-Frias, Will R. Moore
Knowl. Based Syst.1