EDBT 2026 Demo / reviewers in the wild / expert
Chuan-lin Wu
dblp:48/3621
· DBLP profile ↗
44ranked-venue papers
6as first author
0since 2021 · last 1997
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 6 first-authorSoftware engineering, systems software and programming languages · 5Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
22 papers |
Interconnection networks and networks-on-chip · 27% Electronic design automation · 22% Parallel and multicore computing · 15% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 59% Runtime systems and virtual machines · 26% Programming languages and type systems · 13% | |
| Theoretical computer science
2 papers |
Distributed computing theory · 63% Mathematical optimization · 37% | |
| Computer networks
1 paper |
Internet architecture and protocols · 100% |
Topics — the 30 heaviest of 67, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Interconnection networks and networks-on-chip › switching network
multistage interconnection network |
0.0 | 10 | 1992 | Performance Analysis of Multistage Interconnection Network Configurations and Operations · IEEE Trans. Computers 1992 A Fault-Tolerant Mapping Scheme for a Configurable Multiprocessor System · IEEE Trans. Computers 1989 A Conflict-Free Routing Scheme on Multistage Interconnection Networks · IEEE Trans. Computers 1989 |
Electronic design automation › physical design
routing |
0.0 | 2 | 1995 | Routing in a Three-Dimensional Chip · IEEE Trans. Computers 1995 A Microprocessor-Controlled Asynchronous Circuit Switching Network · ISCA 1979 |
Internet architecture and protocols › ATM networks
ATM switch |
0.0 | 1 | 1995 | Design and implementation of a multicast-buffer ATM switch · ICNP 1995 |
Internet architecture and protocols
multicast |
0.0 | 1 | 1995 | Design and implementation of a multicast-buffer ATM switch · ICNP 1995 |
Electronic design automation › physical design › routing › multilayer routing
3d routing |
0.0 | 1 | 1995 | Routing in a Three-Dimensional Chip · IEEE Trans. Computers 1995 |
Electronic design automation › physical design › routing
channel routing |
0.0 | 1 | 1995 | Routing in a Three-Dimensional Chip · IEEE Trans. Computers 1995 |
Electronic design automation
physical design |
0.0 | 1 | 1995 | Routing in a Three-Dimensional Chip · IEEE Trans. Computers 1995 |
Compilers and program optimization › memory optimization
data locality optimization |
0.0 | 1 | 1994 | The Classification, Fusion, and Parallelization of Array Language Primitives · IEEE Trans. Parallel Distributed Syst. 1994 |
Compilers and program optimization › loop transformation
loop fusion |
0.0 | 1 | 1994 | The Classification, Fusion, and Parallelization of Array Language Primitives · IEEE Trans. Parallel Distributed Syst. 1994 |
Compilers and program optimization › deep learning compiler
operator fusion |
0.0 | 1 | 1994 | The Classification, Fusion, and Parallelization of Array Language Primitives · IEEE Trans. Parallel Distributed Syst. 1994 |
Distributed systems › distributed coordination
conflict resolution |
0.0 | 2 | 1992 | Performance Analysis of Multistage Interconnection Network Configurations and Operations · IEEE Trans. Computers 1992 On a Class of Multistage Interconnection Networks · IEEE Trans. Computers 1980 |
Mathematical optimization
nonconvex optimization |
0.0 | 1 | 1993 | Approximation, Dimension Reduction, and Nonconvex Optimization Using Linear Superpositions of Gaussians · IEEE Trans. Computers 1993 |
Hardware reliability and fault tolerance
reconfiguration |
0.0 | 2 | 1989 | A Fault-Tolerant Mapping Scheme for a Configurable Multiprocessor System · IEEE Trans. Computers 1989 Reconfiguration Procedures for a Polymorphic and Partitionable Multiprocessor · IEEE Trans. Computers 1986 |
Parallel and multicore computing
multiprocessor system |
0.0 | 2 | 1988 | A Distributed Resource Management Mechanism for a Partitionable Multiprocessor System · IEEE Trans. Computers 1988 Reconfiguration Procedures for a Polymorphic and Partitionable Multiprocessor · IEEE Trans. Computers 1986 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 1992 | Limitation of superscalar microprocessor performance · MICRO 1992 |
Runtime systems and virtual machines › garbage collection
distributed garbage collection |
0.0 | 1 | 1991 | A Parallel Asynchronous Garbage Collection Algorithm for Distributed Systems · IEEE Trans. Knowl. Data Eng. 1991 |
Runtime systems and virtual machines
garbage collection |
0.0 | 1 | 1991 | A Parallel Asynchronous Garbage Collection Algorithm for Distributed Systems · IEEE Trans. Knowl. Data Eng. 1991 |
Programming languages and type systems
logic programming |
0.0 | 1 | 1991 | A Parallel Execution Model of Logic Programs · IEEE Trans. Parallel Distributed Syst. 1991 |
Distributed systems
distributed algorithms |
0.0 | 1 | 1991 | A Parallel Asynchronous Garbage Collection Algorithm for Distributed Systems · IEEE Trans. Knowl. Data Eng. 1991 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 1991 | Distributed Instruction Set Computer Architecture · IEEE Trans. Computers 1991 |
Parallel and multicore computing › parallel computing › parallel programming languages
parallel logic programming |
0.0 | 1 | 1991 | A Parallel Execution Model of Logic Programs · IEEE Trans. Parallel Distributed Syst. 1991 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 1991 | A Parallel Execution Model of Logic Programs · IEEE Trans. Parallel Distributed Syst. 1991 |
Image and video coding
image compression |
0.0 | 1 | 1990 | Oriented Non-Radial Basis Functions for Image Coding and Analysis · NIPS 1990 |
Parallel and multicore computing › parallel scheduling
communication scheduling |
0.0 | 1 | 1989 | A Conflict-Free Routing Scheme on Multistage Interconnection Networks · IEEE Trans. Computers 1989 |
Interconnection networks and networks-on-chip › routing algorithms
conflict-free routing |
0.0 | 1 | 1989 | A Conflict-Free Routing Scheme on Multistage Interconnection Networks · IEEE Trans. Computers 1989 |
Hardware reliability and fault tolerance › fault-tolerant architecture
fault-tolerant multiprocessor |
0.0 | 1 | 1989 | A Fault-Tolerant Mapping Scheme for a Configurable Multiprocessor System · IEEE Trans. Computers 1989 |
Interconnection networks and networks-on-chip › network-on-chip design
mapping |
0.0 | 1 | 1989 | A Fault-Tolerant Mapping Scheme for a Configurable Multiprocessor System · IEEE Trans. Computers 1989 |
Parallel and multicore computing
parallel algorithms |
0.0 | 1 | 1989 | A Conflict-Free Routing Scheme on Multistage Interconnection Networks · IEEE Trans. Computers 1989 |
Distributed computing theory
mutual exclusion |
0.0 | 1 | 1989 | Token Systems that Self-Stabilize · IEEE Trans. Computers 1989 |
Distributed computing theory
self-stabilization |
0.0 | 1 | 1989 | Token Systems that Self-Stabilize · IEEE Trans. Computers 1989 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.0window control · 0.0classification scheme · 0.0frame inheritance · 0.0data dependency graph · 0.0simulated annealing · 0.0oriented non-radial basis functions · 0.0NP-completeness reduction · 0.0radial basis function networks · 0.0linear projection · 0.0gaussian kernel · 0.0markov chain analysis · 0.0state-transition systems · 0.0state transition system · 0.0quadtree · 0.0postcompiler dependency detection · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 1997 | Evaluation of a memory hierarchy for the MTS multithreaded processorabstractExecuting multiple threads simultaneously on superscalar processor can improve hardware resource utilization and instruction throughput. The Multi-Threaded Superscalar (MTS) processor efficiently achieves the concurrent execution of multiple instruction streams using a VLIW, multiple functional unit architecture. However, the limitations of the memory system may impede the potential performance of the MTS processor. An interactive, parameter-driven simulator of the MTS architecture was developed using SES/workbench. A set of numerical benchmarks was run on it with varying memory system configurations. Assuming a single instruction cache, a single data cache, and one instruction queue per thread, varied parameters included the size of the instruction queues, the number of ports to the instruction cache, main memory latency, and cache hit rates. Based on simulation results, optimal values were chosen for certain parameters. To reasonably utilize the MTS processor, the memory system must provide at least 64 instruction bytes per cycle, though the demands for data access are far less severe. For reasonable memory speeds, this requires a roughly IMB three-ported instruction cache capable of providing 128 bits per port per cycle, as well as an IMB single-ported 32-bit wide data cache. A more realistic multilevel cache hierarchy is proposed. W. Lynn Gallagher, Chuan-lin Wu |
ICPADS | 2 |
| 1995 | A modular growth architecture for an ATM switchabstractA growable ATM switch architecture is desired for constructing a large scale ATM switch. All the current ATM switch architectures are based on uniform connections of unit switch elements. When the switch size becomes large, the uniform switch architecture is no longer suitable for the nonuniform traffic conditions. In this paper, we propose a nonuniform modular growth ATM switch architecture by using a knockout switch as the basic unit. All the internal traffic of the proposed switch has the best delay and throughput performance. An analysis is also presented to study the relationship between the cell-loss performance and the switch parameters. The knockout principle proves to be very efficient for the nonuniform concentration in the proposed switch architecture. The proposed switch architecture is cost-effective compared with the growable switch architecture which employs the uniform connection pattern. Chuan-lin Wu |
ICCCN | 2 |
| 1995 | A novel architecture for an ATM switchabstractMulticast function is essential for an ATM switch. We propose a novel architecture for an output-buffer ATM switch and a shared-buffer ATM switch to realize the multicast function in a more efficient way. In an output-buffer ATM switch, we dedicate a first-in and first-out (FIFO) shared buffer for all multicast cells to increase buffer utilization. In a shared-buffer ATM switch, we dedicate a FIFO address queue for all multicast cells to simplify the design of the control logic. Performance evaluation of the new switch is also provided. Since each multicast cell occupies only one buffer space, the proposed switch achieves a better cell-loss performance under multicast traffic loads without the need for complicated control circuitry. Chuan-lin Wu |
ICCD | 2 |
| 1995 | Design and implementation of a multicast-buffer ATM switchabstractWe propose a multicast-buffer ATM switch, where a dedicated buffer is allocated to store all multicast cells in an ATM switch. We describe the design of the dedicated multicast buffer in an output-buffer ATM switch and in a shared-buffer ATM switch. No additional copy circuit or complicated control circuit is needed to implement the multicast-buffer ATM switch. Our performance evaluation shows that the proposed switch has a low cell-loss ratio. We are implementing an eight-by-eight output buffer ATM switch with the multicast buffer. We also present a window control mechanism to further improve the utilization of the multicast buffer. Therefore, we believe the multicast-buffer ATM switch is a good candidate to support the multicast function for future broadband ISDN applications. Chuan-lin Wu |
ICNP | 2 |
| 1995 | Routing in a Three-Dimensional ChipabstractAs the very large scale integration (VLSI) technology approaches its fundamental scaling limit at about 0.2 /spl mu/m, it is reasonable to consider three-dimensional (3-D) integration to enhance packing density and speed performance. With additional functional units packed into one chip in a 3-D space, computer-aided design (CAD) tools are demanded to ease the complicated design work. This paper presents a 100% completion achievable routing methodology. The routing methodology is based on the two-dimensional (2D) channel routing methodology; thus, it is called a 3-D channel routing methodology. With the routing methodology, a 3-D routing problem is decomposed into two 2D routing subproblems: intra-layer routing that interconnects terminals on the same layer, which can be done by using a 2-D channel router, and inter-layer routing that interconnects terminals on different layers. The inter-layer routing problem is transformed into a 2-D channel routing problem and the transformation is made in some 3-D channels. Detailed discussions are given for the 3-D to 2-D transformation. Optimization of the transformation is shown to be NP-complete. Thus, simulated annealing is used to optimize the transformation.> Chao Chi Tong, Chuan-lin Wu |
IEEE Trans. Computers | 2 |
| 1994 | Design and Evaluation of a Multiprocessor Architecture with Decentralized ControlabstractThis paper presents design and evaluation of a decentralized control mechanism of a variable-topology (phase-reconfigurable) multiprocessor architecture. We design and implement an active crossbar communication processor(XCP) which encodes a set of communication instructions. The new interconnection network approach provides an apt phase-reconfiguration method. The new method is distributed in nature and does not require a global control bus that is a time consuming feature as verified in many existing systems. A conjugate-gradient algorithm for solving a linear system equations is used to perform evaluation and comparison. Hsiao-chen Chung, Chuan-lin Wu, James Rakes, Peter J. Zievers, Yin-Kuan Lin |
ICPP (1) | 2 |
| 1994 | Statement Merge: an Inter-Statement Optimization of Array Language ProgramsabstractArray language provides an effective paradigm for developing portable parallel programs but will only be used if compilers generate efficient code for array operations this paper presents an inter-statement optimization technique which merges related array operations in different statements. A merged statement creates opportunities for previously developed inter-primitive optimizations of array operators The potential advantages of statement merge are the elimination of a program variable, demand driven evaluation, reductions of data movement and loop overhead. The combination of statement merge and inter-primitive optimizations achieves a 3 8 fold execution time speedup in a scientific application Roy Dz-Ching Ju, Chuan-lin Wu, Paul R. Carini |
ICPP (2) | 2 |
| 1994 | The Classification, Fusion, and Parallelization of Array Language PrimitivesabstractWe present a classification scheme for array language primitives that quantifies the variation in parallelism and data locality that results from the fusion of any two primitives. We also present an algorithm based on this scheme that efficiently determines when it is beneficial to fuse any two primitives. Experimental results show that five LINPACK routines report 50% performance improvement from the fusion of array operators.> Roy Dz-Ching Ju, Chuan-lin Wu, Paul R. Carini |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1993 | Reconfigurable Branch Processing Strategy in Super-Scalar MicroprocessorsabstractIn this paper, we develop a model for measuring branch performance on super-scalar processors. This model takes the form of CP I equations for different branch processing strategies. This model is the basis for a proposed design wherein the branch processing strategy in a processor can be reconfigured for optimal performance over a varying application load. There are nine parameters which are required to describe performance of a given branch processing method -four are dependent primarily on the processor architecture, two depend on the application program, and three depend primarily on the branch instance. The highest performance branch prediction algorithm can be determined by a combination of application profiling and runtime statistic gathering Terence M. Potter, Hsiao-chen Chung, Chuan-lin Wu |
ICPP (1) | 3 |
| 1993 | Approximation, Dimension Reduction, and Nonconvex Optimization Using Linear Superpositions of GaussiansabstractThis paper concerns neural network approaches to function approximation and optimization using linear superposition of Gaussians (or what are popularly known as radial basis function (RBF) networks). The problem of function approximation is one of estimating an underlying function f, given samples of the form ((y/sub i/, x/sub i/); i=1,2,...,n; with y/sub i/=f(x/sub i/)). When the dimension of the input is high and the number of samples small, estimation of the function becomes difficult due to the sparsity of samples in local regions. The authors find that this problem of high dimensionality can be overcome to some extent by using linear transformations of the input in the Gaussian kernels. Such transformations induce intrinsic dimension reduction, and can be exploited for identifying key factors of the input and for the phase space reconstruction of dynamical systems, without explicitly computing the dimension and delay. They present a generalization that uses multiple linear projections onto scalars and successive RBF networks (MLPRBF) that estimate the function based on these scaler values. They derive some key properties of RBF networks that provide suitable grounds for implementing efficient search strategies for nonconvex optimization within the same framework.> Avijit Saha, Chuan-lin Wu, Dun-Sung Tang |
IEEE Trans. Computers | 2 |
| 1992 | The Synthesis of Array Functions and Its Use in Parallel Computation
Roy Dz-Ching Ju, Chuan-lin Wu, Paul R. Carini |
ICPP (2) | 2 |
| 1992 | Limitation of superscalar microprocessor performance
Thang Tran, Chuan-lin Wu |
MICRO | 2 |
| 1992 | Performance Analysis of Multistage Interconnection Network Configurations and OperationsabstractA performance evaluation using both analytical and simulation models, of circuit-switching multistage interconnection networks in aspects of configurations and operations is presented. Two configurations of the networks, single and dual, are evaluated. Network operations considered include conflict resolution strategies and communication strategies. Two different conflict resolution strategies, drop and hold, are analyzed. New analyses using Markov chains, are given, and are verified by simulation results. In the single-network configuration, it is shown that the drop strategy is better than the hold strategy for randomly distributed traffic. In the dual network configuration, five different communication strategies are investigated, and the optimum performance level is shown to be dependent on the length of the data transfer time.> Chuan-lin Wu, Manjai Lee |
IEEE Trans. Computers | 1 |
| 1991 | A Benchmark Evaluation of a Multi-threaded RISC Processor Architecture
R. Guru Prasadh, Chuan-lin Wu |
ICPP (1) | 2 |
| 1991 | Microprocessor Architecture with Multi-Bit Scoreboard Concurrency Control
Thang Tran, Chuan-lin Wu |
ICPP (1) | 2 |
| 1991 | Distributed Instruction Set Computer ArchitectureabstractThe Distributed Instruction Set Computer Architecture (DISC) is proposed as a fine-grained multiprocessing computer architecture. DISC uses a parallel instruction set and a distributed control mechanism to explore fine-grained, parallel processing in a multiple-functional-unit system. Multiple instructions are executed in parallel and/or out of order at the highest speed of n instructions/cycle, where n is the number of functional units. Based on this architecture, a hardware system is developed. Extensive studies were conducted on a behavioral DISC system model to investigate the performance level, the effect of program sizes, and the hardware utilization. Simulation showed that a DISC system incorporating 16 functional units can run 7.7 times faster than a single-functional-unit DISC system.> Chuan-lin Wu |
IEEE Trans. Computers | 2 |
| 1991 | A Parallel Asynchronous Garbage Collection Algorithm for Distributed SystemsabstractThe problem of distributed garbage collection is discussed. An algorithm for parallel distributed asynchronous garbage collection is presented. The liveness and safety properties of this method are demonstrated. The algorithm does not require a global clock, complex termination detection methods, or distributed synchronization techniques. A new color code is introduced to distinguish between local cells (black) and those that are exclusively accessible from the remote pointers (gray). The mutator operation is revised to handle a multiple mutator scheme on a given local memory. Simulation results show that the developed distributed and parallel algorithm performs much better than the sequential method as tested on a Balance 8000 computer.> Nader Bagherzadeh, Seng-lai Heng, Chuan-lin Wu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1991 | A Parallel Execution Model of Logic ProgramsabstractA parallel-execution model that can concurrently exploit AND and OR parallelism in logic programs is presented. This model employs a combination of techniques in an approach to executing logic problems in parallel, making tradeoffs among number of processes, degree of parallelism, and combination bandwidth. For interpreting a nondeterministic logic program, this model (1) performs frame inheritance for newly created goals, (2) creates data-dependency graphs (DDGs) that represent relationships among the goals, and (3) constructs appropriate process structures based on the DDGs. (1) The use of frame inheritance serves to increase modularity. In contrast to most previous parallel models that have a large single process structure, frame inheritance facilitates the dynamic construction of multiple independent process structures, and thus permits further manipulation of each process structure. (2) The dynamic determination of data dependency serves to reduce computational complexity. In comparison to models that exploit brute-force parallelism and models that have fixed execution sequences, this model can reduce the number of unification and/or merging steps substantially. In comparison to models that exploit only AND parallelism, this model can selectively exploit demand-driven computation, according to the binding of the query and optional annotations. (3) The construction of appropriate process structures serves to reduce communication complexity. Unlike other methods that map DDGs directly onto process structures, this model can significantly reduce the number of data sent to a process and/or the number of communication channels connected to a process.> Albert C. Chen 0003, Chuan-lin Wu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1990 | Oriented Non-Radial Basis Functions for Image Coding and Analysis
Avijit Saha, Jim Christian, Dun-Sung Tang, Chuan-lin Wu |
NIPS | 4 |
| 1989 | Expert system based automatic network fault management systemabstractAn expert system for network management is designed and prototyped to do network troubleshooting automatically. The expert system employs management information provided by a monitoring mechanism of the network. The whole spectrum of the fault management information is analyzed. The management knowledge derived is categorized into five types: the physical property, the experience in the past, the heuristic rule of thumb, the predictable problems, and the deep knowledge. These knowledge types and related rules are divided into groups to improve reasoning speed. The expert system is composed of a problem manager, a problem analyzer, and many problem solvers. A prototyped expert system, using simulated faults, shows that the expert-system-based fault management system can automatically diagnose problems and take corrective actions.> Show-Way Yeh, Chuan-lin Wu, Hong-Da Sheng, Chaw-Kwei Hung, Rei-Chi Lee |
COMPSAC | 2 |
| 1989 | Token Systems that Self-StabilizeabstractPresents a novel class of mutual exclusion systems, in which processes circulate one token, and each process enters its critical section when it receives the token. Each system in the class is self-stabilizing; i.e. it it starts at any state, possibly one where many tokens exist in the system, it is guaranteed to converge to a good state where exactly one token exists in the system. The systems are better than previous systems in that their state transitions are noninterfering; i.e., if any state transition is enabled at any instant, then it will continue to be enabled until it is executed. This makes the systems easier to implement as delay-insensitive circuits.> Geoffrey M. Brown, Mohamed G. Gouda, Chuan-lin Wu |
IEEE Trans. Computers | 3 |
| 1989 | A Conflict-Free Routing Scheme on Multistage Interconnection NetworksabstractA conflict-free routing scheme is presented for a class of parallel and distributed computing systems. The core of the scheme is a quadtree communication structure. The quadtree structure suggests a general approach to mapping a class of parallel algorithms with intensive communication requirements for selecting data from many different sources and distributing data from a single source. By properly merging messages and efficiently replicating data, the quadtree structure can complete required communications in O(log/sub 4/ M) parallel steps, where M is the network size. It is shown that the size of a quadtree communication structure can be contracted and stretched by adjusting the number of descendent nodes without affecting its conflict-free property. The relationship between the computation/communication ratio of various parallel algorithms and the number of tree levels is presented, and finally, their joint effect on the response time of combining and distributing data messages is examined. This analysis helps determine the optimal adaptation of the quadtree for minimizing the overall algorithm execution time.> Woei Lin, Tsang-Ling Sheu, Chita R. Das, Tse-Yun Feng, Chuan-lin Wu |
IEEE Trans. Computers | 5 |
| 1989 | A Fault-Tolerant Mapping Scheme for a Configurable Multiprocessor SystemabstractA fault-tolerant mapping scheme for a configurable multiprocessor system using multistage interconnection networks is presented. By adapting its interprocessor connections, the multiprocessor system can provide many regular topological configurations suitable for a variety of parallel computation applications. The configurability of the system is achieved by applying a set of configuration procedures to a linear address space of the system. The central idea behind the scheme is the use of two transformations to restore the linear address space in the presence of processor failures. The fault-tolerant mapping scheme is composed of three algorithms. The algorithms adaptively use the two transformations to handle three different types of faults: single faults, double faults, and triple or greater faults. It is shown that when there are a few processor failures, the algorithms can effectively achieve fault-free linear subspaces with graceful degradation.> Woei Lin, Chuan-lin Wu |
IEEE Trans. Computers | 2 |
| 1988 | Adaptive Checkpointing and Rollback in Multiprocessor Systems
Chung-Yang Chiang, Chuan-lin Wu |
ICPP (1) | 2 |
| 1988 | Distributed Instruction Set Computer
Chuan-lin Wu |
ICPP (1) | 2 |
| 1988 | I-NET mechanism for issuing multiple instructionsabstractConventional instruction issuing methods use hardware control mechanism to issue instructions in multiple-functional-unit systems. They reach physical limitations due to the complexity of issuing logic when they intend to issue multiple instruction per cycle. A method called I-NET is proposed to overcome this shortcoming. I-NET uses a postcompiler to detect the data dependencies among instructions. The detected data dependence is then attached to the instruction code to form an instruction unit that can be executed independently with other such units. These precoded instruction units are fetched from the memory and sent to functional units directly. A maximal instruction-issuing rate is achieved by avoiding a lengthy decoding procedure in the hardware. The performance study has shown I-NET to be a very promising method for exploring massive parallelism of a program in a multiple-functional-unit system.> Chuan-lin Wu |
SC | 2 |
| 1988 | A Distributed Resource Management Mechanism for a Partitionable Multiprocessor SystemabstractA resource-management mechanism is presented for a multiprocessor system consisting of a pool of homogeneous processing elements interconnected by multistage networks. The mechanism aims at making effective use of hardware resources of the multiprocessor system in support of high-performance parallel computations. It can create many physically independent subsystems simultaneously without incurring internal fragmentation,. Each subsystem can configure itself to form a desired topology for matching the structure of the parallel computation. The mechanism is distributed in nature; it is divided into three functionally disjoint procedures that can reside in different loci for handling various resource-management tasks concurrently. Simulation results show that, by eliminating internal fragmentation, the mechanism achieves better source utilization than a reference machine.> Woei Lin, Chuan-lin Wu |
IEEE Trans. Computers | 2 |
| 1986 | Operating System Kernel for a Reconfigurable Multiprocessor System
Geoffrey M. Brown, Chuan-lin Wu |
ICPP | 2 |
| 1986 | Fail Safe Distributed Fault Diagnosis of Multiprocessor Systems
Chung-Yang Chiang, Chuan-lin Wu |
ICPP | 2 |
| 1986 | Efficient Execution of Programs with Pipeline Configuration of Reconfigurable Multiprocessor
Karam Mossaad, Chuan-lin Wu |
ICPP | 2 |
| 1986 | Reconfiguration Procedures for a Polymorphic and Partitionable MultiprocessorabstractThis correspondence presents a collection of reconfiguration procedures for a multiprocessor which employs multistage interconnection networks. These procedures are used to dynamically partitipn the multiprocessor into many subsystems, and reconfigure them to form a variety commonly used topologies to match task graphs. By examining the switching capability of the interconnection network, design rules for avoiding connection conflicts are exploited. Then, on the basis of these rules, parallel procedures are designed. With the procedures, a subsystem can be reconfigured in the form of the desired topologies without interfering with other subsystems. In addition, the reconfiguration of a subsystem can be accomplished in constant time, independently of subsystem size. Woei Lin, Chuan-lin Wu |
IEEE Trans. Computers | 2 |
| 1985 | Network Facility for a Reconfigurable Computer Architecture
Manjai Lee, Eric Fiene, Chuan-lin Wu, Geoffrey Brown, Nader Bagherzadeh |
ICDCS | 3 |
| 1985 | Design of Configuration Algorithms of Commonly-Used Topologies for a Multiprocessor : STAR
Woei Lin, Chuan-lin Wu |
ICPP | 2 |
| 1984 | Performance Analysis of Circuit Switching Baseline Interconnection NetworksabstractPerformance evaluation, using both analytical and simulation models, of circuit switching baseline networks is presented. Two configurations of the baseline networks, single and dual, are evaluated. In each configuration, two different conflict resolution strategies, drop and hold, are tried to see the performance difference. Our analytical models are based on a more realistic assumption. New analyses are given and are verified by simulation results. In single network configuration, it is shown that the drop strategy is better than the hold strategy in the case that the data transfer time is longer than 10 cycles under a high request rate. In the dual network configuration, five different communication strategies are investigated and the optimum performance level is shown to be dependent on the length of the data transfer time. Manjai Lee, Chuan-lin Wu |
ISCA | 2 |
| 1983 | Configuring Computation Tree Topologies on a Distributed Computing System
Woei Lin, Chuan-lin Wu |
ICPP | 2 |
| 1982 | Distributed circuit switching starnet
Chuan-lin Wu, Woei Lin, Min-Chang Lin |
ICPP | 1 |
| 1982 | Design of a 2 × 2 fault-tolerant switching elementabstractThis paper describes the architecture of a 2 × 2 fault-tolerant switching element which can be used to modularly construct interconnection networks for multi-processing and local computer networking. The switching element uses distributed control and circuit switching. Its good gate-to-pin ratio can facilitate VLSI implementation. Woei Lin, Chuan-lin Wu |
ISCA | 2 |
| 1982 | Star: A Local Network System for Real-Time Management of Imagery DataabstractOverall architecture of a local computer network, Star, is described. The objective is to accomplish a cost-effective system which provides multiple users a real-time service of manipulating very large volume imagery information and data. Star consists of a reconfigurable communication subnet (Starnet), heterogeneous resource units, and distributed-control software entities. Architectural aspects of a fault-tolerant communication subnet, distributed database management, and a distributed scheduling strategy for configuring desirable computation topology are exploited. A model for comparing cost-effectiveness among Starnet, crossbar, and multiple buses is included. It is concluded that Starnet outperforms the other two when the number of units to be connected is larger than 64. This project serves as a research tool for using current and projected technology to innovate better schemes for parallel image processing. Chuan-lin Wu, Tse-Yun Feng, Min-Chang Lin |
IEEE Trans. Computers | 1 |
| 1981 | Fault-Diagnosis for a Class of Multistage Interconnection NetworksabstractTo study the fault-diagnosis method for a class of multistage interconnection networks a general fault model is first constructed. Specific steps for diagnosing single faults and detecting multiple faults in interconnection networks such as the indirect binary n-cube network and the flip network are then developed. The following results are derived in this study: 1) independent of the network size, only four tests are required for detecting a single fault; 2) the number of tests required for locating a single fault and determining the fault type ranges from 4 to max(12, 6 + 2 ⌈log2(log2N)⌉) except for four types of single faults in the switching elements which cannot be pinpointed at the switching element level where N is the number of inputs/outputs; 3) only four tests are required for locating a single fault if the switching element is designed in such a way that any physical defection of the switching element causes both outputs of the related switching element to be faulty; and 4) multiple faults can be detected by 2(1 + log2N) tests. Tse-Yun Feng, Chuan-lin Wu |
IEEE Trans. Computers | 2 |
| 1981 | The Universality of the Shuffle-Exchange NetworkabstractThis paper has focused on the realization of every arbitrary permutation with the shuffle-exchange network. Permutation properties of shuffle-exchange networks are studied and are used to demonstrate several universal networks. It is concluded that 3(log2 N) –1 passes through a single-stage regular shuffle exchange network are sufficient to realize every arbitrary permutation where N is network size. A routing algorithm is also developed to calculate control settings of the shuffle-exchange switches for the permutation realization. Three optimal universal networks, namely, expanded direct- connection shuffle-exchange network, multiple-pass omega network, and modified shuffle-exchange network are then exploited for better interconnection purposes. In addition, this work specifies the inherent relationship between the shuffle-exchange network and the Benes binary network so that designers can have a broad prospect. Chuan-lin Wu, Tse-Yun Feng |
IEEE Trans. Computers | 1 |
| 1980 | On a Class of Multistage Interconnection NetworksabstractA baseline network and a configuration concept are introduced to evaluate relationships among some proposed multistage interconnection networks. It is proven that the data manipulator (modified version), flip network, omega network, indirect binary n-cube network, and regular SW banyan network (S = F = 2) are topologically equivalent. The configuration concept facilitates developing a homogeneous routing algorithm which allows one-to-one and one- to-many connections from an arbitrary side of a network to the other side. This routing algorithm is extended to full communication which allows connections between terminals on the same side of a network. A conflict resolution scheme is also included. Some practical implications of our results are presented for further research. Chuan-lin Wu, Tse-Yun Feng |
IEEE Trans. Computers | 1 |
| 1980 | The Reverse-Exchange Interconnection NetworkabstractProperties of the reverse-exchange interconnection network are used to develop a reconfiguration scheme and a two-pass structure for enhancing the efficiency of a class of multistage interconnection networks. Functional relationships among a class of multistage interconnection networks are first derived. According to the functional relationships, we propose a reconfiguration scheme which enables a network to accomplish various interconnection functions of other networks. Then the admissible permutations along with related recursive control algorithms of the reverse-exchange interconnection network are specified through a set of theorems. Using the reverse-exchange property, we also prove that the algorithms actually work. Finally, we prove that arbitrary permutations can be realized in two passes (or 2 · 1og2N switching steps where N is the network size). By taking advantage of Benes network control algorithms, a way to control the two-pass structure is also developed. Chuan-lin Wu, Tse-Yun Feng |
IEEE Trans. Computers | 1 |
| 1979 | A Microprocessor-Controlled Asynchronous Circuit Switching NetworkabstractThis paper describes an asynchronous circuit switching network for multiple-processor systems. Several circuit switching networks for various applications have been proposed and constructed. However, there are problems associated with these networks. The asynchronous circuit switching network possesses several features that can solve these problems. A three-stage fully connected topology is utilized to construct the network. Each switching element is functionally and physically identical and this facilitates a cost-effective LSI implementation and software development. The control structure of the switching element and the routing algorithm are re-organized to fit the asynchronous operation. The network takes advantage of low-cost microcomputers to do the distributed routing control and to implement the communication protocols. The graceful degradation characteristic in the network is provided by independent multiple paths existing between any pair of the source and the destination. A network monitor is incorporated to facilitate an adaptive routing strategy and to have the fault diagnostic capability. Three alternatives for the switching element implementation are described to demonstrate the hardware and software tradeoffs. Tse-Yun Feng, Chuan-lin Wu, Dharma P. Agrawal |
ISCA | 2 |
| 1978 | A survey of communication processor systemsabstractIn this paper, various Communication Processor Systems (CPS) architectures are reviewed and classified according to their interconnection organizations. The evolution and future development of the CPS architectures are discussed and the intercommunication subsystem requirement, interconnection networks structure, network hardware and software design aspects are carefully examined. The functional baseline for the processing unit of the CPS is identified and a comparative study of various existing and proposed processing units is included. The critical review is based on the important features such as the reliability improvement, the interrupt handling, the priority encoding, and the optimized repertoire, etc. Finally, general software considerations are also outlined. Dharma P. Agrawal, Tse-Yun Feng, Chuan-lin Wu |
COMPSAC | 3 |