Ding-Ming Kwai

dblp:25/3970 · DBLP profile ↗
← Back
48ranked-venue papers
5as first author
0since 2021 · last 2020
0000-0001-7769-7879ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 46 · 3 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Integrated circuit design · 27% Electronic design automation · 26% Energy-efficient computing · 21%

Topics — the 25 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
hardware verification and test
0.322013
Parametric Delay Test of Post-Bond Through-Silicon Vias in 3-D ICs via Variable Output Thresholding Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Small delay testing for TSVs in 3-D ICs · DAC 2012
Integrated circuit design › 3d integration
through-silicon via
0.322013
Parametric Delay Test of Post-Bond Through-Silicon Vias in 3-D ICs via Variable Output Thresholding Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Small delay testing for TSVs in 3-D ICs · DAC 2012
Memory systems
DRAM
0.212014
DArT: A Component-Based DRAM Area, Power, and Timing Modeling Tool · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Embedded and real-time systems
real-time scheduling
0.212014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Energy-efficient computing › thermal management
thermal-aware scheduling
0.212014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Energy-efficient computing › thermal management
thermal-aware task allocation
0.212014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Energy-efficient computing
thermal management
0.212014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Integrated circuit design › 3d integration
3-D stacked IC
0.212013
Parametric Delay Test of Post-Bond Through-Silicon Vias in 3-D ICs via Variable Output Thresholding Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Electronic design automation › hardware verification and test
design for testability
0.212013
Parametric Delay Test of Post-Bond Through-Silicon Vias in 3-D ICs via Variable Output Thresholding Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Integrated circuit design
3d integration
0.112012
Small delay testing for TSVs in 3-D ICs · DAC 2012
Electronic design automation › hardware verification and test
delay fault testing
0.112012
Small delay testing for TSVs in 3-D ICs · DAC 2012
Parallel and multicore computing › task scheduling
many-core scheduling
0.112014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Electronic design automation › design optimization
power, area and timing modeling
0.112014
DArT: A Component-Based DRAM Area, Power, and Timing Modeling Tool · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Interconnection networks and networks-on-chip
network topology
0.122001
A Unified Formulation of Honeycomb and Diamond Networks · IEEE Trans. Parallel Distributed Syst. 2001
Periodically Regular Chordal Rings · IEEE Trans. Parallel Distributed Syst. 1999
Interconnection networks and networks-on-chip › network topology › loop networks
chordal ring
0.011999
Periodically Regular Chordal Rings · IEEE Trans. Parallel Distributed Syst. 1999
Parallel and multicore computing › parallel architecture
linear array processors
0.011999
Data-Driven Control Scheme for Linear Arrays: Application to a Stable Insertion Sorter · IEEE Trans. Parallel Distributed Syst. 1999
Parallel and multicore computing › parallel algorithms › sorting
parallel sorting
0.011999
Data-Driven Control Scheme for Linear Arrays: Application to a Stable Insertion Sorter · IEEE Trans. Parallel Distributed Syst. 1999
Electronic design automation › physical design
routing
0.011999
Periodically Regular Chordal Rings · IEEE Trans. Parallel Distributed Syst. 1999
Integrated circuit design
VLSI design
0.011999
Data-Driven Control Scheme for Linear Arrays: Application to a Stable Insertion Sorter · IEEE Trans. Parallel Distributed Syst. 1999
Interconnection networks and networks-on-chip › routing algorithms
wormhole routing
0.011999
Periodically Regular Chordal Rings · IEEE Trans. Parallel Distributed Syst. 1999
Interconnection networks and networks-on-chip › network topology
torus network
0.012001
A Unified Formulation of Honeycomb and Diamond Networks · IEEE Trans. Parallel Distributed Syst. 2001
Parallel and multicore computing › parallelizing compiler
dependence graph
0.011992
Data Flow Representation of Iterative Algorithms for Systolic Arrays · IEEE Trans. Computers 1992
High-performance computing
iterative methods
0.011992
Data Flow Representation of Iterative Algorithms for Systolic Arrays · IEEE Trans. Computers 1992
Hardware accelerators and domain-specific architectures
systolic array
0.011992
Data Flow Representation of Iterative Algorithms for Systolic Arrays · IEEE Trans. Computers 1992
Electronic design automation › physical design
VLSI layout
0.011999
Periodically Regular Chordal Rings · IEEE Trans. Parallel Distributed Syst. 1999

Methods — techniques the papers use, named apart from their topics

variable output thresholding · 0.3optimization modeling · 0.2online task migration · 0.2component-based modeling · 0.2circuit simulation · 0.2schmitt-trigger inverter · 0.2oscillation ring · 0.1SPICE simulation · 0.1interleaving · 0.1error correction · 0.1
YearPublicationVenuePosition
2020 Refresh Power Reduction of DRAMs in DNN Systems Using Hybrid Voting and ECC Method
abstract
A deep neural network (DNN) system typically needs dynamic random access memories (DRAMs) for the data buffering. In this paper, an error-correction-code (ECC)-based technique is proposed to reduce the refresh power of DRAMs in the DNN system by extending the refresh period. By taking advantage of the characteristics of weight data of DNNs, a hybrid voting and ECC (VECC) method is used to protect the weight data from data retention fault caused by the refresh period extension. Analysis results show that the VECC method can achieve about 93% refresh power saving with about 0.5% accuracy loss and smaller than 0.5% check bit overhead on AlexNet, ResNet, and VGG19 convolutional neural network CNN models trained by ImageNet data set.
Tsung-Fu Hsieh, Jin-Fu Li 0001, Jenn-Shiang Lai, Chih-Yen Lo, Ding-Ming Kwai, Yung-Fa Chou
ITC-Asia5
2018 A channel-sharable built-in self-test scheme for multi-channel DRAMs
abstract
Various multi-channel dynamic random access memories (MC-DRAMs) have been proposed for the demand of high bandwidth. In this paper, we propose a channel-sharable built-in self-test (BIST) scheme for MC-DRAMs. The BIST can apply test patterns and evaluate test responses for multiple channels simultaneously regardless of the difference of the read/write latency among the channels. Therefore, the proposed BIST can reduce the test time. In our simulation cases show that the proposed BIST scheme can achieve about 11% test time reduction in comparison with an existing conventional shared BIST scheme for a two-channel 1G-bit DRAM by consuming about 0.003% additional area cost.
Kuan-Te Wu, Jin-Fu Li 0001, Chih-Yen Lo, Jenn-Shiang Lai, Ding-Ming Kwai, Yung-Fa Chou
ASP-DAC5
2017 Heterogeneous chip power delivery modeling and co-synthesis for practical 3DIC realization
abstract
Three dimensional IC (3DIC) is becoming practical in today's consumer electronics designs. However, one major problem remains in design synthesis and flow: how to model heterogeneous die(s) with major logic die for power synthesis and signoff. This work provides a realistic model and principle for heterogeneous dies power network for 3DICs. It is based on given abstract or early stage information like bump location and power consumption from the provider. Our work also uses this model to synthesize power network with bottom logic die in the design flow. The result is DRC clean power network without IR and EM violation for all power domains. First, we analyze the location and power consumption of power bump for heterogeneous die(s). Second, according to previous analysis, we decide the stripe location and power sink location of heterogeneous dies model by a clustering method. After the initial model is synthesized, we convert it to a node graph with corresponding resistance of via and metal layer, also nodal voltages. Third, the model is optimized by using Sequential Linear Programming (SLP) to adjust stripe width. It will improve the model iteratively until the target IR-Drop is met. Furthermore, our work will create a pseudo DEF of the proposed model to be incorporated with the commercial tool for verification. We experiment on a real case from design house containing a 3D DRAM stack to demonstrate the effectiveness of this cross-layer realization. Results show that we can save 34% metal layer usage in one of the power domains in our case by using proposed methodology.
Wei-Hsun Liao, Chang-Tzu Lin, Sheng-Hsin Fang, Chien-Chia Huang, Hung-Ming Chen, Ding-Ming Kwai, Yung-Fa Chou
ASP-DAC6
2017 DLL-Assisted Clock Synchronization Method for Multi-Die ICs
abstract
For a multi-die IC, the chip-level clock synchronization problem that aims to establish a global clock signal across multiple functional dies is harder to achieve than its single-die counterpart. In this work, we investigate a process resilient solution for this problem by incorporating Delay-Locked Loops (DLLs). The basic idea is to insert a DLL (which can be generated by a DLL compiler) in each functional die so that the clock latency (from a clock source to the clock ports of a number of FFs) in different dies can be dynamically tuned and equalized. This method has a benefit that the clock network of each die can be designed independently, while the clock skew of the entire chip can still be minimized at run-time, in response to its operating environment. In a preliminary study, experimental results on a pseudo 4-die design demonstrates how the clock skew as high as 233ps initially can be reduced to 34ps after the application of the proposed method.
Chia-Yuan Cheng, Shi-Yu Huang, Ding-Ming Kwai, Yung-Fa Chou
ICCD3
2017 Software-hardware-cooperated built-in self-test scheme for channel-based DRAMs
abstract
Dynamic random access memory (DRAM) is one key component in modern electronic systems. In this paper, we propose a software-hardware-cooperated built-in self-test (SHC-BIST) scheme for the channel-based DRAMs. The testing of DRAMs consists of two major phases: DRAM initialization and DRAM array testing. Typically, the DRAM initialization process is short and executed in the beginning of the DRAM array testing. Thus, it is inefficient to realize it using the dedicated BIST hardware. On the other hand, it is not time efficient if we use the processor (software) to execute the DRAM array testing. Therefore, the SHC-BIST scheme uses a programmable BIST circuit to execute the DRAM array testing and takes advantage of the processor to execute the DRAM initialization and control the programmable BIST circuit such that the test time and hardware cost can be minimized. We verify the SHC-BIST scheme using a system with a LEON3 processor and a multi-channel DRAM.
Tsung-Fu Hsieh, Jin-Fu Li 0001, Kuan-Te Wu, Jenn-Shiang Lai, Chih-Yen Lo, Ding-Ming Kwai, Yung-Fa Chou
ITC-Asia6
2016 A Test Method for Finding Boundary Currents of 1T1R Memristor Memories
abstract
Memristor is a resistive device which is considered as an alternative non-volatile device for future non-volatile memories. For a memristor memory, a reference current is needed for discriminating the high-resistance (ROFF) state from low-resistance (RON) state. The reference current has an impact on the yield and reliability of the memristor memory. In this paper, we propose a test method in associate with a current comparing circuit for finding the boundary currents of ROFF and RON states. Therefore, the user can set a good reference current according to the boundary currents. Simulation results show that if our test method is used to identify the boundary currents, 2.05% and 3.68% memristor cells which may be read incorrectly due to process variation for ROFF/RON = 50 and 3 can be eliminated, respectively.
Tzu-Ying Lin, Yong-Xiao Chen, Jin-Fu Li 0001, Chih-Yen Lo, Ding-Ming Kwai, Yung-Fa Chou
ATS5
2016 A built-in self-repair scheme for DRAMs with spare rows, columns, and bits
abstract
With the shrinking of technology node, the data retention time of DRAM (DRAM) cells is widespread. Thus, the number of the cells with data retention faults is increased. In this paper, therefore, we propose a built-in self-repair (BISR) scheme for DRAMs using redundancies with physical and logical reconfiguration mechanisms. Spare rows and columns with physical reconfiguration mechanism are used to repair functional faults caused by defects. Spare bits with logical reconfiguration mechanism are used to replace data retention faults caused by process variation. Also, a diagnosis algorithm is proposed to identify data retention faults. Simulation results show that the proposed BISR scheme for a DRAM with 2 spare rows, 2 spare columns, and 8 spare bits can provide higher repair yield than a BISR scheme for a DRAM with 3 spare rows and 3 spare columns.
Chih-Sheng Hou, Yong-Xiao Chen, Jin-Fu Li 0001, Chih-Yen Lo, Ding-Ming Kwai, Yung-Fa Chou
ITC5
2014 Intra-channel Reconfigurable Interface for TSV and Micro Bump Fault Tolerance in 3-D RAMs
abstract
Three-dimensional (3-D) integration using through-silicon-via (TSV) is an emerging technology for integrated circuit (IC) design. It has been used in DRAM die stacking extensively. However, yield remains a key issue for volume production of 3-D RAMs. In this paper, we present a point-to-point interconnection structure derived from bus and propose a fault tolerance interface scheme for TSVs and micro bumps to enhance their manufacturing yield in the 3-D RAMs. The interconnection structure is inherently redundant and thus can replace defective TSVs or micro bumps without using repair circuits. Global and local reconfiguration approaches are proposed which benefit distinct situations of the 3-D RAM. Analyses show that the proposed intra-channel reconfigurable interconnection scheme can improve the yield of the 3-D RAM effectively. Compared to the previous solution using an inter-channel reconfigurable interconnection scheme, the yield improvement can be as large as 23% which is very significant.
Kuan-Te Wu, Jin-Fu Li 0001, Yun-Chao Yu, Chih-Sheng Hou, Chi-Chun Yang, Ding-Ming Kwai, Yung-Fa Chou, Chih-Yen Lo
ATS6
2014 BIST-Assisted Tuning Scheme for Minimizing IO-Channel Power of TSV-Based 3D DRAMs
abstract
Three-dimensional dynamic random access memory (3D DRAM) using through-silicon via (TSV) has been acknowledged as one good approach for overcoming the memory wall. However, the IO-channel power of a TSV-based 3D DRAM represents a significant portion of the 3D DRAM power. In this paper, we propose a built-in self-test (BIST) -assisted tuning scheme to adjust the driving capability of programmable drivers to fit the number of stacked 3D DRAM dies such that the IO-channel power can be minimized. A BIST design supporting specific test patterns and test flow for the driver tuning is proposed as well. Simulation results show that about 6.16×10 -- 2 J energy saving can be achieved for a logic-DRAM stack with 150fF/die TSV load under 100s write operations if the proposed BIST-assisted tuning scheme is implemented in the logic die.
Yun-Chao You, Chi-Chun Yang, Jin-Fu Li 0001, Chih-Yen Lo, Chao-Hsun Chen, Jenn-Shiang Lai, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu
ATS7
2014 Improving power delivery network design by practical methodologies
abstract
There are many works on the power network design and prototyping for digital designs, however some usual and practical design concerns are not addressed. In this work, we present a realistic power network design methodology without IR violation certified by state-of-the-art commercial tool. Our work integrates analysis, optimization and synthesis of power network. In particular, we consider thermal effect and power pad's positions during the prototyping of power network. A scenario in placement regarding the violation of design rules is considered and resolved by maximum flow algorithm at the same stage. After the synthesis of initial power network, we generate a sensitivity matrix which is correlated with nodal voltage and resistances of net and via in metal layers. Furthermore, a Sequential Linear Programming(SLP) will be applied to adjust the sensitivity matrix iteratively until the IR drop constraint is satisfied. Our work is experimented on a real design in TSMC 65nm LP process, and the result validates our framework that the IR-Drop can be reduced to 2% of supply voltage.
Chia-Chi Huang, Chang-Tzu Lin, Wei-Syun Liao, Chieh-Jui Lee, Hung-Ming Chen, Chia-Hsin Lee, Ding-Ming Kwai
ICCD7
2014 DArT: A Component-Based DRAM Area, Power, and Timing Modeling Tool
abstract
DRAM renovation calls for a holistic architecture exploration to cope with bandwidth growth and latency reduction need. In this paper, we present DRAM area power timing (DArT), a DRAM area, power, and timing modeling tool, for array assembly and interface customization. Through proper design abstraction, our component-based modeling approach provides increased flexibility and higher accuracy, making DArT suitable for DRAM architecture exploration and performance estimation. We validate the accuracy of DArT with respect to the physical layout and circuit simulation of an industrial 68 nm commodity DRAM device as a reference. The experiment results show that the maximum deviations from the reference design, in terms of area, timing, and power, are 3.2%, 4.92%, and 1.73%, respectively. For an architectural projection by porting it to a 45 nm process, the maximum deviations are 3.4%, 3.42%, and 8.57%, respectively. The combination of modeling performance, flexibility, and accuracy of DArT allows us to easily explore new DRAM architectures in the future, including 3-D stacked DRAM.
Hsiu-Chuan Shih, Pei-Wen Luo, Jen-Chieh Yeh, Shu-Yen Lin, Ding-Ming Kwai, Shih-Lien Lu, Andre Schaefer, Cheng-Wen Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2014 Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization
abstract
3-D many-core processor (3-D MCP) has become an emerging technology to tackle the power wall problem due to rapidly increasing number of transistors. However, when maximizing the throughput of 3-D MCP, which is expressed as a weighted sum of the speeds, due to the inherent heat removal limitation, thermal issues must be taken into consideration. Since the temperature of a core strongly depends on its location in the 3-D IC, a proper task allocation can alleviate the thermal problem and improve the throughput. Nevertheless, conventional techniques require computationally intensive thermal simulation, which prohibits its usage from the online application. In this paper, we propose an efficient online task allocation and task migration algorithm attempting to maximize the throughput of 3-D MCP simultaneously, considering unfinished tasks left from the last scheduling interval and new incoming tasks of this scheduling interval. The results of our experiments show that our proposed method achieves a 20.82X runtime speedup. These results are comparable to the exhaustive solutions obtained from optimization-modeling software LINGO. In addition, on average, our throughput results, with and without consideration of unfinished tasks, are only 4.39% and 0.69% worse, respectively, than that of the exhaustive method. In 128 task-to-core allocations, our method takes only 0.951 ms, which is 59.39 times faster than that of the previous work.
Cody Hao Yu, Chiao-Ling Lung, Yi-Lun Ho, Ruei-Siang Hsu, Ding-Ming Kwai, Shih-Chieh Chang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2014 Application-Independent Testing of 3-D Field Programmable Gate Array Interconnect Faults
abstract
3-D integration has been touted as an approach to reducing the lengths of critical paths in field-programmable gate arrays (FPGAs). A 3-D chip stacks a number of 2-D FPGA bare dies, interconnected by through-silicon vias (TSVs) and micro bumps, to attain a high packing density. However, the technology also introduces new types of defects, such as TSV void and microbump misalignment. Testing the interconnection faults becomes inevitable. In this paper, we present an automatic test pattern generator for open, short, and delay faults on 3-D FPGA interconnects by exploiting the regularity of switch matrix topology and forming repetitive paths with finite steps and with loop-back. The experimental results show that 12 test patterns (TPs) suffice to achieve 100% open fault coverage (FC). To detect all possible neighboring short faults, we need more than 40 TPs, whose number increases only slightly with the height of the 3-D FPGA. The TPs have high delay FC (96%) for 3-D FPGAs with the number of configurable logic blocks ranging from 50 × 50 × 2 to 50 × 50 × 6, demonstrating the scalability of our method.
Yen-Lin Peng, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu
IEEE Trans. Very Large Scale Integr. Syst.2
2013 I-LUTSim: An iterative look-up table based thermal simulator for 3-D ICs
abstract
This work presents an iterative look-up table based thermal simulator, I-LUTSim, to efficiently estimate the temperature profile of three-dimensional integrated circuits. I-LUTSim includes two stages. First, the pre-process stage constructs thermal impulse response tables. Then, the simulation stage iteratively calculates the temperature profile via the table lookup. With this two-stage scheme, the maximum absolute error of I-LUTSim is less than 0.41% compared with that of a commercial tool ANSYS. Moreover, I-LUTSim is at least an order of magnitude faster than a fast matrix solver SuperLU [1] for the full-chip temperature simulation.
Chi-Wen Pan, Yu-Min Lee, Pei-Yu Huang, Chi-Ping Yang, Chang-Tzu Lin, Chia-Hsin Lee, Yung-Fa Chou, Ding-Ming Kwai
ASP-DAC8
2013 Benchmarking for research in power delivery networks of three-dimensional integrated circuits
abstract
Power integrity is generally considered to be one of the major bottlenecks hindering the prevalence of three-dimensional integrated circuits (3D ICs). The higher integration density and smaller footprint result in significantly increased power density, which threatens the system reliability. In view of this, there has been groundswell of interest in academia to model, design or optimize the power delivery networks (PDNs) in 3D ICs. Unfortunately, while several PDN benchmarks exist for 2D PDNs, none is available in the context of 3D. As a consequence, most existing literature resorts to ad-hoc designs by artificially stacking 2D PDNs for experiments, rendering the results less convincing. In this paper, we put forward a set of ten PDN benchmarks that are extracted from industrial 3D designs. These designs are carefully selected such that they cover a wide range of functionality, size, TSV number, tier number and packaging style. We hope that the released benchmarks can facilitate and promote research in 3D PDNs.
Pei-Wen Luo, Chun Zhang 0003, Yung-Tai Chang, Liang-Chia Cheng, Hung-Hsie Lee, Bih-Lan Sheu, Yu-Shih Su, Ding-Ming Kwai, Yiyu Shi 0001
ISPD8
2013 Special session 4C: Hot topic 3D-IC design and test
abstract
Three-dimensional (3D) integration using through silicon via (TSV) is a promising approach to coping with the challenges faced by the current 2D technology. A TSV-based 3D IC is implemented by stacking multiple dies which are vertically connected by TSVs. This may shorten the global interconnects of a 3D IC and greatly improve its performance and power consumption. High bandwidth is achieved by the increase of IO channels provided by the TSVs, which also reduce the unnecessary waste of energy during data movement. In addition, the 3D integration technology shows other advantages over 2D technology, such as high functionality, heterogeneous integration, small form factor, etc. However, there are still challenges that need to be tackled before volume production of 3D ICs using TSV becomes possible, including technology scalability, quality and reliability, yield, thermal management, equipment and infrastructure, and costs. To demonstrate the feasibility of 3D-IC technologies, many academic and industrial institutes have been working on various test vehicles, especially in recent years. An increasing attention also has been attracted around the world in the semiconductor industry by the development of related technologies. In this special session, we will discuss the evolutionary efforts toward the realization of 3-D ICs. As memory dies need to be integrated in most 3D-IC system, we will also address the challenges in the design and test of 3D memories against cross-layer process, voltage and temperature variations while suppressing thermal effect and power consumption, etc. We will share our experiences and show results from some of the test vehicles we have worked on, including processor/memory stacks, analog/logic stacks, logic/logic stacks, etc. We will show a 1,024-bit wide bus chip-to-chip interconnection using 40×40 fine-pitch TSVs to demonstrate ultra-low-power operation. For memories, we will present a 3D-RAM structure using small voltage-swing TSVs, vertical-device-stacking nonvolatile-SRAM and ReRAM, and a 3D vertical-gate NAND flash. To address the challenge of reliability and yield, we will also discuss important test techniques. Demonstration of benefits provided by 3D integration technology will also be shown. Last but not least, we will describe our development plan regarding various types of die stacking using heterogeneous process integration, especially processor/memory stacking that is widely believed to be a key technology in future generations of smart handheld devices.
Jin-Fu Li 0001, Cheng-Wen Wu, Masahiro Aoyagi, Meng-Fan Chang, Ding-Ming Kwai
VTS5
2013 A hybrid ECC and redundancy technique for reducing refresh power of DRAMs
abstract
Dynamic random access memory (DRAM) is one key component in handheld devices. It typically consumes significant portion of the energy of the device even if the device is in standby mode due to the refresh requirement. This paper proposes a hybrid error-correcting code (ECC) and redundancy (HEAR) technique to reduce the refresh power of DRAMs in standby mode. The HEAR circuit consists of a Bose-Chaudhuri-Hocquenghem (BCH) module and an error-bit repair (EBR) module to raise the error correction capability and minimize the adverse effects caused by the ECC technique such that the refresh period can be effectively prolonged and considerable refresh power reduction can be achieved. Analysis results show that the proposed HEAR scheme can achieve 40~70% of energy saving for a 2Gb DDR3 DRAM in standby mode. The area cost of parity data and ECC circuit of HEAR scheme is only about 63 % and 53 % of that of the ECC-only, respectively.
Yun-Chao You, Chih-Sheng Hou, Li-Jung Chang, Jin-Fu Li 0001, Chih-Yen Lo, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu
VTS6
2013 Parametric Delay Test of Post-Bond Through-Silicon Vias in 3-D ICs via Variable Output Thresholding Analysis
abstract
A parametric delay fault could arise in a through-silicon via (TSV) of a 3-D IC due to a manufacturing defect. Identification of such a fault is essential for fault diagnosis, yield-learning, and/or reliability screening. In this paper, we present an innovative design-for-testability technique called variable output thresholding. We discovered that by dynamically switching the output of a TSV from a normal inverter to a Schmitt-Trigger inverter, the parametric delay fault on the TSV can be characterized and detected. SPICE simulation reveals that this technique remains effective even when there is significant process variation. A scalable test infrastructure indicates that the test time is modest at only 17.2 ms for 1024 TSVs and 648.8 ms for 32768 TSVs when the test clock is running at 10 MHz.
Yu-Hsiang Lin, Shi-Yu Huang, Kun-Han Tsai, Wu-Tung Cheng, Stephen K. Sunter, Yung-Fa Chou, Ding-Ming Kwai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2013 Low-Cost Error Tolerance Scheme for 3-D CMOS Imagers
abstract
This paper presents an error tolerance scheme for 3-D CMOS imagers that are constructed by stacking a pixel array of imager sensors, an analog-to-digital converter (ADC) array, and an image signal processor (ISP) array using microbumps$(\mu{\rm bumps})$and through silicon vias (TSVs). To deliver high-quality images in the presence of single or multiple$\mu{\rm bump}$, ADC, or TSV failures, we propose to interleave the connections from pixels to ADCs and recover the corrupted data in the ISPs. Key design parameters, such as the interleaving stride and the grouping ratio are determined by analyzing the employed error correction algorithm. Architectural simulation results demonstrate that the error tolerance scheme enhances the effective yield of an exemplar 3-D imager from 44% to 97%.
Hsiu-Ming Chang 0001, Jiun-Lang Huang, Ding-Ming Kwai, Kwang-Ting Cheng, Cheng-Wen Wu
IEEE Trans. Very Large Scale Integr. Syst.3
2013 In-Situ Method for TSV Delay Testing and Characterization Using Input Sensitivity Analysis
abstract
In this paper, we propose a method and the required architecture for characterizing the propagation delays of the through Silicon vias (TSVs) in a 3-D IC. First of all, every two TSVs are paired up to form an oscillation ring with some peripheral circuits. Their joint performance can thus be measured roughly by the oscillation period of the ring. Next, we utilize a technique called sensitivity analysis to further derive the propagation delay of each individual TSV participating in an oscillation ring-a distilling process. In this process, we perturb the strength of the two TSV drivers, and then measure their effects in terms of the change of the oscillation ring's period. By some following analysis, the propagation delay of each TSV can be revealed. On top of scheme, we also present an architecture that can activate the performance characterization process of each test unit - that consists of two TSVs - one at a time in a proper sequence. The area overhead is only 18.97 equivalent two-input NAND gate per TSV, by which one can gain the ability to profile the capacitances and the propagation delays of the TSVs on a 3-D IC.
Jhih-Wei You, Shi-Yu Huang, Yu-Hsiang Lin, Meng-Hsiu Tsai, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu
IEEE Trans. Very Large Scale Integr. Syst.5
2012 Capturing the phantom of the power grid - on the runtime adaptive techniques for noise reduction
abstract
Power supply noise has become one of the primary concerns in low power designs. To ensure power integrity, designers need to make sure that voltage droop and bounce do not exceed noise margin in all possible scenarios. Since it is very difficult to capture the exact worst corner among the mist of complex functionalities in modern VLSI designs, statistical design methodologies have been adapted, which may bring significant design overhead. In view of this, various runtime techniques have been proposed in literature to suppress power grid noise adaptively. This paper first presents various challenges in power grid designs from an industrial perspective, explains the difficulties in handling them at deign time, and then reviews various runtime techniques to adaptively suppress power supply noise, including sensor-based power gating, re-routable decaps, proactive clock frequency actuator, and PLL based clocking.
Pei-Wen Luo, Yu-Shih Su, Liang-Chia Cheng, Ding-Ming Kwai, Yiyu Shi 0001
ASP-DAC5
2012 Small delay testing for TSVs in 3-D ICs
abstract
In this work, we present a robust small delay test scheme for through-silicon vias (TSVs) in a 3D IC. By changing the output inverter's threshold of a TSV in a testable oscillation ring structure, we can approximate the propagation delay across that TSV, and thereby detecting a small delay fault. SPICE simulation reveals that this Variable Output Thresholding (VOT) technique is still effective even when there is significant process variation in detecting a slow TSV with some resistive open defect that may escape the traditional at-speed test.
Shi-Yu Huang, Yu-Hsiang Lin, Kun-Han Tsai, Wu-Tung Cheng, Stephen K. Sunter, Yung-Fa Chou, Ding-Ming Kwai
DAC7
2012 A built-in self-test scheme for 3D RAMs
abstract
Three-dimensional (3D) random access memory (RAM) using through-silicon vias for inter-die interconnects has been considered as a new approach to overcome the memory wall. In this paper, we propose a built-in self-test (BIST) scheme for 3D RAMs. In the BIST scheme, a clock-domain-crossing-aware test pattern generator is proposed to cope with the clock-domain-crossing issue. An inter-die synchronization mechanism is also proposed to synchronize the BIST circuits in different dies. Furthermore, the BIST circuit provides the high-programmability feature to support the selection of RAMs in a die for testing such that it can support thermal management during the test. We design the proposed BIST scheme in a 3D IC with processor and RAM dies. Experimental results show that the area cost of the BIST circuit is very small. The area overhead of the BIST circuit for four 8192×64-bit RAMs in a die is only 0.45% using TSMC 90nm 1P9M CMOS process technology.
Yun-Chao You, Che-Wei Chou, Jin-Fu Li 0001, Chih-Yen Lo, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu
ITC5
2012 A SAR ADC missing-decision level detection and removal technique
abstract
Capacitor mismatch is the linearity limiter of charge redistribution SAR ADCs. This paper aims at detecting and removing the mismatch induced missing-decision levels (MDLs), i.e., large positive DNLs; these errors lead to information loss that cannot be recovered by external calibration. A switched-capacitor based approach is proposed to avoid DC currents and reduce design overhead; the hardware modification also supports comparator offset compensation to improve calibration quality. Simulation results show that the proposed technique effectively improves the SAR ADC linearity in the presence of capacitor mismatch and comparator offset.
Jiun-Lang Huang, X.-L. Huang, Yung-Fa Chou, Ding-Ming Kwai
VTS4
2012 An MCT-Based Bit-Weight Extraction Technique for Embedded SAR ADC Testing and Calibration
Xuan-Lun Huang, Jiun-Lang Huang, Chang-Yu Chen, Tseng Kuo-Tsai, Ming-Feng Huang, Yung-Fa Chou, Ding-Ming Kwai
J. Electron. Test.8
2011 A self-testing and calibration method for embedded successive approximation register ADC
abstract
This paper presents a self-testing and calibration method for the embedded successive approximation register (SAR) analog-to-digital converter (ADC). We first propose a low cost design-for-test (DfT) technique which tests a SAR ADC by characterizing its digital-to-analog converter (DAC) capacitor array. Utilizing DAC major carrier transition testing, the required analog measurement range is just 4 LSBs; this significantly lowers the test circuitry complexity. Then, we develop a fully-digital missing code calibration technique that utilizes the proposed testing scheme to collect the required calibration information. Simulation results are presented to validate the proposed technique.
Xuan-Lun Huang, Ping-Ying Kang, Hsiu-Ming Chang 0001, Jiun-Lang Huang, Yung-Fa Chou, Yung-Pin Lee, Ding-Ming Kwai, Cheng-Wen Wu
ASP-DAC7
2011 Thermal-aware on-line task allocation for 3D multi-core processor throughput optimization
abstract
Three-dimensional integrated circuit (3D IC) has become an emerging technology in view of its advantages in packing density and flexibility in heterogeneous integration. The multi-core processor (MCP), which is able to deliver equivalent performance with less power consumption, is a candidate for 3D implementation. However, when maximizing the throughput of 3D MCP, due to the inherent heat removal limitation, thermal issues must be taken into consideration. Furthermore, since the temperature of a core strongly depends on its location in the 3D MCP, a proper task allocation helps to alleviate any potential thermal problem and improve the throughput. In this paper, we present a thermal-aware on-line task allocation algorithm for 3D MCPs. The results of our experiments show that our proposed method achieves 16.32X runtime speedup, and 23.18% throughput improvement. These are comparable to the exhaustive solutions obtained from optimization modeling software LINGO. On average, our throughput is only 0.85% worse than that of the exhaustive method. In 128 task-to-core allocations, our method takes only 0.932 ms, which is 57.74 times faster than the previous work.
Chiao-Ling Lung, Yi-Lun Ho, Ding-Ming Kwai, Shih-Chieh Chang 0001
DATE3
2011 A Pre- and Post-bond Self-Testing and Calibration Methodology for SAR ADC Array in 3-D CMOS Imager
abstract
This paper presents a low-cost pre- and post-bond self-testing and calibration methodology for the successive approximation register (SAR) analog-to-digital converter (ADC) array in a three-dimensional (3-D) CMOS imager. The basic idea is to test and calibrate the SAR ADC by measuring the major carrier transitions (MCTs) of the internal digital-to-analog converter (DAC) capacitor array. During the pre-bond stage, when access to the die is very limited, we propose a calibration-oriented testing technique that only determines whether the ADC array can achieve the desired performance after calibration. This substantially reduces the required design-for-test (DfT) circuitry complexity and test time. Then, during the post-bond stage, more thorough characterization on the ADC array is performed, we utilize digital resources from the image signal processor (ISP) die to analyze the measurement results, compute the calibration parameters, and perform the digital calibration. Simulation results are presented to validate the proposed techniques.
Xuan-Lun Huang, Ping-Ying Kang, Jiun-Lang Huang, Yung-Fa Chou, Yung-Pin Lee, Ding-Ming Kwai
ETS6
2011 A built-in self-test scheme for the post-bond test of TSVs in 3D ICs
abstract
Three-dimensional (3D) integration using through silicon via (TSV) has been widely acknowledged as one future integrated-circuit (IC) technology. A 3D IC including multiple dies connected with TSVs offers many benefits over current 2D ICs. However, the testing of 3D ICs is much more difficult than that of 2D ICs. In this paper, we propose a cost-effective built-in self-test circuit (BIST) to test TSVs of a 3D IC. The BIST scheme, arranging the TSVs into arrays similar to memory, has the features of low test/diagnosis time and low silicon area cost. Simulation results show that the area overhead of the BIST circuit implemented with 0.18μm CMOS technology for a 16×32 TSV array in which each TSV cell size is 45μm2is 2.24%. Also, the BIST needs only 130 clock cycles to test the TSV array with stuck-at faults. In comparison with the IEEE 1500-based test approach, the BIST scheme can achieve 85.2% area cost and 93.6% test time reduction.
Yu-Jen Huang, Jin-Fu Li 0001, Ji-Jan Chen, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu
VTS4
2011 Yield Enhancement by Bad-Die Recycling and Stacking With Though-Silicon Vias
abstract
3-D integration provides a means to overcome the difficulties in design and manufacturing of system-on-chip (SOC) and memory products. Introducing a short vertical interconnect, called through-silicon via (TSV), makes it feasible to repair and recycle bad dies by stacking. We propose a method to accomplish this using a dual-TSV hardwired switch (DTHS) in which the via-hole location is programmable. With the DTHS, we activate a spare and establish inter-die routing. The spare is nothing but a good part in another bad die. To be 3-D reparable, the design is partitioned into disjoint parts. The effort for the modification is minor in view of that a typical SOC is readily composed of modules with predefined functions and supply voltages. The DTHS is used: 1) to shut off power connections of both failed and unused parts; 2) to disconnect their signal paths; and 3) to redirect them to the selected good parts in the stacked dies. Despite the speed is degraded due to the extra load incurred by the DTHS, our simulation shows that the increase in delay time can be limited below 100 ps with an over-designed buffer which occupies 0.8% of the area of a 30 μm TSV, using a 65-nm CMOS process. The performance degradation turns out to be a necessary evil, since the increased height of the die stack leads to a thermal conductivity poorer than its 2-D counterpart. The 3-D patch die helps to shorten time-to-market and turn the irreparable dies profitable.
Yung-Fa Chou, Ding-Ming Kwai, Cheng-Wen Wu
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Homogeneous integration for 3D IC with TSV
Ding-Ming Kwai
ASP-DAC1
2010 CAD reference flow for 3D via-last integrated circuits
abstract
Next-decade computing power and interconnect bottle-neck challenge conventional IC design due to the ever increasing demands for high frequency and great bandwidth. Three-dimensional large-scale integration (3D-LSI) provides an opportunity to realize such high performance cores while reducing long latency. In this paper, we present a reference flow for the implementation of 3D via-last ICs in scalable face-to-back bonding style which leverages a mature set of 2D IC physical design tools. The first enabling technology of 3D-LSI is through-silicon via (TSV). Two kinds of TSV diameters are exemplified in the flow, namely, 5¿m and 50¿m. We propose an easy-to-adopt method to address the TSV-aware mixed-sized placement by considering the obstructions generated from adjacent-tier's floorplan, subject to certain TSV alignment constraints. Furthermore, the technique of clock tree synthesis (CTS) for a homogeneous die stack is developed to dramatically reduce the clock latency and skew. The mixed-sized placement and CTS of each tier can be done without iteration. To the best of our knowledge, no work has ever been published in literature discussing CTS for 3D via-last integration in a face-to-back fashion. Finally, to complete the proposed flow 2D timing-driven routing and modified off-line design rule check (DRC) and layout versus schematic (LVS) verification are performed very well.
Chang-Tzu Lin, Ding-Ming Kwai, Yung-Fa Chou, Ting-Sheng Chen, Wen Ching Wu
ASP-DAC2
2010 A Test Integration Methodology for 3D Integrated Circuits
abstract
The three-dimensional (3D) integration technology using through silicon via (TSV) provides many benefits over the 2D integration technology. Although many different manufacturing technologies for 3D integrated circuits (ICs) have been presented, some challenges should be overcome before the volume production of 3D ICs. One of the challenges is the testing of 3D ICs. This paper proposes test integration interfaces for controlling the design-for-test circuits in the dies of a 3D IC. The test integration interfaces can support the pre-bond, known-good stack, and post-bond tests. The minimum number of required test pads of the proposed test interface for pre-bond test using is only four. Furthermore, the test interface is compatible with the IEEE 1149.1 standard for the board-level testing. Simulation results show that the area overhead of the proposed test interfaces for a 3D IC with two dies in which each die implements the function of ITC'99 b19 benchmark is only about 0.15%.
Che-Wei Chou, Jin-Fu Li 0001, Ji-Jan Chen, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu
Asian Test Symposium4
2010 Performance Characterization of TSV in 3D IC via Sensitivity Analysis
abstract
In this paper, we propose a method that can characterize the propagation delays across the Through Silicon Vias (TSVs) in a 3D IC. We adopt the concept of the oscillation test, in which two TSVs are connected with some peripheral circuit to form an oscillation ring. Upon this foundation, we propose a technique called sensitivity analysis to further derive the propagation delay of each individual TSV participating in the oscillation ring-a distilling process. In this process, we perturb the strength of the two TSV drivers, and then measure their effects in terms of the change of the oscillation ring's period. By some following analysis, the propagation delay of each TSV can be revealed. Monte-Carlo analysis of a typical TSV with 30% process variation on transistors shows that the characterization error of this method is only 2.1% with the standard deviation of 8.1%.
Jhih-Wei You, Shi-Yu Huang, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu
Asian Test Symposium3
2010 An error tolerance scheme for 3D CMOS imagers
abstract
A three-dimensional (3D) CMOS imager constructed by stacking a pixel array of backside illuminated sensors, an analog-to-digital converter (ADC) array, and an image signal processor (ISP) array using micro-bumps (μbumps) and through-silicon vias (TSVs) is promising for high throughput applications. However, due to the direct mapping from pixels to ISPs, the overall yield relies heavily on the correctness of the μbumps, ADCs and TSVs -- a single defect leads to the information loss of a tile of pixels. This paper presents an error tolerance scheme for the 3D CMOS imager that can still deliver high quality images in the presence of μbump, ADC, and/or TSV failures. The error tolerance is achieved by properly interleaving the connections from pixels to ADCs so that the corrupted data, if any, can be recovered in the ISPs. A key design parameter, the interleaving stride, is decided by analyzing the employed error correction algorithm. Architectural simulation results demonstrate that the error tolerance scheme enhances the effective yield of an exemplar 3D imager from 46% to 99%.
Hsiu-Ming Chang 0001, Jiun-Lang Huang, Ding-Ming Kwai, Kwang-Ting Cheng, Cheng-Wen Wu
DAC3
2010 On-chip testing of blind and open-sleeve TSVs for 3D IC before bonding
abstract
Pre-bond test is preferred for a three-dimensional integrated circuit (3D IC), since it reduces stacking yield loss and thus saves cost. In this paper, we present two schemes for testing through-silicon vias (TSVs) by performing on-chip screening before wafer thinning and bonding. The first scheme is for blind TSVs, which have one end floating, using a charge-sharing technique commonly seen in DRAM. The second scheme is for open-sleeve TSVs, which have one end shorted to the substrate, using a voltage-dividing technique commonly seen in ROM. By virtue of the inherent capacitive and resistive characteristics, we detect the TSVs out of a specified range as anomalies, taking into account the effects of process variations in the detection circuitry. The statistical design by Monte Carlo simulation using TSMC 65nm low-power process shows that for blind TSVs, the best overkill ratio is below 6%. For open-sleeve TSVs, inherent limitations restrict the applicability, so more work needs to be done in the future. Our implementation enjoys little area overhead, requiring only a simple sense amplifier and a write buffer that are shared among a number of TSVs. Reducing the number of TSVs that share a test module will reduce the test time, but increase the area overhead. For blind TSVs, the parallelism also affects the overkill and escape rates.
Po-Yuan Chen, Cheng-Wen Wu, Ding-Ming Kwai
VTS3
2009 On-Chip TSV Testing for 3D IC before Bonding Using Sense Amplification
abstract
We present a novel testing scheme for TSVs in a 3D IC by performing on-chip TSV monitoring before bonding, using a sense amplification technique that is commonly seen on a DRAM. By virtue of the inherent capacitive characteristics, we can detect the faulty TSVs with little area overhead for the circuit under test.
Po-Yuan Chen, Cheng-Wen Wu, Ding-Ming Kwai
Asian Test Symposium3
2004 Incomplete k-ary n-cube and its derivatives
Behrooz Parhami, Ding-Ming Kwai
J. Parallel Distributed Comput.2
2001 A Unified Formulation of Honeycomb and Diamond Networks
abstract
AbstractÐHoneycomb and diamond networks have been proposed as alternatives to mesh and torus architectures for parallel processing. When wraparound links are included in honeycomb and diamond networks, the resulting structures can be viewed as having been derived via a systematic pruning scheme applied to the links of 2D and 3D tori, respectively. The removal of links, which is performed along a diagonal pruning direction, preserves the network's node-symmetry and diameter, while reducing its implementation complexity and VLSI layout area. In this paper, we prove that honeycomb and diamond networks are special subgraphs of complete 2D and 3D tori, respectively, and show this viewpoint to hold important implications for their physical layouts and routing schemes. Because pruning reduces the node degree without increasing the network diameter, the pruned networks have an advantage when the degree-diameter product is used as a figure of merit. Additionally, if the reduced node degree is used as an opportunity to increase the link bandwidths to equalize the costs of pruned and unpruned networks, a gain in communication performance may result. Index TermsÐCayley graph, k-ary n-cube, network topology, processor array, pruned torus network, VLSI layout. 1
Behrooz Parhami, Ding-Ming Kwai
IEEE Trans. Parallel Distributed Syst.2
2000 etection of SRAM cell stability by lowering array supply voltage
abstract
In this paper, we discuss a design-for-test technique for the detection of cell stability in static random access memory (SRAM). The power supply to the memory array is isolated and independently accessible from an external terminal. By lowering the array supply voltage, the cell stability is degraded, making the defective cells susceptible to noises induced by read/write operations. On-silicon characterization result using 0.18 /spl mu/m CMOS technology is reported. It shows that the weak tailing bits in the statistical distribution can manifest themselves. The implementation of the test mode is inherently low-cost and can be combined with previously proposed methods for an improved detection capability.
Ding-Ming Kwai, Hung-Wen Chang, Hung-Jen Liao, Ching-Hua Chiao, Yung-Fa Chou
Asian Test Symposium1
1999 Data-Driven Control Scheme for Linear Arrays: Application to a Stable Insertion Sorter
abstract
We present a strategy for designing stable insertion sorters based on linear arrays with data-driven control. The novelty of our approach lies in each data item carrying a control tag to specify how it is to be operated upon by a receiving cell and in performing two parallel comparisons within each cell. To assure first-in/first-out handling of equal key values, some data items must be marked to reflect their past histories. Such marking is conveniently carried out by modifying the data item's control tag. It is the combination of the above features that allows us to derive the first single-cycle priority queue that operates in fully pipelined mode, with no broadcasting of data values or control signals. By performing more than two parallel comparisons in each cell, the VLSI implementation cost of our stable sorter can be reduced. We show that highly cost-effective designs can be obtained by selecting an optimal cell size in terms of the number of comparators it contains.
Behrooz Parhami, Ding-Ming Kwai
IEEE Trans. Parallel Distributed Syst.2
1999 Periodically Regular Chordal Rings
abstract
Chordal rings have been proposed in the past as networks that combine the simple routing framework of rings with the lower diameter, wider bisection, and higher resilience of other architectures. Virtually all proposed chordal ring networks are node-symmetric, i.e., all nodes have the same in/out degree and interconnection pattern. Unfortunately, such regular chordal rings are not scalable. In this paper, periodically regular chordal (PRC) ring networks are proposed as a compromise for combining low node degree with small diameter. By varying the PRC ring parameters, one can obtain architectures with significantly different characteristics (e.g., from linear to logarithmic diameter), while maintaining an elegant framework for computation and communication. In particular, a very simple and efficient routing algorithm works for the entire spectrum of PRC rings thus obtained. This flexibility has important implications for key system attributes such as architectural satiability, software portability, and fault tolerance. Our discussion is centered on unidirectional PRC rings with in/out-degree of 2. We explore the basic structure, topological properties, optimization of parameters, VLSI layout, and scalability of such networks, develop packet and wormhole routing algorithms for them, and briefly compare them to competing fixed-degree architectures such as symmetric chordal rings, meshes, tori, and cube-connected cycles.
Behrooz Parhami, Ding-Ming Kwai
IEEE Trans. Parallel Distributed Syst.2
1999 Correction to 'Periodically Regular Chordal Rings'
Behrooz Parhami, Ding-Ming Kwai
IEEE Trans. Parallel Distributed Syst.2
1998 Tight Bounds on the Diameter of Gaussian Cubes
abstract
Gaussian cubes are derived by removing links from a hypercube in a periodic fashion. By varying the partition parameter, one can obtain networks with different characteristics, while maintaining a basic framework for computation and communication. Unfortunately, such networks are in general not regular, making it difficult to derive their topological properties explicitly. In this paper, we study the diameter of Gaussian cubes and show the trade-off between cost and performance.
Ding-Ming Kwai, Behrooz Parhami
Comput. J.1
1998 Pruned Three-Dimensional Toroidal Networks
Ding-Ming Kwai, Behrooz Parhami
Inf. Process. Lett.1
1997 A Class of Fixed-Degree Cayley-Graph Interconnection Networks Derived by Pruning k-ary n-cubes
abstract
We introduce a pruning scheme to reduce the node degree of k-ary n-cube from 2n to 4. The links corresponding to n-2 of the n dimensions are removed from each node. One of the remaining dimensions is common to all nodes and the other is selected periodically from the remaining n-1 dimensions. Despite the removal of a large number of links from the k-ary n-cube, this incomplete version still preserves many of its desirable topological properties. In this paper, we show that this incomplete k-ary n-cube belongs to the class of Cayley graphs, and hence, is node-symmetric. It is 4-connected with diameter close to that of the k-ary n-cube.
Ding-Ming Kwai, Behrooz Parhami
ICPP1
1992 Data Flow Representation of Iterative Algorithms for Systolic Arrays
abstract
An algebraic representation is proposed for regular iterative algorithms that can be described as bundles of data flows with different wavefronts. A form corresponding to a geometric representation such as a dependence graph is obtained by modeling data flows with a generating function of a power series. The main attributes of systolic algorithms and arrays are revealed by a unique dataflow representation. This provides the ability to pipeline two or more systolic arrays solving different subproblems without intermediate buffering. An example is given to show a case in which the technique can be used.>
Chein-Wei Jen, Ding-Ming Kwai
IEEE Trans. Computers2
1989 Multi-dimensional parallel computing structures for regular iterative algorithms
Chein-Wei Jen, Ding-Ming Kwai
Integr.2