Yen-Tai Lai

dblp:68/4168 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Electronic design automation · 99% Integrated circuit design · 1%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
physical design
0.142001
Graph-theory-based simplex algorithm for VLSI layout spacingproblems with multiple variable constraints · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001
Performance-directed compaction for VLSI symbolic layouts · Comput. Aided Des. 1995
Algorithms for floorplan design via rectangular dualization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1988
Electronic design automation › physical design
layout compaction
0.022001
Graph-theory-based simplex algorithm for VLSI layout spacingproblems with multiple variable constraints · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001
Performance-directed compaction for VLSI symbolic layouts · Comput. Aided Des. 1995
Mathematical optimization
linear programming
0.012001
Graph-theory-based simplex algorithm for VLSI layout spacingproblems with multiple variable constraints · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001
Mathematical optimization › linear programming
simplex method
0.012001
Graph-theory-based simplex algorithm for VLSI layout spacingproblems with multiple variable constraints · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001
Electronic design automation › physical design
floorplanning
0.021988
Algorithms for floorplan design via rectangular dualization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1988
An algorithm for building rectangular floor-plans · DAC 1984
Electronic design automation › physical design › VLSI layout
symbolic layout
0.011995
Performance-directed compaction for VLSI symbolic layouts · Comput. Aided Des. 1995
Integrated circuit design
digital circuit design
0.011984
An algorithm for building rectangular floor-plans · DAC 1984
Electronic design automation › physical design
layout design
0.011984
An algorithm for building rectangular floor-plans · DAC 1984

Methods — techniques the papers use, named apart from their topics

longest path algorithm · 0.1graph-theory-based simplex · 0.1graph matching · 0.0bipartite graph algorithms · 0.0rectangular floorplan algorithm · 0.0
YearPublicationVenuePosition
2016 Fast coding unit decision and mode selection for intra-frame coding in high-efficiency video coding
abstract
High‐efficiency video coding (HEVC) is a new video coding compression standard. As the successor to H.264/AVC, it provides better performance and supports higher resolution. However, the encoding complexity increases drastically. One of the major reasons is that the coding unit (CU) in HEVC is multi‐sized and adjustable rather than fixed as in H.264. In addition, the number of prediction modes used in intra‐frame coding is expanded from 9 to 35. The authors analysed the statistical correlations of CU depth to the deviation of pixels in the largest coding unit (LCU) and rate‐distortion cost (RDcost). Accordingly, a fast CU decision method is proposed, which contains two steps: first, the depth to begin searching is determined according to the deviation of the LCU and then splitting the current CU further is decided according to RDcost. For intra‐prediction, we also propose a fast mode selection method to reduce complexity. This method can quickly determine the modes for rate‐distortion optimisation when the combination of most probable modes reveals the pattern direction. Software simulations show that the proposed methods reduce encoding time by more than 50% with an average of 1.4% increase of BD‐rate compared to reference software HM12.0.
Chao-Feng Tseng, Yen-Tai Lai
IET Image Process.2
2013 Power minimization for dynamically reconfigurable FPGA partitioning
abstract
Dynamically reconfigurable FPGA (DRFPGA) implements a given circuit system by partitioning it into stages and then executing each stage sequentially. Traditionally, the number of communication buffers is minimized. In this article, we study the partitioning problem targeting at power minimization for the DRFPGAs that have lookup table (LUT) based logic blocks. We analyze the power consumption caused by the communication buffers in the temporal partitioning. Based on the analysis, we use a flow network to represent a given circuit so that the power consumption of buffers is correctly evaluated and the temporal constraints are satisfied in circuit partitioning. The well known flow-based FBB algorithm is then applied to the network to find the area-balanced partitioning of minimum power consumption. Experimental results show that our method outperforms the conventional partitioning algorithms in terms of power consumption. The problem is then extended to include constraints on the number of communication buffers. We provide a net modeling for this extended problem and present an extension of the FBB algorithm to obtain a power-optimal solution. Experimental results demonstrate the effectiveness of the proposed algorithm in reducing power consumption as compared to the previous partitioning algorithms without exceeding the buffer number limit.
Tzu-Chiang Tai, Yen-Tai Lai
ACM Trans. Embed. Comput. Syst.2
2011 A Performance-Oriented Algorithm with Consideration on Communication Cost for Dynamically Reconfigurable FPGA Partitioning
abstract
Dynamically reconfigurable FPGAs (DRFPGAs) have high logic utilization because of time-multiplexed interconnects and logic. In this article, we propose a performance-oriented algorithm for the DRFPGA partitioning problem. This algorithm partitions a given circuit system into stages such that the upper bound of the execution times of subcircuits is minimized. The communication cost is taken into consideration in the process of searching for the optimal solution. A graph is first constructed to represent the precedence constraints and calculate the number of buffers needed in a partitioning. This algorithm includes three phases. The first phase reduces the problem size by clustering the gates into subsystems that have only one output. Such a subsystem has a large number of intraconnections because the fan-outs of all vertices except for the one output are fed to the vertices inside the subsystem. This phase significantly reduces the computational complexity of partitioning. The second phase finds a partition with optimal performance. Finally, the third phase minimizes the communication cost by using an iterative improvement approach. Experimental results based on the Xilinx architecture show that our algorithm yields better partitioning solutions than traditional approaches.
Tzu-Chiang Tai, Yen-Tai Lai
ACM Trans. Reconfigurable Technol. Syst.2
2009 A Novel Digital Pixel Sensor System
abstract
A new imaging system which competes with the human eyes is proposed. This system includes an asynchronous self-resetting digital pixel sensor array, a readout circuit and a timing control circuit. The features of human eyes, which are wide dynamic range and logarithmic response, are realized by using an adaptive-sampling algorithm. Similar to multiple-sampling, several integration intervals which increase exponentially are supplied for pixels to execute the operation of sampling. Depends on the received illumination, the information of light intensity is transformed to digital codes at pixel level in an adaptive sampling time. An approximated logarithmic response is obtained by combining several piecewise linear responses. This is all controlled by a timing control circuit. Compared to the multiple-sampling scheme, this system offers an excellent image quality with less hardware requirement and much lower power consumption.
Yen-Tai Lai, Chia-Nan Yeh, Chi-Chou Kao
ISCAS1
2008 A novel flash analog-to-digital converter
abstract
In this paper a new ADC architecture of flash type is proposed. This proposed N-bit flash ADC replaces the (2N-1)-to-N encoder with two (2N/2-1)-to-(N/2) encoders to accomplish the encoding of the least significant bits and the most significant bits respectively. A 6-bit ADC of this architecture is implemented. The physical circuit is more compact than the existing ones. Power, processing time and area cost are all minimized. In addition, a new encoding algorithm is proposed to enhance the bubble error tolerance of an ADC. The encoders that have the capability of removing the bubble errors always suffer the problem of long latency. It becomes a bottleneck in the design of high speed flash ADC nowadays. In the proposed 6-bit ADC, the trade-off between bubble tolerance and latency is optimized by applying the proposed encoding algorithm into the encoder circuit design. Delay of 3 gate-levels or fewer is required for processing the encoding and the maximum error induced bubble is 7 LSB. Simulation results demonstrate the benefits introduced above. This new flash ADC offers an excellent choice for modern high speed ADC application.
Chia-Nan Yeh, Yen-Tai Lai
ISCAS2
2006 Low power readout control circuit for high resolution CMOS image sensor
abstract
In this paper, we present a low power, shift-register-based, readout control circuit for high resolution CMOS image sensors. In this circuit, a clock tree with the proposed autonomous clock gating technique is build for the clock supply system. The dynamic power is lowered by decreasing the switch activities and load capacitances. Additionally, from our analysis and simulations, the dynamic power dissipation is demonstrated to increase in proportion to the logarithm of the size of an imager. In high resolution applications, the size of imager is very large. By using this proposed circuit, the dynamic power dissipation of a large imager increases slightly. Besides, in order to reduce the static power dissipation, the leakage current is decreased by using the technique of stack effect.
Chia-Nan Yeh, Yen-Tai Lai
ISCAS2
2005 An efficient algorithm for finding the minimal-area FPGA technology mapping
abstract
Minimum area is one of the important objectives in technology mapping for lookup table-based field-progrmmable gate arrays (FPGAs). Although there is an algorithm that can find an optimal solution in polynomial time for the minimal-area FPGA technology mapping problem without gate duplication, its time complexity can grow exponentially with the number of inputs of the lookup-tables. This article proposes an algorithm with approximate to the area-optimal solution and lower time complexity. The time complexity of this algorithm is proven theoretically to be bounded by O ( n 3 ), where n is the total number of gates in the given circuit. It is shown that except for some cases the proposed algorithm can find an optimal solution of a given problem. We have combined the proposed algorithm with the existing postprocessing procedures which are used to find the gates that can be duplicated on a set of benchmark examples. The experimental results demonstrate the effectiveness of our algorithm.
Chi-Chou Kao, Yen-Tai Lai
ACM Trans. Design Autom. Electr. Syst.2
2004 Area-minimal algorithm for LUT-based FPGA technology mapping with duplication-free restriction
Chi-Chou Kao, Yen-Tai Lai
ASP-DAC2
2003 A technology mapping algorithm for heterogeneous FPGAs
abstract
In this paper, a technology mapping algorithm is proposed for heterogeneous FPGAs. The technology mapping problem is first formulated as a flow network problem. Then, an algorithm based on the min-cost max-flow algorithm is presented to select a proper set of feasible LUTs for various objectives. The objective, the total area composed of LUTs and routing area, are discussed in the paper. This algorithm has been tested on the MCNC benchmark circuits. Compared with other existing LUT-based FPGA mapping algorithms, the algorithm produces better characteristics.
Chi-Chou Kao, Yen-Tai Lai
ASP-DAC2
2001 Graph-theory-based simplex algorithm for VLSI layout spacingproblems with multiple variable constraints
abstract
An efficient algorithm is provided for solving a class of linear programming problems containing a large set of distance constraints of the form x/sub i/-x/sub j//spl ges/k and a small set of multivariable constraints of forms other than x/sub i/-x/sub j//spl ges/k. This class of linear programming formulation is applicable to very large scale integration (VLSI) layout spacing problems, including hierarchy-preserving hierarchical layout compaction, layout compaction with symmetric constraints, layout compaction with attractive and repulsive constraints, performance-driven layout compaction, etc. The longest path algorithm is efficient for solving spacing problems containing only distance constraints. However, it fails to solve problems that involve multiple-variable constraints. The linear programming formulation of a spacing problem requires use of the simplex method, which involves many matrix operations. This can be very time consuming when handling huge constraints systems derived from VLSI layouts. Herein it is found that most of the matrix operations can be replaced with fewer and faster graph operations, creating a more efficient graph-theory-based algorithm. Theoretical analysis shows that the proposed algorithm reduces the computation complexity of the simplex method.
Lih-Yang Wang, Yen-Tai Lai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1997 Hierarchical interconnection structures for field programmable gate arrays
abstract
Field programmable gate arrays (FPGA's) suffer from lower density and lower performance than conventional gate arrays. Hierarchical interconnection structures for field programmable gate arrays are proposed. They help overcome these problems. Logic blocks in a field programmable gate array are grouped into clusters. Clusters are then recursively grouped together. To obtain the optimal hierarchical structure with high performance and high density, various hierarchical structures with the same routability are discussed. The field programmable gate arrays with new architecture can be efficiently configured with existing computer aided design algorithms. The k-way min-cut algorithm is applicable to the placement step in the implementation. Global routing paths in a field programmable gate array can be obtained easily. The placement and global routing steps can be performed simultaneously. Experiments on benchmark circuits show that density and performance are significantly improved.
Yen-Tai Lai, Ping-Tsung Wang
IEEE Trans. Very Large Scale Integr. Syst.1
1995 Performance-directed compaction for VLSI symbolic layouts
Lih-Yang Wang, Yen-Tai Lai, Bin-Da Liu, Tin-Chung Chang
Comput. Aided Des.2
1994 A High Performance FPGA with Hierarchical Interconnection Structure
abstract
To overcome the low speed and low density problems in the FPGAs, we must reduce the number of switches used in the routing paths without sacrificing the routability for an FPGA. A hierarchical interconnection architecture for field programmable gate array (FPGA) is described. In every level of the hierarchy, logic blocks or cluster of logic blocks are connected together with switch blocks. Experiments on benchmark circuits are shown. It can be seen that significant improvement on performance is achieved.>
Ping-Tsung Wang, Kun-Nen Chen, Yen-Tai Lai
ISCAS3
1993 A graph-based simplex algorithm for minimizing the layout size and the delay on timing critical paths
abstract
The layout compaction problem with performance consideration is studied. In our approach, the delay upper-bound on the timing critical paths is reduced first, then the layout size is minimized without increasing that delay bound. Either step is formulated as a linear programming problem, and we propose a graph-based simplex algorithm, which replaces most of the matrix operations with graph operations, to solve the problem. Both theoretical analysis and experimental results show that this algorithm is quite efficient.
Lih-Yang Wang, Yen-Tai Lai, Bin-Da Liu, Ting-Chung Chang
ICCAD2
1993 Layout Compaction with Minimzed Delay Bound on Timing Critical Paths
Lih-Yang Wang, Yen-Tai Lai, Bin-Da Liu, Tin-Chung Chang
ISCAS2
1990 A Theory of Rectangular Dual Graphs
Yen-Tai Lai, Sany M. Leinwand
Algorithmica1
1988 Algorithms for floorplan design via rectangular dualization
abstract
A rectangular floorplan construction problem is approached from a graph-theoretical view. The study is based on a reduction of the rectangular dualization problem to a matching problem on bipartite graphs. This opens the way to applying traditional graph-theoretic methods and algorithms to floorplanning. Another result is a method for generating alternative rectangular duals, such that a proposed floorplan can be optimized by a sequence of iterative transformations. This approach is made more practical than others, by assuming that the given structure graph can be modified to force it to admit a rectangular dual. Algorithms that introduce edges and vertices into the given graph until a rectangular dual can be constructed are also presented.>
Yen-Tai Lai, Sany M. Leinwand
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1984 An algorithm for building rectangular floor-plans
Sany M. Leinwand, Yen-Tai Lai
DAC2