Sao-Jie Chen

dblp:74/4774 · DBLP profile ↗
← Back
50ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0003-1152-171XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 39 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Software engineering, systems software and programming languages · 2Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
16 papers
Electronic design automation · 60% Energy-efficient computing · 21% Hardware reliability and fault tolerance · 6%
Theoretical computer science
1 paper
Automated reasoning and model checking · 67% Mathematical optimization · 33%
Computer graphics and multimedia
2 papers
Image and video coding · 98% Geometric modeling and processing · 2%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
physical design
0.9112020
A Platform of Resynthesizing a Clock Architecture Into Power-and-Area Effective Clock Trees · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Design and Implementation of Block-Based Partitioning for Parallel Flip-Chip Power-Grid Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
An Automatic Optical Simulation-Based Lithography Hotspot Fix Flow for Post-Route Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Energy-efficient computing › dynamic power reduction
clock power reduction
0.412020
A Platform of Resynthesizing a Clock Architecture Into Power-and-Area Effective Clock Trees · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Electronic design automation › physical design › clock network synthesis
clock tree synthesis
0.412020
A Platform of Resynthesizing a Clock Architecture Into Power-and-Area Effective Clock Trees · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Energy-efficient computing
low-power design
0.412020
A Platform of Resynthesizing a Clock Architecture Into Power-and-Area Effective Clock Trees · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Electronic design automation › physical design
routing
0.372010
An Automatic Optical Simulation-Based Lithography Hotspot Fix Flow for Post-Route Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Crosstalk- and performance-driven multilevel full-chip routing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Multilevel full-chip routing for the X-based architecture · DAC 2005
Mathematical optimization › submodular optimization
coverage problem
0.212014
Accelerating Coverage Estimation Through Partial Model Checking · IEEE Trans. Computers 2014
Automated reasoning and model checking
model checking
0.212014
Accelerating Coverage Estimation Through Partial Model Checking · IEEE Trans. Computers 2014
Processor architecture and microarchitecture
instruction set architecture
0.212013
Cycle-Efficient LFSR Implementation on Word-Based Microarchitecture · IEEE Trans. Computers 2013
Image and video coding
rate-distortion optimization
0.112012
Interlayer Bit Allocation for Scalable Video Coding · IEEE Trans. Image Process. 2012
Image and video coding
scalable video coding
0.112012
Interlayer Bit Allocation for Scalable Video Coding · IEEE Trans. Image Process. 2012
Electronic design automation › physical design
power grid analysis
0.112012
Design and Implementation of Block-Based Partitioning for Parallel Flip-Chip Power-Grid Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Integrated circuit design › system-on-chip
system-on-chip design
0.112020
A Platform of Resynthesizing a Clock Architecture Into Power-and-Area Effective Clock Trees · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Hardware reliability and fault tolerance › network fault tolerance
fault-tolerant interconnection network
0.112011
A fault-tolerant NoC scheme using bidirectional channel · DAC 2011
Hardware reliability and fault tolerance › fault-tolerant architecture
fault-tolerant noc
0.112011
A fault-tolerant NoC scheme using bidirectional channel · DAC 2011
Interconnection networks and networks-on-chip › router architecture
network-on-chip router
0.112011
A Bidirectional NoC (BiNoC) Architecture With Dynamic Self-Reconfigurable Channel · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Electronic design automation
design for manufacturability
0.112010
An Automatic Optical Simulation-Based Lithography Hotspot Fix Flow for Post-Route Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Electronic design automation › physical design
optical proximity correction
0.112010
An Automatic Optical Simulation-Based Lithography Hotspot Fix Flow for Post-Route Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Electronic design automation › physical design
post-route optimization
0.112010
An Automatic Optical Simulation-Based Lithography Hotspot Fix Flow for Post-Route Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Electronic design automation › physical design › routing › detailed routing
rip-up and reroute
0.112010
An Automatic Optical Simulation-Based Lithography Hotspot Fix Flow for Post-Route Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Electronic design automation › hardware verification and test
hardware verification
0.112014
Accelerating Coverage Estimation Through Partial Model Checking · IEEE Trans. Computers 2014
Electronic design automation › physical design › routing › VLSI routing
crosstalk-aware routing
0.112005
Crosstalk- and performance-driven multilevel full-chip routing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › physical design › placement
post-placement optimization
0.012004
Fast postplacement optimization using functional symmetries · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Electronic design automation › logic synthesis › logic restructuring
rewiring
0.012004
Fast postplacement optimization using functional symmetries · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Interconnection networks and networks-on-chip › switching network
clos network
0.012001
A Three-Stage One-Sided Rearrangeable Polygonal Switching Network · IEEE Trans. Computers 2001
Interconnection networks and networks-on-chip › nonblocking networks
rearrangeable network
0.012001
A Three-Stage One-Sided Rearrangeable Polygonal Switching Network · IEEE Trans. Computers 2001
Interconnection networks and networks-on-chip
switching network
0.012001
A Three-Stage One-Sided Rearrangeable Polygonal Switching Network · IEEE Trans. Computers 2001
Electronic design automation › physical design › routing › multilayer routing
layer assignment
0.011998
NEWS: a net-even-wiring system for the routing on a multilayer PGA package · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1998
Electronic design automation › physical design › routing
timing-driven routing
0.012005
Crosstalk- and performance-driven multilevel full-chip routing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › physical design
gate sizing
0.012004
Fast postplacement optimization using functional symmetries · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Electronic design automation
logic synthesis
0.012004
Fast postplacement optimization using functional symmetries · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004

Methods — techniques the papers use, named apart from their topics

multiplexer merging · 0.4clock gating · 0.4partial model checking · 0.4mutation analysis · 0.4term-preserving look-ahead transformation · 0.3rate-distortion optimization · 0.1flip-chip locality exploitation · 0.1block-based partitioning · 0.1finite state machine · 0.1cycle-accurate simulation · 0.1bidirectional channel reconfiguration · 0.1search control policy · 0.0hierarchical subgoal organization · 0.0
YearPublicationVenuePosition
2023 A Robust Super-Regenerative Receiver with Optimal Detection on BER Level
abstract
This paper presents a robust super-regenerative receiver (SRR), which is insensitive to the temperature variations. A constant-alpha biasing circuit is proposed to make the start-up oscillation and start-up time of the VCO immune from temperature drifts. The prototype chip is measured in the range of −20 degrees to 90 degrees (Celsius) to verify this functionality of the constant-alpha biasing circuit. Besides, another new perspective on the detection level is proposed to derive the lowest BER. This SRR is fabricated in UMC standard$0.18\mu\mathrm{m}$CMOS process, and the chip area is$1.65\text{mm}\times 1.45\text{mm}$.
Yi-Pei Su, Chao-Yen Huang, Sao-Jie Chen
ISCAS3
2020 A Platform of Resynthesizing a Clock Architecture Into Power-and-Area Effective Clock Trees
abstract
To trigger events for application-specific data transfer among registers in a multimillion-gate system-on-chip (SoC), various kinds of clock signals, selectively driven by different frequency-dependent sources and/or dividers (DIVs), are usually centralized in one or more clock generation modules, where clock gating cells (CGCs), multiplexers (MUXes) and DIVs are used to create the clocks required by different functional operations in an SoC. These modules will introduce uncommon and longer timing paths for clock propagations and further make the clock tree synthesis (CTS) process become more challenging due to the on-chip-variation (OCV) effects. In addition, high volume of switching activities in the increased number of clock logic cells will consume more power. In this article, a novel design platform, merging and replacing of multiple multiplexers and dividers (MRMMD), is developed to intelligently identify those suspicious clock architectures and resynthesize them into a power-and-area effective and less complicated clock structure. Using our resynthesis platform, not only the number of clock-related timing paths and their corresponding logic levels can be reduced, but also the corresponding analysis and implementations of clock skew minimizations during CTS become much easier. The experimental results implemented in TSMC 55- and 28-nm process nodes on optimizing some industrial clock architectures showed that significant reductions of area, power, latency, skew and clock path, logic level, OCV impact, total wire length, and implementation runtime are achieved using our MRMMD platform.
Tung-Liang Lin, Sao-Jie Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2014 Accelerating Coverage Estimation Through Partial Model Checking
abstract
In model checking a system design against a set of properties, coverage estimation is frequently used to measure the amount of system behavior being checked by the properties. A popular coverage estimation method is to mutate the system model and check if the mutation can be detected by the given properties. For each mutation and each property, a full model check is required by some state-of-the-art coverage estimation methods. With such repeated model checking, mutation-based coverage estimation becomes significantly time-consuming. To alleviate this problem, a partial model checking (PMC) technique is proposed to recheck only those system states that were affected by a mutation, thus unnecessary rechecking of a large portion of the system states is avoided and time is saved. The PMC method has been integrated into the State Graph Manipulators model checker. Applying the proposed method to several examples showed that PMC has a saving of 50% to 70% in the coverage estimation time, and a reduction of 90% in mode visits.
Yean-Ru Chen, Jia-Jen Yeh, Pao-Ann Hsiung, Sao-Jie Chen
IEEE Trans. Computers4
2013 Backward probing deadlock detection for networks-on-chip
abstract
To accurately detect deadlocks in Network-on-Chip (NoC) as early as possible, a novel deadlock detection mechanism called Backward-probing Deadlock Detection (BDD) is proposed in this work, which can detect and resolve all existing deadlocks. It was realized using probe systems that generate probes for deadlock detection. A probe system includes a probe System Manager (SM) for turning on probe system, a probe Generator (GEN) for generating probes, a Link Selection (LS) connected to a Switch Allocation (SA), which is used for copying the generated probes, transmitting probes backward, and discarding probes when the probes find that the traversal path is just a congestion not a deadlock or when probe congestion occurs. There is also a TB Calculation (TBC) in LS for TB settings. Finally, a probe comparator (PB Comparator) is used for claiming deadlocks. Note that each port except the local one in a router has its own probe system.
Yean-Ru Chen, Zi-Rong Wangt, Pao-Ann Hsiung, Sao-Jie Chen, Meng-Hsun Tsai
NOCS4
2013 A unified link-layer fault-tolerant architecture for network-based many-core embedded systems
Wen-Chung Tsai, Deng-Yuan Zheng, Yu Hen Hu, Sao-Jie Chen
J. Syst. Archit.4
2013 Cycle-Efficient LFSR Implementation on Word-Based Microarchitecture
abstract
Cycle-efficient implementation of the linear feedback shift register (LFSR) algorithm on a word-based microarchitecture is investigated. This work examines an algorithm transformation method, called term-preserving look-ahead transformation (TePLAT), that transforms the bit-serial LFSR algorithm into a bit parallel format while maintaining the overhead of the original LFSR algorithm. Detailed implementation methodologies as well as extensive simulation results are presented. We apply TePLAT to 25 commonly used LFSRs and test the resulting parallel formulations on two popular word-based microprocessor development platforms: a Texas Instrument C6416 Code Composition Simulator and an ARM-9 Simulator. In all 25 cases, TePLAT transformed LFSR formulations consistently achieve much higher throughput than those of a naïve implementation and a traditional look-ahead transformation-based implementation.
Jui-Chieh Lin, Sao-Jie Chen, Yu Hen Hu
IEEE Trans. Computers2
2012 Congestion-aware scheduling for NoC-based reconfigurable systems
abstract
Network-on-Chip (NoC) is becoming a promising communication architecture in place of dedicated interconnections and shared buses for embedded systems. Nevertheless, it has also created new design issue such as communication congestion and power consumption. A major factor leading to communication congestion is mapping of application tasks to NoC. Latency, throughput, and overall execution time are all affected by task mapping. As a solution, an efficient run-time Congestion-Aware Scheduling (CWS) is proposed for NoC-based reconfigurable systems, which predicts traffic pattern based on the link utilization. The proposed algorithm alleviates the overall congestion, instead of only improving the current packet blocking situation. Our experiment results have demonstrated that compared to other existing congestion-aware algorithm, the proposed CWS algorithm can reduce the average communication latency by 66%, increase the average throughput by 32%, reduce the energy consumption by 23%, and decrease the overall execution by 32%.
Hung-Lin Chao, Yean-Ru Chen, Sheng-Ya Tong, Pao-Ann Hsiung, Sao-Jie Chen
DATE5
2012 Cycle-efficient lineary feedback shift register implementation on word-based micro-architecture
abstract
A novel algorithm transformation method, called term-preserving look-ahead transformation (TePLAT) is proposed to transform the bit-serial linear feedback shift register (LFSR) algorithm into a bit-parallel formulation which promises order of magnitudes improvement of execution speed compared to the traditional look-ahead algorithm transformation approach. TePLAT is applied to 26 commonly used LFSRs and tested on two popular word-based micro-processor development platforms: a Texas Instrument C6416 Code Composition Simulator and an ARM-9 Simulator. In all 26 cases, TePLAT transformed implementations consistently deliver much higher throughput than those implementations based on traditional look-ahead algorithm transformation.
Jui-Chieh Lin, Sao-Jie Chen, Yu Hen Hu
ICASSP2
2012 Optimal Bit-allocation for Wavelet-based Scalable Video Coding
abstract
We investigate the wavelet-based scalable video coding problem and present a solution that takes account of each user's preferred resolution. Based on the preference, we formulate the bit allocation problem of wavelet-based scalable video coding. We propose three methods to solve the problem. The first is an efficient Lagrangian-based method that solves the upper bound of the problem optimally, and the second is a less efficient dynamic programming method that solves the problem optimally. Both methods require knowledge of the user's preference. For the case where the user's preference is unknown, we solve the problem by a min-max approach. Our objective is to find the bit allocation solution that maximizes the worst possible performance. We show that the worst performance occurs when all users subscribe to the same spatial, temporal, and quality resolutions. Thus, the min-max solution is exactly the same as the traditional bit allocation method for a non-scalable wavelet codec. We conduct several experiments on the 2D+t MCTF-EZBC wavelet codec with respect to various subscribers' preferences. The results demonstrate that knowing the users' preferences improves the coding performance of the scalable video codec significantly.
Guan-Ju Peng, Wen-Liang Hwang, Sao-Jie Chen
ICME3
2012 A scalable and fault-tolerant network routing scheme for many-core and multi-chip systems
Wen-Chung Tsai, Kuo-Chih Chu, Yu Hen Hu, Sao-Jie Chen
J. Parallel Distributed Comput.4
2012 Design and Implementation of Block-Based Partitioning for Parallel Flip-Chip Power-Grid Analysis
abstract
Power-grid analysis is one of the critical design steps to ensure circuit reliability and achieve performance targets for very large scale integration systems. With each new technology generation, the circuit size has decreased and the power density has increased. Consequently, power-grid analysis has become ever more complex with greater CPU runtime and memory usage requirements. For a state-of-the-art power-grid design with more than 100-million nodes, it is often desirable to partition the power grid into smaller regions and analyze them in parallel by exploiting the locality of flip-chip packages. However, the traditional area-based partitioning strategy may not be best suited to analyze the DC current and ohmic IR voltage drop of a design that has irregular power rails and nonuniform power consumption because such nonuniformity affects the locality of power supply network and the accuracy of analysis. In this paper, we will present the analysis of a flip-chip design with 136-million nodes and propose a block-based partitioning scheme to improve the accuracy of parallel power-grid analysis.
Chun-Jen Wei, Howard Chen 0001, Sao-Jie Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2012 Interlayer Bit Allocation for Scalable Video Coding
abstract
In this paper, we present a theoretical analysis of the distortion in multilayer coding structures. Specifically, we analyze the prediction structure used to achieve temporal, spatial, and quality scalability of scalable video coding (SVC) and show that the average peak signal-to-noise ratio (PSNR) of SVC is a weighted combination of the bit rates assigned to all the streams. Our analysis utilizes the end user's preference for certain resolutions. We also propose a rate-distortion (R-D) optimization algorithm and compare its performance with that of a state-of-the-art scalable bit allocation algorithm. The reported experiment results demonstrate that the R-D algorithm significantly outperforms the compared approach in terms of the average PSNR.
Guan-Ju Peng, Wen-Liang Hwang, Sao-Jie Chen
IEEE Trans. Image Process.3
2011 A fault-tolerant NoC scheme using bidirectional channel
abstract
A novel Bidirectional Fault-Tolerant NoC (BFT-NoC) architecture capable of mitigating both static and dynamic channel failures is proposed. In a traditional NoC platform, a faulty data channel will force blocked packets to make costly detours, resulting in significant performance hits. In this work, novel fault-tolerance measures for a bidirectional NoC platform are proposed. The dynamically reconfigurable bidirectional channels of the BFT-NoC offer great flexibility to contain data-link permanent or transient faults while incurring negligible performance loss. Potential performance advantages in terms of failure rate reduction and reliability enhancement of the BFT-NoC architecture are carefully analyzed. Extensive experimental results clearly validate the fault-tolerance performance of BFT-NoC at both synthetic and real world network traffic patterns.
Wen-Chung Tsai, Deng-Yuan Zheng, Sao-Jie Chen, Yu Hen Hu
DAC3
2011 A Bidirectional NoC (BiNoC) Architecture With Dynamic Self-Reconfigurable Channel
abstract
A bidirectional channel network-on-chip (BiNoC) architecture is proposed to enhance the performance of on-chip communication. In a BiNoC, each communication channel allows to be dynamically self-reconfigured to transmit flits in either direction. This added flexibility promises better bandwidth utilization, lower packet delivery latency, and higher packet consumption rate. Novel on-chip router architecture is developed to support dynamic self-reconfiguration of the bidirectional traffic flow. This area-efficient BiNoC router delivers better performance and requires smaller buffer size than that of a conventional network-on-chip (NoC). The flow direction at each channel is controlled by a channel direction control (CDC) algorithm. Implemented with a pair of finite state machines, this CDC algorithm is shown to be high performance, free of deadlock, and free of starvation. Extensive cycle-accurate simulations using synthetic and real-world traffic patterns have been conducted to evaluate the performance of the BiNoC. These results exhibit consistent and significant performance advantage over conventional NoC equipped with hard-wired unidirectional channels.
Ying-Cherng Lan, Yueh-Chi Lin, Shih-Hsin Lo, Yu Hen Hu, Sao-Jie Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2010 Cycle efficient scrambler implementation for software defined radio
abstract
The task of efficient implementation of bit-serial scrambler algorithms on a word-parallel software defined radio micro-architecture is considered. By exploiting properties of the exclusive-OR Boolean logic operator, a novel power-of-two look-ahead recursive algorithm transformation method is developed. Together with loop-unrolling, this new method produces an efficient vector scrambler algorithm formulation that realizes orders of magnitude clock-cycle savings compared to state of the art solutions. Using IEEE 802.11a scrambler algorithm as an example, this new formulation is 9 times faster than previously reported results.
Jui-Chieh Lin, Ming-Jung Fan-Chiang, Minja Hsieh, Song-Yen Mao, Sao-Jie Chen, Yu Hen Hu
ICASSP5
2010 QoS aware BiNoC architecture
abstract
A quality-of-service (QoS) aware, bi-directional channel NoC (BiNoC) architecture is proposed to support guarantee-service (GS) traffic while reducing packet delivery latency. By incorporating dynamically self-reconfigured bidirectional communication channels between adjacent routers, BiNoC architecture promises more flexibility for various traffic flow patterns. A novel inter-router communication protocol is proposed that prioritizes bandwidth arbitration in favor of high priority GS traffic flows. Multiple virtual channels with prioritized routing policy are also implemented to facilitate data transmission with QoS considerations. Combining these architectural innovations, the QoS aware BiNoC promises reduced latency of packet delivery and more efficient channel resource utilizations. Cycle-accurate simulations demonstrate significant performance advantage over conventional unidirectional NoC architecture equipped with hard-wired unidirectional channels.
Shih-Hsin Lo, Ying-Cherng Lan, Hsin-Hsien Yeh, Wen-Chung Tsai, Yu Hen Hu, Sao-Jie Chen
IPDPS6
2010 Perfect shuffling for cycle efficient puncturer and interleaver for software defined radio
abstract
This work applies perfect-shuffle on software defined radio's bit-shuffling blocks, such as puncturer and interleaver, which data alignment issue is also investigated. A demonstration of such algorithm on IEEE 802.11a (WiFi) standard is implemented with Texas Instruments®intrinsic perfect-shuffle and inverse perfect-shuffle instructions combined with ANSI C on a TMS320 C64x digital signal processor. Simulation results indicate the performance enhancement using perfect-shuffle technique dramatically improved cycle efficiency comparing to a naive one bit per integer implementation.
Jui-Chieh Lin, Minja Hsieh, Ming-Jung Fan-Chiang, Song-Yen Mao, Chu Yu, Sao-Jie Chen, Yu Hen Hu
ISCAS6
2010 TM-FAR: Turn-Model based Fully Adaptive Routing for Networks on Chip
abstract
A novel Turn-Model based Fully-Adaptive-Routing (TM-FAR) algorithm is proposed for Networks-on-Chip (NoC). TM-FAR retains the deadlock-free property of traditional turn-model based routing algorithms (e.g., XY, Odd-Even), while alleviating restrictions on turn and path selections. Just like the current Virtual-Channel based Fully-Adaptive-Routing (VC-FAR) algorithm, TM-FAR allows full exploitation of all available minimal paths, yet TM-FAR does not use virtual channels. This fully adaptive routing capability of TM-FAR promises improved routing adaptivity and enhanced level of fault-tolerance. Preliminary experimental results indicate that the TM-FAR achieves an averaged delay reduction of 30.08% and a throughput rate increase of 4.54% compared to the state-of-the-art NoC routing algorithm based on the Odd-Even turn model.
Wen-Chung Tsai, Kuo-Chih Chu, Sao-Jie Chen, Yu Hen Hu
VLSI-SoC3
2010 An Automatic Optical Simulation-Based Lithography Hotspot Fix Flow for Post-Route Optimization
abstract
In this paper, an optical simulation-based lithography hotspot fix guidance generator and an automatic hotspot fix flow are proposed. We develop our aerial image simulation engine by enhancing the traditional sum of coherence system method. Subject to the shape changes, a strong correlation between the aerial image intensity difference maps of pre-optical proximity correction (OPC) and post-OPC schemes is found. We collect near a litho hotspot in a pre-OPC layout some fix actions that are local shape changes to optimize the optical intensity. Then, fix guidances will be selected from the collected fix actions by a heuristic algorithm and input to a router for fixing the hotspot. We integrate the fix guidance generation method with a commercial lithography hotspot detection tool to create an automatic post-route optical-simulation-embedded local fix (OSELF) flow and test with industry 65 nm designs. Compared with the commercial flow that uses only local fix, our method has a 1.4x-1.9x fix rate, similar run time, no new design rule check violation, and negligible circuit timing impacts. We also combine our OSELF algorithm with a rip-up and reroute engine, and test on the same designs. Compared to the commercial tool that uses a hybrid (local fix plus reroute) fix flow, our combined flow runs 1.7x-2.9x faster with 45-55% circuit timing impact. Both flows achieve a 100% hotspot fix rate.
Yang-Shan Tong, Sao-Jie Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 Optimal Multiple-Bit Huffman Decoding
abstract
A variable-bit look-ahead Huffman decoding problem is investigated in this paper. The objective is to maximize the decoding throughput rate by exploiting different equivalent state diagrams. The decoding throughput rate is estimated based on the state transition probability of the corresponding Huffman encoding table. We propose an efficient algorithm to search for a heuristic solution that usually yields very good optimization results in polynomial computation time.
Ya-Nan Wen, Guang-Huei Lin, Sao-Jie Chen, Yu Hen Hu
IEEE Trans. Circuits Syst. Video Technol.3
2009 Using XML for VLSI Physical Design Automation
Fong-Ming Shyu, Po-Hsun Cheng, Sao-Jie Chen
ICA3PP3
2009 A 900 MHz to 5.2 GHz Dual-loop Feedback Multi-band LNA
abstract
This paper demonstrates a multi-band low noise amplifier (LNA) in 0.13µm CMOS process, which is configurable with switching capacitor in 900MHz, 1800MHz, 2.4GHz and 5.2GHz bands. A dual-loop feedback technique is used to enhance the performance in noise figure, power gain and power consumption. The noise figures are 2.6dB at 900MHz, 2dB at 1800MHz, 2.1dB at 2.4GHz and 3.5dB at 5.2GHz; and the power gains are 17dB at 900MHz, 21dB at 1800MHz, 26dB at 2.4GHz and 19dB at 5.2GHz in post-simulation. The S11in all of the bands is below −10dB by wideband matching. The LNA consumes 3.92mW from 1 V supply.
Jia-Wei Lin, Da-Tong Yen, Wei-Yi Hu, Chu Yu, Mao-Hsu Yen, Pao-Ann Hsiung, Sao-Jie Chen
ISCAS7
2009 An automatic optical-simulation-based lithography hotspot fix flow for post-route optimization
abstract
In this paper, an automatic lithography hotspot fix guidance generation method based on optical simulation engine and a fix flow utilizing this method are proposed. We find that subject to shape changes of original layout, there is strong correlation between the optical intensity maps of pre-OPC and post-OPC layouts. From this fact, we develop a fix action generation method by optimizing optical intensity in terms of polygon changes near a litho hotspot in a pre-OPC layout. Fix guidances as a subset of fix actions of a hotspot can be selected by a greedy method. The fix guidances are used as input to router to fix the hotspots. We integrate this fix guidance generation method with commercial lithography hotspot detection tool to create a post-route lithography hotspot fixing flow and test with industry 65nm designs. Compared with commercial flow, our method has 1.4X~1.9X fix rate with similar run time. The circuit timing impact is negligible in both our flow and industry flow. And our flow will not introduce any new design rule check (DRC) violation.
Yang-Shan Tong, Chia-Wei Lin, Sao-Jie Chen
ISPD3
2009 BiNoC: A bidirectional NoC architecture with dynamic self-reconfigurable channel
abstract
A Bidirectional channel Network-on-Chip (BiNoC) architecture is proposed to enhance the performance of on-chip communication. The BiNoC allows each communication channel to be dynamically self-configured to transmit flits in either direction in order to better utilize on-chip hardware resources. This added flexibility promises better bandwidth utilization, lower packet delivery latency, and higher packet consumption rate at each on-chip router. In this paper, a novel on-chip router architecture supporting the self-configuring bidirectional channel mechanism is presented. It is shown that the associated hardware overhead is negligible. Cycle-accurate simulation runs on this BiNoC network under synthetic and real-world traffic patterns demonstrate consistent and significant performance advantage over conventional mesh grid NoC architecture equipped with hard-wired unidirectional channels.
Ying-Cherng Lan, Shih-Hsin Lo, Yueh-Chi Lin, Yu Hen Hu, Sao-Jie Chen
NOCS5
2007 Modeling and Automatic Failure Analysis of Safety-Critical Systems Using Extended Safecharts
Yean-Ru Chen, Pao-Ann Hsiung, Sao-Jie Chen
SAFECOMP3
2006 A parameterizable digital-approximated 2D Gaussian smoothing filter for edge detection in noisy image
abstract
This paper describes a 5/spl times/5 window-based, power-of-two approximation algorithm for 2D Gaussian smoothing filter and its generic hardware reference design. With a small change in standard deviation /spl sigma/, the Gaussian smoothing filter is able to provide different levels of smoothing and noise reduction capability, which is highly desirable in the early stage of an image processing flow. By using power-of-two terms, the digital-approximated 2D Gaussian filter can be implemented by simple hardware shifters. A hardware reference design is also proposed to achieve such functionality. Experiments in using the proposed filter with an existing edge detection algorithm show the flexibility and effectiveness of the proposed smoothing mask. The synthesized hardware of such combination shows a 200MHz high operating frequency under UMC 0.18/spl mu/m technology, which proves the proposed filter is able to achieve the real-time high frame-rate and high-resolution video processing requirements.
Pei-Yung Hsiao, Chia-Hsiung Chen, Shin-Shian Chou, Le-Tien Li, Sao-Jie Chen
ISCAS5
2006 Multilevel routing with jumper insertion for antenna avoidance
Tsung-Yi Ho, Yao-Wen Chang, Sao-Jie Chen
Integr.3
2006 RTWPMS: A Real-Time Wireless Physiological Monitoring System
abstract
This paper demonstrates the design and implementation of a real-time wireless physiological monitoring system for nursing centers, whose function is to monitor online the physiological status of aged patients via wireless communication channel and wired local area network. The collected data, such as body temperature, blood pressure, and heart rate, can then be stored in the computer of a network management center to facilitate the medical staff in a nursing center to monitor in real time or analyze in batch mode the physiological changes of the patients under observation. Our proposed system is bidirectional, has low power consumption, is cost effective, is modular designed, has the capability of operating independently, and can be used to improve the service quality and reduce the workload of the staff in a nursing center.
Bor-Shyh Lin, Bor-Shing Lin, Nai-Kuan Chou, Fok-Ching Chong, Sao-Jie Chen
IEEE Trans. Inf. Technol. Biomed.5
2005 Improving End-to-End Performance by Active Queue Management
abstract
Active queue management (AQM) schemes have motivated many researchers to investigate more effective methods to control network congestion. Most AQM schemes are evaluated by their designers on the basis of router-centric metrics, such as queuing delay, link utilization and packet drop ratio. These metrics are important to network operators but they may not reflect the quality of service delivered to end-users. In this paper we propose a method aimed to provide users with better services in terms of end-to-end delay and packet loss ratio. The method captures a significant traffic increase at an early stage and signals TCP sources to slow down. In this way, TCP can quickly adjust the transmission rate, and thereby prevent overloading the network. In addition to RED, another two prominent AQM schemes, namely, random early marking (REM) and adaptive RED (ARED), are compared with the proposed method. Simulation results show that under various network loads and a range of network propagation delays, our method can achieve lower end-to-end delay and packet loss ratio, compared with all aforementioned schemes.
Chin-Fu Ku, Sao-Jie Chen, Jan-Ming Ho, Ray-I Chang
AINA2
2005 Multilevel full-chip routing for the X-based architecture
abstract
As technology advances into the nanometer territory, the interconnect delay has become a first-order effect on chip performance. To handle this effect, the X-architecture has been proposed for high-performance integrated circuits. The X-architecture presents a new way of orienting a chip's microscopic interconnect wires with the pervasive use of diagonal routes. It can reduce the wirelength and via count, and thus improve performance and routability. Furthermore, the continuous increase of the problem size of IC routing is also a great challenge to existing routing algorithms. In this paper, we present the first multilevel framework for full-chip routing using the X-architecture. To take full advantage of the X-architecture, we explore the optimal routing for three-terminal nets on the X-architecture and develop a general X-Steiner tree algorithm based on the delaunay triangulation approach for the X-architecture. The multilevel routing framework adopts a two-stage technique of coarsening followed by uncoarsening, with a trapezoid-shaped track assignment embedded between the two stages to assign long, straight diagonal segments for wirelength reduction. Compared with the state-of-the-art multilevel routing for the Manhattan architecture, experimental results show that our approach reduced wirelength by 18.7% and average delay by 8.8% with similar routing completion rates and via counts.
Tsung-Yi Ho, Chen-Feng Chang, Yao-Wen Chang, Sao-Jie Chen
DAC4
2005 Crosstalk- and performance-driven multilevel full-chip routing
abstract
In this paper, we propose a novel framework for fast multilevel routing considering crosstalk and performance optimization. To handle the crosstalk minimization problem, we incorporate an intermediate stage of layer/track assignment into the multilevel routing framework. For performance-driven routing, we propose a novel minimum-radius minimum-cost spanning tree heuristic for global routing. Compared with the state-of-the-art multilevel routing with the routability mode, the experimental results show that our router achieved a 6.7X runtime speedup, reduced the respective maximum and average crosstalk (coupling length) by about 30% and 24%, reduced the respective maximum and average delay by about 15% and 5%. Compared with the timing-driven mode, the experimental results show that our router still achieved a 5.9X runtime speedup, reduced the respective maximum and average crosstalk by about 35% and 23%, reduced the respective maximum and average delay by about 7% and 10% in comparable routability, and resulted in fewer failed nets.
Tsung-Yi Ho, Yao-Wen Chang, Sao-Jie Chen, D. T. Lee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 Multilevel routing with antenna avoidance
abstract
As technology advances into the nanometer territory, the antenna problem has caused significant impact on routing tools. The antenna effect is a phenomenon of plasma-induced gate oxide degradation caused by charge accumulation on conductors. It directly influences manufacturability and yield of VLSI circuits, especially in deep-submicron technology using high density plasma. Furthermore, the continuous increase of the problem size of IC routing is also a great challenge to existing routing algorithms. In this paper, we propose a novel framework for multilevel full-chip routing with antenna avoidance using a built-in jumper insertion approach. Experimental results show that our approach re-duced antenna-violated gates by about 98 % and also achieved
Tsung-Yi Ho, Yao-Wen Chang, Sao-Jie Chen
ISPD3
2004 Fast postplacement optimization using functional symmetries
abstract
The timing-convergence problem arises because estimations made during logic synthesis may not be met during physical design. In this paper, an efficient rewiring engine is proposed to explore maximal freedom after placement. The most important feature of this approach is that the existing placement solution is left intact throughout the optimization. A linear-time algorithm is proposed to detect functional symmetries in the Boolean network which are then used as the basis for rewiring. Integration with an existing gate-sizing algorithm further proves the effectiveness of our technique. Three applications are demonstrated: delay, power, and reliability optimization.
Chih-Wei Jim Chang, Ming-Fu Hsiao, Bo Hu 0006, Kai Wang 0011, Malgorzata Marek-Sadowska, Chung-Kuan Cheng, Sao-Jie Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2003 A Fast Crosstalk- and Performance-Driven Multilevel Routing System
Tsung-Yi Ho, Yao-Wen Chang, Sao-Jie Chen, D. T. Lee
ICCAD3
2003 Software Platform for Embedded Software Development
Win-Bin See, Pao-Ann Hsiung, Trong-Yen Lee, Sao-Jie Chen
RTCSA4
2002 TCN: Scalable Hierarchical Hypercubes
abstract
Hierarchical hypercubes, such as extended hypercube (EH), hyperweave (HW), and extended hypercube with cross connections (EHC), have been proposed to overcome the scalability limitation of conventional hypercubes through the use of fixed dimension hypercubes of processing elements (PEs) as basic modules interconnected by network controllers (NC) which are themselves interconnected into hypercubes. The scalability of all these three hierarchical hypercube networks is still limited, because the average network communication load in each NC increases as the number of interconnected PEs become very large. In this work a generalization scheme is proposed for improving network scalability, namely transformer cube network (TCN). For illustration purposes, generalized TCN is presented only for EH, though the same scheme can be applied to HW and EHC as well. Several characteristics of TCN, such as topological properties, message routing complexity, fault tolerance, and scalability are analyzed We present a communication algorithm for one-to-one message passing in a fault-free case. Further, the application of TCN to a class of divide-and-conquer problems is shown to have a time complexity of O(log/sub 2/ N), where N is the total number of PEs.
Trong-Yen Lee, Pao-Ann Hsiung, Sao-Jie Chen
ICPADS3
2001 Formal Verification of Embedded Real-Time Software in Component-Based Application Frameworks
abstract
Producing correct software is a major goal for application frameworks that are targeted at embedded real-time systems because incorrect software is of no use and may also cause severe system damage. It is shown how formal verification can be elegantly, seamlessly, and scalably integrated into a component-based object-oriented application framework for embedded real-time systems. Two issues in such technology integration are addressed: (1) the choice of a common system model, and (2) the integration of formal synthesis and model checking. Solutions are provided, respectively, in the form of (1) proposing a new formal object-oriented model (FOOM), and (2) the execution of model checkers within synthesis algorithms. Technically, we propose a compositional software verification framework, in which model checking is employed, with state-space reduction techniques adapted for embedded real-time software. A separate verifier component is proposed for modular integration as illustrated by its implementation in the VERTAF application framework. An example illustrates the success of our approach and the benefits gained through integrating formal verification.
Pao-Ann Hsiung, Win-Bin See, Trong-Yen Lee, Jih-Ming Fu, Sao-Jie Chen
APSEC5
2001 A Three-Stage One-Sided Rearrangeable Polygonal Switching Network
abstract
This paper proposes a three-stage rearrangeable polygonal switching network (PSN) for interconnecting one-sided input-output terminals. In comparing our PSN with a three-stage one-sided Clos switching network of the same size and with the same number of switches, we prove that rearrangeability of a PSN is better than that of a Clos switching network. Also, the switching efficiency of the PSN is explored.
Mao-Hsu Yen, Sao-Jie Chen, Sanko Lan
IEEE Trans. Computers2
2000 Efficient routability check algorithms for segmented channel routing
abstract
The segmented channel-routing problem arises in the context of row-based field programmable gate arrays (FPGAs). Since the K -segment channel-routing problem is NP-complete for K ≥ 2, an efficient algorithm using the weighted bipartite-matching approach is developed for this problem. Connections that;form a maximum clique are chosen first to be routed to the segmented channel. Then, another maximum clique of the remained connections is routed until all connections have been processed. In addition, a powerful “unroutability check” algorithm is uniquely proposed to tell whether the horizontal switches in an interval of the segmented channel are sufficient for routing or not. Hence, we can precisely discriminate the routable and the unroutable ones from all the test cases. As shown in the experiments, average discrimination ratios of 98.8% and 99.4% are obtained for the 2-segmentation and 3-segmentation models, respectively. Moreover, when applying our routing algorithm to the analyzed nonunroutable cases, a routing failure ratio of 1.5% is reported for the 2-segmentation model, compared to Zhu and Wong's 5.9%; also, a routing failure ratio of 0.8% (less than their 4.7%) is obtained for the 3-segmentation model. In total, the routing failure ratio of our routing algorithm is less than 21% of Zhu and Wong's.
Cheng-Hsing Yang, Sao-Jie Chen, Jan-Ming Ho, Chia-Chun Tsai
ACM Trans. Design Autom. Electr. Syst.2
1999 An Automatic Router for the Pin Grid Array Package
abstract
A Pin-Grid-Array (PGA) package router is presented in this paper. Given a chip cavity with a number of I/O pads around its boundary and an equivalent number of pins distributed on the substrate, the objective of the router is to complete the planar interconnection of pad-to-pin nets on one or more layers. This router consists of three phases: layer assignment topological routing, and geometrical routing. Examples tested on a windows-based environment show that our router is efficient and can complete the routing task with less substrate layers. Compared to manual routing, this router features a friendly graphic user interface and can be practically applied to VLSI packaging.
Shuenn-Shi Chen, Jong-Jang Chen, Sao-Jie Chen, Chia-Chun Tsai
ASP-DAC3
1999 An Efficient Two-Level Partitioning Algorithm for VLSI Circuits
abstract
In this paper, a new two-level bipartitioning algorithm TLP, combining a hybrid clustering technique with an iterative improvement based partitioning process, is proposed. The hybrid clustering algorithm consisting of a local bottom-up clustering technique to merge modules and a global top-down ratio-cut technique for decomposition can be used to reduce the partitioning complexity and improve the performance. To generate a high-quality partitioning solution, a module migration based partitioning algorithm MMP is also proposed as the base partitioner for the TLP algorithm. The MMP algorithm implicitly promotes the move of clusters during the module migration processes by paying more attention to the neighbors of moved modules, relaxing the size constraints temporarily during the migration process, and controlling the module migration direction. Experimental results obtained show that the TLP algorithm generates stable and high-quality partitioning results. The TLP algorithm improves the unstable property of module migration based algorithms such as FM and STABLE in terms of the average net cut value. On the other hand, TLP outperforms MELO, GFM/sub t/ and CDIP/sub LA3/ by 23%, 7%, and 10%, respectively and is competitive with hMetis, ML/sub c/ and LSR/MFFS which have generated better results than many recent state-of-the-art partitioning algorithms.
Jong-Sheng Cherng, Sao-Jie Chen, Chia-Chun Tsai, Jan-Ming Ho
ASP-DAC2
1999 An Even Wiring Approach to the Ball Grid Array Package Routing
abstract
An even-wiring router for the BGA package is presented to interconnect each I/O pad of a chip to a corresponding ball distributed on the substrate area. The major phases for the router consist of layer assignment, topological routing, and physical routing. Using this router, we can generate an even distribution of planar and any-angle wires to improve manufacturing yield. We have also conducted various testing examples to verify the efficiency of this router. Experiments show that the router produces very good results, far better than the manual design, thus it can be practically applied to VLSI packaging.
Shuenn-Shi Chen, Jong-Jang Chen, Sao-Jie Chen, Chia-Chun Tsai
ICCD3
1998 NEWS: a net-even-wiring system for the routing on a multilayer PGA package
abstract
Given a die with I/O pads wire bonded onto the multilayer substrates of a pin grid array (PGA) package, a three-step net-even-wiring system (NEWS) is proposed to complete the routing of the bond pads to the corresponding grid pins on one or more layers. First, we performed a maximum-cut partitioning on a net interference graph for the layer assignment step. Second, nets on each layer were converted to planar sketches using a novel insertion sort method. Last, the planar sketches were transformed into a net-even-wired layout using a simplified rubber-band router.
Chia-Chun Tsai, Chwan-Ming Wang, Sao-Jie Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1998 ICOS: an intelligent concurrent object-oriented synthesis methodology for multiprocessor systems
abstract
The design of multiprocessor architectures differs from uniprocessor systems in that the number of processors and their interconnection must be considered. This leads to an enormous increase in the design-space exploration time, which is exponential in the total number of system components. The methodology proposed here, called Intelligent Concurrent Object-Oriented Synthesis (ICOS) methodology, makes feasible the synthesis of complex multiprocessor systems through the application of several techiques that speed up the design process. ICOS is based on Performance Synthesis Methodology (PSM), a recently proposed object-oriented system-level design methodology. Four major techniques: object-oriented design, fuzzy design-space exploration, concurrent design, and intelligent reuse of complete subsystems are integrated in ICOS. First, object-oriented modeling and design, through the use of object-oriented relationships and operators, make the whole design process manageable and maintainable in ICOS. Second, fuzzy comparison applied to the specializations or instances of components reduces the exponential growth of design-space exploration in ICOS. Third, independent components from different design alternatives are synthesized in parallel; this design concurrency shortens the overall design time. Lastly, the resynthesis of complete subsystems can be avoided through the application of learning, thus making the methodology intelligent enough to reuse previous design configurations. Experiments show that all these applied techniques contribute to the synthesis efficiency and the degree of automation in ICOS.
Pao-Ann Hsiung, Chung-Hwang Chen, Trong-Yen Lee, Sao-Jie Chen
ACM Trans. Design Autom. Electr. Syst.4
1997 Hmap: a fast mapper for EPGAs using extended GBDD hash tables
abstract
A fast and efficient algorithm for technology mapping of electrically programmable gate arrays (EPGAs) is proposed. This Hmap algorithm covers the Boolean network with programmed logic modules bottom-up. The covering operation is based on collapsing the fanins of a node to form a bigger supernode such that fewer clusters are needed to be detected. Then Boolean matching is used to detect whether the collapsed supernode can be mapped into a logic module by looking up an extended GBDD hash table. The use of this table look-up matching can shorten the matching time significantly. As shown in the experiments, the average running time of Hmap is 20 times faster than that of MIS-pga2.
Cheng-Hsing Yang, Chia-Chun Tsai, Jan-Ming Ho, Sao-Jie Chen
ACM Trans. Design Autom. Electr. Syst.4
1996 PSM: an object-oriented synthesis approach to multiprocessor system design
abstract
Although multiprocessor systems are becoming a trend today, few synthesis tools currently available can actually automate the design of multiprocessor systems. Performance synthesis methodology (PSM) is an object-oriented system-level synthesis approach to multiprocessor system design. Since PSM was designed specifically for the synthesis of multiprocessor systems, it is not only much more efficient when synthesizing parallel systems, but also produces better parallel systems than currently available uniprocessor system-level synthesis tools. Colored Petri nets used in modeling system components and object modeling technique used in the design process have both contributed to the shortening of system development time and to the reduction of design cost. First, user specification consisting of functional models and performance constraints is translated into architecture models. Then, the system is configured by selecting the method of control, the memory organization, the type of processor, and the type of system interconnection. Finally, a heuristic design space exploration algorithm is used to generate several near-optimal design alternatives. The best architecture is chosen by evaluating the design alternatives using a flexible performance estimation formula that mainly considers system level design features, such as system throughput, utilization, reliability, scalability, fault-tolerance, and cost. Several systems were successfully synthesized using this top-down object-oriented PSM, thus showing its feasibility as a design automation tool for parallel systems.
Pao-Ann Hsiung, Sao-Jie Chen, Tsung-Chien Hu, Shih-Chiang Wang
IEEE Trans. Very Large Scale Integr. Syst.2
1992 An H-V alternating router
abstract
An H-V alternating router based on the concurrent H (horizontal) and V (vertical) tile expansions is presented. The router is modeled by a sequence of alternating H and V corner-stitching space tiles, where the expansion direction is controlled by a heuristic evaluation function using the A* technique and the damping concept. Tile growing is governed by the following three factors: constrained expansion area, limited expansion depth, and oriented expansion direction. All the H-V tile expansion operations can be easily performed on a specially designed net-forest structure. It is shown that this approach generates nearly optimal connection paths with a minimum number of bends and always guarantees a feasible solution if such a path exists. The performance of this router is better than that of H-only tile-expansion routers. This router is also well suited for wiring hierarchical modules with the metal-metal matrix technology and can be extended to multilayer layouts.>
Chia-Chun Tsai, Sao-Jie Chen, Wu-Shiung Feng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1991 Constrained via Minimization with Practical Considerations for Multi-Layer VLSI/PCB Routing Problems
abstract
Article Free Access Share on Constrained via minimization with practical considerations for multi-layer VLSI/PCB routing problems Authors: Sung-Chuan Fang Department of Electrical Engineering, National Taiwan University, Taipei, Taiwan, R.O.C. Department of Electrical Engineering, National Taiwan University, Taipei, Taiwan, R.O.C.View Profile , Kuo-En Chang Department of Information and Computer Education, National Taiwan Normal University, Taipei, Taiwan, R.O.C. Department of Information and Computer Education, National Taiwan Normal University, Taipei, Taiwan, R.O.C.View Profile , Wu-Shiung Feng Department of Electrical Engineering, National Taiwan University, Taipei, Taiwan, R.O.C. Department of Electrical Engineering, National Taiwan University, Taipei, Taiwan, R.O.C.View Profile , Sao-Jie Chen Department of Electrical Engineering, National Taiwan University, Taipei, Taiwan, R.O.C. Department of Electrical Engineering, National Taiwan University, Taipei, Taiwan, R.O.C.View Profile Authors Info & Claims DAC '91: Proceedings of the 28th ACM/IEEE Design Automation ConferenceJune 1991 Pages 60–65https://doi.org/10.1145/127601.127629Published:01 June 1991Publication History 15citation536DownloadsMetricsTotal Citations15Total Downloads536Last 12 Months43Last 6 weeks6 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Sung-Chuan Fang, Kuo-En Chang, Wu-Shiung Feng, Sao-Jie Chen
DAC4
1990 Generalized terminal connectivity problem for multilayer layout scheme
Chia-Chun Tsai, Sao-Jie Chen, Wu-Shiung Feng
Comput. Aided Des.2
1990 GM Plan: a gate matrix layout algorithm based on artificial intelligence planning techniques
abstract
The CMOS gate matrix layout problem is formulated and solved as an artificial intelligence planning problem in which a plan (the solution algorithm) is to be generated to achieve a goal (the gate matrix layout). The overall goal consists of many subgoals, each of which corresponds to the placement of a gate to a slot, and to the routing of associated nets connecting to that gate. As different nets compete for track (resource) usage, these subgoals interact (interfere) with each other, rendering suboptimal solutions. Here, such interaction among subgoals is managed with two artificial intelligence planning techniques: hierarchical subgoal organization and domain-independent search control policies. The subgoal hierarchy facilitates an object classification of the subgoals into priority classes according to a proposed distance measure of connectivity. Two search control policies (general problem-solving heuristics)-most-constraint (MC) and least impact (LI)-are used to guide the search process. A planning-based gate matrix layout algorithm, called GM Plan, which combines the gate placement and net routing into a single, incremental, problem-solving loop has been developed using these techniques. Encouraging results have been observed in a number of test examples.>
Yu Hen Hu, Sao-Jie Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2