Jing-Yang Jou

dblp:26/5117 · DBLP profile ↗
← Back
86ranked-venue papers
7as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 85 · 6 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 An Evaluation and Architecture Exploration Engine for CNN Accelerators through Extensive Dataflow Analysis
abstract
Systolic array is one of the popular convolutional neural network accelerator architectures due to its high computation efficiency. Nevertheless, the huge design space and complicated interactions among different design parameters make it hard to find the best configuration for various applications. To overcome this issue, this paper presents an evaluation and design space exploration engine, NNeed, for systolic-array CNN accelerators through extensive dataflow analysis. It uses a highly configurable hardware template to describe accelerator operations in detail. The rapid evaluation provides PPA results, pipeline stage analysis, external memory access statistics, and so on. NNeed explores the 9-dimensional design space and supports multiple objective functions for design optimization. Experimental results show that NNeed can generate an accelerator configuration with up to 23% and 50% improvement in performance and energy as compared with a typical handcrafted design.
Shan-Hui Chou, Ting-Yun Hsiao, Jing-Yang Jou, Juinn-Dar Huang
VLSI-SoC3
2021 1st-Order to 2nd-Order Threshold Logic Gate Transformation with an Enhanced ILP-based Identification Method
abstract
This paper introduces a method to enhance an integer linear programming (ILP)-based method for transforming a 1st-order threshold logic gate (1-TLG) to a 2nd-order TLG (2-TLG) with lower area cost. We observe that for a 2-TLG, most of the 2nd-order weights (2-weights) are zero. That is, in the ILP formulation, most of the variables for the 2-weights could be set to zero. Thus, we first propose three sufficient conditions for transforming a 1-TLG to a 2-TLG by extracting 2-weights. These extracted weights are seen to be more likely non-zero. Then, we simplify the ILP formulation by eliminating the non-extracted 2-weights to speed up the ILP solving. The experimental results show that, to transform a set of 1-TLGs to 2-TLGs, the enhanced method saves an average of 24% CPU time with only an average of 1.87% quality loss in terms of the area cost reduction rate.
Li-Cheng Zheng, Hao-Ju Chang, Yung-Chih Chen, Jing-Yang Jou
ASP-DAC4
2021 Performance-driven Routing Methodology with Incremental Placement Refinement for Analog Layout Design
abstract
Analog layout is often considered as a difficult task because many layout-dependent effects will impact final circuit performance. In the literature, many automation techniques have been proposed for analog placement and routing respectively. However, very few works are able to consider the two steps simultaneously to obtain the best performance and cost after layout. Most of the routing-aware placement techniques optimize the layout results based on an assumed routing result, which may be quite different to the final layout. In this work, we proposed an automatic two-step layout methodology for analog circuits to alleviate the performance loss during layout process. Instead of using a rough routing prediction during placement stage, a crossing-aware global routing technique is first performed to provide an accurate routing resource estimation of the given compact placement. Then, the improved CDL-based layout migration technique is adopted to do a fast adjustment on the placement and routing to reduce the difference between estimation and final layout while keeping the optimality of the given placement. As shown in the experimental results, the proposed methodology is able to improve the accuracy of routing resource estimation thus improving the final layout quality and circuit performance.
Hao-Yu Chi, Han-Chung Chang, Chih-Hsin Yang, Chien-Nan Jimmy Liu, Jing-Yang Jou
DATE5
2016 Chain-based pin count minimization for general-purpose digital microfluidic biochips
abstract
Minimizing the number of external control pins is one of the most important optimization objectives in digital microfluidic biochip (DMFB) designs especially as the chip size gets even bigger. So far, only few works focus on this issue for general-purpose DMFBs. In this paper, we present a pin count minimization algorithm based on sophisticated electrode chaining on regular or irregular electrode arrays. The key idea of the proposed method is that actuation information can be implied from previous neighborhood electrodes to later ones throughout a chain. Experimental results show that the pin count reduction can be near 50% in large DMFBs.
Yung-Chun Lei, Chen-Shing Hsu, Juinn-Dar Huang, Jing-Yang Jou
ASP-DAC4
2016 Resource-aware functional ECO patch generation
An-Che Cheng, Iris Hui-Ru Jiang, Jing-Yang Jou
DATE3
2016 Wave digital filter based analog circuit emulation on FPGA
abstract
Unlike well accepted FPGA emulation for digital circuits, there is no winning emulation solution for analog and mixed-signal (AMS) circuits. This paper presents an analog circuit emulation based on wave digital filters (WDFs), which covers the entire flow of transforming an AMS circuit from SPICE netlist to hardware implementation in FPGA. More specifically, it presents the theoretical support of how to map linear and nonlinear circuit components to WDF. The detail implementation of each WDF component in FPGA is not elaborated due to the page limit. Experiments show that there is a virtually perfect match between FPGA emulation and HSPICE simulations on two small but representative analog circuits, indicating high accuracy of the proposed emulation, and the FPGA-based WDF emulation can process analog signal sampled at as high as 512KHz, which is adequate for a variety of biomedical sensing applications.
Yen-Lung Chen, Chien-Nan Jimmy Liu, Jing-Yang Jou, Sudhakar Pamarti, Lei He 0001
ISCAS5
2015 A Cache Hierarchy Aware Thread Mapping Methodology for GPGPUs
abstract
The recently proposed GPGPU architecture has added a multi-level hierarchy of shared cache to better exploit the data locality of general purpose applications. The GPGPU design philosophy allocates most of the chip area to processing cores, and thus results in a relatively small cache shared by a large number of cores when compared with conventional multi-core CPUs. Applying a proper thread mapping scheme is crucial for gaining from constructive cache sharing and avoiding resource contention among thousands of threads. However, due to the significant differences on architectures and programming models, the existing thread mapping approaches for multi-core CPUs do not perform as effective on GPGPUs. This paper proposes a formal model to capture both the characteristics of threads as well as the cache sharing behavior of multi-level shared cache. With appropriate proofs, the model forms a solid theoretical foundation beneath the proposed cache hierarchy aware thread mapping methodology for multi-level shared cache GPGPUs. The experiments reveal that the three-staged thread mapping methodology can successfully improve the data reuse on each cache level of GPGPUs and achieve an average of 2.3× to 4.3× runtime enhancement when compared with existing approaches.
Bo-Cheng Lai, Hsien-Kai Kuo, Jing-Yang Jou
IEEE Trans. Computers3
2015 Scalable Global Power Management Policy Based on Combinatorial Optimization for Multiprocessors
abstract
Multiprocessors have become the main architecture trend in modern systems due to the superior performance; nevertheless, the power consumption remains a critical challenge. Global power management (GPM) aims at dynamically finding the power state combination that satisfies the power budget constraint while maximizing the overall performance (or vice versa). Due to the increasing number of cores in a multiprocessor system, the scalability of GPM policies has become critical when searching satisfactory state combinations within acceptable time. This article proposes a highly scalable policy based on combinatorial optimization with theoretical proofs, whereas previous works take exhaustive search or heuristic methods. The proposed policy first applies an optimum algorithm to construct a state combination table in pseudo--polynomial time using dynamic programming. Then, the state combination is assigned to cores with minimum transition cost in linear time by mapping to the network flow problem. Simulation results show that the proposed policy achieves better system performance for any given power budget when compared to the state-of-the-art heuristic. Furthermore, the proposed policy demonstrates its prominent scalability with 125 times faster policy runtime for 512 cores.
Gung-Yu Pan, Jed Yang, Jing-Yang Jou, Bo-Cheng Lai
ACM Trans. Embed. Comput. Syst.3
2014 A read-write aware DRAM scheduling for power reduction in multi-core systems
abstract
The demand of high performance and low power has increased the importance of power efficiency in multi-core systems. In modern multi-core architectures, DRAM has dominated the power consumption and therefore reordering based DRAM scheduling has been intensively studied to reduce the power. However, the benefit of reordering is not fully explored by the previous studies. To further reduce the power, this paper proposes the read-write reordering and the read-write aware throttling. When compared to the existing work, the proposed techniques reduce 10% more DRAM power with less performance degradation.
Chih-Yen Lai, Gung-Yu Pan, Hsien-Kai Kuo, Jing-Yang Jou
ASP-DAC4
2014 A learning-on-cloud power management policy for smart devices
abstract
Energy consumption poses severe limitations for smart devices, urging the development of effective and efficient power management policies. State-of-the-art learning-based policies are autonomous and adaptive to the environment, but they are subject to costly computational overhead and lengthy convergence time. As smart devices are connected to Internet, this paper proposes the Learning-on-Cloud (LoC) policy to exploit cloud computing for power management. Sophisticated learning engines are offloaded from local devices to the cloud with minimal communication data, thus the runtime overhead is reduced. The learning data are shared between many devices with the same model, hence the convergence rate is raised. With one thousand devices connecting to the cloud, the LoC agent is able to converge within a few iterations; the energy saving is better than both of the greedy and the learning-based policies with less latency penalty. By implementing the LoC policy as an Android App, the measured overhead is only 0.01% of the system time.
Gung-Yu Pan, Bo-Cheng Lai, Sheng-Yen Chen, Jing-Yang Jou
ICCAD4
2014 Efficient Coverage-Driven Stimulus Generation Using Simultaneous SAT Solving, with Application to SystemVerilog
abstract
SystemVerilog provides powerful language constructs for verification, and one of them is the covergroup functional coverage model. This model is designed as a complement to assertion verification, that is, it has the advantage of defining cross-coverage over multiple coverage points. In this article, a coverage-driven verification (CDV) approach is formulated as a simultaneous Boolean satisfiability (SAT) problem that is based on covergroups. The coverage bins defined by the functional model are converted into Conjunction Normal Form (CNF) and then solved together by our proposed simultaneous SAT algorithm PLNSAT to generate stimuli for improving coverage. The basic PLNSAT algorithm is then extended in our second proposed algorithm GPLNSAT, which exploits additional information gleaned from the structure of SystemVerilog covergroups. Compared to generating stimuli separately, the simultaneous SAT approaches can share learned knowledge across each coverage target, thus reducing the overall solving time drastically. Experimental results on a UART circuit and the largest ITC benchmark circuits show that the proposed algorithms can achieve 10.8x speedup on average and outperform state-of-the-art techniques in most of the benchmarks.
An-Che Cheng, Chia-Chih Yen, Celina G. Val, Sam Bayless, Alan J. Hu, Iris Hui-Ru Jiang, Jing-Yang Jou
ACM Trans. Design Autom. Electr. Syst.7
2014 Reducing Contention in Shared Last-Level Cache for Throughput Processors
abstract
Deploying the Shared Last-Level Cache (SLLC) is an effective way to alleviate the memory bottleneck in modern throughput processors, such as GPGPUs. A commonly used scheduling policy of throughput processors is to render the maximum possible thread-level parallelism. However, this greedy policy usually causes serious cache contention on the SLLC and significantly degrades the system performance. It is therefore a critical performance factor that the thread scheduling of a throughput processor performs a careful trade-off between the thread-level parallelism and cache contention. This article characterizes and analyzes the performance impact of cache contention in the SLLC of throughput processors. Based on the analyses and findings of cache contention and its performance pitfalls, this article formally formulates the aggregate working-set-size-constrained thread scheduling problem that constrains the aggregate working-set size on concurrent threads. With a proof to be NP-hard, this article has integrated a series of algorithms to minimize the cache contention and enhance the overall system performance on GPGPUs. The simulation results on NVIDIA's Fermi architecture have shown that the proposed thread scheduling scheme achieves up to 61.6% execution time enhancement over a widely used thread clustering scheme. When compared to the state-of-the-art technique that exploits the data reuse of applications, the improvement on execution time can reach 47.4%. Notably, the execution time improvement of the proposed thread scheduling scheme is only 2.6% from an exhaustive searching scheme.
Hsien-Kai Kuo, Bo-Cheng Lai, Jing-Yang Jou
ACM Trans. Design Autom. Electr. Syst.3
2014 Scalable Power Management Using Multilevel Reinforcement Learning for Multiprocessors
abstract
Dynamic power management has become an imperative design factor to attain the energy efficiency in modern systems. Among various power management schemes, learning-based policies that are adaptive to different environments and applications have demonstrated superior performance to other approaches. However, they suffer the scalability problem for multiprocessors due to the increasing number of cores in a system. In this article, we propose a scalable and effective online policy called MultiLevel Reinforcement Learning (MLRL). By exploiting the hierarchical paradigm, the time complexity of MLRL is O ( n lg n ) for n cores and the convergence rate is greatly raised by compressing redundant searching space. Some advanced techniques, such as the function approximation and the action selection scheme, are included to enhance the generality and stability of the proposed policy. By simulating on the SPLASH-2 benchmarks, MLRL runs 53% faster and outperforms the state-of-the-art work with 13.6% energy saving and 2.7% latency penalty on average. The generality and the scalability of MLRL are also validated through extensive simulations.
Gung-Yu Pan, Jing-Yang Jou, Bo-Cheng Lai
ACM Trans. Design Autom. Electr. Syst.2
2013 Cache Capacity Aware Thread Scheduling for Irregular Memory Access on many-core GPGPUs
abstract
On-chip shared cache is effective to alleviate the memory bottleneck in modern many-core systems, such as GPGPUs. However, when scheduling numerous concurrent threads on a GPGPU, a cache capacity agnostic scheduling scheme could lead to severe cache contention among threads and thus significant performance degradation. Moreover, the diverse working sets in irregular applications make the cache contention issue an even more serious problem. As a result, taking cache capacity into account has become a critical scheduling issue of GPGPUs. This paper formulates a Cache Capacity Aware Thread Scheduling Problem to capture the impact of cache capacity as well as different architectural considerations. With a proof to be NP-hard, this paper has proposed two algorithms to perform the cache capacity aware thread scheduling. The simulation results on Nvidia's Fermi configuration have shown that the proposed scheduling scheme can effectively avoid cache contention, and achieve an average of 44.7% cache miss reduction and 28.5% runtime enhancement. The paper also shows the runtime can be enhanced up to 62.5% for more complex applications.
Hsien-Kai Kuo, Ta-Kan Yen, Bo-Cheng Lai, Jing-Yang Jou
ASP-DAC4
2012 Thread affinity mapping for irregular data access on shared Cache GPGPU
abstract
Memory Coalescing and on-chip shared Cache are two effective techniques to alleviate the memory bottleneck in modern GPGPUs. These two techniques are very useful on applications with regular memory accesses. However, they become ineffective on concurrent threads with large numbers of uncoordinated accesses and the potential performance benefit could be significantly degraded. This paper proposes a thread affinity mapping methodology to coordinate the irregular data accesses on shared cache GPGPUs. Based on the proposed affinity metrics, threads are congregated into execution groups which are able to fully exploit the memory coalescing and data sharing within an application. An average of 3.5x runtime speedup is achieved on a Fermi GPGPU. The speedup scales with the sizes of test cases, which makes the proposed methodology an effective and promising solution for the continually increasing complexities of applications in the future many-core systems.
Hsien-Kai Kuo, Kuan-Ting Chen, Bo-Cheng Lai, Jing-Yang Jou
ASP-DAC4
2011 Equivalence checking of scheduling with speculative code transformations in high-level synthesis
abstract
This paper presents a formal method for equivalence checking between the descriptions before and after scheduling in high-level synthesis (HLS). Both descriptions are represented by finite state machine with datapaths (FSMDs) and are then characterized through finite sets of paths. The main target of our proposed method is to verify scheduling employing code transformations-such as speculation and common subexpression extraction (CSE), across basic block (BB) boundaries-which have not been properly addressed in the past. Nevertheless, our method can verify typical BB-based and path-based scheduling as well. The experimental results demonstrate that the proposed method can indeed outperform an existing state-of-the-art equivalence checking algorithm.
Chi-Hui Lee, Che-Hua Shih, Juinn-Dar Huang, Jing-Yang Jou
ASP-DAC4
2011 Design-for-debug layout adjustment for FIB probing and circuit editing
abstract
While the technology node continually and aggressively scales, the resolution of FIB techniques does not scale as fast. Thus, the percentage of nets which can be observed or repaired through FIB probing or circuit editing is significantly decreased for advanced process technologies, which limits the candidates that can be physically examined through the FIB techniques during the debugging process. This paper introduces a design-for-debug framework which can adjust the layout to increase the FIB observable rate and the FIB repairable rate for its signals. The layout adjustment is made through pre-defined simple operations subject to the design rules and the timing constraints. Hence, the proposed framework does not require a complicated router as its core and can be applied in conjunction with any commercial APR tool. The experimental result based on an 90 nm technology has demonstrated that the proposed DFD framework can effectively increase the FIB observable and repairable rates under different parameter settings while the overall area and circuit performance remain the same.
Kuo-An Chen, Tsung-Wei Chang, Meng-Chen Wu, Mango Chia-Tso Chao, Jing-Yang Jou, Sonair Chen
ITC5
2009 Multiple-Fault Diagnosis Using Faulty-Region Identification
abstract
The fault diagnosis has become an increasing portion of todaypsilas IC-design cycle and significantly determines productpsilas time-to-market. However, the failure behaviors from the defective chips may not be fully represented by the single fault model. In this paper, we propose a fault-diagnosis framework targeting multiple stuck-at faults. This framework first reports a minimal suspect region, in which all real faults are topologically covered. Next, a proposed ranking method is applied to sieve out the real faults from the candidates within the suspect region. The experimental results show that the proposed diagnosis framework can effectively locate the multiple stuck-at faults within a neighborhood, which may generate erroneous signals cancelling one another and are difficult to be diagnosed based on a single-fault-model method.
Meng-Jai Tasi, Mango Chia-Tso Chao, Jing-Yang Jou, Meng-Chen Wu
VTS3
2009 Accurate Rank Ordering of Error Candidates for Efficient HDL Design Debugging
abstract
When hardware description languages (HDLs) are used in describing the behavior of a digital circuit, design errors (or bugs) almost inevitably appear in the HDL code of the circuit. Existing approaches attempt to reduce efforts involved in this debugging process by extracting a reduced set of error candidates. However, the derived set can still contain many error candidates, and finding true design errors among the candidates in the set may still consume much valuable time. Adebuggingprioritymethod was proposed to speed up the error-searching process in the derived error candidate set. The idea is to display error candidates in an order that corresponds to an individual's degree of suspicion. With this method, error candidates are placed in a rank order based on their probability of being an error. The more likely an error candidate is a design error (or a bug), the higher the rank order that it has. With the displayed rank order, circuit designers should find design errors quicker than with blind searching when searching for design errors among all the derived candidates. However, the currently used confidence score (CS) for deriving thedebuggingpriorityhas some flaws in estimating the likelihood of correctness of error candidates due to themaskingerrorsituation. This reduces the degree of accuracy in establishing adebuggingpriority. Therefore, the objective of this work is to develop a new probabilistic confidence score (PCS) that takes themaskingerrorsituation into consideration in order to provide a more reliable and accuratedebuggingpriority. The experimental results show that our proposed PCS achieves better results in estimating the likelihood of correctness and can indeed suggest adebuggingprioritywith better accuracy, as compared to the CS.
Tai-Ying Jiang, Chien-Nan Jimmy Liu, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Performance-constrained voltage assignment in multiple supply voltage SoC floorplanning
abstract
Using voltage island methodology to reduce power consumption for System-on-a-Chip (SoC) designs has become more and more popular recently. Currently this approach has been considered either in system-level architecture or postplacement stage. Since hierarchical design and reusable intellectual property (IP) are widely used, it is necessary to optimize floorplanning/placement methodology considering voltage islands generation to solve power and critical path delay problems. In this article, we propose a floorplanning methodology considering voltage islands generation and performance constraints. Our method is flexible and can be extended to hierarchical design. The experimental results on some MCNC benchmarks show that our method is effective in meeting performance constraints and can simultaneously consider the tradeoff between power routing cost and total power dissipation.
Meng-Chen Wu, Ming-Ching Lu, Hung-Ming Chen, Jing-Yang Jou
ACM Trans. Design Autom. Electr. Syst.4
2009 Automatic Verification Stimulus Generation for Interface Protocols Modeled With Non-Deterministic Extended FSM
abstract
Verifying if an integrated component is compliant with certain interface protocol is a vital issue in component-based system-on-a-chip (SoC) designs. For simulation-based verification, generating massive constrained simulation stimuli is becoming crucial to achieve a high verification quality. To further improve the quality, stimulus biasing techniques are often used to guide the simulation to hit design corners. In this paper, we model the interface protocol with the non-deterministic extended finite-state machine (NEFSM), and then propose an automatic stimulus generation approach based on it. This approach is capable of providing numerous biasing strategies. Experiment results demonstrate the high controllability and efficiency of our stimulus generation scheme.
Che-Hua Shih, Juinn-Dar Huang, Jing-Yang Jou
IEEE Trans. Very Large Scale Integr. Syst.3
2007 A Precise Bandwidth Control Arbitration Algorithm for Hard Real-Time SoC Buses
abstract
On an SoC bus, contentions occur while different IP cores request the bus access at the same time. Hence an arbiter is mandatory to deal with the contention issue on a shared bus system. In different applications, IPs may have real-time and/or bandwidth requirements. It is very difficult to design an arbitration algorithm to simultaneously meet these two requirements. In this paper, we propose an innovative arbitration algorithm, RB_lottery, to meet both of the requirements. It can provide not only the hard real-time guarantee but also the precise bandwidth controllability. The experimental results show that RBJottery outperforms several well-known existing arbitration algorithms.
Bu-Ching Lin, Geeng-Wei Lee, Juinn-Dar Huang, Jing-Yang Jou
ASP-DAC4
2007 Hybrid Wordlength Optimization Methods of Pipelined FFT Processors
abstract
Quickly and accurately predicting the performance based on the requirements for IP-based system implementations optimizes the design and reduces the design time and overall cost. This study describes a novel hybrid method for the word-length optimization of pipelined FFT processors that is the arithmetic kernel of OFDM-based systems. This methodology utilizes the rapid computing of statistical analysis and the accurate evaluation of simulation-based analysis to investigate a speedy optimization flow. A statistical error model for varying word-lengths of PE stages of an FFT processor was developed to support this optimization flow. Experimental results designate that the word-length optimization employing the speedy flow reduces the percentage of the total area of the FFT processor that increases with an increasing FFT length. Finally, the proposed hybrid method requires a shorter prediction time than the absolute simulation-based method does and achieves more accurate outcomes than a statistical calculation does.
Cheng-Yeh Wang, Chih-Bin Kuo, Jing-Yang Jou
IEEE Trans. Computers3
2007 Observability Analysis on HDL Descriptions for Effective Functional Validation
abstract
Simulation-based functional validation is still one of the primary approaches for verifying designs described in hardware description languages. Traditional code coverage metrics do not address the observability issue and may overestimate the extent of functional validation. Observability-based code coverage metric (OCCOM) is the first code coverage metric considering the essential observability issue. However, tags can only be observed or unobserved, providing only two levels of measurement (i.e., 1 and 0). Errors with lower opportunities to be observed may still be judged as observable, thus misleading the verification results. Therefore, instead of extending tag coverage, we develop a probabilistic observability measure and its efficient computation algorithm. Besides being used as a new OCCOM, our new measure can point out hard-to-observe points for inserting assertions to prevent bugs from hiding behind these points. Experimental results show that the detection of the injected errors and the degree of our observability measure are strongly related. The results also show that our fine-grained observability measure is less likely to overestimate the extent of validation with reasonable computation time.
Tai-Ying Jiang, Chien-Nan Jimmy Liu, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 A real-time and bandwidth guaranteed arbitration algorithm for SoC bus communication
abstract
In shared SoC bus systems, arbiters are usually adopted to solve bus contentions with various kinds of arbitration algorithms. We propose an arbitration algorithm, RT/spl I.bar/lottery, which is designed to meet both hard real-time and bandwidth requirements. For fast evaluation and exploration, we use high abstract-level models in our system simulation environment to generate parameters for our configurable arbiter. The experimental results show that RT/spl I.bar/lottery can meet all hard real-time requirements and perform very well in bandwidth allocation. The results also show that RT/spl I.bar/lottery outperforms several commonly-used arbitration algorithms today.
Chien-Hua Chen, Geeng-Wei Lee, Juinn-Dar Huang, Jing-Yang Jou
ASP-DAC4
2006 FSM-based transaction-level functional coverage for interface compliance verification
abstract
Interface compliance verification plays a very important role in modern SoC designs. In order to perform a quantitative analysis of simulation completeness, adequate coverage metrics are mandatory. In this paper, we propose a finite state machine (FSM) based transaction-level functional coverage methodology for interface compliance verification. A language, state-oriented language (SOL), is developed to specify functional transactions mainly at the higher FSM level instead of lower logic or signal level. By utilizing SOL, it is simple and rigorous to specify interesting transactions from the specification FSM of the target interface protocol. Experimental results show that the proposed methodology can effectively improve the verification quality as well as increase the efficiency of regression verification.
Man-Yun Su, Che-Hua Shih, Juinn-Dar Huang, Jing-Yang Jou
ASP-DAC4
2006 An Optimum Algorithm for Compacting Error Traces for Efficient Design Error Debugging
abstract
Diagnosing counterexamples with error traces has acted as one of the most critical steps in functional verification. Unfortunately, error traces are normally very lengthy such that designers need to spend considerable effort to understand them. To alleviate the designers' burden for debugging, we present a SAT-based algorithm for reducing the lengths of error traces. The algorithm performs the paradigm of the binary search algorithm to halve the search space recursively. Furthermore, it applies a novel theorem to guarantee gaining the shortest lengths for the error traces. Based on the optimum algorithm, we develop two robust heuristics to handle real designs. Experimental results demonstrate that our approaches greatly surpass previous work and, indeed, have promising solutions.
Chia-Chih Yen, Jing-Yang Jou
IEEE Trans. Computers2
2006 RLC Coupling-Aware Simulation and On-Chip Bus Encoding for Delay Reduction
abstract
This paper shows that the worst case switching pattern that incurs the longest bus delay while considering the RLC effect is quite different from that while considering the RC effect alone. It implies that the existing encoding schemes based on the RC model may not improve or possibly worsen the delay when the inductance effects become dominant. A bus-invert method is also proposed to reduce the on-chip bus delay based on the RLC model. Simulation results show that the proposed encoding scheme significantly reduces the worst case coupling delay of the inductance-dominated buses
Shang-Wei Tu, Yao-Wen Chang, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 Reliable crosstalk-driven interconnect optimization
abstract
As technology advances apace, crosstalk becomes a design metric of comparable importance to area and delay. This article focuses mainly on the crosstalk issue, specifically on the impacts of physical design and process variation on crosstalk. While the feature size shrinks below 0.25μ m , the impact of process variation on crosstalk increases rapidly. Hence, a crosstalk insensitive design is desirable in the deep submicron regime. In this article, crosstalk sensitivity is referred to as the influence of process variation on crosstalk in a circuit. We show that the lower bound of crosstalk sensitivity grows quadratically, while that of crosstalk increases linearly. Therefore, designers should also consider crosstalk sensitivity, when optimizing other design objectives such as crosstalk, area, and delay. According to our modeling, these objectives are all in posynomial forms, and thus the multi-objective optimization problem can optimally be solved by Lagrangian relaxation. Experimental results show that our method is effective and efficient. For instance, a circuit of 2856 gates and 5272 wires is optimized using 13-minute runtime and 2.8-MB memory on a Pentium III 1.0 GHz PC with 256-MB memory. In particular, by relaxing Lagrange multipliers to the critical paths, it takes only two iterations for all solutions to converge to the global optimal, which is much more efficient than related previous work. This relaxation scheme provides a key insight into the rapid convergence in Lagrangian relaxation.
Iris Hui-Ru Jiang, Song-Ra Pan, Yao-Wen Chang, Jing-Yang Jou
ACM Trans. Design Autom. Electr. Syst.4
2005 An observability measure to enhance statement coverage metric for proper evaluation of verification completeness
abstract
Simulation based validation approaches are still the primary workhorse for solving the verification problem of getting the initial HDL description correct, especially for large scaled designs. However, most of existing code coverage metrics do not address obsevability issue [2]. Therefore, we intend to provide additional observability measures to statement coverage metric for more proper and realistic evaluation of verification completeness for a HDL design. As compared to OCCOM [1,2,3], our approach estimates a real probabilistic likelihood of propagating erroneous effects without any unreasonable assumptions and can always provide lower bound estimation.
Tai-Ying Jiang, Chien-Nan Jimmy Liu, Jing-Yang Jou
ASP-DAC3
2005 Communication-driven task binding for multiprocessor with latency insensitive network-on-chip
abstract
Network-on-Chip is a new design paradigm for designing core based System-on-Chip. It features high degree of reusability and scalability. In this paper, we propose a switch which employs the latency insensitive concepts and applies the round-robin scheduling techniques to achieve high communication resource utilization. Based on the assumptions of the 2D-mesh network topology constructed by the switch, this work not only models the communication and the contention effect of the network, but develops a communication-driven task binding algorithm that employs the divide and conquer strategy to map applications onto the multiprocessor system-on-chip. The algorithm attempts to derive a binding of tasks such that the overall system throughput is maximized. To compare with the task binding without consideration of communication and contention effect, the experimental results demonstrate that the overall improvement of the system throughput is 20% for 844 test cases.
Liang-Yu Lin, Cheng-Yeh Wang, Pao-Jui Huang, Chih-Chieh Chou, Jing-Yang Jou
ASP-DAC5
2005 An efficient heterogeneous tree multiplexer synthesis technique
abstract
In this paper, a novel strategy for designing the heterogeneous tree multiplexer is proposed. The authors build the multiplexer delay model by curve fitting and then formulate the heterogeneous tree multiplexer design problem as a special type of optimization problem called mixed-integer nonlinear programming (MINLP). A new design parameter, the switch size in each stage, is introduced to improve the speed of the heterogeneous tree multiplexer. The proposed strategy can determine the multiplexer architecture and the switch size in each stage simultaneously. Three optimization methods are provided to synthesize the heterogeneous tree multiplexer according to the design specifications.
Hsu-Wei Huang, Cheng-Yeh Wang, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 Optimal design of high fan-in multiplexers via mixed-integer nonlinear programming
Hsu-Wei Huang, Cheng-Yeh Wang, Jing-Yang Jou
ASP-DAC3
2004 On compliance test of on-chip bus for SOC
Hue-Min Lin, Chia-Chih Yen, Che-Hua Shih, Jing-Yang Jou
ASP-DAC4
2004 Layout techniques for on-chip interconnect inductance reduction
Shang-Wei Tu, Jing-Yang Jou, Yao-Wen Chang
ASP-DAC2
2004 Graph Automorphism-Based Algorithm for Determining Symmetric Inputs
abstract
We propose a graph automorphism-based algorithm for computing maximal sets of symmetric inputs of circuits. It can be used to identify nonsymmetric inputs in a circuit and enhance the efficiency of input matching, library binding, as well as logic verification problems. We conduct the experiments on some benchmarks. The experimental results demonstrate that our approach distinguishes more non-symmetric inputs than that of previous work.
Chen-Ling Chou, Chun-Yao Wang, Geeng-Wei Lee, Jing-Yang Jou
ICCD4
2004 Verification on Port Connections
abstract
In a system-on-a-chip (SOC) design, several to hundreds of design blocks or intellectual properties (IPs) are integrated to form a complex function. Prior to verify the functionality of the integrated IPs, it is very important to ensure the correctness of the port connections among these IPs. This work addresses the problem of verification on port connections while IPs are integrated into a larger block or a system, and presents a new connection model and the corresponding error model for port connections. An algorithm providing the minimum pattern set and a general verification flow used to verify port connections are also proposed.
Geeng-Wei Lee, Juinn-Dar Huang, Jing-Yang Jou, Chun-Yao Wang
ITC3
2004 Simultaneous floor plan and buffer-block optimization
abstract
As technology advances and the number of interconnections among modules rapidly increases, timing closure, and design convergence are the most important concerns. Hence, it is desirable to consider interconnect optimization as early as possible. Previous work for this issue can be classified into two directions: wire planning and buffer-block planning for interconnect-driven floorplanning. Wire planning for interconnect-driven floorplanning does not consider buffer insertion, and buffer-block planning for interconnect-driven floorplanning cannot overcome the limitation of a bad initial floorplan. In this paper, we first address simultaneous floorplanning and buffer-block planning (i.e., integrating buffer-block planning into floorplanning) for interconnect optimization. We adopt simulated annealing to refine a floorplan so that buffers can be inserted more effectively. In each iteration, we construct a routing tree for each net, allocate buffers for all nets, introduce corresponding buffer blocks into the intermediate floorplan, and invoke Lagrangian relaxation to optimize area and satisfy timing requirements. Further, in order to reduce the problem size, we present supermodule partitioning which partitions modules into supermodules. Experimental results show that our method of integrating buffer-block planning into floorplanning can significantly improve the interconnect delay and reduce the number of buffers needed. Based on a set of MCNC benchmark circuits, our approach achieves an average success rate of 86.1% of nets meeting timing constraints, inserts only 272 buffers on average, and consumes an average extra area of only 0.28% over the given floorplan, compared with the average success rate of 62.6%, 1123 buffers, and extra area of 1.05% resulted from a famous recent work presented at ICCAD'99.
Iris Hui-Ru Jiang, Yao-Wen Chang, Jing-Yang Jou, Kai-Yuan Chao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2003 An efficient IP-level power model for complex digital circuits
abstract
In this paper, we propose an efficient IP-Level power model with a small lookup table for complex CMOS circuits. The table has only one dimension that maps the zero-delay charging and discharging capacitance into the real power consumption of pattern pairs but still has high accuracy. In order to improve the efficiency of the characterization process, the Monte Carlo approach is used during the estimation of the average power to skip the samples that will not increase the accuracy too much. The experimental result shows the table sizes are only up to 107 entries for ISCAS'85 benchmark circuits and the estimation error is only 2.99% on average using the lookup table.
Chih-Yang Hsu, Chien-Nan Jimmy Liu, Jing-Yang Jou
ASP-DAC3
2003 Simultaneous floorplanning and buffer block planning
abstract
As technology advances and the number of interconnections among modules rapidly increases, timing closure and design convergence are the most important concerns. Hence, it is desirable to consider interconnect optimization as early as possible. In this paper, we first address simultaneous floorplanning and buffer block planning (i.e., integrating buffer block planning into floorplanning) for interconnect optimization. Experimental results show that our method can significantly improve the interconnect delay and reduce the number of buffers needed.
Iris Hui-Ru Jiang, Yao-Wen Chang, Jing-Yang Jou, Kai-Yuan Chao
ASP-DAC3
2003 An automatic interconnection rectification technique for SoC design integration
abstract
This paper presents an automatic interconnection rectification (AIR) technique to correct the misplaced interconnection occurred in the integration of a SoC design automatically. The experimental results show that the AIR can correct the misplaced interconnection and therefore accelerates the integration verification of a SoC design.
Chun-Yao Wang, Shing-Wu Tung, Jing-Yang Jou
ASP-DAC3
2003 Automatic interconnection rectification for SoC design verification based on the port order fault model
abstract
Embedded cores are being increasingly used in large system-on-a-chip (SoC) designs. The high complexity of SoC designs lead the design verification to be a challenge for system integrators. This paper presents an automatic interconnection rectification (AIR) technique based on the port order fault model to detect, diagnose, and correct the misplacements of interconnection that occurred in the integration of a SoC design automatically. The experiments are conducted on combinational and sequential benchmarks. Experimental results show that the AIR can correct the misplaced interconnection exactly within reasonable efforts and, therefore, accelerates the integration verification of SoC designs.
Chun-Yao Wang, Shing-Wu Tung, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2002 Effective Error Diagnosis for RTL Designs in HDLs
abstract
We propose an effective approach to diagnose multiple design errors in HDL designs with only one erroneous test case. Error candidates will be greatly reduced while ensuring that true erroneous statements are included in. The probability of correctness for each potential erroneous statement will be estimated such that the most suspected statements are reported first. Experiments show that the size of error candidates is indeed small and the estimation for the probability of correctness for potential error candidates is accurate.
Tai-Ying Jiang, Chien-Nan Jimmy Liu, Jing-Yang Jou
Asian Test Symposium3
2002 On automatic-verification pattern generation for SoC withport-order fault model
abstract
Embedded cores are being increasingly used in the design of large system-on-a-chip (SoC). Because of the high complexity of SoC, the design verification is a challenge for system integrators. To reduce the verification complexity, the port-order fault (POF) model has been used for verifying core-based designs (Tang and Jou, 1998). In this paper, we present an automatic-verification pattern generation (AVPG) for SoC design verification based on the POF model and perform experiments on combinational and sequential benchmarks. Experimental results show that our AVPG can efficiently generate verification patterns with high POF coverage.
Chun-Yao Wang, Shing-Wu Tung, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2002 An automorphic approach to verification pattern generation for SoC design verification using port-order fault model
abstract
Embedded cores are being increasingly used in the design of large system-on-a-chip (SoC). Because of the high complexity of SoC, the design verification is a challenge for system integrators. To reduce the verification complexity, the port-order fault (POF) model was proposed. It has been used for verifying core-based designs and the corresponding verification pattern generation has been developed. Here, the authors present an automorphic technique to improve the efficiency of the automatic verification pattern generation (AVPG) for SoC design verification based on the POF model. On average, the size of pattern sets obtained on the ISCAS-85 and MCNC benchmarks are 45% smaller and the run time decreases 16% as compared with the previous results of AVPG.
Chun-Yao Wang, Shing-Wu Tung, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2001 An efficient design-for-verification technique for HDLs
abstract
Due to the high complexity of modern circuit designs, verification has become the major bottleneck of the entire design process. There is an emerging need for a practical solution to reduce the verification time. In manufacturing test, a well-known technique, design-for-testability, is often used to reduce the testing time. By inserting some extra circuits on the hard-to-test points, the testability can be improved and the testing time can be reduced. In this paper, we apply the similar idea to functional verification and propose an efficient design-for-verification (DFV) technique to help users reduce the verification time. The conditions for hard-to-control (HTC) codes in a HDL design are clearly defined, and an efficient algorithm to detect them automatically is proposed. Besides the HTC detection, we also propose an algorithm that can eliminate those HTC points with minimum number of DFV points. By the help of those DFV points, the number of required test patterns to reach the same coverage can be greatly reduced especially for deep-sequential designs.
Chien-Nan Jimmy Liu, I-Ling Chen, Jing-Yang Jou
ASP-DAC3
2001 An Improved AVPG Algorithm for SoC Design Verification Using Port Order Fault Model
abstract
Embedded cores are being increasingly used in the design of large system-on-a-chip (SoC). Because of the high complexity of SoC the design verification is a challenge for the system integrator. To reduce the verification complexity, the port order fault (POF) model has been used for verifying core-based designs and the corresponding verification pattern generation has been developed. Here we present an automorphic technique to improve the efficiency of the automatic verification pattern generation (AVPG) for SoC design verification based on POF model. On average, the size, of pattern sets obtained on the ISCAS-85 and MCNC benchmarks are 45% smaller and the run time decreases 16% as compared with the results of AVPG.
Chun-Yao Wang, Shing-Wu Tung, Jing-Yang Jou
Asian Test Symposium3
2001 Converter-free multiple-voltage scaling techniques for low-powerCMOS digital design
abstract
Recent research has shown that voltage scaling is a very effective technique for low-power design. This paper describes a voltage scaling technique to minimize the power consumption of a combinational circuit. First, the converter-free multiple-voltage (CFMV) structures are proposed, including the p-type, the n-type, and the two-way CFMV structures. The CFMV structures make use of multiple supply voltages and do not require level converters. In contrast, previous works employing multiple supply voltages need level converters to prevent static currents, which may result in large power consumption. In addition, the CFMV structures group the gates with the same supply voltage in a cluster to reduce the complexity of placement and routing for the subsequent physical layout stage. Next, we formulated the problem and proposed an efficient heuristic algorithm to solve it. The heuristic algorithm has been implemented in C and experiments were performed on the ISCAS85 circuits to demonstrate the effectiveness of our approach.
Yi-Jong Yeh, Sy-Yen Kuo, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2001 Unified functional decomposition via encoding for FPGA technology mapping
abstract
Functional decomposition has recently been adopted for look-up table (LUT)-based field-programmable gate array (FPGA) technology mapping with good results. In this paper we propose a novel method to unify functional single-output and multiple-output decomposition. We first address a compatible class encoding algorithm to minimize the number of compatible classes in the image function. After applying the encoding algorithm, we can therefore improve the decomposability in the subsequent decomposition of the image function. The above encoding algorithm is then extended to encode multiple-output functions through the construction of a hyperfunction. Common subexpressions among these multiple-output functions can be extracted during the decomposition of the hyperfunction. Consequently, we can handle multiple-output decomposition in the same manner as single-output decomposition. Experimental results show that our algorithms are promising.
Jie-Hong Roland Jiang, Jing-Yang Jou, Juinn-Dar Huang
IEEE Trans. Very Large Scale Integr. Syst.2
2000 A new method for constructing IP level power model based on power sensitivity
Heng-Liang Huang, Jiing-Yuan Lin, Wen-Zen Shen, Jing-Yang Jou
ASP-DAC4
2000 Collaboration between Industry and Academia in Test Research
Kwang-Ting Cheng, Vishwani D. Agrawal, Jing-Yang Jou, Li-C. Wang, Chi-Feng Wu, Shianling Wu
Asian Test Symposium3
2000 A novel approach for functional coverage measurement in HDL
abstract
While the coverage-driven functional verification is getting popular, a fast and convenient coverage measurement tool is necessary. In this paper, we propose a novel approach for functional coverage measurement based on the VCD files produced by the simulators. The usage flow of the proposed dumpfile-based coverage analysis is much easier and smoother than that of existing instrumentation-based coverage tools. No pre-processing tool is required and no extra code will be inserted into the source code. Most importantly, the flexibility in choosing coverage metrics and measured code regions is increased. Only one simulation run is needed for any kind of coverage reports. By conducting some experiments on real examples, it shows very promising results in terms of the performance and the accuracy of coverage reports.
Chien-Nan Jimmy Liu, Chen-Yi Chang, Jing-Yang Jou, Ming-Chih Lai, Hsing-Ming Juan
ISCAS3
2000 Optimal reliable crosstalk-driven interconnect optimization
abstract
Article Optimal reliable crosstalk-driven interconnect optimization Share on Authors: Iris Hui-Ru Jiang Department of Electronics Engineering, National Chiao Tung University, Hsinchu 30010, Taiwan Department of Electronics Engineering, National Chiao Tung University, Hsinchu 30010, TaiwanView Profile , Song-Ra Pan Department of Computer and Information Science, National Chiao Tung University, Hsinchu 30010, Taiwan Department of Computer and Information Science, National Chiao Tung University, Hsinchu 30010, TaiwanView Profile , Yao-Wen Chang Department of Computer and Information Science, National Chiao Tung University, Hsinchu 30010, Taiwan Department of Computer and Information Science, National Chiao Tung University, Hsinchu 30010, TaiwanView Profile , Jing-Yang Jou Department of Electronics Engineering, National Chiao Tung University, Hsinchu 30010, Taiwan Department of Electronics Engineering, National Chiao Tung University, Hsinchu 30010, TaiwanView Profile Authors Info & Claims ISPD '00: Proceedings of the 2000 international symposium on Physical designMay 2000 Pages 128–133https://doi.org/10.1145/332357.332388Online:01 May 2000Publication History 6citation191DownloadsMetricsTotal Citations6Total Downloads191Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Iris Hui-Ru Jiang, Song-Ra Pan, Yao-Wen Chang, Jing-Yang Jou
ISPD4
2000 Crosstalk-driven interconnect optimization by simultaneous gate andwire sizing
abstract
Noise, as well as area, delay, and power, is one of the most important concerns in the design of deep submicrometer integrated circuits. Currently existing algorithms do not handle simultaneous switching conditions of signals for noise minimization. In this paper, we model not only physical coupling capacitance, but also simultaneous switching behavior for noise optimization. Based on Lagrangian relaxation, we present an algorithm which can optimally solve the simultaneous noise, area, delay, and power optimization problem by sizing circuit components. Our algorithm, with linear memory requirement and linear runtime, is very effective and efficient. For example, for a circuit of 6144 wires and 3512 gates, our algorithm solves the simultaneous optimization problem using only 2.1-MB memory and 19.4-min runtime to achieve the precision of within 1% error on a SUN Spare Ultra-I workstation.
Iris Hui-Ru Jiang, Yao-Wen Chang, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2000 On computing the minimum feedback vertex set of a directed graph bycontraction operations
abstract
Finding the minimum feedback vertex set (MFVS) in a graph is an important problem for a variety of computer-aided design (CAD) applications and graph reduction plays an important role in solving this intractable problem. This paper is largely concerned with three new and powerful reduction operations. Each of these operations defines a new class of graphs, strictly larger than the class of contractible graphs [Levy and Low (1988)] in which the MFVS can be found in polynomial-time complexity. Based on these operations, an exact algorithm run on branch and bound manner is developed. This exact algorithm uses a good heuristic to find out an initial solution and a good bounding strategy to prune the solution space. To demonstrate the efficiency of our algorithms, we have implemented our algorithms and applied them to solving the partial scan problem in ISCAS'89 benchmarks. The experimental results show that if our three new contraction operations are applied, 27 out of 31 circuits in ISCAS'89 benchmarks can be fully reduced. Otherwise, only 12 out of 31 can be fully reduced. Furthermore, for all ISCAS'89 benchmarks our exact algorithm can find the exact cutsets in less than 3 s (CPU time) on a SUN-UltraII workstation. Therefore, the new contraction operations and our algorithms are demonstrated to be very effective in the partial scan application.
Hen-Ming Lin, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2000 ALTO: an iterative area/performance tradeoff algorithm for LUT-based FPGA technology mapping
abstract
In this paper, we propose an iterative area/performance tradeoff algorithm for look-up table (LUT)-based field programmable gate array (FPGA) technology mapping. First, it finds an area-optimized, performance-considered initial network by a modified area optimization technique. Then, an iterative algorithm consisting of several resynthesizing techniques is applied to trade the area for the performance in the network gracefully. Experimental results show that this approach can efficiently provide a complete set of mapping solutions from the area-optimized one to the performance-optimized one for the given design. Furthermore, these two extreme solutions produced by our algorithm outperform the results provided by most existing algorithms. Therefore, our algorithm is very useful for the timing-driven, LUT-based FPGA synthesis.
Juinn-Dar Huang, Jing-Yang Jou, Wen-Zen Shen
IEEE Trans. Very Large Scale Integr. Syst.2
1999 Hierarchical Floorplan Design on the Internet
abstract
With the proliferation of transistor count in VLSI design, more and more design groups try to figure out a way to efficiently combine their designs. The Internet features distributed computing and resource sharing. Consequently, a hierarchical floorplan design can be adequately solved in the Internet environment. In this paper, we address the problem of area minimization floorplan design in the Internet environment. We propose a novel algorithm, RMG algorithm. Taking advantage of the Internet, the RMG algorithm reduces the computing time by shortening the critical path in the floorplan tree. With creating floorplan design in the Internet environment, it can be seen that the Internet has advantages for electronic design automation.
Jiann-Horng Lin, Jing-Yang Jou, Iris Hui-Ru Jiang
ASP-DAC2
1999 Noise-Constrained Performance Optimization by Simultaneous Gate and Wire Sizing Based on Lagrangian Relaxation
abstract
Noise, as well as area, delay, and power, is one of the most important concerns in the design of deep sub-micron ICs. Currently existing algorithms do not handle simultaneous switching conditions of signals for noise minimization. In this paper, we model not only physical coupling capacitance, but also simultaneous switching behavior for noise optimization. Based on Lagrangian relaxation, we present an algorithm which can optimally solve the simultaneous noise, area, delay, and power optimization problem by sizing circuit components. Our algorithm, with linear memory requirement overall and linear runtime per iteration, is very effective and efficient. For example, for a circuit of 6144 wires and 3512 gates, our algorithm solves the simultaneous optimization problem using only 1.8 MB memory and 47 minute runtime to achieve the precision of within 1% error on a SUN Sparc Ultra-I workstation. 1 Introduction With decreasing feature sizes, higher clock rates, and increasing interconnect...
Iris Hui-Ru Jiang, Jing-Yang Jou, Yao-Wen Chang
DAC2
1999 Computing Minimum Feedback Vertex Sets by Contraction Operations and its Applications on CAD
abstract
Finding the minimum feedback vertex set (MFVS) in a graph is an important problem for a variety of CAD applications, and graph reduction plays an important role in solving this intractable problem. This paper is largely concerned with three new and powerful reduction operations. Each of these operations defines a new class of graphs which is strictly larger than the class of contractible graphs, in which the MFVS can be found with polynomial-time complexity. Based on these operations, an exact algorithm run in a branch-and-bound manner is developed. This exact algorithm uses a good heuristic to find an initial solution and a good bounding strategy to prune the solution space. We have implemented our algorithms and applied them to solving the partial scan problem in the ISCAS89 benchmarks. Experimental results show that, for all ISCAS89 benchmarks, our exact algorithm can find the exact cutsets in less than three seconds of CPU time on a Sun Ultra II workstation.
Hen-Ming Lin, Jing-Yang Jou
ICCD2
1999 An Efficient Functional Coverage Test for HDL Descriptions at RTL
abstract
Until now, simulation has been the primary approach for the functional verification of register transfer level (RTL) circuit descriptions written in a hardware description language (HDL). A finite state machine (FSM) coverage test can find all the bugs in a FSM design. However, this is impractical for large designs because of the state explosion problem. In this paper, we modify the higher-level FSM models used in other applications to replace the FSM model in the FSM coverage test. The state transition graphs (STGs) can be significantly reduced in this model, so that the complexity of the test becomes acceptable even for large designs. This model can be easily extracted from the original HDL code automatically, with little computation overhead. Experimental results show that it is indeed a promising functional test for FSMs.
Chien-Nan Jimmy Liu, Jing-Yang Jou
ICCD2
1999 Two-level logic minimization for low power
abstract
In this paper we present a complete Boolean method for reducing the power consumption in two-level combinational circuits. The two-level logic optimizer performs the logic minimization for low power targeting static PLA, general logic gates, and dynamic PLA implementations. We modify the espresso algorithm by adding our heuristics, which bias logic minimization toward lowering power dissipation. In our heuristics, signal probabilities and transition densities are two important parameters. The experimental results are promising.
Jyh-Mou Tseng, Jing-Yang Jou
ACM Trans. Design Autom. Electr. Syst.2
1999 A structure-oriented power modeling technique for macrocells
abstract
To characterize the power consumption of a macrocell, a general method involves recording the power consumption of all possible input transition events in the look-up tables. However, though this approach is accurate, the size of the table becomes very large. In this paper, we propose a new power modeling technique that takes advantage of the structural information of a macrocell. In this approach, a subset of primary inputs and internal nodes in the macrocell are selected as the state variables to build a state transition graph (STG). These state variables can model the steady-state transitions completely. Moreover, by selecting the characterization patterns properly, the STG can also model the glitch power in the macrocell accurately. To further simplify the complexity of the STG, an incomplete power modeling technique is presented. Without losing much accuracy, the property of compatible patterns is exploited for a macrocell to further reduce the number of edges in the corresponding STG. Experimental results show that our modeling techniques can provide SPICE-like accuracy, while the size of the look-up table is significantly reduced.
Jiing-Yuan Lin, Wen-Zen Shen, Jing-Yang Jou
IEEE Trans. Very Large Scale Integr. Syst.3
1998 Verification Pattern Generation for Core-Based Design Using Port Order Fault Model
abstract
The lack of information about core's internal structure means designers must rely solely on the test set distributed by the core provider. Sometimes the stuck at fault (SAF) model and automatic test pattern generation (ATPG) are used to generate test vectors for those pre-defined blocks. However, a SAF test set could waste lots of time to verify the pre-verified internal structure of the cores. Therefore, in order to reduce the core-based design verification time, we should adopt the connectivity-based port order fault (POF) model instead of the stuck at fault model. In this paper, we compare the POF model with the SAF model and propose a method that the POF test set for functional verification can be generated by using the SAF-based ATPG tools with proper assignment of don't care terms in inputs.
Shing-Wu Tung, Jing-Yang Jou
Asian Test Symposium2
1998 Compatible Class Encoding in Hyper-Function Decomposition for FPGA Synthesis
abstract
Recently, functional decomposition has been adopted for LUT based FPGA technology mapping with good results. In this paper, we propose a novel method for functional multiple-output decomposition. We first address a compatible class encoding method to minimize the compatible classes in the image function. After the encoding algorithm is applied, the decomposability will be improved in the subsequent decomposition of the image function. The above encoding algorithm is then extended to encode multiple-output functions through the construction of a hyper-function. Common sub-expressions among these multiple-output functions can be extracted during the decomposition of the hyper-function. Therefore, we can handle the multiple-output decomposition in the same manner as the single-output decomposition. Experimental results show that our algorithms are very promising.
Jie-Hong Roland Jiang, Jing-Yang Jou, Juinn-Dar Huang
DAC2
1998 On circuit clustering for area/delay tradeoff under capacity and pin constraints
abstract
In this paper, we propose an iterative area/delay tradeoff algorithm to solve the circuit clustering problem under the capacity constraint. It first finds an initial delay-considered area-optimized clustering solution by a delay-oriented depth first-search procedure. Then, an iterative procedure consisting of several reclustering techniques is applied to gradually trade the area for the performance. We then show that this algorithm can be easily extended to solve the clustering problem subject to both capacity and pin constraints. Experimental results show that our algorithm can provide a complete set of clustering solutions from the area-optimized one to the delay-optimized one for a given circuit. Furthermore, compared to the existing delay-optimized algorithms, this algorithm achieves almost the same performance but with much less area overhead. Therefore, this algorithm is very useful for solving the timing-driven circuit clustering problem.
Juinn-Dar Huang, Jing-Yang Jou, Wen-Zen Shen, Hsien-Ho Chuang
IEEE Trans. Very Large Scale Integr. Syst.2
1997 BDD based lambda set selection in Roth-Karp decomposition for LUT architecture
abstract
Field Programmable Gate Arrays (FPGAs) are important devices for rapid system prototyping. Roth-Karp decomposition is one of the most popular decomposition techniques for Look-Up Table (LUT)-based FPGA technology mapping. In this paper, we propose a novel algorithm based on Binary Decision Diagrams (BDDs) for selecting good lambda set variables in Roth-Karp decomposition to minimize the number of consumed configurable logic blocks (CLBs) in FPGAs. The experimental results on a set of benchmarks show that our algorithm can produce much better results than those of the previous approach (Wen-Zen Shen et al., 1995).
Jie-Hong Roland Jiang, Jing-Yang Jou, Juinn-Dar Huang, Jung-Shian Wei
ASP-DAC2
1997 A power driven two-level logic optimizer
abstract
In this paper we present Boolean techniques for reducing the power consumption in two-level combinational circuits. The two-level logic optimizer performs the logic minimization for low power targeting static PLA, general logic gates and dynamic PLA implementations. We modify Espresso algorithm by adding our heuristics that bias the logic minimization toward lowering the power dissipation. In our heuristics, signal probabilities and transition densities are two important parameters. The experimental results are promising.
Jyh-Mou Tseng, Jing-Yang Jou
ASP-DAC2
1997 A power modeling and characterization method for macrocells using structure information
abstract
To characterize a macrocell, a general method is to store the power consumption of all possible transition events at primary inputs in the lookup tables. Though this approach is very accurate, the lookup tables could be huge for the macrocells with many inputs. We present a new power modeling method which takes advantage of the structure information of macrocells and selects a minimum number of primary inputs or internal nodes in a macrocell as state variables to build a state transition graph (STG). Those state variables can completely model the transitions of all internal nodes and the primary outputs. By carefully deleting some state variables, we further introduce an incomplete power modeling technique which can simplify the STG without losing much accuracy. In addition, we exploit the property of the compatible patterns of a macrocell to further reduce the number of edges in the corresponding STG. Experimental results show that our modeling techniques can provide SPICE-like accuracy and can reduce the size of the lookup table significantly, compared to the general approach.
Jiing-Yuan Lin, Wen-Zen Shen, Jing-Yang Jou
ICCAD3
1997 Power Driven Partial Scan
abstract
The power consumption and testability are two of major considerations in modern VLSI design. A full-scan method had been used widely in the past to improve the testability of sequential circuits. Due to the lower overheads incurred, the partial-scan design has gradually become popular. The authors propose a partial scan selection strategy that bases on the structural analysis approach and considers the area and power overheads simultaneously. A powerful sample-and-search algorithm is used to find the solution that minimizes the user-specified cost function in term of power and area overheads. The experimental results show that the sample-and-search algorithm can effectively find the best solution of the specified cost function for almost all circuits, and the saving of overheads on average for each specific cost function is significant.
Jing-Yang Jou, Ming-Chang Nien
ICCD1
1997 Gauss-elimination-based generation of multiple seed-polynomial pairs for LFSR
abstract
This paper presents a new and efficient strategy of pseudorandom pattern generation (PRPG) for IC testing. It uses a general programmable LFSR (P-LFSR) to offer multiple-seed and multiple-polynomial PRPG. The deterministic pattern set generated by an ATPG tool or supplied by the designers is used to guide the generation of pseudorandom patterns. A novel application of the Gauss-elimination procedure is proposed to find the seeds as well as the polynomials. With an intelligent heuristic to further utilize the essential faults, this approach becomes very efficient, even for the random pattern resistant (RPR) circuits. Experiments are conducted on the ISCAS-85 benchmarks and the full scan version of the ISCAS-89 benchmarks. For all benchmark circuits, complete fault coverage is achieved with good balance on the hardware overhead and the test lengths as compared to other schemes.
Liren Huang, Jing-Yang Jou, Sy-Yen Kuo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1996 An Efficient PRPG Strategy By Utilizing Essential Faults
abstract
One major drawback of the LFSR-based BIST is its low fault coverage. To obtain the complete fault coverage, multiple seeds and multiple polynomials are usually required. One way to find the seeds and polynomials for the LFSR was utilizing the Gauss-elimination procedure. In this approach, the test patterns which are generated by LFSR are modeled as a set of multivariable linear equations. It is created from a given deterministic test set. The corresponding seed and polynomial are then obtained from the solution of this equations set. However, given the original deterministic test set without don't cares, it were not acceptable on the random pattern resistant circuits. In this paper, we allow the test patterns to have don't care values. With an intelligent heuristic of further utilizing the essential faults, this approach becomes much more efficient even for the random pattern resistant circuits. The experimental results on the ISCAS-85 and the ISCAS-89 benchmarks show that a significant improvement can be obtained both on the hardware overhead and the test length.
Liren Huang, Jing-Yang Jou, Sy-Yen Kuo
Asian Test Symposium2
1996 Easily Testable Data Path Allocation Using Input/Output Registers
abstract
Most existing behavioral synthesis systems concentrate on area and performance optimization, while ignoring other design qualities such as testability. In this paper/sup /spl Dagger//, we present three algorithms for register, module, and interconnection allocation of behavioral synthesis respectively to improve testability in data path allocation without assuming any specific test strategy. By using primary input/output registers effectively, the proposed algorithms produce RTL designs with better testability, while incur low or even no hardware overhead. Four benchmarks are synthesized using the proposed approaches and the results are compared with the best results of similar works in the literature. It shows that our approaches give both higher fault coverage and lower hardware overhead.
Liren Huang, Jing-Yang Jou, Sy-Yen Kuo, Wen-Bin Liao
Asian Test Symposium2
1996 An iterative area/performance trade-off algorithm for LUT-based FPGA technology mapping
abstract
In this paper, we propose an iterative area/performance trade-off algorithm for LUT-based FPGA technology mapping. First, it finds an area-optimized performance-considered initial network by a modified area optimization technique. Then, an iterative algorithm consisting of several resynthesizing techniques is applied to trade the area for the performance in the network gracefully. Experimental results show that this approach can provide a complete set of mapping solutions from the area-optimized one to the performance-optimize one for the given design. Furthermore, these two extreme solutions, the area-optimized one and the performance-optimized one, produced by our algorithm outperform the results of most existing algorithms. Therefore, our algorithm is very useful for the timing driven FPGA synthesis.
Juinn-Dar Huang, Jing-Yang Jou, Wen-Zen Shen
ICCAD2
1996 A power modeling and characterization method for the CMOS standard cell library
abstract
In this paper, we propose power consumption models for complex gates and transmission gates, which are extended from the model of basic gates proposed in Lin et al., (1994). We also describe an accurate power characterization method for CMOS standard cell libraries which accounts for the effects of input slew rate, output loading, and logic state dependencies. The characterization methodology separates the power consumption of a cell into three components, e.g., capacitive feedthrough power, short-circuit power, and dynamic power. For each component, power equation is derived from SPICE simulation results where the netlist is extracted from cell's layout. Experimental results on a set of ISCAS'85 benchmark circuits show that the power estimation based on our power modeling and characterization provides within 7% error of SPICE simulation on average while the CPU time consumed is more than two orders of magnitude less.
Jiing-Yuan Lin, Wen-Zen Shen, Jing-Yang Jou
ICCAD3
1995 An effective BIST design for PLA
abstract
In this paper, we describe a new design of built-in self test for programmable logic arrays (PLAs). The idea is to use a simple deterministic test pattern generator to generate test patterns such that each cross point in the AND array can be evaluated one after another. The simplest multiple input signature register which uses X/sup Q/+1 as its characteristic polynomial is used to evaluate the test results, where Q is the number of outputs. The final signature can be further compressed into only ONE bit. Instead of determining the probability of fault detection only, in this design, the fault detection capability is analyzed using the stuck-at fault, and the contact fault models. It is shown that all these modeled faults can be detected. This design is shown to give a better trade-off between the cost and the performance of built-in self test designs for PLAs.
Jing-Yang Jou
Asian Test Symposium1
1995 Compatible class encoding in Roth-Karp decomposition for two-output LUT architecture
abstract
Roth-Karp decomposition is one of the most popular techniques for LUT-based FPGA technology mapping because it can decompose a node into a set of nodes with fewer numbers of fanins. In this paper, we show how to formulate the compatible class encoding problem in Roth-Karp decomposition as a symbolic-output encoding problem in order to exploit the feature of the two-output LUT architecture. Based on this formulation, we also develop an encoding algorithm to minimize the number of LUT's required to implement the logic circuit. Experimental results show that our encoding algorithm can produce promising results in the logic synthesis environment for the two-output LUT architecture.
Juinn-Dar Huang, Jing-Yang Jou, Wen-Zen Shen
ICCAD2
1992 A functional fault model for sequential machines
abstract
A fault model at the state transition level is proposed for finite state machines. In this model, a fault causes the destination state of a state transition to be faulty. Analysis shows that a test set that detects all single-state-transition (SST) faults will also detect most multiple-state-transition (MST) faults in practical finite state machines. The quality of the test set generated for SST faults is close to that of the sequences derived from the checking experiment. It is also shown that the upper bound of the length of the SST fault test is 2MN/sup 2/ for an N-state M-transition machine, while that of the checking sequence is exponential. An automatic test generation algorithm and a test generation system, FTG, based on the model show that the test set generated for SST faults achieves high single stuck-at-fault coverage as well as high transistor fault coverage for multilevel implementations of the machine.>
Kwang-Ting Cheng, Jing-Yang Jou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1991 Timing-Driven Partial Scan
abstract
A partial scan approach that aims to reduce both area overhead and performance degradation caused by test logic is presented. Given a target speed and an initial design that meets the target, the algorithm selects a minimum set of scan flip-flops, if they exist, that (1) will break all sequential cycles and (2) will not violate the performance requirement after the scan logic is added. If such a set does not exist, the algorithm will find a set of scan flip-flops in which (1) all sequential cycles are broken and (2) the total area increase caused by the scan logic and the subsequent performance optimization is minimized, For circuits synthesized by automatic synthesis tools, the authors suggest a novel design flow, which selects/inserts the partial scan logic after area optimization, but before performance optimization. For meeting both performance and testability requirements, this design flow produces designs with less area increase than the traditional design flow, which considers testability and adds test logic after performance optimization. Experimental results on the ISCA'89 sequential circuits are presented as well as comparisons between the proposed method and previous methods.>
Jing-Yang Jou, Kwang-Ting Cheng
ICCAD1
1990 A Single-State-Transition Fault Model for Sequential Machines
abstract
A fault model in the state transition level of finite state machines is studied. In this model, called a single-state-transition (SST) fault model, a fault causes a state transition to go to a wrong destination state while leaving its input/output label intact. An analysis is given to show that a test set that detects all SST faults will also detect most multiple-state-transition (MST) faults in practical finite state machines. It is shown that, for an N-state M-transaction machine, the length of the SST fault test set is upper-bounded by 2*M*N/sup 2/ while the length is exponential in terms of N for a checking experiment. Experimental results show that the test set generated for SST faults achieves not only a high single stuck-at fault coverage but also a high transistor fault coverage for a multilevel implementation of the machine.>
Kwang-Ting Cheng, Jing-Yang Jou
ICCAD2
1990 Functional test generation for finite state machines
abstract
A functional test generation method for finite-state machines is described. A functional fault model, called the single-transition fault model, on the state transition level is used. In this model, a fault causes a single transition to a wrong destination state. A fault-collapsing technique for this fault model is also described. For each state transition, a small subset of states is selected as the faulty destination states so that the number of modeled faults for test generation is minimized. On the basis of this fault model, the authors developed an automatic test generation algorithm and built a test generation system. The effectiveness of this method is shown by experimental results on a set of benchmark finite-state machines. A 100% stuck-at fault coverage is achieved by the proposed method for several machines, and a very high coverage (>97%) is also obtained for other machines. In comparison with a gate-level test generator STG3, the test generation time is speeded up by a factor of 100.>
Kwang-Ting Cheng, Jing-Yang Jou
ITC2
1989 OPAM: an efficient output phase assignment for multilevel logic minimization
abstract
When a multiple-output function (z/sub 1/, z/sub 2/, . . ., z/sub m/) of multilevel logic is realized by complex gates, the option often exists to realize either z/sub i/ or its complement for each output. An efficient output phase assignment for the multilevel logic minimization (OPAM) is presented. The results of this study show that the proposed algorithm further reduces the literal count of the optimized network obtained by MIS (a multilevel logic minimization system).>
Chin-Long Wey, Sin-Min Chang, Jing-Yang Jou
ICCD3
1988 BECOME: Behavior Level Circuit Synthesis Based on Structure Mapping
Ruey-Sing Wei, Steven G. Rothweiler, Jing-Yang Jou
DAC3
1988 A testable PLA design with low overhead and ease of test generation
abstract
The author presents a hybrid programmable-logic array (PLA) design-for-testability technique that requires negligible hardware overhead and still preserves the property of ease of test generation. The key idea is to further utilize the 'don't care' assignment by introducing the control of both true and complement bits of some inputs to meet the requirement of distance-2 test sets. This approach is applied to the BARNEW PLA, and results support the claim that the hardware overhead of this technique is negligible and the ease of test generation is preserved.>
Jing-Yang Jou
ICCD1
1988 Fault-Tolerant Algorithms and Architectures for Real Time Signal Processing
Jing-Yang Jou, Jacob A. Abraham
ICPP (1)1
1988 Fault-Tolerant FFT Networks
abstract
Two concurrent error detection (CED) schemes are proposed for N-point fast Fourier transform (FFT) networks that consists of log/sub 2/N stages with N/2 two-point butterfly modules for each stage. The method assumes that failures are confined to a single complex multiplier or adder or to one input or output set of lines. Such a fault model covers a broad class of faults. It is shown that only a small overhead ratio, O(2/log/sub 2/N) of hardware, is required for the networks to obtain fault-secure results in the first scheme. A novel data retry technique is used to locate the faulty modules. Large roundoff errors can be detected and treated in the same manner as functional errors. The retry technique can also distinguish between the roundoff errors and functional errors that are caused by some physical failures. In the second scheme, a time-redundancy method is used to achieve both error detection and location. It is sown that only negligible hardware overhead is required. However, the throughput is reduced to half that of the original system, without both error detection and location, because of the nature of time-redundancy methods.>
Jing-Yang Jou, Jacob A. Abraham
IEEE Trans. Computers1
1986 Fault-tolerant matrix arithmetic and signal processing on highly concurrent computing structures
abstract
Hardware for executing matrix arithmetic and signal processing algorithms at high speeds is in great demand in many real-time and scientific applications. With the advent of VLSI technology, large numbers of processing elements which cooperate with each other at high speed have become economically feasible. Since any functional error in a high-performance system may seriously jeopardize the operation of the system and its data integrity, some level of fault tolerance must be incorporated in order to ensure that the results of long computations are valid. Since the major computational requirements for many important real-time signal processing tasks can be reduced to a common set of basic matrix operations, the development of a unified fault-tolerant scheme for matrix operations can solve the problems of both reliable signal processing and reliable matrix operations. Earlier work proposed a low-cost checksum scheme for fault-tolerant matrix operations on multiple processor systems. However, this scheme can only correct errors in matrix multiplication; it can detect, but not correct, errors in matrix-vector multiplication, LU decomposition, matrix inversion, etc. In order to solve these problems with the checksum scheme, a very general matrix encoding scheme is proposed in this paper to achieve fault-tolerant matrix arithmetic and signal processing with linear arrays, which are believed to hold the most promise in VLSI computing structures for their flexibility, low cost, and applicability to most of the interesting algorithms. This proposed technique is, therefore, a very cost-effective encoding technique to achieve fault-tolerant matrix arithmetic and signal processing on highly concurrent VLSI computing structures.
Jing-Yang Jou, Jacob A. Abraham
Proc. IEEE1