EDBT 2026 Demo / reviewers in the wild / expert
Juinn-Dar Huang
dblp:43/453
· DBLP profile ↗
51ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0001-5961-7863ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 48 · 10 first-author · 7 since 2021Software engineering, systems software and programming languages · 4Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Information-aware optimization for enhanced feature retention in medical image segmentation
Yung-Han Chen, Juinn-Dar Huang |
Neurocomputing | 2 |
| 2026 | Expertfuse: A huffman tree-based gradual expert integration framework for MoE models
Yi-Zeng Fang, Juinn-Dar Huang |
Neural Networks | 2 |
| 2025 | An Accurate and Compact Design Integrating Seven Common Nonlinear Functions in Deep LearningabstractDeep learning models based on Transformer architectures are undergoing significant transformations as large language models (LLMs) become more widely adopted and computationally demanding. All these models rely on various nonlinear activation functions, such as Softmax, LayerNorm, and other nontrivial activation functions, whose hardware implementations typically consume substantial area and power, mainly due to the complexity of underlying operations, like inverse square root, division, and exponential functions. To address these challenges, this paper proposes a compact and highly shared architecture that can efficiently and accurately compute reciprocals and inverse square roots via algorithm/hardware co-design. Besides, we propose a novel two-staged method for exponential function computation. The above techniques lead to an efficient implementation of LayerNorm and Softmax. Furthermore, those techniques also enable accurate implementations of other common activation functions, such as Sigmoid, Swish, GELU, and Tanh, with a minor extra hardware cost. The proposed 7-in-1 integrated design merely takes 5450μm2using a TSMC 28nm process. Experimental results also demonstrate that the proposed design achieves much higher Softmax accuracy than existing approaches. Jye-En Wu, Ting-Wei Hu, Chih-Yao Liang, Juinn-Dar Huang |
ISCAS | 4 |
| 2024 | A Hardware-Friendly Alternative to Softmax Function and Its Efficient VLSI Implementation for Deep Learning ApplicationsabstractThe Softmax function holds an essential role in most machine learning algorithms. Conventional realization of Softmax necessitates computationally intensive exponential operations and divisions, thereby posing formidable challenges in developing low-cost hardware implementations. This paper presents a promising hardware-friendly alternative, Squaremax, which gets rid of complex exponential operations. The function definition is extremely simple and can thus be efficiently implemented in both software and hardware. Experimental results show that Squaremax consistently attains comparable or superior accuracy over several popular models. Besides, this paper also proposes an efficient hardware architecture design of Squaremax. It requires no functional units for exponential and logarithmic operations, and is even lookup table (LUT) free. It adopts a flexible 16-bit fixed-point Q format for I/O to better preserve the output precision, which leads to higher model accuracy. Moreover, it yields substantial improvements in speed, area, and power, as well as achieves remarkable area and power efficiency of 664 G/mm2and 1396 G/W in a 40nm process. Therefore, hardware-friendly Squaremax is a very promising alternative to complex Softmax in both software and hardware for deep learning applications, and the proposed hardware architecture design and efficient LUT-free implementation do achieve a notable improvement in speed, area, and power. Meng-Hsun Hsieh, Xuan-Hong Li, Yu-Hsiang Huang, Pei-Hsuan Kuo, Juinn-Dar Huang |
ISCAS | 5 |
| 2024 | Diabetic foot ulcers segmentation challenge report: Benchmark and analysisabstractMonitoring the healing progress of diabetic foot ulcers is a challenging process. Accurate segmentation of foot ulcers can help podiatrists to quantitatively measure the size of wound regions to assist prediction of healing status. The main challenge in this field is the lack of publicly available manual delineation, which can be time consuming and laborious. Recently, methods based on deep learning have shown excellent results in automatic segmentation of medical images, however, they require large-scale datasets for training, and there is limited consensus on which methods perform the best. The 2022 Diabetic Foot Ulcers segmentation challenge was held in conjunction with the 2022 International Conference on Medical Image Computing and Computer Assisted Intervention, which sought to address these issues and stimulate progress in this research domain. A training set of 2000 images exhibiting diabetic foot ulcers was released with corresponding segmentation ground truth masks. Of the 72 (approved) requests from 47 countries, 26 teams used this data to develop fully automated systems to predict the true segmentation masks on a test set of 2000 images, with the corresponding ground truth segmentation masks kept private. Predictions from participating teams were scored and ranked according to their average Dice similarity coefficient of the ground truth masks and prediction masks. The winning team achieved a Dice of 0.7287 for diabetic foot ulcer segmentation. This challenge has now entered a live leaderboard stage where it serves as a challenging benchmark for diabetic foot ulcer segmentation. Moi Hoon Yap, Bill Cassidy, Michal Byra, Ting-Yu Liao, Huahui Yi, Adrian Galdran, Yung-Han Chen, Raphael Brüngel, Sven Koitka, Christoph M. Friedrich, Yu-Wen Lo, Ching-Hui Yang, Kang Li 0004, Qicheng Lao, Miguel Ángel González Ballester, Gustavo Carneiro 0001, Yi-Jen Ju, Juinn-Dar Huang, Joseph Pappachan, Neil D. Reeves, Vishnu Chandrabalan, Darren Dancey, Connah Kendrick |
Medical Image Anal. | 18 |
| 2023 | An Evaluation and Architecture Exploration Engine for CNN Accelerators through Extensive Dataflow AnalysisabstractSystolic array is one of the popular convolutional neural network accelerator architectures due to its high computation efficiency. Nevertheless, the huge design space and complicated interactions among different design parameters make it hard to find the best configuration for various applications. To overcome this issue, this paper presents an evaluation and design space exploration engine, NNeed, for systolic-array CNN accelerators through extensive dataflow analysis. It uses a highly configurable hardware template to describe accelerator operations in detail. The rapid evaluation provides PPA results, pipeline stage analysis, external memory access statistics, and so on. NNeed explores the 9-dimensional design space and supports multiple objective functions for design optimization. Experimental results show that NNeed can generate an accelerator configuration with up to 23% and 50% improvement in performance and energy as compared with a typical handcrafted design. Shan-Hui Chou, Ting-Yun Hsiao, Jing-Yang Jou, Juinn-Dar Huang |
VLSI-SoC | 4 |
| 2023 | Machine Learning Assisted Circuit Sizing Approach for Low-Voltage Analog Circuits with Efficient Variation-Aware OptimizationabstractLow-power analog design is a hot topic for various power efficient applications. Sizing low-power analog circuits is not easy because the increasing uncertainties from low-voltage techniques magnify process variation effects on the design yield. Simulation-based approaches are often adopted for analog circuit sizing because of its high accuracy and adaptability in different cases. However, if process variation is also considered, the huge number of simulations becomes almost infeasible for large circuits. Although there are some recent works that adopt machine learning (ML) techniques to speed up the optimization process, the process variation effects are still hard to be considered in those approaches. Using the popular evolutionary algorithm (EA) as an example, this paper proposes an ML-assisted prediction model to speed up the variation-aware circuit sizing technique for low-voltage analog circuits. By predicting the likelihood for a design that has worse performance, the enhanced EA process is able to skip many unnecessary simulations to reduce the convergence time. Moreover, a novel force-directed model is proposed to guide the optimization toward better yield. Based on the performance of prior circuit samples in the EA optimization, the proposed force model is able to predict the likelihood of a design that has better yield without time-consuming Monte Carlo simulations. Compared with prior works, the proposed approach significantly reduces the number of simulations in the yield-aware EA optimization, which helps to generate practical low-voltage designs with high reliability and low cost. Ling-Yen Song, Chih-Yun Chou, Tung-Chieh Kuo, Chien-Nan Jimmy Liu, Juinn-Dar Huang |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2022 | Fast Variation-aware Circuit Sizing Approach for Analog Design with ML-Assisted Evolutionary AlgorithmabstractEvolutionary algorithm (EA) based on circuit simulation is one of the popular approaches for analog circuit sizing because of its high accuracy and adaptability on different cases. However, if process variation is also considered, the huge number of simulations becomes almost infeasible for large circuits. Although there are some recent works that adopt machine learning (ML) techniques to speed up the optimization process, the variation effects are still hard to be considered in those approaches. In this paper, we propose a fast variation-aware evolutionary algorithm for analog circuit sizing with a ML-assisted prediction model. By predicting the likelihood for a design that has worse performance, our EA process is able to skip many unnecessary simulations to reduce the convergence time. Moreover, a novel force-directed model is proposed to guide the optimization toward better yield. Based on the performance of prior circuit samples in the EA optimization, the proposed force model is able to predict the likelihood of a design that has better yield without time-consuming Monte Carlo simulations. Compared with prior works, the proposed approach significantly reduces the number of simulations in the yield-aware EA optimization, which helps to generate more practical designs with high reliability and low cost. Ling-Yen Song, Tung-Chieh Kuo, Ming-Hung Wang, Chien-Nan Jimmy Liu, Juinn-Dar Huang |
ASP-DAC | 5 |
| 2021 | Design for Restricted-Area and Fast Dilution using Programmable Microfluidic Device based Lab-on-a-ChipabstractMicrofluidic lab-on-a-chip has emerged as a new technology for implementing biochemical protocols on small-sized portable devices targeting low-cost medical diagnostics. Among various efforts of fabrication of such chips, programmable microfluidic device (PMD) is a relatively new technology for implementation of flow-based lab-on-a-chips. A PMD chip is suitable for automation due to its symmetric nature. In order to implement a bioprotocol on such a reconfigurable device, it is crucial to automate sample preparation on a chip as well. Sample preparation, which is a front-end process to produce the desired target concentrations of the input reagent fluid, plays a pivotal role in every bioassay or bioprotocol. In this paper, first, a method referred as dilution algorithm in two steps (DATS) is proposed, which needs only two diluting operations for any target concentration to achieve. Then, we present another method called as dilution algorithm on a small dilution area (DASDA), which needs less area compared to that by DATS. Finally, we propose the heuristic for efficient dilution of biochemical fluids using a PMD chip referred as dilution algorithm in a restricted dilution area (DARDA) that produces more accurate (with less error) target concentration value on a restricted area of the PMD chip in a shorter mixing time. Simulation results reveal that DARDA outperforms a start-of-the-art dilution algorithm applicable for PMD chips in terms of three performance parameters namely mixing time, mixing area and error in target concentration. Shuaijie Ying, Sudip Roy 0001, Juinn-Dar Huang, Shigeru Yamashita |
DSD | 3 |
| 2021 | Diagnosis for Reconfigurable Single-Electron Transistor Arrays with a More Generalized Defect ModelabstractSinge-Electron Transistor (SET) is considered as a promising candidate of low-power devices for replacement or co-existence with Complementary Metal-Oxide-Semiconductor (CMOS) transistors/circuits. In this work, we propose a diagnosis approach for SET array under a more generalized defect model. With the more generalized defect model, the diagnosis approach will become more practical but complicated. We conducted experiments on a set of SET arrays with different dimensions and defect rates. The experimental results show that our approach only has 3.8% false-negative rate and 0.7% misjudged-category rate on average without reporting any false-positive edge when the defect rate is 4%. Therefore, the proposed diagnosis approach can diagnose the defective SET arrays and elevate the reliability of the SET arrays in the synthesis flow. Chia-Cheng Wu, Yi-Hsiang Hu, Chia-Chun Lin, Yung-Chih Chen, Juinn-Dar Huang, Chun-Yao Wang |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2020 | High-Speed Power-Efficient Coarse-Grained Convolver Architecture using Depth-First Compression SchemeabstractConvolutional neural networks (CNNs) have been playing an important role in various applications, e.g., computer vision. Since CNN computations require numerous multiply-accumulate (MAC) operations, how to get them done efficiently is a crucial issue for CNN hardware accelerators. In this paper, we propose a high-speed power-efficient convolver architecture for CNN acceleration. A 3×3 convolver is asked to produce an output every cycle and is commonly accomplished by summing up the results of nine parallel multiplications, which requires ten carry-propagation adders (CPAs) in total. However, the proposed coarse-grained convolver can break the boundary between multipliers and reduce all partial products in a more global way. Consequently, it requires only one CPA to generate the final outcome. It also features a globally delay-optimized partial product reduction tree and a depth-first compression scheme for both area and power minimization. The proposed convolver has been implemented using TSMC 40nm technology. Compared to a conventional 3×3 convolver baseline design, our design can reduce area and power by 15.8% and 26.5% respectively at the clock rate of 1GHz. Juinn-Dar Huang |
ISCAS | 3 |
| 2020 | Storage-Aware Algorithms for Dilution and Mixture Preparation With Flow-Based Lab-on-ChipabstractLab-on-chip (LoC) technology has emerged as one of the major driving forces behind the recent surge in biochemical protocol automation. Dilution and mixture preparation with fluids in a desired ratio, constitute basic steps in sample preparation for which several LoC-based architectures and algorithms are known. The optimization of cost and time for such protocols requires proper sequencing of fluidic mix-and-split steps, and storage-units for holding intermediate-fluids to be reused in the later steps. However, practical design constraints often limit the amount of on-chip storage in microfluidic LoC architectures and thus can badly affect the performance of the algorithms. Consequently, results generated by previous work may not be useful (in the case they require more storage-units than available) or more expensive than necessary (in the case when storage-units are available but not used, e.g., to further reduce the number of mix/split operations or reactant-cost). In this paper, we propose new algorithms for dilution and mixing with continuous-flow-based LoCs that explicitly take care of storage constraints while optimizing reactant-cost and time of sample preparation. We present a symbolic formulation of the problem that captures the degree of freedom in algorithmic steps satisfying the specified storage constraints. Solvers based on Boolean satisfiability are used to achieve the optimization goals. The experimental results show the efficiency and effectiveness of the solution as well as a variety of applications where the proposed methods would prove beneficial. Sukanta Bhattacharjee, Robert Wille, Juinn-Dar Huang, Bhargab B. Bhattacharya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Forecast-Based Sample Preparation Algorithm for Unbalanced Splitting Correction on DMFBsabstractSample preparation is regarded as one of essential processing steps in most biochemical assays. In the past decade, numerous techniques have been presented to deal with sample preparation under the (1:1) mixing model on digital microfluidic biochips (DMFBs) for various optimization goals. However, most of previous works assumed that mixing-then-splitting would get two identical output droplets, which is not always true due to unbalanced splitting. As a consequence, those works may fail to provide correct solutions at the presence of unbalanced splitting. Several methods have been proposed to deal with this issue. Nevertheless, some of them rely on hypotheses that may not be practical, while the others demand extra reactants or special hardware. In this paper, we propose a new probability-based sample preparation algorithm for unbalanced splitting correction. Our new algorithm not only guarantees a correct solution, but requires neither extra reactants nor on-chip special hardware. Experimental results show that the effect of unbalanced splitting can be eliminated only at the cost of 20% more operation steps. That is, the proposed algorithm is both reliable and efficient. Ling-Yen Song, Yi-Ling Chen 0005, Yung-Chun Lei, Juinn-Dar Huang |
ICCD | 4 |
| 2019 | Design Automation for Dilution of a Fluid Using Programmable Microfluidic Device-Based BiochipsabstractMicrofluidic lab-on-a-chip has emerged as a new technology for implementing biochemical protocols on small-sized portable devices targeting low-cost medical diagnostics. Among various efforts of fabrication of such chips, a relatively new technology is a programmable microfluidic device (PMD) for implementation of flow-based lab-on-a-chip. A PMD chip is suitable for automation due to its symmetric nature. In order to implement a bioprotocol on such a reconfigurable device, it is crucial to automate a sample preparation on-chip as well. In this article, we propose a dilution PMD algorithm (namely DPMD ) and its architectural mapping scheme (namely generalized architectural mapping algorithm ( GAMA )) for addressing fluidic cells of such a device to perform dilution of a reagent fluid on-chip. We used an optimization function that first minimizes the number of mixing steps and then reduces the waste generation and further reagent requirement. Simulation results show that the proposed DPMD scheme is comparative to the existing state-of-the-art dilution algorithm. The proposed design automation using the architectural mapping scheme reduces the required chip area and, hence, minimizes the valve switching that, in turn, increases the life span of the PMD-chip. Ankur Gupta 0002, Juinn-Dar Huang, Shigeru Yamashita, Sudip Roy 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2018 | Storage-aware sample preparation using flow-based microfluidic Labs-on-ChipabstractRecent advances in microfluidics have been the major driving force behind the ubiquity of Labs-on-Chip (LoC) in biochemical protocol automation. The preparation of dilutions and mixtures of fluids is a basic step in sample preparation for which several algorithms and chip-architectures are well known. Dilution and mixing are implemented on biochips through a sequence of basic fluid-mixing and splitting operations performed in certain ratios. These steps are abstracted using a mixing graph. During this process, on-chip storage-units are needed to store intermediate fluids to be used later in the sequence. This allows to optimize the reactant-costs, to reduce the sample-preparation time, and/or to achieve the desired ratio. However, the number of storage-units is usually limited in given LoC architectures. Since this restriction is not considered by existing methods for sample preparation, the results that are obtained are often found to be useless (in the case when more storage-units are required than available) or more expensive than necessary (in the case when storage-units are available but not used, e.g., to further reduce the number of mixing operations or reactant-cost). In this paper, we present a storage-aware algorithm for sample preparation with flow-based LoCs which addresses these issues. We present a SAT-based approach to construct a mixing graph that enables the best usage of available storage-units while optimizing sample-preparation cost and/or time. Experimental results on several test cases reveal the scope, effectiveness, and the flexibility of the proposed method. Sukanta Bhattacharjee, Robert Wille, Juinn-Dar Huang, Bhargab B. Bhattacharya |
DATE | 3 |
| 2018 | A Comprehensive Security System for Digital Microfluidic BiochipsabstractDigital microfluidic biochips (DMFBs) have become popular in the healthcare industry recently because of its lowcost, high-throughput, and portability. Users can execute the experiments on biochips with high resolution, and the biochips market therefore grows significantly. However, malicious attackers exploit Intellectual Property (IP) piracy and Trojan attacks to gain illegal profits. The conventional approaches present defense mechanisms that target either IP piracy or Trojan attacks. In practical, DMFBs may suffer from the threat of being attacked by these two attacks at the same time. This paper presents a comprehensive security system to protect DMFBs from IP piracy and Trojan attacks. We propose an authentication mechanism to protect IP and detect errors caused by Trojans with CCD cameras. By our security system, we could generate secret keys for authentication and determine whether the bioassay is under the IP piracy and Trojan attacks. Experimental results demonstrate the efficacy of our security system without overhead of the bioassay completion time. Juinn-Dar Huang, Hailong Yao 0002, Tsung-Yi Ho |
ITC-Asia | 2 |
| 2018 | Versatile Ring-Based Architecture and Synthesis Flow for General-Purpose Digital Microfluidic BiochipsabstractDigital microfluidic biochip (DMFB) is a tiny device that can carry out a rich set of bioassays without the need of bulky equipment. However, designing a good general-purpose DMFB architecture is still considered a big challenge today. NP-hard synthesis problems make on-line synthesis virtually impossible on exiting array-based architectures. In this paper, we first elaborate on the major concerns in a DMFB design flow, from the aspects of both synthesis and physical design. We then propose a versatile ring-based architecture VERBA and its corresponding fast one-pass synthesis flow. Experimental results show that VERBA incorporated with the proposed synthesis flow is a better solution than existing architectures and synthesis algorithms especially for real-time cyber-physical systems. Juinn-Dar Huang, Chia-Hung Liu, Wei-Hao Yang |
VLSI-SoC | 1 |
| 2018 | Concentration-Resilient Mixture Preparation with Digital Microfluidic Lab-on-ChipabstractSample preparation plays a crucial role in almost all biochemical applications, since a predominant portion of biochemical analysis time is associated with sample collection, transportation, and preparation. Many sample-preparation algorithms are proposed in the literature that are suitable for execution on programmable digital microfluidic (DMF) platforms. In most of the existing DMF-based sample-preparation algorithms, a fixed target ratio is provided as input, and the corresponding mixing tree is generated as output. However, in many biochemical applications, target mixtures with exact component proportions may not be needed. From a biochemical perspective, it may be sufficient to prepare a mixture in which the input reagents may lie within a range of concentration factors. The choice of a particular valid ratio, however, strongly impacts solution-preparation cost and time. To address this problem, we propose a concentration-resilient ratio-selection method from the input ratio space so that the reactant cost is minimized. We propose an integer linear programming--based method that terminates very fast while producing the optimum solution, considering both uniform and weighted cost of reagents. Experimental results reveal that the proposed method can be used conveniently in tandem with several existing sample-preparation algorithms for improving their performance. Sukanta Bhattacharjee, Yi-Ling Chen 0005, Juinn-Dar Huang, Bhargab B. Bhattacharya |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | Defect-aware synthesis for reconfigurable single-electron transistor arraysabstractAs fabrication process exploits even deeper submicron technology, power consumption is becoming one of the most critical obstacles in electronic circuit and system designs nowadays. Meanwhile, the leakage power is dominating the power consumption. Various emerging nanodevices have been developed to tackle the leakage power issue in recent years. The single-electron transistor (SET) is regarded as one of the most promising devices since several works have successfully demonstrated that it can operate with only few electrons at room temperature. Therefore, the reconfigurable SET array has been proposed to continue Moore's Law due to its ultra-low power consumption. Nevertheless, most existing synthesis algorithms assume the given SET array is defect-free. Hence, mapping a correct synthesis outcome onto a faulty SET array still yields an erroneous result. In this paper, we propose a new synthesis algorithm that guarantees the correct functionality in the presence of defects. Furthermore, the proposed technique can sometimes benefit from those defects to further reduce the mapping area. In certain cases, the required area in a faulty SET array is even smaller than that in a fault-free one. Experimental results show that our new algorithm can synthesize moderately large circuits in a reasonable runtime and achieve an area reduction of 14% as compared to the prior art. Juinn-Dar Huang, Yi-Hang Chen, Jia-Shin Lu |
VLSI-SoC | 1 |
| 2017 | Dilution and Mixing Algorithms for Flow-Based Microfluidic BiochipsabstractAlbeit sample preparation is well-studied for digital microfluidic biochips, very few prior work addressed this problem in the context of continuous-flow microfluidics from an algorithmic perspective. In the latter class of chips, microvalves and micropumps are used to manipulate on-chip fluid flow through microchannels in order to execute a biochemical protocol. Dilution of a sample fluid is a special case of sample preparation, where only two input reagents (commonly known as sample and buffer) are mixed in a desired volumetric ratio. In this paper, we propose a satisfiability-based dilution algorithm assuming the generalized mixing models supported by an N-segment, continuous-flow, rotary mixer. Given a target concentration and an error limit, the proposed algorithm first minimizes the number of mixing operations, and subsequently, reduces reagent-usage. Simulation results demonstrate that the proposed method outperforms existing dilution algorithms in terms of mixing steps (assay time) and waste production, and compares favorably with respect to reagent-usage (cost) when 4- and 8-segment rotary mixers are used. Next, we propose two variants of an algorithm for handling the open problem of k-reagent mixture-preparation (k ≥ 3) with an N-segment continuous-flow rotary mixer, and report experimental results to evaluate their performance. A software tool called flow-based sample preparation algorithm has also been developed that can be readily used for running the proposed algorithms. Sukanta Bhattacharjee, Sudip Poddar, Sudip Roy 0001, Juinn-Dar Huang, Bhargab B. Bhattacharya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Chain-based pin count minimization for general-purpose digital microfluidic biochipsabstractMinimizing the number of external control pins is one of the most important optimization objectives in digital microfluidic biochip (DMFB) designs especially as the chip size gets even bigger. So far, only few works focus on this issue for general-purpose DMFBs. In this paper, we present a pin count minimization algorithm based on sophisticated electrode chaining on regular or irregular electrode arrays. The key idea of the proposed method is that actuation information can be implied from previous neighborhood electrodes to later ones throughout a chain. Experimental results show that the pin count reduction can be near 50% in large DMFBs. Yung-Chun Lei, Chen-Shing Hsu, Juinn-Dar Huang, Jing-Yang Jou |
ASP-DAC | 3 |
| 2016 | Area Minimization Synthesis for Reconfigurable Single-Electron Transistor Arrays with Fabrication ConstraintsabstractPower dissipation has become a pressing issue of concern in the designs of most electronic system as fabrication processes enter even deeper submicron regions. More specifically, leakage power plays a dominant role in system power dissipation. An emerging circuit design style, the reconfigurable single-electron transistor (SET) array, has been proposed for continuing Moore's Law due to its ultra-low leakage power consumption. Recently, several works have been proposed to address the issues related to automated synthesis for the reconfigurable SET array. Nevertheless, all of those existing approaches consider mandatory fabrication constraints of SET array merely in late synthesis stages. In this article, we propose a synthesis algorithm, featuring input-variable ordering and dynamic product term ordering, for area minimization. The fabrication constraints are taken into account at every synthesis stage of proposed flow to guarantee better synthesis outcomes. We also develop a simulated annealing-based postprocess to find a proper phase assignment of each input variable for further area reduction. Experimental results show that our new methodology can achieve up to 29% area reduction as compared to existing state-of-the-art techniques. Yi-Hang Chen, Juinn-Dar Huang |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2015 | Volume-oriented sample preparation for reactant minimization on flow-based microfluidic biochips with multi-segment mixers
Chi-Mei Huang, Chia-Hung Liu, Juinn-Dar Huang |
DATE | 3 |
| 2015 | Reactant Minimization in Sample Preparation on Digital Microfluidic BiochipsabstractSample preparation plays an essential role in most biochemical reactions. Raw reactants are diluted to solutions with desirable concentration values in this process. Since the reactants, like infant's blood, DNA evidence collected from crime scenes, or costly reagents, are extremely valuable, their usage should be minimized whenever possible. In this paper, we propose a two-phased reactant minimization algorithm (REMIA), for sample preparation on digital microfluidic biochips. In the former phase, REMIA builds a reactant-minimized interpolated dilution tree with specific leaf nodes for a target concentration. Two approaches are developed for tree construction; one is based on integer linear programming (ILP) and the other is heuristic. The ILP one guarantees to produce an optimal dilution tree with minimal reactant consumption, whereas the heuristic one ensures runtime efficiency. Then, REMIA constructs a forest consisting of exponential dilution trees to produce those aforementioned specific leaf nodes with minimal reactant consumption in the latter phase. Experimental results show that REMIA achieves a reduction of reactant usage by 32%-52% as compared with three existing state-of-the-art sample preparation approaches. Besides, REMIA can be easily extended to solve the sample preparation problem with multiple target concentrations, and the extended version also effectively lowers the reactant consumption further. Chia-Hung Liu, Ting-Wei Chiang, Juinn-Dar Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | Reactant Minimization for Sample Preparation on Microfluidic Biochips With Various Mixing ModelsabstractSample preparation is one of the essential processes for most on-chip biochemical applications. During this process, raw reactants are diluted to specific concentration values. Current sample preparation algorithms are generally created for digital microfluidic biochips with the (1:1) mixing model. For other biochip architectures supporting multiple mixing models, such as flow-based microfluidic biochips, there is still no dedicated solution yet. Hence, in this paper, we propose the first sample preparation method dedicated to microfluidic biochips with various mixing models, named tree pruning and grafting (TPG) algorithm. It starts with a dilution tree created by regarding the (1:1) mixing model only, and then applies TPG through a bottom-up dynamic programming strategy to obtain a solution with minimal reactant consumption. Experimental results show that our algorithm can save reactant amount by up to 69% against the well-known bit-scanning method on a biochip with a four-segment mixer. Even compared with the state-of-the-art reactant minimization algorithm, it still achieves a reactant reduction of 37%. Therefore, it is convincing that the TPG algorithm is a promising sample preparation solution for biochip architectures that support various mixing models. Chia-Hung Liu, Kuo-Cheng Shen, Juinn-Dar Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | Area minimization synthesis for reconfigurable single-electron transistor arrays with fabrication constraintsabstractAs fabrication processes exploit even deeper submicron technology, power dissipation has become a crucial issue for most electronic circuit and system designs nowadays. In particular, leakage power is becoming a dominant source of power consumption. Recently, the reconfigurable single-electron transistor (SET) array has been proposed as an emerging circuit design style for continuing Moore's Law due to its ultra-low power consumption. Several automated synthesis approaches have been developed for the reconfigurable SET array in the past few years. Nevertheless, all of those existing methods consider fabrication constraints, which are mandatory, merely in late synthesis stages. In this paper, we propose a synthesis algorithm, featuring both variable reordering and product term reordering, for area minimization. In addition, our algorithm takes those mandatory fabrication constraints into account in early stages for better outcomes. Experimental results show that our new method can achieve an area reduction of up to 24% as compared to current state-of-the-art techniques. Yi-Hang Chen, Juinn-Dar Huang |
DATE | 3 |
| 2013 | Sample preparation for many-reactant bioassay on DMFBs using common dilution operation sharingabstractSample preparation is an essential processing step in most biochemical applications. Various reactants are mixed together to produce a solution with the target concentration. Since reactants generally take a notable part of the cost in a bioassay, their usage should be minimized whenever possible. In this paper, we propose an algorithm, CoDOS, to prepare the target solution with many reactants using common dilution operation sharing on digital microfluidic biochips (DMFBs). CoDOS first represents the given target concentration as a recipe matrix, and then identifies rectangles in the matrix, where each rectangle indicates an opportunity of dilution operation sharing for reactant minimization. Experimental results demonstrate that CoDOS can achieve up to 27% of reactant saving as compared with the bit-scanning method in single-target sample preparation. Moreover, even if CoDOS is not developed for multi-target sample preparation, it still outperforms the recent state-of-the-art algorithm, RSMA. Hence, it is convincing that CoDOS is a better alternative for many-reactant sample preparation. Chia-Hung Liu, Hao-Han Chang, Tung-Che Liang, Juinn-Dar Huang |
ICCAD | 4 |
| 2013 | Reactant and Waste Minimization in Multitarget Sample Preparation on Digital Microfluidic BiochipsabstractSample preparation is one of essential processes in biochemical reactions. Raw reactants are diluted in this process to achieve given target concentrations. A bioassay may require several different target concentrations of a reactant. Both the dilution operation count and the reactant usage can be minimized if multiple target concentrations are considered simultaneously during sample preparation. Hence, in this paper, we propose a multitarget sample preparation algorithm that extensively exploits the ideas of waste recycling and intermediate droplet sharing to reduce both reactant usage and waste amount for digital microfluidic biochips. Experimental results show that our waste recycling algorithm can reduce the waste and operation count by 48% and 37%, respectively, as compared to an existing state-of-the-art multitarget sample preparation method if the number of target concentrations is ten. The reduction can be up to 97% and 73% when the number of target concentrations goes even higher. Juinn-Dar Huang, Chia-Hung Liu, Huei-Shan Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2012 | Thermal-aware logic block placement for 3D FPGAs considering lateral heat dissipation (abstract only)abstractThree-dimensional (3D) integration is an attractive and promising technology to keep Moore's Law alive, whereas the thermal issue also presents a critical challenge for 3D integrated circuits. Meanwhile, accurate thermal analysis is very time-consuming and thus can hardly be incorporated into most of placement algorithms generally performing numerous iterative refinement steps. As a consequence, in this paper, we first present a fine-grained grid-based thermal model for the 3D regular FPGA architecture and also highlight that lateral heat dissipation paths can no longer be assumed negligible. Then we propose two fast thermal-aware placement algorithms for 3D FPGAs, Standard Deviation (SD) and MineSweeper (MS), in which rapid thermal evaluation instead of slow detailed analysis is utilized. Moreover, both take the lateral heat dissipation into consideration and focus on distributing heat sources more evenly within a layer in a 3D FPGA to avoid creating hotspots. Experimental results show that SD and MS achieve 12.1%/7.6% reduction in maximum temperature and 82%/56% improvement in temperature deviation compared with a classical thermal-unaware placement method only at the cost of minor increase in wirelength and delay. Moreover, MS merely consumes 4% more runtime for producing thermal-aware placement solutions. Juinn-Dar Huang, Ya-Shih Huang, Mi-Yu Hsu, Han-Yuan Chang |
FPGA | 1 |
| 2012 | Reactant minimization during sample preparation on digital microfluidic biochips using skewed mixing treesabstractSample preparation is an indispensable process to biochemical reactions. Original reactants are usually diluted to the solutions with desirable concentrations. Since the reactants, like infant's blood, DNA evidence collected from a crime scene, or costly reagents, are extremely valuable, the usage of reactant must be minimized in the sample preparation process. In this paper, we propose the first reactant minimization approach, REMIA, during sample preparation on digital microfluidic biochips (DMFBs). Given a target concentration, REMIA constructs a skewed mixing tree to guide the sample preparation process for reactant minimization. Experimental results demonstrate that REMIA can save about 31%~52% of reactant usage on average compared with three existing sample preparation methods. Besides, REMIA can be extended to tackle the sample preparation problem with multiple target concentrations, and the extended version also successfully decreases the reactant usage further. Juinn-Dar Huang, Chia-Hung Liu, Ting-Wei Chiang |
ICCAD | 1 |
| 2011 | Throughput optimization for latency-insensitive system with minimal queue insertionabstractAs fabrication process exploits even deeper submicron technology, global interconnect delay is becoming one of the most critical performance obstacles in system-on-chip (SoC) designs nowadays. Recent years latency-insensitive system (LIS), which enables multicycle communication to tolerate variant interconnect delay without substantially modifying pre-designed IP cores, has been proposed to conquer this issue. However, imbalanced interconnect latency and communication back-pressure residing in an LIS still degrade system throughput. In this paper, we present a throughput optimization technique with minimal queue insertion. We first model a given LIS as a quantitative graph (QG), which can be further compacted using the proposed techniques, so that much bigger problems can be handled. On top of QG, the optimal solution with minimal queue size can be achieved through integer linear programming based on the proposed constraint formulation in an acceptable runtime. The experimental results show that our approach can deal with moderately large systems in a reasonable runtime and save about 28% of queues compared to the prior art. Juinn-Dar Huang, Yi-Hang Chen, Ya-Chien Ho |
ASP-DAC | 1 |
| 2011 | Equivalence checking of scheduling with speculative code transformations in high-level synthesisabstractThis paper presents a formal method for equivalence checking between the descriptions before and after scheduling in high-level synthesis (HLS). Both descriptions are represented by finite state machine with datapaths (FSMDs) and are then characterized through finite sets of paths. The main target of our proposed method is to verify scheduling employing code transformations-such as speculation and common subexpression extraction (CSE), across basic block (BB) boundaries-which have not been properly addressed in the past. Nevertheless, our method can verify typical BB-based and path-based scheduling as well. The experimental results demonstrate that the proposed method can indeed outperform an existing state-of-the-art equivalence checking algorithm. Chi-Hui Lee, Che-Hua Shih, Juinn-Dar Huang, Jing-Yang Jou |
ASP-DAC | 3 |
| 2011 | Architectural exploration of 3D FPGAs towards a better balance between area and delayabstractThe emerging 3D technology, which stacks multiple dies within a single chip and utilizes through-silicon vias (TSVs) as vertical connections, is considered a promising solution for achieving better performance and easy integration. Similarly, a generic 2D FPGA architecture can evolve into a 3D one by extending its signal switching scheme from 2D to 3D by means of TSVs. However, replacing all 2D switch boxes (SBs) by 3D ones with full vertical connectivity is found both area-consuming and resource-squandering. Therefore, it is possible to greatly reduce the footprint with only minor delay increase by properly tailoring the structure and deployment strategy of 3D SB. In this paper, we perform a comprehensive architectural exploration of 3D FPGAs. Various architectural alternatives are proposed and then evaluated thoroughly to pick out the most appropriate ones with a better balance between area and delay. Finally, we recommend several configurations for generic 3D FPGA architectures, which can save up to 52% area with virtually no delay penalty. Chia-I Chen, Bau-Cheng Lee, Juinn-Dar Huang |
DATE | 3 |
| 2009 | CriAS: a performance-driven criticality-aware synthesis flow for on-chip multicycle communication architectureabstractIn deep submicron era, wire delay is no longer negligible and is dominating the system performance. Several state-of-the-art architectural synthesis flows have been proposed for the distributed register architectures to cope with the increasing wire delay by allowing on-chip multicycle communication. In this paper, we present a new performance-driven criticality-aware synthesis flow CriAS targeting regular distributed register architectures. CriAS features a hierarchical binding strategy and a coarse-grained placer for minimizing the number of critical global data transfers. The key ideas are to take time criticality as the major concern at earlier binding stages before the detailed physical placement information is available, and to preserve the locality of closely related critical components in the later placement phase. The experimental results show that 19% overall performance improvement can be achieved on average as compared to the previous work. Chia-I Chen, Juinn-Dar Huang |
ASP-DAC | 2 |
| 2009 | Simultaneous data transfer routing and scheduling for interconnect minimization in multicycle communication architectureabstractIn deep submicron technology, wire delay is no longer negligible and is gradually becoming a dominant factor of system performance. Several state-of-the-art architectural synthesis flows have already adopted the distributed register architecture to cope with the increasing wire delay by allowing multicycle communication. In this paper, we formulate channel and register allocation within a refined regular distributed register architecture, named RDR-GRS, as a problem of simultaneous data transfer routing and scheduling for minimizing global interconnect resources. We also present an innovative algorithm with both spatial and temporal considerations. It features both a concentration-oriented path router gathering wire-sharable data transfers and a channel-based time scheduler resolving contentions for wires in a channel, which are in spatial and temporal domain, respectively. The experimental results show that the proposed algorithm can significantly outperform existing related works. Yu-Ju Hong, Ya-Shih Huang, Juinn-Dar Huang |
ASP-DAC | 3 |
| 2009 | Reducing fault dictionary size for million-gate large circuitsabstractIn general, fault dictionary is prevented from practical applications in fault diagnosis due to its extremely large size. Several previous works are proposed for the fault dictionary size reduction. However, some of them fail to bring down the size to an acceptable level, and others might not be able to handle today's million-gate circuits due to their high time and space complexity. In this article, an algorithm is presented to reduce the size of pass-fail dictionary while still preserving high diagnostic resolution. The proposed algorithm possesses low time and space complexity by avoiding constructing the huge distinguishability table, which inevitably boosts up the required computation complexity. Experimental results demonstrate that the proposed algorithm is capable of handling industrial million-gate large circuits in a reasonable amount of runtime and memory. Yu-Ru Hong, Juinn-Dar Huang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2009 | Automatic Verification Stimulus Generation for Interface Protocols Modeled With Non-Deterministic Extended FSMabstractVerifying if an integrated component is compliant with certain interface protocol is a vital issue in component-based system-on-a-chip (SoC) designs. For simulation-based verification, generating massive constrained simulation stimuli is becoming crucial to achieve a high verification quality. To further improve the quality, stimulus biasing techniques are often used to guide the simulation to hit design corners. In this paper, we model the interface protocol with the non-deterministic extended finite-state machine (NEFSM), and then propose an automatic stimulus generation approach based on it. This approach is capable of providing numerous biasing strategies. Experiment results demonstrate the high controllability and efficiency of our stimulus generation scheme. Che-Hua Shih, Juinn-Dar Huang, Jing-Yang Jou |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | A multicycle communication architecture and synthesis flow for Global interconnect Resource SharingabstractIn deep submicron technology, wire delay is no longer negligible and is gradually dominating the system latency. Some state-of-the-art architectural synthesis flows adopt the distributed register (DR) architecture to cope with this increasing latency. The DR architecture, though allows multicycle communication, introduces extra overhead on interconnect resource. In this paper, we propose the regular distributed register - global resource sharing (RDR-GRS) architecture to enable global sharing of interconnects and registers. Based on the RDR-GRS architecture, we further define the channel and register allocation problem as a path scheduling problem of data transfers. A formal and flexible formulation of this problem is then presented and optimally solved by Integer Linear Programming (ILP). Experimental results show that RDR-GRS/ILP can averagely reduce 58% wires and 35% registers compared to the previous work. Wei-Sheng Huang, Yu-Ru Hong, Juinn-Dar Huang, Ya-Shih Huang |
ASP-DAC | 3 |
| 2007 | Fault Dictionary Size Reduction for Million-Gate Large CircuitsabstractIn general, fault dictionary is prevented from practical applications for its extremely large size. Several previous works are proposed for the fault dictionary size reduction. However, they might not be able to handle today's million-gate circuits due to the high time and space complexity. In this paper, we propose an algorithm to significantly reduce the size of fault dictionary while still preserving high diagnostic resolution. The proposed algorithm possesses extremely low time and space complexity by avoiding constructing the huge distinguishability table, which inevitably boosts up the required computation complexity. Experimental results demonstrate that the proposed algorithm is fully capable of handling industrial million-gate large circuits in a reasonable amount of runtime and memory. Yu-Ru Hong, Juinn-Dar Huang |
ASP-DAC | 2 |
| 2007 | A Precise Bandwidth Control Arbitration Algorithm for Hard Real-Time SoC BusesabstractOn an SoC bus, contentions occur while different IP cores request the bus access at the same time. Hence an arbiter is mandatory to deal with the contention issue on a shared bus system. In different applications, IPs may have real-time and/or bandwidth requirements. It is very difficult to design an arbitration algorithm to simultaneously meet these two requirements. In this paper, we propose an innovative arbitration algorithm, RB_lottery, to meet both of the requirements. It can provide not only the hard real-time guarantee but also the precise bandwidth controllability. The experimental results show that RBJottery outperforms several well-known existing arbitration algorithms. Bu-Ching Lin, Geeng-Wei Lee, Juinn-Dar Huang, Jing-Yang Jou |
ASP-DAC | 3 |
| 2006 | A real-time and bandwidth guaranteed arbitration algorithm for SoC bus communicationabstractIn shared SoC bus systems, arbiters are usually adopted to solve bus contentions with various kinds of arbitration algorithms. We propose an arbitration algorithm, RT/spl I.bar/lottery, which is designed to meet both hard real-time and bandwidth requirements. For fast evaluation and exploration, we use high abstract-level models in our system simulation environment to generate parameters for our configurable arbiter. The experimental results show that RT/spl I.bar/lottery can meet all hard real-time requirements and perform very well in bandwidth allocation. The results also show that RT/spl I.bar/lottery outperforms several commonly-used arbitration algorithms today. Chien-Hua Chen, Geeng-Wei Lee, Juinn-Dar Huang, Jing-Yang Jou |
ASP-DAC | 3 |
| 2006 | FSM-based transaction-level functional coverage for interface compliance verificationabstractInterface compliance verification plays a very important role in modern SoC designs. In order to perform a quantitative analysis of simulation completeness, adequate coverage metrics are mandatory. In this paper, we propose a finite state machine (FSM) based transaction-level functional coverage methodology for interface compliance verification. A language, state-oriented language (SOL), is developed to specify functional transactions mainly at the higher FSM level instead of lower logic or signal level. By utilizing SOL, it is simple and rigorous to specify interesting transactions from the specification FSM of the target interface protocol. Experimental results show that the proposed methodology can effectively improve the verification quality as well as increase the efficiency of regression verification. Man-Yun Su, Che-Hua Shih, Juinn-Dar Huang, Jing-Yang Jou |
ASP-DAC | 3 |
| 2004 | Verification on Port ConnectionsabstractIn a system-on-a-chip (SOC) design, several to hundreds of design blocks or intellectual properties (IPs) are integrated to form a complex function. Prior to verify the functionality of the integrated IPs, it is very important to ensure the correctness of the port connections among these IPs. This work addresses the problem of verification on port connections while IPs are integrated into a larger block or a system, and presents a new connection model and the corresponding error model for port connections. An algorithm providing the minimum pattern set and a general verification flow used to verify port connections are also proposed. Geeng-Wei Lee, Juinn-Dar Huang, Jing-Yang Jou, Chun-Yao Wang |
ITC | 2 |
| 2001 | Unified functional decomposition via encoding for FPGA technology mappingabstractFunctional decomposition has recently been adopted for look-up table (LUT)-based field-programmable gate array (FPGA) technology mapping with good results. In this paper we propose a novel method to unify functional single-output and multiple-output decomposition. We first address a compatible class encoding algorithm to minimize the number of compatible classes in the image function. After applying the encoding algorithm, we can therefore improve the decomposability in the subsequent decomposition of the image function. The above encoding algorithm is then extended to encode multiple-output functions through the construction of a hyperfunction. Common subexpressions among these multiple-output functions can be extracted during the decomposition of the hyperfunction. Consequently, we can handle multiple-output decomposition in the same manner as single-output decomposition. Experimental results show that our algorithms are promising. Jie-Hong Roland Jiang, Jing-Yang Jou, Juinn-Dar Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2000 | ALTO: an iterative area/performance tradeoff algorithm for LUT-based FPGA technology mappingabstractIn this paper, we propose an iterative area/performance tradeoff algorithm for look-up table (LUT)-based field programmable gate array (FPGA) technology mapping. First, it finds an area-optimized, performance-considered initial network by a modified area optimization technique. Then, an iterative algorithm consisting of several resynthesizing techniques is applied to trade the area for the performance in the network gracefully. Experimental results show that this approach can efficiently provide a complete set of mapping solutions from the area-optimized one to the performance-optimized one for the given design. Furthermore, these two extreme solutions produced by our algorithm outperform the results provided by most existing algorithms. Therefore, our algorithm is very useful for the timing-driven, LUT-based FPGA synthesis. Juinn-Dar Huang, Jing-Yang Jou, Wen-Zen Shen |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1998 | Compatible Class Encoding in Hyper-Function Decomposition for FPGA SynthesisabstractRecently, functional decomposition has been adopted for LUT based FPGA technology mapping with good results. In this paper, we propose a novel method for functional multiple-output decomposition. We first address a compatible class encoding method to minimize the compatible classes in the image function. After the encoding algorithm is applied, the decomposability will be improved in the subsequent decomposition of the image function. The above encoding algorithm is then extended to encode multiple-output functions through the construction of a hyper-function. Common sub-expressions among these multiple-output functions can be extracted during the decomposition of the hyper-function. Therefore, we can handle the multiple-output decomposition in the same manner as the single-output decomposition. Experimental results show that our algorithms are very promising. Jie-Hong Roland Jiang, Jing-Yang Jou, Juinn-Dar Huang |
DAC | 3 |
| 1998 | On circuit clustering for area/delay tradeoff under capacity and pin constraintsabstractIn this paper, we propose an iterative area/delay tradeoff algorithm to solve the circuit clustering problem under the capacity constraint. It first finds an initial delay-considered area-optimized clustering solution by a delay-oriented depth first-search procedure. Then, an iterative procedure consisting of several reclustering techniques is applied to gradually trade the area for the performance. We then show that this algorithm can be easily extended to solve the clustering problem subject to both capacity and pin constraints. Experimental results show that our algorithm can provide a complete set of clustering solutions from the area-optimized one to the delay-optimized one for a given circuit. Furthermore, compared to the existing delay-optimized algorithms, this algorithm achieves almost the same performance but with much less area overhead. Therefore, this algorithm is very useful for solving the timing-driven circuit clustering problem. Juinn-Dar Huang, Jing-Yang Jou, Wen-Zen Shen, Hsien-Ho Chuang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1997 | BDD based lambda set selection in Roth-Karp decomposition for LUT architectureabstractField Programmable Gate Arrays (FPGAs) are important devices for rapid system prototyping. Roth-Karp decomposition is one of the most popular decomposition techniques for Look-Up Table (LUT)-based FPGA technology mapping. In this paper, we propose a novel algorithm based on Binary Decision Diagrams (BDDs) for selecting good lambda set variables in Roth-Karp decomposition to minimize the number of consumed configurable logic blocks (CLBs) in FPGAs. The experimental results on a set of benchmarks show that our algorithm can produce much better results than those of the previous approach (Wen-Zen Shen et al., 1995). Jie-Hong Roland Jiang, Jing-Yang Jou, Juinn-Dar Huang, Jung-Shian Wei |
ASP-DAC | 3 |
| 1996 | An iterative area/performance trade-off algorithm for LUT-based FPGA technology mappingabstractIn this paper, we propose an iterative area/performance trade-off algorithm for LUT-based FPGA technology mapping. First, it finds an area-optimized performance-considered initial network by a modified area optimization technique. Then, an iterative algorithm consisting of several resynthesizing techniques is applied to trade the area for the performance in the network gracefully. Experimental results show that this approach can provide a complete set of mapping solutions from the area-optimized one to the performance-optimize one for the given design. Furthermore, these two extreme solutions, the area-optimized one and the performance-optimized one, produced by our algorithm outperform the results of most existing algorithms. Therefore, our algorithm is very useful for the timing driven FPGA synthesis. Juinn-Dar Huang, Jing-Yang Jou, Wen-Zen Shen |
ICCAD | 1 |
| 1995 | Lambda Set Selection in Roth-Karp Decomposition for LUT-Based FPGA Technology Mappingabstractpartition tends to produce better results.However, to the best of our knowledge, finding a good input partition in Roth-Karp decomposition has not been formally addressed in previous research.In this paper, we propose a new heuristics to solve this problem.Roth-Karp decomposition is a classical decomposition method.Because it can reduce the number of input variables of a function, it becomes one of the most popular techniques used in LUT-based FPGA technology mapping.However, the lambda set selection problem, which can dramatically affect the decomposition quality in Roth-Karp decomposition, has not been formally addressed before.In this paper, we propose a new heuristic-based algorithm to solve this problem.The experimental results show that our algorithm can efficiently produce outputs with better decomposition quality than that produced by other algorithms without using lambda set selection strategy.This paper is organized as follows.Section 2 briefly introduces some terminologies used in this paper.Section 3 describes Roth-Karp decomposition and a classical implementation for it.In Section 4, our newly developed heuristics for selecting a good λ set is given in detail.Section 5 shows some experimental results and concluding remarks are given in Section 6.2. Preliminaries m cover of a cover C, denoted as C m , is the set containing all m cubes in C.The corresponding compatibility graph is illustrated in Fig. 1. Wen-Zen Shen, Juinn-Dar Huang, Shih-Min Chao |
DAC | 2 |
| 1995 | Compatible class encoding in Roth-Karp decomposition for two-output LUT architectureabstractRoth-Karp decomposition is one of the most popular techniques for LUT-based FPGA technology mapping because it can decompose a node into a set of nodes with fewer numbers of fanins. In this paper, we show how to formulate the compatible class encoding problem in Roth-Karp decomposition as a symbolic-output encoding problem in order to exploit the feature of the two-output LUT architecture. Based on this formulation, we also develop an encoding algorithm to minimize the number of LUT's required to implement the logic circuit. Experimental results show that our encoding algorithm can produce promising results in the logic synthesis environment for the two-output LUT architecture. Juinn-Dar Huang, Jing-Yang Jou, Wen-Zen Shen |
ICCAD | 1 |