EDBT 2026 Demo / reviewers in the wild / expert
TingTing Hwang
dblp:56/4092 · also Ting Ting Hwang, Tingting Hwang
· DBLP profile ↗
95ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-7206-560XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 94 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 19 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Using OFF-set only for Corrupting Circuit to Resist Structural Attack in CAC LockingabstractCorrupt-and-Correct (CAC) Logic Lockings [1]–[4] are state-of-the-art hardware security techniques designed to protect IC/IP designs from IP piracy, reverse engineering, overproduction, and unauthorized use. Although these techniques are resilient to SAT-based attacks, they remain vulnerable to structural attacks, which exploit structural traces left by the synthesis tool to recover the original form. In this paper, we will propose a novel method that uses only the OFF-set to corrupt the circuit. This approach helps the added circuitry better merge with the original circuit, thereby thwarting structural attacks while maintaining resilience to SAT-based attacks. Additionally, we demonstrate that our proposed method can incur less area overhead compared to previous locking methods in HIID [5]. Compared to SFLL-rem [4], our method can achieve comparable area overhead while effectively resisting structural attacks, including Valkyrie [6] and SPI attacks [7]. Hsiang-Chun Cheng, TingTing Hwang |
DATE | 3 |
| 2025 | Expanding and Obfuscating In-Cone Trees to Resist SAT Attack in Logic LockingabstractHardware security is crucial to protect the confidentiality and integrity of circuit designs. One of the techniques used for this purpose is logic locking, which safeguards against piracy, overuse, and reverse engineering. Logic locking protects a circuit by introducing extra key gates to obfuscate the circuit’s functionality such that the circuit operates the correct function only when the correct key is applied. Recently, researchers have found that obfuscating a point function, such as anand-tree, within a circuit can effectively resist the powerful SAT-based attack method. Although the obfuscation techniques are effective in securing circuit designs, they may suffer from two issues: First, the tree size in the circuit may not be sufficient to achieve the desired level of security, making the locked circuit vulnerable to SAT attack. Second, the obfuscation can be defeated in a single iteration if the SAT solver finds a specific input, referred to as the remove-all distinguishing input pattern (DIP). In this article, we address the two issues. A split-compensate operation is proposed to expand an obfuscated tree. Moreover, by selecting internal variables using a fault-based method, our method mitigates the remove-all DIP issue, and thereby increase the average-case SAT iteration count. The experimental evaluation confirms that the proposed methods are effective in defending against SAT attacks on benchmarks from MCNC, ISCAS’85, EPFL, and ITC’99. In addition, SAT attack fails to break the majority of the benchmarks within a 48 h runtime. Furthermore, our method can defend against removal attack, SPI attack, and Valkyrie attack as well. Li-Nung Hsu, Yung-Chih Chen, TingTing Hwang |
IEEE Trans. Reliab. | 4 |
| 2024 | A Hybrid Approach to Reverse Engineering on Combinational CircuitsabstractReverse engineering is a process that converts low-level description to high-level one. In this paper, we propose a hybrid approach consisting of structural analysis and black-box testing to reverse engineering on combinational circuits. Our approach is able to convert combinational circuits from gate-level netlist to Register-Transfer Level (RT-level) design accurately and efficiently. We developed our approach and participated in Problem A of the 2022 CAD Contest @ ICCAD. The revised version of our program successfully converted most cases and achieved higher scores than the 1stplace team in the contest. Wuqian Tang, Yi-Ting Li, Kai-Po Hsu, Kuan-Ling Chou, You-Cheng Lin, Chia-Feng Chien, Tzu-Li Hsu, Yung-Chih Chen, Ting-Chi Wang, Shih-Chieh Chang 0001, TingTing Hwang, Chun-Yao Wang |
DATE | 11 |
| 2024 | A Cost-Driven Chip Partitioning Method for Heterogeneous 3D IntegrationabstractThree-dimensional integration circuit (3D IC) offers significant benefits in terms of performance and cost. Existing research in through-silicon via (TSV)-based 3D IC partitioning has focused on minimizing the number of TSVs to reduce costs. Partitioning methods based on heterogeneous integration have emerged as viable approaches for cost optimization. Leveraging mature processes to manufacture not timing-critical blocks can yield cost benefits. Nevertheless, none of the previous 3D partitioning work has focused on reducing the overall cost, including both design and manufacturing costs, for heterogeneous 3D integration. Moreover, throughput constraints have not been considered. This article presents a cost-aware integer linear programming-based formulation and a heuristic algorithm that partition the functional blocks in the design into different technological groups. Each group of functional blocks will be implemented using a particular process technology, and then integrated into a 3D IC. Our results show that 3D heterogeneous integration chip implementation can reduce overall cost while satisfying various timing constraints. Cheng-Hsien Lin, Kuan-Ting Chen, Yi-Yu Liu, Allen C.-H. Wu, TingTing Hwang |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2023 | Expanding In-Cone Obfuscated Tree for Anti SAT AttackabstractLogic locking is a hardware security technology to protect circuit designs from overuse, piracy, and reverse engineering. It protects a circuit by inserting key gates to hide the circuit functionality, so that the circuit is functional only when a correct key is applied. In recent years, encrypting the point function, e.g., AND-tree, in a circuit has been shown to be promising to resist SAT attack. However, the encryption technique may suffer from two problems: First, the tree size may not be large enough to achieve desired security. Second, SAT attack could break the encryption in one iteration when it finds a specific input pattern, called remove-all DIP. Thus, in this paper, we present a new method for constructing the obfuscated tree. We first apply the sum-of-product transformation to find the largest AND-tree in a circuit, and then insert extra variables with the proposed split-compensate operation to further enlarge the AND-tree and mitigate the remove-all DIP issue. The experimental results show that the proposed obfuscated tree can effectively resist SAT attack. Li-Nung Hsu, Yung-Chih Chen, TingTing Hwang |
DATE | 4 |
| 2021 | A Dynamic Link-latency Aware Cache Replacement Policy (DLRP)abstractMultiprocessor system-on-chips (MPSoCs) in modern devices have mostly adopted the non-uniform cache architecture (NUCA) [1], which features varied physical distance from cores to data locations and, as a result, varied access latency. In the past, researchers focused on minimizing the average access latency of the NUCA. We found that dynamic latency is also a critical index of the performance. A cache access pattern with long dynamic latency will result in a significant cache performance degradation without considering dynamic latency. We have also observed that a set of commonly used neural network application kernels, including the neural network fully-connected and convolutional layers, contains substantial accessing patterns with long dynamic latency. This paper proposes a hardware-friendly dynamic latency identification mechanism to detect such patterns and a dynamic link-latency aware replacement policy (DLRP) to improve cache performance based on the NUCA. Yen-Hao Chen, Allen C.-H. Wu, TingTing Hwang |
ASP-DAC | 3 |
| 2017 | Communication driven remapping of processing element (PE) in fault-tolerant NoC-based MPSoCsabstractWe propose a remapping algorithm to tolerate the failures of Processing Elements (PEs) on Multiprocessor System-on-Chip. A new graph modeling is proposed to precisely define the increase of communication cost among PEs after remapping. Our method can be used not only to repair faults but also to improve the communication cost of given initial mapping results. Experimental results show that under multiple failures, the communication cost by our method is 43.59% less on average compared with that by previous work [1] using the same number of spare PEs. Moreover, the communication cost is further reduced by 4.16% after applying our method to initial mappings produced by NMAP [2]. Chia-Ling Chen, Yen-Hao Chen, TingTing Hwang |
ASP-DAC | 3 |
| 2017 | Architectural evaluations on TSV redundancy for reliability enhancementabstractThree-dimensional Integrated Circuits (3D-ICs) is a next-generation technology that could be a solution to overcome the scaling problem. It stacks dies with Through-Silicon Vias (TSVs) so that signals can be transmitted through dies vertically. However, researchers have noticed that the aging effect due to the electormigration (EM) may result in faulty TSVs and affect the chip lifetime [1]. Several redundant TSV architectures have been proposed to address this issue. By replacing the faulty TSV with redundant TSVs which are added at design time, chips can achieve better reliability and longer lifetime. In this paper, we will study the tradeoff of various redundant TSV architectures in terms of effectiveness and cost. To allow the measurement of reliability more realistically, we propose a new standard, repair rate, to appraise the redundant TSV architectures. Moreover, to design a more flexible and efficient structure, we enhance the ring-based design [2] that can adjust the size of the TSV block and TSV redundancy. Yen-Hao Chen, Chien-Pang Chiu, Russell Barnes, TingTing Hwang |
DATE | 4 |
| 2017 | A Novel Cache-Utilization-Based Dynamic Voltage-Frequency Scaling Mechanism for Reliability EnhancementsabstractWe propose a cache architecture using a 7T/14T SRAM (Fujiwara et al., 2009) and a control mechanism for reliability enhancements. Our control mechanism differs from conventional dynamic voltage-frequency scaling (DVFS) methods in that it considers not only the cycles per instruction behaviors but also the cache utilization. To measure cache utilization, a novel metric is proposed. The experimental results show that our proposed method achieves 1000 times less bit-error occurrences compared with conventional DVFS methods under the ultralow-voltage operation. Moreover, the results indicate that our proposed method surprisingly not only incurs no performance and energy overheads but also achieves on average a 2.10% performance improvement and a 6.66% energy reduction compared with conventional DVFS methods. Yen-Hao Chen, Yi-Lun Tang, Yi-Yu Liu, Allen C.-H. Wu, TingTing Hwang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | A novel cache-utilization based dynamic voltage frequency scaling (DVFS) mechanism for reliability enhancements
Yen-Hao Chen, Yi-Lun Tang, Yi-Yu Liu, Allen C.-H. Wu, TingTing Hwang |
DATE | 5 |
| 2016 | Thermal-aware dynamic page allocation policy by future access patterns for Hybrid Memory Cube (HMC)
Wei-Hen Lo, Kai-zen Liang, TingTing Hwang |
DATE | 3 |
| 2016 | Architecture of Ring-Based Redundant TSV for Clustered FaultsabstractThree-dimensional integrated circuits (3-D-ICs) that employ the through-silicon vias (TSVs) vertically stacking multiple dies provide many benefits, such as high density, high bandwidth, and low power. However, the fabrication and bonding of TSVs may fail because of many factors, such as the winding level of the thinned wafers, the surface roughness and cleanness of silicon dies, and bonding technology. To improve the yield of 3-D-ICs, many redundant TSV (RTSV) architectures were proposed to repair 3-D-ICs with faulty TSVs. These methods reroute the signals of faulty TSVs to other regular TSV or RTSV. In practice, the faulty TSVs may cluster because of imperfect bonding technology. To resolve the problem of clustered TSV faults, router-based RTSV architecture was the first proposed to pay attention to it. Their method enables faulty TSVs to be repaired by RTSVs that are farther apart. However, to repair some rarely occurring defective patterns, their method requires too much area. In this paper, we propose a ring-based RTSV architecture to utilize the area more efficiently as well as to maintain high yield. Assume that the size of the TSVs is 1 μm. Simulation results show that for a given number of TSVs (8 × 8) and TSV failure rate (1%), our design achieves 58.9% area reduction of MUXes per signal, 54.6% total area reduction per signal, and 50.54% total wire length reduction while the yield of our ring-based RTSV architectures can still maintain 98.47%-99% as compared with the router-based design. Furthermore, the shifting length of our ring-based RTSV architecture is at most 1, which guarantees at most one MUX-delay timing overhead of each signal. Wei-Hen Lo, Kang Chi, TingTing Hwang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Architecture of ring-based redundant TSV for clustered faults
Wei-Hen Lo, Kang Chi, TingTing Hwang |
DATE | 3 |
| 2014 | Fault-tolerant TSV by using scan-chain test TSVabstractIn order to increase the yield of 3-D IC, fault-tolerance technique to recover failed TSV is essential. In this paper, an architecture of TSV recovery by using scan-chain test TSV is proposed. With the architecture, only a small amount of redundant TSVs is required to be inserted. Extra TSV area that occurs by our method is much less than that of other methods. Moreover, a 3-D IC scan-chain optimization algorithm is proposed taking into consideration the locations of functional TSVs as well as test TSVs, so that the number of total TSVs including test TSV and extra redundant TSV of a 3-D IC design is effectively reduced. Fu-Wei Chen, Hui-Ling Ting, TingTing Hwang |
ASP-DAC | 3 |
| 2014 | Compaction-free compressed cache for high performance multi-core systemabstractCompressed cache was used in shared last level cache (LLC) to increase the effective capacity. However, because of various data compression sizes, fragmentation problem of storage is inevitable in this cache design. When it happens, usually, a compaction process is invoked to make contiguous storage space. This compaction process induces extra cycle penalty and degrades the effectiveness of compressed cache design. In this paper, we propose a compaction-free compressed cache architecture which can completely eliminate the time for executing compaction. Based on this cache design, we demonstrate that our results, compared with the conventional cache, have system performance improvement by 16% and energy reduction by 16%. Compared with the work by Alameldeen et al. [1], our design has 5% more performance improvement and 3% more energy reduction. Compared with the work by Sardashti et al. [2], our design has 3% more performance improvement and 2% more energy reduction. Po-Yang Hsu, Pei-Lan Lin, TingTing Hwang |
ICCAD | 3 |
| 2014 | Clock-Tree Synthesis with Methodology of Reuse in 3D-ICabstractIP reuse methodology has been used extensively in SoC (system-on-chip) design. In this reuse methodology, while design and implementation costs are saved, manufacturing cost is not. To further reduce the cost, this reuse concept has been proposed at mask and die level in three-dimensional integrated circuits (3D-IC). In order to achieve manufacturing reuse, in this article, we propose a new methodology for designing a global clock tree in 3D-IC. The objective is to extend an existing clock tree in 2D IC to 3D IC, taking into consideration the wirelength, clock skew, and the number of TSVs. Compared with NNG- and 3D-MMM-based methods, our proposed method reduces the wirelength of the new die and the skew of the global 3D clock tree on average, 5.85% and 2.3%, and 76.92% and 48.7%, respectively. In more than two die design, the average improvements of the wirelength and clock skew of our method as compared with the 3D-MMM-based method are 4.23% and 46.84%, respectively. Fu-Wei Chen, TingTing Hwang |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2014 | Stacking Signal TSV for Thermal Dissipation in Global Routing for 3-D ICabstractWith no further shrinking of device size, 3-D chip stacking by through-silicon-via (TSV) has been identified as an effective way to achieve better performance in speed and power. However, such solution inevitably encounters challenges in thermal dissipation since stacked dies generate a significant amount of heat per unit volume. We leverage an integrated design methodology of stacked-signal-TSVs to minimize temperature. Based on this structure, a three-stage TSV locating algorithm in global routing is designed. We demonstrate that our results, compared with baseline circuits, have 17% temperature reduction with 3% wiring overhead and no performance loss calculated by 3-D Elmore delay model. Compared with the paper by Cong and Zhang where additional thermal TSVs are inserted, our experimental results have in average 23% less TSVs with the same temperature constraint. Compared with the paper by Pathak and Lim, where movable signal TSVs are relocated to reduce temperature in hotspot regions, our result has 8% more temperature reduction with the same number of signal TSVs. Po-Yang Hsu, Hsien-Te Chen, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | Utilizing Circuit Structure for Scan Chain DiagnosisabstractScan chain diagnosis has become a critical issue to yield loss in modern technology. In this paper, we present a scan chain partitioning algorithm and a scan chain reordering algorithm to improve scan chain fault diagnosis resolution. In our scan chain partition algorithm, we take into consideration not only logic dependency but also the controllability between scan flip-flops. After the partition step, the ordering of scan cells is performed to decrease the range of suspect faulty scan cells by a bipartite matching reordering algorithm. The experimental results show that our method can reduce the number of suspect scan cells from 378-31 to at most 3 for most cases of ITC'99 benchmarks. Wei-Hen Lo, Ang-Chih Hsieh, Chien-Ming Lan, Min-Hsien Lin, TingTing Hwang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2013 | Stacking signal TSV for thermal dissipation in global routing for 3D ICabstractWith no further shrink of device size, three dimensional (3D) chip stacking by Through-Silicon-VIA (TSV) has been identified as an effective way to achieve better performance in speed and power. However, such solution inevitably encounters challenges in thermal dissipation since stacked dies generate significant amount of heat per unit volume. We leverage an integrated architecture of stacked-signal-TSVs to minimize temperature with small wiring overhead. Based on the structure of stacked signal TSV, a two-stage TSV locating algorithm in global routing is designed. By this TSV locating algorithm, we demonstrate that our stacking signal TSV structure is able to reduce 17% temperature with 4% wiring overhead and 3% performance loss calculated by 3D Elmore delay model. Compared to a previous work by Cong and Zhang [1] where additional thermal TSVs are inserted, our experimental results have in average 23% less TSVs than Cong and Zhang's [1] with the same temperature constraint. Po-Yang Hsu, Hsien-Te Chen, TingTing Hwang |
ASP-DAC | 3 |
| 2013 | Utilizing circuit structure for scan chain diagnosisabstractScan chain diagnosis has become a critical issue to yield loss in modern technology. In this paper, we present a scan chain partitioning algorithm and a scan chain reordering algorithm to improve scan chain fault diagnosis resolution. In our scan chain partition algorithm, we take into consideration not only logic dependency but also the controllability between scan flip flops. After partition step, the ordering of scan cells is performed to decrease the range of suspect faulty scan cells by a bipartite matching reordering algorithm. The experimental results show that our method can reduce the number of suspect scan cells from 378-31 to at most 3 for most cases of ITC'99 benchmarks. Wei-Hen Lo, Ang-Chih Hsieh, Chien-Ming Lan, Min-Hsien Lin, TingTing Hwang |
ETS | 5 |
| 2013 | Thread-criticality aware dynamic cache reconfiguration in multi-core systemabstractTo alleviate high energy dissipation of cache memory, some research has proposed to reconfigure cache parameters such as cache capacity, number of way associative, and cache line size during program phase changes. However, none of previous research on cache reconfiguration takes thread criticality into consideration. In this paper, we dynamically predict thread criticality of a parallel application and tune our cache memory architecture accordingly. The experimental results show that our method not only reduces 42% energy consumption, but also improves the system performance by 4% compared to the baseline cache setting without reconfiguration. Compared with the work by Chen et al. [1] where cache capacity is configured based on its hit count, our method yields extra 16% energy reduction and 7% performance improvement. Compared with the work by Gordon-Ross et al. [2] where cache always select the configuration with the minimum energy consumption for the current interval, our result has 8% more energy reduction and 12% more performance improvement. Po-Yang Hsu, TingTing Hwang |
ICCAD | 2 |
| 2013 | Thermal-aware memory mapping in 3D designs
Ang-Chih Hsieh, TingTing Hwang |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2012 | Clock tree synthesis with methodology of re-use in 3D ICabstractIP reuse methodology has been used extensively in SoC (System on Chip) design. In this reuse methodology, while design and implementation cost is saved, manufacturing cost is not. To further reduce the cost, this reuse concept has been proposed at mask and die level in three-dimension integrated circuit (3D IC). In order to achieve manufacturing reuse, in this paper, we propose a new methodology to design a global clock tree in 3D IC. The objective is to extend an existing clock tree in 2D IC to 3D IC taking into consideration the wirelength, clock skew and the number of TSVs. Compared with NNG-based method, our proposed method reduces the wirelength of the new die and skew of the global 3D clock tree, on an average, 47.16% and 5.85%, respectively. Fu-Wei Chen, TingTing Hwang |
DAC | 2 |
| 2012 | A Physical-Location-Aware X-Bit Redistribution for Maximum IR-Drop ReductionabstractTo guarantee that an application-specific integrated circuit (ASIC) meets its timing requirement, at-speed scan testing becomes an indispensable procedure for verifying the performance of ASIC. However, at-speed scan test suffers the test-induced yield loss. Because the switching-activity in test mode is much higher than that in normal mode, the switching-induced large current drawn causes severe IR drop and increases gate delay. X-filling is the most commonly used technique to reduce IR-drop effect during at-speed test. However, the effectiveness of X-filling depends on the number and the characteristic of X-bit distribution. In this paper, we propose a physical-location-aware X-identification which redistributes X-bits so that the maximum switching-activity is guaranteed to be reduced after X-filling. We estimate IR-drop using RedHawk tool and the experimental results on ITC'99 show that our method has an average of 9.42% more reduction of maximum IR-drop as compared to a previous work which redistributes X-bits evenly in all test vectors. Fu-Wei Chen, Shih-Liang Chen, Yung-Sheng Lin, TingTing Hwang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | TSV Redundancy: Architecture and Design Issues in 3-D ICabstract3-D technology provides many benefits including high density, high bandwidth, low-power, and small form-factor. Through Silicon Via (TSV), which provides communication links for dies in vertical direction, is a critical design issue in 3-D integration. Just like other components, the fabrication and bonding of TSVs can fail. A failed TSV can severely increase the cost and decrease the yield as the number of dies to be stacked increases. A redundant TSV architecture with reasonable cost is proposed in this paper. Based on probabilistic models, some interesting findings are reported. First, the number of failed TSVs in a tier is usually less than 2 when the number of TSVs in a tier is less than 1000 and less than 5 when the number of TSVs in a tier is less than 10000. Assuming that there are at most 2-5 failed TSVs in a tier. With one redundant TSV allocated to one TSV block, our proposed structure leads to 90% and 95% recovery rates for TSV blocks of size 50 and 25, respectively. Finally, analysis on overall yield shows that the proposed design can successfully recover most of the failed chips and increase the yield of TSV to 99.4%. Ang-Chih Hsieh, TingTing Hwang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Run-Time Reconfiguration of Expandable Cache for Embedded SystemsabstractExpandable cache proposed by Bournoutian and Orailoglu is very efficient in reducing miss rate and energy consumption with small area overhead. However, the original expandable cache with only one expansion scheme may lead to thrashing problems. In this work, based on the structure of expandable cache, we will introduce a new cache design which has many expansion schemes to fit different run-time program behaviors. The expansion scheme of our proposed cache is dynamically changed by executing configuration instructions which are inserted at compile time. The experimental results of SPEC CPU2000 have shown that our proposed cache design effectively improves the miss rate by 14.74% as compared with the original expandable cache. In terms of energy improvement ratio, our method is 5.62% higher than that of expandable cache when the baseline is set as the energy consumption of 2-way set-associative cache. Ang-Chih Hsieh, TingTing Hwang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | A physical-location-aware fault redistribution for maximum IR-drop reductionabstractTo guarantee that an application specific integrated circuits (ASIC) meets its timing requirement, at-speed scan testing becomes an indispensable procedure for verifying the performance of ASIC. However, at-speed scan test suffers the test-induced yield loss. Because the switching activity in test mode is much higher than that in normal mode, the switching-induced large current drawn causes severe IR drop and increases gate delay. X-filling is the most commonly used technique to reduce IR-drop effect during at-speed test. However, the effectiveness of X-filling depends on the number and the characteristic of X-bit distribution. In this paper, we propose a physical-location-aware X-identification which redistributes faults so that the maximum switching activity is guaranteed to be reduced after X-filling. The experimental results on ITC'99 show that our method has an average of 8.54% more reduction of maximum IR-drop as compared to a previous work which re-distributes X-bits evenly in all test vectors. Fu-Wei Chen, Shih-Liang Chen, Yung-Sheng Lin, TingTing Hwang |
ASP-DAC | 4 |
| 2011 | Enhanced Heterogeneous Code Cache management scheme for Dynamic Binary TranslationabstractRecently, Dynamic Binary Translation (DBT) technology has gained much attentions on embedded systems due to its various capabilities. However, the memory resource in embedded systems is often limited. This leads to the overhead of code retranslation and causes significant performance degradation. To reduce this overhead, Heterogeneous Code Cache (HCC), is proposed to split the code cache among SPM and main memory to avoid code re-translation. Although HCC is effective in handling applications with large working sets, it ignores the execution frequencies of program segments. Frequently executed program segments can be stored in main memory and suffer from large access latency. This causes significant performance loss. To address this problem, an enhanced Heterogeneous Code Cache management scheme which considers program behaviors is proposed in this paper. Experimental results show that the proposed management scheme can effectively improve the access ratio of SPM from 49.48% to 95.06%. This leads to 42.68% improvement of performance as compared with the management scheme proposed in the previous work. Ang-Chih Hsieh, Chun-Cheng Liu, TingTing Hwang |
ASP-DAC | 3 |
| 2011 | A new architecture for power network in 3D ICabstractProviding high vertical interconnection density between device tiers, through silicon via (TSV) offers a promising solution in 3D IC to reduce the length of global interconnection. However, some design issues hinder TSV from volumes of adoption, such as IR drop, thermal dissipation, current delivery per package pin and various voltage domains among tiers. To tackle these problems, the design of power network plays an important role in 3D IC. A new integrated architecture of stacked-TSV and power distributed network (STDN) is proposed in this paper. Our new STDN serves triple roles: power network to deliver larger current and reduce IR drop, thermal network to reduce temperature, and decoupling capacitor network to reduce power noise. As well, it helps to alleviate the limitation of the number of IO power pins. For both single and multiple power domains, the proposed STDN architecture demonstrates good performance in 3D floorplan, IR drop, power noise, temperature, area and even the total length of signal connections for selected MCNC benchmarks. Hsien-Te Chen, Hong-Long Lin, TingTing Hwang |
DATE | 4 |
| 2011 | Memory Mapping and Task Scheduling Techniques for Computation Models of Image Processing on Many-Core PlatformsabstractMany-core technology is proposed as a solution to improve the performance of modern computer systems. To obtain good performance on a many-core system, exploiting parallelism in arithmetic level is not enough. Due to the contention of shared hardware resource, the speedup ratio of a many-core system is usually much lower than the number of processor units. In this paper, the contention of shared memory resource is addressed. An algorithm is developed to perform memory mapping and task scheduling for many-core systems. According to experimental results, the proposed algorithm can effectively improve the performance by 64.77% in average. 43.55X speedup ratio can be achieved when 48 processor units are activated. The performance loss ratio is less than 10%. Ang-Chih Hsieh, Yi-Ta Wu, Shau-Yin Tseng, TingTing Hwang |
ICPP | 4 |
| 2011 | Through-Silicon Via Planning in 3-D FloorplanningabstractIn this paper, we will study floorplanning in 3-D integrated circuits (3D-ICs). Although literature is abundant on 3D-IC floorplanning, none of them consider the areas and positions of signal through-silicon vias (TSVs). In previous research, signal TSVs are viewed as points during the floorplanning stage. Ignoring the areas, positions and connections of signal TSVs, previous research estimates wirelength by measuring the half-perimeter wirelength of pins in a net only. Experimental results reveal that 29.7% of nets possess signal TSVs that cannot be put into the white space within the bounding boxes of pins. Moreover, the total wirelength is underestimated by 26.8% without considering the positions of signal TSVs. The considerable error in wirelength estimation severely degrades the optimality of the floorplan result. Therefore, in this paper, we will propose a two-stage 3-D fixed-outline floorplaning algorithm. Stage one simultaneously plans hard macros and TSV-blocks for wirelength reduction. Stage two improves the wirelength by reassigning signal TSVs. Experimental results show that stage one outperforms a post-processing TSV planning algorithm in successful rate by 57%. Compared to the post-processing TSV planning algorithm, the average wirelength of our result is shorter by 22.3%. In addition, stage two further reduces the wirelength by 3.45% without any area overhead. Ming-Chao Tsai, Ting-Chi Wang, TingTing Hwang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | TSV redundancy: Architecture and design issues in 3D ICabstract3D technology provides many benefits including high density, high band-with, low-power, and small form-factor. Through Silicon Via (TSV), which provides communication links for dies in vertical direction, is a critical design issue in 3D integration. Just like other components, the fabrication and bonding of TSVs can fail. A failed TSV may cause a number of known-good-dies that are stacked together to be discarded. This can severely increase the cost and decrease the yield as the number of dies to be stacked increases. A redundant TSV architecture with reasonable cost for ASICs is proposed in this paper. Design issues including recovery rate and timing problem are addressed. Based on probabilistic models, some interesting findings are reported. First, the probability that three or more TSVs are failed in a tier is less than 0.002%. Assumption of that there are at most two failed TSVs in a tier is sufficient to cover 99.998% of all possible faulty free and faulty cases. Next, with one redundant TSV allocated to one TSV block, limiting the number of TSVs in each TSV block to be no greater than 50 and 25 leads to 90% and 95% recovery rates when 2 failed TSVs are assumed. Finally, analysis on overall yield shows that the proposed design can successfully recover most of the failed chips and increase the yield of TSV bonding to 99.99%. This can effectively reduce the cost of manufacturing 3D ICs. Ang-Chih Hsieh, TingTing Hwang, Ming-Tung Chang, Min-Hsiu Tsai, Chih-Mou Tseng, Hung-Chun Li |
DATE | 2 |
| 2010 | A Physical-Location-Aware X-Filling Method for IR-Drop Reduction in At-Speed Scan TestabstractThe IR-drop problem during test mode exacerbates delay defects and results in false failures. In this paper, we take theX-fillingapproach to reduce the IR-drop effect during an at-speed test. The main difference between our approach and the previous X-filling approaches lies in two aspects. The first one is that we take the spatial information into consideration in our approach. The second one is how X-filling is performed. We propose a backward-propagation technique instead of a forward-propagation approach taken in previous work. The experimental results show that our approach can reduce 21.1% of the maximum IR-drop in the best case and 9.1% on the average as compared to previous work. Wen-Wen Hsieh, Shih-Liang Chen, I-Sheng Lin, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2010 | Reconfigurable ECO Cells for Timing Closure and IR Drop MinimizationabstractUnused spare cells occur inevitably in traditional engineering change order (ECO) design flow. It results in inefficient area usage, more leakage, and more IR drop impacts. To tackle these problems, a reconfigurable cell is proposed, which serves the dual purposes of decoupling capacitance and spare cell in this paper. Before ECO is applied, these cells are preplaced as decoupling capacitors. When ECO is applied, these cells are configured as functional cells. To demonstrate the efficiency of our configurable cell, we propose an algorithm for timing closure and IR drop minimization. Compared with traditional ECO flow, our method shows 15% reduction in maximum IR drop and 9% reduction in leakage before applying ECO, and 7% reduction in maximum IR drop after applying ECO, with 10% area of spare cells. In addition, we show that there remain less unsolved timing-violation paths after applying our ECO timing optimization flow due to less IR drop and free selection of ECO gate type. Hsien-Te Chen, Chieh-Chun Chang, TingTing Hwang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Thermal-aware post compilation for VLIW architecturesabstractDevelopment of a thermal management method to reduce hotspots and to balance the temperature distribution has become an important issue. In this paper, we propose a static thermal management technique at compiler level. The target machine is a VLIW architecture where the compiler is required to schedule instructions to achieve instruction level parallelism (ILP). Two technique are proposed. The first one is register binding to balance the temperature of the register file by taking both spatial and temporal thermal information into consideration. The second one is forwarding methods including forwarding-aware architecture and instruction scheduling to reduce the access count of register file. The experimental results show that by combining the two techniques, the peak temperature reduction can reach 7.89 (°C) in the best case and 7.22 (°C) in average with only 0.9% performance penalty in average. Wen-Wen Hsieh, TingTing Hwang |
ASP-DAC | 2 |
| 2009 | New spare cell design for IR drop minimization in Engineering Change OrderabstractUnused spare cells occur inevitably in traditional ECO design flow. It results in inefficient area usage, more leakage, and more IR drop impacts. To tackle these problems, a reconfigurable cell is proposed which serves the dual purposes of decoupling capacitance and spare cell in this paper. Before Engineering Change Order (ECO) is applied, these cells are pre-placed as decoupling capacitors. When ECO is applied, these cells are configured as functional cells. To demonstrate the efficiency of our configurable cell, we propose an algorithm for timing closure and IR drop minimization. Compared with traditional ECO flow, our method shows 16% reduction in maximum IR drop and 56% reduction in leakage before applying ECO, and 8% reduction in maximum IR drop after applying ECO, with 10% area of spare cells. In addition, we show that there are less unsolved ECO timing paths left after applying our ECO timing optimization algorithm due to free selection of ECO gate type. Hsien-Te Chen, Chieh-Chun Chang, TingTing Hwang |
DAC | 3 |
| 2009 | Thermal-aware memory mapping in 3D designsabstractDRAM is usually used as main memory for program execution. The thermal behavior of a memory block in a 3D SIP is affected not only by the power behavior but also the heat dissipating ability of that block. The power behavior of a block is related to the applications run on the system while the heat dissipating ability is determined by the number of tier and the position the block locates. Therefore, a thermal-aware memory allocator should consider the following two points. First, allocator should consider not only the power behavior of a memory block but also the physical location during memory mapping, second, the changing temperature of a physical block during execution of programs. In this paper, we will propose a memory mapping algorithm taking into consideration the above-mentioned two points. Our technique can be classified as static thermal management to be applied to embedded software designs. Experiments show that our method can reduce temperature of memory system by 17.2degC as compared to a straightforward mapping in the best case, and 13.4degC in average. Ang-Chih Hsieh, TingTing Hwang |
DATE | 2 |
| 2009 | A physical-location-aware X-filling method for IR-drop reduction in at-speed scan testabstractIR-drop problem during test mode exacerbates delay defects and results in false failures. In this paper, we take the X-filling approach to reduce IR-drop effect during at-speed test. The main difference between our approach and the previous X-filling methods [7]–[9] lies in two aspects. The first one is that we take the spatial information into consideration in our approach. The second one is how X-filling is performed. We propose a backward-propagation approach instead of a forward-propagation approach taken in previous work. The experimental results show that we have 42.81% reduction for the worst IR-drop and 45.71% reduction in the average IR-drop as compared to random fill method. Wen-Wen Hsieh, I-Sheng Lin, TingTing Hwang |
DATE | 3 |
| 2009 | Leakage reduction, delay compensation using partition-based tunable body-biasing techniquesabstractIn recent years, fabrication technology of CMOS has scaled to nanometer dimensions. As scaling progresses, several new challenges follow. Among them, the most noticeable two are process variations and leakage current of the circuit. To tackle the problems of process variations and leakage current, an effective way is to use a body-biasing technique. In substance, using the RBB technique can minimize leakage current but increase the delay of a gate. Contrary to RBB, the FBB technique decreases the delay but increases leakage current of a gate. In the previous work, a single body-biasing is applied to the whole circuit. In a slow circuit, since the FBB is applied to the whole circuit, the leakage current of all gates in the circuit increases dramatically. On the other hand, in a fast circuit, RBB is applied to decrease the leakage current. However, without violating the timing specification, the value of body-biasing is restricted by the critical paths, and the saving of leakage current is limited. In this article, we propose a design flow to partition the circuit into subcircuits so that each subcircuit can be applied its individual RBB or FBB. Experiments show that our method is able to save leakage current from 42% to 47% as compared to designs not using a body-biasing technique. Under process variations, our method can save 42% to 49% leakage on fast circuits and 20% to 35% on slow circuits. Po-Yuan Chen, Chiao-Chen Fang, TingTing Hwang, Hsi-Pin Ma |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2009 | Skew-aware polarity assignment in clock treeabstractIn modern sequential VLSI designs, clock tree plays an important role in synchronizing different components in a chip. To reduce peak current and power/ground noises caused by clock network, assigning different signal polarities to clock buffers is proposed in previous work. Although peak current and power/ground noises are minimized by signal polarities assignment, an assignment without timing information may increase the clock skew significantly. As a result, a timing-aware signal polarities assigning technique is necessary. In this article, we propose a novel signal polarities assigning technique which can not only reduce peak current and power/ground noises simultaneously but also render the clock skew in control. The experimental result shows that the clock skew produced by our algorithm is 94% of original clock skew in average while the clock skews produced by three algorithms (Partition, MST, Matching) in the absence of post clock tuning steps in the previous work are 235%, 272%, and 283%, respectively. Moreover, our algorithm is as efficient as the three algorithms of the previous work in reducing peak current and power/ground noises. Po-Yuan Chen, Kuan-Hsien Ho, TingTing Hwang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2008 | Transition-aware decoupling-capacitor allocation in power noise reductionabstractDynamic power noises may not only degrade the circuit performance but also reduce the noise margin which may result in the functional errors in integrated circuit. Decoupling capacitor (decap) allocation is one of the most effective way in reducing serious dynamic power noises (hotspots). To allocate decap before placement, we observed that not only locations but also rising time of functional cells are required to accurately predict power noises. Compared to a previous work which only takes neighborhood relation into consideration, our method is more efficient in reducing hotspots. Furthermore, to reduce the hotspots after placement, instead of only using the empty space as proposed in the previous work, we move out cells in the area with serious power noise area (hot area). The obtained empty space can be used to accommodate decaps to further reduce the hotspots. The experimental result shows, compared to the previous work [1], our estimation function to allocate decap before placement is 23% better in reducing power noises. Moreover, compared to a method which fills decaps to all remaining empty space, our cell move algorithm can almost eliminate all the remaining hot grid nodes and hot cells. In summary, compared to the original circuits (without decap), about 60% of hotspots can be removed using our prediction function before placement, and most of the remaining hotspots are removed by our cell moving step after placement. Po-Yuan Chen, Che-Yu Liu, TingTing Hwang |
ICCAD | 3 |
| 2008 | Synthesis of a novel timing-error detection architectureabstractDelay variation can cause a design to fail its timing specification. Ernst et al. [2003] observe that the worst delay of a design is least probable to occur. They propose a mechanism to detect and correct occasional errors while the design can be optimized for the common cases. Their experimental results show significant performance (or power) gain as compared with the worst-case design. However, the architecture in Ernst et al. [2003] suffers the short path problem, which is difficult to resolve. In this article, we propose a novel error-detecting architecture to solve the short path problem. Our experimental results show considerable performance gain can be achieved with reasonable area overhead. Yu-Shih Su, Po-Hsien Chang, Shih-Chieh Chang 0001, TingTing Hwang |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2007 | Skew aware polarity assignment in clock treeabstractIn modern sequential VLSI designs, clock tree plays an important role in synchronizing different components in a chip. To reduce peak current and power/ground noises caused by clock network, assigning different signal polarities to clock buffers is proposed in previous work. Although peak current and power/ground noises are minimized by signal polarities assignment, an assignment without timing information may increase the clock skew significantly. As a result, a timing-aware signal polarities assigning technique is necessary. In this paper, we propose a novel signal polarities assigning technique which can not only reduce peak current and power/ground noises simultaneously but also render the clock skew in control. The experimental result shows that the clock skew produced by our algorithm is 94% of original clock skew in average while the dock skews produced by three algorithms (Partition. MST, Matching) [5] are 235%, 272%, and 283%, respectively. Moreover, our algorithm is as efficient as the three algorithms of [5] in reducing peak current and power/ground noises. Po-Yuan Chen, Kuan-Hsien Ho, TingTing Hwang |
ICCAD | 3 |
| 2007 | A Bus-Encoding Scheme for Crosstalk Elimination in High-Performance Processor DesignabstractA crosstalk effect leads to increases in delay and power consumption and, in the worst-case scenario, to inaccurate results. With the scale down of technology to deep-submicrometer level, the crosstalk effect between adjacent wires becomes more and more serious, particularly between long on-chip buses. In this paper, we propose a deassembler/assembler technique to eliminate undesirable crosstalk effects on bus transmission. By taking advantage of the prefetch process, where the instruction/data fetch rate is always higher than the instruction/data commit rate, the proposed method incurs almost no penalty in terms of dynamic instruction count. In addition, when the bus width is 128 b, the required number of extra bus wires is only 7 as compared to the 85 extra bus wires needed in the work of Victor and Keutzer. Wen-Wen Hsieh, Po-Yuan Chen, Chun-Yao Wang, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | Performance-Driven Crosstalk Elimination at Postcompiler Level-The Case of Low-Crosstalk Op-Code AssignmentabstractSignificant advances in very large-scale integration process technology have scaled the feature size down. One effect of this scaling down is that coupling capacitances have grown reciprocal in the square of the scaling factor. This crosstalk effect will not only increase the power consumption but also lengthen the propagation delay. Since the data sequences on an instruction bus are known during the compile time, this paper presents two compiler algorithms, rescheduling and renaming, for performance improvement by eliminating crosstalk effects on an instruction bus. The results show that our crosstalk-eliminating postcomplier algorithms significantly reduce the dynamic instruction overhead from 11.50% to 0.52% by eliminating the 4middotC crosstalk. Due to the effective 4middotC crosstalk elimination, our proposed method can improve the instruction fetch time up to 9.59% Wu-An Kuo, Yi-Ling Chiang, TingTing Hwang, Allen C.-H. Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | Crosstalk-Aware Domino-Logic SynthesisabstractWe propose a logic synthesis flow to synthesize a domino-cell network with less crosstalk effect. Crosstalk-immunity property of or gate and relations between wire adjacency and cell I/O are exploited in technology mapping. Meanwhile, a metric to measure the crosstalk sensitivity of domino cells in synthesis level is proposed. Experimental results demonstrate that the crosstalk sensitivity of the synthesized domino-cell network is greatly reduced by 52% using our synthesis flow as compared with conventional methodology. Furthermore, after placement and routing are performed, the ratio of the number of crosstalk-immune wire pairs to the number of total wire pairs is about 24% using our methodology as compared to 9% using conventional techniques, and the maximum wire coupling can be greatly reduced from 95% to 60% Yi-Yu Liu, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Energy-aware scheduling and simulation methodologies for parallel security processors with multiple voltage domains
Yung-Chia Lin, Yi-Ping You, Chung-Wen Huang, Jenq Kuen Lee, Wei-Kuan Shih, TingTing Hwang |
J. Supercomput. | 6 |
| 2007 | A functionality-directed clustering technique for low-power MTCMOS design - computation of simultaneously discharging currentabstractMultithreshold CMOS (MTCMOS) is a circuit style that can effectively reduce leakage power consumption. Sleep transistor sizing is the key issue when a MTCMOS circuit is designed. If the size of sleep transistor is large enough, the circuit performance can surely be maintained but the area and dynamic power consumption of the sleep transistor may increase. On the other hand, if the sleep transistor size is too small, there will be significant performance degradation because of the increased resistance to ground. Previous approaches [Kao et al. 1998; Anis et al. 2002] to designing sleep transistor size are based mainly on mutually-exclusive discharge patterns. However, these approaches considered only the topology of a circuit (i.e., interconnections of nodes in the circuit-graph saving the functionality of node). We observed that any two possible simultaneously switching gates may not discharge at the same time in terms of functionality. Thus, we propose an algorithm to determine how to cluster cells to share sleep transistors, while taking both topology and functionality into consideration. Moreover, one placement refinement algorithm that takes clustering information into account will be presented. At the logic level, the results show that the proposed clustering method can achieve an average of 22% reduction in terms of the number of unit-size sleep transistors as compared to a method that does not consider functionality. At the physical level, two placement results are discussed. The first is produced by a traditional placement tool plus topology check (functionality check) for insertion of sleep transistors. It shows that the functionality check algorithm produces 9% less chip area as compared with the topology check algorithm. The second result is produced by a placement refinement algorithm where the initial placement is done in the first placement experiment. It shows that the placement refinement algorithm achieves 5% more reduction in area at the expense of 4% increase in wire length. Totally, around 14% reduction is achieved by utilizing the clustering information. Ang-Chih Hsieh, Tzu-Teng Lin, Tsuang-Wei Chang, TingTing Hwang |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2006 | Switching-activity driven gate sizing and Vth assignment for low power designabstractPower consumption has gained much saliency in circuit design recently. One design problem is modeled as "under a timing constraint, to minimize power as much as possible". Previous research regarding this problem focused on either minimizing dynamic power by gate sizing, or reducing leakage power by dual threshold voltage assignment on non-critical path. However, given a timing constraint, an optimization algorithm must be able to utilize gate sizing and threshold-voltage assignment interchangeably, in order to minimize total power consumption including dynamic and leakage power in active mode and leakage power in idle mode. We find that switching-activity of a gate plays an important role in making decision as to choosing gate sizing or threshold-voltage assignment for performance improvement. For high switching-activity gates, threshold-voltage assignment should be used while for low switching-activity gates, gate sizing should be utilized. We develop an algorithm to perform gate sizing and threshold-voltage assignment simultaneously taking switching activity into consideration. The results show that under the same timing constraint, our circuits have 16.26%, and 18.53%, improvement of total power as compared to the original circuits for the cases where the percentage of active time are 100%, and 50%, respectively. Yu-Hui Huang, Po-Yuan Chen, TingTing Hwang |
ASP-DAC | 3 |
| 2006 | Crosstalk-aware domino logic synthesisabstractWe propose a logic synthesis flow which utilizes the functionality of circuit to synthesize a domino-cell network which will have more wires crosstalk-immune to each other For that purpose, techniques of output phase flipping and crosstalk-aware technology mapping are used. Meanwhile, metric to measure the crosstalk sensitivity of domino cells in synthesis level is proposed. Experimental results demonstrate that the crosstalk sensitivity of the synthesized domino-cell network is greatly reduced by 51% using our synthesis flow as compared with conventional methodology. Furthermore, after placement and routing are performed, the ratio of the number of crosstalk-immune wire pairs to the number of total wire pairs is about 25% using our methodology as compared to 9% using conventional techniques. Yi-Yu Liu, TingTing Hwang |
DATE | 2 |
| 2006 | Performance-driven crosstalk elimination at post-compiler levelabstractSignificant advances in VLSI process technology have scaled the feature size down. One effect of this scaling down is that coupling capacitances have grown reciprocal in the square of the scaling factor. This crosstalk effect will not only increase the power consumption but also lengthen the propagation delay. Since the data sequences on an instruction bus are known during the compile time, this paper presents two post-compiler algorithms, rescheduling and renaming, for performance improvement by eliminating crosstalk effects on an instruction bus. The results show that our crosstalk-eliminating post-complier algorithms significantly reduce the dynamic instruction overhead from 11.50% to 0.52% by eliminating the 4-C crosstalk. Due to the effective 4-C crosstalk elimination, our proposed method can improve the instruction fetch time up to 9.59%. Wu-An Kuo, Yi-Ling Chiang, TingTing Hwang, Allen C.-H. Wu |
ISCAS | 3 |
| 2006 | Decomposition of instruction decoders for low-power designsabstractDuring the execution of processor instruction, decoding the instructions is a major task in identifying instructions and generating control signals for data paths. In this article, we propose two instruction decoder decomposition techniques for low-power designs. First, by tracing program execution sequences, we propose an algorithm that explores the relations between frequently executed instructions. Second, we propose a two-stage low-power decomposition structure for decoding instructions. Experimental results demonstrate that our proposed techniques achieve an average of 34.18% in power reduction and 12.93% in critical-path delay reduction for the instruction decoder. Wu-An Kuo, TingTing Hwang, Allen C.-H. Wu |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2006 | Crosstalk minimization in logic synthesis for PLAsabstractWe propose a maximum crosstalk effect minimization algorithm that takes logic synthesis into consideration for PLA structures. To minimize the crosstalk effect, a technique for permuting wire is used which contains the following steps. First, product terms are partitioned into long and short sets, and then the product terms in the long and short sets are interleaved. After that, we take advantage of the crosstalk immunity of product terms in the long set to further reduce the maximum coupling capacitance of the PLA. Finally, synthesis techniques such as local and global transformations are taken into consideration to search for a better result. The experiments demonstrate that our algorithm can effectively minimize the maximum coupling capacitance of a circuit by 51% as compared with the original area-minimized PLA without crosstalk effect minimization. Yi-Yu Liu, Kuo-Hua Wang, TingTing Hwang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2006 | A power-driven multiplication instruction-set design method for ASIPsabstractThis paper presents a novel power-driven multiplication instruction-set design method for application-specific instruction-set processors (ASIPs). Based on a dual-and-configurable-multiplier structure, our proposed method devises a multiplication instruction set for low-power ASIPs. Our method exploits the execution sequences of multiplication instructions and effective bit widths of variables to reduce power consumed by redundant multiplication bits while minimizing the multiplication execution time. Experimental results on a set of DSP programs demonstrate that our proposed method achieves significant power reduction (up to 18.53%) and execution time improvement (up to 10.43%) with 18% area overhead. Wu-An Kuo, TingTing Hwang, Allen C.-H. Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Instruction buffering for nested loops in low-power designabstractSeveral loop-buffering techniques were proposed for reducing power consumption of embedded processors. Although the schemes are effective in reducing power, they work for unnested loops (or the inner-most loop in nested loops) only. In this paper, we propose a stack-based controller which can handle sequential loops being nested in a loop of all styles and the if-then-else construct inside of a loop. Our experiments by power estimator Wattch show that the reduction in energy consumption using our technique is up to 36% improvement of the design without buffering technique and has 25% more improvement when compared to the results which handle inner-most loop only at the fetch and decode stages. Chi Ta Wu, Ang-Chih Hsieh, TingTing Hwang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2005 | Functionality directed clustering for low power MTCMOS designabstractMulti-Threshold CMOS (MTCMOS) is a circuit style that can effectively reduce leakage power consumption. Sleep transistor sizing is the key issue when MTCMOS circuit is designed. If the sleep transistor size is too large, the circuit performance can be maintained but the dynamic power consumption of the sleep transistor will increase. On the other hand, if the sleep transistor size is too small, there will be significant performance degradation because of the increased resistance to ground. Previous approach [1, 2] designed the sleep transistor size based on mutual exclusive discharge patterns. However, these approaches considered only topology of a circuit. We observed that two possible simultaneous switching gates may not discharge at the same time in terms of functionality. Thus, we propose an algorithm to determine how to cluster cells to share sleep transistors taking both topology and functionality into consideration. The results show that the proposed method can achieve on the average 18% reduction ratio in terms of the number of sleep transistors as compared to the method without considering functionality. Tsuang-Wei Chang, TingTing Hwang, Sheng-Yu Hsu |
ASP-DAC | 2 |
| 2005 | Low-power techniques for network security processorsabstractIn this paper, we present several techniques for low-power design, including a descriptor-based low-power scheduling algorithm, design of dynamic voltage generator, and dual threshold voltage assignments, for network security processors. The experiments show that the proposed methods and designs provide the opportunity for network security processors to achieve the goals of both high performance and low power. Yi-Ping You, Chun-Yen Tseng, Yu-Hui Huang, Po-Chiun Huang, TingTing Hwang, Sheng-Yu Hsu |
ASP-DAC | 5 |
| 2004 | Low power design using dual threshold voltage
Yen-Te Ho, TingTing Hwang |
ASP-DAC | 2 |
| 2004 | Decomposition of Instruction Decoder for Low Power DesignabstractMicroprocessors have been used in wide-ranged applications.During the execution of instructions, instruction decoding is a major task for identifying instructions and generating control signals for data-paths.By exploiting program behaviors, we propose a novel instruction-decoding approach for power minimization.Using the proposed instruction-decoding structure, we present a partitioning method that decomposes the instruction-decoding circuit into two sub-circuits according to the execution frequencies of instructions.Using our proposed decoding structure, only one sub-circuit will be activated when executing an instruction.Experimental results have demonstrated that our proposed approach achieves on an average of 26.71% and 15.69% power reductions for the instruction decoder and the control unit, respectively. Wu-An Kuo, TingTing Hwang, Allen C.-H. Wu |
DATE | 2 |
| 2004 | Crosstalk Minimization in Logic Synthesis for PLAabstractWe propose a maximum crosstalk minimization algorithm taking logic synthesis into consideration for PLA structure. To minimize the crosstalk, technique of permuting wire is used which includes the following steps. First, product lines are partitioned into long set and short set, and then product lines in long set and short set are interleaved. By interleaving algorithm, an upper bound on the maximum coupling capacitance of the product lines can be derived. Then, we take advantage of crosstalk immunity of product lines in long set to further reduce the maximum crosstalk effect of the PLA. Finally, synthesis techniques such as local transformation and global transformation are taken into consideration to search for a better result. The experiments demonstrate that our algorithm can effectively minimize the maximum crosstalk effect of a circuit by 48% as compared with the original area-minimized PLA without crosstalk minimization. Yi-Yu Liu, Kuo-Hua Wang, TingTing Hwang |
DATE | 3 |
| 2003 | G-MAC: An Application-Specific MAC/Co-Processor Synthesizer
Alex C.-Y. Chang, Wu-An Kuo, Allen C.-H. Wu, TingTing Hwang |
DATE | 4 |
| 2003 | Decomposition of Extended Finite State Machine for Low Power Design
MingHung Lee, TingTing Hwang, Shi-Yu Huang |
DATE | 2 |
| 2003 | A Custom-Cell Identification Method for High-Performance Mixed Standard/Custom-Cell Designs
Jennifer Y.-L. Lo, Wu-An Kuo, Allen C.-H. Wu, TingTing Hwang |
DATE | 4 |
| 2003 | Compiler optimization on VLIW instruction scheduling for low powerabstractIn this article, we investigate compiler transformation techniques regarding the problem of scheduling VLIW instructions aimed at reducing power consumption of VLIW architectures in the instruction bus. The problem can be categorized into two types: horizontal scheduling and vertical scheduling. For the case of horizontal scheduling, we propose a bipartite-matching scheme for instruction scheduling. We prove that our greedy bipartite-matching scheme always gives the optimal switching activities of the instruction bus for given VLIW instruction scheduling policies. For the case of vertical scheduling, we prove that the problem is NP-hard, and we further propose a heuristic algorithm to solve the problem. Our experiment is performed on Alpha-based VLIW architectures and an ATOM simulator, and the compiler incorporated in our proposed schemes is implemented based on SUIF and MachSUIF. Experimental results of horizontal scheduling optimization show an average 13.30% reduction with four-way issue architecture and an average 20.15% reduction with eight-way issue architecture for transitional activities of the instruction bus as compared with conventional list scheduling for an extensive set of benchmarks. The additional reduction for transitional activities of the instruction bus from horizontal to vertical scheduling with window size four is around 4.57 to 10.42%, and the average is 7.66%. Similarly, the additional reduction with window size eight is from 6.99 to 15.25%, and the average is 10.55%. Chingren Lee, Jenq Kuen Lee, TingTing Hwang, Shi-Chun Tsai |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2002 | A technology mapping algorithm for CPLD architecturesabstractIn this paper, we propose a technology mapping algorithm for CPLD architectures. Our algorithm proceeds in two phases: mapping for single-output PLAs and packing for multiple-output PLAs. In the mapping phase we propose a look-up-table (LUT) based mapping algorithm. We will take advantage of existing LUT mapping algorithms for area and depth minimization. Benchmark results show that our algorithm produce better results in terms of area and depth as compared to TEMPLA. Shih-Liang Chen, TingTing Hwang, C. L. Liu 0001 |
FPT | 2 |
| 2002 | Logic transformation for low-power synthesisabstractIn this article we present a new approach to the problem of local logic transformation for reducing power dissipation in logic circuits. The proposed approach overcomes one of the critical limitations common to the previous approaches of local logic transformations for low power, namely, a sequential greedy transformation that identifies signals with high switching activities and then resynthesizes the signals one by one. Instead, we identify a set of signal lines as a group for logic transformation, and determine an order of transformation of the signals with the maximum reduction of power dissipation in the circuit. As a practically feasible solution to this problem, we develop a power model called a finite state input transition (FIT) model, which allows the efficient measurement of the change of power dissipation of the circuit for every possible sequence of logic transformations among the signal lines. Experimental results show that the proposed approach performs an extensive local logic transformation, reducing power consumption by 33% on average without any increase of circuit delay. Ki-Wook Kim, Taewhan Kim 0001, TingTing Hwang, C. L. Liu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2001 | A construction of minimal delay Steiner tree using two-pole delay modelabstractIn this paper, we will study the construction of a Steiner routing tree for a given net with the objective of minimizing the delay of the routing tree. Previous researches adopt Elmore delay model to compute delay. However, with the advancement of IC technology, a more accurate delay model is required. Therefore, in this paper, we will use two-pole delay model to compute the cost function of a Steiner tree. Moreover, we propose a new algorithm to construct the Steiner tree. Our algorithm takes into consideration the net topology, the total wire length and the longest path from the source to sink. Experimental results show that our algorithm is very effective and efficient as compared to [8]. LiYi Lin, Yi-Yu Liu, TingTing Hwang |
ASP-DAC | 3 |
| 2001 | Binary decision diagram with minimum expected path lengthabstractWe present methods to generate a Binary Decision Diagram (BDD) with minimum expected path length. A BDD is a generic data structure which is widely used in several fields. One important application is the representation of Boolean functions. A BDD representation enables us to evaluate a Boolean function: Simply traverse the BDD from the root node to the terminal node and retrieve the value in the terminal node. For a BDD with minimum expected path length will be also minimized the evaluation time for the corresponding Boolean function. Three efficient algorithms for constructing BDDs with minimum expected path length are proposed. Yi-Yu Liu, Kuo-Hua Wang, TingTing Hwang, C. L. Liu 0001 |
DATE | 3 |
| 2001 | Architecture driven circuit partitioningabstractIn this paper, we propose an architecture driven partitioning algorithm for netlists with multiterminal nets. Our target architecture is a multifield-programmable gate array (FPGA) emulation system with folded-Clos network for board routing. Our goal is to minimize the number of FPGA chips used and maximize routability. To that end, we introduce a new cost function: the average number of pseudoterminals per net in a multiway cut. Experimental result shows that our algorithm is very effective in terms of the number of chips used and routability as compared to other methods. Chau-Shen Chen, TingTing Hwang, C. L. Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | A Clustering Based Linear Ordering Algorithm for K-Way Spectral PartitioningabstractThe spectral method can lead to a high quality of multi-way partition due to its ability to capture global netlist information. For spectral partition, n netlist modules are mapped to n points in d-dimensional space, and then a linear ordering of these n modules is constructed to be used as a basis for partitioning. In this paper, we propose two clustering based linear ordering algorithms taking into consideration the objective function presented. Shiuann-Shiuh Lin, Wen-Hsin Chen, Wen-Wei Lin, TingTing Hwang |
ASP-DAC | 4 |
| 1999 | Logic Transformation for Low Power SynthesisabstractIn this paper we present a new approach to the problem of local logic transformation for reduction of power dissipation in logic circuits. Based on the finite-state input transition (FIT) power dissipation model, we introduce a cost function which accounts for the effects of input capacitance, input slew rate, internal parasitic capacitance of logic gates, interconnect capacitance, as well as switching power. Our approach provides an efficient way of estimating estimating the global effect of local logic transformations in logic circuits. In our approach, the FIT model for the transitive fanout cells of a locally transformed subcircuit can be reused to measure the global power dissipation by varying the input probabilities of the transitive fanout cells. Local logic transformation is carried our based on compatible sets of permissible functions (CSPF). Experimental results show that local logic transformation based on CSPF using our cost function can reduce power consumption by about 36% on average without increase in the worst-case circuit delay. Ki-Wook Kim, TingTing Hwang, C. L. Liu 0001 |
DATE | 3 |
| 1999 | On determining sensitization criterion in an iterative gate sizing processabstractSince only sensitizable paths contribute to the delay of a circuit, false paths must be excluded in optimizing the delay of the circuit. However, just identifying sensitizable paths in the first place is not sufficient since during the optimization process, false paths may become sensitizable, and sensitizable paths false. In addition, some paths may shuttle frequently as false paths and sensitizable ones. That is, a thrashing phenomenon may occur. To lessen the thrashing phenomenon of critical paths, a loose sensitization criterion is proposed in this paper to identify both exact sensitizable paths and shuttle paths. Moreover, the selection of sensitization criterion should be related to the specified delay constraint of a given circuit. Taking into account the tightness of sensitization criterion and delay constraint together, we define in this paper a thrashing coefficient which is used to guide the selection of sensitization criterion in the performance optimization process. Moreover, to save time for path sensitizability analysis in the optimization process, we find a special type of false paths, function-false paths, which can be put aside since they never become sensitizable during the process. How-Rern Lin, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | A Re-engineering Approach to Low Power FPGA Design Using SPFDabstractIn this paper, we present a method to re-synthesize Look-Up Table (LUT) based Field Programmable Gate Arrays (FPGAs) for low power design after technology mapping, placement and routing are performed. We use Set of Pairs of Functions to be Distinguished (SPFD) to express functional permissibility of each signal. Using different propagations of SPFD to fan-in signals, we change the functionality of a PLB (Programmable Logic Block) which drives large loading into one with low transition density. Experimental results show that our method can reduce on average 12% power consumption compared to the original circuits without affecting placement and routing. Jan-Min Hwang, Feng-Yi Chiang, TingTing Hwang |
DAC | 3 |
| 1998 | Architecture driven circuit partitioningabstractb thw paper, we propose an architecture driven partitioning algorithm for nethsts tith mtiti-termind nets.Our target architecture is a mtiti-FPGA emulation system with folded-Clos network for board routing.Our gord is to minimize the number of FPGA chips used and maximize the routabfity.To that end, we introduce a new cost function: the average number of pseudo terrninrds per net in a mdti-way cut.Experiment restit shows that our flg~ rithm is very effective in terms of the number of chips used and the routabfity as compared to other methods. 1 Chau-Shen Chen, TingTing Hwang, C. L. Liu 0001 |
ICCAD | 2 |
| 1998 | Layout Driven Selection and Chaining of Partial Scan Flip-Flops
Chau-Shen Chen, TingTing Hwang |
J. Electron. Test. | 2 |
| 1997 | Low Power FPGA Design - A Re-engineering ApproachabstractIn this paper, technology mapping algorithmsfor minimizing power consumption in FPGA design are studied.The technology mapping problem for power minimizationhas been shown to be NP-complete.Furthermore, thereare other important objectives, such as the number of PLBs(Programmable Logic Blocks), the number of levels and soon, that should also be optimized simultaneously.We proposea transformational approach in which we start with amapping solution which optimizes certain objective(s) (e.g., the number of PLBs.)The mapping solution is then transformedto reduce the power consumption while keeping thenumber of PLBs fixed.Our algorithm explores the possibilitiesof transforming the functionality of the PLBs so thatthe switching densities of the output edges of the PLBs willbe reduced, leading to a reduction in total power consumption.Our transformational approach can also be viewed as are-engineering approach in which power reduction is achievedthrough re-routing after the PLBs have been placed, utilizingeffectively the capability of a PLB to realize any booleanfunction of up to k variables. Chau-Shen Chen, TingTing Hwang, C. L. Liu 0001 |
DAC | 2 |
| 1997 | Net assignment for the FPGA-based logic emulation system in the folded-Clos network structureabstractIn this paper, we study the net assignment problem for a logic emulation system in the folded-Clos network interconnection, also referred to as the "partial crossbar interconnection structure". Net assignment of two-terminal nets in this interconnection structure is guaranteed to be completed in polynomial time. However, net assignment of multiterminal nets becomes NP-complete. A previous paper by Butts et al. (1992) has proposed a simple heuristic to perform net assignment for multiterminal nets. Its results showed that it failed to complete routing of all nets for many cases. It is inadequate to have a net assignment algorithm which does not guarantee an exact solution, as the failure of interconnecting field programmable gate arrays (FPGA's) will result in the failure of mapping to the computing engine as a whole and will result in redoing the previous steps, e.g., partitioning of circuits. Therefore, we propose an exact algorithm to solve the net assignment problem. The exact algorithm will find a solution if one exists. However, the exact algorithm may take exponential time. Accordingly, a two-phase approach is taken in this paper. A time-efficient heuristic method is used first. The exact solver will be called only if the heuristic fails to deliver a solution. Shiuann-Shiuh Lin, Yuh-Ju Lin, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1997 | Boolean matching for incompletely specified functionsabstractBoolean matching is to check the equivalence of two functions under input permutation and input/output phase assignment. In this paper, we address Boolean matching problems for incompletely specified functions. We formulate the searching of input variable mapping between two target functions as a logic equation by using multiple-valued functions. Based on this equation, a Boolean matching algorithm is proposed. Delay and power dissipation can also be taken into consideration when this method is used for technology mapping. Experimental results on a set of benchmarks show that our algorithm is indeed very effective in solving the Boolean matching problem for incompletely specified functions. Kuo-Hua Wang, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1996 | Layout Driven Selecting and Chaining of Partial ScanabstractIn an era of sub-micron technology, routing is becoming a dominant factor in area, timing, and power consumption. In this paper, we study the problem of selecting and chaining of scan flip-flops with the objective of achieving minimum routing area overhead. Most of previous work on partial scan has put emphasis on selecting as few scan flip-flops as possible to break all cycles in S-graph. However, the flip-flops that break more cycles are often the ones that have more fanins and fanouts. The area adjacent to these nodes is often congestive in layout. Such selections will cause layout congestion and increase in number of tracks to chain the scan flip-flops. We propose a matching-based algorithm to perform simultaneously the selecting and chaining of scan flip-flop taking layout information into account. Experimental results show that our approach outperforms the traditional one in final layout area. Chau-Shen Chen, Kuang-Hui Lin, TingTing Hwang |
DAC | 3 |
| 1996 | Technology mapping for TLU FPGAs based on decomposition of binary decision diagramsabstractThis paper proposes an efficient algorithm for technology mapping targeting table look-up (TLU) blocks. It is capable of minimizing either the number of TLUs used or the depth of the produced circuit. Our approach consists of two steps. First a network of super nodes, is created. Next a Boolean function of each super node with an appropriate don't care set is decomposed into a network of TLUs. To minimize the circuit's depth, several rules are applied on the critical portion of the mapped circuit. Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1996 | Exploiting communication complexity for Boolean matchingabstractBoolean matching is to check the equivalence of two functions under input permutation and input/output phase assignment. A straightforward implementation takes time complexity O(n!2/sup n/2), where n is the number of variables. Various signatures of variables were used to prune impossible permutations by many researchers. In this paper, based on communication complexity, we also propose two signatures, cofactor and equivalence signatures, which are general forms of many existing signatures. These signatures are used to develop an efficient Boolean matching algorithm which is based on checking structural equivalence of OBDD's. Experimental results on a set of benchmarks show that our algorithm is indeed very effective in solving Boolean matching problem. Kuo-Hua Wang, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1996 | Low power realization of finite state machines - a decomposition approachabstractWe present in this article a new approach to the synthesis problem for finite state machines with the reduction of power dissipation as a design objective. A finite state machine is decomposed into a number of coupled submachines. Most of the time, only one of the submachines will be activated which, consequently, could lead to substantial savings in power consumption. The key steps in our approach are: (1) decomposition of a finite state machine into submachines so that there is a high probability that state transitions will be confined to the smaller of the submachines most of the time, and (2) synthesis of the coupled submachines to optimize the logic circuits. Experimental results confirmed that our approach produced very good results (in particular, for finite state machines with a large number of states.) Sue-Hong Chow, Yi-Cheng Ho, TingTing Hwang, C. L. Liu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 1995 | Power recduction by gate sizing with path-oriented slack calculationabstractNo abstract available. How-Rern Lin, TingTing Hwang |
ASP-DAC | 2 |
| 1995 | Boolean Matching for Incompletely Specified FunctionsabstractBoolean matching is to check the equivalence of two functions under input permutation and input/output phase assignment.In this paper, we will address Boolean matching problem for incompletely speci ed functions.We will formulate the searching of input variable mapping between two target functions as a logic equation by using multiple-valued function.Based on this equation, a Boolean matching algorithm will be proposed.Delay and power dissipation can also be taken into consideration when this method is used for technology mapping.Experimental results on a set of benchmarks show that our algorithm is indeed very eective in solving Boolean matching problem for incompletely speci ed functions. Kuo-Hua Wang, TingTing Hwang |
DAC | 2 |
| 1995 | Combining technology mapping and placement for delay-minimization in FPGA designsabstractWe combine technology mapping and placement into a single procedure, M.Map, for the design of RAM-based FPGAs. Iteratively, M.Map maps several subnetworks of a Boolean network into a number of CLBs on the layout plane simultaneously. For every output node of the unmapped portion of the Boolean network, many ways of mapping are possible. The choice of which mapping to be used depends not only on the location of the CLB into which the output node will be mapped but also on its interconnection with those already mapped CLBs. To deal with such a complicated interaction among multiple output nodes of a Boolean network, multiple ways of mappings and multiple number of CLBs, any greedy algorithm will be insufficient. Therefore, we use a bipartite weighted matching algorithm in finding a solution that takes the global information into consideration. With the availability of the partial placement information, M.Map is able to minimize the routing delay in addition to the number of CLBs. Experimental results on a set of benchmarks demonstrate that M.Map is indeed effective and efficient.> Chau-Shen Chen, Yu-Wen Tsay, TingTing Hwang, Allen C.-H. Wu, Youn-Long Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1994 | Dynamical identification of critical paths for iterative gate sizingabstractSince only sensitizable paths contribute to the delay of a circuit, false paths must be excluded in optimizing the delay of the circuit. Just identifying false paths in the first place is not sufficient since during iterative optimization process, false paths may become sensitizable, and sensitizable paths false. In this paper, we examine cases for false path becoming sensitizable and sensitizable becoming false. Based on these conditions, we adopt a so-called loose sensitization criterion which is used to develop an algorithm for dynamically identification of sensitizable paths. By combining gate sizing and dynamically identification of sensitizable paths, an efficient performance optimization tool is developed. Results on a set of circuits from ISCAS benchmark set demonstrate that our tool is indeed very effective in reducing circuit delay with less number of gate sized as compared with other methods. How-Rern Lin, TingTing Hwang |
ICCAD | 2 |
| 1994 | State Assignment for Power and Area MinimizationabstractWe address the problem of state assignment to minimize both area and power dissipation for finite state machine designs. We propose a new matching-based state-assignment algorithm which considers area and state transitions simultaneously. Experimental results on a set of benchmarks demonstrate that our approach is indeed very effective in minimizing both area and power dissipation.> Kuo-Hua Wang, Wen-Sing Wang, TingTing Hwang, Allen C.-H. Wu, Youn-Long Lin |
ICCD | 3 |
| 1994 | Logic synthesis for field-programmable gate arraysabstractIn this paper, we consider the problem of configuring Field Programmable Gate Arrays (FPGA's) so that some given function is computed by the device. Obtaining the information necessary to configure a FPGA entails both logic synthesis and logic embedding. Due to the very constrained nature of the embedding process, this problem differs from traditional multilevel logic synthesis in that the structure (or lack thereof) of the synthesized logic is much more important. Furthermore, a metric-like literal count is much less important. We present a communication complexity-based decomposition technique that appears to be more suitable for FPGA synthesis than other multilevel logic synthesis methods. The key is that our logic optimization technique based on reducing communication complexity is good enough to allow a simple technology mapping to work well for FPGA devices.> TingTing Hwang, Robert Michael Owens, Mary Jane Irwin, Kuo-Hua Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1994 | Performance-driven interconnection optimization for microarchitecture synthesisabstractThis paper addresses the interconnection synthesis problem in microarchitecture-level designs. With emphasis on the speed of data movement operations, we propose algorithms that take into consideration the effect of each data-transfer-to-bus binding on the data transfer delay time. The delay time is calculated as a function of both data source load and data carrier (bus) load. By balancing loads among hardware components, the data transfer delay time (hence the total execution time) is shortened. We consider two types of problems: resource-constrained binding and performance-constrained binding. Two integer linear programming (ILP) formulations are derived to optimally solve the problems. In order to speed up the computation, a bipartite weighted matching method for the resource-constrained binding and a greedy merging method for the performance-constrained binding are also proposed. Both the ILP formulation generators and the heuristics have been programmed. Experimental results indicate that the proposed algorithms are indeed very effective in optimizing the performance aspect of the interconnection design.> Yi-Min Jiang, Tsing-Fa Lee, TingTing Hwang, Youn-Long Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1993 | Combining technology mapping and placement for delay-optimization in FPGA designsabstractWe combine technology mapping and placement into a single procedure, M.map, for the design of RAM-based FPGAs. Iteratively, M.map maps several subnetworks of the Boolean network into a number of CLBs on the layout plane simultaneously. For every output node of the un-mapped portion of the Boolean network, many ways of mapping are possible. The choice depends on the location of the CLB into which the output node will be mapped as well as the interconnection with those already mapped CLBs. To deal with such a complicated interaction among multiple output nodes, multiple ways of mappings and multiple CLBs, any greedy algorithms will be insufficient. Instead, we use a bipartite weighted matching algorithm to find a globally optimum solution. With the availability of the partial placement information. M.map is able to minimize the routing delay in addition to the number of CLBs. Experimental results on a set of benchmarks demonstrate that M.map is indeed very effective in minimizing the real delay (after routing) as well as the number of CLBs. Chau-Shen Chen, Yu-Wen Tsay, TingTing Hwang, Allen C.-H. Wu, Youn-Long Lin |
ICCAD | 3 |
| 1992 | ELM-A Fast Addition Algorithm Discovered by a ProgramabstractA new addition algorithm, ELM, is presented. This algorithm makes use of a tree of simple processors and requires O(log n) time, where n is the number of bits in the augend and addend. The sum itself is computed in one pass through the tree. This algorithm was discovered by a VLSI CAD tool, FACTOR, developed for use in synthesizing CMOS VLSI circuits.> Thomas P. Kelliher, Robert Michael Owens, Mary Jane Irwin, TingTing Hwang |
IEEE Trans. Computers | 4 |
| 1992 | Efficiently computing communication complexity for multilevel logic synthesisabstractA new method for computing the communication complexity of a given partitioning whose running time is O(pq), where p is the number of implicants (cubes) in the minimum covering of the function and q is the number of different overlapping of those cubes, is presented. Two heuristics for finding a good partition which give encouraging results are presented. Together, these two techniques allow a much larger class of functions to be synthesized. Two heuristic partitioning methods have been tested for certain circuits from the MCNC benchmark set. Using either heuristic, 11 out of 14 examples actually achieve the optimal solutions. A prototype program designed using the above techniques was developed and tested for circuits from the MCNC benchmark set. The experiment shows that the new symbolic manipulation technique is several orders of magnitude faster than an old version.> TingTing Hwang, Robert Michael Owens, Mary Jane Irwin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1990 | Logic synthesis for programmable logic devicesabstractThe use of communication complexity based logic synthesis when configuring programmable logic devices (PLDs) is discussed. Configuration of a PLD involves the two processes of logic synthesis and logic embedding. Since the allowable PLD logic primitives usually include a very large number of gates, the processes of logic synthesis and technology mapping cannot be completely decoupled as they normally are in traditional logic synthesis systems. The proposed communication-complexity-based logic synthesis tool has the advantage of not completely decoupling these two processes. It is more suited to PLD configuring than other multilevel logic synthesis methods.> TingTing Hwang, Robert Michael Owens, Mary Jane Irwin |
ICCD | 1 |
| 1990 | Exploiting communication complexity for multilevel logic synthesisabstractA multilevel logic synthesis technique based on minimizing communication complexity is presented. This approach is believed to be viable because, for many types of circuits, the area needed is dominated by interconnections. By minimizing communication complexity and interconnect, area is reduced. This approach performs especially well for functions that are hierarchically decomposable (e.g., adders, parity generators, comparators, etc.). Unlike many other multilevel logic synthesis techniques, a lower bound can be computed to determine how well the synthesis was performed. A new multilevel logic synthesis program based on the techniques described for reducing communication complexity is presented.> TingTing Hwang, Robert Michael Owens, Mary Jane Irwin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1989 | Multi-Level Logic Synthesis Using Communication ComplexityabstractWe present a new multi-level logic synthesis technique based on minimizing communication complexity. Intuitively, we believe this approach is viable because for many types of circuits lower bounds on the area needed to implement those circuits have been obtained considering only communication complexity. It performs especially well for functions which are hierarchically decomposable (e.g., adders, parity generators, comparators, etc.). Unlike many other multi-level logic synthesis techniques, a lower bound can be computed to determine how well the synthesis was performed. We also present a new multi-level logic synthesis program based on the techniques described for reducing communication complexity. TingTing Hwang, Robert Michael Owens, Mary Jane Irwin |
DAC | 1 |