Charles H.-P. Wen

dblp:37/838 · also Charles Hung-Pin Wen · DBLP profile ↗
← Back
68ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0003-4623-9941ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 57 · 5 first-author · 21 since 2021Computer networks · 7Software engineering, systems software and programming languages · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SOFA-H: Post-Synthesis Area Optimization via Functionally Encoded, Net-Driven Subgraph Mining and SAT-Based Hypercell Remapping
abstract
Synthesized netlists often leave substantial room for area optimization due to the limited function diversity in standard cell libraries, which frequently results in recurring logic patterns that could be compacted through cell combination-referred to as hypercells in this work. While prior studies have demonstrated the potential of hypercell-based optimization, most lack efficient and scalable mining strategies. We present SOFA-H, a post-synthesis framework that extracts and remaps hypercells for maximum area reduction. SOFA-H (i) mines fanout-induced subgraphs and canonically encodes them using P-Representatives, (ii) selects an optimal set of hypercells with non-overlapping replacements via a one-shot weighted MaxSAT formulation, and (iii) supports high input, multi-output cells with scalable runtime. Evaluated on the EPFL benchmark suite synthesized using FreePDK45 and ASAP7, SOFA-H achieves average area reductions of 12.2% and 7.4%, respectively, and runs $380 \times$ faster on average at ASAP7 compared to the state-of-the-art method. These results demonstrate that the extracted hypercells offer a scalable and effective path to closing the area gap left by conventional synthesis.
Jimmy Y.-C. Lee, Yen-Ju Su, Jiun-Cheng Tsai, Aaron C.-W. Liang, Charles H.-P. Wen, Hsuan-Ming Huang
ASP-DAC5
2026 DN-FF: A SEU-Tolerant Flip-Flop Design for Advanced Technology Nodes
abstract
Single-event upsets (SEUs) pose a critical reliability threat in advanced automotive and space electronics. While existing SEU-tolerant latch designs, such as those based on C-elements and unique modules, often fail to meet the stringent space radiation standards (linear energy transfer (LET)$= 60~\text {MeV} \cdot \text {cm}^{2}$/mg) at advanced fin field-effect transistor (FinFET) technology nodes, triple modular redundancy (TMR) achieves sufficient tolerance but incurs significant overhead. To address these limitations, this brief introduces DN-FF, a novel detection-node flip-flop (DN-FF) architecture that leverages reduced node spacing in modern processes for complete SEU immunity with significantly reduced overhead compared to TMR. Incorporating strategically placed detection nodes (DNs) and a dedicated detection circuit (DC), DN-FF achieves robust radiation-hardness while significantly reducing physical area, delay, and power consumption compared to traditional TMR-based solutions. The experimental results demonstrate that DN-FF reduces area by 8.2%, delay by 18.4%, and power by 15.9%, delivering a 37% improvement in the overall area–delay–power quality (ADPQ) metric. These advantages make DN-FF a compact, high-performance, and reliable solution for demanding automotive and aerospace applications.
Lowry P.-T. Wang, Charles H.-P. Wen
IEEE Trans. Very Large Scale Integr. Syst.2
2025 ResCap: Fast-yet-Accurate Capacitance Extraction for Standard Cell Design by Physics-Guided Machine Learning
abstract
In the field of VLSI design, accurate capacitance extraction is essential for ensuring optimal performance of integrated circuits, especially in standard cell designs. Conventional techniques, such as the 2.5D model and 3D field solver, either suffer from inaccuracies or are computationally intensive. To address these challenges, we present ResCap, an innovative approach that synergizes physics-guided linear models with advanced machine learning techniques. Rather than directly predicting the target capacitance, our method starts by employing physical principles to estimate the initial capacitance value, ensuring that predictions are grounded in well-established physical laws. Subsequently, machine learning is applied to predict residual values, thereby refining the initial estimates. This approach not only enhances accuracy and generalization but also reduces dependency on extensive training datasets. Experimental results demonstrate that ResCap significantly outperforms conventional methods on industrial standard cell designs under 4nm process technology, achieving high accuracy with an average error of 0.06% in delay and 0.16% in power. Notably, ResCap exhibits no outliers (error > 1%), whereas the conventional 2.5D extraction tool shows significant outliers of 10.95% in delay and 50.4% in power. Furthermore, our framework demonstrates remarkable efficiency, reducing extraction time by 215x compared to field solvers.
Jiun-Cheng Tsai, Hsuan-Ming Huang, Wei-Min Hsu, Pei-Ting Lee, Jen-Hang Yang, Heng-Liang Huang, Yen-Ju Su, Charles H.-P. Wen
ASP-DAC8
2025 CoP&R: Co-Optimizing Place-and-Route for Standard Cell Layout via MCTS and AllSAT
abstract
Standard cell layout design at advanced technology nodes faces a massive combinatorial explosion of transistor placement possibilities, especially when targeting optimal performance, power, and area (PPA). In this work, we propose a novel framework that integrates AllSAT-based pruning and Monte Carlo Tree Search (MCTS) to tackle this challenge efficiently. Our method first employs an AllSAT formulation that incorporates routing constraints and layout heuristics to exhaustively enumerate only the legal and promising placement solutions. This dramatically reduces the solution space while preserving high-quality candidates. We then apply a guided MCTS algorithm to explore the reduced space and identify optimal or near-optimal placements under given objectives such as total wire length (TWL). Experimental results on a diverse set of standard cells demonstrate the effectiveness of our approach. The AllSAT filtering improves average solution routability from 1.1% to 62.7%, while reducing the total placement space by over 99.9%. On top of that, our MCTS achieves a 62.2× runtime speedup over brute-force exploration, with only a 0.2% degradation in TWL quality. These results confirm that our AllSAT+MCTS framework offers a scalable and practical solution for high-quality standard cell layout synthesis.
Yen-Ju Su, Jiun-Cheng Tsai, Hsuan-Ming Huang, Aaron C.-W. Liang, Han-Ya Tsai, Wei-Min Hsu, Jen-Hang Yang, Charles H.-P. Wen
ICCAD8
2025 MuSTNet: SAT-based Exact Multi-Stage Transistor Network Synthesis with Placement Awareness
abstract
Optimizing power, performance, and area in IC designs remains a key focus. However, the limited functionality of standard cell libraries restricts further optimization. A promising solution involves developing customized complex gates that integrate multiple basic gate functions into a single gate at the transistor level. While prior research has extensively explored transistor network optimization of the complex gates, most studies still focus on 1-stage networks, limiting the potential for deeper optimization. Furthermore, existing methodologies often neglect considering transistor placement during network synthesis, potentially leading to suboptimal area even if the transistor count is reduced. To address these limitations, we propose MuSTNet, a SAT-based exact synthesis framework that minimizing transistor networks by incorporating two key innovations: (1) multi-stage hierarchy for deeper optimization, and (2) transistor placement constraints to simultaneously minimize both transistor count and physical area. Experimental results demonstrate that MuSTNet surpasses previous studies, achieving an 18% reduction in transistor count for 4-input P-class functions and a 12.9% reduction for multi-output functions. When applied to an industrial library benchmark, MuSTNet yields a 6.3% area reduction compared to the approach neglecting placement constraints. Moreover, MuSTNet has been applied to complex gate generation, allowing simultaneous functional and topological optimization at the transistor level. Compared to traditional cell-level design, this approach reduces transistor count by 12.3% and area by 21.4%, demonstrating its potential in advanced IC design.
Jiun-Cheng Tsai, Wei-Min Hsu, Kuei-Lin Wu, Hsuan-Ming Huang, Jen-Hang Yang, Heng-Liang Huang, Yen-Ju Su, Charles H.-P. Wen
ICCAD8
2025 Enhancing Timing Predictability in Automotive Electronics: Addressing Aging and Temperature Distributions
abstract
The reliability of automotive electronics is heavily affected by aging and temperature variations, which can degrade performance or cause failures. Traditional timing-prediction methods assuming a constant temperature can introduce a 17% error compared to models that incorporate temperature distribution over 30 years of aging, resulting in inaccurate reliability assessments. To address this limitation, this paper introduces the timing-prediction framework based on the Aging- and Temperature Distribution-Aware Timing Model (ATD-Timing Model) that integrates functional behaviors, aging effects, and temperature distributions for accurate analysis. By leveraging a two-phase machine-learning approach for cell-delay prediction, our ATD-Timing Model achieves an average error of 0.58% at the cell level compared to SPICE. Furthermore, after eliminating false paths at the circuit level, the average critical delay error of our timing-prediction framework is reduced to 2.22% compared to SPICE, while computational efficiency is significantly improved with an average speedup of 203.97×. This enables accurate and efficient automotive circuit reliability analysis under realistic conditions.
Jeffery Y.-C. Chen, Jason W.-Y. Cheng, Aaron C.-W. Liang, Charles H.-P. Wen, Gung-Yu Pan, Hen-Ming Lin
ITC4
2025 Designing Radiation-Hardened D Flip-Flop with Reduced Latency and Area Using Filtering Buffer
abstract
In safety-critical applications such as biomedical and automotive electronics, soft errors induced by radiation particles pose a significant reliability concern. This paper presents a novel Filtering Buffer-based D Flip-Flop (FB-DFF) designed to enhance radiation hardening while minimizing area and timing overheads. The FB-DFF integrates a Filtering Buffer (FB) and an Auto Delay Element (ADE) to effectively filter out transient errors and ensure the correct input signal is latched. The Single-Master Dual-Slaves (SMDS) architecture further prevents error propagation to the output. Experimental results demonstrate that FB-DFF can resist radiation particles below 77 LET, providing full protection against both Single Event Transient (SET) and Single Event Upset (SEU). Compared to existing designs, FB-DFF reduces area overhead by 44% and timing overhead by 144% at the chip level, making it a highly efficient solution for radiation-hardened applications. The proposed FB-DFF offers a balanced trade-off between radiation hardening capability, area, and timing performance, making it suitable for integration into various safety-critical systems.
Nelson M.-C. Wu, Lowry P.-T. Wang, Chia-Wei Liang, Charles H.-P. Wen, Herming Chiueh
VTS4
2025 Machine-Learning-Based Ranking of Cell Layout Delay Considering Layout-Dependent Effects
abstract
Cell layout generation plays a crucial role in design automation. The generated layout must not only adhere to design rules but also exhibit optimized performance in terms of factors such as delay, power, area, and cost. However, prior works often rely on metrics that fail to consider the layout-dependent effects (LDEs). Furthermore, evaluating the actual performance using commercial tools can be excessively time-consuming, especially when iteratively optimizing cell layouts. Therefore, this work proposes a new machine-learning(ML)-based ranking model to enable rapid performance ranking between layout candidates of standard cells. This model incorporates all LDEs in feature extraction, generating an ordered list of cell layouts, and evaluating only the top-Kcandidates for performance. The experiments show that this approach successfully identifies the optimal layout from ten benchmark cells, which are most used in intellectual property (IP) cores, in a sub-5 nm fin field-effect transistor (FinFET) industrial standard cell library, achieving a$348\times $speedup over the conventional flow.
Ya-Rou Hsu, Aaron C.-W. Liang, Han-Ya Tsai, Yen-Ju Su, Charles H.-P. Wen, Hsuan-Ming Huang
IEEE Trans. Very Large Scale Integr. Syst.5
2024 MAXCell: PPA-Directed Multi-Height Cell Layout Routing Optimization using Anytime MaXSAT with Constraint Learning
abstract
To optimize power, performance, and area (PPA) of IC designs, standard cell has evolved from basic to complicated designs, resulting in complex multi-height structures. Although extensive research on single-height cell automatic synthesis, multi-height cell studies are still limited due to the extremely large solution space. In this paper, we present MAXCell, a PPA-directed standard cell layout optimization framework for both single-height and multi-height designs using anytime MaxSAT with constraint learning. This framework incorporates two novel techniques: (1) learning additional constraints from the original constraints database to accelerate convergence during problem-solving and (2) integrating a genetic algorithm with a ranking model to dynamically guide the router towards the PPA goal directly during optimization. Experimental results indicate that MAXCell outperforms previous studies that target wire length optimization, achieving a 5.5% power reduction in evaluations of 33 multi-bit flip-flop designs beyond 4nm technology. Furthermore, compared to an industrial library designed by experienced engineers, MAXCell provides a 3.5% power optimization benefit and drastically reduces the delivery time from multiple days to a mere few hours (21.6X faster). This emphasizes its efficiency and its potential in modern integrated circuit design.
Jiun-Cheng Tsai, Wei-Min Hsu, Yun-Ting Hsieh, Yu-Ju Li, C. N. Ho, Hsuan-Ming Huang, Jen-Hang Yang, Heng-Liang Huang, Aaron C.-W. Liang, Charles H.-P. Wen
ICCAD11
2024 LESER-2: Detailed Consideration in Latch Design under Process Migration for Prevention of Single-Event Double-Node Upsets
abstract
Single-Event Double-Node Upsets (SEDUs) increasingly compromise the integrity of storage cells like latches or flip-flops, especially as technology scales down to sub-65nm nodes. Traditional RHBD (Radiation-Hardened by Design) strategies, such as DICE and TMR, crafted to counter Single-Event Node Upset (SEU), are proving to be inadequate to SEDU. Consequently, recent research suggests the use of expanded areas to shield circuits from SEDUs, yet this often results in disproportional overhead. In response, the LESER methodology emerged as an effective measure, introducing minimal redundancy while still guaranteeing full SEDU resilience. However, having mitigated SEDUs across different latches within ASAP7 (a predictive process), LESER calls for further improvement in three pivotal issues: (1) spacing impact, (2) dummy gates, and (3) device configuration. Thus, LESER-2 has been developed, targeting these three challenges in latches designed with two industrial process nodes. LESER-2 presents a dual-level architecture, encompassing both device and circuit strata. At the device level, LESER-2 replaces the virtual process model card with industrial ones, significantly enhancing simulation accuracy concerning heavy ion impacts on transistors. Additionally, circuit-level layout modification within LESER-2 are meticulously calibrated to mitigate SEDUs in vulnerable node pairs. Empirical tests validate the LESER-2 modifications applied to latches under two industrial process nodes, accomplishing a 100% rate of soft error prevention while incurring an average area overhead of 16.5%, effectively resolving the three aforementioned issues.
Alan S.-M. Liu, Lowry P.-T. Wang, Charles H.-P. Wen, Herming Chiueh
ITC3
2024 Temperature-Insensitive Soft-Error-Tolerant Flip-Flop Design For Automotive Electronics
abstract
Many existing soft-error-tolerant flip-flop designs (e.g., MDAD-FF, SETU-TOFF, SEDR-FF) apply delayed latching to mitigate strikes of radiation particles. However, according to AEC-Q100 (Grade 1), automotive electronics are permitted to operate at temperatures between −40°C to 125°C, resulting in two reliability issues: (1) protection failure and (2) timing degradation. At −40°C, these rad-hard FF designs are capable of providing a worst-case delay of only 113 ps, ineffective in protecting against 77-LET particles (which require 200 ps in 45 nm process). At 125°C, however, the performance of these FF designs may degrade to 386 ps, resulting in more timing violations. Therefore, RAV-FF is proposed to address these two issues by incorporating a MOSFET capacitance (MCAP) to generate sufficient delay to delay clock and a current-control transistor (CC) to stabilize delay at different temperature corners. Experimental results indicate that RAV-FF provides effective soft-error protection in the temperature range of −40°C to 125°C by ensuring a delay of at least 200 ps with only 3% variation.
Ralf E.-H. Yee, Nicholas Y.-J. Su, Lowry P.-T. Wang, Charles H.-P. Wen, Herming Chiueh
VTS4
2024 PS-IPS: Deploying Intrusion Prevention System with machine learning on programmable switch
Alan Y.-P. Lee, Michael I.-C. Wang, Chi-Hsiang Hung, Charles H.-P. Wen
Future Gener. Comput. Syst.4
2023 Preventing Single-Event Double-Node Upsets by Engineering Change Order in Latch Designs
abstract
Single-event-induced soft errors are serious issues in advanced nano-scale technology, causing malfunctions in systems. As the size of technology node decreases to sub-65nm with closer transistor spacing, single-event double-node upsets (SEDU) occur more frequently than single-event upsets (SEU). Previous studies handled SEDU by incorporating protection mechanisms in cell designs or modifying the physical layout. However, they have massive area overhead and SEDU cannot be fully prevented. In this paper, we propose a LESER framework to reconstruct the latch design, achieving 100% SEDU tolerance. Based on the concept of engineering change order (ECO), LESER contains a two-level analysis process to prevent SEDU with minimum modification on layout, including 1) device level and 2) circuit level. The device level extracts the current source model by TCAD simulation and the circuit level reconstructs the layout with scanning process, double-node injection test, and layout modification approach. Experiments show that the reconstructed design can achieve a 100% soft error protection rate with the costs of an increment of 6.4% in area, 1% in timing and power penalty. The results indicate that LESER can fully prevent SEDU by reconstructing the latch with minimum performance penalties.
Sam M.-H. Hsiao, Amy H.-Y. Tsai, Lowry P.-T. Wang, Aaron C.-W. Liang, Charles H.-P. Wen, Herming Chiueh
ITC5
2022 SEM-latch: a lost-cost and high-performance latch design for mitigating soft errors in nanoscale CMOS process
abstract
Soft errors (primarily single-event transients (SET) and single-event upsets (SEU)) are receiving increased attention due to the increasing prevalence of automotive and biomedical electronics. In recent years, several latch designs have been developed for SEU/SET protection, but each has its own issues regarding timing, area, and power. Therefore, we propose a novel soft-error mitigating latch design, called SEM-Latch, which extends QUATRO and incorporates a speed path whereas embedding a reference voltage generator (RVG) for simultaneously improving timing, area, and power in 45nm CMOS process. SEM-Latch effectively reduces the power, area, and PDAP (product of delay, area, and power) by an average of 1.4%, 12.5%, and 8.7%, respectively, in comparison to a previous latch (HPST) with equivalent SEU protection. Furthermore, in comparison to AMSER-Latch, SEM-Latch reduces area, timing overhead and PDAP by 27.2%, 48.2%, and 60.2%, respectively, to provide 99.9999% particle rejection rate for SET protection.
Zhong-Li Tang, Chia-Wei Liang, Ming-Hsien Hsiao, Charles H.-P. Wen
DAC4
2022 Timing-Critical Path Analysis in Circuit Designs Considering Aging with Signal Probability
abstract
Aging is an important determinant for the reliability of circuit designs and has been addressed by a number of protection techniques based on static timing analysis (STA). The timing reported by STA, however, is often too optimistic without considering the functional behavior of the circuit. Furthermore, signal probability has also been found to be a significant factor in the aging effect. As such, we present in this paper a timing-critical path analysis that takes function and aging into account as well as signal probability. Functional timing analysis (FTA) eliminates the false paths and generates more accurate timing. Furthermore, machine learning can be used to build models for predicting the timing of each cell for various aging lifetimes and signal probabilities. Experimental results indicate that there can be a difference of up to 6% on path delay between STA and FTA. The path ranks also differ for most of the benchmark circuits after considering aging with signal probability, resulting in the delay differences of up to 6.12 %. In conclusion, it is necessary to consider function, aging, and signal probability simultaneously when analyzing timing-critical paths in a circuit design.
Jiun-Cheng Tsai, Aaron C.-W. Liang, Charles H.-P. Wen
ITC-Asia3
2022 Existence of Single-Event Double-Node Upsets (SEDU) in Radiation-Hardened Latches for Sub-65nm CMOS Technologies
abstract
A single-event double-node upset (SEDU) may appear to result in an erroneous state of the storage element due to the scalability of transistor features. Therefore, SEDU must be well addressed from the perspective of circuit reliability, especially for safety-critical electronics. Some previous studies claimed to protect against SEDUs in 65-nm process technologies, but were not thoroughly verified. To better understand the technology-scaling impact, we re-examine SEDU in different advanced technologies (including 7-nm finFET, 45-nm bulk CMOS, and 65-nm bulk CMOS). An integrated multi-level framework is developed with the current-source modeling derived from the device-level TCAD simulation, combined with voltage calculation derived from the circuit-level SPICE simulation. To adequately capture the probability of errors occurring in the latch design under all possible scenarios, this paper also considers a variety of environmental factors, such as strike angles, temperature variation, and technology nodes. Also, three classical latch designs (i.e., TMR, DICE, and HLR) have been implemented in different technologies and well calibrated for experiments. According to experiment results, it is evident that SEDU is highly dependent on both the physical layout of the design as well as its design style. DICE is found to be the most susceptible to SEDU in all three manufacturing technologies, whereas TMR and HLR can be immune to SEDU in the 45-nm and 65-nm technologies due to a lack of sufficient charge to upset more than two nodes. It is, therefore, essential to consider both the physical layout and the manufacturing technology employed for ensuring the robustness of a radiation-hardened design against particle strikes.
Sam M.-H. Hsiao, Lowry P.-T. Wang, Aaron C.-W. Liang, Charles H.-P. Wen
ITC4
2022 Enabling Malware Detection with Machine Learning on Programmable Switch
abstract
Malware detection is an important issue for network security, especially for the Internet of Things (IoT) network. Traditional network intrusion detection system (NIDS), running on external host servers, are not scalable for ever-increasing IoT traffic and waste time on transmitting data back and forth. Here, we propose a novel architecture called on-switch malware detector that utilizes the programmable switch and the machine-learning technique to achieve better performance on detecting malicious flows in the network. The on-switch malware detector mainly consists of four components: (1) packet forwarder, (2) feature extractor, (3) flow director, and (4) neural-network detector. According to the experimental results, the on-switch malware detection has a 99.57% shorter response time than a conventional signature-based NIDS; meanwhile its processing capacity increases by 800 times. As a result, the on-switch malware detector efficiently overcomes the shortcomings of conventional NIDSs, making it a better fit for the IoT network.
Hsin-Fu Chang, Michael I.-C. Wang, Chi-Hsiang Hung, Charles H.-P. Wen
NOMS4
2022 P4SF: A High-Performance Stateful Firewall on Commodity P4-Programmable Switch
abstract
This paper presents a high-performance stateful firewall called P4SF that runs on a commodity P4-programmable switch and uses an extended finite state machine to provide match-state-action in the forwarding plane for stateful processing while significantly reducing the controller’s workload. P4SF is composed of three key blocks (i.e. Match Block, State Block, Action Block) that are responsible for reading/writing flow states, maintaining state transitions, and forwarding packets. Preemptive data caching is also realized into a buffer called State Pre-Fetch in P4SF for hiding transmission delay during state updates of flows. As a result, P4SF is successfully exercised on a commodity P4-programmable switch, and can be scaled to support 384,000 entries (120,000 under the three-way handshake in TCP connections) for TCP flows, achieving the 100Gb/s linerate speed for packet forwarding.
Linyih Teng, Chi-Hsiang Hung, Charles H.-P. Wen
NOMS3
2022 Rad-Hard Designs by Automated Latching-Delay Assignment and Time-Borrowable D-Flip-Flop
abstract
As the safety-critical applications (e.g., automotive and medical electronics) emerge, various techniques of radiation hardening by design (RHBD) are proposed to deal with soft errors. Among all RHBD techniques, Built-In Soft-Error Resilience (BISER) is the first one to apply the delayed latching to separate input signals on all flip-flops for error detection. However, the delay values induced by BISER extend the setup time of all flip-flops, and may fail to meet the timing specification of the design. For minimizing such delay impact on the setup time of each flip-flop, we propose the Automated Latching-Delay Assignment (ALDA) to transfer partial values to the CK-Q delay. Later, Time-Borrowable D-Flip-Flop (TBD-FF) as well as a modified design flow is also proposed to realize the delay assignment by ALDA and to complete the design hardening. Experiments show that ALDA together with TBD-FF effectively protects five benchmark circuits against soft errors, and optimally avoids the timing violations caused by the prior delayed-latching solutions.
Dave Y.-W. Lin, Charles H.-P. Wen
IEEE Trans. Computers2
2022 Roadrunner+: An Autonomous Intersection Management Cooperating with Connected Autonomous Vehicles and Pedestrians with Spillback Considered
abstract
The recent emergence of Connected Autonomous Vehicles (CAVs) enables the Autonomous Intersection Management (AIM) system, replacing traffic signals and human driving operations for improved safety and road efficiency. When CAVs approach an intersection, AIM schedules their intersection usage in a collision-free manner while minimizing their waiting times. In practice, however, there are pedestrian road-crossing requests and spillback problems, a blockage caused by the congestion of the downstream intersection when the traffic load exceeds the road capacity. As a result, collisions occur when CAVs ignore pedestrians or are forced to the congested road. In this article, we present a cooperative AIM system, named Roadrunner+ , which simultaneously considers CAVs, pedestrians, and upstream/downstream intersections for spillback handling, collision avoidance, and efficient CAV controls. The performance of Roadrunner+ is evaluated with the SUMO microscopic simulator. Our experimental results show that Roadrunner+ has 15.16% higher throughput than other AIM systems and 102.53% higher throughput than traditional traffic signals. Roadrunner+ also reduces 75.62% traveling delay compared to other AIM systems. Moreover, the results show that CAVs in Roadrunner+ save up to 7.64% in fuel consumption, and all the collisions caused by spillback are prevented in Roadrunner+.
Michael I.-C. Wang, Charles H.-P. Wen, H. Jonathan Chao
ACM Trans. Cyber Phys. Syst.2
2022 A General and Automatic Cell Layout Generation Framework With Implicit Learning on Design Rules
abstract
Design rule (DR) is the most critical challenge for generating a cell layout automatically in the advanced process technologies (e.g., finFET-EUV). Previous works explicitly encode the complicated DRs into routing constraints and automation scripts, which may not be general and efficient for addressing the DR problem. Therefore, an automatic cell layout generation (ACLG) framework is proposed and adopts three implicit-learning techniques [i.e., guidance learning (EGL), DR learning (DRL), and mistake-driven learning (MDL)], which jointly discover the knowledge of complex DRs from the existing layouts in the cell library. EGL learns the geometry behavior of the target metals from the legal cell layouts. DRL learns the DRs from layout patterns. MDL learns the routing constraints iteratively from the encountered mistakes during the layout generation (LG). These three implicit-learning techniques are combined into ACLG and developed into four core stages to cope with the DR challenge in a more general and efficient way. The experimental results demonstrate that ACLG effectively solves all the DR violations (DRVs) in an advanced finFET-EUV process (100% success rate on fixing DRVs) and successfully yields DRC-clean cell layouts for 13 benchmark cells. In addition, the proposed DRL technique is more efficient than the commercial DRC tool in excluding the illegal layout solution space. The number of iterations for a generated legal cell layout is reduced by 30% on average with DRL (3.31) compared to the commercial DRC tool (4.69). Moreover, the total runtime of generating legal layouts for 13 benchmark cells is further improved by$2.57\times $on average. Since DRL not only reduces the iterations of refining the DRVs in the generated cells but also speedups the process of DR checking (DRC) efficiently.
Aaron C.-W. Liang, Charles H.-P. Wen, Hsuan-Ming Huang
IEEE Trans. Very Large Scale Integr. Syst.2
2021 Generating Layouts of Standard Cells by Implicit Learning on Design Rules for Advanced Processes
abstract
For the advanced process technologies (e.g, finFET with EUV), the design rules (DRs) are the most challenging issue to the generation of cell layouts and all DR violations must be solved in a legal cell layout. However, most of previous works apply explicit encoding on the selected DRs into the routing engine and cannot accommodate the rapid growth on the size and complexity of DRs as the processes continue to advance. Therefore, in this paper, we propose two implicit-learning techniques, (1) experience-guidance learning (EGL) and (2) constraint-driven learning (CDL) for effectively solving such two problems of DRs, and meanwhile develop an automatic cell-layout generation (ACLG) framework for efficiently generating legal cell layouts. The experimental results show that in a finFET-EUV process [1], EGL and CDL successfully reduce all DR violations on eight target cells where each case takes averagely three minutes. As a result, without manual effort, ACLG is capable of generating legal layouts of standard cells by implicit learning on DRs of advanced processes.
Aaron C.-W. Liang, Hsuan-Ming Huang, Charles H.-P. Wen
DATE3
2021 AMSER-FF: Area-Minimized Soft-Error-Recoverable Flip-Flop for Radiation Hardening
abstract
Among various radiation hardening by de-signs (RHBD), triple-modular redundancy (TMR) is frequently used to correct soft errors. One of the well-known TMR solutions is called $\Delta{\mathrm TMR}$, which makes three copies of flop-flops (FFs) and inserts different delay buffers in front of data inputs of the second and third FFs for capturing soft errors. However, such solution has shortcomings in the area/power overhead and the imbalanced rise/fall delay. As a result, a novel TMR flip-flop design called area-minimized soft-error-recoverable flip-flop (AMSER-FF) is proposed in this paper to reduce the area/power overhead and timing degradation. Each AMSER-FF embeds a reference voltage generator (RVG), which generates balanced rise/fall delay and has fixed area utilization instead of inserting delay buffers for correcting errors. Experimental results show that AMSER-FF achieves lower area overhead, power overhead and less timing degradation than $\Delta\mathrm{TMR}$.
John Z.-L. Tang, Dave Y.-W. Lin, Ralf E.-H. Yee, Charles H.-P. Wen
ITC-Asia4
2021 A Delay-Adjustable, Self-Testable Flip-Flop for Soft-Error Tolerability and Delay-Fault Testability
abstract
As the demand of safety-critical applications (e.g., automobile electronics) increases, various radiation-hardened flip-flops are proposed for enhancing design reliability. Among all flip-flops, Delay-Adjustable D-Flip-Flop (DAD-FF) is specialized in arbitrarily adjusting delay in the design to tolerate soft errors induced by different energy levels. However, due to a lack of testability on DAD-FF, its soft-error tolerability is not yet verified, leading to uncertain design reliability. Therefore, this work proposes Delay-Adjustable, Self-Testable Flip-Flop (DAST-FF), built on top of DAD-FF with two extra MUXs (one for scan test and the other for latching-delay verification) to achieve both soft-error tolerability and testability. Meanwhile, a built-in self-test method is also developed on DAST-FFs to verify the cumulative latching delay before operation. The experimental result shows that for a design with 8,802 DAST-FFs, the built-in self-test method only takes 946 ns to ensure the soft-error tolerability. As to the testability, the enhanced scan capability can be enabled by inserting one extra transmission gate into DAST-FF with only 4.5 area overhead.
Dave Y.-W. Lin, Charles H.-P. Wen
ACM Trans. Design Autom. Electr. Syst.2
2020 SDPTA: Soft-Delay-aware Pattern-based Timing Analysis and Its Path-Fixing Mechanism
abstract
In modern VLSI design flow, timing analysis is crucial for verifying whether a circuit design can operate without errors. Soft-delay effect (SDE), which is a kind of degraded soft error, will make system failed though the circuit has passed the typical timing analysis. Therefore, we propose a soft-delay-aware timing analysis which takes SDE into consideration. Additionally, a path-fixing mechanism is also proposed to fix up the violated paths automatically. Experimental results show that only 1.05% area budget is required averagely that all violated paths can be fixed up. In summary, SDPTA and the path-fixing mechanism are capable of reducing SDE to general circuits without other manual effort.
Gary K.-C. Huang, Dave Y.-W. Lin, John Z.-L. Tang, Charles H.-P. Wen
ATS4
2020 SAFCast: Smart Inter-Datacenter Multicast Transfer with Deadline Guarantee by Store-And-Forwarding
abstract
With the increasing demand of network services, many researches employ a Software Defined Networking architecture to manage large-scale inter-datacenter networks. Some of existing services such as backup, data migration and update need to replicate data to multiple datacenters by multicast before deadline. Recent works set up minimum-weight Steiner tree as routing paths for multicasting transfers to reduce bandwidth waste; meanwhile, using deadline-aware scheduling guarantees deadlines of requests. However, there is an issue of bandwidth competition among those works. SAFCast as a new algorithm is proposed for multicasting transfers and deadline-aware scheduling. In SAFCast, we develop a tree pruning process and make datacenters employ the store-and-forwarding mechanism to improve the issue of bandwidth competition. Meanwhile, more requests can be accepted by SAFCast. Our experimental result shows that SAFCast outperforms DDCCast in the acceptance rate by 16.5%. In addition, given a volume-to-price function for revenue, SAFCast can achieve 10% more profit than DDCCast. As a result, SAFCast is a better choice for cloud providers with the effective deadline guarantee and more efficient multicasting transfers.
Hsueh-Hong Kang, Chi-Hsiang Hung, Charles H.-P. Wen
INFOCOM3
2020 Speeding Up Functional Timing Analysis by Concise Formulation of Timed Characteristic Functions
abstract
Functional timing analysis (FTA) is a renowned method of finding the true critical delay for the design under interest. By constructing conjunctive normal form (CNF) clauses based on temporal and function constraints, false paths can be identified through the satisfiability (SAT) solving. As a result, the critical delay estimated by FTA is more accurate than that by conventional static timing analysis (STA). However, FTA suffers from the extremely long formulation and computation time, as the number of the clauses in CNF grows exponentially with the increasing size of the design. Due to the reconvergent effect, thousands of clauses can be redundantly formulated for one pin. Even worse, most of them are found useless but seriously lengthen the computation time. Therefore, to avoid ineffective computation in FTA, three novel techniques are proposed: 1) encoding duplication removal (EDR) for removing duplicated functional literals; 2) redundant state propagation (RSP) for propagating temporal states to identify redundant clauses; and 3) temporal footprint identification (TFI) for combining clauses that represent constraints with the same behavior. The experiments show that under a given timing constraint, 94% clauses and 95% literals can be pruned averagely (99% clauses and 99% literals under the best case), resulting in $15.3 \times $ speedup ($72.99 \times $ under the best case) for formulation and SAT solving. As a result, the proposed techniques (EDR, RSP, and TFI) are proven effective to reduce useless computation and improve the overall performance of FTA.
Denny C.-Y. Wu, Aaron C.-W. Liang, Charles H.-P. Wen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 DAD-FF: Hardening Designs by Delay-Adjustable D-Flip-Flop for Soft-Error-Rate Reduction
abstract
For the safety-critical applications such as biomedical and automobile electronics, the system failure induced by soft errors becomes a major issue of reliability. However, most of the commercial cell libraries do not include radiation-hardened components to build a safety-critical design. Therefore, a delay-adjustable D-flip-flop (DAD-FF) is proposed together with a design flow to construct a radiation-hardened system by automation. To enable such radiation-hardened design into the current design flow, DAD-FF is characterized as a general cell and compiled as a patch in the NanGate FreePDK45 bulk 45-nm open cell library, as an example. The experimental results show that DAD-FF is capable of reducing 1.3 × 1010X soft errors with respect to the standard flip-flop (STD-FF) and resisting over 99.999997% strikes of heavy ions. Meanwhile, four radiation-hardened benchmark circuits are synthesized with DAD-FF cell, and further used to prove the effectiveness against soft errors compared to a prior work, built-in soft-error resilience (BISER), with 18% area and 40% timing improvement. To sum up, DADFF is elaborated from the modeling at the device-level to the validation at the system-level and exhibits its strong robustness to soft errors.
Dave Y.-W. Lin, Charles H.-P. Wen
IEEE Trans. Very Large Scale Integr. Syst.2
2019 FAE: Autoencoder-Based Failure Binning of RTL Designs for Verification and Debugging
abstract
As the Register Transfer Level (RTL) designs are more complicated, debugging becomes a major bottleneck in the design process. To make debugging more efficient, failure binning aims at grouping failure traces caused by the same error source together so that designers can focus on one bug at one time. However, as there are multiple bugs in a design, behaviors exhibited by failure traces are diverse and severely confuse designers. One error source may result in different appearances subject to different activation conditions. In addition, different error sources may also exhibit similar appearances among the limited number of failure traces. In this work, we propose an autoencoder-based failure binning engine name FAE for debugging RTL designs more efficiently. The autoencoders extract meaningful representations from the sparse and high-dimensional feature space to the latent space with good properties for clustering. Superior to prior works, FAE provides confidence ranks between bins and in a bin to clearly guide designers during debugging. Experimental results show that FAE can drive bins of higher purity under an acceptable number of bins than prior works, dropping only few less-informative failures. Evaluated by three common metrics for clustering, FAE also achieves averagely 13.1% improvement in purity, 25.0% improvement in NMI and 18.2% improvement in ARI, respectively. As a result, the proposed autoencoder-based engine, FAE, applies machine learning to extract useful information from diverse failure traces and is effective on failure binning with more focused debugging.
Cheng-Hsien Shen, Aaron C.-W. Liang, Charles C.-H. Hsu, Charles H.-P. Wen
ITC4
2019 Dynamic Switch Migration in Distributed Software-Defined Networks to Achieve Controller Load Balance
abstract
Multiple distributed controllers have been used in software-defined networks (SDNs) to improve scalability and reliability, where each controller manages one static partition of the network. In this paper, we show that dynamic mapping between switches and controllers can improve efficiency in managing traffic load variations. In particular, we propose balanced controller (BalCon) and BalConPlus, two SDN switch migration schemes to achieve load balance among SDN controllers with small migration cost. BalCon is suitable for the scenarios where the network does not require a serial processing of switch requests. For other scenarios, BalConPlus is more suitable, as it is immune to the switch migration blackout and does not cause any service disruption. Simulations demonstrate that BalCon and BalConPlus significantly reduce the load imbalance among SDN controllers by migrating only a small number of switches with low computation overhead. We also build a prototype testbed based on the open-source SDN framework RYU to verify the practicality and effectiveness of BalCon and BalConPlus. Experiment confirms the results of the simulations. It also shows that BalConPlus is immune to switch migration blackout, an adverse effect in the baseline BalCon.
Yang Xu 0010, Marco Cello, Michael I.-C. Wang, Anwar Elwalid, Gordon T. Wilfong, Charles H.-P. Wen, Mario Marchese, H. Jonathan Chao
IEEE J. Sel. Areas Commun.6
2018 Skew-Aware Functional Timing Analysis Against Setup Violation for Post-Layout Validation
abstract
Beyond the deep sub-micron era, clock skew is becoming an indispensable factor in post-layout timing and contributes significant delay during signal propagation in paths. Although Functional Timing Analysis (FTA) can provide accurate timing by identifying functionally false paths, clock skew is not yet considered. Therefore, we are motivated to propose a skew-aware functional-timing-analysis engine (named Sk-FTA) for better post-layout validation of designs. In particular, Sk-FTA can save more cost in checking setup-time violations and filters false alarms. Our experimental results show that given three different clock networks, Sk-FTA induces more accurate delay from removing functionally false paths (e.g. 60% less for s13207 under clock tree 3) when comparing with the skew-aware static timing analysis (Sk-STA). Moreover, Sk-FTA eminently yields fewer setup-time violations than Sk-STA does on benchmark circuits. In particular, for vga lcd, all 512 setup-time violations reported by Sk-STA are proved redundant and thus removed, manifesting the power of Sk-FTA.
Pin-Ru Jhao, Denny C.-Y. Wu, Charles H.-P. Wen
ITC-Asia3
2018 Improving Quality of Experience of Service-Chain Deployment for Multiple Users
abstract
The fifth generation (5G) mobile communication network aims at providing high-rate, low-latency services. When a user subscribes a chain of service functions (a.k.a. service chain) from the telecom providers, a Service Level Agreement (SLA) is specified according to his requirement. Deploying service chains optimally has always been a big issue. Several previous works have presented various strategies of service-chain deployment for optimizing either latency or computational resources; however, over-optimization of latency or computational resource is not necessarily equivalent to improvement on quality of experience. Therefore, in this paper, we formally formulate this problem of optimizing quality of experience with the queuing theory and mixed-integer linear programming. In addition, we propose an efficient algorithm named “QoE-driven Service-Chain Deployment with Latency Prediction” for deploying a service chain for a user in practice. According to the experiments, our algorithm reduces > 99% rejections and > 99% waiting time, notably elevating the quality of experience for users.
Michael I.-C. Wang, Charles H.-P. Wen, H. Jonathan Chao
IWQoS2
2017 FASIC: A Fast-Recovery, Adaptively Spanning In-Band Control Plane in Software-Defined Network
abstract
To operate a reliable SDN network, the control plane should be robust in order to guarantee the correctness of traffic forwarding. For in-band control, the traffic flows along with the data plane without a dedicated control network. Thus, the failure of data plane may isolate the switches from control plane and lead to erroneous forwarding behavior. In this paper, we design a control plane, FASIC, which can bootstrap the control network autonomously and adaptively span the control plane with the control-link switching mechanism. By implementing our design with Open vSwitch (OVS), OpenDayLight OVSDB Integration (ODL-OVSDB) and Floodlight software, we demonstrate that FASIC is an in-band control solution which is feasible under support of the available SDN and OpenFlow resources. Compared with standard OVS, the experimental result shows that FASIC reduces 87.83% of control plane downtime under hard link failure and incurs zero downtime under congestion. The preemptive control-link switching can prevent the control-link failure when it suffers from heavy traffic load.
Yu-Lun Su, Michael I.-C. Wang, Yao-Tsung Hsu, Charles H.-P. Wen
GLOBECOM4
2017 Coupling-Aware Functional Timing Analysis for Tighter Bounds: How Much Margin Can We Relax?
abstract
As the manufacturing technology keeps scaling, crosstalk noise induces a greater impact on timing. Although many previous works proposed techniques like timing correlation, functional correlation and path refinement to consider crosstalk noise during static timing analysis, they often suffered from descendant problems including overly-pessimistic timing bound, bad aggressor selection and false paths. Therefore, in this paper, a coupling-aware functional timing analysis tool named CA-FTA is proposed to tame the three problems stated above and to derive a tighter timing bound for the true longest path. Experimental results show that CA-FTA averagely reduces timing bounds of those obtained from crosstalk STA (with path refinement) by 7.26% on several ISCAS'85 and ISPD-2012 benchmark circuits (in particular, by 34.6% for b19). As a result, CA-FTA successfully solves the aggressor-selection problem as well as the false-path problem and relaxes more margins when considering crosstalk noise.
Jack S.-Y. Lin, Louis Y.-Z. Lin, Ryan H.-M. Huang, Charles H.-P. Wen
ACM Great Lakes Symposium on VLSI4
2017 Radiation-Hardened Designs for Soft-Error-Rate Reduction by Delay-Adjustable D-Flip-Flops
abstract
For reducing soft error rate (SER) in system-level failures, this paper proposes a radiation-hardened design by Delay-Adjustable D Flip-Flop (DAD-FF), which can be generally applied to sequential circuits such as shift registers. DAD-FF, modified from the Built-In Soft-Error Resilience (BISER) latch, can be easily integrated in the CAD flow and its delay can be adjusted to reject particle strikes with the maximum energy level. As a result, at the device level, DAD-FF eliminates 99.999997% soft errors by heavy ions on a satellite orbiting at a height of 720 km, and shows greater reduction on SER (e.g. 1.3E10X in the best case) than the standard DFF (STD-FF) through TCAD and SPICE simulation. Moreover, a real chip was also fabricated in a CMOS 90nm technology and performed the experiment of radiation exposure in UCL, Belgium. The laboratory measurement indicates that at the system level, the radiation-hardened design by DAD-FFs achieves 15.69X and 2.62X improvements on the overall SER, compared with those by STD-FFs and DICEs, respectively.
Yuwen Lin, Charles H.-P. Wen, Herming Chiueh
ACM Great Lakes Symposium on VLSI2
2017 TVM: Tabular VM migration for reducing hop violations of service chains in cloud datacenters
abstract
Network Function Virtualization is a recent and significant development applied in cloud datacenters, which can help service providers manage network services more flexibly and economically than using middle-boxes. User requests may demand different numbers of services in a chain, where the related flows can be directed to different physical machines (PMs) hosting required virtual network functions (VNFs) in a designated order. Sometimes, with the limited resources in current PMs, the deployment of service chains may result in great hop counts and cause hop (latency) violations against the service level agreements (SLAs) of user requests. Therefore, we propose a dynamic service-chain management approach called Tabular VM Migration (TVM) to reduce hop violations of service chains in cloud datacenters. Different from other VM management policies, TVM considers the order of a set of VMs in a service chain and migrates key VMs which may cause hop violations onto the selected PMs. Thus, hop violations are reduced with few migrations. Simulation results indicate that TVM performs better than other common VM management policies, and reduce at most 16.68% of total number of hop violations induced by a static deployment approach (SOVWin).
Ying-Feng Wu, Yu-Lun Su, Charles H.-P. Wen
ICC3
2017 Accelerating functional timing analysis with encoding duplication removal and redundant state propagation
abstract
Functional timing analysis (FTA) emerges for better timing closure than static timing analysis (STA) by providing the true delay of the circuit as well as its input pattern. For Satisfiability(SAT)-based FTA, a search problem for circuit delay can be expressed by clauses corresponding to circuit consistency function (CCF) and timed characteristic function (TCF). In particular, the clause number tends to grow exponentially as the circuit size increases, lengthening runtime for FTA. However, when formulating TCF, numerous clauses and literals are found useless. Therefore, two key techniques are proposed: (1) Encoding Duplication Removal (EDR) for removing those literals that are previously encoded in CCF but now duplicated in TCF, and (2) Redundant State Propagation (RSP) for propagating redundant states of nodes to help prune TCF clauses. Experiments indicate that under the worst-case delay of each benchmark circuit, EDR and RSP successfully reduce averagely 49% of clauses, 65% of literals, and 52% runtime on seven benchmark circuits for FTA.
Denny C.-Y. Wu, Pin-Ru Jhao, Charles H.-P. Wen
ICCAD3
2017 Speeding up power verification by merging equivalent power domains in RTL design with UPF
abstract
Low-power becomes a critical issue for modern VLSI designs. Unified Power Format (UPF) was invented for power management and enables the low-power design flow. In the UPF specification, controlling cells (including isolation cells, level shifter and retention cells) need to be placed properly to prevent unpredictable errors. Therefore, many commercial EDA tools support to examine the correctness of inserted cells and search missing/uncovered ones. However, such overall verification takes a long time for complex designs due to numerous power domains. Considering many of these power domains are equivalent and can be further merged, three strategies are proposed to explore (1) intra-scope domain equivalence, (2) inter-scope domain equivalence and (3) behavior-driven domain equivalence for RTL designs with UPF. For a case study on the OpenFire processor, the number of power domains is reduced from 4000+ to 500+, thus saving 77% time on signal checking in power verification.
Charles Chia-Hao Hsu, Charles H.-P. Wen
ITC-Asia2
2016 Speed binning with high-quality structural patterns from functional timing analysis (FTA)
abstract
In the nanometer era where the operating speed of a chip decides its price, design companies rely on high-qualty speed binning approaches to maxmizie their profits. The conventional speed binning approach is legacy (i.e. structural) since functional tests are too expensive to derive. Besides legacy and functional tests, recent studies tried to apply the notion of delay testing for deriving speed-binning patterns; however, all of them could not determine the number of patterns required for speed-binning nor taking process variation into consideration. Therefore, in this paper, we propose a speed-binning pattern generation (SBPG) method to deterministically generate a high-quality pattern set for speed binning. This SBPG mainly consists of two core techniques: (1) empirical variation sampling (EVS) and (2) functional timing analysis (FTA), which efficiently derives few high-quality patterns from a small number of learning samples. SBPG achieves a satisfactory accuracy (> 99% on average) for five benchmark circuits under various conditions of process variation, and is shown to be an efficient solution for speed binning.
Louis Y.-Z. Lin, Charles H.-P. Wen
ASP-DAC2
2016 Fast-yet-accurate variation-aware current and voltage modelling of radiation-induced transient fault
Hsuan-Ming Huang, Yuwen Lin, Charles H.-P. Wen
DATE3
2016 Layout-Based Soft Error Rate Estimation Framework Considering Multiple Transient Faults - From Device to Circuit Level
abstract
This paper investigated the soft errors caused by particle strikes, such as high-energy neutrons, extending beyond the deep submicrometer era. Considering the structure of the layout and resulting nuclear reactions, multiple transient faults (MTFs) tend to be induced more frequently than do single transient faults (STFs), due to the effects of technology scaling. This means that the soft error rates (SER) are beyond traditional netlist-based STF analysis, which can result in serious mis-estimations. This paper proposes a layout-based soft error estimation framework, which takes into account MTFs from the device level to the circuit level. This framework comprises two systems: 1) generation and 2) propagation. In the generation system, transient faults are modeled through nuclear reactions, charge collection, and voltage transformation at the device level. The propagation system abstracts these effects from the device level to the circuit level, taking into account three masking mechanisms associated with the propagation of transient faults. Experiment results demonstrate that the SER can be underestimated by an average of 15.72% if only single (rather than multiple) transient faults are taken into account. Our results indicate that netlist-based analysis for the estimation of SERs is no longer sufficient, due to the overwhelming influence of the structural layout. Thus, using benchmark c432, a tighter layout will result in an SER 34% higher than that generated in a looser layout.
Hsuan-Ming Huang, Charles H.-P. Wen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 TA-FTA: transition-aware functional timing analysis with a four-valued encoding
abstract
Timing analysis becomes profound for modern VLSI designs. Functional timing analysis (FTA) has emerged to eliminate false paths and provide better timing closure than traditional static timing analysis (STA). However, signal transitions effect, such as multiple input switching (MIS), which changes the pin-to-pin delay of a gate as well as the overall circuit delay, has not yet been considered in FTA. Therefore, a Transition-Aware FTA (TA-FTA) engine using a novel four-valued encoding for calculating true delay under the signal-transition effect is developed in this work. However, timing analysis becomes sophisticated once the signal-transition effect is concerned. Therefore, two techniques, cone separation and filtering (CSF) and quadratic dynamic search (QDS), are also proposed to speed up TA-FTA by more than two orders in time. Experimental results shows that after considering the MIS effect, in the benchmark circuits, the delay reported by our TA-FTA increases by 23% on average and by 38% for the worst case.
Jasper C. C. Chang, Ryan H.-M. Huang, Louis Y.-Z. Lin, Charles H.-P. Wen
DAC4
2015 An online thermal-constrained task scheduler for 3D multi-core processors
Chien-Hui Liao, Charles H.-P. Wen, Krishnendu Chakrabarty
DATE2
2015 SWF: Segmented Wildcard Forwarding for flow migration in OpenFlow datacenter networks
abstract
Networking and performance become major issues in cloud computing. Software Defined Networking (SDN) enables flow-level management and makes routing more effective and flexible for datacenter networks (DCNs). A prior dynamic routing algorithm named Flow Migration shows good results in improving throughput and avoiding congestion through simulation. However, as Flow Migration is applied to a SDN-based DCN, tremendous overhead occurs in modifying flow tables of switches. Therefore, a novel Segmented Wildcard Forwarding (SWF) scheme is proposed to facilitate Flow Migration. SWF leverages the symmetry of a Portland topology and applies wildcard matching for balancing network loading and lowering table-modification cost for flows. Experimental results show that on a Portland 4-8-8 datacenter network, the proposed SWF scheme achieves higher throughput over than the traditional SPB does, and even reduces ~70% modification cost used by the original implementation of Flow Migration. As a result, Segmented Wildcard Forwarding (SWF) is proven a mechanism that practically realizes dynamic routing of SDN but also improves network performance as well as cost for flow migration, simultaneously.
Kuan-Tsen Kuo, Charles H.-P. Wen, Cheng Suo, I-Chen Tsai
ICC2
2015 A Determinate Radiation Hardened Technique for Safety-Critical CMOS Designs
Ryan H.-M. Huang, Dennis K.-H. Hsu, Charles H.-P. Wen
J. Electron. Test.3
2015 Demystifying Iddq Data With Process Variation for Automatic Chip Classification
abstract
Iddq testing is an integral component of test suites for the screening of unreliable devices. As the scale of silicon technology continues shrinking, Iddq values and associated fluctuations increase. In addition, increased design complexity makes defect-induced leakage currents difficult to differentiate from full-chip currents. Consequently, traditional Iddq methods result in more test escapes and yield loss. This brief proposes a new test method, called σ-Iddq to provide the following: 1) Iddq analysis with process-parameter deduction and 2) the algorithm for automatic chip-classification called collective analysis without the need to manually determine threshold values. We randomly inserted a number of multiple defects into samples of ISCAS'89 and IWSL'05 benchmark circuits. Experimental results demonstrate that the proposed σ-Iddq method can achieve higher classification accuracy than single-threshold Iddq testing or AIddq in a 45-nm technology. The overall classification accuracy of the collective analysis achieve averaged 99.28% and 99.70% on σ-Iddq data from process-parameter deductions with average-case search and multilevel search, respectively, demonstrating that the influence of process variation and design scaling can be significantly reduced to enable a better identification of defective chips.
Chia-Ling Chang, Charles H.-P. Wen
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Thermal-Constrained Task Scheduling on 3-D Multicore Processors for Throughput-and-Energy Optimization
abstract
Thermal-constrained task scheduler for throughput optimization on 3-D multicore processors (3-D MCPs) has been studied extensively. However, these throughput-optimized strategies often ignore energy consumption and overuse thermal simulations. Therefore, in this brief, a new strategy named thermal-aware mapping and VoltagE scaling (TAMVES) is proposed to optimize throughput and energy consumption while satisfying thermal constraints (in terms of both peak temperature and temperature gradient) simultaneously. Layer-by-layer task-to-core mapping and thermal-and-energy-aware voltage scaling are incorporated in TAMVES to reduce peak temperature and temperature gradient without extensive thermal simulation. Furthermore, idle time slots are also utilized by voltage scaling for minimizing energy consumption. Our experimental results show that under thermal constraints, TAMVES outperforms a previous work (3-D Wave) by 35.30% averagely on throughput. In addition, TAMVES that features three-order faster speed under timing constraints outperforms 3-D Wave for saving 51.17% more energy and reducing 8.37% more peak temperature and 5.67% more temperature gradient. As a result, TAMVES has proven itself an effective task scheduler that optimizes throughput and energy on 3-D MCPs under thermal constraints.
Chien-Hui Liao, Charles H.-P. Wen
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Suppressing test inflation in shared-memory parallel Automatic Test Pattern Generation
abstract
Multi-core machines enable the possibility of parallel computing in Automatic Test Pattern Generation (ATPG). With sufficient computing power, previously proposed parallel ATPG has reached near linear speedup. However, test inflation in parallel ATPG yet arises as a critical problem and limits its practicality. Therefore, we developed a parallel ATPG system that incorporates (1) concurrent interruption (CI), (2) ripple compaction (RC) and (3) fan-in-cone based fault ordering (FIC) to deal with such problem. Concurrent interruption aborts test generation on simultaneously detected faults by fault simulation. Ripple compaction combines tests for different faults while fan-in-cone based fault ordering strategically arranges the fault list to reduce the number of test generations and thus speeds up the ATPG process. According to our experiments, the proposed parallel ATPG system effectively reduces 11% pattern count and achieves ~0% test inflation while maintaining an average of 6.5X speedup with no attenuation in fault coverage on experimental circuits.
Jerry C. Y. Ku, Ryan H.-M. Huang, Louis Y.-Z. Lin, Charles H.-P. Wen
ASP-DAC4
2014 Advanced Soft-Error-Rate (SER) Estimation with Striking-Time and Multi-Cycle Effects
abstract
Soft error rate (SER) has become a critical reliability issue for CMOS designs due to continuous technology scaling. However, the striking-time and multi-cycle effects have not been properly considered in SER for advanced CMOS designs. Therefore, in this paper, the striking-time and multi-cycle effects are formulated into the problem of SER estimation, and then a SER analysis framework is proposed, accordingly. Experimental results show that SERs on the benchmark circuits are seriously underestimated when ignoring both effects. Moreover, SERs increase more on those high-performance or low-power CMOS designs. New treatment to SER needs to be explored in the future.
Ryan H.-M. Huang, Charles H.-P. Wen
DAC2
2014 CASTA: CUDA-Accelerated Static Timing Analysis for VLSI Designs
abstract
General-purpose computing on graphics processing unit (GPGPU) enables the possibility of parallel computing for Static Timing Analysis (STA) of VLSI designs. However, memory access and synchronization between massively many cores become challenges to parallelizing STA. In this work, we developed a fast CUDA-Accelerated STA engine (named CASTA) that incorporates four novel techniques including Table-Index Remapping (TIR), Texture-Accelerated Rendering (TAR), Cell Levelization & Type Sorting (CLTS) and Timing-Table Restructuring(TTR) to enable high parallelism. Cell Levelization & Type Sorting (CLTS) levelizes cells and sort their types in order to efficiently access the same timing library. Timing-Table Restructuring (TTR) modifies the data structure for timing signals of cells to increase memory throughput. Table-Index Remapping (TIR) re-maps the axes of timing tables to retrieve data more efficiently while Texture-Accelerated Rendering (TAR) expands look-up tables (LUTs) to avoid extrapolation and stores LUTs in the texture for speed. As a result, our experimental result indicates that CASTA successfully enables high parallelism and outperforms a commercial tool by a three-order speedup on average over several benchmark circuits.
Hunta H.-W. Wang, Louis Y.-Z. Lin, Ryan H.-M. Huang, Charles H.-P. Wen
ICPP4
2013 Synthesizing multiple scan chains by cost-driven spectral ordering
abstract
Power cost and wire cost are two most critical issues in scan-chain optimization for modern VLSI testing. Many previous works used layout-based partitioning and greedy heuristics to synthesize multiple scan chains, making themselves suffer from (1) cost-metric monotonicity and (2) crossing-edge problem. Therefore, in this paper, we propose cost-driven spectral ordering including (1) cost-driven k-way spectral partitioning and (2) greedy non-crossing 2-opt ordering to resolve two above problems, respectively. Experiments show that different cost metrics can be properly addressed in k-way spectral partitioning. Moreover, our cost-driven spectral ordering achieves averagely 9% mixed (power-and-wire) reduction than two previous works on benchmark circuits, which evidently demonstrates its effectiveness on multiple scan-chain synthesis.
Louis Y.-Z. Lin, Christina C.-H. Liao, Charles H.-P. Wen
ASP-DAC3
2013 Flow-and-VM Migration for Optimizing Throughput and Energy in SDN-Based Cloud Datacenter
abstract
Minimizing energy consumption and improving performance in data centers are critical to cost-saving for cloud operators, but traditionally, these two optimization objectives are treated separately. Therefore, this paper presents an unified solution combining two strategies, flow migration and VM migration, to maximize throughput and minimize energy, simultaneously. Traffic-aware flow migration (FM) is first incorporated in dynamic reroute (DENDIST), evolving into DENDIST-FM, in a software-defined network (SDN) for improving throughput and avoiding congestion. Second, given energy and topology information, VM migration (ETA-VMM) can help reduce traffic loads and meanwhile save energy. Our experimental result indicates that compared to previous works, the proposed method can improve throughput by 42.5% on average with only 2.2% energy overhead. Accordingly, the unified flow-and-VM migration solution has been proven effective for optimizing throughput and energy in SDN-based cloud data centers.
Wei-Chu Lin, Chien-Hui Liao, Kuan-Tsen Kuo, Charles H.-P. Wen
CloudCom (1)4
2013 Process-variation-aware Iddq diagnosis for nano-scale CMOS designs - the first step
abstract
Along with the shrinking CMOS process and rapid design, scaling, both Iddq values and their variation of chips increase. As a result, the defect leakages become less significant when compared to the full-chip currents, making them more in-distinguishable for traditional Iddq diagnosis. Therefore, in this paper, a new approach called σ-Iddq duagnosis is proposed for reinterpreting original data and diagnosing failing chips, intelligently. The overall flow consists of two key components, (1) σ-Iddq transformation and (2) defect-syndrome matching: σ-Iddq transformation first manifests defect leakages by excluding both the process-variation and design-scaling impacts. Later, defect-syndrome matching applies data mining with a pre-built library to identify types and locations of defects on the fly. Experimental result show that an average of 93.68% accuracy with a resolution of 1.75 defect suspects can be achieved on ISCAS'89 and IWLS'05 benchmark circuits using a 45nm technology, demonstrating the effectiveness of σIddq diagnosis.
Chia-Ling Chang, Charles H.-P. Wen, Jayanta Bhadra
DATE2
2013 CASSER: A Closed-Form Analysis Framework for Statistical Soft Error Rate
abstract
CMOS designs in the deep submicrometer era require statistical methods to accurately estimate the circuit soft error rate (SER). However, process variation increases the complexity of statistical characteristics related to transient faults, leading to considerable uncertainty in the behavior of soft errors. Regardless of the methods used, current statistical SER (SSER) frameworks invariably involve a tradeoff between accuracy and efficiency. This paper presents accurate cell models in first-order closed form to overcome this problem, thereby enabling the analysis of SSERs in a block-based fashion similar to statistical static timing analysis. These cell models are derived as a closed form in the proposed framework named CASSER, and remain precise under the assumption of a normal distribution for the process parameters. Experimental results demonstrate the efficiency (> 2-order times faster than the latest framework) and accuracy ( error) of CASSER in estimating circuit SERs, when compared with the Monte Carlo SPICE simulation.
Austin C.-C. Chang, Ryan H.-M. Huang, Charles H.-P. Wen
IEEE Trans. Very Large Scale Integr. Syst.3
2013 Fast Scan-Chain Ordering for 3-D-IC Designs Under Through-Silicon-Via (TSV) Constraints
abstract
This brief addresses the problem of scan-chain ordering under a limited number of through-silicon vias (TSVs), and proposes a fast two-stage algorithm to compute a final order of scan flip-flops. To enable 3-D optimization, a greedy algorithm, multiple fragment heuristic, is modified and combined with a dynamic closest-pair data structure, FastPair, to derive a good initial solution in stage one. Stage two initiates two local refinement techniques, 3-D planarization and 3-D relaxation, to reduce the wire (and/or power) cost and to relax the number of TSVs in use to meet TSV constraints, respectively. Experimental results show that the proposed algorithm results in comparable performance (in terms of wire cost only, power cost only, and both wire-and-power cost) to a genetic-algorithm method but runs two-order faster, which makes it practical for TSV-constrained scan-chain ordering for 3-D-IC designs.
Christina C.-H. Liao, Allen W.-T. Chen, Louis Y.-Z. Lin, Charles H.-P. Wen
IEEE Trans. Very Large Scale Integr. Syst.4
2012 An intelligent analysis of Iddq data for chip classification in very deep-submicron (VDSM) CMOS technology
abstract
Iddq testing has been a critical integral component in test suites for screening unreliable devices. As the silicon technology keeps shrinking, Iddq values and their variation increase as well. Moreover, along with rapid design scaling, defect-induced leakage currents become less significant when compared to full-chip current and also make themselves less distinguishable. Traditional Iddq methods become less effective and cause more test escapes and yield loss. Therefore, in this paper, a new test method named σ-Iddq testing is proposed and integrates (1) a variation-aware full-chip leakage estimator and (2) a clustering algorithm to classify chip without using threshold values. Experimental result shows that σ-Iddq testing achieves a higher classification accuracy in a 45 nm technology when compared to a single-threshold Iddq testing. As a result, both the process-variation and design-scaling impacts are successfully excluded and thus the defective chips can be identified intelligently.
Chia-Ling Chang, Chia-Ching Chang, Hui-Ling Chan, Charles H.-P. Wen, Jayanta Bhadra
ASP-DAC4
2012 D2ENDIST: Dynamic and disjoint ENDIST-based layer-2 routing algorithm for cloud datacenters
abstract
This paper presents an improved layer-2 routing algorithm, called dynamic and disjoint edge node divided spanning tree (D2ENDIST), to overcome the issues of the single path route and unbalanced link utilization in cloud datacenters. D2ENDIST consists of two key schemes: (1) disjoint ENDIST routing and (2) reroute by dynamic reweights. The former scheme can provide multi-path routes, thereby reducing traffic congestion. The latter scheme can balance the traffic load and improve the link utilization. Our experimental results show that the proposed scheme can enhance system throughput by 25% subject to the constraint of very short failure recovery time compared to the existing ENDIST scheme.
Gen-Hen Liu, Charles H.-P. Wen, Li-Chun Wang 0001
GLOBECOM2
2012 Statistical Soft Error Rate (SSER) Analysis for Scaled CMOS Designs
abstract
This article re-examines the soft error effect caused by radiation-induced particles beyond the deep submicron regime. Considering the impact of process variations, voltage pulse widths of transient faults are found no longer monotonically diminishing after propagation, as they were formerly. As a result, the soft error rates in scaled electronic designs escape traditional static analysis and are seriously underestimated. In this article we formulate the statistical soft error rate (SSER) problem and present two frameworks to cope with the aforementioned sophisticated issues. The table-lookup framework captures the change of transient-fault distributions implicitly by using a Monte-Carlo approach, whereas the SVR-learning framework does the task explicitly by using statistical learning theory. Experimental results show that both frameworks can more accurately estimate SERs than static approaches do. Meanwhile, the SVR-learning framework outperforms the table-lookup framework in both SER accuracy and runtime.
Huan-Kai Peng, Hsuan-Ming Huang, Yu-Hsin Kuo, Charles H.-P. Wen
ACM Trans. Design Autom. Electr. Syst.4
2011 Diagnosing Multiple Byzantine Open-Segment Defects Using Integer Linear Programming
Chen-Yuan Kao, Chien-Hui Liao, Charles H.-P. Wen
J. Electron. Test.3
2010 Monte-Carlo-based statistical soft error rate (SSER) analysis for the deep sub-micron era
abstract
Variation in the deep sub-micron eras has made soft error rates (SERs) more statistical and difficult to capture using static analysis. Therefore, this paper presents a Monte-Carlo based SER analysis considering the statistical impact due to variation. Quasirandom sequences are also incorporated for fast convergence of SER accuracy and time efficiency. Experiments show that the proposed framework yields more accurate SERs compared to static analysis. On top of 106X speedup compared to Monte Carlo SPICE simulation, an additional 2.4X speedup can also be observed in the proposed framework after applying quasirandom sequences.
Yu-Shin Kuo, Huan-Kai Peng, Charles H.-P. Wen
ISCAS3
2009 On soft error rate analysis of scaled CMOS designs - A statistical perspective
abstract
This paper re-examines the soft error effect caused by cosmic radiation in sub 90nm technologies. Considering the impact of process variation, a number of statistical natures of transient faults are found more sophisticated than their static ones. We apply the state-of-the-art statistical learning algorithm to tackle the complexity of these natures and build compact yet accurate generation and propagation models for transient fault distributions. A statistical analysis framework for soft error rate (SER) is also proposed on the basis of these models. Experimental results show that the proposed framework can obtain improved SER estimation compared to the static approaches.
Huan-Kai Peng, Charles H.-P. Wen, Jayanta Bhadra
ICCAD2
2009 Speeding up bounded sequential equivalence checking with cross-timeframe state-pair constraints from data learning
abstract
A learning-and-filtering algorithm is proposed to uncover cross-timeframe state-pair constraints for speeding up SAT solving of bounded sequential equivalence checking (BSEC) problems. First, relaxed Boolean functions for flip-flop states at respective timeframes are learned from a small number of simulation data to derive the initial set of the state-pair candidates. Next, each candidate is examined and removed if both values in such a candidate have coordinately appeared during the simulation. Then, the validity of the remaining candidates is checked against the corresponding augmented circuit. Last, only the true constraints are annotated to the BSEC problems to facilitate SAT solving. All benchmark circuits are synthesized under 10 configurations to produce different BSEC problems. Experimental results show that the new SAT solving runs 2-order faster in average compared to using MiniSAT 2.0 only. Moreover, given a time bound, the total number of timeframes can increase by 8X-20X on 4 larger circuits after applying the proposed framework.
Chia-Ling Chang, Charles H.-P. Wen, Jayanta Bhadra
ITC2
2009 Portable simulation/emulation stimulus on an industrial-strength SoC
abstract
Reuse of system-on-chip (SoC) verification stimuli across various design models is a challenging problem. However, if used effectively, it significantly reduces verification time and quickly increases confidence in the robustness of a design. We use pseudo-random stimuli to drive tests on an SoC using simulation BFMs and reuse them on emulation-BFMs. Initial results on a Power Architecture¿Technology-based SoC demonstrate about a 100x speedup on the emulator vis-a¿-vis the simulator.
Rohit Srivastava, Javier Ruiz, Charles H.-P. Wen, Mrinal Bose, Jayanta Bhadra
ITC4
2007 An incremental learning framework for estimating signal controllability in unit-level verification
abstract
Unit-level verification is a critical step to the success of full-chip functional verification for microprocessor designs. In the unit-level verification, a unit is first embedded in a complex software that emulates the behavior of surrounding units, and then a sequence of stimuli is applied to measure the functional coverage. In order to generate such a sequence, designers need to comprehend the relationship between boundaries at the unit under verification and at the inputs to the emulation software. However, figuring out this relationship can be very difficult. Therefore, this paper proposes an incremental learning framework that incorporates an ordered-binary-decision-forest(OBDF) algorithm, to automate estimating the controllability of unit-level signals and to provide full-chip level information for designers to govern these signals. Mathematical analysis shows that the proposed OBDF algorithm has lower model complexity and lower error variance than the previous algorithms. Meanwhile, a commercial microprocessor core is also applied to demonstrate that controllability of input signals on the load/store unit in the microprocessor core can be estimated automatically and information about how to govern these signals can also be extracted successfully.
Charles H.-P. Wen, Li-C. Wang, Jayanta Bhadra
ICCAD1
2006 Simulation-based functional test justification using a decision-digram-based Boolean data miner
abstract
In simulation-based functional verification, composing and debugging testbenches can be tedious and time-consuming. A simulation data-mining approach, called TTPG (C. Wen, L-C Wang et al., 2005), was proposed as an alternative for functional test pattern generation. However, the core of simulation data-mining approach is Boolean learning, which tries to extract the simplified view of the design functionality according to the given bit-level simulation data. In this work1, an efficient data-mining engine is presented based on decision-diagram(DD)-based learning approaches. We compare the DD-based learning approaches to other known methods, such as the Nearest Neighbor method and support vector machine. We demonstrate that the proposed Boolean data miner is efficient for practical use. Finally, that the TTPG methodology incorporated with the Boolean data miner can achieve a high fault coverage (95.36%) on the OpenRISC 1200 microprocessor concludes the effectiveness of the proposed approach.
Charles H.-P. Wen, Onur Guzey, Li-C. Wang
ICCD1
2006 Simulation-Based Functional Test Generation for Embedded Processors
abstract
Deterministic functional test pattern generation has been a long-standing open problem, which is an important problem to be solved for both design verification and manufacturing testing. One key in developing a practical functional test pattern generation approach is to avoid the exponential growth of the test generation complexity in terms of the design size. This work proposes a novel functional test generation approach where simulation results are used to guide the generation of additional tests. Our methodology avoids the complexity growth issue by converting some modules in a design into simpler and more efficient models. Then, these models are used to facilitate the actual test generation process. We develop two sets of techniques to achieve these conversions: Boolean learning for random logic and arithmetic learning for datapath modules. We demonstrate the effectiveness and discuss the. limitations of these techniques through experiments on benchmark circuits. Last, we validate the overall test generation methodology based on the OpenRISC 1200 microprocessor
Charles H.-P. Wen, Li-C. Wang, Kwang-Ting Cheng
IEEE Trans. Computers1
2005 Simulation-based target test generation techniques for improving the robustness of a software-based-self-test methodology
abstract
Software-based self-test (SBST) was previously proposed as an on-chip functional test methodology. Achieving desired full-chip functional fault coverage has always been a challenge because random test program generation (RTPG) alone may not be sufficient. This work investigates the potential of using target test program generation (TTPG) to supplement the RTPG method. The proposed TTPG method utilizes simulation results to develop learned models for the surrounding modules of the block under test. Then, the learned models replace the surrounding modules around the block in the actual test generation process. Because the learned models are much simpler to handle, this method minimizes the cost of functional TPG. For developing the simulation-based learning scheme, we divide the surrounding modules into two categories: Boolean and arithmetic. We apply different techniques for each category and explain their applicability and limitations. The feasibility and effectiveness of the proposed simulation-based TTPG method in the context of supplementing RTPG for achieving high fault coverage in SBST of a RISC pipelined microprocessor design is demonstrated as well
Charles H.-P. Wen, Li-C. Wang, Kwang-Ting Cheng, Wei-Ting Liu, Ji-Jan Chen
ITC1
2005 On A Software-Based Self-Test Methodology and Its Application
abstract
Software-based self-test (SBST) was originally proposed for cost reduction in SOC test environment. Previous studies have focused on using SBST for screening logic defects. SBST is functional-based and hence, achieving a high full-chip logic defect coverage can be a challenge. This raises the question of SBST's applicability in practice. In this paper, we investigate a particular SBST methodology and study its potential applications. We conclude that the SBST methodology can be very useful for producing speed binning tests. To demonstrate the advantage of using SBST in at-speed functional testing, we develop a SBST framework and apply it to an open source microprocessor core, named OpenRISC 1200. A delay path extraction methodology is proposed in conjunction with the SBST framework. The experimental results demonstrate that our SBST can produce tests for a high percentage of extracted delay paths of which less than half of them would likely be detected through traditional functional test patterns. Moreover, the SBST tests can exercise the functional worst-case delays which could not be reached by even 1M of traditional verification test patterns. The effectiveness of our SBST and its current limitations are explained through these experimental findings.
Charles H.-P. Wen, Li-C. Wang, Kwang-Ting Cheng, Wei-Ting Liu, Ji-Jan Chen
VTS1