EDBT 2026 Demo / reviewers in the wild / expert
Ulf Schlichtmann
dblp:07/6841
· DBLP profile ↗
277ranked-venue papers
8as first author
129since 2021 · last 2026
0000-0003-4431-7619ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 267 · 8 first-author · 125 since 2021Software engineering, systems software and programming languages · 68 · 2 first-author · 27 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained TransformerabstractPrefix adders are widely used in compute-intensive applications for their high speed. However, designing optimized prefix adders is challenging due to strict design rules and an exponentially large design space. We introduce PrefixGPT, a generative pre-trained Transformer (GPT) that directly generates optimized prefix adders from scratch. Our approach represents an adder's topology as a two-dimensional coordinate sequence and applies a legality mask during generation, ensuring every design is valid by construction. PrefixGPT features a customized decoder-only Transformer architecture. The model is first pre-trained on a corpus of randomly synthesized valid prefix adders to learn design rules and then fine-tuned to navigate the design space for optimized design quality. Compared with existing works, PrefixGPT not only finds a new optimal design with a 7.7% improved area-delay product (ADP) but exhibits superior exploration quality, lowering the average ADP by up to 79.1%. This demonstrates the potential of GPT-style models to first master complex hardware design principles and then apply them for more efficient design optimization. Ruogu Ding, Ulf Schlichtmann, Weikang Qian |
AAAI | 3 |
| 2026 | Quantifying Compiler-induced Reliability Loss in Software-Implemented Hardware Fault ToleranceabstractCompiler mechanisms for Software-Implemented Hardware Fault Tolerance (SIHFT) offer a cost-effective solution for reliability, paving the way towards the adoption of Commercial Off-The-Shelf (COTS) components in safety-critical environments. However, default compiler optimizations can remove the SIHFT-induced redundancy and checks. For this reason, the use of compiler optimizations was discouraged in the literature. This article presents a comprehensive study of the reliability degradation introduced by LLVM’s O2 optimization pipeline when using a state-of-the-art SIHFT tool. We quantify, via RTL fault injection, the impact of O2 at different optimization stages, which identified a data corruption rate increase by up to $48 \times$. We also propose a static exploration methodology to identify the LLVM passes that harm the reliability. Then, we remove these harmful passes from the optimization pipeline, demonstrating how to tune optimization pipelines to make SIHFT successful even in the presence of compiler optimizations. Davide Baroffio, Johannes Geier, Federico Reghenzani, Ulf Schlichtmann, William Fornaciari |
ASP-DAC | 4 |
| 2026 | Built-In Self-Test for Locating Leakage Defects on Continuous-Flow Microfluidic Chips
Jiahui Peng, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
ASP-DAC | 4 |
| 2026 | Accessible Ratio-Specific Mixing: Single-Pressure-Driven Multi-Reagent Mixer Design and Synthesis for 3D-Printed MicrofluidicsabstractPrecise reagent mixing in user-defined ratios is a fundamental requirement in many microfluidic applications, including diagnostics, chemical synthesis, and biological assays. However, existing solutions for ratio-specific mixing often rely on complex active components, such as multiple pressure sources, flow controllers, or on-chip valves, making them costly, bulky, and unsuitable for portable or low-resource settings. In this work, we present a mixer design and a synthesis method for generating 3D-printable microfluidic devices that achieve ratiospecific mixing using only a single constant pressure source. Our method decomposes the desired mixing ratio into additive subcomponents, each represented by a dedicated inlet channel with a tailored length to enforce the correct hydraulic resistance. The method outputs a complete microfluidic layout, ready for direct fabrication via 3D printers. We validate our approach through numerical simulations and physical prototyping across eight diverse mixing scenarios. Results show that the achieved mixing ratios closely resemble the target, demonstrating the method’s accuracy and robustness. This work enables low-cost, portable, and accessible microfluidic devices for ratio-specific solution delivery, broadening the scope of microfluidics in settings where simplicity, reproducibility, and affordability are critical. Yushen Zhang, Debraj Kundu, Tsun-Ming Tseng, Sudip Roy 0001, Shigeru Yamashita, Ulf Schlichtmann |
ASP-DAC | 6 |
| 2026 | Multi-Partner Project: Advancing European Semiconductor and Chiplet Innovation Through the Bavarian Chip Design CenterabstractEurope’s semiconductor industry relies heavily on Asian and US manufacturers. The EU Chips Act seeks to strengthen Europe’s capabilities across the semiconductor value chain. Aligned with this goal, the Bavarian Chip Design Center (BCDC) supports local chip design, manufacturing, and talent development, with a focus on RISC-V computing and heterogeneous integration. Within BCDC, the Technical University of Munich and Fraunhofer are developing a chiplet-based architecture optimized for low-power edge AI. The system integrates two chiplets, combining a security-enhanced RISC-V core and AI accelerators, connected via a chiplet-optimized serial interface that supports encrypted data. The chiplets are mounted on a custom interposer with low-capacitance wires for efficient data transmission. System-and component-level development is currently ongoing, with a tapeout in 22 nm FD-SOI planned for 2027. The overall goal is to deliver a proof of concept for a small-scale energy-efficient chiplet system that demonstrates Bavaria’s and Europe’s capability to drive innovation in novel chip design fields. Hussam Amrouch, Jehaan Joseph, Michael Schirmer, Johannes Geier, Ulf Schlichtmann, Michael Meidinger, Thomas Wild, Andreas Herkersdorf, Jens Nöpel, Georg Sigl, Carsten Trinitis, Aswathy Nedumpalli Sankaranarayanan, Martin Schulz 0001, Andreas Korb, Konrad Hohentanner |
DATE | 5 |
| 2026 | Multi-Partner Project: A Holistic and Open-Source Approach to Efficient, Secure and Reliable AI Hardware Deployment in DI-EDAIabstractArtificial Intelligence (AI) has demonstrated strong capabilities across various domains over the past decade. Edge and specifically mission-critical applications, such as automotive and aerospace, require both high performance and efficiency without compromises in security and reliability. This stems from tightly constrained power consumption, failures that can have catastrophic consequences and devices that may be physically accessible to malicious actors. AI algorithm deployment to hardware also presents significant barriers, requiring specialized knowledge and expensive development tools. The DI-EDAI project aims to offer a holistic approach for connecting high-level AI algorithms with hardware implementations while tackling the aforementioned issues. Unlike other approaches that address individual aspects of the AI deployment flow, we investigate solutions across multiple layers of the design stack. Through our work we develop efficient hardware, map AI algorithms to hardware while simultaneously ensuring security and reliability. Furthermore, we leverage AI-techniques to assist with Electronic Design Automation (EDA) workflows for design optimization, verification and implementation. Our open source approach aims to reduce entry barriers, promote transparency and education, and spark innovation. This paper presents the current state of the DI-EDAI project at midterm, highlighting our latest contributions, identifying limitations in existing state-of-the-art approaches, and outlining ongoing work to address these gaps. Georgios Sotiropoulos, Felix Frombach, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Henrik Iver Thorøe, Vincent Meyers, Mehdi Baradaran Tahoori, Zeynep Demirdag, Mohammed Bakr Sikal, Hassan Nassar, Heba Khdr, Jörg Henkel, Christopher Wolters, Philipp van Kempen, Johannes Geier, Ulf Schlichtmann, Batuhan Sesli, Muhammad Sabih, Jakob Wittmann, Frank Hannig, Jürgen Teich, Lukas Steiner, Norbert Wehn, Mohamed Shelkamy Ali, Philipp Schmitz, Wolfgang Kunz, Stefan Koegler, Georg Sigl |
DATE | 17 |
| 2026 | G-PathGen: An Efficient GPU-Parallel k-Critical Path Generation AlgorithmabstractCritical path generation (CPG) plays a key role in many circuit timing analysis (CTA) applications. As the design complexity continues to increase, CPG runtime has become a major bottleneck in many timing-driven applications. To mitigate this runtime challenge, several CPU-based algorithms have been introduced by both the CTA and parallel computing communities, but they remain slow for large CPG problems. While GPU-accelerated solutions exist, they are often inexact and incur significant overhead from iterative CPU–GPU data transfers, limiting their practical use in CTA applications. To overcome this challenge, we propose G-PathGen, an exact GPU-parallel CPG algorithm targeting CTA applications. G-PathGen introduces efficient kernel algorithms for generating critical paths in parallel and dynamically adjusts the generated path count to maximize GPU utilization while minimizing redundant work. Compared to a state-of-the-art GPU solution, G-PathGen is 1.6 × –243.8 × faster when generating one million critical paths on industrial circuit graphs. Che Chang, Yi-Hua Chung, Cheng-Hsiang Chiu, Wan-Luan Lee, Boyang Zhang 0007, Ulf Schlichtmann, Ing-Chao Lin, Xiangyao Yu, Tsung-Wei Huang |
ICS | 6 |
| 2026 | Invited: Navigating the Frontier of Optimality and Complexity: Advanced Design Automation for Wavelength-Routed ONoCsabstractThe sustained growth of high-performance computing (HPC) workloads places increasing pressure on on-chip communication, where data movement, bandwidth, and power budget have become primary limiting factors. Optical networks-on-chip (ONoCs), particularly wavelength-routed ONoCs (WRONoCs), offer a transformative path toward ultra-high-speed and energy-efficient communication. The design of a WRONoC involves a complex interplay between two main design aspects: logic connection design (topological design) and layout synthesis (physical design). Existing design methodologies either separate the two design aspects into two sequential steps, which is computationally affordable but prone to suboptimality, or perform concurrent optimization that is theoretically holistic but computationally intractable. Such inefficiencies limit the scalability and generalizability of WRONoCs. To address these challenges, this paper presents two advanced design methodologies that have demonstrated their effectiveness in synthesizing high-performance WRONoCs. By pre-integrating layout constraints into the topological design phase, the logic connections created by both methodologies can be compatible with physical design, thereby resolving the long-standing conflict between solution quality and synthesis overhead. Zhidan Zheng, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
ISPD | 4 |
| 2026 | Path-Driven Washing and Drying Co-Optimization in Continuous-Flow Lab-on-ChipsabstractRapid advances in microfluidics technologies have facilitated the emergence of highly integrated lab-on-a-chip (LoC) biochip systems. With such a coin-sized biochip, complicated bioassay procedures can be executed efficiently without any human intervention. To ensure the correctness of assay outcomes, however, cross-contamination among different fluid samples and reagents needs to be dealt with separately during assay execution. As a consequence, washing operations have to be introduced and a washing path network needs to be established on the chip to remove the residues left behind in flow channels/devices. Also, chip drying after washing operations is crucial for maintaining some properties (e.g., pH values) of the subsequent reagents, so that precision degradation caused by residual buffer fluids can be avoided for those concentration-sensitive assays. To realize optimized assay procedures, we consider both washing operations and chip drying for the first time and propose an integer linear programming (ILP)-based path-driven washing and drying cooptimization method called PathDriver-WD for continuous-flow LoC biochip systems. The proposed method includes the following four key techniques: 1) The necessity of contamination removals and channel drying is analyzed systemically to avoid unnecessary washing and drying operations, 2) washing and drying operations are integrated with the regular removal of excess fluids, so that extra channel occupation can be minimized, 3) practical computation models are adopted to evaluate the durations of different washing and drying operations, and 4) optimized washing/drying paths and time windows are computed and assigned so that the completion time of assays can be minimized. Simulation results on multiple benchmarks demonstrate that the proposed method leads to highly efficient washing and drying procedures as well as minimized assay completion time. Xing Huang 0001, Zhiwen Yu 0001, Bin Guo 0001, Hanbin Ma, Tsung-Yi Ho, Ulf Schlichtmann, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | ParaVOM: Parallel-Execution-Aware Validation and Optimization for Multilayered Continuous-Flow Microfluidic BiochipsabstractMultilayered continuous-flow microfluidic biochips are rapidly advancing platforms for delicate bio-applications. The high complexity of biochip structures and application protocols drives the growing demand for design automation solutions. Current research enables the automatic synthesis of the physical layout and the scheduling and binding protocols of biochips, showcasing the significant potential of microfluidic design automation for improved resource utilization and reduced bioassay completion time. However, state-of-the-art synthesis methods primarily focus on device and operation levels, assuming flow paths are always available and neglecting interactions of flow and control channels. This creates a critical gap in the synthesis process, causing performance degradation, resource redundancy, or even infeasible designs. This work bridges this gap with a two-stage approach. Firstly, we perform a mathematical model to synthesize a high-level protocol that specifies the paths and execution orders of fluid transportation operations. Specifically, we construct flow paths based on the fluidic architecture of a given biochip design and optimize scheduling schemes to minimize the completion time of a given bioassay. Next, we perform a simulation-based synthesis of control channel pressurization sequences to realize the high-level protocol. Experimental results confirm that the proposed approach efficiently validates flow paths for feasible designs, identifies conflicting design features in infeasible designs, and improves the design efficiency and quality: compared to the original designs, it reduces the average number of control channels by 49%, and compared to the preliminary work, it reduces the average fluid transportation time by 19% and the average program run time by 38%. Meng Lian 0001, Shucheng Yang, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | HLSRewriter: Efficient Refactoring and Optimization of C/C++ Code with LLMs for High-Level SynthesisabstractIn High-Level Synthesis (HLS), refactoring a standard C/C++ code into its HLS-compatible version (HLS-C) still requires significant human effort. While various program scripts have been introduced to automate this process, the resulting code still contains many HLS-incompatible issues that need to be manually refactored and optimized by developers. Since Large Language Models (LLMs) have the ability to automate code generation, they can also be used for automated code refactoring and optimization in HLS. However, due to the limited training of LLMs, considering hardware and software simultaneously, hallucinations may occur when using LLMs for HLS, leading to synthesis failures. To address these challenges, we introduce HLSRewriter , an LLM-aided code refactoring and optimization framework that takes regular C/C++ code as input and automatically generates its corresponding optimized HLS-C code for hardware synthesis with minimal human intervention. To mitigate LLM hallucinations, a step-wise reasoning process is employed to analyze and detect HLS-incompatible errors. Afterwards, a repair library containing reference templates is efficiently created by scanning the HLS tool manual, followed by cooperation with a Retrieval-Augmented Generation (RAG) paradigm to guide the LLMs toward correct refactoring. In addition, a pipeline-aware decomposition strategy is introduced to progressively break down complex loop structures into smaller tasks with a balanced trade-off between latency and area, thereby enabling efficient pipelining and parallel execution. To further improve hardware efficiency, a bit width adjuster module is incorporated into this framework to optimize the precision of floating-point variables. Moreover, LLM-aided HLS optimization strategies are introduced to add/tune hardware directives in HLS-C code, thereby enhancing the performance of the final synthesized hardware. Experimental results demonstrate that the proposed LLM-aided framework can achieve higher refactoring pass rates and superior hardware performance in 24 real-world tasks compared with traditional approaches and the direct application of LLMs for code refactoring and optimization. The codes are open-sourced at this link: https://github.com/code-source1/catapult . Kangwei Xu, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ulf Schlichtmann, Bing Li 0005 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | Dynamic Topology-Aware Flow Path Construction and Scheduling Optimization for Multilayered Continuous-Flow Microfluidic BiochipsabstractMultilayered continuous-flow microfluidic biochips are highly valued for their miniaturization and high bio-application throughput. However, challenges arise as the dynamic connections of channels, adjusted to satisfy varying demands of fluid transportation at different moments, complicate the execution of bioassays. The existing methods often focus on device binding and operation scheduling during high-level synthesis but overlook the topological connections within the microfluidic network. This oversight leads to mismanagement of conflicts between fluid transportations and erroneous assumptions about constant flow velocities, resulting in decreased accuracy and efficiency or even infeasibility of bioassay execution. To address this problem, we mathematically model the flow velocity that varies according to the dynamic changes of the topological connections between the on-chip components during the execution of the bioassay. Further integrating the flow velocity model into the high-level synthesis, we propose a quadratic programming (QP) method that constructs flow paths and optimizes scheduling schemes to minimize the bioassay completion time. Experimental results confirm that, compared with the state-of-the-art approach, our method shortened the bioassay completion time by an average of 40.9%. Meng Lian 0001, Shucheng Yang, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
ASP-DAC | 5 |
| 2025 | CPONoC: Critical Path-aware Physical Implementation for Optical Networks-on-ChipabstractOptical networks-on-chips (ONoCs), which adopt optical waveguides, microring resonators (MRRs), and the wavelength division multiplexing (WDM) scheme to transmit optical signals, serve as promising solutions for integrating multi- and many-core systems to provide high-bandwidth, low-latency, and low-power on-chip communication. To minimize the insertion loss of a wavelength-routed ONoC (WRONoC) during physical implementation, existing studies either adopt conventional standard cell placement techniques or maximally avoid waveguide crossings; however, all of them ignore the fact that the critical path suffering from the maximum insertion loss dominates the overall power efficiency and system performance. In this work, we propose CPONoC, a critical-path-aware physical implementation tool for WRONoCs. Different from existing studies, CPONoC focuses on minimizing the insertion loss of the critical path using an iterative crossing-aware force-directed method, and it is compatible with different representative logic schemes and input configurations. Compared to the state-of-the-art design automation tools, CPONoC achieves an average reduction of 9.6% in maximum insertion loss. Zhidan Zheng, Shao-Yun Fang, Tsun-Ming Tseng, Ulf Schlichtmann |
ASP-DAC | 5 |
| 2025 | An Efficient General-Purpose Optical Accelerator for Neural NetworksabstractGeneral-purpose optical accelerators (GOAs) have emerged as a promising platform to accelerate deep neural networks (DNNs) due to their low latency and energy consumption. Such an accelerator is usually composed of a given number of interleaving Mach-Zehnder-Interferometers (MZIs). This interleaving architecture, however, has a low efficiency when accelerating neural networks of various sizes due to the mismatch between weight matrices and the GOA architecture. In this work, a hybrid GOA architecture is proposed to enhance the mapping efficiency of neural networks onto the GOA. In this architecture, independent MZI modules are connected with microring resonators (MRRs), so that they can be combined to process large neural networks efficiently. Each of these modules implements a unitary matrix with inputs adjusted by tunable coefficients. The parameters of the proposed architecture are searched using genetic algorithm. To enhance the accuracy of neural networks, selected weight matrices are expanded to multiple unitary matrices applying singular value decomposition (SVD). The kernels in neural networks are also adjusted to use up the on-chip computational resources. Experimental results show that with a given number of MZIs, the mapping efficiency of neural networks on the proposed architecture can be enhanced by 21.87%, 21.20%, 24.69%, and 25.52% for VGG16 and Resnet18 on datasets Cifar10 and Cifar100, respectively. The energy consumption and computation latency can also be reduced by over 67% and 21%, respectively. Sijie Fei, Amro Eldebiky, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann |
ASP-DAC | 5 |
| 2025 | HyperG: Multilevel GPU-Accelerated k-way Hypergraph PartitionerabstractHypergraph partitioning plays a critical role in computer-aided design (CAD) because it allows us to break down a large circuit into several manageable pieces that facilitate efficient CAD algorithm designs. However, as circuit designs continue to grow in size, hypergraph partitioning becomes increasingly time-consuming. Recent research has introduced parallel hypergraph partitioners using multi-core CPUs to reduce the long runtime. However, the speedup of existing CPU parallel hypergraph partitioners is typically limited to a few cores. To overcome these challenges, we propose HyperG, a GPU-accelerated multilevel k-way hypergraph partitioning algorithm. HyperG introduces an innovative balanced group coarsening and a sequence-based refinement algorithm to accelerate both the coarsening and uncoarsening stages. Experimental results show that HyperG outperforms both the state-of-the-art sequential and CPU-based parallel partitioners with an average speedup of 133× and 4.1× while achieving comparable partitioning quality. Wan-Luan Lee, Dian-Lun Lin, Cheng-Hsiang Chiu, Ulf Schlichtmann, Tsung-Wei Huang |
ASP-DAC | 4 |
| 2025 | 3M-DeSyn: Design Synthesis for Multi-Layer 3D-Printed Microfluidics with Timing and Volumetric Controlabstract3D printing has revolutionized microfluidic device fabrication, enabling rapid prototyping and intricate geometries. However, designing lab-on-a-chip systems remains challenging. In microfluidic devices, precise control over fluid behavior is crucial, requiring careful attention to both timing and volume. Current state-of-the-art design automation tools for microfluidics have limitations, particularly in addressing the specific challenges of 3D-printed microfluidics and user-defined timing and volumetric constraints, and no design synthesis tool exists targeting these domains. We present 3M-DeSyn, a novel design synthesis method for 3D-printed microfluidics that incorporates timing and volumetric constraints and outputs print-ready 3D modeling files. It automates the design process, allowing users to specify schematics and desired flow control parameters. The underlying methodology is based on mathematical modeling of fluidic behavior and utilizes constraint optimization programming to find optimized solutions. Experimental results show significant improvements in design time while enabling rapid development of custom microfluidic systems. Yushen Zhang, Dragan Raseta, Tsun-Ming Tseng, Ulf Schlichtmann |
ASP-DAC | 4 |
| 2025 | A Backup Resource Customization and Allocation Method for Wavelength-Routed Optical Networks-on-Chip TopologiesabstractWavelength-routed networks-on-chip (WRONoCs) are known for providing high-speed and low-power communication. Despite those advantages, the key components, microring resonators (MRRs), are prone to process and thermal variations, which cause signals to fail to reach their intended destinations. Thus, several WRONoC fault-tolerant methods propose to prepare a constant number of backups, which often leads to inefficient resource allocation, i.e. insufficient backups for the signals that are prone to errors, while more than enough backups for the signals that are barely affected, resulting in much power waste. In this work, we propose a dynamical backup resource allocation method for reliability maximization and power minimization in WRONoCs. Precisely, our method starts with accurately modeling the WRONoC faults, which considers the deviation of an MRR's default behavior as a Gaussian Distribution. Since signal paths consist of different numbers of MRRs, and the signals have different probabilities of deviating from their designated paths, our method customizes the number of backup paths for every signal and automatically allocates the minimum resources to optimize the reliability. Zhidan Zheng, You-Jen Chang, Liaoyuan Cheng, Tsun-Ming Tseng, Ulf Schlichtmann |
ASP-DAC | 5 |
| 2025 | Special Session - Hardware-Software Co-Design for Machine Learning Systems Made Open-SourceabstractChip technologies are crucial for the digital transformation of industry and society. Machine Learning (ML) and Artificial Intelligence (AI) are increasingly shaping both daily life and industrial applications, with AI hardware playing a vital role in enabling efficient and scalable ML deployment. However, significant challenges remain in bridging the gap between ML algorithm development and hardware implementation, particularly for edge ML applications where efficiency, power constraints, and adaptability are critical. In such resource-constrained environments, hardware-software co-design becomes essential to achieve the necessary trade-offs between performance, energy efficiency, and system responsiveness. One of the key bottlenecks in ML hardware development is the lack of seamless integration between ML toolchains and electronic design automation (EDA) tools for hardware synthesis and mapping. Current solutions often require extensive manual optimization and costly proprietary software, limiting accessibility and innovation. Open-source tools can play a transformative role in democratizing ML hardware design, fostering collaboration, and addressing the growing shortage of skilled professionals. This paper covers key aspects of hardware-software co-design for ML systems, such as ML algorithms, hardware design, compiler technologies and system security, with a focus on open-source solutions. We highlight the critical need for open-source toolchains that connect ML model development with hardware synthesis and optimization and present solutions for custom hardware, as well as FPGA accelerators. Mehdi Baradaran Tahoori, Vincent Meyers, Mahboobe Sadeghipourrudsari, Huashuangyang Xu, Jürgen Becker 0001, Tanja Harbaum, Felix Frombach, Julian Höfer, Georgios Sotiropoulos, Jörg Henkel, Zeynep Demirdag, Heba Khdr, Hassan Nassar, Ulf Schlichtmann, Johannes Geier, Philipp van Kempen, Georg Sigl, Stefan Koegler, Matthias Probst, Jürgen Teich, Frank Hannig, Muhammad Sabih, Batuhan Sesli, Norbert Wehn, Lukas Steiner, Wolfgang Kunz, Mohamed Shelkamy Ali |
CODES+ISSS | 14 |
| 2025 | Process-Variation-Aware Design Optimization for Wavelength-Routed Optical Networks-on-ChipabstractWith low latency and collision-free communication, wavelength-routed optical networks-on-chip (WRONoC) become an effective solution to the growing demands for multi-core communications. Microring resonators (MRRs), the primary optical components in WRONoC, are susceptible to process variation. Under process variation, the MRR’s transmission spectrum shifts, results reduced signal power and increased crosstalk. However, the impacts caused by process variation have not yet been considered in WRONoC designs. In this work, we propose a methodology to optimize the MRR radii and signal wavelengths to counter process variation. Specifically, we construct analytical models of expected signal power under MRR process variation and develop optimization methods to maximize the expected signal transmission power in WRONoC. Results show up to 7.51 dB improvement in worst-case expected signal transmission power compared to designs that do not consider process variation. Liaoyuan Cheng, Mengchu Li, Tsun-Ming Tseng, Martin Schottenloher, Ulf Schlichtmann |
DAC | 5 |
| 2025 | iG-kway: Incremental k-way Graph Partitioning on GPUabstractRecent advances in GPU-accelerated graph partitioning have achieved significant performance gains but remain limited to full graph partitioning, lacking support for incremental updates. This limitation is critical in CAD applications, where circuit graphs undergo iterative, incremental modifications during optimization. We present iG-kway, the first GPU-based incremental k-way graph partitioner. iG-kway features an incrementality-aware data structure and a refinement kernel that efficiently updates only affected vertices with minimal quality loss. Experiments show that iG-kway delivers up to $84 \times$ speedup over the state-of-the-art G-kway with comparable partitioning quality. Wan-Luan Lee, Shui Jiang, Dian-Lun Lin, Che Chang, Boyang Zhang 0007, Yi-Hua Chung, Ulf Schlichtmann, Tsung-Yi Ho, Tsung-Wei Huang |
DAC | 7 |
| 2025 | FT-MUX: A Fault-Tolerant Microfluidic Multiplexer DesignabstractContinuous-flow microfluidic chips are multilayered miniaturized platforms to manipulate small volumes of fluids with valves. There are two types of channels on a chip: flow channels for the reaction of fluids, and control channels for the actuation of valves. Multiplexers (MUXes) are essential microfluidic components for individually addressing many flow channels with few control channels. As the integration scale of microfluidic chips increases, the reliability of MUXes becomes a critical concern, as a single defective control channel in a MUX will affect a large part of the flow channels addressed by the MUX. This paper formally analyzes and identifies the design rules for a MUX to tolerate n defective control channels, and model the fault-tolerant MUX (FT-MUX) design problem as a binary constant weight code problem to minimize resource overheads. We demonstrate that FT-MUX improves resource efficiency by up to hundreds of times compared to the conventional fault-tolerant design method. Besides, given no less than 10 control channels, FT-MUX tolerates at least one defective control channel and addresses even more flow channels with equal or fewer resources than a standard MUX. The advantages become more significant as the integration scale increases. Mengchu Li, Jiahui Peng, Tsun-Ming Tseng, Ulf Schlichtmann |
DAC | 4 |
| 2025 | AutoRE: Bayesian-Optimization-based Automatic Reliability Enhancement Tool for Flow-based Microfluidic BiochipsabstractAs an emerging platform for biochemical experiments, flow-based microfluidic biochips are currently suffering from malfunctions caused by manufacturing defects, thereby having low yield. While many related studies have been conducted and reliability quantification models have been published, layout optimization methods are yet lacking. In this paper, we propose AutoRE, the first tool to automatically enhance reliability by optimizing layouts. AutoRE varies the layout within a certain range without changing its topology, and adopts Bayesian optimization (BO) to identify the most reliable variant. Experimental results demonstrate that AutoRE can efficiently and effectively improve the reliability across all testcases by around 40% on average. Siyuan Liang 0002, Yushen Zhang, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann, Tsung-Yi Ho |
DAC | 5 |
| 2025 | LA-MTL: Latency-Aware Automated Multi-Task LearningabstractMulti-Task Learning (MTL) aims to unify a variety of tasks into a single network for improved training and inference efficiency. This is particularly attractive for real-time applications that require simultaneous execution of multiple workloads in resource-constrained embedded environments. However, most MTL approaches focus on enhancing parameters efficiency and overall tasks metrics, often lacking explicit inference latency awareness in the optimization loop. The design space exploration should not compromise on the parameters efficiency or task accuracy objectives in order to meet latency requirements. To address this, we propose LA-MTL, an automated layer-level MTL policy search that incorporates a novel analytical latency factor (ALF). By accounting for local and global latencies during the MTL policy search, we derive solutions that balance task metrics, parameters efficiency and latency constraints. LA-MTL search on ResNet34 yields solutions with up to 50% lower latency on the Jetson AGX Orin while maintaining competitive metrics in semantic segmentation and depth estimation tasks with a +/-2 p.p., on the CityScapes dataset. Additionally, we achieve a superior parameters efficiency, surpassing the state-of-theart MTL parameters reduction by over 20 p.p. Experiments on benchmark datasets (CityScapes, NYUv2) demonstrate the effectiveness of our approach across various backbones including ResNet34, MobileNetV2, and MobileOne in its expanded form. Code is available at https://github.com/shamvbs/LA-MTL.1 Shambhavi Balamuthu Sampath, Sami Sawani, Moritz Thoma, Lukas Frickenstein, Pierpaolo Morì, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Ulf Schlichtmann, Claudio Passerone, Walter Stechele |
DAC | 9 |
| 2025 | SuperFast: Fast Supernet Training Using Initial KnowledgeabstractOnce-for-all based neural architecture search (NAS) proposes to train a supernet once and extract specialized subnets from it for efficient deployment. This decoupling between training and search enables easy multi-target deployment without retraining. Nevertheless, the initial training cost has remained extremely high, with SOTA approaches like ElasticViT and NASViT taking more than 72 and 83 GPU days respectively. While other approaches have tried to accelerate the training by warming up the largest model in the search space, we argue that this is suboptimal, and knowledge is easier scaled upward than downward. Hence, we propose SuperFast, a simple, plug and play workflow, that (I.) pretrains a subnet of the supernet search space, and (II.) distributes its knowledge within the supernet before the training. SuperFast offers a substantial acceleration in the supernet training, resulting in a significantly better accuracy vs. training-cost trade-off. Using SuperFast on both ElasticViT and NASViT supernets achieves the baseline’s accuracy $1.4 \times$ and $\mathbf{1. 8} \times$ faster on the ImageNet dataset. Moreover, for a given time budget, SuperFast improves accuracy vs. latency trade-offs for subnets, gaining 4.0 p.p. for the $20-50 \mathrm{~ms}$ range on Pixel 6. Code available in https://github.com/MoritzTho/SuperFast. Moritz Thoma, Emad Aghajanzadeh, Shambhavi Balamuthu Sampath, Pierpaolo Morì, Nael Fasfous, Alexander Frickenstein, Manoj Rohit Vemparala, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DAC | 9 |
| 2025 | Rapid Fault Injection Simulation by Hash-Based Differential Fault Effect Equivalence ChecksabstractAssessing a computational system's resilience to hardware faults is essential for safety and security-related systems. Fault Injection (FI) simulation is a valuable tool that can increase confidence in computational systems and guide hardware and software design decisions in the early stages of development. However, simulating hardware at low levels of abstraction, such as Register Transfer Level (RTL), is costly, and minimizing the effort required for large-scale FI campaigns is a significant objective. This work introduces Hash-based Differential Fault Effect Equivalence Checks to automatically terminate experiments early based on predicting their outcome. We achieve this by matching observed fault effects to ones already encountered in previous experiments. We generate these hashes from differentials computed by repurposing existing fast boot checkpoints from a state-of-the-art acceleration method. By integrating these approaches in an automated manner, we can accelerate a large-scale FI simulation of a CPU at RTL. We reduce the average simulation time by a factor of up to 25 compared to a factor of around 2 to 5 for state-of-the-art techniques. While maintaining 100 % accuracy, we can recover the faulty state through the stored differentials. Johannes Geier, Leonidas Kontopoulos, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DATE | 4 |
| 2025 | Multi-Partner Project: Advancing the EDA Tools Landscape for the European RISC-V Ecosystem in TRISTANabstractThe TRISTAN project aims to expand and industrialize the European RISC-V ecosystem to compete effectively with existing commercial alternatives. This initiative specifically targets the critical challenges in the development of Electronic Design Automation (EDA) tools, essential for RISC-V-based solutions, by leveraging the synergy between the open-source community and industrial solutions. This paper presents an overview of the current landscape of TRISTAN's EDA flow, highlighting specific tools and methodologies that streamline the early design phases of RISC-V-based systems. We explore the unique features of these tools, emphasizing how they complement each other to strengthen the overall design process. Fatma Jebali, Caaliph Andriamisaina, Mathieu Jan, Wolfgang Ecker, Florian Egert, Bernhard Fischer, Alessio Burrello, Daniele Jahier Pagliari, Sara Vinco, Giuseppe Tagliavini, Ingo Feldner, Andreas Mauderer, Axel Sauer, Arnór Kristmundsson, Alexander Schober, Téo Bernier, Matti Käyrä, Ulf Schlichtmann, Rocco Jonack |
DATE | 18 |
| 2025 | Loading-Aware Mixing-Efficient Sample Preparation on Programmable Microfluidic DeviceabstractSample preparation, where a certain number of reagents must be mixed in a specific volumetric ratio, is an integral step for various bio-assays. A programmable microfluidic device (PMD) is an advanced flow-based microfluidic biochip (FMB) platform, that considered to be very effective for sample preparation. However, the impact of mixer placement, reagents' distribution, and mixing time on the automation of sample preparation has not yet been investigated. We consider a mixing efficiency model controlled by the number of alternations “μ” of reagents along the mixing circulation path and propose a loading-aware placement strategy that maximizes the mixing efficiency. We use satisfiability modulo theories (SMT) and propose a one-pass strategy for placing the mixers and the reagents, that successfully enhance the loading and mixing efficiencies. Debraj Kundu, Tsun-Ming Tseng, Shigeru Yamashita, Ulf Schlichtmann |
DATE | 4 |
| 2025 | CorrectBench: Automatic Testbench Generation with Functional Self-Correction using LLMs for HDL DesignabstractFunctional simulation is an essential step in digital hardware design. Recently, there has been a growing interest in leveraging Large Language Models (LLMs) for hardware testbench generation tasks. However, the inherent instability associated with LLMs often leads to functional errors in the generated testbenches. Previous methods do not incorporate automatic functional correction mechanisms without human intervention and still suffer from low success rates, especially for sequential tasks. To address this issue, we propose CorrectBench, an automatic testbench generation framework with functional self-validation and self-correction. Utilizing only the RTL specification in natural language, the proposed approach can validate the correctness of the generated testbenches with a success rate of 88.85 %. Furthermore, the proposed LLM-based corrector employs bug information obtained during the self-validation process to perform functional self-correction on the generated testbenches. The comparative analysis demonstrates that our method achieves a pass ratio of 70.13 % across all evaluated tasks, compared with the previous LLM-based testbench generation framework's 52.18% and a direct LLM-based generation method's 33.33%. Specifically in sequential circuits, our work's performance is 62.18 % higher than previous work in sequential tasks and almost 5 times the pass ratio of the direct method. The codes and experimental results are open-sourced at the link: https://github.com/AutoBench/CorrectBench. Ruidi Qiu, Grace Li Zhang, Rolf Drechsler, Ulf Schlichtmann, Bing Li 0005 |
DATE | 4 |
| 2025 | Multi-Partner Project: Open-Source Design Tools for Co-Development of AI Algorithms and AI Chips: (Initial Stage)abstractChip technologies are crucial for the digital transformation of industry and society. Artificial Intelligence (AI) is playing an increasingly important role in both our daily lives and in industry. The development of advanced AI chip designs, essential for the successful deployment of AI, is of critical importance for innovation and competitiveness. However, challenges arise from the complexity of hardware development, expensive access to state-of-the-art design tools, and a global shortage of hardware experts. In addition to cost optimization, computational power, and energy consumption, security and trustworthiness are becoming increasingly important. This project aims to address these challenges in AI chip design by enabling efficient hardware development. We are developing a seamless transition between software-based AI model development and optimization, and efficient hardware implementation, while considering security, trustworthiness, and energy efficiency. An open-source approach plays a key role, facilitating access for small and medium-sized enterprises (SMEs) and expanding the community involved in AI chip design to help mitigate the shortage of skilled professionals. Mehdi Baradaran Tahoori, Jürgen Becker 0001, Jörg Henkel, Wolfgang Kunz, Ulf Schlichtmann, Georg Sigl, Jürgen Teich, Norbert Wehn |
DATE | 5 |
| 2025 | SRing: A Sub-Ring Construction Method for Application-Specific Wavelength-Routed Optical NoCsabstractWavelength-routed optical networks-on-chip (WR-ONoCs) attract ever-increasing attention for supporting high-speed communications with low power and latency. Among all WRONoC routers, optical ring routers attract much interest for their simple structures. However, current designs of ring routers have overlooked the customization problem: when adapting to applications that have specific communication requirements, current designs suffer high propagation loss caused by long worst-case signal paths and high splitter usage in power distribution networks (PDN). To address those problems, we propose a novel customization method to generate application-specific ring routers with multiple sub-rings, SRing. Instead of sequentially connecting all nodes in a large ring, we cluster the nodes and connect them with sub-ring waveguides to reduce the path length. Besides, we propose a mixed integer linear programming model for wavelength assignment to reduce the number of PDN splitters. We compare SRing to three state-of-the-art ring router design methods for six applications. Experimental results show that SRing can greatly reduce the length of the longest signal path, the worst-case insertion loss, and the number of splitters in the PDN, significantly improving the power efficiency. Zhidan Zheng, Meng Lian 0001, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
DATE | 5 |
| 2025 | Large Language Models (LLMs) for Verification, Testing, and Design
Chandan Kumar Jha 0001, Muhammad Hassan 0001, Khushboo Qayyum, Sallar Ahmadi-Pour, Kangwei Xu, Ruidi Qiu, Jason Blocklove, Luca Collini, Andre Nakkab, Ulf Schlichtmann, Grace Li Zhang, Ramesh Karri, Bing Li 0005, Siddharth Garg, Rolf Drechsler |
ETS | 10 |
| 2025 | A Lifetime Extension Framework for Communication-Intensive Systems Based on Wavelength-Routed Optical Networks-on-Chip
Zhidan Zheng, Liaoyuan Cheng, Jeng-De Chang, Tsun-Ming Tseng, Ing-Chao Lin, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | Accurate Fault Detection for Wavelength-Routed Optical Networks-on-Chip Under Thermal Variation
Zhidan Zheng, Liaoyuan Cheng, Tsun-Ming Tseng, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | WROXIM: A Network-Level Simulation Platform for Wavelength-Routed Optical Networks-on-ChipabstractTo meet the increasing demand for high-speed communication in many-cores systems, optical networks-on-chip (ONoCs) have gained attention for their ability to deliver low-latency and high-bandwidth data transmission. As a specific type of ONoCs, wavelength-routed ONoCs (WRONoCs) offer exclusive benefits such as collision-free and arbitration-free communication between cores. To guide the design and optimization of WRONoCs, many simulators have been developed to model WRONoC behavior at the device and circuit levels, offering detailed insights into photonic components. These tools have significantly advanced photonic design. As WRONoCs move closer to practical applications, network-level simulation becomes increasingly important for evaluating end-to-end communication behavior. Despite the need, existing simulators rarely support WRONoC modeling at the network level, leaving a critical gap in current toolchains. To address that, we propose the first network-level WRONoC simulation platform, WROXIM, adapted from an open-source NoC simulator, Noxim. It supports cycle-accurate modeling of optical communication behaviors and retains full compatibility with traditional electrical NoC simulations. To model realistic communication behaviors, WROXIM incorporates both optical and electrical components, including serializers, deserializers, waveguides, buffers, and processing elements (PEs). This modeling approach allows the simulator to capture end-to-end data movement from a PE through the optical interconnect to another PE. Given an application with multiple tasks, WROXIM supports task-driven workloads and provides key performance metrics such as latency, throughput, and energy consumption, offering a platform for design exploration and verification for WRONoCs. Jeng-De Chang, Zhidan Zheng, Liaoyuan Cheng, Liu-Xuan-Wei Zhang, Tsun-Ming Tseng, Ing-Chao Lin, Ulf Schlichtmann |
ICCAD | 7 |
| 2025 | FAB: Fast and Demand-Aware Bandwidth Allocation Method for Wavelength-Routed Optical Networks-on-ChipabstractWavelength-routed optical networks-on-chip (WRONoC) is a promising solution for achieving high-bandwidth, low-power on-chip communications. Previous work has aimed to reduce transmission latency in WRONoC by bandwidth allocation considering application communication demands. This is achieved by optimizing core-to-port mapping and configuring the radii of microring resonators (MRRs), which are the key routing components in WRONoC. However, previous work suffers from slow runtime and limited solution quality and fails to obtain feasible solutions for large applications given WRONoC topology. To address these, we propose a novel graph-based approach, FAB, to fast and efficiently reduce transmission latency in WRONoC. Specifically, we formulate core-to-port mapping and MRR radii configuration as two-stage weighted subgraph matching problems. In each stage, we construct graphs to model the demands or the bandwidth potential between vertices. By introducing strict weight-based matching strategies and progressively relaxing infeasible cases, FAB can quickly achieve optimized high-quality solutions. Experimental results demonstrate that FAB reduces worst-case transmission latency by up to 87.5% and accelerates optimization by up to 500,000× compared to the state-of-the-art method. Liaoyuan Cheng, Mengchu Li, Zhidan Zheng, Tsun-Ming Tseng, Ulf Schlichtmann |
ICCAD | 5 |
| 2025 | HLSTester: Efficient Testing of Behavioral Discrepancies with LLMs for High-Level Synthesis
Kangwei Xu, Bing Li 0005, Grace Li Zhang, Ulf Schlichtmann |
ICCAD | 4 |
| 2025 | Automated Graph-level Passes for TinyML Fault ToleranceabstractDeploying Machine Learning (ML) applications on Microcontroller Unit (MCU)-type devices, known as TinyML, poses significant challenges due to constrained resources. Consequently, this restricts the integration of fault tolerance mechanisms. Many redundancy techniques consume substantial amounts of already limited resources. To address this, Algorithm-Based Fault Tolerance (ABFT) methods for ML workloads focus on resource-intensive neural network operators at the kernel-level. However, this abstraction level introduces challenges, as these operators are often encapsulated within vendor-specific proprietary libraries. This work introduces automated and universal graph-level Data Flow Graph (DFG) passes supporting common kernel-level ABFT methods, along with a sophisticated redundancy mechanism called Dual Module Redundancy Island (DMRland). These DFG passes are implemented in the Tensor Virtual Machine (TVM) compiler framework and evaluated on a CPU-only execution platform with the MLPerf™Tiny benchmark. Our experimental results demonstrate that the proposed approach achieves competitive fault resilience in comparison to kernel-level ABFT methods, resulting in a reduction by two orders of magnitude for misclassification with combined ABFT and DMRland methods. Further, the run-time overhead (RTO) remains, on average, 19% lower than for comparable kernel-based implementations. Johannes Kappes, Johannes Geier, Phillipp van Kempen, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
IJCNN | 5 |
| 2025 | Enabling Machine Learning for Power Modeling via Artificial Netlist GenerationabstractPrecise power estimation is increasingly important in developing integrated circuits for both server and edge applications. Recent research focuses on using data-driven methods, like machine learning, to enable power estimation in various design stages, but it requires diverse and large datasets to ensure adequate training. However, the limited number of freely available circuit designs limits the potential of such an approach. One solution could be artificially generating netlists with realistic power behavior to enable accurate training.In this paper, we enhance an existing artificial netlist generation flow [1] for the development of power estimation tools. By enhancing the technology mapping algorithm to take the switching characteristics of individual gates into account, we exert direct influence on the power characteristics of the generated netlists. Philipp Fengler, Sani R. Nassif, Ulf Schlichtmann |
ISCAS | 4 |
| 2025 | COSMO: COmpressed Sensing for Models and Logging Optimization in MCU Performance ScreeningabstractIn safety-critical applications, microcontrollers must meet stringent quality and performance standards, including the maximum operating frequency$F_{\max}$. Machine learning models have proven effective in estimating$F_{\max}$by utilizing data from on-chip ring oscillators. Previous research has shown that increasing the number of ring oscillators on board can enable the deployment of simple linear regression models to predict$F_{\max}$. However, the scarcity of labeled data that characterize this context poses a challenge in managing high-dimensional feature spaces; moreover, a very high number of ring oscillators is not desirable due to technological reasons. By modeling$F_{\max}$as a linear combination of the ring oscillators’ values, this paper employs Compressed Sensing theory to build the model and perform feature selection, enhancing model efficiency and interpretability. We explore regularized linear methods with convex/non-convex penalties in microcontroller performance screening, focusing on selecting informative ring oscillators. This permits reducing models’ footprint while retaining high prediction accuracy. Our experiments on two real-world microcontroller products compare Compressed Sensing with two alternative feature selection approaches: filter and wrapped methods. In our experiments, regularized linear models effectively identify relevant ring oscillators, achieving compression rates of up to 32:1, with no substantial loss in prediction metrics. Nicolò Bellarmino, Riccardo Cantoro, Sophie M. Fosson, Martin Huch, Tobias Kilian, Ulf Schlichtmann, Giovanni Squillero |
IEEE Trans. Computers | 6 |
| 2025 | Deep Learning Strategies for Labeling and Accuracy Optimization in Microcontroller Performance ScreeningabstractIn safety-critical applications, microcontrollers must be compliant with the required quality constraints and performance standards, particularly in terms of the maximum operating frequency$(F_{\max })$. Machine learning (ML) models have proven effective in estimating$F_{\max }$by utilizing data extracted from on-chip ring oscillators (ROs), making them a valuable instrument for performance screening. However, the cost of obtaining labeled samples and the stringent accuracy needed by the model create hard challenges in this context. In order to address these, we explored three deep-learning (DL)-based key strategies: 1) semi-supervised learning with deep feature extractors: we leverage the abundance of unlabeled production data in a semi-supervised approach. Deep feature extractor models are employed to transform data into higher-dimensional spaces. These feature embeddings enable accurate performance prediction using simple linear regression, with a fraction of labeled data to reach baseline performances; 2) intrafamily transfer learning: when introducing new microcontroller products, with slightly different characteristics but the same set of ROs, previously trained deep feature extractors can be used, in a transfer learning fashion. This permits the use of significantly fewer labeled data compared to traditional methods; and 3) interfamily transfer learning: we extend the previous transfer learning concept to new microcontroller products with completely distinct characteristics. We aim to demonstrate that adapting the features set and fine-tuning DL feature extractors initially trained on specific legacy product data permits to yield better performance. Our research aims to provide a holistic framework for DL-based microcontroller performance screening to address the challenge of limited labeled data. The proposed methodologies significantly improve prediction accuracy and reduce the dependency on a large number of labeled samples, thus enhancing the efficiency and efficacy of ML-based microcontroller screening. The proposed framework enables models reuse, serving as a valuable baseline when new products are released. Nicolò Bellarmino, Riccardo Cantoro, Martin Huch, Tobias Kilian, Ulf Schlichtmann, Giovanni Squillero |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | A Scalable 2T-1FeFET-Based Content Addressable Memory Design for Energy Efficient Data SearchabstractContent addressable memory (CAM) is widely used in advanced machine learning models and data-intensive applications for associative search tasks, thanks to the highly parallel pattern matching capability. Most state-of-the-art CAM designs primarily aim to reduce the CAM cell area by utilizing nonvolatile memories (NVMs). However, there has been limited research on optimizing the design and energy efficiency of NVM-based CAMs for practical deployment in edge devices and AI hardware. This article introduces a general compact and energy efficient CAM design scheme that minimizes design overhead by using only one NVM device per cell. Our proposed CAM design realizes both binary CAM (BCAM) and multibit CAM (MCAM) by leveraging the binary and multilevel storage property of NVM devices without additional cell overheads. Additionally, we propose an adaptive matchline (ML) precharge and discharge scheme to further optimize search energy by significantly reducing the ML voltage swing. Ferroelectric field-effect transistors (FeFETs) serve as representative NVMs in our proposed design, and we present a 2T-1FeFET CAM array incorporating a sense amplifier that implements the proposed ML scheme. Evaluation results show that our proposed 2T-1FeFET BCAM design achieves energy efficiency improvements of$6.64\times $/$4.74\times $/$9.14\times $/$3.02\times $compared to CMOS/ReRAM/STT-MRAM/2FeFET BCAM arrays, while 2T-1FeFET MCAM design achieves$8.25\times $/$5.68\times $/$56.35\times $better-energy efficiency compared to ReRAM/3T-1FeFET/1FeFET-1R MACM arrays. Benchmarking results demonstrate that our BCAM/MCAM approach provides$3.2\times $/$3.7\times $and$2.0\times $/$2.2\times $energy-delay product improvement over the 2T-2R and 2FeFET CAM in accelerating query processing applications. Jiahao Cai, Hamza Errahmouni Barkam, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | Efficient Model Switching in RRAM-Based DNN AcceleratorsabstractResistive random access memory (RRAM) has emerged as a promising technology for deep neural network (DNN) accelerators, but programming every weight in a DNN onto RRAM cells for inference can be both time-consuming and energy-intensive, especially when switching between different DNN models. This article introduces a hardware-aware multimodel merging (HA3M) framework designed to minimize the need for reprogramming by maximizing weight reuse, while taking into account the hardware constraints of the accelerator. The framework includes three key approaches: 1) crossbar (XB)-aware model mapping (XAMM); 2) block-based layer matching (BLM); and 3) multimodel retraining (MMR). XAMM reduces the XB usage of the preprogrammed model on RRAM XBs while preserving the model’s structure. BLM reuses preprogrammed weights in a block-based manner, ensuring the inference process remains unchanged. MMR then equalizes the block-based matched weights across multiple models. Experimental results show that the proposed framework significantly reduces programming cycles in multi-DNN switching scenarios while maintaining or even enhancing accuracy, and eliminating the need for reprogramming. Fang-Yi Gu, Ing-Chao Lin, Bing Li 0005, Ulf Schlichtmann, Grace Li Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Manufacturing Cycle Time Optimization for Inkjet-Printed ElectronicsabstractInkjet-printed electronics has attracted considerable attention for low-cost mass production. High-density inkjet-printed designs can benefit from printing and drying in batches to avoid defects due to undesired ink redistribution and ink merging. The state-of-the-art approach decomposes the design into small objects, assigns the objects to different layers to be printed in different iterations, and minimizes the number of layers to reduce the number of iterations. However, it overlooks the differences in the printing and drying time between different layers and thus cannot properly model the impact of different layer assignment solutions on the manufacturing cycle time. In this work, we propose a row-based printing model that simulates the inkjet-printing mechanism and an integral Gaussian drying model that evaluates the local evaporation rate to approximate the printing and drying process of inkjet-printed manufacturing. Based on these models, we propose a mixed-integer-linear programming (MILP) method called the manufacturing model to minimize the manufacturing cycle time and avoid defects by optimally assigning objects to different iterations to be printed and dried in batches. Experimental results confirm that, compared with the preliminary work, manufacturing using our optimized solutions required up to 42.7% less time. Meng Lian 0001, Hu Peng, Mengchu Li, Yushen Zhang, Tsun-Ming Tseng, Bernhard Wolfrum, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | Combinatorial-Coding-Based High-Performance Microfluidic Control Multiplexer: Design, Synthesis, and AdaptationabstractFlow-based microfluidic biochips have emerged as a promising platform for biochemical experiments. These chips contain transportation channels and operational devices that are controlled by microvalves, which are actuated by external controllers. As the complexity of experiments conducted on these chips continues to increase, control multiplexers (MUXes) have become essential for actuating a large number of valves. However, current binary-coding-based MUXes do not fully utilize the coding capacity and suffer from reliability issues due to long total length of channels and high control channel density. In this article, we propose the combinatorial coding, a novel MUX coding strategy, along with an algorithm to synthesize combinatorial-coding-based MUXes (CoMUXes) of arbitrary sizes with the theoretical maximum coding capacity. We also develop a simplification method to reduce the number of valves and the total length of control channels in CoMUXes, thereby improving their reliability. Additionally, we develop a reliability-aware adaptation method to reliably integrate the CoMUXes into the main functional part of the designs. We compare CoMUX with state-of-the-art MUXes under different control demands with up to$10 \times 2^{13}$independent control channels. Experimental results show that CoMUXes can reliably address more independent control channels with fewer resources. For instance, when the number of control channels to be controlled is up to$10 \times 2^{13}$, compared to a state-of-the-art MUX, the optimized CoMUX reduces the number of required flow channels by 44% and the number of valves by 90%. The proposed adaptation method is also tested to be capable of significantly reducing area usage, total length of control channels, and the risk of having defects. Siyuan Liang 0002, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann, Tsung-Yi Ho |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | A Homogeneous FeFET-Based Time-Domain Compute-in-Memory Fabric for Matrix-Vector Multiplication and Associative SearchabstractMatrix-vector multiplication (MVM) and content-based search are two key operations in many machine learning workloads. This article proposes a ferroelectric FET (FeFET) time-domain compute-in-memory (TD-CiM) array that can accelerate both operations in a homogeneous fabric. We demonstrate that 1) the AND and xor/XNOR logic functions required by MVM and content-based search can be realized using a single compute-in-memory (CiM) cell composed of 2FeFETs connected in series; 2) an inverter chain-based TD-CiM array along with a two-phase time-domain computation principle of the TD-CiM can be employed to implement the MVM and content-based search functions; 3) a signal delay-to-digital output conversion can be implemented by associating a loading capacitor with each stage of the inverter chain-based TD-CiM array, ensuring the full digital compatibility; and 4) the proposed 2FeFET cell and inverter chain-based TD-CiM array are robust against FeFET variation according to our comprehensive theoretical and experimental validation. We show how the FeFET TD-CiM can be exploited to accelerate hyperdimensional computing (HDC) and adjusted to process different tasks through dynamic and fine-grained resource allocation. HDC application benchmarking results show that the proposed FeFET-based TD-CiM offers on average$106\times $/$63\times $energy reduction/speedup compared to GPU-based implementation. With more than 8500 TOPS/W energy-efficiency, the proposed FeFET-based TD-CiM exhibits huge potential as a processing fabric for various memory-intensive applications. Xunzhao Yin, Qingrong Huang, Hamza Errahmouni Barkam, Franz Müller 0001, Shan Deng, Alptekin Vardar, Sourav De 0002, Zhouhang Jiang, Mohsen Imani, Ulf Schlichtmann, Xiaobo Sharon Hu, Cheng Zhuo, Thomas Kämpfe, Kai Ni 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2025 | ZuSE-KI-Mobil: AI Chip Design Platform for Automotive and Industrial Applications
Shaown Mojumder, Simon Friedrich, Emil Matús, Matthias Lüders, Martin Friedrich, Oliver Renke, Holger Blume, Markus Kock, Gregor Schewior, Darius Grantz, Jens Benndorf, Julian Höfer, Patrick Schmidt 0003, Jürgen Becker 0001, Nael Fasfous, Pierpaolo Morì, Hans-Jörg Vögel, Samira Ahmadifarsani, Leonidas Kontopoulos, Ulf Schlichtmann, Yun-Jin Li, Gerhard P. Fettweis |
IEEE Trans. Very Large Scale Integr. Syst. | 20 |
| 2024 | Logic Design of Neural Networks for High-Throughput and Low-Power ApplicationsabstractNeural networks (NNs) have been successfully deployed in various fields. In NNs, a large number of multiply-accumulate (MAC) operations need to be performed. Most existing digital hardware platforms rely on parallel MAC units to accelerate these MAC operations. However, under a given area constraint, the number of MAC units in such platforms is limited, so MAC units have to be reused to perform MAC operations in a neural network. Accordingly, the throughput in generating classification results is not high, which prevents the application of traditional hardware platforms in extreme-throughput scenarios. Besides, the power consumption of such platforms is also high, mainly due to data movement. To overcome this challenge, in this paper, we propose to flatten and implement all the operations at neurons, e.g., MAC and ReLU, in a neural network with their corresponding logic circuits. To improve the throughput and reduce the power consumption of such logic designs, the weight values are embedded into the MAC units to simplify the logic, which can reduce the delay of the MAC units and the power consumption incurred by weight movement. The retiming technique is further used to improve the throughput of the logic circuits for neural networks. In addition, we propose a hardware-aware training method to reduce the area of logic designs of neural networks. Experimental results demonstrate that the proposed logic designs can achieve high throughput and low power consumption for several high-throughput applications. Kangwei Xu, Grace Li Zhang, Ulf Schlichtmann, Bing Li 0005 |
ASPDAC | 3 |
| 2024 | Late Breaking Results: Efficient Built-in Self-Test for Microfluidic Large-Scale Integration (mLSI)abstractControl channels on microfluidic large-scale integration (mLSI) chips are prone to blockage and leakage defects. In this work, we propose a built-in self-test (BIST) method that drastically improves the test efficiency. Given n to-be-tested control channels, we reduced the number of test patterns for blockage and leakage tests from [EQUATION] to 1, and from ⌈log2(n + 1)⌉ to ⌈log2(χ(G) + 1)⌉, respectively, where χ(G) denotes the vertex chromatic number of a graph G consisting of n vertices. We fabricated our design and demonstrated the feasibility and efficiency of our method. Mengchu Li, Hanchen Gu, Yushen Zhang, Siyuan Liang 0002, Hudson Gasvoda, Rana Altay, Ismail Emre Araci, Tsun-Ming Tseng, Tsung-Yi Ho, Ulf Schlichtmann |
DAC | 10 |
| 2024 | LaMUX: Optimized Logic-Gate-Enabled High-Performance Microfluidic Multiplexer DesignabstractAfter decades of development, flow-based microfluidic biochips have become an increasingly attractive platform for biochemical experiments. The fluid transportation and the on-chip device operation are controlled by microvalves, which are driven by external pneumatic controllers. To meet the increasingly complex experimental demands, the number of microvalves has significantly increased, making it necessary to adopt multiplexers (MUXes) for the actuation of microvalves. However, existing MUX designs have limited coding capacities, resulting in area overhead and excessive chip-to-world interface. This paper proposes a novel gate structure for modifying the current MUX architecture, along with a mixed coding strategy that achieves the maximum coding capacity within the modified MUX architecture. Additionally, an efficient synthesis tool for the mixed-coding-based MUXes (LaMUXes) is presented. Experimental results demonstrate that the LaMUX is exceptionally efficient, substantially reducing the usage of pneumatic controllers and microvalves compared to existing MUX designs. Siyuan Liang 0002, Yushen Zhang, Rana Altay, Hudson Gasvoda, Mengchu Li, Ismail Emre Araci, Tsun-Ming Tseng, Ulf Schlichtmann, Tsung-Yi Ho |
DAC | 8 |
| 2024 | Multi-Resonance Mesh-Based Wavelength-Routed Optical Networks-on-ChipabstractWavelength-routed optical networks-on-chip (WRONoCs) are well-known for providing high-speed and collision-free communication in multi-core processors. Previous work was unable to simultaneously reduce the design complexity and total optical power consumption of WRONoC. Besides, in current designs, each microring resonator (MRR), which is the key component of WRONoC, is configured to demultiplex to one specific wavelength. This significantly increases the MRR usage and the insertion loss. In this work, we adapt different types of ONoC routers into the mesh-based template. To reduce MRR usage, we take advantage of an important feature of MRR, multi-resonance, so that a single MRR can demultiplex signals on multiple wavelengths. To this end, we propose an efficient design method that synthesizes mesh-based WRONoCs using multi-resonance MRRs and existing optical routers to reduce total power consumption. The experimental results show that our method outperforms state-of-the-art design methods in significantly reducing MRR usage and optical power. Zhidan Zheng, Liaoyuan Cheng, Kanta Arisawa, Alexandre Truppel, Shigeru Yamashita, Tsun-Ming Tseng, Ulf Schlichtmann |
DAC | 8 |
| 2024 | Computational and Storage Efficient Quadratic Neurons for Deep Neural NetworksabstractDeep neural networks (DNNs) have been widely deployed across diverse domains such as computer vision and natural language processing. However, the impressive accomplishments of DNNs have been realized alongside extensive computational demands, thereby impeding their applicability on resource-constrained devices. To address this challenge, many researchers have been focusing on basic neuron structures, the fundamental building blocks of neural networks, to alleviate the computational and storage cost. In this work, an efficient quadratic neuron architecture distinguished by its enhanced utilization of second-order computational information is introduced. By virtue of their better expressivity, DNNs employing the proposed quadratic neurons can attain similar accuracy with fewer neurons and computational cost. Experimental results have demonstrated that the proposed quadratic neuron structure exhibits superior computational and storage efficiency across various tasks when compared with both linear and non-linear neurons in prior work. Chuangtao Chen 0001, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ulf Schlichtmann, Bing Li 0005 |
DATE | 5 |
| 2024 | A FeFET-based Time-Domain Associative Memory for Multi-bit Similarity ComputationabstractThe exponential growth of data across various domains of human society necessitates the rapid and efficient data processing. In many contemporary data-intensive applications, similarity computation (SC) is one of the most fundamental and indispensable operations. In recent years, In-memory computing (IMC) architectures have been designed to accelerate SC by reducing data movement costs, however, they encounter challenges with signal domain conversion, variation sensitivity, and limited precision. This paper proposes a ferroelectric FET (FeFET) based time-domain (TD) associative memory (AM) for energy efficient SC. Such TD design can convert its output (i.e., time interval) to digits with relatively simple sensing circuitry thus saves large amount of area and energy compared with conventional IMC designs that process analog voltage/current signals. The variable-capacitance (VC) delay chain structure in our design supports quantitative SC and enhances robustness against variations. Furthermore, by exploiting multi-domain ferroelctric FET (FeFET), our design is capable of performing SC on vectors with multi-bit element, enabling support for higher-precision algorithms. Simulation results show that the proposed TD-AM achieves 13.8x/1.47x energy saving of our design compared to CMOS/NVM based TD-IMC designs. Additionally, our design exhibits good robustness in monte carlo simulation with variation extracted from experimental measurements. Investigation on precision of hyperdimensional computing (HDC) show that higher element precision reduces the size of HDC model when considering to achieve same accuracy, indicating an improved efficiency. Benchmarkings against GPU demonstrate in general 2/3 orders of magnitude speedup/energy efficiency improvement of our design. Our proposed multi-bit TD-AM promises energy-efficient quantitative SC for diverse intensive data processing application, especially in energy-constrained scenarios. Qingrong Huang, Hamza Errahmouni Barkam, Jianyi Yang 0003, Thomas Kämpfe, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Mohsen Imani, Cheng Zhuo, Xunzhao Yin |
DATE | 9 |
| 2024 | PathDriver-Wash: A Path-Driven Wash Optimization Method for Continuous-Flow Lab-on-a-Chip SystemsabstractRapid advances in microfluidics technologies have facilitated the emergence of highly integrated lab-on-a-chip (LoC) biochip systems. With such coin-sized biochips, complicated bioassay procedures can be executed efficiently without any human intervention. To ensure the correctness and precision of assay outcomes, however, cross-contamination among different fluid samples/reagents needs to be dealt with separately during assay execution. As a consequence, wash operations have to be introduced and a wash path network needs to be established on the chip to remove the residues left in flow channels. To realize optimized assay procedures with efficient wash operations, we propose PathDriver-Wash in this paper, a path-driven wash optimization method for continuous-flow LoC biochip systems. The proposed method includes the following three key techniques: 1) The necessity of contamination removals is analyzed systemically to avoid unnecessary wash operations, 2) wash operations are integrated with the regular removal of excess fluids, so that extra path occupations caused by wash can be minimized, and 3) optimized wash paths and time windows are computed and assigned to wash operations, so that the completion time of assays can be minimized. Experimental results demonstrate that the proposed method leads to highly efficient wash procedures as well as minimized assay completion times. Xing Huang 0001, Zhiwen Yu 0001, Bin Guo 0001, Tsung-Yi Ho, Ulf Schlichtmann, Krishnendu Chakrabarty |
DATE | 6 |
| 2024 | ScanCamouflage: Obfuscating Scan Chains with Camouflaged Sequential and Logic GatesabstractScan chain is a commonly used technique in testing integrated circuits as it provides observability and controllability of the internal states of circuits. However, its presence can make circuits vulnerable to attacks and potentially result in confidential internal data leakage. In this paper, we propose a novel technique for obfuscating scan chains using camouflaged flip-flops, which are designed with the same layout as the original flip-flops but have the actual functionality of a buffer. Furthermore, we employ camouflaged logic gates interconnected in special configurations to increase the difficulty of SAT attack. Experimental results demonstrate that circuits with only a small number of flip-flops can already be protected by the proposed technique while incurring only a minimal area overhead. Tarik Ibrahimpasic, Grace Li Zhang, Michaela Brunner, Georg Sigl, Bing Li 0005, Ulf Schlichtmann |
DATE | 6 |
| 2024 | OplixNet: Towards Area-Efficient Optical Split-Complex Networks with Real-to-Complex Data Assignment and Knowledge DistillationabstractHaving the potential for high speed, high throughput, and low energy cost, optical neural networks (ONN s) have emerged as a promising candidate for accelerating deep learning tasks. In conventional ONNs, light amplitudes are modulated at the input and detected at the output. However, the light phases are still ignored in conventional structures, although they can also carry information for computing. To address this issue, in this paper, we propose a framework called OplixNet to compress the areas of ONNs by modulating input image data into the amplitudes and phase parts of light signals. The input and output parts of the ONN s are redesigned to make full use of both amplitude and phase information. Moreover, mutual learning across different ONN structures is introduced to maintain the accuracy. Experimental results demonstrate that the proposed framework significantly reduces the areas of ONNs with the accuracy within an acceptable range. For instance, 75.03 % area is reduced with a 0.33% accuracy decrease on fully connected neural network (FCNN) and 74.88% area is reduced with a 2.38% accuracy decrease on ResNet-32. Ruidi Qiu, Amro Eldebiky, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ulf Schlichtmann, Bing Li 0005 |
DATE | 6 |
| 2024 | Seal5: Semi-Automated LLVM Support for RISC-V ISA Extensions Including AutovectorizationabstractThe RISC-V instruction set architecture (ISA) is popular for its extensibility, allowing easy integration of custom vendor-defined instructions tailored to specific applications. However, a quick exploration of instruction candidates fails due to the lack of tools to auto-generate embedded software toolchain support. In particular, exploiting SIMD instructions to accelerate typical DSP and machine learning workloads needs specialized integration. This work establishes a semi-automated flow to generate LLVM compiler support for custom instructions based on a C-style ISA description language. The implemented Seal5 tool is capable of generating support for functionalities ranging from baseline assembler-level support, over builtin functions to compiler code generation patterns for scalar as well as vector instructions, while requiring no deeper compiler know-how. This paper focuses primarily on a novel pattern generator approach for the optimized code generation for SIMD instructions, including support for autovectorization. The auto generated LLVM toolchain reduces development times drastically while performing similarly or better compared to the existing, manually implemented Core-V reference LLVM toolchain on a wide variety of benchmarks. Seal5 further allows the addition of compiler code generation support for the Core-VSIMD instructions, which is not yet available in the reference toolchain. Additionally, Seal5 facilitates a quick exploration of custom instruction candidates as demonstrated for a cryptography extension. Philipp van Kempen, Mathis Salmen, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DSD | 4 |
| 2024 | Minimizing Worst-Case Data Transmission Cycles in Wavelength-Routed Optical NoC through Bandwidth AllocationabstractWith the rapid development of integrated photonic technology, wavelength-routed optical networks-on-chip (WRONoC) is emerging as a high-potential computing architecture due to its low power consumption, high bandwidth, and conflict-free communication advantages. Previous works utilize the multi-resonance properties of the microring resonator (MRR), the key component in WRONoC, to transmit multiple signals on different wavelengths in one transmission path, thereby achieving parallel communication. However, they do not consider the demands of the actual application communication bandwidth. If communications with high bandwidth demands are not allocated with highly parallel transmission paths, they may become the bottleneck of the network, resulting in increased transmission cycles and overall data transmission time. In this work, we propose an optimization strategy that allocates signal wavelengths for each communication based on its actual bandwidth demand to reduce communication time. Specifically, based on the actual bandwidth demands and topology structure, we first map the communication nodes in the target application to the ports of a WRONoC topology. Next, we optimize the radii of MRRs in the topology and allocate the signal wavelengths to each transmission path to minimize the worst-case data transmission cycle. Experimental results show that, compared to methods that only consider communication parallelism, our strategy can reduce the worst-case of data transmission cycles by over five times, thereby significantly decreasing the time required for data transmission. Liaoyuan Cheng, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
ICCAD | 4 |
| 2024 | BasisN: Reprogramming-Free RRAM-Based In-Memory-Computing by Basis Combination for Deep Neural NetworksabstractDeep neural networks (DNNs) have made breakthroughs in various fields including image recognition and language processing. DNNs execute hundreds of millions of multiply-and-accumulate (MAC) operations. To efficiently accelerate such computations, analog in-memory-computing platforms have emerged leveraging emerging devices such as resistive RAM (RRAM). However, such accelerators face the hurdle of being required to have sufficient on-chip crossbars to hold all the weights of a DNN. Otherwise, RRAM cells in the crossbars need to be reprogramed to process further layers, which causes huge time/energy overhead due to the extremely slow writing and verification of the RRAM cells. As a result, it is still not possible to deploy such accelerators to process large-scale DNNs in industry. To address this problem, we propose the BasisN framework to accelerate DNNs on any number of available crossbars without reprogramming. BasisN introduces a novel representation of the kernels in DNN layers as combinations of global basis vectors shared between all layers with quantized coefficients. These basis vectors are written to crossbars only once and used for the computations of all layers with marginal hardware modification. BasisN also provides a novel training approach to enhance computation parallelization with the global basis vectors and optimize the coefficients to construct the kernels. Experimental results demonstrate that cycles per inference and energy-delay product were reduced to below 1% compared with applying reprogramming on crossbars in processing large-scale DNNs such as DenseNet and ResNet on ImageNet and CIFAR100 datasets, while the training and hardware costs are negligible. Amro Eldebiky, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ing-Chao Lin, Ulf Schlichtmann, Bing Li 0005 |
ICCAD | 6 |
| 2024 | RABER: Reliability-Aware Bayesian-Optimization-based Control Layer Escape Routing for Flow-based MicrofluidicsabstractAfter decades of development, flow-based microfluidic biochips have become one of the most promising platforms for biochemical experiments. Control ports, which are remarkably area-consuming punch holes, are interfaces to external pneumatic controllers. To prevent the inserted outer catheters from hindering microscopic observation during experiments, control ports are placed on chip boundaries in practice. In this paper, we propose a practical and novel control layer escape routing methodology, which efficiently connects microvalves to user-specified boundaries. Particularly, the proposed methodology groups certain microvalves, and constructs a tree to connect them with the same control port, which is regarded as the root of the tree. Clustering more microvalves into the same group can reduce the usage of control ports, but will lead to more intensive connections among microvalves, which becomes larger obstacles for the routing of other microvalves outside the group, thereby reducing the routability. To derive an optimized tradeoff between the control port usage and the routability, we adapt a hierarchical clustering algorithm with a dynamically changing threshold that ascertains the closeness of the microvalves. We also adopt the Bayesian optimization (BO) to determine the optimized routing order for better routing results. Additionally, we propose a fault-tolerant structure as an option for users, which only occupies little area around control channels, and significantly improves the reliability against blockage defects. Experimental results demonstrate that the proposed methodology can efficiently connect all microvalves to user-specified boundaries, significantly reduce control port usage, shorten control channels, and improve reliability compared to baseline methods. Siyuan Liang 0002, Rongliang Fu, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann, Tsung-Yi Ho |
ICCAD | 5 |
| 2024 | Aging-Aware Energy-Efficient Task Deployment of Heterogeneous Multicore SystemsabstractHeterogeneous multicore systems, which consist of high-performance and power-efficient cores, are emerging to satisfy the various demands on performance and power consumption. On the other hand, as CMOS technology continues to shrink in size, the aging effect, which can cause performance degradation or timing failures, has become a non-negligible threat to lifetime reliability. To overcome the challenges under the aging effect, various approaches have been proposed in previous studies. Most previous studies, however, did not consider the different characteristics of big and little cores. In addition, most of them do not consider critical tasks with the strict timing requirements present in real-time applications, resulting in early system failure. Therefore, considering different characteristics of cores and the presence of critical tasks, we propose an aging-aware task deployment framework for real-time systems. In this framework, for high-performance big cores, we propose a novel asymmetric aging-aware strategy. This strategy finds an energy-efficient task-to-core assignment to reserve some healthy cores at the early system life stage. The reserved cores are kept idle with the lowest voltage and can execute critical tasks at the late system life stage, extending the system lifetime. Meanwhile, the non-reserved cores use lower voltages to execute tasks, reducing the aging effect. For energy-efficient little cores, we adopt the symmetric aging-aware strategy to balance out the aging effect of each little core. With a balanced aging effect, the utilization of little cores is improved. In addition, we propose voltage/frequency boosting and task migration techniques to increase the number of cores that can meet the task timing constraints. Compared to the state of the art, the proposed framework can achieve 1.10x lifetime improvement and 5% energy reduction. Yu-Guang Chen, Chieh-Shih Wang, Ing-Chao Lin, Zheng-Wei Chen, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | CorrectNet+: Dealing With HW Non-Idealities in In-Memory-Computing Platforms by Error Suppression and CompensationabstractThe last decade has witnessed the breakthrough of deep neural networks (DNNs) in many fields. With the increasing depth of DNNs, hundreds of millions of multiply-and-accumulate (MAC) operations need to be executed. To accelerate such operations efficiently, analog in-memory computing platforms based on emerging devices, e.g., resistive RAM (RRAM), have been introduced. These acceleration platforms rely on analog properties of the devices and thus suffer from process variations. Consequently, weights in neural networks configured into these platforms can deviate from the nominal trained values, which may lead to feature errors and a significant degradation of the inference accuracy. Besides, additional HW aspects represent key controlling factors for such computing platforms, namely, the limited RRAM cell programmable conductance levels, which limits the number of bits stored in one RRAM cell, the ADC noise converting analog values to digital domain and the ADC power scaling with the number of bits of its output. To address these points, in this article, we propose a framework to enhance the robustness of neural networks under variations. First, an enhanced Lipschitz constant regularization is adopted during neural network training to suppress the amplification of errors propagated through network layers. Additionally, the quantization setting of a NN model is optimized considering robustness against weight variations and total ADC power consumption. Afterward, error compensation is introduced at necessary locations determined by reinforcement learning (RL) to rescue the feature maps with remaining errors. Experimental results demonstrate that inference accuracy of neural networks can be recovered from as low as 1.69% under variations back to more than 95% of their original accuracy at the highest level of variations and reducing total ADC power consumption by 55% while the training and hardware cost are negligible. Amro Eldebiky, Grace Li Zhang, Georg Böcherer, Bing Li 0005, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Flexible Generation of Fast and Accurate Software Performance Simulators From Compact Processor DescriptionsabstractTo find optimal solutions for modern embedded systems, designers frequently rely on the software performance simulators. These simulators combine an abstract functional description of a processor with a nonfunctional timing model to accurately estimate the processor’s timing while maintaining high simulation speeds. However, current performance simulators either inflexibly target specific processors or sacrifice accuracy or simulation speed. This article presents a new approach to the software performance simulation, combining flexibility with highly accurate estimates and high simulation speed. A code generator converts a compact structural description of the target processor’s pipeline into sets of timing constraints, describing the processor’s instruction execution. Based on these, it generates corresponding scheduling functions and timing variables, representing the availability of the modeled pipeline. The performance estimator uses these components to approximate the processor’s timing based on an instruction trace provided by an instruction set simulator. Results for the state-of-the-art CV32E40P and CVA6 RISC-V processors show an average relative error of 0.0015% and 3.88%, respectively, over a large set of benchmarks. Our approach reaches an average simulation speed of 24 and 15 million instructions per second (MIPS), respectively. Conrad Foik, Robert Kunzelmann, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Control-Logic Synthesis of Fully Programmable Valve Array Using Reinforcement LearningabstractFully programmable valve array (FPVA) biochips have emerged as a promising alternative for traditional application-specific microfluidic platforms thanks to their advantages in terms of flexibility and reconfigurability. By regularly deploying microvalves along vertical and horizontal flow channels, microfluidic modules with different sizes and shapes can be constructed dynamically on the chip, thereby enabling the automatic execution of various assay procedures in biology and biochemistry. The above advantages, however, result largely from the large-scale integration of valves as well as accurate control of their switchings, leading to very complicated control-logic design of such chips. In this article, we propose an reinforcement learning (RL)-based synthesis flow for the control-logic design of fully programmable valve array (FPVA) biochips, taking multichannel switching and control-cost minimization into consideration simultaneously. By employing a double deep$Q$-network (DDQN) and two Boolean-logic simplification techniques, control logics with both high-switching efficiency and low-fabrication cost can be constructed automatically. Furthermore, the solution space of multichannel-switching combinations is reduced to improve the search efficiency of the proposed method. Experimental results on multiple benchmarks demonstrate that the proposed synthesis flow leads to better-design solutions compared with the state-of-the-art techniques. Xing Huang 0001, Huayang Cai, Wenzhong Guo, Genggeng Liu, Tsung-Yi Ho, Krishnendu Chakrabarty, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | A Hardware Friendly Variation-Tolerant Framework for RRAM-Based Neuromorphic ComputingabstractEmerging resistive random access memory (RRAM) attracts considerable interest in computing-in-memory by its high efficiency in multiply-accumulate operation, which is the key computation in the neural network (NN). However, due to the imperfect fabrication, RRAM cells suffer from the variations, which make the values in RRAM cells deviate from the target values so that the accuracy of the RRAM-based NN accelerator degrades significantly. Moreover, in a practical hardware design of RRAM-based NN accelerators, if the number of wordlines and bitlines in a crossbar array activated at the same time increases, ADCs with a high resolution are required and the power consumption of ADC increases. This paper proposes a novel methodology to mitigate the impact of variations in RRAM-based neural network accelerators. The methodology includes a unary-based non-uniform quantization method and a variation-aware operation unit (OU) based framework. The unary-based non-uniform quantization method equalizes the significance of weights stored in each RRAM cell to reduce the impact of variations. The variation-aware OU-based framework activates only RRAM cells in the same OU at the same time, which reduces the power consumption of ADCs. Additionally, the framework introduces three methods, including OU skipping, OU recombination, and OU compensation, to further mitigate the impact of variations. The experiments show that the proposed approach outperforms the state-of-the-art among four NN models on two datasets with 2-bit cell resolution. Fang-Yi Gu, Cheng-Han Yang, Ing-Chao Lin, Da-Wei Chang, Darsen D. Lu, Ulf Schlichtmann |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | A 3D Hybrid Optical-Electrical NoC Using Novel Mapping Strategy Based DCNN Dataflow AccelerationabstractA large number of multiply-accumulate operations and memory accesses required in deep convolutional neural networks (DCNN) leads to high latency and energy consumption (EC), that hinder their further applications. Dataflow-based acceleration schemes reduce memory accesses by leveraging reusable data in DCNNs. Row Stationary (RS) dataflow is a more advanced dataflow. In the convolutional layer acceleration of RS dataflow, the flexibility of mapping from logical processing element (LPE) sets to physical PE sets is relatively poor. The utilization of processing elements (PEs) is low. In this paper, a novel mapping strategy based on genetic algorithm (GAMS) with the goal of optimizing EC is proposed. GAMS is designed to address the energy inefficiencies faced when mapping RS dataflow. A 3D hybrid optical-electrical Network-on-Chip (3DHOENoC) is proposed to further improve the communication efficiency, energy efficiency and the processing speed of DCNN. Simulation and evaluation results show that GAMS can achieve better mapping flexibility, higher PEs utilization and 15.9% improvement of execution speed on average. In addition, the execution time (ET) performance of processing the DCNN can be further improved by adopting the 3DHOENoC architecture with better communication parallelism. Bowen Zhang 0004, Huaxi Gu, Grace Li Zhang, Yintang Yang, Ziteng Ma, Ulf Schlichtmann |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2023 | CompaSeC: A Compiler-Assisted Security Countermeasure to Address Instruction Skip Fault Attacks on RISC-VabstractFault-injection attacks are a risk for any computing system executing security-relevant tasks, such as a secure boot process. While hardware-based countermeasures to these invasive attacks have been found to be a suitable option, they have to be implemented via hardware extensions and are thus not available in most Commonly used Off-The-Shelf (COTS) components. Software Implemented Hardware Fault Tolerance (SIHFT) is therefore the only valid option to enhance a COTS system's resilience against fault attacks. Established SIHFT techniques usually target the detection of random hardware errors for functional safety and not targeted attacks. Using the example of a secure boot system running on a RISC-V processor, in this work we first show that when the software is hardened by these existing techniques from the safety domain, the number of vulnerabilities in the boot process to single, double, triple, and quadruple instruction skips cannot be fully closed. We extend these techniques to the security domain and propose Compiler-assisted Security Countermeasure (CompaSeC). We demonstrate that CompaSeC can close all vulnerabilities for the studied secure boot system. To further reduce performance and memory overheads we additionally propose a method for CompaSeC to selectively harden individual vulnerable functions without compromising the security against the considered instruction skip faults. Johannes Geier, Lukas Auer, Daniel Mueller-Gritschneder, Uzair Sharif, Ulf Schlichtmann |
ASP-DAC | 5 |
| 2023 | Enabling Inter-Product Transfer Learning on MCU Performance ScreeningabstractIn safety-critical applications, microcontrollers must meet strict quality and performance standards, including the maximum operating frequency$(F_{\max})$. Machine learning (ML) models can estimate$F_{\max}$using data from on-chip ring oscillators (ROs), making them suitable for performance screening. However, when new products are introduced, existing ML models may no longer be suitable and require updating. Training a new model from scratch is challenging due to limited data availability. Acquiring$F_{\max}$data is time-consuming and costly, resulting in a small labeled dataset. However, a large amount of data from legacy products may be available, along with existing ML models. In order to address the scarcity of labeled data, this paper proposes using deep learning feature extractors trained on specific MCU product data and fine-tuning them for new devices, in a Transfer Learning fashion. Experimental results show that these models can extract useful general features for performance prediction. As a result, they achieve better performance with significantly less labeled data compared to traditional shallow learning approaches. Nicolò Bellarmino, Riccardo Cantoro, Martin Huch, Tobias Kilian, Ulf Schlichtmann, Giovanni Squillero |
ATS | 5 |
| 2023 | PowerPruning: Selecting Weights and Activations for Power-Efficient Neural Network AccelerationabstractDeep neural networks (DNNs) have been successfully applied in various fields. A major challenge of deploying DNNs, especially on edge devices, is power consumption, due to the large number of multiply-and-accumulate (MAC) operations. To address this challenge, we propose PowerPruning, a novel method to reduce power consumption in digital neural network accelerators by selecting weights that lead to less power consumption in MAC operations. In addition, the timing characteristics of the selected weights together with all activation transitions are evaluated. The weights and activations that lead to small delays are further selected. Consequently, the maximum delay of the sensitized circuit paths in the MAC units is reduced even without modifying MAC units, which thus allows a flexible scaling of supply voltage to reduce power consumption further. Together with retraining, the proposed method can reduce power consumption of DNNs on hardware by up to 73.9% with only a slight accuracy loss. Richard Petri 0002, Grace Li Zhang, Yiran Chen 0001, Ulf Schlichtmann, Bing Li 0005 |
DAC | 4 |
| 2023 | CorrectNet: Robustness Enhancement of Analog In-Memory Computing for Neural Networks by Error Suppression and Compensation
Amro Eldebiky, Grace Li Zhang, Georg Böcherer, Bing Li 0005, Ulf Schlichtmann |
DATE | 5 |
| 2023 | VE-FIDES: Designing Trustworthy Supply Chains Using Innovative Fingerprinting ImplementationsabstractThe project VE-FIDES will contribute with a solution based on an innovative multi-level fingerprinting approach to secure electronics supply chains against the threats of malicious modification, piracy, and counterfeiting. Hardware-fingerprints are derived from minuscule, unavoidable process variations using the technology of Physical Unclonable Functions (PUFs). The derived fingerprints are processed to a system fingerprint enabling unique identification, not only of single components but also on PCB level. With the proposed concept, we show how the system fingerprint can enhance the trustworthiness of the overall system. For this purpose, the complete system including tiny sensors, a Secure Element and its interface to the application is considered in VE-FIDES. New insights into methodologies to derive component and system fingerprints are gained. These techniques for the verification of system integrity are complemented by methods for preventing reverse engineering. Two application scenarios are in the focus of VE-FIDES: Industrial control systems and an automotive use case are considered, giving insights to a wide spectrum of requirements for products built from components provided by international supply chains. Bernhard Lippmann, Joel Hatsch, Stefan Seidl, Detlef Houdeau, Niranjana Papagudi Subrahmanyam, Malek Safieh, Anne Passarelli, Aliza Maftun, Michaela Brunner, Tim Music, Michael Pehl, Tauseef Siddiqui, Ralf Brederlow, Ulf Schlichtmann, Bjoern Driemeyer, Maurits Ortmanns, Robert Hesselbarth, Matthias Hiller |
DATE | 15 |
| 2023 | Extended Abstract: Monitoring-based Thermal Management for Mixed-Criticality SystemsabstractWith a rapidly growing number of functions in embedded real-time systems, it becomes inevitable to integrate tasks of different safety integrity levels (SILs) into one mixed-criticality system. Here, it is important to not only isolate shared architectural resources, as tasks executing on different cores may also interfere via the processor's thermal manager. In order to prevent a scenario where best-effort tasks cause deadline violations for critical tasks, we propose a thermal management strategy that guarantees a sufficient thermal isolation between tasks of different SILs, and simultaneously reduces the run-time of best-effort tasks by up to 45 % compared to the state of the art without incurring any real-time violations for critical tasks. Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann |
DATE | 6 |
| 2023 | Efficient Software-Implemented HW Fault Tolerance for TinyML Inference in Safety-critical ApplicationsabstractTinyML research has mainly focused on optimizing neural network inference in terms of latency, code-size and energy-use for efficient execution on low-power micro-controller units (MCUs). However, distinctive design challenges emerge in safety-critical applications, for example in small unmanned autonomous vehicles such as drones, due to the susceptibility of off-the-shelf MCU devices to soft-errors. We propose three new techniques to protect TinyML inference against random soft errors with the target to reduce run-time overhead: one for protecting fully-connected layers; one adaptation of existing algorithmic fault tolerance techniques to depth-wise convolutions; and an efficient technique to protect the so-called epilogues within TinyML layers. Integrating these layer-wise methods, we derive a full-inference hardening solution for TinyML that achieves run-time efficient soft-error resilience. We evaluate our proposed solution on MLPerf-Tiny benchmarks. Our experimental results show that competitive resilience can be achieved compared with currently available methods, while reducing run-time overheads by ~120% for one fully-connected neural network (NN); ~20% for the two CNNs with depth-wise convolutions; and ~2% for standard CNN. Additionally, we propose selective hardening which reduces the incurred run-time overhead further by ~2x for the studied CNNs by focusing exclusively on avoiding mispredictions. Uzair Sharif, Daniel Mueller-Gritschneder, Rafael Stahl, Ulf Schlichtmann |
DATE | 4 |
| 2023 | Class-based Quantization for Neural NetworksabstractIn deep neural networks (DNNs), there are a huge number of weights and multiply-and-accumulate (MAC) operations. Accordingly, it is challenging to apply DNNs on resource- constrained platforms, e.g., mobile phones. Quantization is a method to reduce the size and the computational complexity of DNNs. Existing quantization methods either require hardware overhead to achieve a non-uniform quantization or focus on model-wise and layer-wise uniform quantization, which are not as fine-grained as filter-wise quantization. In this paper, we propose a class-based quantization method to determine the minimum number of quantization bits for each filter or neuron in DNNs individually. In the proposed method, the importance score of each filter or neuron with respect to the number of classes in the dataset is first evaluated. The larger the score is, the more important the filter or neuron is and thus the larger the number of quantization bits should be. Afterwards, a search algorithm is adopted to exploit the different importance of filters and neurons to determine the number of quantization bits of each filter or neuron. Experimental results demonstrate that the proposed method can maintain the inference accuracy with low bit-width quantization. Given the same number of quantization bits, the proposed method can also achieve a better inference accuracy than the existing methods. Grace Li Zhang, Huaxi Gu, Bing Li 0005, Ulf Schlichtmann |
DATE | 5 |
| 2023 | SteppingNet: A Stepping Neural Network with Incremental Accuracy EnhancementabstractDeep neural networks (DNNs) have successfully been applied in many fields in the past decades. However, the in-creasing number of multiply-and-accumulate (MAC) operations in DNNs prevents their application in resource-constrained and resource-varying platforms, e.g., mobile phones and autonomous vehicles. In such platforms, neural networks need to provide ac-ceptable results quickly and the accuracy of the results should be able to be enhanced dynamically according to the computational resources available in the computing system. To address these challenges, we propose a design framework called SteppingNet. SteppingNet constructs a series of sub nets whose accuracy is incrementally enhanced as more MAC operations become avail-able. Therefore, this design allows a trade-off between accuracy and latency. In addition, the larger sub nets in SteppingNet are built upon smaller subnets, so that the results of the latter can directly be reused in the former without recomputation. This property allows SteppingNet to decide on-the-fly whether to enhance the inference accuracy by executing further MAC operations. Experimental results demonstrate that SteppingNet provides an effective incremental accuracy improvement and its inference accuracy consistently outperforms the state-of-the-art work under the same limit of computational resources. Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Huaxi Gu, Bing Li 0005, Ulf Schlichtmann |
DATE | 7 |
| 2023 | XRing: A Crosstalk-Aware Synthesis Method for Wavelength-Routed Optical Ring RoutersabstractWavelength-routed optical networks-on-chip (WR-ONoCs) are well-known for supporting high-bandwidth communications with low power and latency. Among all WRONoC routers, optical ring routers have attracted great research interest thanks to their simple structure, which looks like concentric cycles formed by waveguides. Current ring routers are designed manually. When the number of network nodes increases or the position of network nodes changes, it can be difficult to manually determine the optimal design options. Besides, current ring routers face two problems. First, some signal paths in the routers can be very long and suffer high insertion loss; second, to connect the network nodes to off-chip lasers, waveguides in the power distribution network (PDN) have to intersect with the ring waveguides, which causes additional insertion loss and crosstalk noise. In this work, we propose XRing, which is the first design automation method to automatically synthesize optical ring routers based on the number and position of network nodes. In particular, XRing optimizes the waveguide connections between the network nodes with a mathematical modelling method. To reduce insertion loss and crosstalk noise, XRing constructs efficient shortcuts between the network nodes that suffer long signal paths and creates openings on ring waveguides so that the PDN can easily access the network nodes without causing waveguide crossings. The experimental results show that XRing outperforms other WRONoC routers in reducing insertion loss and crosstalk noise. In particular, more than 98% of signals in XRing do not suffer first-order crosstalk noise, which significantly enhances the signal quality. Zhidan Zheng, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
DATE | 4 |
| 2023 | Memory Latency Distribution-Driven Regulation for Temporal Isolation in MPSoCs
Ahsan Saeed, Denis Hoornaert, Dakshina Dasari, Dirk Ziegenbein, Daniel Mueller-Gritschneder, Ulf Schlichtmann, Andreas Gerstlauer, Renato Mancuso 0001 |
ECRTS | 6 |
| 2023 | Semi-Supervised Deep Learning for Microcontroller Performance ScreeningabstractIn safety-critical applications, microcontrollers must satisfy strict quality constraints and performances in terms of Fmax(the maximum operating frequency). Data extracted from on-chip ring oscillators (ROs) can model the Fmaxof integrated circuits using machine learning models. Those models are suitable for the performance screening process. Acquiring data from the ROs is a fast process that leads to many unlabeled data. Contrarily, the labeling phase (i.e., acquiring Fmax) is a time-consuming and costly task, that leads to a small set of labeled data. This paper presents deep-learning-based methodologies to cope with the low number of labeled data in microcontroller performance screening. We propose a method that takes advantage of the high number of unlabeled samples in a semi-supervised learning fashion. We derive deep feature extractor models that project data into higher dimensional spaces and use the data feature embedding to face the performance prediction problem with simple linear regression. Experiments showed that the proposed models outperformed state-of-the-art methodologies in terms of prediction error and permitted us to use a significantly smaller number of devices to be characterized, thus reducing the time needed to build ML models by a factor of six with respect to baseline approaches. Nicolò Bellarmino, Riccardo Cantoro, Martin Huch, Tobias Kilian, Ulf Schlichtmann, Giovanni Squillero |
ETS | 5 |
| 2023 | GAT-based Concentration Prediction for Random Microfluidic Mixers with Multiple Input Flow RatesabstractMicrofluidic biochips have emerged with significant promise and versatility in automating a variety of biochemical protocols. Accurate preparation of fluid samples with microfluidic mixers is an essential component of these protocols, where concentration prediction and generation are critical. Recently, machine learning models have been adopted in concentration prediction, which demonstrate great potential in enhancing the efficiency and scalability over the traditional finite element analysis (FEA) methods. However, the state-of-the-art machine learning-based method can only predict the concentration of microfluidic mixers with fixed input flow rates, but suffers poor prediction accuracy for multiple input flow rates. To address this issue, this paper proposes a new concentration prediction method based on the graph attention networks (GAT). By modeling each channel of the mixer as a graph node in a GAT, the proposed method efficiently and accurately predicts the generated concentration of random microfluidic mixers with multiple input flow rates. Experimental results show that compared with the state-of-the-art method, the proposed GAT-based simulation method obtains a reduction of 85% in terms of errors of predicted concentration, which validates the effectiveness of the proposed GAT model. Weiqing Ji, Hailong Yao 0002, Tsung-Yi Ho, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | SOAER: Self-Obstacle Avoiding Escape Routing for Paper-Based Digital Microfluidic BiochipsabstractIn paper-based digital microfluidic biochips (P-DMFBs), conductive electrodes and control lines are printed on the same side of the photo paper, which introduces a critical design challenge on the so-called control interference issue. This introduces a distinct escape routing problem, named Self-Obstacle Avoiding Escape Routing (SOAER). In the SOAER problem, each electrode has a specific set of routing obstacles of its own, which are forbidden to be crossed over by the control line of the electrode. Based on an enhanced network flow model, this paper proposes an effective SOAER routing method for P-DMFBs. Experimental results show that compared with the state-of-the-art method, SOAER obtains 49x speedup in runtime. Our proposed method also shows the efficiency and effectiveness of the overall system. The success rate is up to 100% and the runtime is decreased significantly. Weiqing Ji, Xingcheng Yao, Hailong Yao 0002, Tsung-Yi Ho, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | A Novel and Efficient Block-Based Programming for ReRAM-Based Neuromorphic ComputingabstractReRAM-based accelerators have emerged as promising accelerators for deep neural networks (DNNs). How-ever, programming every ReRAM cell to its corresponding conductance before inference can be time-consuming and energy-intensive using existing one-by-one/row-by-row programming mechanisms. Although a two-phase multi-row programming scheme has been proposed to enhance programming efficiency, there are situations where multiple rows cannot be programmed together and only row-by-row programming can be employed. Therefore, this paper proposes a new block-based programming architecture for ReRAM crossbars that enables precise control of wordline and bitline transistors. In addition, a block-based programming framework, including the approximation phase and the fine-tuning phase, along with a multi-line programming algorithm and a programming-aware model retraining are proposed to reduce programming cycles and energy consumption. Experimental results demonstrate that our proposed method can reduce programming cycles and energy consumption by 46%-49 % and 63 % -64 %, respectively, compared to the state of the art. Additionally, the area and power overhead are negligible. Wei-Lun Chen, Fang-Yi Gu, Ing-Chao Lin, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann |
ICCAD | 6 |
| 2023 | NeuroEscape: Ordered Escape Routing via Monte-Carlo Tree Search and Neural NetworkabstractOrdered escape routing is a critical stage for printed circuit board design. State-of-the-art solutions to ordered escape routing are either heuristic and non-optimal, or trapped in exponential time complexity. In this work, for the first time, we prove that ordered escape routing is not only NP-hard, but also hard to approximate in polynomial time, indicating the limitation of optimal algorithms. We further present NeuroEscape, an efficient ordered escape routing method, which is based on reinforcement learning with a Monte-Carlo tree search (MCTS) and heuristic rollouts for design space exploration. A neural policy model is incorporated to further enhance the MCTS process. Theoretical results show that the number of samples required to train the model is upper bounded by a polynomial. Experimental results show that NeuroEscape solves 69% more testcases, with an average acceleration of 19.4x compared with state-of-the-art methods. Zhiyang Chen 0006, Tsung-Yi Ho, Ulf Schlichtmann, Datao Chen, Hailong Yao 0002 |
ICCAD | 3 |
| 2023 | ARMM: Adaptive Reliability Quantification Model of Microfluidic Designs and its Graph-Transformer-Based ImplementationabstractAfter decades of development, flow-based microfluidic biochips have become a revolutionary platform for biochemical experiments. To meet the increasingly complex experimental demands, the length and density of channels in these chips grow significantly, which brings about higher defect probabilities. Till now, several methods have been proposed to improve the yield of these increasingly complex chips. However, the effectiveness of these methods cannot be properly evaluated, since there has been no method that systematically analyzes the reliability of a microfluidic design. In this paper, we propose the first mathematical models to quantify the reliability of a microfluidic design by calculating the probability of blockage and leakage defects happening to the design. Besides, we propose a graph-transformer-based method to speed up the calculation, so that designers can have a fast and accurate evaluation of the reliability of a microfluidic design at any scale. Siyuan Liang 0002, Meng Lian 0001, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann, Tsung-Yi Ho |
ICCAD | 5 |
| 2023 | FXT-Route: Efficient High-Performance PCB Routing with Crosstalk Reduction Using Spiral Delay LinesabstractIn high-performance printed circuit boards (PCBs), adding serpentine delay lines is the most prevalent delay-matching technique to balance the delays of time-critical signals. Serpentine topology, however, can induce simultaneous accumulation of the crosstalk noise, resulting in erroneous logic gate triggering and speed-up effects. The state-of-the-art approach for crosstalk alleviation achieves waveform integrity by enlarging wire separation, resulting in an increased routing area. We introduce a method that adopts spiral delay lines for delay matching to mitigate the speed-up effect by spreading the crosstalk noise uniformly in time. Our method avoids possible routing congestion while achieving a high density of transmission lines. We implement our method by constructing a mixed-integer-linear programming (MILP) model for routing and a quadratic programming (QP) model for spiral synthesis. Experimental results demonstrate that our method requires, on average, 31% less routing area than the original design. In particular, compared to the state-of-the-art approach, our method can reduce the magnitude of the crosstalk noise by at least 69%. Meng Lian 0001, Yushen Zhang, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
ISPD | 5 |
| 2023 | A Multilabel Active Learning Framework for Microcontroller Performance ScreeningabstractIn safety-critical applications, microcontrollers have to be tested to satisfy strict quality and performance constraints. It has been demonstrated that on-chip ring oscillators can be used as speed monitors to reliably predict the performances. However, any machine-learning (ML) model is likely to be inaccurate if trained on an inadequate dataset, and labeling data for training is quite a costly process. In this article, we present a methodology based on active learning to select the best samples to be included in the training set, significantly reducing the time and cost required. Moreover, since different speed measurements are available, we designed a multilabel technique to take advantage of their correlations. Experimental results demonstrate that the approach halves the training-set size, with respect to a random-labeling, while it increases the predictive accuracy, with respect to standard single-label ML models. Nicolò Bellarmino, Riccardo Cantoro, Martin Huch, Tobias Kilian, Raffaele Martone, Ulf Schlichtmann, Giovanni Squillero |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | BRoCoM: A Bayesian Framework for Robust Computing on Memristor CrossbarabstractMemristor crossbar arrays are considered to be a promising platform for neuromorphic computing. To deploy a trained neural network (NN) model on memristor crossbars, memristors need to be programmed to the corresponding weight values. In fact, due to device-based process variation and noise, deviations of the stored weights from the trained weights are inevitable, thereby causing the degradation of the actual inference performance. This article proposes a unified Bayesian inference-based framework, BRoCoM, which connects device nonidealities and algorithmic training together for robust computing on memristor crossbars. BRoCoM is able to incorporate different levels of nonidealities into prior weight distribution, and transform robustness optimization to Bayesian NN (BNN) training, the weights of NNs are optimized to accommodate uncertainties and minimize inference degradation. Experimental results confirm the capability of the proposed BRoCoM to achieve stable inference performance while tolerating the nonideal effects of process variation and noise. Qingrong Huang, Grace Li Zhang, Xunzhao Yin, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | Design Automation for Continuous-Flow Lab-on-a-Chip Systems: A One-Pass ParadigmabstractOwing to the high complexity of chip architecture and assay protocol, considerable effort has been directed toward the design automation of continuous-flow microfluidics over the past decade. Existing methods, however, perform the corresponding design tasks, including binding, scheduling, placement, and routing separately, leading to serious gaps between different steps and potentially even cause design failure. To overcome these drawbacks, in this article, we propose a one-pass design paradigm for continuous-flow microfluidic lab-on-a-chip systems, integrating all the design steps into an “organic whole,” which has never been considered in prior work. With the proposed paradigm, all the design tasks can be synchronized seamlessly and performed in a combined manner, thereby eliminating the gaps between design steps. Consequently, optimized biochip architectures can be generated without any design adjustments and modifications. The experimental results demonstrate the effectiveness of the proposed automation flows. Xing Huang 0001, Youlin Pan, Wenzhong Guo, Lu Wang 0014, Qingshan Li, Robert Wille, Tsung-Yi Ho, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2023 | Integrated Test Module Design for Microfluidic Large-Scale IntegrationabstractMicrofluidic large-scale integration (mLSI) is a promising lab-on-a-chip platform for high-throughput bio-applications. Due to the high integration scale and the small feature size, control channels on mLSI chips are prone to blockage and leakage defects, which may lead to faulty behavior of valves and erroneous experimental results. Thus, mLSI chips need to be tested before usage. Current mLSI-tests are mostly performed in a straightforward way by testing each valve individually, which is very time consuming and error prone. As the integration scale of mLSI chips keeps increasing, there is a pressing demand for more efficient test approaches. This work proposes the first built-in-self-test (BIST) method for mLSI with an integrated test module design. Instead of testing individual valves, the proposed method directly tests the control channels and thus greatly improves the test efficiency. Only${}({n}/{2})$and$\lceil \log _{2}(n+1)\rceil $test operations are required to test the blockage and leakage defects, respectively, of$n$control channels. The proposed test module consumes moderate area overhead and the test method is easy to operate. Neither specialized software nor external pressure sensors are required for carrying out the tests. Experiments show that our test approach is sensitive enough to detect defects that have a feature size as small as 10$\mu \text{m}$and that are several centimeters away from the test module. Mengchu Li, Yushen Zhang, Ju Young Lee, Hudson Gasvoda, Ismail Emre Araci, Tsun-Ming Tseng, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | Ferroelectric Ternary Content Addressable Memories for Energy-Efficient Associative SearchabstractA fast and efficient search function across the database has been a core component for a number of data-intensive tasks in machine learning, IoT applications, and inference. However, the conventional digital machines implementing the search functionality with repetitive arithmetic operations suffer from the energy efficiency and performance degradation due to the significant data transfer between the storage and processing units in the Von Neumann architecture. Ternary content addressable memories (TCAMs) are an essential hardware form of computing-in-memory (CiM) designs that aim to overcome the data transfer bottlenecks by implementing the parallel associative search function within the memory blocks. While most state-of-the-art TCAM designs focus on improving the information density by harnessing compact nonvolatile memories (NVMs), little efforts have been spent on optimizing the energy efficiency of the NVM-based TCAM. In this article, by exploiting the ferroelectric FET (FeFET) as a representative NVM, we propose an NOR-type 2FeFET-1T and an NAND-type 2FeFET-2T TCAM designs that enable highly energy-efficient associative search by reducing the associated precharge overheads. We then propose a hybrid ferroelectric NAND-NOR (HFNN) TCAM design to further improve the energy efficiency. An HFNN-based segmented architecture is proposed to reduce the search delay and energy by search operation pipeline. Evaluation results suggest that the proposed 2FeFET-1T, 2FeFET-2T and HFNN TCAM design consume$3.03\times $,$8.08\times $, and$226.92\times $less search energy than the conventional 16T complementary metal oxide semiconductor (CMOS) TCAM, respectively. Application benchmarking shows that our proposed 2FeFET-1T/2FeFET-2T/HFNN TCAM can save, on average, 45.2%/50.6%/57.5% the GPU energy consumption as compared to the conventional GPU. Xunzhao Yin, Yu Qian 0002, Mohsen Imani, Kai Ni 0004, Chao Li 0065, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | Training PPA Models for Embedded Memories on a Low-data DietabstractSupervised machine learning requires large amounts of labeled data for training. In power, performance, and area (PPA) estimation of embedded memories, every new memory compiler version is considered independently of previous compiler versions. Since the data of different memory compilers originate from similar domains, transfer learning may reduce the amount of supervised data required by pre-training PPA estimation neural networks on related domains. We show that provisioning times of PPA models for new compiler versions can be reduced significantly by exploiting similarities among different compilers, versions, and technology nodes. Through transfer learning, we shorten the time to provision PPA models for new compiler versions, which speeds up time-critical periods of the design cycle. Using only 901 training samples (10%) is sufficient to achieve an almost worst-case (98th percentile) estimation error of 2.67% and allows us to shorten model provisioning times from 40 days to less than one week without sacrificing accuracy. To enable a diverse set of source domains for transfer learning, we devise a new, application-independent method for overcoming structural domain differences through domain equalization that attains competitive results when compared to domain-free transfer. A high degree of automation necessitates the efficient assessment of the best source domains. We propose using various metrics to accurately identify four of the five best among 45 datasets with low computational effort. Felix Last, Ulf Schlichtmann |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | Performance Screening Using Functional Path Ring OscillatorsabstractThe testing of integrated circuits is an important topic, particularly in safety-critical applications. This is especially true for microcontrollers (MCUs) used in the automotive industry. A critical test is the performance screening in which the maximum clock frequency of the MCU is determined. For this performance screening, indirect monitors, such as ring oscillators (ROs), are used. This article presents a holistic overview of the functional path RO from the pre-silicon to the post-silicon. The implementation of such ROs is presented, as the associated advantages in terms of area consumption, leakage, and routing. In the post-silicon phase, the functional path RO frequencies are correlated with the MCU performance using machine learning approaches. Tobias Kilian, Daniel Tille, Martin Huch, Markus Hanel, Ulf Schlichtmann |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | Differentially Evolving Memory Ensembles: Pareto Optimization based on Computational Intelligence for Embedded Memories on a System LevelabstractAs the relative power, performance, and area (PPA) impact of embedded memories continues to grow, proper parameterization of each of the thousands of memories on a chip is essential. When the parameters of all memories of a product are optimized together as part of a single system, better trade-offs may be achieved than if the same memories were optimized in isolation. However, challenges such as a sparse solution space, conflicting objectives, and computationally expensive PPA estimation impede the application of common optimization heuristics. We show how the memory system optimization problem can be solved through computational intelligence. We apply a Pareto-based Differential Evolution to ensure unbiased optimization of multiple PPA objectives. To ensure efficient exploration of a sparse solution space, we repair individuals to yield feasible parameterizations. PPA is estimated efficiently in large batches by pre-trained regression neural networks. Our framework enables the system optimization of thousands of memories while keeping a small resource footprint. Evaluating our method on a tractable system, we find that our method finds diverse solutions which exhibit less than 0.5% distance from known global optima. Felix Last, Ceren Yeni, Ulf Schlichtmann |
ASP-DAC | 3 |
| 2022 | Energy efficient data search design and optimization based on a compact ferroelectric FET content addressable memoryabstractContent Addressable Memory (CAM) is widely used for associative search tasks in advanced machine learning models and data-intensive applications due to the highly parallel pattern matching capability. Most state-of-the-art CAM designs focus on reducing the CAM cell area by exploiting the nonvolatile memories (NVMs). There exists only little research on optimizing the design and energy efficiency of NVM based CAMs for practical deployment in edge devices and AI hardware. In this paper, we propose a general compact and energy efficient CAM design scheme that alleviates the design overhead by employing just one NVM device in the cell. We also propose an adaptive matchline (ML) precharge and discharge scheme that further optimizes the search energy by fully reducing the ML voltage swing. We consider Ferroelectric field effect transistors (FeFETs) as the representative NVM, and present a 2T-1FeFET CAM array including a sense amplifier implementing the proposed ML scheme. Evaluation results suggest that our proposed 2T-1FeFET CAM design achieves 6.64×/4.74×/9.14×/3.02× better energy efficiency compared with CMOS/ReRAM/STT-MRAM/2FeFET CAM arrays. Benchmarking results show that our approach provides 3.3×/2.1× energy-delay product improvement over the 2T-2R/2FeFET CAM in accelerating query processing applications. Jiahao Cai, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
DAC | 6 |
| 2022 | GNN-based concentration prediction for random microfluidic mixersabstractRecent years have witnessed significant advances brought by microfluidic biochips in automating biochemical processing. Accurate preparation of fluid samples with microfluidic mixers is a fundamental step in various biomedical applications, where concentration prediction and generation are critical. Finite element analysis (FEA) is the most commonly used simulation method for accurate concentration prediction of a given biochip design, such as COMSOL. However, the FEA simulation process is time-consuming with poor scalability for large biochip sizes. This paper proposes a new concentration prediction method based on the graph neural networks (GNN), which efficiently and accurately predicts the generated concentration by random microfluidic mixers of different sizes. Experimental results show that compared with the state-of-the-art method, the proposed GNN-based simulation method obtains a reduction of 88% in terms of errors of predicted concentration, which validates the effectiveness of the proposed GNN model. Weiqing Ji, Xingzhuo Guo, Shouan Pan, Tsung-Yi Ho, Ulf Schlichtmann, Hailong Yao 0002 |
DAC | 5 |
| 2022 | Contamination-Free Switch Design and Synthesis for Microfluidic Large-Scale IntegrationabstractMicrofluidic large-scale integration (mLSI) biochips have developed rapidly in recent decades. The gap between design efficiency and application complexity has led to a growing interest in mLSI design automation. The state-of-the-art design automation tools for mLSI focus on the simultaneous co-optimisation of the flow and control layers but neglect potential contamination between different fluid reagents and products. Microfluidic switches, as fluid routers at the intersection of flow paths, are especially prone to contamination. State-of-the-art tools design the switches as spines with junctions, which aggregate the contamination problem. In this work, we present a contamination-free microfluidic switch design and a synthesis method to generate application-specific switches that can be employed by physical design tools for mLSI. We also propose a scheduling and binding method to transport the fluids with least time and fewest resources. To reduce the number of pressure inlets, we consider pressure sharing between valves within the switch. Experimental results demonstrate that our methods show advantages in avoiding contamination and improving transportation efficiency over conventional methods. Duan Shen, Yushen Zhang, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
DATE | 5 |
| 2022 | RRAM-based Neuromorphic Computing: Data Representation, Architecture, Logic, and ProgrammingabstractRRAM crossbars provide a promising hardware plat-form to accelerate matrix-vector multiplication in deep neural networks (DNNs). To exploit the efficiency of RRAM crossbars, extensive research ex-amining architecture, data representation, logic de-sign as well as device programming should be conducted. This extensive scope of research aspects is enabled and required by the versatility of RRAM cells and their organization in a computing system. These research aspects affect or benefit each other. Therefore, they should be considered systematically to achieve an efficient design in terms of design complexity and computational performance in accelerating DNNs. In this paper, we illustrate study exam-ples on these perspectives on RRAM crossbars, in-cluding data representation with pulse widths, archi-tecture improvement, implementation of logic functions using RRAM cells, and efficient programming of RRAM devices for accelerating DNNs. Grace Li Zhang, Shuhang Zhang, Hai Li 0001, Ulf Schlichtmann |
DSD | 4 |
| 2022 | Test, Reliability and Functional Safety Trends for Automotive System-on-ChipabstractThis paper encompasses three contributions by industry professionals and university researchers. The contributions describe different trends in automotive products, including both manufacturing test and run-time reliability strategies. The subjects considered in this session deal with critical factors, from optimizing the final test before shipment to market to in-field reliability during operative life. Francesco Angione, Davide Appello, Joseph Aribido, Jyotika Athavale, Nicolò Bellarmino, Paolo Bernardi 0002, Riccardo Cantoro, Corrado De Sio, Tommaso Foscale, Gabriele Gavarini, Juan-David Guerrero-Balaguera, Martin Huch, Giusy Iaria, Tobias Kilian, Riccardo Mariani, Raffaele Martone, Annachiara Ruospo, Ernesto Sánchez 0001, Ulf Schlichtmann, Giovanni Squillero, Matteo Sonza Reorda, Luca Sterpone, Vincenzo Tancorre, Roberto Ugioli |
ETS | 19 |
| 2022 | Reducing Routing Overhead by Self-Enabling Functional Path Ring OscillatorsabstractAutomotive Microcontrollers (MCUs) are extensively tested to guarantee zero-defect quality. Performance screening is one of the critical factors to ensure that MCUs meet quality requirements. Ring Oscillator (RO) structures are used for this performance screening. Such RO structures usually cause routing overhead on the chip. The routing overhead increases, especially when many ROs are implemented. This paper presents a novel self-enabling technique that significantly reduces the routing overhead for functional path ROs. We present a proof of concept on a large automotive MCU. The routing overhead can be reduced by over 80% compared to traditional approaches. Tobias Kilian, Markus Hanel, Daniel Tille, Martin Huch, Ulf Schlichtmann |
ETS | 5 |
| 2022 | CorePerfDSL: A Flexible Processor Description Language for Software Performance SimulationabstractInstruction set simulators (ISSs) model the functional behavior of embedded processors for early software development. While they offer high simulation speeds, an ISS usually does not model the timing behavior of the processor accurately. Existing software performance simulators are typically either specific to a certain microarchitecture or their description language mixes functional and microarchitectural aspects.In this paper, we introduce CorePerfDSL, an architecture description language (ADL) specifically designed to model the timing behavior of processor microarchitectures for software performance estimation. CorePerfDSL is clearly separated from any functional description of the modelled processor by a generic trace definition. As such, it is well suited to generate performance simulators that can be paired with an existing ISS, which supplies an execution trace. In addition, CorePerfDSL provides a high degree of flexibility, supporting the fast generation of models for various microarchitecture variants, which can be used for rapid architectural exploration. We demonstrate the flexibility of CorePerfDSL by describing several variants of a single-issue five-stage RISC-V microarchitecture and estimate their performances for a software benchmark program. Conrad Foik, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
FDL | 3 |
| 2022 | CoMUX: Combinatorial-Coding-Based High-Performance Microfluidic Control Multiplexer DesignabstractFlow-based microfluidic chips are one of the most promising platforms for biochemical experiments. Transportation channels and operation devices inside these chips are controlled by microvalves, which are driven by external pressure sources. As the complexity of experiments on these chips keeps increasing, control multiplexers (MUXes) become necessary for the actuation of the enormous number of valves. However, current binary-coding-based MUXes do not take full advantage of the coding capacity and suffer from the reliability problem caused by the high control channel density. In this work, we propose a novel MUX coding strategy, named Combinatorial Coding, along with an algorithm to synthesize combinatorial-coding-based MUXes (CoMUXes) of arbitrary sizes with the proven maximum coding capacity. Moreover, we develop a simplification method to reduce the number of valves and control channels in CoMUXes and thus improve their reliability. We compare CoMUX with the state-of-the-art MUXes under different control demands with up to 10 × 213 independent control channels. Experiments show that CoMUXes can reliably control more independent control channels with fewer resources. For example, when the number of the to-be-controlled control channels is up to 10 × 213, compared to a state-of-the-art MUX, the optimized CoMUX reduces the number of required flow channels by 44% and the number of valves by 90%. Siyuan Liang 0002, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann, Tsung-Yi Ho |
ICCAD | 4 |
| 2022 | Microcontroller Performance Screening: Optimizing the Characterization in the Presence of Anomalous and Noisy DataabstractIn safety-critical applications, microcontrollers must satisfy strict quality constraints and performances in terms of $F_{\max}$, that is, the maximum operating frequency. It has been demonstrated that data extracted from on-chip speed monitors can model the $F_{\max}$ of integrated circuits by means of machine learning models, and that those models are suitable for the performance screening process. However, while acquiring data from these monitors is quite an accurate process, the labelling is time-consuming, costly, and may be subject to different measurements errors, impairing the final quality. This paper presents a methodology to cope with anomalous and noisy data in the context of the multi-label regression problem of microcontroller performance screening. We used outlier detection based on Inter Quartile Range (IQR) and Z-score and imputation techniques to detect errors in the labels and to avoid to drop incomplete samples, building higher-quality training set for our models, optimizing the devices characterization phase. Experiments showed that the proposed methodology increases the performance of existing models, making them more robust. These techniques permitted us to use a significantly smaller number of samples (about one third of the devices available for characterization), thus making the costly data acquisition process more efficient. Nicolò Bellarmino, Riccardo Cantoro, Martin Huch, Tobias Kilian, Ulf Schlichtmann, Giovanni Squillero |
IOLTS | 5 |
| 2022 | Aging Aware Retraining for Memristor-based Neuromorphic ComputingabstractMemristor-based crossbars, which can achieve 1-2 orders of magnitude energy efficiency improvement over digital machines, have been introduced to accelerate the neural networks of machine learning tasks. Due to the high voltage pulses repeatedly applied onto memristors during programming and online tuning, the effective resistance ranges of the memristors actually decrease as a result of aging, which eventually impair the inference accuracy of the neural network running on the memristor-based crossbar. In this paper, we propose an algorithm-hardware co-design framework combining aging aware retraining and gradient sparsification to mitigate the impact of aging and extend the lifetime of the crossbar. Experimental results show that the proposed method can effectively increase the inference accuracy by up to 16% even with severe aging, while the crossbar lifetime can be extended by up to $2.7\times$. Wenwen Ye, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
ISCAS | 4 |
| 2022 | A Path Selection Flow for Functional Path Ring Oscillators using Physical Design DataabstractA lot of effort and money is invested in testing to ensure zero-defect quality of automotive microcontrollers. One crucial test is the performance screening. Indirect structures such as Ring Oscillators (ROs) are used for this. Here, the quality of the performance screening strongly depends on the quality and selection of the RO structures used. This paper proposes a path selection and implementation method to provide a set of functional path ROs with good representativeness for the whole chip. In addition, a simulation-based validation is presented, which is used to improve the selection process continually. The proposed path selection is validated by simulation and on silicon. The results show a high diversity and good coverage of the chip parameters with the selected functional path ROs, providing good conditions for a high-quality performance screening. Tobias Kilian, Markus Hanel, Daniel Tille, Martin Huch, Ulf Schlichtmann |
ITC | 5 |
| 2022 | Memory Utilization-Based Dynamic Bandwidth Regulation for Temporal Isolation in Multi-CoresabstractTemporal isolation is one of the key challenges for co-running mixed-criticality applications on Commercial Off-The-Shelf (COTS) multi-core platforms. In particular, the main memory subsystem is one of the most prominent causes of interference and loss of isolation. Existing mechanisms for memory bandwidth regulation are limited to conservative bandwidth reservation, use pessimistic worst-case execution time (WCET) estimations or require dedicated hardware that is not feasible in COTS multi-core platforms.In this paper, we propose a novel mechanism for memory interference control that uses feedback-based control to dynamically regulate memory accesses of individual cores in a multicore platform. Our mechanism directly regulates the source of interference by leveraging information about memory utilization, acquired from existing hardware performance counters provided by modern COTS-based memory controllers. The proposed solution is implemented on Linux as a loadable kernel module. The results of evaluating our approach with real and synthetic benchmarks on a COTS multi-core (NXP S32V234) platform demonstrate that it is able to provide temporal isolation with up to 4x and 2x more overall throughput for non-real-time applications compared to static and dynamic memory bandwidth-based regulation approaches, respectively, while maintaining guarantees for applications running on the real-time core. Ahsan Saeed, Dakshina Dasari, Dirk Ziegenbein, Varun Rajasekaran, Falk Rehm, Michael Pressler, Arne Hamann 0001, Daniel Mueller-Gritschneder, Andreas Gerstlauer, Ulf Schlichtmann |
RTAS | 10 |
| 2022 | An FPGA-based Approach to Evaluate Thermal and Resource Management Strategies of Many-core ProcessorsabstractThe continuous technology scaling of integrated circuits results in increasingly higher power densities and operating temperatures. Hence, modern many-core processors require sophisticated thermal and resource management strategies to mitigate these undesirable side effects. A simulation-based evaluation of these strategies is limited by the accuracy of the underlying processor model and the simulation speed. Therefore, we present, for the first time, an field-programmable gate array (FPGA)-based evaluation approach to test and compare thermal and resource management strategies using the combination of benchmark generation, FPGA-based application-specific integrated circuit (ASIC) emulation, and run-time monitoring. The proposed benchmark generation method enables an evaluation of run-time management strategies for applications with various run-time characteristics. Furthermore, the ASIC emulation platform features a novel distributed temperature emulator design, whose overhead scales linearly with the number of integrated cores, and a novel dynamic voltage frequency scaling emulator design, which precisely models the timing and energy overhead of voltage and frequency transitions. In our evaluations, we demonstrate the proposed approach for a tiled many-core processor with 80 cores on four Virtex-7 FPGAs. Additionally, we present the suitability of the platform to evaluate state-of-the-art run-time management techniques with a case study. Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann |
ACM Trans. Archit. Code Optim. | 6 |
| 2022 | Flow-Based Microfluidic Biochips With Distributed Channel Storage: Synthesis, Physical Design, and Wash OptimizationabstractSystem-architecture design optimization of flow-based microfluidic biochips has been extensively investigated over the past decade. Most of the prior work, however, is still based on chip architectures with dedicated storage units and this, not only limits the performance of biochips, but also increases their fabrication cost. To overcome this limitation, a distributed channel-storage architecture can be implemented, where fluid samples can be cached temporarily in flow channels instead of using a dedicated storage. This new concept of fluid storage, however, requires a careful arrangement of fluid samples to enable the channels to fulfill the dual functions of transportation and caching. Moreover, to avoid cross-contamination between different fluidic flows, wash operations are necessary to remove the residue left in flow channels. In this article, we formulate the first practical system level design and wash optimization problem for microfluidic biochips with distributed channel storage architecture, considering high-level synthesis, physical design, and wash optimization simultaneously, and present a top-down design flow to solve this problem systematically. Given the protocol of a biochemical application and the corresponding design requirements, our goal is to generate a chip architecture with low fabrication cost. Meanwhile the biochemical application can be executed efficiently with an optimized wash scheme. Experimental results on multiple benchmarks confirm that our approach leads to short completion time of biochemical applications, low chip cost, as well as high wash efficiency. Xing Huang 0001, Wenzhong Guo, Zhisheng Chen 0002, Bing Li 0005, Tsung-Yi Ho, Ulf Schlichtmann |
IEEE Trans. Computers | 6 |
| 2022 | MiniControl 2.0: Co-Synthesis of Flow and Control Layers for Microfluidic Biochips With Strictly Constrained Control PortsabstractRecent advances in continuous-flow microfluidics have enabled highly integrated lab-on-a-chip biochips. These chips can execute complex biochemical applications precisely and efficiently within a tiny area, but they require a large number of control ports and the corresponding control logic to generate required pressure patterns for flow control, which, consequently, offset their advantages and prevent their wide adoption. In this article, we propose the first flow-control layer co-synthesis flow called MiniControl, for continuous-flow microfluidic biochips under strict constraints for control ports, incorporating high-level synthesis, physical design, and control system design simultaneously, which has never been considered in previous work. With the maximum number of allowed control ports specified in advance, this synthesis flow aims to generate biochip architectures with high execution efficiency and the corresponding control systems with optimized timing performance. Besides, the overall cost of a biochip can be reduced and the tradeoff between a control system and execution efficiency of biochemical applications can be evaluated for the first time. The experimental results demonstrate that MiniControl leads to high execution efficiency, low platform cost, as well as excellent timing performance, while strictly satisfying the given control-port constraints. Xing Huang 0001, Tsung-Yi Ho, Genggeng Liu, Lu Wang 0014, Qingshan Li, Wenzhong Guo, Bing Li 0005, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2022 | PathDriver+: Enhanced Path-Driven Architecture Design for Flow-Based Microfluidic BiochipsabstractContinuous-flow microfluidic biochips have attracted high research interest over the past years. Inside such a chip, fluid samples of milliliter volumes are efficiently transported between devices (e.g., mixers, heaters, etc.) to automatically perform various laboratory procedures in biology and biochemistry. Each transportation task, however, requires an exclusive flow path composed of multiple contiguous microchannels during its execution period. Excess/waste fluids, in the meantime, should be discarded by independent flow paths connected to waste ports. All these paths are etched in a very tiny chip area using multilayer soft lithography and driven by flow ports connecting with external pressure sources, forming a highly integrated chip architecture that determines the final performance of biochips. In this article, we propose a new and practical design flow called PathDriver+ (PD+) for the architecture design of microfluidic biochips, integrating the actual fluid manipulations into both high-level synthesis and physical design, which has never been considered in prior work. With this design flow, highly efficient chip architectures with a flow-path network that enables the actual fluid transportation and removal can be constructed automatically. Meanwhile, fluid volume management between devices and flow-path minimization are realized for the first time, thus, ensuring the correctness of assay outcomes while reducing the complexity of chip architectures. Additionally, diagonal channel routing is implemented to fundamentally improve the chip performance. The tradeoff between the numbers of channel intersections and fluidic ports is evaluated to further reduce the fabrication cost of biochips. The experimental results on multiple benchmarks confirm that the proposed design flow leads to high assay execution efficiency and low overall chip cost. Xing Huang 0001, Youlin Pan, Grace Li Zhang, Bing Li 0005, Wenzhong Guo, Tsung-Yi Ho, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Crosstalk-Aware Automatic Topology Customization and Optimization for Wavelength-Routed Optical NoCsabstractOptical network-on-chip (ONoC) is an emerging upgrade for electronic network-on-chip (ENoC). As a kind of ONoC, wavelength-routed ONoC (WRONoC) shows ultrahigh bandwidth and ultralow latency in data communication. Manually designed WRONoC topologies typically reserve all to all links. This causes the waste of resources. Topology customization for each individual communication network can save resources, but requires automation for efficient design. The state-of-the-art design automation method is not efficient and does not support crosstalk analysis and signal-to-noise ratio (SNR) optimization. Moreover, the state of the art does not consider the physical locations of the data sending/receiving ports, causing unavoidable detours and crossings in the physical layout. In this work, we present FAST+: an automatic topology customization and optimization method. Compared with the state of the art, FAST+ operates much more efficiently and proposes a concrete router-level crosstalk-analysis method and a novel SNR optimization algorithm. This work also provides solutions to avoid detours and crossings in the physical layout. When SNR optimization is not enabled, experimental results show that FAST+ runs thousands times faster than the state of the art on average while providing multiple better or equally good topologies regarding resource usage and the worst case insertion loss. When SNR optimization is enabled, FAST+ provides$1.75\times $better worst case SNR on average after the optimization while not sacrificing resource usage and the worst case insertion loss. Xiao Moyuan, Tsun-Ming Tseng, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Contamination-Aware Synthesis for Programmable Microfluidic DevicesabstractProgrammable microfluidic devices (PMDs) have emerged as a new software-controlled architecture for next-generation flow-based biochips. These devices can be dynamically reconfigured to perform different bioassays flexibly and efficiently owing to their 2-D regularly arranged valve structure. However, PMDs are confronted with critical contamination issues due to the matrix-like structure with intersecting channels. In this article, a block-flushing method is proposed for contamination removal, based on which an overall contamination-aware synthesis flow is proposed. In the proposed block-flushing approach, contaminated areas are first collected according to specific patterns and then flushed as a whole to increase washing efficiency. Then, the synthesis flow integrating the block-flushing method is further optimized such that functional bioassay operations and washing operations can be performed simultaneously for higher efficiency. Experimental results demonstrate that the proposed washing approach reduces the washing time by 28% on commonly used bioassays. Equipped with the proposed washing method, our contamination-aware synthesis flow effectively reduces 30% of the completion time of the bioassays compared with the baseline method. Hui-Chieh Yu, Yu-Huei Lin, Zhiyang Chen 0006, Bing Li 0005, Xing Huang 0001, Ulf Schlichtmann, Tsung-Yi Ho, Hailong Yao 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | VirtualSync+: Timing Optimization With Virtual SynchronizationabstractIn digital circuit designs, sequential components such as flip-flops are used to synchronize signal propagations. Logic computations are aligned at and thus isolated by flip-flop stages. Although this fully synchronous style can reduce design efforts significantly, it may affect circuit performance negatively, because sequential components can only introduce delays into signal propagations but never accelerate them. In this article, we propose a new timing model, VirtualSync+, in which signals, specially those along critical paths, are allowed to propagate through several sequential stages without flip-flops. Timing constraints are still satisfied at the boundary of the optimized circuit to maintain a consistent interface with existing designs. By removing clock-to-q delays and setup time requirements of flip-flops on critical paths, the performance of a circuit can be pushed even beyond the limit of traditional sequential designs. In addition, we further enhance the optimization with VirtualSync+ by fine-tuning with commercial design tools, e.g., design compiler from Synopsys, to achieve more accurate result. To achieve this fine-tuning, we first optimize the circuits by reallocating sequential components with sequential and combinational components as delay units. Afterward, the removal locations of flip-flops with respect to the circuits under optimization are extracted and the corresponding wave-pipelining timing constraints compatible with commercial design tools are established. These timing constraints are then incorporated into the optimization flow of commercial tools to generate the optimized circuits. The experimental results demonstrate that circuit performance can be improved by up to 4% (average 1.5%) compared with that after extreme retiming and sizing, while the increase of area is still negligible. This timing performance is enhanced beyond the limit of traditional sequential designs. It also demonstrates that compared with those after retiming and sizing, the circuits with VirtualSync+ can achieve better timing performance under the same area cost or smaller area cost under the same clock period, respectively. Grace Li Zhang, Bing Li 0005, Xing Huang 0001, Xunzhao Yin, Cheng Zhuo, Masanori Hashimoto, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2021 | Connection-based Processing-In-Memory Engine Design Based on Resistive CrossbarsabstractDeep neural networks have successfully been applied to various fields. The efficient deployment of neural network models emerges as a new challenge. Processing-in-memory (PIM) engines that carry out computation within memory structures are widely studied for improving computation efficiency and data communication speed. In particular, resistive memory crossbars can naturally realize the dot-product operations and show great potential in PIM design. The common practice of a current-based design is to map a matrix to a crossbar, apply the input data from one side of the crossbar, and extract the accumulated currents as the computation results at the orthogonal direction. In this study, we propose a novel PIM design concept that is based on the crossbar connections. Our analysis on star-mesh network transformation reveals that in a crossbar storing both input data and weight matrix, the dot-product result is embedded within the network connection. Our proposed connection-based PIM design leverages this feature and discovers the latent dot-products directly from the connection information. Moreover, in the connection-based PIM design, the output current range of resistive crossbars can easily be adjusted, leading to more linear conversion to voltage values, and the output circuitry can be shared by multiple resistive crossbars. The simulation results show that our design can achieve on average 46.23% and 33.11% reductions in area and energy consumption, with a merely 3.85% latency overhead compared with current-based designs. Shuhang Zhang, Hai Li 0001, Ulf Schlichtmann |
ASP-DAC | 3 |
| 2021 | Robustness of Neuromorphic Computing with RRAM-based Crossbars and Optical Neural NetworksabstractRRAM-based crossbars and optical neural networks are attractive platforms to accelerate neuromorphic computing. However, both accelerators suffer from hardware uncertainties such as process variations. These uncertainty issues left unaddressed, the inference accuracy of these computing platforms can degrade significantly. In this paper, a statistical training method where weights under process variations and noise are modeled as statistical random variables is presented. To incorporate these statistical weights into training, the computations in neural networks are modified accordingly. For optical neural networks, we modify the cost function during software training to reduce the effects of process variations and thermal imbalance. In addition, the residual effects of process variations are extracted and calibrated in hardware test, and thermal variations on devices are also compensated in advance. Simulation results demonstrate that the inference accuracy can be improved significantly under hardware uncertainties for both platforms. Grace Li Zhang, Bing Li 0005, Ying Zhu 0008, Yiyu Shi 0001, Xunzhao Yin, Cheng Zhuo, Huaxi Gu, Tsung-Yi Ho, Ulf Schlichtmann |
ASP-DAC | 10 |
| 2021 | Light: A Scalable and Efficient Wavelength-Routed Optical Networks-On-Chip TopologyabstractWavelength-routed optical networks-on-chip (WRONoCs) are known for delivering collision- and arbitration-free on-chip communication in many-cores systems. While appealing for low latency and high predictability, WRONoCs are challenged by scalability concerns due to two reasons: (1) State-of-the-art WRONoC topologies use a large number of microring resonators (MRRs) which result in much MRR tuning power and crosstalk noise. (2) The positions of master and slave nodes in current topologies do not match realistic layout constraints. Thus, many additional waveguide crossings will be introduced during physical implementation, which degrades the network performance. In this work, we propose an N x (N - 1) WRONoC topology: Light with a 4 x 3 router Hash as the basic building block, and a simple but efficient approach to configure the resonant wavelength for each MRR. Experimental results show that Light outperforms state-of-the-art topologies in terms of enhancing signal-to-noise ratio (SNR) and reducing insertion loss, especially for large-scale networks. Furthermore, Light can be easily implemented onto a physical plane without causing external waveguide crossings. Zhidan Zheng, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
ASP-DAC | 4 |
| 2021 | Bayesian Inference Based Robust Computing on Memristor CrossbarabstractMemristor based crossbars are a promising platform for neural network acceleration. To deploy a trained network model on a memristor crossbar, memristors need to be programmed to realize the trained weights of the network. However, due to process and dynamic variations, deviation of weights from the trained value is inevitable and inference accuracy thus degrades. In this paper, we propose a unified Bayesian inference based framework which connects hardware variations and algorithmic training together for robust computing on memristor crossbars. The framework incorporates different levels of variations into priori weight distribution, and transforms robustness optimization to Bayesian neural network training, where weights of neural networks are optimized to accommodate variations and minimize inference degradation. Simulation results with the proposed framework confirm stable inference accuracy under process and dynamic variations. Qingrong Huang, Grace Li Zhang, Xunzhao Yin, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo |
DAC | 6 |
| 2021 | FAST: A Fast Automatic Sweeping Topology Customization Method for Application-Specific Wavelength-Routed Optical NoCsabstractOptical network-on-chip (ONoC) is an emerging upgrade for electronic network-on-chip (ENoC). As a kind of ONoC, wavelength-routed optical network-on-chip (WRONoC) shows ultra-high bandwidth and ultra-low latency in data communication. Manually designed WRONoC topologies typically reserve all to all communications. Topologies customized for application-specific networks can save resources, but require automation for their efficient design. The state-of-the-art design automation method proposes an integer-linear-programming (ILP) model. The runtime for solving the ILP model increases exponentially with the growth of communication density. Besides, the locations of the physical ports are not taken into consideration in the model. This causes unavoidable detours and crossings in physical layout. In this work, we present FAST: an automatic topology customization and optimization method combining ILP and a sweeping technique. FAST overcomes the runtime problem and provides multiple topology variations with different port orders for physical layout. Experimental results show that FAST is thousands times faster when tackling dense communications and ten to thousands times faster when tackling sparse communications while providing multiple better or equivalent topologies regarding resource usage and the worst-case insertion loss. Xiao Moyuan, Tsun-Ming Tseng, Ulf Schlichtmann |
DATE | 3 |
| 2021 | Energy-Aware Designs of Ferroelectric Ternary Content Addressable MemoryabstractTernary content addressable memories (TCAMs) are a special form of computing-in-memory (CiM) circuits that aim to address the so-called memory wall issues by merging the parallel search function with memory blocks. Due to the content addressing nature, TCAMs have been widely utilized for search intensive tasks in low-power, data analytic applications, such as IP routers, associative memories, and learning models. While most state-of-the-art TCAM designs focus on improving the TCAM density by harnessing compact nonvolatile memories (NVMs), little efforts have been spent on reducing and optimizing the energy consumption of the NVM based TCAM. In this paper, by exploiting the Ferroelectric FET (FeFET) as a representative NVM, we propose two compact and energy-aware designs of ferroelectric TCAMs for low power applications. We first introduce a novel 2FeFET based XOR-like gate structure that can also be adopted to other NVMs, and then leverage the structure to propose two TCAM designs that achieve high energy efficiency by either reducing the associated precharge overhead (2FeFET-1T cell), or eliminating the precharge phase typically required by TCAMs (2FeFET-2T cell). We evaluate and compare the designs w.r.t area, search energy and delay at array level with other existing designs, and benchmark the proposed TCAM designs in an associative memory based GPU architecture. The results suggest that the proposed 2FeFET-1T/2FeFET-2T TCAM design consumes 3.03X/8.08X less search energy than the conventional 16T CMOS TCAM, while the proposed design cell area is only 32.1%/39.3% of the latter. Compared with the state-of-the-art 2FeFET only TCAM array, our proposed designs still achieve 1.79X and 4.79X search energy reduction, respectively. Moreover, our proposed designs can achieve, on average, 45.2%/51.5% energy saving compared with the conventional GPU based architecture at the application level. Yu Qian 0002, Zhenhao Fan, Chao Li 0065, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
DATE | 9 |
| 2021 | Hardware-Software Codesign of Weight Reshaping and Systolic Array Multiplexing for Efficient CNNsabstractThe last decade has witnessed the breakthrough of deep neural networks (DNNs) in various fields, e.g., image/speech recognition. With the increasing depth of DNNs, the number of multiply-accumulate operations (MAC) with weights explodes significantly, preventing their applications in resource-constrained platforms. The existing weight pruning method is considered to be an effective method to compress neural networks for acceleration. However, weights after pruning usually exhibit irregular patterns. Implementing MAC operations with such irregular weight patterns on hardware platforms with regular designs, e.g., GPUs and systolic arrays, might result in an underutilization of hardware resources. To utilize the hardware resource efficiently, in this paper, we propose a hardware-software codesign framework for acceleration on systolic arrays. First, weights after unstructured pruning are reorganized into a dense cluster. Second, various blocks are selected to cover the cluster seamlessly. To support the concurrent computations of such blocks on systolic arrays, a multiplexing technique and the corresponding systolic architecture is developed for various CNNs. The experimental results demonstrate that the performance of CNN inferences can be improved significantly without accuracy loss. Jingyao Zhang 0002, Huaxi Gu, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann |
DATE | 5 |
| 2021 | An Efficient Programming Framework for Memristor-based Neuromorphic ComputingabstractMemristor-based crossbars are considered to be promising candidates to accelerate vector-matrix computation in deep neural networks. Before being applied for inference, mem-ristors in the crossbars should be programmed to conductances corresponding to the network weights after software training. Existing programming methods, however, adjust conductances of memristors individually with many programming-reading cycles. In this paper, we propose an efficient programming framework for memristor crossbars, where the programming process is partitioned into the predictive phase and the fine-tuning phase. In the predictive phase, multiple memristors are programmed simultaneously with a memristor programming model and IR-drop estimation. To deal with the programming inaccuracy resulting from process variations, noise and IR-drop and move conductances to target values, memristors are fine-tuned afterwards to reach a specified programming accuracy. Simulation results demonstrate that the proposed method can reduce the number of programming-reading cycles by up to 94.77% and 90.61% compared to existing one-by-one and row-by-row programming methods, respectively. Grace Li Zhang, Bing Li 0005, Xing Huang 0001, Shuhang Zhang, Florin Burcea, Helmut E. Graeb, Tsung-Yi Ho, Hai Li 0001, Ulf Schlichtmann |
DATE | 10 |
| 2021 | Exploiting Active Learning for Microcontroller Performance PredictionabstractSpeed monitors provide on-chip measurements of the the performance of integrated circuits. In recent years, they have been extensively used to predict Fmaxof microcontrollers for speed binning and performance screening during production test. However, while the use of machine learning is getting increasingly popular, the models may become significantly inaccurate if not trained on the appropriate devices. Previous research has demonstrated how to predict performance from speed-monitor data using corner-lot wafers. We show how to extend this approach to select the best corner-lot wafers to label when preparing the training set, thus significantly reducing the time and cost required for the process. Nicolò Bellarmino, Riccardo Cantoro, Martin Huch, Tobias Kilian, Raffaele Martone, Ulf Schlichtmann, Giovanni Squillero |
ETS | 6 |
| 2021 | Reliable Memristor-based Neuromorphic Design Using Variation- and Defect-Aware TrainingabstractThe memristor crossbar provides a unique opportunity to develop a neuromorphic computing system (NCS) with high scalability and energy efficiency. However, the reliability issues that arise from the immature fabrication process and physical device limitations, i.e., variations and stuck-at-faults (SAF), dramatically prevent its wide application in practice. Specifically, variations make the programmed weights deviate from their expected values. On the other hand, defective mem-ristors cannot even represent the weights effectively. In this work, we propose a variation- and defect-aware framework to improve the reliability of memristor-based NCS while minimizing the inference performance loss. We propose to develop analytical weight models to characterize the non-ideal effects of variations and SAFs, which can then be incorporated into a Bayesian neural network as priori and constraint. We then convert the reliability improvement to the neural network training for optimal weights that can accommodate variations and defects across the chips, which does not require computation-intensive retraining or cost-expensive testing. Extensive experimental results with the proposed framework confirm its effective capability of improving the reliability of NCS, while significantly mitigating the inference accuracy degradation under even severe variations and SAFs. Grace Li Zhang, Xunzhao Yin, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo |
ICCAD | 5 |
| 2021 | BigIntegr: One-Pass Architectural Synthesis for Continuous-Flow Microfluidic Lab-on-a-Chip SystemsabstractThe emergence of continuous-flow microfluidics has led to a revolution in biochemistry and biomedicine. On such a microscale lab-on-a-chip system, complex biochemical assays, e.g., DNA analysis and drug discovery, can be executed efficiently without any human intervention. Owing to the high complexity of chip architecture and assay protocol, considerable effort has been directed towards the design automation of such chips over the past decade. Existing methods, however, perform the corresponding design tasks including binding, scheduling, placement, and routing separately, leading to serious gaps between different steps and may even cause design failure. To overcome these drawbacks, in this paper, we propose a one-pass architecture synthesis flow called BigIntegr, for continuous-flow microfluidic lab-on-a-chip, integrating all the design steps into an “organic whole”, which has never been considered in prior work. With the proposed BigIntegr, the aforementioned design tasks can be synchronized seamlessly and performed in a combined manner, thereby eliminating the gaps between design steps. As a result, biochip architectures with both high efficiency and low cost can be generated without any design adjustments and modifications. Experimental results on multiple benchmarks demonstrate the effectiveness of the proposed automation flow. Xing Huang 0001, Youlin Pan, Wenzhong Guo, Robert Wille, Tsung-Yi Ho, Ulf Schlichtmann |
ICCAD | 7 |
| 2021 | Manufacturing Cycle-Time Optimization Using Gaussian Drying Model for Inkjet-Printed ElectronicsabstractInkjet-printed electronics have attracted considerable attention for low-cost mass production. To avoid undesired device behavior due to accidental ink merging and redistribution, high-density designs can benefit from layering and drying in batches. The overall manufacturing cycle-time, however, now becomes dominated by the cumulative drying time of these individual layers. The state-of-the-art approach decomposes the whole design, arranges the modified objects in different layers, and minimizes the number of layers. Fewer layers imply a reduction in the number of printing iterations and thus a higher manufacturing efficiency. Nevertheless, printing objects with significantly different drying dynamics in the same layer leads to a reduction of manufacturing efficiency, since the longest drying object in a given layer dominates the time required for this layer to dry. Consequently, an accurate estimation of the individual layers' drying time is indispensable to minimize the manufacturing cycle-time. To this end, we propose the first Gaussian drying model to evaluate the local evaporation rate in the drying process. Specifically, we estimate the drying time depending on the number, area, and distribution of the objects in a given layer. Finally, we minimize the total drying time by assigning to-be-printed objects to different layers with mixed-integer-linear programming (MILP) methods. Experimental results demonstrate that our Gaussian drying model closely approximates the actual drying process. In particular, comparing the non-optimized fabrication to the optimized results demonstrates that our method is able to reduce the drying time by 39%. Tsun-Ming Tseng, Meng Lian 0001, Mengchu Li, Philipp Rinklin, Leroy Grob, Bernhard Wolfrum, Ulf Schlichtmann |
ICCAD | 7 |
| 2021 | Peripheral Circuitry Assisted Mapping Framework for Resistive Logic-In-Memory ComputingabstractIn-memory computing has been applied in different fields due to its superior speed and energy efficiency. Among a variety of memory technologies that have been explored, resistive memory has widely been adopted for various purposes, including Processing-In-Memory (PIM) for neural networks and Logic-In-Memory (LIM) for general logic operations. PIM has intensively been studied in recent years, while the progress in developing LIM computing falls behind. LIM computing is usually implemented based on MAGIC operations, which require inputs to be aligned regularly along rows or columns in a memory crossbar. As the intermediate data generated during the logic execution are normally scattered across the memory crossbar, alignment operations are inserted to align the data, which often costs numerous cycles and dominates the overall latency. In current MAGIC-based designs, alignment operations induce a significant overhead in either area or latency. Therefore, the Area-Latency-Product (ALP), known as a key metric for circuit performance, still has significant optimization potential in LIM computing. In this work, we leverage peripheral circuitry to conduct alignment operations and propose a novel mapping framework to optimize the latency and area costs. Intermediate data are read out, processed in peripheral circuits, then in parallel written back into target cells of the memory crossbar. The approach eliminates the use of redundant memory cells, leading to area reduction. Moreover, it enables simultaneous alignments of multiple intermediate data, which can decrease the overall latency significantly. Based on simulation results, our proposed mapping framework can achieve around 93% ALP reductions on average compared with prior designs with merely 2.13% total area overhead. Shuhang Zhang, Hai Li 0001, Ulf Schlichtmann |
ICCAD | 3 |
| 2021 | ToPro: A Topology Projector and Waveguide Router for Wavelength-Routed Optical Networks-on-ChipabstractTo meet the ever-increasing requirements of on-chip communication, the trend is towards wavelength-routed optical networks-on-chip (WRONoCs), which support high-speed communication with low power. A typical WRONoC design flow consists of two consecutive steps: topological design and physical design. Current physical design tools interpret the input topology as a pure logic scheme and perform placement and routing for all network components from scratch. Due to the large design complexity and the layout constraints, additional waveguide crossings in the synthesized layouts are hardly avoidable, which results in an increase in insertion loss and crosstalk noise and thus degrades the network performance. In this work, we propose a physical design tool, ToPro, which retains the interconnection among the optical switching elements by projecting the structure of a WRONoC topology onto the physical plane, and focuses on the waveguide routing to the IP-cores. To avoid the increase in insertion loss and crosstalk noise, ToPro removes the extra crossings and long detours of waveguides by changing the routing order of nets. The experimental results demonstrate the superiority of ToPro in time- and energy-efficiency. For example, compared to a state-of-the-art design automation tool, ToPro synthesizes a network with 16 IP-cores with a 17% reduction on the worst-case insertion loss and decreases the synthesis time from more than six days to less than one second. Zhidan Zheng, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
ICCAD | 4 |
| 2021 | Relative-Scheduling-Based High-Level Synthesis for Flow-Based Microfluidic BiochipsabstractThe rapid development of microfluidic biochips requires matching automated synthesis methods. In particular, high-level synthesis methods for microfluidic chips need to consider various bio-constraints regarding time and device occupancy. Recent biochemical applications show that the duration of some bio-operations cannot be predicted in advance, which increases the likelihood that current synthesis methods would waste on-chip resources or even violate given bio-constraints. In this work, we present a relative-scheduling-based high-level synthesis method to optimize the bio-assay schedules and the usage of on-chip devices considering bio-operations with indeterminate durations. Experimental results show that our method significantly reduces the total execution time of bioassays without violating bio-constraints when using the same on-chip resources. Fangda Zuo, Mengchu Li, Tsun-Ming Tseng, Tsung-Yi Ho, Ulf Schlichtmann |
ICCAD | 5 |
| 2021 | A Scalable Design Flow for Performance Monitors Using Functional Path Ring OscillatorsabstractThe automotive industry sets high reliability standards for microcontroller (MCUs). To increase reliability, the automotive MCU manufacturers are looking for accurate performance screening. One of these performance screening mechanisms are functional path ring oscillators (RO). In this paper, a scalable and efficient method for creating functional path ring oscillators is presented. Implementation data demonstrate that functional path RO monitors show a significant advantage in area and power consumption over comparable performance screening methods. Tobias Kilian, Heiko Ahrens, Daniel Tille, Martin Huch, Ulf Schlichtmann |
ITC | 5 |
| 2021 | A Distributed Hardware Monitoring System for Runtime Verification on Multi-Tile MPSoCsabstractExhaustive verification techniques do not scale with the complexity of today’s multi-tile Multi-processor Systems-on-chip (MPSoCs). Hence, runtime verification (RV) has emerged as a complementary method, which verifies the correct behavior of applications executed on the MPSoC during runtime. In this article, we propose a decentralized monitoring architecture for large-scale multi-tile MPSoCs. In order to minimize performance and power overhead for RV, we propose a lightweight and non-intrusive hardware solution. It features a new specialized tracing interconnect that distributes and sorts detected events according to their timestamps. Each tile monitor has a consistent view on a globally sorted trace of events on which the behavior of the target application can be verified using logical and timing requirements. Furthermore, we propose an integer linear programming-based algorithm for the assignment of requirements to monitors to exploit the local resources best. The monitoring architecture is demonstrated for a four-tiled MPSoC with 20 cores implemented on a Virtex-7 field-programmable gate array (FPGA). Marcel Mettler, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Trans. Archit. Code Optim. | 3 |
| 2021 | DCSA: Distributed Channel-Storage Architecture for Flow-Based Microfluidic BiochipsabstractFlow-based microfluidic biochips have attracted much attention in the EDA community due to their miniaturized size and execution efficiency. Previous research, however, still follows the traditional computing model with a dedicated storage unit, which actually becomes a bottleneck of the performance of biochips. In this article, we propose a distributed channel-storage architecture (DCSA) to cache fluid samples inside flow channels temporarily. Since distributed storage can be accessed more efficiently than a dedicated storage unit and channels can switch between the roles of transportation and storage easily, biochips with this architecture can achieve a higher execution efficiency even with fewer resources. Furthermore, we also address the flow-path planning that enables the manipulation of actual fluid transportation/caching on a chip. The simulation results confirm that the execution efficiency of a bioassay can be improved significantly, while the number of valves in the biochip can be reduced accordingly. Also, flow paths for transportation tasks can be constructed and planned automatically with minimum extra resources. Xing Huang 0001, Bing Li 0005, Hailong Yao 0002, Paul Pop, Tsung-Yi Ho, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2021 | REPAIR: Control Flow Protection based on Register Pairing Updates for SW-Implemented HW Fault ToleranceabstractSafety-critical embedded systems may either use specialized hardware or rely on Software-Implemented Hardware Fault Tolerance (SIHFT) to meet soft error resilience requirements. SIHFT has the advantage that it can be used with low-cost, off-the-shelf components such as standard Micro-Controller Units. For this, SIHFT methods apply redundancy in software computation and special checker codes to detect transient errors, so called soft errors, that either corrupt the data flow or the control flow of the software and may lead to Silent Data Corruption (SDC). So far, this is done by applying separate SIHFT methods for the data and control flow protection, which leads to large overheads in computation time. This work in contrast presents REPAIR, a method that exploits the checks of the SIHFT data flow protection to also detect control flow errors as well, thereby, yielding higher SDC resilience with less computational overhead. For this, the data flow protection methods entail duplicating the computation with subsequent checks placed strategically throughout the program. These checks assure that the two redundant computation paths, which work on two different parts of the register file, yield the same result. By updating the pairing between the registers used in the primary computation path and the registers in the duplicated computation path using the REPAIR method, these checks also fail with high coverage when a control flow error, which leads to an illegal jumps, occurs. Extensive RTL fault injection simulations are carried out to accurately quantify soft error resilience while evaluating Mibench programs along with an embedded case-study running on an OpenRISC processor. Our method performs slightly better on average in terms of soft error resilience compared to the best state-of-the-art method but requiring significantly lower overheads. These results show that REPAIR is a valuable addition to the set of known SIHFT methods. Uzair Sharif, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Maximizing the Communication Parallelism for Wavelength-Routed Optical Networks-On-ChipsabstractEnabled by recent development in silicon photonics, wavelength-routed optical networks-on-chips (WRONoCs) emerge as an appealing next-generation architecture for the communication in multiprocessor system-on-chip. WRONoCs apply a passive routing mechanism that statically reserves all data transmission paths at design time, and are thus able to avoid the latency and energy overhead for arbitration, compared to other ONoC architectures. Current research mostly assumes that in a WRONoC topology, each initiator node sends one bit at a time to a target node. However, the communication parallelism can be increased by assigning multiple wavelengths to each path, which requires a systematic analysis of the physical parameters of the silicon microring resonators and the wavelength usage among different paths. This work proposes a mathematical modeling method to maximize the communication parallelism of a given WRONoC topology, which provides a foundation for exploiting the bandwidth potential of WRONoCs. Experimental results show that the proposed method significantly outperforms the state-of-the-art approach, and is especially suitable for application-specific WRONoC topologies. Mengchu Li, Tsun-Ming Tseng, Mahdi Tala, Ulf Schlichtmann |
ASP-DAC | 4 |
| 2020 | Investigating the Inherent Soft Error Resilience of Embedded Applications by Full-System SimulationabstractIt has long been acknowledged that some applications feature inherent resilience against soft errors, e.g., the impact of soft errors on multimedia applications is often non-visible to humans. In this paper we investigate the inherent resilience of two typical embedded applications using a case study of a control system and a robot arm. Both studies were enabled by our mixed-mode fault injection simulator ETISS-ML, which allows RTL-accurate fault injection while being able to simulate very long scenarios, e.g. robot movements of several seconds. Our results indicate that full simulation of the embedded system and its environment are required to classify whether the system can tolerate the impact of a soft error. This is due to the fact that it is hard to predict the impact of a certain output deviation without investigating the change in the system behavior taking into account the control loop. Based on this classification method we hope to be able to exploit this resilience for lowering the cost of error detection mechanisms in future research. Uzair Sharif, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ASP-DAC | 3 |
| 2020 | Timing Resilience for Efficient and Secure CircuitsabstractIn this paper, we will cover several techniques that can enhance the resilience of timing of digital circuits. Using post-silicon tuning components, the clock arrival times at flip-flops can be modified after manufacturing to balance delays between flip-flops. The actual delay properties of flip-flops will be examined to exploit the natural flexibility of such components. Wave-pipelining paths spanning several flip-flop stages can be integrated into a synchronous design to improve the circuit performance and to reduce area. In addition, with this technique, it cannot be taken for granted anymore that all the combinational paths in a circuit work with respect to one clock period. Therefore, a netlist alone does not represent all the design information. This feature enables the potential to embed wave-pipelining paths into a circuit to increase the complexity of reverse engineering. In order to replicate a design, attackers therefore have to identify the locations of the wave-pipelining paths, in addition to the netlist extracted from reverse engineering. Therefore, the security of the circuit against counterfeiting can be improved. Grace Li Zhang, Michaela Brunner, Bing Li 0005, Georg Sigl, Ulf Schlichtmann |
ASP-DAC | 5 |
| 2020 | Transport-Free Module Binding for Sample Preparation using Microfluidic Fully Programmable Valve ArraysabstractMicrofluidic fully programmable valve array (FPVA) biochips have emerged as general-purpose flow-based microfluidic lab-on-chips (LoCs). An FPVA supports highly re-configurable on-chip components (modules) in the two-dimensional grid-like structure controlled by some software programs, unlike application-specific flow-based LoCs. Fluids can be loaded into or washed from a cell with the help of flows from the inlet to outlet of an FPVA, whereas cell-to-cell transportation of discrete fluid segment(s) is not precisely possible. The simplest mixing module to realize on an FPVA-based LoC is a four-way mixer consisting of a 2 × 2 array of cells working as a ring-like mixer having four valves. In this paper, we propose a design automation method for sample preparation that finds suitable placements of mixing operations of a mixing tree using four-way mixers without requiring any transportation of fluid(s) between modules. We also propose a heuristic that modifies the mixing tree to reduce the sample preparation time. We have performed an extensive simulation and examined several parameters to determine the performance of the proposed solution. Gautam Choudhary, Sandeep Pal, Debraj Kundu, Sukanta Bhattacharjee, Shigeru Yamashita, Bing Li 0005, Ulf Schlichtmann, Sudip Roy 0001 |
DATE | 7 |
| 2020 | A Pulse-width Modulation Neuron with Continuous Activation for Processing-In-Memory EnginesabstractProcessing-in-memory engines have successfully been applied to accelerate deep neural networks. For improving computing efficiency, spiking-based designs are widely explored. However, spiking-based designs quantize inter-layer signals naturally, leading to performance loss. In addition, the spike mismatch effect makes digital processing necessary, impeding direct signal transfer between layers and thus resulting in longer latency. In this paper, we propose a novel neuron design based on pulse width modulation, avoiding the quantization step and bypassing spike mismatch via the continuous activation. The computation latency and circuit complexity can significantly be reduced due to the absence of quantization and digital processing steps, while keeping a competitive performance. Simulation results show that the proposed neuron design can achieve > 100× speedup compared with spiking-based designs. The area and power consumption can be reduced up to 74.87% and 25.63%. Shuhang Zhang, Bing Li 0005, Hai Li 0001, Ulf Schlichtmann |
DATE | 4 |
| 2020 | Statistical Training for Neuromorphic Computing using Memristor-based Crossbars Considering Process Variations and NoiseabstractMemristor-based crossbars are an attractive platform to accelerate neuromorphic computing. However, process variations during manufacturing and noise in memristors cause significant accuracy loss if not addressed. In this paper, we propose to model process variations and noise as correlated random variables and incorporate them into the cost function during training. Consequently, the weights after this statistical training become more robust and together with global variation compensation provide a stable inference accuracy. Simulation results demonstrate that the mean value and the standard deviation of the inference accuracy can be improved significantly, by even up to 54% and 31%, respectively, in a two-layer fully connected neural network. Ying Zhu 0008, Grace Li Zhang, Bing Li 0005, Yiyu Shi 0001, Tsung-Yi Ho, Ulf Schlichtmann |
DATE | 7 |
| 2020 | Reliable and Robust RRAM-based Neuromorphic ComputingabstractRRAM-based crossbars are a promising hardware platform to accelerate computations in neural networks. Before such a crossbar can be used as an accelerator for neural networks, RRAM cells should be programmed to target resistances to represent weights in neural networks. However, this process degrades the valid range of the resistances of RRAM cells from the fresh state, called aging effect. Therefore, after a certain number of programming iterations, these RRAM cells cannot be programmed reliably anymore, affecting the classification accuracy of neural networks negatively. In addition, process variations during manufacturing and noise during programming of RRAM cells also lead to significant accuracy degradation. To solve the problems described above, in this paper, we introduce a software/hardware codesign framework to reduce the aging effect in RRAM crossbars. To counter process variations and noise, we first model them as random variables and then modify the computations in software training considering these variables. Simulation results show that the lifetime of RRAM crossbars can be extended by up to 11 times with the codesign framework and the mean value and the standard deviation of the inference accuracy under process variations and noise can be improved significantly. Grace Li Zhang, Bing Li 0005, Ying Zhu 0008, Shuhang Zhang, Yiyu Shi 0001, Tsung-Yi Ho, Hai Li 0001, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 9 |
| 2020 | PathDriver: A Path-Driven Architectural Synthesis Flow for Continuous-Flow Microfluidic BiochipsabstractContinuous-flow microfluidic biochips have attracted high research interest over the past years. Inside such a chip, fluid samples of milliliter volumes are efficiently transported between devices (e.g., mixers, etc.) to automatically perform various laboratory procedures in biology and biochemistry. Each transportation task, however, requires an exclusive flow path composed of multiple contiguous microchannels during its execution period. Excess/waste fluids, in the meantime, should be discarded by independent flow paths connected to waste ports. All these paths are etched in a very tiny chip area using multilayer soft lithography and driven by flow ports connecting with external pressure sources, forming a highly integrated chip architecture that dominates the performance of biochips. In this paper, we propose a practical synthesis flow called PathDriver for the design automation of microfluidic biochips, integrating the actual fluid manipulations into both high-level synthesis and physical design, which has never been considered in prior work. Given the protocols of biochemical applications, PathDriver aims to generate highly efficient chip architectures with a flow-path network that enables the manipulation of actual fluid transportation and removal. Additionally, fluid volume management between devices and flow-path minimization are realized for the first time, thus ensuring the correctness of assay outcomes while reducing the complexity of chip architectures. Experimental results on multiple benchmarks demonstrate the effectiveness of the proposed synthesis flow. Xing Huang 0001, Youlin Pan, Grace Li Zhang, Bing Li 0005, Wenzhong Guo, Tsung-Yi Ho, Ulf Schlichtmann |
ICCAD | 7 |
| 2020 | Overview of 2020 CAD Contest at ICCADabstractThe "CAD Contest at ICCAD" is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has been publishing many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The 2020 CAD Contest hits a record high of 186 teams from all over the world, which represents more than 50% growth compared to last year. The contest keeps enhancing impact and boosting EDA research. Ing-Chao Lin, Ulf Schlichtmann, Tsung-Wei Huang, Mark Po-Hung Lin |
ICCAD | 2 |
| 2020 | PSION 2: Optimizing Physical Layout of Wavelength-Routed ONoCs for Laser Power ReductionabstractOptical Networks-on-Chip (ONoCs) are becoming increasingly attractive for intra-chip communications due to their low power-perbit requirements and high bandwidth. Wavelength-Routed ONoCs (WRONoCs), a subtype of ONoCs, further reduce network latency. Recently, tools to design WRONoCs have been developed, but these tools are still incomplete as they do not yet consider key design aspects such as the type of laser source used and the impact of the laser Power Distribution Network (PDN) on the laser power consumption. In this work we propose the first design automation tool to combine awareness of both on-chip and off-chip lasers with optimization of both the logical topology and the physical layout of WRONoCs for application-specific designs. Compared to previous works, the incorporation of the type of laser and the PDN into the optimization process combined with a new Generic Routing Unit (GRU) placement method leads to a laser power reduction of up to 20%. Alexandre Truppel, Tsun-Ming Tseng, Ulf Schlichtmann |
ICCAD | 3 |
| 2020 | Countering Variations and Thermal Effects for Accurate Optical Neural NetworksabstractOptical neural networks (ONNs) have emerged as a promising high-performance computing platform to accelerate deep neural networks. In ONNs, phases of light are modulated through Mach-Zehnder Interferometers (MZIs), and MZIs are connected in a gridlike layout to implement multiply-accumulate operations. However, ONNs are very sensitive to process variations and thermal effects. This sensitivity leads to a significant degradation of inference accuracy of ONNs and thus renders them unusable in practice. In this paper, we propose a framework to calibrate process variations and counter thermal effects by power compensation. Experimental results demonstrate that the proposed framework can recover the inference accuracy under variations and thermal effects, e.g., from as low as 11.05% back to 74.11% for LeNet-5 on Cifar10, so that ONNs can achieve an inference accuracy similar to the accuracy after software training while providing their high bandwidth in neuromorphic computing. Ying Zhu 0008, Grace Li Zhang, Bing Li 0005, Xunzhao Yin, Cheng Zhuo, Huaxi Gu, Tsung-Yi Ho, Ulf Schlichtmann |
ICCAD | 8 |
| 2020 | Machine Learning based Performance Prediction of Microcontrollers using Speed MonitorsabstractDuring the manufacturing process, electronic devices are thoroughly tested for defects. However, testing for well-known fault models, such as stuck-at and transition delay, may not be sufficient for an effective performance screening. In modern devices, Design-for-Testability features embedded at design time can allow the tester to apply stimuli and measure different critical parameters. We propose to use some of these structures, namely the speed monitors, to predict the maximum operating speed, and screen out under-performing devices. We design a complete methodology, from the extraction of robust labels, through a machine-learning algorithm, down to a post-processing step, able to meet the quality standards imposed by industry. Experimental results using real production data demonstrate the feasibility of the approach. Riccardo Cantoro, Martin Huch, Tobias Kilian, Raffaele Martone, Ulf Schlichtmann, Giovanni Squillero |
ITC | 5 |
| 2020 | Machine learning and structural characteristics for reverse engineering
Johanna Baehr 0001, Alessandro Bernardini, Georg Sigl, Ulf Schlichtmann |
Integr. | 4 |
| 2020 | Test Generation for Flow-Based Microfluidic Biochips With General ArchitecturesabstractFlow-based microfluidic biochips have become a promising platform for complex biochemical assays. As the integration of such chips is increasing, a flexible general reconfigurable platform, fully programmable valve array (FPVA), has emerged. Such a 2-D array comprises regularly arranged valves using which flow-networks with different geometry, size, and connectivity can be constructed dynamically. However, the test generation for such arrays becomes challenging due to the large number of potential flow-networks and transportation paths that can be configured on-chip. In this article, we propose a strategy to generate efficient test patterns for FPVAs based on the concepts of test paths and cuts. These patterns together can cover multiple faults in both flow and control layers. We also introduce the concept of test trees and multiple cuts for a test pattern to deal with faults in FPVAs with multiple ports. Moreover, the proposed method can be applied to generate test patterns for traditional flow-based biochips with predefined architectures. The simulation results demonstrate that defects in FPVAs can be detected reliably by a limited number of test patterns generated by the proposed method. For traditional biochips with predefined architectures, these patterns also exhibit an improved test efficiency. Bing Li 0005, Bhargab B. Bhattacharya, Krishnendu Chakrabarty, Tsung-Yi Ho, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | An Efficient Fault-Tolerant Valve-Based Microfluidic Routing Fabric for Droplet Barcoding in Single-Cell AnalysisabstractSingle-cell analysis is used to gain insights into diseases, such as cancer. Advances in microfluidic solutions have enabled the efficient classification and analysis of a heterogeneous population of cells. Recently, a hybrid microfluidic platform was proposed for concurrent single-cell analysis on thousands of heterogeneous cells. In this design, barcoding droplets are routed using a valve-based routing fabric to label the input cells. However, prior work overlooked defects that are likely to occur during chip fabrication and system integration and the fault tolerance of this routing fabric remains a major concern. We address the above limitation and introduce a low-overhead design technique for guaranteeing the tolerance of single faults, while maintaining the efficiency of the cell-analysis platform. We show that the proposed method is optimal in that it minimizes the overhead in terms of fabric size. Yasamin Moradi, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | PSION+: Combining Logical Topology and Physical Layout Optimization for Wavelength-Routed ONoCsabstractOptical networks-on-chip (ONoCs) are a promising solution for high-performance multicore integration with better latency and bandwidth than traditional electrical NoCs. Wavelength-routed ONoCs (WRONoCs) offer yet additional performance guarantees. However, WRONoC design presents new EDA challenges which have not yet been fully addressed. So far, most topology analysis is abstract, i.e., overlooks layout concerns, while for layout the tools available perform place and route (P&R) but no topology optimization. Thus, a need arises for a novel optimization method combining both aspects of WRONoC design. In this article, such a method, PSION+, is laid out. This new procedure uses a linear programming model to optimize a WRONoC physical layout template to optimality. This template-based optimization scheme is a new idea in this area that seeks to minimize problem complexity while keeping design flexibility. A simple layout template format is introduced and explored. Finally, multiple model reduction techniques to reduce solver run-time are also presented and tested. When compared to the state-of-the-art design procedure, results show a decrease in maximum optical insertion loss of 41%. Alexandre Truppel, Tsun-Ming Tseng, Davide Bertozzi, José Carlos Alves, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Integrated Control-Fluidic Codesign Methodology for Paper-Based Digital Microfluidic BiochipsabstractPaper-based digital microfluidic biochips (P-DMFBs) have recently emerged as a promising low-cost and fast-responsive platform for biochemical assays. In P-DMFBs, electrodes and control lines are printed on a piece of photograph paper using an inkjet printer and carbon nanotubes (CNTs) conductive ink. Compared with traditional digital microfluidic biochips (DMFBs), P-DMFBs enjoy significant advantages, such as faster in-place fabrication with printer and ink, lower costs, and better disposability. Since electrodes and CNT control lines are printed on the same side of this paper, a critical design challenge for P-DMFB is to prevent control interference between moving droplets and the voltages on CNT control lines. Control interference may result in unexpected droplet movements and thus incorrect assay outputs. To address this design challenge, a control-fluidic codesign methodology is proposed in this paper, along with two demonstrative design flows integrating both fluidic design and control design, i.e., the droplet-oriented codesign flow and the electrode-oriented codesign flow. The droplet-oriented flow is suitable for designing biochips with sparse electrodes and relatively larger number of droplets, whereas the electrode-oriented flow is suitable for biochips with dense electrodes and smaller number of droplets. The computational simulation results of real-life bioassays demonstrate the effectiveness of the proposed codesign flows. Qin Wang 0005, Ulf Schlichtmann, Yici Cai, Weiqing Ji, Zeyan Li 0001, Haena Cheong, Oh-Sun Kwon, Hailong Yao 0002, Tsung-Yi Ho, Kwanwoo Shin, Bing Li 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | TimingCamouflage+: Netlist Security Enhancement With Unconventional TimingabstractWith recent advances in reverse engineering, attackers can reconstruct a netlist to counterfeit chips by opening the die and scanning all layers of authentic chips. This relatively easy counterfeiting is made possible by the use of the standard simple clocking scheme, where all combinational blocks function within one clock period, so that a netlist of combinational logic gates and flip-flops is sufficient to duplicate a design. In this article, we propose to invalidate the assumption that a netlist completely represents the function of a circuit with unconventional timing. With the introduced wave-pipelining (WP) paths, attackers have to capture gate and interconnect delays during reverse engineering, or to test a huge number of combinational paths to identify the WP paths. To hinder the test-based attack, we construct false paths with WP to increase the counterfeiting challenge. The experimental results confirm that WP true paths and false paths can be constructed in benchmark circuits successfully with only a negligible cost, thus thwarting the potential attack techniques. Grace Li Zhang, Bing Li 0005, Meng Li 0004, Bei Yu 0001, David Z. Pan, Michaela Brunner, Georg Sigl, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2020 | Multicontrol: Advanced Control-Logic Synthesis for Flow-Based Microfluidic BiochipsabstractFlow-based microfluidic biochips are one of the most promising platforms used in biochemical and pharmaceutical laboratories due to their high efficiency and low costs. Inside such a chip, fluids of nanoliter volumes are transported between devices for various operations, such as mixing and detection. The transportation channels and corresponding operation devices are controlled by microvalves driven by external pressure sources. Since assigning an independent pressure source to every microvalve would be impractical due to high costs and limited system dimensions, states of microvalves are switched by a control logic using time multiplexing. Existing control-logic designs, however, still switch only a single control channel per operation, leading to a low efficiency. In this article, we present the first automatic synthesis approach for a control logic that is able to switch multiple control channels simultaneously. Moreover, we propose the first fault-aware design in control logic by introducing backup control paths to maintain the correct function even when manufacturing defects occur. The construction of control logic is achieved by a highly efficient framework based on particle swarm optimization, Boolean logic simplification, grid routing, together with mixing multiplexing. The simulation results demonstrate that the proposed multichannel switching mechanism leads to fewer valve-switching times and lower total logic cost, while realizing fault tolerance for all control channels. Ying Zhu 0008, Xing Huang 0001, Bing Li 0005, Tsung-Yi Ho, Qin Wang 0005, Hailong Yao 0002, Robert Wille, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2020 | Machine Learning Approaches for Efficient Design Space Exploration of Application-Specific NoCsabstractIn many Multi-Processor Systems-on-Chip (MPSoCs), traffic between cores is unbalanced. This motivates the use of an application-specific Network-on-Chip (NoC) that is customized and can provide a high performance at low cost in terms of power and area. However, finding an optimized application-specific NoC architecture is a challenging task due to the huge design space. This article proposes to apply machine learning approaches for this task. Using graph rewriting, the NoC Design Space Exploration (DSE) is modelled as a Markov Decision Process (MDP). Monte Carlo Tree Search (MCTS), a technique from reinforcement learning, is used as search heuristic. Our experimental results show that—with the same cost function and exploration budget—MCTS finds superior NoC architectures compared to Simulated Annealing (SA) and a Genetic Algorithm (GA). However, the NoC DSE process suffers from the high computation time due to expensive cycle-accurate SystemC simulations for latency estimation. This article therefore additionally proposes to replace latency simulation by fast latency estimation using a Recurrent Neural Network (RNN). The designed RNN is sufficiently general for latency estimation on arbitrary NoC architectures. Our experiments show that compared to SystemC simulation, the RNN-based latency estimation offers a similar speed-up as the widely used Queuing Theory (QT). Yet, in terms of estimation accuracy and fidelity, the RNN is superior to QT, especially for high-traffic scenarios. When replacing SystemC simulations with the RNN estimation, the obtained solution quality decreases only slightly, whereas it suffers significantly when QT is used. Marcel Mettler, Daniel Mueller-Gritschneder, Thomas Wild, Andreas Herkersdorf, Ulf Schlichtmann |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2020 | Predicting Memory Compiler Performance Outputs Using Feed-forward Neural NetworksabstractTypical semiconductor chips include thousands of mostly small memories. As memories contribute an estimated 25% to 40% to the overall power, performance, and area (PPA) of a product, memories must be designed carefully to meet the system’s requirements. Memory arrays are highly uniform and can be described by approximately 10 parameters depending mostly on the complexity of the periphery. Thus, to improve PPA utilization, memories are typically generated by memory compilers. A key task in the design flow of a chip is to find optimal memory compiler parametrizations that, on the one hand, fulfill system requirements while, on the other hand, they optimize PPA. Although most compiler vendors also provide optimizers for this task, these are often slow or inaccurate. To enable efficient optimization in spite of long compiler runtimes, we propose training fully connected feed-forward neural networks to predict PPA outputs given a memory compiler parametrization. Using an exhaustive search-based optimizer framework that obtains neural network predictions, PPA-optimal parametrizations are found within seconds after chip designers have specified their requirements. Average model prediction errors of less than 3%, a decision reliability of over 99%, and productive usage of the optimizer for successful, large volume chip design projects illustrate the effectiveness of the approach. Felix Last, Max Haeberlein, Ulf Schlichtmann |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2019 | Machine learning and structural characteristics for reverse engineeringabstractIn the past years, much of the research into hardware reverse engineering has focused on the abstraction of gate level netlists to a human readable form. However, none of the proposed methods consider a realistic reverse engineering scenario, where the netlist is physically extracted from a chip. This paper analyzes how errors caused by this extraction and the later partitioning of the netlist affect the ability to identify the functionality. Current formal verification based methods, which compare against a golden model, are incapable of dealing with such erroneous netlists. Two new methods are proposed, which focus on the idea that structural similarity implies functional similarity. The first approach uses fuzzy structural similarity matching to compare the structural characteristics of an unknown design against designs in a golden model library using machine learning. The second approach proposes a method for inexact graph matching using fuzzy graph isomorphisms, based on the functionalities of gates used within the design. For realistic error percentages, both approaches are able to match more than 90% of designs correctly. This is an important first step for hardware reverse engineering methods beyond formal verification based equivalence matching. Johanna Baehr 0001, Alessandro Bernardini, Georg Sigl, Ulf Schlichtmann |
ASP-DAC | 4 |
| 2019 | SeRoHAL: generation of selectively robust hardware abstraction layers for efficient protection of mixed-criticality systemsabstractA major challenge in mixed-criticality system design is to ensure safe behavior under the influence of hardware errors while complying with cost and performance constraints. SeRoHAL generates hardware abstraction layers with software-based safety mechanisms to handle errors in peripheral interfaces. To reduce performance and memory overheads, SeRoHAL can select protection mechanisms, depending on the criticality of the hardware accesses. Petra R. Kleeberger, Juana Rivera, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ASP-DAC | 4 |
| 2019 | Cross-Layer Resilience: Challenges, Insights, and the Road AheadabstractResilience to errors in the underlying hardware is a key design objective for a large class of computing systems, from embedded systems all the way to the cloud. Sources of hardware errors include radiation, circuit aging, variability induced by manufacturing and operating conditions, manufacturing test escapes, and early-life failures. Many publications have suggested that cross-layer resilience, where multiple error resilience techniques from different layers of the system stack cooperate to achieve cost-effective resilience, is essential for designing cost-effective resilient digital systems. This paper presents a comprehensive overview of cross-layer resilience by addressing fundamental cross-layer resilience questions, by summarizing insights derived from recent advances in cross-layer resilience research, and by discussing future cross-layer resilience challenges. Eric Cheng, Daniel Mueller-Gritschneder, Jacob A. Abraham, Pradip Bose, Alper Buyuktosunoglu, Deming Chen, Hyungmin Cho, Yanjing Li, Uzair Sharif, Kevin Skadron, Mircea R. Stan, Ulf Schlichtmann, Subhasish Mitra |
DAC | 12 |
| 2019 | MiniControl: Synthesis of Continuous-Flow Microfluidics with Strictly Constrained Control PortsabstractRecent advances in continuous-flow microfluidics have enabled highly integrated lab-on-a-chip biochips. These chips can execute complex biochemical applications precisely and efficiently within a tiny area, but they require a large number of control ports and the corresponding control logic to generate required pressure patterns for flow control, which, consequently, offset their advantages and prevent their wide adoption. In this paper, we propose the first synthesis flow called MiniControl, for continuous-flow microfluidic biochips (CFMBs) under strict constraints for control ports, incorporating high-level synthesis and physical design simultaneously, which has never been considered in previous work. With the maximum number of allowed control ports specified in advance, this synthesis flow generates a biochip architecture with high execution efficiency. Moreover, the overall cost of a CFMB can be reduced and the tradeoff between control logic and execution efficiency of biochemical applications can be evaluated for the first time. Experimental results demonstrate that MiniControl leads to high execution efficiency and low overall platform cost, while satisfying the given control port constraint strictly. Xing Huang 0001, Tsung-Yi Ho, Wenzhong Guo, Bing Li 0005, Ulf Schlichtmann |
DAC | 5 |
| 2019 | ACCESS: HW/SW Co-Equivalence Checking for Firmware OptimizationabstractCustomizing embedded computing platforms to specific application domains often necessitates optimizing the firmware and/or the HW/SW interface under tight resource constraints. Such optimizations frequently alter the communication between the firmware and the peripheral devices, possibly compromising functional correctness of the input/output behavior of the embedded system. This paper proposes a formal HW/SW co-equivalence checking technique for verifying correct I/O behavior of peripherals under a modified firmware. We demonstrate the great promise of our approach on RTL implementations of several open-source peripherals. In our experiments we successfully prove or disprove correctness of firmware optimizations for an industrial driver software. In addition, we also found a subtle bug in one of the peripherals and several undocumented preconditions for correct device behavior. Michael Schwarz 0010, Raphael Stahl, Daniel Mueller-Gritschneder, Ulf Schlichtmann, Dominik Stoffel, Wolfgang Kunz |
DAC | 4 |
| 2019 | Fault Localization in Programmable Microfluidic DevicesabstractProgrammable Microfluidic Devices (PMDs) have revolutionized the traditional biochemical experiment flow. Test algorithms for PMDs have recently been proposed. Test patterns can be generated algorithmically. But an algorithm for fault localization once some faults have been identified is not yet available. When testing a PMD, once a test pattern fails it is unknown where the stuck valve is located. The stuck valve can be any one valve out of many valves forming the test pattern. In this paper, we propose an effective algorithm for the localization of stuck-at-0 faults and stuck-at-1 faults in a PMD. The stuck valve is localized either exactly or within a very small set of candidate valves. Once the locations of faulty valves are known, it becomes possible to continue to use the PMD by resynthesizing the application. Alessandro Bernardini, Bing Li 0005, Ulf Schlichtmann |
DATE | 4 |
| 2019 | Physical Synthesis of Flow-Based Microfluidic Biochips Considering Distributed Channel StorageabstractFlow-based microfluidic biochips (FBMBs) have attracted much attention over the past decade. On such a micrometer-scale platform, various biochemical applications, also called bioas-says, can be processed concurrently and automatically. To improve execution efficiency and reduce fabrication cost, a distributed channel-storage architecture (DCSA) can be implemented on this platform, where fluid samples can be cached temporarily in flow channels close to components. Although DCSA can improve the execution efficiency of FBMBs significantly, it requires a careful arrangement of fluid samples to enable the channels to fulfill the dual functions of transportation and caching. In this paper, we formulate the first flow-layer physical design problem considering DCSA, and propose a top-down synthesis algorithm to generate efficient solutions considering execution efficiency, washing, and resource usage simultaneously. Experimental results demonstrate that the proposed algorithm leads to a shorter execution time, less flow-channel length, and a higher efficiency of on-chip resource utilization for biochemical applications compared with a direct approach to incorporate distributed storage into existing frameworks. Zhisheng Chen 0002, Xing Huang 0001, Wenzhong Guo, Bing Li 0005, Tsung-Yi Ho, Ulf Schlichtmann |
DATE | 6 |
| 2019 | Towards Reliable and Secure Post-Quantum Co-Processors based on RISC-VabstractIncreasingly complex and powerful Systems-on-Chips (SoCs), connected through a 5G network, form the basis of the Internet-of-Things (IoT). These technologies will drive the digitalization in all domains, e.g. industry automation, automotive, avionics, and healthcare. A major requirement for all above domains is the long-term (10 to 30 years) secure communication between the SoCs and the cloud over public 5G networks. The foreseeable breakthrough of quantum computers represents a risk for all communication. In order to prepare for such an event, SoCs must integrate secure quantum-computer-resistant cryptography which is reliable and protected against SW and HW attacks. Empowering SoCs with such strong security poses a challenging problem due to limited resources, tight performance requirements and long-term life-cycles. While current works are focused on efficient implementations of post-quantum cryptography, implementation-security and reliability aspects for SoCs are still largely unexplored. To this end, we present three contributions. First, we present a RISC-V co-processor for post-quantum security, able to support lattice-based cryptography. Second, we use HW/SW co-design techniques to accelerate the NTT transformation and hash generation. Third, we perform the fault analysis of the implementation. We show that our coprocessor achieves high reliability and security capabilities while preserving good performance. Tim Fritzmann, Uzair Sharif, Daniel Mueller-Gritschneder, Cezar Reinbrecht, Ulf Schlichtmann, Martha Johanna Sepúlveda |
DATE | 5 |
| 2019 | Block-Flushing: A Block-based Washing Algorithm for Programmable Microfluidic DevicesabstractProgrammable Microfluidic Devices (PMDs) have emerged as a new architecture for next-generation flow-based biochips. These devices can be dynamically reconfigured to execute different bioassays flexibly and efficiently owing to their two-dimensional regularly-arranged valve structure. During execution of a bioassay or between the execution of multiple bioassays, some areas on the PMD, however, become contaminated and must be cleaned by washing them with a buffer flow before they are reused. In this paper, we propose a novel block-based washing technique called block flushing. In this method, contaminated areas are first collected according to given patterns and flushed as a whole to increase washing efficiency. Simulation results show that with this technique the proposed method can achieve on average 28% improvement in reducing washing time compared with two other baseline solutions. Yu-Huei Lin, Tsung-Yi Ho, Bing Li 0005, Ulf Schlichtmann |
DATE | 4 |
| 2019 | SRAM Design Exploration with Integrated Application-Aware Aging AnalysisabstractOn-Chip SRAMs are an integral part of safety-critical System-on-Chips. At the same time however, they are also most susceptible to reliability threats such as Bias Temperature Instability (BTI), originating from the continuous trend of technology shrinking. BTI leads to a significant performance degradation, especially in the Sense Amplifiers (SAs) of SRAMs, where failures are fatal, since the data of a whole column is destroyed. As BTI strongly depends on the workload of an application, the aging rates of SAs in a memory array differ significantly and the incorporation of workload information into aging simulations is vital. Especially in safety-critical systems precise estimation of application specific reliability requirements to predict the memory lifetime is a key concern. In this paper we present a workload-aware aging analysis for On-Chip SRAMs that incorporates the workload of real applications executed on a processor. According to this workload, we predict the performance degradation of the SAs in the memory. We integrate this aging analysis into an aging-aware SRAM design exploration framework that generates and characterizes memories of different array granularity to select the most reliable memory architecture for the intended application. We show that this technique can mitigate SA degradation significantly depending on the environmental conditions and the application workload. Alexandra Listl, Daniel Mueller-Gritschneder, Ulf Schlichtmann, Sani R. Nassif |
DATE | 3 |
| 2019 | Aging-aware Lifetime Enhancement for Memristor-based Neuromorphic ComputingabstractMemristor-based crossbars have been applied successfully to accelerate vector-matrix computations in deep neural networks. During the training process of neural networks, the conductances of the memristors in the crossbars must be updated repetitively. However, memristors can only be programmed reliably for a given number of times. Afterwards, the working ranges of the memristors deviate from the fresh state. As a result, the weights of the corresponding neural networks cannot be implemented correctly and the classification accuracy drops significantly. This phenomenon is called aging, and it limits the lifetime of memristor-based crossbars. In this paper, we propose a co-optimization framework combining software training and hardware mapping to reduce the aging effect. Experimental results demonstrate that the proposed framework can extend the lifetime of such crossbars up to 11 times, while the expected accuracy of classification is maintained. Shuhang Zhang, Grace Li Zhang, Bing Li 0005, Hai Li 0001, Ulf Schlichtmann |
DATE | 5 |
| 2019 | VOM: Flow-Path Validation and Control-Sequence Optimization for Multilayered Continuous-Flow Microfluidic BiochipsabstractMultilayered valve-based continuous-flow microfluidic biochips are a rapidly developing platform for delicate bio-applications. Due to the high complexity of the biochip structure and the application protocols, there is an increasing demand for design automation approaches. Current research has enabled automated generation of biochip physical designs, operation scheduling, and binding protocols, which has demonstrated the potential for better resource utilization and execution time reduction. However, the state-of-the-art high-level synthesis methods are on operation- and device-level. They assume fluid transportation paths to be always available but overlook the physical layout of the control and flow channels. This mismatch leads to a gap in the complete synthesis flow, and can result in performance drop, waste of resources due to redundancy or even infeasible designs. This work proposes to bridge this gap with a simulation-based approach, which takes a biochip design and a high-level protocol as inputs, and synthesizes channel-level pressurization protocols to support dynamic construction of valid fluid transportation paths. Experimental results show that the proposed method can efficiently validate and optimize the flow paths for feasible designs and protocols, detect redundant resource usage, and locate the conflicts for infeasible designs and protocols. It opens up a new direction to improve the performance and the feasibility of customized biochip synthesis. Mengchu Li, Tsun-Ming Tseng, Yanlu Ma, Tsung-Yi Ho, Ulf Schlichtmann |
ICCAD | 5 |
| 2019 | Overview of 2019 CAD Contest at ICCADabstractThe “CAD Contest at ICCAD” is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has published many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The contest keeps enhancing impact and boosting EDA research. Ulf Schlichtmann, Sabya Das, Ing-Chao Lin, Mark Po-Hung Lin |
ICCAD | 1 |
| 2019 | Cloud Columba: Accessible Design Automation Platform for Production and Inspiration: Invited PaperabstractDesign automation for continuous-flow microfluidic large-scale integration (mLSI) biochips has made remarkable progress over the past few years. Nowadays a biochip containing up to hundreds of components can be automatically synthesized within a few minutes. However, the current advanced design automation tools are mostly developed for research use, which focus essentially on the algorithmic performance but overlook the accessibility. Therefore, we have started the Cloud Columba project since 2017 to provide users from different backgrounds with easy access to the state-of-the-art design automation approaches. Without being limited by the computing power of their end devices, users just need to formulate their design requests in a high abstraction level, based on which the cloud server will automatically synthesize a customized manufacturing-ready biochip design, which can be viewed and stored using simply a web browser. With the computer-synthesized designs, Cloud Columba supports application developers to explore a wider range of possibilities, and algorithm developers to validate and improve their ideas based on a practical foundation. Tsun-Ming Tseng, Mengchu Li, Yushen Zhang, Tsung-Yi Ho, Ulf Schlichtmann |
ICCAD | 5 |
| 2019 | Wavelength-Routed Optical NoCs: Design and EDA - State of the Art and Future Directions: Invited PaperabstractWavelength-routed optical network-on-chip (WRONoC) design consists of topological and physical synthesis. It covers many interacting design aspects such as wavelength assignment, message routing, network construction, component placement, and waveguide routing. Due to the high complexity of the design problem, current manual design usually trades optimality for scalability and feasibility, which results in performance degradation and waste of resources. In this paper, we will present an overview of the existing design automation approaches that have demonstrated their effectiveness in customizing and optimizing application-specific WRONoC designs, and of the potential design automation directions to address a wider range of design challenges. We will also discuss the advantages of comprehensive optimization considering multiple design aspects simultaneously, and the possible barriers that need to be removed to achieve this goal. Tsun-Ming Tseng, Alexandre Truppel, Mengchu Li, Mahdi Nikdast, Ulf Schlichtmann |
ICCAD | 5 |
| 2019 | PSION: Combining Logical Topology and Physical Layout Optimization for Wavelength-Routed ONoCsabstractOptical Networks-on-Chip (ONoCs) are a promising solution for high-performance multi-core integration with better latency and bandwidth than traditional Electrical NoCs. Wavelength-routed ONoCs (WRONoCs) offer yet additional performance guarantees. However, WRONoC design presents new EDA challenges which have not yet been fully addressed. So far, most topology analysis is abstract, i.e., overlooks layout concerns, while for layout the tools available perform Place & Route (P&R) but no topology optimization. Thus, a need arises for a novel optimization method combining both aspects of WRONoC design. In this paper such a method, PSION, is laid out. When compared to the state-of-the-art design procedure, results show a 1.8x reduction in maximum optical insertion loss. Alexandre Truppel, Tsun-Ming Tseng, Davide Bertozzi, José Carlos Alves, Ulf Schlichtmann |
ISPD | 5 |
| 2019 | Synthesis of a Cyberphysical Hybrid Microfluidic Platform for Single-Cell AnalysisabstractSingle-cell genomics is used to advance our understanding of diseases, such as cancer. Microfluidic solutions have recently been developed to classify cell types or perform single-cell biochemical analysis on preisolated types of cells. However, new techniques are needed to efficiently classify cells and conduct biochemical experiments on multiple cell types concurrently. Nondeterministic cell-type identification, system integration, and design automation are major challenges in this context. To overcome these challenges, we present a hybrid microfluidic platform that enables complete single-cell analysis on a heterogeneous pool of cells. We combine this architecture with an associated design-automation and optimization framework, referred to as co-synthesis (CoSyn). The proposed framework employs real-time resource allocation to coordinate the progression of concurrent cell analysis. Besides this framework, a probabilistic model based on a discrete-time Markov chain is also deployed to investigate protocol settings, where experimental conditions, such as sonication time, vary probabilistically among cell types. Simulation results show that CoSyn efficiently utilizes platform resources and outperforms baseline techniques. Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Synthesis of Reconfigurable Flow-Based Biochips for Scalable Single-Cell ScreeningabstractSingle-cell screening is used to sort a stream of cells into clusters (or types) based on prespecified biomarkers, thus supporting type-driven biochemical analysis. Reconfigurable flow-based microfluidic biochips (RFBs) can be utilized to screen hundreds of heterogeneous cells within a few minutes, but they are overburdened with the control of a large number of valves. To address this problem, we present a pin-constrained RFB design methodology for single-cell screening. The proposed design is analyzed using computational fluid dynamics simulations, mapped to an RC-lumped model, and combined with intervalve connectivity information to construct a high-level synthesis framework, referred to as cell sorter using multiplexed control (Sortex). Simulation results show that Sortex significantly reduces the number of control pins and fulfills the timing requirements of single-cell screening. Mohamed Ibrahim 0002, Aditya Sridhar, Krishnendu Chakrabarty, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | EffiTest2: Efficient Delay Test and Prediction for Post-Silicon Clock Skew Configuration Under Process VariationsabstractAt nanometer manufacturing technology nodes, process variations affect circuit performance significantly. This trend leads to a large timing margin and thus overdesign in the traditional worst-case circuit design flow. To combat this pessimism, post-silicon clock tuning buffers can be deployed to balance timing slacks of consecutive combinational paths in individual chips by tuning clock skews after manufacturing. A challenge of this method is that path delays of each chip with timing failures should be measured to gather the information for clock skew configuration. However, current methods for delay measurement rely on path-wise frequency stepping, which requires much time from expensive testers. In this paper, we propose an efficient delay test framework (EffiTest2) to solve the post-silicon testing problem by testing only representative paths with delay alignment using the already-existing tunable buffers in the circuit. Experimental results demonstrate that EffiTest2 can reduce the number of frequency stepping iterations by more than 94% with only a slight yield loss. Grace Li Zhang, Bing Li 0005, Yiyu Shi 0001, Jiang Hu 0001, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | On enabling diagnosis for 1-Pin Test fails in an industrial flowabstractThe 1-Pin Test concept has proven to be beneficial for test cost reduction. By compacting test responses into a signature and reading them out at test end, test parallelism can be increased significantly. This reduces the test time and thus test cost. Especially cost-sensitive devices, e.g. IoT end nodes, profit. A drawback of this method is the limited capability of diagnosis due to the lack of cycle-accurate PASS/FAIL information. In this paper, we present a new approach to tackle this challenge. It enables the use of an industrial diagnosis flow for fails that occurred during 1-Pin Test. For this purpose, we propose failing vector and failing cycle analysis techniques. Our approach is fault model independent and not limited to a single fault assumption. We mitigate the aliasing problem by masking. The effectiveness of our approach is shown on an investigation of real silicon fails in industrial designs. Daniel Tille, Benedikt Gottinger, Ulrike Pfannkuchen, Helmut E. Graeb, Ulf Schlichtmann |
ASP-DAC | 5 |
| 2018 | PlanarONoC: concurrent placement and routing considering crossing minimization for optical networks-on-chipabstractOptical networks-on-chips (ONoCs) have become a promising solution for the on-chip communication of multi-and many-core systems to provide superior communication bandwidths, efficiency in power consumption, and latency performance compared to electronic NoCs. Serving as the critical part of ONoCs, an optical router composed of waveguides and photonic switching elements (PSEs) routes signals between two hubs or between a hub and a memory controller. Many studies focus on developing efficient architectures of optical routers, while their physical implementation that can seriously deteriorate the quality of the architectures is rarely addressed. The existing automatic place-and-route tools suffer from considerable insertion loss due to many waveguide crossings outside of PSEs, which leads to huge power consumption of laser sources. By observing that the logic schemes of most optical routers are actually planar, we develop a concurrent PSE placement and waveguide routing flow, called PlanarONoC, that guarantees optimal solutions in terms of crossings for planar logic schemes. Experimental results show that the proposed flow reduces the maximum insertion loss by 37% on average, guarantees no waveguide crossing outside of PSEs, and performs much more efficient compared to the state-of-the-art work. Yu-Kai Chuang, Kuan-Jung Chen, Kun-Lin Lin, Shao-Yun Fang, Bing Li 0005, Ulf Schlichtmann |
DAC | 6 |
| 2018 | Design-for-testability for continuous-flow microfluidic biochipsabstractFlow-based microfluidic biochips are gaining traction in the microfluidics community since they enable efficient and low-cost biochemical experiments. These highly integrated lab-on-a-chip systems, however, suffer from manufacturing defects, which cause some chips to malfunction. To test biochips after manufacturing, air pressure is applied to input ports of a chip and predetermined test vectors are used to change the states of microvalves in the chip. Pressure meters are connected to the output ports to measure pressure values, which are compared with expected values to detect errors. To reduce the cost of the test platform, the number of pressure sources and meters should be reduced. We propose a design-for-testability (DFT) technique that enables a test procedure with only a single pressure source and a single pressure meter. Furthermore, the valves inserted for DFT share control channels with valves in the original chip so that no additional control signals are required. Simulation results demonstrate that this technique can generate efficient chip architectures for single-source single-meter test in all experiment cases successfully to reduce test cost, while the performance of these chips in executing applications is still maintained. Bing Li 0005, Tsung-Yi Ho, Krishnendu Chakrabarty, Ulf Schlichtmann |
DAC | 5 |
| 2018 | Columba S: a scalable co-layout design automation tool for microfluidic large-scale integrationabstractMicrofluidic large-scale integration (mLSI) is a promising platform for high-throughput biological applications. Design automation for mLSI has made much progress in recent years. Columba and its succeeding work Columba 2.0 proposed a mathematical modeling method that enables automatic design of manufacturing-ready chips within minutes. However, current approaches suffer from a huge computation load when the designs become larger. Thus, in this work, we propose Columba S with a focus on scalability. Columba S applies a new architectural framework and a straight channel routing discipline, and synthesizes multiplexers for efficient and reconfigurable valve control. Experiments show that Columba S is able to generate mLSI designs with more than 200 functional units within three minutes, which enables the design of a platform for large and complex applications. Tsun-Ming Tseng, Mengchu Li, Daniel Nestor Freitas, Amy Mongersun, Ismail Emre Araci, Tsung-Yi Ho, Ulf Schlichtmann |
DAC | 7 |
| 2018 | Virtualsync: timing optimization by synchronizing logic waves with sequential and combinational components as delay unitsabstractIn digital circuit designs, sequential components such as flip-flops are used to synchronize signal propagations. Logic computations are aligned at and thus isolated by flip-flop stages. Although this fully synchronous style can reduce design efforts significantly, it may affect circuit performance negatively, because sequential components can only introduce delays into signal propagations instead of accelerating them. In this paper, we propose a new timing model, VirtualSync, in which signals, specially those along critical paths, are allowed to propagate through several sequential stages without flip-flops. Timing constraints are still satisfied at the boundary of the optimized circuit to maintain a consistent interface with existing designs. By removing clock-to-q delays and setup time requirements of lip-lops on critical paths, the performance of a circuit can be pushed even beyond the limit of traditional sequential designs. Experimental results demonstrate that circuit performance can be improved by up to 11.5% (average 3.1%) compared with that after thorough sizing and retiming, while the increase of area is still negligible. Grace Li Zhang, Bing Li 0005, Masanori Hashimoto, Ulf Schlichtmann |
DAC | 4 |
| 2018 | Fault-tolerant valve-based microfluidic routing fabric for droplet barcoding in single-cell analysisabstractHigh-throughput single-cell genomics is used to gain insights into diseases such as cancer. Motivated by this important application, microfluidics has emerged as a key technology for developing comprehensive biochemical procedures for studying DNA, RNA, proteins, and many other cellular components. Recently, a hybrid microfluidic platform has been proposed to efficiently automate the analysis of a heterogeneous sequence of cells. In this design, a valve-based routing fabric based on transposers is used to label/barcode the target cells. However, the design proposed in prior work overlooked defects that are likely to occur during chip fabrication and system integration. We address the above limitation by investigating the fault tolerance of the valve-based routing fabric. We develop a theory of failure assessment and introduce a design technique for achieving fault tolerance. Simulation results show that the proposed method leads to a slight increase in the fabric size and decrease in cell-analysis throughput, but this is only a small price to pay for the added assurance of fault tolerance in the new design. Yasamin Moradi, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ulf Schlichtmann |
DATE | 4 |
| 2018 | ETISS-ML: A multi-level instruction set simulator with RTL-level fault injection support for the evaluation of cross-layer resiliency techniquesabstractETISS is an instruction set simulator (ISS) for Virtual Prototypes (VPs) modeled with SystemC/TLM. In this paper, we propose the extension ETISS-ML, which enables a multi-level simulation that switches between ISS-level and register transfer level (RTL) to accurately evaluate the impact of soft errors in the pipeline of a RISC processor. ETISS-ML achieves close-to-RTL-accurate fault injection simulation results with close-to-ISS simulation performance with a speed up gain up to 100x compared to RTL. For this, we propose an approach to dynamically determine the length of the RTL simulation period. The high simulation performance of ETISS-ML enables an ultra-efficient and accurate evaluation of cross-layer resiliency techniques for embedded applications, which requires running a large number of fault injections for long simulation scenarios. This is demonstrated on a case study of a Microcontroller Unit (MCU) executing a control algorithm for adaptive cruise control. Daniel Mueller-Gritschneder, Martin Dittrich, Josef Weinzierl, Eric Cheng, Subhasish Mitra, Ulf Schlichtmann |
DATE | 6 |
| 2018 | TimingCamouflage: Improving circuit security against counterfeiting by unconventional timingabstractWith recent advances in reverse engineering, attackers can reconstruct a netlist to counterfeit chips by opening the die and scanning all layers of original chips. This relatively easy counterfeiting is made possible by the use of the standard simple clocking scheme where all combinational blocks function within one clock period. In this paper, we propose a method to invalidate the assumption that a netlist completely represents the function of a circuit. With the help of wave-pipelining paths, this method forces attackers to capture delay information from manufactured chips, which is a very challenging task because we also introduce false paths. Experimental results confirm that wave-pipelining paths and false paths can be constructed in benchmark circuits successfully with only a negligible cost, while the potential attack techniques can be thwarted. Grace Li Zhang, Bing Li 0005, Bei Yu 0001, David Z. Pan, Ulf Schlichtmann |
DATE | 5 |
| 2018 | An efficient fault-tolerant valve-based microfluidic routing fabric for single-cell analysisabstractSingle-cell analysis is used to gain insights into diseases such as cancer. Recently, a hybrid microfluidic platform was proposed for concurrent single-cell analysis on thousands of heterogeneous cells. In this design, barcoding droplets are routed using a valve-based routing fabric to label the input cells. The fault-tolerance of this routing fabric has also been studied and a design technique for implementing a fault-tolerant crossbar has been proposed. However, prior work leads to a significant increase in fabric size and a decrease in cell-analysis performance. We address the above drawbacks and introduce a low-overhead design technique for achieving fault-tolerance, while maintaining the efficiency of the cell-analysis platform. We show that the proposed method is optimal in that it minimizes the overhead in terms of fabric size. We also show that the new design outperforms the previous solution in terms of cell-analysis performance. Yasamin Moradi, Krishnendu Chakrabarty, Ulf Schlichtmann |
ETS | 3 |
| 2018 | Automated Redirection of Hardware Accesses for Host-Compiled Software SimulationabstractFor host-compiled software simulation it is required that accesses from the target software to memory-mapped hardware are identified, so that they can be redirected to a virtual prototype. This is straight-forward if the software uses a hardware abstraction layer as interface. If such an interface is not used by existing or third-party source code, the rewriting of the code for host-compiled simulation involves a considerable manual effort. In this paper, we present a method to automate this process with the help of a symbolic execution engine. With our approach the time to adjust software for host-compilation is significantly reduced. We show that the most memory-mapped hardware accesses are correctly rewritten in a real-world application by comparing recorded access traces on a virtual prototype. Additionally, a test suite has been developed to cover edge-cases. Rafael Stahl, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
FDL | 3 |
| 2018 | Automatic Design of Microfluidic DevicesabstractThis overview paper summarizes the content of a tutorial given at the 2018 edition of the Forum on specification & Design Languages. The aim of the tutorial was to introduce the technology of microfluidic devices, which gained significant interest in the recent past, as well as corresponding design challenges to a community focused on design automation and corresponding specification/design languages. By this, the overview presents a starting point for researchers and engineers interested in getting involved in this area. Robert Wille, Bing Li 0005, Rolf Drechsler, Ulf Schlichtmann |
FDL | 4 |
| 2018 | Wavefront-MCTS: multi-objective design space exploration of NoC architectures based on Monte Carlo tree searchabstractApplication-specific MPSoCs profit immensely from a custom-fit Network-on-Chip (NoC) architecture in terms of network performance and power consumption. In this paper we suggest a new approach to explore application-specific NoC architectures. In contrast to other heuristics, our approach uses a set of network modifications defined with graph rewriting rules to model the design space exploration as a Markov Decision Process (MDP). The MDP can be efficiently explored using the Monte Carlo Tree Search (MCTS) heuristics. We formulate a weighted sum reward function to compute a single solution with a good trade-off between power and latency or a set of max reward functions to compute the complete Pareto front between the two objectives. The Wavefront feature adds additional efficiency when computing the Pareto front by exchanging solutions between parallel MCTS optimization processes. Comparison with other popular search heuristics demonstrates a higher efficiency of MCTS-based heuristics for several test cases. Additionally, the Wavefront-MCTS heuristics allows complete tracability and control by the designer to enable an interactive design space exploration process. Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ICCAD | 3 |
| 2018 | CustomTopo: a topology generation method for application-specific wavelength-routed optical NoCsabstractOptical network-on-chip (NoC) is a promising platform beyond electronic NoCs. In particular, wavelength-routed optical network-on-chip (WRONoC) is renowned for its high bandwidth and ultra-low signal delay. Current WRONoC topology generation approaches focus on full-connectivity, i.e. all masters are connected to all slaves. This assumption leads to wasted resources for application-specific designs. In this work, we propose CustomTopo: a general solution to the topology generation problem on WRONoCs that supports customized connectivity. CustomTopo models the topology structure and its communication behavior as an integer-linear-programming (ILP) problem, with an adjustable optimization target considering the number of add-drop filters (ADFs), the number of wavelengths, and insertion loss. The time for solving the ILP problem in general positively correlates with the network communication densities. Experimental results show that CustomTopo is applicable for various communication requirements, and the resulting customized topology enables a remarkable reduction in both resource usage and insertion loss. Mengchu Li, Tsun-Ming Tseng, Davide Bertozzi, Mahdi Tala, Ulf Schlichtmann |
ICCAD | 5 |
| 2018 | Performance and accuracy in soft-error resilience evaluation using the multi-level processor simulator ETISS-MLabstractSoft errors are a major safety concern in many devices, e.g., in automotive, industrial, control or medical applications. Ideally, safety-critical systems should be resilient against the impact of soft errors, but at a low cost. This requires to evaluate the soft error resilience, which is typically done by extensive fault injection. In this paper, we present ETISS-ML, a multi-level processor simulator, which manages to achieve both accuracy and performance for fault simulation by intelligently switching the level of abstraction between an Instruction Set Simulator (ISS) and an RTL simulator. For a given software testcase and fault scenario, the software is first executed in ISS-mode until shortly before the fault injection. Then ETISS-ML switches to RTL-mode for accurate fault simulation. Whenever the impact of the fault is propagated completely out of the processor's micro-architecture, the simulation can switch back to ISS-mode. This paper describes the methods needed to preserve accuracy during both of these switches. Experimental results show that ETISS-ML obtains near to ISS performance with RTL accuracy. It is also shown that ETISS-ML can be used as the processor model in SystemC / TLM virtual prototypes (VPs) and, hence, allows to investigate the impact of soft errors at system level. Daniel Mueller-Gritschneder, Uzair Sharif, Ulf Schlichtmann |
ICCAD | 3 |
| 2018 | Multi-channel and fault-tolerant control multiplexing for flow-based microfluidic biochipsabstractContinuous flow-based biochips are one of the promising platforms used in biochemical and pharmaceutical laboratories due to their efficiency and low costs. Inside such a chip, fluid volumes of nanoliter size are transported between devices for various operations, such as mixing and detection. The transportation channels and corresponding operation devices are controlled by microvalves driven by external pressure sources. Since assigning an independent pressure source to every microvalve would be impractical due to high costs and limited system dimensions, states of microvalves are switched using a control logic by time multiplexing. Existing control logic designs, however, still switch only a single control channel per operation – leading to a low efficiency. In this paper, we propose the first automatic synthesis approach for a control logic that is able to switch multiple control channels simultaneously to reduce the overall switching time of valve states. In addition, we propose the first fault-aware design in control logic to introduce redundant control paths to maintain the correct function even when manufacturing defects occur. Compared with the existing direct connection method, the proposed multi-channel switching mechanism can reduce the switching time of valve states by up to 64%. In addition, all control paths for fault tolerance have been realized. Ying Zhu 0008, Bing Li 0005, Tsung-Yi Ho, Qin Wang 0005, Hailong Yao 0002, Robert Wille, Ulf Schlichtmann |
ICCAD | 7 |
| 2018 | Emulation of an ASIC Power, Temperature and Aging Monitor System for FPGA PrototypingabstractTechnology scaling has enabled the fabrication of Multi-Processor Systems-on-Chips (MPSoCs), which satisfy the ever growing demand for performance, while continuously reducing the chip size. Thus, scaling has also led to new challenges such as increasing power densities, which critically influence the chip temperatures and accelerate device degradation due to aging. Runtime power management can be utilized to counter these reliability threats to increase the lifetime of a system. For the development of runtime power management strategies monitoring data for power, temperature and aging is required. In this paper we propose a real-time power, temperature and aging monitor system (eTAPMon) for FPGA prototypes of MPSoCs. The monitor system emulates data characterized from the target ASIC design. The emulation approach models the behavior of ASIC power monitors based on an instruction-level energy model, temperature monitors based on a linear regression model obtained from thermal offline simulations and aging monitors based on a critical path model to compute the decreasing timing margin due to aging. An accelerated aging emulation is possible to predict aged ASIC behavior. Hence, this FPGA emulation enables the early evaluation of runtime power management strategies. Alexandra Listl, Daniel Mueller-Gritschneder, Fabian Kluge, Ulf Schlichtmann |
IOLTS | 4 |
| 2018 | Efficient Fault Injection for Embedded Systems: As Fast as Possible but as Accurate as NecessaryabstractWhen used for safety-critical applications, embedded systems must behave safely at all times - even in the presence of random hardware faults. To ensure this, fault effect simulation by simulation-based fault injection is an integral part of embedded system development. The high complexity of embedded systems results in low simulation performance if all details of the system are simulated. Not simulating all details, i.e. increasing the simulation abstraction level, speeds up fault injection but can result in less accuracy in predicting the fault impacts on the system behavior. To achieve high accuracy and high simulation performance at the same time, we avoid simulation of details unrelated to the injected fault. For this, we divide the set of faults that can occur in an embedded system into three subsets. For each subset, we select the fault injection abstraction level of the embedded processor model that is as accurate as necessary but as fast as possible. The considered levels are host-compiled simulation, instruction set simulation and register transfer level simulation. For additional speed-up, the abstraction level can be switched during the fault injection simulation between register transfer and instruction set level. The fault set for host-compiled simulation can be reduced by static program analysis. Our results show that adapting the abstraction level to the fault set achieves high performance of the fault injection simulation. Petra R. Maier, Uzair Sharif, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
IOLTS | 4 |
| 2018 | Thermal-Aware Placement and Routing for 3D Optical Networks-on-ChipsabstractMany-core chip architectures integrate tens to hundreds of processor cores on a single chip. Recent development of photonic interconnects has made Optical Networks-on-Chips (ONoCs) an attractive technology to overcome the drawbacks of electrical networks-on-chips. With ultra-high bandwidth, low latency, and great energy efficiency, ONoCs enable the designer to build scalable systems. However, photonic devices are sensitive to temperature fluctuations, and hence, require proactive management. This paper first calculates the thermal distribution from cell distribution using an approximated Green's function and proposes a post-placement algorithm to reduce the number of photonic devices in the hotspots. The paper then improves the routing algorithm considering bending loss and temperature variations. Experimental results also verify the efficiency and effectiveness of our algorithm. Fengxian Jiao, Sheqin Dong, Bei Yu 0001, Bing Li 0005, Ulf Schlichtmann |
ISCAS | 5 |
| 2018 | Columba 2.0: A Co-Layout Synthesis Tool for Continuous-Flow Microfluidic BiochipsabstractContinuous-flow microfluidic large-scale integration (mLSI) shows increasing importance in biological/chemical fields, thanks to its advantages in miniaturization and high throughput. Current mLSI is designed manually, which is time-consuming and error-prone. In recent years, design automation research for mLSI has evolved rapidly, aiming to replace manual labor by computers. However, previous design automation approaches used to design each microfluidic layer separately and over-simplify the layer interactions to various degrees, which resulted in a gap between realistic requirements and automatically generated designs. In this paper, we propose a module model library to accurately model microfluidic components involving layer interactions; and we propose a co-layout synthesis tool, Columba, which generates AutoCAD-compatible designs that fulfill all designs rules and can be directly used for mask fabrication. Columba takes plain-text netlist descriptions as inputs, and performs simultaneous placement and routing for multiple layers while ensuring the planarity of each layer. We validate Columba by fabricating two of its output designs. Columba is the first design automation tool that can seamlessly synchronize with the manufacturing flow. Tsun-Ming Tseng, Mengchu Li, Daniel Nestor Freitas, Travis McAuley, Bing Li 0005, Tsung-Yi Ho, Ismail Emre Araci, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2018 | Design-Phase Buffer Allocation for Post-Silicon Clock Binning by Iterative LearningabstractAt submicrometer manufacturing technology nodes, process variations affect circuit performance significantly. To counter these variations, engineers are reserving more timing margin to maintain yield, leading to an unaffordable overdesign. Most of these margins, however, are wasted after manufacturing, because process variations cause only some chips to be really slow, while other chips can easily meet given timing specifications. To reduce this pessimism, we can reserve less timing margin and tune failed chips after manufacturing with clock buffers to make them meet timing specifications. With this post-silicon clock tuning, critical paths can be balanced with neighboring paths in each chip specifically to counter the effect of process variations. Consequently, chips with timing failures can be rescued and the yield can thus be improved. This is specially useful in high-performance designs, e.g., high-end CPUs, where clock binning makes chips with higher performance much more profitable. In this paper, we propose a method to determine where to insert post-silicon tuning buffers during the design phase to improve the overall profit with clock binning. This method learns the buffer locations with a Sobol sequence iteratively and reduces the buffer ranges afterward with tuning concentration and buffer grouping. Experimental results demonstrate that the proposed method can achieve a profit improvement of about 14% on average and up to 26%, with only a small number of tuning buffers inserted into the circuit. Grace Li Zhang, Bing Li 0005, Jinglan Liu, Yiyu Shi 0001, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | Fault Injection for Test-Driven Development of Robust SoC FirmwareabstractRobustness against errors in hardware must be considered from the very beginning of safety-critical system-on-chip firmware design. Therefore, we present fault injection for test-driven development (TDD) of robust firmware. As TDD is based on instant feedback to the designer, fault injection must execute within few minutes. In contrast to state-of-the-art approaches, we avoid long simulation scenarios and runtimes by injecting faults at the unit level and utilizing host-compiled simulation. Further, three static bit-level analyses of firmware source code and hardware specification reduce the fault set significantly. This accelerates fault injection by several orders of magnitude and enables robustness-aware TDD. Petra R. Maier, Veit Kleeberger, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | Graph-Grammar-Based IP-Integration (GRIP) - An EDA Tool for Software-Defined SoCsabstractIn modern system-on-chip (SoC) designs, IP-reuse is considered a driving force to increase productivity. To support various designs, a huge amount of Intellectual Property (IP) hardware blocks have been developed. The integration of those IPs into an SoC may require significant effort—up to days or weeks depending on experience and complexity. This article presents a novel approach to significantly reduce the design effort to bring-up a working SoC design by automatic IP integration as part of a library-based Software-defined SoC flow. In detail, the IP-supplier prepares a HW-accelerated software library (HASL) for the SoC architect, who wants to use the IP in an SoC design. As a key point of our approach, integration knowledge is encoded in the library as a set of integration rules. These rules are defined in the machine-readable standardized IP-XACT format by the IP supplier, who has a good knowledge of the IP’s hardware details. The library preparation step on the IP supplier’s side is also partly automated in the proposed flow, including a partial generation of configurable HW drivers, schedulers, and the software library functions. For the SoC architect, we have developed the graph-grammar-based IP-integration (GRIP) tool. The software application is developed using the functions supplied in the HASL. According to the calls to the HASL functions, the GRIP tool automatically integrates IP-blocks using the rule information supplied with the library and runs a full Design Space Exploration. For this, the SoC architecture and rules are transformed into the graph domain to apply graph rewriting methods. The GRIP tool is model-driven and based on the Eclipse Modeling Framework. With code generation techniques, SoC candidate architectures can be transformed to hardware descriptions for the target platform. The HW/SW interfaces between SW library functions and IP blocks can be automatically generated for bare-metal or Linux-based applications. The approach is demonstrated with two case-studies on the Xilinx Zynq-based ZedBoard evaluation board using a HASL for computer vision. It can yield 10×-150× performance improvement for the bare-metal application versions and 4×--7× performance improvement for the Linux-based application versions, when executed on an optimized HW-accelerated SoC architecture compared to a non HW-accelerated SoC. The effort for IP integration is comparable to using a software library, hence, providing a significant advantage over a manual IP integration. Munish Jassi, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2018 | Automated Phase-Noise-Aware Design of RF Clock Distribution Circuits
Dimo Martev, Sven Hampel, Ulf Schlichtmann |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Hamming-distance-based valve-switching optimization for control-layer multiplexing in flow-based microfluidic biochipsabstractFlow-based microfluidic biochips have progressed significantly in the past decade. Thanks to innovations in multilayer soft lithography (MSL) fabrication technology, the integration of thousands of microvalves along with large-scale networks of microchannels on a chip has been enabled. This progress has even been compared to the evolution of VLSI circuits following Moore's Law. In flow-based microfluidic biochips, microvalves are critical components to control the fluidic transportation for complex operations. To activate the open/close states of a microvalve, off-chip control pins are required. Due to the tremendous increase of the number of microvalves, a software-programmable microfluidic platform has been proposed to reduce the number of off-chip control pins, which integrates a microfluidic multiplexer on a separate control layer to control the array of microvalves. The multiplexer needs to be switched when the states of microvalves are changed between every two adjacent time slots. High switching frequency will make the multiplexer vulnerable and decrease the chip's reliability. We observe that different switching orders of microvalves lead to different switching frequencies of a multiplexer. Based on this observation, this paper proposes the first Hamming-distance-based switching order optimization method for microvalves to enhance the reliability of the multiplexer. Experimental results show that our method can significantly reduce the switching frequency of multiplexer, and the solution is very close to the theoretical optimal lower bound. Qin Wang 0005, Shiliang Zuo, Hailong Yao 0002, Tsung-Yi Ho, Bing Li 0005, Ulf Schlichtmann, Yici Cai |
ASP-DAC | 6 |
| 2017 | Component-Oriented High-level Synthesis for Continuous-Flow Microfluidics Considering Hybrid-SchedulingabstractTechnological innovations in continuous-flow microfluidics require updated automated synthesis methods. As new microfluidic components and biochemical applications are constantly introduced, the current functionality-based application mapping methods and the fixed-time-slot scheduling methods are insufficient to solve the new design challenges. In this work, we propose a component-oriented general device concept that enables precise description of operations and devices, and adapts well to technological updates. Applying this concept, we propose a layering algorithm together with a mathematical modeling method to synthesize binding and hybrid-scheduling solutions that support both fixed schedule and real-time decisions. We also consider potential chip layout and optimize the number of flow channels among devices to save routing efforts. Experimental results demonstrate that our solution fully utilizes the chip resources and can handle operations with different requirements. Mengchu Li, Tsun-Ming Tseng, Bing Li 0005, Tsung-Yi Ho, Ulf Schlichtmann |
DAC | 5 |
| 2017 | Transport or Store?: Synthesizing Flow-based Microfluidic Biochips using Distributed Channel StorageabstractFlow-based microfluidic biochips have attracted much attention in the EDA community due to their miniaturized size and execution efficiency. Previous research, however, still follows the traditional computing model with a dedicated storage unit, which actually becomes a bottleneck of the performance of biochips. In this paper, we propose the first architectural synthesis framework considering distributed storage constructed temporarily from transportation channels to cache fluid samples. Since distributed storage can be accessed more efficiently than a dedicated storage unit and channels can switch between the roles of transportation and storage easily, biochips with this distributed computing architecture can achieve a higher execution efficiency even with fewer resources. Experimental results confirm that the execution efficiency of a bioassay can be improved by up to 28% while the number of valves in the biochip can be reduced effectively. Bing Li 0005, Hailong Yao 0002, Paul Pop, Tsung-Yi Ho, Ulf Schlichtmann |
DAC | 6 |
| 2017 | CoSyn: Efficient single-cell analysis using a hybrid microfluidic platformabstractSingle-cell genomics is used to advance our understanding of diseases such as cancer. Microfluidic solutions have recently been developed to classify cell types or perform single-cell biochemical analysis on pre-isolated types of cells. However, new techniques are needed to efficiently classify cells and conduct biochemical experiments on multiple cell types concurrently. System integration and design automation are major challenges in this context. To overcome these challenges, we present a hybrid microfluidic platform that enables complete single-cell analysis on a heterogeneous pool of cells. We combine this architecture with an associated design-automation and optimization framework, referred to as Co-Synthesis (CoSyn). The proposed framework employs real-time resource allocation to coordinate the progression of concurrent cell analysis. Simulation results show that CoSyn efficiently utilizes platform resources and outperforms baseline techniques. Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ulf Schlichtmann |
DATE | 3 |
| 2017 | Testing microfluidic Fully Programmable Valve Arrays (FPVAs)abstractFully Programmable Valve Array (FPVA) has emerged as a new architecture for the next-generation flow-based microfluidic biochips. This 2D-array consists of regularly-arranged valves, which can be dynamically configured by users to realize microfluidic devices of different shapes and sizes as well as interconnections. Additionally, the regularity of the underlying structure renders FPVAs easier to integrate on a tiny chip. However, these arrays may suffer from various manufacturing defects such as blockage and leakage in control and flow channels. Unfortunately, no efficient method is yet known for testing such a general-purpose architecture. In this paper, we present a novel formulation using the concept of flow paths and cut-sets, and describe an ILP-based hierarchical strategy for generating compact test sets that can detect multiple faults in FPVAs. Simulation results demonstrate the efficacy of the proposed method in detecting manufacturing faults with only a small number of test vectors. Bing Li 0005, Bhargab B. Bhattacharya, Krishnendu Chakrabarty, Tsung-Yi Ho, Ulf Schlichtmann |
DATE | 6 |
| 2017 | A Method for Phase Noise Analysis of RF CircuitsabstractIn this paper we present a method for analysis of phase noise in logic circuits. This method allows the design and verification of phase noise critical circuits using a digital toolchain, significantly reducing the design time and effort compared to the traditional approach using analog tools such as SPICE simulation. It is based on a set of pre-characterized standard cells and the generated phase noise is estimated using a lookup table approach. Comparison of the estimation results with back-annotated analog simulations in 28 nm CMOS technology show that the error of the estimation is within 7.2 % of the actual phase noise, and the runtime is reduced by three orders of magnitude. Dimo Martev, Sven Hampel, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | Sortex: Efficient timing-driven synthesis of reconfigurable flow-based biochips for scalable single-cell screeningabstractSingle-cell screening is used to sort a stream of cells into clusters (or types) based on pre-specified biomarkers, thus supporting type-driven biochemical analysis. Reconfigurable flow-based microfluidic biochips (RFBs) can be utilized to screen hundreds of heterogeneous cells within a few minutes, but they are overburdened with the control of a large number of valves. To address this problem, we present a pin-constrained RFB design methodology for single-cell screening. The proposed design is analyzed using computational fluid dynamics simulations, mapped to an RC-lumped model, and combined with a high-level synthesis framework, referred to as Sortex. Simulation results show that Sortex significantly reduces the number of control pins and fulfills the timing requirements of single-cell screening. Mohamed Ibrahim 0002, Aditya Sridhar, Krishnendu Chakrabarty, Ulf Schlichtmann |
ICCAD | 4 |
| 2017 | Methodology for automated phase noise minimization in RF circuit interconnect treesabstractWe present a methodology for phase noise minimization of interconnects in radio frequency circuits, integrated into a commercial digital tool chain. Accurate estimates of the produced phase noise are derived using a lookup table approach, eliminating the need for analog simulations. A dynamic programming algorithm is utilized to produce the optimal tree structure. The tree is automatically translated into a netlist and placed and routed within the VLSI flow. Back annotated simulations in 28 nm technology show that the obtained results are within 2.2 dB of the actual phase noise, while significantly reducing the design time compared to traditional manual design. To the best of our knowledge, this is the first work on the automation of buffer insertion for phase noise minimization. Dimo Martev, Sven Hampel, Ulf Schlichtmann |
ISCAS | 3 |
| 2017 | The extendable translating instruction set simulator (ETISS) interlinked with an MDA framework for fast RISC prototypingabstractThis paper describes the Extendable Translating Instruction Set Simulator (ETISS). In addition to binary translation, ETISS features a plugin mechanism that allows to quickly include new functionality into the translation stage, the simulation loop, during accesses to the memory or whenever an interrupt is received. ETISS targets to become an advanced industrial-strength ISS with special focus on virtual prototypes (VPs) written in SystemC/TLM. In this paper, we will show examples of ETISS Plugins which include tracing tools, SystemC interfaces, closey-coupled peripherals or triggers for fault injection. A major drawback of developing a new binary translator such as ETISS is its lack of support for a variety of instruction set architectures (ISAs). At the moment ETISS supports the open-source OpenRISC orlk and partly RISC-V ISAs. Yet, in order to overcome this problem, we developed a toolchain to generate the binary translation stage for different ISAs following the MDA concept based on meta-modeling and code generation. It is planned to make ETISS available as an open-source tool to the research community. Daniel Mueller-Gritschneder, Keerthikumara Devarajegowda, Martin Dittrich, Wolfgang Ecker, Marc Greim, Ulf Schlichtmann |
RSP | 6 |
| 2017 | An Efficient Two-Phase ILP-Based Algorithm for Precise CMOS RFIC Layout GenerationabstractWith advancing process technologies and booming Internet of Things markets, millimeter-wave CMOS RFICs have evolved rapidly and been widely applied in recent years. The performance of CMOS RFICs is very sensitive to the chip layout, and a tiny variation of the microstrip length can cause a large impact to the circuit performance. This results in a time-consuming tuning process including much simulation effort for chip design, which becomes the major bottleneck for time to market. This paper introduces a progressive integer-linear-programming-based method consisting of two phases: 1) global layout generation and 2) iterative validation. In the global layout generation phase, we focus on the most critical constraints such as layout planarity and device connection relations to determine the topology of the final design. This provides a basis for constructing the accurate model in the iterative validation phase. The layouts generated by applying our method can satisfy very stringent routing requirements of microstrip lines, including spacing/noncrossing rules, precise length, and bend number minimization, within a given layout area. The resulting RFIC layouts excel in both performance and area with much fewer bends compared with the simulation-tuning based manual layout, while the layout generation time is significantly reduced from weeks to a few minutes. Tsun-Ming Tseng, Bing Li 0005, Ching-Feng Yeh, Hsiang-Chieh Jhan, Zuo-Min Tsai, Mark Po-Hung Lin, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2016 | Reliability, adaptability and flexibility in timing: Buy a life insurance for your circuitsabstractAt nanometer manufacturing technology nodes, process variations affect circuit performance significantly. In addition, performance deterioration of circuits due to aging effects is also increasing. Consequently, a large timing margin is required to maintain yield. To combat the pessimism and the resulting overdesign, aging analysis with highlevel models, on-chip timing margin monitoring and tuning, and flexible delay models of flip-flops can be deployed. This paper gives an overview of the state of the art of applying these techniques to improve the health of circuits. Ulf Schlichtmann, Masanori Hashimoto, Iris Hui-Ru Jiang, Bing Li 0005 |
ASP-DAC | 1 |
| 2016 | Columba: co-layout synthesis for continuous-flow microfluidic biochipsabstractContinuous-flow microfluidics have evolved rapidly in the last decades, due to their advantages in effective and accurate control. However, complex control results in complicated valve actuations. As a result, sophisticated interactions between control and flow layers substantially raise the design difficulty. Previous work on design automation for microfluidics neglects the interactions between the control and flow layers and designs each layer separately, which leads to unrealistic designs. We propose the first planarity-guaranteed architectural model, and the first physical-design module models for important microfluidic components, which have modelled the interactions between both control and flow layers, while reducing the design difficulty. Based on the above, we propose the co-layout synthesis tool called Columba, which considers the pressure sharing among different valves, and routes channels in an any-angled manner. Experimental results show that complicated designs considering layer interactions can be synthesized for the first time. Tsun-Ming Tseng, Mengchu Li, Bing Li 0005, Tsung-Yi Ho, Ulf Schlichtmann |
DAC | 5 |
| 2016 | Novel CMOS RFIC layout generation with concurrent device placement and fixed-length microstrip routingabstractWith advancing process technologies and booming IoT markets, millimeter-wave CMOS RFICs have been widely developed in recent years. Since the performance of CMOS RFICs is very sensitive to the precision of the layout, precise placement of devices and precisely matched microstrip lengths to given values have been a labor-intensive and time-consuming task, and thus become a major bottleneck for time to market. This paper introduces a progressive integer-linear-programming-based method to generate high-quality RFIC layouts satisfying very stringent routing requirements of microstrip lines, including spacing/non-crossing rules, precise length, and bend number minimization, within a given layout area. The resulting RFIC layouts excel in both performance and area with much fewer bends compared with the simulation-tuning based manual layout, while the layout generation time is significantly reduced from weeks to half an hour. Tsun-Ming Tseng, Bing Li 0005, Ching-Feng Yeh, Hsiang-Chieh Jhan, Zuo-Min Tsai, Mark Po-Hung Lin, Ulf Schlichtmann |
DAC | 7 |
| 2016 | EffiTest: efficient delay test and statistical prediction for configuring post-silicon tunable buffersabstractAt nanometer manufacturing technology nodes, process variations significantly affect circuit performance. To combat them, post-silicon clock tuning buffers can be deployed to balance timing budgets of critical paths for each individual chip after manufacturing. The challenge of this method is that path delays should be measured for each chip to configure the tuning buffers properly. Current methods for this delay measurement rely on path-wise frequency stepping. This strategy, however, requires too much time from expensive testers. In this paper, we propose an efficient delay test framework (EffiTest) to solve the post-silicon testing problem by aligning path delays using the already-existing tuning buffers in the circuit. In addition, we only test representative paths and the delays of other paths are estimated by statistical delay prediction. Experimental results demonstrate that the proposed method can reduce the number of frequency stepping iterations by more than 94% with only a slight yield loss. Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann |
DAC | 3 |
| 2016 | Sieve-valve-aware synthesis of flow-based microfluidic biochips considering specific biological execution limitations
Mengchu Li, Tsun-Ming Tseng, Bing Li 0005, Tsung-Yi Ho, Ulf Schlichtmann |
DATE | 5 |
| 2016 | Sampling-based buffer insertion for post-silicon yield improvement under process variability
Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann |
DATE | 3 |
| 2016 | Hardware-Accelerated Software Library Drivers Generation for IP-Centric SoC DesignsabstractIn recent years, the semiconductor industry has been witnessing an increasing reuse of hardware IPs for System-on-Chip (SoC) designs and embedded computing systems on FPGA platforms with hard-core processors. The IP-reuse comes with an increasing complexity at the hardware-software (HW-SW) interface. The efforts required to access the HW through the increasingly complex HW-SW interface diminishes the potential IP-reuse productivity gain. In our work, we are proposing hierarchical drivers for accessing IP-subsystems and its generation for enabling easier SW application adaptation to HW-changes and faster design space exploration (DSE) on a targeted HW-accelerated SW libraries. At the lowest level, closest to the HW, is the hardware abstraction layer (HAL), these are the platform-specific register-access drivers. At the next layer are the drivers to access the registers and bit-fields of each IP component of the IP-library. Next are the IP-subsystems drivers. At the top-layer, closest to the SW, is the simple scheduler with SW interface library that provides access functions to the SW application. The drivers generator uses the HW knowledge of IPs and IP-subsystems encoded in IP-XACT for generating the drivers for both operating system (OS) and non-OS based applications. For the OS-based applications, user-space drivers are generated, as well as device tree source (DTS) and drivers mapping in the kernel-space. In a case study, we have validated our methodology while performing DSE for a video processing application targeted to an IP-library, both as non-OS and with OS on Xilinx Zynq-based FPGA. Munish Jassi, Uzair Sharif, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 4 |
| 2016 | Where formal verification can help in functional safety analysisabstractFormal techniques seem to be a way to cope with the exploding complexity of functional safety analysis. Here, the overall fault propagation probability to a certain safety-point in the design must be analyzed. As a consequence, the careful verification of the design is no longer sufficient. In addition, the propagation of all possible faults potentially showing up at all of the design's internal nodes must be validated. But this is not only a complexity challenge. Safety standards have a probabilistic view on functional safety analysis results and aspects such as different fault and pattern probability must be considered and related to requirements such as confidence level and maximum FIT rate. Following an overview on verification challenges around functional safety analysis, we introduce our innovative concept on how formal formal techniques can substantially simplify industrial functional safety analysis ows. Alessandro Bernardini, Wolfgang Ecker, Ulf Schlichtmann |
ICCAD | 3 |
| 2016 | Control-fluidic CoDesign for paper-based digital microfluidic biochipsabstractPaper-based digital microfluidic biochips (P-DMFBs) have recently emerged as a promising low-cost and fast-responsive platform for biochemical assays. In P-DMFBs, electrodes and control lines are printed on a piece of photo paper using inkjet printer and conductive ink of carbon nanotubes (CNTs). Compared with traditional digital microfluidic biochips (DMFBs), P-DMFBs enjoy notable advantages, such as faster in-place fabrication with printer and ink, lower costs, better disposability, etc. Because electrodes and CNT control lines are printed on the same side of a paper, a new design challenge for P-DMFB is to prevent the interference between moving droplets and the voltages on CNT control lines. These interactions may result in unexpected droplet movements and thus incorrect assay outputs. To address the new challenges in automated design of P-DMFBs, this paper proposes the first control-fluidic codesign flow, which simultaneously adjusts the control line routing and fluidic droplet scheduling to achieve an optimized solution. As the control line routing may not be able to address all the interferences between moving droplets and the voltages on control lines, droplet rescheduling is performed to effectively deal with the remaining interferences in the routing solution. Computational simulation results on real-life bioassays show that the proposed codesign method successfully eliminates all the interferences, while a state-of-the-art maze routing method cannot solve any of the benchmarks without conflicts. Qin Wang 0005, Zeyan Li 0001, Haena Cheong, Oh-Sun Kwon, Hailong Yao 0002, Tsung-Yi Ho, Kwanwoo Shin, Bing Li 0005, Ulf Schlichtmann, Yici Cai |
ICCAD | 9 |
| 2016 | From biochips to quantum circuits: computer-aided design for emerging technologiesabstractWhile previous decades have witnessed impressive accomplishments in the design and realization of conventional computing devices, physical boundaries and cost restrictions led to an increasing interest in alternative technologies (often referred to as Beyond CMOS or More than Moore technologies). In addition, these accomplishments also triggered many “complementary” applications and led to technologies providing an additional value to the conventional logic (often referred to as More than Moore). This led to a variety of emerging technologies such as Quantum Computation, Optical Circuits, or Microfluidic Biochips out of which many are considered very promising and some even entered the market recently. This poses new challenges to researchers and engineers working in computer-aided design. In this tutorial paper, we provide an overview on the main concepts of selected emerging technologies as well as the resulting design methods. To this end, we review the respective technological background and introduce the correspondingly used circuit models. Based on that, we show how computer-aided design has to adapt the common design tasks and review recently proposed solutions. Robert Wille, Bing Li 0005, Ulf Schlichtmann, Rolf Drechsler |
ICCAD | 3 |
| 2016 | PieceTimer: a holistic timing analysis framework considering setup/hold time interdependency using a piecewise modelabstractIn static timing analysis, clock-to-q delays of flip-flops are considered as constants. Setup times and hold times are characterized separately and also used as constants. The characterized delays, setup times and hold times, are applied in timing analysis independently to verify the performance of circuits. In reality, however, clock-to-q delays of flip-flops depend on both setup and hold times. Instead of being constants, these delays change with respect to different setup/hold time combinations. Consequently, the simple abstraction of setup/hold times and constant clock-to-q delays introduces inaccuracy in timing analysis. In this paper, we propose a holistic method to consider the relation between clock-to-q delays and setup/hold time combinations with a piecewise linear model. The result is more accurate than that of traditional timing analysis, and the incorporation of the interdependency between clock-to-q delays, setup times and hold times may also improve circuit performance. Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann |
ICCAD | 3 |
| 2016 | PLATON: A Force-Directed Placement Algorithm for 3D Optical Networks-on-ChipabstractOptical Networks-on-Chip (ONoCs) are a promising technology to further increase the bandwidth and decrease the power consumption of today's multicore systems. To determine the laser power consumption of an ONoC, the physical design of the system is indispensible. The only place and route tool for 3D ONoCs already proposed in the literature badly scales with the increasing number of optical devices. Thus, within this contribution we present the first force-directed placement algorithm for 3D optical NoCs. Our algorithm decreases the runtime up to 99.7\% compared to the state-of-the-art placer. Using our algorithm large topologies can be placed within a short runtime. Anja von Beuningen, Ulf Schlichtmann |
ISPD | 2 |
| 2016 | Efficient handling of the fault space in functional safety analysis utilizing formal methodsabstractCircuit robustness can be increased with selective Flip-Flop hardening. Finding candidate sets of Flip-Flops for optimal selective hardening requires costly fault simulations, in particular if we consider safety properties stating that a bad state should never be reached in future. We present a fully symbolic formal method that gives a rigorous robustness measure without the need of extensive fault simulation and that can be applied in early design stages for selective hardening. Using Formal Verification, we define, compute and measure a set of “critical transitions”. The Markov Property is not required for the proposed method. Alessandro Bernardini, Wolfgang Ecker, Ulf Schlichtmann |
VLSI-SoC | 3 |
| 2016 | PROTON+: A Placement and Routing Tool for 3D Optical Networks-on-Chip with a Single Optical LayerabstractOptical Networks-on-Chip (ONoCs) are a promising technology to overcome the bottleneck of low bandwidth of electronic Networks-on-Chip. Recent research discusses power and performance benefits of ONoCs based on their system-level design, while layout effects are typically overlooked. As a consequence, laser power requirements are inaccurately computed from the logic scheme but do not consider the layout. In this article, we propose PROTON+, a fast tool for placement and routing of 3D ONoCs minimizing the total laser power. Using our tool, the required laser power of the system can be decreased by up to 94% compared to a state-of-the-art manually designed layout. In addition, with the help of our tool, we study the physical design space of ONoC topologies. For this purpose, topology synthesis methods (e.g., global connectivity and network partitioning) as well as different objective function weights are analyzed in order to minimize the maximum insertion loss and ultimately the system’s laser power consumption. For the first time, we study optimal positions of memory controllers. A comparison of our algorithm to a state-of-the-art placer for electronic circuits shows the need for a different set of tools custom-tailored for the particular requirements of optical interconnects. Anja von Beuningen, Luca Ramini, Davide Bertozzi, Ulf Schlichtmann |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2016 | Multivariate Modeling of Variability Supporting Non-Gaussian and Correlated ParametersabstractProcess variations and atomic-level fluctuations increasingly pose challenges to the design and analysis of integrated circuits by introducing variability. Although several approaches have been proposed to deal with the inherent statistical nature of circuit design, we consider them incomplete with two important aspects often being insufficiently addressed: 1) non-Gaussian distributions and 2) highly correlated parameters. To address these points, we propose a fully multivariate and non-Gaussian approach based on an arbitrary model. A subset of the model parameters is treated as a multidimensional random variable, which is represented by a combination of generalized lambda distributions and Spearman rank correlation matrices-a very general approach with nearly arbitrary freedom in distribution shapes and parameter correlations. In our application scenarios, we show that such a model is able to fully and accurately capture variability in device compact models and standard cell performance models. Finally, we present adapted analysis methods making use of these models in circuit simulations and in efficient gate level analyses of digital circuits with high accuracy. André Lange, Christoph Sohrmann, Roland Jancke, Joachim Haase, Binjie Cheng, A. Asenov, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2016 | Reliability-Aware Synthesis With Dynamic Device Mapping and Fluid Routing for Flow-Based Microfluidic BiochipsabstractIn flow-based biochips, peristaltic pumps consisting of valves are essential to generate circulation flow in a mixer. When a peristaltic pump is activated, the related valves for peristalsis are required to be actuated for many times. However, the roles of valves in traditional chips are fixed, and therefore the valves for peristalsis can wear out much faster than the valves for guiding fluid transportation. This could lead to a reduced lifetime of the chip, because the whole chip function can be affected when just a few or even only a single valve wears out. In this paper, we propose a valve-centered architecture with virtual valves, based on which we introduce a valve-role-changing concept to balance the valve actuations. By switching a valve into different roles, microfluidic components such as mixers, storages, and flow channels can be formed dynamically during the assay process, which enables us to balance the utilization of valves, and synthesize designs that support various kinds of operations. Compared with our preliminary work, we further decrease the largest number of valve actuation as well as the number of valves by the revised dynamic device mapping and fluid path routing. For dynamic device mapping, we introduce a virtual-boundary concept to generate devices at better places while connections between devices are still guaranteed. For fluid path routing, we accurately model valve actuation resulting from our valve-actuation-aware routing, and revise the results by rip-up and reroute. In addition to performance, we improve the reliability of our method by assuring fluid paths from devices to chip boundaries. Experiments show that the new method can be eight times better than the traditional method, and outperforms our preliminary work for large cases even with fewer valves. Tsun-Ming Tseng, Bing Li 0005, Mengchu Li, Tsung-Yi Ho, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2015 | GRIP: grammar-based IP integration and packaging for acceleration-rich SoC designsabstractIncreased hardware IP reuse is required to meet the productivity demands for the future complex Systems-on-Chip (SoC). Nowadays, IP integration is enabled using standardized meta-data formats such as IP-XACT. We present a new concept called grammar-based IP integration and packaging (GRIP), which additionally encodes design integration knowledge into a set of graph re-writing rules using standard IP-XACT. These GRIP rules are packaged into a domain-specific library of IP blocks. The library can be supplied by an IP provider along to an SoC architect. An integration tool can automatically use the GRIP rules to search the design space using the integration knowledge of the IP provider. The tool generates all design alternatives with different trade-offs for the SoC architect. We demonstrate the GRIP approach on a computer vision IP library for FPGA-based SoCs. Eighteen functional design alternatives are automatically generated within a few hours using IP integration knowledge encoded by the GRIP rules. Munish Jassi, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DAC | 3 |
| 2015 | Reliability-aware synthesis for flow-based microfluidic biochips by dynamic-device mappingabstractOn flow-based biochips, valves that are used to form peristaltic pumps wear out much earlier than valves for transportation since the former are actuated more often, which leads to a reduced lifetime of the chip. In this paper, we introduce a valve-role-changing concept to avoid always using the same valves for peristalsis. Based on this, we generate dynamic devices from a valve-centered architecture to distribute the valve actuation activities evenly and reduce the largest number of valve actuations with even fewer valves. In addition, we propose in situ on-chip storages, which can overlap with other devices, so that less area is needed compared with dedicated storages on traditional chips. Moreover, our method provides good support for assays requiring different volumes and ratios of samples. Experiments show that compared with traditional designs, the largest number of valve actuations can be reduced by 72.97% averagely, while the number of valves is reduced by 10.62%. Tsun-Ming Tseng, Bing Li 0005, Tsung-Yi Ho, Ulf Schlichtmann |
DAC | 4 |
| 2015 | Timing verification for adaptive integrated circuits
Rohit Kumar 0001, Bing Li 0005, Yiren Shen, Ulf Schlichtmann, Jiang Hu 0001 |
DATE | 4 |
| 2015 | Beyond GORDIAN and Kraftwerk: EDA Research at TUMabstractAt the Institute for Electronic Design Automation of Technische Universität München (TUM), founded in 1975 by Prof. Kurt Antreich as Germany's first university institute dedicated to EDA, a broad range of research has been performed in the past 40 years. We describe here the research activities that Prof. Antreich undertook in addition to his research on physical design, as well as the physical design research undertaken in the past decade after his official retirement. Ulf Schlichtmann |
ISPD | 1 |
| 2015 | Runtime Adaptation of Application Execution under Thermal and Power Constraints in Massively Parallel Processor ArraysabstractMassively Parallel Processor Arrays (MPPAs) can be nicely used in portable devices such as tablets and smartphones. However, applications running on mobile platforms require a certain performance level or quality (e.g., high-resolution image processing) that need to be satisfied while adhering to a certain power budget and temperature threshold. As a solution to the aforementioned challenges, we consider a resource-aware computing paradigm to exploit runtime adaptation without violating any thermal and/or power constraint in a programmable MPPA. For estimating the power consumption, we developed a mathematical model based on the post-synthesis implementation of an MPPA in different CMOS technologies while the temperature variation was emulated. We showcase our hardware/software mechanism to load new, on-the-fly configurations into the accelerator, considering quality/throughput tradeoffs for image processing applications. The results show that the average power consumption of a Sobel and Laplace operators using different number of processing elements amounts to 1.24 mW and 10.35 mW, respectively. Furthermore, only 1.64 μs are necessary for configuring a class of MPPA running at 550 MHz. Éricles Sousa, Frank Hannig, Jürgen Teich, Qingqing Chen 0004, Ulf Schlichtmann |
SCOPES | 5 |
| 2015 | A Cross-Layer Approach to Measure the Robustness of Integrated CircuitsabstractThe demands on system robustness and its immunity against perturbations are getting increasingly important. Nearly everybody has an intuitive understanding of what robustness means, but there is no proper way how to measure robustness of integrated circuits already during the design phase. Therefore, a general cross-layer robustness model and methods to quantitatively measure robustness are presented. Moreover, these methods are refined to predict the robustness against degradation of digital circuits due to aging effects. Martin Barke, Ulf Schlichtmann |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2015 | Statistical Timing Analysis and Criticality Computation for Circuits With Post-Silicon Clock Tuning ElementsabstractPost-silicon clock tuning elements are widely used in high-performance designs to mitigate the effects of process variations and aging. Located on clock paths to flip-flops, these tuning elements can be configured through the scan chain so that clock skews to these flip-flops can be adjusted after manufacturing. Owing to the delay compensation across consecutive register stages enabled by the clock tuning elements, higher yield and enhanced robustness can be achieved. These benefits are, nonetheless, attained by increasing die area due to the inserted clock tuning elements. For balancing performance improvement and area cost, an efficient timing analysis algorithm is needed to evaluate the performance of such a circuit. So far this evaluation is only possible by Monte Carlo simulation which is very time-consuming. In this paper, we propose an alternative method using graph transformation, which computes a parametric minimum clock period and is more than 104times faster than Monte Carlo simulation while maintaining a good accuracy. This method also identifies the gates that are critical to circuit performance, so that a fast analysis-optimization flow becomes possible. Bing Li 0005, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | ILP-Based Alleviation of Dense Meander Segments With Prioritized Shifting and Progressive Fixing in PCB RoutingabstractLength-matching is an important technique to balance delays of bus signals in high-performance printed circuit board (PCB) routing. Existing routers, however, may generate very dense meander segments. Signals propagating along these meander segments exhibit a speedup effect due to crosstalk between the segments of the same wire, thus leading to mismatch of arrival times even under the same physical wire length. In this paper, we present a post-processing method to enlarge the width and the distance of meander segments and hence distribute them more evenly on the board so that crosstalk can be reduced. In the proposed framework, we model the sharing of available routing areas after removing dense meander segments from the initial routing, as well as the generation of relaxed meander segments and their groups for wire length compensation. This model is transformed into an ILP problem and solved for a balanced distribution of wire patterns. In addition, we adjust the locations of long wire segments according to wire priorities to swap free spaces toward critical wires that need much length compensation. To reduce the problem space of the ILP model, we also introduce a progressive fixing technique so that wire patterns are grown gradually from the edge of the routing toward the center area. Experimental results show that the proposed method can expand meander segments significantly even under very tight area constraints, so that the speedup effect can be alleviated effectively in high-performance PCB designs. Tsun-Ming Tseng, Bing Li 0005, Tsung-Yi Ho, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2014 | Workload- and Instruction-Aware Timing Analysis: The missing Link between Technology and System-level ResilienceabstractIn today's design of resilient embedded systems, logic circuit components play a key role. Many possible design choices at the gate level, such as implementation architecture or synthesis constraints, are vital for the resilience of the entire system. Hence, EDA algorithms at this level have to support exposing technology characteristics (such as process variations or aging) for consideration on higher levels of abstraction. Similarly, key parameters from system level, such as workload or executed processor instructions, have to be considered at lower levels for accurate analysis of, e.g., degradation effects. Circuit-level timing analysis plays a key role in this context as it provides key metrics such as achievable frequency, available timing margins and timing violation vulnerabilities of the analyzed circuit. We present an enhanced static timing analysis which links technology-level effects to system-level and vice versa. Specifically, we discuss the accurate and efficient consideration of system workload and impact of executed instructions on circuit timing. Veit Kleeberger, Petra R. Maier, Ulf Schlichtmann |
DAC | 3 |
| 2014 | Safety Evaluation of Automotive Electronics Using Virtual Prototypes: State of the Art and Research ChallengesabstractIntelligent automotive electronics significantly improved driving safety in the last decades. With the increasing complexity of automotive systems, dependability of the electronic components themselves and of their interaction must be assured to avoid any risk to driving safety due to unexpected failures caused by internal or external faults. Jan-Hendrik Oetjens, Nico Bannow, Markus Becker 0001, Oliver Bringmann 0001, Andreas Burger, Moomen Chaari, Samarjit Chakraborty, Rolf Drechsler, Wolfgang Ecker, Kim Grüttner, Thomas Kruse, Christoph Kuznik, Hoang Minh Le 0001, Andreas Mauderer, Wolfgang Müller 0003, Daniel Mueller-Gritschneder, Frank Poppen, Hendrik Post, Sebastian Reiter 0003, Wolfgang Rosenstiel, S. Roth, Ulf Schlichtmann, Andreas von Schwerin, Bogdan-Andrei Tabacaru, Alexander Viehl |
DAC | 22 |
| 2014 | Probabilistic standard cell modeling considering non-Gaussian parameters and correlationsabstractVariability continues to pose challenges to integrated circuit design. With statistical static timing analysis and high-yield estimation methods, solutions to particular problems exist, but they do not allow a common view on performance variability including potentially correlated and non-Gaussian parameter distributions. In this paper, we present a probabilistic approach for variability modeling as an alternative: model parameters are treated as multi-dimensional random variables. Such a fully mul-tivariate statistical description formally accounts for correlations and non-Gaussian random components. Statistical characterization and model application are introduced for standard cells and gate-level digital circuits. Example analyses of circuitry in a 28 nm industrial technology illustrate the capabilities of our modeling approach. André Lange, Christoph Sohrmann, Roland Jancke, Joachim Haase, Ingolf Lorenz, Ulf Schlichtmann |
DATE | 6 |
| 2014 | Special session: How secure are PUFs really? On the reach and limits of recent PUF attacksabstractJust over a decade ago, Physical Unclonable Functions (PUFs) have been introduced as a new cryptographic and security primitive in a number of seminal publications. Due to their assumed security and cost advantages, they have attracted substantial attention both from the security industry and the academic community, and are also gaining ground in commercial applications. Nevertheless, a number of recent works have presented successful attacks on PUF core properties, such as their digital and physical unclonability. How strong and relevant are these attacks, and how secure are PUFs really? This question is addressed in a dedicated hot topic session at DATE 2014. This paper provides a short and easily accessible overview of the session. Ulrich Rührmair, Ulf Schlichtmann, Wayne P. Burleson |
DATE | 2 |
| 2014 | Connecting different worlds - Technology abstraction for reliability-aware design and TestabstractThe rapid shrinking of device geometries in the nanometer regime requires new technology-aware design methodologies. These must be able to evaluate the resilience of the circuit throughout all System on Chip (SoC) abstraction levels. To successfully guide design decisions at the system level, reliability models, which abstract technology information, are required to identify those parts of the system where additional protection in the form of hardware or software coun-termeasures is most effective. Interfaces such as the presented Resilience Articulation Point (RAP) or the Reliability Interchange Information Format (RIIF) are required to enable EDA-assisted analysis and propagation of reliability information. The models are discussed from different perspectives, such as design and test. Ulf Schlichtmann, Veit Kleeberger, Jacob A. Abraham, Adrian Evans, Christina Gimmler-Dumont, Michael Glaß, Andreas Herkersdorf, Sani R. Nassif, Norbert Wehn |
DATE | 1 |
| 2014 | Deterministic Synthesis of Hybrid Application-Specific Network-on-Chip TopologiesabstractNetworks-on-Chip (NoCs) enable cost-efficient and effective communication between the processing elements inside modern systems-on-chip (SoCs). NoCs with regular topologies such as meshes, tori, rings, and trees are well suited for general-purpose many core SoCs. These topologies might prove suboptimal for SoCs with predefined application characteristics and traffic patterns. Such SoCs benefit from application-specific NoC topologies, designed and optimized according to the application characteristics. This paper proposes a synthesis approach for creating hybrid, application-specific NoCs from an input floorplan and a set of use cases, describing the applications running on the SoC. The method considers latency, port count, and link length constraints. It produces hybrid topologies that utilize both NoC routers and shared buses. Furthermore, the proposed approach can insert intermediate relay routers that act as bridges or repeaters and help to reduce the cost further. Finally, the approach creates a deadlock-free routing of the communication flows by either finding deadlock-free paths or by inserting virtual channels. The benefits of the proposed method are demonstrated by comparing it to state-of-the-art approaches on a generic and an industrial SoC examples. Vladimir Todorov, Daniel Mueller-Gritschneder, Helmut Reinig, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2013 | Memory access reconstruction based on memory allocation mechanism for source-level simulation of embedded softwareabstractTo date, there still lacks a way to accurately simulate data memory accesses in source-level simulation (SLS) of host-compiled embedded SW. The difficulty lies in that the accessed addresses for the load and store instructions can not be statically determined. Without knowing those addresses, the source code can not be annotated appropriately for data cache simulation. In this paper, we show an approach that is capable of resolving the accessed memory addresses based on the memory allocation mechanism. Applying this approach, the source code can be annotated to perform precise data cache simulation. The novelty of our methodology is that it is the first of its kind to take the memory allocation mechanism into account and thus can handle all the stack, data, heap and text sections. Moreover, a method is also proposed to handle pointer dereferences. In experiments, SLS with our approach yields almost identical cache miss rate and pattern when compared to the reference simulation. Kun Lu 0005, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ASP-DAC | 3 |
| 2013 | Reliability challenges for electric vehicles: from devices to architecture and systems softwareabstractToday, modern high-end cars have close to 100 electronic control units (ECUs) that are used to implement a variety of applications ranging from safety-critical control to driver assistance and comfort-related functionalities. The total sum of these applications is several million lines of software code. The ECUs are connected to different sensors and actuators and communicate via a variety of communication buses like CAN, FlexRay and now also Ethernet. In the case of electric vehicles, both the amount and the importance of such electronics and software are even higher. Here, a number of hydraulic or pneumatic controls are replaced by corresponding software-implemented controllers in order to reduce the overall weight of the car and hence to improve its driving range. Until recently, most of the software and system design in the automotive domain -- as in many other domains -- relied on an always correctly functioning or a zero-defect hardware implementation platform. However, as the device geometries of integrated circuits continue to shrink, this assumption is increasingly not true. Incorporating large safety margins in the design process results in very pessimistic design and expensive processors. Further, the processors in cars -- in contrast to those in many consumer electronics devices like mobile phones -- are exposed to harsh environments, extreme temperature variations, and often, strong electromagnetic fields. Hence, their reliability is even more questionable and must be explicitly accounted for in all layers of design abstraction -- starting from circuit design to architecture design, to software design and runtime management and monitoring. In this paper we outline some of these issues, currently followed practices, and the challenges that lie ahead of us in the automotive and electric vehicles domain. Georg Georgakos, Ulf Schlichtmann, Reinhard Schneider 0001, Samarjit Chakraborty |
DAC | 2 |
| 2013 | Predicting future product performance: modeling and evaluation of standard cells in FinFET technologiesabstractWith continued scaling of CMOS technology it becomes increasingly difficult to maintain reliable circuits. Early predictive technology and design exploration help to understand major effects of variability sources and their impact on circuit performances. With each new technology basic circuit blocks have to be redesigned to appropriately evaluate the impact of technology scaling. Therefore, this paper presents an approach which is able to find the optimal sizing of basic circuit blocks considering process variation. We utilize this approach to predict the impact of scaling in FinFET technologies and the influence of process variations in future technology nodes. Veit Kleeberger, Helmut E. Graeb, Ulf Schlichtmann |
DAC | 3 |
| 2013 | Fast cache simulation for host-compiled simulation of embedded softwareabstractHost-compiled simulation has been proposed for software performance estimation, because of its high simulation speed. However, the simulation speed may be significantly lowered due to the cache simulation overhead. In this paper, we propose an approach that can reduce much of the cache simulation overhead, while still calculating cache misses precisely. For instruction cache, we statically analyze possible cache conflicts and perform cache conflicts aware annotation for host-compiled simulation. Within loops, the conflicts are dynamically captured by tagging the basic blocks instead of performing the expensive cache simulation. In this way, a vast majority of the cache accesses can be saved from simulation. For data cache, aggregated cache simulation is used for a large data block. Further, the data locality can be bound by considering the data allocation principle of a program. Experiments show that our approach improves the speed of host-compiled simulation by one order of magnitude, while providing the cache miss numbers with high accuracy. Kun Lu 0005, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DATE | 3 |
| 2013 | Analytical timing estimation for temporally decoupled TLMs considering resource conflictsabstractTransaction level models (TLMs) can use temporal decoupling to increase the simulation speed. However, there is a lack of modeling support to time the temporally decoupled TLMs. In this paper, we propose a timing estimation mechanism for TLMs with temporal decoupling. This mechanism features an analytical model and novel delay formulas. Concepts such as resource usage and availability are used to derive the delay formulas. Based on them, a fast scheduling algorithm resolves resource conflicts and dynamically determines the timing of concurrent transaction sequences. Experiments show that the delay estimation formulas are capable of capturing the timing effects of resource conflicts. At the same time, the overhead of the scheduling algorithm is very low, hence the simulation speed remains high. Kun Lu 0005, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DATE | 3 |
| 2013 | A virtual prototyping platform for real-time systems with a case study for a two-wheeled robotabstractIn today's real-time system design, a virtual prototype can help to increase both the design speed and quality. Developing a virtual prototyping platform requires realistic modeling of the HW system, accurate simulation of the real-time SW, and integration with a reactive real-time environment. Such a VP simulation platform is often difficult to develop. In this paper, we propose a case-study of autonomous two-wheeled robot to show how to develop a virtual prototyping platform rapidly in SystemC/TLM to adequately aid in the design of this instable system with hard real-time constraints. Our approach is an integration of four major model components. Firstly, an accurate physical model of the robot is provided. Secondly, a virtual world is modeled in Java that offers a 3D environment for the robot to move in. Thirdly, the embedded control SW is developed. Finally, the overall HW system is modeled in SystemC at transaction level. This HW model wraps the physical model, interacts with the virtual world, and simulates the real-time SW by integrating an Instruction Set Simulator of the embedded CPU. By integrating these components into a platform, designers can efficiently optimize the embedded SW architecture, explore the design space and check real-time conditions for different system parameters such as buffer sizes, CPU frequency or cache configurations. Daniel Mueller-Gritschneder, Kun Lu 0005, Erik Wallander, Marc Greim, Ulf Schlichtmann |
DATE | 5 |
| 2013 | A spectral clustering approach to application-specific network-on-chip synthesisabstractModern System-on-Chip (SoC) design relies heavily on efficient interconnects like Networks-on-Chip (NoCs). They provide an effective, flexible and cost efficient way of communication exchange between the individual processing elements of the SoC. Therefore, the choice of topology and design of the NoC itself plays a crucial role in the performance of the system. Depending on the field of application, standard topologies like meshes, fat-trees, and tori might be suboptimal in terms of power consumption, latency and area. This calls for a custom topology design methodology, which is based on the requirements imposed by the application, function and the use-cases of the SoC in question. This work proposes a fast approach, which uses spectral clustering and cluster ensembles to partition the system using normalized cuts and insert the necessary routers. Then, by using delay-constrained minimum spanning trees, links between the individual routers are created, such that any present latency constraints are satisfied at minimum cost. Results from applying the methodology to a smartphone SoC are presented. Vladimir Todorov, Daniel Mueller-Gritschneder, Helmut Reinig, Ulf Schlichtmann |
DATE | 4 |
| 2013 | Post-route refinement for high-frequency PCBs considering meander segment alleviationabstractIn this paper, we propose a post-processing framework which iteratively refines the routing results from an existing PCB router by removing dense meander segments. By swapping and detouring dense meander segments the proposed method can effectively alleviate accumulating crosstalk noise, while respecting pre-defined area constraints. Experimental results show more than 85% reduction of the meander segments and hence the noise cost. Tsun-Ming Tseng, Bing Li 0005, Tsung-Yi Ho, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 4 |
| 2013 | PROTON: an automatic place-and-route tool for optical networks-on-chipabstractOptical Networks-on-Chip (ONoCs) are considered a promising way of improving power and bandwidth limitations in next generation multi- and many-core integrated systems. Today, most related research acknowledges the key role of the physical layer in assessing ONoC topologies (e.g., insertion loss), but overlooks the placement and routing stage in the design process, hence applying physical design considerations to topology logic schemes. Such a mismatch is fundamentally due to the lack of mature CAD tools for placement and routing of optical NoCs. The objective of this work is to bridge this gap: We propose PROTON, a fast tool for automatic placement and routing of ONoC topologies, which can support designers in quantifying the degradation of design quality metrics when moving from topology logic schemes to their physical implementation. This gap is especially relevant for Wavelength-Routed ONoCs (WRONoCs), where logic schemes typically make unrealistic assumptions about the placement of initiators and targets. For this reason, we put PROTON to work with the most promising WRONoC topologies and explore their physical design space given the placement and routing constraints of a 3D stacked system. We also compare automatically generated layouts with handcrafted ones reported in the literature for the same topologies and target system, and prove an insertion loss improvement by up to 150x. With PROTON the exploration of the physical design space of ONoC topologies is possible as well as their scalability analysis considering the layout. Anja Boos, Luca Ramini, Ulf Schlichtmann, Davide Bertozzi |
ICCAD | 3 |
| 2013 | Post-route alleviation of dense meander segments in high-performance printed circuit boardsabstractLength-matching is an important technique to balance delays of bus signals in high-performance PCB routing. Existing routers, however, may generate dense meander segments with small distance. Signals propagating across these meander segments exhibit a speedup effect due to crosstalks between the segments of the same wire, thus leading to mismatch of arrival times even with the same physical wire length. In this paper, we propose a post-processing method to enlarge the width and the distance of meander segments and distribute them more evenly on the board so that the crosstalks can be reduced. In the proposed framework, we model the sharing combinations of available routing areas after removing dense meander segments from the initial routing, as well as the generation of relaxed meander segments and their groups in subareas. Thereafter, this model is transformed into an ILP problem and solved efficiently. Experimental results show that the proposed method can extend the width and the distance of meander segments about two times even under very tight area constraints, so that the crosstalks and thus the speedup effect can be alleviated effectively in high-performance PCB designs. Tsun-Ming Tseng, Bing Li 0005, Tsung-Yi Ho, Ulf Schlichtmann |
ICCAD | 4 |
| 2013 | A greedy approach for latency-bounded deadlock-free routing path allocation for application-specific NoCsabstractCustom network-on-chip (NoC) structures have improved power and area metrics compared to regular NoC topologies for application-specific systems-on-a-chip (SoCs). The synthesis of an application-specific NoC is a combinatorial problem. This paper presents a novel heuristic for solving the routing path allocation step. Its main advantages are the support of realistic nonlinear cost estimation and the ability to handle latency constraints, which guarantee high performance of processing elements sensitive to communication delays. Additionally, the method generates deadlock-free routing by avoiding cycles in the channel dependency graph. The NoC is constructed sequentially in a greedy manner by selecting the routing path for each communication flow in such a way that the additional NoC HW resources are kept minimal. The routing path is found using a binary search cheapest bounded path (BSCBP) algorithm. The method is highly efficient and provides a NoC routing path allocation for a smart phone SoC with 25 processing elements and 96 flows in less than a minute. Pritpal S. Multani, Daniel Mueller-Gritschneder, Vladimir Todorov, Ulf Schlichtmann |
NOCS | 5 |
| 2013 | On Timing Model Extraction and Hierarchical Statistical Timing AnalysisabstractIn this paper, we investigate the challenges of applying statistical static timing analysis in hierarchical design flow, where modules supplied by IP vendors are used to hide design details for IP protection and to reduce the complexity of design and verification. For the three basic circuit types, combinational, flip-flop-based, and latch-controlled, we propose methods for extracting timing models that contain interfacing and compressed internal constraints. Using these compact timing models, the runtime of full-chip timing analysis can be reduced, while circuit details from IP vendors are not exposed. We also propose a method for reconstructing correlation between modules during full-chip timing analysis. This correlation cannot be incorporated into timing models because it depends on the layout of the corresponding modules in the chip. In addition, we investigate how to apply the extracted timing models with the reconstructed correlation to evaluate the performance of the complete design. Experiments demonstrate that using the extracted timing models and reconstructed correlation full-chip timing analysis can be several times faster than applying the flattened circuit directly, while the accuracy of statistical timing analysis is still well maintained. Bing Li 0005, Ning Chen 0006, Yang Xu 0019, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | Characterization of the bistable ring PUFabstractThe bistable ring physical(ly) unclonable function (BR-PUF) is a novel electrical intrinsic PUF design for physical cryptography. FPGA prototyping has provided a proof-of-concept, showing that the BR-PUF could be a promising candidate for strong PUFs. However, due to the limitations (device resources, placement and routing) of FPGA prototyping, the effectiveness of a practical ASIC implementation of the BR-PUF could not be validated. This paper characterizes the BR-PUF further through transistor-level simulations. Based on process variation, mismatch, and noise models provided or suggested by industry, these simulations are able to provide predictions on the figures-of-merit of ASIC implementations of the BR-PUF. This paper also suggests a more secure way of using the BR-PUF based on its supply voltage sensitivity. Qingqing Chen 0004, György Csaba, Paolo Lugli, Ulf Schlichtmann, Ulrich Rührmair |
DATE | 4 |
| 2012 | Current source modeling for power and timing analysis at different supply voltagesabstractThis paper presents a new current source model (CSM) that allows to model noise on supply nets originating from CMOS logic cells. It also captures the influence of dynamic supply voltage changes on power consumption and cell delay. The CSM models n/pMOS blocks separately to reduce the complexity of model components. Compared with other CSMs, only two-dimensional tables are needed. This results in low characterization times and high simulation speed. Moreover, no re-characterization is needed for different supply voltages. The model is tested in a SPICE simulator. A reduction in transient simulation time by up to 53X was observed in the results, while the error in delay and current consumption was typically less than 3 percent. Christoph Knoth, Hela Jedda, Ulf Schlichtmann |
DATE | 3 |
| 2012 | Accurately timed transaction level models for virtual prototyping at high abstraction levelabstractTransaction level modeling (TLM) improves the simulation performance by raising the abstraction level. In the TLM 2.0 standard based on OSCI SystemC, a single transaction can transfer a large data block. Due to such high abstraction, a great amount of information becomes invisible and thus timing accuracy can be degraded heavily. We present a methodology to accurately time such block transactions and achieve high simulation performance at the same time. First, before abstraction, a profiling process is performed on an instruction set simulator (ISS). Driver functions that implement the transfer of the data blocks are simulated. Several techniques are employed to trace the exact start and end of the driver functions as well as HW usages. Thus, a profile library of those driver functions can be constructed. Then, the application programs are host-compiled and use a single transaction to transfer a data block. A strategy is presented that efficiently estimates the timing of block transactions based on the profile library. It is the first method that takes into account caching effects that influence the timing of block transactions. Moreover, it ensures overall timing accuracy when integrated in other SW timing tools for full system simulation. Experimental results show that the block transactions are accurately timed, with average error less than 1%. At the same time, the simulation gain can be up to three orders of magnitude. Kun Lu 0005, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DATE | 3 |
| 2012 | Automated construction of a cycle-approximate transaction level model of a memory controllerabstractTransaction level (TL) models are key to early design exploration, performance estimation and virtual prototyping. Their speed and accuracy enable early and rapid System-on-Chip (SoC) design evaluation and software development. Most devices have only register transfer level (RTL) models that are too complex for SoC simulation. Abstracting these models to TL ones, however, is a challenging task, especially when the RTL description is too obscure or not accessible. This work presents a methodology for automatically creating a TL model of an RTL memory controller component. The device is treated as a black box and a multitude of simulations is used to obtain results, showing its timing behavior. The results are classified into conditional probability distributions, which are reused within a TL model to approximate the RTL timing behavior. The presented method is very fast and highly accurate. The resulting TL model executes approximately 1200 times faster, with a maximum measured average timing offset error of 7.66%. Vladimir Todorov, Daniel Mueller-Gritschneder, Helmut Reinig, Ulf Schlichtmann |
DATE | 4 |
| 2012 | Schedulability Analysis for Processors with Aging-Aware Autonomic Frequency ScalingabstractWith the rapid progress in semiconductor technology and the shrinking of device geometries, the resulting processors are increasingly becoming prone to effects like aging and soft errors. As a processor ages, its electrical characteristics degrade, i.e., the switching times of its transistors increase. Hence, the processor cannot continue error-free operation at the same clock frequency and/or voltage for which it was originally designed. In order to mitigate such effects, recent research proposes to equip processors with special circuitry that automatically adapts its clock frequency in response to changes in its circuit-level timing properties (arising from changes in its electrical characteristics). From the point of view of tasks running on these processors, such autonomic frequency scaling(AFS) processors become slower as they gradually age. This leads to additional execution delay for tasks, which needs to be analyzed carefully, particularly in the context of hard real time or safety-critical systems. Hence, for real-time systems based on AFS processors, the associated schedulability analysis should be aging-aware which is a relatively unexplored topic so far. In this paper we propose a schedulability analysis framework that accounts such aging-induced degradation and changes in timing properties of the processor, when designing hard real-time systems. In particular, we address the schedulability and task mapping problem by taking a lifetime constraint of the system into account. In other words, the system should be designed to be fully operational (i.e., meet all deadlines) till a given minimum period of time (i.e., its lifetime). The proposed framework is based on an aging model of the processor which we discuss in detail. In addition to studying the effects of aging on the schedulability of real-time tasks, we also discuss its impact on task mapping and resource dimensioning. Alejandro Masrur, Philipp H. Kindt, Martin Becker 0001, Samarjit Chakraborty, Veit Kleeberger, Martin Barke, Ulf Schlichtmann |
RTCSA | 7 |
| 2012 | Statistical Timing Analysis for Latch-Controlled Circuits With Reduced Iterations and Graph TransformationsabstractLevel-sensitive latches are widely used in high-performance designs. For such circuits, efficient statistical timing analysis algorithms are needed to take increasing process variations into account. The existing methods for solving this problem are still computationally expensive and can only provide the yield at a given clock period. In this paper, we propose a method combining reduced iterations and graph transformations. The reduced iterations extract setup time constraints and identify a subgraph for the following graph transformations handling the constraints from nonpositive loops. The combined algorithms are very efficient, more than ten times faster than other existing methods, and result in a parametric minimum clock period, which, together with the hold-time constraints, can be used to compute the yield at any given clock period very easily. Bing Li 0005, Ning Chen 0006, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Control-Flow-Driven Source Level Timing Annotation for Embedded Software Models on Transaction LevelabstractInstrumented software models feature a combination of software functionality as well as timing information to model execution times on embedded processors. They aim to replace instruction set simulators in virtual prototypes (VP) of embedded systems to improve simulation efficiency. In this work, a novel control flow mapping algorithm is presented to automatically generate timing annotations for instrumented software models. The method is based on the analysis of loop and control dependency properties of basic code blocks in the binary and source code control flow. With these properties, the method can find suitable positions to annotate the timing delay statements of binary code basic blocks into the source code. It shows high accuracy even in the case that the binary code is optimized during compilation. The paper also presents the novel idea of adding timing control statements into the source code to improve timing accuracy. The error in runtime estimation was found to be below 6\% for standard test programs. A case study for a VP shows a gain in simulation efficiency of three orders of magnitude compared to an ISS based model. Daniel Mueller-Gritschneder, Kun Lu 0005, Ulf Schlichtmann |
DSD | 3 |
| 2011 | Comprehensive Generation of Hierarchical Placement Rules for Analog Integrated CircuitsabstractThis paper presents a new method to automatically generate hierarchical placement rules, which are crucial for a successful analog placement. The method is based on a novel symmetry computation method, introducing the structural signal flow graph. Five types of proximity, matching and symmetry constraints are determined. According to the priority of the constraint types, a constraint requirement graph and a hierarchical partitioning of the circuit into matching, proximity and symmetry groups is then automatically computed. Based on experimental results with a state-of-the-art placement tool, we show that the new approach generates more placement rules and can lead to better circuit performance and parametric yield according to post-layout simulation. Michael Eick, Martin Strasser, Kun Lu 0005, Ulf Schlichtmann, Helmut E. Graeb |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2010 | Fast statistical timing analysis of latch-controlled circuits for arbitrary clock periodsabstractLatch-controlled circuits have a remarkable advantage in timing performance as process variations become more relevant for circuit design. Existing methods of statistical timing analysis for such circuits, however, still need improvement in runtime and their results should be extended to provide yield information for any given clock period. In this paper, we propose a method combining a simplified iteration and a graph transformation algorithm. The result of this method is in a parametric form so that the yield for any given clock period can easily be evaluated. The graph transformation algorithm handles the constraints from nonpositive loops effectively, completely avoiding the heuristics used in other existing methods. Therefore the accuracy of the timing analysis is well maintained. Additionally, the proposed method is much faster than other existing methods. Especially for large circuits it offers about 100 times performance improvement in timing verification. Bing Li 0005, Ning Chen 0006, Ulf Schlichtmann |
ICCAD | 3 |
| 2010 | Aging analysis at gate and macro cell levelabstractAging, which can be regarded as a time-dependent variability, has until recently not received much attention in the field of electronic design automation. This is changing because increasing reliability costs threaten the continued scaling of ICs. We investigate the impact of aging effects on single combinatorial gates and present methods that help to reduce the reliability costs by accurately analyzing the performance degradation of aged circuits at gate and macro cell level. Dominik Lorenz, Martin Barke, Ulf Schlichtmann |
ICCAD | 3 |
| 2010 | Automatic generation of hierarchical placement rules for analog integrated circuitsabstractThis paper presents a new method to automatically generate hierarchical placement rules, which are crucial for a successful analog placement. Michael Eick, Martin Strasser, Helmut E. Graeb, Ulf Schlichtmann |
ISPD | 4 |
| 2010 | Towards Electrical, Integrated Implementations of SIMPL Systems
Ulrich Rührmair, Qingqing Chen 0004, Martin Stutzmann, Paolo Lugli, Ulf Schlichtmann, György Csaba |
WISTP | 5 |
| 2009 | On hierarchical statistical static timing analysisabstractStatistical static timing analysis deals with the increasing variations in manufacturing processes to reduce the pessimism in the worst case timing analysis. Because of the correlation between delays of circuit components, timing model generation and hierarchical timing analysis face more challenges than in static timing analysis. In this paper, a novel method to generate timing models for combinational circuits considering variations is proposed. The resulting timing models have accurate input-output delays and are about 80% smaller than the original circuits. Additionally, an accurate hierarchical timing analysis method at design level using pre-characterized timing models is proposed. This method incorporates the correlation between modules by replacing independent random variables to improve timing accuracy. Experimental results show that the correlation between modules strongly affects the delay distribution of the hierarchical design and the proposed method has good accuracy compared with Monte Carlo simulation, but is faster by three orders of magnitude. Bing Li 0005, Ning Chen 0006, Manuel Schmidt, Walter Schneider 0001, Ulf Schlichtmann |
DATE | 5 |
| 2009 | Digital design at a crossroads How to make statistical design methodologies industrially relevantabstractStatistical analysis is generally seen as the next EDA technology for timing and power sign-off. Research into this field has seen significant activity started about five years ago. Recently, interest appears to have fallen off somewhat. Also, while a lot of focus has been put on research fundamentals, extremely few applications in industry have been reported so far. Therefore, a group including Infineon Technologies as a leading semiconductor IDM and various universities and research institutes, as well as an EDA provider has tackled key challenges to enable statistical design in industry in a publicly funded project called ldquoSigma65rdquo. Sigma65 strives to provide key foundations to allow a change from traditional deterministic design methods to future design methods driven by statistical considerations. The project starts with statistical modeling and optimization of library components and ranges to statistical techniques for designing ICs on gate level and higher levels. In this paper, we present some results of this project, demonstrating how the interaction between industrial perspective, research institutions and EDA provider enables solutions which are applicable already in the near future. After an overview of the industrial perspective of the current situation in dealing with variations recent results on both statistical timing and power analysis will be given. In addition, recent research advances on fast yield estimation concerning parametric timing yield will be given. Ulf Schlichtmann, Manuel Schmidt, Harald Kinzelbach, Michael Pronath, Volker Gloeckel, Manfred Dietrich, Uwe Eichler, Joachim Haase |
DATE | 1 |
| 2009 | Timing model extraction for sequential circuits considering process variationsabstractAs semiconductor devices continue to scale down, process variations become more relevant for circuit design. Facing such variations, statistical static timing analysis is introduced to model variations more accurately so that the pessimism in traditional worst case timing analysis is reduced. Because all delays are modeled using correlated random variables, most statistical timing methods are much slower than corner based timing analysis. To speed up statistical timing analysis, we propose a method to extract timing models for flip-flop and latch based sequential circuits respectively. When such a circuit is used as a module in a hierarchical design, the timing model instead of the original circuit is used for timing analysis. The extracted timing models are much smaller than the original circuits. Experiments show that using extracted timing models accelerates timing verification by orders of magnitude compared to previous approaches using flat netlists directly. Accuracy is maintained, however, with the mean and standard deviation of the clock period both showing usually less than 1% error compared to Monte Carlo simulation on a number of benchmark circuits. Bing Li 0005, Ning Chen 0006, Ulf Schlichtmann |
ICCAD | 3 |
| 2009 | Aging analysis of circuit timing considering NBTI and HCIabstractWe present an aging analysis flow able to calculate the degraded circuit timing. To the best of our knowledge it is the first approach on gate level so far capable of analyzing the impact of the two dominant drift-related aging effects - NBTI and HCI - on complex digital circuits. The aging-aware gate model used to compute the aged circuit timing provides not just the cell delay degradation, but also the degradation of the output slope. To get more accurate results, the individual workload of a gate can be considered. Dominik Lorenz, Georg Georgakos, Ulf Schlichtmann |
IOLTS | 3 |
| 2008 | Sizing Rules for Bipolar Analog Circuit DesignabstractThis paper presents sizing rules for basic building blocks in analog bipolar circuit design. Sizing rules efficiently capture design knowledge on the technology-specific level of transistor-pair groups. This reduces the effort for and improves the resulting quality of analog circuit synthesis. We present a hierarchical library of transistor-pair groups as basic building blocks for analog bipolar circuits. Sizing rules are constraints associated to these building blocks that must be satisfied to guarantee the function and robustness of each block. Results of applications like circuit sizing or design centering show that the use of sizing rules leads to improved and robust results. Tobias Massier, Helmut E. Graeb, Ulf Schlichtmann |
DATE | 3 |
| 2008 | Deterministic analog circuit placement using hierarchically bounded enumeration and enhanced shape functionsabstractThe analog placement algorithm Plantage, presented in this paper, generates placements for analog circuits with comprehensive placement constraints. Plantage is based on a hierarchically bounded enumeration of basic building blocks, using B*-trees. The practically relevant solution space is thereby enumerated quasi-complete. The sets of possible placements of the basic building blocks are represented and combined in a new efficient way, using enhanced shape functions. The result of Plantage is the Pareto front of placements with respect to different aspect ratios. The whole approach is deterministic, in contrast to existing analog placement algorithms. Martin Strasser, Michael Eick, Helmut E. Graeb, Ulf Schlichtmann, Frank M. Johannes |
ICCAD | 4 |
| 2008 | A random and pseudo-gradient approach for analog circuit sizing with non-uniformly discretized parametersabstractMany methods for analog circuit sizing are available as commercial, in-house and academic tools. They are based on continuous optimization, e.g., of transistor geometries, although the subsequent layout step requires values on a pre-defined grid. In addition, sizing of transistors for bipolar and RF circuits frequently necessitates the use of multiples of predefined values for the design parameters. This paper presents a novel method for solving this type of discrete optimization problem. An iterative approach is presented, which is based on pseudo-gradients and a randomized calculation of search regions and steps. Experimental comparisons with simulated annealing and a continuous sizing approach with subsequent discretization clearly show the effectivity and efficiency of the presented method. Michael Pehl, Tobias Massier, Helmut E. Graeb, Ulf Schlichtmann |
ICCD | 4 |
| 2008 | Abacus: fast legalization of standard cell circuits with minimal movementabstractStandard cell circuits consist of millions of standard cells, which have to be aligned overlap-free to the rows of the chip. Placement of these circuits is done in consecutive steps. First, a global placement is obtained by roughly spreading the cells on the chip, while considering all relevant objectives like wirelength, and routability. After that, the global placement is legalized, i.e., the cell overlap is removed, and the cells are aligned to the rows. To preserve the result of global placement, cells should be moved as little as possible during legalization Peter Spindler, Ulf Schlichtmann, Frank M. Johannes |
ISPD | 2 |
| 2008 | The Sizing Rules Method for CMOS and Bipolar Analog Integrated Circuit SynthesisabstractThis paper presents the sizing rules method for basic building blocks in analog CMOS and bipolar circuit design. It consists of the development of a hierarchical library of transistor-pair groups as basic building blocks for analog CMOS and bipolar circuits, the derivation of a hierarchical generic list of constraints that must be satisfied to guarantee the function and robustness of each block, and the development of a reliable automatic recognition procedure of building blocks in a circuit schematic. Sizing rules efficiently capture design knowledge on the technology-specific level of transistor-pair groups. This reduces the effort and improves the resulting quality for analog circuit synthesis. Results of applications like circuit sizing, design centering, response surface modeling, or analog placement show the benefits of the sizing rules method. Tobias Massier, Helmut E. Graeb, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | Kraftwerk2 - A Fast Force-Directed Quadratic Placement Approach Using an Accurate Net ModelabstractThe force-directed quadratic placer ldquoKraftwerk2,rdquo as described in this paper, is based on two main concepts. First, the force that is necessary to distribute the modules on the chip is separated into the following two components: a hold force and a move force. Both components are implemented in a systematic manner. Consequently, Kraftwerk2 converges such that the module overlap is reduced in each placement iteration. The second concept of Kraftwerk2 is to use the ldquoBound2Boundrdquo net model, which accurately represents the half-perimeter wirelength (HPWL) in the quadratic cost function. Aside from these features, this paper presents additional details about Kraftwerk2. An approach to remove halos (free space) around large modules is described, and a method to control the module density is presented. In order to choose the important tradeoff between runtime and quality, a systematic quality control is shown. Furthermore, plots demonstrating the convergence of Kraftwerk2 are presented. Results using various benchmark suites demonstrate that Kraftwerk2 offers both high quality and excellent computational efficiency. Peter Spindler, Ulf Schlichtmann, Frank M. Johannes |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Trade-off design of analog circuits using goal attainment and "Wave Front" sequential quadratic programmingabstractOne of the main tasks in analog design is the sizing of the circuit parameters, such as transistor lengths and widths, in order to obtain optimal circuit performances, such as high gain or low power consumption. In most cases one performance can only be optimized at cost of others, therefore a sizing must aim at an optimal trade-off between the important circuit performances. This paper presents a new deterministic method to calculate the complete range of performance trade-offs, the so-called Pareto-optimal front, of a given circuit topology. Known deterministic methods solve a set of constrained multi-objective optimization problems independently of each other. The presented method minimizes a set of goal attainment (GA) optimization problems simultaneously. In a parallel algorithm, the individual GA optimization processes compare and exchange their iterative solutions. This leads to a significant improvement in the efficiency and quality of analog trade-off design Daniel Mueller-Gritschneder, Helmut E. Graeb, Ulf Schlichtmann |
DATE | 3 |
| 2006 | A CPPLL hierarchical optimization methodology considering jitter, power and locking timeabstractIn this paper, a hierarchical optimization methodology for charge pump phase-locked loops (CPPLLs) is proposed. It has the following features: 1) A comprehensive and efficient behavioral modeling of the PLL enables fast simulations and includes the important PLL performances jitter, power and locking time, as well as stability constraints for the nonlinear locking process and the linear lock-in state; 2) Behavioral modeling of the PLL building blocks addresses as behavioral-level parameters: current and jitter of the charge pump (CP), gain, current and jitter of the voltage controlled oscillator (VCO), as well as R, C's of the loop filter (LF). It enables a proper propagation of PLL specifications down to the circuit-level design parameters; 3) An accurate and efficient performance space exploration technique on circuit level provides the feasible regions of the behavioral-level parameters of the building block by multidimensional Pareto-optimal fronts. This enables a first-time-successful top-down optimization process. Experimental results show the efficacy and efficiency of the presented method. The methodology can be applied to other large-scale analog circuits. Daniel Mueller-Gritschneder, Helmut E. Graeb, Ulf Schlichtmann |
DAC | 4 |
| 2006 | DFM/DFY design for manufacturability and yield - influence of process variations in digital, analog and mixed-signal circuit designabstractThe concepts of design for manufacturability and design for yield DFM/DFY are bringing together domains that co-existed mostly separated until now $circuit design, physical design and manufacturing process. New requirements like SoC, mixed analog/digital design and deep-submicron technologies force to a mutual integration of all levels. A major challenge coming with new deep-submicron technologies is to design and verify integrated circuits for high yield. Random and systematic defects as well as parametric process variations have a large influence on quality and yield of the designed and manufactured circuits. With further shrinking of process technology, the on-chip variation is getting worse for each technology node. For technologies larger than 180nm feature sizes, variations are mostly in a range of below 10%. Here an acceptable yield range is achieved by regular but error-prone re-shifts of the drifting process. However, shrinking technologies down to 90nm, 65nm and below cause on-chip variations of more than 50%. It is understandable that tuning the technology process alone is not enough to guarantee sufficient yield and robustness levels any more. Redesigns and, therefore, respins of the whole development and manufacturing chain lead to high costs of multiple manufacturing runs. All together the risk to miss the given market window is extremely high. Thus, it becomes inevitable to have a seamless DFM/DFY concept realized for the design phase of digital, analog, and mixed-signal circuits. New DFY methodologies are coming up for parametric yield analysis and optimization and have recently been made available for the industrial design of individual analog blocks on transistor level up to 1500 transistors. The transfer of yield analysis and yield optimization techniques to other abstraction levels - both for digital as well as for analog - is a big challenge. Yield analysis and optimization is currently applied to individual circuit blocks and not to the overall chip yielding on the one hand often too pessimistic results - best/worst case and OCV (on chip variation) factor - for the digital parts. On the other hand for analog often very high efforts are spent to design individual blocks with high robustness (>6sigma). For abstraction to higher digital levels first approaches like statistical static timing analysis (SSTA) are under development. For the analog parts a strategy to develop macro models and hierarchical simulation or behavioral simulation methodologies is required that includes low-level statistical effects caused by local and global process variation of the individual devices Markus Bühler, Jürgen Koehl, Jeanne Bickford, Jason Hibbeler, Ulf Schlichtmann, Ralf Sommer, Michael Pronath, Andreas Ripp |
DATE | 5 |
| 2006 | Fast evaluation of analog circuit structures by polytopal approximationsabstractIn this paper we present a method for the fast evaluation of circuit structures. It is part of a methodology for the structural synthesis of analog circuits which generates a large number of different circuit structures. Goal of the presented methods is to find circuit structures, which fit best the design goals. Based on implicit analog circuit specifications, as well as explicit performance specifications given by the designer, the presented method approximates the feasible region of parameters by a polytope. This polytopal approximation of the performance capabilities can be calculated and visualized or the feasibility of the specification can be tested by linear programming. The method has been validated on a set of operational amplifier structures Daniel Mueller-Gritschneder, Guido Stehr, Helmut E. Graeb, Ulf Schlichtmann |
ISCAS | 4 |
| 2005 | Deterministic approaches to analog performance space exploration (PSE)abstractPerformance space exploration (PSE) determines the range of feasible performance values of a circuit block for a given topology and technology. In this paper, we present two deterministic approaches for PSE. One approximates the feasible performance space based on linearized circuit models and is suitable for investigating a large number of performances. The other one computes discretizations of the Pareto front of competing performances. In addition, a motivation and application of PSE using a hierarchical design example is presented. Daniel Mueller-Gritschneder, Guido Stehr, Helmut E. Graeb, Ulf Schlichtmann |
DAC | 4 |
| 2004 | Extremely Low-Power LogicabstractFor extremely low-power logic, three very new and promising techniques will be described. The first are methods on circuit and system level for reduced supply voltages. In large logic blocks, interconnect becomes a main issue, that could be solved by on-chip optical interconnect. Nano-devices will also be presented, as a possibility to compute with nearly zero power, and compared to future 10 nanometers transistors. Christian Piguet, Jacques Gautier, Christoph Heer, Ian O'Connor, Ulf Schlichtmann |
DATE | 5 |
| 2004 | Design Methodology Innovations Address Manufacturing Technology Challenges: Power and PerformanceabstractSemiconductor design has benefited tremendously from process technology scaling in the past, especially for power consumption and performance. This era is coming to an end. Continued improvement in these key metrics requires even more innovation in design methodology and design automation than in the past. Power consumption increasingly is becoming the most important bottleneck in the design of ICs in advanced process technologies. An evaluation of the use ultra-low threshold voltage (V/sub th/) devices for power reduction in an advanced process technology is described. It turns out that in contrast to older process technologies, this approach increasingly is becoming less suitable for industrial usage in advanced process technologies. Thus, design methodologies are described which can reduce power consumption by optimizations in logic design, specifically by utilizing multiple levels of supply voltage V/sub dd/ and threshold voltage V/sub th/. The next major challenge on the horizon is increasing variability, both in the manufacturing process and in operating conditions. The need for statistical approaches to counter rising variability is described. Ulf Schlichtmann |
DSD | 1 |
| 2002 | Power Crisis in SoC Design: Strategies for Constructing Low-Power, High-Performance SoC DesignsabstractThis special panel session brings together several leading technologists to discuss the challenges and solutions in constructing SoC designs that achieve their performance goals within a very tight power budget. These challenges are addressed from the often conflicting perspectives of semiconductor design teams and commercial solutions providers of EDA construction tools, EDA analysis tools and semiconductor IP (SIP). K. Brock, C. Edwards, R. Lannoo, Ulf Schlichtmann, Antun Domic, Jacques Benkoski, David Overhauser, M. Kliment |
DATE | 4 |
| 2002 | Systems Are Made from Transistors: UDSM Technology Creates New Challenges for Library and IC DevelopmentabstractThe progress of silicon process technology relentlessly marches on. Moore's law still holds, the number of transistors that can be integrated on an IC doubles approximately every 18 months. The inability of system designs to keep up with this ever increasing number of available transistors has been debated for a long time, many solutions have been proposed. Now, as 130 nm processes enter volume production, 90 nm yields first engineering samples, and 65 nm processes are being developed, the design productivity crisis is exacerbated by the fact that very difficult design challenges are inherent in Ultra-Deep Submicron (UDSM) technologies. They threaten the approach of abstracting technological features away at higher levels, thus endangering design productivity even more. This paper outlines current challenges, presents approaches to address them and proposes further areas for research. Ulf Schlichtmann |
DSD | 1 |
| 1999 | Functional multiple-output decomposition with application to technology mapping for lookup table-based FPGAsabstractFunctional decomposition is an important technique for technology mapping to look up table-based FPGA architectures. We present the theory of and a novel approach to functional disjoint decomposition of multiple-output functions, in which common subfunctions are extracted during technology mapping. While a Boolean function usually has a very large number of subfunctions, we show that not all of them are useful for multiple-output decomposition. We use a partition of the set of bound set vertices as the basis to computepreferabledecomposition functions, which are sufficient for an optimal multiple-output decomposition. We propose several new algorithms that deal with central issues of functional multiple-output decomposition. First, an efficient algorithm to solve the variable partitioning problem is described. Second, we show how to implicitly compute all preferable functions of a single-output function and how to identify all common preferable functions of a multiple-output function. Due to implicit computation in the crucial steps, the algorithm is very efficient. Experimental results show significant reductions in area. Bernd Wurth, Ulf Schlichtmann, Klaus Eckl, Kurt Antreich |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 1992 | Characterization of Boolean Functions for Rapid Matching in FPGA Technology Mapping
Ulf Schlichtmann, Franc Brglez, Michael Hermann |
DAC | 1 |