VLDB 2026 Research / reviewers in the wild / expert
Rasit Onur Topaloglu
dblp:75/6468
· DBLP profile ↗
34ranked-venue papers
13as first author
14since 2021 · last 2026
0000-0001-8759-6959ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 13 first-author · 14 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results - Feature-Aware Trojan Alteration to Evade ML-Based Detection
Lizi Zhang, Navid Nader Tehrani, Azadeh Davoodi, Rasit Onur Topaloglu |
VTS | 4 |
| 2025 | ReBERT: LLM for Gate-Level to Word-Level Reverse EngineeringabstractIn this paper, we introduce ReBERT, a specialized large language model (LLM) based on BERT, fine-tuned specifically for grouping bits into words within gate-level netlists. By treating the netlist as a form of language, we encode bits and their fan-in cones into sequences that capture structural dependencies. A novel contribution is augmenting BERT's embedding with a tree-based embedding strategy which mirrors the hierarchical nature of circuit designs in hardware. Leveraging the powerful representational learning capabilities of LLMs, we interpret hardware circuits at a higher level of abstraction. We evaluate ReBERT on various hardware designs, demonstrating that it significantly outperforms a state-of-the-art work based on partial structural matching in recovering word-level groupings. Our improvements are on average between 12.2% to 218.1% depending on degree of corrupting the structural patterns. Lizi Zhang, Azadeh Davoodi, Rasit Onur Topaloglu |
DATE | 3 |
| 2023 | ObfusX: Routing obfuscation with explanatory analysis of a machine learning attack
Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu |
Integr. | 3 |
| 2023 | Florets for Chiplets: Data Flow-aware High-Performance and Energy-efficient Network-on-Interposer for CNN Inference TasksabstractRecent advances in 2.5D chiplet platforms provide a new avenue for compact scale-out implementations of emerging compute- and data-intensive applications including machine learning. Network-on-Interposer (NoI) enables integration of multiple chiplets on a 2.5D system. While these manycore platforms can deliver high computational throughput and energy efficiency by running multiple specialized tasks concurrently, conventional NoI architectures have a limited computational throughput due to their inherent multi-hop topologies. In this paper, we propose Floret, a novel NoI architecture based on space-filling curves (SFCs). The Floret architecture leverages suitable task mapping, exploits the data flow pattern, and optimizes the inter-chiplet data exchange to extract high performance for multiple types of convolutional neural network (CNN) inference tasks running concurrently. We demonstrate that the Floret architecture reduces the latency and energy up to 58% and 64%, respectively, compared to state-of-the-art NoI architectures while executing datacenter-scale workloads involving multiple CNN tasks simultaneously. Floret achieves high performance and significant energy savings with much lower fabrication cost by exploiting the data-flow awareness of the CNN inference tasks. Lukas Pfromm, Rasit Onur Topaloglu, Janardhan Rao Doppa, Ümit Y. Ogras, Anantharaman Kalyanaraman, Partha Pratim Pande |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | Muzzle the Shuttle: Efficient Compilation for Multi-Trap Trapped-Ion Quantum ComputersabstractTrapped-ion systems can have a limited number of ions (qubits) in a single trap. Increasing the qubit count to run meaningful quantum algorithms would require multiple traps where ions need to shuttle between traps to communicate. The existing compiler has several limitations, which result in a high number of shuttle operations and degraded fidelity. In this paper, we target this gap and propose compiler optimizations to reduce the number of shuttles. Our technique achieves a maximum reduction of 51.17% in shuttles (average ~ 33%) tested over 125 circuits. Furthermore, the improved compilation enhances the program fidelity up to 22.68X with a modest increase in the compilation time. Abdullah Ash-Saki, Rasit Onur Topaloglu, Swaroop Ghosh |
DATE | 2 |
| 2022 | A Shuttle-Efficient Qubit Mapper for Trapped-Ion Quantum ComputersabstractTrapped-ion (TI) quantum computer is one of the forerunner quantum technologies. Execution of a quantum gate in multiple trap TI system may frequently involve ions from two different traps, hence one of the ions needs to be shuttled (moved) between traps to be co-located, degrading fidelity, and increasing the program execution time. The choice of initial mapping influences the number of shuttles. The existing Greedy policy neglects the depth of the program at which a gate is present. Intuitively, the contribution of the late-stage gates to the initial mapping is less since the ions might have already shuttled to a different trap to satisfy other gate operations. In this paper, we target this gap and propose a new program adaptive policy especially for programs with considerable depth and high number of qubits (valid for practical-scale quantum programs). Our technique achieves an average reduction of 9% shuttles/program (with 21.3% at best) for 120 random circuits and enhances the program fidelity up to 3.3X (1.41X on average). Suryansh Upadhyay, Abdullah Ash-Saki, Rasit Onur Topaloglu, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Quantum Machine Learning for Material Synthesis and Hardware Security (Invited Paper)abstractUsing quantum computing, this paper addresses two scientifically-pressing and day to day-relevant problems, namely, chemical retrosynthesis which is an important step in drug/material discovery and security of semiconductor supply chain. We show that Quantum Long Short-Term Memory (QLSTM) is a viable tool for retrosynthesis. We achieve 65% training accuracy with QLSTM whereas classical LSTM can achieve 100%. However, in testing we achieve 80% accuracy with the QLSTM while classical LSTM peaks at only 70% accuracy! We also demonstrate an application of Quantum Neural Network (QNN) in the hardware security domain, specifically in Hardware Trojan (HT) detection using a set of power and area Trojan features. The QNN model achieves detection accuracy as high as 97.27%. Collin Beaudoin, Satwik Kundu, Rasit Onur Topaloglu, Swaroop Ghosh |
ICCAD | 3 |
| 2022 | Optimization of Quantum Read-Only Memory CircuitsabstractQuantum computing is a rapidly expanding field with applications ranging from optimization all the way to complex machine learning tasks. Quantum memories, while lacking in practical quantum computers, have the potential to bring quantum advantage. In quantum machine learning applications for example, a quantum memory can simplify the data loading process and potentially accelerate the learning task. Quantum memory can also store intermediate quantum state of qubits that can be reused for computation. However, the depth, gate count and compilation time of quantum memories such as, Quantum Read Only Memory (QROM) scale exponentially with the number of address lines making them impractical in state-of-the-art Noisy Intermediate-Scale Quantum (NISQ) computers beyond 4-bit addresses. In this paper, we propose techniques such as, pre-decoding logic and qubit reset to reduce the depth and gate count of QROM circuits to target wider address ranges such as, 8-bits. The proposed approach reduces the number of gates and depth count by at least 2X compared to the naive implementation at only 36% qubit overhead. A reduction in circuit depth and gate count as high as 75X and compilation time by 85X at the cost of a maximum of 2.28X qubit overhead is observed. Experimentally, the fidelity with the proposed pre-decoding circuit compared to existing optimization approach is also higher (as much as 73% compared to 40.8%) under reduced error rates. Koustubh Phalak, Mahabubul Alam, Abdullah Ash-Saki, Rasit Onur Topaloglu, Swaroop Ghosh |
ICCD | 4 |
| 2021 | ObfusX: Routing Obfuscation with Explanatory Analysis of a Machine Learning AttackabstractThis is the first work that incorporates recent advancements in "explainability" of machine learning (ML) to build a routing obfuscator called ObfusX. We adopt a recent metric---the SHAP value---which explains to what extent each layout feature can reveal each unknown connection for a recent ML-based split manufacturing attack model. The unique benefits of SHAP-based analysis include the ability to identify the best candidates for obfuscation, together with the dominant layout features which make them vulnerable. As a result, ObfusX can achieve better hit rate (97% lower) while perturbing significantly fewer nets when obfuscating using a via perturbation scheme, compared to prior work. When imposing the same wirelength limit using a wire lifting scheme, ObfusX performs significantly better in performance metrics (e.g., 2.4 times more reduction on average in percentage of netlist recovery). Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu |
ASP-DAC | 3 |
| 2021 | Logic Synthesis Meets Machine Learning: Trading Exactness for GeneralizationabstractLogic synthesis is a fundamental step in hardware design whose goal is to find structural representations of Boolean functions while minimizing delay and area. If the function is completely-specified, the implementation accurately represents the function. If the function is incompletely-specified, the implementation has to be true only on the care set. While most of the algorithms in logic synthesis rely on SAT and Boolean methods to exactly implement the care set, we investigate learning in logic synthesis, attempting to trade exactness for generalization. This work is directly related to machine learning where the care set is the training set and the implementation is expected to generalize on a validation set. We present learning incompletely-specified functions based on the results of a competition conducted at IWLS 2020. The goal of the competition was to implement 100 functions given by a set of care minterms for training, while testing the implementation using a set of validation minterms sampled from the same function. We make this benchmark suite available and offer a detailed comparative analysis of the different approaches to learning. Shubham Rai, Walter Lau Neto, Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi, Masahiro Fujita 0004, Guilherme B. Manske, Matheus F. Pontes, Leomar S. da Rosa Jr., Marilton S. de Aguiar, Paulo F. Butzen, Po-Chun Chien, Yu-Shan Huang, Hoa-Ren Wang, Jie-Hong Roland Jiang, Jiaqi Gu 0002, Zheng Zhao 0003, Zixuan Jiang, David Z. Pan, Brunno Abreu, Isac de Souza Campos, Augusto Andre Souza Berndt, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi, Aditya Lohana, Akash Kumar 0001, Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu, Jordan Dotzel, Yichi Zhang 0006, Hanyu Wang 0005, Zhiru Zhang, Valerio Tenace, Pierre-Emmanuel Gaillardon, Alan Mishchenko, Satrajit Chatterjee |
DATE | 32 |
| 2021 | A Survey and Tutorial on Security and Resilience of Quantum ComputingabstractPresent-day quantum computers suffer from various noises or errors such as, gate error, relaxation, dephasing, readout error, and crosstalk. Besides, they offer a limited number of qubits with restrictive connectivity. Therefore, quantum programs running these computers face resilience issues and low output fidelities. The noise in the cloud-based access of quantum computers also introduce new modes of security and privacy issues. Furthermore, quantum computers face several threat models from insider and outsider adversaries including input tampering, program misallocation, fault injection, Reverse Engineering (RE) and Cloning. This paper provides an overview of various assets embedded in quantum computers and programs, vulnerabilities and attack models and the relation between resilience and security. We also cover countermeasures against the reliability and security issues and present future outlook for security of quantum computing. Abdullah Ash-Saki, Mahabubul Alam, Koustubh Phalak, Aakarshitha Suresh, Rasit Onur Topaloglu, Swaroop Ghosh |
ETS | 5 |
| 2021 | Sampling-Based Approximate Logic Synthesis: An Explainable Machine Learning ApproachabstractRecent years have seen promising studies on machine learning (ML) techniques applied to approximate logic synthesis (ALS), especially based on logic reconstruction from samples of input-output pairs. This “sampling-based ALS” supports integration with conventional logic synthesis and optimization techniques, as well as synthesis for a constrained input space (e.g., when primary input values are restricted using Boolean relations). To achieve an effective sampling-based ALS, for the first time, this paper proposes the use of adaptive decision trees (ADTs), and in particular variations guided by explainable ML. We adopt SHAP importance, which is a feature importance metric derived from a recent advance in explainable ML to guide the training of ADTs. We also include approximation techniques for ADT which are specifically designed for ALS, including don't-care bit assertion and instantiation. Comprehensive experiments show that we can achieve 39%-42% area reduction with 0.20%-0.22% error rate on average, based on 15 logic functions in the IWLS'20 benchmark suite. Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu |
ICCAD | 3 |
| 2021 | Quantum-Classical Hybrid Machine Learning for Image Classification (ICCAD Special Session Paper)abstractImage classification is a major application domain for conventional deep learning (DL). Quantum machine learning (QML) has the potential to revolutionize image classification. In any typical DL-based image classification, we use convolutional neural network (CNN) to extract features from the image and multi-layer perceptron network (MLP) to create the actual decision boundaries. QML models can be useful in both of these tasks. On one hand, convolution with parameterized quantum circuits (Quanvolution) can extract rich features from the images. On the other hand, quantum neural network (QNN) models can create complex decision boundaries. Therefore, Quanvolution and QNN can be used to create an end-to-end QML model for image classification. Alternatively, we can extract image features separately using classical dimension reduction techniques such as, Principal Components Analysis (PCA) or Convolutional Autoen-coder (CAE) and use the extracted features to train a QNN. We review two proposals on quantum-classical hybrid ML models for image classification namely, Quanvolutional Neural Network and dimension reduction using a classical algorithm followed by QNN. Particularly, we make a case for trainable filters in Quanvolution and CAE-based feature extraction for image datasets (instead of dimension reduction using linear transformations such as, PCA). We discuss various design choices, potential opportunities, and drawbacks of these models. We also release a Python-based framework to create and explore these hybrid models with a variety of design choices. Mahabubul Alam, Satwik Kundu, Rasit Onur Topaloglu, Swaroop Ghosh |
ICCAD | 3 |
| 2021 | Split Compilation for Security of Quantum CircuitsabstractAn efficient quantum circuit (program) compiler aims to minimize the gate-count - through efficient instruction translation, routing, gate, and cancellation - to improve run-time and noise. Therefore, a high-efficiency compiler is paramount to enable the game-changing promises of quantum computers. To date, the quantum computing hardware providers are offering a software stack supporting their hardware. However, several third-party software toolchains, including compilers, are emerging. They support hardware from different vendors and potentially offer better efficiency. As the quantum computing ecosystem becomes more popular and practical, it is only prudent to assume that more companies will start offering software-as-a-service for quantum computers, including high-performance compilers. With the emergence of third-party compilers, the security and privacy issues of quantum intellectual properties (IPs) will follow. A quantum circuit can include sensitive information such as critical financial analysis and proprietary algorithms. Therefore, submitting quantum circuits to untrusted compilers creates opportunities for adversaries to steal IPs. In this paper, we present a split compilation methodology to secure IPs from untrusted compilers while taking advantage of their optimizations. In this methodology, a quantum circuit is split into multiple parts that are sent to a single compiler at different times or to multiple compilers. In this way, the adversary has access to partial information. With analysis of over 152 circuits on three IBM hardware architectures, we demonstrate the split compilation methodology can completely secure IPs (when multiple compilers are used) or can introduce factorial time reconstruction complexity while incurring a modest overhead (~ 3% to ~ 6% on average). Abdullah Ash-Saki, Aakarshitha Suresh, Rasit Onur Topaloglu, Swaroop Ghosh |
ICCAD | 3 |
| 2020 | Explainable DRC Hotspot Prediction with Random Forest and SHAP Tree ExplainerabstractWith advanced technology nodes, resolving design rule check (DRC) violations has become a cumbersome task, which makes it desirable to make predictions at earlier stages of the design flow. In this paper, we show that the Random Forest (RF) model is quite effective for the DRC hotspot prediction at the global routing stage, and in fact significantly outperforms recent prior works, with only a fraction of the runtime to develop the model. We also propose, for the first time, to adopt a recent explanatory metric-the SHAP value-to make accurate and consistent explanations for individual DRC hotspot predictions from RF. Experiments show that RF is 21%-60% better in predictive performance on average, compared with promising machine learning models used in similar works (e.g. SVM and neural networks) while exhibiting good explainability, which makes it ideal for DRC hotspot prediction. Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu |
DATE | 3 |
| 2018 | Editorial for TODAES Special Issue on Internet of Things System Performance, Reliability, and SecurityabstractNo abstract available. Rasit Onur Topaloglu, Farinaz Koushanfar |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2017 | Editorial for JETC Special Issue on Alternative Computing SystemsabstractNo abstract available. Rasit Onur Topaloglu, Naveen Verma |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2016 | ICCAD-2016 CAD contest in pattern classification for integrated circuit design space analysis and benchmark suiteabstractLayout pattern classification has been utilized in recent years in integrated circuit design towards various goals such as design space analysis, design rule generation, and systematic yield optimization. There is a need for open source or academic solutions as very limited vendors are available to provide this functionality. Speed and accuracy are key aspects to target in the solutions. Given a circuit layout and various markers, contestants are asked to provide a reduced set of representative layout clips around these markers. Each such representative clip identifies a class and has an associated set of one or more unique layout markers. Rasit Onur Topaloglu |
ICCAD | 1 |
| 2014 | Design and technology co-optimization near single-digit nodesabstractAs we approach single-digit nodes, traditional design for manufacturability is augmented through several methodologies and design paradigms such as design-technology co-optimization (DTCO), systematic yield limiters optimization (SYLO), and design retargeting. We discuss triple-patterning and spacer-based multiple patterning and their design implications as these technologies will be necessary to cruise us to single-digit nodes. With the help of DTCO, there seems to be a clear path to sub-10nm with or without extreme ultra-violet lithography. Lars Liebmann, Rasit Onur Topaloglu |
ICCAD | 2 |
| 2014 | ICCAD-2014 CAD contest in design for manufacturability flow for advanced semiconductor nodes and benchmark suiteabstractWe introduce the fill optimization problem and benchmarks. We provide two new hotspot definitions, slot line deviation and outliers, both of which pertain to yield. We provide the inputs, expected output, as well as objectives and constraints of the problem. Rasit Onur Topaloglu |
ICCAD | 1 |
| 2014 | Guest Editorial Special Section on Optical InterconnectsabstractIncreasing delay of interconnects over devices at advanced nodes indicates a need for drastic change for interconnect. Such drastic changes have traditionally been possible through material changes. However, an alternative change in design may be near. In particular, on-chip optical interconnections are on the verge of being introduced to be able to cope with system-scale performance roadmaps. Thereby, photon-based information transfer may take the place of charge-based transfer that has been utilized within and across semiconductor chips for decades. Such a change would trigger significant updates to the design flows and circuitry. Rasit Onur Topaloglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Design with FinFETs: design rules, patterns, and variabilityabstractFinFETs have proven to be the device of choice for the next few technology generations. Consequently, design rules and limitations related to FinFETs need to be carefully understood. We present restricted and gridded design rules related to FinFETs. We also present results which indicate that a complete front-end-of-line (FEOL) and middle-of-line (MOL) of a memory with controllers can be designed from a fixed set of patterns. We also indicate sources of variations in FinFETs. Rasit Onur Topaloglu |
ICCAD | 1 |
| 2012 | Block-level 3D IC design with through-silicon-via planningabstractSince re-designing and re-optimizing existing logic, memory, and IP blocks in a 3D fashion significantly increases design cost, near-term three-dimensional integrated circuit (3D IC) design will focus on reusing existing 2D blocks. One way to reuse 2D blocks in the 3D IC design is to first perform 3D floorplanning, insert signal through-silicon vias (TSVs) for 3D inter-block connections, and then route the blocks. In this paper, we propose algorithms (finding signal TSV locations, assigning TSVs to whitespace blocks, and manipulating whitespace blocks) for post-floorplanning signal TSV planning in the block-level 3D IC design. Experimental results show that our signal TSV planner outperforms the state-of-the-art TSV-aware 3D floorplanner by 7% to 38% with respect to wirelength. In addition, our multiple TSV insertion algorithm outperforms a single TSV insertion algorithm by 27% to 37%. Dae Hyun Kim 0004, Rasit Onur Topaloglu, Sung Kyu Lim |
ASP-DAC | 2 |
| 2012 | Efficient pattern relocation for EUV blank defect mitigationabstractBlank defect mitigation is a critical step for extreme ultraviolet (EUV) lithography. Targeting the defective blank, a layout relocation method, to shift and rotate the whole layout pattern to a proper position, has been proved to be an effective way to reduce defect impact. Yet, there is still no published work about how to find the best pattern location to minimize the impact from the buried defects with reasonable defect model and considerable process variation control. In this paper, we successfully present an algorithm that can optimally solve this pattern relocation problem. Experimental results validate our method, and the relocation results with full scale layouts generated from Nangate Open Cell Library has shown great advantages with competitive runtimes compared to the existing commercial tool. Hongbo Zhang 0001, Yuelin Du, Martin D. F. Wong, Rasit Onur Topaloglu |
ASP-DAC | 4 |
| 2011 | Applications driving 3D integration and corresponding manufacturing challengesabstractThree dimensional (3D) semiconductor circuit integration has been an active area of research recently. Part of the driving force behind this interest has been applications. In this paper, we identify applications that drive 3D integration and point out the challenges they bring. In particular, we focus on through silicon via-based (TSV-based) 3D integration. TSV-based 3D integration opens up a new genre, and new opportunities for semiconductor integrated circuits. After a brief overview of TSV-based 3D technology overview, we identify driver applications that are dominant in this transition. We then point out challenges in the manufacturing area, i.e., thermal, reliability, EDA, cost, and test. Rasit Onur Topaloglu |
DAC | 1 |
| 2011 | Self-aligned double patterning decomposition for overlay minimization and hot spot detectionabstractSelf-aligned double patterning (SADP) lithography is a promising technology which can reduce the overlay and print 2D features for sub-32nm process. Yet, how to decompose a layout to minimize the overlay and perform hot spot detection is still an open problem. In this paper, we present an algorithm that can optimally solve the SADP decomposition problem. For a decomposable layout, our algorithm guarantees to find a decomposition solution that minimizes overlay. For a non-decomposable layout our algorithm guarantees to find all hot spots. Experimental results validate our method, and decomposition results for Nangate Open Cell Library and larger testcases are also provided with competitive runtimes. Hongbo Zhang 0001, Yuelin Du, Martin D. F. Wong, Rasit Onur Topaloglu |
DAC | 4 |
| 2011 | GPU programming for EDA with OpenCLabstractGraphical processing unit (GPU) computing has been an interesting area of research in the last few years. While initial adapters of the technology have been from image processing domain due to difficulties in programming the GPUs, research on programming languages made it possible for people without the knowledge of low-level programming languages such as OpenGL develop code on GPUs. Two main GPU architectures from AMD (former ATI) and NVIDIA acquired grounds. AMD adapted Stanford's Brook language and made it into an architecture-agnostic programming model. NVIDIA, on the other hand, brought CUDA framework to a wide audience. While the two languages have their pros and cons, such as Brook not being able to scale as well and CUDA having to account for architectural-level decisions, it has not been possible to compile one code on another architecture or across platforms. Another opportunity came with the introduction of the idea of combining one or more CPUs and GPUs on the same die. Eliminating some of the interconnection bandwidth issues, this combination makes it possible to offload tasks with high parallelism to the GPU. The technological direction towards multicores for CPU-only architectures also require a programming methodology change and act as a catalyst for suitable programming languages. Hence, a unified language that can be used both on multiple core CPUs as well as GPUs and their combinations has gained interest. Open Computing Language (OpenCL), developed originally by the Khronos Group of Apple and supported by both AMD and NVIDIA, is seen as the programming language of choice for parallel programming. In this paper, we provide a motivation for our tutorial talk on usage of OpenCL for GPUs and highlight key features of the language. We provide research directions on OpenCL for EDA. In our tutorial talk, we use EDA as our application domain to get the readers started with programming the rising language of parallelism, OpenCL. Rasit Onur Topaloglu, Benedict R. Gaster |
ICCAD | 1 |
| 2008 | Chip Optimization Through STI-Stress-Aware Placement Perturbations and Fill InsertionabstractStarting at the 65-nm node, stress engineering to improve the performance of transistors has been a major industry focus. An intrinsic stress source-shallow trench isolation (STI)-has not been fully utilized up to now for circuit performance improvement. In this paper, we present a new methodology that combines detailed placement and active-layer fill insertion to exploit STI stress for performance improvement. We conduct process simulation of a 65-nm production STI technology to generate mobility and delay impact models for STI stress. We then utilize these models to perform STI-stress-aware delay analysis of critical paths using Simulation Program with Integrated Circuit Emphasis (SPICE). We present our timing-driven optimization of STI stress in standard cell designs, using detailed placement perturbation and active-layer fill insertion to improve complementary metal-oxide-semiconductor performance. We assess the proposed analysis and optimization on small designs implemented with a 65-nm production cell library and a standard synthesis place-and-route flow. Our stress-aware timing analysis improves the clock frequency by 4.68% to 6.31% over traditional worst case analysis, and our optimization improves clock frequency by 2.44% to 5.26%. The frequency improvement through exploitation of STI stress comes at practically zero cost in terms of design area and wire length. Andrew B. Kahng, Rasit Onur Topaloglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | Exploiting STI stress for performanceabstractStarting at the 65 nm node, stress engineering to improve performance of transistors has been a major industry focus. An intrinsic stress source -shallow trench isolation -has not been fully utilized up to now for circuit performance improvement. In this paper, we present a new methodology that combines detailed placement and active-layer fill insertion to exploit STI stress for performance improvement. We perform process simulation of a production 65 nm STI technology to generate mobility and delay impact models for STI stress. Based on these models, we are able to perform STI stress-aware delay analysis of critical paths using SPICE. We then present our timing-driven optimization of STI stress in standard cell designs, using detailed placement perturbation to optimize PMOS performance and active-layer fill insertion to optimize NMOS performance. We assess our optimization on small designs implemented with a 65 nm production cell library and a standard synthesis, place and route flow. Our timing-driven optimization of STI stress impacts can improve clock frequency by between 7% to 11%. The frequency improvement through exploitation of STI stress comes at practically zero cost in terms of design area and wirelength. Andrew B. Kahng, Rasit Onur Topaloglu |
ICCAD | 3 |
| 2006 | Interconnect Matching Design Rule Inferring and Optimization through Correlation ExtractionabstractNew back-end design for manufacturability rules have brought guarantee rules for interconnect matching. These rules indicate a certain capacitance matching guarantee given spacing between interconnects and interconnect area. Yet, the number of these rules is so few that they are of limited value in circuit or interconnect optimization. A method to infer additional guarantees from the provided guarantees is necessary so that optimization can be optimal. In this paper, we target two problems. First, we present a methodology to infer additional matching guarantees through extracting correlation information from the given limited set of matching guarantees in the design manual. In order to achieve this, we propose a multi-function variant of multi-variate Newton-Raphson method to extract parameters of the proposed dimension-and distance-based process correlation model for interconnects. We propose to use the extracted correlation information to infer a continuum of matching rules through simulation with proposed modifications to the standard capacitance extraction procedure. Secondly, we show how to directly incorporate the inferred interconnect matching guarantees for accurate interconnect optimization in a flexible geometric programming construction. We show how much resource savings are possible through inferring of new matching rules. Applying the inferred mismatch guarantees allows a geometric programming-based H-tree optimization to reduce the clock tree resources 27% on average and up to 56%. Rasit Onur Topaloglu, Andrew B. Kahng |
ICCD | 1 |
| 2006 | Early, Accurate and Fast Yield Estimation through Monte Carlo-Alternative Probabilistic Behavioral Analog System SimulationsabstractMonte Carlo analysis has so far been the corner stone for analog statistical simulations. Fast and accurate simulations are necessary for stringent time-to-market, design for manufacturability and yield concerns in the analog domain. Although Monte Carlo attains accuracy, it does so with a sacrifice in run-time for analog simulations. In this paper, we propose a fast and accurate probabilistic simulation method alternative to Monte Carlo using deterministic sampling and weight propagation. We furthermore propose accuracy improvement algorithms and a fast yield calculation method. The proposed method shows accuracy improvement combined with a 100-fold reduction in run-time with respect to a 1000-sample Monte Carlo analysis Rasit Onur Topaloglu |
VTS | 1 |
| 2005 | Forward discrete probability propagation method for device performance characterization under process variationsabstractProcess variations are becoming influential at the device level in deep sub-micron and sub-wavelength design regimes, whereas they used to be a few generations away only influential at circuit level. Process variations cause device performance parameters, such as current or output resistance, to acquire a probability distribution. Estimation of these distributions has been accomplished using Monte Carlo techniques so far. The large number of samples needed by Monte Carlo methods adversely affects the possibility of integrating probabilistic device performance at the circuit level due to run-time inefficiency. In this paper, we introduce a novel technique called Forward Discrete Probability Propagation (FDPP). This method discretizes the probability distributions and effectively propagates these probabilities across a device formula hierarchy, such as the one present in the SPICE3v3 model. Consequently, probability distributions for process parameters are propagated to the device level. It is shown in the paper that with far fewer number of samples, comparable accuracy to a Monte Carlo method is achieved. Rasit Onur Topaloglu, Alex Orailoglu |
ASP-DAC | 1 |
| 2005 | A DFT approach for diagnosis and process variation-aware structural test of thermometer coded current steering DACsabstractA design for test (DFT) hardware is proposed to increase the controllability of a thermometer coded current steering digital to analog converter. A procedure is introduced to reduce the diagnosis and structural test time from quadratic to linear using the proposed DFT hardware. To evaluate the applicability of the proposed technique, principal component analysis is used to create virtual process variations to simulate in lieu of semiconductor fabrication data. An architecture specific soft fault model is suggested for the diagnosis problem. Random errors according to the fault model are introduced in the virtual test environment on top of the process variations and it is shown that diagnosis of a fault is possible with high accuracy with the proposed method. The same technique employing principal component analysis is furthermore used to provide process variation-aware reference test comparison values for a structural test of the DAC. The structural test provides a mechanism to test for even unmodeled manufacturing faults. The process variation-aware test values help detect defects even under process variations. The proposed DFT hardware and method are low cost and quite suitable for a built-in self diagnosis and test implementation. Rasit Onur Topaloglu, Alex Orailoglu |
DAC | 1 |
| 2004 | On mismatch in the deep sub-micron era - from physics to circuits
Rasit Onur Topaloglu, Alex Orailoglu |
ASP-DAC | 1 |