Peter Domanski

dblp:306/8164 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-5283-2712ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Focus Session: Do Agentic LLMs Change the Paradigm of Hardware Test Generation?
abstract
Technology scaling and increasing System-on-Chip (SoC) complexity exacerbate reliability challenges arising from both structural defects and runtime-dependent failures, including Silent Data Corruptions (SDCs) that evade traditional error detection mechanisms. Structural testing remains essential for detecting modeled faults such as stuck-at faults; however, it is inherently limited in capturing failures that arise under dynamic operating conditions. In contrast, functional testing can expose workload-dependent failures, albeit at the cost of high testing overhead and largely unguided workload generation. This paper presents an agentic testing framework that integrates Large Language Models (LLMs) with Reinforcement Learning (RL) and Tree-structured Parzen Estimators (TPE) to guide functional workload generation and Automatic Test Pattern Generation (ATPG) settings under user-defined constraints. The proposed approach leverages feedback-driven optimization to steer test generation toward failure-prone behaviors while reducing reliance on manual expertise. Experimental evaluation on a RISC-V processor core demonstrates that the method outperforms manually generated workloads for functional testing, while experiments on six benchmark circuits show test quality comparable to expert-generated ATPG scripts for structural testing, with improved efficiency and scalability.
Farshad Firouzi, Agastya Seth, Peter Domanski, Bahareh J. Farahani, Sanmitra Banerjee, Jonti Talukdar, Krishnendu Chakrabarty
DATE4
2026 DiabLLM: An LLM-Based Framework for Blood Glucose Prediction in Type 1 Diabetes
abstract
Accurate Blood Glucose (BG) prediction is essential for enabling glycemic control in individuals with Type 1 Diabetes Mellitus (T1DM), particularly within Smart and Connected Health (SCH) systems that integrate Continuous Glucose Monitoring (CGM) and automated insulin delivery. The adaptability of Large Language Models (LLMs) provides a promising foundation for unified, fine-tunable forecasting models. We introduce DiabLLM, a framework based on two recent LLM-based architectures: Time-LLM, which incorporates a lightweight projection layer and alignment techniques to transform time-series data into embeddings interpretable by pre-trained LLMs, and Chronos, which employs time-series-aware tokenization and quantization to convert continuous inputs into discrete sequences for forecasting. Both models process 30-minute sequences of six historical BG values and predict 30- and 45-minute horizons. Experimental results on the OhioT1DM and D1NAMO datasets demonstrate that DiabLLM outperforms state-of-the-art baselines, including a Deep Reinforcement Learning model and an ensemble of LSTM, GRU, and WaveNet, achieving up to 27% improvement in RMSE and 37% in MAE. To enhance robustness to noisy and missing input data, a denoising autoencoder was employed for input reconstruction, yielding improved predictive performance. In addition, knowledge distillation was shown to significantly compress the model, making it a practical candidate for efficient deployment on resource-constrained edge devices without compromising accuracy.
Amirhossein Mahmoudi, Ghazal Farahani, Peter Domanski, Bahareh J. Farahani, Farshad Firouzi, Krishnendu Chakrabarty
IEEE J. Biomed. Health Informatics3
2026 TIDE-S: Telemetry Informed Delay Testing With Optimized Sensor Placement
abstract
Silent data corruption (SDC) refers to undetected errors that yield incorrect results without triggering system alerts or error logs. Existing test methodologies are inadequate for capturing dynamic voltage fluctuations that occur under realistic workload conditions, thereby limiting their effectiveness for detecting SDCs. We present TIDE-S, a telemetry-informed delay testing (TIDE) framework that integrates presilicon sensor placement strategies to improve telemetry accuracy. By evaluating different sensor allocation schemes—uniform,$K$-means, and energy-aware clustering—TIDE-S improves the spatial granularity of voltage observation, enabling more accurate correlation between voltage fluctuations and path delay behavior. The combined telemetry- and sensor-aware methodology significantly improves the detection of timing-sensitive SDCs. The proposed framework incurs minimal infrastructure overhead by leveraging standard pad-based voltage observation, making it practical for real system-on-chip (SoC) designs. We demonstrate the effectiveness of TIDE-S across multiple RISC-V-based SoCs and a diverse set of real-world benchmarks, showing quantifiable improvements in voltage estimation, slack prediction, and test coverage.
Deepesh Sahoo, Eduardo Ortega, Peter Domanski, Farshad Firouzi, Krishnendu Chakrabarty
IEEE Trans. Very Large Scale Integr. Syst.3
2025 LLM-Aided In-Field Workload Generation for Detecting Silent Data Corruptions at Scale
abstract
Computational integrity is crucial in large-scale data centers where Silent Data Corruptions (SDCs) pose a growing reliability challenge. SDCs can lead to incorrect computation results not captured by traditional error detection mechanisms, making their detection and mitigation essential. However, existing post-manufacturing and in-field testing methods, such as opportunistic and ripple testing, face significant scalability challenges due to high computational costs and test times. We propose an LLM-aided approach for generating targeted test cases to detect SDCs. As a case study, we focus on the functional blocks of a RISC-V CV32E40P processor core. Our method generates targeted test cases that maximize voltage droops in given hardware modules, such as functional units, increasing the likelihood of triggering SDCs in-field. Additionally, our approach is architecture-aware and layout-aware, enhancing fault activation and enabling automated optimization of generated test cases. Experimental evaluations demonstrate that the proposed method significantly improves SDC detection efficiency by reducing the number of required test cases while preserving high test coverage. Compared to commonly used test cases, the proposed approach increases average voltage droops by up to 38%. By integrating LLM-aided test case generation, the proposed approach achieves voltage droops of up to 9% relative to the supply voltage, improving the effectiveness of in-field SDC detection and mitigation strategies.
Peter Domanski, Deepesh Sahoo, Eduardo Ortega, Farshad Firouzi, Krishnendu Chakrabarty
ITC1
2025 TIDE: Telemetry-Informed Delay Testing for Silent Data Corruption *
abstract
Silent Data Corruption (SDC) is caused by undetected errors that yield incorrect results without triggering system alerts or error logs. Existing test methodologies are inadequate for capturing dynamic voltage fluctuations that occur under realistic workload conditions, thereby limiting their effectiveness for detecting SDCs. To address these limitations, we introduce Telemetry-Informed Delay Testing (TIDE), a novel methodology that enhances SDC detection by leveraging telemetry sensors to monitor voltage fluctuations and their impact on timing integrity. By incorporating dynamic, workload-aware test generation, the proposed framework overcomes key limitations of traditional approaches and facilitates early detection of SDCs. The effectiveness of TIDE is demonstrated through case studies conducted on two RISC-V-based SoCs and multiple workloads.
Deepesh Sahoo, Eduardo Ortega, Peter Domanski, Farshad Firouzi, Krishnendu Chakrabarty
ITC3
2025 Silent Data Corruption: Advancing Detection, Diagnosis, and Mitigation Strategies
abstract
Silent Data Corruptions (SDCs) pose a critical challenge to computer system reliability, arising from vulnerabilities across different layers of the computing stack. This paper addresses this challenge through three complementary contributions that systematically target SDCs from hardware manufacturing to application-level resilience. First, we analyze timing failures caused by random process variations in advanced technology nodes, revealing that extreme slow paths at lower voltages are dominated by single weak transistors—insights crucial for manufacturing and in-field testing. Second, we introduce an LLM-driven framework that generates targeted functional test programs to induce SDCs, demonstrating its effectiveness in stressing hardware, uncovering latent vulnerabilities, and increasing energy consumption in a given device under test (DUT), making it a valuable tool for in-field testing. Third, as machine learning continues to drive advancements across critical domains such as healthcare, finance, and autonomous systems, ensuring its reliability is paramount. However, the susceptibility of these applications to SDCs threatens their reliability and robustness. To address this, we propose Fidelity-Q, a novel fault injection methodology to evaluate the impact of SDCs on Quantized Neural Networks (QNNs), showing that lower-bit quantization increases error susceptibility. Collectively, these contributions provide a comprehensive approach to identifying, analyzing, and mitigating SDCs across the computing stack, from hardware testing to machine learning applications.
Peter Domanski, Mukarram Ali Faridi, Gabriel Kaunang, Wilson Pradeep, Adit D. Singh, Alfian Amrizal, Yanjing Li, Farshad Firouzi, Krishnendu Chakrabarty
VTS1
2025 ChipMnd: LLMs for Agile Chip Design
abstract
The increasing complexity of semiconductor design, along with stringent performance, power, and time-to-market requirements, has outpaced the capabilities of traditional Electronic Design Automation (EDA) methodologies. Conventional design workflows rely on manual intervention for critical tasks such as hardware description, synthesis optimization, and verification, leading to inefficiencies and scalability limitations. Large Language Models (LLMs) present a transformative approach by automating key stages of the design pipeline, enabling intelligent synthesis tuning, test generation, and security analysis. This paper introduces ChipMind, an LLM-driven framework comprising specialized agents and modules for digital and analog chip design. ChipMind integrates AI-driven methodologies to enhance design efficiency, accelerate prototyping, and optimize key design trade-offs, thereby addressing fundamental challenges in modern semiconductor development.
Farshad Firouzi, David Z. Pan, Jiaqi Gu 0002, Bahareh J. Farahani, Jayeeta Chaudhuri, Ziang Yin, Pingchuan Ma 0012, Peter Domanski, Krishnendu Chakrabarty
VTS8
2024 Training Large Language Models for System-Level Test Program Generation Targeting Non-functional Properties
abstract
System-Level Test (SLT) has been an integral part of integrated circuit test flows for over a decade and continues to be significant. Nevertheless, there is a lack of systematic approaches for generating test programs, specifically focusing on the non-functional aspects of the Device under Test (DUT). Currently, test engineers manually create test suites using commercially available software to simulate the end-user environment of the DUT. This process is challenging and laborious and does not assure adequate control over non-functional properties. This paper proposes to use Large Language Models (LLMs) for SLT program generation. We use a pre-trained LLM and fine-tune it to generate test programs that optimize non-functional properties of the DUT, e.g., instructions per cycle. Therefore, we use Gem5, a microarchitectural simulator, in conjunction with Reinforcement Learning-based training. Finally, we write a prompt to generate C code snippets that maximize the instructions per cycle of the given architecture. In addition, we apply hyperparameter optimization to achieve the best possible results in inference.
Denis Schwachhofer, Peter Domanski, Steffen Becker 0001, Stefan Wagner 0001, Matthias Sauer 0002, Dirk Pflüger, Ilia Polian
ETS2
2023 Learn to Tune: Robust Performance Tuning in Post-Silicon Validation
abstract
Post-silicon validation is a crucial yet challenging problem primarily due to the increasing complexity of the semi-conductor value chain. Existing techniques cannot keep up with the rapid increase in the complexity of designs. Therefore, post-silicon validation is becoming an expensive bottleneck. Robust performance tuning is relevant to compensate impacts of process variations and non-ideal design implementations. We propose a novel approach based on Deep Reinforcement Learning and Learn to Optimize. The method automatically learns flexible tuning strategies tailored to specific circuits. Additionally, it addresses high-dimensional tuning tasks, including mixed data types and dependencies, e.g., on operating conditions. In this work, we introduce Learn to Tune and demonstrate its appealing properties in post-silicon validation, e.g., lower computational cost or faster time-to-optimize, allowing a more efficient adaption of the tuning to changing tuning conditions than classical methods.
Peter Domanski, Dirk Pflüger, Raphaël Latty
ETS1
2022 Intelligent Methods for Test and Reliability
abstract
Test methods that can keep up with the ongoing increase in complexity of semiconductor products and their underlying technologies are an essential prerequisite for maintaining quality and safety of our daily lives and for continued success of our economies and societies. There is a huge potential how test methods can benefit from recent breakthroughs in domains such as artificial intelligence, data analytics, virtual/augmented reality, and security. The Graduate School on “Intelligent Methods for Semiconductor Test and Reliability” (GS-IMTR) at the University of Stuttgart is a large-scale, radically interdisciplinary effort to address the scientific-technological challenges in this domain. It is funded by Advantest, one of the world leaders in automatic test equipment. In this paper, we describe the overall philosophy of the Graduate School and the specific scientific questions targeted by its ten projects.
Hussam Amrouch, Jens Anders, Steffen Becker 0001, Maik Betka, Gerd Bleher, Peter Domanski, Nourhan Elhamawy, Thomas Ertl, Athanasios Gatzastras, Paul R. Genssler, Sebastian Hasler, Martin Heinrich, André van Hoorn, Hanieh Jafarzadeh, Ingmar Kallfass, Florian Klemme, Steffen Koch 0001, Ralf Küsters, Andrés Lalama, Raphaël Latty, Yiwen Liao, Natalia Lylina, Zahra Paria Najafi-Haghi, Dirk Pflüger, Ilia Polian, Jochen Rivoir, Matthias Sauer 0002, Denis Schwachhofer, Steffen Templin, Christian Volmer, Stefan Wagner 0001, Daniel Weiskopf, Hans-Joachim Wunderlich, Bin Yang 0009
DATE6
2021 ORSA: Outlier Robust Stacked Aggregation for Best- and Worst-Case Approximations of Ensemble Systems
abstract
In recent years, the usage of ensemble learning in applications has grown significantly due to increasing computational power allowing the training of large ensembles in reasonable time frames. Many applications, e.g., malware detection, face recognition, or financial decision-making, use a finite set of learning algorithms and do aggregate them in a way that a better predictive performance is obtained than any other of the individual learning algorithms. In the field of Post-Silicon Validation for semiconductor devices (PSV), data sets are typically provided that consist of various devices like, e.g., chips of different manufacturing lines. In PSV, the task is to approximate the underlying function of the data with multiple learning algorithms, each trained on a device-specific subset, instead of improving the performance of arbitrary classifiers on the entire data set. Furthermore, the expectation is that an unknown number of subsets describe functions showing very different characteristics. Corresponding ensemble members, which are called outliers, can heavily influence the approximation. Our method aims to find a suitable approximation that is robust to outliers and represents the best or worst case in a way that will apply to as many types as possible. A ‘softmax’ or ‘soft-min’ function is used in place of a maximum or minimum operator. A Neural Network (NN) is trained to learn this ‘soft-function’ in a two-stage process. First, we select a subset of ensemble members that is representative of the best or worst case. Second, we combine these members and define a weighting that uses the properties of the Local Outlier Factor (LOF) to increase the influence of non-outliers and to decrease outliers. The weighting ensures robustness to outliers and makes sure that approximations are suitable for most types.
Peter Domanski, Dirk Pflüger, Raphaël Latty, Jochen Rivoir
ICMLA1