Ilya Wagner

dblp:85/4054 · DBLP profile ↗
← Back
16ranked-venue papers
11as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 11 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Electronic design automation · 46% Memory systems · 21% Processor architecture and microarchitecture · 17%
Network and information security
1 paper
Blockchain and cryptocurrency security · 100%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation › hardware verification and test
hardware verification
0.452015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Microprocessor Verification via Feedback-Adjusted Markov Models · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Shielding against design flaws with field repairable control logic · DAC 2006
Electronic design automation › hardware verification and test › design validation
post-silicon validation
0.322015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Dacota: Post-silicon validation of the memory subsystem in multi-core designs · HPCA 2009
Memory systems
cache coherence
0.212015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Memory systems
memory consistency
0.212015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Memory systems › memory consistency
memory consistency verification
0.212015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Processor architecture and microarchitecture
multicore design
0.212015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Electronic design automation
hardware verification and test
0.232009
Dacota: Post-silicon validation of the memory subsystem in multi-core designs · HPCA 2009
Shielding against design flaws with field repairable control logic · DAC 2006
StressTest: an automatic approach to test generation via activity monitors · DAC 2005
Processor architecture and microarchitecture
microprocessor design
0.232008
Using Field-Repairable Control Logic to Correct Design Errors in Microprocessors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Shielding against design flaws with field repairable control logic · DAC 2006
StressTest: an automatic approach to test generation via activity monitors · DAC 2005
Electronic design automation › hardware verification and test › debugging
design error correction
0.122008
Using Field-Repairable Control Logic to Correct Design Errors in Microprocessors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Shielding against design flaws with field repairable control logic · DAC 2006
Hardware reliability and fault tolerance › memory reliability
cache reliability
0.112011
Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011
Hardware reliability and fault tolerance › error correction
error-correcting codes
0.112011
Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011
Energy-efficient computing
voltage scaling
0.112011
Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011
Blockchain and cryptocurrency security › smart contract security
vulnerability detection
0.112008
Testudo: Heavyweight security analysis via statistical sampling · MICRO 2008
Electronic design automation › hardware verification and test
test generation
0.122007
StressTest: an automatic approach to test generation via activity monitors · DAC 2005
Microprocessor Verification via Feedback-Adjusted Markov Models · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Electronic design automation › hardware verification and test
processor verification
0.112007
Microprocessor Verification via Feedback-Adjusted Markov Models · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Electronic design automation › hardware verification and test › functional verification
simulation-based verification
0.112007
Microprocessor Verification via Feedback-Adjusted Markov Models · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Distributed systems
fault tolerance
0.112006
Shielding against design flaws with field repairable control logic · DAC 2006
Electronic design automation › hardware verification and test › test generation
random test generation
0.112005
StressTest: an automatic approach to test generation via activity monitors · DAC 2005
Hardware reliability and fault tolerance
soft errors
0.012011
Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011
Processor architecture and microarchitecture
chip multiprocessor
0.012009
Dacota: Post-silicon validation of the memory subsystem in multi-core designs · HPCA 2009
Electronic design automation › hardware verification and test › design validation
functional validation
0.012009
Dacota: Post-silicon validation of the memory subsystem in multi-core designs · HPCA 2009

Methods — techniques the papers use, named apart from their topics

tagged memory · 0.2statistical sampling · 0.2distributed debugging · 0.2graph cycle detection · 0.2cache-based logging · 0.2markov model · 0.1closed-loop feedback · 0.1variable-strength ECC · 0.1distributed checking algorithm · 0.1data-coloring · 0.1
YearPublicationVenuePosition
2026 Innovative Practices Session: Recent Developments in IEEE Draft Standards P1687, P1687.2, and P2929
Adam Cron, Stephen K. Sunter, Sankaran Menon, Jeff Rearick, Ilya Wagner, Tapan Chakraborty, Martin Keim
VTS5
2024 Functional State Extraction using Scan DFT
abstract
Functional state extraction of both sequential logic and arrays in a design is widely used in the industry to debug logic and timing bugs. Typically, this is done by repurposing logic and array test DFT that is already present in the design. In this paper, we describe the logic state extraction methodology as practiced in Intel servers, which is referred to as "scan dump". First, we motivate the state extraction problem and provide contextual background. Secondly, we describe design-for-test (DFT) implementation in support of scan dump for state extraction. Thirdly, we describe the pre-silicon validaion methodology for scan dump. Fourthly, we describe a set of best practices for implementing the scan dump feature. Finally, we describe an actual debug example using scan dump.
Ilya Wagner, Pankaj Pant, Arani Sinha
ITC1
2015 Post-Silicon Validation of Multiprocessor Memory Consistency
abstract
Shared-memory chip-multiprocessor (CMP) architectures define memory consistency models that establish the ordering rules for memory operations from multiple threads. Validating the correctness of a CMP's implementation of its memory consistency model requires extensive monitoring and analysis of memory accesses while multiple threads are executing on the CMP. In this paper, we present a low overhead solution for observing, recording and analyzing shared-memory interactions for use in an emulation and/or post-silicon validation environment. Our approach leverages portions of the CMP's own data caches, augmented only by a small amount of hardware logic, to log information relevant to memory accesses. After transferring this information to a central memory location, we deploy our own analysis algorithm to detect any possible memory consistency violations. We build on the property that a violation corresponds to a cycle in an appropriately defined graph representing memory interactions. The solution we propose allows a designer to choose where to run the analysis algorithm: 1) on the CMP itself; 2) on a separate processor residing on the validation platform; or 3) off-line on a separate host machine. Our experimental results show an 83% bug detection rate, in our testbed CMP, over three distinct memory consistency models, namely: relaxed-memory order, total-store order, and sequential consistency. Finally, note that our solution can be disabled in the final product, leading to zero performance overhead and a per-core area overhead that is smaller than the size of a physical integer register file in a modern processor.
Biruk Mammo, Valeria Bertacco, Andrew DeOrio, Ilya Wagner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2011 Distributed hardware matcher framework for SoC survivability
abstract
Modern systems on chip (SoCs) are rapidly becoming complex high-performance computational devices, featuring multiple general purpose processor cores and a variety of functional IP blocks, communicating with each other through on-die fabric. While modular SoC design provides power savings and simplifies the development process, it also leaves significant room for a special type of hardware bugs, interaction errors, to slip through pre- and post-silicon verification. Consequently, hard to fix silicon escapes may be discovered late in production schedule or even after a market release, potentially causing costly delays or recalls. In this work we propose a unified error detection and recovery framework that incorporates programmable features into the on-die fabric of an SoC, so triggers of escaped interaction bugs can be detected at runtime. Furthermore, upon detection, our solution locks the interface of an IP for a programmed time period, thus altering interactions between accesses and bypassing the bug in a manner transparent to software. For classes of errors that cannot be circumvented by this in-hardware technique our framework is programmed to propagate the error detection to the software layer. Our experiments demonstrate that the proposed framework is capable of detecting a range of interaction errors with less than 0.01% performance penalty and 0.45% area overhead.
Ilya Wagner, Shih-Lien Lu
DATE1
2011 Energy-efficient cache design using variable-strength error-correcting codes
abstract
Voltage scaling is one of the most effective mechanisms to improve microprocessors' energy efficiency. However, processors cannot operate reliably below a minimum voltage, Vccmin, since hardware structures may fail. Cell failures in large memory arrays (e.g., caches) typically determine Vccmin for the whole processor. We observe that most cache lines exhibit zero or one failures at low voltages. However, a few lines, especially in large caches, exhibit multi-bit failures and increase Vccmin. Previous solutions either significantly reduce cache capacity to enable uniform error correction across all lines, or significantly increase latency and bandwidth overheads when amortizing the cost of error-correcting codes (ECC) over large lines.
Alaa R. Alameldeen, Ilya Wagner, Zeshan Chishti, Wei Wu 0024, Chris Wilkerson, Shih-Lien Lu
ISCA2
2009 Caspar: Hardware patching for multicore processors
abstract
Ensuring correctness of execution of complex multi-core processor systems deployed in the field remains to this day an extremely challenging task. The major part of this effort is concentrated on design verification, where different pre- and post-silicon techniques are used to guarantee that devices behave exactly as stated in the specification. Unfortunately, the performance of even state-of-the-art validation tools lags behind the growing complexity of multi-core designs. Therefore, subtle bugs still slip into released components, causing incorrect computational results, or even compromising the security of the end-user systems. In this work we present Caspar - an approach for in-the-field patching of the memory subsystem hardware in multi-core chips. Caspar relies on a checkpointing system, which periodically logs the state of the chip, and a novel error detection and recovery scheme, which uses a simplified mode of operation to bypass cache coherence and consistency errors. The implementation of Caspar employs hardware detectors: on-die programmable circuits to identify system's configurations that may lead to bugs, and to trigger recovery and bypass. Our experimental results show that Caspar can be used effectively to detect and bypass a variety of memory subsystem bugs, with as little as 2% performance impact and 6% area overhead during bug-free operation.
Ilya Wagner, Valeria Bertacco
DATE1
2009 Dacota: Post-silicon validation of the memory subsystem in multi-core designs
abstract
The number of functional errors escaping design verification and being released into final silicon is growing, due to the increasing complexity and shrinking production schedules of modern processor designs. Recent trends towards chip multiprocessors (CMPs) are exacerbating the problem because of their complex and sometimes non-deterministic memory subsystems, prone to subtle but devastating bugs. This deteriorating situation calls for high-efficiency, high-coverage results in functional validation, results that are be achieved by leveraging the performance of post-silicon validation, that is, those verification tasks that are executed directly on prototype hardware. The orders-of-magnitude faster testing in post-silicon enables designers to achieve much higher coverage before customer release, but only if the limitations of this technology in diagnosis and internal node observability could be overcome. In this work, we unlock the full performance of post-silicon validation through Dacota, a new high-coverage solution for validating memory operation ordering in CMPs. When activated, Dacota reconfigures a portion of the cache storage to log memory accesses using a compact data-coloring scheme. Logs are periodically aggregated and checked by a distributed algorithm running in-situ on the CMP to verify correct memory operation ordering. When the design is ready for customer shipment, Dacota can be deactivated, releasing all cache storage, and only leaving a small silicon area footprint, less than 0.01% (three orders of magnitude smaller than previous solutions). We found experimentally that Dacota is effective in exposing memory subsystem bugs, and it delivers its high coverage capabilities at a 26% performance slowdown (only during validation) for real-world applications.
Andrew DeOrio, Ilya Wagner, Valeria Bertacco
HPCA2
2008 MCjammer: Adaptive Verification for Multi-core Designs
abstract
The challenge of verification of multi-core and multi-processor designs grows dramatically with each new generation of systems produced today. Validation of memory coherence of such systems, which include multiple levels of cache and complex protocols, constitutes a major fraction of this task. Unfortunately, current tools are incapable of addressing these challenges, allowing bugs, which cause unpredictable software behavior and wrong computation results, to slip into hardware. In this work we present a scalable approach to the verification of memory coherence protocols in large multi-core and multi-processor systems. We accomplish this task through a distributed network of cooperating agents, which feed the processors with stimuli, each agent attempting to accomplish its own verification goals and support other agents on theirs as well. The agents can dynamically change the stimuli based on coverage and pressure observed during simulation. Since each agent has a minimal knowledge of the entire system, their communication and decision process is greatly simplified. Moreover, since the agents' view of the system is linear in the number of nodes in it, our approach can be efficiently scaled to target large multi-core systems. Experimental results on two common coherence protocols and a range of multi-core configurations demonstrate that our technique can reach high levels of coverage of the system-level protocol much faster than a constrained-random generator.
Ilya Wagner, Valeria Bertacco
DATE1
2008 Reversi: Post-silicon validation system for modern microprocessors
abstract
Verification remains an integral and crucial phase of todaypsilas microprocessor design and manufacturing process. Unfortunately, with soaring design complexities and decreasing time-to-market windows, todaypsilas verification approaches are incapable of fully validating a microprocessor before its release to the public. Increasingly, post-silicon validation is deployed to detect complex functional bugs in addition to exposing electrical and manufacturing defects. This is due to the significantly higher execution performance offered by post-silicon methods, compared to pre-silicon approaches. Validation in the post-silicon domain is predominantly carried out by executing constrained-random test instruction sequences directly on a hardware prototype. However, to identify errors, the state obtained from executing tests directly in hardware must be compared to the one produced by an architectural simulation of the designpsilas golden model. Therefore, the speed of validation is severely limited by the necessity of a costly simulation step. In this work we address this bottleneck in the traditional flow and present a novel solution for post-silicon validation that exposes its native high performance. Our framework, called Reversi, generates random programs in such a way that their correct final state is known at generation time, eliminating the need for architectural simulations. Our experiments show that Reversi generates tests exposing more bugs faster, and can speed up post-silicon validation by 20x compared to traditional flows.
Ilya Wagner, Valeria Bertacco
ICCD1
2008 Testudo: Heavyweight security analysis via statistical sampling
abstract
Heavyweight security analysis systems, such as taint analysis and dynamic type checking, are powerful technologies used to detect security vulnerabilities and software bugs. Traditional software implementations of these systems have high instrumentation overhead and suffer from significant performance impacts. To mitigate these slowdowns, a few hardware-assisted techniques have been recently proposed. However, these solutions incur a large memory overhead and require hardware platform support in the form of tagged memory systems and extended bus designs. Due to these costs and limitations, the deployment of heavyweight security analysis solutions is, as of today, limited to the research lab. In this paper, we describe Testudo, a novel hardware approach to heavyweight security analysis that is based on statistical sampling of a programpsilas dataflow. Our dynamic distributed debugging reduces the memory overhead to a small storage space by selectively sampling only a few tagged variables to analyze during any particular execution of the program. Our system requires only small hardware modifications: it adds a small sample cache to the main processor and extends the pipeline registers to propagate analysis tags. To gain high analysis coverage, we rely on a population of users to run the program, sampling a different random set of variables during each new run. We show that we can achieve high coverage analysis at virtually no performance impact, even with a reasonably-sized population of users. In addition, our approach even scales to heavyweight debugging techniques by keeping per-user runtime overheads low despite performing traditionally costly analyses. Moreover, the low hardware cost of our implementation allows it to be easily distributed across large user populations, leading to a higher level of security analysis coverage than previously.
Joseph L. Greathouse, Ilya Wagner, David A. Ramos, Gautam Bhatnagar, Todd M. Austin, Valeria Bertacco, Seth Pettie
MICRO2
2008 Using Field-Repairable Control Logic to Correct Design Errors in Microprocessors
abstract
Functional correctness is a vital attribute of any hardware design. Unfortunately, due to extremely complex architectures, widespread components, such as microprocessors, are often released with latent bugs. The inability of modern verification tools to handle the fast growth of design complexity exacerbates the problem even further. In this paper, we propose a novel hardware-patching mechanism, called the field-repairable control logic (FRCL), that is designed for in-the-field correction of errors in the design's control logic-the most common type of defects, as our analysis demonstrates. Our solution introduces an additional component in the processor's hardware, a state matcher, that can be programmed to identify erroneous configurations using signals in the critical control state of the processor. Once a flawed configuration is ldquomatched,rdquo the processor switches into a degraded mode, a mode of operation which excludes most features of the system and is simple enough to be formally verified, yet still capable to execute the full instruction-set architecture at one instruction at a time. Once the program segment exposing the design flaw has been executed in a degraded mode, we can switch the processor back to its full-performance mode. In this paper, we analyze a range of approaches to selecting signals comprising the processor's critical control state and evaluate their effectiveness in representing a variety of design errors. We also introduce a new metric (average specificity per signal) that encodes the bug-detection capability and amount of control state of a particular critical signal set. We demonstrate that the FRCL can support the detection and correction of multiple design errors with a performance impact of less than 5% as long as the incidence of the flawed configurations is below 1% of dynamic instructions. In addition, the area impact of our solution is less than 2% for the two microprocessor designs that we investigated in our experiments.
Ilya Wagner, Valeria Bertacco, Todd M. Austin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2007 Engineering trust with semantic guardians
abstract
The ability to guarantee the functional correctness of digital integrated circuits and, in particular, complex microprocessors, is a key task in the production of secure and trusted systems. Unfortunately, this goal remains today an unfulfilled challenge, as even the most straightforward practical designs are released with latent bugs. Patching techniques can repair some of these escaped bugs, however, they often incur a performance overhead, and most importantly, they can only be deployed after an escaped bug has been exposed at the customer site. In this paper we present a novel approach to guaranteeing correct system operation by deploying a semantic guardian component. The semantic guardian is an additional control logic block which is included in the design, and can switch the microprocessor's mode of operation from its normal, high-performance but error-prone mode, to a secure, formally verified safe mode, guaranteeing that the execution will be functionally correct. We explore several frameworks where a selective use of the safe mode can enhance the overall functional correctness of a processor. Additionally, we observe through experimentation that semantic guardians facilitate the trade-off between the design validation effort and the performance and area cost of the final secure product. The experimental results show that the area cost and performance overheads of a semantic guardian can be as small as 3.5% and 5%, respectively
Ilya Wagner, Valeria Bertacco
DATE1
2007 Microprocessor Verification via Feedback-Adjusted Markov Models
abstract
The challenge of verifying a modern microprocessor design is an overwhelming one: Increasingly complex microarchitectures combined with heavy time-to-market pressure have forced microprocessor vendors to employ immense verification teams in the hope of finding the most critical bugs in a timely manner. Unfortunately, too often, size does not seem to matter in verification, as design schedules continue to slip and microprocessors find their way to the marketplace with design errors. In this paper, we describe a novel closed-loop simulation-based approach to hardware verification and present a tool called StressTest that uses our methods to locate hard-to-find corner-case design bugs and performance problems. StressTest is based on a Markov-model-driven random instruction generator with activity monitors. The model is generated from the user-specified template files and is used to generate the instructions sent to the design under test (DUT). In addition, the user specifies key activity nodes within the design that should be stressed and monitored throughout the simulation. The StressTest engine then uses closed-loop feedback techniques to transform the Markov model into one that effectively stresses the user-selected points of interest. In parallel, StressTest monitors the correctness of the DUT response and, if the design behaves against expectation, it reports a bug and a trace leading to it. Using two microarchitectures as example testbeds, we demonstrate that StressTest finds more bugs with less effort than open-loop random instruction test generation techniques
Ilya Wagner, Valeria Bertacco, Todd M. Austin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2006 Depth-driven verification of simultaneous interfaces
abstract
The verification of modern computing systems has grown to dominate the cost of system design, often with limited success as designs continue to be released with latent bugs. This trend is accelerated with the advent of highly integrated system-on-a-chip (SoC) designs, which feature multiple complex subcomponents connected by simultaneously active interfaces. In this paper, we introduce a closed-loop feedback technique targeting the verification of multiple components connected by parallel interfaces. We utilize an environment with hierarchical Markov models, where top-level submodels specify overarching simulation goals of the system, while lower-level submodels specify the detailed component-level input generation. Test accuracy is improved through the use of depth-driven random test generation. The approach allows users to specify correctness properties and key activity nodes in the design to be exercises. We examine three nontrivial designs, two microprocessors and a chip-multiprocessor router switch, and we demonstrate that our technique finds many more bugs than constrained-random test generation technique and reduces the simulation effort in half, compared to previous Markov-model based solutions.
Ilya Wagner, Valeria Bertacco, Todd M. Austin
ASP-DAC1
2006 Shielding against design flaws with field repairable control logic
abstract
Correctness is a paramount attribute of any microprocessor design; however, without novel technologies to tame the increasing complexity of design verification, the amount of bugs that escape into silicon will only grow in the future. In this paper, we propose a novel hardware patching mechanism that can detect design errors which escaped the verification process, and can correct them directly in the field. We accomplish this goal through a simple field-programmable state matcher, which can identify erroneous configurations in the processor's control state and switch the processor into formally-verified degraded performance mode, once a "match" occurs. When the instructions exposing the design flaw are committed, the processor is switched back to normal mode. We show that our approach can detect and correct infrequently-occurring errors with almost no performance impact and has approximately 2% area overhead.
Ilya Wagner, Valeria Bertacco, Todd M. Austin
DAC1
2005 StressTest: an automatic approach to test generation via activity monitors
abstract
The challenge of verifying a modern microprocessor design is an overwhelming one: Increasingly complex micro-architectures combined with heavy time-to-market pressure have forced microprocessor vendors to employ immense verification teams in the hope of finding the most critical bugs in a timely manner. Unfortunately, too often size doesn't seem to matter for verification teams, as design schedules continue to slip and microprocessors find their way to the marketplace with design errors. In this paper, we describe a simulationbased random test generation tool, called StressTest, that provides assistance in locating hard-to-find corner-case design bugs and performance problems. StressTest is based on a Markov-model-driven random instruction generator with activity monitors. The model is generated from the userspecified template programs and is used to generate the instructions sent to the design under test (DUT). In addition, the user specifies key activity points within the design that should be stressed and monitored throughout the simulation. The StressTest engine then uses closed-loop feedback techniques to transform the Markov model into one that effectively stresses the points of interest. In parallel, StressTest monitors the correctness of the DUT response to the supplied stimuli, and if the design behaves unexpectedly, a bug and a trace that leads to it are reported. Using two micro-architectures as example testbeds, we demonstrate that StressTest finds more bugs with less effort than open-loop random instruction test generation techniques.
Ilya Wagner, Valeria Bertacco, Todd M. Austin
DAC1