Andrew DeOrio

dblp:27/3783 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
3since 2021 · last 2025
0000-0001-5653-5109ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 10 first-authorHuman-computer interaction and ubiquitous computing · 5 · 2 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Electronic design automation · 47% Memory systems · 20% Hardware reliability and fault tolerance · 12%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software testing › mutation testing
mutant detection
0.612022
On the use of mutation analysis for evaluating student test suite quality · ISSTA 2022
Software testing
mutation testing
0.612022
On the use of mutation analysis for evaluating student test suite quality · ISSTA 2022
Electronic design automation › hardware verification and test › design validation
post-silicon validation
0.322015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Dacota: Post-silicon validation of the memory subsystem in multi-core designs · HPCA 2009
Electronic design automation
hardware verification and test
0.332009
Inferno: Streamlining Verification With Inferred Semantics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Dacota: Post-silicon validation of the memory subsystem in multi-core designs · HPCA 2009
Event-driven gate-level simulation with GP-GPUs · DAC 2009
Processor architecture and microarchitecture
multicore design
0.322015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
DRAIN: distributed recovery architecture for inaccessible nodes in multi-core chips · DAC 2011
Hardware reliability and fault tolerance › fault-tolerant architecture
fault-tolerant noc
0.222011
DRAIN: distributed recovery architecture for inaccessible nodes in multi-core chips · DAC 2011
Vicis: a reliable network for unreliable silicon · DAC 2009
Memory systems
cache coherence
0.212015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Electronic design automation › hardware verification and test
hardware verification
0.212015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Memory systems
memory consistency
0.212015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Memory systems › memory consistency
memory consistency verification
0.212015
Post-Silicon Validation of Multiprocessor Memory Consistency · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Software testing › mutation testing
fault injection
0.212022
On the use of mutation analysis for evaluating student test suite quality · ISSTA 2022
Electronic design automation
physical design
0.112010
Electronic design automation for social networks · DAC 2010
Electronic design automation › physical design
placement and routing
0.112010
Electronic design automation for social networks · DAC 2010
Electronic design automation › hardware verification and test › hardware verification
assertion-based verification
0.112009
Inferno: Streamlining Verification With Inferred Semantics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Electronic design automation
boolean satisfiability
0.112009
Human computing for EDA · DAC 2009
Electronic design automation › hardware verification and test › functional verification
constrained random verification
0.112009
Inferno: Streamlining Verification With Inferred Semantics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Performance modeling and evaluation › simulation
discrete-event simulation
0.112009
Event-driven gate-level simulation with GP-GPUs · DAC 2009
Electronic design automation › hardware verification and test › logic simulation
gate-level simulation
0.112009
Event-driven gate-level simulation with GP-GPUs · DAC 2009
GPUs and heterogeneous computing
GPU computing
0.112009
Event-driven gate-level simulation with GP-GPUs · DAC 2009
Electronic design automation › hardware verification and test
logic simulation
0.112009
Event-driven gate-level simulation with GP-GPUs · DAC 2009
Hardware reliability and fault tolerance
permanent fault tolerance
0.122012
A Reliable Routing Architecture and Algorithm for NoCs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012
Vicis: a reliable network for unreliable silicon · DAC 2009
Processor architecture and microarchitecture
many-core architecture
0.012011
DRAIN: distributed recovery architecture for inaccessible nodes in multi-core chips · DAC 2011
Electronic design automation
logic synthesis
0.012010
Electronic design automation for social networks · DAC 2010
Processor architecture and microarchitecture
chip multiprocessor
0.012009
Dacota: Post-silicon validation of the memory subsystem in multi-core designs · HPCA 2009
Electronic design automation › hardware verification and test › design validation
functional validation
0.012009
Dacota: Post-silicon validation of the memory subsystem in multi-core designs · HPCA 2009
Distributed systems › distributed system verification
protocol verification
0.012009
Inferno: Streamlining Verification With Inferred Semantics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Interconnection networks and networks-on-chip
router architecture
0.012009
Vicis: a reliable network for unreliable silicon · DAC 2009

Methods — techniques the papers use, named apart from their topics

mutation analysis · 1.1all-pairs grading · 1.1graph cycle detection · 0.2cache-based logging · 0.2reconfigurable architecture · 0.1distributed routing algorithm · 0.1distributed recovery · 0.1graph analysis · 0.1EDA tooling · 0.1human computation · 0.1gamification · 0.1event-driven simulation · 0.1
YearPublicationVenuePosition
2025 Instructor-Written Hints as Automated Test Suite Quality Feedback
abstract
Mutation testing measures a test suite's ability to detect bugs by inserting bugs into the code and seeing if the tests behave differently. Mutation testing has recently seen increased adoption in industrial and open-source software but sees limited use in education. Some instructors use manually-constructed mutants to evaluate student tests and provide general automated feedback. Additional tutoring requires more intensive instructor interaction such as in office hours, which requires substantial resources at scale. Prior work suggests that students benefit from frequent, actionable feedback, and our work focuses on the challenge of leveraging automation to give students high-quality feedback when they need it.
James Perretta, Andrew DeOrio, Arjun Guha, Jonathan Bell 0001
SIGCSE (1)2
2022 On the use of mutation analysis for evaluating student test suite quality
abstract
A common practice in computer science courses is to evaluate student-written test suites against either a set of manually-seeded faults (handwritten by an instructor) or against all other student-written implementations (“all-pairs” grading). However, manually seeding faults is a time consuming and potentially error-prone process, and the all-pairs approach requires significant manual and computational effort to apply fairly and accurately. Mutation analysis, which automatically seeds potential faults in an implementation, is a possible alternative to these test suite evaluation approaches. Although there is evidence in the literature that mutants are a valid substitute for real faults in large open-source software projects, it is unclear whether mutants are representative of the kinds of faults that students make. If mutants are a valid substitute for faults found in student-written code, and if mutant detection is correlated with manually-seeded fault detection and faulty student implementation detection, then instructors can instead evaluate student test suites using mutants generated by open-source mutation analysis tools.
James Perretta, Andrew DeOrio, Arjun Guha, Jonathan Bell 0001
ISSTA2
2021 Teaching TAs to Teach: Strategies for TA Training
abstract
"The only thing that scales with undergrads is undergrads". As Computer Science course enrollments have grown, there has been a necessary increase in the number of undergraduate and graduate teaching assistants (TAs, and UTAs). TA duties often extend far beyond grading, including designing and leading lab or recitation sections, holding office hours and creating assignments. Though advanced students, TAs need proper pedagogical training to be the most effective in their roles. Training strategies have widely varied from no training at all, to semester-long prep courses. We will explore the challenges of TA training across both large and small departments. While much of the effort has focused on teams of undergraduates, most presenters have used the same tools and strategies with their graduate students. Training for TAs should not just include the mechanics of managing a classroom, but culturally relevant pedagogy. The panel will focus on the challenges of providing "just in time", and how we manage both intra-course training and department or campus led courses.
Michael Ball 0001, Andrew DeOrio, Justin Hsia, Adam Blank
SIGCSE2
2020 A Longitudinal View of Gender Balance in a Large Computer Science Program
abstract
Computer Science has a persistent lack of women's participation. In order to best effect change, we require a more fine-grain analysis of the gender disparity as it changes throughout an undergraduate Computer Science curriculum. In this paper, we use a quantitative approach to highlight, with greater specificity, the point in an undergraduate career where gender balance changes. We also examine the role of grades in students' decisions to stay in the course sequence. Our goal is to enable targeted interventions that will make Computer Science a more welcoming discipline. Our study examines 30,890 unique student records over ten years at a large, public research institution. The records include students who took a Computer Science course over the past ten years. The dataset contains information about gender, majors, minors, academic level, and GPA. The dataset also includes a record from each course taken by each student and their final grade. We observed a modest increase in women's participation in all Computer Science courses over the past ten years. Despite this increase, the gender disparity is still large. Through our analysis, we found that women consistently choose not to continue through the Computer Science sequence at a higher rate than men. This higher attrition could be linked to women receiving lower grades in most introductory CS courses despite having the same or higher GPAs than men. Our results reveal specific areas where intervention can be the most effective in changing the stubborn gender disparity in Computer Science.
Amy Baer, Andrew DeOrio
SIGCSE2
2020 Teaching TAs To Teach: Strategies for TA Training
abstract
"The only thing that scales with undergrads is undergrads". As Computer Science course enrollments have grown, there has been a necessary increase in the number of undergraduate and graduate teaching assistants (TAs, and UTAs). TA duties often extend far beyond grading, including designing and leading lab or recitation sections, holding office hours and creating assignments. Though advanced students, TAs need proper pedagogical training to be the most effective in their roles. Training strategies have widely varied from no training at all, to semester-long prep courses. We will explore the challenges of TA training across both large and small departments. While much of the effort has focused on teams of undergraduates, most presenters have used the same tools and strategies with their graduate students. Training for TAs should not just include the mechanics of managing a classroom, but culturally relevant pedagogy. The panel will focus on the challenges of providing "just in time", and how we manage both intra-course training and department or campus led courses.
Michael Ball 0001, Justin Hsia, Heather Pon-Barry, Andrew DeOrio, Adam Blank
SIGCSE4
2019 Gender-balanced TAs from an Unbalanced Student Body
abstract
Increasing participation of women and underrepresented minorities is a key challenge in the field of Computer Science Education. Balanced representation of these groups among teaching assistants in Computer Science courses influences recruitment and retention of underrepresented students. At the same time, the status-quo reduced participation of these students makes it more difficult to hire instructional staff from underrepresented groups. In this paper, we describe our experience evaluating candidates with teaching-demonstration videos, followed by in-person interviews, to hire a gender-balanced set of undergraduate TAs for a large-scale CS2 course. Our research goal is to quantitatively assess gender balance throughout the hiring process. Our initial applicant pool is just one-sixth women, but we found that women applicants perform better in our application process than men, resulting in a gender-balanced course staff without making hiring decisions based on the gender of applicants. We show that our approach results in a more gender-balanced teaching staff than hiring based on applicant GPA. We also use course-evaluation data to demonstrate that women perform as well as men as teaching assistants in CS2, and that the overall quality of our teaching assistants has remained high after the hiring-process change.
Amir Kamil, James Juett, Andrew DeOrio
SIGCSE3
2015 Post-Silicon Validation of Multiprocessor Memory Consistency
abstract
Shared-memory chip-multiprocessor (CMP) architectures define memory consistency models that establish the ordering rules for memory operations from multiple threads. Validating the correctness of a CMP's implementation of its memory consistency model requires extensive monitoring and analysis of memory accesses while multiple threads are executing on the CMP. In this paper, we present a low overhead solution for observing, recording and analyzing shared-memory interactions for use in an emulation and/or post-silicon validation environment. Our approach leverages portions of the CMP's own data caches, augmented only by a small amount of hardware logic, to log information relevant to memory accesses. After transferring this information to a central memory location, we deploy our own analysis algorithm to detect any possible memory consistency violations. We build on the property that a violation corresponds to a cycle in an appropriately defined graph representing memory interactions. The solution we propose allows a designer to choose where to run the analysis algorithm: 1) on the CMP itself; 2) on a separate processor residing on the validation platform; or 3) off-line on a separate host machine. Our experimental results show an 83% bug detection rate, in our testbed CMP, over three distinct memory consistency models, namely: relaxed-memory order, total-store order, and sequential consistency. Finally, note that our solution can be disabled in the final product, leading to zero performance overhead and a per-core area overhead that is smaller than the size of a physical integer register file in a modern processor.
Biruk Mammo, Valeria Bertacco, Andrew DeOrio, Ilya Wagner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Machine learning-based anomaly detection for post-silicon bug diagnosis
abstract
The exponentially growing complexity of modern processors intensifies verification challenges. Traditional pre-silicon verification covers less and less of the design space, resulting in increasing post-silicon validation effort. A critical challenge is the manual debugging of intermittent failures on prototype chips, where multiple executions of a same test do not yield a consistent outcome. We leverage the power of machine learning to support automatic diagnosis of these difficult, inconsistent bugs. During post-silicon validation, lightweight hardware logs a compact measurement of observed signal activity over multiple executions of a same test: some may pass, somemay fail. Our novel algorithm applies anomaly detection techniques similar to those used to detect credit card fraud to identify the approximate cycle of a bug's occurrence and a set of candidate root-cause signals. Compared against other state-of-the-art solutions in this space, our new approach can locate the time of a bug's occurrence with nearly 4x better accuracy when applied to the complex OpenSPARC T2 design.
Andrew DeOrio, Qingkun Li, Matthew Burgess, Valeria Bertacco
DATE1
2012 Bridging pre- and post-silicon debugging with BiPeD
abstract
The growing complexity of modern chips has caused an increasing share of the verification effort to shift towards post-silicon validation. This phase is challenged by poor observability, limited off-chip bandwidth, and complex, concurrent communication interfaces. Furthermore, pre-silicon verification and post-silicon validation methodologies are very different and share little information between them. As as result, the diagnosis and debugging of post-silicon failures is very much an ad-hoc and time-consuming task that is largely unable to leverage the vast body of design knowledge available in pre-silicon.
Andrew DeOrio, Valeria Bertacco
ICCAD1
2012 Comprehensive online defect diagnosis in on-chip networks
abstract
We propose a comprehensive yet low-cost solution for online detection and diagnosis of permanent faults in on-chip networks. Using error syndrome collection and packet/flit-counting techniques, high-resolution defect diagnosis is feasible in both datapath and control logic of the on-chip network without injecting any test traffic or incurring significant performance overhead.
Amirali Ghofrani, Ritesh Parikh, Saeed Shamshiri, Andrew DeOrio, Kwang-Ting Cheng, Valeria Bertacco
VTS4
2012 A Reliable Routing Architecture and Algorithm for NoCs
abstract
Aggressive transistor scaling continues to drive increasingly complex digital designs. The large number of transistors available today enables the development of chip multiprocessors that include many cores on one die communicating through an on-chip interconnect. As the number of cores increases, scalable communication platforms, such as networks-on-chip (NoCs), have become more popular. However, as the sole communication medium, these interconnects are a single point of failure so that any permanent fault in the NoC can cause the entire system to fail. Compounding the problem, transistors have become increasingly susceptible to wear-out related failures as their critical dimensions shrink. As a result, the on-chip network has become a critically exposed unit that must be protected. To this end, we present Vicis, a fault-tolerant architecture and companion routing protocol that is robust to a large number of permanent failures, allowing communication to continue in the face of permanent transistor failures. Vicis makes use of a two-level approach. First, it attempts to work around errors within a router by leveraging reconfigurable architectural components. Second, when faults within a router disable a link's connectivity, or even an entire router, Vicis reroutes around the faulty node or link with a novel, distributed routing algorithm for meshes and tori. Tolerating permanent faults in both the router components and the reliability hardware itself, Vicis enables graceful performance degradation of networks-on-chip.
Andrew DeOrio, David Fick, Valeria Bertacco, Dennis Sylvester, David T. Blaauw, Gregory K. Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2011 ARIADNE: Agnostic Reconfiguration in a Disconnected Network Environment
abstract
Extreme transistor technology scaling is causing increasing concerns in device reliability: the expected lifetime of individual transistors in complex chips is quickly decreasing, and the problem is expected to worsen at future technology nodes. With complex designs increasingly relying on Networks-on-Chip (NoCs) for on-chip data transfers, a NoC must continue to operate even in the face of many transistor failures. Specifically, it must be able to reconfigure and reroute packets around faults to enable continued operation, i.e., generate new routing paths to replace the old ones upon a failure. In addition to these reliability requirements, NoCs must maintain low latency and high throughput at very low area budget. In this work, we propose a distributed reconfiguration solution named Ariadne, targeting large, aggressively scaled, unreliable NoCs. Ariadne utilizes up*/down* for fast routing at high bandwidth, and upon any number of concurrent network failures in any location, it reconfigures to discover new resilient paths to connect the surviving nodes. Experimental results show that Ariadne provides a 40%-140% latency improvement (when subject to 50 faults in a 64-node NoC) over other on-chip state-of-the-art fault tolerant solutions, while meeting the low area budget of on-chip routers with an overhead of just 1.97%.
Konstantinos Aisopos, Andrew DeOrio, Li-Shiuan Peh, Valeria Bertacco
PACT2
2011 DRAIN: distributed recovery architecture for inaccessible nodes in multi-core chips
abstract
As transistor dimensions continue to scale deep into the nanometer regime, silicon reliability is becoming a chief concern. At the same time, transistor counts are scaling up, enabling the design of highly integrated chips with many cores and a complex interconnect fabric, often a network on chip (NoC). Particularly problematic is the case when the accumulation of permanent hardware faults leads to disconnected cores in the system. In order to maintain correct system operation, it is necessary to salvage the data from these isolated nodes.
Andrew DeOrio, Konstantinos Aisopos, Valeria Bertacco, Li-Shiuan Peh
DAC1
2011 Post-silicon bug diagnosis with inconsistent executions
abstract
The complexity of modern chips intensifies verification challenges, and an increasing share of this verification effort is shouldered by post-silicon validation. Focusing on the first silicon prototypes, post-silicon validation poses critical new challenges such as intermittent failures, where multiple executions of a same test do not yield a consistent outcome. These are often due to on-chip asynchronous events and electrical effects, leading to extremely time-consuming, if not unachievable, bug diagnosis and debugging processes. In this work, we propose a methodology called BPS (Bug Positioning System) to support the automatic diagnosis of these difficult bugs. During post-silicon validation, lightweight BPS hardware logs a compact encoding of observed signal activity over multiple executions of the same test: some passing, some failing. Leveraging a novel post-analysis algorithm, BPS uses the logged activity to diagnose the bug, identifying the approximate manifestation time and critical design signals. We found experimentally that BPS can localize most bugs down to the exact root signal and within about 1,000 clock cycles of their occurrence.
Andrew DeOrio, Daya Shanker Khudia, Valeria Bertacco
ICCAD1
2011 Functional correctness for CMP interconnects
abstract
As transistor counts continue to scale, modern designs are transitioning towards large chip multi-processors (CMPs). In order to match the advancing performance of CMPs, on-chip interconnects are becoming increasingly complex, commonly deploying advanced network-on-chip (NoC) structures. Ensuring the correct operation of these system-level infrastructures has become increasingly problematic and, in order to avoid the potential for functional design errors manifesting into the final product, there is a need for mechanisms to safeguard communication integrity at runtime. In this paper, we propose SafeNoC, an end-to-end error detection and recovery solution to ensure the functional correctness of CMP interconnects. SafeNoC augments the existing interconnect with a simple, lightweight checker network that is guaranteed to deliver messages correctly. For each data message sent over the primary NoC, a look-ahead signature is transmitted over the checker network and is used to detect errors in the corresponding data message. If a functional communication bug is detected, a novel recovery algorithm reconstructs the data that was in flight at the time of the error occurrence, ensuring that it reaches the intended destination. In our experiments, we found that SafeNoC can recover from a wide variety of errors, with almost no performance impact in the absence of errors. A lightweight solution, SafeNoC occupies a 2.41% area overhead in a 64-core CMP, 7× smaller than common retransmission-based approaches.
Rawan Abdel-Khalek, Ritesh Parikh, Andrew DeOrio, Valeria Bertacco
ICCD3
2011 Gate-Level Simulation with GPU Computing
abstract
Functional verification of modern digital designs is a crucial, time-consuming task impacting not only the correctness of the final product, but also its time to market. At the heart of most of today’s verification efforts is logic simulation, used heavily to verify the functional correctness of a design for a broad range of abstraction levels. In mainstream industry verification methodologies, typical setups coordinate the validation effort of a complex digital system by distributing logic simulation tasks among vast server farms for months at a time. Yet, the performance of logic simulation is not sufficient to satisfy the demand, leading to incomplete validation processes, escaped functional bugs, and continuous pressure on the EDA industry to develop faster simulation solutions. In this work we propose GCS, a solution to boost the performance of logic simulation, gate-level simulation in particular, by more than a factor of 10 using recent hardware advances in Graphic Processing Unit (GPU) technology. Noting the vast available parallelism in the hardware of modern GPUs and the inherently parallel structures of gate-level netlists, we propose novel algorithms for the efficient mapping of complex designs to parallel hardware. Our novel simulation architecture maximizes the utilization of concurrent hardware resources while minimizing expensive communication overhead. The experimental results show that our GPU-based simulator is capable of handling the validation of industrial-size designs while delivering more than an order-of-magnitude performance improvements on average, over the fastest multithreaded simulators commercially available.
Debapriya Chatterjee, Andrew DeOrio, Valeria Bertacco
ACM Trans. Design Autom. Electr. Syst.2
2010 Electronic design automation for social networks
abstract
Online social networks are a growing internet phenomenon: they connect millions of individuals through sharing of common interests, political and religious views, careers, etc. Social networking websites are observing an ever-increasing number of regular users, who rely on this virtual medium to connect with friends and share in the community. As a result, they have become the repository of a vast amount of demographic information, which could deliver valuable insights to businesses and individuals. However, as of today, this data is for the most part still untapped, partly because of the complexity entailed by analyzing some of these vast social connectivity graphs. Another area that deals with large data sets is Electronic Design Automation (EDA), the result of increasingly complex computer systems. The powerful tools used to deal with these data sets open many possibilities for social networks. In this work we propose to study interesting aspects of social networks by deploying some of the solutions commonly used in EDA.
Andrew DeOrio, Valeria Bertacco
DAC1
2009 Event-driven gate-level simulation with GP-GPUs
abstract
Logic simulation is a critical component of the design tool flow in modern hardware development efforts. It is used widely -- from high-level descriptions down to gate-level ones -- to validate several aspects of the design, particularly functional correctness. Despite development houses investing vast resources in the simulation task, particularly at the gate-level, it is still far from achieving the performance demands required to validate complex modern designs.
Debapriya Chatterjee, Andrew DeOrio, Valeria Bertacco
DAC2
2009 Human computing for EDA
abstract
Electronic design automation is a field replete with challenging -- and often intractable -- problems to be solved over very large instances. As a result, the field of design automation has developed a staggering expertise in approximations, abstractions and heuristics as a means to side-step the NP-hard nature of these problems. Approximations and heuristics are at heart a natural application of human reasoning. In this work we propose to harness human potential to solve some of these problems. Specifically, we propose FunSAT, a massively multi-player puzzle game for SAT solving. FunSAT leverages visual pattern recognition skills, abstract perception and intuitive strategy skills of humans to solve complex SAT instances. Players are motivated by the puzzle-solving challenges of the game and by its social interaction aspects.
Andrew DeOrio, Valeria Bertacco
DAC1
2009 Vicis: a reliable network for unreliable silicon
abstract
Process scaling has given designers billions of transistors to work with. As feature sizes near the atomic scale, extensive variation and wearout inevitably make margining uneconomical or impossible. The ElastIC project seeks to address this by creating a large-scale chip-multiprocessor that can self-diagnose, adapt, and heal. Creating large, flexible designs in this environment naturally lends itself to the repetitive nature of network-on-chip (NoC), but the loss of a single link or router will result in complete network failure. In this work we present Vicis, an ElastIC-style NoC that can tolerate the loss of many network components due to wearout induced hard faults. Vicis uses the inherent redundancy in the network and its routers in order to maintain correct operation while incurring a much lower area overhead than previously proposed N-modular redundancy (NMR) based solutions. Each router has a built-in-self-test (BIST) that diagnoses the locations of hard fault and runs a number of algorithms to best use ECC, port swapping, and a crossbar bypass bus to mitigate them. The routers work together to run distributed algorithms to solve network-wide problems as well, protecting the networking against critical failures in individual routers. In this work we show that with stuck-at fault rates as high as 1 in 2000 gates, Vicis will continue to operate with approximately half of its routers still functional and communicating.
David Fick, Andrew DeOrio, Valeria Bertacco, David T. Blaauw, Dennis Sylvester
DAC2
2009 GCS: High-performance gate-level simulation with GPGPUs
abstract
In recent years, the verification of digital designs has become one of the most challenging, time consuming and critical tasks in the entire hardware development process. Within this area, the vast majority of the verification effort in industry relies on logic simulation tools. However, logic simulators deliver limited performance when faced with vastly complex modern systems, especially synthesized netlists. The consequences are poor design coverage, delayed product releases and bugs that escape into silicon. Thus, we developed a novel GPU-accelerated logic simulator, called GCS, optimized for large structural netlists. By leveraging the vast parallelism offered by GP-GPUs and a novel netlist balancing algorithm tuned for the target architecture, we can attain an order-of-magnitude performance improvement on average over commercial logic simulators, and simulate large industrial-size designs, such as the OpenSPARC processor core design.
Debapriya Chatterjee, Andrew DeOrio, Valeria Bertacco
DATE2
2009 A highly resilient routing algorithm for fault-tolerant NoCs
abstract
Current trends in technology scaling foreshadow worsening transistor reliability as well as greater numbers of transistors in each system. The combination of these factors will soon make long-term product reliability extremely difficult in complex modern systems such as systems on a chip (SoC) and chip multiprocessor (CMP) designs, where even a single device failure can cause fatal system errors. Resiliency to device failure will be a necessary condition at future technology nodes. In this work, we present a network-on-chip (NoC) routing algorithm to boost the robustness in interconnect networks, by reconfiguring them to avoid faulty components while maintaining connectivity and correct operation. This distributed algorithm can be implemented in hardware with less than 300 gates per network router. Experimental results over a broad range of 2D-mesh and 2D-torus networks demonstrate 99.99% reliability on average when 10% of the interconnect links have failed.
David Fick, Andrew DeOrio, Gregory K. Chen, Valeria Bertacco, Dennis Sylvester, David T. Blaauw
DATE2
2009 Dacota: Post-silicon validation of the memory subsystem in multi-core designs
abstract
The number of functional errors escaping design verification and being released into final silicon is growing, due to the increasing complexity and shrinking production schedules of modern processor designs. Recent trends towards chip multiprocessors (CMPs) are exacerbating the problem because of their complex and sometimes non-deterministic memory subsystems, prone to subtle but devastating bugs. This deteriorating situation calls for high-efficiency, high-coverage results in functional validation, results that are be achieved by leveraging the performance of post-silicon validation, that is, those verification tasks that are executed directly on prototype hardware. The orders-of-magnitude faster testing in post-silicon enables designers to achieve much higher coverage before customer release, but only if the limitations of this technology in diagnosis and internal node observability could be overcome. In this work, we unlock the full performance of post-silicon validation through Dacota, a new high-coverage solution for validating memory operation ordering in CMPs. When activated, Dacota reconfigures a portion of the cache storage to log memory accesses using a compact data-coloring scheme. Logs are periodically aggregated and checked by a distributed algorithm running in-situ on the CMP to verify correct memory operation ordering. When the design is ready for customer shipment, Dacota can be deactivated, releasing all cache storage, and only leaving a small silicon area footprint, less than 0.01% (three orders of magnitude smaller than previous solutions). We found experimentally that Dacota is effective in exposing memory subsystem bugs, and it delivers its high coverage capabilities at a 26% performance slowdown (only during validation) for real-world applications.
Andrew DeOrio, Ilya Wagner, Valeria Bertacco
HPCA1
2009 Inferno: Streamlining Verification With Inferred Semantics
abstract
Understanding designers' intentions and accurately verifying a design are major obstacles for verification engineers today. Currently available debugging tools, such as waveform viewers, are unwieldy, often requiring the user to search through millions of cycles of logic simulation data to locate a problem. In this paper, we present Inferno, a novel solution capable of automatically extracting semantic information from a design's interface from simulation information. The semantic structure of an interface's communication protocol is presented to the user as a set of transactions, that is, monolithic communication units that have typically been observed several times during the logic simulation. Transactions can graphically be presented to the user and used as an aid to understand and validate the communication protocol of a design's interface. In addition, approved transactions can also be encoded as assertions expressed in a hardware description language (HDL) and used in constrained-random simulation to certify that the interface protocol adheres to the set of observed (and user-approved) transactions. Moreover, we developed a new closed-loop verification methodology based on Inferno, called transactional verification, which leverages approved transactions to describe correct design behavior. In our methodology, transactions are concurrently extracted during a constraint-based random simulation: the anomalous ones are flagged as potentially buggy and presented to the user for inspection. In the experimental results, we evaluate the performance and the quality of the results of Inferno on a broad range of testbench designs and several of their interfaces, including a number of communication intellectual properties and the OpenSPARC T1 8-core processor from Sun.
Andrew DeOrio, Adam Bauserman, Valeria Bertacco, Beth Isaksen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2008 Post-silicon verification for cache coherence
abstract
Modern processor designs are extremely complex and difficult to validate during development, causing a growing portion of the verification effort to shift to post-silicon, after the first few hardware prototypes become available. Extremely slow simulation speeds during pre-silicon verification result in functional errors escaping into silicon, a problem that is further exacerbated by the growing complexity of the memory subsystem in multi-core platforms. In this work we present CoSMa, a novel technology offering high coverage functional post-silicon validation of cache coherence protocols in multi-core systems. It enables the detection and diagnosis of functional errors in the memory subsystem by recording at runtime a compact encoding of the operations occurring at each cache line and checking their correctness at regular intervals. We leverage the systempsilas existing memory resources to store the required activity, thus minimizing area overhead. When the system is finally ready for customer shipment, CoSMa can be completely disabled, eliminating any performance or memory overhead. We reproduce in our experiments a set of coherence protocol bugs based on published errata documents of commercial multi-core designs, and show that CoSMa is highly effective in detecting them.
Andrew DeOrio, Adam Bauserman, Valeria Bertacco
ICCD1