Peter Jamieson

dblp:62/4735 · also Peter A. Jamieson · DBLP profile ↗
← Back
38ranked-venue papers
19as first author
10since 2021 · last 2025
0000-0002-3741-0201ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 9 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 12 · 9 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Gate-Breaker: An LLM-Powered Netlist-to-RTL Reverse Engineering Tool
abstract
The escalating sophistication of hardware intellectual property (IP) theft, a multi-billion dollar problem for the semiconductor industry, demands novel approaches to both understanding attack vectors and fortifying defenses. This paper evaluates the potential of Large Language Models (LLMs) to reverse engineer Register Transfer Level (RTL) designs from gate-level netlists. We introduce a framework for netlist-to-RTL conversion, leveraging pattern recognition and code generation capabilities of modern LLMs. Our evaluation of four available LLM models across 156 circuit benchmarks reveals that LLMs can indeed recover functional RTL with reasonable accuracy until the RTL benchmarks become our classified 4th quartile of code complexity. We also provide a similarity metrics-based methodology to evaluate and ascertain the quality of reverse engineering. Among the evaluated models, OpenAI's O3-Mini emerged as the best performer with 75.1 % overall success rate. While performance degrades significantly on the most complex 4th quartile (53.4 % success rate), when O3-Mini does succeed on these challenging designs, it maintains a high Abstract Syntax Tree (AST) similarity of 0.924 and achieves a moderate 0.871 Control Flow Graph (CFG) similarity.
Md. Omar Faruque, Peter Jamieson, Ahmad Patooghy, Abdel-Hameed A. Badawy
IPCCC2
2024 Pre-Conference Workshop: Let's Play- Improving our Teaching in the Medium of Board Games
abstract
This workshop paper focuses on exposing faculty to how the medium of board games provides an exceptional space for faculty development for our teaching and student learning. The “Let's Play” intervention [1] takes the form of a workshop where participants have opportunities to experience role reversal through being a learner again. Participants become active learners by playing board games that help them remember the experience of being a learner again. By choosing different types and styles of games, we can provide a space for the participants to discuss broader teaching practices, such as the importance of technical vocabulary, scaffolding ideas as we teach them, and the benefits of student-centered learning approaches. Another critical aspect of this intervention is that we hope to use role reversal to remind teachers how hard it is to learn in the hope that teachers will have more empathy for their learners. In this paper, we describe our workshop structure and pertaining literature and ideas on why board games are part of this medium. NOTE that many past participants are looking for how to use games in the classroom. This workshop does not address that aspect of the medium even though we have experience in this space [2].
Peter Jamieson, Karen C. Davis, Eric James Rapos
FIE1
2024 The Seeker's Dilemma: Realistic Formulation and Benchmarking for Hardware Trojan Detection
abstract
This work focuses on advancing the security field in the hardware design space by formally defining the problem of Hardware Trojan (HT) detection. The goal is to model HT detection more closely to the real world, i.e., describing the problem as "The Seeker’s Dilemma" (an extension of Hide&Seek on a graph), where a detecting agent is unaware of whether HTs infect circuits or not. Using this problem formulation, we create a benchmark that consists of a mixture of HT-free and HT-infected restructured circuits while preserving their original functionalities. The restructured circuits are randomly infected by HTs, causing a situation where the defender is uncertain if a circuit is infected. Our innovative dataset will help the community better judge the detection quality of different methods by comparing their success rates in circuit classification. We use our benchmark to evaluate three state-of-the-art HT detection tools to show baseline results for this approach. We use Principal Component Analysis to assess the strength of our benchmark, where we observe that some restructured HT-infected circuits are mapped closely to HT-free circuits, leading to significant label misclassification by detectors.
Amin Sarihi, Ahmad Patooghy, Abdel-Hameed A. Badawy, Peter Jamieson
IPCCC4
2024 Trojan playground: a reinforcement learning framework for hardware Trojan insertion and detection
Amin Sarihi, Ahmad Patooghy, Peter Jamieson, Abdel-Hameed A. Badawy
J. Supercomput.3
2023 With ChatGPT, Do We have to Rewrite Our Learning Objectives - CASE Study in Cybersecurity
abstract
With the emergence of Artificial Intelligent chatbot tools such as ChatGPT and code writing AI tools such as GitHub Copilot, educators need to question what and how we should teach our courses and curricula in the future. In reality, automated tools may result in certain academic fields being deeply reduced in the number of employable people. In this work, we make a case study of cybersecurity undergrad education by using the lens of “Understanding by Design” (UbD). First, we provide a broad understanding of learning objectives (LOs) in cybersecurity from a computer science perspective. Next, we dig a little deeper into a curriculum with an undergraduate emphasis on cybersecurity and examine the major courses and their LOs for our cybersecurity program at Miami University. With these details, we perform a thought experiment on how attainable the LOs are with the above-described tools, asking the key question “what needs to be enduring concepts?” learned in this process. If an LO becomes something that the existence of automation tools might be able to do, we then ask “what level is attainable for the LO that is not a simple query to the tools?”. With this exercise, we hope to establish an example of how to prompt ChatGPT to accelerate students in their achievements of LOs given the existence of these new AI tools, and our goal is to push all of us to leverage and teach these tools as powerful allies in our quest to improve human existence and knowledge.
Peter Jamieson, Suman Bhunia, Dhananjai M. Rao
FIE1
2022 Hardware Trojan Insertion Using Reinforcement Learning
abstract
This paper utilizes Reinforcement Learning (RL) as a means to automate the Hardware Trojan (HT) insertion process to eliminate the inherent human biases that limit the development of robust HT detection methods. An RL agent explores the design space and finds circuit locations that are best for keeping inserted HTs hidden. To achieve this, a digital circuit is converted to an environment in which an RL agent inserts HTs such that the cumulative reward is maximized. Our toolset can insert combinational HTs into the ISCAS-85 benchmark suite with variations in HT size and triggering conditions. Experimental results show that the toolset achieves high input coverage rates (100% in two benchmark circuits) that confirms its effectiveness. Also, the inserted HTs have shown a minimal footprint and rare activation probability.
Amin Sarihi, Ahmad Patooghy, Peter Jamieson, Abdel-Hameed A. Badawy
ACM Great Lakes Symposium on VLSI3
2021 Personalizing Online Computer Engineering Resources and Labs for Digital, Embedded, and Computer System Courses
abstract
The immediate and new challenges of the current Covid-19 pandemic have made it hard for all of us; in this work in progress as an innovative practice, we look to both leverage the challenges and share our work so that others might see some benefit to these times and improve their courses. In particular, our focus is on creating automated tools to physically create exams and sample code for Digital Systems and Computer Architecture courses. Additionally, we focus on shifting, traditional in-person labs to online, personalized formats for Digital Systems and Embedded Systems so that both educator and learner can still provide/experience virtual computer engineering education. We focus on three courses (Digital System Design, Computer Architecture/Organization, and Embedded System Design) as they are fundamentally driven by the implementation and execution of “algorithms”. From this starting point, we have created tools to generate sample code and exams, and have found means to virtualize labs and hands-on activities. In particular, we have created Python tools that allow educators to personalize code and problems, create these codes/problems (as text files or incorporated in word documents), and email these documents to students. This provides the means to create problems and code examples that are different from their peers and can be assessed on a per individual basis to alleviate some of the challenges with live and proctored exams. Additionally, we have found tools and methods for students to virtually perform the hands-on portion of these three subjects without the need for traditional lab equipment. This requires students to spend less than 100 USD worth of equipment and software. Our goal is to share these resources and our methodologies to help in this time of crisis. Additionally, these tools and methods have forced us to innovate our teaching, and we will, likely, use these tools and methods in the future. We share these tools in hope that the computer engineering education community will join this process to help us all improve our student's education.
Peter Jamieson, Ricardo S. Ferreira 0001, José A. M. Nacif
FIE1
2021 Is It Time to Include High-Level Synthesis Design in Digital System Education for Undergraduate Computer Engineers?
abstract
We ask the question, “should High-level Synthesis (HLS) design be part of an undergraduate computer engineering education while learning digital system design?”. Current trends in industry include an increasing demand for engineers who can build FPGA systems. FPGAs, just like other chips, continue to improve in terms of complexity, speed, available resources, and new features. The design complexity for using an FPGA continues to grow and Hardware Description Languages (HDLs), though a step up in design efficiency compared to schematic design, is a low-level approach akin to assembly language for programmers, and HDLs limits the productivity of an engineer. HLS tools, such as Legup, Intel HLS, and Xilinx's Vivado attempt to provide designers with a higher-level design abstraction providing a means to describe their computation in high-level languages - such as C. As these tools become more mainstream in industry, when should education follow? In this work, we explore how HLS tools might be used by an undergraduate by looking at exemplar designs, a simple RISC-V processor and a basic C loop, and implementing the design in both HDL and HLS. We then analyze the FPGA cost of each implementation. Next, we provide a philosophical discussion based on this experience on what the pros and cons of moving students to HLS design abstraction level are.
Isaac Nelson, Ricardo S. Ferreira 0001, José A. M. Nacif, Peter Jamieson
ISCAS4
2021 TRAVERSAL: A Fast and Adaptive Graph-Based Placement and Routing for CGRAs
abstract
Coarse grain reconfigurable architectures (CGRAs) are an emerging hybrid computational architecture that has the parallel customization benefits of low-level logic devices, such as FPGAs and ASICs, while the relative coarseness of these architectures makes CGRAs easier to design for, which is more similar to the traditional processor. In the process of mapping designs to CGRAs, flexible, fast, and adaptive placement and routing (P&R) is fundamental in order to implement efficient run-time reconfigurable frameworks. It is well-known that P&R is an NP-complete problem, and thus, solutions rely on heuristics to achieve quality results with acceptable execution times. CGRA P&R has different constraints compared to traditional VLSI P&R, e.g., path latency balancing and modulo scheduling of loops. In this work, we propose a graph-based P&R approach that uses graph traversals to map designs to CGRAs. Additionally, we parallelize our approach with a graph-based greedy heuristic that executes on a GPU. We compare our proposed P&R approach with the CGRA-ME framework, which implements simulated annealing and integer linear programming placement algorithms. Our results show that this new approach can generate optimal mappings and improve the execution run-time up to several orders of magnitude. Furthermore, considering spatial mapping at the millisecond scale, our GPU approach is one order of magnitude faster compared to the state-of-the-art tool VPR.
Michael Canesche, Marcelo de Matos Menezes, Westerley Carvalho, Frank Sill, Peter Jamieson, José A. M. Nacif, Ricardo S. Ferreira 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 You Only Traverse Twice: A YOTT Placement, Routing, and Timing Approach for CGRAs
abstract
Coarse-grained reconfigurable architecture (CGRA) mapping involves three main steps: placement, routing, and timing. The mapping is an NP-complete problem, and a common strategy is to decouple this process into its independent steps. This work focuses on the placement step, and its aim is to propose a technique that is both reasonably fast and leads to high-performance solutions. Furthermore, a near-optimal placement simplifies the following routing and timing steps. Exact solutions cannot find placements in a reasonable execution time as input designs increase in size. Heuristic solutions include meta-heuristics, such as Simulated Annealing (SA) and fast and straightforward greedy heuristics based on graph traversal. However, as these approaches are probabilistic and have a large design space, it is not easy to provide both run-time efficiency and good solution quality. We propose a graph traversal heuristic that provides the best of both: high-quality placements similar to SA and the execution time of graph traversal approaches. Our placement introduces novel ideas based on “you only traverse twice” (YOTT) approach that performs a two-step graph traversal. The first traversal generates annotated data to guide the second step, which greedily performs the placement, node per node, aided by the annotated data and target architecture constraints. We introduce three new concepts to implement this technique: I/O and reconvergence annotation, degree matching, and look-ahead placement. Our analysis of this approach explores the placement execution time/quality trade-offs. We point out insights on how to analyze graph properties during dataflow mapping. Our results show that YOTT is 60.6 , 9.7 , and 2.3 faster than a high-quality SA, bounding box SA VPR, and multi-single traversal placements, respectively. Furthermore, YOTT reduces the average wire length and the maximal FIFO size (additional timing requirement on CGRAs) to avoid delay mismatches in fully pipelined architectures.
Michael Canesche, Westerley Carvalho, Lucas Reis, Matheus Aguilar de Oliveira, Salles V. G. Magalhães, Peter Jamieson, José A. M. Nacif, Ricardo S. Ferreira 0001
ACM Trans. Embed. Comput. Syst.6
2019 READY: A Fine-Grained Multithreading Overlay Framework for Modern CPU-FPGA Dataflow Applications
abstract
In this work, we propose a framework called REconfigurable Accelerator DeploY (READY), the first framework to support polynomial runtime mapping of dataflow applications in high-performance CPU-FPGA platforms. READY introduces an efficient mapping with fine-grained multithreading onto an overlay architecture that hides the latency of a global interconnection network. In addition to our overlay architecture, we show how this system helps solve some of the challenges for FPGA cloud computing adoption in high-performance computing. The framework encapsulates dataflow descriptions by using a target independent, high-level API, and a dataflow model that allows for explicit spatial and temporal parallelism. READY directly maps the dataflow kernels onto the accelerator. Our tool is flexible and extensible and provides the infrastructure to explore different accelerator designs. We validate READY on the Intel Harp platform, and our experimental results show an average 2x execution runtime improvement when compared to an 8-thread multi-core processor.
Lucas B. da Silva, Ricardo S. Ferreira 0001, Michael Canesche, Marcelo de Matos Menezes, Maria D. Vieira, Jeronimo Costa Penha, Peter Jamieson, José A. M. Nacif
ACM Trans. Embed. Comput. Syst.7
2017 WIP - A modification to the case study method to teach students to read academic papers
abstract
Most learners are introduced to academic papers in their graduate work. Typically, most of us learn to read and understand these papers by reading many of them and listening to more senior colleagues and teachers describe and interpret these papers in the class or reading groups. As our experience grows we become better at this skill. There is a need to read papers in a particular research area to provide an understanding of what has been done in the area, and what questions remain unanswered. We describe a work-in-progress to help teach students to read academic papers and build a base understanding of a field under the broad framework of case studies adopted and developed by business educators. This approach requires students to prepare for a discussion on a particular case in a class, mostly, independent of the instructor. We have applied a number of these best practices from case study teaching and applied a modified approach to a sequence of papers a research area during a course. We applied our modified technique in 2015 with 10 students, and students commented on how the approach helped them, significantly, in being able to read and understand academic papers.
Peter Jamieson
FIE1
2016 Improved method for creating criterion maps for automatic mind map analysis
abstract
In this work, we continue our study on analyzing student created mind maps automatically by providing a new methodology to select the technical vocabulary that students use in their mind maps. The basis of our previous experiments is an instructor chooses a set of twenty words used within the course that will be the set of words to test in mind maps. The instructor then creates their own mind map with this set of words, which is called the criterion map. Next, students create their mind maps using the same twenty words, and the criterion map and student map are analyzed with each other using various algorithms to produce metrics that quantify how similar the two maps are. When this activity is repeated longitudinally over a semester we can show that students are learning if their metrics of similarity are improving over time. One challenge, however, is which twenty words should be selected by the teacher. Similarly, will the set of twenty words impact the quality of observed learning. In 2011 and 2012, we collected our mind map data based on twenty words selected with no methodology (random). In 2013 and 2014, we created a methodology where approximately forty words are initially chosen, and these forty words are reduced down to 20 by creating a larger mind map and picking the words that have low connectivity. The hypothesis here is that less connectivity in the mind map will make it easier for the student to create their own quality maps. Our results show that this new methodology improves the arithmetic average of one of our best comparison metrics for all data points by a worse case of 2.4% better and best case 75% better.
Amber Franklin, Ryan Sunderhaus, Chris Bell, Peter Jamieson
FIE4
2016 A framework to help analyze if creating a game to teach a learning objective is worth the work
abstract
Video games are a popular technology adopted by educators to help teach ideas. The benefits are due to pedagogically beneficial characteristics of such games including their ability to adapt to the learner, allow failure, and entertain and engage players. However, designing a video game is a significant effort that takes time and may not even teach the desired learning objective(s). In this work, we provide a framework that can be used by educators to help determine if the effort needed to create a video game is worth it for a given learning objective(s). Our framework blends four pedagogical ideas so that educators can consider if their game is worth the design effort; these pedagogical tools/theories include: (1) Bloom's taxonomy; (2) the Substitution Augmentation Modification Redefinition (SAMR) Model; (3) Wiggins & McTighe course design approach and filter for learning objectives; (4) what we call, pedagogical logistics. With this framework, we analyze two games we have created, and we determine if the games we created were actually worth the effort. The overall goal is to create a framework and show how it can be used to help other researchers determine if their video game idea is worth creating.
Peter Jamieson, Lindsay D. Grace
FIE1
2015 VerilogTown: cars, crashes and hardware design
abstract
VerilogTown is a game about cars, crashes and hardware design. The game is designed to help teach and practice the hardware description language, Verilog. The game uses the metaphor of traffic signals to help players understand and practice the code needed to implement combinational and sequential logic in digital circuits. Borrowing from the emerging space of human computation games, player solutions in the games can be directly transferred to embedded system for real world use.
Lindsay D. Grace, Peter Jamieson, Naoki Mizuno
Advances in Computer Entertainment2
2015 Evaluating metrics for automatic mind map assessment in various classes
abstract
Over the past three years, we have been studying how automated evaluation of student mind maps (when compared to an expert map) shows student learning for a variety of metrics. The goal of this work is to build a system that would then allow students to evaluate their understanding of the terminology in a respective field. The weakness of our studies, so far, is that our focus group to study these metrics consists of a single course in the field of Computer Engineering, and though this class has been used over multiple years to demonstrate the feasibility of our approach in a longitudinal study, a broader study needs to be done. In this work, we show how our current metrics perform across three additional fields; specifically, we have collaborated with instructors in speech pathology, communications, and political science (in addition to our traditional class in computer engineering). We then use our methodology to determine if these courses and a term long mind map exercise have similar results than previously reported and are these results evidence of student learning. Our results show that our existing metrics have similar results for one of the three new courses. However, in the two other courses, the data shows no evidence of learning based on the mind map exercise. Each of the instructors of these courses describes their experience with the activity. Additionally, we evaluate the construction of the expert maps in each course to understand if there is a graph-based structural reason why we the results might be different. We conclude by suggesting our methodology is good for courses where terminology is clearly defined and is used and studied throughout the semester, and describe some future directions for this research.
Amber Franklin, Tuo Li 0004, Peter Jamieson, Julie Semlak, Walter Vanderbush
FIE3
2015 More missing the Boat - Arduino, Raspberry Pi, and small prototyping boards and engineering education needs them
abstract
In this work, we describe a range of prototyping boards such as Arduino, Raspberry Pi, and BeagleBone Black, and we show how these devices are being used in our ECE curriculum in a range of courses for projects. We describe the continuing challenges we have with adopting such technology from an educational standpoint, and some best practices/techniques we have learned and adopted to include these devices in our courses. We believe integrating these devices into our course flow is of a huge benefit to both our curriculum and our students.
Peter Jamieson, Jeff Herdtner
FIE1
2014 A methodology for identifying and placing heterogeneous cluster groups based on placement proximity data (abstract only)
abstract
Due to the rapid growth in the size of designs and Field Programmable Gate Arrays (FPGAs), CAD run-time has increased dramatically. Reducing FPGA design compilation times without degrading circuit performance is crucial. In this work, we describe a novel approach for incremental design flows that both identifies tightly grouped FPGA logic blocks and then uses this information during circuit placement. Our approach reduces placement run-time on average by more than 17% while typically maintaining the design's critical path delay and marginally increasing its minimum channel width and wire length on average. Instead of following the traditional approach of evaluating a circuit's pre-placement netlist, this new algorithm analyzes designs post-placement to detect proximity data. It uses this information to non-aggressively extract heterogeneous cluster groupings from the design, which we call "gems," that consist of two to seventeen clusters. We modified VPR's simulated annealing placement algorithm to use our Singularity Placer, which first crushes each cluster grouping into a "singularity," to be treated as a single cluster. We then run the annealer over this condensed circuit, followed by an expansion of the singularities, and a second annealing phase for the entire expanded circuit.
Farnaz Gharibian, Lesley Shannon, Peter Jamieson
FPGA3
2014 Identifying and placing heterogeneously-sized cluster groupings based on FPGA placement data
abstract
Field Programmable Gate Arrays (FPGAs) CAD flow run-time has increased due to the rapid growth in size of designs and FPGAs. Researchers are trying to find new ways to improve compilation time without degrading design performance. In this paper, we present a novel approach that identifies tightly grouped FPGA logic blocks and then uses this information during circuit placement. Our approach is an orthogonal optimization applicable in incremental design and physical optimization, and reduces placement run-time. Specifically, we present a new algorithm that analyzes designs post-placement to extract medium-grained super-clusters that consist of two to seventeen clusters, which we call “gems”. We modified VPR's simulated annealing placement algorithm to place our mixture of gems and clusters. Our new “Singularity Annealing” algorithm first crushes each cluster grouping into a “singularity” (treated as a single cluster). Then, the Singularity Annealer is run over this condensed circuit to obtain an initial placement, followed by an expansion of the singularities. Finally, we run a second low-temperature annealing phase on the entire expanded circuit. Our results show that our system reduces placement run-time on average by 17% while maintains the designs critical path delay, and increases designs channel width, and wirelength by 2% and 6.3%, respectively. We have also presented a test case to show the re-usability of gems in an incremental design example.
Farnaz Gharibian, Lesley Shannon, Peter Jamieson
FPL3
2013 Supergenes in a genetic algorithm for heterogeneous FPGA placement
abstract
Supergenes are an addition to a genetic algorithm's genome that duplicate genes in the genome, represent local optimizations, and have the potential to be expressed overriding the duplicated gene. We introduce supergenes in a genetic algorithm for FPGA placement where a placement algorithm places a mix of fine-grain components and medium-grain components (where a medium-grain component is 2 to 10 times the size of a finegrain component). This is the first placement algorithm, to our knowledge, that can deal with such a mix of components. Our results show that supergenes improve a placement metric (clock speed of the FPGA) by approximately 10%. We also show and explore mutation operators on supergenes, and we experimentally demonstrate that the expression of a supergene can be effectively controlled via a binary function for our placement problem.
Peter Jamieson, Farnaz Gharibian, Lesley Shannon
IEEE Congress on Evolutionary Computation1
2013 More graph comparison techniques on mind maps to provide students with feedback
abstract
One of the limiting aspects in education research is the techniques available for determining if a student has learned something. In this work, the goal is to extend our exploration of how mind maps can be automatically analyzed using their graph properties to reflect student learning. In particular, a set of student mind maps are created three times during a class in both 2011 and 2012 on digital system design using a common technical vocabulary. These mind maps are analyzed by extracting graph metrics by comparison with a criterion mind map, which is an expert created mind map. The metrics are derived from traditional graph metrics (average degree and graph density), three sets of difference metrics analyzed with a internally created tool, and a graph metric invented for comparing proteins. The results of this exploratory analysis is that five of the six metrics can be used to evaluate if a student is learning and connecting the vocabulary in a given subject over time. Additionally, these five metrics are correlated to one another. This result is promising, but we emphasize that these metrics do not correlate directly to class performance based on student grades over the course, and therefore, the current goal for this measurement technique is to be used to provide the student with automated feedback on their mind maps as related to the technical vocabulary of a course. This work extends our original work by increasing the number of graph metrics that are used to automatically analyze student maps to a criterion map. The idea is to find a number of graph metrics that can then be combined to help analyze a students mind map and provide them with useful feedback. Even though our results show that compare 5 metrics and each metric can be used to observe student improvement, each of these metrics differs in how the metric can be interpreted and related to the process of learning. Therefore, our goal is to find a number of these metrics so that they can be combined to provide the student with a variety of feedback results to help them understand their errors in terms of the structure of their mind map.
Peter Jamieson
FIE1
2013 Analyzing System-Level Information's Correlation to FPGA Placement
abstract
One popular placement algorithms for Field-Programmable Gate Arrays (FPGAs) is called Simulated Annealing (SA). This algorithm tries to create a good quality placement from a flattened design that no longer contains any high-level information related to the original design hierarchy. Placement is an NP-hard problem, and as the size and complexity of designs implemented on FPGAs increases, SA does not scale well to find good solutions in a timely fashion. In this article, we investigate if system-level information can be reconstructed from a flattened netlist and evaluate how that information is realized in terms of its locality in the final placement. If there is a strong relationship between good quality placements and system-level information, then it may be possible to divide a large design into smaller components and improve the time needed to create a good quality placement. Our preliminary results suggest that the locality property of the information embedded in the system-level HDL structure (i.e. “module”, “always”, and “if” statements) is greatly affected by designer HDL coding style. Therefore, a reconstructive algorithm, called Affinity Propagation, is also considered as a possible method of generating a meaningful coarse-grain picture of the design.
Farnaz Gharibian, Lesley Shannon, Peter Jamieson, Kevin Chung
ACM Trans. Reconfigurable Technol. Syst.3
2012 Using modern graph analysis techniques on mind maps to help quantify learning
abstract
In this work, we use a graph analysis tool to measure how student-created mind maps reflect learning. Mind maps consist of words and connections between words, and this visual tool helps illustrate how an individual understands how these words connect together in a field. From an analysis standpoint, mind maps are graphs consisting of nodes connected by edges. In the fall of 2011, students created three mind maps over the duration of a digital system design course, and at each of the three intervals, these mind maps were created with the same 20 terms that were introduced throughout the course. Each student's mind maps were then digitally encoded and analyzed using a modern graph analysis tool called GraphCrunch II. Our results show that a simple analysis of graph density is a poor indicator of learning since this metric does not capture a graph's structure, and it is this structure that reflects meaning and understanding by the learner. Instead, a metric called relative graphlet frequency distance (RGF-distance), which is calculated by comparing a golden mind map (expert created mind map) to each of the students mind maps, is used to analyze each students understanding of how these words relate. Our results show that learner's mind maps decrease in RGF-distance over the period of the course, and this means that the students are building graphs more similar to that of the golden model. We, also, see that the RGF-distance over the set of students compared to their grades on an exam or overall grade in the course has some correlation, meaning that these mind maps relate to grades in terms of the learners understanding of vocabulary, but the correlation is not strong. The ultimate goal of these tools is to provide the learners with a method of getting automatic feedback on their understanding as well as learning progress in particular topics.
Peter Jamieson
FIE1
2012 The VTR project: architecture and CAD for FPGAs from verilog to routing
abstract
To facilitate the development of future FPGA architectures and CAD tools -- both embedded programmable fabrics and pure-play FPGAs -- there is a need for a large scale, publicly available software suite that can synthesize circuits into easily-described hypothetical FPGA architectures. These circuits should be captured at the HDL level, or higher, and pass through logical and physical synthesis. Such a tool must provide detailed modelling of area, performance and energy to enable architecture exploration. As software flows themselves evolve to permit design capture at ever higher levels of abstraction, this downstream full-implementation flow will always be required. This paper describes the current status and new release of an ongoing effort to create such a flow - the 'Verilog to Routing' (VTR) project, which is a broad collaboration of researchers. There are three core tools: ODIN II for Verilog Elaboration and front-end hard-block synthesis, ABC for logic synthesis, and VPR for physical synthesis and analysis. ODIN II now has a simulation capability to help verify that its output is correct, as well as specialized synthesis at the elaboration step for multipliers and memories. ABC is used to optimize the 'soft' logic of the FPGA. The VPR-based packing, placement and routing is now fully timing-driven (the previous release was not) and includes new capability to target complex logic blocks. In addition we have added a set of four large benchmark circuits to a suite of previously-released Verilog HDL circuits. Finally, we illustrate the use of the new flow by using it to help architect a floating-point unit in an FPGA, and contrast it with a prior, much longer effort that was required to do the same thing.
Jonathan Rose, Jason Luu, Chi Wai Yu, Opal Densmore, Jeffrey B. Goeders, Andrew Somerville, Kenneth B. Kent, Peter Jamieson, Jason Helge Anderson
FPGA8
2011 Early project based learning improvements via a "star trek engineering room" game framework, and competition
abstract
In this work, we show how providing a constrained project framework for a second year digital design course improves the number of working student projects from 55% to 86%. Instead of an open-ended project as in previous years, we introduce an optional project framework, called "Redhawk Duels". Redhawk Duels is a game framework in which students design control algorithms and interfaces for a virtual ship. Once a competition begins, two opposing groups and their respective ships attempt to incapacitate the opposing ship by finding the opponent, shooting them, and budgeting their energy accordingly. Fifteen of the twenty-one groups in the 2010 class participated in Redhawk Duels for their final project, and 86% of these projects were working and demonstrated with sufficient complexity. The remaining six groups chose to implement open-ended projects and had a 66% success rate. This rate is similar to the 55% success rate of the 2009 class which were all open-ended projects. We surveyed the students involved to see how they felt the project helped them and how much they enjoyed the activity. The results show that the students strongly agree that participating in the framework motivated them and will help them in future engineering design projects.
Peter Jamieson
FIE1
2011 VPR 5.0: FPGA CAD and architecture exploration tools with single-driver routing, heterogeneity and process scaling
abstract
The VPR toolset has been widely used in FPGA architecture and CAD research, but has not evolved over the past decade. This article describes and illustrates the use of a new version of the toolset that includes four new features: first, it supports a broad range of single-driver routing architectures, which have superior architectural and electrical properties over the prior multidriver approach (and which is now employed in the majority of FPGAs sold). Second, it can now model, for placement and routing a heterogeneous selection of hard logic blocks. This is a key (but not final) step toward the incluion of blocks such as memory and multipliers. Third, we provide optimized electrical models for a wide range of architectures in different process technologies, including a range of area-delay trade-offs for each single architecture. Finally, to maintain robustness and support future development the release includes a set of regression tests for the software. To illustrate the use of the new features, we explore several architectural issues: the FPGA area efficiency versus logic block granularity, the effect of single-driver routing, and a simple use of the heterogeneity to explore the impact of hard multipliers on wiring track count.
Jason Luu, Ian Kuon, Peter Jamieson, Ted Campbell, Andy Gean Ye, Wei Mark Fang, Kenneth B. Kent, Jonathan Rose
ACM Trans. Reconfigurable Technol. Syst.3
2010 Odin II - An Open-Source Verilog HDL Synthesis Tool for CAD Research
abstract
In this work, we present Odin II, a framework for Verilog Hardware Description Language (HDL) synthesis that allows researchers to investigate approaches/improvements to different phases of HDL elaboration that have not been previously possible. Odin II's output can be fed into traditional back-end flows for both FPGAs and ASICs so that these improvements can be better quantified. Whereas the original Odin [1] provided an open source synthesis tool, Odin II's synthesis framework offers significant improvements such as a unified environment for both front-end parsing and netlist flattening. Odin II also interfaces directly with VPR [2], a common academic FPGA CAD flow, allowing an architectural description of a target FPGA as an input to enable identification and mapping of design features to custom features. Furthermore, Odin II can also read the netlists from downstream CAD stages into its netlist data-structure to facilitate analysis. Odin II can be used for a wide range of experiments; in this paper, we show three specific instances of how Odin II can be used by ASIC and FPGA researchers for more than basic synthesis. Odin II is open source and released under the MIT License.
Peter Jamieson, Kenneth B. Kent, Farnaz Gharibian, Lesley Shannon
FCCM1
2010 Odin II: an open-source verilog HDL synthesis tool for FPGA cad flows (abstract only)
abstract
Odin II is a high-level Verilog Hardware Description Language synthesis tool. This tool is a significant improvement on the original Odin for a number of reasons including Odin II does both front-end parsing and netlist flattening, Odin II interfaces with VPR architecture description of an FPGA to help identify and use available hard circuits, and Odin II can read in netlists from downstream stages in the VPR 5.0 CAD flow into its netlist data-structure. Odin II is open source and is released under the MIT License. Odin II source code, regression benchmarks, and more documentation can be found at http://www.users.muohio.edu/jamiespa/ODIN II/.
Peter Jamieson, Kenneth B. Kent
FPGA1
2010 Finding System-Level Information and Analyzing Its Correlation to FPGA Placement
abstract
One of the more popular placement algorithms for Field Programmable Gate Arrays (FPGAs) is called Simulated Annealing (SA). This algorithm tries to create a good quality placement from a flattened design that no longer contains any high-level information related to the original design hierarchy. Unfortunately, placement is an NP-hard problem and as the size and complexity of designs implemented on FPGAs increases, SA does not scale well to find good solutions in a timely fashion. As modern FPGAs can be used to implement Systems- and Networks-on-Chip, designers are required to spend an increasing amount of time waiting for place and route tools to complete that is not being matched by an increase in the power of computing work stations. In this paper, we investigate if system-level information can be reconstructed from a flattened netlist and evaluate how that information is realized in terms of its locality in the final placement. If there is a strong relationship between good quality placements and system-level information, then it may be possible to divide a large design into smaller components and improve the time needed to create a good quality placement. Our preliminary results suggest that the locality property of the information embedded in the system-level HDL structure (i.e. “module”, “always”, and “if” statements) is greatly affected by both the designer and the design itself. A reconstructive algorithm, called affinity propagation, is also considered as a possible method of generating a meaningful coarse grain picture of the design.
Farnaz Gharibian, Lesley Shannon, Peter Jamieson
FPL3
2010 Benchmarking and evaluating reconfigurable architectures targeting the mobile domain
abstract
We present the GroundHog 2009 benchmarking suite that evaluates the power consumption of reconfigurable technology for applications targeting the mobile computing domain. This benchmark suite includes seven designs; one design targets fine-grained FPGA fabrics allowing for quick state-of-the-art evaluation, and six designs are specified at a high level allowing them to target a range of existing and future reconfigurable technologies. Each of the six designs can be stimulated with the help of synthetically generated input stimuli created by an open-source tool included in the downloadable suite. Another tool is included to help verify the correctness of each implemented design. To demonstrate the potential of this benchmark suite, we evaluate the power consumption of two modern industrial FPGAs targeting the mobile domain. Also, we show how an academic FPGA framework, VPR 5.0, that has been updated for power estimates can be used to estimates the power consumption of different FPGA architectures and an open-source CAD flow mapping to these architectures.
Peter Jamieson, Tobias Becker, Peter Y. K. Cheung, Wayne Luk, Tero Rissa, Teemu Pitkänen
ACM Trans. Design Autom. Electr. Syst.1
2010 Enhancing the Area Efficiency of FPGAs With Hard Circuits Using Shadow Clusters
abstract
There is a dramatic logic density gap between field-programmable gate arrays (FPGAs) and application-specific integrated circuits, and this gap is the main reason FPGAs are not cost-effective in high-volume applications. Modern FPGAs narrow this gap by including “hard” circuits such as memories and multipliers, which are very efficient when they are used. However, if these hard circuits are not used, they go wasted (including the very expensive programmable routing that surrounds the logic), and have a negative impact on logic density. In this paper, we present an architectural concept, called shadow clusters, which seeks to mitigate this loss. A shadow cluster is a standard FPGA logic “cluster” (typically consisting of a group of lookup tables and flip-flops) that is placed “behind” every hard circuit, and can programmably, through simple, small multiplexers, replace the hard circuit in the event it is not needed. A shadow cluster is effective because the largest area cost, by far, in an FPGA is for the programmable routing that connects the logic. The shadow cluster area cost is small, and yet it enables more consistent employment of the programmable routing across applications with varying demand for hard circuits. We introduce new terminology to describe the economics of hard circuits on FGPAs, and provide a scientific way to measure the area effectiveness. We measure the area efficiency of FPGAs with and without shadow clusters, and show that a modern commercial architecture (with a fixed ratio of multipliers to soft logic) would gain 4.7% in area efficiency by employing shadow clusters. Indeed, every architecture we studied under “reasonable” conditions never showed a loss of area efficiency. Furthermore, we show that most area-efficient architecture that employs the shadow cluster concept is 12.5% better than the most area-efficient architecture without shadow clusters.
Peter Jamieson, Jonathan Rose
IEEE Trans. Very Large Scale Integr. Syst.1
2009 Benchmarking Reconfigurable Architectures in the Mobile Domain
abstract
In this paper, we introduce GroundHog 2009 benchmarking suite that can be used to evaluate the power consumption of reconfigurable technology implementing applications targeting the mobile computing domain. This benchmark suite includes seven designs; one design targets fine-grained FPGA fabrics, and six designs are specified at a high level, which allows them to target a range of reconfigurable technologies. Each of the six designs can be stimulated with synthetically generated input stimuli created by a tool included in the suite. Additionally, another tool can help verify the correctness of each implemented design. Finally,we use our benchmark suite to evaluate the power consumption of two modern FPGAs targeting the mobile domain.
Peter Jamieson, Tobias Becker, Wayne Luk, Peter Y. K. Cheung, Tero Rissa, Teemu Pitkänen
FCCM1
2009 VPR 5.0: FPGA cad and architecture exploration tools with single-driver routing, heterogeneity and process scaling
abstract
The VPR toolset [6, 7] has been widely used to perform FPGA architecture and CAD research, but has not evolved over the past decade to include many architectural features now present in modern FPGAs. This paper describes a new version of the toolset that includes four significant features: first, it now supports a broad range of single-driver routing architectures [29, 4, 16]. Single-driver routing has significantly different architectural and electrical properties from the multi-driver approach previously modelled, and is now employed in the majority of FPGAs sold. Second, the new release can now model a heterogeneous selection of hard logic blocks, which could include the hard memory and multipliers that are now ubiquitous in FPGAs. Third, we provide optimized electrical models of a wide range of architectures in different process technologies, including a range of area-delay tradeoffs for each single architecture. Prior releases of VPR did not publish even one architecture file with accurate resistance and capacitance parameters. Finally, to maintain robustness and to support future development the release includes a set of regression tests to check functionality and quality of result of the output of the tools.
Jason Luu, Ian Kuon, Peter Jamieson, Ted Campbell, Andy Gean Ye, Wei Mark Fang, Jonathan Rose
FPGA3
2008 Towards benchmarking energy efficiency of reconfigurable architectures
abstract
Energy research in reconfigurable architectures often involves legacy benchmarks such as the MCNC benchmarks. These benchmarks, however, are not well-suited for assessing energy consumption of reconfigurable technology, since they lack realistic input stimuli. This paper reviews and categorises a range of computation system benchmarks, and shows that there are no comprehensive benchmarks targeting reconfigurable architectures that would stimulate energy or power research. We review existing energy research in the field which involves microbenchmarks, in-house designs, or legacy benchmark suites used to evaluate power optimisations.
Tobias Becker, Peter Jamieson, Wayne Luk, Peter Y. K. Cheung, Tero Rissa
FPL2
2007 Architecting Hard Crossbars on FPGAs and Increasing their Area Efficiency with Shadow Clusters
abstract
We explore the architecture of on-chip hard crossbars in FPGAs and show that the area efficiency of such FPGAs can be improved when combined with shadow clusters (which are soft-logic LUT-based clusters that are architected to sit "behind" the multiplier), as an exemplar of an application circuit that appears less commonly in the designs targeting FPGAs. The metric that we seek to improve is the "frequency" that the need for hard crossbars must appear in the FPGA's target application suite for the inclusion of the hard crossbar to appear to be area-neutral. For example, we show that this break-even point for a hard 32 full-way crossbar changes from 32% of benchmarks needing to require crossbars to 9% for FPGAs with shadow clusters.
Peter Jamieson, Jonathan Rose
FPT1
2006 Enhancing the area-efficiency of FPGAs with hard circuits using shadow clusters
abstract
There is a dramatic logic density gap between FPGAs and ASICs, and this gap is the main reason FPGAs are not cost-effective in high volume applications. Modern FPGAs narrow this gap by including "hard" circuits such as memories and multipliers, which are very efficient when they are used. However, if these hard circuits are not used, they go wasted (including the very expensive programmable routing that surrounds the logic) and have a negative impact on logic density. In this paper we propose a new architectural concept, called shadow clusters, that seeks to mitigate this loss. A shadow cluster is a standard FPGA logic "cluster" that is placed "behind" every hard circuit and can programmably, through simple, small multiplexers, replace the hard circuit in the event it isn't needed. The authors measure the area-efficiency of FPGAs with and without shadow clusters and show that a modern commercial architecture (with a fixed ratio of multipliers to soft logic) would gain 4.7% in area-efficiency by employing shadow clusters. Indeed, every architecture we studied under "reasonable" conditions never showed a loss of area-efficiency. Furthermore, we show that most area-efficient architecture that employs the shadow cluster concept is 12.5 % better than the most area-efficient architecture without shadow clusters
Peter Jamieson, Jonathan Rose
FPT1
2005 A Verilog RTL Synthesis Tool for Heterogeneous FPGAs
abstract
Modern heterogeneous FPGAs contain "hard" specific-purpose structures such as blocks of memory and multipliers in addition to the completely flexible "soft" programmable logic and routing. These hard structures provide major benefits, yet raise interesting questions in FPGA CAD and architecture. To develop high-quality CAD mapping algorithms for these structures, and indeed to measure the quality of proposed new structures in the architectural domain, it is essential to have a flexible tool at the RTL synthesis level that permits heterogeneous FPGA CAD and architecture experimentation. In this paper we present a synthesis tool, called Odin, and an algorithm that permits flexible targeting of hard structures in FPGAs. Odin maps Verilog designs to two different FPGA CAD flows: Altera's Quartus, and the academic VPR CAD flow. We have expended significant effort to make the quality of this tool comparable to an industrial front-end synthesis tool, and we present mapping results for our benchmarks that show the quality of our results.
Peter Jamieson, Jonathan Rose
FPL1
2002 CableS: Thread Control and Memory Management Extensions for Shared Virtual Memory Clusters
abstract
Clusters of high-end workstations and PCs are currently used in many application domains to perform large-scale computations or as scalable servers for I/O bound tasks. Although clusters have many advantages, their applicability in emerging areas of applications has been limited. One of the main reasons for this is the fact that clusters do not provide a single system image and thus are hard to program. In this work we address this problem by providing a single-cluster image with respect to thread and memory management. We implement our system, CableS (Cluster enabled threads), on a 32-processor cluster interconnected with a low-latency, high-bandwidth system area network and conduct an early exploration of the costs involved in providing the extra functionality. We demonstrate the versatility :of Cables with a wide range of applications and show that clusters can be used to support applications that have been written for more expensive tightly-coupled systems, With very little effort on the programmer side: (a) We run legacy pthreads applications without any major modifications. (b) We use a public domain OpenMP compiler (OdinMP) to translate OpenMP programs to pthreads and execute them on our system, with no or few modifications to the translated pthreads source code. (c) We provide an implementation of the M4 macros for our pthreads system and run the SPLASH-2 applications. We also show that the overhead introduced by the extra functionality of CableS affects the parallel section of applications that have been tuned for the shared memory abstraction only in cases where the data placement is affected by operating system (WindowsNT) limitations in virtual memory mappings granularity.
Peter Jamieson, Angelos Bilas
HPCA1