VLDB 2026 Research / reviewers in the wild / expert
Phillip H. Jones
dblp:91/7033 · also Phillip H. Jones III
· DBLP profile ↗
39ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-8220-7552ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 11 · 5 since 2021Theory of computation · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MLTL Multi-type: A Typed Logic for Cyber-Physical SystemsabstractModern cyber-physical systems-of-systems (CPSoS) operate in complex systems-of-systems that must seamlessly work together to control safety- or mission-critical functions. Linear Temporal Logic (LTL) and Mission-time Linear Temporal logic (MLTL) intuitively express CPSoS requirements for automated system verification and validation. However, both LTL and MLTL presume that all signals populating the variables in a formula are sampled over the same rate and type (e.g., time or distance), and agree on a standard “time” step. Formal verification of CPSoS needs validate-able requirements expressed over (sub-)system signals of different types, such as signals sampled at different timescales, distances, or levels of abstraction, expressed in the same formula. Previous works developed more expressive logics to account for types (e.g., timescales) by sacrificing the intuitive simplicity of LTL. However, a legible direct one-to-one correspondence between a verbal and formal specification will ease validation, reduce bugs, increase productivity, and linearize the workflow from a project’s conception to actualization. Validation includes both transparency for human interpretation, and tractability for automated reasoning, as CPSoS often run on resource-limited embedded systems. To address these challenges, we introduced Mission-time Linear Temporal Logic Multi-type (Hariharan et al., Numerical Software Verification Workshop, 2022), a logic building on MLTL. MLTLM enables writing formal requirements over finite input signals (e.g., sensor signals and local computations) of different types, while maintaining the same simplicity as LTL and MLTL. Furthermore, MLTLM maintains a direct correspondence between a verbal requirement and its corresponding formal specification. Additionally, reasoning a formal specification in the intended type (e.g., hourly for an hourly rate, and per second for a seconds rate) will use significantly less memory in resource-constrained hardware. This article extends the previous work with (1) many illustrated examples on types (e.g., time and space) expressed in the same specification, (2) proofs omitted for space in the workshop version, (3) proofs of succinctness of MLTLM compared to MLTL, and (4) a minimal translation to MLTL of optimal length. Gokul Hariharan, Brian Kempa, Tichakorn Wongpiromsarn, Phillip H. Jones, Kristin Y. Rozier |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | R2U2 Playground: Visualization of a Real-time, Temporal Logic Runtime Monitor
Alexis A. Aurandt, Kristin Y. Rozier, Phillip H. Jones |
FMCAD | 3 |
| 2025 | Scalable MLTL Runtime Monitoring and Satisfiability via Bit-Vector Encoding
Christopher Johannsen, Phillip H. Jones, Kristin Y. Rozier, Tichakorn Wongpiromsarn |
FMCAD | 2 |
| 2024 | Multimodal Model Predictive Runtime Verification for Safety of Autonomous Cyber-Physical Systems
Alexis A. Aurandt, Phillip H. Jones, Kristin Y. Rozier, Tichakorn Wongpiromsarn |
FMICS | 2 |
| 2023 | R2U2 Version 3.0: Re-Imagining a Toolchain for Specification, Resource Estimation, and Optimized Observer Generation for Runtime Verification in Hardware and SoftwareabstractAbstract R2U2 is a modular runtime verification framework capable of monitoring sets of specifications in real time and in resource-constrained environments. Such environments demand that a runtime monitor be fast, easily integratable, accessible to domain experts, and have predictable resource requirements. Version 3.0 adds new features to R2U2 and its associated suite of tools that meet these needs including a new front-end compiler that accepts a custom specification language, a GUI for resource estimation, and improvements to R2U2’s internal architecture. Christopher Johannsen, Phillip H. Jones, Brian Kempa, Kristin Y. Rozier, Pei Zhang 0009 |
CAV (3) | 2 |
| 2023 | Case Studies in Applying Design Thinking to Course Design in Computer EngineeringabstractThis Innovative Practice Full Paper describes case studies from an instructional design process based on design thinking, illustrating tools used during stages of design. Instructional teams investigated the potential relevance of design thinking in engineering course design in electrical and computer engineering. Two teams of educators used a design thinking process in the redesign of two computer engineering courses, one in embedded systems and one in computer organization and architecture. The process of applying design thinking methods and tools was led by a facilitator with expertise in design thinking and electrical and computer engineering. The process leveraged specific tools and collaboration. This paper presents examples from each course, focusing on the design thinking tools used by the instructors and team members, highlighting what design thinking looks like when applied in this setting, and giving specific examples. The purpose is to suggest strategies and provide information and guidance for educators to use tools in their own course design efforts. Diane T. Rover, Henry Duwe, Phillip H. Jones, Nick Fila, Mani Mina |
FIE | 3 |
| 2023 | Impossible Made Possible: Encoding Intractable Specifications via Implied Domain Constraints
Christopher Johannsen, Brian Kempa, Phillip H. Jones, Kristin Y. Rozier, Tichakorn Wongpiromsarn |
FMICS | 3 |
| 2022 | Defining and Supporting a Debugging Mindset in Computer Engineering CoursesabstractWhile it is commonly held that debugging is a critical activity for engineers, particularly computer engineers, it is rarely a core component in engineering curriculum. Often it is either considered an innate skill or one to be developed indirectly through coursework. We argue that debugging is more authentically a constellation of mindsets or intrinsic beliefs, values, and dispositions that orient behavior. We define the debugging mindset, describe a series of activities to support the development of such a mindset, and offer evidence of a debugging mindset among students in two computer engineering project courses. Henry Duwe, Diane T. Rover, Phillip H. Jones, Nick Fila, Mani Mina |
FIE | 3 |
| 2021 | An Efficient Hardware Architecture for Sparse Convolution using Linear Feedback Shift RegistersabstractDeep convolutional neural networks (CNNs) have shown remarkable success in many computer vision tasks. However, their intensive storage, bandwidth and computational requirements limit their deployment to embedded platforms. Although several research efforts have shown that pruning redundant weights could significantly reduce storage and computations, working with sparse weights remains challenging. The irregular computation of sparse weights and the overhead of managing their representation limit the efficiency of the underlaying hardware. To address these issues, we propose a hardware-friendly pruning algorithm that generates structured sparse weights. In this algorithm, locations of non-zero weights are derived on-chip in real-time using Linear Feedback Shift Registers (LFSRs) to eliminate the overhead of managing sparse weight representations. In this paper, we also propose a hardware inference engine for sparse convolution on FPGAs. It uses LFSRs to localize non-zero weights within weights tensors and avoids copying sparse weights indices by generating them on-chip. Experimental results show that the proposed pruning method can reduce the size of VGG16, ResNet50, and InceptionV3 models by 80%, 76% and 65% with less than 2% accuracy loss. Experiments also demonstrate that our accelerator can achieve 456-534 effective GOP/s for the modern CNNs on Xilinx ZCU102, which provides a 1.2-2.7× speedup over previous sparse CNN accelerators on FPGAs. Murad Qasaimeh, Joseph Zambreno, Phillip H. Jones |
ASAP | 3 |
| 2021 | Learning and Professional Development Through Integrated Reflective Activities in Electrical and Computer Engineering CoursesabstractThis Research-to-Practice Full Paper describes the implementation of integrated reflective activities in two computer engineering courses. Reflective activities contribute to student learning and professional development. Instructional team members have been examining the need and opportunities to deepen learning by integrating reflective activities into problem-solving experiences. We implemented reflective activities using a coordinated framework for a modified Kolbian cycle. The framework consists of reflection-for-action, reflection-in-action, reflection-on-action, and composted reflections. Reflection-for-action takes place before the experience and involves thinking about and planning future actions. Reflection-in-action takes place during the experience while actively problem-solving. Reflection-on-action takes place after the problem-solving experience. Composting involves revisiting past experiences and reflections to inform future planning. We describe the reflective activities in the context of the coordinated framework, including strategies to support reflection and increase the likelihood of engagement and success. We conclude with an analysis of the activities using the CPREE framework for reflection pathways. Diane T. Rover, Henry Duwe, Mani Mina, Nick Fila, Phillip H. Jones, Lindsey S. Sleeth |
FIE | 5 |
| 2021 | Benchmarking vision kernels and neural network inference accelerators on embedded platforms
Murad Qasaimeh, Kristof Denolf, Alireza Khodamoradi, Michaela Blott, Jack Lo, Lisa Halder, Kees A. Vissers, Joseph Zambreno, Phillip H. Jones |
J. Syst. Archit. | 9 |
| 2020 | ParaHist: FPGA Implementation of Parallel Event-Based Histogram for Optical Flow CalculationabstractIn this paper, we present an FPGA-based architecture for histogram generation to support event-based camera optical flow calculation. Our proposed histogram generation mechanism reduces memory and logic resources by storing the time difference between consecutive events, instead of the absolute time of each event. Additionally, we explore the trade-off between system resource usage and histogram accuracy as a function of the precision at which time is encoded. Our results show that across three event-based camera benchmarks we can reduce the encoding of time from 32 to 7 bits with a loss of only approximately 3% in histogram accuracy. In comparison to a software implementation, our architecture shows a significant speedup. Mohammad Pivezhandi, Phillip H. Jones, Joseph Zambreno |
ASAP | 2 |
| 2020 | Introducing Autonomy in an Embedded Systems Course ProjectabstractThis Research-to-Practice Full Paper presents the redesign of a course project to promote student professional formation in engineering in the Electrical and Computer Engineering Department at Iowa State University. This is part of a larger effort to redesign core courses in the sophomore and junior years through a collaborative instructional model and pedagogical approaches that promote professional formation. A required sophomore course on embedded computer systems has been assessed and revised over multiple semesters. The redesign of the project was initiated with the purpose of promoting student professional formation, interest, autonomy and innovation, and it was undertaken using a collaborative process. This paper describes the course, final project, redesign process, assessment, results and future work. Several conclusions from the research may be useful to other educators. A small change to the course project yielded positive effects in interest and autonomy and may influence longer term effects of the project. There was evidence of difference in engagement with the project. The difference observed was not only due to option selected by students but why students selected the option. Diane T. Rover, Nick Fila, Phillip H. Jones, Mani Mina |
FIE | 3 |
| 2019 | Analyzing the Energy-Efficiency of Vision Kernels on Embedded CPU, GPU and FPGA PlatformsabstractThis paper presents a benchmark of the energy efficiency of a wide range of vision kernels on three commonly used hardware accelerators for embedded vision applications: ARM57 CPU, Jetson TX2 GPU and ZCU102 FPGA, using their vendor optimized vision libraries: OpenCV, VisionWorks and xfOpenCV. Our results show that the GPU achieves an energy/frame reduction ratio of 1.1-3.2× compared to CPU and FPGA for simple kernels. While for more complicated kernels, the FPGA outperforms the others with energy/frame reduction ratios of 1.2-22.3×. It is also observed that the FPGA performs increasingly better as a vision kernel's complexity grows. Murad Qasaimeh, Joseph Zambreno, Phillip H. Jones, Kristof Denolf, Jack Lo, Kees A. Vissers |
FCCM | 3 |
| 2018 | A Runtime Configurable Hardware Architecture for Computing Histogram-Based Feature DescriptorsabstractFeature description is an essential component of many computer vision applications. It encodes the visual contents of images in a manner that is robust against various image transformations. Computing these descriptors is computationally expensive, which causes a performance bottleneck in many embedded vision systems. Although many hardware architectures have been proposed to accelerate feature description computation, most target a single feature description algorithm under specific constraints. The lack of flexibility of such implementations increases development effort if deployed applications need to be modified or upgraded. In this paper, we propose a software configurable hardware architecture capable of computing different types of histogram-based feature descriptors without the need for re-synthesizing the hardware. The architecture takes advantage of data streaming to reduce the computational complexity of computing this class of descriptor. To illustrate the efficiency of our architecture, we deploy two of the most commonly used descriptors (SIFT and HOG) and compare their quality with software implementations. The architecture is also evaluated in terms of execution speed and resource usage and compared with dedicated hardware architectures. Our flexible architecture shows a speed up of 3× and 5× compared to state-of-the-art dedicated hardware architectures for SIFT and HOG, with resource usage overheads [LUTs, FFs, and DSPs] of [1.1×, 15×, and 1.6×] and [6.4×, 7×, and 32×] for SIFT and HOG, respectively. Murad Qasaimeh, Joseph Zambreno, Phillip H. Jones |
FPL | 3 |
| 2018 | ARMOR: A Recompilation and Instrumentation-Free Monitoring Architecture for Detecting Memory ExploitsabstractSoftware written in programming languages that permit manual memory management, such as C and C++, are often littered with exploitable memory errors. These memory bugs enable attackers to leak sensitive information, hijack program control flow, or otherwise compromise the system and are a critical concern for computer security. Many runtime monitoring and protection approaches have been proposed to detect memory errors in C and C++ applications, however, they require source code recompilation or binary instrumentation, creating compatibility challenges for applications using proprietary or closed source code, libraries, or plug-ins. This paper introduces a new approach for detecting heap memory errors that does not require applications to be recompiled or instrumented. We show how to leverage the calling convention of a processor to track all dynamic memory allocations made by an application during runtime. We also present a transparent tracking and caching architecture to efficiently verify program heap memory accesses. Performance simulations of our architecture using SPEC benchmarks and real-world application workloads show our architecture achieves hit rates over 95 percent for a 256-entry cache, resulting in only 2.9 percent runtime overhead. Security analysis using a software prototype shows our architecture detects 98 percent of heap memory errors from selected test cases in the Juliet Test Suite and real-world exploits. Alex Grieve, Phillip H. Jones, Joseph Zambreno |
IEEE Trans. Computers | 3 |
| 2017 | An embedded scalable linear model predictive hardware-based controller using ADMMabstractModel predictive control (MPC) is a popular advanced model-based control algorithm for controlling systems that must respect a set of system constraints (e.g. actuator force limitations). However, the computing requirements of MPC limits the suitability of deploying its software implementation into embedded controllers requiring high update rates. This paper presents a scalable embedded MPC controller implemented on a field-programmable gate array (FPGA) coupled with an on-chip ARM processor. Our architecture implements an Alternating Direction Method of Multipliers (ADMM) approach for computing MPC controller commands. All computations are performed using floating-point arithmetic. We introduce a software/hardware (SW/HW) co-design methodology, for which the ARM software can configure on-chip Block RAM to allow users to (1) configure the MPC controller for a wide range of plants, and (2) update at runtime the desired trajectory to track. Our hardware architecture has the flexibility to compromise between the amount of hardware resources used (regarding Block RAMs and DSPs) and the controller computing speed. For example, this flexibility gives the ability to control plants modeled by a large number of decision variables (i.e. a plant model using many Block RAMs) with a small number of computing resources (i.e. DSPs) at the cost of increased computing time. The hardware controller is verified using a Plant-on-Chip (PoC), which is configured to emulate a mass-spring system in real-time. A major driving goal of this work is to architect an SW/HW platform that brings FPGAs a step closer to being widely adopted by advanced control algorithm designers for deploying their algorithms into embedded systems. Pei Zhang 0009, Joseph Zambreno, Phillip H. Jones |
ASAP | 3 |
| 2017 | The design and integration of a software configurable and parallelized coprocessor architecture for LQR control
Pei Zhang 0009, Aaron Mills, Joseph Zambreno, Phillip H. Jones |
J. Parallel Distributed Comput. | 4 |
| 2016 | Evidence-based planning to broaden the participation of women in electrical and computer engineeringabstractThe percentages of women in undergraduate electrical and computer engineering programs at Iowa State University averages below the national average. An external assessment of diversity and inclusion provided an impetus for faculty, staff and administrators to discuss issues, focus on specific areas, and collaborate on planning. In particular, the department has teamed up with the university's Program for Women in Science and Engineering to better integrate their programs with departmental activities. This has resulted in an enhanced student experience model being designed for undergraduate ECE women. The model leverages effective practices including learning communities, leadership and professional development, academic support and advising for the ISU Engineering Basic Program, academic preparation for the ECE field, and state and national resources for inclusive ECE career awareness, recruiting and teaching. The WI-ECSEL Initiative has been designed to improve diversity and inclusion in Iowa State's electrical, computer, and software engineering programs; improve educational pathways including transfer transitions from community colleges; provide a supportive and integrated student experience; establish a community of practice for faculty; and use research to inform practice. Diane T. Rover, Joseph Zambreno, Mani Mina, Phillip H. Jones, Lora Leigh Chrystal |
FIE | 4 |
| 2016 | RAMPS: A Reconfigurable Architecture for Minimal Perfect SequencingabstractThe alignment of many short sequences of DNA, called reads, to a long reference genome is a common task in molecular biology. When the problem is expanded to handle typical workloads of billions of reads, execution time becomes critical. In this paper we present a novel reconfigurable architecture for minimal perfect sequencing (RAMPS). While existing solutions attempt to align a high percentage of the reads using a small memory footprint, RAMPS focuses on performing fast exact matching. Using the human genome as a reference, RAMPS aligns short reads hundreds of thousands of times faster than current software implementations such as SOAP2 or Bowtie, and about a thousand times faster than GPU implementations such as SOAP3. Whereas other aligners require hours to preprocess reference genomes, RAMPS can preprocess the reference human genome in a few minutes, opening the possibility of using new reference sources that are more genetically similar to the newly sequenced data. Chad Nelson, Kevin Townsend, Osama G. Attia, Phillip H. Jones, Joseph Zambreno |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | A Reconfigurable Architecture for the Detection of Strongly Connected ComponentsabstractThe Strongly Connected Components (SCCs) detection algorithm serves as a keystone for many graph analysis applications. The SCC execution time for large-scale graphs, as with many other graph algorithms, is dominated by memory latency. In this article, we investigate the design of a parallel hardware architecture for the detection of SCCs in directed graphs. We propose a design methodology that alleviates memory latency and problems with irregular memory access. The design is composed of 16 processing elements dedicated to parallel Breadth-First Search (BFS) and eight processing elements dedicated to finding intersection in parallel. Processing elements are organized to reuse resources and utilize memory bandwidth efficiently. We demonstrate a prototype of our design using the Convey HC-2 system, a commercial high-performance reconfigurable computing coprocessor. Our experimental results show a speedup of as much as 17× for detecting SCCs in large-scale graphs when compared to a conventional sequential software implementation. Osama G. Attia, Kevin Townsend, Phillip H. Jones, Joseph Zambreno |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2015 | A software configurable coprocessor-based state-space controllerabstractWe present a software configurable coprocessor-based state-space controller that can control physical processes representable by a linear state-space model. Our proposed architecture has distinct advantages over purely software or purely hardware approaches. It differs from other hardware controllers in that it is not hardwired to control one or a small range of plant types (e.g. only electric motors). Via software, an embedded systems engineer can easily reconfigure the controller to suit a wide range of controls applications that can be represented as a state-space linear model. Additionally, we introduce a novel design methodology to help bridge the gap between controls and embedded system engineering. Control of the well-understood inverted pendulum on a cart is used as an illustrative example of how the proposed hardware accelerator architecture supports our envisioned design methodology for helping bridge the gap between controls and embedded software engineering. Aaron Mills, Pei Zhang 0009, Sudhanshu Vyas, Joseph Zambreno, Phillip H. Jones |
FPL | 5 |
| 2015 | A Fault-Aware Toolchain Approach for FPGA Fault ToleranceabstractAs the size and density of silicon chips continue to increase, maintaining acceptable manufacturing yields has become increasingly difficult. Recent works suggest that lithography techniques are reaching their limits with respect to enabling high yield fabrication of small-scale devices, thus there is an increasing need for techniques that can tolerate fabrication time defects. One candidate technology to help combat these defects is reconfigurable hardware. The flexible nature of reconfigurable devices, such as Field Programmable Gate Arrays (FPGAs), makes it possible for them to route around defective areas of a chip after the device has been packaged and deployed into the field. This work presents a technique that aims to increase the effective yield of FPGA manufacturing by re-claiming a portion of chips that would be ordinarily classified as unusable. In brief, we propose a modification to existing commercial toolchain flows to make them fault aware. A phase is added to identify faults within the chip. The locations of these faults are then used by the toolchain to avoid faults during the placement and routing phase. Specifically, we have applied our approach to the Xilinx commercial toolchain flow and evaluated its tolerance to both logic and routing resource faults. Our findings show that, at a cost of 5--10% in device frequency performance, the modified toolchain flow can tolerate up to 30% of logic resources being faulty and, depending on the nature of the target application, can tolerate 1--30% of the device's routing resources being faulty. These results provide strong evidence that commercial toolchains not designed for the purpose of tolerating faults can still be greatly leveraged in the presence of faults to place and route circuits in an efficient manner. Adwait Gupte, Sudhanshu Vyas, Phillip H. Jones |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2014 | Cache design for mixed criticality real-time systemsabstractShared caches in mixed criticality systems are a source of interference for safety critical tasks. Shared memory not only leads to worst-case execution time (WCET) pessimism, but also affects the response time of safety critical tasks. In this paper, we present a criticality aware cache design which implements a Least Critical (LC) cache replacement policy, where a least recently used non-critical cache line is replaced during a cache miss. The cache acts as a Least Recently Used (LRU) cache if there are no critical lines or if all cache lines are critical in a set. In our design, data within a certain address space is given higher preference in the cache. These critical address spaces are configured using critical address range (CAR) registers. The new cache design was implemented in a Leon3 processor core, a 32bit processor compliant with the SPARC V8 architecture. Experimental results are presented that illustrate the impact of the Least Critical cache replacement policy on the response time of critical tasks, and on overall application performance as compared to a conventional LRU cache policy. N. G. Chetan Kumar, Sudhanshu Vyas, Ron Cytron, Christopher D. Gill, Joseph Zambreno, Phillip H. Jones |
ICCD | 6 |
| 2014 | A high performance systolic architecture for k-NN classificationabstractThis paper describes the architecture of the winning entry to the 2014 Memocode Design Contest, in the maximum performance category. This year's Memocode design contest asks contestants to find the 10 nearest neighbors between 1,000 testing points and 10,000,000 training points. Instead of using Euclidean distance, the contest uses Mahalanobis distance. The contest has 2 awards: the maximum performance award and the cost adjusted performance award. Our implementation uses a brute force approach that calculates the distance between every testing point to every training point. We use the Convey HC-2ex, a FPGA-based platform. However, the theory applies to software implementations as well. At the time of publication, our runtime is 0.54 seconds. Kevin Townsend, Phillip H. Jones, Joseph Zambreno |
MEMOCODE | 2 |
| 2013 | Hardware architectural support for control systems and sensor processingabstractThe field of modern control theory and the systems used to implement these controls have shown rapid development over the last 50 years. It was often the case that those developing control algorithms could assume the computing medium was solely dedicated to the task of controlling a plant, for example, the control algorithm being implemented in software on a dedicated Digital Signal Processor (DSP), or implemented in hardware using a simple dedicated Programmable Logic Device (PLD). As time progressed, the drive to place more system functionality in a single component (reducing power, cost, and increasing reliability) has made this assumption less often true. Thus, it has been pointed out by some experts in the field of control theory (e.g., Astrom) that those developing control algorithms must take into account the effects of running their algorithms on systems that will be shared with other tasks. One aspect of the work presented in this article is a hardware architecture that allows control developers to maintain this simplifying assumption. We focus specifically on the Proportional-Integral-Derivative (PID) controller. An on-chip coprocessor has been implemented that can scale to support servicing hundreds of plants, while maintaining microsecond-level response times, tight deterministic control loop timing, and allowing the main processor to service noncontrol tasks. In order to control a plant, the controller needs information about the plant's state. Typically this information is obtained from sensors with which the plant has been instrumented. There are a number of common computations that may be performed on this sensor data before being presented to the controller (e.g., averaging and thresholding). Thus in addition to supporting PID algorithms, we have developed a Sensor Processing Unit (SPU) that off-loads these common sensor processing tasks from the main processor. We have prototyped our ideas using Field Programmable Gate Array (FPGA) technology. Through our experimental results, we show our PID execution unit gives orders of magnitude improvement in response time when servicing many plants, as compared to a standard general software implementation. We also show that the SPU scales much better than a general software implementation. In addition, these execution units allow the simplifying assumption of dedicated computing medium to hold for control algorithm development. Sudhanshu Vyas, Adwait Gupte, Christopher D. Gill, Ron Cytron, Joseph Zambreno, Phillip H. Jones |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2012 | Design and evaluation of a delay-based FPGA Physically Unclonable FunctionabstractA new Physically Unclonable Function (PUF) variant was developed on an FPGA, and its quality evaluated. It is conceptually similar to PUFs developed using standard SRAM cells, except it utilizes general FPGA reconfigurable fabric, which offers several advantages. Comparison between our approach and other PUF designs indicates that our design is competitive in terms of repeatability within a given instance, and uniqueness between instances. The design can also be tuned to achieve desired response characteristics which broadens the potential range of applications. Aaron Mills, Sudhanshu Vyas, Michael Patterson, Christopher Sabotta, Phillip H. Jones, Joseph Zambreno |
ICCD | 5 |
| 2012 | Shepard: A fast exact match short read alignerabstractThe mapping of many short sequences of DNA, called reads, to a long reference genome is an common task in molecular biology. The task amounts to a simple string search, allowing for a few mismatches due to mutations and inexact read quality. While existing solutions attempt to align a high percentage of the reads using small memory footprints, Shepard is concerned with only exact matches and speed. Using the human genome, Shepard is on the order of hundreds of thousands of times faster than current software implementations such as SOAP2 or Bowtie, and about 60 times faster than GPU implementations such as SOAP3. Shepard contains two components: a software program to preprocess a reference genome into a hash table, and a hardware pipeline for performing fast lookups. The hash table has one entry for each unique 100 base pair sequence that occurs in the reference genome, and contains the index of last occurrence and the number of occurrences. To reduce the hash table size, a minimal perfect hash table is used. The hardware pipeline was designed to perform hash table lookups very quickly, on the order of 600 million lookups per second, and was implemented on a Convey HC-1 high performance reconfigurable computing system. Shepard streams all of the short reads through a custom hardware pipeline and writes the alignment data (index of last occurrence and number of occurrences) to a binary results array. Chad Nelson, Kevin Townsend, Bhavani Satyanarayana Rao, Phillip H. Jones, Joseph Zambreno |
MEMOCODE | 4 |
| 2011 | Circumventing a ring oscillator approach to FPGA-based hardware Trojan detectionabstractRing oscillators are commonly used as a locking mechanism that binds a hardware design to a specific area of silicon within an integrated circuit (IC). This locking mechanism can be used to detect malicious modifications to the hardware design, also known as a hardware Trojan, in situations where such modifications result in a change to the physical placement of the design on the IC. However, careful consideration is needed when designing ring oscillators for such a scenario to guarantee the integrity of the locking mechanism. This paper presents a case study in which flaws discovered in a ring oscillator-based Trojan detection scheme allowed for the circumvention of the security mechanism and the implementation of a large and diverse set of hardware Trojans, limited only by hardware resources. Justin Rilling, David Graziano, Jamin Hitchcock, Tim Meyer, Xinying Wang 0004, Phillip H. Jones, Joseph Zambreno |
ICCD | 6 |
| 2010 | An evaluation of a slice fault aware tool chainabstractAs FPGA sizes and densities grow, their manufacturing yields decrease. This work looks toward reclaiming some of this lost yield. Several previous works have suggested fault aware CAD tools for intelligently routing around faults. In this work we evaluate such an approach quantitatively with respect to some standard benchmarks. We also quantify the trade-offs between performance and fault tolerance in such a method. Leveraging existing CAD tools, we show up to 30% of slices being faulty can be tolerated. Such approaches could potentially allow manufacturers to sell larger chips with manufacturing faults as smaller chips using a nomenclature that appropriately captures the reduction in logic resources. Adwait Gupte, Phillip H. Jones |
DATE | 2 |
| 2010 | CANSCID-CUDAabstractThe 2010 MEMOCODE Hardware Software Co-design challenge is to implement a Deep Packet Inspection architecture, called the CANSCID - Combined Architecture for Stream Categorization and Intrusion Detection. In this short paper, we present the design details of our submission, that utilizes a Graphical Processing Unit (GPU) to accelerate the parallel regular expression matching. The target line rate of 500 Mbps is met on all of the 25 mandatory and 10 optional patterns. The design is developed using the NVIDIA CUDA framework and tested on the Tesla GPU. Michael Steffen, Veerendra Allada, Phillip H. Jones, Joseph Zambreno |
MEMOCODE | 3 |
| 2010 | Team [Ii][Ss][Uu][0-2]{4} design overview: MEMOCODE 2010 design contestabstractThis paper describes the architecture of a high-speed regular expression matching system implemented by the [Ii][Ss][Uu][0-2]{4} team from Iowa State University for the 2010 MEMOCODE competition. The purpose of this system is to detect malicious patterns in high-speed network data streams. The core functionality is implemented on a Stratix III 260 FPGA, and software running on a Xeon processor is used to transfer data to/from main memory and the FPGA. An interesting aspect of this architecture is the novel use of context switching resources to avoid buffering packets of connections whose classification are pending. The implemented solution detects malicious patterns at over 500 Mbps, and is estimated to scale to support well over 400 rules on our Stratix III FPGA. Sudhanshu Vyas, Pooja Mhapsekar, Aditya Ashok, Moinuddin Sayed, Avinash Srinivasa, Gunjan Pandey, Adam Jackson, Matthew Nelson, Anand Saggi, Harini Sundararaman, Phillip H. Jones |
MEMOCODE | 11 |
| 2009 | Towards Hardware Support for Common Sensor Processing TasksabstractSensor processing is a common task within many embedded system domains, such as in control systems, the sensor feedback is used for actuator control. In this paper we have surveyed several embedded system domains, and extracted kernels of computation that are common across applications within a given domain, or across domains. We have shown that adding architectural support for executing these common kernels of computation can yield an overall better system performance. We present a light weight, simplified prototype of a sensor processing unit (SPU) that offloads these computations from the main arithmetic logic unit (ALU) of an embedded processor, and that accesses sensor data in a low latency manner. Our SPU prototype shows an average speed up factor of 2.48 over executing these kernels on an embedded PowerPC processor. A large portion of this speed up is due to our low latency method for accessing sensor data. Isolating our speed up to purely computation still shows an average speed up factor of 1.38 for these kernels. Adwait Gupte, Phillip H. Jones |
RTCSA | 2 |
| 2007 | Changing Output Quality for Thermal ManagementabstractA growing number of embedded computing systems are used outside of environmentally controlled locations. In locations such as remote parts of deserts, deep ocean floors, and outer space, it is not only difficult to predict environmental effects on a system, they also allow very limited accessibility once a system is deployed. Therefore, it is often necessary for system parameters to be over provisioned to guarantee correct functionality under worst case environmental conditions. This often leads to an end system that is suboptimal for typical conditions. Phillip H. Jones, James Moscola, Young H. Cho, John W. Lockwood |
FCCM | 1 |
| 2007 | Adaptive Thermoregulation for Applications on Reconfigurable DevicesabstractA biological organism's ability to sense and adapt to its environment is essential to its survival. Likewise, environmentally aware computing systems avail themselves to a longer operational life and a wider range of applications than traditional systems. In this paper, we propose a novel circuit design methodology that allows parameterizable hardware to self-regulate its temperature. We apply this methodology to an image recognition system on an Xilinx Virtex 4 FX100 field programmable gate array (FPGA). The image recognition system sustains a safe operational temperature by automatically adjusting its frequency and output quality. The circuit sacrifices output performance and quality to lower its internal temperature as the ambient temperature increases, and can leverage cooler temperatures by increasing output performance and quality. Furthermore, the circuit will shutdown if the ambient temperature becomes too hot for the device to function properly. A performance evaluation of our adaptive circuit under various thermal conditions shows up to a 4× factor increase in performance and a 2× factor increase in quality over a system without dynamic thermal control. Phillip H. Jones, James Moscola, Young H. Cho, John W. Lockwood |
FPL | 1 |
| 2006 | A Thermal Management and Profiling Method for Reconfigurable Hardware ApplicationsabstractGiven large circuit sizes, high clock frequencies, and possibly extreme operating environments, Field Programmable Gate Arrays (FPGAs) are capable of heating beyond their designed thermal limits. As new circuits are developed for FPGAs and deployed remotely, engineers are challenged to determine in advance if the device will operate within recommended thermal ranges. The amount of power consumed by the circuit depends on how an algorithm is compiled into hardware, how the circuit is placed and routed, and the patterns of data that pass through the system. The amount of heat that can be dissipated depends on the thermal transfer characteristics of the package, the air flow that passes over the package, and the ambient temperature of the remote systems. Rather than designing a system to handle unreasonable worst-case situations, we have implemented a thermal management system that continuously monitors the temperature of the FPGA and reprograms the device if the temperate approaches the outer limits of safe operating conditions. Our system measures the junction temperature of a Xilinx Virtex FPGA using a built-in thermal diode. Using the temperature monitoring mechanism, we have studied the steady-state and transient conditions of multiple benchmark circuits implemented in an FPGA logic on the Field-programmable Port Extender (FPX) development platform. We observed properties of these benchmark circuits that enable us to predict power and thermal characteristics for real applications. We propose a Dynamic Thermal Management (DTM) strategy for FPGAs based on temperature feedback. Phillip H. Jones, John W. Lockwood, Young H. Cho |
FPL | 1 |
| 2006 | An adaptive frequency control method using thermal feedback for reconfigurable hardware applicationsabstractReconfigurable circuits running in field programmable gate arrays (FPGAs) can be dynamically optimized for power based on computational requirements and thermal conditions of the environment. In the past, FPGA circuits were typically small and operated at a low frequency. Few users were concerned about high-power consumption and the heat generated by FPGA devices. The current generation of FPGAs, however, use extensive pipelining techniques to achieve high data processing rates and dense layouts that can generate significant amounts of heat. FPGA circuits can be synthesized that can generate more heat than the package can dissipate. For FPGAs that operate in controlled environments, heatsinks and fans can be mounted to the device to extract heat from the device. When FPGA devices do not operate in a controlled environment, however, changes to ambient temperature due to factors such as the failure of a fan or a reconfiguration of bitfile running on the device can drastically change the operating conditions. A protection mechanism is needed to ensure the proper operation of the FPGA circuits when such a change occurs. To address these issues, we have devised a reconfigurable temperature monitoring system that gives feedback to the FPGA circuit using the measured junction temperature of the device. Using this feedback, we designed a novel dual frequency switching system that allows the FPGA circuits to maintain the highest level of performance for a given maximum junction temperature. Our working system has been implemented and deployed on the field programmable port extender (FPX) platform at Washington University in St. Louis. Our experimental results with a scalable image correlation circuit show up to a 2.4times factor increase in performance as compared to a system without thermal feedback. Our circuit ensures that the device performs the maximum required computation while always operating within a safe temperature range Phillip H. Jones, Young H. Cho, John W. Lockwood |
FPT | 1 |
| 2004 | The Effects of an ARMOR-Based SIFT Environment on the Performance and Dependability of User ApplicationsabstractFew, distributed software-implemented fault tolerance (SIFT) environments have been experimentally evaluated using substantial applications to show that they protect both themselves and the applications from errors. We present an experimental evaluation of a SIFT environment used to oversee spaceborne applications as part of the Remote Exploration and Experimentation (REE) program at the Jet Propulsion Laboratory. The SIFT environment is built around a set of self-checking ARMOR processes running on different machines that provide error detection and recovery services to themselves and to the REE applications. An evaluation methodology is presented in which over 28,000 errors were injected into both the SIFT processes and two representative REE applications. The experiments were split into three groups of error injections, with each group successively stressing the SIFT error detection and recovery more than the previous group. The results show that the SIFT environment added negligible overhead to the application's execution time during failure-free runs. Correlated failures affecting a SIFT process and application process are possible, but the division of detection and recovery responsibilities in the SIFT environment allows it to recover from these multiple failure scenarios. Only 28 cases were observed in which either the application failed to start or the SIFT environment failed to recognize that the application had completed. Further investigations showed that assertions within the SIFT processes-coupled with object-based incremental checkpointing-were effective in preventing system failures by protecting dynamic data within the SIFT processes. Keith Whisnant, Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Phillip H. Jones, David A. Rennels, Raphael R. Some |
IEEE Trans. Software Eng. | 4 |
| 2002 | NFTAPE: Networked Fault Tolerance and Performance EvaluatorabstractThe NFTAPE is a software implemented, highly flexible fault injection environment for conducting automated fault/error injection-based dependability characterization. NFTAPE: (1) enables a user: (i) to specify a fault/error injection plan, (ii) to carry out injection experiments, and (iii) to collect the experimental results for analysis; (2) targets assessment of a broad set of dependability metrics, e.g., availability, reliability, coverage; (3) operates in a distributed environment; (4) can be configured to implement a variety of fault/error injection strategies and thus to serve multiple users and target systems; (5) imposes minimal disturbance of target systems. David T. Stott, Phillip H. Jones, M. Hamman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 2 |