EDBT 2026 Demo / reviewers in the wild / expert
Kenneth B. Kent
dblp:k/KennethBKent · also Kenneth Blair Kent
· DBLP profile ↗
86ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0003-2764-823XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 3 first-author · 13 since 2021Software engineering, systems software and programming languages · 23 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ExpertoRhythm: Morphology-Aware Learning for Waveform Reconstruction and Cuffless Blood Pressure Estimation from Single-Channel PPG
Amir Arjomand, Kenneth B. Kent, Georgiy Krylov |
AIME (1) | 2 |
| 2026 | Reconfigurable acceleration for database systems: Taxonomy, techniques, and research challengesabstractDatabase query processing and optimization are critical components of modern database management systems (DBMS) that efficiently process user queries. In big data application scenarios, the movement of large volumes of data influences performance, power efficiency, and reliability, which are the three essential aspects of a computing system. Large-scale data centers require an exceptionally efficient server and storage infrastructure. The systems currently employed for managing and processing big data are increasingly showing inefficiency, both in terms of energy usage and scalability, primarily due to the constraints imposed by existing CPU architectures. A significant challenge in Database Management Systems (DBMS) is the growing disparity between the speeds of processors and memory access, which results in notable performance bottlenecks. This paper presents a comprehensive survey of reconfigurable acceleration in database systems, offering a structured taxonomy that categorizes existing work based on query types, integration models, and hardware/software co-design strategies. We examine key acceleration techniques across relational operators, indexing, join algorithms, and compression, highlighting their trade-offs in performance, scalability, and adaptability. Furthermore, we identify current limitations in programmability, data movement, and workload variability, and outline open research challenges including dynamic reconfiguration, hybrid architectures, and compiler support. This taxonomy-driven perspective aims to guide both researchers and practitioners in navigating the design space and pushing the boundaries of FPGA-accelerated data processing. Geetesh More, Suprio Ray, Kenneth B. Kent |
J. Syst. Archit. | 3 |
| 2025 | DGSim: A Scalable Framework For Simulating Energy Consumption Of Household AppliancesabstractTraditional household energy simulation tools often struggle to efficiently scale with increasing numbers of appliances and households, leading to high computational overhead and synchronization delays. In this work, we introduce DGSim, a scalable framework for simulating household energy consumption that supports extensive simulations with millions of instances while allowing detailed customization of appliances, usage patterns, and community demographics. DGSim incorporates stochastic techniques, parallel processing, chunking, and dynamic core allocation to improve execution efficiency. Our evaluation demonstrates that DGSim efficiently handles large workloads, utilizing configurable simulations to adapt to diverse household structures and energy consumption behaviors. By integrating flexible configurations into large-scale experiments, DGSim enables detailed analysis for demand-side management and residential energy modeling. Bhavani Sai Prasad Addala, Mohammad Mehabadi Mohammadi, Kenneth B. Kent |
ECMS | 3 |
| 2025 | Comparing Client- & Server-Side AEAD Encryption in Software-Defined Storage SystemsabstractTo provide data confidentiality and establish data integrity with minimal performance overhead, “Authenticated Encryption with associated data” (AEAD) ciphers have become mandatory for implementing transport security in TLS 1.3. However, these ciphers have yet to find wide-range adoption within the domain of security at rest. This is problematic since software-defined storage (SDS) systems provide even smallscale organizations with a cost-effective method of storing large amounts of data. At present, AEAD ciphers at rest are most commonly used to encrypt objects in the object stores of AWS, Azure, and Google Cloud. However, these ciphers are not applied to client-side encryption for other storage formats, such as block or file storage, resulting in asymmetric security guarantees across storage services. On the server-side, AEAD encryption has, to the best of our knowledge, also not been widely adopted. Since neither AEAD encryption on the client- nor server-side has seen a wide-range adoption, this paper compares the benefits and downsides of employing AEAD encryption on the client- or server-side within SDS systems. To establish this comparison, we implemented a client-side and server-side encryption approach in the widely adopted Ceph SDS system. Our study demonstrates that incorporating AEAD ciphers can be achieved with a write performance loss of less than $5 \%$, while ensuring a higher security level than traditional encryption-at-rest approaches. Most importantly, we managed to achieve these security gains for data stored in all available storage formats in Ceph. David Mohren, Minh Tien Truong, Brett Kelly, Kenneth B. Kent |
PST | 4 |
| 2025 | Lightweight 1D UNet-CPCA Regression Model for Energy-Efficient Blood Pressure Estimation from Raw PPG Signals
Amir Arjomand, Kenneth B. Kent, Georgiy Krylov |
RSP | 2 |
| 2025 | Early Detection of Unsupported Compilations during Prototyping in a Template-based Just-in-Time Compiler
Michael C. Goodyear, Scott Ryan Young, Marius Pirvu, Harpreet Kaur 0003, Kenneth B. Kent |
RSP | 5 |
| 2025 | FALCON: FPGA Accelerated Lightweight Updatable Learned Index
Geetesh More, Suprio Ray, Kenneth B. Kent |
SSDBM | 3 |
| 2025 | Efficient security interface for high-performance Ceph storage systemsabstractCeph portrays a resilient clustered storage solution with supporting object, block, and file storage capabilities with no single point of failure. Despite these qualifications, data confidentiality defines a concern in the system, as authentication and access control are the only data protection security services in Ceph. CephArmor was proposed as a third-party security interface to protect data confidentiality by adding an extra protection layer to data at rest. Despite the added layer, the initial design of the API needed to be more efficient in addressing security and performance simultaneously. In this study, we propose a new architectural design to address the associated issues with the preliminary prototype. Comprehensive performance and security analysis verify the improvement of the proposed method compared to the initial approach. The benchmark result has indicated a 37% improvement on average in IOPS, elapsed time, and bandwidth for the write benchmark compared to the initial model. Fatemeh Khoda Parast, Seyed Alireza Damghani, Brett Kelly, Yang Wang 0006, Kenneth B. Kent |
Future Gener. Comput. Syst. | 5 |
| 2025 | Understanding Serverless Inference in Mobile-Edge Networks: A Benchmark ApproachabstractAlthough the emerging serverless paradigm has the potential to become a dominant way of deploying cloud-service tasks across millions of mobile and IoT devices, the overhead characteristics of executing these tasks on such a volume of mobile devices remain largely unclear. To address this issue, this paper conducts a deep analysis based on the OpenFaaS platform—a popular open-source serverless platform for mobile edge environments—to investigate the overhead of performing deep learning inference tasks on mobile devices. To thoroughly evaluate the inference overhead, we develop a performance benchmark, namedESBench, whereby a set of comprehensive experiments are conducted with respect to a bunch of simulated mobile devices associated with an edge cluster. Our investigation reveals that the performance of deep learning inference tasks is significantly influenced by the model size and resource contention in mobile devices, leading to up to$3\times$degradation in performance. Moreover, we observe that the network environment can negatively impact the performance of mobile inference, increasing the CPU overhead under poor network conditions. Based on our findings, we further propose some recommendations for designing efficient serverless platforms and resource management strategies as well as for deploying serverless computing in the mobile edge environment. Yanying Lin, Shuaipeng Wu, Kenneth B. Kent, Kejiang Ye, Yang Wang 0006 |
IEEE Trans. Cloud Comput. | 5 |
| 2025 | VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture ExplorationabstractThis work details the capabilities of a major new release of the Verilog-to-Routing (VTR) open source FPGA CAD tool flow. Enhancements include generalizations of VTR’s architecture modeling language and optimizers to enable a more diverse set of programmable routing fabrics, FPGAs with embedded hard Networks-on-Chip (NoCs) and three-dimensional 3D FPGA systems that leverage stacked silicon integration. The new Parmys logic synthesis flow improves language coverage and result quality, and the physical implementation flow includes a more efficient placement engine, floorplanning constraints to guide placement, the ability to perform single-stage (flat) routing to improve quality, and parallel routing algorithms to reduce CPU time. This release also includes new architecture captures of recent commercial devices (Xilinx’s 7-series and Altera’s Stratix 10) and new benchmark suites (Titanium25 and Hermes) to aid FPGA architecture investigation. Verilog language coverage is greatly improved with the new Parmys logic synthesis flow, enabling more designs to be used with VTR. Finally, the placement and routing engines have beeenbeen sped up by 4 \(\times\) and 2.2 \(\times\) vs. VTR 8, respectively, leading to an overall physical implementation flow CPU time reduction of 48% with better result quality on average compared to VTR 8. Mohamed A. Elgammal, Amin Mohaghegh, Soheil Gholami Shahrouz, Fatemehsadat Mahmoudi, Fahrican Kosar, Kimia Talaei, Joshua Fife, Daniel Khadivi, Kevin E. Murray, Andrew Boutros, Kenneth B. Kent, Jeffrey B. Goeders, Vaughn Betz |
ACM Trans. Reconfigurable Technol. Syst. | 11 |
| 2025 | Corrigendum: VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture ExplorationabstractThis is a corrigendum for the article “VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture Exploration” published in ACM Trans. Reconfig. Technol. Syst. 18, 3, Article 39 (August 2025), 53 pages. Mohamed A. Elgammal, Amin Mohaghegh, Soheil Gholami Shahrouz, Fatemehsadat Mahmoudi, Fahrican Kosar, Kimia Talaei, Joshua Fife, Daniel Khadivi, Kevin E. Murray, Andrew Boutros, Kenneth B. Kent, Jeffrey B. Goeders, Vaughn Betz |
ACM Trans. Reconfigurable Technol. Syst. | 11 |
| 2024 | Learned Index Acceleration with FPGAs: A SMART ApproachabstractIndexes in database systems such as B+trees and hash tables retrieve data quickly. Much research has been conducted on the faster index lookup in recent years. A learned index is one such area of study. Learned index approaches, such as radix spline (RS) [1] can achieve significant performance improvement over traditional indexing techniques. However, query performance with learned indexes is limited by the constraints imposed by CPU architecture. This paper introduces a novel methodology that leverages the benefits of learned indexes and FPGAs. We term this approach as the Selective Mathematical operation AcceleRaTion (SMART) with an FPGA for an end-to-end acceleration of learned indexes. As a hybrid of CPU and FPGA approaches, the SMART model of index acceleration surpasses the throughput of CPU-based implementations while preserving the data structure storage on the CPU. Our proposed FPGA-based RS learned index architecture (Figure 1) consists of two major stages: Build and Lookup. The Build model$(SMART-RS_{Build})$accelerates the build stage on an FPGA, while the Lookup model$(SMART-RS_{Lookup})$executes the FPGA-based lookup acceleration. The build stage is accelerated by offloading the computationally intensive interpolation operation onto an FPGA. In this stage, the index data is stored on a CPU while the interpolation operation is executed on an FPGA. As shown in Figure 1, input to the FPGA-based$SMART-RS_{Build}$are key, max spline error, and previously stored CDF point. Here, the$X$and$Y$coordinates are the bounding box of a spline. The FPGA-based build-stage accelerator will output the orientation type: clockwise (CW), counter-clockwise (CCW), or collinear. Based on the orientation obtained, upper and lower limits are set and the previous CDF point is stored as the next spline point. RadixSpline is constructed over a sorted data set. The module AddKeyToSpline iterates over the input sorted data set and will create the array of resultant splines. The decision-making in deciding the orientation type is offloaded to the module SMART-@$RS_{Build-Interp}$on the FPGA. The output from the FPGA is the type of orientation that will make the following decision during spline construction: Store the input key and its position in the dataset in the spline points array (SplinePoints). • Store the spline index in the radix table (Radix TABLE). • Update the upper and lower error bounds. In the lookup stage acceleration, the radix table and set of spline points are offloaded to our lookup-stage accelerator (SMART@$RS_{Lookup}$), where they are stored in BRAMs/SRLs as partitioned arrays. With our approach, Speedup of 5.5× as compared to a CPU -based RS index. • Specific compute-intensive operations were identified, thereby avoiding the need for a full-scale FPGA implementation. • Complexities associated with debugging RTL-related issues were reduced. • Controlled on-chip and off-chip memory resource usage. • More accurate comparison between CPU and FPGA implementations. Geetesh More, Suprio Ray, Kenneth B. Kent |
FCCM | 3 |
| 2024 | Enhancing the VTR Flow: Integration of ABC9 via Yosys for Better Technology Mapping and OptimizationabstractThis study optimizes the Verilog-to-Routing (VTR) flow, an open-source Computer-Aided Design (CAD) tool. It utilizes ODIN II and Parmys for synthesis, ABC for technology mapping, and Versatile Place and Route for packing, placement, and routing. The ABC9 optimizations, integrated as a pass within the Yosys Open Synthesis Suite, enhance technology mapping and optimization stages and outperform the traditional ABC tool for large, complex designs. These optimizations improve timing behavior in multi-clock designs and include a delay model for Field Programmable Gate Array (FPGA) hard blocks. Various benchmarks assess the effectiveness of the workflow across different design complexities and FPGA architectures, including the utilization of hard blocks. Navid Jafarof, Kenneth B. Kent |
RSP | 2 |
| 2024 | Open Set Dandelion Network for IoT Intrusion DetectionabstractAs Internet of Things devices become widely used in the real-world, it is crucial to protect them from malicious intrusions. However, the data scarcity of IoT limits the applicability of traditional intrusion detection methods, which are highly data-dependent. To address this, in this article, we propose the Open-Set Dandelion Network (OSDN) based on unsupervised heterogeneous domain adaptation in an open-set manner. The OSDN model performs intrusion knowledge transfer from the knowledge-rich source network intrusion domain to facilitate more accurate intrusion detection for the data-scarce target IoT intrusion domain. Under the open-set setting, it can also detect newly-emerged target domain intrusions that are not observed in the source domain. To achieve this, the OSDN model forms the source domain into a dandelion-like feature space in which each intrusion category is compactly grouped and different intrusion categories are separated, i.e., simultaneously emphasising inter-category separability and intra-category compactness. The dandelion-based target membership mechanism then forms the target dandelion. Then, the dandelion angular separation mechanism achieves better inter-category separability, and the dandelion embedding alignment mechanism further aligns both dandelions in a finer manner. To promote intra-category compactness, the discriminating sampled dandelion mechanism is used. Assisted by the intrusion classifier trained using both known and generated unknown intrusion knowledge, a semantic dandelion correction mechanism emphasises easily-confused categories and guides better inter-category separability. Holistically, these mechanisms form the OSDN model that effectively performs intrusion knowledge transfer to benefit IoT intrusion detection. Comprehensive experiments on several intrusion datasets verify the effectiveness of the OSDN model, outperforming three state-of-the-art baseline methods by 16.9%. The contribution of each OSDN constituting component, the stability and the efficiency of the OSDN model are also verified. Jiashu Wu, Kenneth B. Kent, Jerome Yen, Cheng-Zhong Xu 0001, Yang Wang 0006 |
ACM Trans. Internet Techn. | 3 |
| 2023 | Extending Memory Compatibility with Yosys Front-End in VTR FlowabstractVerilog-to-routing (VTR) is an open source Computer Aided Design (CAD) framework that is widely used for research purposes. VTR provides researchers with a comprehensive set of benchmarks and FPGA architectures to test, compare and verify their designs and CAD algorithms. VTR employs Odin, Yosys and a combination of Yosys+Odin as its front-ends for elaborating digital designs written in Verilog. Yosys is a standalone synthesis framework that is maintained separately. In order to take advantage of Yosys' latest upgrades, it is essential to upgrade VTR in tandem with Yosys to keep up with its latest changes. In our study, we investigated the integration of the latest available version of Yosys into VTR. In the course of this research, we encountered a challenge stemming from a memory incompatibility in newer versions of Yosys. Our primary research question became how to effectively facilitate netlist conversion from Yosys to Odin to overcome this incompatibility issue. We then showcase how this upgrade affects our benchmark suite, highlighting the notable changes in circuit quality and the tool performance. Our post-upgrade evaluations reveal that the STA (static timing analysis) time of placement and the placement time itself exhibited improvements of up to 11.53% and 10.43%, respectively. Alireza Azadi, Amir Arjomand, Kenneth B. Kent |
RSP | 3 |
| 2023 | The Impact of Heterogeneous Logic on Adders and Multipliers in VTRabstractThis paper presents an extension of the Verilog-to-Routing (VTR) Computer-Aided Design (CAD) tool, focusing specifically on the utilization of heterogeneous logic for both multipliers and adders. We build upon the heterogeneous logic implementation of adders in VTR by addressing bugs and refining the implementation. This research aims to uncover the optimal configuration in terms of critical path delay (CPD) and device area for hard and soft logic implementations within different architectures and designs. Our findings show that a balanced approach, instead of the exclusive utilization of either hard or soft logic, can yield superior results. Extensive testing reveals significant improvements in speed and potential reductions in device size. The results validate the implementation's success and suggest future directions for enhancing the use of different types of logic in VTR. Navid Jafarof, Kenneth B. Kent |
RSP | 2 |
| 2023 | Adaptive Bi-Recommendation and Self-Improving Network for Heterogeneous Domain Adaptation-Assisted IoT Intrusion DetectionabstractAs Internet of Things (IoT) devices become prevalent, using intrusion detection to protect IoT from malicious intrusions is of vital importance. However, the data scarcity of IoT hinders the effectiveness of traditional intrusion detection methods. To tackle this issue, in this article, we propose the adaptive bi-recommendation and self-improving network (ABRSI) based on unsupervised heterogeneous domain adaptation (HDA). The ABRSI transfers enrich intrusion knowledge from a data-rich network intrusion source domain to facilitate effective intrusion detection for data-scarce IoT target domains. The ABRSI achieves fine-grained intrusion knowledge transfer via adaptive bi-recommendation matching. Matching the bi-recommendation interests of two recommender systems (RSs) and the alignment of intrusion categories in the shared feature space form a mutual-benefit loop. Besides, the ABRSI uses a self-improving mechanism, autonomously improving the intrusion knowledge transfer from four ways. A hard pseudo label (PL) voting mechanism jointly considers RS decision and label relationship information to promote more accurate hard PL assignment. To promote diversity and target data participation during intrusion knowledge transfer, target instances failing to be assigned with a hard PL will be assigned with a probabilistic soft PL, forming a hybrid pseudo-labeling strategy. Meanwhile, the ABRSI also makes soft pseudo-labels globally diverse and individually certain. Finally, an error knowledge learning mechanism is utilized to adversarially exploit factors that causes detection ambiguity and learns through both current and previous error knowledge, preventing error knowledge forgetfulness. Holistically, these mechanisms form the ABRSI model that boosts IoT intrusion detection accuracy via HDA-assisted intrusion knowledge transfer. Comprehensive experiments on several intrusion data sets demonstrate the state-of-the-art performance of the ABRSI method, outperforming its counterparts by 9.2%, and also verify the effectiveness of ABRSI constituting components and ABRSI’s overall efficiency. Jiashu Wu, Yang Wang 0006, Cheng-Zhong Xu 0001, Kenneth B. Kent |
IEEE Internet Things J. | 5 |
| 2023 | A comprehensive survey of cryptography key management systems
Subhabrata Rana, Fatemeh Khoda Parast, Brett Kelly, Yang Wang 0006, Kenneth B. Kent |
J. Inf. Secur. Appl. | 5 |
| 2023 | Koios 2.0: Open-Source Deep Learning Benchmarks for FPGA Architecture and CAD Researchabstractthe prevalence of deep learning (DL) in many applications, researchers are investigating different ways of optimizing field-programmable gate array (FPGA) architecture and CAD to achieve better quality-of-results (QoRs) on DL-based workloads. In this optimization process, benchmark circuits are an essential component; the QoR achieved on a set of benchmarks is the main driver for architecture and CAD design choices. However, current academic benchmark suites are inadequate, as they do not capture any designs from the DL domain. This work presents the second version of our suite of DL acceleration benchmark circuits for FPGA architecture and CAD research, called Koios. This suite of 40 circuits covers a wide variety of accelerated neural networks, design sizes, implementation styles, abstraction levels, and numerical precisions. These benchmarks include 32 DL designs and eight synthetic (proxy) benchmarks. The Koios benchmarks are larger, more data parallel, more heterogeneous, more deeply pipelined, and utilize more FPGA architectural features compared to existing open-source benchmarks. This enables researchers to pinpoint architectural inefficiencies for this class of workloads and optimize CAD tools on more representative benchmarks that stress the CAD algorithms in different ways. In this article, we describe the Koios designs, compare their characteristics to prior FPGA benchmark suites, and present results of running them through the verilog-to-routing (VTR) flow using a recent FPGA architecture model. Finally, we present case studies showing how exploration of DL-optimized FPGA architecture and CAD algorithms can be performed using our new benchmark suite. Aman Arora 0001, Andrew Boutros, Seyed Alireza Damghani, Karan Mathur, Vedant Mohanty, Tanmay Anand, Mohamed A. Elgammal, Kenneth B. Kent, Vaughn Betz, Lizy Kurian John |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2022 | SIMD support to improve eclipse OpenJ9 performance on the AArch64 platformabstractJust-in-time (JIT) compilers achieve application portability and improved management of large code-bases by abstracting the architecture specific details from programmers. Eclipse OMR and Eclipse OpenJ9 invest extensively in JIT technology to efficiently execute architecture-neutral Java bytecode. OMR is a robust language runtime builder, and OpenJ9 is a managed language runtime that consumes OMR. The targeted domains of OMR and OpenJ9 include AArch64, a 64-bit version of the ARM architecture. AArch64 is a popular member of the embedded computing market, where computing infrastructure resources (e.g., CPU, memory) are constrained. The concept of SIMD (Single Instruction, Multiple Data) instructions primarily evolved to accelerate the performance of multimedia applications such as motion video, real-time physics and graphics where repetitive operations were involved on large arrays of numbers. This paper discusses the steps taken to add SIMD support to OMR for AArch64. The implementation of advanced SIMD and floating-point instructions are also discussed, which cover vectorized mathematical operations, including addition, subtraction, multiplication, and division for supported data-types. We validate our implementation through relevant OMR tril tests, present two microbenchmarks VectorizationMicrobenchmark and Sepia Tone Filter and a set of standard benchmarks, which leverage the OpenJ9 autovectorization process in AArch64. The AArch64 vectorized operations are evaluated against non-vectorized, but similar operations using Eclipse OpenJ9. We demonstrate an improvement of up to four times in execution speed of certain vector arithmetic operations. Md Alvee Noor, Kenneth B. Kent, Kazuhiro Konno, Daryl Maier |
CF | 2 |
| 2022 | Odin-II Partial Technology Mapping for Yosys Coarse-grained Netlists in VTRabstractThe Verilog-to-routing (VTR) front-end interface for Verilog compilation, Odin-II, lacks complete support for the Verilog-2005 standard. However, Odin-II provides complex partial mapping for balancing soft logic and hard blocks. Yosys, an open framework for RTL synthesis, provides extensive support for HDLs. However, the Yosys flow forces the user to decide the discrete circuit implementation manually. This research proposes improving device utilization and simplifying the flow by automating complex logic decisions with architecture awareness. According to VTR architectures, hard/soft logic trade-off decisions and heterogeneous logic inference have become available for such coarse-grained BLIF files. Yosys+Odin-II demonstrates promising results by lowering resource consumption, shrinking final circuit footprint, and reducing overall routed wire length, while other criteria remain approximately the same. Seyed Alireza Damghani, Kenneth B. Kent |
FCCM | 2 |
| 2022 | Yosys+Odin-II: The Odin-II Partial Mapper with Yosys Coarse-grained Netlists in VTRabstractVerilog-to-routing (VTR) provides users with an entire flow from the Verilog circuit description to a final FPGA programming configuration. The VTR front-end interface for Verilog compilation, Odin-II, lacks support for the Verilog-2005 standard. However, Odin-II provides complex partial mapping for balancing soft logic and hard blocks. Yosys, an open framework for RTL synthesis, provides extensive support for HDLs. However, the Yosys flow forces the user to decide the discrete circuit implementation manually. The approach taken by Yosys is to map all discrete components into available hard blocks or to explode them in low-level logic when not available. This research proposes improving device utilization and simplifying the flow by automating complex logic decisions with architecture awareness. Yosys, as a front-end for generating coarse-grained netlists, and Odin-II, as BLIF elaborator and partial mapper, are integrated into the VTR flow. According to VTR architectures, hard/soft logic trade-off decisions and heterogeneous logic inference have become available for such coarse-grained BLIF files. Yosys+Odin-II demonstrates promising results by lowering resource consumption, shrinking final circuit footprint, and reducing overall routed wire length, while other criteria remain approximately the same. The overall VTR flow runtime is imperceptibly reduced with Yosys+Odin-II, while the capability of the VTR flow to synthesize more complex designs is improved, and the control over an intelligent partial mapper is provided for users. Seyed Alireza Damghani, Kenneth B. Kent |
FPGA | 2 |
| 2022 | Machine Learning-Based Hard/Soft Logic Trade-offs in VTRabstractCircuit optimization, in any application, is of high importance since it not only improves the efficiency of the intended purpose but also enhances the quality of the final product. It enables the circuit designer to cater to the specific needs of the customer. For circuit optimization to occur, we need to elaborate these circuits on a primary level and perform synthesis operations. Previous research shows that the investigation of improvements to different Hardware Description Language (HDL) elaboration phases, was completely closed source. Verilog To Routing (VTR) is an open-source Electronic Design Automation (EDA) tool. ODIN II is the VTR synthesizer that parses the input Verilog, elaborates its Abstract Syntax Tree (AST), performs the partial mapping according to the architecture file, and performs optimizations such as unused logic removal. To that end, the hard versus soft logic trade-off aims to optimize the performance of the circuit. This project focuses on using machine learning approaches to make synthesis tools intelligent enough to decide this ratio on their own, without the need for human intervention, and based on some predefined criteria. This paper discusses the criteria for having less latency or less critical path delay in the circuit. Also, it aims at providing this level of intelligence at an earlier stage in the VTR pipeline to make better use of this information. Ritwik Sinha, Seyed Alireza Damghani, Kenneth B. Kent |
RSP | 3 |
| 2022 | Cloud computing security: A survey of service-based models
Fatemeh Khoda Parast, Chandni Sindhav, Seema Nikam, Hadiseh Izadi Yekta, Kenneth B. Kent, Saqib Hakak |
Comput. Secur. | 5 |
| 2022 | Benchmarking and learning garbage collection delays for resource-restricted graphical user interfacesabstractAbstract Tablets, smartphones, and wearables have limited resources. Applications on these devices employ a graphical user interface (GUI) for interaction with users. Language runtimes for GUIs employ dynamic memory management using garbage collection (GC). However, GC policies and algorithms are designed for data centers and cloud computing, but they are not necessarily ideal for resource‐constrained embedded devices. In this article, we present GUI GC, a JavaFX GUI benchmark, which we use to compare the performance of the four GC policies of the Eclipse OpenJ9 Java runtime on a resource‐constrained environment. Overall, our experiments suggest that the default policy Gencon registered significantly lower execution times than its counterparts. The region‐based policy, Balanced, did not fully utilize blocking times; thus, using GUI GC, we conducted experiments with explicit GC invocations that measured significant improvements of up to 13.22% when multiple CPUs were available. Furthermore, we created a second version of GUI GC that expands on the number of controllable load‐stressing dimensions; we conducted a large number of randomly configured experiments to quantify the performance effect that each knob has. Finally, we analyzed our dataset to derive suitable knob configurations for desired runtime, GC, and hardware stress levels. Harry McCarthy, Abigail M. Y. Koay, Michael Dawson 0001, Kenneth B. Kent, Panos Patros |
Softw. Pract. Exp. | 4 |
| 2022 | Deadlock Avoidance Algorithms for Recursion-Tree Modeled Requests in Parallel ExecutionsabstractWe present an extension of the bankers algorithm to resolve deadlock for programs whose resource-request graph can be modeled as a recursion tree for parallel execution. Our algorithm implements the bankers logic, with the key difference being that some properties of the tree are fully exploited to improve the resource utilization and safety check in deadlock avoidance. For an n-node tree modeled program making requests to m types of resources, our recursion-tree based algorithm can obtain a time complexity of O(mn loglogn) on average in safety check while reducing the conservativeness in resource utilization. We reap these benefits by proposing a concept of the resource critical tree and leverage it to localize the maximum claim associated with each node in the tree. To tackle the case when the tree model is not statically known, we relax the definition of a local maximum claim by sacrificing some resource utilization. With this trade-off, the algorithm can resolve the deadlock and achieve more efficient safety checks within time of O(m loglogn). Our empirical studies on a two-dimensional integration problem on sparse grids show that the proposed algorithms can reduce resource utilization conservativeness and improve avoidance performance by minimizing the number of safety checks. Yang Wang 0006, Kenneth B. Kent, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 4 |
| 2022 | The State of the Art of Metadata Managements in Large-Scale Distributed File Systems - Scalability, Performance and AvailabilityabstractFile system metadata is the data in charge of maintaining namespace, permission semantics and location of file data blocks. Operations on the metadata can account for up to 80% of total file system operations. As such, the performance of metadata services significantly impacts the overall performance of file systems. A large-scale distributed file system (DFS) is a storage system that is composed of multiple storage devices spreading across different sites to accommodate data files, and in most cases, to provide users with location independent access interfaces. Large-scale DFSs have been widely deployed as a substrate to a plethora of computing systems, and thus their metadata management efficiency is crucial to a massive number of applications, especially with the advent of the Big Data age, which poses tremendous pressure on underlying storage systems. This paper reports the state-of-the-art research on metadata services in large-scale distributed file systems, which is conducted from three indicative perspectives that are always used to characterize DFSs: high-scalability, high-performance, and high-availability, with special focus on their respective major challenges as well as their developed mainstream technologies. Additionally, the paper also identifies and analyzes several existing problems in the research, which could be used as a reference for related studies. Yang Wang 0006, Kenneth B. Kent, Lingfang Zeng, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Koios: A Deep Learning Benchmark Suite for FPGA Architecture and CAD ResearchabstractWith the prevalence of deep learning (DL) in many applications, researchers are investigating different ways of optimizing FPGA architecture and CAD to achieve better quality-of-results (QoR) on DL-based workloads. In this optimization process, benchmark circuits are an essential component; the QoR achieved on a set of benchmarks is the main driver for architecture and CAD design choices. However, current academic benchmark suites are inadequate, as they do not capture any designs from the DL domain. This work presents a new suite of DL acceleration benchmark circuits for FPGA architecture and CAD research, called Koios. This suite of 19 circuits covers a wide variety of accelerated neural networks, design sizes, implementation styles, abstraction levels, and numerical precisions. These designs are larger, more data parallel, more heterogeneous, more deeply pipelined, and utilize more FPGA architectural features compared to existing open-source benchmarks. This enables researchers to pin-point architectural inefficiencies for this class of workloads and optimize CAD tools on more realistic benchmarks that stress the CAD algorithms in different ways. In this paper, we describe the designs in our benchmark suite, present results of running them through the Verilog-to-Routing (VTR) flow using a recent FPGA architecture model, and identify key insights from the resulting metrics. On average, our benchmarks have 3.7× more netlist primitives, 1.8× and 4.7× higher DSP and BRAM densities, and 1.7× higher frequency with 1.9× more near-critical paths compared to the widely-used VTR suite. Finally, we present two example case studies showing how architectural exploration for DL-optimized FPGAs can be performed using our new benchmark suite. Aman Arora 0001, Andrew Boutros, Daniel Rauch, Aishwarya Rajen, Aatman Borda, Seyed Alireza Damghani, Samidh Mehta, Sangram Kate, Pragnesh Patel, Kenneth B. Kent, Vaughn Betz, Lizy Kurian John |
FPL | 10 |
| 2021 | Energy and Runtime Performance Optimization of Node.js Web RequestsabstractThe Node.js framework uses an event-driven model with a single-threaded event loop and provides asynchronous and non-blocking I/O operations. As with other programs, Node.js web applications take advantage of underlying resources, including CPUs, which can incorporate the dynamic voltage and frequency scaling (DVFS) technique. Using CPU DVFS, the applications can increase their runtime performance, at the expense of the system's energy consumption. Thus, software code that utilizes the CPU DVFS technique efficiently should lead to “green” and high-performing applications with respect to the business logic. To this end, we build a CPU frequency scaling/energy aware system to enable CPU frequency control within Node.js applications and measure the energy consumption of specific tasks. We also build a benchmark suite to analyze the energy consumption and runtime performance of different requests based on the CPU frequency impact and collect information and patterns, as we scale the CPU frequency. The analysis aims to provide data and knowledge on the CPU frequency “suitability” and impact in order to create a model for CPU frequency scaling on Node.js web applications and achieve an efficient and sustainable runtime performance. Maria Patrou, Kenneth B. Kent, Joran Siu, Michael Dawson 0001 |
IC2E | 2 |
| 2021 | Heterogeneous Logic Implementation for Adders in VTRabstractVerilog-to-Routing (VTR) is a Field-Programmable Gate Array (FPGA) Computer-Aided Design (CAD) tool. It is composed of three tools, namely ODIN II, ABC and VPR with each performing distinctive optimizations at different stages of the design flow. The elaboration and hard block synthesis stage of VTR is the core responsibility of the sub-project ODIN II. This work enables ODIN II to use fewer hard adders in the circuit by allowing soft logic implementation alongside hard logic for circuits featuring addition operations. This is particularly useful in scenarios where a sufficient number of hard blocks are not available. The results of applying our modifications to ODIN II as well as the entire VTR flow have been analysed. The results reveal the potential of current adder optimizations to achieve up to 17% performance gains in terms of critical path delays. Another effect of the optimization is the implications on the resulting device size. Some future prospects in this respect are also outlined in this paper. Harpreet Kaur 0003, Georgiy Krylov, Seyed Alireza Damghani, Kenneth B. Kent |
RSP | 4 |
| 2020 | Hard and Soft Logic Trade-offs for Multipliers in VTRabstractThis paper discusses improvements to the Verilog- To-Routing (VTR) Computer Aided Design (CAD) tool, that enables synthesis of Verilog circuits to a Field Programmable Gate Array (FPGA) architecture, previously impossible due to device size limitations imposed by device growth. The proposed solution allows reducing device sizes required for well known circuits, through exploring the space/performance trade-off question at a finer granularity at early CAD stages. Results of as much as 2.63 times increase in performance and a 48% reduction in device size have been achieved for some circuits. Georgiy Krylov, Jean-Philippe Legault, Kenneth B. Kent |
DSD | 3 |
| 2020 | Desired Footprint by Technology Mapping Modification using a Genetic Algorithm in Odin IIabstractTechnology mapping is the transformation of a general Boolean logic network into a functional equivalent K-LUT network that can be implemented by the target FPGA device. Because an FPGA architecture is pre-determined, technology mapping is limited to the available resources. However, circuits can be optimized before the low-level synthesis phase. Odin II, part of the Verilog-to-routing project, is responsible for synthesis and elaboration. In the partial mapping phase of Odin II, some modifications are still possible for high-level modules-adder, multiplier-when there is no hard block available. When Odin II performs partial mapping to create soft logic, we can choose which implementation of a high-level module works best with respect to the desired goals: area versus speed. In this paper, we describe a method to modify circuit characteristics based on placement criteria. More specifically, after partial mapping circuit components during Verilog HDL code synthesis, there are still potential modifications in soft-logic circuit generation. We propose using a genetic algorithm during synthesis to adjust soft-logic circuit implementation in order to achieve the desired synthesis goal. We show that the approach provides promising results for a marginal cost in runtime. Seyed Alireza Damghani, Jean-Philippe Legault, Kenneth B. Kent |
RSP | 3 |
| 2020 | VTR 8: High-performance CAD and Customizable FPGA Architecture ModellingabstractDeveloping Field-programmable Gate Array (FPGA) architectures is challenging due to the competing requirements of various application domains and changing manufacturing process technology. This is compounded by the difficulty of fairly evaluating FPGA architectural choices, which requires sophisticated high-quality Computer Aided Design (CAD) tools to target each potential architecture. This article describes version 8.0 of the open source Verilog to Routing (VTR) project, which provides such a design flow. VTR 8 expands the scope of FPGA architectures that can be modelled, allowing VTR to target and model many details of both commercial and proposed FPGA architectures. The VTR design flow also serves as a baseline for evaluating new CAD algorithms. It is therefore important, for both CAD algorithm comparisons and the validity of architectural conclusions, that VTR produce high-quality circuit implementations. VTR 8 significantly improves optimization quality (reductions of 15% minimum routable channel width, 41% wirelength, and 12% critical path delay), run-time (5.3× faster) and memory footprint (3.3× lower). Finally, we demonstrate VTR is run-time and memory footprint efficient, while producing circuit implementations of reasonable quality compared to highly-tuned architecture-specific industrial tools—showing that architecture generality, good implementation quality, and run-time efficiency are not mutually exclusive goals. Kevin E. Murray, Oleg Petelin, Jia Min Wang, Mohamed Eldafrawy, Jean-Philippe Legault, Eugene Sha, Aaron Graham, Jean Wu, Matthew J. P. Walker, Hanqing Zeng, Panagiotis Patros, Jason Luu, Kenneth B. Kent, Vaughn Betz |
ACM Trans. Reconfigurable Technol. Syst. | 14 |
| 2020 | Optimizing FPGA Logic Block Architectures for ArithmeticabstractHardened adder and carry logic is widely used in commercial field-programmable gate arrays (FPGAs) to improve the efficiency of arithmetic functions. There are many design choices and complexities associated with such hardening, including circuit design, FPGA architectural choices, and the computer-aided design (CAD) flow. However, these choices have not been studied much and hence we explore a number of possibilities. We also highlight front-end elaboration optimization that helps ameliorate the restrictions placed on logic synthesis by hardened arithmetic. We show that hard adders and carry chains increase the performance of simple adders by a factor of 4 or more, but on larger benchmark designs that contain arithmetic improve the overall performance by 15%. Our results also show that for complete application circuits simple hardened ripple-carry adders perform as well as more complex carry-lookahead adders. Our best non-fracturable lookup table (non-fLUT) architecture with hardened arithmetic yields 12% better area-delay product than architectures without hardened arithmetic. We also investigate the impact of fLUTs and their interaction with hardened arithmetic. We find that fLUTs offer significant (12%-15%) area reduction, which is complementary to the delay reduction of hardened arithmetic. Therefore, our best fLUT architectures which use two bits of hardened arithmetic achieve 25% better area-delay product than non-fLUT architectures without hardened arithmetic. Kevin E. Murray, Jason Luu, Matthew J. P. Walker, Conor McCullough, Safeen Huda, Charles Chiasson, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2019 | Improving Digital Circuit Simulation with Batch-Parallel Logic EvaluationabstractIntegrated circuit simulators reproduce the behavior and functionality of the underlying circuits. They are part of FPGA CAD flow tools and they ensure the correctness of the circuits after the various conversions and optimizations occurring in the previous stages. During this procedure a graph with dependencies across nodes is created for each circuit design. Large circuits, and thus graphs, require more time to be simulated, making a parallel approach necessary. We explore a new solution-batch-parallel simulation in which the circuit output is calculated by worker threads that process batches of input vectors. The threads traverse and calculate their assigned nodes in parallel taking into consideration the intra-node dependencies. Furthermore, a node calculation analysis is performed and used to achieve work balance across threads. We apply this technique on the open-source Odin II framework and compare it with the existing approaches. The batch-parallel simulation is compared with the two existing approaches, single-threaded and multi-threaded, under various configurations, considering the number of threads and the batch sizes. The results demonstrate performance gains against the existing approaches in the majority of the benchmarks used for specific metrics, such as simulation elapsed time. Maria Patrou, Jean-Philippe Legault, Aaron Graham, Kenneth B. Kent |
DSD | 4 |
| 2019 | Toward Efficient Processing of Spatio-Temporal Workloads in a Distributed In-Memory SystemabstractLocation-based services (LBS) are a widely adopted technology that produces large volumes of spatio-temporal data at high velocity. Spatial data is also being generated from many other geo-spatial applications. To address the challenge of data volume, a number of big spatial data management systems have emerged that are based on the MapReduce paradigm. Recent projects have developed spatial data systems using Spark's distributed in-memory architecture. These projects, which include GeoSpark, SpatialSpark, and LocationSpark, do not support the high update rates required by LBS applications. Alternatively, systems such as MD-HBase support data updates, but are hindered by the performance characteristics of HBase, which is a disk-oriented framework. We present DISTIL+, a distributed spatio-temporal data processing system designed for high velocity location data. Our system achieves high update throughput and low query latency by leveraging the APGAS (Asynchronous Partitioned Global Address Space) architecture to build a multi-level distributed in-memory index. We present extensive experimental evaluation of our system, comparing several indexing and data placement schemes, as well as competing systems. Our results show that DISTIL+ excels at supporting high throughput location updates, and low latency spatio-temporal range queries and kNN queries, while offering better performance than existing approaches. Puya Memarzia, Maria Patrou, Suprio Ray, Virendrakumar C. Bhavsar, Kenneth B. Kent |
MDM | 6 |
| 2019 | Scaling Parallelism Under CPU - Intensive Loads in Node.jsabstractAn increasing number of applications are using Node.js, a framework for asynchronous I/O, event-driven, server-side JavaScript. The backbone of Node.js is the single-threaded event loop. Therefore, computationally intensive tasks are bound to the performance of a single core. Modules with different characteristics have been built to provide parallelism and scaling. We evaluate the performance of some representative Node.js multiprocess and multi-thread techniques focusing on their scaling behavior on different environments. We present computation metrics using a compute-intensive task as a constant. Finally, we use statistical analysis to identify similarities and differences in performance with the end goal of providing recommendations on deployment. Maria Patrou, Kenneth B. Kent, Michael Dawson 0001 |
PDP | 2 |
| 2019 | Verilog Loop Unrolling, Module Generation, Part-Select and Arithmetic Right Shift Support in Odin IIabstractVerilog is a hardware description language (HDL) that supports the specification of hardware circuitry and control logic for production, simulation and testing. The subset of the specification used for production is called synthesizable. Verilog-to-Routing (VTR) is a Computer-Aided Design (CAD) flow. It transforms synthesizable Verilog into a placed and routed configuration for a Field Programmable Gate Array (FPGA) architecture specified in XML. The front end of the VTR CAD flow is Odin II. Odin II parses Verilog files and uses them to create a netlist consisting of inputs, outputs, nodes, and connections. Odin II is an open-source research project, and full Verilog language coverage is a work in progress. This work extends Odin II's Verilog support to files containing the arithmetic right shift operator (>>>) and both the + : and - : part-select operators. It also adds support for simple for loops, while loops and loop-based module generation. Dynamic looping constructs are not synthesizable, so all looping constructs are processed before the netlist is generated. This paper will present the missing language features that were implemented, the scope of their implementation, the architecture of the solution, testing and finally the efficiency of the contributions. Scott Ryan Young, Alexandrea Demmings, Nasrin Eshraghi Ivari, Jean-Philippe Legault, Kenneth B. Kent |
RSP | 5 |
| 2019 | A multi-granularity locking scheme for java packedobjects based on a concurrent multiway treeabstractSummary In this paper, we develop a multi‐granularity locking scheme for Java PackedObjects, an experimental enhancement introduced in IBM's J9 Java Virtual Machine. The packed object model organizes data in a multi‐tier manner in which object data can be nested in the container object instead of being pointed to by an object reference, as in the traditional Java object model. This new object data model creates new challenges for multi‐tier data synchronization, requiring concurrent locks on the multi‐tier data of different granularities for maintaining consistency. This is different from the traditional Java synchronization model. In this paper we make use of a concurrent multiway tree to represent the containing and ordering relationship between PackedObjects at different tiers and develop an efficient multi‐granularity locking scheme allowing multiple threads to concurrently manipulate the concurrent multiway tree for synchronization operations. In the evaluation, we compare our new tree‐based multitierSync with the previous multitierSync approaches based on linked‐lists (optimized‐list‐based and lazy‐list‐based). The experimental results show that the tree‐based MultitierPackedSync outperforms the list‐based approaches considerably in different workloads, and the higher the workload, the better the performance gains achieved by the tree‐based MultitierPackedSync. Kenneth B. Kent, Eric E. Aubanel, Stephen A. MacKay, Tobi Agila |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | ROCO: Using a Solid State Drive Cache to Improve the Performance of a Host-Aware Shingled Magnetic Recording Drive
Wenguo Liu 0004, Lingfang Zeng, Dan Feng 0001, Kenneth B. Kent |
J. Comput. Sci. Technol. | 4 |
| 2018 | NUMA Awareness: Improving Thread and Memory ManagementabstractMany Java Virtual Machines (JVM) recognize Non-uniform Memory Access (NUMA) systems and use memory and threads from the available nodes in a distributed manner. However, such a design might not benefit all applications having different memory and thread patterns. A design for a node-isolated memory and thread policy is proposed, called NumaVM. A node-heap resize functionality to retrieve memory, based mostly on each node's memory information, is further described. Additionally, different modes regarding hardware and thread characteristics are used in order to find an optimal one and identify the application attributes that can benefit from specific modes, based on the underlying hardware. Maria Patrou, Kenneth B. Kent, Gerhard W. Dueck, Charlie Gracie, Aleksandar Micic |
SEAA | 2 |
| 2018 | DISTIL: a distributed in-memory data processing system for location-based servicesabstractLocation-based services (LBS) have become an ubiquitous technology and spatio-temporal data generated by LBS is characterized by high volume and velocity. In recent times several projects, such as GeoSpark, SpatialSpark and LocationSpark, have focused on developing spatial data systems that take advantage of the distributed in-memory data processing capability of Spark. However, most of these systems assume immutable spatial data, and they do not support high throughput location data updates that are common in LBS. On the other hand, a few HBase-based systems, such as MD-HBase, have been proposed that support data updates. However, these systems do not take advantage of any distributed in-memory query processing frameworks. Maria Patrou, Puya Memarzia, Suprio Ray, Virendrakumar C. Bhavsar, Kenneth B. Kent, Gerhard W. Dueck |
SIGSPATIAL/GIS | 6 |
| 2018 | Towards Trainable Synthesis for Optimized Circuit Deployment on FPGAabstractField Programmable Gate Arrays (FPGAs) utilize multiple programmable elements and non-programmable blocks. After synthesizing an input Hardware Design Language (HDL) design into a circuit, optimizations are used to discover a satisfactory deployment on a target FPGA. HDLs' compound operations, such as addition, can be implemented in various ways and thus, multiple but functionally equivalent circuits can be synthesized. To leverage this, we propose a methodology that first enables configurable synthesis of compound operations. Second, it trains the system using a set of HDL files and architectures to optimize target performance objectives, such as critical path length and power. We prototyped our technique in the open source Verilog-To-Routing (VTR) tool. We subsequently produced two configuration files targeting different deployment objectives; experimental results with the VTR Verilog benchmarks revealed significant improvements. Jean-Philippe Legault, Panagiotis Patros, Kenneth B. Kent |
RSP | 3 |
| 2017 | Dynamically Compiled Artifact Sharing for CloudsabstractPlatform as a Service (PaaS) clouds provide part of the hardware/software stack and related services to tenant applications. Increased load is handled elastically by scaling, which either modifies the number of instances an application has available on the cloud or increases their available resources. However, because all these instances run inside isolated containers, experience gained by the first instance of an application cannot be easily shared with subsequent scaled instances. This results in both increased startup time and response timeout errors for the scaled instances as well as increased performance interference for any co-located applications; reacquiring this experience is a time-consuming and resource-intensive process. We propose a scalable and secure technique to share dynamically compiled artifacts produced by the first execution instance of an application and otherwise created for intra-OS sharing only with subsequent scaled or restarted instances as a solution to these problems. Our solution abides by the usual PaaS limitations and uses a distributed and containerized cloud service, which we experimentally show to be scalable on a Docker Swarm running on top of a 6-VM cluster; also, we discuss the results of a usability survey for the service's GUI conducted with expert subjects. The effectiveness of the DCAS technique was experimentally tested on an isolated installation of the PaaS software Cloudy Foundry; we measured significant reductions in both the startup time and response errors of scaled out instances as well as performance interference to co-located tenants during scaling. Panagiotis Patros, Dayal Dilli, Kenneth B. Kent, Michael Dawson 0001 |
CLUSTER | 3 |
| 2017 | Investigating the Effect of Garbage Collection on Service Level Objectives of CloudsabstractPlatform as a Service (PaaS) clouds abstract large parts of the hardware/software stack to its tenant clients and provide it as a service. In this paper, we highlight the lack of scientific literature on the problem of Service Level Objective (SLO) satisfaction effects on clouds due to Garbage Collection (GC). To this end, we propose and implement CloudGC, a configurable PaaS application framework that aims in stressing the GC component of the underlying runtime. We use our CloudGC to experimentally evaluate the performance of the four GC policies (Gencon, Balanced, Optavgpause and Optthroughput) available in the IBM J9 Java Runtime, running on top of a local and isolated installation of the PaaS software Cloud Foundry. Panagiotis Patros, Kenneth B. Kent, Michael Dawson 0001 |
CLUSTER | 2 |
| 2017 | A Region-Based Approach to Pipeline Parallelism in Java Programs on MulticoresabstractAs multicore architectures dominate mainstream computing platforms, migrating legacy applications into their parallel representation becomes a viable approach to reaping the benefits of multicore computing. In this paper we present a dataflow analysis tool that assists programmers to exploit the coarse-grained pipeline parallelism in stream-like Java applications on multicores. With this tool, programmers can partition a source Java program into a set of regions, which as pipeline stages, are connected via data channels to execute on multicores. To this end, we propose a simple yet effective framework that leverages JVMTI (JVM Tool Interface) and Java agent techniques to track the data communication patterns among different regions, whereby a stream graph of the program is constructed. The graph is further used by the framework and programmers to re-factor the Java application into a pipelined program so that the potential of the multicores can be fully utilized. This procedure can be repeated in several rounds to progressively improve the performance. By applying this tool to several selected benchmarks, we demonstrate the effectiveness of the approach in terms of the performance improvements of some stream-like Java applications. Yang Wang 0006, Kenneth B. Kent |
PDP | 2 |
| 2017 | Simulation-based circuit-activity estimation for FPGAs containing hard blocksabstractFPGAs are electronic devices that are programmable and can functionally perform equivalently to a number of other circuits. FPGAs are used for both rapid and cheap prototyping of new circuit designs as well as for replacing outdated chip models. Due to their complexity, circuits cannot be practically designed by hand; instead, specialized Computer Aided Design (CAD) software performs this complex task. A major concern for devices is power requirements, which can have adverse effects on both the environment and users. The power requirements of a circuit can be directly connected with its activity, which can be estimated by the CAD tools. In this work, we focus on the open source Verilog-To-Routing (VTR) CAD software and propose an improved activity estimation tool using VTR's synthesizer (Odin II) that extends beyond the capabilities of its current estimator (ACE2), such as proper black box activity propagation and support for circuits containing no clocks or more than one clock. Our results are experimentally evaluated with VTR's FPGA architectures and benchmark circuits. Sean Seeley, Vidya Sankaranaryanan, Zack Deveau, Panagiotis Patros, Kenneth B. Kent |
RSP | 5 |
| 2017 | Naplus: a software distributed shared memory for virtual clusters in the cloudabstractSummary Virtual clusters (VCs) have exhibited various advantages over traditional cluster computing platforms by virtue of their extensibility, reconfigurability, and maintainability. As such, they have become a major execution environment for cloud‐based cluster applications. However, compared with traditional clusters, their distributed‐memory programming paradigm still remains largely unchanged, which implies that cluster applications cannot be efficiently deployed in VCs, especially when virtual machines (VMs) are running in different physical hosts. Recently, some efforts have been made to improve inter‐VM communication, resulting in many studies on how cluster applications could take advantages of VCs. However, most of them mainly focus on the situation that the VMs are all coresident on the same physical machine where the message passing mechanism is usually optimized away by exploiting the host's shared memory. In this paper, we present a design and implementation of Naplus, a kernel‐based virtual machine approach to the inter‐VM communications that are across different physical hosts. Naplus is based on Nahanni, a mechanism for shared‐memory communication in virtual environments. As such, it not only inherits the major merits of Nahanni with respect to flexible data structures and efficient synchronization but also achieves a shared‐memory paradigm among VMs. With Naplus, we enable the size of shared space to be maximized as large as the sum of each machine's local memory to accommodate cluster applications with large memory footprints. We prototype Naplus in a dual‐host system where an empirical study is conducted to show the effectiveness of the Naplus approach in achieving distributed shared memory for VCs in data centers. Copyright © 2017 John Wiley & Sons, Ltd. Lingfang Zeng, Yang Wang 0006, Kenneth B. Kent, Ziliang Xiao |
Softw. Pract. Exp. | 3 |
| 2017 | Toward cost-effective replica placements in cloud storage systems with QoS-awarenessabstractSummary In this paper, we propose a simulation model to study real‐world replication workflows for cloud storage systems. With this model, we present three new methods to maximize the storage space usage during replica creation, and two novel QoS aware greedy algorithms for replica placement optimization. By using a simulation method, our algorithms are evaluated, through a comparison with the existing placement algorithms, to show that (i) a more evenly distributed replicas for a data set can be achieved by using round‐robin methods in replica creation phase and (ii) the two proposed greedy algorithms, namedGS_QoSandGS_QoS_C1, not only have more economical results than those from Chenet al., but also guarantee the QoS for clients. Copyright © 2016 John Wiley & Sons, Ltd. Lingfang Zeng, Yang Wang 0006, Kenneth B. Kent, David Bremner, Cheng-Zhong Xu 0001 |
Softw. Pract. Exp. | 4 |
| 2017 | CosaFS: A Cooperative Shingle-Aware File SystemabstractIn this article, we design and implement a cooperative shingle-aware file system, called CosaFS , on heterogeneous storage devices that mix solid-state drives (SSDs) and shingled magnetic recording (SMR) technology to improve the overall performance of storage systems. The basic idea of CosaFS is to classify objects as hot or cold objects based on a proposed Lookahead with Recency Weight scheme. If an object is identified as a hot (small) object, then it will be served by SSD. Otherwise, cold (large) objects are stored on SMR. For an SMR, large objects can be accessed in large sequential blocks, rendering the performance of their accesses comparable with that of accessing the same large sequential blocks as if they were stored on a hard drive. Small objects, such as inodes and directories, are stored on the SSD where “seeks” for such objects are nearly free. With thorough empirical studies, we demonstrate that CosaFS, as a cooperative shingle-aware file system, with metadata separation and cache-assistance, is a very effective way to handle the disk-based data demanded by the shingled writes and outperforms the device- and host-side shingle-aware file systems in terms of throughput, IOPS, and access latency as well. Lingfang Zeng, Zehao Zhang, Yang Wang 0006, Dan Feng 0001, Kenneth B. Kent |
ACM Trans. Storage | 5 |
| 2016 | A Low Disk-Bound Transaction Logging System for In-memory Distributed Data StoresabstractTransaction logging and snapshotting are techniques used to deliver durability to the data in in-memory data stores. Absolute durability guarantees are delivered to a system by sequentially recording the transaction logs and snapshots to a non-volatile disk. Recent advancements in database restoration techniques have given rise to lock-free fuzzy snapshots. Still the transaction log that completes the fuzzy snapshots is not lock-free. In addition to locking, the major overhead behind the transaction logging technique is the bottleneck involved in storing the logs to a persistent but slower disk. This paper concentrates on implementing an in-memory transaction logging system with a lesser disk dependency. This logging system mainly targets the distributed in-memory data stores that are transaction replicated, eventually consistent and fault tolerant to crash failures. By making logging in-memory, the performance will be improved, but during the crash fails, the state may be lost. On recovery, we restore the current state partially from the locally available fuzzy snapshot and the remaining from the non-failed nodes in the distributed replica. ZooKeeper, a distributed data store that offers distributed coordination as its major service is used to implement and test our research. On average, a 30 times write performance improvement has been achieved with this approach guaranteeing sufficient durability in replicated mode. Dayal Dilli, Kenneth B. Kent, Yang Wang 0006, Cheng-Zhong Xu 0001 |
CLUSTER | 2 |
| 2016 | Automatic detection and elision of reset sub-circuitsabstractElectronic circuits are too complex to be designed by hand so hardware languages, like Verilog, and Computer Aided Design (CAD) tools are used for these purposes. Two main types of circuits are Application-Specific Integrated Circuits (ASICs) and Field Programmable Gate Arrays (FPGAs). ASICs require a reset sub-circuit to initialize their state; however, such a procedure is not necessary for FPGAs that support power-on reset. We propose and evaluate a tool that automatically detects and elides reset sub-circuits as part of the Verilog-to-Routing (VTR) CAD flow and in particular Odin II. Also, our tool can be used to decide if a reset sub-circuit has been properly implemented and can point towards memory components that have not been initialized. Such a tool is the first to our knowledge. Our tests with the VTR Verilog benchmarks and other Verilog circuits showed significant reductions in resource consumption on the target FPGAs as much as 25.9% shorter critical path, 90.39% shorter maximum net, 62.84% fewer used logic blocks and 30.87% fewer used routing elements. Also, our approach yielded significant reductions in the execution time of the placement-and-routing algorithm for the elided circuits that were as high as 4.5 times faster VPR execution and 3.35 times faster overall VTR flow execution. Panagiotis Patros, Kenneth B. Kent |
RSP | 2 |
| 2016 | Inter-JVM SharingabstractSome Java programs lend themselves to being run many times and create the same fixed objects every time. Many of these common objects are Strings. To exploit this trend, we have modified IBM's J9 Java virtual machine (JVM) to allow the same String objects to share (reuse) their internal char[] (character) arrays in each JVM. The first instance of the Java program runs to completion and then sets up the Strings for sharing, so that subsequent instances of the same program can use the char[] arrays that it created instead of recreating them. String sharing will not provide benefit in all applications, but for those that fit the pattern, as exemplified by the Eclipse and H2 benchmarks, we were able to achieve significant heap saving with negligible impact on performance. Copyright © 2015 John Wiley & Sons, Ltd. Adam Richard, Lai Nguyen, Peter Shipton, Kenneth B. Kent, Azden Bierbrauer, Konstantin Nasartschuk, Marcel Dombrowski |
Softw. Pract. Exp. | 4 |
| 2015 | Hard block reduction and synthesis improvements in Odin IIabstractA Field-Programmable Gate Array (FPGA) is an integrated circuit that allows users to program product features and functions after manufacturing. Verilog-to-Routing (VTR) is an open source CAD tool for conducting FPGA architecture and CAD research and development. As one of the core tools of VTR, Odin II is responsible for Verilog elaboration and hard block synthesis. This project describes the improvements in Odin II on three aspects: for loop support, abstract syntax tree (AST) simplification and hard block reduction. This work allows elaboration of a for loop statement by modifying the Abstract Syntax Tree (AST). There are different alternatives to simplify an AST, and this paper demonstrates three ways: simplifying expressions with variables, reducing parameters with values and using shift operations to replace multiplications or divisions. For a circuit design, some hard blocks in the netlist have the same high-level function. This project further provides a method to reduce redundant hard blocks. Each implementation is tested with designed testing cases or sets of benchmarks, and the results of running them through Odin II and VTR are demonstrated. Kenneth B. Kent |
RSP | 2 |
| 2015 | Improving J9 virtual machine with LTTng for efficient and effective tracingabstractSummary The ability to observe the internal operation of the J9 virtual machine is essential for effective performance tuning. To this end, tracing is an important method, which is the action of recording events from a running system with minimum performance overhead for online or off‐line analysis. In this paper, we propose the integration of LTTng, an effective open‐source tracing toolset, with J9 to improve its tracing functions. With this integration, the tracing component is not only decoupled from the virtual machine but also performed efficiently at both user and kernel levels to achieve a high‐throughput result. To validate the integration and its impact performance, some empirical study results based on SpecJBB2005 and SQLBenchmark (supported by instrumented MariaDB) are also presented. Copyright © 2014 John Wiley & Sons, Ltd. Yang Wang 0006, Kenneth B. Kent, Graeme Johnson |
Softw. Pract. Exp. | 2 |
| 2015 | WaFS: A Workflow-Aware File System for Effective Storage Utilization in the CloudabstractWe present WaFS, a user-level file system, and a related scheduling algorithm for scientific workflow computation in the cloud. WaFS’s primary design goal is to automatically detect and gather the explicit and implicit data dependencies between workflow jobs, rather than high-performance file access. Using WaFS’s data, a workflow scheduler can either make effective cost-performance tradeoffs or improve storage utilization. Proper resource provisioning and storage utilization on pay-as-you-go clouds can be more cost effective than the uses of resources in traditional HPC systems. WaFS and the scheduler controls the number of concurrent workflow instances at runtime so that the storage is well used, while the total makespan (i.e., turnaround time for a workload) is not severely compromised. We describe the design and implementation of WaFS and the new workflow scheduling algorithm based on our previous work. We present empirical evidence of the acceptable overheads of our prototype WaFS and describe a simulation-based study, using representative workflows, to show the makespan benefits of our WaFS-enabled scheduling algorithm. Yang Wang 0006, Paul Lu, Kenneth B. Kent |
IEEE Trans. Computers | 3 |
| 2015 | XCollOpts: A Novel Improvement of Network Virtualizations in Xen for I/O-Latency Sensitive Applications on MulticoresabstractIt has long been recognized that the Credit scheduler selectively favors CPU-bound applications whereas for I/O-latency sensitive workloads, such as those related to stream-based audio/video services, it only exhibits tolerable, or even worse, unacceptable performance. The reasons behind this phenomenon are the poor understanding (to some degree) of the virtual machine scheduling as well as the network I/O virtualizations. In order to address these problems and make the system more responsive to the I/O-latency sensitive applications, in this paper, we present XCollOpts which performs a collection of novel optimizations to improve the Credit scheduler and the underlying I/O virtualizations in multicore environments, each from two perspectives. To optimize the schedule, in XCollOpts, we first pinpoint the Imbalanced Multi-Boosting problem among the cores thereby minimizing the system response time by load balancing the BOOST VCPUs. Then, we describe the Premature Preemption problem and address it by monitoring the received network packets in the driver domain and deliberately preventing it from being prematurely preempted during the packet delivery. However, these optimizations on the scheduling strategies cannot be fully exploited if the performance issues of the underlying supportive communication mechanisms are not considered. To this end, we make two further optimizations for the network I/O virtualizations, namely, Multi-Tasklet Pairs and Optimized Small Data Packet. Our empirical studies show that with XCollOpts, we can significantly improve the performance of the latency-sensitive applications at a cost of relatively small system overhead. Lingfang Zeng, Yang Wang 0006, Dan Feng 0001, Kenneth B. Kent |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2014 | Monetary-and-QoS Aware Replica Placements in Cloud-Based Storage SystemsabstractThis paper proposes a replication cost model and two greedy algorithms, named GS QoS and GS QoS C1, for replication placements in cloud-based storage systems. The model aims to minimize replication cost with full consideration of quality of user access to storage nodes. Our two algorithms employ a utility measurement to guide placement procedures. Our final experimental results show that 1) GS QoS outperforms GS QoS C1, 2) both algorithms have more economical results than those from existing greedy replica placement algorithm. Lingfang Zeng, Yang Wang 0006, Xiang Cui, Tan Wee Kiat, David Bremner, Kenneth B. Kent |
CloudCom | 7 |
| 2014 | On Hard Adders and Carry Chains in FPGAsabstractHardened adder and carry logic is widely used in commercial FPGAs to improve the efficiency of arithmetic functions. There are many design choices and complexities associated with such hardening, including circuit design, FPGA architectural choices, and the CAD flow. There has been very little study, however, on these choices and hence we explore a number of possibilities for hard adder design. We also highlight optimizations during front-end elaboration that help ameliorate the restrictions placed on logic synthesis by hardened arithmetic. We show that hard adders and carry chains, when used for simple adders, increase performance by a factor of four or more, but on larger benchmark designs that contain arithmetic, improve overall performance by roughly 15%. We measure an average area increase of 5% for architectures with carry chains but believe that better logic synthesis should reduce this penalty. Interestingly, we show that adding dedicated inter-logic-block carry links or fast carry look-ahead hardened adders result in only minor delay improvements for complete designs. Jason Luu, Conor McCullough, Safeen Huda, Charles Chiasson, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz |
FCCM | 7 |
| 2014 | System-on-chip processor using different FPGA architectures in the VTR CAD flowabstractField Programmable Gate Arrays (FPGA) are often the go to choice for system prototyping and comparison. Circuit design and the impact of hardware architecture can be measured and experimented with using short iteration times. The Verilog To Routing (VTR) CAD flow offers a framework for synthesis and experimentation with customizable FPGA architectures. This paper describes the implemented ability to use the VTR flow for tests and experiments with an ARM processor. This includes different possible FPGA architectures for supporting the ARM processor. A thorough set of experiments is performed which aims to determine the impact of hard block memories, multipliers and adders. The results suggest that 2 bit adder units and 36*36 multipliers offer a good choice of parameters. Konstantin Nasartschuk, Kenneth B. Kent |
RSP | 3 |
| 2014 | VTR 7.0: Next Generation Architecture and CAD System for FPGAsabstractExploring architectures for large, modern FPGAs requires sophisticated software that can model and target hypothetical devices. Furthermore, research into new CAD algorithms often requires a complete and open source baseline CAD flow. This article describes recent advances in the open source Verilog-to-Routing (VTR) CAD flow that enable further research in these areas. VTR now supports designs with multiple clocks in both timing analysis and optimization. Hard adder/carry logic can be included in an architecture in various ways and significantly improves the performance of arithmetic circuits. The flow now models energy consumption, an increasingly important concern. The speed and quality of the packing algorithms have been significantly improved. VTR can now generate a netlist of the final post-routed circuit which enables detailed simulation of a design for a variety of purposes. We also release new FPGA architecture files and models that are much closer to modern commercial architectures, enabling more realistic experiments. Finally, we show that while this version of VTR supports new and complex features, it has a 1.5× compile time speed-up for simple architectures and a 6× speed-up for complex architectures compared to the previous release, with no degradation to timing or wire-length quality. Jason Luu, Jeffrey B. Goeders, Michael Wainberg, Andrew Somerville, Thien Yu, Konstantin Nasartschuk, Miad Nasr, Tim Liu, Nooruddin Ahmed, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz |
ACM Trans. Reconfigurable Technol. Syst. | 11 |
| 2013 | Visual exploration of changing FPGA architectures in the VTR projectabstractDeveloping applications for Field Programmable Gate Array (FPGA) devices utilizes Computer Aided Design (CAD) flows. The transition from a high level Verilog hardware description to the optimized structure of programmed soft logic blocks and routing structure includes stages such as Verilog synthesis, hardware mapping, logical synthesis, packing, placement and routing. The VTR CAD flow is a collaborative project consisting of Odin II (University of New Brunswick), ABC (University of California, Berkeley) and VPR (University of Toronto), which offers an FPGA CAD flow for research and experimentation purposes. This paper describes developments in the visualization and simulation modules of Odin II, the first stage of the CAD flow. The contributions include new netlist visualization possibilities as well as an extended netlist simulator capable of simulating circuits with multiple clocks and providing extended generic structure simulation abilities. This results in the possibility to explore and simulate a larger set of new FPGA architectures and evaluate them using the VTR flow. Konstantin Nasartschuk, Rainer Herpers, Kenneth B. Kent |
RSP | 3 |
| 2012 | FPGA Based Real-Time Tracking Approach with Validation of Precision and PerformanceabstractThis paper presents the implementation and validation of a tracking approach for image processing in hardware. It compares the implementation for the addressed problem on a Field Programmable Gate Array (FPGA) with a software implementation for a General Purpose Processor (GPP) architecture. For both solutions the implementation costs for their development is an important aspect in the validation. This research project is motivated by the MI6 project of the Computer Vision research group, which is located at the Bonn-Rhein-Sieg University of Applied Sciences. The intent of the MI6 project is the tracking of a user in an immersive environment. In this research work the development and validation of a detection system for BLOBs on a Cyclone II FPGA from Altera has been implemented. The analysis of the hardware solution showed a similar precision for the tracking as the software approach. One problem is the large increase of allocated resources when extending the system to process more objects. The implementation of the tracking approach in hardware required much more effort than the software solution. The design of high level problems in hardware for this case are more expensive than the software implementation. The results of the implementation indicate that a mixed hardware/software solution would provide optimal results. Alexander Bochem, Kenneth B. Kent, Rainer Herpers |
DSD | 2 |
| 2012 | Analyzing Bus Load Data Using an FPGA and a MicrocontrollerabstractIn this paper we present the design, implementation, and testing of an evaluation tool for the ongoing development of the Prosthetic Device Communication Protocol (PDCP) which is an open protocol and is featured in the University of New Brunswick's most recent prosthetic limb research project, the UNB Hand System. This prosthetic device utilizes the CAN bus hardware with the PDCP for passing command and data messages between modules within the prosthetic limb system. The PDCP allows abstraction of the underlying bus system and allows different network topologies depending on particular needs. To be able to analyze communication in the CAN layers as well as in the PDCP layer we present our own solutions utilizing an FPGA for CAN bus bandwidth load monitoring and a microcontroller for PDCP monitoring and analysis. Marcel Dombrowski, Kenneth B. Kent, Yves G. Losier, Adam W. Wilson, Rainer Herpers |
DSD | 2 |
| 2012 | The VTR project: architecture and CAD for FPGAs from verilog to routingabstractTo facilitate the development of future FPGA architectures and CAD tools -- both embedded programmable fabrics and pure-play FPGAs -- there is a need for a large scale, publicly available software suite that can synthesize circuits into easily-described hypothetical FPGA architectures. These circuits should be captured at the HDL level, or higher, and pass through logical and physical synthesis. Such a tool must provide detailed modelling of area, performance and energy to enable architecture exploration. As software flows themselves evolve to permit design capture at ever higher levels of abstraction, this downstream full-implementation flow will always be required. This paper describes the current status and new release of an ongoing effort to create such a flow - the 'Verilog to Routing' (VTR) project, which is a broad collaboration of researchers. There are three core tools: ODIN II for Verilog Elaboration and front-end hard-block synthesis, ABC for logic synthesis, and VPR for physical synthesis and analysis. ODIN II now has a simulation capability to help verify that its output is correct, as well as specialized synthesis at the elaboration step for multipliers and memories. ABC is used to optimize the 'soft' logic of the FPGA. The VPR-based packing, placement and routing is now fully timing-driven (the previous release was not) and includes new capability to target complex logic blocks. In addition we have added a set of four large benchmark circuits to a suite of previously-released Verilog HDL circuits. Finally, we illustrate the use of the new flow by using it to help architect a floating-point unit in an FPGA, and contrast it with a prior, much longer effort that was required to do the same thing. Jonathan Rose, Jason Luu, Chi Wai Yu, Opal Densmore, Jeffrey B. Goeders, Andrew Somerville, Kenneth B. Kent, Peter Jamieson, Jason Helge Anderson |
FPGA | 7 |
| 2012 | Improving memory support in the VTR flowabstractVTR is an open source academic FPGA CAD flow designed for exploration of hypothetical FPGA architectures. High level language and architectural feature support are of continuing importance to researchers using the VTR flow to explore these architectures. In this paper we will discuss the extension of VTR to support implicit memories and the elaboration of explicit and implicit memories into soft logic, contributing to improved language support as well as available elaboration options for memories. The addition of full support for architecture-aware memory splitting and padding during elaboration is also detailed and validated which introduces a range of memory elaboration options which were previously not possible. This paper will also detail two experiments designed to validate and explore these new capabilities using the VTR flow. These experiments pave the way for further explorations using these new capabilities. Application of these new features to future architectural explorations will also be discussed. Andrew Somerville, Kenneth B. Kent |
FPL | 2 |
| 2012 | Visualization support for FPGA architecture explorationabstractField Programmable Gate Arrays (FPGA) are used in many fields of research, e.g. to create prototypes of hardware or in applications where hardware functionality has to be changed more frequently. Boolean circuits, which can be implemented by FPGAs are the compiled result of hardware description languages such as Verilog or VHDL. Odin II is a tool, which supports developers in the research of FPGA based applications and FPGA architecture exploration by providing a framework for compilation and verification. In combination with the tools ABC, T-VPACK and VPR, Odin II is part of a CAD flow, which compiles Verilog source code that targets specific hardware resources. This paper describes the development of a graphical user interface as part of Odin II. The goal is to visualize the results of these tools in order to explore the changing structure during the compilation and optimization processes, which can be helpful to research new FPGA architectures and improve the work flow. Konstantin Nasartschuk, Rainer Herpers, Kenneth B. Kent |
RSP | 3 |
| 2012 | Editorial to the Special Issue of Rapid System Prototyping'10
Kenneth B. Kent, Jérôme Hugues |
Softw. Pract. Exp. | 1 |
| 2011 | A framework for verifying functional correctness in Odin IIabstractFPGA architecture exploration is a topic of great interest to hardware researchers. By synthesizing many hardware descriptions with different architecture specifications, it is possible to compare the generated circuits and draw a conclusion about those specifications. In order to be confident in results obtained from this exploration, it is necessary to verify that the circuits have been compiled correctly. In this paper, we outline the simulator implemented in the VTR CAD tool flow. We detail the features of the simulator and cover its value to researchers performing architecture exploration. Joseph C. Libby, Ashley Furrow, Paddy O'Brien, Kenneth B. Kent |
FPT | 4 |
| 2011 | VPR 5.0: FPGA CAD and architecture exploration tools with single-driver routing, heterogeneity and process scalingabstractThe VPR toolset has been widely used in FPGA architecture and CAD research, but has not evolved over the past decade. This article describes and illustrates the use of a new version of the toolset that includes four new features: first, it supports a broad range of single-driver routing architectures, which have superior architectural and electrical properties over the prior multidriver approach (and which is now employed in the majority of FPGAs sold). Second, it can now model, for placement and routing a heterogeneous selection of hard logic blocks. This is a key (but not final) step toward the incluion of blocks such as memory and multipliers. Third, we provide optimized electrical models for a wide range of architectures in different process technologies, including a range of area-delay trade-offs for each single architecture. Finally, to maintain robustness and support future development the release includes a set of regression tests for the software. To illustrate the use of the new features, we explore several architectural issues: the FPGA area efficiency versus logic block granularity, the effect of single-driver routing, and a simple use of the heterogeneity to explore the impact of hard multipliers on wiring track count. Jason Luu, Ian Kuon, Peter Jamieson, Ted Campbell, Andy Gean Ye, Wei Mark Fang, Kenneth B. Kent, Jonathan Rose |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2010 | Odin II - An Open-Source Verilog HDL Synthesis Tool for CAD ResearchabstractIn this work, we present Odin II, a framework for Verilog Hardware Description Language (HDL) synthesis that allows researchers to investigate approaches/improvements to different phases of HDL elaboration that have not been previously possible. Odin II's output can be fed into traditional back-end flows for both FPGAs and ASICs so that these improvements can be better quantified. Whereas the original Odin [1] provided an open source synthesis tool, Odin II's synthesis framework offers significant improvements such as a unified environment for both front-end parsing and netlist flattening. Odin II also interfaces directly with VPR [2], a common academic FPGA CAD flow, allowing an architectural description of a target FPGA as an input to enable identification and mapping of design features to custom features. Furthermore, Odin II can also read the netlists from downstream CAD stages into its netlist data-structure to facilitate analysis. Odin II can be used for a wide range of experiments; in this paper, we show three specific instances of how Odin II can be used by ASIC and FPGA researchers for more than basic synthesis. Odin II is open source and released under the MIT License. Peter Jamieson, Kenneth B. Kent, Farnaz Gharibian, Lesley Shannon |
FCCM | 2 |
| 2010 | Odin II: an open-source verilog HDL synthesis tool for FPGA cad flows (abstract only)abstractOdin II is a high-level Verilog Hardware Description Language synthesis tool. This tool is a significant improvement on the original Odin for a number of reasons including Odin II does both front-end parsing and netlist flattening, Odin II interfaces with VPR architecture description of an FPGA to help identify and use available hard circuits, and Odin II can read in netlists from downstream stages in the VPR 5.0 CAD flow into its netlist data-structure. Odin II is open source and is released under the MIT License. Odin II source code, regression benchmarks, and more documentation can be found at http://www.users.muohio.edu/jamiespa/ODIN II/. Peter Jamieson, Kenneth B. Kent |
FPGA | 2 |
| 2009 | Customizable bit-width in an OpenMP-based circuit design toolabstractAs transistor density grows, increasingly complex hardware designs are implemented. In order to manage this complexity, hardware design can be performed at a higher level of abstraction. High level synthesis enables the automatic conversion of algorithms into hardware implementations, abstracting away the underlying complexities of hardware from the designer. A number of high level synthesis tools have recently been developed, including an OpenMP to Handel-C translator. Improvements to the translator, including a new compiler directive allowing customizable register width, are described. Using a set of benchmark tests, the OpenMP to Handel-C translator is evaluated on several criteria, with the goal of evaluating the variable bit-width effects and identifying further areas for improvement. Timothy F. Beatty, Eric E. Aubanel, Kenneth B. Kent |
FPGA | 3 |
| 2009 | An embedded implementation of the Common Language Infrastructure
Joseph C. Libby, Kenneth B. Kent |
J. Syst. Archit. | 2 |
| 2008 | Automatic Identification of Parallelism in Handel-CabstractHigh level hardware design languages are making it possible for people with little background in hardware design to create their own custom hardware. This allows software designers to begin looking beyond general purpose computing into the realm of customized hardware in order to increase the performance of their applications. The ease with which hardware can be developed using hardware definition languages comes with a cost. Developers accustomed to working in software environments may have issues dealing with some of the more complex facets of hardware design, such as exploiting parallelism. This work aims to alleviate some of the frustration that may occur when attempting to identify and exploit parallelism in a hardware design by providing a set of tools that can automatically identify parallelism in Handel-C hardware designs. Joseph C. Libby, Farnaz Gharibian, Kenneth B. Kent |
DSD | 3 |
| 2008 | A quantitative analysis of the .NET common language runtime
Joshua R. Dick, Kenneth B. Kent, Joseph C. Libby |
J. Syst. Archit. | 2 |
| 2008 | Editorial: Embedded systems - new challenges and future directionsabstractEmbedded Systems- Fabiano Hessel, Kenneth B. Kent, Dionisios N. Pnevmatikatos |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2007 | An Embedded Implementation of the Microsoft Common Language InfrastructureabstractThe common language infrastructure (CLI) provides a framework for managing and executing applications. Developers designing applications for the CLI need not worry about the underlying architecture as it is abstracted from view by the CLI framework. This abstraction, while a boon for developers, leads to degraded performance. It is because of these inefficiencies that the CLI is not well suited for developing embedded applications. It would, however, be beneficial for developers to be able to develop embedded applications using the CLI. In order to address the issues caused by the extra layer of abstraction added by the CLI, an embedded processor is designed and implemented. This processor is capable of natively executing the CLI instruction set which effectively removes the performance problems caused by the extra layer of abstraction. Joseph C. Libby, Kenneth B. Kent |
DSD | 2 |
| 2007 | Analysis of Variable Reordering on the QMDD Representation of Quantum CircuitsabstractIn previous research, a novel structure was discussed for representing the matrices that can be built from an n-variable r-valued reversible/quantum circuit. This structure, called a QMDD, takes on a form similar to that of a reduced-ordered-binary-decision-diagram (ROBDD). It is known that the order of variables used for developing an ROBDD from a binary logic circuit is relevant to the size and structure of that ROBDD. This paper determines what effect, if any, variable order has on the QMDD structure and proposes a simple heuristic for choosing a 'good' variable order. Sharon Van Schaick, Kenneth B. Kent |
DSD | 2 |
| 2007 | High Performance Software-Hardware Network Intrusion Detection SystemabstractNetwork intrusion detection systems (NIDS) and quality of service (QoS) demands have been steadily increasing over the past few years. Current solutions using software become inefficient running on high speed high volume networks and will end up dropping packets. Hardware solutions are available and result in much higher efficiency but present problems such as flexibility and cost. Our proposed system uses a modified version of Snort, a robust widely deployed open-sourced NIDS. It has been found that Snort spends at least 30%-60% of its processing time doing pattern matching. Our proposed system runs Snort in software until it gets to the pattern matching function and then offloads that processing to the field programmable gate array (FPGA). The software can then go on to other processing while it waits for the results from the FPGA. The hardware is able to process data at upto 1.7 GB/s on one Xilinx XC2VP100 FPGA. The design is scaleable and will allow for multiple FPGAs to be used in parallel to increase the processing speed even further. Ryan B. Proudfoot, Kenneth B. Kent, Eric E. Aubanel |
FPT | 2 |
| 2006 | Periodic licensing of FPGA based intellectual propertyabstractAs Field Programmable Gate Arrays (FPGA) gain popularity and become more prevalent in consumer products, the desire to have expiring FPGA Intellectual Property(IP) will also rise. Up to this point, the sale of Intellectual Property (IP) targeting FPGA-based consumer products have not been tremendously profitable for the creators of this IP. The sale of the products containing this IP however have been. This is due in part to the way the IP is licensed. This research investigates the feasibility of physically enforced periodic licensing of FPGA IP. The goal is to design a hardware architecture to support licensable FPGA IP cores targeting consumer products. This work describes a method of licensing IP on FPGAs based on techniques derived from software licensing schemes. Current software and hardware licensing techniques are described in detail, including a survey of current research in the fields of FPGA security, secure memory technologies, and cryptography. A licensing architecture for FPGA IP is proposed, and an implementation on a Xilinx Virtex 2 FPGA demonstrates that expiration of FPGA based IP can be achieved. Future work includes the development of a hardware architecture for consumer products that supports licensable IP cores as well as their delivery. Nathaniel Couture, Kenneth B. Kent |
FPGA | 2 |
| 2006 | Periodic licensing of FPGA based intellectual propertyabstractThis work describes a method of licensing IP on FPGAs based on techniques derived from software licensing schemes. Current software and hardware licensing techniques are described in detail, including a survey of current research in the fields of FPGA security, secure memory technologies, and cryptography. A licensing architecture for FPGA IP is proposed, and an implementation on a Xilinx Vertex 2 FPGA demonstrates that expiration of FPGA based IP can be achieved. Future work includes the development of a hardware architecture for consumer products that supports licensable IP cores as well as their delivery Nathaniel Couture, Kenneth B. Kent |
FPT | 2 |
| 2006 | A systolic array technique for determining common approximate substringsabstractA new technique that makes use of a systolic array structure is proposed for solving the common approximate substring (CAS) problem. This approach extends the technique introduced in (Kent et al., 2006) from the computation of the edit-distance between two strings to the more encompassing CAS problem. The technique presented is validated and analyzed through simulation Kenneth B. Kent, Jacqueline E. Rice |
ISCAS | 1 |
| 2005 | Hardware-Based Implementation of the Common Approximate Substring AlgorithmabstractAn implementation of an algorithm for string matching, commonly used in DNA string analysis, using configurable technology is proposed. The design of the circuit allows for pipelining to provide a performance increase. The proposal is unique in that we suggest a design that is specific to certain parameters of the problem, but may be reused for any particular instance of the problem that matches these parameters. The use of a field programmable gate array allows the implementation to be instance specific, thus ensuring maximal usage of the hardware. Analysis and preliminary results based on a prototype implementation are presented. Kenneth B. Kent, Sharon Van Schaick, Jacqueline E. Rice, Patricia A. Evans |
DSD | 1 |
| 2005 | Configurable hardware solutions for computing autocorrelation coefficients: a case study (abstract only)abstractThere are many computationally intensive problems in the area of digital design and logic synthesis. Some of these have no "good" solution; that is, simply by their definition they have exponential run-times. In order to overcome this, we examine the possibility of a configurable hardware solution to speed up one such problem. The computation of the problem is carried out on a Field Programmable Gate Array (FPGA), where the problem is encoded in such a way that within certain parameters, the design of the solution need not be changed for working with a variety of benchmark circuits. This saves considerably on compilation and configuration times. The use of configurable hardware, however, still allows for reconfiguration in situations where the parameters change significantly enough to require an altered approach. We examine two implementations of the problem, which in this case consists of computing the autocorrelation coefficients for a Boolean function. Jacqueline E. Rice, Kenneth B. Kent, Troy Ronda, Zhao Yong |
FPGA | 2 |
| 2002 | Hardware Architecture for Java in a Hardware/Software Co-Design of the Virtual MachineabstractThis paper discusses the hardware architecture used in the hw/sw co-design of a Java virtual machine. The paper briefly outlines the partitioning of instructions and support for the virtual machine. Discussion concerning the hardware architecture follows focusing on the special requirements that must be considered for the target environment. A comparison is performed between this design and that of picoJava, a stand-alone processor for Java. The paper concludes with benchmark results for this architecture compared with software execution. Kenneth B. Kent, Micaela Serra |
DSD | 1 |