Michel A. Kinsy

dblp:73/7116 · DBLP profile ↗
← Back
49ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0002-1432-6939ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 7 first-author · 7 since 2021Security and privacy · 5 · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1
YearPublicationVenuePosition
2026 PrivSpike: A Privacy-Preserving Inference Framework for Deep Spiking Neural Networks Using Homomorphic Encryption
Nges Brian Njungle, Eric Jahns, Milan Stojkov, Michel A. Kinsy
ACNS (2)4
2026 FHEON: A Configurable Framework for Developing Privacy-Preserving Encrypted Neural Networks
abstract
The widespread adoption of Machine Learning as a Service raises critical privacy and security concerns, particularly about data confidentiality and trust in both cloud providers and the machine learning models provided. Homomorphic Encryption (HE) has emerged as a promising solution to these problems, allowing computations on encrypted data without decryption. Despite its potential, existing works that integrate HE into neural networks are often limited to specific architectures or classes. This leaves a wide gap in providing a framework for easy development of HE-friendly privacy-preserving neural network models similar to what we have in the broader field of machine learning. In this paper, we present FHEON, an open-source configurable framework for developing privacy-preserving neural network models for inference using the CKKS scheme of HE. FHEON introduces optimized and configurable implementations of privacy-preserving neural network layers, including convolution layers, average pooling layers, ReLU activation functions, and fully connected layers. These layers are configured using standard parameters such as input channels, output channels, kernel size, stride, and padding to support arbitrary convolution neural network (CNN) architectures. Furthermore, FHEON provides utility functions that ease usage and adoption. We assess the performance of FHEON using several CNN architectures, including LeNet-5, VGG-11, VGG-16, ResNet-20, and ResNet-34. FHEON maintains encrypted-domain accuracies within +-1% of their plaintext counterparts for ResNet-20 and LeNet-5 models. Notably, on a consumer-grade CPU, the models built on FHEON achieved 98.5% accuracy with a latency of 13 seconds on MNIST using LeNet-5, and 92.2% accuracy with a latency of 403 seconds on CIFAR-10 using ResNet-20. Though configurable, FHEON outperform all state-of the-art HE inference works in both latency and memory utilization. Additionally, FHEON operates within a practical memory budget requiring not more than 42.3 GB for VGG-16.
Nges Brian Njungle, Eric Jahns, Michel A. Kinsy
Proc. Priv. Enhancing Technol.3
2026 SentinelTouch: A Lightweight Privacy-Preserving Biometric-Fingerprinting Authentication and Identification System Based on Neural Networks and Homomorphic Encryption
abstract
Biometric fingerprint authentication and identification systems are increasingly deployed, yet widespread adoption in cloud and server-based platforms remains hindered by privacy and security concerns. Unlike passwords, compromised fingerprints are immutable, making their secure storage and computation paramount. Homomorphic Encryption (HE) offers strong privacy guarantees for fingerprint data processing by enabling computation directly on encrypted data. However, the high dimensionality of fingerprint images and the complexity of the neural networks needed for accurate recognition creates significant bottlenecks, which hinder the practical deployment of HE in this domain. We introduce SentinelTouch, an open-source framework for privacy-preserving fingerprint authentication and identification that delivers both efficiency and accuracy in HE environments. Our key insight is a twofold optimization: (1) a preprocessing pipeline that reduces fingerprint image dimensions to as low as 28x28 while preserving most of its discriminative details, and (2) the design of a lightweight, HE-friendly neural network that generalizes effectively on this compact data. We evaluate two deployment pipelines: (1) a full-privacy pipeline, where encrypted images are processed entirely under HE settings, achieving user identification in a one-to-many setting in just 16 seconds. (2) A hybrid pipeline, where only encrypted embeddings are processed under HE settings, achieving one-to-many user identification in 284 milliseconds. Our results show a 10x and 2.5x speedup over the current state-of-the-art results in both pipelines, respectively. Across the SOKOTO and PolyU datasets, SentinelTouch achieves Rank-1 accuracies within +-0.1% of the leading encrypted systems. This work demonstrates the practicality of end-to-end privacy-preserving fingerprint identification and authentication systems, offering HE security guarantees and utilizing neural networks, without compromising accuracy.
Nges Brian Njungle, Eric Jahns, Mishel Jyothis Paul, Michel A. Kinsy
Proc. Priv. Enhancing Technol.4
2025 TRACED: Trust-Aware Clustering and Dynamic Role Management for Secure Edge Systems
abstract
Edge and IoT environments face growing security threats due to their distributed and dynamic nature, where malicious nodes can compromise clustering, coordination, and data integrity. We introduce TRACED, a fully distributed framework for trust-aware clustering and role assignment in adversarial environments. TRACED continuously evaluates node trustworthiness via lightweight challenge-response mechanisms, heartbeat monitoring, and behavioral analysis, updating trust scores with exponential moving averages. Nodes dynamically transition between worker and cluster head roles based on trust thresholds, capability, and resilience metrics. The framework integrates a Sybil protection layer using unique ID checks, publickey validation, and rate limiting to prevent fake node infiltration. Additionally, TRACED features robust heartbeat-based failure detection, adaptive role reassignment, and detailed metrics logging for post-analysis. Through extensive simulations across multiple scenarios, TRACED demonstrated resilience against Sybil attacks, dynamically isolated misbehaving nodes, and maintained functional clustering without central coordination. These results confirm that TRACED enables secure, adaptive, and fully decentralized clustering in volatile edge and IoT environments.
Luigi Mastromauro, Muslum Ozgur Ozmen, Michel A. Kinsy
AICCSA3
2025 AQUILA: A Flexible Architecture Guideline for Building Custom Distributed Systems Testbeds
Luigi Mastromauro, Edwin Kayang, Mishel Jyothis Paul, Eric Jahns, Muslum Ozgur Ozmen, Michel A. Kinsy
EUC6
2025 Gotta Hash 'Em All! Accelerating Hash Functions for Zero-Knowledge Proof Applications
abstract
Collision-resistant cryptographic hash functions (CRHs) are crucial for security, particularly for message authentication in Zero-knowledge Proof (ZKP) applications. However, traditional CRHs like SHA-2 or SHA-3, while optimized for CPUs, generate large circuits, rendering them inefficient in the ZK domain. Conversely, ZK-friendly hashes are designed for circuit efficiency but struggle on conventional hardware, often orders of magnitude slower than standard hashes due to their reliance on expensive finite field arithmetic. To bridge this performance gap, we present HashEmAll, a novel collection of FPGA-based realizations for three prominent ZK-friendly hashes: Griffin, Rescue-Prime, and Reinforced Concrete. Each offers distinct optimization pro les, with both area-optimized and latency-optimized variants available, allowing users to tailor hardware selection to specific application constraints regarding resource utilization and performance.Our extensive evaluation shows that latency-optimized HashEmAll designs outperform CPU implementations by at least 10×, with the leading design achieving a 23× speedup. These gains are coupled with lower power consumption and compatibility with accessible FPGAs. Importantly, the highly parallel and pipelined architecture of HashEmAll enables significantly better practical scaling than CPU-based approaches towards building real-world ZKP applications, such as data commitments with Merkle Trees, by mitigating the hashing bottleneck for large trees. This highlights the suitability of HashEmAll for real-world ZKP applications involving large-scale data authentication. We also highlight the ability to translate the HashEmAll methodology to various ZK-friendly hash functions and different field sizes.
Nojan Sheybani, Tengkai Gong, Anees Ahmed, Nges Brian Njungle, Michel A. Kinsy, Farinaz Koushanfar
ICCAD5
2025 R-Visor: An Extensible Dynamic Binary Instrumentation and Analysis Framework for Open Instruction Set Architectures
abstract
Binary instrumentation tools are widely used to facilitate the development of hardware and software systems. Traditionally, these tools are designed around a fixed Instruction Set Architecture (ISA) specification. However there is a shift in the architectural community towards open ISAs, whose key feature is the ability to add custom ISA extensions. The lack of extensibility in traditional binary instrumentation tools limits their capacity to adapt to these evolving ISAs, thus hindering their ability to analyze and modify binaries built for open ISAs.
Edwin Kayang, Mishel Jyothis Paul, Eric Jahns, Muslum Ozgur Ozmen, Milan Stojkov, Kevin Rudd, Michel A. Kinsy
LCTES7
2025 A Safety-Centric Analysis and Benchmarks of Modern Open-Source Homomorphic Encryption Libraries
Nges Brian Njungle, Milan Stojkov, Michel A. Kinsy
SECRYPT3
2024 AMAZE: Accelerated MiMC Hardware Architecture for Zero-Knowledge Applications on the Edge
abstract
Collision-resistant, cryptographic hash (CRH) functions have long been an integral part of providing security and privacy in modern systems. Certain constructions of zero-knowledge proof (ZKP) protocols aim to utilize CRH functions to perform cryptographic hashing. Standard CRH functions, such as SHA2, are inefficient when employed in the ZKP domain, thus calling for ZK-friendly hashes, which are CRH functions built with ZKP efficiency in mind. The most mature ZK-friendly hash, MiMC, presents a block cipher and hash function with a simple algebraic structure that is well-suited, due to its achieved security and low complexity, for ZKP applications. Although ZK-friendly hashes have improved the performance of ZKP generation in software, the underlying computation of ZKPs, including CRH functions, must be optimized on hardware to enable practical applications. The challenge we address in this work is determining how to efficiently incorporate ZK-friendly hash functions, such as MiMC, into hardware accelerators, thus enabling more practical applications. In this work, we introduce AMAZE, a highly hardware-optimized open-source framework for computing the MiMC block cipher and hash function. Our solution has been primarily directed at resource-constrained edge devices; consequently, we provide several implementations of MiMC with varying power, resource, and latency profiles. Our extensive evaluations show that the AMAZE-powered implementation of MiMC outperforms standard CPU implementations by more than 13×. In all settings, AMAZE enables efficient ZK-friendly hashing on resource-constrained devices. Finally, we highlight AMAZE's underlying open-source arithmetic backend as part of our end-to-end design, thus allowing developers to utilize the AMAZE framework for custom ZKP applications.
Anees Ahmed, Nojan Sheybani, Davi Moreno, Nges Brian Njungle, Tengkai Gong, Michel A. Kinsy, Farinaz Koushanfar
ICCAD6
2022 NeuroFabric: Hardware and ML Model Co-Design for A Priori Sparse Neural Network Training
abstract
Sparse Deep Neural Networks (DNN) offer a large improvement in model storage requirements, execution latency and execution throughput. DNN pruning is contingent on knowing model weights, so networks can be pruned only after training. A priori sparse neural networks have been proposed as a way to extend sparsity benefits to the training process as well. Selecting a topology a priori is also beneficial for hardware accelerator specialization, lowering power, chip area, and latency.We present NeuroFabric, a hardware-ML model co-design approach that jointly optimizes a sparse neural network topology and a hardware accelerator configuration. NeuroFabric replaces dense DNN layers with cascades of sparse layers with a specific topology. We present an efficient and data-agnostic method for sparse network topology optimization, and show that parallel butterfly networks with skip connections achieve the best accuracy independent of sparsity or depth. We also present a multi-objective optimization framework that finds a Pareto frontier of hardware-ML model configurations over six objectives: accuracy, parameter count, throughput, latency, power, and hardware area.
Mihailo Isakov, Michel A. Kinsy
ICCD2
2022 A Taxonomy of Error Sources in HPC I/O Machine Learning Models
abstract
I/O efficiency is crucial to productivity in scientific computing, but the growing complexity of HPC systems and applications complicates efforts to understand and optimize I/O behavior at scale. Data-driven machine learning-based I/O throughput models offer a solution: they can be used to identify bottlenecks, automate I/O tuning, or optimize job scheduling with minimal human intervention. Unfortunately, current state-of-the-art I/O models are not robust enough for production use and underperform after being deployed. We analyze four years of application, scheduler, and storage system logs on two leadership-class HPC platforms to understand why I/O models underperform in practice. We propose a taxonomy consisting of five categories of I/O modeling errors: poor application and system modeling, inadequate dataset coverage, I/O contention, and I/O noise. We develop litmus tests to quantify each category, allowing researchers to narrow down failure modes, enhance I/O throughput models, and improve future generations of HPC logging and analysis tools.
Mihailo Isakov, Mikaela Currier, Eliakin Del Rosario, Sandeep Madireddy, Prasanna Balaprakash, Philip H. Carns, Robert B. Ross, Glenn K. Lockwood, Michel A. Kinsy
SC9
2021 Distributed Memory Guard: Enabling Secure Enclave Computing in NoC-based Architectures
abstract
Emerging applications, like cloud services, are demanding more computational power, while also giving rise to various security and privacy challenges. Current multi-/many-core chip designs boost performance by using Networks-on-Chip (NoC) based architectures. Although NoC-based architectures significantly improve communication concurrency, they have thus far lack adequate security mechanisms such as enforceable process isolation. On the other hand, new security-aware architectures that protect applications and sensitive services in isolated execution environments, i.e., enclaves, have not been extended to provide comprehensive protection for NoC platforms. These enclave-based architectures (i) lack secure enclave-device interaction, (ii) cannot include unmodifiable third-party IP, or (iii) provide flexible enclave memory management.To address these design challenges, we introduce a new hardware security primitive, the Distributed Memory Guard, and design the first security architecture that protects sensitive services in NoC-based enclaves. We provide evaluation of this reference architecture and highlight the fact that one can design a scalable (i.e., NoC-based) and secure (i.e., enclave-based) architecture with minimal hardware complexity and system performance overhead.
Ghada Dessouky, Mihailo Isakov, Michel A. Kinsy, Pouya Mahmoody, Miguel Mark, Ahmad-Reza Sadeghi, Emmanuel Stapf, Shaza Zeitouni
DAC3
2021 xBGAS: A Global Address Space Extension on RISC-V for High Performance Computing
abstract
The tremendous expansion of data volume has driven the transition from monolithic architectures towards systems integrated with discrete and distributed subcomponents in modern scalable high performance computing (HPC) systems. As such, multi-layered software infrastructures have become essential to bridge the gap between heterogeneous commodity devices. However, operations across synthesized components with divergent interfaces inevitably lead to redundant software footprints and undesired latency. Therefore, a scalable and unified computing platform, capable of supporting efficient interactions between individual components, is desirable for largescale data-intensive applications. In this work, we introduce the Extended Base Global Address Space, or xBGAS, microarchitecture extension to the RISC-V instruction set architecture (ISA) for scalable high performance computing. The xBGAS extension provides native ISA-level support for direct accesses to remote shared memory by mapping remote data objects into a system's extended address space. We perform both software and hardware evaluations of the xBGAS design. The results show that xBGAS reduces instruction count generated by interprocess communication by 69.26% on average. Overall, xBGAS achieves an average performance gain of 21.96% (up to 37.29%) across the tested workloads.
Xi Wang 0009, John D. Leidel, Brody Williams, Alan Ehret, Miguel Mark, Michel A. Kinsy, Yong Chen 0001
IPDPS6
2021 Security Threat Analyses and Attack Models for Approximate Computing Systems: From Hardware and Micro-architecture Perspectives
abstract
Approximate computing (AC) represents a paradigm shift from conventional precise processing to inexact computation but still satisfying the system requirement on accuracy. The rapid progress on the development of diverse AC techniques allows us to apply approximate computing to many computation-intensive applications. However, the utilization of AC techniques could bring in new unique security threats to computing systems. This work does a survey on existing circuit-, architecture-, and compiler-level approximate mechanisms/algorithms, with special emphasis on potential security vulnerabilities. Qualitative and quantitative analyses are performed to assess the impact of the new security threats on AC systems. Moreover, this work proposes four unique visionary attack models, which systematically cover the attacks that build covert channels, compensate approximation errors, terminate normal error resilience mechanisms, and propagate additional errors. To thwart those attacks, this work further offers the guideline of countermeasure designs. Several case studies are provided to illustrate the implementation of the suggested countermeasures.
Pruthvy Yellu, Landon Buell, Miguel Mark, Michel A. Kinsy, Dongpeng Xu 0001, Qiaoyan Yu
ACM Trans. Design Autom. Electr. Syst.4
2020 Remote Atomic Extension (RAE) for Scalable High Performance Computing
abstract
Emerging data-intensive applications such as graph analytics, machine learning, and data-driven scientific computing are driving the evolution of high-performance computing (HPC) systems from monolithic to scaled-out, heterogeneous, and complex architectures. In these systems, enormous data sets are mapped to discrete nodes to improve the performance of the system by using distributed storage and computing resources. As such, these data distributions induce frequent cross-node data transactions which challenge the performance of large-scale systems. Global atomic operations are one emerging class of the remote data operations that enable lock-free remote shared data operations. However, the cross-node read-modify-write operations consist of multiple distinct data operations and specific atomicity management, which induces a large amount of overhead. As such, these global atomic operations require an efficient communication methodology Existing advanced compo-nents, such as network interface controllers, network fabrics, network-on-chip (NoC) interconnects, are architected together to improve the system performance. However, complex software infrastructures are needed to provide integration between each discrete component. As a result, the redundant software routines across distinct devices induce a large amount of overhead that causes performance degradationIn this paper, we propose a remote atomic extension (RAE) design that provides inherent ISA-level instructions and micro-architecture support for remote atomic operations based on the RISC-V instruction set architecture (ISA). We design a toolchain and evaluate the RAE infrastructure via simulation. Our experiment results show that RAE eliminates 89.71% of the redundant software instructions used for remote atomic accesses and improves the performance by 17.61% on average (up to 23.35%), compared with the OpenSHMEM.
Xi Wang 0009, Brody Williams, John D. Leidel, Alan Ehret, Michel A. Kinsy, Yong Chen 0001
DAC5
2020 Design-flow Methodology for Secure Group Anonymous Authentication
abstract
In heterogeneous distributed systems, computing devices and software components often come from different providers and have different security, trust, and privacy levels. In many of these systems, the need frequently arises to (i) control the access to services and resources granted to individual devices or components in a context-aware manner and (ii) establish and enforce data sharing policies that preserve the privacy of the critical information on end users. In essence, the need is to authenticate and anonymize an entity or device simultaneously, two seemingly contradictory goals. The design challenge is further complicated by potential security problems, such as man-in-the-middle attacks, hijacked devices, and counterfeits. In this work, we present a system design flow for a trustworthy group anonymous authentication protocol (GAAP), which not only fulfills the desired functionality for authentication and privacy, but also provides strong security guarantees.
Rashmi S. Agrawal 0001, Lake Bu, Eliakin Del Rosario, Michel A. Kinsy
DATE4
2020 Fast Arithmetic Hardware Library For RLWE-Based Homomorphic Encryption
abstract
With billions of devices connected over the internet, the rise of sensor-based electronic devices have led to cloud computing being used as a commodity technology service. These sensor-based devices are often small and limited by power, storage, or compute capabilities, and hence, they achieve these capabilities via cloud services. However, this gives rise to data privacy issues as sensitive data is stored and computed over the cloud, which at most times, is a shared resource. Homomorphic encryption can be used along with cloud services to perform computations on encrypted data, guaranteeing data privacy. While about a decade’s work on improving homomorphic encryption has ensured its practicality, it is still several magnitudes slower than expected, making it expensive and infeasible to use. In this work, we propose a first-of-its-kind FPGA-based arithmetic hardware library that focuses on accelerating the key arithmetic operations involved in Ring Learning with Error (RLWE) based homomorphic encryption. We design and implement the FPGAbased Residue Number System (RNS), Chinese Remainder Theorem (CRT), modulo inverse and modulo reduction operations as a first step. For all of these operations, we include a hardware cost efficient serial, and a fast parallel implementation in the library. A modular and parameterized design approach helps in easy customization, provides flexibility to extend these operations for use in most homomorphic encryption applications, and fits well into emerging FPGA-equipped cloud architectures.
Rashmi S. Agrawal 0001, Lake Bu, Michel A. Kinsy
FCCM3
2020 Towards Programmable All-Digital True Random Number Generator
abstract
Random number generator (RNG) is a core component in many applications such as scientific research, testing and diagnosis, gaming, and cryptosystems (e.g., obfuscation, encryption, and authentication). Although, there are various RNG designs targeting specific application goals such as low-power, high-throughput, stronger security guarantees, a universal programmable RNG design has remained elusive. Indeed, it is a challenge to have only one RNG unit in a system with multiple compute modules with different randomness requirements. In this work, we aim to provide a practical solution to this design challenge by proposing a multi-purpose true random number generator (TRNG), which can be configured in real time to generate random sequences with different requirements. Such a programmable TRNG is able to supply random bits to multiple modules with different demands. The proposed TRNG is a highly convenient multi-purpose hardware primitive that can be deployed in many designs as it provides a tunable physical entropy source and a dynamic cost-performance trade-off.
Rashmi S. Agrawal 0001, Lake Bu, Eliakin Del Rosario, Michel A. Kinsy
ACM Great Lakes Symposium on VLSI4
2020 Quantum-Proof Lightweight McEliece Cryptosystem Co-processor Design
abstract
Due to the rapid advances in the development of quantum computers and their susceptibility to errors, there is a renewed interest in error correction algorithms. In particular, error correcting code-based cryptosystems have reemerged as a highly desirable coding technique. This is due to the fact that most classical asymmetric cryptosystems will fail in the quantum computing era. However, code-based cryptosystems are still secure against quantum computers, since the decoding of linear codes remains NP-hard even on these computing systems. One such code-based cryptosystem was proposed by McEliece. The classic McEliece cryptosystem uses binary Goppa code, which is known for its good code rate and error correction capability. However, its key generation and decoding procedures have a high computation complexity. In this work, we propose the design of a public-key encryption and decryption coprocessor based on a new variant of the McEliece cryptosystem. This co-processor takes advantage of non-binary Orthogonal Latin Square Code to achieve much smaller computation complexity and key size. We also propose a hardware-cost efficient, fully-parameterized FPGA-based implementation of the co-processor to perform fast encoding and decoding operations. When compared to an existing classic McEliece cryptosystem, we observe a speed up of about 3.3 ×.
Rashmi S. Agrawal 0001, Lake Bu, Michel A. Kinsy
ICCD3
2020 HPC I/O throughput bottleneck analysis with explainable local models
abstract
With the growing complexity of high-performance computing (HPC) systems, achieving high performance can be difficult because of I/O bottlenecks. We analyze multiple years' worth of Darshan logs from the Argonne Leadership Computing Facility's Theta supercomputer in order to understand causes of poor I/O throughput. We present Gauge: a data-driven diagnostic tool for exploring the latent space of supercomputing job features, understanding behaviors of clusters of jobs, and interpreting I/O bottlenecks. We find groups of jobs that at first sight are highly heterogeneous but share certain behaviors, and analyze these groups instead of individual jobs, allowing us to reduce the workload of domain experts and automate I/O performance analysis. We conduct a case study where a system owner using Gauge was able to arrive at several clusters that do not conform to conventional I/O behaviors, as well as find several potential improvements, both on the application level and the system level.
Mihailo Isakov, Eliakin Del Rosario, Sandeep Madireddy, Prasanna Balaprakash, Philip H. Carns, Robert B. Ross, Michel A. Kinsy
SC7
2020 Addressing a New Class of Reliability Threats in 3-D Network-on-Chips
abstract
Network-on-chips (NoCs) are vulnerable to transient and permanent faults caused by thermal violations, aging effects, component wear out, or even transient fault sources. Although some of these faults are addressed by previous research, we show that there are reliability threats in 3-D NoCs that go beyond the reliability issues investigated in 2-D interconnect networks. First, we highlight one such class of reliability threats and discuss their manifestations in 3-D NoCs. Second, we propose a thermal, reliability, and performance-aware routing algorithm to tackle: 1) previously established fault models and 2) the new highlighted class of reliability threats in partially connected 3-D NoCs. The proposed routing algorithm takes into account the states of routers and both the horizontal and through silicon via (TSV) links, along with the temperatures of routers and cores. It then routes the packets around failed or overheated links and routers, achieving lower latencies by avoiding misrouting. To achieve this, the proposed routing algorithm uses the concept of vertical link announcement to inform nodes in the network of the working condition of vertical links. We evaluate the proposed routing algorithm under a wide range of working conditions using the access Noxim NoC simulator. Results show that the proposed routing algorithm: 1) is able to tolerate almost any number and pattern of vertical link failures; 2) is reliable against the newly identified reliability threats; and 3) improves the latency and temperature distribution of the network compared to previously proposed routing algorithms.
Ebadollah Taheri, Mihailo Isakov, Ahmad Patooghy, Michel A. Kinsy
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 BRISC-V: An Open-Source Architecture Design Space Exploration Toolbox
abstract
In this work, we introduce the BRISC-V Toolbox, a register-transfer level (RTL) tool for architecture design space exploration. The BRISC-V Toolbox is an open-source, parameterized, synthesizable set of RTL modules for designing RISC-V based single and multi-core architecture systems. The toolbox is designed with a high degree of modularity. It provides highly parameterized, composable RTL modules for fast and accurate exploration of different RISC-V based core complexities, multi-level caching and memory organizations, system topologies, router architectures and routing schemes. BRISC-V can be used for both RTL simulation and FPGA based emulation. The hardware modules are implemented in synthesizable Verilog using no vendor-specific blocks. The toolbox includes a GCC RISC-V compiler tool-chain to assist in developing software for the cores and a web based system configuration graphical user interface (GUI). The BRISC-V Toolbox supports a myriad of RISC-V architectures, ranging from a simple single cycle processor to a multi-core SoC with a complex memory hierarchy and a network-on-chip. The modules are designed to support incremental additions and modifications. The module interfaces are carefully designed to enable an arbitrary pipeline depth and changes to a single module without impacting the rest of the system. The BRISC-V platform allows researchers to quickly instantiate complete working RISC-V multicore systems with synthesizable RTL correctness and make targeted modifications to fit their needs.
Sahan Bandara, Alan Ehret, Donato Kava, Michel A. Kinsy
FPGA4
2019 Open-Source FPGA Implementation of Post-Quantum Cryptographic Hardware Primitives
abstract
The following topics are dealt with: field programmable gate arrays; learning (artificial intelligence); convolutional neural nets; logic design; system-on-chip; power aware computing; cryptography; cloud computing; reconfigurable architectures; neural nets.
Rashmi S. Agrawal 0001, Lake Bu, Alan Ehret, Michel A. Kinsy
FPL4
2019 Secure Computing Systems Design Through Formal Micro-Contracts
abstract
Two enduring concepts in computer system design are abstraction levels and layered composition. The design generally takes a layered approach where each layer implements a different abstraction of the system. The layers communicate through interfaces that are designed to support functional specification of the system as a whole. Traditionally, these layers and interfaces have primarily focused on functionality and efficiency --- performance, power and area. Security-related issues are often overlooked or deferred until later in the design cycle or applied as add-ons when some security features are explicitly required. The challenge with this approach is that chasing security implications of certain design decisions along the multiple layers is a complex and error-prone task. Therefore, in this work, we are introducing the notion of "security micro-contracts" or simply "micro-contracts". We propose a novel secure computer systems design approach through minimal contracts --- micro-contracts --- between adjacent layers. These contracts have strict structures that contain security-relevant details of each connected layer and the secure-properties that have to be preserved to assure confidentiality, integrity and availability of the data of interest. Micro-contracts may be used as (i) basic formalism for proving security properties of computing systems both in the software and hardware layers and across them or (ii) run time security policy checks.
Michel A. Kinsy, Novak Boskov
ACM Great Lakes Symposium on VLSI1
2019 Security Threats in Approximate Computing Systems
abstract
Approximate computing systems improve energy efficiency and computation speed at the cost of reduced accuracy on system outputs. Existing efforts mainly explore the feasible approximation mechanisms and their implementation methods. There is limited work that investigates the security threats brought by approximate computing. To fill this gap, we first analyze the approximate mechanisms used in approximate system, software, storage, and arithmetic circuits, and then propose potential attacks that will challenge the integrity and security of approximate systems. Some illustrative examples are provided accordingly to showcase the consequences of the proposed new attacks.
Pruthvy Yellu, Novak Boskov, Michel A. Kinsy, Qiaoyan Yu
ACM Great Lakes Symposium on VLSI3
2019 A secure and robust scheme for sharing confidential information in IoT systems
Lake Bu, Mihailo Isakov, Michel A. Kinsy
Ad Hoc Networks3
2019 Bulwark: Securing implantable medical devices communication channels
Lake Bu, Mark G. Karpovsky, Michel A. Kinsy
Comput. Secur.3
2019 Design Space Exploration of Neural Network Activation Function Circuits
abstract
The widespread application of artificial neural networks has prompted researchers to experiment with field-programmable gate array and customized ASIC designs to speed up their computation. These implementation efforts have generally focused on weight multiplication and signal summation operations, and less on activation functions used in these applications. Yet, efficient hardware implementations of nonlinear activation functions like exponential linear units (ELU), scaled ELU (SELU), and hyperbolic tangent (tanh), are central to designing effective neural network accelerators, since these functions require lots of resources. In this paper, we explore efficient hardware implementations of activation functions using purely combinational circuits, with a focus on two widely used nonlinear activation functions, i.e., SELU and tanh. Our experiments demonstrate that neural networks are generally insensitive to the precision of the activation function. The results also prove that the proposed combinational circuit-based approach is very efficient in terms of speed and area, with negligible accuracy loss on the MNIST, CIFAR-10, and IMAGE NET benchmarks. Synopsys design compiler synthesis results show that circuit designs for tanh and SELU can save between ${\times 3.13\sim \times 7.69}$ and ${ {\times 4.45\sim \times 8.45}}$ area compared to the look-up table/memory-based implementations, and can operate at 5.14 GHz and 4.52 GHz using the 28-nm SVT library, respectively. The implementation is available at: https://github.com/ThomasMrY/ActivationFunctionDemo.
Tao Yang 0032, Yadong Wei, Zhijun Tu, Haolun Zeng, Michel A. Kinsy, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2018 Weighted Group Decision Making Using Multi-identity Physical Unclonable Functions
abstract
To enable next-generation distributed and connected computing systems, we must address the context-aware chip authentication challenge. An important remaining gap in the design of these systems is the enabling of multi-personality authentication to support applications or schemes requiring a single device to own manifold legitimate identities. In this work, we propose a Multi-identity Physical Unclonable Function (Mi-PUF) assisted weighted group decision making scheme. The Mi-PUF approach enables individual devices to be authenticated and associated with multiple identities in order to hold different number of ballots. Hence, devices with higher impact in a decision making network will have more weight than the less influential ones. Besides the introduction of the scheme, the design and FPGA implementation details of the Mi-PUF are explored and presented.
Lake Bu, Michel A. Kinsy
FPL2
2018 ClosNets: Batchless DNN Training with On-Chip a Priori Sparse Neural Topologies
abstract
The deployment of deep neural network (DNN) models is generally hindered by their training time. DNN training throughput is commonly limited by the fully-connected layers. This is due to their large size and low data reuse. Large batch sizes are often used to mitigate some of the effects. Increasing batch size can however hurt model accuracy, creating a tradeoff between accuracy and efficiency. We tackle the problem of training DNNs in on-chip memory, allowing us to train models without the use of batching. Pruning and quantizing dense layers can greatly reduce network size, allowing models to fit on the chip, but can only be applied after training. We propose a fully-connected but sparse layer that reduces the memory requirements of DNNs without sacrificing accuracy. We replace a dense matrix with a sparse matrix product with a predetermined topology. This allows us to: (1) train significantly smaller networks without a loss in accuracy, and (2) store weights without having to store connection indices. We therefore achieve significant training speedups due to the fast access to on-chip weights, smaller network size, and a reduced amount of computation per epoch.
Mihailo Isakov, Alan Ehret, Michel A. Kinsy
FPL3
2018 Hardening AES Hardware Implementations Against Fault and Error Inject Attacks
abstract
The Advanced Encryption Standard (AES) enables secure transmission of confidential messages. Since its invention, there have been many proposed attacks against the scheme. For example, one can inject errors or faults to acquire the encryption keys. It has been shown that the AES algorithm itself does not provide a protection against these types of attacks. Therefore, additional techniques like error control codes (ECCs) have been proposed to detect active attacks. However, not all the proposed solutions show the adequate efficacy. For instance, linear ECCs have some critical limitations, especially when the injected errors are beyond their fault detection or tolerance capabilities. In this paper, we propose a new method based on a non-linear code to protect all four internal stages of the AES hardware implementation. With this method, the protected AES system is able to (a) detect all multiplicity of errors with a high probability and (b) correct them if the errors follow certain patterns or frequencies. Results shows that the proposed method provides much higher security and reliability to the AES hardware implementation with minimal overhead.
Lake Bu, Michel A. Kinsy
ACM Great Lakes Symposium on VLSI2
2018 NoSync: Particle Swarm Inspired Distributed DNN Training
Mihailo Isakov, Michel A. Kinsy
ICANN (2)2
2017 PreNoc: Neural Network based Predictive Routing for Network-on-Chip Architectures
abstract
In this paper, we introduce a neural network based predictive routing algorithm for on-chip networks which uses anticipated global network state and congestion information to efficiently route network traffic. The core of the algorithm is a multi-layer neural network machine learning approach where the inputs are level of occupancy of virtual channels, average latency for a particular router to be selected for route computation, the probability of virtual channel allocation, and the probability of winning switch arbitration at the crossbar. The algorithm lends itself to both node routing and source routing. To evaluate the PreNoc routing algorithm, we simulate both synthetic traffic and real application traces using a cycle-accurate simulator. In most test cases, the proposed approach outperforms current deterministic and adaptive routing techniques in terms of latency and throughput. The hardware overhead for supporting the new routing algorithm is minimal.
Michel A. Kinsy, Shreeya Khadka, Mihailo Isakov
ACM Great Lakes Symposium on VLSI1
2017 Crosstalk Free Coding Systems to Protect NoC Channels against Crosstalk Faults
abstract
Reliability of modern multicore and many-core chips is tightly coupled with the reliability of their on-chip networks. Communication channels in current Network-on-Chips (NoCs) are extremely susceptible to crosstalk faults. In this work, we propose a set of rules for generating classes of crosstalk free coding systems to protect communication channels in NoCs against crosstalk faults. Codewords generated through these rules are free of '101' and '010' bit patterns, which are the main sources of crosstalk faults in NoC communication channels. The proposed rules determine: (1) the weights of different bit positions in a coding system to reach crosstalk free codings, and (2) how the coding might be utilized in an NoC to prevent crosstalk generating bit patterns in NoC channels. Using the proposed set of rules, designers can obtain coding systems which are crosstalk free for any widths of communication channels. Compared to conventional Forbidden Pattern Free (FPF) systems, the proposed methodology is able to provide unique representation to any input values at the lower bound of the codeword lengths. Analyses show that the proposed rules, along with the proposed encoding/decoding mechanisms, are effective in preventing forbidden pattern coding systems for network-on-chips of any arbitrary channel width.
Kimia Soleimani, Ahmad Patooghy, Nasim Soltani, Lake Bu, Michel A. Kinsy
ICCD5
2017 Adaptive Manycore Architectures for Big Data Computing
abstract
This work presents a cross-layer design of an adaptive manycore architecture to address the computational needs of emerging big data applications within the technological constraints of power and reliability. From the circuits end, we present links with reconfigurable repeaters that allow single-cycle traversals across multiple hops, creating fast single-cycle paths on demand. At the microarchitecture end, we present a router with bi-directional links, unified virtual channel (VC) structure, and the ability to perform self-monitoring and self-configuration around faults. We present our vision for self-aware manycore architectures and argue that machine learning techniques are very appropriate to efficiently control various configurable on-chip resources in order to realize this vision. We provide concrete learning algorithms for core and NoC reconfiguration; and dynamic power management to improve the performance, energy-efficiency, and reliability over static designs to meet the demands of big data computing. We also discuss future challenges to push the state-of-the-art on fully adaptive manycore architectures.
Janardhan Rao Doppa, Ryan Gary Kim, Mihailo Isakov, Michel A. Kinsy, Hyoukjun Kwon, Tushar Krishna
NOCS4
2016 Fault-Aware Load-Balancing Routing for 2D-Mesh and Torus On-Chip Network Topologies
abstract
Routing algorithm design for on-chip networks (OCNs) has become increasingly challenging due to high levels of integration and complexity of modern systems-on-chip (SoCs). The inherent unreliability of components, embedded oversized IP blocks, and finegrained voltage-frequency islands (VFIs) management among others, raise several challenges in OCNs: (a) network topologies become irregular or asymmetric making circular route dependencies that lead to deadlock hard to detect; and (b) routing algorithms that lack strong load-balancing properties often saturate prematurely. In order to address the aforementioned deadlock and loadbalancing problems, we propose the traffic balancing oblivious routing (TBOR) algorithm. It is a two-phase routing algorithm consisting of: (1) construction of the weighted acyclic channel dependency graph (CDG) for the OCN to efficiently maximize available resource utilization; and (2) channel ordering across turn models to keep the underlying CDG cycle-free to guarantee deadlock-freedom using one or more turn-models. Channel bandwidth utilization and traffic balancing are achieved through static virtual channel allocation according to residual bandwidth of healthy links. In addition, we introduce in this work two schemes of different granularity of fault detection and analysis while guaranteeing in-order packet delivery by assigning a unique path to each flow. Extensive experiments demonstrate the proposed routing methodology outperforms previous algorithms.
Pengju Ren, Michel A. Kinsy, Nanning Zheng 0001
IEEE Trans. Computers2
2016 A Deadlock-Free and Connectivity-Guaranteed Methodology for Achieving Fault-Tolerance in On-Chip Networks
abstract
To improve the reliability of on-chip network based systems, we design a deadlock-free routing technique that is more resilient to component failures and guarantees a higher degree of node connectivity. The routing methodology consists of three key steps. First, we determine the maximal connected subgraph of the faulty network by checking whether the defective components happen to be the cut vertices and bridges of the network topology. A precise fault diagnosis mechanism is used to identify partial defective routers. Second, we construct an acyclic channel dependency graph that breaks all cycles and preserves connectivity of the maximal connected subgraph. This is done through the cycle-breaking and connectivity guaranteed (CBCG) algorithm. Finally, we introduce a fault-tolerant adaptive routing scheme that can be used with or without virtual channels for network congestion avoidance and high-throughput routing. The simulation results show both the effectiveness and robustness of the proposed approach. For an 8 × 8 2D-Mesh with 40 percent of link damage, full connectivity and deadlock freedom are still archived without disabling any faultless router in 98.18 percent of the simulations. In a 2D-Torus, the simulation percentage is even higher (99.93 percent). The hardware overhead for supporting the introduced features is minimal. An on-line implementation of CBCG using TSMC 65nm library has only 0.966 and 1.139 percent area overhead for the 8 × 8 and 16 × 16 2D-Meshes.
Pengju Ren, Xiaowei Ren, Sudhanshu Sane, Michel A. Kinsy, Nanning Zheng 0001
IEEE Trans. Computers4
2013 MARTHA: architecture for control and emulation of power electronics and smart grid systems
abstract
This paper presents a novel Multicore Architecture for Real-Time Hybrid Applications (MARTHA) with time-predictable execution, low computational latency, and high performance that meets the requirements for control, emulation and estimation of next-generation power electronics and smart grid systems. Generic general-purpose architectures running real-time operating systems (RTOS) or quality of service (QoS) schedulers have not been able to meet the hard real-time constraints required by these applications. We present a framework based on switched hybrid automata for modeling power electronics applications. Our approach allows a large class of power electronics circuits to be expressed as switched hybrid models which can be executed on a single hardware platform.
Michel A. Kinsy, Ivan Celanovic, Omer Khan, Srini Devadas
DATE1
2013 Heracles: a tool for fast RTL-based design space exploration of multicore processors
abstract
This paper presents Heracles, an open-source, functional, parameterized, synthesizable multicore system toolkit. Such a multi/many-core design platform is a powerful and versatile research and teaching tool for architectural exploration and hardware-software co-design. The Heracles toolkit comprises the soft hardware (HDL) modules, application compiler, and graphical user interface. It is designed with a high degree of modularity to support fast exploration of future multicore processors of di erent topologies, routing schemes, processing elements (cores), and memory system organizations. It is a component-based framework with parameterized interfaces and strong emphasis on module reusability. The compiler toolchain is used to map C or C++ based applications onto the processing units. The GUI allows the user to quickly con gure and launch a system instance for easy factorial development and evaluation. Hardware modules are implemented in synthesizable Verilog and are FPGA platform independent. The Heracles tool is freely available under the open-source MIT license at: http://projects.csail.mit.edu/heracles
Michel A. Kinsy, Michael Pellauer, Srini Devadas
FPGA1
2013 Optimal and Heuristic Application-Aware Oblivious Routing
abstract
Conventional oblivious routing algorithms do not take into account resource requirements (e.g., bandwidth, latency) of various flows in a given application. As they are not aware of flow demands that are specific to the application, network resources can be poorly utilized and cause serious local congestion. Also, flows, or packets, may share virtual channels in an undetermined way; the effects of head-of-line blocking may result in throughput degradation. In this paper, we present a framework for application-aware routing that assures deadlock freedom under one or more virtual channels by forcing routes to conform to an acyclic channel dependence graph. In addition, we present methods to statically and efficiently allocate virtual channels to flows or packets, under oblivious routing, when there are two or more virtual channels per link. Using the application-aware routing framework, we develop and evaluate a bandwidth-sensitive oblivious routing scheme that statically determines routes considering an application's communication characteristics. Given bandwidth estimates for flows, we present a mixed integer-linear programming (MILP) approach and a heuristic approach for producing deadlock-free routes that minimize maximum channel load. Our framework can be used to produce application-aware routes that target the minimization of latency, number of flows through a link, bandwidth, or any combination thereof. Our results show that it is possible to achieve better performance than traditional deterministic and oblivious routing schemes on popular synthetic benchmarks using our bandwidth-sensitive approach. We also show that, when oblivious routing is used and there are more flows than virtual channels per link, the static assignment of virtual channels to flows can help mitigate the effects of head-of-line blocking, which may impede packets that are dynamically competing for virtual channels. We experimentally explore the performance tradeoffs of static and dynamic virtual channel allocation on bandwidth-sensitive and traditional oblivious routing methods.
Michel A. Kinsy, Myong Hyon Cho, Keun Sup Shim, Mieszko Lis, G. Edward Suh, Srini Devadas
IEEE Trans. Computers1
2011 Heracles: Fully Synthesizable Parameterized MIPS-Based Multicore System
abstract
Heracles is an open-source complete multicore system written in Verilog. It is fully parameterized and can be reconfigured and synthesized into different topologies and sizes. Each processing node has a fully bypassed, 7-stage pipelined microprocessor running the MIPS-III ISA, a 4-stage input-buffer, virtual-channel router, and a local variable-size shared memory. Our design is highly modular with clear interfaces between the core, the memory hierarchy, and the on-chip network. In the baseline design, the microprocessor is attached to two caches, one instruction cache and one data cache, which are oblivious to the global memory organization. The memory system in Heracles can be configured as one single global shared memory (SM), or distributed shared memory (DSM), or any combination thereof. Each core is connected to the rest of the network of processors by a parameterized, realistic, wormhole router. We show different topology configurations of the system, and their synthesis results on the Xilinx Virtex-5 LX330T FPGA board. We also provide a small MIPS cross-compiler tool chain to assist in developing software for Heracles.
Michel A. Kinsy, Michael Pellauer, Srini Devadas
FPL1
2011 HAsim: FPGA-based high-detail multicore simulation using time-division multiplexing
abstract
In this paper we present the HAsim FPGA-accelerated simulator. HAsim is able to model a shared-memory multicore system including detailed core pipelines, cache hierarchy, and on-chip network, using a single FPGA. We describe the scaling techniques that make this possible, including novel uses of time-multiplexing in the core pipeline and on-chip network. We compare our time-multiplexed approach to a direct implementation, and present a case study that motivates why high-detail simulations should continue to play a role in the architectural exploration process.
Michael Pellauer, Michael Adler, Michel A. Kinsy, Angshuman Parashar, Joel S. Emer
HPCA3
2011 Time-Predictable Computer Architecture for Cyber-Physical Systems: Digital Emulation of Power Electronics Systems
abstract
The smart grid concept is a good example of a complex cyber-physical system (CPS) that exhibits intricate interplay between control, sensing, and communication infrastructure on one side, and power processing and actuation on the other side. The more extensive use of computation, sensing, and communication, tightly coupled with power processing, calls for a fundamental reassessment of some of the prevailing paradigms in the real-time control and communication abstractions. Today these abstractions are mostly thought of as embedded systems, and the overall framework needs to be reformed in order to fully realize the potential of the emerging field of cyber-physical systems. This paper details the design and application of a new ultrahigh speed real-time emulation platform for Hardware-in-the-Loop (HiL) testing and design of high-power power electronics systems. Our real-time hardware emulation for HiL systems is based on a reconfigurable, heterogeneous, multicore processor architecture that emulates power electronics, and includes a circuit compiler that translates graphic system models into processor executable machine code. We present the hardware architecture, and describe the process of power electronic circuit compilation. This approach yields real-time execution on the order of 1μs simulation time step (including input/output latency) for a broad class of power electronics converters. To the best of our knowledge, no current academic or industrial HiL system has such a fast emulation response time. We present HiL experimental results for three representative systems: a variable speed induction motor drive, a utility grid connected photovoltaic converter system, and a hybrid electric vehicle motor drive.
Michel A. Kinsy, Omer Khan, Ivan Celanovic, Dusan Majstorovic, Nikola L. Celanovic, Srini Devadas
RTSS1
2011 Brief announcement: distributed shared memory based on computation migration
abstract
Share on Brief announcement: distributed shared memory based on computation migration Authors: Mieszko Lis Massachusetts Institute of Technology, Cambridge, MA, USA Massachusetts Institute of Technology, Cambridge, MA, USAView Profile , Keun Sup Shim Massachusetts Institute of Technology, Cambridge, MA, USA Massachusetts Institute of Technology, Cambridge, MA, USAView Profile , Myong Hyon Cho Massachusetts Institute of Technology, Cambridge, MA, USA Massachusetts Institute of Technology, Cambridge, MA, USAView Profile , Christopher W. Fletcher Massachusetts Institute of Technology, Cambridge, MA, USA Massachusetts Institute of Technology, Cambridge, MA, USAView Profile , Michel Kinsy Massachusetts Institute of Technology, Cambridge, MA, USA Massachusetts Institute of Technology, Cambridge, MA, USAView Profile , Ilia Lebedev Massachusetts Institute of Technology, Cambridge, MA, USA Massachusetts Institute of Technology, Cambridge, MA, USAView Profile , Omer Khan Massachusetts Institute of Technology, Cambridge, MA, USA Massachusetts Institute of Technology, Cambridge, MA, USAView Profile , Srinivas Devadas Massachusetts Institute of Technology, Cambridge, MA, USA Massachusetts Institute of Technology, Cambridge, MA, USAView Profile Authors Info & Claims SPAA '11: Proceedings of the twenty-third annual ACM symposium on Parallelism in algorithms and architecturesJune 2011 Pages 253–256https://doi.org/10.1145/1989493.1989530Online:04 June 2011Publication History 3citation154DownloadsMetricsTotal Citations3Total Downloads154Last 12 Months8Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Mieszko Lis, Keun Sup Shim, Myong Hyon Cho, Christopher W. Fletcher, Michel A. Kinsy, Ilia A. Lebedev, Omer Khan, Srini Devadas
SPAA5
2009 Oblivious Routing in On-Chip Bandwidth-Adaptive Networks
abstract
Oblivious routing can be implemented on simple router hardware, but network performance suffers when routes become congested. Adaptive routing attempts to avoid hot spots by re-routing flows, but requires more complex hardware to determine and configure new routing paths. We propose onchip bandwidth-adaptive networks to mitigate the performance problems of oblivious routing and the complexity issues of adaptive routing. In a bandwidth-adaptive network, the bisection bandwidth of network can adapt to changing network conditions. We describe one implementation of a bandwidth-adaptive network in the form of a two-dimensional mesh with adaptive bidirectional links, where the bandwidth of the link in one direction can be increased at the expense of the other direction. Efficient local intelligence is used to reconfigure each link, and this reconfiguration can be done very rapidly in response to changing traffic demands. We compare the hardware designs of a unidirectional and bidirectional link and evaluate the performance gains provided by a bandwidth-adaptive network in comparison to a conventional network under uniform and bursty traffic when oblivious routing is used.
Myong Hyon Cho, Mieszko Lis, Keun Sup Shim, Michel A. Kinsy, Tina Wen, Srini Devadas
PACT4
2009 Application-aware deadlock-free oblivious routing
abstract
Conventional oblivious routing algorithms are either not application-aware or assume that each flow has its own private channel to ensure deadlock avoidance. We present a framework for application-aware routing that assures deadlock-freedom under one or more channels by forcing routes to conform to an acyclic channel dependence graph. Arbitrary minimal routes can be made deadlock-free through appropriate static channel allocation when two or more channels are available. Given bandwidth estimates for flows, we present a mixed integer-linear programming (MILP) approach and a heuristic approach for producing deadlock-free routes that minimize maximum channel load. The heuristic algorithm is calibrated using the MILP algorithm and evaluated on a number of benchmarks through detailed network simulation. Our framework can be used to produce application-aware routes that target the minimization of latency, number of flows through a link, bandwidth, or any combination thereof.
Michel A. Kinsy, Myong Hyon Cho, Tina Wen, G. Edward Suh, Marten van Dijk, Srini Devadas
ISCA1
2009 Static virtual channel allocation in oblivious routing
abstract
Most virtual channel routers have multiple virtual channels to mitigate the effects of head-of-line blocking. When there are more flows than virtual channels at a link, packets or flows must compete for channels, either in a dynamic way at each link or by static assignment computed before transmission starts. In this paper, we present methods that statically allocate channels to flows at each link when oblivious routing is used, and ensure deadlock freedom for arbitrary minimal routes when two or more virtual channels are available. We then experimentally explore the performance trade-offs of static and dynamic virtual channel allocation for various oblivious routing methods, including DOR, ROMM, Valiant and a novel bandwidth-sensitive oblivious routing scheme (BSORM). Through judicious separation of flows, static allocation schemes often exceed the performance of dynamic allocation schemes.
Keun Sup Shim, Myong Hyon Cho, Michel A. Kinsy, Tina Wen, Mieszko Lis, G. Edward Suh, Srini Devadas
NOCS3
2008 Diastolic arrays: throughput-driven reconfigurable computing
abstract
Diastolic arrays are arrays of processing elements that communicate exclusively through First-In First-Out (FIFO) queues. FIFO virtualization units enable relaxed timing of data transfers, and include hardware support to guarantee bandwidth and buffer space for all data transfers, which may follow composite paths through the network. We show that the architecture of diastolic arrays enables efficient synthesis from high-level specifications of communicating finite state machines so average throughput is maximized. Preliminary results are presented on an H.264 decoding benchmark.
Myong Hyon Cho, Chih-Chi Cheng, Michel A. Kinsy, G. Edward Suh, Srini Devadas
ICCAD3
2007 Storing Efficiently Bioinformatics Workflows
abstract
We propose an efficient storage strategy to record bioinformatics workflows. Our approach presents scientists with a flexible design model that distinguishes the scientific aim as a design protocol expressed against an ontology from the implementations (s), scientific workflows composed of bioinformatics services, and their execution. The storage strategy presented in this paper allows efficient access and constitutes the framework for reasoning on scientific protocols and experimental data.
Michel A. Kinsy, Zoé Lacroix
BIBE1